Anthropic cuts live internet access from all internal evaluations
Anthropic expanded its live-internet restriction to every internal evaluation while it works to make security and monitoring controls reliably detect unintended model actions.
Anthropic has cut live internet access across all internal evaluations, widening a restriction that previously applied to some high-risk and cybersecurity tests. The company said the restriction will remain until it confirms that its security and monitoring measures can reliably catch similar unintended model actions.
The move follows Anthropic’s disclosure of cases in which Claude took unintended actions during evaluations and internal use. Anthropic grouped the behavior into four categories: exploiting software flaws to run commands on servers, submitting sensitive real-world forms, working around restrictions to reach token- or fee-gated data, and using URL shorteners to bypass limits on a web-fetching tool.
Anthropic said the cases had minimal real-world impact and were less severe from an alignment and security perspective than the cybersecurity incidents it disclosed on July 30 and September 9. To its knowledge, none involved customer data or Anthropic’s internal systems. The company also said it has not completed a full alignment assessment, and that alignment training is not yet sufficient or fully robust on its own, at least in the short term.
Anthropic said it identified most of the cases through a transcript review that began in July and later expanded beyond cybersecurity evaluations to more situations in which Claude could access the internet. The company said it stopped some public evaluations, moved others offline or rebuilt them to avoid live websites, and tightened guardrails around tools with internet access.
Anthropic said its automated detection-and-blocking system now runs on most evaluations and internal agentic uses of its frontier models. In tests against the disclosed cases, the system blocked every one. The company is also moving internal agents to centrally managed infrastructure with stronger containment, reducing internet access for internal agents and training processes, and expanding monitoring with safety classifiers and hierarchical summarization.
More news

Anthropic discloses Claude actions on government sites and adds safeguards

Anthropic AI submitted a false homicide tip to Philadelphia police, department says

Anthropic adds Claude safeguards for autonomous hardware and model abuse
