Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

Anthropic says it disrupted Claude misuse across cyberattacks and weapons research

Anthropic says it disrupted Claude misuse in seven harm areas as AI took on more of the cyberattack workflow. Most incident details and actor attributions remain the company's account.

D
Sep 14, 2026 · 3 min read

Anthropic says it disrupted malicious uses of Claude in seven harm areas between December 2025 and August 2026. They included cyber operations, surveillance, weapons research, biological misuse, fraud, influence operations and illicit model distillation.

According to the company’s September threat-intelligence report, a majority of the cyber operations used AI for direct execution or orchestration, not just chatbot-style advice. Anthropic described multi-agent workflows that handled parts of reconnaissance, exploitation and data exfiltration, while human operators chose targets and reviewed stolen material.

That account points to connected AI tasks spanning an attack workflow rather than isolated requests for advice. The public evidence reviewed for this story does not independently establish most of the incidents, their scope or Anthropic’s actor attributions.

In one case, Anthropic linked an internally designated actor, GTG-20006, to Russian state-connected espionage activity consistent with public reporting about Midnight Blizzard. The company said the actor automated infrastructure acquisition, phishing, persistence, command-and-control activity and exfiltration while targeting more than 20 organizations. Neither those numbers nor the Claude-specific actor link has been independently verified.

Microsoft Threat Intelligence separately reported overlapping activity by Storm-2945, which it assesses as a Midnight Blizzard sub-cluster. Microsoft said the group manipulated hospitality-network traffic, used device-code phishing, delivered malware and used AI to support a significant portion of its CaptiveCrunch operation. Microsoft thanked Anthropic and OpenAI for support, but its report does not establish Anthropic’s full GTG-20006 scope or independently verify Claude-specific telemetry.

Anthropic also said suspected ShinyHunters affiliates used AI agents for credential discovery, intrusion, data theft and extortion. It said customer API keys involved in the activity were taken from customer environments, not through a breach of Anthropic’s systems. Those details remain Anthropic’s account.

The report describes six conventional-weapons cases: three in China, two in Russia and one in Yemen. Anthropic said four involved weapons software and two involved procurement or intelligence gathering. The work touched a guided rocket, torpedo interception, drone-swarm software and targeting software, according to the company. The public evidence reviewed for this story does not independently verify the cases or attributions.

Anthropic presented five other cases in which model use could have supported biological-weapons development. The company said it was not asserting that the scientists involved intended harm, and it withheld their institutions, countries and some technical details. The information Anthropic made public is therefore insufficient to independently check the underlying activity.

In response, Anthropic said it banned accounts tied to the reported activity, incorporated investigative findings into safeguards and shared intelligence with authorities and industry partners where appropriate. It also said it launched classifiers intended to detect and block high-yield-explosives and weapons-development traffic and placed stronger restrictions on dual-use biological research queries in newer models. The reviewed evidence includes no public measurements of those safeguards’ effectiveness.

Anthropic also said it disrupted illicit model-distillation campaigns associated with seven China-based labs. The company described proxy accounts, stolen credentials and replayed user exchanges as access methods. Its countermeasures, it said, include network attribution based on metadata, extraction-detection classifiers, request blocking, account bans, controls that limit exposure of internal reasoning and identity verification when abuse signals appear.

More news