Anthropic: Claude misused for weapons, scams, espionage

Anthropic: Claude misused for weapons, scams, espionage

Anthropic published its September 2026 threat intelligence report on September 10, detailing how its Threat Intelligence team identified and disrupted threat actors that misused Claude between December 2025 and August 2026 across seven harm areas, including cyber operations, influence operations, surveillance, and illicit model distillation. The team disrupted operations in which threat actors tried to use Claude for malicious activity, used what they learned to strengthen safeguards, and shared intelligence with authorities and industry partners where appropriate. The cases involve suspected state-sponsored groups, financially motivated criminals, and politically motivated individuals from at least 10 countries.

AI threats diagram showing 8 malicious uses: cyber operations, influence, surveillance, biological misuse, scams, illicit ...

Iranian Surveillance and Weapons Development Highlighted

Among the most significant cases disclosed in the report, an actor in Yemen turned to Claude Code to build rocket guidance software and returned for advice after an apparently failed test flight. The report provides no evidence of an operational weapon. This case represents one of the first documented instances of AI being used in conventional weapons development efforts.

Missile guidance system diagram showing GPS satellite, inertial sensors, ground control station, radar targeting, and flig...

Iranian security units allegedly used Claude to expand surveillance operations targeting dissidents. The report also describes a consultant who built a system for Mali’s spy agency aimed at monitoring approximately 25 million SIMs. The Mali case shows a limit to account bans, as Anthropic’s bans left the local deployment in place, because software built with Claude can run elsewhere after model access is cut off.

Massive Scale AI-Powered Scams and Influence Operations

In one operation, Claude powered more than 4,700 dating app personas that exchanged 2.36 million messages with at least 25,000 users over two weeks in April 2026. Humans handled the video calls. This demonstrates how threat actors are using AI to achieve unprecedented scale in romance scams while maintaining the human touch needed to sustain deception.

Fake personas network visualization showing 4,700 fraudulent accounts, 2.36M messages, and 25,000 users in disinformation ...

Another actor drew on roughly 8,400 Telegram posts to imitate an activist’s writing style in live conversations with his contacts. The report also describes an operation that automated malware rebuilding after security detections.

Russian State-Linked Cyber Espionage

A Russian state-nexus espionage actor, identified as GTG-20006, used AI-driven workflows to automate operations from infrastructure development, with AI’s role in cyber operations becoming increasingly autonomous as Claude is used not only as an assistant but also to execute or coordinate parts of attacks. In several cases, multi-agent AI systems carried out reconnaissance, exploitation and data exfiltration, while humans largely remained involved in selecting targets and reviewing stolen information.

Anthropic said the actor “developed AI-driven workflows to research, then register domains, and then configure the hosting infrastructure used to send phishing emails.” One breach of an enterprise software company took only hours from first access to bulk data theft, while another compromise escalated from a single stolen developer token to full administrative control of a victim’s cloud environment in roughly three hours.

Biological Misuse Attempts Blocked

Anthropic flagged five biology cases involving scientists for possible weapons applications, while saying it “does not assert that they intended harm.” In May 2026, Anthropic’s biological safety classifier blocked a request to help author a grant application for chikungunya gain-of-function research intended for a military research institute, routed through an evasion platform.

Unprecedented AI Model Distillation by Chinese Tech Giants

The report revealed massive illicit distillation campaigns by major Chinese AI companies attempting to extract Claude’s capabilities. Anthropic attributed to Alibaba the largest distillation campaign it has measured: chain-of-thought distillation of Opus 4.6 and 4.7 peaking at nearly 3 million exchanges per day from more than 3,500 fraudulent accounts, with over 151 million exchanges observed between May and July 2026 and the harvested transcripts used to train Qwen 3.5, 3.6, and 3.7.

Distillation campaign scales across threat actors: DarkCrew, LemonGrove, DeepScav, HydraLabs, and Alibaba daily model exch...

Moonshot AI silently forwarded customer requests to Claude instead of its Kimi models, relaying almost 300,000 requests over ten days through 5,380 fraudulent accounts and using a cross-session replay attack on Claude’s thinking signatures to extract reasoning traces, with over 23 million exchanges observed between May and July 2026. The report also identified similar activity from DeepSeek using the same replay technique.

Key Findings on AI “Uplift” in Malicious Activity

The report attempts to measure uplift, a term used to describe the AI capability boost, or how much more harm was caused with AI versus without AI, viewing uplift through the lens of speed, scale, and depth. The risk from AI adoption is more pronounced across the cyber kill chain, where adversaries can operate faster, across a broader and deeper surface area, with fewer resources.

Anthropic found that Opus 5 outperformed models in the Mythos family on a simulated weapons guidance task. However, the synthetic data and simulations do not establish how much more effective an attacker would become in practice.

Models Involved and Safeguard Effectiveness

Claude Haiku, Sonnet, and Opus models were used in the misuse cases, while none of the misuse cases involved the use of Claude Fable or Mythos-class models, with the exception of one illicit distillation case. This suggests that the additional safeguards implemented on Anthropic’s more advanced model families have been effective at preventing malicious use.

The report observes how AI has moved from being an assistant to becoming an orchestrator with enhanced autonomous operations planned by bad actors, with actors embedding AI inside workflows that execute reconnaissance, exploitation, credential theft, data processing and exfiltration.

Key Facts

  • Report covers December 2025 to August 2026, documenting threat actors from at least 10 countries
  • Over 151 million exchanges stolen by Alibaba for training Qwen models between May and July 2026
  • 4,700 AI-powered dating app personas sent 2.36 million messages to 25,000+ users in two weeks
  • Yemen-based actor used Claude Code for rocket guidance software development
  • Mali surveillance system designed to monitor approximately 25 million SIMs
  • Russian state-linked actor escalated from stolen token to full cloud admin control in three hours
  • Five biology cases flagged for possible weapons applications
  • Claude Fable and Mythos models remained largely protected from misuse

Sources

Sources

  1. Anthropic Details Disrupted Claude Misuse Across Seven Harm Areas – Unite.AI
  2. Countering misuse of AI: September 2026 / Anthropic \ Anthropic
  3. Anthropic details Claude misuse in deception, surveillance, and malware | The Rundown AI
  4. Anthropic Report Reveals Growing Misuse of Claude AI in Cyberattacks and Espionage | Daily Pioneer
  5. Detecting and countering misuse of AI: September 2026 Published
  6. Anthropic Claude Misuse Report 2026: Missiles, Cyberattacks & Espionage