Anthropic has published its threat intelligence report for September 2026, covering activity it disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation.

The conventional weapons section describes six cases — "three in China, two in Russia, and one in Yemen." One of them is the subject of this piece.

What Was Built

Case GTG-17002:

"We identified a China-based actor who used Claude's chat, coding, and agentic work tools to design, build, and iterate on a Chinese-language suite of about 16 modules for electronic warfare, using the electromagnetic spectrum to detect, jam, or deceive an opponent's radar and communications, and for suppressing an opponent's air defenses."

Note the shape of this. It is not somebody asking a chatbot a dangerous question. It is a software project — about 16 modules, built "from the underlying logic to the user interface," iterated through 12 versions.

What the finished suite did:

"The software suite the actor built analyzed an opponent's radars, surface-to-air missile sites, command posts, and communications nodes, then computed their detection coverage, assessed the effectiveness of jamming, ranked targets by value and vulnerability, including which to suppress first, and determined how best to assign jammer sorties to targets across multi-day campaigns."

It also "modeled specific engagement envelopes, including those of Patriot and THAAD-class systems."

The Sentence That Matters

"Mid-project, we observed the actor change the simulation's default scenario to 12 targets in Taiwan. The targets included a command bunker in Taiwan, an early warning radar site, Patriot and Tien Kung batteries, major air bases, and a regional combatant command headquarters."

Sit with the structure of that for a second, because it's the actual lesson.

Every module described above is dual-use on its face. Modeling radar detection coverage is what you do to evaluate your own radar network. Assessing jamming effectiveness is standard defensive electronic warfare analysis. Ranking assets by value and vulnerability is, stripped of context, an optimization problem that appears in insurance, logistics and infrastructure planning.

None of the software announced what it was for.

The intent arrived as a default scenario — a configuration setting. Twelve real places.

That is a genuinely hard detection problem, and it's worth being honest about why. The signal wasn't in the capability, the code, or the requests. It was in the smallest, most easily changed, least reviewed part of the system. A config value can be swapped in seconds and swapped back just as fast.

Who

"Based on our investigation, we assess the actor is a China-based defense and military-industrial researcher. Account-level metadata and content flagged by our safeguards indicated the actor was linked to PRC research institutions, including the PLA Academy of Military Sciences."

Anthropic banned the accounts linked to the actor.

The attribution language is worth reading carefully. "We assess" is analytic judgment, not proof. "Indicated" is weaker than "confirmed." The report is being appropriately careful, and summaries of it should be too.

Three Things Most Summaries Get Wrong

1. It wasn't "consumer AI."

One sentence in the report changes the framing entirely: "The actor also ran a self-hosted model on an internal network alongside Claude and connected the software suite to this model through a tool-use integration."

This was a hybrid stack. A self-hosted model on an internal network doing part of the work, a frontier model doing the part the self-hosted one couldn't, wired together through tool use. Describing it as somebody misusing a chatbot understates the sophistication and, more importantly, overstates how much cutting off one vendor accomplishes.

2. The software did not rank the Taiwanese sites for attack.

The report says the suite ranked targets and which to suppress first. It separately says the default scenario was changed to 12 Taiwan targets. It does not join those two statements, and neither should we. The difference between "a tool capable of ranking targets was configured with Taiwanese locations" and "a tool ranked Taiwanese locations for attack" is real, and the first is what's actually documented.

3. Safeguards worked partially, not cleanly.

Across these cases the report describes actors who "split their work across many sessions to conceal the full nature of their programs." In the parallel Yemen case, it says plainly: "Our safeguards blocked many of their requests, but not all of them."

A safeguard that evaluates each request on its own merits cannot see a program distributed across hundreds of individually innocuous requests. That's not a tuning problem. It's an architectural one.

The Other Cases, Briefly

For context on what else is in this section:

  • Yemen (GTG-87001). A weapons engineering cell used Claude Code in place of human software engineers to develop guidance software for a guided rocket, a ballistic missile with a stated range goal above 2,000 km, and a multi-variant missile set including a hypersonic glide vehicle. They ran multiple Claude instances at once with assigned roles — one writing code, one researching, one reviewing. They test-fired a guided rocket; it appears to have failed, and "within hours, the actors returned to Claude to work out why it failed." They had also already built an offline simulation toolkit that "does not rely on Claude."

  • Russia (GTG-27005). Freelance actors built a full-stack autonomous FPV drone swarm, designed "for autonomous lethal engagement" where the onboard model "could select targets (including a 'person' target class) and issue detonation commands without a human in the loop." Trained on scraped Ukrainian combat footage.

  • China (GTG-17001). An actor drafted a fire control specification and a 200-page acquisition proposal for an anti-torpedo system, repeatedly having Claude role-play a hostile expert reviewer to critique its own drafts.

Anthropic says it has "recently launched a new set of classifiers designed to better detect and block traffic related to high-yield explosives and weapons development."

About the Source

I want to be direct about something, because the alternative is pretending it isn't there.

This report is published by Anthropic. Anthropic also makes the AI model used to help assemble this newsletter. So what you are reading is a company's account of its own product being misused, disclosed on its own schedule, in its own framing, with no independent verification available to anyone outside the company.

Publishing it is clearly better than not publishing it. Most of what's in this report would otherwise be invisible — Anthropic notes that historically "this kind of work has been uncovered by governments, United Nations panels, and outside investigators."

But voluntary self-disclosure is a different evidentiary category from investigative journalism or a regulator's finding, and right now the industry has only the first one. There's no obligation to publish, no standard for what counts as a reportable case, no external audit of what was left out, and no way to know how many cases were not disclosed.

There's also an uncomfortable structural feature. A sentence like "our product was used to build a 16-module electronic warfare suite" is not neutral marketing copy for a company selling that product. I don't believe that's the motive here. I do think it's a good argument for why these reports should eventually be produced by someone who doesn't sell the model.

For Parents and Teachers

This one is harder to bring into a classroom than most stories we cover, and the temptation is either to skip it or to make it frightening. Neither is right.

The useful lesson isn't about weapons. It's about dual use, and it's one of the most important ideas a young person can learn about technology.

Almost every module in that suite has a legitimate, boring, defensive version. The capability was neutral. What made it dangerous was a configuration file — a setting somebody typed in, which could have been typed differently.

That generalizes far past military software. The same facial recognition that unlocks a phone identifies a protester. The same location tracking that finds a lost child follows an ex-partner. The same summarizer that helps with homework generates spam. In almost every case the technology didn't change. The target list did.

The question to teach isn't "is this tool good or bad?" It's "who gets to fill in the blanks, and who checks what they filled in?"

That question works on an AI targeting suite and it works on a school's new attendance app, which is precisely why it's worth teaching.

Disclosure: Anthropic, the company that makes the AI model used to help assemble this newsletter, is the publisher of the threat intelligence report this piece is based on. It seemed better to say so than to leave it out.

Source: "Detecting and countering misuse of AI: September 2026," Anthropic: https://www.anthropic.com/threat-intelligence-report-september-2026

Want practical AI guidance for parents and educators every week? Subscribe: https://www.aibyage.com/?modal=signup&utm_source=beehiiv&utm_medium=newsletter