Anthropic is now measuring whether its AI models can help with two of the most sensitive tasks in national security: intelligence targeting and the use of conventional weapons. According to Anthropic, this work is part of a broader push to understand exactly where frontier models start to become dangerous, rather than guessing. What stands out here is the shift from vague worry to structured measurement.
This matters because the debate around AI and weapons has mostly run on assumptions. Anthropic is trying to replace those assumptions with evidence.
What the research is actually doing
The core idea is straightforward. Instead of asking whether an AI “could be dangerous” in the abstract, Anthropic reports it’s building concrete evaluations that probe specific capabilities tied to real-world military and intelligence tasks.
That means testing whether a model can:
- Support intelligence targeting, meaning the process of identifying and prioritizing targets from raw data.
- Provide meaningful uplift on conventional weapons, the non-nuclear, non-biological category that covers most military hardware.
- Do these things better than existing tools a person could already access.
The last point is the one to watch. A capability only counts as a real risk if the model adds something a determined person couldn’t already get elsewhere. Measuring that “uplift” is harder than it sounds, and it’s the difference between a scary demo and an actual threat.
Why this is different from a red-team headline
Plenty of AI safety stories are one-off stunts: someone jailbreaks a model, gets an alarming answer, and posts it. This is a different exercise.
Anthropic is treating weapons and targeting capability as something you benchmark over time, the same way labs track math or coding performance. That gives them a repeatable yardstick. As models get more capable, they can re-run the same tests and see whether the risk line has moved.
This fits directly into how Anthropic frames its Responsible Scaling Policy, where specific capability thresholds are supposed to trigger stronger safeguards before a model ships. You can’t enforce a threshold you can’t measure. That’s the gap this research is trying to close.
What it means for practitioners
If you build with or deploy frontier models, there are a few practical takeaways.
- Capability evaluations are becoming a product gate, not an afterthought. Expect national-security-relevant testing to shape what future models will and won’t do out of the box.
- “Uplift over baseline” is the metric that counts. When you assess any model for risk, the useful question isn’t “can it answer this,” it’s “does it beat what’s already freely available.” That framing applies well beyond weapons.
- Domain experts are part of the loop. Serious evaluation in this space needs people who actually understand targeting and munitions, not just prompt engineers. If your organization does risk assessment, that mix of expertise matters.
The limits worth noting
Anthropic is candid that measuring these capabilities is genuinely difficult. Real military tasks depend on classified data, physical resources, and operational context that a chatbot doesn’t have, so a model scoring well on a text-based proxy isn’t the same as a battlefield capability. There’s also the ever-present risk that a test underestimates what a motivated user could coax out of a system.
The honest read is that this is early, imperfect, and evolving. But it’s the right direction.
As models keep climbing in capability, the labs that can prove where the danger lines sit, and show they’re watching them, will set the terms for how this technology gets governed. Expect more of these evaluations, and more scrutiny of them, in the months ahead. Full details are available in Anthropic’s original write-up.