Alan Turing is famous for the test of whether a machine can pass for a human. The harder test he faced in his lifetime was breaking Nazi Germany’s Enigma code. Now two frontier AI models have passed a version of that one too. TechCrunch AI reports that two cryptanalysts used OpenAI’s newest model, GPT-6 Astra, and Anthropic’s Claude Opus 5 to crack two separate Enigma messages that had gone unsolved for decades.
This isn’t a benchmark score or a lab demo. Real archival puzzles held out against human experts for years, and they’re now solved and checked by one of the field’s recognized record-keepers.
🔐 The Problem: Messages Nobody Could Break
During World War II, Turing and his team built the Bombe, an early electromechanical machine that let Britain read Enigma traffic at scale. A handful of archived messages still resisted every attempt, though. According to TechCrunch AI, that’s usually because of mistranscriptions or mistakes made by the original German operators.
Those errors make these leftovers much harder than a textbook Enigma problem. You can’t just brute-force the machine settings. You need context: who sent the message, when, from which unit, and what they probably said.
Frode Weierud keeps track of what’s left. He’s a retired electrical engineer who runs Crypto Cellar, a website with records, resources and a database of Enigma messages.
🤖 The Solution: Two Very Different Approaches
Carter Leffen and GPT-6 Astra (almost no guidance). Leffen, a developer, simply told Astra to search a database of Enigma messages, pick an unbroken one and decode it. According to TechCrunch AI, the model then:
- Did its own archival research
- Found context clues in the historical record
- Built a working Enigma simulator
- Recovered the plaintext of a message that had stumped researchers since 2005
Leffen then had Astra build an interactive website explaining the whole problem.
Jack Willis and Claude Opus 5 (heavy guidance). On September 21, Willis, a cybersecurity executive, contacted Weierud to say he had broken a different unsolved message with Claude Opus 5. TechCrunch AI notes that Willis gave Claude “significantly more guidance.” The break came from a known signature: the name of a particular officer, which gave the model a foothold into the ciphertext.
✅ The Result: Expert Validation
Weierud checked Leffen’s solution last week and said it left him in “awe.” His assessment of Astra is worth quoting in full:
GPT-6 Astra is behaving like a very professional cryptanalyst and archive researcher. What it has achieved in two days would take a human researcher weeks or even months. Personally, I spent several weeks researching the Bundesarchiv files GPT-6 Astra refers to.
There’s a strange footnote, though. Astra’s logs mention archived messages in a “private collection” that Weierud doesn’t host. He still doesn’t know whether the model actually got to them. His guesses are that another researcher shared them somewhere online, or that the model reached them through the German government’s public archives.
🧭 Why This Matters
What stands out to me is that the hard part wasn’t the math. Enigma simulators have existed for years. The breakthrough came from the grunt work around the math: digging through archives, connecting context and forming hypotheses. That’s the long, tedious research that used to take specialists weeks.
A few takeaways for practitioners:
- Agentic research is getting real. Astra went from a one-line instruction to a verified result without anyone steering it step by step.
- Guidance still pays off. Willis’s more hands-on approach with Claude also worked. Expert direction plus a strong model is still a powerful combination.
- Watch where your agents go. The “private collection” mystery shows how far these models will go to answer a question. If you deploy agents, logging and source auditing aren’t optional.
🔭 What’s Next
Weierud says just seven unbroken Enigma messages remain, plus one message whose plaintext is known but whose cipher still hasn’t been cracked. With two falling to AI in quick succession, I wouldn’t bet on that list lasting long.
The wider signal matters more than the codes. Plenty of fields have their own “unbroken messages”: stalled research problems, messy archives and cold cases. Those are next. Full details are available in the original TechCrunch AI report.