The AI That Built Its Own Bombe Changes the Conversation About Tool-Use
A Korean research team claims an unreleased model called 'GPT-6 Astra' deciphered a 1941 Enigma message by writing its own simulator and Bombe decoder in two days. The result is less interesting than the method: if verified, it marks a shift from pattern-matching to autonomous tool construction.
The result is boring. The method is not.
An Enigma message from July 1941, sitting undecoded since its 2005 online publication, reportedly fell to an AI in two days. The headline number — 85 years of unsolved ciphertext — makes for good copy. But the actual engineering question hiding underneath is sharper: did the system reason through the problem, or did it pattern-match its way to a pre-built answer?
According to a report published on the 24th by Cryptocellar, the crypto-history research platform run by Frode Weierud, a team member named Carter Reffer fed a 1941 German military message into a model calling itself “GPT-6 Astra.” Two days later, the system had produced a working Enigma simulator, a Python-and-C++ implementation of Turing’s famous Bombe decoder, and a plaintext. The message, labeled MVUEH, was an 82-character ciphertext sent from a German military radio station to the SS Totenkopf Division’s supply unit. The decoded text turned out to be nearly identical to a sibling message, SIPVX, sent the same day — which is exactly what you would expect from a logistics radio network retransmitting the same orders.
If you have spent time around cryptanalysis, the result itself is unremarkable. Every Enigma message with a known crib is solvable by a modern laptop in under a second. The mathematical structure has been public since 1940. The 85-year gap is a gap of attention, not of difficulty. Computers have been breaking Enigma traffic routinely since the late 1990s.
What the two days actually covered
Here is where the report gets more interesting. The team claims the system was not handed a step-by-step procedure. It identified MVUEH as the most promising target from a batch of undecoded messages, spotted the likely connection to SIPVX, proposed the geographic crib “ROSENOW ROSENOW” as a common element, and then wrote the tools itself. That means generating a rotor-logic simulator, encoding the Bombe’s null-testing algorithm in C++, and running a large-scale search. The team also notes two specific quirks that complicated the decode: a left-rotor behavior anomaly at the 72nd character, and transcription errors in the original typewritten document that had persisted since the 1940s.
The researchers estimate that a human analyst working through the same pipeline would need several months. They call it a paradigm shift in historical crypto-decoding. That is a strong claim, and it deserves scrutiny.
The model you cannot find
The biggest problem with the story is the model’s name. “GPT-6 Astra” does not correspond to any publicly announced product from OpenAI, Anthropic, Google, or Meta. No press release, no API documentation, no benchmark suite goes by that label. The Korean source, AI Times, reports it as fact without linking to a primary technical paper or a code repository. Cryptocellar has not posted a verifiable git commit, a hash of the generated C++ source, or a step-by-step log of the two-day session.
Until someone publishes the generated code, the rotor settings, and the prompt/response transcript, the claim sits in the category of “plausible but unconfirmed.” The methodology described is historically correct — cribs, Bombe null tests, rotor-order guessing are the standard toolkit. A competent coding agent could plausibly produce working implementations of all three. But the gap between “a coding agent wrote a functional Bombe” and “an autonomous reasoning system solved a novel cryptographic problem” is enormous, and the report blurs that line by leading with the 85-year framing.
Why the tool-use distinction matters outside cryptography
Strip the Enigma shell off and the underlying capability question is one that software engineers, researchers, and policy makers have been arguing about for eighteen months. Current frontier models are excellent at pattern matching over learned contexts: given a codebase, they refactor; given a dataset, they classify. They are far less proven at tool construction under constraint — deciding what to build, writing it from scratch, debugging it against physical or logical specifications, and iterating without a human in the loop.
If the Cryptocellar report holds up, the two-day session demonstrates all of those steps in sequence, applied to a domain where the “correct tool” is not a neural network layer but a mechanical simulator with 26-contact rotors and a permutation network. That is a different cognitive profile than the LLM benchmarks we track (MMLU, HumanEval, SWE-bench). It is closer to what robotics researchers call autonomous setup: the system has to know enough about the physics to build the right instrument before it can even start measuring.
That has direct consequences. In scientific research, most experiments require custom instrumentation. In materials science, a researcher designs a crystal lattice probe; in drug discovery, someone scripts a docking pipeline. If a language model can move from “here is the problem” to “here is a working tool that tests the hypothesis” without a human writing the scaffolding, the bottleneck in discovery shifts from writing code to framing the question correctly.
The ADFGVX footnote and what to watch next
The report mentions that, just the prior week, the same model reportedly decoded a 108-year-old German ADFGVX message from World War I, a polyalphabetic cipher that resisted manual attack for decades. Two distinct cipher systems, two different toolchains (Enigma rotors vs. Latin-square polyalphabetic tables), in the span of a week. If both checks out, the pattern is less about one lucky solve and more about a general procedure: identify the cipher family, build the matching simulator, enumerate the key space, and validate against known-plaintext constraints.
That procedure generalizes. It applies to one-time pad analysis, to SIGINT traffic from the Cold War era, to any archived communication where the encoding scheme is known but the keys are lost. The practical value is not breaking modern AES; it is recovering history at machine speed.
What would convince a skeptic
Three things. First, a public repository with the generated C++ Bombe source, the Python rotor logic, and the full command log. Second, a reproducibility test: hand the same prompt to a different model, say Claude or Gemini, and see whether the same pipeline emerges. Third, a cryptographic audit of the plaintext against known archival documents from the Totenkopf Division’s 1941 supply correspondence, confirming that “ROSENOW” and “Waschbusch” are genuine period locations and operator names.
Until those three boxes are checked, the story is a strong lead, not a confirmed result. But the shape of the claim — an AI that builds its own measurement instruments before taking the measurement — is the one worth taking seriously, because it is the exact capability that current benchmark suites do not test. The Enigma is just the first exam. The real question is whether the exam format holds for the next one.
Correction: The model name “GPT-6 Astra” is as reported in the source. No independent confirmation of its existence or specifications was available at time of writing.