← All posts / Research

Two Ciphers, One Model: GPT-6 Astra Cracks a 108-Year-Old WWI Code and an 83-Year-Old Enigma Message in the Same Week

In a single week, OpenAI's GPT-6 Astra deciphered a 1918 ADFGVX radio message that had resisted codebreakers for 108 years and an Enigma-encrypted Wehrmacht dispatch from 1941 — verifying its work against HMS Canterbury's original logs and publishing every step.

Two Ciphers, One Model: GPT-6 Astra Cracks a 108-Year-Old WWI Code and an 83-Year-Old Enigma Message in the Same Week

Ask a professional cryptanalyst what it takes to break a century-old military cipher and you will hear about frequency analysis, archive work, and long unglamorous hours of trial and error. Ask GPT-6 Astra, and the answer this week turned out to be: about ten hours of autonomous work per cipher, a talent for questioning assumptions that humans baked in decades ago, and a habit of checking its own answers against primary historical sources. In the span of a few days, OpenAI’s frontier model has reportedly solved two encrypted messages that had sat in the “unsolved” column for generations — a 1918 German ADFGVX radio dispatch and a 1941 Enigma-encrypted Wehrmacht transmission — and in doing so offered one of the cleanest public demonstrations yet of what agentic AI can do to a problem domain that was supposed to be the exclusive province of experts.

The 108-year-old message from the Black Sea

The first breakthrough concerns a German radio message transmitted on November 27, 1918, weeks after the Armistice, when the German military was still passing traffic about Allied naval movements in the Black Sea. The message was encoded with the ADFGVX cipher, a First World War field encryption that combines a 6×6 Polybius square with a columnar transposition keyed by a codeword. It sat on Scienceblogs.de’s relatively famous list of 50 unsolved ciphers — a roster that ranges from cryptograms published by serial killers to the Voynich manuscript — as one of more than a dozen WWI German radio messages that had eluded decoding.

A developer who publishes under the name Prinz picked the message from the list and set GPT-6 Astra on it. The model worked through the mechanics methodically. ADFGVX is only as strong as its keyword: the same ciphertext, re-keyed, yields a completely different table. The known keys of the era are documented in J. Rives Childs’s The History and Principles of German Military Ciphers, 1914–1918, and hundreds of contemporary messages have been decoded with them, including by the codebreaking expert George Lasry. This one had not.

Astra’s answer was that the message used “TRUPPENVERSCHIEBUNG” — German for “troop movement” — as the key. Working the transposition by hand-checkable steps (reorder the key alphabetically, write the ciphertext under it in rows of nineteen, read the columns in key order), the decode fell out as:

EIN ENGLISCHER KREUZER EINLIEG X SEWASTOPOL X S4STEN X EIN GESCHWADER DER X ALLIIERTEN FOLGT 26STEN X

Or, in English: “An English cruiser arrived at Sevastopol on the ?4th. An Allied squadron follows on the 26th.”

The model checked its own homework

What elevates this from party trick to genuine result is what Astra did next. Rather than declaring victory, the model went looking for corroboration — and found it. The original logs of HMS Canterbury, the British light cruiser, record its arrival in Sevastopol on November 24, 1918, matching the “?4th” in the decoded text (the garbled digit was likely a transmission error). The Allied squadron’s arrival on November 26 appears in the historical record as well, just below line 11 of the log page the model consulted. Both checkable claims in a 170-character message survived contact with primary sources.

Astra also offered a hypothesis for why the message had gone unsolved for 108 years: cryptanalysts had assumed “TRUPPENVERSCHIEBUNG” was only used as a key starting December 9, 1918 — ten days after this message was transmitted. The model’s insight was simply to question that assumption. Whether the discrepancy reflects sloppy key discipline in the chaos of November 1918 or an error in the historical record is unknown; either way, a century of failed attempts apparently rested on a wrong premise about a calendar date.

The Enigma break, three days earlier

The WWI solve was not even the week’s first. On September 17, The Decoder reported on a case study by Carter Leffen, a product development coach at Bloomberg LP in New York, who used the “GPT-6 Astra Extra High” variant to crack an 82-character Enigma message from July 10, 1941. A German soldier in the town of Rosenow had encrypted a short dispatch asking for marching orders and requesting an immediate radio reply (“Sofort Funkantwort”). For 83 years, nobody had decrypted it.

The Enigma machine offered roughly 159 quintillion possible daily settings; brute force was never on the table. Known keys from the same day didn’t fit, and automated searches came up empty. The breakthrough came from a classic cryptanalytic move, executed with machine patience: a second, already-decrypted message from the same day contained the town name “Rosenow” twice in a row. Leffen and his AI agents guessed the name might also appear in the unsolved message and used it as a probable-word attack — aided by the Enigma’s famous weakness that it never encrypts a letter as itself, which instantly rules out many candidate positions.

At one position, everything lined up. The remaining 68 characters produced coherent German, complete with operator typos (“BTTE” for “BITTE”) that Leffen considers evidence of authenticity — a fabricated decryption would more likely be clean. All code, search data, and a working 3D Enigma simulator were published on the project page for independent verification, building on prior cryptanalytic work by Frode Weierud, Geoff Sullivan, and Olaf Ostwald.

Why this matters beyond trivia

It is tempting to file these stories under “fun weekend demo.” That would be a mistake, for three reasons.

First, the skill being demonstrated is not decryption — it’s research autonomy. In both cases, the hard part was not the mathematics. ADFGVX and Enigma are thoroughly understood systems; the algorithms to attack them exist in textbooks. What Astra did was manage an open-ended research process: search historical archives, identify which assumptions previous solvers had made, hypothesize an alternative, test it, and then cross-validate the output against independent primary sources. That is the workflow of a junior cryptanalyst — or, increasingly, of an agent in any evidence-heavy domain from legal discovery to materials science. Leffen’s own assessment is telling: he says he spent “99 times more effort into building the website that describes the problem and the solution than into actually cracking the code.”

Second, the verification step is the story. Every AI-generated “breakthrough” claim now carries an epistemic tax: is this real, or is it a confident hallucination? Astra’s instinct to check the decoded message against HMS Canterbury’s actual logs — and to flag the mismatched digit as a probable transmission typo rather than paper over it — is exactly the behavior that separates a trustworthy research agent from a plausible-sounding one. The typos preserved in the Enigma plaintext serve the same function: imperfection as authenticity.

Third, the security implications cut both ways. Historical ciphers are a safe testing ground, but the same agentic loop — archive search, hypothesis, code generation, parallel experiments, self-check — applies to cryptanalysis of systems still in use. The week’s other headlines include researchers factoring RSA-260 with an agent swarm and AI models breaking out of testing environments. A model that can autonomously reason its way through a transposition cipher while questioning century-old assumptions about key schedules is a model whose offensive security ceiling deserves serious attention.

The limits

Both results come with caveats. Prinz writes “to my knowledge” the 1918 message had never been decoded before — an important qualifier, since classified or simply unpublished solutions may exist. The Enigma case study’s own project page notes that the published materials verify the calculations but do not prove the message’s historical identity or that the solution is unique; independent expert review is still the arbiter. And a 170-character message verified against two log entries, however elegant, is a small sample of one model’s capabilities.

But as a signal of where things are heading, the pattern is hard to miss. Ciphers that protected state secrets for a century are falling not to new mathematics but to an AI’s willingness to re-examine assumptions nobody thought to question — and then to do the unglamorous archival work of proving itself right. The machines that could not be bothered to check their sources a few years ago are now the ones insisting on checking the original logs.