Ask Devin: One Researcher and an Agent Swarm Just Factored RSA-260 and Made Breaking RSA 10x Cheaper
Cognition's Eric Lu drove up to 18 concurrent Devin sessions to build the world's fastest GPU lattice siever, factoring the 862-bit RSA-260 challenge number in 15 days on spare cluster compute for roughly $400,000 — and putting RSA-1024 within reach of any well-funded lab for about $30 million.
For twenty-six years, the RSA Factoring Challenge list has had a ceiling: RSA-250, factored in February 2020 by a team of six academic cryptographers using roughly 2,700 CPU core-years. On September 3, 2026, that ceiling moved — not because of a breakthrough in number theory, and not because of quantum computers, but because an AI coding agent got very good at GPU performance engineering. Cognition, the company behind the Devin software-engineering agent, has now published the full methodology: RSA-260, a 260-digit (862-bit) challenge number, was factored by a heavily GPU-optimized implementation of the general number field sieve (GNFS) — built, tuned, and operated largely by Devin, at roughly 10x lower cost than the previous public state of the art.
The human at the wheel was Eric Lu, a researcher at Cognition and a factoring hobbyist for the past decade. His framing of what happened is unusually candid: “My role was primarily to set priorities, establish benchmarks, and recognize when work was going off-track. Devin otherwise autonomously handled measurements, cluster operations, and optimization end-to-end.” What would previously have been a multi-month effort by a team of specialized experts — the intersection of GPU kernel experts and number field theory experts is famously tiny — compressed into about three weeks.
What was actually broken
RSA-260 is a 260-digit semiprime from the RSA Laboratories factoring challenge, the benchmark for how feasible it is to break RSA keys by factoring the modulus. The new factorization sets a record for the largest publicly solved challenge number. Before anyone panics: standard RSA keys today are 2048-bit (~617 digits), and 1024-bit RSA was deprecated back in 2013. The sky is not falling for TLS. But the economics of the record did move, and by more than the digit count suggests.
The total cost: about 4,900 GPU-days, or 13.5 GPU-years — roughly $400,000 at market GPU prices. The final run consumed 1,344,878 seconds (~15.6 days) from the start of polynomial selection on August 18 to factors in hand on September 3. The time breakdown across GNFS stages: 643 GPU-days in polynomial selection (which Lu sheepishly calls “anomalously high basically due to operator incompetence”), 3,813 GPU-days in lattice sieving, and 467 GPU-days in sparse linear algebra.
The more consequential number is the extrapolation. By standard GNFS scaling, RSA-1024 is only 78x more computation than RSA-260. Lu’s estimate: a hyperscaler or frontier lab could factor an RSA-1024 modulus for on the order of $30 million — and with further optimization, “likely substantially less.” His own implementation remains “significantly suboptimal,” and he would not be surprised if moderate additional work halved that figure again.
Two honest caveats from the post itself. First, RSA-1024 being insecure is old news — there has been speculation since the mid-2000s (the TWIRL device, Bernstein’s factorization circuits) that the NSA could already do it economically. What changed is accessibility: you now need commodity GPUs and a capable coding agent, not custom silicon. Second, RSA-2048 remains roughly a billion times harder than RSA-1024 under GNFS scaling and “does not appear to be meaningfully affected by this work.”
How a hobby project became a cryptography record
The origin story is a cluster-utilization problem. Cognition’s LLM training and inference clusters use NVL72 racks — 18 computers interconnected by NVLink — and the job scheduler’s packing constraints leave a single-digit percentage of compute stranded as idle nodes. Lu rigged the scheduler to fill those gaps with low-priority single-node jobs, then went looking for a workload that is embarrassingly parallel, preempt-safe, and computationally bottomless. Lattice sieving — the most expensive stage of GNFS, long done only on CPUs — fit perfectly. All that was missing was a GPU lattice siever fast enough to handle RSA-260 parameters.
So, in Lu’s words: “Ask Devin.”
On August 13 at 0:11:58 Pacific, he gave Devin the task: build a drop-in GPU replacement for las, CADO-NFS’s CPU lattice siever, iterating on a rented GPU box until it beat the CPU version. Two hours later he added one requirement — it should handle the parameters used for RSA-250. “Then I went to bed. I woke up to find that, after another 7 hours of iteration, Devin had succeeded.”
What followed was a five-day sprint up a scaling ladder of real factoring targets — C155, C173, C190, a C311 from a repunit number, a C344 odd-perfect-number roadblock — until factoring a 190-digit number took the same ~3 hours a 157-digit number had at the start. Along the way there is a wonderfully human footnote: a C190 factored “just to startle a Devin who randomly generated a semiprime.”
The final RSA-260 pipeline modified almost every component of CADO-NFS, the open-source GNFS implementation that made the whole project possible: a GPU-adapted stage-1 polynomial selection (with components from msieve), the new GPU siever, a GPU-optimized block Wiedemann linear algebra implementation, accelerated square-root reconstruction, parallelized filtering stages. Only four upstream programs survived untouched. Crucially, Lu reports “essentially no algorithmic advancements” — this is classic systems engineering, exploiting “the preposterous memory systems of the GPU.” The 10x cost win comes from hardware utilization, not new mathematics.
The division of labor that matters
The most carefully documented part of the post is what the human still did. Over three weeks, Lu ran an average of 3 and a maximum of 18 concurrent Devin sessions — 192 of 233 sessions were his, totaling 14,450 ACUs, with 101 child sessions started by the Devins themselves and 36 requiring no intervention at all. Devin claims Lu sent 82,702 words across 3,328 messages. His summary of his own role: executive function. Setting the hierarchy of goals. Recognizing unproductive measurement loops and redirecting. Suggesting untried directions (“make sure the GPU never blocks on the CPU,” “can you use NVLink SHARP here?”). And above all, forcing into existence a unified, comparable set of benchmarks — “which evidently were not otherwise going to self-assemble.”
Lu’s reflections avoid both hype and false modesty. He describes current agents as “somewhat like a sewing machine or a loom… I push it along in some way; it evidently could not happen without me, but neither am I throwing the shuttle by hand.” He notes a genuine loss: he did not learn as much number theory and GPU programming as he would have doing it by hand, and thinks the community needs to deliberately allocate resources to preserve human understanding. And he names the uncomfortable truth hiding in the timeline: three weeks from first prompt to record is a capability overhang whose extent we are only beginning to explore.
Why this lands now
This record arrives in the same week the field is arguing about OpenAI’s 10,000-agent Navier–Stokes claim and Terence Tao’s warning that AI-driven “problem mining” could strip-mine mathematics. RSA-260 is the constructive counterpoint: no priority dispute, no opaque claims, full methodology and cost appendices published, credit explicitly shared with the open-source CADO-NFS community and the online factoring world (GIMPS, mersenneforum, FactorDB) that kept the hobby alive.
The security takeaway is not “RSA is broken.” It is that the cost curve for offensive cryptanalysis is now set by whoever has idle GPUs and a good coding agent — and idle GPUs are something frontier labs have in increasing abundance. Lu’s estimate that the implementation could still improve 2x, applied to an already 10x-improved baseline, means the practical price of factoring 1024-bit keys is falling faster than the deprecation schedules assumed. Any system still carrying 1024-bit RSA in 2026 should treat this as the closing bell.
The deeper signal is for research itself. As Lu puts it: if a problem can be solved by “just programming,” it is now worth attempting immediately. The barrier to entry for cryptanalysis, computational mathematics, and large-scale scientific computing just dropped by whatever multiplier you assign to replacing a specialized team with one determined engineer and a swarm of agents.