← All posts / Tools

Can AI Design Circuit Boards Yet? EEBench Says the Answer Is Already Partly Yes

EEBench, a new benchmark that grades AI agents on real circuit-board design in declarative atopile code with SPICE simulation and real component tolerances, puts Claude Opus 5 on top at 61.6% — while OpenAI's GPT-5.5 and GPT-5.6 Sol trail at 42.3% and 39.4% and GPT-6 Astra remains unscored.

Can AI Design Circuit Boards Yet? EEBench Says the Answer Is Already Partly Yes

When OpenAI put a demo of GPT-6 Astra working on a circuit board in KiCad on the front page of its launch post, it did more than show off a model. It put a question on the table that the electronics world has been quietly asking for two years: how do we measure whether the circuits coming out of these models are actually any good? On September 4, 2026, the team behind EEBench published their answer — a benchmark that grades AI agents on real hardware design, with deterministic SPICE simulation, real manufacturer parts, and a leaderboard that already separates the frontier pack in ways coding benchmarks cannot.

The September 1 results are striking. Claude Opus 5 leads EEBench V1 at 61.6% across 13 tasks, followed by Grok 4.6 at 57.1%, Claude Fable 5.1 at 56.4%, Claude Fable 5 at 54.3%, and Claude Opus 4.8 Max at 51.4%. OpenAI’s models sit further down the table: GPT-5.5 scored 42.3%, and GPT-5.6 Sol scored 39.4%. There is no GPT-6 Astra score yet — the EEBench team says that after watching it manipulate a PCB in KiCad, they would very much like to find out. A few months ago, the authors write, they would not have expected models to do this well.

Why circuit design needed a new yardstick

The gap EEBench fills is methodological. Current models, the team observes, know much more about electronics than their output in conventional design tools tends to show — they have read textbooks, datasheets, application notes and enormous amounts of code. The bottleneck is the interface. Ask an agent to operate a graphical CAD tool and it spends most of its effort clicking around and tracking menus; a large fraction of its context window is consumed by coordinates, UI state, and application chrome rather than electrical reasoning.

EEBench’s solution is atopile, an open-source hardware description language where the circuit lives in declarative code. The agent works directly on components, connections and electrical constraints — change the design, build it, run a simulation, inspect what failed, all without leaving the project. The benchmark, in other words, spends less time testing computer use and more time testing electronics. The approach has an obvious analogy: it is giving a hardware agent what a compiler and test suite give a coding agent, except the tests measure voltages and component behavior instead of program output.

The tasks: where textbook answers go to die

What makes EEBench genuinely hard is that it models the messy parts of real engineering. One public task is based on a residential energy meter. When its 5 V supply disappears, the circuit must keep the processor alive for another 20 milliseconds so it can save the accumulated reading — and the protected rail must stay above the processor’s 3.0 V brownout threshold throughout the window.

Most models intuitively reach the right base conclusion: add a capacitor. The real capacitor is what makes the task interesting. A ceramic part may deliver far less than its advertised capacitance once voltage sits across it. Parts carry tolerances. More capacitance costs more money, takes up board space, and slows the rail’s recharge when power returns. A design that works on nominal values can fail with the parts that actually arrive.

The benchmark’s grading is fully deterministic. It builds the submitted design, constructs the circuit graph and bill of materials, and runs SPICE simulations and design checks where each requirement produces a measurement against a limit. The energy-meter task measures the protected rail through dropout and recovery; other tasks measure gain, thresholds, ripple, transient response, and behavior at component tolerance corners, with the harness rebuilding the SPICE deck for worst-case corner simulations.

The blog post documents one illustrative failure. A submitted design used 22 µF nominally. At 4.7 V bias, the grader found only 11.4 µF of effective capacitance — far below the roughly 545 µF requirement. The source built successfully; the circuit still failed the job, with the protected rail falling below the 3 V threshold after just 0.85 milliseconds, a factor of twenty short of the 20 ms specification.

The harder analog tasks demand more. An agent may have to synthesize a multiple-feedback low-pass filter around an op-amp, solve the resistor and capacitor ratios for the required poles, and keep gain, cutoff frequency and Q inside their limits after every component is pushed to its worst-case tolerance corner. EEBench uses real manufacturer parts, with specifications extracted from datasheets and carried into the SPICE model. The agent has to find a combination that works across tolerance corners while choosing parts that exist, can be ordered, and are reasonably priced. That three-way trade-off between electrical performance, cost and supply is, as the authors put it, what electrical engineering eventually boils down to.

The leaderboard: an Anthropic advantage, a Grok surge, an OpenAI gap

Two results stand out beyond the raw ordering. First, Anthropic’s models have consistently done well in this environment — the top spot and three of the top five positions belong to Claude variants. Second, xAI has gone further than anyone in embracing the benchmark: the company included EEBench in the Grok 4.6 model card, in its “engineering acceleration” section alongside evaluations for 3D modeling and parametric CAD. Their published run put Grok 4.6 at 60.0% with high reasoning effort, slightly above the 57.1% figure on EEBench’s own September 1 leaderboard. xAI attributes the performance to high-quality engineering data and reinforcement learning in domain-specific environments including computer-aided design, and the EEBench result fits that story.

The OpenAI gap is the leaderboard’s most consequential absence. GPT-5.5’s 42.3% and GPT-5.6 Sol’s 39.4% sit roughly 15 to 22 points behind Claude Opus 5. Whether GPT-6 Astra — demonstrated manipulating a board in KiCad at launch — closes that distance is now one of the benchmark’s most watched open questions.

From benchmark to training environment

The most forward-looking part of the EEBench design is its second life as an RL environment. Once you have a simulation harness that can grade a circuit, the same checks can serve as reward signals during post-training. A failed run contains useful information: which voltage missed its limit, which operating corner failed, whether the model solved the problem with an unnecessarily expensive design. That gives a training loop far more to work with than a model’s verbal assurance that a schematic looks plausible.

EEBench V1 covers analog and digital design through simulation. It does not yet test layout, manufacturing, or board bring-up — the authors want to add those later, and concentrate today on the requirements-design-verification loop where engineering work can already be graded objectively. The full methodology and a sample result explorer are public, and the benchmark is built and funded by the team behind atopile, which pays for the public runs and does not sell benchmark scores.

So, can it design a circuit board?

The EEBench team’s own verdict is carefully hedged but unmistakably positive: for a useful and growing set of circuit problems, the answer is already yes. Nobody should ask a model to design a pacemaker and blindly install the result. But with OpenAI choosing a PCB for one of Astra’s first public demos, xAI publishing EEBench in a model card, and Elon Musk saying Grok 4.7 is weeks away after training on a large collection of SpaceX engineering data, hardware capability is moving from curiosity to competitive frontier. The benchmark authors promise to keep adding harder tasks as models improve. For an industry that has measured AI almost exclusively through text and code until now, EEBench is an early map of territory that is about to matter a great deal more.