← All posts / Research

40 Measurements, 4 Interventions: How GPT-5.6 Sol Learned to Calibrate MIT's Qubits Overnight

A new MIT case study shows GPT-5.6 Sol running through Codex autonomously characterizing a never-calibrated six-qubit superconducting chip — discovering resonators, fitting coherence times, and needing human help on only 4 of 40 target measurements. The agents are slower than PhD students, but they work overnight, and that changes the economics of a quantum lab.

40 Measurements, 4 Interventions: How GPT-5.6 Sol Learned to Calibrate MIT's Qubits Overnight

On September 9, 2026, OpenAI published a case study that reads less like a model demo and more like a shift in how a physics laboratory actually runs. Researchers at MIT’s Engineering Quantum Systems Group (EQuS) connected GPT-5.6 Sol — running through Codex on the Ultra reasoning level — directly to the software that controls a dilution refrigerator, and let it characterize a superconducting-qubit chip that had never been measured before. The agent found the resonators, calibrated the readout pulses, measured coherence times, and updated the lab database itself. Out of 40 target measurements on the chip’s four fixed-frequency qubits, human researchers intervened to improve only four.

The paper, “Case Study: Agentic Calibration of Superconducting Qubits,” is authored by Beatriz Yankelevich, Eva Zhang, JonLuca DeCaro, Jeffrey A. Grover, and William D. Oliver — with the MIT authors having developed the agentic measurement infrastructure and performed the experiments, and the OpenAI authors providing advice on the infrastructure and comments on the text. It is the most concrete published account yet of an AI agent operating real laboratory hardware end-to-end, not just writing analysis code for humans to run.

Why qubit calibration is the perfect agentic grind

Superconducting-qubit experiments have an unusual property: once a chip is fabricated, packaged, and cooled to millikelvin temperatures, nearly everything is controlled by software. Pulses generated by room-temperature instruments travel down cables into the refrigerator, interact with the quantum circuits, and return as signals that are amplified, digitized, and analyzed. That makes the lab an unusually good testbed for agents — the actuator problem reduces to API calls.

But calibration is not a fixed script. The measurements form an interdependent chain: resonator spectroscopy yields frequencies that define the pulse-calibration steps, which in turn gate the coherence measurements. An incorrect value early in the process poisons everything downstream. Each chip also behaves differently despite identical design, because fabrication imperfections and environmental interactions shift every parameter. Calibrating a novel multi-qubit experiment can take months and thousands of preliminary measurements — and, as the paper dryly notes, only a handful of those measurements ever make it into the publication.

The chip in this study was deliberately modest: six uncoupled qubits (four fixed-frequency, two tunable), each with a coupled readout resonator, of a type EQuS regularly produces to benchmark its fabrication process. The point was not scale. The point was whether a general-purpose agent could work through the same workflow a human researcher would, using the lab’s existing orchestration software.

How the agent was set up — and what it actually did

EQuS did not build a bespoke harness. The agent ran through the Codex app with nothing beyond a simple in-house Jupyter MCP server, giving it access to live measurement parameters, programs, plots, raw data, logs, and the measurement database. It could analyze and plot data in the notebook and keep its own laboratory notes in Markdown files.

The real engineering effort went into context. After several months of iteration, the researchers converged on a recipe: experimental setup details, chip designs, access to the orchestration software’s source code, and — critically — measurement-specific “skills.” Each skill documents the template code for execution and analysis, which calibrations must be completed first, tips for choosing good parameters, common physical reasons the measurement might fail, and example plots of both successful and failed runs.

With that context, the agent’s behavior looks strikingly like a competent student’s. In resonator discovery, it correctly identified six small resonator features against large background fluctuations, noting: “[these measurements] isolate six regularly spaced RF-stable candidates at approximately 7.2920, 7.3770, 7.4715, 7.5580, 7.6580, and 7.7520 GHz. That spacing is consistent with a six-resonator multiplexed design. I’ll provisionally assign them in ascending frequency to r1-r6.” It then zoomed in on each resonator, ran 2D power-frequency sweeps to find the “punchout” behavior, and — in a line that could come from a careful postdoc — explained it was “taking one higher-stat narrow trace per resonator there to get defensible centers and linewidths before the flux maps; this avoids fitting κ from the low-SNR tail of the 2D sweep.”

It then worked through the standard suite on the fixed-frequency qubits: cavity calibration, qubit spectroscopy, power Rabi, readout optimization, T1 and T2 echo/Ramsey fits, chi shifts, and effective qubit temperature (one qubit came out at 53.04 mK ± 2.69 mK, consistent with the mixing-chamber stage of the fridge). Across the 40 target measurements, researchers intervened on only four.

The honest limitations

The paper is refreshingly candid about where the agent falls short. Anecdotally, agents take longer than experienced researchers even when they converge on correct parameters. They often “lack experimental intuition” — pursuing an incorrect line of investigation, or missing an obvious physical reason a measurement failed. And because each measurement takes minutes and runs serially, physical acquisition is the rate-limiting step: agent swarms cannot brute-force the problem with parallel exploration. The authors suggest model fine-tuning may be needed to confer that experimental intuition.

There is also a scale caveat that matters. Six uncoupled qubits is a benchmark chip, not a processor. Once qubits are coupled, calibration complexity grows sharply — desirable and undesirable interactions both must be accounted for — and novel experiments lack the well-defined skill documentation that made this demonstration work. When signals were weak, noisy, or physically ambiguous, the agent struggled, and Yankelevich retains tighter control over genuinely novel measurements, giving agents narrower goals.

Why “slower but overnight” changes the economics

The immediate value proposition is not speed. OpenAI acknowledges researchers may still find the best settings faster than current models. The value is that the agent does not need supervision. Characterizing one of these standard chips occupies a researcher for about a day (fixed-frequency) or a week (tunable). The EQuS group now routinely uses agents to characterize chips overnight, while researchers interpret results, design experiments, and plan next steps. Since multiple chips can be measured simultaneously in one refrigerator — limited mainly by cabling and control-electronics ports — agent labor increases characterization throughput without increasing staff time.

That matters because of how fabrication-driven the field is: groups often must characterize many nearly identical chips to find the best one for the real experiment. Automating the weeding-out process is unglamorous and enormously valuable.

The broader signal for AI-in-science is the pattern, not the qubits. This is the same recipe playing out across laboratories: a general-purpose model, domain-specific skill documentation written by experts, tool access to real instrumentation, and a human who steps in when the data looks wrong. A week earlier, the world was arguing about whether 10,000 agents “solving” Navier–Stokes counts as mathematics; this case study is the quieter, more replicable version of the trend — agents embedded in the daily workflow of an experimental science, doing the measurements nobody wants to do.

Yankelevich’s own summary is the one to quote: “I’ve built infrastructure to guide agents through several parts of my work — measurement, theory, and chip design — and now it’s really starting to pay off. I can have multiple agents working on different problems at once, and I spend most of my time on higher-level work — interpreting results, devising experiments, planning next steps for the agents, reading, and writing.”

That is not a lab replaced by AI. It is a lab where the graduate students finally get to do the scientist’s part of the job.