Anthropic's Model Hardware Standard Gives AI Agents Hands: MHS Is MCP for the Physical World
Anthropic's new Model Hardware Standard (MHS) lets AI agents safely discover, operate, and orchestrate lab and factory equipment — from Genentech liquid handlers to QuEra quantum lasers — cutting integration time from months to hours.
On August 27, 2026, Anthropic opened a research preview of the Model Hardware Standard (MHS) — a shared specification for AI agents to safely operate physical devices. If that sounds familiar, it should: in November 2024 Anthropic open-sourced the Model Context Protocol (MCP), which became the de facto standard for connecting AI models to software tools and data. MHS is the sequel, and it points at something much bigger: the physical world. Where MCP gave agents access to APIs and documents, MHS gives them hands — a standardized way to discover, understand, control, and orchestrate microscopes, liquid handlers, robotic arms, plate readers, and even the laser systems inside quantum computers.
The first group of users is drawn from scientific research labs and advanced manufacturers, and the early results Anthropic published alongside the announcement are striking: integration work that normally takes weeks or months collapsed to hours, agents recovering from hardware errors without human help, and one quantum-computing pilot where an overnight agent loop produced a laser-recovery controller that beats a months-long human engineering effort.
The integration problem MHS kills
Every lab and factory floor has the same nightmare: devices that don’t talk to each other. Each instrument ships with its own vendor software, its own programming interface, often its own language — MATLAB here, Python there, C# elsewhere. Getting two instruments to coordinate requires bespoke “translator” code written by specialists, and there has been no standardized way to then hand control of those devices to an AI agent, let alone do it safely. Setting up an automated system typically takes weeks to months.
MHS addresses this with a standardized driver — software that translates between a computer and a hardware device. The driver exposes a small set of primitives, “read” (get temperature) and “write” (set temperature), that any device can understand, and makes each device discoverable on a network in a standard format, so agents and devices can find each other without a bespoke translator in between.
The clever part is how MHS handles physical context that an agent can’t infer from code alone. The driver carries tags, written in natural language, describing machine characteristics — the weight of a robot arm, for instance, which matters enormously for manipulating it safely. Historically this knowledge lived in paper manuals or in the heads of operators. The driver uses these tags to auto-generate a reference file describing what a device can measure, what can be adjusted, and which safety limits are enforced. Once connected, an agent can control hardware through three mechanisms — MCP, a command line interface, and code files (APIs) — and orchestrate multiple devices in a single line of code. For long-running or high-speed tasks, the agent can chain driver commands into deterministic scripts that execute without step-by-step reasoning. The standard is model-agnostic: any agent harness can use it.
MHS began as a collaboration between Alek Kemeny on Anthropic’s Beneficial Deployments team and Arco Bast, a postdoctoral scientist at HHMI Janelia Research Campus who had built a shared-memory dictionary to make his brain-imaging rig’s instruments communicate. Anthropic plans to open-source the standard after the preview.
What the pilots actually did
Genentech ran the BCA protein assay — a routine measure of protein concentration requiring a liquid handler, a robotic arm, and a plate reader — with Claude orchestrating all three. Asked to optimize pipetting flow rates, Claude ran trial transfers, scored itself against an expert’s “ground truth” plate using RMSE, and converged on ~140 µL/s for water and 10 µL/s for viscous BSA solutions, parameters Genentech’s automation experts confirmed were reasonable. Claude also recovered on its own from tip-pickup failures and fluid-detection errors — a capability most scientific instruments simply lack.
The pilot also exposed the limits. When bubbles foamed up in protein samples, Claude’s instinct was to retry in the same well with different parameters, which only agitated the fluid further. It didn’t understand the physics until researchers told it to move to a clean well and cut mixing cycles. Those lessons were then codified into reusable “skills.”
Carnegie Mellon University built drivers from scratch for a CyBio Felix liquid handler, a Varioskan LUX plate reader, a Spinnaker robotic arm, and monitoring cameras — spread across three computers with fundamentally incompatible interfaces (a directory-watcher scheduler, a legacy Windows ActiveX/COM interface, and a GUI-only program with no API at all, which MHS drives the way a person would). Total time from raw equipment to a completed serial-dilution dose-response curve, including one autonomous rerun: eight hours, versus the weeks a vendor-built setup typically takes — and roughly 3× faster experiment execution. When the team induced six failure conditions (missing plate, rotated plate, busy reader, disconnected camera, unreachable device, active e-stop), the system correctly blocked all six before any device moved. On science quality: the agent rejected its first dose-response run (R² < 0.9 due to saturation), decided on its own to discard the plate, compressed the concentration range from 200 to 100 µg/mL, and produced a clean R² > 0.98 fit with zero human input.
QuEra Computing, the neutral-atom quantum computer company, handed Claude control of the titanium-sapphire laser system that must hold its frequency to roughly one part in a trillion. A bespoke recovery script — built over months by a four-person team — relocked the laser 58% of the time in ~150 seconds. QuEra then ran an overnight loop of four Claude roles (hypothesize, edit, execute, review) that iterated hundreds of times unattended. By morning the rewritten controller recovered the lock in ~6 seconds with 96% success; in a later blind test across 700 trials it succeeded 695 times (99.3%), with even the hardest disturbances taking 10–14 seconds versus 5–10 minutes for a human. The agent also retuned the servo loop’s 12 PID parameters over 363 experiments and 16 unattended hours, cutting residual noise from 15.7 mV to 1.55 mV — roughly 10× quieter than the human specialist’s tune — and holding the lock for 19 hours straight where the manual tune dropped ~1.6 times per hour.
Other partners: the University of Washington’s Baker and Pinglay labs connected six instruments in under a week (including hand-written drivers) for remote monitoring, agent-supervised qPCR, and collision-free plate handoffs between a LeRobot arm and a liquid handler; HHMI Janelia unified a seven-program microscopy rig into one shared state dictionary, shrinking hardware integration from days to minutes; and Tetsuwan Scientific used MHS in its ResearchOS automated biology platform for a citizen-science qPCR project profiling fecal contamination in California’s San Pedro Creek — during which Claude spotted bubbles in a reagent tube, scanned the network for a compatible centrifuge, and spun them out on its own.
An ecosystem forming on day one
Hardware vendors are already building in support: AWS (via its Strands Robots library), Automata (LINQ), Danaher, Doosan Robotics, MBF Bioscience (ScanImage, which runs laser-scanning microscopes in hundreds of neuroscience labs), QIAGEN, Tecan (Fluent liquid handlers), and Universal Robots. On the open-source side, Hugging Face is adding MHS to LeRobot and Raspberry Pi is enabling integration across a range of products.
Why it matters — and what to watch
MCP’s lesson is that standards compound: once connecting an agent to a tool became boring, an entire ecosystem of agentic software formed. MHS aims for the same dynamic in labs and factories, where the bottleneck has never been model intelligence but integration cost and safety. Anthropic is explicit about the gaps: Claude learns physics from text and images, so spatial reasoning still requires expert oversight; MHS only works with devices that have a programmable interface; and the company is building a physical safety roadmap alongside the preview, with safety evaluations and open-source release to follow.
The direction, though, is unmistakable. The most valuable pilot results aren’t the speedups — they’re the behaviors: an agent rejecting its own data quality, rerunning an experiment unprompted, blocking unsafe states before hardware moves, and turning exploratory trial-and-error into deterministic, inspectable scripts. That’s the anatomy of an autonomous lab. For an industry that has spent two years automating keyboards, MHS is the first serious, vendor-backed attempt to standardize the other side of the loop — machines that move.
Sources
- [1] https://www.anthropic.com/news/model-hardware-standard-research-preview
- [2] https://arstechnica.com/ai/2026/08/anthropics-new-hardware-standard-lets-ai-agents-control-the-physical-world/
- [3] https://www.cnbc.com/2026/08/27/anthropic-pushes-into-physical-world-with-new-standard-to-help-ai-agents-operate-machines.html
- [4] https://qz.com/anthropic-model-hardware-standard-ai-robots-lab-equipment-082826
- [5] https://mlq.ai/news/anthropic-previews-model-hardware-standard-for-ai-controlled-lab-and-factory-equipment/