← All posts / Tools

Anthropic's Model Hardware Standard Gives AI Agents Hands in the Lab

Anthropic's new Model Hardware Standard (MHS) lets AI agents orchestrate microscopes, liquid handlers, and robotic arms through one shared interface — cutting integration from months to hours.

Anthropic's Model Hardware Standard Gives AI Agents Hands in the Lab

On August 27, 2026, Anthropic opened a research preview of the Model Hardware Standard (MHS) — a shared specification for letting AI agents safely operate physical devices — to a first group of scientific research labs and advanced manufacturers. It is the company’s most concrete step yet from chatbots and code into physical AI, and it arrives with an unusually detailed set of early results from Genentech, the University of Washington, Carnegie Mellon, HHMI Janelia, QuEra Computing, and Tetsuwan Scientific.

The problem MHS actually solves

Walk into any modern biology or materials lab and you’ll find instruments from a dozen vendors: liquid handlers, plate readers, robotic arms, microscopes, lasers. Each comes with its own programming interface, its own data format, and often its own language — one rig described in Anthropic’s announcement spans MATLAB detectors, Python cameras, and C# electrophysiology. Integrating them takes specialists weeks or months of bespoke glue code, and once connected, there is still no standard way for an AI agent to share data with them or operate them safely.

MHS attacks this with three pieces:

  1. A standardized driver. Software that translates between the OS and each hardware device using simple primitives — “read” (get temperature) or “write” (set temperature) — that any device can act on. Devices become discoverable in a standard format, so agents and instruments can find each other across a network without a bespoke translator in between.
  2. Natural-language device descriptions. The driver carries tags where users write machine characteristics in plain language — the weight of a robot arm, for example, or safety limits that must be enforced. From these tags the driver auto-generates a reference file describing what a device can measure, what can be adjusted, and what limits apply. Crucially, this captures knowledge that has traditionally lived in paper manuals, local files, or a technician’s head.
  3. Three control mechanisms. Agents reach hardware via MCP (the Model Context Protocol), the command line, or code files (APIs), which work together to orchestrate multiple devices “via a single line of code.” For long-running or high-speed tasks, agents can chain driver commands into deterministic scripts the devices execute themselves.

The design is deliberately model-agnostic — it works with any agent harness, not just Claude — and applies to any device with a programmable interface. It builds on MCP, the open protocol Anthropic released in 2024 that became the de facto standard for connecting agents to software tools. Think of MHS as MCP’s sibling for atoms instead of bits.

What early partners actually measured

The research-preview partners published numbers, and they’re the most convincing part of the announcement:

  • Genentech used MHS with Claude to automate the BCA protein assay, coordinating a liquid handler, robotic arm, and plate reader. Claude autonomously optimized pipetting flow rates against an expert’s “ground truth” transfers — converging on ~140 µL/s for water (0.016 RMSE) and 10 µL/s for viscous BSA protein solution (0.181 RMSE). It also recovered on its own from tip-pickup failures and fluid-detection errors, a capability most scientific instruments simply lack.
  • University of Washington (Baker and Pinglay labs) connected six instruments in under a week, including writing the drivers. A LeRobot-based open-source robotic arm coordinated with a liquid handler for collision-free plate handoffs, and an agent-supervised qPCR watched amplification curves and halted the reaction at exactly the right moment. PhD student Zihao Song’s verdict: the time he used to spend monitoring qPCR curves at 4 a.m. now goes to planning experiments.
  • Carnegie Mellon University ran serial-dilution dose-response experiments roughly three times faster than before, orchestrating a CyBio Felix liquid handler, Varioskan LUX plate reader, Spinnaker robotic arm, and monitoring cameras spread across three computers with fundamentally incompatible interfaces (an ActiveX/COM-era scripting interface, a directory-watcher job queue, and a GUI-only reader with no API at all). Total integration time: about eight hours versus weeks for a vendor-built setup. When the first run produced a saturated curve (R² < 0.9), the agent independently discarded the plate and reran with a compressed concentration range, achieving R² > 0.98 with zero human input.
  • HHMI Janelia unified a two-photon microscopy rig that previously required launching seven vendor programs in a fixed order. MHS puts each device’s variables into a shared-memory state dictionary; adding a new camera went from a multi-day project to minutes. The rig’s origin story is notable — MHS grew out of a collaboration between Anthropic’s Alek Kemeny and Janelia postdoc Arco Bast, whose custom shared-memory dictionary became the seed of the standard.
  • QuEra Computing handed Claude control of laser stabilization in its neutral-atom quantum computers. The existing expert-built recovery script worked 58% of the time in ~150 seconds per attempt. After an overnight unattended optimization loop, Claude’s rewritten decision-tree controller relocked the laser in ~6 seconds with 96% success — and a blind test across 700 induced disturbances hit 99.3%. Claude then tuned 12 interdependent PID parameters over 16 unattended hours, cutting residual noise from 15.7 mV to 1.55 mV and beating a human specialist’s tune on a 220 kHz resonance by three orders of magnitude.
  • Tetsuwan Scientific used MHS to orchestrate a qPCR workflow profiling fecal contamination in California’s San Pedro Creek — when a camera detected bubbles in viscous master mix, the system scanned the network for MHS-connected devices that could help and spun the tube in a centrifuge to fix it. Over 9,143 test dispenses, their MHS-refined compiler model predicted liquid-transfer precision ~12% more accurately than the manufacturer’s own spec sheet.

Hardware vendors are already lining up

Also significant: vendors are building MHS support into their equipment natively. AWS will support it through Strands Robots; Automata is adding it to LINQ; Danaher is exploring it for smart instruments and autonomous labs; Doosan Robotics is testing it on robot arms for automated QA; MBF Bioscience is building a driver for ScanImage, the software running laser-scanning microscopes in hundreds of neuroscience labs; QIAGEN has a proof-of-concept on QIAsymphony Connect; Tecan is adding support to Fluent liquid handlers; Universal Robots has had early access. On the open-source side, Hugging Face is adding MHS support to LeRobot and Raspberry Pi is enabling integration across several products after successful tests with a Camera MHS Driver.

Honest limits

Anthropic is unusually candid about where MHS and current models fall short. Claude still struggles with physical intuition: at Genentech, when bubbles caused runtime errors, its instinct was to retry in the same well with different parameters — which only agitated the liquid further until researchers explained the physics. At QuEra, Claude sometimes paused overnight waiting for human approval of slightly risky actions, and couldn’t troubleshoot problems rooted in hardware rather than code. MHS also doesn’t yet work with devices that lack any programming interface, and the standard won’t be open-sourced until the preview produces safety evaluations and deployment guidance — Anthropic says it is building a physical-safety roadmap to extend its safeguards policy to misuse risks in the physical world.

Why it matters

The MCP analogy is the right frame. Before MCP, every agent-to-tool connection was a bespoke integration; after, it became a commodity. MHS wants to do the same for the physical world, where the integration problem is even worse and the payoffs — round-the-clock autonomous experiments, self-tuning instruments, labs that generate their own scientific data — are larger. If the vendor list above is any indication, the industry is ready for a standard rather than another walled garden. The fact that it’s model-agnostic and MCP-native gives it a real shot at becoming infrastructure rather than a Claude feature.

The research preview is application-only for now, with open-source release promised after safety evaluations. For labs and manufacturers, the pitch is simple: integration work that took weeks or months now takes hours or minutes — and your instruments get an agent that never sleeps.