Caterpillar and FieldAI are collaborating on physical AI for construction, manufacturing and industrial operations. The deeper idea is bigger than adding autonomy to one machine: FieldAI is building robot-agnostic foundation models designed to operate across different embodiments, while Caterpillar brings heavy-equipment engineering, operational data and large-scale industrial environments. Early applications include autonomous inspection, situational awareness, digital twins and operational optimization using NVIDIA accelerated computing and Omniverse technologies. The important question is whether one risk-aware autonomy stack can transfer across machines, tasks and changing jobsites without requiring every deployment to be engineered from scratch.

Physical AI Gets More Interesting When the Machine Weighs 30 Tons

AI in robotics is easy to demonstrate on a small platform.

A robot dog walks through a room.

A humanoid picks up a box.

A mobile robot follows a route.

Heavy industry is different.

Construction sites change constantly.

Machines operate around people, vehicles, dust, mud, slopes, temporary structures and unfinished terrain.

The cost of a bad decision is much higher.

That is why Caterpillar’s new collaboration with FieldAI is interesting.

The goal is not simply to make one excavator autonomous.

It is to test whether a general-purpose autonomy layer can work across complex industrial environments and multiple types of machines.

Caterpillar Announced the Collaboration on September 2

Caterpillar announced the FieldAI collaboration on September 2, 2026.

The companies say they will work together on physical AI, autonomy, robotics and digital twins.

Caterpillar brings heavy-industry engineering, operational knowledge and data from real jobsites and manufacturing environments.

FieldAI brings robot foundation models and autonomy software designed for unstructured environments.

NVIDIA technologies provide part of the simulation and accelerated-computing layer.

The collaboration is early.

Caterpillar has not announced a specific autonomous excavator, loader or truck product tied to the deal.

The announcement is about a technology direction and development partnership.

FieldAI’s Main Claim Is One Brain Across Many Robots

FieldAI describes its platform as a general-purpose robot brain.

The company’s Field Foundation Models are designed to work across different robot embodiments rather than being built around one fixed machine.

That is a large claim.

Industrial robotics has historically been highly machine-specific.

A system is configured for one arm.

One vehicle.

One sensor arrangement.

One environment.

FieldAI wants to move toward a shared autonomy layer that can transfer more knowledge across platforms.

If that works, adding autonomy to a new machine could become less like writing a new control stack from zero and more like adapting a common intelligence layer.

Robot-Agnostic Does Not Mean Hardware-Agnostic

The phrase robot-agnostic needs a careful interpretation.

A foundation model cannot ignore physics.

An excavator does not move like a quadruped.

A wheeled inspection robot does not have the same sensing or control limits as a haul truck.

Actuators differ.

Mass differs.

Braking distance differs.

Sensor placement differs.

Safety envelopes differ.

The useful meaning of robot-agnostic is that the higher-level autonomy architecture can be reused across embodiments while lower-level interfaces and constraints remain machine-specific.

The model may share reasoning and perception.

The machine still keeps its own physics.

FieldAI Builds Around Uncertainty, Not Perfect Maps

Traditional autonomy often depends on structured assumptions.

Known maps.

Defined lanes.

Preplanned routes.

Controlled environments.

FieldAI positions its system around the opposite problem.

Construction and industrial sites change.

A path that existed yesterday may be blocked today.

Material piles move.

Temporary barriers appear.

People enter the scene.

Lighting changes.

The company says its Field Foundation Models combine data-driven AI with physics-based reasoning and uncertainty awareness.

The goal is to let robots operate when the world cannot be perfectly preprogrammed.

A Belief World Model Is the Core of FieldAI’s Current Architecture

FieldAI describes its EDGE platform around a Belief World Model.

The concept is that the robot maintains an internal representation of what it thinks is happening in the environment and how uncertain those beliefs are.

That matters because robots never observe the physical world perfectly.

A camera can be blocked.

LiDAR can have incomplete coverage.

A moving object may disappear behind equipment.

The system has to reason about partial information.

Risk-aware autonomy is less about pretending uncertainty does not exist and more about making uncertainty part of the decision process.

Heavy Equipment Makes Risk Awareness Essential

A lightweight indoor robot can often stop quickly.

Heavy equipment cannot.

Momentum matters.

Blind spots matter.

Terrain matters.

A machine may need several meters to stop safely.

That means perception confidence cannot be treated as an abstract model score.

Uncertainty has physical consequences.

If the autonomy stack is unsure whether an area is clear, the safest action may be to slow down, stop or request human input.

That is why risk-aware modeling is more meaningful in heavy industry than a generic claim that a model can “see” the environment.

The First Applications Are Not Full Autonomous Earthmoving

Caterpillar lists four early application areas.

Autonomous inspection.

Digital twins.

Situational awareness.

Operational optimization.

That list is revealing.

The first value does not require a machine to perform every construction task without a human.

Inspection alone can be valuable.

A robot can repeatedly capture site conditions.

A digital twin can update from those observations.

AI can identify changes or risks.

Simulation can test better workflows.

This is a more practical path than jumping directly to full autonomy.

Autonomous Inspection Is the Lowest-Risk Entry Point

Inspection is one of the strongest early use cases for physical AI.

A robot can travel through a site and collect visual, depth and other sensor data.

It can revisit the same areas repeatedly.

Humans do not have to enter every difficult or hazardous location.

The task is valuable even if the robot never moves a bucket of dirt.

That creates an adoption path.

First, the system observes.

Then it builds trust.

Then autonomy can expand into more consequential actions.

Every Inspection Run Can Also Build a Digital Twin

FieldAI’s collaboration with NVIDIA shows how inspection becomes more than inspection.

The company says robots collect multimodal data during normal missions.

Vision.

Depth.

LiDAR.

Other sensors.

That data can be transformed into high-fidelity digital reconstructions using NVIDIA Omniverse technologies.

The result is a digital twin that can evolve as the physical site changes.

A robot performing ordinary work becomes a mobile reality-capture system.

A Living Digital Twin Is More Useful Than a One-Time Scan

Traditional digital twins can become stale.

A construction site changes every day.

A factory layout changes.

Equipment moves.

Temporary structures appear.

If the digital model takes weeks or months to rebuild, it may describe a site that no longer exists.

FieldAI argues that robots can update the model continuously as a byproduct of operations.

That turns the digital twin from a project deliverable into an ongoing data layer.

The Real-to-Sim Loop Is the More Important NVIDIA Connection

NVIDIA Omniverse is not only being used to make attractive 3D models.

FieldAI describes a real-to-sim pipeline.

Robots operate in the physical world.

Their sensors capture the site.

Omniverse reconstruction tools turn that data into simulation-ready environments.

Those environments can then be loaded into Isaac Sim and Isaac Lab.

Autonomy can be tested against a reconstruction of the real site.

The results can feed back into future deployments.

Real operations create simulation.

Simulation improves the robot.

The robot returns to the real world.

That Loop Creates a Data Flywheel

Every deployment can create more training and validation data.

Every site adds different terrain.

Different obstacles.

Different lighting.

Different machine layouts.

Different failure cases.

If that data enters a common model-development pipeline, the autonomy system can improve across deployments.

FieldAI calls this a data flywheel.

The strategic advantage is obvious.

The more industrial environments the system sees, the harder it becomes for a new competitor to reproduce the same diversity of real-world experience.

Caterpillar Brings a Different Kind of Scale

FieldAI already deploys robotics systems across industrial environments.

Caterpillar brings another scale entirely.

Construction equipment.

Mining equipment.

Factories.

Dealer networks.

Long-lived machines.

Operational data accumulated across decades.

A partnership with a major equipment manufacturer gives an autonomy company access to use cases that are difficult to reproduce in a robotics lab.

The challenge is also larger.

A technology that works on one inspection robot is not automatically ready for a fleet of heavy machines.

Caterpillar Already Has Its Own Autonomy History

Caterpillar is not entering autonomy for the first time.

The company has long deployed autonomous mining systems and machine-control technologies.

That existing background matters.

FieldAI is not replacing an empty stack.

The collaboration adds a newer foundation-model approach to an organization that already understands industrial automation, safety engineering and large-scale equipment operations.

The interesting question is whether general-purpose models can expand autonomy beyond tightly engineered environments.

Mining Autonomy Is Easier Than a Chaotic Construction Site in One Important Way

Large mining operations can be highly structured.

Routes can be mapped.

Traffic rules can be controlled.

Access can be restricted.

Machines can operate in managed areas.

Construction sites are often less predictable.

People move through the environment.

Layouts change quickly.

Temporary materials appear.

Multiple contractors work simultaneously.

That makes general autonomy harder.

FieldAI’s value proposition is specifically aimed at those environments where traditional automation becomes expensive to configure or brittle when conditions change.

Traditional Automation Is Strongest When the World Stays the Same

Automation works beautifully in repeatable environments.

A factory robot can perform the same motion thousands of times.

A conveyor follows a fixed path.

A machine-vision camera inspects the same product.

The system becomes harder to scale when the world changes faster than engineers can reconfigure it.

Foundation-model robotics is trying to reduce that reconfiguration burden.

Instead of encoding every situation manually, the autonomy system learns broader patterns and adapts.

That promise is powerful.

It is also much harder to validate.

Generalization Is the Whole Bet

The economic argument for general-purpose autonomy depends on reuse.

If every new site requires months of custom engineering, deployment does not scale well.

If one model can transfer useful behavior from one environment to another, integration cost can fall.

The model does not have to work perfectly with zero adaptation.

It only has to reduce the amount of machine-specific and site-specific engineering enough to change the economics.

That is what Caterpillar and FieldAI are really testing.

Digital Twins Can Make Validation More Site-Specific

A general model creates a safety problem.

How do you know it will behave correctly at this particular site?

Digital twins provide one answer.

Capture the real environment.

Reconstruct it.

Run simulated missions.

Introduce edge cases.

Test routes.

Change layouts.

Evaluate policies before the robot operates physically.

The simulation does not replace real validation.

It gives engineers another layer where failures can be found before they happen around real equipment.

Simulation Is Only Useful If It Represents the Hard Parts

A clean synthetic environment is easy to simulate.

Real sites are messy.

Dust.

Debris.

Uneven terrain.

Partial visibility.

Temporary structures.

Dynamic obstacles.

FieldAI’s approach is interesting because the simulation environment begins with data captured from actual deployments.

That can preserve site-specific complexity that a manually built digital twin might miss.

The closer the simulation is to the operating environment, the more useful it becomes for regression testing and scenario analysis.

NVIDIA Supplies the Infrastructure, Not the Robot Intelligence

The partnership includes NVIDIA accelerated computing and Omniverse technologies.

The roles should stay clear.

FieldAI provides the autonomy models and robot-intelligence layer.

Caterpillar provides industrial machines, engineering and operational expertise.

NVIDIA provides simulation, reconstruction and computing infrastructure used in the development pipeline.

That three-layer structure matters because physical AI increasingly depends on an ecosystem rather than one company building everything.

Operational Optimization May Be Valuable Before Full Autonomy

Caterpillar also lists operational optimization as an early use case.

A digital twin can show how machines, people and materials move through a site.

Simulation can test alternative layouts.

AI can identify bottlenecks.

Maintenance data can be connected to site conditions.

A company may gain meaningful productivity improvements without removing human operators.

That makes physical AI easier to deploy incrementally.

Observation and optimization first.

More autonomous execution later.

Situational Awareness Is Another Intermediate Layer

Situational awareness sits between sensing and control.

A robot or machine can identify what is happening around it and surface useful information to people.

Potential hazards.

Blocked routes.

Changing site conditions.

Equipment positions.

Areas that need inspection.

The system does not need authority to perform every action.

It can make the environment more legible to operators and supervisors.

That can improve decision-making while keeping humans inside the control loop.

The Human-Machine Boundary Is Part of Caterpillar’s Message

Caterpillar frames the collaboration around combining human expertise with AI-powered machines.

That language matters.

Heavy industry will not switch from human-operated equipment to fully autonomous fleets overnight.

Mixed environments are more likely.

Some machines autonomous.

Some remote-controlled.

Some human-operated.

Robots performing inspections.

AI systems providing recommendations.

Humans handling exceptions and high-consequence decisions.

The hard problem is not only autonomous capability.

It is coordination across all of those modes.

A Universal Robot Brain Still Needs Machine-Specific Safety Layers

Even if FieldAI’s model generalizes across platforms, safety cannot be universal in the same way.

A large articulated truck has different stopping behavior from a small quadruped.

An excavator has a rotating upper structure and a large working envelope.

A robotic arm has joint limits and collision zones.

Each machine needs hard constraints close to the control layer.

The foundation model can choose a goal.

The equipment still needs deterministic limits on what is physically allowed.

This Is Where Physical AI Differs From Software Agents

A software agent can make a mistake and often retry.

A physical machine may not get a second chance.

A collision cannot be undone.

A slope failure cannot be rolled back.

A person entering a machine path changes the risk instantly.

That makes uncertainty calibration, fail-safe behavior and verification central to industrial autonomy.

The model cannot simply be impressive.

The complete system has to be predictable enough to earn operational trust.

FieldAI Says It Already Operates Across Hundreds of Sites

FieldAI says its technology has been tested and deployed across hundreds of complex industrial environments.

The company also says its deployments span multiple continents and several robot types.

Those are company-reported deployment claims.

They are useful because they show the technology is beyond a single laboratory demo.

They should not be converted into an independent measure of reliability.

Scale of deployment is not the same thing as verified safety performance.

The Company Has Raised More Than $400 Million

FieldAI announced in 2025 that it had raised $405 million across two funding rounds.

Its investor list includes major venture and strategic investors.

That capital gives the company resources to build expensive robotics infrastructure, collect field data and support deployments.

Physical AI is capital intensive.

Models need data.

Robots need hardware.

Sites need integration.

Safety testing takes time.

A large funding base matters because the product has to survive the gap between a research result and an industrial platform.

The Caterpillar Partnership Is More Important Than a Humanoid Demo

Robotics headlines often focus on humanoids because they are visually striking.

Caterpillar’s partnership points toward a different future.

The most valuable robot may be the machine already designed for the job.

An excavator already has the right body for digging.

A haul truck already has the right body for moving material.

A quadruped may be better for inspection.

The autonomy layer does not need every robot to look human.

It needs to understand enough different embodiments to use the right machine for each task.

Embodiment-Agnostic AI Could Reduce the Pressure to Build One Universal Robot

The idea of one humanoid doing everything is attractive because it simplifies the hardware story.

One body.

Many tasks.

FieldAI is pursuing another kind of universality.

Many bodies.

One higher-level intelligence architecture.

That may fit industrial reality better.

Factories and jobsites already contain specialized machines.

Replacing all of them with humanoids would be expensive and often unnecessary.

Adding a reusable autonomy layer to existing machine categories could be a faster route to general-purpose physical AI.

The Real Product May Become the Autonomy Layer

If robot intelligence becomes portable, the strategic value shifts.

The robot body becomes one part of the system.

The autonomy layer carries perception, world modeling, planning, risk reasoning and learned experience.

That starts to resemble what operating systems did for computers.

Different hardware.

A common software layer.

Robotics is not there yet.

But Caterpillar working with a company that explicitly describes itself as one brain for many robots shows where the industry wants to go.

What Caterpillar and FieldAI Have Actually Confirmed

Caterpillar and FieldAI announced a collaboration on September 2, 2026.

The companies say they will work on physical AI, robotics, autonomy and digital twins for complex jobsites and manufacturing environments.

Caterpillar lists autonomous inspections, site and facility digital twins, enhanced situational awareness and operational optimization as early application areas.

The collaboration combines Caterpillar’s industry expertise and operational data with FieldAI’s robot-agnostic foundation models.

The companies also say NVIDIA accelerated computing and Omniverse technologies will support the work.

FieldAI describes its models as combining data-driven AI, physics-based reasoning and uncertainty awareness.

No specific fully autonomous Caterpillar production machine was announced as part of the release.

What We Should Not Claim

We should not say Caterpillar has launched a fully autonomous excavator powered by FieldAI.

It has not announced that.

We should not say one FieldAI model already controls every Caterpillar machine.

The collaboration is broader and earlier than that.

We should not treat robot-agnostic as meaning the physical differences between machines disappear.

We should not claim the system eliminates human operators.

Caterpillar explicitly frames the work around human expertise and AI-powered machines.

We should not turn FieldAI’s deployment scale into an independent reliability statistic.

And we should not claim digital twins prove real-world safety.

Simulation is one validation layer, not a substitute for physical testing.

The Bigger Shift Is From Autonomous Machines to Portable Autonomy

The first generation of industrial autonomy was machine-specific.

One autonomous haul truck.

One warehouse robot.

One programmed arm.

Caterpillar and FieldAI are testing a more ambitious idea.

Build an autonomy layer that can move between machines, learn from many sites and use the physical world itself to improve simulation and validation.

If that works, the important product is no longer one autonomous machine.

It is portable autonomy.

A common intelligence layer that can understand enough about different machines to make more of the industrial world programmable.

That is a much bigger shift than putting AI inside an excavator.

Anthropic’s Model Hardware Standard is an early attempt to give AI agents a common way to discover, understand and operate programmable physical devices. Instead of writing a bespoke integration for every microscope, liquid handler, robot arm or laser controller, MHS introduces standardized drivers, simple read/write primitives, machine-readable device descriptions and device-level safety limits. Early pilots show why this matters: agents have coordinated multiple lab instruments, adapted microscopy settings in real time and helped develop a quantum-laser recovery controller that later succeeded in 695 of 700 blind trials. The important caveat is equally physical: current models still misunderstand real-world failures, so MHS is being tested as a research preview with expert oversight before Anthropic plans to open-source it.

AI Agents Have Learned to Use Software Tools — Physical Machines Are the Next Problem

AI agents are already becoming comfortable inside software.

They can open files.

Run commands.

Use APIs.

Call databases.

Control browsers.

Chain multiple tools together.

Physical equipment is much harder.

A microscope may use one vendor API.

A robotic arm may use another SDK.

A camera may expose a different interface.

A laser controller may need custom scripts written by someone who understands the hardware.

Even when every machine is technically programmable, connecting them into one reliable workflow can take weeks or months.

Anthropic’s Model Hardware Standard, or MHS, is an attempt to make that layer look more uniform to AI agents.

MHS Is Not a Robot Model

MHS is not a new Claude model.

It is not a robot operating system.

It is not a general-purpose replacement for every industrial control protocol.

Anthropic describes it as a shared specification for AI agents to operate physical devices safely.

The current version is a limited research preview being tested with scientific labs, robotics companies, electronics firms and manufacturers.

The standard is designed to work with devices that already expose a programmable interface.

It is also model-agnostic.

Anthropic says any agent harness can access MHS using standard mechanisms, including Model Context Protocol.

That separation is important.

The intelligence layer and the hardware interface are not supposed to be the same thing.

The Basic Problem Is That Every Machine Speaks Its Own Language

A research lab rarely buys every instrument from one vendor.

One camera comes with Python bindings.

A detector may use MATLAB.

Another controller may be wrapped in C#.

A microscope may require a proprietary application.

A robotic arm may expose its own SDK.

Scientists then build glue code between all of them.

The result can work.

It is also fragile.

The integration knowledge often lives in scripts, local documentation or in the memory of the person who built the rig.

MHS tries to move that knowledge into a common hardware-facing layer.

The Core of MHS Is a Standardized Driver

Anthropic describes the MHS driver as the translation layer between a computer and a hardware device.

Instead of forcing the agent to understand every vendor-specific interface directly, the driver presents a simpler set of standardized operations.

Anthropic gives basic examples such as read and write.

Read might mean get the temperature.

Write might mean set the temperature.

Real devices obviously expose richer behavior than two verbs.

The point is the abstraction.

The agent works through a predictable interface even when the underlying hardware is different.

Discovery Matters as Much as Control

Controlling a device is only useful if the agent knows the device exists.

MHS makes connected hardware discoverable in a standard format.

That means an agent can find available equipment across a network instead of depending on a custom integration written specifically for one workflow.

This sounds similar to software tool discovery.

The difference is consequence.

A mistaken software tool call may produce a bad file.

A mistaken physical call can move a robot, damage a sample or push an optical system outside a safe operating range.

Discovery therefore has to carry more than a function name.

MHS Tries to Give the Agent Physical Context the Code Does Not Contain

A hardware API can tell software which function moves a robotic arm.

It may not tell the model how heavy the arm is.

How far it can safely travel.

Which movement risks a collision.

Which setting can damage a sample.

Anthropic says MHS drivers can include natural-language tags describing machine characteristics that are difficult to infer from code alone.

The user can enter this information directly or let an agent interview them about the hardware setup.

The driver then produces a reference description of what the machine can measure, what can be adjusted and which safety limits should be enforced.

This converts some physical knowledge from tacit expertise into explicit machine context.

The Safety Boundary Is Supposed to Live at the Device Layer Too

One of the stronger design choices is that safety is not left entirely to the language model.

MHS can enforce device-level limits.

In a microscopy example at HHMI Janelia, the researcher describes using those limits to prevent the agent from applying excessive laser power that could bleach fluorescent molecules and degrade the sample.

That is the correct direction for physical AI.

Do not ask a probabilistic model to remember every safety rule on every turn.

Put critical limits closer to the machine.

The model can decide what it wants to do.

The hardware interface still decides what it is allowed to do.

MCP Is One Control Path — Not the Whole Standard

It would be easy to describe MHS as MCP for robots.

That is too simple.

Anthropic says MHS can expose hardware control through three mechanisms: MCP, a command-line interface and code files through APIs.

Those paths serve different timing requirements.

An agent can reason interactively through MCP.

It can invoke commands through a CLI.

For long-running or faster operations, it can package sequences into code so the device can execute without waiting for the model to reason at every step.

That last part is especially important for real machines.

Physical control loops often cannot pause while an LLM thinks.

The Interesting Pattern Is Agent Exploration Followed by Deterministic Code

Anthropic describes an MHS experiment in which Claude adjusted a laser, observed the result through a camera and repeated the process while learning how the beam responded.

Then the agent packaged what it learned into code.

The final alignment procedure could run as a deterministic script with one command.

That pattern may be more useful than keeping the agent permanently in the lowest-level control loop.

Let the model explore.

Let it search.

Let it infer a better procedure.

Then compile the useful behavior into inspectable deterministic software.

The agent becomes a system designer rather than a permanent joystick.

This Is a Different Vision of Physical AI

The popular image of physical AI is a humanoid robot controlled continuously by a large model.

MHS points toward another architecture.

The AI does not have to directly generate every motor command.

It can operate at a higher level.

Read state.

Select a procedure.

Adjust parameters.

Call a deterministic routine.

Observe the result.

Escalate when something unexpected happens.

This hierarchy is closer to how complex automation already works.

The model adds flexible reasoning around the deterministic machinery instead of replacing every controller with a chatbot.

A University of Washington Demo Connected a Robot Arm and a Liquid Handler

One research-preview example came from the University of Washington Baker and Pinglay labs.

The team connected a liquid handler and an open-source robotic arm through MHS.

The liquid handler dispensed reagents into a plate.

The robotic arm waited until that step finished.

Then it removed the completed plate and loaded the next one.

Claude Code coordinated the sequence.

According to the researchers, repeated tests completed without the two instruments colliding.

The agent observed completion signals before triggering the next device.

The important result is not that an arm moved a plate.

Industrial automation has done that for decades.

The interesting part is that two heterogeneous devices were orchestrated through one agent-facing layer.

The Same Lab Connected Six Instruments in Under a Week

The University of Washington researcher says connecting six instruments through MHS took less than a week, including time spent writing drivers.

That is not a universal benchmark.

It is one early case.

But it illustrates the problem MHS is trying to solve.

The same researcher describes previous automation attempts involving weeks of evaluating platforms, chasing vendor support and building glue code.

If standardized drivers can be reused across labs, integration work can compound.

One team writes a robust driver.

Another team uses it instead of starting from zero.

That is how a hardware standard can become more valuable than a one-off automation demo.

Genentech Used MHS to Coordinate Three Pieces of Lab Equipment

Genentech tested MHS on a proof-of-concept laboratory workflow involving a liquid handler, robotic arm and microplate reader.

Claude acted as the orchestration layer across the three devices.

Anthropic’s page says the time from non-automated equipment readiness to a completed dilution curve, including an autonomous rerun, was eight hours.

The comparison given by the participating team is that a vendor-delivered automated setup would typically take multiple weeks.

This is a partner-reported result from an early proof of concept.

It should not be read as a guarantee that MHS compresses every lab-integration project to eight hours.

The Failure Case Is More Important Than the Successful Demo

The Genentech pilot also exposed a weakness.

Claude encountered runtime errors caused by bubbles forming during mixing.

Its initial response was software-like.

Retry the operation.

Change parameters.

Try again.

But the physical system behaved differently.

Retrying agitated the liquid further and created more bubbles.

The researchers had to guide Claude toward a gentler physical correction.

Anthropic highlights this example itself.

That is important.

A model trained through text and images can understand an API while still misunderstanding what matter, friction, fluid and force are doing in the real world.

A Hardware Error Is Not Always a Software Error

Software agents are trained by an environment where many failures can be solved with another command.

Retry the request.

Restart the service.

Change the parameter.

Re-run the test.

Physical systems do not always forgive that strategy.

A liquid can foam.

A sample can degrade.

A motor can collide.

A laser can damage material.

A machine can overheat.

The difference is irreversibility.

An agent operating real hardware needs a richer model of cause and effect than an agent fixing a compiler error.

MHS standardizes access.

It does not automatically give the model physical intuition.

HHMI Janelia Shows Why One Shared State Can Change a Microscope Rig

MHS began partly from a real integration problem at HHMI Janelia Research Campus.

One researcher was working with a brain-imaging rig made from lasers, motorized focusers, specialized cameras and other devices from different vendors.

The early idea was to place the rig’s state in a standardized shared-memory dictionary.

That evolved into MHS.

Another Janelia project describes a microscopy setup previously spread across seven vendor programs.

With the rig exposed through a common state layer, an agent can make higher-level decisions without separately learning seven different control systems.

Agentic Microscopy Is About Choosing What to Observe Next

Microscopy makes the value of adaptive agents easier to see.

Traditional experiments often start with fixed parameters.

Image this region.

At this speed.

At this resolution.

For this long.

But biological systems change while the experiment is running.

A fixed setting may miss the interesting event.

In the Janelia work, the agent can enter at decision points and choose acquisition parameters based on what the system is observing.

That is more than automating a button press.

The experiment becomes closed-loop.

Observe.

Analyze.

Decide what to measure next.

Then change the acquisition plan.

MHS Does Not Remove the Physics

A standard interface cannot repeal physical trade-offs.

Imaging faster can reduce coverage.

More light can damage a sample.

Higher precision can cost time.

A robot can only move within its mechanical limits.

The Janelia researchers explicitly note that MHS does not remove these trade-offs.

What it changes is how quickly an experiment can move through the parameter space.

That is the right way to describe the value.

MHS does not make hardware infinitely capable.

It makes hardware easier for agents and humans to coordinate.

The QuEra Pilot Is the Strongest Demonstration

The most striking MHS example comes from QuEra Computing.

QuEra builds neutral-atom quantum computers.

The machines depend on extremely precise lasers.

If a laser loses its frequency lock, quantum operations can begin to fail.

A human expert may need five to ten minutes to restore the lock.

QuEra had already spent months building a bespoke automated recovery script.

That script worked about 58% of the time and took roughly 150 seconds per attempt.

Then the team gave Claude controlled access to the laser system through MHS.

Claude Did Not Simply Operate the Laser — It Redesigned the Recovery Procedure

QuEra used multiple Claude instances in an iterative loop.

One proposed a hypothesis.

Another modified the recovery script.

Another ran the updated script against the live laser and logged the result.

Another reviewed the logs and decided what to try next.

The cycle repeated hundreds of times overnight.

By morning, Anthropic reports that recovery took around six seconds and succeeded 96% of the time in the development run.

The important point is that the model was not only executing a fixed procedure.

It was searching for a better one.

The Blind Test Reached 695 Successful Recoveries Out of 700

After the development loop, QuEra tested the finished script against randomized induced disturbances with no agent controlling the test.

Across 700 trials, the controller recovered the correct lock 695 times.

That is the reported 99.3% success rate.

The hardest disturbances took roughly 10 to 14 seconds.

Simpler ones took less.

Again, the result needs precise wording.

This is a QuEra and Anthropic pilot on one laser system.

It is not evidence that MHS makes arbitrary hardware 99.3% reliable.

But it is strong evidence for a particular workflow: use the agent to discover a better control strategy, then deploy the resulting deterministic procedure.

The Production Artifact Was Inspectable Code, Not an Autonomous Agent

The most reassuring detail is what QuEra ended up with.

The relock controller became a deterministic, inspectable script that could run without an AI agent controlling it.

That matters for engineering.

Critical systems often need repeatability.

A deterministic controller can be reviewed.

Versioned.

Tested.

Rolled back.

The agent can still be useful during development and optimization.

But the final artifact can be ordinary software.

That architecture may be one of the most practical ways to bring frontier models into physical systems without handing them permanent unrestricted control.

Some Tasks Still Keep the Agent in the Loop

Not every task can be compiled into one fixed script.

QuEra also used Claude to tune interdependent laser parameters as environmental conditions changed.

That workflow remained adaptive.

The agent repeatedly measured system behavior, changed parameters and evaluated the result.

This shows why MHS needs both deterministic and agentic modes.

Stable procedures can become code.

Dynamic optimization can keep the model involved.

The hard engineering problem is deciding which category a task belongs in and where human approval should enter.

Research-Preview Partners Extend Far Beyond Anthropic

Anthropic lists a broad set of companies experimenting with MHS.

AWS plans support through Strands Robots.

Automata is adding MHS support to its lab-automation platform.

Doosan Robotics is testing it with robotic arms.

MBF Bioscience is building a driver for ScanImage.

QIAGEN is testing instrument troubleshooting.

Tecan is adding support for Fluent liquid handlers.

Universal Robots has had early access.

Hugging Face is adding MHS support in LeRobot.

Raspberry Pi is enabling integrations across products following tests with a Camera MHS Driver.

The list matters because a standard only becomes useful when hardware vendors actually implement it.

The Hugging Face and Raspberry Pi Links Make MHS Bigger Than Lab Automation

Hugging Face adding MHS support to LeRobot connects the standard to an open robotics ecosystem.

Raspberry Pi support points in another direction: low-cost programmable hardware.

If MHS eventually becomes an open standard with reusable drivers, the addressable hardware could extend well beyond expensive scientific instruments.

Cameras.

Robot arms.

Sensors.

Embedded devices.

Education rigs.

Prototype machines.

That broader future is still speculative.

The current research preview is focused on controlled environments.

But the choice to stay model-agnostic and work through programmable interfaces makes the architecture more general.

MHS Is Not Open Source Yet

This point should not be blurred.

Anthropic says it plans to open-source MHS.

It has not done that yet.

The current release is a limited research preview available by application.

Anthropic says the preview period will be used to test the standard, build safety evaluations and develop deployment best practices.

When MHS is eventually opened, Anthropic says it plans to publish findings from the preview alongside guidance.

So today we can analyze the architecture and the reported pilots.

We cannot treat MHS as a mature, fully open ecosystem with stable public implementations everywhere.

Programmable Hardware Is a Hard Requirement

MHS does not magically connect to every physical machine.

Anthropic says it currently requires hardware with a programmable interface.

That can be an API.

An SDK.

A software interface.

Machines with no accessible control layer still need manufacturers to expose one or build compatible drivers.

This sounds obvious.

It is actually a major deployment constraint.

Industrial and scientific environments contain enormous amounts of legacy equipment.

A new standard can reduce integration work only after there is something to integrate with.

The Physical Safety Problem Is Larger Than the Interface Problem

Standardizing the interface is technically difficult.

Standardizing safe behavior is harder.

Different devices have different hazards.

A microscope may risk sample damage.

A robot arm can collide.

A laser can exceed safe power.

An industrial machine may interact with people.

A liquid handler can contaminate a workflow.

The same abstract command—write a new value—can have completely different physical consequences.

Anthropic says it is using the research preview to build additional safety evaluations and a physical-safety roadmap.

That work may ultimately matter more than the convenience of the driver format itself.

Human Approval Still Appears in the Hard Cases

The QuEra pilot is explicit about this limitation.

Claude sometimes stopped and waited for human confirmation when it considered an action even slightly risky.

That meant experiments could pause overnight.

The team also had to provide extensive context describing the goal and how the experiment should be conducted.

This is not a fully autonomous machine intelligence that walks into an unfamiliar lab and figures everything out.

It is a constrained agent operating inside a carefully prepared environment.

That is still useful.

It is also much more realistic.

MHS Could Become the Missing Layer Between MCP and the Physical World

MCP standardized one important idea: agents need a consistent way to discover and call software tools and data sources.

MHS extends the same philosophy toward physical devices.

Not by replacing MCP.

By giving MCP and other agent-control mechanisms a cleaner hardware layer underneath.

The stack could look like this:

User goal.

Agent.

MCP or another harness.

MHS driver.

Device safety limits.

Vendor hardware.

Sensor feedback.

Then the loop returns upward.

If that architecture works, an agent can reason across many machines without learning every vendor protocol from scratch.

The USB Analogy Is Useful — but Only Up to a Point

It is tempting to call MHS USB for AI agents.

The analogy helps because USB made many peripherals discoverable through common expectations.

MHS wants to create a common layer between agents and machines.

But physical automation is more complicated.

A keyboard is relatively standardized.

A quantum laser, microscope and six-axis robot arm are not interchangeable devices.

Their safety constraints, timing and capabilities are completely different.

So MHS is less about making machines identical.

It is about making their differences legible through a common interface.

The Bigger Shift Is From AI Using Tools to AI Operating Environments

Software agents operate inside digital environments.

MHS points toward agents operating physical environments.

That changes the stakes.

An AI that can read a database is useful.

An AI that can inspect a sensor, move a robot arm, change a microscope setting and react to a machine fault can participate in an entire workflow.

The benefit could be enormous.

So is the need for constraints.

The future of physical AI will not be defined only by how smart the model is.

It will be defined by the interfaces, safety limits, deterministic fallbacks and human checkpoints wrapped around it.

What Anthropic Has Actually Demonstrated

Anthropic has opened MHS as a limited research preview.

The company says the standard works with programmable hardware and is model-agnostic.

MHS uses standardized drivers, discoverable device descriptions, read/write-style primitives and device-level safety limits.

Agents can operate hardware through MCP, CLI and code APIs.

Research-preview partners have used MHS with robotic arms, liquid handlers, microscopes, cameras and quantum-laser systems.

At QuEra, an agent-assisted development loop produced a deterministic recovery controller that later succeeded in 695 of 700 blind trials.

At the University of Washington, a robot arm and liquid handler were coordinated without collisions in repeated demo runs.

Hugging Face and Raspberry Pi are among the organizations adding support.

Anthropic says it plans to open-source the standard after the research-preview phase.

What We Should Not Claim Yet

We should not say MHS is already an open-source standard.

It is not.

We should not say it works with every physical device.

Hardware needs a programmable interface or compatible driver.

We should not say Claude understands physical systems as reliably as software.

Anthropic explicitly describes current spatial and physical reasoning limitations.

We should not generalize QuEra’s 99.3% result to other machines.

We should not say every MHS workflow is autonomous.

Human approvals and expert oversight remain part of the pilots.

And we should not describe device-level safety limits as proof that physical AI is solved.

The research preview exists partly because it is not.

The Most Important MHS Idea May Be Where the Intelligence Stops

The exciting part of MHS is obvious.

An AI agent can touch the real world.

The more important design question is where we stop letting it improvise.

Anthropic’s strongest examples repeatedly move between two modes.

Agentic exploration when flexibility matters.

Deterministic execution when repeatability matters.

Device-level limits when safety matters.

Human approval when uncertainty becomes too high.

That layered model is more interesting than the fantasy of a fully autonomous robot scientist.

The future may not be an AI controlling every machine directly.

It may be an AI that knows when to reason, when to call a tool, when to compile what it learned into code—and when the hardware should simply refuse.

Robots Are Starting to Share a Planning Layer

Robotics is beginning to separate planning from physical execution more clearly.

One robot may have wheels.

Another may have two arms.

Another may be a humanoid.

Their bodies are different, but a higher-level model can still reason about the same shared task.

Google introduced Gemini Robotics ER 2 on July 30, 2026 as an embodied-reasoning model for robotics.

Its job is not to directly generate every motor command.

It sits above lower-level control systems.

The model can watch the physical environment, understand what is happening, plan a multi-step task, decide which tool or robot capability should be used next and then hand the motor execution to a Vision-Language-Action model or another robotics API.

ER 2 also adds multi-robot coordination.

That gives one reasoning layer a way to work across different machines in the same environment.

The important change is architectural.

A robot fleet no longer has to be understood only as several independent machines.

It can also be treated as a collection of physical capabilities coordinated by one shared reasoning system.

Gemini Robotics ER 2 Is the High-Level Brain, Not the Motor Controller

Google describes Gemini Robotics ER 2 as a high-level brain for robots.

That wording defines its place in the stack.

The model receives text, images, video and audio.

It reasons about the physical environment.

It can choose tools.

It can plan a sequence.

But the lower-level movement can remain inside another system.

A VLA model can translate vision and language into motor action.

A navigation API can move a mobile robot.

A manipulator API can control an arm.

A grasping system can handle the final contact with an object.

ER 2 coordinates those capabilities instead of replacing every controller underneath them.

This creates a hierarchy.

The user states the goal.

ER 2 decides what the goal requires.

The robot APIs or VLA models execute the physical operations.

That separation also lets the same reasoning architecture work with different robot bodies.

The model can call whatever lower-level interface the developer exposes as a tool.

Robot Capabilities Become Tools the Model Can Orchestrate

Gemini Robotics ER 2 uses a tool-oriented architecture.

Developers can expose low-level robot controls as callable functions.

A navigation function can move to a location.

A camera function can inspect the scene.

A manipulator function can position an arm.

A VLA model can perform a learned physical behavior.

The reasoning model does not need every physical detail encoded inside one prompt.

It can select from the available tools as the task progresses.

Google’s developer documentation calls this task orchestration.

The model combines spatial reasoning with custom robotics APIs to complete longer sequences.

That makes robot control look increasingly like software-agent orchestration.

The difference is the output.

A software tool may edit a file.

A robotics tool may move a machine.

The planning structure is similar: understand the current state, choose the next capability, receive the result and continue.

Continuous Video Gives the Planner a Live View of Progress

ER 2 adds a stronger temporal layer through continuous video.

A physical task unfolds over time.

The robot moves.

An object changes position.

A container becomes full.

A tool reaches the intended location.

The reasoning system needs to know when one step is actually complete before it moves to the next.

Google says ER 2 can watch continuous video feeds to track progress and verify task completion.

That gives the planner a live source of state rather than only a sequence of isolated images.

The model can observe what the robot is currently doing while it is already reasoning about the following step.

This changes the rhythm of the control loop.

Perception and planning can overlap.

The robot does not always have to stop completely before the reasoning layer begins the next decision.

The model can keep watching the task as it unfolds.

Moment Finding Helps the Model Decide When a Step Is Complete

One of the new evaluation areas is moment finding.

The model receives a video and has to identify the moment when a defined event occurs.

In robotics, that can mean the moment a subtask has been completed.

Google gives examples such as tightening a light bulb or tying a trash bag.

ER 2 can use the continuous video to decide when the required state has been reached.

Google reports 91.3 percent accuracy on its moment-finding evaluation with a mean absolute distance of 0.96 seconds.

Google also reports that the model executes this evaluation four times faster than the larger comparison category used in its release.

Those are Google’s published benchmark results for its test setup.

The important architectural point is that timing becomes part of embodied reasoning.

The model is not only asking what is in the scene.

It is asking when the scene has reached the state required for the next action.

Planning Can Continue While the Robot Is Already Acting

Google designed the ER 2 control loop so reasoning and action can overlap.

The model can think about the next step while the current physical action is still being performed.

This matters because robotics operates in real time.

A long pause between every decision can change the feel of a multi-step workflow.

ER 2’s streaming version integrates with the Gemini Live API through a bidirectional streaming endpoint designed for latency-sensitive robotics work.

The streaming endpoint can receive continuous audio and video while using function calling to control robot tools.

That gives the planning system a persistent connection to the physical process.

The robot acts.

Video continues arriving.

The model updates its understanding.

The next tool call can be prepared from the current state.

This is closer to an ongoing control conversation than a sequence of isolated prompt-response requests.

Google Provides Two ER 2 Endpoints for Different Control Loops

Gemini Robotics ER 2 currently has two model endpoints.

The standard preview endpoint is gemini-robotics-er-2-preview.

It supports multimodal input, function calling, structured outputs, code execution, search grounding and other Gemini capabilities.

The streaming endpoint is gemini-robotics-er-2-streaming-preview.

It is optimized for the Live API and low-latency robotics workflows.

Both accept text, images, video and audio as input.

Both produce text output that can include reasoning results and tool calls interpreted by the application.

Google documents an input context window of 131,072 tokens and an output limit of 65,536 tokens for the current preview endpoints.

The two-endpoint structure gives developers a choice.

A task that does not require continuous low-latency control can use the standard model.

A robot that needs ongoing video and audio interaction can use the streaming path.

Multi-Robot Collaboration Adds Another Level Above One Machine

The largest conceptual change in ER 2 is multi-robot collaboration.

Google demonstrates different robot bodies working on one shared task.

The reasoning layer can understand that the machines have different physical capabilities.

It can then coordinate a handoff between them.

Google’s release shows Apptronik’s Apollo 2 working with a Franka FR3 Duo system.

The two robots do not need identical bodies.

One can be better positioned for one part of the workflow.

Another can handle the next part.

ER 2 provides the shared semantic plan between them.

This turns robot diversity into something the planner can reason about.

The fleet does not need to behave as several copies of one machine.

Different forms can contribute different capabilities to the same objective.

A Shared Semantic Plan Lets Different Robot Bodies Work Together

Multi-robot work requires more than sending the same command twice.

Different machines may represent movement differently.

One may navigate through a room.

Another may stay fixed at a workstation.

One may have a humanoid arm configuration.

Another may use a dual-arm manipulator.

The high-level instruction therefore has to be translated into body-specific execution.

ER 2 handles the shared semantic layer.

The model understands the task in terms of objects, goals, sequence and physical relationships.

Each robot’s lower-level controller handles the mechanics of its own body.

That separation allows the planner to say what should happen without requiring the same motor representation across every machine.

The result resembles a team structure.

The shared plan defines the mission.

Each robot contributes the physical actions its own hardware is designed to perform.

Apollo 2 Shows How a Humanoid Can Become One Capability in the Team

Apptronik’s Apollo 2 is one of the robot platforms shown in Google’s multi-robot work.

Apollo 2 is a modular humanoid platform available with bipedal and wheeled-base configurations.

Apptronik describes it as a physical platform designed for embodied AI, with perception, manipulation and mobility systems that can work with higher-level models.

Apptronik and Google DeepMind already have a research partnership around Gemini Robotics.

Apptronik’s Robot Park also collects real-world task data from Apollo 2 fleets for the development of future robotics models.

Inside a multi-robot plan, the humanoid does not need to become the entire system.

It can be one physical participant.

The reasoning model can assign the part of the task that fits Apollo’s body and then coordinate another machine for a different step.

Franka's Dual-Arm Platform Adds a Different Physical Skill Set

The Franka FR3 Duo represents another kind of robot.

It uses two collaborative robotic arms rather than a humanoid body.

Franka has demonstrated Gemini Robotics with its dual-arm systems for real-time planning and multi-step manipulation.

This gives the shared planner another physical form to work with.

A dual-arm station can specialize in coordinated manipulation around a fixed workspace.

A mobile or humanoid platform can approach the task from another position.

The high-level model can reason about the complete workflow while the local robot controller handles the detailed motion of each arm.

That is the advantage of separating embodied reasoning from body-specific execution.

The reasoning model can remain common.

The machines underneath can stay specialized.

Spot Shows the Same Planning Pattern Through Robot APIs

Boston Dynamics Spot demonstrates the same architecture from another direction.

Google used ER 2 to orchestrate Spot APIs for navigation and manipulator movement.

The user can give a natural-language task.

ER 2 can decide which Spot capabilities need to be called and in what order.

Boston Dynamics Spot robot viewed indoors from above
A high-level reasoning model can orchestrate existing robot APIs while the robot's own control systems continue handling navigation and movement. This is an independent photograph of Spot, not the ER 2 demonstration.

Boston Dynamics has also documented earlier work connecting Gemini Robotics to Spot through a tool layer built on the Spot SDK.

The reasoning model acted like a high-level operator.

Spot’s existing autonomy, navigation and manipulation systems remained underneath it.

ER 2 extends that orchestration model with real-time streaming and stronger progress understanding.

The important part is that the robot API becomes a tool interface.

A mature robot platform does not have to discard its own control stack.

The higher-level model can coordinate the capabilities that already exist.

Spatial Reasoning Connects Language to Physical Coordinates

Planning in the physical world also requires spatial grounding.

Google’s robotics API documentation exposes capabilities for pointing to objects, tracking them in video, detecting bounding boxes and planning trajectories.

A model can receive an image and return normalized coordinates for visible objects.

Those coordinates can then be passed to another robotics system.

This creates a bridge between language and geometry.

The user says which object matters.

The model identifies it.

A robot controller receives the spatial output.

The physical system acts on that location.

Spatial reasoning becomes another shared service in the planning layer.

It is not tied to one manipulator.

Any compatible lower-level system can use the structured location information the model produces.

Tool Calls Can Reach Software Services as Well as Robot Hardware

ER 2 can also call non-robot tools.

Google says the model can use services such as Google Search or user-defined functions.

That widens the planning loop.

A robot task may require information that is not visible in the room.

The model can retrieve the information.

Then it can continue the physical workflow.

The same agent can therefore move between digital and physical tools.

Search can answer a question.

A database can provide a parameter.

A robot API can move a machine.

A VLA model can execute manipulation.

ER 2 coordinates the sequence.

This is another reason embodied reasoning is becoming part of the broader agent architecture.

The distinction between software tools and physical tools remains important at the execution layer.

At the planning layer, both can appear as capabilities available to the same task.

ER 2 Is Available to Developers as a Preview

Gemini Robotics ER 2 is currently distributed as a preview model.

Google makes the standard endpoint available through the Gemini API and Google AI Studio.

Google also lists Gemini Enterprise Agent Platform as a private-preview distribution path.

The model card identifies ER 2 as a Vision-Language Model based on Gemini 3.5 Flash.

Google says it was trained on Gemini 3.5 data together with additional embodied-reasoning datasets.

The current release is therefore a developer platform as well as a research result.

Developers can connect their own robot APIs.

They can stream multimodal input.

They can test spatial reasoning, progress understanding and task orchestration.

The preview status is important because it defines where the technology sits today.

The architecture is accessible.

The ecosystem is still actively developing around it.

Robots Are Starting to Plan Together, Not Just Act Alone

Gemini Robotics ER 2 shows a different direction for robotics.

The intelligence layer does not have to live entirely inside one robot body.

A high-level model can watch the environment.

Understand the shared task.

Plan several steps.

Track progress through continuous video.

Decide when one step is complete.

Call robot APIs or VLA models.

Coordinate different machines.

Keep reasoning while physical actions continue.

Apollo 2, Franka’s dual-arm platform and Spot illustrate three different robot forms that can participate in this kind of architecture.

The bodies stay different.

Their lower-level controllers stay different.

The planning layer provides a common semantic structure above them.

That is the important shift.

The future robot team may not be a fleet of identical machines running the same behavior.

It can be a collection of specialized bodies coordinated by a shared reasoning system.

The robot stops being the only unit of intelligence.

The workflow becomes the unit.

That is the upgrade.

AI Agents Are Moving From Software Tools Into Physical Equipment

AI agents have already learned how to work with software.

They can call APIs, search databases, edit files, use developer tools and coordinate workflows across applications.

The next interface is physical equipment.

A robot arm has position, speed, payload and safety limits.

A camera has exposure, focus and image data.

A laser system has tunable parameters and sensor readings.

A manufacturing machine may expose its own commands, status values and control software.

The Model Hardware Standard, or MHS, is designed to give AI agents a common way to understand that equipment.

Anthropic opened MHS as a limited research preview on August 27, 2026 after developing it with HHMI Janelia Research Campus and a group of partners across science, robotics, electronics and manufacturing.

The central idea is straightforward.

A physical device should be able to describe what it is, what it can measure, what can be changed and which operating limits must be enforced.

An AI agent should then be able to discover that description and interact with the device through a common interface.

That takes the agent architecture beyond software tools.

The tool can now be a physical machine.

MHS Adds a Standard Driver Between the Agent and the Device

The main technical layer in MHS is the driver.

A driver is software that translates between a higher-level system and the specific hardware underneath.

Operating systems already use this pattern.

A printer driver turns generic print operations into commands a specific printer understands.

A graphics driver connects a software graphics stack to a GPU.

MHS applies a related idea to AI-controlled equipment.

The device gets an MHS driver.

That driver exposes the machine through a standard structure instead of requiring the agent to understand every vendor-specific interface directly.

Anthropic describes simple primitives such as read and write.

Read can request a value such as temperature, position or another sensor state.

Write can change a supported parameter.

The actual machine may still use a proprietary SDK, serial protocol, command-line tool or control application underneath.

The MHS layer translates the common operation into whatever the device requires.

That separation matters.

The agent can reason at the level of the task.

The driver handles the device-specific control path.

The same agent architecture can therefore interact with several types of equipment while each machine retains its own implementation underneath.

A Device Can Describe Itself in a Machine-Readable Reference

Control commands are only one part of physical operation.

An agent also needs context.

A robot may have a maximum payload.

A motor may have a permitted speed range.

A camera may expose several imaging modes.

A temperature controller may have upper and lower operating limits.

Some of that information exists in code.

Some exists in manuals.

Some exists as operating knowledge held by the people who work with the equipment.

MHS adds a structured way to bring that information into the interface.

Anthropic says MHS drivers support tags where users can describe machine characteristics in natural language.

The system can then produce a reference file describing the device.

That reference can include what the machine measures, which parameters can be adjusted and which safety limits are enforced.

The result is more than a list of functions.

It is a device model.

The agent receives both verbs and context.

Read this value.

Set this parameter.

This part moves.

This value has a defined range.

This machine has a physical characteristic the agent should know before operating it.

That turns hardware documentation into part of the agent’s working interface.

Discovery Lets Agents Find Devices Through a Common Format

A shared hardware interface becomes more useful when devices can also be discovered consistently.

MHS is designed to make connected equipment discoverable in a standard format across a network.

That gives an agent a way to learn what hardware is available before it starts building a workflow.

A system might expose a camera.

Another node might expose a robotic arm.

A third might expose a sensor controller.

The agent can inspect the available devices, read their reference information and determine which capabilities fit the task.

This resembles what has happened in software-agent systems.

A software agent can discover tools exposed by an MCP server.

Each tool has a name, a description and an input structure.

MHS extends the same general idea into physical devices.

The available capability is not only “search database” or “write file.”

It can be “read camera,” “move arm,” “set controller value” or another machine-specific operation represented through the common hardware layer.

Discovery changes orchestration.

The agent does not have to begin with one hard-coded machine.

It can begin with the equipment that is present.

MCP Becomes One of the Ways an Agent Reaches MHS Hardware

MHS does not replace the Model Context Protocol.

It can sit behind it.

Anthropic says MHS is model-agnostic and can be accessed by agent harnesses through standard protocols such as MCP.

That creates two distinct layers.

MCP gives an AI application a standard way to discover and call tools.

MHS gives physical equipment a standard way to describe and expose its hardware capabilities.

The layers can connect.

An MCP tool can represent an action that ultimately reaches an MHS driver.

The agent sees a callable operation.

MHS handles the relationship between that operation and the physical device.

This is why MHS is a different topic from MCP even though the two standards can work together.

MCP standardizes model-to-tool interaction.

MHS standardizes the hardware side of the tool.

The distinction becomes visible when a workflow crosses from software into the physical world.

The agent may use MCP to reach the tool.

The tool may use MHS to reach the machine.

The result is a stack rather than one protocol trying to describe everything.

Command Line and Code Files Give MHS More Than One Control Path

Anthropic describes three control mechanisms for MHS: MCP, command-line interfaces and code files or APIs.

That gives the agent several ways to operate the same hardware layer.

Interactive reasoning can use MCP calls.

A developer or operator can use command-line access.

A long-running workflow can be packaged into code.

The code path is important because an AI agent does not need to reason through every low-level step forever.

The agent can explore a device, learn a sequence and then turn that sequence into deterministic code.

Anthropic demonstrated this pattern with laser alignment.

Claude adjusted the laser, observed the result through a camera and repeated the process while learning the relationship between control changes and the observed beam.

It then packaged the learned sequence into a script.

The final operation could run as a deterministic command rather than requiring live reasoning for every adjustment.

That creates two operating modes.

Reason while discovering the procedure.

Execute code after the procedure has been defined.

MHS provides one hardware interface underneath both.

The Standard Can Coordinate Several Machines as One Workflow

The main architectural value appears when a task needs more than one device.

A production cell may contain several robot arms.

An imaging system may combine cameras, motors, sensors and optical equipment.

A test setup may use measurement devices on different computers.

MHS gives the agent one structured control surface across those devices.

The agent can read state from one machine.

Wait for a condition.

Trigger a second device.

Inspect another sensor.

Adjust a parameter.

Continue to the next stage.

Anthropic describes MHS as supporting parallel operation across multiple instruments and says device commands can be chained together in code files.

The standard therefore sits above the individual machine.

Its unit of work can become the complete physical workflow.

That is similar to what software agents already do across several applications.

The difference is that the state now includes real-world positions, measurements and machine status.

The workflow exists in both software and physical space.

Universal Robots Tested MHS Across Multiple Cobots

Manufacturing provides a direct example.

Universal Robots joined the MHS research preview and tested the standard on its collaborative robots.

The company described an experiment involving four separate robot applications.

An agent running on Claude Opus 4.8 coordinated them as one cell and handed payloads between the robots.

The relevant part is the interface structure.

A compact articulated robotic manipulator being controlled by hand
A robotic manipulator illustrates the physical-control layer that MHS is designed to standardize for AI agents. The pictured manipulator is a generic device and is not represented as an MHS-enabled product.

Each robot application still runs on real industrial hardware.

The existing robot-control and safety architecture remains underneath.

MHS adds a layer where the agent can discover the devices, understand declared operating information and orchestrate the applications together.

Universal Robots also describes machine bounds, interlocks and emergency-stop information being represented through the MHS layer while the robot platform’s own safety architecture remains active.

That is how a hardware standard can participate in an industrial system.

The device controller remains underneath the MHS layer.

It provides an additional interface above it.

The agent works with the standardized representation.

The robot controller continues operating the machine.

QuEra Used MHS to Connect an Agent to Precision Quantum Hardware

Quantum computing shows another type of physical system.

QuEra Computing used MHS to connect an AI agent to parts of the laser-control system inside its neutral-atom quantum-computing hardware.

The task was laser stabilization.

A neutral-atom quantum system depends on lasers maintaining highly precise operating frequencies.

QuEra used MHS as the hardware interface while Claude developed and tested control logic for restoring the laser to its target lock.

In a later blind test reported by Anthropic and QuEra, the resulting controller recovered the lock 99.3 percent of the time.

QuEra also reports that the workflow reduced recovery from minutes of specialist work to seconds in the test environment.

The point for MHS is not the quantum algorithm.

It is the hardware loop.

Read physical state.

Change controls.

Observe the result.

Evaluate whether the target condition has been reached.

Package the resulting procedure into code.

That same loop can exist in many forms of precision equipment.

MHS gives the AI system a standard place to connect to it.

Raspberry Pi Shows How MHS Can Reach Smaller Hardware

MHS is not limited to large industrial or scientific systems.

Anthropic says Raspberry Pi is enabling MHS integration across a number of its products after successful tests with a Camera MHS Driver.

A camera is a useful small-scale example because it combines hardware state with continuous data.

The device has configuration values.

It has an image stream.

It may be connected to motors, sensors or other peripherals.

An agent can use the camera as both an input and part of a control loop.

The MHS driver provides the structured hardware interface.

A Raspberry Pi can provide the programmable compute environment around it.

That makes the standard relevant to prototyping, robotics and edge systems as well as large installations.

The physical-AI stack can therefore scale down.

It does not require a factory floor.

A developer can start with a programmable board, a camera and another controllable device.

The same general driver model can then grow with the hardware system.

Hugging Face Is Connecting MHS to the LeRobot Ecosystem

Robotics libraries add another layer to the ecosystem.

Anthropic says Hugging Face is adding MHS support to LeRobot.

LeRobot is Hugging Face’s open robotics framework for models, datasets and tools used with real-world robots.

Its current documentation describes a hardware-agnostic Python interface for controlling robots, recording datasets and deploying learned policies.

MHS can add a standardized agent-to-hardware interface to that environment.

The two systems have different roles.

LeRobot focuses on robot learning, datasets, policies and hardware abstractions.

MHS focuses on exposing physical equipment to AI agents through a common device interface.

Combining them creates a path where a robot-learning stack and an agent orchestration stack can share the same physical equipment.

An agent can discover the machine.

A policy can control learned motion.

The hardware interface can expose state and permitted operations.

That is another sign that MHS is being positioned as a layer rather than a complete robotics framework.

It can connect to existing robotics software instead of replacing it.

Model-Agnostic Design Separates the Hardware Standard From One AI Model

MHS was introduced by Anthropic, but the specification is designed to be model-agnostic.

Anthropic says any agent harness can access MHS through standard protocols such as MCP.

That matters for the architecture.

A hardware installation can last much longer than one model generation.

A robotic system, camera rig or manufacturing cell may operate for years.

AI models can change much faster.

A model-agnostic hardware interface separates those lifecycles.

The driver describes the machine.

The agent harness connects to the driver.

The model used by that harness can change independently.

That keeps the machine integration focused on the physical device rather than one model API.

The same principle is common in computing standards.

The same USB connector can work across different processors.

The same Ethernet framing can work across different operating systems.

MHS is applying a similar separation to agent-controlled hardware.

The hardware contract lives in one layer.

The intelligence using that contract lives in another.

Physical Limits Can Be Declared as Part of the Interface

A physical interface needs more than capability descriptions.

It also needs operating boundaries.

MHS drivers can include information about safety limits that will be enforced.

Universal Robots describes bounds, interlocks and emergency-stop information being included in the MHS representation while the robot’s own safety architecture remains responsible underneath.

That layered design matters.

The agent sees what it is allowed to request.

The MHS layer can expose the permitted operating envelope.

The hardware controller still applies the machine’s own safety mechanisms.

The result is not one safety system replacing another.

It is a hierarchy.

The model works inside the capabilities it has been given.

The MHS driver describes and constrains the interface.

The device’s native controller remains responsible for its own hardware-level protections.

This is especially relevant when an AI system controls equipment with movement, energy or other physical effects.

The machine interface has to describe not only what can be done.

It also has to define the boundaries around how those actions can be performed.

The Research Preview Is Building the Standard Before Open Source Release

MHS is not yet a generally available open standard.

Anthropic currently describes it as a limited research preview.

Access is by application.

The company says the preview is being used with partners across science, robotics, electronics and manufacturing to build evaluations and deployment practices before the standard is released as open source.

The official MHS site uses the same framing.

The project began with Anthropic and HHMI Janelia Research Campus.

The current phase expands testing across outside organizations.

That status is important because it tells us where the technology sits.

The architecture exists.

Drivers are being tested on real devices.

Partners are building integrations.

The specification is still being developed through the research-preview process.

Anthropic says it plans to open-source MHS after that work.

For developers watching the standard, the present moment is therefore about the interface design and the early ecosystem rather than a finished universal hardware layer.

The direction is visible before the final open-source release.

MHS Creates a Hardware Layer Beside the Software Tool Layer

The larger architecture becomes clear when MHS is placed beside the other agent standards emerging around AI.

MCP gives models access to software tools and data.

App Intents, AppFunctions and operating-system action frameworks expose application capabilities.

Agent plugin formats package reusable agent extensions.

MHS addresses a different boundary.

It connects the agent stack to physical equipment.

That means future agent systems can have several layers of tools.

A database tool.

A browser tool.

A code tool.

An application action.

A camera.

A robot arm.

A controller.

A sensor.

To the agent, each one can become a structured capability.

The underlying implementation can remain very different.

Software stays software.

Hardware stays hardware.

The standard defines the interface where the agent meets each system.

That is why MHS can sit next to MCP rather than competing with it.

The software tool layer and the physical-device layer are becoming separate parts of the same agent architecture.

AI Agents Are Getting a Common Interface to the Physical World

The Model Hardware Standard turns one broad idea into an engineering structure.

A physical device gets a driver.

The driver exposes common operations such as reading and writing state.

The device can publish a reference describing its characteristics and operating limits.

Agents can discover connected hardware.

MCP, command-line interfaces and code files can provide control paths into the same system.

Several devices can be orchestrated as one workflow.

Model choice remains separate from the hardware standard.

Existing machine controllers and safety systems remain underneath the MHS layer.

Early integrations now span robotics, Raspberry Pi hardware and quantum-computing equipment.

Anthropic is testing the specification with partners before a planned open-source release.

This extends the agent stack in a specific direction.

The model can now work with software resources and physical equipment through separate structured interfaces.

Physical equipment can be described as a structured, discoverable capability too.

The agent gets a common interface.

The machine keeps its own implementation.

MHS connects the two.

That is the upgrade.