Lenovo Used IFA 2026 to Push Qira Into a Much Bigger Role
Lenovo’s IFA 2026 update for Lenovo & Motorola Qira is about expanding where the same personal AI can show up. At Lenovo Innovation World in Berlin on September 3, Lenovo announced that Qira is reaching more eligible PCs, more Motorola devices, a wearable, and a wider set of connected app experiences. The direction is clear: Lenovo wants one intelligence to follow the user across the devices and services that already make up a normal day. That turns Qira from a feature attached to one machine into a broader personal layer that can stay useful while the user moves between screens, applications, and moments.
The Core Idea Is One Personal AI Across Multiple Devices
Lenovo describes Qira as personal ambient intelligence built to work across compatible Lenovo and Motorola devices. The same intelligence appears as Lenovo Qira on Lenovo products and Motorola Qira on Motorola products, with a shared experience designed to carry context across PCs, smartphones, tablets, and wearables. That continuity is the foundation of the whole system. A conversation, saved document, meeting context, or personal knowledge item can remain useful as the user moves between supported devices. The value is not only having AI on several products. It is having one personal AI identity that is designed to remain consistent across them.
IFA Expands Qira to More 16GB Lenovo PCs
One of the most practical IFA announcements is broader PC eligibility. Lenovo says Qira support is expanding to eligible Lenovo PCs with 16GB of memory, extending access beyond the higher-memory configurations that defined earlier availability. That matters because it brings the personal-AI layer into a larger part of Lenovo’s PC portfolio. More compatible systems can now participate in the same cross-device experience, giving Qira a broader hardware base. For users, the important idea is continuity: the same personal knowledge and context can become available on more of the devices they already use.
Motorola Qira Is Expanding With Android 17
Lenovo also says Motorola Qira will expand to eligible Motorola devices as they receive Android 17. The company names select edge, signature, and razr devices in the rollout. This connects the phone more tightly to the same personal-AI identity used on compatible PCs and other Lenovo devices. A phone is a natural place for context because it is with the user throughout the day, capturing conversations, messages, schedules, photos, and quick requests. Bringing Qira deeper into Motorola’s Android lineup gives Lenovo another important surface for keeping the personal experience continuous.
Moto Watch Ultra Becomes Qira’s First Wearable
Lenovo identifies the new moto watch ultra as the first wearable to support Motorola Qira. That is a meaningful expansion because a wearable gives personal AI a different kind of presence. It is closer to the user throughout the day and can become another surface for timely information, quick interactions, and connected experiences. In Lenovo’s broader vision, Qira is not tied to one screen size or one category. The intelligence is the common layer. Moving from PC to phone to wearable makes that idea easier to understand: the device changes, but the personal AI remains part of the same system.

The Bigger IFA Upgrade Is Moving From Devices Into Apps
Lenovo’s new announcement becomes especially interesting when Qira moves beyond hardware and into everyday applications and services. Lenovo says supported Qira experiences can connect with Gmail, Google Calendar, Google Contacts, Slack, Outlook Mail, and Outlook Calendar through its collaboration with Workato. That gives Qira a path from understanding the user’s request to working with information and actions inside the services that already organize email, meetings, contacts, and collaboration. The personal AI can therefore become useful not only because it knows the user’s context, but because it can connect that context to the tools that help complete the next step.
Workato Is the Connectivity Layer Behind the New App Experiences
Lenovo’s separate Workato announcement explains the architecture more clearly. Workato is serving as the connectivity layer powering the underlying Model Context Protocol server infrastructure used for these connected Qira experiences. That detail shows how Lenovo is turning the personal-AI idea into an integration system. Qira can remain the user-facing intelligence while a dedicated connectivity layer helps it reach supported applications and workflows. This separation gives the system a clean structure: personal context and natural-language interaction at the Qira layer, with approved app connectivity underneath.
MCP Gives Qira a Standard Way to Reach More Tools
Model Context Protocol gives AI systems a structured way to connect with tools and data. In Lenovo’s implementation, the Workato-powered MCP layer helps Qira connect natural-language intent with supported services. This creates a useful division of roles. Qira handles the personal context, conversation, and understanding of what the user wants to accomplish. The integration layer exposes approved application capabilities that can help carry out the request. That architecture also gives Lenovo room to expand the ecosystem over time, because new tools can be connected through the same underlying pattern rather than requiring a completely separate personal-AI experience.
The Integration List Extends Beyond Email and Calendars
Lenovo’s Workato release names Microsoft 365, Google Workspace, Trello, Asana, and Discord among the initial integrations described for the collaboration. The Qira expansion release also calls out Gmail, Google Calendar, Google Contacts, Slack, Outlook Mail, and Outlook Calendar. Together, those examples show the range Lenovo is targeting: communication, planning, project management, collaboration, contacts, and personal organization. The point is not the number of logos. The important change is that one personal AI can begin to understand a goal and then connect with different supported services that each handle part of the user’s day.
A Natural-Language Request Can Become an App Action
Lenovo gives concrete examples of how this connected model is meant to work. A user can ask Qira to create a Trello task for a website redesign project, or ask it to share meeting notes with a team in Discord. The value is the direct path between the request and the supported action. The user stays in the Qira experience while the integration layer connects that intent to the service that can complete the next step. This is where personal AI starts to feel more agentic: understanding a request is only the beginning, while completing a useful action across a connected service becomes part of the experience.
Catch Me Up Becomes More Useful When Apps Are Connected
Lenovo specifically highlights Catch Me Up as part of the new connected experience. If an important update arrives through email, a group conversation, or a calendar change, Qira can bring relevant information together and help the user continue from the same interface. That fits the larger ambient-intelligence idea: useful context can follow the user rather than remaining limited to the application where it first appeared. A personal AI becomes more valuable when it can connect information from the places the user already works and communicate what matters in a way that fits the current moment.
The Official Demo Shows Presence, Action and Perception
Lenovo’s official Qira demo describes the system around three ideas: Presence, Action, and Perception. Presence means the same intelligence can remain available across devices. Action is about orchestrating supported tasks across apps and devices. Perception is the ability to build useful knowledge around the user and the context they choose to share. The IFA expansion strengthens all three ideas by adding more device surfaces and a broader application layer. The demo also shows features such as context-aware suggestions, live transcription, a personal knowledge base, cross-device control, local AI, Live Mode, and Catch Me Up.
Qira Can Build a Personal Knowledge Base
Lenovo’s product page emphasizes a personal knowledge base where users can add documents and saved memories that Qira can use to support later interactions. This gives the system continuity that goes beyond a single chat session. A file saved earlier, a remembered preference, or information collected on another supported device can become part of the context available when the user asks for help later. Personal AI becomes more useful when it can build on previous information with the user’s control, because the system can support a continuing workflow rather than treating every request as an isolated moment.
Local AI Keeps Part of the Experience Close to the Device
Qira is also designed to use local AI capabilities on compatible devices. Lenovo’s official materials describe experiences that can work directly on the PC, including offline interactions and local creative features. That local layer complements the connected app ecosystem. Some work can remain close to the device, while connected services can contribute when the task benefits from information or actions available elsewhere. This hybrid model fits Lenovo’s broader AI direction: local hardware, personal context, and connected services can work together as parts of one experience rather than separate products.
User Permission and Control Stay Central to the Design
Lenovo repeatedly frames Qira around user permission, choice, and control. Its IFA announcement says the connected experiences are designed to maintain user permission and control as Qira works across supported services. The Qira product page also describes personal data as being stored on the device and cloud connections being used when needed. That makes control part of the product architecture. For a personal AI that is designed to remember context and connect across multiple services, keeping the user at the center is essential to the experience Lenovo is building.
Lenovo Is Building Toward a Qira Marketplace
The Workato collaboration is described as the first phase of a multi-stage agreement. Lenovo says the work begins with Qira for consumers and is intended to extend toward a broader Lenovo Qira marketplace. That points to a platform strategy rather than a fixed set of integrations. If the ecosystem keeps growing, Qira can become a common personal-AI layer connected to a wider range of services while preserving the same device-to-device identity. The marketplace idea also gives developers and service providers a clearer place in Lenovo’s long-term vision for connected personal AI.
IFA 2026 Makes Qira Feel More Like an Ecosystem Than an App
Qira started the year as Lenovo’s cross-device personal ambient intelligence, and the IFA update makes that concept much more concrete. More PC configurations can participate. More Motorola devices are joining. A wearable becomes part of the system. Workato-powered app connections give the intelligence a way to reach the tools people already use. The result looks less like one application and more like an ecosystem layer that spans hardware, software, and services. That is the bigger story behind the IFA announcement: Lenovo is turning Qira into connective tissue across its personal-computing world.
The Upgrade Feeling
The most compelling part of Lenovo Qira is not a single AI trick. It is the idea that the same personal intelligence can move with the user, carry context across devices, and connect that context to supported actions inside everyday applications. IFA 2026 expands that vision in exactly the places that matter: more devices, more surfaces, and more services. If Lenovo keeps building on this architecture, the upgrade may feel less like opening another assistant and more like having one intelligence already present across the digital environment you use every day.
Zoho’s Catalyst 3.0 is built around a problem that appears after AI has already generated the application: someone still has to create the database, wire authentication, provision compute, deploy services, inspect logs, manage permissions and promote the finished system into production. Catalyst now combines Agent Skills, a Dynamic MCP server and a non-interactive CLI so coding agents such as Claude Code and Codex can operate the cloud platform from inside the coding workflow. The more interesting design choice is the guardrail: agents can work in development, but Zoho says production promotion remains a manual human action.
AI Solved the First Half of Vibe Coding Faster Than the Second
Vibe coding made application creation feel deceptively complete.
Describe the interface. Ask for a database-backed feature. Generate the API route. Fix the error. Refresh the page.
A working prototype can appear in hours.
Then the application leaves the prompt window.
It needs authentication, persistent data, secrets, storage, functions, background jobs, hosting, logs, permissions and a production environment.
The developer discovers that writing the code was only one layer of shipping software.
This is the problem Zoho is targeting with Catalyst 3.0. The September 2 release is less interesting as another AI coding feature than as an attempt to make the cloud itself understandable and operable by the coding agent.
Catalyst 3.0 Is a Cloud Platform, Not a New Coding Model
Catalyst is Zoho’s Platform-as-a-Service.
It provides managed application infrastructure rather than a model that competes with Claude, GPT or Gemini.
Catalyst 3.0 adds an agent-facing layer on top of that platform. Zoho’s current documentation describes integrations for coding agents including Claude Code, Codex and GitHub Copilot through its AI Plugin, while the broader Catalyst 3.0 page also lists tools such as Cursor and Gemini.
The design is intentionally model-agnostic.
Your coding agent generates and modifies the application. Catalyst supplies the application services and exposes platform actions the agent can invoke.
That separation matters. The model remains replaceable. The infrastructure becomes the persistent execution environment.
The Next Vibe-Coding Bottleneck Is Infrastructure Context
A coding model can know what a database is and still use a specific cloud platform incorrectly.
Every platform has its own naming, service boundaries, deployment rules, SDK initialization patterns, authentication conventions, CLI and production restrictions.
Without current platform context, an agent may generate code that is generally reasonable but wrong for the actual environment.
This is similar to the documentation problem in ordinary AI coding, except the consequences extend beyond code. A mistaken infrastructure action can create the wrong resource, deploy the wrong component or configure the wrong service.
Catalyst 3.0 tries to reduce that gap by teaching the agent the platform and giving it controlled ways to act on it.
Agent Skills Tell the Model How Catalyst Is Supposed to Be Used
The first layer is Agent Skills.
Catalyst’s current Plugin documentation lists 15 Product Skill Files and four AI Architect Skill Files.
The product skills cover platform-specific areas including functions, AppSail, Slate hosting, authentication, Data Store, NoSQL, cache, Stratus object storage, Signals, browser automation, Zia services, MCP, SDK usage and pricing.
Instead of expecting the model to infer the correct Catalyst architecture from general training knowledge, a skill gives it an explicit implementation playbook.
That moves some decision-making out of probabilistic memory and into maintained platform instructions.
For vibe coding, that can be more important than a larger model. A smaller model with the right current instructions can sometimes make a better platform decision than a stronger model guessing from stale context.
The Four Architect Skills Solve a Different Problem
Product skills explain individual services. Architecture requires choosing between them.
Should this workload be a serverless function or an AppSail service? Should data live in the relational store or NoSQL? Does the application need object storage? How should authentication be wired?
The four AI Architect Skill Files are intended to feed popular coding agents with platform context in an optimized form.
This is a useful distinction. A cloud agent needs both vocabulary and judgment.
Knowing every available service does not automatically tell the model which service belongs in the design. Catalyst’s orchestration layer is attempting to narrow that decision space before the agent starts creating infrastructure.
MCP Gives the Agent Hands, Not Just Documentation
Skills can tell the model what to do. Model Context Protocol gives it a way to do it.
Zoho’s Catalyst MCP server exposes platform operations to compatible AI clients.
The release gives simple examples such as creating a database table or adding a column directly from the developer’s coding environment instead of opening the Catalyst console.
That changes the role of the coding assistant. It is no longer only producing commands for the human to copy. The agent can be authorized to execute supported infrastructure operations itself.
The convenience is obvious. The risk is also obvious.
Once a coding model can change cloud state, permissions and environment separation become part of the AI system design.
Catalyst Uses Dynamic MCP Instead of Loading Every Tool at Once
Catalyst currently documents more than 100 MCP tools.
Loading every tool definition into every coding session would be wasteful. It could also make tool selection harder.
Zoho’s Dynamic MCP server uses tool discovery. When a prompt requires an action, the model queries the server directory and only the tools needed for that operation are loaded and enabled.
Zoho argues that this reduces tool overload and unnecessary hallucination while also reducing manual MCP configuration.
The broader design idea is worth watching beyond Catalyst.
As agent toolboxes grow, tool discovery may become as important as tool execution. An agent with 500 tools does not necessarily need 500 schemas in context. It needs a reliable way to find five.
The Dynamic Server Is Also a Context-Efficiency Strategy
Tool schemas consume context. Large tool descriptions consume more.
If an agent session carries dozens or hundreds of tool definitions that will never be called, part of the model’s working context is spent describing capabilities irrelevant to the current task.
Dynamic discovery turns that into an on-demand problem.
A prompt about a data table should surface data tools. A deployment prompt should surface deployment tools. A storage task should surface storage operations.
Zoho positions this as a way to optimize token use.
The company’s exact efficiency claims should be treated as product claims rather than universal guarantees. But the architectural direction is sensible: agent infrastructure needs context management, not only API access.
The Non-Interactive CLI Solves a Very Old Automation Problem
Command-line tools were built for humans long before coding agents became common.
A normal CLI often pauses: select an organization, choose a project, confirm a runtime, answer yes or no, pick a component.
Those prompts are friendly when a person is sitting at the terminal. They are friction when an agent is trying to execute a multi-step workflow autonomously.
Catalyst’s non-interactive mode lets required answers be supplied through flags, environment variables or programmatic arguments.
If required information is still missing, the command exits with an error instead of waiting indefinitely for human input.
That small change makes ordinary CLI operations much more agent-compatible.
Non-Interactive Does Not Mean Unrestricted
Removing interactive prompts could sound like removing safety.
Zoho documents the opposite intention.
Its Agent Skills page says destructive commands are disabled in non-interactive mode. The agent can also operate with its own scoped collaborator permissions.
That means autonomy is not supposed to come from giving the model an administrator account and telling it to be careful.
The intended model is narrower: give the agent only the permissions it needs, prevent some destructive paths entirely, log what it does and keep production separated.
This is much closer to how automation should be designed for a probabilistic actor.
The Orchestration Layer Decides Between CLI and MCP
Catalyst exposes more than one execution mechanism. Some operations fit the CLI. Others fit MCP.
Leaving that choice entirely to the language model creates another place for inconsistent behavior.
Zoho says orchestration is built into the Skill so the coding assistant can be routed deterministically down a CLI or MCP path.
The word deterministic should be interpreted carefully. The model is still interpreting the user’s request and the surrounding workflow is still agentic.
What Zoho is making deterministic is the platform-routing logic once a particular implementation path is selected.
That reduces one category of model improvisation. It does not turn the entire development process into deterministic software.
The Cloud Becomes Part of the Coding Conversation
Traditional development separates several interfaces.
The editor contains code. The terminal contains deployment commands. The cloud console contains infrastructure. The monitoring console contains logs. The database console contains tables.
The developer jumps between them and mentally keeps the state synchronized.
Catalyst 3.0 is trying to collapse more of those operations into the agent conversation.
A developer can ask for a feature. The coding agent can write the code. The skill can recommend the relevant Catalyst service. MCP can create supporting resources. The CLI can perform supported project operations.
The result is a more continuous prompt-to-infrastructure workflow.
But Zoho Draws a Hard Line Before Production
The strongest design decision in Catalyst 3.0 may be what the agent is not allowed to do.
Zoho says development and production environments are decoupled. Code moves to production through manual promotion only.
Its Agent Skills page states the rule even more directly: the agent never touches production.
That is significant.
The goal is not maximum autonomy. The goal is bounded autonomy.
An agent can build, configure development resources, deploy and test within the development workflow. But the final transition to the environment used by real users remains a human-controlled step.
That one boundary changes the risk model substantially.
Catalyst’s Existing Environment Model Makes That Guardrail Concrete
The production restriction is not only marketing language around AI.
Catalyst already has separate Development and Production environments.
The documentation says new projects begin in Development. Resource creation, configuration, testing, CLI actions and API-driven changes are reflected there. Changes do not appear in the live application until they are deployed to Production.
The production environment also restricts many direct modifications. For example, the documentation says new functions or Signals rules generally cannot simply be created directly in Production.
The agent guardrail therefore builds on an environment model that already existed. AI is being inserted into the safer side of that boundary rather than redefining the boundary around AI.
Manual Promotion Is Slower Than Full Autonomy — That Is the Point
A fully autonomous demo is more impressive.
Tell the agent to build an app. Watch it create resources. Watch it deploy publicly. Open the URL. Done.
That is also the workflow with the largest blast radius.
A production deployment can expose bad code to users. A schema change can affect live data. A misconfigured permission can become a security issue. A runaway resource can become a cost problem.
Manual promotion introduces friction exactly where friction is useful.
The developer has a natural checkpoint to inspect what the agent created before real users inherit it.
For production software, removing every click is not necessarily progress.
Scoped Permissions Matter More Once the Agent Can Change Infrastructure
Coding assistants already operate with filesystem and shell permissions. Cloud access raises the stakes.
A scoped collaborator model lets an organization treat an agent more like a constrained service identity than a trusted human administrator.
Catalyst documents detailed project profiles and permissions for development and production capabilities. Permissions can govern access to data stores, logs, settings, migrations and other components.
The exact safe configuration will depend on the application.
The principle is broader: AI agents should receive capabilities according to the task, not according to convenience.
If an agent only needs to create development tables, it should not automatically receive authority over billing or production migration.
Audit Logs Turn Agent Actions Into Reviewable Events
Agentic development creates a provenance problem.
When something changes, who changed it? The developer? A script? The coding agent? Which tool did it call? When?
Catalyst’s governance story includes application logs, platform logs and MCP tool-call logs. Zoho says these records can be used to reconstruct what an agent did and when.
The existing Audit Logs system also records configuration events such as adding a Data Store column, changing an event rule or deleting a cron job.
This matters because conversational interfaces can otherwise hide operational detail.
A natural-language prompt is not a sufficient audit record for the side effects that followed.
Reversible Changes Are an Antidote to Confident Agent Mistakes
AI agents often fail with confidence.
A tool call can be syntactically valid and still be the wrong operational decision.
Zoho says changes after launch are versioned, attributable and reversible.
That is the correct direction for agent-controlled infrastructure.
If the system assumes mistakes will eventually occur, rollback becomes a first-class capability.
The same philosophy already exists in source control. We do not trust every code edit simply because it compiled. We preserve history.
Agent infrastructure needs the same mentality. Autonomy becomes safer when actions leave evidence and can be undone.
The Benchmark Says Skills and MCP Help — but It Is Zoho’s Benchmark
Zoho publishes a task-completion comparison for three models on the Catalyst 3.0 page.
In the company’s test, completion without Catalyst Skills ranged from 25% to 55%. With Skills plus Zoho MCP, the reported completion rates were 90%, 92% and 95% depending on the model.
The table also reports fewer human interventions and, in several cases, fewer retries or tool calls.
That is an encouraging result. It is not independent validation.
Zoho is testing its own platform, skills and tool layer. The public page does not establish that every real-world application will see the same improvement.
The useful conclusion is narrower: in Zoho’s own evaluated tasks, supplying platform-specific skills and tools made the tested agents much more successful at completing Catalyst workflows.
The Benchmark Also Shows Why Bigger Models Are Not the Whole Answer
The interesting pattern in Zoho’s table is not which model wins.
All three reported models improve sharply when given the Catalyst-specific context and tools.
That supports a broader lesson in agent engineering.
Capability is not only model intelligence. It is model plus instructions plus tools plus permissions plus environment.
A coding model can be excellent at reasoning and still fail because it does not know the platform’s exact workflow.
Giving it a current skill file and the right API may create a larger practical improvement than switching to a slightly stronger model.
Vibe coding is gradually becoming systems engineering around the model.
Catalyst 3.0 Is Trying to Productize the Agent Harness
A serious coding agent needs more than an LLM.
It needs context, tool discovery, execution, authentication, permissions, logging, recovery and environment boundaries.
Developers can build that harness themselves. Many teams already do.
Catalyst 3.0 packages a version of that harness around one cloud platform.
The Agent Skills provide platform context. Dynamic MCP provides discoverable actions. The non-interactive CLI provides automation-friendly commands. The cloud supplies managed services. The production boundary supplies a human checkpoint.
This is why the release matters beyond Zoho.
Cloud platforms are beginning to redesign themselves around agents as first-class operators.
This Is Different From Infrastructure-as-Code
Infrastructure-as-code already lets developers describe cloud resources in version-controlled files.
Catalyst 3.0 does not make that idea obsolete.
The agent-ready approach attacks a different layer of friction.
Instead of requiring the developer to know the provider syntax and construct every resource declaration directly, the coding agent can translate application intent into supported platform operations.
The danger is obvious. Generated infrastructure can become harder to understand than generated application code.
That makes exportability, logging and review important.
A convenient agent interface should not become an excuse to stop knowing what infrastructure exists.
A Full-Stack Platform Reduces Integration Work by Reducing Choice
Catalyst includes frontend hosting, serverless functions, compute, authentication, relational and NoSQL data, storage, events and other managed services.
Putting those services together reduces the number of providers an agent has to understand. That can make orchestration easier.
It also creates platform dependence.
A system built deeply around one provider’s authentication, data services, event model and deployment workflow is not automatically portable to another cloud.
This is not unique to Zoho. It is the trade-off of integrated PaaS platforms in general.
Vibe coders should understand the bargain: you exchange some infrastructure flexibility for a smaller operational surface and a more opinionated development path.
The Platform Still Cannot Decide What Your Production Architecture Should Be
Agent Skills can recommend Catalyst patterns. They cannot understand every business constraint automatically.
A production architecture still depends on data sensitivity, latency, availability requirements, regulatory obligations, traffic shape, recovery objectives, cost, team capability and external dependencies.
A vibe-coded prototype may work perfectly on one service arrangement and still need architectural changes as usage grows.
Catalyst can reduce plumbing. It cannot remove architecture as a discipline.
The agent can propose the system. The developer still needs to know what promises the system must keep.
Cloud Cost Is Another Place Where Agent Autonomy Needs Limits
Infrastructure actions can create financial side effects.
Provision more resources. Increase storage. Invoke more functions. Move more data.
Those decisions can change a cloud bill.
Catalyst uses a pay-as-you-go model and documents budget alerts and ceilings.
That fits the same bounded-autonomy pattern as manual production promotion.
An agent should not only be limited by what it is technically allowed to create. The environment should also expose economic guardrails.
Vibe coding makes resource creation easier. That makes cost visibility more important, not less.
What Zoho Has Actually Shipped
Zoho announced Catalyst 3.0 on September 2, 2026 and says it is available for immediate use.
The current AI Plugin documentation lists a Dynamic MCP Server, 15 Product Skill Files and four AI Architect Skill Files.
Catalyst’s Dynamic MCP documentation says the service exposes more than 100 MCP tools through on-demand discovery.
The non-interactive CLI supports agent-friendly command execution without waiting for interactive prompts.
The platform documents separate development and production environments.
Zoho says production promotion remains manual and that agents do not directly touch production.
The company also documents scoped permissions, audit logging and production restrictions.
And Zoho publishes a vendor benchmark showing 90–95% task completion for three tested models when Catalyst Skills and MCP were added to its evaluation.
What We Should Not Claim Yet
We should not claim Catalyst makes arbitrary AI-generated code production-ready automatically.
We should not claim Dynamic MCP eliminates hallucinations.
We should not claim the 90–95% completion figures will reproduce across every application or coding agent.
We should not call Zoho’s benchmark independent.
We should not say a non-interactive CLI is equivalent to fully autonomous deployment.
We should not claim the agent can freely modify production; Zoho’s current positioning says the opposite.
And we should not claim using one full-stack platform removes the need for architecture, security review, testing, monitoring or cost management.
The Better Vibe-Coding Stack Has a Human Gate at the End
There is a temptation to judge agentic development by how little the human has to do.
Zero clicks. Zero confirmation. Zero review.
That may be the wrong metric.
A useful coding agent should remove repetitive implementation work. A useful platform should make the agent capable of operating the development environment. A safe production workflow should still know when to stop.
Catalyst 3.0’s most interesting idea is not that an AI can create a database table from Claude Code. Many platforms will eventually support that.
The more important idea is the boundary around the capability.
Give the agent platform knowledge. Give it tools. Give it a non-interactive automation path. Give it scoped permissions. Record its actions. Let it build aggressively in development. Then make a human decide when that work becomes production.
Maybe mature vibe coding is not unlimited autonomy. Maybe it is knowing exactly where autonomy should end.
AI coding agents are good at reading files, but large software systems are defined by relationships between files: calls, imports, inheritance, API consumers, execution flows and cross-repository dependencies. GitNexus indexes a codebase into a precomputed knowledge graph and exposes that structure to agents such as Claude Code, Cursor, Codex and Windsurf through MCP. The interesting idea is not another coding model. It is giving the model a structural map before it edits. GitNexus can surface callers, trace execution paths, estimate blast radius and map a git diff to affected flows — potentially reducing the amount of blind exploration an agent has to do as a vibe-coded project grows.
Vibe Coding Works Best Before the Codebase Has a Memory
A small project is unusually friendly to AI coding.
There may be ten files.
A route is easy to find.
A component imports one service.
A database helper lives in an obvious folder.
If an agent needs context, it can open a few files and reconstruct the system quickly.
That changes as the project grows.
A function that looks local may be called from twelve places.
A type change may affect an API handler, a background job and a test helper.
One frontend component may depend on a response shape that is produced three services away.
The codebase develops memory.
Relationships accumulate faster than any one file can explain them.
That is where vibe coding starts becoming dangerous.
The model can still write code.
The harder problem is knowing what the edit is connected to.
The File Tree Is Not the Architecture
Most coding agents begin with a filesystem.
Folders.
Files.
Names.
Search results.
That is useful, but architecture is not a directory listing.
The important questions are relational.
Who calls this function?
Which implementation satisfies this interface?
What endpoint returns the field this component reads?
Which services depend on this package?
What execution path reaches this database write?
Those answers may cross many files and repositories.
A file tree tells the agent where code lives.
It does not automatically tell the agent how the system behaves.
GitNexus is built around that distinction.
Its core idea is to index the relationships before the agent needs them.
GitNexus Is Not Another Coding Model
GitNexus does not replace Claude Code, Cursor, Codex, Windsurf or another coding agent.
It acts as context infrastructure around them.
The project indexes a repository into a graph containing software entities and the relationships between them.
Functions become nodes.
Classes become nodes.
Files become nodes.
Calls, imports, inheritance and other relationships become edges.
The resulting graph can then be queried through Model Context Protocol tools.
That means the coding model does not need to be retrained to understand one specific repository.
Instead, it receives structured answers about the repository at the moment it needs them.
For a vibe coder, that is a useful architectural pattern.
Keep the agent you already like.
Improve what the agent knows before it edits.
The Index Is Built Before the Prompt Arrives
The most important implementation choice is precomputation.
GitNexus does not wait for every agent request and then ask the model to explore the repository from scratch.
Its current documentation describes a multi-stage indexing pipeline.
It walks the repository structure.
It parses source code with Tree-sitter.
It extracts functions, classes, methods and interfaces.
It resolves imports and call relationships.
It groups related symbols into functional clusters.
It traces execution processes.
It builds search indexes.
The work is done ahead of the coding task.
When the agent later asks what depends on a symbol, much of the structural analysis has already happened.
That turns repository understanding from repeated exploration into reusable infrastructure.
Tree-Sitter Turns Source Files Into Syntax the Graph Can Reason About
Text search sees characters.
A parser sees structure.
GitNexus uses Tree-sitter parsers to extract language constructs from source code.
That lets it distinguish a function declaration from a string containing the same text.
It can identify classes, methods, interfaces and imports rather than treating every match as equivalent.
This matters because software relationships are defined by syntax and semantics, not simply by word similarity.
If an agent searches for “save,” it may find hundreds of unrelated text matches.
If a graph knows that one save method is called by a particular service method, the answer becomes much more specific.
Parsing is therefore the first step from code search toward code intelligence.
Resolution Is Where a Code Graph Becomes More Useful Than a Diagram
A graph is only useful if its edges mean something.
GitNexus says it resolves imports, function calls, inheritance, constructor inference and receiver types across files using language-aware logic.
Its current documentation gives a chained example such as user.address.getCity().save(), where the system attempts to resolve the receiver at each hop.
That is much more useful than drawing boxes around files.
The agent can ask about a symbol and receive callers or downstream relationships.
The graph becomes executable context.
But the word “resolve” also needs a limit.
Static analysis can be highly reliable when relationships are explicit.
Runtime reflection, dynamically generated code, unusual metaprogramming and some dynamic imports can still make static resolution incomplete.
Deterministic analysis does not mean omniscience.
GitNexus Tries to Discover Modules Instead of Trusting Folder Names
Large repositories are often organized imperfectly.
A folder named utils may contain authentication logic.
Billing code may span several packages.
A feature may have grown across layers over years.
GitNexus applies Leiden community detection to the relationship graph to group symbols into functional clusters.
The idea is that strongly connected symbols may reveal a real subsystem even when the repository structure does not describe it cleanly.
That can help an agent understand that a change belongs to an authentication flow or ingestion pipeline rather than merely to one folder.
This is an inferred architectural view.
It is not the same as a human-written architecture document.
But for an undocumented or fast-growing vibe-coded project, discovering communities from actual code relationships can be valuable.
The Most Useful Question May Be: What Breaks If I Change This?
Vibe coding often optimizes for the first edit.
Change the function.
Refresh the page.
If it works, continue.
The risk appears later when the same function has hidden dependents.
GitNexus exposes an impact-analysis tool designed around blast radius.
Given a symbol, it can trace downstream or upstream dependencies by depth and attach confidence to the relationships it returns.
That changes the workflow.
Before editing a shared service, the agent can ask what depends on it.
Before changing a return type, it can identify consumers.
Before renaming a symbol, it can inspect where the graph expects that symbol to participate.
The goal is not to predict every bug.
It is to turn dependency awareness into a normal pre-edit step.
detect_changes Moves Impact Analysis From a Symbol to a Git Diff
The current GitNexus CLI and MCP documentation includes a detect_changes tool.
Instead of starting with one manually selected function, the tool maps changed lines in the git diff to affected processes and graph entities.
That is interesting because real edits rarely touch exactly one symbol.
A vibe-coding session may change a component, a schema and an endpoint in one pass.
A diff-aware graph can ask a broader question:
Which execution flows are affected by everything that changed?
That moves the graph closer to review infrastructure.
The agent can inspect the likely blast radius after an edit instead of waiting for a failing test or a user report to reveal the connection.
trace Answers a Different Question: How Are These Two Things Connected?
Impact analysis expands outward.
Tracing tries to find a path.
GitNexus documents a trace tool that finds a directed path between two symbols using call and class-member relationships.
That can be useful when a developer knows the beginning and end of a behavior but not the middle.
A button triggers a request.
The request eventually writes to a database.
Where is the chain?
Without a graph, an agent may search one function, open its caller, follow an import, inspect another file and repeat.
A path query can compress that exploration into one structural answer.
The model still needs source code to understand what each step does.
But it no longer has to discover every step blindly.
The Current Repository Documents 17 MCP Tools
GitNexus exposes its code intelligence through MCP.
The current GitHub README documents 17 tools: 15 per-repository tools and two group-level tools.
They include general tools such as query, context, impact, trace, detect_changes and cypher.
The newer toolset also includes more specific checks such as route_map, shape_check and api_impact.
The public Akon Labs landing page still shows a smaller seven-tool count.
That appears to be a documentation/version mismatch rather than two different fundamental products.
For a fast-moving developer tool, this is worth noting.
The repository is the better source for the current CLI and MCP surface.
The marketing page is better for the broader product positioning.
route_map and shape_check Point Toward a More Practical Kind of Code Intelligence
A graph becomes more useful when it answers developer questions rather than simply exposing graph theory.
The current GitNexus tool list includes route_map, which maps API routes to callers and handlers.
It also includes shape_check, designed to compare API response shapes against the properties consumers access.
Those tools illustrate where code graphs can become useful for vibe coding.
A model may happily change an API response from user.name to user.displayName.
The backend still compiles.
The frontend may not.
A structural tool can surface that consumer relationship before the edit is treated as finished.
The value is not the graph visualization itself.
The value is converting graph relationships into checks that match how software actually breaks.
One Graph Can Span Multiple Repositories
The problem becomes harder when a system is split across repositories.
An API lives in one repo.
A web frontend lives in another.
A mobile app consumes the same endpoint from a third.
A shared schema package may live somewhere else.
Per-repository search creates artificial boundaries.
A breaking API change does not care which Git repository owns the consumer.
Akon Labs positions the managed and enterprise GitNexus system around unified multi-repository graphs.
Cross-repository edges can connect a service to downstream consumers.
For organizations, that may be the feature with the largest potential value.
The blast radius of a change can escape the repository long before it escapes the architecture.
For a Solo Vibe Coder, the Same Problem Arrives Earlier Than Expected
Multi-repo complexity sounds like an enterprise problem.
It is not only an enterprise problem.
A solo project can become structurally large very quickly with AI.
Agents write code faster than traditional manual development.
That means technical debt can also accumulate faster.
A person who would normally add five files in a week can generate fifty.
A prototype can become a product before the original author has built a reliable mental model of the new system.
This is one of the paradoxes of vibe coding.
AI reduces the cost of adding code.
That increases the value of tools that explain the code already added.
GitNexus is interesting because it targets the second half of that equation.
The Basic Local Workflow Is Deliberately Small
The current GitNexus README presents a two-command quick start.
Run gitnexus analyze from the repository.
Then run gitnexus setup.
The first command builds the index and installs agent-context integration.
The second configures MCP for supported coding tools.
The project currently documents integrations with Claude Code, Cursor, Codex, Antigravity, OpenCode, CodeBuddy, Qoder and Windsurf, with different levels of hooks and skills depending on the client.
For Claude Code and Codex, the repository documents deeper hook integration that can add graph context around tool calls and warn when the index becomes stale after repository changes.
The important point is that the graph is designed to sit inside the normal coding workflow rather than require a separate manual analysis session.
Local Indexing Is a Major Part of the Appeal
Source-code context can be sensitive.
A tool that improves coding by uploading the whole repository to another service creates a new trust decision.
GitNexus’s local CLI is designed to build and query the index on the developer’s machine.
The project documentation says local indexing and storage can operate without sending repository data to a remote GitNexus service.
Its browser mode similarly runs the graph in the browser for smaller repositories, while a managed enterprise offering also exists.
Those modes should not be conflated.
A local CLI workflow and a hosted SaaS workflow have different privacy boundaries.
For developers choosing the tool specifically because of local code intelligence, the deployment mode matters as much as the feature list.
The Web UI Is Convenient, but It Has a Different Scale Profile
GitNexus also provides a visual browser interface.
The README describes it as useful for quick exploration, demos and one-off analysis.
The local CLI is the recommended path for daily development and larger repositories.
The browser version uses in-memory WebAssembly storage and is constrained by browser memory, while the CLI uses native persistent storage.
This distinction is useful because the graph visualization is likely what attracts many developers first.
But the practical coding-agent value comes from persistent indexing and MCP access.
The pretty graph is the interface.
The reusable structural context is the infrastructure.
Language Support Is Broad, but Not Every Language Has the Same Depth
The current repository documents support across TypeScript, JavaScript, Python, Java, Kotlin, C#, Go, Rust, PHP, Ruby, Swift, C, C++ and Dart.
The capability matrix is not identical for each language.
Some languages have import resolution, named-binding tracking, inheritance analysis, type annotations, constructor inference and framework detection.
Others support only a subset.
Optional control-flow and program-dependence analysis is currently documented for TypeScript and JavaScript, with other languages planned.
This is another reason to avoid describing GitNexus as universally exact.
The quality of a code graph depends on what the parser and resolver can understand for a particular language and framework.
A TypeScript application may expose more structure than a codebase dominated by dynamic runtime behavior.
Deterministic Is Better Than Similarity for Some Questions — Not All Questions
Akon Labs contrasts GitNexus with embedding-based retrieval.
That comparison is strongest for questions with exact structural answers.
Who imports this module?
Which known callers invoke this function?
Which class implements this interface?
What static path connects these symbols?
Those are graph questions.
Semantic retrieval solves a different problem.
Where is code related to billing?
Which file discusses retry behavior?
What implementation is conceptually similar to this one?
Those may benefit from embeddings.
The current GitNexus query tool itself uses hybrid search that combines lexical and semantic retrieval.
So the useful lesson is not “graphs replace embeddings.”
It is that structural questions should not be answered only by similarity search when the codebase contains explicit relationships that can be resolved.
Akon Labs Reports a Large Benchmark Gain — but It Is Still a Vendor Benchmark
Akon Labs has published a DeepSWE benchmark comparing the same coding-agent setup with GitNexus, with a simpler Graphify extraction layer, and with no retrieval graph.
The benchmark reports 113 tasks across 89 open-source projects and 3,471 trials.
In that setup, GitNexus achieved a 68.37% pass rate.
The bare setup achieved 36.99%.
The company also reports fewer steps, fewer output tokens and lower average cost per trial with GitNexus.
Those results are interesting because the comparison tries to hold the model and agent scaffold constant.
They are not independent validation.
Akon Labs designed and published the benchmark implementation and is evaluating its own product.
The numbers should therefore be described as company-reported benchmark results, not universal proof that every agent becomes nearly twice as capable.
The Often-Repeated “51% Cheaper” Number Needs a Denominator
Product Hunt and Akon Labs describe GitNexus as making coding-agent runs roughly 51% cheaper in their public benchmark.
The detailed benchmark page makes the denominator clearer.
Average cost per trial falls from $0.6631 for the bare model to $0.6008 with GitNexus.
That is about 9.4% lower per attempt.
The much larger saving appears when cost is divided by successfully solved tasks.
Akon Labs reports $0.88 per solved task with GitNexus versus $1.79 for the bare setup.
That is roughly a 51% reduction in cost per successful fix.
Both measurements are legitimate.
They answer different questions.
A careful article should not describe a 9.4% per-run saving as a 51% cheaper run.
The Benchmark Is Also Narrower Than the Product’s Biggest Claim
Akon Labs explicitly acknowledges this limitation.
DeepSWE consists of individual bugs and feature requests in open-source repositories.
That can measure how well an agent navigates an unfamiliar codebase while fixing one issue.
It does not directly test the product’s most ambitious claims.
Cross-repository impact analysis.
Large organizational graphs.
Pre-merge blast-radius checks across services.
Long-lived graph reuse in a production team.
Those require different benchmarks.
In other words, the published evaluation tests the code-reading advantage.
It does not independently establish the full enterprise value proposition.
That does not weaken the benchmark.
It defines what the benchmark can actually support.
There Is an Important Licensing Catch for Commercial Vibe Coders
GitNexus is frequently described on the Akon Labs site and Product Hunt as open source.
The current GitHub repository is publicly accessible and its code can be inspected.
However, the repository license is PolyForm Noncommercial 1.0.0.
That license explicitly limits the granted software rights to noncommercial purposes, with specified exceptions.
That is materially different from permissive licenses such as MIT, Apache-2.0 or BSD.
A developer building a commercial product should not assume that the free repository can be used commercially simply because the source is visible.
Akon Labs also offers commercial enterprise options.
For TUF, the accurate wording is that GitNexus is source-available with a noncommercial license, while the company markets it using the “open source” label.
Commercial users should read the current license or obtain appropriate commercial terms.
A Code Graph Cannot Replace Tests
Knowing the blast radius is not the same as proving the change works.
Static relationships cannot model every runtime behavior.
External services may return unexpected data.
Configuration may change code paths.
Reflection can hide relationships.
Feature flags can alter execution.
A database migration can fail even when the call graph is correct.
A user can click a sequence nobody anticipated.
GitNexus should therefore sit before and beside testing, not replace it.
A better vibe-coding loop would be:
understand the graph,
estimate impact,
make the edit,
run tests,
exercise the interface,
inspect the diff,
then recheck affected flows.
The graph helps the agent ask better questions.
Tests still determine whether the software behaves correctly.
The Best Use Case Is Not Generating More Code
Most AI coding products sell speed.
Generate faster.
Refactor faster.
Ship faster.
GitNexus is more interesting when viewed as a braking system.
Before editing, ask what depends on this.
After editing, ask which flows changed.
Before renaming, inspect callers.
Before changing an API, inspect consumers.
That may add a step to the prompt.
But it can remove much more expensive steps later.
The strongest vibe-coding tools may not be the ones that generate the most code.
They may be the ones that prevent the agent from confidently changing code it does not understand.
What GitNexus Currently Documents
The current public GitNexus repository documents a local indexing pipeline based on Tree-sitter and graph construction.
It documents knowledge-graph entities and relationships for code structure, clustering and execution processes.
It exposes that context through MCP to multiple coding agents.
The repository currently lists 17 MCP tools, including query, context, impact, trace, detect_changes, route_map, shape_check and api_impact.
It documents support across a broad set of programming languages with different levels of analysis depth.
Akon Labs reports multi-repository graph capabilities in its commercial platform.
The company’s DeepSWE benchmark reports a 68.37% pass rate with GitNexus versus 36.99% for the bare setup under its published configuration.
And the current repository is licensed under PolyForm Noncommercial 1.0.0.
What We Should Not Claim Yet
We should not claim GitNexus understands every runtime dependency.
We should not claim static analysis can always resolve dynamic imports, reflection or generated code.
We should not claim the graph replaces tests.
We should not claim the 68.37% benchmark result will reproduce for every model, language or private codebase.
We should not call the vendor benchmark independent validation.
We should not say every GitNexus run is 51% cheaper; the large reduction in the published benchmark is cost per successful solved task, while cost per trial falls by about 9.4%.
We should not say all 17 MCP tools are equally supported by every editor.
And we should not describe the current repository as permissively open source for commercial use.
The license is explicitly noncommercial.
The Bigger Vibe-Coding Upgrade Is Giving the Agent a Model of the System
AI made code generation cheap.
That changes what becomes expensive.
Understanding.
Review.
Dependency tracking.
Regression discovery.
Architecture.
A small vibe-coded project can survive on file search because the developer and the model can still reconstruct the system quickly.
A large one cannot.
At some point, the question stops being “Can the AI write this function?”
The harder question becomes “Does the AI know what this function belongs to?”
GitNexus is one answer to that problem.
It turns a codebase from a pile of files into a queryable relationship map and gives that map to the coding agent through MCP.
The model still writes the code.
The tests still have to pass.
The developer still owns the result.
But the agent no longer has to navigate entirely by guesswork.
That may be one of the most important upgrades for vibe coding as projects stop being prototypes and start becoming software.
Sonos Is Opening Its Speakers to ChatGPT and Gemini Through MCP — Here’s How It Works
ChatGPT can answer a question.
Gemini can plan something.
Now Sonos is opening a route for AI assistants to reach outside the chat window and control the audio system in your home.
The new piece is called Sonos 27mcp.
It is an official Model Context Protocol server hosted by Sonos. The company says any large language model capable of using an MCP server may connect and, once authorized by the Sonos user, control that user’s Sonos system.
Sonos staff explicitly names Claude, ChatGPT, Gemini and self-hosted AI systems as examples.
That does not mean ChatGPT is being installed inside every Sonos speaker.
The conversation can stay in the AI interface you already use. MCP becomes the bridge. Sonos remains the system that performs the audio action.
The first Early Access window for Sonos 27mcp starts September 8, 2026 for US-English users, with a wider rollout to follow.
The important shift is not that an AI can press Play. Smart speakers have handled playback commands for years. The shift is that the reasoning interface and the physical speaker system no longer have to be the same product.
Sonos 27mcp Connects AI Assistants to the Sonos System
Sonos has published a hosted MCP endpoint for the new system.
The server address is:
That endpoint is the technical bridge between an MCP-capable AI and the user’s Sonos environment.
The basic chain is:
AI assistant → MCP server → Sonos authorization → Sonos system.
The assistant handles the conversation. The MCP layer exposes the tools or actions Sonos chooses to make available. The Sonos system performs the resulting command.
This is a different architecture from a traditional smart speaker where one assistant is tightly bound to the device.
Here, the AI model can live somewhere else entirely. The speaker becomes the physical endpoint for an external reasoning system.
This Is Different From Putting ChatGPT Inside a Speaker
The distinction matters because the headline can otherwise sound more dramatic than the actual architecture.
Sonos is not saying that ChatGPT becomes the built-in operating system of every Sonos speaker.
Sonos is also not announcing an exclusive OpenAI integration.
The company describes 27mcp as an open connection point for LLMs that can use MCP.
That means the structure is closer to:
conversation happens in ChatGPT, Gemini, Claude or another compatible AI → the AI calls Sonos through MCP → Sonos carries out the action.
The model and the speaker remain separate systems. MCP gives them a standardized way to work together after the Sonos owner authorizes the connection.
That is a much more useful way to understand the announcement than treating it as another voice-assistant swap.
The AI Can Discover and Control the Sonos System
Sonos says 27mcp allows a compatible AI to discover and control the user’s Sonos system.
Discovery is important.
A multi-room Sonos setup is not one anonymous speaker. It can contain different rooms, grouped products and different playback states across the home.
An AI agent needs some understanding of that environment before it can act usefully.
Sonos gives an example of asking an assistant to find something new and play it in the office without leaving the conversation.
That interaction contains several steps even though the request feels simple to the user.
The AI has to interpret what “something new” means in context. It has to identify the office as a valid Sonos destination. Then it has to make the appropriate Sonos action through MCP.
The user sees one conversation. Behind it is a tool chain.
MCP Is the Bridge Between Reasoning and Action
Model Context Protocol is useful here because it separates reasoning from execution.
The language model does not need to contain Sonos control code inside the model itself. Instead, the AI can connect to an MCP server that describes the tools available to it.
The simplified flow is:
user request → AI interprets intent → MCP exposes Sonos actions → AI selects an action → Sonos executes it.
That is the same broad pattern behind the current push toward AI agents.
A chatbot can already explain how to change the music. A tool-connected agent can ask the actual system to change it.
MCP gives Sonos a standardized interface for that second step.
The interesting part is not audio alone. It is the transition from conversational intelligence to an authorized physical action in the home.
Authorization Comes Before Control
A compatible AI does not automatically gain access to a Sonos system.
Sonos states that the LLM can control the system once it has been authorized by a Sonos user.
That creates an important boundary:
MCP compatibility ≠ automatic device access.
The actual chain is:
compatible AI + user authorization → access to the Sonos MCP interface → Sonos actions.
Sonos has not published every security and permission detail for the Early Access experience yet, so the article should not invent them.
There is no basis to claim that every AI gets every Sonos capability. There is also no basis to claim that authorization is permanent, universal or shared across every model.
What Sonos has confirmed is the principle: the owner has to authorize the connection before an outside LLM can control the system.
ChatGPT, Gemini and Claude Are Examples — Not an Exclusive List
The recognizable names make this announcement easy to explain.
Sonos staff explicitly lists Claude, ChatGPT and Gemini when describing the kinds of AI people may already use.
But the protocol is broader than those brands. Sonos says any LLM capable of using an MCP server may connect. That can include systems the user runs themselves.
So the product strategy is not:
build one integration for ChatGPT, then another integration for Gemini, then another integration for Claude.
The MCP approach is closer to:
publish one standardized interface → let compatible AI clients connect to it.
That does not guarantee identical support in every AI product. Each client still needs to support MCP and the relevant authorization flow. But it changes the integration model from one assistant at a time to a protocol-based connection.
Sonos 27mcp Starts Early Access on September 8
The announcement is current, but the feature is not generally available to everyone today.
Sonos says Sonos 27mcp enters Early Access on September 8, 2026.
The first release is for US-English users. The company says a larger rollout will follow.
That distinction matters.
The correct framing is:
announced now → Early Access September 8 → broader rollout later.
It is not:
ChatGPT can already control every Sonos system worldwide today.
Early Access also means the experience can still change as Sonos gathers feedback and expands the platform.
The MCP server itself is already publicly identified, but access to the consumer experience follows Sonos’s rollout schedule.
The Bigger Idea Is Multi-Room Control Through an AI Conversation
Sonos becomes more interesting when the system contains several rooms.
A normal app interface makes the user think in controls: select room, select source, choose music, start playback, group another room, adjust volume.
An AI interface can let the user start with intent instead.
For example:
put something relaxed downstairs.
That request is not a complete list of device commands. The AI has to interpret “relaxed.” It has to understand what “downstairs” corresponds to in the Sonos system. Then it has to translate that intent into actual playback and room actions.
The final system can therefore look like:
natural-language intent → reasoning → room selection → content selection → playback action.
That is more significant than adding another button to the Sonos app.

Sonos Is Also Building a Separate LLM Voice Assistant
Sonos 27mcp is only one part of Sonos 27.
The company is also introducing Sonos 27voice.
These two systems should not be confused.
Sonos 27mcp connects an external MCP-capable AI to Sonos.
Sonos 27voice is Sonos’s own next-generation voice assistant for voice-enabled Sonos products.
Sonos says 27voice uses large language model technology to support more natural interactions than traditional command-based assistants. It also adds a new “ask me anything” domain.
So Sonos is pursuing two paths at the same time.
One path lets outside AI systems control Sonos. The other upgrades the intelligence of the assistant Sonos provides itself.
That makes Sonos 27 less about one new assistant and more about opening several AI entry points into the same speaker system.
27voice Can Understand Requests That Are Less Precise
Traditional voice assistants work best when the user already knows the command structure.
Sonos 27voice is designed for less precise language.
Sonos says it can understand implicit or vague requests, handle natural dialogue and ask clarifying questions.
The company gives examples such as describing an album by its cover or referring indirectly to an artist instead of naming them exactly.
That changes the interaction model.
The user does not always have to translate a thought into a perfectly structured command first. The assistant can do more interpretation before choosing the audio action.
This is where LLM technology matters more than it does in a simple “play/pause” command. The model is being used to resolve language and context before the speaker system acts.
27voice Can Handle Chained Commands
Sonos says 27voice can handle complex, chained requests.
That means one spoken request can contain several actions or references instead of requiring a separate command for each step.
Sonos provides examples that combine choosing an artist, selecting a room and grouping another room in the same request.
The architecture becomes:
one natural-language request → multiple interpreted actions.
This matters for a multi-room system because real user intent often crosses several controls at once.
A person may not think: first choose the track, then change the room, then group another speaker.
They think: play this there, and include that room too.
LLM-based parsing gives Sonos a way to map that single thought onto several device operations.
27voice Adds an “Ask Me Anything” Domain
Sonos is also expanding beyond direct audio control.
Its support documentation says Sonos 27voice includes an “ask me anything” domain.
The examples include general information questions as well as follow-up conversation.
That moves the speaker from a narrow command interface toward a broader conversational interface.
The device can still control music and the Sonos system. But it can also answer questions unrelated to playback.
The important boundary is that this is Sonos 27voice, not Sonos 27mcp.
27voice is the built-in conversational assistant path. 27mcp is the external-AI tool path.
Both are part of the same broader Sonos 27 platform, but they solve different problems.
Sonos Says Some Everyday 27voice Work Can Be Handled Locally
Sonos staff says 27voice handles everyday tasks quickly and locally, and reaches for a larger model when a question calls for it.
That suggests a hybrid architecture rather than sending every interaction through the same reasoning path.
At the same time, Sonos support documentation notes that 27voice requires a cloud connection for basic playback or information requests on a portable Sonos product when it is in Bluetooth mode.
Those statements are not enough to reconstruct the full internal architecture.
They do show that “local” does not mean the entire 27voice experience is offline.
The safe interpretation is narrower:
Sonos says some everyday processing is handled locally, while the broader service still relies on cloud connectivity for parts of the experience.
The exact routing rules, model identities and thresholds have not been fully published.
27voice Can Reach Beyond Music
Sonos support documentation lists third-party smart-home integrations for 27voice, including Philips Hue lighting and Lutron home automations.
That expands the conversational surface beyond audio.
A Sonos speaker can therefore become one place where the user talks not only about music but also about supported home actions.
This does not make Sonos a universal smart-home operating system. The supported integrations are still bounded by what Sonos exposes and what third-party services support.
But the direction is clear.
The speaker is becoming a conversational control point for more than the speaker itself.
That makes the combination of 27voice and 27mcp particularly interesting. One opens Sonos to outside AI systems. The other expands what Sonos’s own assistant can understand and control.
Custom Agents Push the Idea Further
Sonos has also announced Custom Agents.
Sonos staff describes these as user-built agents that can have their own summon phrase, voice, personality and model choice while living in the same speaker ecosystem.
This is another reason not to think of Sonos 27 as a single-assistant launch.
The platform is moving toward multiple intelligence layers.
A user might have Sonos 27voice for the standard Sonos experience. An external AI could reach the system through MCP. Custom Agents could provide separate personalities or task-oriented experiences.
Sonos has not published every implementation detail yet, so the article should not treat Custom Agents as a fully defined developer platform today.
But the announced direction is clear enough to describe: the same speaker hardware can become an endpoint for different agent experiences.
One Speaker Can Become a Front End for Multiple AI Systems
This is the larger platform change.
For years, smart speakers were closely identified with one assistant.
Alexa speaker. Google Assistant speaker. Siri speaker.
Sonos 27 points toward a looser relationship.
The hardware remains Sonos. The audio system remains Sonos. But the intelligence interacting with that hardware can come from several places.
The stack can look like:
speaker hardware
↓
Sonos 27 platform
↓
Sonos 27voice / Sonos 27mcp / Custom Agents
↓
different AI models and interfaces.
The model no longer has to be the identity of the device. The speaker can instead become a physical interface for multiple intelligence layers.
That is a more flexible architecture than tying every capability to one permanently embedded assistant.
What Sonos Has Confirmed — and What It Has Not
Sonos has confirmed that Sonos 27mcp is an official hosted Model Context Protocol server.
The published endpoint is mcp.ws.sonos.com/mcp.
Sonos says any LLM capable of using an MCP server may connect and, once authorized by a Sonos user, control that user’s Sonos system.
Sonos staff names Claude, ChatGPT, Gemini and self-hosted AI systems as examples.
Early Access for 27mcp starts September 8, 2026, beginning with US-English users.
Sonos has also confirmed that 27voice uses LLM technology, supports natural dialogue, chained commands and an “ask me anything” domain, and is planned for Early Access in fall 2026.
What Sonos has not said is equally important.
It has not said ChatGPT is installed inside Sonos speakers.
It has not announced an exclusive OpenAI or Google partnership for 27mcp.
It has not said every Sonos function will be exposed to every MCP client.
It has not published the complete permission model, internal model-routing logic or every security detail for the Early Access system.
Those gaps should remain gaps.
The Real Change Is Conversation → Tool → Speaker
The interesting part of Sonos 27mcp is not that an AI can press Play.
Smart speakers have handled playback commands for years.
The change is where the decision can come from.
A user can stay inside an AI conversation.
The model interprets the request.
MCP gives that model an authorized route into the Sonos system.
Sonos performs the physical action in the home.
That creates a new chain:
conversation → reasoning → tool call → room → speaker.
Alongside Sonos 27voice and Custom Agents, the speaker starts to look less like a device permanently tied to one assistant and more like an endpoint for different AI systems.
The speaker still produces the sound. Sonos still controls the system. But the intelligence deciding what should happen can now come from somewhere else.
Agent Plugins 1.0 Standardizes the Package Around Agent Extensions
AI agent extensions already had reusable parts before Agent Plugins 1.0.
Agent Skills can package instructions, scripts and reference material.
MCP servers can expose tools and data.
The problem was not that those components did not exist.
The problem was that different agent clients could expect different packaging around them.
Agent Plugins 1.0 adds a shared package format.
Version 1.0.0 defines one directory structure that compatible clients can inspect and load.
At the root is plugin.json.
Skills go under skills/.
MCP server configuration goes in mcp.json.
Client-specific additions can live in namespaced directories.
The standard therefore sits around the components rather than replacing them.
An Agent Skill remains an Agent Skill.
An MCP server remains an MCP server.
Agent Plugins defines how those pieces travel together as one portable extension package.
AWS, Cursor, Microsoft, OpenAI and Vercel are represented on the initial Technical Steering Committee, and the specification is developed publicly.
The result is a new layer in the agent stack: not another tool protocol, but a common package contract for reusable agent capabilities.
The Smallest Plugin Starts With plugin.json
Every conforming Agent Plugin begins with one required file.
plugin.json.
The manifest identifies the package and the Agent Plugins specification version it targets.
For version 1.0.0, the required schema identifier is the canonical Agent Plugins 1.0 manifest schema.
The minimal manifest can contain only that schema reference and a plugin name.
Additional metadata can describe version, author, homepage, repository, license, keywords and client extension data.
The important part is predictability.
A compatible client knows where to look first.
It does not need to guess whether the package uses one vendor’s root manifest or another vendor’s directory convention.
It checks plugin.json.
It validates the declared format.
Then it discovers the component types it supports.
That sequence gives the package a stable identity before any skill or MCP server is loaded.
The manifest is small because Agent Plugins 1.0 deliberately keeps most component details in their own standard locations instead of turning plugin.json into a large universal configuration file.
Skills Live in a Fixed skills Directory
Agent Plugins 1.0 defines Agent Skills as one of its two portable component types.
The package uses a fixed skills/ directory.
Each immediate child directory can contain one SKILL.md file and the related resources defined by the Agent Skills specification.
Scripts can live beside it.
References can live beside it.
Assets can live inside the skill structure defined by the underlying Skills format.
Agent Plugins does not redefine how SKILL.md works.
The Agent Skills specification remains the source of truth for the skill itself.
Agent Plugins only defines where compatible clients should discover those skills inside the portable package.
That separation is useful.
The skill author keeps one skill format.
The plugin package adds a predictable location around it.
A client that supports Agent Skills can inspect skills/ and load the skills it understands.
A client that supports another portable component can continue loading that component independently.
The package becomes a container for reusable agent behavior without creating a new instruction language.
MCP Configuration Gets One Portable Root File
The second portable component type in Agent Plugins 1.0 is MCP server configuration.
The standard places it in mcp.json at the plugin root.
That file can describe local stdio MCP servers or remote MCP endpoints using supported transports.
The current specification includes stdio, Streamable HTTP and legacy HTTP+SSE configuration.
Each server entry declares its type explicitly.
Agent Plugins does not redefine the MCP wire protocol.
The Model Context Protocol specification still defines MCP behavior and lifecycle semantics.
Agent Plugins defines the portable configuration used to locate and connect to MCP servers that travel with the plugin.
A compatible client can then translate that portable description into its own native runtime configuration.
This is another example of the packaging layer staying separate from the component standard underneath it.
MCP defines how the agent talks to the server.
Agent Plugins defines how the package tells compatible clients that the server is part of the extension.
Skills and MCP Servers Can Travel in the Same Directory
The format becomes more interesting when the two component types are combined.
A plugin can ship a skill that explains how an agent should perform a task.
The same package can also ship an MCP server configuration that gives the agent the tools needed to perform that task.
AWS gives the example of packaging a deployment runbook with its tool integration.
The instruction layer and the tool layer remain separate, but they travel together.
That lets an extension author design one coherent capability.
The skill can contain the workflow.
The MCP server can expose the external operations.
The root manifest gives the package an identity.
The compatible client discovers whichever components it supports.
This creates a useful package boundary.
The plugin is no longer only a prompt bundle or only a tool connection.
It can be a reusable unit containing both instructions and executable integrations while preserving the standards that define each part.
A Fixed Directory Layout Creates the Portable Contract
The core directory layout is intentionally small.
A typical package can look like this:
plugin.json at the root.
skills/ for Agent Skills.

mcp.json for MCP servers.
Optional reverse-domain directories for client-specific extensions.
LICENSE and documentation can sit beside them.
The fixed locations are part of the portability contract.
plugin.json cannot move the skills directory somewhere else.
MCP configuration cannot be embedded into an arbitrary core field.
Compatible clients discover the same portable components from the same locations.
That gives authors one predictable structure to publish.
It gives client implementers one predictable structure to load.
The format is therefore less about inventing new component types and more about removing packaging ambiguity around existing ones.
Portability comes from agreeing on where the pieces go and how the root manifest identifies the package.
Client-Specific Features Use Reverse-Domain Namespaces
Portable standards still need room for clients to add their own capabilities.
Agent Plugins 1.0 handles that with reverse-domain extension namespaces.
A client can place manifest data under an extension key such as com.example.client.
It can also use a top-level directory with the same namespace.
Other clients ignore namespaces they do not implement.
GitHub Copilot uses this pattern with com.github.copilot.
VS Code documentation says Copilot-specific custom agents, slash commands, rules and hooks can live in that namespace while the same package keeps its portable skills and MCP configuration in the standard locations.
This creates two layers inside one directory.
The portable core is shared.
Client-specific behavior stays namespaced.
An extension author therefore does not have to choose between portability and every client-specific capability.
The common parts can remain portable.
Extra behavior can stay attached to the package for clients that understand it.
The Standard Leaves Installation and Distribution to Clients
Agent Plugins 1.0 does not try to standardize the entire plugin ecosystem.
The official documentation says distribution, installation, permissions, user experience and client-specific capabilities remain under each client’s control.
That scope is deliberate.
The shared format defines the package.
Clients still decide how users discover it.
A client can use a marketplace.
Another can install from a local directory.
Another can integrate plugin management into an enterprise policy system.
The standard does not need every client to have the same interface.
It only needs the compatible parts of the package to remain recognizable.
This creates an interoperability floor rather than one universal agent application model.
The package can travel.
The installation experience can still differ.
That gives vendors room to design their own client behavior while keeping a shared structure underneath.
Compatible Clients Can Adopt Component Types Incrementally
A client does not need to implement every possible plugin capability at once.
The Agent Plugins compatible-clients page documents support by component type and MCP transport.
Current listed clients include VS Code, Cursor, GitHub Copilot, ChatGPT and Codex, Kiro, Hermes Agent and OpenClaw.
All of those listed clients support Agent Skills.
Their MCP transport support varies by client.
The specification says clients can ignore component types they do not support.
That makes adoption incremental.
A client can begin with skills.
It can add MCP support.
It can support more transports later.
The plugin package does not have to change its identity each time.
This is useful for an emerging standard because the ecosystem can converge gradually.
Portability does not require every client to expose exactly the same feature set.
It requires the common pieces to remain discoverable and interpretable when they are supported.
GitHub Copilot and VS Code Already Use the 1.0 Format
GitHub announced general availability of Agent Plugins 1.0 support in VS Code, Copilot CLI, the GitHub Copilot SDK and the GitHub Copilot app on August 12, 2026.
GitHub says the format lets authors build a plugin once and use it across compatible clients.
For GitHub’s own tools, the portable Agent Skills and MCP configuration stay in the standard locations.
Copilot-specific files can live under com.github.copilot.
GitHub also says existing Copilot plugin formats remain supported, so the 1.0 format adds a new portable path rather than requiring every older package to be rewritten immediately.
This is an important stage for a specification.
It is not only a document.
A major agent client family is loading the format in production tools.
The common package contract is therefore already connected to real installation and runtime workflows.
VS Code Can Detect the Package From Its Schema
VS Code’s current documentation shows how a client identifies Agent Plugins 1.0.
A root plugin.json with the canonical Agent Plugins schema uses Agent Plugins semantics.
VS Code then discovers skills from skills/ and MCP configuration from mcp.json.
The editor also understands client-specific namespaces it owns and ignores namespaces owned by other clients.
That gives the package a clear loading path.
The schema identifies the portable format.
The fixed directories identify the components.
The namespace identifies client-specific additions.
The runtime can then decide which supported parts to activate.
This layered discovery model is what turns a folder into a standard package rather than a collection of loosely related files.
PLUGIN_ROOT and PLUGIN_DATA Give Packaged Files Stable References
Portable packages also need a way to refer to their own files.
Agent Plugins 1.0 defines placeholders such as PLUGIN_ROOT and PLUGIN_DATA for supported configuration fields.
PLUGIN_ROOT refers to the plugin’s installed root.
PLUGIN_DATA refers to client-managed writable state that can persist across plugin updates.
That distinction is useful.
Packaged scripts and configuration can reference files that ship with the plugin.
Writable runtime state can be placed in a separate client-managed location.
The plugin does not need to know the exact absolute installation path on every operating system or client.
The runtime expands the portable placeholder.
This is another small part of the format that supports movement between clients and machines.
The package describes relationships relative to itself instead of hard-coding one installation environment.
Versioning Lets Clients Know Which Contract a Plugin Targets
The manifest schema also acts as a version selector.
Agent Plugins 1.0.0 uses a canonical schema identifier.
A compatible client reads that identifier before interpreting the package.
If the client supports the declared Agent Plugins version, it can apply the corresponding validation and discovery rules.
That creates a path for the format to evolve without making the package ambiguous.
The plugin states which contract it expects.
The client states which contracts it implements.
The optional version metadata for the plugin itself can then be used separately for plugin updates and cache freshness.
There are therefore two version concepts.
The Agent Plugins specification version defines the package format.
The plugin’s own version identifies the extension release.
Keeping those concepts separate makes updates easier to reason about.
The Project Uses Public Governance Across Several Companies
Agent Plugins is also designed as a multi-company project.
The official site says the initial Technical Steering Committee includes core maintainers from Amazon, Cursor, Microsoft, OpenAI and Vercel.
GitHub participated in refining the 1.0 proposal, and Google joined as a core maintainer when the format launched according to GitHub’s August announcement.
Technical proposals and decisions are developed publicly.
The specification and schemas are openly licensed.
That governance structure matters because the package is supposed to travel between clients operated by different companies.
A portability format becomes more useful when several implementers can participate in defining the shared contract.
The common directory is therefore not presented as one client’s private plugin layout exported to everyone else.
It is being developed as an open interoperability project around components that already cross agent ecosystems.
Version 1.0.0 Is Public While the Normative Page Still Says Working Draft
There is one status detail worth stating precisely.
AWS, GitHub and the Agent Plugins project describe version 1.0.0 as publicly available.
The official specification page currently identifies the normative document as Spec Version 1.0.0 and labels its status Working Draft.
Those statements can coexist.
The 1.0.0 format is available for implementation and is already supported by real clients.
At the same time, the normative documentation still carries a Working Draft status label.
For developers, the practical rule is to target the published 1.0.0 schema and current normative specification used by supporting clients.
For editorial accuracy, it is better to preserve both facts instead of simplifying the status into either “just a draft” or “fully frozen forever.”
The implementation ecosystem is active.
The specification remains an openly developed technical standard.
Agent Plugins 1.0 Turns Extension Packaging Into a Shared Layer
The agent ecosystem now has several different standards solving different problems.
Agent Skills define reusable instructions and resources.
MCP defines model-to-tool and model-to-data connections.
Agent Plugins defines how skills and MCP server configurations can be packaged together for compatible clients.
That distinction is the important part.
The format does not need to replace the components below it.
It gives them a common container.
plugin.json identifies the extension.
skills/ contains portable skills.
mcp.json contains portable MCP configuration.
Reverse-domain namespaces carry client-specific additions.
Clients decide how installation, distribution, permissions and user experience work.
Current implementations already span VS Code, Cursor, GitHub Copilot, ChatGPT and Codex, Kiro and additional agents listed by the project.
The result is a packaging layer that can sit above several reusable agent technologies.
The skills stay skills.
The MCP servers stay MCP servers.
The package becomes portable.
That is the upgrade.
AI Agents Are Moving From Software Tools Into Physical Equipment
AI agents have already learned how to work with software.
They can call APIs, search databases, edit files, use developer tools and coordinate workflows across applications.
The next interface is physical equipment.
A robot arm has position, speed, payload and safety limits.
A camera has exposure, focus and image data.
A laser system has tunable parameters and sensor readings.
A manufacturing machine may expose its own commands, status values and control software.
The Model Hardware Standard, or MHS, is designed to give AI agents a common way to understand that equipment.
Anthropic opened MHS as a limited research preview on August 27, 2026 after developing it with HHMI Janelia Research Campus and a group of partners across science, robotics, electronics and manufacturing.
The central idea is straightforward.
A physical device should be able to describe what it is, what it can measure, what can be changed and which operating limits must be enforced.
An AI agent should then be able to discover that description and interact with the device through a common interface.
That takes the agent architecture beyond software tools.
The tool can now be a physical machine.
MHS Adds a Standard Driver Between the Agent and the Device
The main technical layer in MHS is the driver.
A driver is software that translates between a higher-level system and the specific hardware underneath.
Operating systems already use this pattern.
A printer driver turns generic print operations into commands a specific printer understands.
A graphics driver connects a software graphics stack to a GPU.
MHS applies a related idea to AI-controlled equipment.
The device gets an MHS driver.
That driver exposes the machine through a standard structure instead of requiring the agent to understand every vendor-specific interface directly.
Anthropic describes simple primitives such as read and write.
Read can request a value such as temperature, position or another sensor state.
Write can change a supported parameter.
The actual machine may still use a proprietary SDK, serial protocol, command-line tool or control application underneath.
The MHS layer translates the common operation into whatever the device requires.
That separation matters.
The agent can reason at the level of the task.
The driver handles the device-specific control path.
The same agent architecture can therefore interact with several types of equipment while each machine retains its own implementation underneath.
A Device Can Describe Itself in a Machine-Readable Reference
Control commands are only one part of physical operation.
An agent also needs context.
A robot may have a maximum payload.
A motor may have a permitted speed range.
A camera may expose several imaging modes.
A temperature controller may have upper and lower operating limits.
Some of that information exists in code.
Some exists in manuals.
Some exists as operating knowledge held by the people who work with the equipment.
MHS adds a structured way to bring that information into the interface.
Anthropic says MHS drivers support tags where users can describe machine characteristics in natural language.
The system can then produce a reference file describing the device.
That reference can include what the machine measures, which parameters can be adjusted and which safety limits are enforced.
The result is more than a list of functions.
It is a device model.
The agent receives both verbs and context.
Read this value.
Set this parameter.
This part moves.
This value has a defined range.
This machine has a physical characteristic the agent should know before operating it.
That turns hardware documentation into part of the agent’s working interface.
Discovery Lets Agents Find Devices Through a Common Format
A shared hardware interface becomes more useful when devices can also be discovered consistently.
MHS is designed to make connected equipment discoverable in a standard format across a network.
That gives an agent a way to learn what hardware is available before it starts building a workflow.
A system might expose a camera.
Another node might expose a robotic arm.
A third might expose a sensor controller.
The agent can inspect the available devices, read their reference information and determine which capabilities fit the task.
This resembles what has happened in software-agent systems.
A software agent can discover tools exposed by an MCP server.
Each tool has a name, a description and an input structure.
MHS extends the same general idea into physical devices.
The available capability is not only “search database” or “write file.”
It can be “read camera,” “move arm,” “set controller value” or another machine-specific operation represented through the common hardware layer.
Discovery changes orchestration.
The agent does not have to begin with one hard-coded machine.
It can begin with the equipment that is present.
MCP Becomes One of the Ways an Agent Reaches MHS Hardware
MHS does not replace the Model Context Protocol.
It can sit behind it.
Anthropic says MHS is model-agnostic and can be accessed by agent harnesses through standard protocols such as MCP.
That creates two distinct layers.
MCP gives an AI application a standard way to discover and call tools.
MHS gives physical equipment a standard way to describe and expose its hardware capabilities.
The layers can connect.
An MCP tool can represent an action that ultimately reaches an MHS driver.
The agent sees a callable operation.
MHS handles the relationship between that operation and the physical device.
This is why MHS is a different topic from MCP even though the two standards can work together.
MCP standardizes model-to-tool interaction.
MHS standardizes the hardware side of the tool.
The distinction becomes visible when a workflow crosses from software into the physical world.
The agent may use MCP to reach the tool.
The tool may use MHS to reach the machine.
The result is a stack rather than one protocol trying to describe everything.
Command Line and Code Files Give MHS More Than One Control Path
Anthropic describes three control mechanisms for MHS: MCP, command-line interfaces and code files or APIs.
That gives the agent several ways to operate the same hardware layer.
Interactive reasoning can use MCP calls.
A developer or operator can use command-line access.
A long-running workflow can be packaged into code.
The code path is important because an AI agent does not need to reason through every low-level step forever.
The agent can explore a device, learn a sequence and then turn that sequence into deterministic code.
Anthropic demonstrated this pattern with laser alignment.
Claude adjusted the laser, observed the result through a camera and repeated the process while learning the relationship between control changes and the observed beam.
It then packaged the learned sequence into a script.
The final operation could run as a deterministic command rather than requiring live reasoning for every adjustment.
That creates two operating modes.
Reason while discovering the procedure.
Execute code after the procedure has been defined.
MHS provides one hardware interface underneath both.
The Standard Can Coordinate Several Machines as One Workflow
The main architectural value appears when a task needs more than one device.
A production cell may contain several robot arms.
An imaging system may combine cameras, motors, sensors and optical equipment.
A test setup may use measurement devices on different computers.
MHS gives the agent one structured control surface across those devices.
The agent can read state from one machine.
Wait for a condition.
Trigger a second device.
Inspect another sensor.
Adjust a parameter.
Continue to the next stage.
Anthropic describes MHS as supporting parallel operation across multiple instruments and says device commands can be chained together in code files.
The standard therefore sits above the individual machine.
Its unit of work can become the complete physical workflow.
That is similar to what software agents already do across several applications.
The difference is that the state now includes real-world positions, measurements and machine status.
The workflow exists in both software and physical space.
Universal Robots Tested MHS Across Multiple Cobots
Manufacturing provides a direct example.
Universal Robots joined the MHS research preview and tested the standard on its collaborative robots.
The company described an experiment involving four separate robot applications.
An agent running on Claude Opus 4.8 coordinated them as one cell and handed payloads between the robots.
The relevant part is the interface structure.

Each robot application still runs on real industrial hardware.
The existing robot-control and safety architecture remains underneath.
MHS adds a layer where the agent can discover the devices, understand declared operating information and orchestrate the applications together.
Universal Robots also describes machine bounds, interlocks and emergency-stop information being represented through the MHS layer while the robot platform’s own safety architecture remains active.
That is how a hardware standard can participate in an industrial system.
The device controller remains underneath the MHS layer.
It provides an additional interface above it.
The agent works with the standardized representation.
The robot controller continues operating the machine.
QuEra Used MHS to Connect an Agent to Precision Quantum Hardware
Quantum computing shows another type of physical system.
QuEra Computing used MHS to connect an AI agent to parts of the laser-control system inside its neutral-atom quantum-computing hardware.
The task was laser stabilization.
A neutral-atom quantum system depends on lasers maintaining highly precise operating frequencies.
QuEra used MHS as the hardware interface while Claude developed and tested control logic for restoring the laser to its target lock.
In a later blind test reported by Anthropic and QuEra, the resulting controller recovered the lock 99.3 percent of the time.
QuEra also reports that the workflow reduced recovery from minutes of specialist work to seconds in the test environment.
The point for MHS is not the quantum algorithm.
It is the hardware loop.
Read physical state.
Change controls.
Observe the result.
Evaluate whether the target condition has been reached.
Package the resulting procedure into code.
That same loop can exist in many forms of precision equipment.
MHS gives the AI system a standard place to connect to it.
Raspberry Pi Shows How MHS Can Reach Smaller Hardware
MHS is not limited to large industrial or scientific systems.
Anthropic says Raspberry Pi is enabling MHS integration across a number of its products after successful tests with a Camera MHS Driver.
A camera is a useful small-scale example because it combines hardware state with continuous data.
The device has configuration values.
It has an image stream.
It may be connected to motors, sensors or other peripherals.
An agent can use the camera as both an input and part of a control loop.
The MHS driver provides the structured hardware interface.
A Raspberry Pi can provide the programmable compute environment around it.
That makes the standard relevant to prototyping, robotics and edge systems as well as large installations.
The physical-AI stack can therefore scale down.
It does not require a factory floor.
A developer can start with a programmable board, a camera and another controllable device.
The same general driver model can then grow with the hardware system.
Hugging Face Is Connecting MHS to the LeRobot Ecosystem
Robotics libraries add another layer to the ecosystem.
Anthropic says Hugging Face is adding MHS support to LeRobot.
LeRobot is Hugging Face’s open robotics framework for models, datasets and tools used with real-world robots.
Its current documentation describes a hardware-agnostic Python interface for controlling robots, recording datasets and deploying learned policies.
MHS can add a standardized agent-to-hardware interface to that environment.
The two systems have different roles.
LeRobot focuses on robot learning, datasets, policies and hardware abstractions.
MHS focuses on exposing physical equipment to AI agents through a common device interface.
Combining them creates a path where a robot-learning stack and an agent orchestration stack can share the same physical equipment.
An agent can discover the machine.
A policy can control learned motion.
The hardware interface can expose state and permitted operations.
That is another sign that MHS is being positioned as a layer rather than a complete robotics framework.
It can connect to existing robotics software instead of replacing it.
Model-Agnostic Design Separates the Hardware Standard From One AI Model
MHS was introduced by Anthropic, but the specification is designed to be model-agnostic.
Anthropic says any agent harness can access MHS through standard protocols such as MCP.
That matters for the architecture.
A hardware installation can last much longer than one model generation.
A robotic system, camera rig or manufacturing cell may operate for years.
AI models can change much faster.
A model-agnostic hardware interface separates those lifecycles.
The driver describes the machine.
The agent harness connects to the driver.
The model used by that harness can change independently.
That keeps the machine integration focused on the physical device rather than one model API.
The same principle is common in computing standards.
The same USB connector can work across different processors.
The same Ethernet framing can work across different operating systems.
MHS is applying a similar separation to agent-controlled hardware.
The hardware contract lives in one layer.
The intelligence using that contract lives in another.
Physical Limits Can Be Declared as Part of the Interface
A physical interface needs more than capability descriptions.
It also needs operating boundaries.
MHS drivers can include information about safety limits that will be enforced.
Universal Robots describes bounds, interlocks and emergency-stop information being included in the MHS representation while the robot’s own safety architecture remains responsible underneath.
That layered design matters.
The agent sees what it is allowed to request.
The MHS layer can expose the permitted operating envelope.
The hardware controller still applies the machine’s own safety mechanisms.
The result is not one safety system replacing another.
It is a hierarchy.
The model works inside the capabilities it has been given.
The MHS driver describes and constrains the interface.
The device’s native controller remains responsible for its own hardware-level protections.
This is especially relevant when an AI system controls equipment with movement, energy or other physical effects.
The machine interface has to describe not only what can be done.
It also has to define the boundaries around how those actions can be performed.
The Research Preview Is Building the Standard Before Open Source Release
MHS is not yet a generally available open standard.
Anthropic currently describes it as a limited research preview.
Access is by application.
The company says the preview is being used with partners across science, robotics, electronics and manufacturing to build evaluations and deployment practices before the standard is released as open source.
The official MHS site uses the same framing.
The project began with Anthropic and HHMI Janelia Research Campus.
The current phase expands testing across outside organizations.
That status is important because it tells us where the technology sits.
The architecture exists.
Drivers are being tested on real devices.
Partners are building integrations.
The specification is still being developed through the research-preview process.
Anthropic says it plans to open-source MHS after that work.
For developers watching the standard, the present moment is therefore about the interface design and the early ecosystem rather than a finished universal hardware layer.
The direction is visible before the final open-source release.
MHS Creates a Hardware Layer Beside the Software Tool Layer
The larger architecture becomes clear when MHS is placed beside the other agent standards emerging around AI.
MCP gives models access to software tools and data.
App Intents, AppFunctions and operating-system action frameworks expose application capabilities.
Agent plugin formats package reusable agent extensions.
MHS addresses a different boundary.
It connects the agent stack to physical equipment.
That means future agent systems can have several layers of tools.
A database tool.
A browser tool.
A code tool.
An application action.
A camera.
A robot arm.
A controller.
A sensor.
To the agent, each one can become a structured capability.
The underlying implementation can remain very different.
Software stays software.
Hardware stays hardware.
The standard defines the interface where the agent meets each system.
That is why MHS can sit next to MCP rather than competing with it.
The software tool layer and the physical-device layer are becoming separate parts of the same agent architecture.
AI Agents Are Getting a Common Interface to the Physical World
The Model Hardware Standard turns one broad idea into an engineering structure.
A physical device gets a driver.
The driver exposes common operations such as reading and writing state.
The device can publish a reference describing its characteristics and operating limits.
Agents can discover connected hardware.
MCP, command-line interfaces and code files can provide control paths into the same system.
Several devices can be orchestrated as one workflow.
Model choice remains separate from the hardware standard.
Existing machine controllers and safety systems remain underneath the MHS layer.
Early integrations now span robotics, Raspberry Pi hardware and quantum-computing equipment.
Anthropic is testing the specification with partners before a planned open-source release.
This extends the agent stack in a specific direction.
The model can now work with software resources and physical equipment through separate structured interfaces.
Physical equipment can be described as a structured, discoverable capability too.
The agent gets a common interface.
The machine keeps its own implementation.
MHS connects the two.
That is the upgrade.
The Screen Is No Longer the Only Way Into an App
An application has traditionally presented itself through a visible interface.
Open the app.
Find the screen.
Choose the item.
Tap the action.
The data and logic behind the interface already existed, but the screen was the main way a person reached them.
AI platforms are adding another path.
An app can describe its data in structured types. It can register actions that other system experiences can discover. It can index content for semantic retrieval. It can expose functions that an authorized agent can invoke. A server can publish resources and tools through a standard protocol.
That changes the architectural role of the application.
The UI remains one interface.
The app’s structured content and capabilities become another.
Apple is doing this through App Intents and App Entities. Android is doing it through AppFunctions. Windows has App Actions and App Content Search. MCP provides a related server-side pattern for resources and tools.
The application is becoming something an AI system can understand as data and capabilities, not only as pixels on a screen.
An App Already Contains More Than Its Interface Shows
A note-taking app is not only a list of note cards.
Behind the screen are notes, titles, timestamps, folders, tags and actions such as create, edit and search.
A music app contains songs, albums, playlists and playback actions.
A travel app may contain destinations, reservations, itineraries and operations for creating or updating a plan.
A graphical interface turns that internal model into buttons, lists and screens that people can navigate.
AI integration frameworks let the application expose selected parts of the same model in a structured form.
The system can know that an object is a note rather than only seeing text inside a rectangle.
It can know that “play song” is an action with a song parameter.
It can know that a search result represents one app entity with an identifier that can be passed into another action.
That is the core shift.
The underlying application model is becoming addressable outside the UI.
Apple App Intents Turns Actions and Data Into System-Readable Types
Apple’s App Intents framework is a direct example of this architecture.
Apple says App Intents makes an app’s actions and data available outside the app in a structured form that Apple Intelligence and system experiences can discover.
An AppIntent represents an action.
An AppEntity represents an app-specific data object or concept.
An AppEnum represents a defined set of values used by the app.
A music application can expose an action for playing content and entities for songs or albums.
A messaging application can expose actions such as drafting or sending a message and entities representing contacts or conversations.
The operating system receives more than a UI label.
It receives typed information about what the app can do and what kinds of objects it contains.
That structure can then participate in Siri, Spotlight, Shortcuts, widgets and Apple Intelligence.
The app continues to have its own interface.
App Intents gives the same application model another system-facing representation.
App Entities Give AI a Structured Representation of App Data
The data side becomes clearer with AppEntity.
Apple defines AppEntity as an interface for making an app-specific type or concept discoverable by Apple Intelligence and experiences such as Siri and Shortcuts.
The developer decides which data objects should be represented.
A trail app might expose a trail entity.
A media app might expose songs, albums or playlists.
A messaging app can expose the types Apple Intelligence needs to resolve contacts and message-related references.
The entity has an identity and declared properties.
That matters because an AI system can refer to the underlying object instead of treating the visible text as the entire meaning of the item.
Once the system resolves an entity, that entity can become an input to an action.
The data object and the operation become connected.
This is the point where an app begins to act like a structured data source for AI.
The application is not publishing its entire internal database.
It is defining the app-specific objects that the system is allowed to understand and use.
App Schemas Give Different Apps a Shared Vocabulary
Apple expanded this direction in 2026 with App Schemas for Apple Intelligence and Siri AI.
A schema gives well-known actions and content types a recognizable structure.
Instead of every messaging app inventing a completely different way to describe a message, contact or send action, the app can conform its intents and entities to a system-defined schema domain.
Apple says these schemas help Apple Intelligence interpret natural language, understand app capabilities and work with content through recognized structures.
This adds a semantic layer above the individual application.
The app still owns its own data model.
The schema tells the system how selected parts of that model map into a shared concept.
That makes cross-app orchestration easier to express.
A user can describe an objective in natural language.
The system can resolve which application has the relevant entity and which action applies to it.
The app becomes part of a wider vocabulary rather than an isolated collection of screens.
Spotlight Can Turn App Content Into Semantic Context
Structured app data can also become searchable by meaning.
Apple’s current App Intents guidance says developers can index app entities so Apple Intelligence can use Spotlight’s semantic search to find app content, including cases where the user describes the content without using its exact name.
That changes what search means inside the application ecosystem.
The system does not need to begin only from a literal keyword.
It can search for an app entity through semantic relationships and then use that object as context for an action.
A piece of app data therefore gains two lives.
It remains part of the app’s own database and UI.
It can also become an indexed entity available to system intelligence through the framework the developer adopts.
This is one reason the phrase “data layer” fits the change.
AI features need context.
The app can provide selected context through structured, indexed entities rather than requiring the system to infer everything from the visual interface.
Android AppFunctions Makes App Capabilities Discoverable to Agents
Android is building a similar model through AppFunctions.
Google describes AppFunctions as an Android platform API with a Jetpack library that lets apps expose functionality to the Android intelligence system.
An app defines discrete functions representing capabilities it wants to make available.
Those functions can provide services, data and actions to a registry built into Android.
Authorized callers can discover the functions and execute them to fulfill a user request.
Google gives examples such as creating a note, editing a note, listing notes or playing media.
That moves the app from a screen-centered integration to a callable-capability integration.
The system does not have to reproduce the user’s taps one by one to know that the app can create a note.
The app can publish a typed function for that operation.
The UI remains available for direct interaction.
AppFunctions adds a second entry point designed for system agents and assistants.
Android Treats AppFunctions as an On-Device Equivalent of MCP Tools
Google explicitly connects AppFunctions to the Model Context Protocol.
Its documentation describes AppFunctions as the mobile equivalent of tools in MCP and says the API simplifies Android MCP integration.
That comparison explains the architecture.
In MCP, a server can expose tools with names, descriptions and input schemas so a language model can discover and invoke them.
In Android, an application can expose AppFunctions so authorized agents can discover and invoke capabilities on the device.
The execution location is different.
The underlying idea is similar.
The application describes what can be called.
The agent sees a structured tool rather than a sequence of UI coordinates.
This makes the app itself part of the tool layer available to device intelligence.
Google’s current documentation says AppFunctions is available from Android 16 and remains in an experimental preview during the 2026 rollout.
The architecture is already clear: app capabilities are becoming explicit machine-readable functions.
A Function Can Return Data as Well as Perform an Action
The word “function” can sound as if AppFunctions is only about automation.
It also creates a data path.
A note application can expose a function for listing notes.
A travel application can expose a function that returns itinerary information.
A media application can expose data needed to select content before another function plays it.
The result of a callable function becomes structured context for the agent.
That means an application can contribute both verbs and nouns.
Create is a verb.
The note returned by a list function is data.
Play is a verb.
The selected song or album is data.
This is the same relationship visible in Apple’s App Intents architecture: entities represent the objects and intents represent the actions.
AI orchestration becomes more useful when both sides are available.
The system can identify the object, then choose the operation that applies to it.
Windows App Actions Gives Applications a Callable Surface
Windows has its own version of this idea through App Actions.
Microsoft describes an App Action as an atomic unit of functionality that a Windows application can implement and register so it can be accessed from other apps and experiences.
The registered action includes a definition describing inputs and outputs.
Windows or another application can discover the registered action and invoke it when it fits the workflow.
Microsoft lists examples such as translating text, processing an image, file operations and other reusable tasks.
The architectural point is that the capability is registered outside the normal navigation path of the provider app.
The function becomes a system-discoverable building block.
Another application can ask Windows what actions are available for a set of inputs.
The provider app supplies the behavior.
The calling experience composes that behavior into its own workflow.
That gives Windows applications a callable surface alongside their visible UI.
Windows App Actions Can Carry Typed Entities Between Applications
App Actions also has a data model around the callable action.
Microsoft’s action definitions can describe entity types and the inputs and outputs attached to an action.
An entity can represent content such as a file or another supported data object with properties that the system can understand.
The action therefore has more than a name.
It has a contract.
The caller can know what type of input the action accepts and what type of output it returns.
That structure is relevant to AI-backed experiences because model-generated intent still has to become a concrete software operation.
A structured entity provides the bridge.
The AI or calling app resolves the user’s request into an object and an operation.
The registered App Action performs the operation on the defined type.
Once again, the app is being represented as data plus capabilities rather than only as a navigation tree.
Windows App Content Search Can Turn an App's Knowledge Into a Semantic Index
Windows is also adding a data-oriented AI layer through App Content Search.
Microsoft’s experimental AppContentIndexer API lets an application index text and images for both lexical and semantic search.
The index stores app-defined identifiers alongside vector representations used for semantic retrieval.
A query can therefore return content that is related in meaning even when the exact words differ.
Microsoft also describes the index as a local knowledge source that can support retrieval-augmented generation when paired with a language model.
That creates another version of the app-as-data-layer idea.
The application’s own content can become a knowledge base for its AI features.
The developer does not have to build the entire embedding and vector-store layer from the ground up for this API path.
The application supplies the content and identifiers.
The platform supplies semantic indexing and retrieval.
The AI feature can then use the retrieved app content as context.
MCP Extends the Same Pattern Beyond the Device
The Model Context Protocol generalizes this architecture beyond one operating system.
MCP defines standard primitives that let a server provide context and capabilities to an AI client.
Resources provide structured data or content.
Tools provide executable functions that a model can invoke.
Prompts provide reusable interaction templates.
The current MCP specification describes the protocol as a standardized way to integrate LLM applications with external data sources and tools.
That means a service can expose its application layer directly to an AI system without requiring the AI to operate the visible website as if it were a person.
A database query can become a tool.
A file can become a resource.
An API operation can become a callable function with a schema.
The same application may therefore have several interfaces at once.
A web UI for people.
A mobile UI for people.
An API for software.
An MCP server for AI clients.
AI Needs Structured Context as Another Input Surface
A visual interface is designed to communicate with a person.
Text labels, icons, cards, buttons and page hierarchy tell the user what is available.
AI systems can also work with visual interfaces, while structured integration gives them another source of meaning.
An entity can state what object is being referenced.
A function schema can state which parameters an action accepts.
A semantic index can retrieve content by meaning.
A tool result can return structured information from a backend.
That changes how context reaches the model.
Instead of reconstructing the entire application state from pixels or copied text, the AI system can receive the pieces the application deliberately exposes.
Apple App Entities, Android AppFunctions, Windows entities and MCP resources all express versions of that idea.
The application becomes a provider of structured context.
The UI and the AI Interface Can Use the Same Application Logic
This architecture does not require the application to split into two separate products.
The visible interface and the AI-facing interface can call the same underlying business logic.
A note app may already have one internal function that creates a note.
The button in the UI can call it.
An AppFunction can call it.
An App Intent can wrap it.
A server tool can expose a related backend operation.
The entry point changes while the product logic remains shared.
That gives the application several interfaces around the same capability.
The UI becomes one client of the application layer.
AI systems can become another client when the developer chooses to expose that capability.
The architectural shift is therefore not from UI to AI.
It is from one interface to several interfaces built around the same application state and operations.
Apps Are Starting to Publish Their Own AI Vocabulary
The common pattern across these platforms is vocabulary.
An application tells the system what kinds of things it contains.
It tells the system what actions it can perform.
It defines the parameters those actions accept.
It can provide descriptions that help an agent understand when a function applies.
It can index selected content so semantic search can resolve a broad request to a specific object.
That vocabulary becomes part of the AI integration surface.
A calendar app can define events and scheduling actions.
A notes app can define notes and editing actions.
A media app can define songs and playback actions.
A travel app can define trips, locations and itinerary operations.
The application is teaching the platform how to refer to its own domain.
That is more structured than exposing a collection of screens.
It is an application-level language for data and behavior.
The App Is Becoming a Screen, a Data Source and a Tool Provider
The application interface is expanding.
The screen remains the direct place where people browse, inspect and control the product.
App entities and semantic indexes can expose selected content as structured context.
App intents, AppFunctions and App Actions can expose selected capabilities as callable operations.
MCP can extend the same model to server-side data and tools.
These layers can exist together.
A person can open the app and use the UI.
A system search can surface an app entity.
An AI assistant can invoke an authorized app action.
A model can retrieve relevant app content from a semantic index.
A remote AI client can call a server-side tool.
That is why the app is becoming a data layer for AI as well as a screen.
Its data model is becoming discoverable.
Its capabilities are becoming callable.
Its content is becoming retrievable by meaning.
The application is gaining another interface: one designed for software and AI systems to understand directly.
That is the upgrade.