ThinkCentre X Ultra Brings Agentic AI Into a 1.6L Desktop
Lenovo introduced the ThinkCentre X Ultra at Innovation World during IFA 2026 on September 3, 2026.
The new system is built around a simple idea: substantial local AI capability does not need a large desktop tower. ThinkCentre X Ultra fits into a 1.6-liter chassis measuring 183 × 183 × 51mm, yet Lenovo positions it as a new class of desktop for the agentic AI era.
That combination makes the launch interesting. The system is not only compact. It is designed around high-memory local AI, developer tooling and a cluster-ready architecture that can connect several units together.
Up to AMD Ryzen AI Max+ PRO 495 Powers the System
At the top of the configuration range, ThinkCentre X Ultra uses AMD Ryzen AI Max+ PRO 495.
AMD lists the Ryzen AI Max+ PRO 495 with 16 Zen 5 CPU cores and 32 threads, while Lenovo pairs the processor with integrated Radeon 8065S graphics and an NPU rated at up to 55 TOPS.
That creates a compact platform with CPU, GPU and NPU resources available inside one system. For local AI development, those compute engines can support different parts of the workflow while keeping the machine small enough to sit almost anywhere on a desk.
Up to 128GB of Unified Memory Is the Real AI Headline
ThinkCentre X Ultra supports up to 128GB of onboard LPDDR5X unified memory.
For local AI, memory capacity is one of the most important parts of the hardware story. A larger memory pool gives the system room for bigger models, longer working contexts and more demanding agent workflows.
Lenovo also allows up to 96GB of that unified memory to be allocated as dedicated graphics memory for the integrated Radeon 8065S graphics. That gives the graphics engine a very large working pool for AI workloads while keeping the platform inside a compact integrated design.
The Memory Runs at Up to 8533MHz Across Four Channels
Lenovo specifies the onboard LPDDR5X memory at up to 8533MHz with four-channel support.
The combination of capacity and bandwidth is designed to keep large local workloads moving efficiently through the system. For AI developers, that matters because model execution involves constant movement of weights, context and intermediate data between memory and compute resources.
ThinkCentre X Ultra is therefore not simply a small office PC with extra memory. Its memory architecture is central to the local AI role Lenovo has designed for it.
Four ThinkCentre X Ultra Systems Can Become One AI Platform
The standout feature is the cluster-ready architecture.
Lenovo says up to four ThinkCentre X Ultra systems can be connected into a unified platform. The goal is to expand the compute and memory resources available to AI workloads beyond one compact desktop.
That changes the product from a single small PC into a building block. A developer can start with one system and use several systems together when the workflow grows.
The Cluster Is Designed for Larger AI Models
Lenovo explicitly connects the four-system architecture with the ability to run larger AI models.
Each ThinkCentre X Ultra brings its own compute and memory resources into the broader platform. For teams experimenting with local generative AI, that creates a path from one compact workstation toward a more substantial local compute environment.
The most interesting part is the form factor: the expansion happens by adding another 1.6L system rather than moving immediately to a large traditional server or workstation footprint.
Longer Context Windows Are Part of the Cluster Story
Lenovo also says the clustered platform can support longer context windows.
Long context is increasingly important for agentic AI. Coding agents may need to work across large repositories, research agents may process many documents, and business agents may need a substantial amount of project material available during one workflow.
ThinkCentre X Ultra is designed to give those workloads access to more local compute and memory as the deployment scales from one system to several.
Multi-Agent Workflows Are a First-Class Target
Lenovo is positioning ThinkCentre X Ultra directly for multi-agent workflows.
Instead of one assistant performing one task, agentic systems can use several specialized agents working in parallel. One can plan, another can analyze documents, another can write code and another can prepare a final result.
The cluster-ready design gives those concurrent workloads a local hardware platform that can grow with the number of agents and the amount of work being handled.
The Platform Can Handle Multiple AI Requests at Once
Lenovo also highlights simultaneous AI requests as part of the system’s scaling story.
That is useful for shared local AI environments where several applications, agents or users may need model inference at the same time. A multi-system ThinkCentre X Ultra setup can provide a broader local compute pool for those requests.
This is where the four-node concept becomes more than a spec-sheet feature. It gives local AI a way to become a shared service inside a compact business or development environment.
AMD Ryzen AI Developer Center Is Integrated
ThinkCentre X Ultra is integrated with AMD Ryzen AI Developer Center.
Lenovo says this gives users access to preconfigured AI tools, models and workflows across Windows and Linux. That software layer is important because powerful hardware becomes much more useful when developers can reach working tools and models quickly.
The integration is designed to shorten the path from opening the system to experimenting with local AI applications.
Windows and Linux Are Both Part of the Developer Story
Lenovo supports Windows 11 as well as Linux options for ThinkCentre X Ultra.
The specification list includes Windows 11 Pro and Home, Linux AMD AI OS and Ubuntu certification. Combined with AMD Ryzen AI Developer Center, that gives developers flexibility in how they build local AI projects.
A Windows-focused team can stay inside its existing environment, while Linux-oriented AI developers can work with the toolchains they already use.
Up to 8TB of High-Speed SSD Storage Fits Inside
ThinkCentre X Ultra supports up to two 4TB M.2 SSDs, creating up to 8TB of internal solid-state storage in the compact chassis.
Local AI projects can quickly accumulate models, datasets, embeddings, source repositories and generated assets. Large internal storage gives developers room to keep more of that material close to the compute platform.
It also reinforces the idea that ThinkCentre X Ultra is intended to operate as a serious local AI workstation rather than only as a thin client for cloud services.
10GbE Gives the Desktop High-Speed Wired Networking
Lenovo includes 10-gigabit Ethernet in the ThinkCentre X Ultra port selection.
The rear panel includes a 10GbE RJ-45 connection, and the optional punch-out port can also be configured with another 10GbE interface. High-speed wired networking is a natural fit for a desktop designed around local AI, large files and multi-system workflows.
It gives the small chassis connectivity that matches the scale of the compute and memory inside it.
Thunderbolt 4 and Modern Display Outputs Expand the Workspace
The rear I/O also includes two Thunderbolt 4 ports, DisplayPort 2.1 and HDMI 2.1.
That gives ThinkCentre X Ultra a broad set of options for displays, high-speed peripherals and external workflows. The front adds two USB-C ports and a headset connection, keeping frequently used ports within easy reach.
For a system that can act as both a local AI node and a daily workstation, that balance of compute and connectivity makes the compact design more versatile.
Wi-Fi 7 Adds High-Speed Wireless Connectivity
ThinkCentre X Ultra also supports Wi-Fi 7 and Bluetooth 5.4.
That gives the desktop modern wireless connectivity alongside its high-speed wired networking. For flexible office layouts, development labs and creative workspaces, the system can fit into different network arrangements without turning its small footprint into a cabling project.
The result is a compact machine that can sit quietly in a workspace while staying connected to modern peripherals and infrastructure.
Adaptive Lighting Turns System Activity Into Visual Feedback
Lenovo adds a visual touch with Adaptive Lighting.
The feature transforms system activity into real-time visual feedback, giving users a quick way to see the state of the machine at a glance. That is especially fitting for an AI workstation that may continue processing local workloads while the user is focused on something else.
The lighting becomes part of the interface between the physical machine and the background compute activity happening inside it.
The Thermal Design Is Built for Sustained Work
Lenovo designed the cooling system around sustained AI workloads while keeping the chassis compact.
The company says the thermal design supports reliable and quiet operation during extended workloads. That is important for a desktop intended to sit directly in a workspace and continue running local inference, agent tasks or development workloads over longer periods.
The engineering goal is clear: keep the local AI capability close to the user without giving up the compact 1.6L form factor.

Enterprise Features Sit Alongside the AI Hardware
ThinkCentre X Ultra also includes Lenovo ThinkShield, AMD PRO technologies and AMD DASH manageability.
That positions the system for professional environments where local AI hardware needs to fit into existing device-management practices. Lenovo also lists discrete TPM 2.0, TCG certification and FIPS 140-2 certification among the platform’s security features.
The AI workstation is therefore designed as part of a managed business fleet as well as a high-performance local development machine.
A 2kg Starting Weight Keeps the System Truly Compact
ThinkCentre X Ultra starts at 2kg while fitting into a chassis just over seven inches wide and deep.
That physical scale is part of what makes the four-system idea interesting. Several nodes can provide a substantial local AI platform without requiring the footprint normally associated with multiple full-size workstations.
For development teams or offices where desk and lab space matter, the form factor becomes part of the compute strategy.
Lenovo Plans Availability From November 2026
Lenovo says the ThinkCentre X Ultra will be available starting in November 2026.
That puts the product on a near-term path from IFA announcement to commercial availability. For developers and businesses building more local AI into their workflows, the system represents a new option that combines compact hardware, large unified memory and a multi-node scaling model.
The launch also expands Lenovo’s ThinkCentre family further into dedicated local AI infrastructure.
The Bigger Idea Is a Modular Local AI Desktop
ThinkCentre X Ultra is most interesting when viewed as a modular local AI building block.
One 1.6L machine can serve as a compact AI workstation. Several can become a larger platform for models, contexts, agents and concurrent requests. The same product therefore spans individual development and small-scale local AI infrastructure.
That is a useful direction for personal and business AI because it gives compute a physical form that can grow in small, manageable steps.
The Upgrade Feeling
Lenovo ThinkCentre X Ultra takes the idea of a mini PC much further than simple space saving.
Up to 128GB of unified memory, Ryzen AI Max+ PRO 495, Radeon 8065S graphics, AMD Ryzen AI Developer Center and a four-system cluster-ready architecture turn the tiny chassis into a serious local AI platform.
The upgrade is the ability to start small and scale physically. One box can be a powerful local AI workstation. Four boxes can become a broader platform for larger models, longer contexts and multiple agents working at the same time.
That makes the ThinkCentre X Ultra feel less like a miniature desktop and more like a new modular form of local AI infrastructure.
Project Zenith Starts With a Different Kind of Windows PC
Microsoft announced Project Zenith on September 4, 2026 as a ready-to-code Windows experience built around developer-class devices.
The idea is simple: combine high-memory hardware with a Windows 11 setup that already reflects the way developers work. Instead of treating the operating system, development tools and local AI hardware as separate layers that only meet after setup, Project Zenith brings them together from the beginning.
Microsoft says the first Project Zenith systems will arrive with AMD Ryzen AI Halo, followed by additional devices from OEM and silicon partners in the coming months.
That makes Project Zenith more than one reference machine. It is a Windows experience intended to appear across a broader class of developer-focused PCs.
64GB+ of Unified Memory Is Part of the Baseline
Microsoft defines the Project Zenith hardware class around at least 64GB of unified memory and memory bandwidth of 250GB/s or more.
Those two numbers explain why the project is closely tied to local AI development. Modern coding models and agent workflows can require large working sets, especially when they are handling source code, project context, tools and multiple steps at the same time.
A large unified memory pool gives the system more room to keep model data and application state close to the compute hardware. Project Zenith uses that hardware foundation as the starting point for the Windows developer experience rather than treating it as an optional upgrade later.
30B+ Parameter Models Can Run Locally and Unmetered
Microsoft says Project Zenith devices are designed to run models with more than 30 billion parameters locally and unmetered.
That changes the role of the developer PC. A machine can become a place where coding models run continuously as part of the local workflow, giving developers another compute layer alongside cloud services.
Local inference is especially interesting for iterative development. A coding agent can be available while the developer edits files, tests ideas and moves through a project. The computer becomes both the development environment and part of the AI execution environment.
AMD Ryzen AI Halo Is the First Hardware Platform
Project Zenith will first become available with AMD Ryzen AI Halo.
AMD introduced Ryzen AI Halo as a developer platform for local AI and agentic workloads. The platform is built around high-capacity unified memory and a software stack designed to help developers run AI models directly on the device.
The partnership gives Project Zenith a hardware platform that already targets the same core idea: a developer computer with enough local AI capacity to become an active part of the application-building workflow.
Microsoft also says more Project Zenith devices from OEM and silicon partners are planned, so the Windows experience is designed to extend beyond one hardware family.
The Windows Setup Is Ready for Development From the Start
The software side of Project Zenith is just as important as the hardware.
Microsoft says these devices ship with a preconfigured Windows environment for development and a curated set of tools covering languages, runtimes, source control and productivity.
Windows Terminal and Visual Studio Code are pinned to the taskbar by default. That detail captures the overall philosophy of the project: the first screen a developer sees should already feel like a development machine.
Project Zenith turns setup into part of the product experience instead of leaving every developer to rebuild the same baseline manually.
File Explorer Is Preconfigured for Coding Work
Project Zenith also adjusts Windows itself for development.
Microsoft says File Explorer is configured to show file extensions, hidden files, the full path in the title bar and the details pane. Long-path support is enabled as well.
These settings make project structure more visible and put technical file information closer to the surface. For developers moving between repositories, build directories, configuration files and generated assets, that creates a more direct working environment from the first boot.
Search, Start and the Taskbar Follow the Same Developer Baseline
The developer configuration continues through Search, Start and the taskbar.
Microsoft says Command Palette is enabled in Search and Start, while the broader Project Zenith setup is designed around a focused developer workspace.
The result is a Windows experience where common development entry points are already present and easy to reach. Project Zenith keeps the familiar Windows shell while tuning the default environment around coding, navigation and command-driven workflows.
WSL Is Part of the Core Development Story
Windows Subsystem for Linux has become a central part of Microsoft’s developer platform, and Project Zenith builds directly on that foundation.
WSL lets developers run Linux environments and tools alongside Windows. Microsoft open-sourced WSL in 2025 and has continued integrating it more deeply into the operating system.
For Project Zenith, that means the local AI workstation can support Windows-native development and Linux-first toolchains from the same machine. A developer can work across ecosystems while keeping the hardware and operating-system experience unified.
WSL Containers Bring Linux Containers Into Windows
Microsoft is also bringing WSL containers into the developer experience.
WSL containers provide a built-in way to create, run and interact with Linux containers directly on Windows. That gives Project Zenith another important layer for modern software development because containerized workflows are common across AI, backend services, tooling and deployment pipelines.
The combination of Windows, WSL and containers gives developers several execution environments on one workstation, all sitting on top of the same high-memory local AI hardware.
Project Zenith Is Designed for the Agent Era
Microsoft connects Project Zenith directly to agentic software development.
Coding agents increasingly work across files, tools, terminals and multi-step plans. They can keep running while a developer continues other work, and they can call local models as part of that process.
Project Zenith gives those workflows a natural home: high-memory hardware for local inference, Windows tools for development, WSL for Linux workflows and platform capabilities for building and running agents.
The workstation becomes an environment where the developer and the agent can work side by side.
Microsoft Execution Containers Add an OS-Level Agent Foundation
At Build 2026, Microsoft introduced Microsoft Execution Containers, or MXC, as a policy-driven execution layer for agents across Windows and WSL.
Project Zenith devices benefit from those Windows platform investments from day one. MXC gives developers a way to define the environment an agent can use, while Windows applies those policies at runtime.
For developer-class hardware, this creates a useful pairing: local AI compute can run directly on the machine, while the operating system provides dedicated primitives for agent execution and management.
Agent Identity and Manageability Are Built Into the Windows Direction
Microsoft is also building agent identity and enterprise manageability into Windows.
The company has described a model where agent activity can be associated with a dedicated local or cloud-backed identity, while tools such as Microsoft Entra and Intune can participate in management.
Project Zenith inherits that broader Windows platform direction. For developers building agentic applications, the machine is therefore positioned as both a local compute platform and an operating-system environment designed specifically for the way agents execute real work.
The Hardware and Software Are Being Designed as One Developer Experience
This is the most important part of Project Zenith.
The project connects device memory, memory bandwidth, local AI models, developer tools, Windows settings, WSL, containers and agent platform features into one baseline.
A developer workstation has traditionally been assembled layer by layer. Project Zenith takes a more integrated approach: the device class and the Windows configuration are planned together.
That makes the hardware specifications meaningful beyond benchmarks. The memory and compute are there to support the software experience Microsoft is building around them.

Project Zenith Builds on Windows Developer Configurations
Microsoft has already been moving toward a developer-optimized Windows baseline through Windows Developer Configurations.
At Build 2026, the company made those configurations generally available through WinGet, with a setup that can prepare tools and developer-focused Windows settings through one command.
Project Zenith takes that idea into a device-class experience. Instead of beginning with a general PC and applying a developer configuration later, the new systems are intended to arrive with the development experience already in place.
It is the same direction expressed through hardware, operating-system defaults and local AI capacity together.
Local AI Gives the Developer PC a New Role
The ability to run 30B+ parameter models locally gives the workstation a role that extends beyond editing and compiling code.
The same machine can host coding intelligence, agent sub-tasks and other model-driven tools directly on the device. That creates a richer local development loop where code, context, tools and inference can all live close to the project.
Cloud models remain part of modern development, and Project Zenith adds another powerful layer: substantial local model capacity that is available directly from the workstation.
The Experience Still Leaves Room for Personalization
Project Zenith provides a curated starting point while preserving the ability to extend and personalize the environment.
Microsoft says developers can continue choosing the tools, languages and frameworks that fit their work. The project is about beginning from a strong developer baseline rather than defining one fixed workflow.
That balance matters because software development is deeply personal. One developer may live in Visual Studio Code and WSL, another may add specialized IDEs, local model runtimes or custom terminal tools. Project Zenith gives each of them a prepared foundation to build on.
More OEM and Silicon Partners Are Coming
AMD Ryzen AI Halo is the starting point, and Microsoft says Project Zenith will expand to more devices from OEM and silicon partners in the coming months.
That gives the project room to become a broader Windows developer hardware category. Different devices can offer different physical designs and performance tiers while keeping the same ready-to-code promise.
The shared idea is consistent: developer-class hardware, a prepared Windows environment, strong local AI capability and Windows platform support for modern agent workflows.
The Upgrade Feeling
Project Zenith is interesting because it treats the developer PC as a complete system rather than a blank machine waiting to be configured.
Microsoft is pairing high-memory local AI hardware with a Windows experience that already understands coding, WSL, containers, agents and the tools developers reach for first.
The result is a new kind of starting point: turn on the machine, open the development environment and begin building with substantial local AI compute already part of the workstation.
That is the upgrade. The PC is becoming both the place where software is written and one of the places where the intelligence inside that software can run.
NVIDIA PAIR is a free, open-source virtual inference router that helps compatible computers on the same local network share independent AI inference workloads through one familiar local interface.
The Short Version
NVIDIA PAIR is a new open-source virtual inference router built to help people use several compatible computers on the same local network for local AI. Instead of sending every independent inference job to one machine, PAIR can discover participating systems, check which models and engines are available, and route work to an eligible computer. The result is a cleaner way to make more of the AI hardware already available at home or in a personal workspace.
A New Layer for Personal AI
Personal AI is quickly moving beyond one chat window and one model call at a time. Developers and power users are now running research agents, coding agents, organization tools and multiple local sessions at once. PAIR gives that growing activity a shared routing layer. Applications can continue using familiar local interfaces while PAIR handles where each independent request should run across the available machines.
PAIR Works With Familiar Local AI Engines
NVIDIA designed PAIR to work with existing local inference services rather than asking users to rebuild their software stack. The beta supports Ollama and LM Studio, two widely used tools for running models locally. PAIR sits in front of those engines and presents compatible proxy endpoints, giving applications a familiar connection while the routing layer manages placement across participating systems.
One Local Endpoint, Multiple Compute Options
From the application’s point of view, the workflow stays simple. A request arrives through the local endpoint, PAIR identifies the engine and model it needs, and then selects an eligible node. The application keeps seeing one connection while the routing happens behind the scenes. This is an elegant approach because it adds flexibility without forcing every agent or desktop tool to learn a completely new cluster interface.
Independent AI Jobs Can Run Across Different Machines
PAIR is especially useful when a workload creates several independent inference requests. A lead agent can assign research, coding, verification and summarization jobs to different subagents, and PAIR can place those requests on different available systems. That opens the door to more parallel local AI activity using hardware that might otherwise be sitting unused.
The Same Model Can Be Available on Several Nodes
Users can prepare the same model on multiple participating machines. When several nodes have that model available, PAIR has more eligible places to route incoming requests. This creates a simple way to expand service capacity for the models a user runs most often, especially in workflows where many agents may call the same model during one larger task.
Different Machines Can Host Different Models
PAIR also supports a more specialized setup. One computer can host one set of models while another holds a different set. The router checks model availability as part of its placement decision, allowing a personal AI network to become more organized. A workstation can be prepared for one class of workload while another system is ready for a different model or task.
PAIR Discovers Systems on the Local Network
The software uses local-network discovery through mDNS to find nearby compatible systems. Users can also add a node manually by IP address. Once the desired computers are paired, PAIR can treat them as part of the same trusted local group and keep track of which systems are currently ready to contribute AI capacity.
Pairing Is Designed to Be Simple
Connecting machines starts with a six-digit pairing PIN. After the pairing step, PAIR establishes certificate-based trust between cluster members. NVIDIA combines this straightforward setup flow with mutual TLS for most peer communication, giving the local cluster a secure foundation without turning setup into a complex infrastructure project.
The Scheduler Watches the State of Each Node
PAIR continuously tracks useful routing signals across participating systems. It checks whether a node is online and ready, whether the required inference engine is enabled, whether the requested model is present, the amount of queued work, and GPU utilization. These signals help the router choose an available destination for each new independent request.
Home Hardware Can Join and Leave Dynamically
One of PAIR’s most practical ideas is elastic participation. A compatible laptop, gaming PC, workstation or DGX Spark can contribute capacity while it is available and then simply leave the active pool when it is powered down or moved elsewhere. That makes PAIR a natural fit for real personal hardware, where devices are used for many different things throughout the day.
Support Starts With GeForce RTX 20 Series and Newer
NVIDIA says the PAIR beta supports systems with GeForce RTX 20 Series GPUs and newer. It also supports NVIDIA RTX PRO workstation GPUs based on Turing or newer architectures and NVIDIA DGX Spark. That gives the beta access to a broad range of existing RTX hardware rather than focusing only on the newest desktop systems.
Apple M4 and Newer Systems Are Included
The supported-hardware list also includes Apple M4 or newer silicon. That cross-platform support makes the concept especially interesting for users who already have a mixed collection of computers. PAIR can provide one routing layer across compatible systems even when those systems are not all built around the same desktop platform.
Windows, Linux and macOS Are Supported
The beta is available for supported Windows, Linux and macOS systems. NVIDIA provides both graphical and terminal interfaces, so PAIR can fit desktop workflows as well as more technical setups. The terminal option also makes it practical to include machines that are used primarily as compute nodes.
NVIDIA Demonstrated a Major Multi-Agent Speedup
NVIDIA demonstrated PAIR with a five-subagent workload using Hermes Desktop and Ollama. In the company’s test, the workload completed in 18 minutes on a single RTX Spark laptop. With a three-device PAIR cluster, it completed in 8 minutes and 48 seconds. The demo gives a concrete example of how parallel local AI requests can benefit when more compatible systems are available to serve them.
Multi-Agent Workflows Are a Natural Match
Agent systems naturally create the kind of workload PAIR is designed to organize. A lead agent can delegate separate tasks to specialized subagents, and those subagents can produce many model calls during one larger job. PAIR gives those independent calls more places to run, turning a collection of local computers into a more coordinated environment for agentic AI.
Research Agents Can Spread Work Across the Network
A research workflow can divide a topic into several branches, ask different subagents to collect evidence, and then bring the results together. With PAIR, those independent inference requests can be routed across multiple available systems. This is a strong example of how personal AI can move from a single-machine workflow toward a more flexible local compute network.
Coding Agents Can Benefit From More Available Capacity
Coding assistants increasingly combine planning, code generation, testing, review and documentation. When those activities are handled by several subagents, the number of local model calls can grow quickly. PAIR gives developers a way to bring additional computers into that workflow while keeping the application connected through a familiar local interface.
The Main PC Can Stay Focused on the User
PAIR can also help users make better use of a second PC or workstation while keeping the primary computer focused on interactive work. New inference jobs can be routed toward another eligible system with available capacity. For people who already own several capable machines, that makes local AI feel less tied to whichever computer happens to be in front of them.
Local-First Architecture Keeps the Experience Close to Home
NVIDIA designed PAIR around local-network operation. Participating systems discover one another on the LAN, and the routing layer is built to keep prompts, data and inference traffic within the user’s local environment when the local application and inference stack are configured that way. This fits neatly with the appeal of local AI: more direct control over where personal compute runs.
Mutual TLS Protects Most Peer Communication
After systems are paired, PAIR uses certificate trust and mutual TLS for most communication between cluster members. That gives the local compute group authenticated connections between participating nodes while keeping the overall setup approachable for personal use. Security is integrated into the pairing and routing design rather than being left as a separate manual project.
Getting Started Follows a Familiar Local-AI Flow
The setup process is straightforward: install PAIR on the participating computers, discover or add the systems, pair them, enable a supported inference engine, and prepare the models needed by the workload. Compatible applications can then connect through PAIR’s local endpoint. The structure feels close to a normal local-AI setup, with the routing layer adding access to more machines.
PAIR Can Help Install Engines and Prepare Models
NVIDIA’s getting-started documentation says PAIR can help install and start supported inference engines and initiate model downloads on participating nodes. That makes it more than a passive traffic layer. It can also help users prepare the machines that will provide local inference capacity, reducing some of the repetitive setup work across a multi-computer environment.
Open Source Gives Developers a Clear View of the Project
NVIDIA released Personal AI Router as an open-source project under the Apache License 2.0. Developers can inspect the code, study the architecture, report issues and contribute improvements. For a tool that coordinates AI work across several personal machines, that openness is valuable because the routing logic and project direction are visible to the community.
PAIR Creates a Home Inference Fabric
The clearest way to think about PAIR is as a home inference fabric. One local entry point can coordinate independent AI jobs across several available systems. Applications keep using familiar interfaces while the router handles placement. This creates a clean bridge between today’s local model tools and a future where personal AI regularly uses more than one computer.
Personal AI Is Expanding From One Session to Many
PAIR arrives at a useful moment. Local AI is expanding from one user talking to one model toward multiple agents and background sessions working at the same time. As that pattern grows, the ability to coordinate several computers becomes increasingly useful. PAIR gives NVIDIA users an early look at what a more distributed personal AI environment can feel like.
Who Will Get the Most From PAIR
PAIR is especially appealing for AI enthusiasts, developers, creators and power users who already own more than one capable computer. It also fits people experimenting with local research agents, coding agents, personal automation and multi-agent workflows. The more independent AI jobs a workflow creates, the more useful an organized pool of local compute can become.
Why This Launch Matters
PAIR makes spare local AI capacity easier to use. It connects familiar inference engines, familiar application interfaces and existing personal hardware through one open-source routing layer. That is a meaningful step because it makes multi-computer local AI feel more like a normal desktop capability and less like a specialized infrastructure project.
The Upgrade Feeling
NVIDIA PAIR has a simple but powerful idea behind it: the computers already around you can work together more intelligently for local AI. A gaming PC, workstation, laptop or DGX Spark can become part of the same personal inference network, with PAIR deciding where independent jobs should run. For multi-agent workflows, that turns existing hardware into a more flexible and coordinated AI environment — exactly the kind of upgrade that can change how personal compute feels in everyday use.
NVIDIA PAIR turns compatible computers on the same local network into a shared inference pool for AI apps and agents. It does not merge GPU memory or split one model across machines. Instead, PAIR discovers eligible Windows, Linux and macOS systems, tracks whether Ollama or LM Studio is ready, checks whether the requested model is present and routes each independent inference request to one available node. That architecture is especially useful for multi-agent workflows where several subagents make model calls at the same time. In NVIDIA’s configuration-specific Hermes demo, a three-device PAIR cluster completed a five-subagent workload in 8 minutes 48 seconds versus 18 minutes on one RTX Spark laptop. The more important shift is that local AI is starting to become distributed software: the gaming PC, workstation and laptop already in a home can act as separate workers behind one local endpoint.
Your Second PC Just Became an AI Worker
Local AI usually begins with one machine.
One GPU.
One inference server.
One model queue.
That works until the agent stops behaving like a chatbot.
A multi-agent system can create several model calls at once.
One subagent researches.
Another checks documents.
Another verifies an answer.
Another writes code.
Another summarizes the result.
If every request targets the same GPU, the jobs queue behind one another while another capable PC in the house may be doing nothing.
NVIDIA PAIR is designed around that mismatch.
It turns several compatible computers on one local network into a shared pool for independent AI inference jobs.
PAIR Was Announced at IFA 2026
NVIDIA announced Personal AI Router, or PAIR, on September 3, 2026 as part of its IFA push around local agents.
The software is available in beta for supported Windows, Linux and macOS systems.
NVIDIA describes it as a free, open-source virtual inference router.
PAIR works with Ollama and LM Studio at launch.
The goal is to let existing AI applications keep talking to a familiar local endpoint while PAIR decides which participating machine should actually run each request.
That is a software-routing problem more than a new model problem.
PAIR Is Not a New Inference Engine
This distinction is important.
PAIR does not replace Ollama.
It does not replace LM Studio.
It does not execute the model itself.
A supported inference engine still loads and runs the model on the selected computer.
PAIR sits in front of those engines.
It discovers machines.
Tracks their readiness.
Checks model availability.
Routes requests.
Then returns the response to the application that made the call.
The agent sees one local service.
PAIR handles placement behind it.
One Endpoint Hides Several Machines
The abstraction is simple.
An AI application connects to a local endpoint.
PAIR presents Ollama-compatible and OpenAI-compatible proxy interfaces.
The application sends a request as if it were talking to one local engine.
PAIR reads the engine and model requirements, chooses one eligible node, forwards the request and streams the answer back.
The application does not need to discover every machine itself.
That is the part that makes the cluster usable.
The complexity moves from the agent into the router.
The Devices Stay Separate
NVIDIA uses the phrase personal AI cluster.
That can create the wrong mental picture.
PAIR does not fuse several PCs into one giant computer.
Each device remains an independent machine.
Each GPU keeps its own memory.
Each inference engine keeps its own model files.
PAIR simply sends different independent requests to different systems.
That distinction defines what PAIR can accelerate and what it cannot.
PAIR Does Not Pool VRAM
Two 24 GB GPUs do not become one 48 GB GPU through PAIR.
NVIDIA explicitly says PAIR does not combine GPUs into a larger logical accelerator.
It does not pool VRAM.
If a model requires more memory than one machine can provide, PAIR cannot make that model fit by borrowing memory from another node.
The full model still has to fit on the individual computer selected to run the request.
That makes PAIR a routing layer, not distributed tensor-parallel inference.
PAIR Does Not Shard One Model Across Machines
The same rule applies to model execution.
PAIR does not split one inference request across several computers.
One request goes to one eligible node and stays there for its lifetime.
Another independent request can go to a different node.
That means the gain comes from concurrency.
Many calls at once.
Not one giant call spread across many GPUs.
This is why multi-agent workflows are the natural target.
Agents Create Exactly the Kind of Work PAIR Can Parallelize
A single chatbot conversation is often sequential.
Prompt.
Answer.
Next prompt.
An agentic workflow can be much wider.
A lead agent decomposes one task into several independent jobs.
Those jobs can run at the same time.
Each one may trigger its own inference request.
That is where a single local GPU becomes a queue.
PAIR gives the inference layer the same parallel structure as the agent workflow.
Several subagents can make progress simultaneously on different machines.
NVIDIA’s Demo Cut One Five-Agent Workload From 18 Minutes to 8:48
NVIDIA demonstrated PAIR with Hermes Desktop and Ollama.
Hermes created five specialist subagents for a synthetic household-inbox task.
On one RTX Spark laptop using Qwen 3.6 35B A3B, NVIDIA says the workload took 18 minutes on average.
A three-device PAIR cluster containing an RTX Spark laptop, a DGX Spark and an RTX 5090 desktop completed the same workload in 8 minutes 48 seconds on average.
That is a large difference.
It is also not a universal benchmark.
The 8:48 Result Has a Big Asterisk
NVIDIA explicitly labels the demonstration unofficial and configuration-specific.
The result depends on how parallel the workload is.
Which model is running.
The inference-engine configuration.
The hardware mix.
Network conditions.
Whether nodes are available.
A different task can scale differently.
A sequential workflow may gain almost nothing.
So the correct conclusion is not “PAIR makes AI twice as fast.”
The demo shows that routing independent calls across several ready devices can reduce queueing substantially when the workload exposes enough parallel work.
PAIR Watches Which Machines Are Actually Available
A home cluster is not a datacenter.
A laptop closes.
A gaming PC becomes busy.
A workstation goes to sleep.
A machine may have the right inference engine but not the requested model.
PAIR is designed around that instability.
It maintains a live view of the participating systems and decides whether each one can accept a new request.
The available pool can change while the cluster is running.
The Scheduler Checks More Than Whether a PC Is Online
NVIDIA says PAIR currently considers several factors for each request.
Is the paired node online and ready?
Is the required inference engine enabled?
Is the exact requested model present?
How many jobs are already active?
Is the GPU busy with another graphics-intensive workload?
Those signals let the router avoid sending work blindly.
A connected machine is not automatically an eligible machine.
Your Gaming PC Can Leave the Pool When You Need It
This is one of the more practical design choices.
A gaming PC may be an excellent AI worker while nobody is using it.
Then a game launches.
GPU utilization rises.
The machine should stop behaving like spare inference capacity.
PAIR can route new requests elsewhere when another system is more appropriate.
The same logic applies when a laptop sleeps or leaves the network.
Local AI capacity becomes elastic instead of permanently reserved.
The Requested Model Still Has to Exist Somewhere
PAIR cannot route a model request to a machine that cannot serve it.
The inference engine has to be running.
The model has to be available on that node.
The hardware has to have enough memory to load it.
That creates an important operational detail.
If only one computer has a particular model, every request for that model will still converge on that one machine.
The cluster only becomes useful for that workload when several eligible nodes can serve the requests.
Every Node Does Not Need the Same Model Library
The opposite is also true.
NVIDIA says machines in a PAIR cluster do not have to hold identical model collections.
One system can host one model.
Another can host a different model.
PAIR can route requests according to where the requested model exists.
Replicating the same model across several nodes increases the number of machines that can serve that request.
Different model placement can turn the home cluster into a small heterogeneous inference pool.
That Makes Storage Part of Local AI Scaling
Distributed inference sounds like a GPU story.
It is also a storage story.
If the same large model is copied to three machines so all three can answer requests, the model consumes storage three times.
Local agents may use several models.
Embedding models.
Language models.
Vision models.
Rerankers.
The more redundancy a user wants across the cluster, the more local disk capacity is consumed.
PAIR does not eliminate that trade-off.
It makes the trade-off manageable through routing.
PAIR Supports RTX 20 Series and Newer
NVIDIA says the beta supports GeForce RTX 20 Series GPUs and newer.
RTX PRO workstation GPUs using Turing architecture and newer are also included.
DGX Spark is supported.
That gives PAIR access to hardware accumulated across several generations rather than requiring only the latest flagship card.
A household may already own much of the compute before installing the software.
That is central to the pitch:
use the machines that are already there.
Apple M4 and Newer Can Join Too
The most interesting compatibility decision may be Apple silicon.
NVIDIA says PAIR supports Macs with M4 or newer silicon.
That means the personal cluster is not limited to NVIDIA-only client hardware.
A Windows RTX desktop and a supported Mac can participate in the same routing system.
The inference engine and model still have to work on the node.
But the router itself is designed around mixed operating systems and mixed hardware.
That makes PAIR closer to a software-defined home inference layer than an RTX-only cluster manager.
Windows, Linux and macOS Nodes Can Mix
PAIR supports Windows 11, Linux and macOS.
The open-source repository says x64 and arm64 are supported across the three operating-system families, with Windows on Arm marked experimental.
Nodes using different operating systems can be paired together.
That flexibility matters because personal hardware is messy.
One person may have a gaming desktop on Windows, a Linux workstation and a MacBook.
PAIR is designed to treat that mixture as potential inference capacity rather than forcing the user to standardize the entire household.
Ollama and LM Studio Are the First Backends
At launch, PAIR supports Ollama and LM Studio.
That choice lowers integration friction.
Both are already widely used for local model inference.
Applications that already know how to talk to those local interfaces do not need a custom distributed-computing API.
PAIR proxies the familiar interface instead.
NVIDIA says agent harness changes are not required for compatible workflows.
That is strategically important.
A new router is much easier to adopt when users do not also have to replace the tools around it.
The Router Can Discover Machines Automatically
PAIR uses local-network discovery through mDNS to find nearby participating systems.
A node can also be added by IP address.
The user then approves a pairing request.
The goal is to make adding compute closer to connecting a local device than configuring a traditional compute cluster.
That matters because the target environment is not an enterprise datacenter run by cluster administrators.
It is a home, studio or small workstation network.
The Pairing Layer Uses mTLS
NVIDIA says node-to-node communication is blocked until pairing is established.
After pairing, communications use mutual TLS with generated certificates.
That provides encrypted traffic and authentication between cluster members.
The security model still depends on the local environment.
NVIDIA’s repository explicitly warns users to read the security documentation before deploying PAIR on an untrusted or shared network.
A local network is not automatically a trusted network.
Local Does Not Automatically Mean Every Byte Stays Private
NVIDIA describes PAIR as designed for private local inference.
Prompts, files and agent context can remain on the home network rather than going to a cloud inference service.
The open-source repository adds an important qualifier.
That statement holds when the configured client, model source, inference engine and participating nodes are all local.
An agent can still choose to call a cloud service.
An application can still transmit data elsewhere.
PAIR keeps its routing local.
It cannot guarantee the behavior of every other component in the workflow.
No Internet Is Required for Operation
NVIDIA’s PAIR product page says internet connectivity is not required for operation.
Internet access is required for downloading models.
Once the needed software and model files are present, the routing layer can operate across the local network.
That matters for privacy.
It also matters for resilience.
A local agent workflow can continue using local compute even when external connectivity is unavailable, as long as the application itself does not depend on cloud services.
The Best Workload Is Wide, Not Long
PAIR helps when several independent inference calls exist at the same time.
It helps much less when one long model call dominates the task.
If step two cannot start until step one finishes, another idle GPU has nothing useful to do.
This is basic parallel computing.
The workload has to expose parallelism before a router can exploit it.
Multi-agent systems happen to create that parallelism naturally.
That is why PAIR arrived at the same moment local agents are becoming more ambitious.
The Scheduler Is Still Simple
The current PAIR beta should not be mistaken for a mature datacenter scheduler.
NVIDIA’s repository says the shipped scheduling policy combines queued work with a coarse, smoothed GPU-utilization signal.
It does not yet consider every useful variable.
GPU model.
Available memory.
Model warmness.
Estimated request cost.
NVIDIA says those are areas it may improve.
That means today’s version can route intelligently enough to be useful while still leaving substantial room for better scheduling.
Mixed Hardware Is Supported — but Similar Machines May Be Easier
PAIR can connect heterogeneous systems.
That does not mean every mixed cluster will balance perfectly.
A small laptop GPU and a large workstation GPU can have very different inference speeds.
A Mac and an RTX desktop may use different runtime paths.
The current scheduler does not fully model every performance difference.
NVIDIA’s own repository says PAIR can be a better fit for similar machines than a highly mixed cluster under the current policy.
The hardware support is broad.
The scheduling intelligence is still evolving.
This Is More Like a Load Balancer Than a Supercomputer
The simplest analogy is a load balancer.
Requests arrive.
The router looks at available workers.
One worker receives each request.
Other requests can go to other workers.
The system becomes more capable under concurrent demand because work is spread across several machines.
PAIR is not trying to recreate an HPC fabric inside the house.
It is making local inference routing simple enough that several personal computers can behave like a small service pool.
The Agent Does Not Need to Know Which Computer Answered
This abstraction is what makes distributed local AI practical.
The agent asks for a model response.
PAIR decides where it runs.
The response returns through the same interface.
NVIDIA exposes Jobs and metrics views so the user can inspect which node actually handled each request.
But the agent itself does not have to carry cluster topology.
That separation allows the agent developer to focus on task decomposition while the router handles resource placement.
The Source Code Is Public Under Apache 2.0
NVIDIA has published PAIR on GitHub.
The repository is licensed under Apache License 2.0.
Developers can inspect the implementation, report issues and contribute to discovery, pairing, routing, engine integration and the user experience.
That matters for a tool sitting between private local data and inference engines.
The routing layer is not an opaque cloud service.
Users and developers can examine how the system is built.
Third-party inference engines and models can still carry their own separate licenses and terms.
PAIR Turns Old Hardware Into Capacity Instead of E-Waste — Sometimes
A previous-generation RTX machine may no longer be the user’s main PC.
PAIR gives compatible hardware another possible role.
Serve local inference when idle.
That does not mean keeping every old computer powered on is automatically efficient.
Electricity use still matters.
Older hardware may perform poorly per watt.
The practical value depends on how often the extra capacity is needed.
But the architecture creates an option that did not exist in the normal one-PC local-AI model:
reuse existing machines as temporary workers.
The Cloud Is Still Better for Some Jobs
PAIR is not an argument that every AI workload should move home.
Cloud systems offer larger accelerators.
Large memory pools.
High-bandwidth interconnects.
Managed availability.
Frontier models that may not fit on local machines.
The local cluster solves a different problem.
Private data.
Existing hardware.
No per-token local inference charge.
Low network dependency.
Parallel agent workloads that fit on the available nodes.
The likely future is hybrid.
Local when local makes sense.
Cloud when scale wins.
PAIR Could Change How People Buy Their Next PC
The most interesting long-term effect may be behavioral.
Today, a laptop and desktop are usually treated as separate computers.
PAIR gives them a second identity.
Members of one personal compute pool.
That changes the value of idle hardware.
A workstation in another room is no longer disconnected from the agent running on the main PC.
A supported Mac can contribute.
A gaming desktop can contribute when not gaming.
The home network becomes part of the AI architecture.
This Is the Software Layer RTX Spark Was Missing
RTX Spark gives NVIDIA a powerful local-AI hardware platform with large unified memory and strong inference capability.
PAIR solves a different layer.
What happens when a user owns more than one capable device?
Instead of treating each computer as an isolated AI island, PAIR lets the machines contribute to one routing pool.
That makes the local-AI story larger than one expensive laptop or desktop.
The unit of compute starts to become the network.
What NVIDIA Has Actually Confirmed
NVIDIA announced PAIR on September 3, 2026.
PAIR is a free, open-source virtual inference router available in beta.
It supports compatible Windows, Linux and macOS systems.
At launch it works with Ollama and LM Studio.
Supported hardware includes GeForce RTX 20 Series and newer, RTX PRO workstation GPUs, DGX Spark and Apple M4 or newer silicon.
PAIR discovers and pairs systems on the local network, tracks node readiness and routes each independent inference request to one eligible machine.
It does not pool GPU memory or shard one request across several machines.
NVIDIA’s Hermes demonstration showed 18 minutes on one RTX Spark laptop versus 8 minutes 48 seconds on a three-device PAIR cluster for that specific five-subagent configuration.
The GitHub project is licensed under Apache 2.0.
What We Should Not Claim
We should not say two GPUs become one larger GPU.
PAIR does not pool VRAM.
We should not say one large model can be split across several PCs.
PAIR does not shard a single inference request.
We should not say the demo proves a universal two-times speedup.
NVIDIA explicitly says it is configuration-specific and unofficial.
We should not say every connected machine can serve every request.
The engine, exact model and sufficient memory have to be available.
We should not say local routing guarantees total privacy if the agent or another component still calls cloud services.
And we should not describe the current scheduler as a datacenter-grade optimizer.
NVIDIA says it still uses a relatively simple policy.
The Bigger Shift Is That Personal AI Is Becoming Distributed
The first local-AI wave asked whether one PC could run a useful model.
The next question is what happens when several machines can.
PAIR’s answer is not to weld the hardware together.
It is to make the software smart enough to route around the house.
One laptop handles one subagent.
A desktop handles another.
A workstation picks up a third request.
The agent sees one local endpoint.
The user sees several machines finally doing something useful at the same time.
That is a subtle shift.
Local AI stops being a property of one PC.
It becomes a property of the network.
ASUS is showcasing new ProArt P16 and P14 laptops at IFA 2026 powered by NVIDIA RTX Spark. The platform combines a Blackwell RTX GPU, Grace CPU, up to 128 GB of unified memory and up to 1 petaflop of FP4 AI performance. ASUS and NVIDIA say supported local workflows can run LLMs up to 120 billion parameters, work with up to 1 million tokens of context, generate 4K AI video and handle 90 GB+ 3D scenes. The ProArt P16 packages that architecture into a chassis just 12.9 mm thick.
A 120B Model Sounds Like Server Hardware — ASUS Is Putting That Class of Workload in a 12.9 mm Laptop
A 120-billion-parameter AI model sounds like something that belongs in a server rack.
ASUS is putting that class of local workload into a laptop only 12.9 mm thick.
The new ProArt P16 uses NVIDIA’s RTX Spark platform with up to 128 GB of unified memory and up to 1 petaflop of FP4 AI performance. ASUS and NVIDIA say the platform can run large language models with as many as 120 billion parameters locally, including supported agent workflows with context windows up to 1 million tokens.
The important part is not simply that the laptop has a fast GPU.
It is that the CPU and GPU are designed around one unusually large memory pool.
Large local models need compute, but before the accelerator can process a model, the weights and working state need somewhere to live.
That memory architecture is what makes the 120B figure interesting in a laptop this thin.
The 120B Number Needs One Important Phrase: “Up To”
The ProArt P16 does not ship with one specific 120-billion-parameter model installed by default.
The claim is about the class of workload the RTX Spark platform is designed to support.
ASUS says RTX Spark enables creators and developers to run LLMs with up to 120 billion parameters locally.
That means the chain is:
hardware platform → available compute and memory → compatible model and software → local inference.
The exact experience can vary substantially from one model to another.
Parameter count does not tell you model architecture, quantization, context length, runtime overhead or inference speed.
So “120B” is best treated as a supported upper workload class under compatible conditions, not as a promise that every 120B model will behave identically on the machine.
Why 128 GB of Unified Memory Changes the Equation
On a conventional PC, system memory and GPU memory are usually separate pools.
The CPU works primarily from system RAM.
A discrete GPU works primarily from its own VRAM.
That separation matters when a model becomes very large.
If the model or its active working set does not fit comfortably in the GPU’s available memory, the workflow can become more complicated and data movement becomes part of the problem.
RTX Spark changes that arrangement.
ASUS says the platform provides up to 128 GB of high-bandwidth unified memory shared across the integrated system.
The CPU and GPU can therefore work from a much larger common memory pool instead of treating a smaller discrete VRAM allocation as the only fast memory available to the accelerator.
For large local AI, that is one of the most important changes in the machine.
A 120B Model Is a Memory Problem Before It Becomes a Speed Problem
A useful way to understand the scale is to look at the model weights themselves.
A 120-billion-parameter model at FP16 would require roughly 240 GB for the weights alone if every parameter used two bytes.
At FP8, the same simple calculation is roughly 120 GB.
At FP4, four bits per parameter works out to roughly 60 GB for the raw weights.
Those are illustrative calculations, not the exact memory footprint of every model.
Real inference also needs room for runtime state, caches, context, temporary buffers and software overhead.
But the comparison explains why FP4 support and a 128 GB unified pool can matter together.
The accelerator does not only need to be fast enough.
The system needs enough usable memory to hold the model and the rest of the inference workload at the same time.
RTX Spark Combines a Blackwell GPU and Grace CPU
RTX Spark is a highly integrated NVIDIA platform rather than a conventional laptop pairing of an unrelated CPU and discrete graphics chip.
ASUS says the superchip combines an NVIDIA Blackwell RTX GPU with 6,144 CUDA cores, fifth-generation Tensor Cores with FP4 precision and an NVIDIA Grace CPU with up to 20 cores.
The two sides are connected through NVIDIA NVLink-C2C.
That matters because the architecture is being designed around shared AI and graphics workloads from the start.
The CPU handles general-purpose work and system orchestration.
The Blackwell GPU provides the parallel compute and Tensor Core acceleration used by many AI workloads.
The shared memory architecture connects those pieces into one system instead of forcing every large workload through the assumptions of a traditional discrete-GPU PC.
One Petaflop Does Not Mean Every AI Task Runs at One Petaflop
ASUS and NVIDIA advertise up to 1 petaflop of AI performance for RTX Spark.
That number needs context.
It refers to the platform’s peak FP4 AI compute capability under suitable workloads.
It does not mean every application receives one petaflop of real-world performance.
Actual throughput depends on the model, precision, quantization format, software stack, context length, batch size, memory behavior and how well the workload maps to the hardware.
The useful point is that RTX Spark combines very low-precision AI compute with a large unified memory pool.
Compute and memory solve different parts of the problem.
The petaflop figure describes how much mathematical work the accelerator can theoretically process.
The 128 GB figure describes how much working data the system can keep available.
Large local models need both.
The Bigger Shift Is Running the Agent on Your Own PC
ASUS is positioning RTX Spark around personal agents as much as traditional generative AI.
The workflow can move from:
prompt → remote service → cloud model → response
toward:
prompt → local model → local files and tools → local action.
That does not mean every agent should run locally.
It means the user has enough on-device compute to move more advanced workloads onto the PC when the model and software support it.
ASUS explicitly says RTX Spark is designed to let creators and developers explore advanced AI applications directly on their PCs without relying exclusively on cloud processing.
That phrase matters.
The local machine becomes a serious inference target rather than only a thin client for a remote model.
“Local” Does Not Mean the Cloud Disappears
The local-versus-cloud story should not be framed as an all-or-nothing choice.
ASUS describes a hybrid workflow.
Local RTX-powered AI can handle workloads on the PC.
Cloud services can still be used for burst capacity, frontier-scale models or services that are only offered remotely.
That gives the user another execution option.
A model that fits the machine and matches the task can run locally.
A larger or specialized service can still run in the cloud when that makes more sense.
The practical change is therefore not “no cloud required for everything.”
It is that the cloud no longer has to be the only place where a demanding AI workflow can happen.
Token-Free Local AI Does Not Mean AI Has No Cost
ASUS uses the phrase token-free local AI when describing workloads that run on the user’s own hardware.
The practical meaning is straightforward.
If the model is running locally, there is no remote API provider billing that local inference by generated or processed token.
That is different from saying the AI is free.
The hardware has a purchase cost.
The computer uses electricity.
Some software or models may have their own licensing terms.
Maintenance and storage still exist.
So the useful distinction is:
local inference → no per-token cloud inference charge for that local workload.
That can make repeated experimentation, local agents and iterative creative workflows easier to budget because usage is tied to owned compute capacity rather than a metered remote API.
RTX Spark Supports Agent Workflows With Up to One Million Tokens of Context
NVIDIA says RTX Spark can run 120-billion-parameter LLMs with up to 1 million tokens of context using local agents.
Context length matters because an agent often needs more than a short conversation.
A long-context workflow can include documents, source code, research material, project history, tool outputs and intermediate state.
The larger the active context becomes, the more memory the runtime may need.
That links the million-token claim back to the same architectural theme.
The 128 GB unified pool is not only useful for model weights.
It can also provide room for the working state around the model.
The “up to” still matters: the specific model and software stack must support the context length being used.
Local Agents Need More Than the Model Weights
Loading a model is only the beginning of an agent workflow.
The system may also need:
model weights,
context,
KV cache,
tool state,
application memory,
temporary inference buffers,
and sometimes multiple models or encoders.
That is why raw parameter count is not enough to predict whether a local workflow will fit comfortably.
A 120B model can consume a large amount of memory before the first response is generated.
Then the context and runtime add more.
Unified memory gives the system a larger shared workspace for that full chain.
The important question becomes less “how much VRAM does the GPU have?” and more “how much usable memory can the AI workload access across the system?”
RTX Spark Is Also Targeting Local 4K AI Video Generation
The platform is not designed only for text models.
ASUS and NVIDIA also highlight 4K AI video generation as an RTX Spark workload.
That broadens the meaning of local AI on the ProArt machines.
A language model primarily works with tokens and model state.
Generative video has to create and transform large sequences of image data.
3D workloads add geometry, textures, materials and rendering state.
Different applications stress the machine in different ways, but the same combination of large shared memory and GPU acceleration can support all of them.
That is why ASUS is positioning ProArt P16 as a creator system rather than a laptop built around one chatbot.
The Same Platform Can Work With 90 GB+ 3D Scenes
NVIDIA also says RTX Spark can render ultra-large 3D scenes larger than 90 GB.
Again, the number points back to memory capacity.
A large scene can contain geometry, textures, simulation data and rendering resources that would exceed the VRAM capacity of many conventional laptop GPUs.
A large unified memory pool changes the ceiling.
The graphics and AI accelerator can work with a much larger data set without treating a small discrete VRAM allocation as the only high-performance workspace.
That does not guarantee identical performance for every 90 GB scene.
Scene structure, application behavior and rendering settings still matter.
But it shows why memory architecture is central to RTX Spark’s pitch.
All of This Is Going Into a 12.9 mm Chassis
The ProArt P16 is where the architecture becomes visually surprising.
ASUS says the new model is 12.9 mm thick and weighs 1.77 kg.
It also carries a 99 Wh battery.
That puts a platform designed for large local AI, 4K AI video and heavy creator workflows into a chassis closer to an ultrathin creator notebook than a conventional mobile workstation.
ASUS describes it as the thinnest 16-inch RTX Spark laptop and the thinnest 16-inch ProArt it has produced.
The 120B figure attracts attention.
The physical packaging is what makes the story unusual.
Large local AI is moving into hardware that is designed to travel.
The Smaller ProArt P14 Is 13.9 mm Thick
ASUS is also putting RTX Spark into the ProArt P14.
The company lists the P14 at 13.9 mm thick and 1.48 kg, with a 90 Wh battery.
ASUS calls it the lightest ProArt laptop it has produced.
That matters because RTX Spark is not being limited to one 16-inch flagship chassis.
The platform is being packaged into a smaller portable form factor as well.
Both laptops support up to 128 GB of unified memory depending on configuration.
The result is a local-AI platform spanning more than one size class rather than one oversized demonstration machine.
The Display Still Targets Creator Work
The AI hardware is only one side of the ProArt P16.
ASUS also equips the system with a Lumina Pro OLED display aimed at creator workflows.
The P16 supports configurations up to 4K at 120 Hz with variable refresh rate and NVIDIA G-SYNC.
ASUS lists color accuracy below Delta E 1 and HDR peak brightness up to 1,600 nits.
The P14 reaches up to 3K resolution.
Those specifications matter because many of the workloads ASUS is describing — video generation, editing, 3D rendering and image creation — eventually become visual work.
The machine is being positioned as an AI-capable creator laptop rather than an inference appliance with a keyboard attached.
There Is Also a Desktop Version: ProArt GR1X
ASUS is extending the same RTX Spark platform into the ProArt GR1X Mini PC.
The compact system measures 150 × 150 × 51 mm.
ASUS describes it as an always-on agentic AI computer for creators.
The GR1X includes 10GbE wired networking, Wi-Fi 7, Bluetooth 5.4 and support for up to four 4K displays.
The form factor changes the role.
P16 and P14 are mobile systems.
GR1X can stay on a desk and act as a persistent local AI and creator machine.
That creates two ways to use the same architectural idea:
portable local AI → laptop.
always-on local AI → mini PC.
An Always-On Local Agent Is Different From a Chatbot Tab
An always-on local AI machine can support a different workflow from opening a browser tab whenever a task appears.
A local agent can remain available alongside the user’s files, applications and tools.
Architecturally, that can support workflows such as:
new input arrives → agent processes it locally → tool performs an action → result remains on the machine.
The specific automations still depend on the agent software and permissions.
The GR1X does not automatically perform every possible agent task out of the box.
But an always-on system with local inference capacity gives developers and creators a persistent place to run those workflows without depending on a remote model for every step.
The Real Story Is Memory Architecture, Not Just Another AI Performance Number
AI PCs have spent several years being marketed around accelerator-performance numbers.
RTX Spark makes another number unusually important:
128 GB of unified memory.
A large model has to fit somewhere before the GPU can accelerate it.
A long-context agent needs room for more than the model.
Large video and 3D workflows also consume substantial memory.
That creates a simple chain:
large workload → needs large working set → unified memory provides space → GPU accelerates the computation.
The 1-petaflop figure matters.
The Blackwell Tensor Cores matter.
But the memory architecture is what lets the system point that compute at workloads that would otherwise be difficult to fit inside a slim laptop.
What ASUS and NVIDIA Have Confirmed — and What They Have Not
ASUS and NVIDIA have confirmed the core specifications and workload claims behind RTX Spark.
The ProArt P16, P14 and GR1X use NVIDIA RTX Spark.
The platform combines a Blackwell RTX GPU with 6,144 CUDA cores, a Grace CPU with up to 20 cores and fifth-generation Tensor Cores with FP4 support.
It offers up to 1 petaflop of AI performance and up to 128 GB of unified memory.
NVIDIA says RTX Spark can run 120B-parameter LLMs with up to 1 million tokens of context in supported local-agent workflows, generate 4K AI video and work with 90 GB+ 3D scenes.
ASUS lists the P16 at 12.9 mm and 1.77 kg and the P14 at 13.9 mm and 1.48 kg.
What those sources do not establish is that every 120B model will run at the same speed, every model supports a million-token context, local inference is always faster than cloud inference, or every workload reaches 1 petaflop.
Those claims should not be added.
The Number That Makes 120B Local AI Plausible Is 128 GB
The attention-grabbing number is 120 billion parameters.
The number that makes that class of workload plausible is 128 gigabytes.
Large AI models need somewhere to live before the accelerator can run them.
RTX Spark gives the CPU and GPU access to a large unified memory pool, combines that with Blackwell Tensor Core acceleration and FP4 compute, then packages the system inside a laptop only 12.9 mm thick.
That changes the shape of local AI.
A demanding model no longer automatically implies a rack-mounted server or a remote API.
More of that work can move onto the machine sitting in front of the user.
model → unified memory → local inference → local agent.
The cloud does not disappear.
But it no longer has to be the only place where serious AI work happens.
AI Is Moving From One Location to Several
For a long time, most consumer AI had one obvious home.
The model ran in a data center. The phone or laptop sent a request. The server produced the result. The device displayed it.
That architecture is still important, but it is no longer the only one.
Modern laptops, phones and tablets now include CPUs, GPUs and NPUs capable of running useful AI models directly on the device. At the same time, cloud infrastructure continues to scale into larger models, longer context windows and more demanding multimodal workloads.
The result is a split computing stack.
Some AI can stay close to the user. Some AI can run remotely. The application can decide which execution path fits the job.
Microsoft reflects this directly in its Windows AI guidance, which includes local APIs, Foundry Local, Windows ML and cloud services as parts of the same broader development environment. Apple is moving in the same direction through its Foundation Models framework, on-device system models and Private Cloud Compute.
That changes the question.
Local AI and cloud AI are no longer two competing ideas.
They are becoming two places where the same product can think.
Local AI Brings the Model Closer to the User
Local AI begins with proximity.
The model is running on the same device as the file, camera, microphone, application or user interaction it is working with.
That can make certain AI features feel immediate.
OCR can read text from a local image. A search tool can index files on the machine. A camera feature can process a live feed. A writing tool can summarize or transform text. A small assistant can answer from information already available on the device.
Microsoft’s Windows AI APIs are designed around this kind of execution on supported hardware. The platform exposes local capabilities such as OCR, image description, summarization and access to local models. Apple’s Foundation Models framework likewise gives developers access to an on-device model for tasks such as summarization, entity extraction, refinement and structured generation.
The common idea is simple.
The AI feature can use the machine itself as part of the inference infrastructure.
The laptop is no longer only a window into AI.
It can be one of the places where the AI actually runs.
Low-Latency Tasks Fit Naturally on the Device
Some AI interactions happen often enough that speed becomes part of the product experience.
A camera effect is active continuously. A search box may respond dozens of times in one session. OCR may run whenever a document appears. A text tool may make small changes repeatedly while the user is writing.
Local execution gives these tasks a short path.
The application can send the input directly to the local model without waiting for a remote request to travel across the network first.
That is why many on-device AI features are small, frequent and interactive.
Microsoft has been building Windows AI around this pattern on supported PCs. Apple uses the same direction with on-device Foundation Models. The model can sit close to the application and respond as part of the normal interface.
This changes how AI can be designed.
Instead of saving AI for a large command, developers can use it inside smaller moments throughout the application.
The model becomes part of typing, searching, viewing, organizing and navigating.
That is one of the most important effects of local AI: it makes AI easier to weave into the software itself.
Local Processing Creates a Strong Privacy Architecture
Local AI also creates a useful privacy model for supported workflows.
When inference happens entirely on the device, the input can remain on the machine for that part of the process.
Microsoft states that supported Windows AI APIs process data locally, and its Foundry Local documentation describes inference paths where inputs and outputs stay on the device. Apple’s on-device Foundation Models are designed around the same principle: supported generative work can happen on the user’s hardware.
That creates new possibilities for applications working with personal content.
A document assistant can process local notes. A photo tool can understand images already stored on the device. A search feature can index local files. A writing tool can work with text before any cloud request becomes part of the experience.
This architecture is useful because it gives developers another way to design privacy-conscious products.
The application can decide that some steps belong on the device by default.
Then larger cloud services can be added around those local capabilities when the product wants additional scale.
Privacy becomes part of system design, not only a policy written after the product is built.
Offline AI Makes the Device More Independent
A local model can also continue working when the network is not part of the moment.
That makes offline AI useful in travel, field work, aircraft, remote locations and any workflow where connectivity changes throughout the day.
Microsoft describes Foundry Local as supporting local inference after the required model is available on the machine. Apple likewise exposes on-device Foundation Models for supported generative tasks.
The practical effect is easy to understand.
The user can open the application and keep using the supported AI feature even when the device is disconnected.
A local summarizer can work with a document. A classification model can organize content. OCR can continue reading text. A local assistant can keep working with information stored on the machine.
That is a meaningful change in how AI software behaves.
The application does not have to treat an internet connection as the beginning of every intelligent action.
Instead, the device can carry some of its own intelligence with it.
That makes AI feel more like a built-in computing capability and less like a remote service the machine has to reach before anything useful can happen.
Cloud AI Gives Applications Access to a Much Larger Compute Pool
Cloud AI has a different strength: scale.
A server platform can combine large accelerator fleets, high-capacity memory, fast networking and centralized model infrastructure. That gives applications access to models and workloads far beyond what a portable battery-powered device is designed to carry locally.
This becomes useful for long context, complex reasoning, large multimodal inputs, agentic workflows and other tasks that benefit from a larger compute envelope.
Microsoft’s local-versus-cloud guidance treats model size and complexity as major architectural inputs. Apple does the same in its Foundation Models work, pairing on-device models with Private Cloud Compute for workloads that can use larger server-side capability.
The cloud therefore expands the ceiling of the application.
The device can remain thin and portable while the product reaches infrastructure that may contain far more memory and compute than any laptop or phone.
This is why cloud AI remains central even as local AI improves.
The local device adds immediacy and proximity.
The cloud adds scale.
Longer Context Windows Create a Different Kind of AI Experience
One of the easiest ways to see the difference between device-scale and server-scale AI is context.
Context determines how much information a model can work with inside one interaction.
Apple’s WWDC26 developer material provides a concrete example inside its own stack. Apple describes an on-device system model with roughly a 4K context window and a Private Cloud Compute model with roughly 32K context in that specific framework.
Those figures are Apple-specific, but the architecture is useful to understand.
The on-device model is designed to be available locally. The server model can draw on a larger resource envelope and work with substantially more context.
That difference changes the kinds of experiences an application can create.
A local model can handle short transformations, extraction and immediate assistance. A larger cloud model can take in longer documents, larger conversations or more complex multimodal material.
The product does not have to choose one forever.
It can use the smaller local model for frequent everyday interactions, then move a larger request to the cloud when the task expands.
That is where the two execution paths begin to look complementary rather than separate.
Centralized Cloud Models Can Evolve Across an Entire Service
Cloud AI also gives providers a powerful deployment model.
The model lives in centralized infrastructure.
When the provider updates that model, expands the serving stack or adds new capabilities, the change can become available across the service without moving the full model weights onto every user’s device.
That creates a fast path for platform evolution.
A cloud service can scale capacity, introduce a newer model family, extend context, add tools or improve multimodal processing from the server side.
Developers can then expose those capabilities through the same application interface.
This is one reason cloud AI is especially useful for products that serve many users and need access to large shared infrastructure.
The application can remain relatively lightweight while the provider manages the deeper compute environment centrally.
Local AI and cloud AI therefore distribute responsibility differently.
The local path places more intelligence directly inside the device.
The cloud path places more intelligence inside the service.
Modern applications can use both.
Private Cloud Compute Shows That Cloud AI Can Have a Purpose-Built Privacy Architecture
Cloud AI does not have to mean one generic server model.
Apple’s Private Cloud Compute shows how a provider can build a remote AI architecture around specific privacy and security requirements.
Apple documents PCC as infrastructure designed so user data sent for a request is used for the computation and is not retained after the response, with additional verification and security mechanisms around the system.
That makes PCC an important example because it widens the architecture choices available to developers and platform designers.
An application can keep supported work fully on-device. It can use a purpose-built private cloud architecture for larger requests. It can also connect to other server models where the product design calls for them.
The important shift is choice.
Privacy-sensitive design is not limited to one execution location.
It can influence how the local path is built and how the remote path is built.
That gives modern AI systems more flexibility than the old binary of ‘device equals private’ and ‘cloud equals remote.’
The architecture itself can carry the privacy model.
Hybrid AI Lets the Application Route Work to the Right Place
The most interesting architecture is often the one that uses both locations.
A hybrid AI application can keep lightweight work on the device and send larger work to server infrastructure when the task expands.
The local side might handle OCR, indexing, classification, short summarization, image understanding or quick text generation. The cloud side might handle longer context, deeper reasoning, larger multimodal inputs or an agentic workflow that needs more compute.
Apple is moving directly toward this model through its Foundation Models framework and Private Cloud Compute. Microsoft is doing the same across Windows AI, Foundry Local, Windows ML and cloud services.
This creates a routing layer inside the product.
The user asks for one thing.
The application decides where each part should run.
Some work can happen immediately on the device. More demanding work can move to the server. The result comes back into the same interface.
That is a major architectural change.
AI products are beginning to manage compute location the way modern systems already manage storage, networking and graphics resources.
The location becomes part of the software design.
AI PCs Make Local Execution a Larger Product Category
Local AI is becoming more important because consumer hardware is changing underneath it.
AI PCs increasingly include NPUs designed for neural-network workloads. CPUs and GPUs continue to improve. Memory capacity is rising. Operating systems are exposing local AI APIs. Model developers are creating smaller models designed to run efficiently on endpoint hardware.
Those changes reinforce one another.
Better hardware makes more local AI possible.
More local AI gives developers a reason to target the hardware.
More software gives users a reason to care about the NPU, GPU and local model stack inside the machine.
Microsoft’s Windows AI work is part of that cycle. Apple’s Foundation Models are another example on its platforms.
The device is becoming an AI execution target in its own right.
That does not shrink the role of cloud AI.
It expands the total AI system.
Instead of one remote model doing everything, the product gains another compute layer close to the user.
The Future AI Stack Is Local, Cloud and Everything Between Them
The long-term shift is not difficult to see.
AI is becoming distributed.
The phone can run a model. The laptop can run a model. The operating system can expose local AI services. The cloud can run larger models. Private server architectures can handle sensitive remote workloads. Applications can route between those places as the task changes.
That gives developers more ways to design the experience.
Fast, repetitive and personal interactions can stay close to the device. Offline features can remain available while the network disappears. Larger reasoning tasks can use server-scale compute. Long context can move to infrastructure with more memory. The same application can combine all of those paths without turning them into separate products.
Microsoft and Apple are already building software frameworks around this model.
That is the important signal.
Local AI is no longer a small alternative to cloud AI.
Cloud AI is no longer the only place where intelligence lives.
They are becoming layers of the same computing stack.
The device thinks.
The cloud thinks.
The application decides how to connect them.
That is the upgrade.