ThinkCentre X Ultra Brings Agentic AI Into a 1.6L Desktop

Lenovo introduced the ThinkCentre X Ultra at Innovation World during IFA 2026 on September 3, 2026.

The new system is built around a simple idea: substantial local AI capability does not need a large desktop tower. ThinkCentre X Ultra fits into a 1.6-liter chassis measuring 183 × 183 × 51mm, yet Lenovo positions it as a new class of desktop for the agentic AI era.

That combination makes the launch interesting. The system is not only compact. It is designed around high-memory local AI, developer tooling and a cluster-ready architecture that can connect several units together.

Up to AMD Ryzen AI Max+ PRO 495 Powers the System

At the top of the configuration range, ThinkCentre X Ultra uses AMD Ryzen AI Max+ PRO 495.

AMD lists the Ryzen AI Max+ PRO 495 with 16 Zen 5 CPU cores and 32 threads, while Lenovo pairs the processor with integrated Radeon 8065S graphics and an NPU rated at up to 55 TOPS.

That creates a compact platform with CPU, GPU and NPU resources available inside one system. For local AI development, those compute engines can support different parts of the workflow while keeping the machine small enough to sit almost anywhere on a desk.

Up to 128GB of Unified Memory Is the Real AI Headline

ThinkCentre X Ultra supports up to 128GB of onboard LPDDR5X unified memory.

For local AI, memory capacity is one of the most important parts of the hardware story. A larger memory pool gives the system room for bigger models, longer working contexts and more demanding agent workflows.

Lenovo also allows up to 96GB of that unified memory to be allocated as dedicated graphics memory for the integrated Radeon 8065S graphics. That gives the graphics engine a very large working pool for AI workloads while keeping the platform inside a compact integrated design.

The Memory Runs at Up to 8533MHz Across Four Channels

Lenovo specifies the onboard LPDDR5X memory at up to 8533MHz with four-channel support.

The combination of capacity and bandwidth is designed to keep large local workloads moving efficiently through the system. For AI developers, that matters because model execution involves constant movement of weights, context and intermediate data between memory and compute resources.

ThinkCentre X Ultra is therefore not simply a small office PC with extra memory. Its memory architecture is central to the local AI role Lenovo has designed for it.

Four ThinkCentre X Ultra Systems Can Become One AI Platform

The standout feature is the cluster-ready architecture.

Lenovo says up to four ThinkCentre X Ultra systems can be connected into a unified platform. The goal is to expand the compute and memory resources available to AI workloads beyond one compact desktop.

That changes the product from a single small PC into a building block. A developer can start with one system and use several systems together when the workflow grows.

The Cluster Is Designed for Larger AI Models

Lenovo explicitly connects the four-system architecture with the ability to run larger AI models.

Each ThinkCentre X Ultra brings its own compute and memory resources into the broader platform. For teams experimenting with local generative AI, that creates a path from one compact workstation toward a more substantial local compute environment.

The most interesting part is the form factor: the expansion happens by adding another 1.6L system rather than moving immediately to a large traditional server or workstation footprint.

Longer Context Windows Are Part of the Cluster Story

Lenovo also says the clustered platform can support longer context windows.

Long context is increasingly important for agentic AI. Coding agents may need to work across large repositories, research agents may process many documents, and business agents may need a substantial amount of project material available during one workflow.

ThinkCentre X Ultra is designed to give those workloads access to more local compute and memory as the deployment scales from one system to several.

Multi-Agent Workflows Are a First-Class Target

Lenovo is positioning ThinkCentre X Ultra directly for multi-agent workflows.

Instead of one assistant performing one task, agentic systems can use several specialized agents working in parallel. One can plan, another can analyze documents, another can write code and another can prepare a final result.

The cluster-ready design gives those concurrent workloads a local hardware platform that can grow with the number of agents and the amount of work being handled.

The Platform Can Handle Multiple AI Requests at Once

Lenovo also highlights simultaneous AI requests as part of the system’s scaling story.

That is useful for shared local AI environments where several applications, agents or users may need model inference at the same time. A multi-system ThinkCentre X Ultra setup can provide a broader local compute pool for those requests.

This is where the four-node concept becomes more than a spec-sheet feature. It gives local AI a way to become a shared service inside a compact business or development environment.

AMD Ryzen AI Developer Center Is Integrated

ThinkCentre X Ultra is integrated with AMD Ryzen AI Developer Center.

Lenovo says this gives users access to preconfigured AI tools, models and workflows across Windows and Linux. That software layer is important because powerful hardware becomes much more useful when developers can reach working tools and models quickly.

The integration is designed to shorten the path from opening the system to experimenting with local AI applications.

Windows and Linux Are Both Part of the Developer Story

Lenovo supports Windows 11 as well as Linux options for ThinkCentre X Ultra.

The specification list includes Windows 11 Pro and Home, Linux AMD AI OS and Ubuntu certification. Combined with AMD Ryzen AI Developer Center, that gives developers flexibility in how they build local AI projects.

A Windows-focused team can stay inside its existing environment, while Linux-oriented AI developers can work with the toolchains they already use.

Up to 8TB of High-Speed SSD Storage Fits Inside

ThinkCentre X Ultra supports up to two 4TB M.2 SSDs, creating up to 8TB of internal solid-state storage in the compact chassis.

Local AI projects can quickly accumulate models, datasets, embeddings, source repositories and generated assets. Large internal storage gives developers room to keep more of that material close to the compute platform.

It also reinforces the idea that ThinkCentre X Ultra is intended to operate as a serious local AI workstation rather than only as a thin client for cloud services.

10GbE Gives the Desktop High-Speed Wired Networking

Lenovo includes 10-gigabit Ethernet in the ThinkCentre X Ultra port selection.

The rear panel includes a 10GbE RJ-45 connection, and the optional punch-out port can also be configured with another 10GbE interface. High-speed wired networking is a natural fit for a desktop designed around local AI, large files and multi-system workflows.

It gives the small chassis connectivity that matches the scale of the compute and memory inside it.

Thunderbolt 4 and Modern Display Outputs Expand the Workspace

The rear I/O also includes two Thunderbolt 4 ports, DisplayPort 2.1 and HDMI 2.1.

That gives ThinkCentre X Ultra a broad set of options for displays, high-speed peripherals and external workflows. The front adds two USB-C ports and a headset connection, keeping frequently used ports within easy reach.

For a system that can act as both a local AI node and a daily workstation, that balance of compute and connectivity makes the compact design more versatile.

Wi-Fi 7 Adds High-Speed Wireless Connectivity

ThinkCentre X Ultra also supports Wi-Fi 7 and Bluetooth 5.4.

That gives the desktop modern wireless connectivity alongside its high-speed wired networking. For flexible office layouts, development labs and creative workspaces, the system can fit into different network arrangements without turning its small footprint into a cabling project.

The result is a compact machine that can sit quietly in a workspace while staying connected to modern peripherals and infrastructure.

Adaptive Lighting Turns System Activity Into Visual Feedback

Lenovo adds a visual touch with Adaptive Lighting.

The feature transforms system activity into real-time visual feedback, giving users a quick way to see the state of the machine at a glance. That is especially fitting for an AI workstation that may continue processing local workloads while the user is focused on something else.

The lighting becomes part of the interface between the physical machine and the background compute activity happening inside it.

The Thermal Design Is Built for Sustained Work

Lenovo designed the cooling system around sustained AI workloads while keeping the chassis compact.

The company says the thermal design supports reliable and quiet operation during extended workloads. That is important for a desktop intended to sit directly in a workspace and continue running local inference, agent tasks or development workloads over longer periods.

The engineering goal is clear: keep the local AI capability close to the user without giving up the compact 1.6L form factor.

Interior view of a compact small-form-factor desktop computer
Compact desktop engineering is central to the ThinkCentre X Ultra story. This generic open-license interior view is illustrative and is not a ThinkCentre X Ultra component photo. Original image by Dllu / Wikimedia Commons, adapted by That Upgrade Feeling with TUF watermark/branding.

Enterprise Features Sit Alongside the AI Hardware

ThinkCentre X Ultra also includes Lenovo ThinkShield, AMD PRO technologies and AMD DASH manageability.

That positions the system for professional environments where local AI hardware needs to fit into existing device-management practices. Lenovo also lists discrete TPM 2.0, TCG certification and FIPS 140-2 certification among the platform’s security features.

The AI workstation is therefore designed as part of a managed business fleet as well as a high-performance local development machine.

A 2kg Starting Weight Keeps the System Truly Compact

ThinkCentre X Ultra starts at 2kg while fitting into a chassis just over seven inches wide and deep.

That physical scale is part of what makes the four-system idea interesting. Several nodes can provide a substantial local AI platform without requiring the footprint normally associated with multiple full-size workstations.

For development teams or offices where desk and lab space matter, the form factor becomes part of the compute strategy.

Lenovo Plans Availability From November 2026

Lenovo says the ThinkCentre X Ultra will be available starting in November 2026.

That puts the product on a near-term path from IFA announcement to commercial availability. For developers and businesses building more local AI into their workflows, the system represents a new option that combines compact hardware, large unified memory and a multi-node scaling model.

The launch also expands Lenovo’s ThinkCentre family further into dedicated local AI infrastructure.

The Bigger Idea Is a Modular Local AI Desktop

ThinkCentre X Ultra is most interesting when viewed as a modular local AI building block.

One 1.6L machine can serve as a compact AI workstation. Several can become a larger platform for models, contexts, agents and concurrent requests. The same product therefore spans individual development and small-scale local AI infrastructure.

That is a useful direction for personal and business AI because it gives compute a physical form that can grow in small, manageable steps.

The Upgrade Feeling

Lenovo ThinkCentre X Ultra takes the idea of a mini PC much further than simple space saving.

Up to 128GB of unified memory, Ryzen AI Max+ PRO 495, Radeon 8065S graphics, AMD Ryzen AI Developer Center and a four-system cluster-ready architecture turn the tiny chassis into a serious local AI platform.

The upgrade is the ability to start small and scale physically. One box can be a powerful local AI workstation. Four boxes can become a broader platform for larger models, longer contexts and multiple agents working at the same time.

That makes the ThinkCentre X Ultra feel less like a miniature desktop and more like a new modular form of local AI infrastructure.

ASUS is showcasing new ProArt P16 and P14 laptops at IFA 2026 powered by NVIDIA RTX Spark. The platform combines a Blackwell RTX GPU, Grace CPU, up to 128 GB of unified memory and up to 1 petaflop of FP4 AI performance. ASUS and NVIDIA say supported local workflows can run LLMs up to 120 billion parameters, work with up to 1 million tokens of context, generate 4K AI video and handle 90 GB+ 3D scenes. The ProArt P16 packages that architecture into a chassis just 12.9 mm thick.

A 120B Model Sounds Like Server Hardware — ASUS Is Putting That Class of Workload in a 12.9 mm Laptop

A 120-billion-parameter AI model sounds like something that belongs in a server rack.

ASUS is putting that class of local workload into a laptop only 12.9 mm thick.

The new ProArt P16 uses NVIDIA’s RTX Spark platform with up to 128 GB of unified memory and up to 1 petaflop of FP4 AI performance. ASUS and NVIDIA say the platform can run large language models with as many as 120 billion parameters locally, including supported agent workflows with context windows up to 1 million tokens.

The important part is not simply that the laptop has a fast GPU.

It is that the CPU and GPU are designed around one unusually large memory pool.

Large local models need compute, but before the accelerator can process a model, the weights and working state need somewhere to live.

That memory architecture is what makes the 120B figure interesting in a laptop this thin.

The 120B Number Needs One Important Phrase: “Up To”

The ProArt P16 does not ship with one specific 120-billion-parameter model installed by default.

The claim is about the class of workload the RTX Spark platform is designed to support.

ASUS says RTX Spark enables creators and developers to run LLMs with up to 120 billion parameters locally.

That means the chain is:

hardware platform → available compute and memory → compatible model and software → local inference.

The exact experience can vary substantially from one model to another.

Parameter count does not tell you model architecture, quantization, context length, runtime overhead or inference speed.

So “120B” is best treated as a supported upper workload class under compatible conditions, not as a promise that every 120B model will behave identically on the machine.

Why 128 GB of Unified Memory Changes the Equation

On a conventional PC, system memory and GPU memory are usually separate pools.

The CPU works primarily from system RAM.

A discrete GPU works primarily from its own VRAM.

That separation matters when a model becomes very large.

If the model or its active working set does not fit comfortably in the GPU’s available memory, the workflow can become more complicated and data movement becomes part of the problem.

RTX Spark changes that arrangement.

ASUS says the platform provides up to 128 GB of high-bandwidth unified memory shared across the integrated system.

The CPU and GPU can therefore work from a much larger common memory pool instead of treating a smaller discrete VRAM allocation as the only fast memory available to the accelerator.

For large local AI, that is one of the most important changes in the machine.

A 120B Model Is a Memory Problem Before It Becomes a Speed Problem

A useful way to understand the scale is to look at the model weights themselves.

A 120-billion-parameter model at FP16 would require roughly 240 GB for the weights alone if every parameter used two bytes.

At FP8, the same simple calculation is roughly 120 GB.

At FP4, four bits per parameter works out to roughly 60 GB for the raw weights.

Those are illustrative calculations, not the exact memory footprint of every model.

Real inference also needs room for runtime state, caches, context, temporary buffers and software overhead.

But the comparison explains why FP4 support and a 128 GB unified pool can matter together.

The accelerator does not only need to be fast enough.

The system needs enough usable memory to hold the model and the rest of the inference workload at the same time.

RTX Spark Combines a Blackwell GPU and Grace CPU

RTX Spark is a highly integrated NVIDIA platform rather than a conventional laptop pairing of an unrelated CPU and discrete graphics chip.

ASUS says the superchip combines an NVIDIA Blackwell RTX GPU with 6,144 CUDA cores, fifth-generation Tensor Cores with FP4 precision and an NVIDIA Grace CPU with up to 20 cores.

The two sides are connected through NVIDIA NVLink-C2C.

That matters because the architecture is being designed around shared AI and graphics workloads from the start.

The CPU handles general-purpose work and system orchestration.

The Blackwell GPU provides the parallel compute and Tensor Core acceleration used by many AI workloads.

The shared memory architecture connects those pieces into one system instead of forcing every large workload through the assumptions of a traditional discrete-GPU PC.

One Petaflop Does Not Mean Every AI Task Runs at One Petaflop

ASUS and NVIDIA advertise up to 1 petaflop of AI performance for RTX Spark.

That number needs context.

It refers to the platform’s peak FP4 AI compute capability under suitable workloads.

It does not mean every application receives one petaflop of real-world performance.

Actual throughput depends on the model, precision, quantization format, software stack, context length, batch size, memory behavior and how well the workload maps to the hardware.

The useful point is that RTX Spark combines very low-precision AI compute with a large unified memory pool.

Compute and memory solve different parts of the problem.

The petaflop figure describes how much mathematical work the accelerator can theoretically process.

The 128 GB figure describes how much working data the system can keep available.

Large local models need both.

The Bigger Shift Is Running the Agent on Your Own PC

ASUS is positioning RTX Spark around personal agents as much as traditional generative AI.

The workflow can move from:

prompt → remote service → cloud model → response

toward:

prompt → local model → local files and tools → local action.

That does not mean every agent should run locally.

It means the user has enough on-device compute to move more advanced workloads onto the PC when the model and software support it.

ASUS explicitly says RTX Spark is designed to let creators and developers explore advanced AI applications directly on their PCs without relying exclusively on cloud processing.

That phrase matters.

The local machine becomes a serious inference target rather than only a thin client for a remote model.

“Local” Does Not Mean the Cloud Disappears

The local-versus-cloud story should not be framed as an all-or-nothing choice.

ASUS describes a hybrid workflow.

Local RTX-powered AI can handle workloads on the PC.

Cloud services can still be used for burst capacity, frontier-scale models or services that are only offered remotely.

That gives the user another execution option.

A model that fits the machine and matches the task can run locally.

A larger or specialized service can still run in the cloud when that makes more sense.

The practical change is therefore not “no cloud required for everything.”

It is that the cloud no longer has to be the only place where a demanding AI workflow can happen.

Token-Free Local AI Does Not Mean AI Has No Cost

ASUS uses the phrase token-free local AI when describing workloads that run on the user’s own hardware.

The practical meaning is straightforward.

If the model is running locally, there is no remote API provider billing that local inference by generated or processed token.

That is different from saying the AI is free.

The hardware has a purchase cost.

The computer uses electricity.

Some software or models may have their own licensing terms.

Maintenance and storage still exist.

So the useful distinction is:

local inference → no per-token cloud inference charge for that local workload.

That can make repeated experimentation, local agents and iterative creative workflows easier to budget because usage is tied to owned compute capacity rather than a metered remote API.

RTX Spark Supports Agent Workflows With Up to One Million Tokens of Context

NVIDIA says RTX Spark can run 120-billion-parameter LLMs with up to 1 million tokens of context using local agents.

Context length matters because an agent often needs more than a short conversation.

A long-context workflow can include documents, source code, research material, project history, tool outputs and intermediate state.

The larger the active context becomes, the more memory the runtime may need.

That links the million-token claim back to the same architectural theme.

The 128 GB unified pool is not only useful for model weights.

It can also provide room for the working state around the model.

The “up to” still matters: the specific model and software stack must support the context length being used.

Local Agents Need More Than the Model Weights

Loading a model is only the beginning of an agent workflow.

The system may also need:

model weights,

context,

KV cache,

tool state,

application memory,

temporary inference buffers,

and sometimes multiple models or encoders.

That is why raw parameter count is not enough to predict whether a local workflow will fit comfortably.

A 120B model can consume a large amount of memory before the first response is generated.

Then the context and runtime add more.

Unified memory gives the system a larger shared workspace for that full chain.

The important question becomes less “how much VRAM does the GPU have?” and more “how much usable memory can the AI workload access across the system?”

RTX Spark Is Also Targeting Local 4K AI Video Generation

The platform is not designed only for text models.

ASUS and NVIDIA also highlight 4K AI video generation as an RTX Spark workload.

That broadens the meaning of local AI on the ProArt machines.

A language model primarily works with tokens and model state.

Generative video has to create and transform large sequences of image data.

3D workloads add geometry, textures, materials and rendering state.

Different applications stress the machine in different ways, but the same combination of large shared memory and GPU acceleration can support all of them.

That is why ASUS is positioning ProArt P16 as a creator system rather than a laptop built around one chatbot.

The Same Platform Can Work With 90 GB+ 3D Scenes

NVIDIA also says RTX Spark can render ultra-large 3D scenes larger than 90 GB.

Again, the number points back to memory capacity.

A large scene can contain geometry, textures, simulation data and rendering resources that would exceed the VRAM capacity of many conventional laptop GPUs.

A large unified memory pool changes the ceiling.

The graphics and AI accelerator can work with a much larger data set without treating a small discrete VRAM allocation as the only high-performance workspace.

That does not guarantee identical performance for every 90 GB scene.

Scene structure, application behavior and rendering settings still matter.

But it shows why memory architecture is central to RTX Spark’s pitch.

All of This Is Going Into a 12.9 mm Chassis

The ProArt P16 is where the architecture becomes visually surprising.

ASUS says the new model is 12.9 mm thick and weighs 1.77 kg.

It also carries a 99 Wh battery.

That puts a platform designed for large local AI, 4K AI video and heavy creator workflows into a chassis closer to an ultrathin creator notebook than a conventional mobile workstation.

ASUS describes it as the thinnest 16-inch RTX Spark laptop and the thinnest 16-inch ProArt it has produced.

The 120B figure attracts attention.

The physical packaging is what makes the story unusual.

Large local AI is moving into hardware that is designed to travel.

The Smaller ProArt P14 Is 13.9 mm Thick

ASUS is also putting RTX Spark into the ProArt P14.

The company lists the P14 at 13.9 mm thick and 1.48 kg, with a 90 Wh battery.

ASUS calls it the lightest ProArt laptop it has produced.

That matters because RTX Spark is not being limited to one 16-inch flagship chassis.

The platform is being packaged into a smaller portable form factor as well.

Both laptops support up to 128 GB of unified memory depending on configuration.

The result is a local-AI platform spanning more than one size class rather than one oversized demonstration machine.

The Display Still Targets Creator Work

The AI hardware is only one side of the ProArt P16.

ASUS also equips the system with a Lumina Pro OLED display aimed at creator workflows.

The P16 supports configurations up to 4K at 120 Hz with variable refresh rate and NVIDIA G-SYNC.

ASUS lists color accuracy below Delta E 1 and HDR peak brightness up to 1,600 nits.

The P14 reaches up to 3K resolution.

Those specifications matter because many of the workloads ASUS is describing — video generation, editing, 3D rendering and image creation — eventually become visual work.

The machine is being positioned as an AI-capable creator laptop rather than an inference appliance with a keyboard attached.

There Is Also a Desktop Version: ProArt GR1X

ASUS is extending the same RTX Spark platform into the ProArt GR1X Mini PC.

The compact system measures 150 × 150 × 51 mm.

ASUS describes it as an always-on agentic AI computer for creators.

The GR1X includes 10GbE wired networking, Wi-Fi 7, Bluetooth 5.4 and support for up to four 4K displays.

The form factor changes the role.

P16 and P14 are mobile systems.

GR1X can stay on a desk and act as a persistent local AI and creator machine.

That creates two ways to use the same architectural idea:

portable local AI → laptop.

always-on local AI → mini PC.

An Always-On Local Agent Is Different From a Chatbot Tab

An always-on local AI machine can support a different workflow from opening a browser tab whenever a task appears.

A local agent can remain available alongside the user’s files, applications and tools.

Architecturally, that can support workflows such as:

new input arrives → agent processes it locally → tool performs an action → result remains on the machine.

The specific automations still depend on the agent software and permissions.

The GR1X does not automatically perform every possible agent task out of the box.

But an always-on system with local inference capacity gives developers and creators a persistent place to run those workflows without depending on a remote model for every step.

The Real Story Is Memory Architecture, Not Just Another AI Performance Number

AI PCs have spent several years being marketed around accelerator-performance numbers.

RTX Spark makes another number unusually important:

128 GB of unified memory.

A large model has to fit somewhere before the GPU can accelerate it.

A long-context agent needs room for more than the model.

Large video and 3D workflows also consume substantial memory.

That creates a simple chain:

large workload → needs large working set → unified memory provides space → GPU accelerates the computation.

The 1-petaflop figure matters.

The Blackwell Tensor Cores matter.

But the memory architecture is what lets the system point that compute at workloads that would otherwise be difficult to fit inside a slim laptop.

What ASUS and NVIDIA Have Confirmed — and What They Have Not

ASUS and NVIDIA have confirmed the core specifications and workload claims behind RTX Spark.

The ProArt P16, P14 and GR1X use NVIDIA RTX Spark.

The platform combines a Blackwell RTX GPU with 6,144 CUDA cores, a Grace CPU with up to 20 cores and fifth-generation Tensor Cores with FP4 support.

It offers up to 1 petaflop of AI performance and up to 128 GB of unified memory.

NVIDIA says RTX Spark can run 120B-parameter LLMs with up to 1 million tokens of context in supported local-agent workflows, generate 4K AI video and work with 90 GB+ 3D scenes.

ASUS lists the P16 at 12.9 mm and 1.77 kg and the P14 at 13.9 mm and 1.48 kg.

What those sources do not establish is that every 120B model will run at the same speed, every model supports a million-token context, local inference is always faster than cloud inference, or every workload reaches 1 petaflop.

Those claims should not be added.

The Number That Makes 120B Local AI Plausible Is 128 GB

The attention-grabbing number is 120 billion parameters.

The number that makes that class of workload plausible is 128 gigabytes.

Large AI models need somewhere to live before the accelerator can run them.

RTX Spark gives the CPU and GPU access to a large unified memory pool, combines that with Blackwell Tensor Core acceleration and FP4 compute, then packages the system inside a laptop only 12.9 mm thick.

That changes the shape of local AI.

A demanding model no longer automatically implies a rack-mounted server or a remote API.

More of that work can move onto the machine sitting in front of the user.

model → unified memory → local inference → local agent.

The cloud does not disappear.

But it no longer has to be the only place where serious AI work happens.