NVIDIA Is Bringing the Open Model Hub Into Its AI Platform

NVIDIA announced on September 3, 2026 that it has agreed to acquire Hugging Face. The announcement immediately connects two very different but highly complementary parts of modern AI development. NVIDIA builds the accelerated computing platforms used across training, inference, graphics, robotics, and large-scale AI infrastructure. Hugging Face has become a central place where developers discover models, datasets, applications, libraries, and deployment options. NVIDIA’s stated goal is to scale the Hugging Face platform, strengthen its infrastructure, and expand access to AI for developers and institutions around the world.

Hugging Face Has Become One of AI’s Main Discovery Layers

The scale NVIDIA highlighted helps explain why this deal matters to developers. According to NVIDIA’s announcement, more than 18 million developers, researchers, and creators use Hugging Face. The platform hosts more than 3 million models, around 500,000 datasets, and about 1 million applications, while more than 200,000 companies use the platform to discover, evaluate, customize, and deploy AI. Those numbers make Hugging Face more than a repository. It acts as a discovery layer across the open-model ecosystem, connecting model creators, application builders, inference providers, researchers, and organizations in one shared environment.

The Most Important Promise Is That the Hub Stays Open

NVIDIA’s announcement makes openness a central part of the acquisition. The company says Hugging Face will remain an open platform for the entire AI ecosystem. Developers will continue choosing the models they want, the frameworks they want, the cloud platforms they want, the inference service providers they want, and the computing platforms they want. NVIDIA goes one step further and says NVIDIA compute will not be required to build on or deploy through Hugging Face. That statement preserves one of the Hub’s defining characteristics: it is designed to connect many models, tools, and providers rather than force every workflow into one stack.

Model Choice Remains at the Center of the Experience

Hugging Face’s own documentation describes the Hub as a reference platform for open machine learning and as a collaboration layer for models, datasets, and applications. Model repositories can contain weights, configuration files, documentation, evaluation information, and version history. Developers can browse, compare, download, fine-tune, and integrate models using a wide range of libraries. NVIDIA’s commitment to preserve model choice means this workflow is expected to remain broad. A developer can continue selecting the model that fits the task rather than treating the platform as a catalog tied to one model family.

Framework Choice Is Part of the Same Open Design

Modern AI development rarely uses one framework for every job. Teams move between Transformers, PyTorch-based workflows, optimized inference runtimes, local engines, orchestration tools, and custom application code. NVIDIA explicitly says developers will continue choosing their preferred frameworks on Hugging Face. That is important because the Hub has grown partly by serving as a common meeting point between many software ecosystems. A model can live in one repository while being discovered, tested, downloaded, or deployed through different tools. Keeping that flexibility gives the combined platform room to support many different development styles.

Cloud Choice Also Remains Flexible

NVIDIA’s statement also preserves cloud choice. Hugging Face already supports dedicated endpoints and deployment workflows across different infrastructure environments, and its documentation describes the Hub as a collaboration layer rather than a single-cloud destination. That flexibility matters because AI teams often choose deployment environments based on workload, geography, scale, organizational requirements, or existing architecture. NVIDIA’s announcement says those choices will remain available after the acquisition, keeping the Hub positioned as a common layer that can connect models to multiple infrastructure paths.

Inference Provider Choice Is Especially Important

Hugging Face’s Inference Providers system already gives developers one interface for running models through a broad set of serverless inference partners. Its current documentation lists providers including Cerebras, Cohere, DeepInfra, fal, Fireworks, Groq, Replicate, Scaleway, Together, and others alongside Hugging Face’s own inference services. Developers can use the Hugging Face SDK with a selected provider or let the client route automatically. NVIDIA’s promise that inference-provider choice will continue is therefore directly connected to a major part of the Hub’s current design.

The Hub Is More Than a Model Download Page

Hugging Face repositories are Git-based and support versioning, commit history, diffs, branches, collaboration, and integrations. The platform also hosts datasets and Spaces, giving developers a way to move from a model file to evaluation data, demonstrations, applications, and interactive experiments without leaving the same ecosystem. That broader structure is one reason the acquisition reaches beyond model hosting. NVIDIA is not only gaining a catalog of weights. It is bringing a large developer workflow, collaboration surface, and model-discovery network closer to its own AI software and infrastructure ecosystem.

Datasets Are a Major Part of the Platform

Datasets are another important layer. Hugging Face’s documentation describes datasets on the Hub as repositories that can include training, evaluation, and testing data together with Dataset Cards and browser-based viewers. Developers can search by task, language, license, and other attributes, then access datasets through the Hub or programmatically through the datasets library. NVIDIA’s acquisition therefore connects compute not only to finished models, but also to the data workflows used to evaluate, adapt, and build AI systems. That gives the combined ecosystem a wider development surface from experimentation through deployment.

Laptop displaying programming code
Illustrative developer-workflow image. Pexels source via Wikimedia Commons, CC0 1.0. TUF branding/watermark required for publication.

Spaces Turn Models Into Working Applications

Hugging Face Spaces add an application layer on top of the Hub. Developers can build interactive demos with Gradio, Docker, or static HTML, link models and datasets, and publish working AI experiences that other users can test directly in the browser. This turns model discovery into something more tangible. Instead of only reading a model card, a user can often interact with an application built around the model. In the context of NVIDIA’s acquisition, Spaces give the platform a visible application layer that sits between model repositories and full production deployment.

NVIDIA Already Works With Hugging Face Models Today

The two ecosystems are already connected technically. NVIDIA’s NIM documentation includes Hugging Face as a model source, using Hugging Face repository identifiers and tokens when needed. NVIDIA NeMo Platform documentation also describes workflows for deploying supported Hugging Face models. That means the acquisition is not starting from zero integration. Developers already move models between Hugging Face repositories and NVIDIA deployment tooling. Bringing the organizations together can make those paths easier to coordinate while the Hub continues supporting other infrastructure options.

The Combination Connects Discovery With Accelerated Deployment

One of the clearest opportunities is the connection between discovering a model and running it efficiently. Hugging Face is where many developers begin: search for a model, inspect the model card, compare alternatives, test an application, download weights, or call an inference provider. NVIDIA’s software stack begins to matter when teams want accelerated training, optimized inference, or larger-scale deployment. Bringing those layers closer creates a more continuous path from finding a model to experimenting with it and then moving into optimized production infrastructure when that is the right fit.

Open Models Can Reach More Production Paths

Hugging Face’s current inference architecture already supports several routes: hosted inference providers, managed endpoints, and local endpoints such as llama.cpp, Ollama, vLLM, LiteLLM, and Text Generation Inference. NVIDIA’s announcement says those kinds of choices will remain open. That means the Hub can continue serving as a model layer that feeds many production paths. NVIDIA can add stronger integrations and infrastructure around its own stack without removing the broader routing model that developers already use.

The Developer Experience Is the Real Center of the Deal

The announcement is easy to describe as a connection between a hardware company and a model platform, but the developer workflow is the more useful lens. Hugging Face sits where developers search, compare, evaluate, share, and collaborate. NVIDIA sits where many teams optimize and scale AI workloads. The value of the combination comes from reducing the distance between those activities. A developer can begin with an open model, evaluate it, test it in an application, choose an inference path, and move toward deployment while staying inside a connected ecosystem.

Hugging Face Keeps Its Role as a Multi-Provider Layer

The most distinctive part of NVIDIA’s announcement is that Hugging Face is expected to continue supporting many providers rather than becoming an NVIDIA-only front end. Developers can keep choosing cloud platforms, inference services, frameworks, and compute. That preserves the Hub’s usefulness as neutral connective tissue across AI workflows. NVIDIA can still improve integration with its own software and infrastructure, but the platform can remain valuable to developers who use different hardware or deployment environments.

NVIDIA Gets Closer to Where Model Decisions Begin

Most infrastructure decisions happen after a team has already chosen or narrowed down a model. Hugging Face sits much earlier in that process. Developers use the Hub to discover what exists, inspect how a model was built, compare versions, read documentation, test demos, and identify compatible deployment options. By acquiring Hugging Face, NVIDIA moves closer to that starting point. It gains a direct connection to the layer where millions of developers begin deciding what models and tools they want to use.

Hugging Face Gets a Larger Infrastructure Partner

NVIDIA’s stated plan is to scale Hugging Face’s platform and strengthen its infrastructure. For a service hosting millions of AI artifacts and serving developers around the world, infrastructure is a central part of the product experience. Faster access, stronger deployment paths, larger-scale services, and deeper optimization can all make the Hub more useful as model sizes and application demands grow. NVIDIA’s accelerated computing experience gives the combined organization a substantial technical foundation for expanding those capabilities.

The Open Model Ecosystem Becomes More Connected

AI development is increasingly distributed across model creators, dataset builders, application developers, cloud platforms, inference services, local runtimes, and hardware vendors. Hugging Face already connects many of those pieces. NVIDIA adds another large layer of infrastructure and software to that network. If the company follows the open-platform commitments in its announcement, the result can be a more connected ecosystem where developers keep choosing their tools while gaining more direct paths into accelerated infrastructure when they need it.

This Is a Platform Story More Than a Hardware Story

NVIDIA is best known for accelerated computing, but this acquisition is fundamentally about software distribution and developer access. Hugging Face is one of the places where open models become discoverable, testable, shareable, and deployable. That makes the deal important even for developers who are not thinking about GPUs at the moment they open the Hub. The strategic connection is between the place where AI assets are organized and the infrastructure that can help run them at scale.

The Upgrade Feeling

The biggest idea here is not that NVIDIA now owns another AI company. It is that the company is moving closer to one of the main places where developers discover and use open models. Hugging Face brings the models, datasets, apps, repositories, collaboration tools, and provider connections. NVIDIA brings accelerated infrastructure and a large software stack built around AI deployment. The most encouraging part of the announcement is the promise that the Hub will remain open and multi-provider. If that stays true in practice, developers gain a tighter bridge between open-model discovery and production-scale AI without giving up the choices that made Hugging Face useful in the first place.

Intel’s Hot Chips 2026 roadmap shows how the company thinks agentic AI changes hardware architecture across the stack. Diamond Rapids targets enterprise-scale orchestration with up to 256 cores, 16 memory channels, PCIe 6.0 and CXL 3.0. Crescent Island targets inference economics with up to 480 GB of LPDDR5X on a 350-watt air-cooled PCIe GPU. Wildcat Lake brings a smaller hybrid-AI design to mainstream clients and edge systems with integrated Xe3 graphics and an NPU rated at up to 17 TOPS. The three products are different, but the strategy is one: agents need orchestration, inference and local execution to work as a coordinated system rather than as one giant accelerator doing everything.

Agentic AI Changes the Hardware Question

Most AI-chip discussions ask one question.

How much acceleration can one processor deliver?

Agentic AI makes that question too narrow.

An agent does not simply generate one answer.

It can plan.

Call tools.

Run code.

Search memory.

Trigger other models.

Use databases.

Interact with the edge.

Repeat the loop.

That creates several different compute jobs inside one workflow.

Intel’s Hot Chips 2026 announcements are interesting because the company is not presenting one universal AI processor.

It is dividing the work across three architectures.

Diamond Rapids for large-scale orchestration and general compute.

Crescent Island for inference.

Wildcat Lake for mainstream client and edge execution.

The hardware is being designed around a system of agents rather than one model call.

Intel Presented Three Architectures, Not One AI Chip

Intel’s August 24 Hot Chips 2026 announcement centers on Diamond Rapids, Crescent Island and Wildcat Lake.

They sit in very different power and deployment classes.

Diamond Rapids is a next-generation Xeon server processor.

Crescent Island is a datacenter GPU optimized for inference.

Wildcat Lake is a Core Series 3 SoC for mainstream laptops and intelligent edge platforms.

Intel describes the portfolio as a way to scale agentic AI from the rack to the edge.

That framing matters because agent workloads naturally spread across infrastructure.

The orchestration may happen on a CPU.

Heavy generation may happen on a GPU.

A local model may stay on the endpoint.

A real system can use all three.

Diamond Rapids Is the Orchestration Layer

Intel positions Diamond Rapids as the compute foundation for enterprise-scale agentic AI.

That does not mean the CPU replaces accelerators.

The role is broader.

Agents need general-purpose compute for scheduling, tool execution, business logic, memory management, networking, databases and coordination.

Those jobs are not always best handled by a GPU.

Intel is betting that the CPU remains central even when most of the attention goes to AI accelerators.

The more agents an enterprise runs, the more orchestration work exists around the models themselves.

Up to 256 Cores Changes the Scale of General-Purpose Work

Intel says Diamond Rapids will offer up to 256 new cores.

That is a major increase in server-side general-purpose compute density.

Agent workloads can create a large number of parallel tasks.

One process is waiting on a model.

Another is running a database query.

Another is executing a tool.

Another is validating output.

Another is handling security policy.

A high-core-count CPU can absorb that surrounding work while accelerators focus on matrix-heavy inference.

The important point is not simply “256 cores.”

It is what those cores are being asked to coordinate.

The 1.28 GB Last-Level Cache Is Part of the Story

Intel lists up to 1.28 GB of last-level cache for Diamond Rapids.

Large cache capacity can reduce how often workloads have to reach external memory for frequently reused data.

That matters in server environments where many software layers are active at once.

Agentic systems may include model runtimes, retrieval systems, databases, orchestration frameworks and network services.

Not every workload benefits equally from a large cache.

But the specification shows that Intel is designing for high-density server workloads where keeping more data closer to the cores can matter.

Memory Bandwidth Is Becoming an AI Constraint

AI discussions often focus on arithmetic.

Data movement can be just as important.

Diamond Rapids supports 16 memory channels at up to 12,800 MT/s, according to Intel’s Hot Chips material.

More memory channels increase aggregate bandwidth.

That helps workloads that need to feed many cores with data continuously.

Agentic systems may create memory pressure from retrieval, context preparation, databases and concurrent services even when the largest model itself runs on a separate accelerator.

The CPU still needs a fast path to data.

PCIe 6.0 and CXL 3.0 Are About the Rest of the Rack

A modern AI server is not one processor.

It is a network of processors, accelerators, memory devices, NICs and storage.

Diamond Rapids includes 128 lanes of PCIe Gen 6 and support for CXL 3.0.

Those interfaces are important because they determine how much external hardware a CPU can connect to and how quickly data can move between components.

CXL can also support more flexible memory architectures.

For agentic AI, the server platform has to move information between many specialized devices without turning I/O into the bottleneck.

Foveros Direct and UCIe Show Where Intel Wants Packaging to Go

Intel is also emphasizing advanced packaging.

Diamond Rapids uses Foveros Direct 3D and UCIe-S interconnect.

UCIe is an industry standard for connecting chiplets inside a package.

The strategic idea is modularity.

Instead of building one huge monolithic die containing every function, designers can combine smaller compute blocks, I/O blocks and other components.

That can make future processors easier to scale and specialize.

Agentic AI increases the incentive for that approach because different workloads want different types of compute inside the same platform.

APX and AMX Keep the CPU Relevant to AI

Diamond Rapids also adds new Advanced Performance Extensions and enhanced Advanced Matrix Extensions.

AMX is specifically designed to accelerate matrix operations.

That means Intel is not treating the CPU only as a traffic controller.

Some AI work can remain directly on the CPU.

Smaller inference tasks.

Preprocessing.

Postprocessing.

Classical machine learning.

Vector and matrix operations embedded inside larger applications.

The architecture is trying to keep general-purpose compute useful even as specialized accelerators become more important.

Crescent Island Is About Inference Economics

Crescent Island attacks a different problem.

Inference cost.

Once a model is trained, the business problem becomes serving it repeatedly.

Agents can multiply inference demand because one user request may trigger several model calls.

A planning step.

A tool-selection step.

A verification step.

A second model.

A retry.

One agent workflow can consume far more tokens than a simple chatbot response.

Crescent Island is designed around that continuous serving workload.

480 GB of LPDDR5X Is the Headline Crescent Island Number

Intel says Crescent Island supports up to 480 GB of LPDDR5X memory.

That is a striking amount of memory for a PCIe accelerator.

Large memory capacity can allow bigger models to fit on one device.

It can also support longer context windows or more concurrent workloads.

Those are exactly the problems agent systems create.

One model may be large.

Several agents may need to run at once.

Each may carry substantial context.

Capacity becomes part of inference economics.

LPDDR5X Is an Unusual Choice for a Datacenter GPU

High-end AI accelerators often use HBM because it delivers very high memory bandwidth.

Crescent Island instead uses LPDDR5X.

That choice reflects a different target.

Intel is optimizing for lower power, large capacity and easier deployment rather than pursuing the highest possible accelerator class.

The trade-off is important.

Not every inference workload needs maximum bandwidth.

Some enterprises may value capacity, power and deployment simplicity more than peak performance.

Crescent Island is aimed at that middle ground.

350 Watts Keeps the GPU Inside Existing Air-Cooled Infrastructure

Intel lists Crescent Island as a 350-watt air-cooled PCIe card.

That matters because many enterprise datacenters are not designed for extreme accelerator power density.

A very high-power GPU may require liquid cooling or major rack redesign.

A 350-watt PCIe card can fit more naturally into existing air-cooled systems.

Intel is explicitly selling deployment economics.

The question is not only how fast the GPU is.

It is how much infrastructure has to change before an enterprise can use it.

32 Xe Cores and 256 XMX Engines Target Sustained Inference

Crescent Island uses 32 Xe cores and 256 XMX engines based on Xe3P.

XMX engines provide matrix acceleration for AI workloads.

Intel says the design is optimized for sustained inference performance and token throughput.

That distinction matters.

Training hardware is often judged by how quickly it can build a model.

Inference hardware is judged by how cheaply and reliably it can serve that model again and again.

Agentic AI pushes infrastructure toward the second problem.

Agents Make Concurrency a First-Class Requirement

A chatbot can serve one user request at a time.

An enterprise agent platform may run many agents simultaneously.

Each agent may generate several model calls.

That creates concurrency.

The accelerator has to keep many streams of work moving without wasting capacity.

Memory capacity, scheduling and token throughput all become important.

Crescent Island is designed around that environment.

The goal is not one spectacular benchmark run.

It is keeping many inference jobs moving efficiently across a shared accelerator.

Wildcat Lake Moves the Same Idea to the Edge

Wildcat Lake sits at the other end of the stack.

Intel launched it as Core Series 3 for price-sensitive laptops and intelligent edge platforms.

The chip combines x86 CPU cores, Xe3 graphics with XMX acceleration and an NPU rated at up to 17 TOPS.

That is far smaller than datacenter AI hardware.

It does not need to compete with it.

The point is to keep selected AI work close to the user or machine while larger work can move to cloud or enterprise infrastructure.

Two Performance Cores and Four Efficiency Cores Show the Target Market

Intel lists Wildcat Lake with two performance cores and four efficiency cores.

This is not a flagship workstation design.

It is meant to bring a right-sized AI platform to more affordable systems.

That matters strategically.

AI adoption becomes much larger when local acceleration stops being limited to premium devices.

A smaller chip can put hybrid AI into mainstream laptops, embedded systems and edge platforms where cost matters more than maximum compute.

The 17-TOPS NPU Is for Hybrid AI, Not Frontier Models

Intel says Wildcat Lake’s NPU provides up to 17 TOPS.

That is useful for supported local workloads.

Noise suppression.

Vision.

Small language models.

Local classification.

Background AI features.

It is not enough to treat the device as a replacement for datacenter inference.

Intel uses the phrase Hybrid AI for a reason.

Some tasks stay local.

Others move outward.

The edge becomes one tier in a larger system.

Wildcat Lake Is Intel’s First Processor to Use UCIe

Intel says Wildcat Lake marks the first use of UCIe in an Intel processor.

That is notable because UCIe is not only a datacenter technology.

Intel is bringing chiplet-style modular packaging into mainstream client hardware.

The company says the approach enables more cost-effective multi-chip package designs.

If that strategy scales, future processors can mix compute blocks more flexibly across product tiers instead of redesigning every chip from scratch.

The Same Packaging Idea Runs From Server to Client

This is one of the most coherent parts of Intel’s roadmap.

Diamond Rapids uses UCIe-S and advanced 3D packaging.

Wildcat Lake also uses UCIe.

The products are very different.

The packaging philosophy is shared.

Build systems from modular pieces.

Combine general-purpose compute with specialized acceleration.

Scale the architecture up or down depending on the market.

Agentic AI gives Intel a narrative that connects those pieces across the entire product line.

18A Connects the Three Products to Intel Foundry

Intel says the architectures are underpinned by its Foundry technology, including the Intel 18A process family.

Diamond Rapids uses the performance- and power-enhanced 18A-P variant.

Wildcat Lake is built on Intel 18A.

That gives the roadmap another strategic layer.

Intel is not only trying to sell processors.

It is trying to prove that its process technology and advanced packaging can support competitive AI hardware across servers, accelerators and client devices.

The AI roadmap doubles as a manufacturing roadmap.

This Is Intel’s Version of Heterogeneous AI

The most important word in Intel’s strategy is heterogeneous.

Different compute engines do different work.

CPU.

GPU.

NPU.

Real-time or edge processing.

Specialized matrix acceleration.

Open chiplet interconnect.

Fast memory and I/O.

The agentic AI story gives all of those components a role.

Instead of asking which chip wins, Intel is asking how many different processors can cooperate inside one workload.

That is a system-level argument.

Agentic AI Makes Orchestration More Expensive Than Chatbots

One chatbot prompt may create one inference job.

An agent can create a chain.

Plan.

Search.

Call a model.

Use a tool.

Read the result.

Call another model.

Verify.

Retry.

Log.

Escalate.

Each step creates compute and data movement.

At enterprise scale, the surrounding work can become substantial.

That is why Intel keeps emphasizing orchestration.

The more autonomous the workflow becomes, the more important the infrastructure around the model becomes.

Inference Cost Can Become the Real Scaling Limit

Frontier model training gets the headlines.

Enterprise AI spending often happens during inference.

Every production request consumes resources.

Agents multiply that consumption.

If a workflow triggers five or ten model calls instead of one, token economics change quickly.

Crescent Island is Intel’s answer to that problem.

The company is explicitly talking about token throughput, power and cooling rather than only peak compute.

That is a sign the market is moving from experimentation toward operational economics.

Edge Execution Reduces Latency and Data Movement

Wildcat Lake covers the other side of the cost problem.

Not every task should leave the device.

Sending audio, images or sensor data to the cloud creates bandwidth and latency.

A local NPU can handle selected tasks immediately.

The device can then send only the result or escalate the difficult work.

That can reduce network traffic and improve responsiveness.

Hybrid AI is therefore partly an economic architecture.

Use expensive centralized compute only when the workload needs it.

Intel Is Also Preparing an Agentic-AI Infrastructure Narrative Beyond Hot Chips

Intel’s August 26 AI Infra Summit preview extends the same strategy.

The company says agentic infrastructure is moving toward heterogeneous compute, disaggregated inference and intelligent orchestration.

It also highlights hybrid architectures that combine cloud-scale reasoning with edge-based inference and control.

That language is consistent with the Hot Chips hardware roadmap.

Diamond Rapids, Crescent Island and Wildcat Lake are not isolated product announcements.

They are components inside the same infrastructure thesis.

The Open-Standards Message Is Strategic

Intel repeatedly emphasizes UCIe and open infrastructure.

That is not accidental.

The AI accelerator market is dominated by tightly integrated hardware and software stacks.

Intel’s alternative argument is interoperability.

Open chiplet standards.

Heterogeneous compute.

Multiple accelerators.

Software that can span architectures.

Whether the ecosystem delivers that smoothly is another question.

But the strategic goal is clear: make openness part of the reason enterprises consider Intel hardware.

None of This Proves Intel Has Won the AI Hardware Race

A roadmap is not a market result.

Intel has announced specifications and architecture.

Customers still have to deploy the products.

Software has to mature.

Performance has to hold up under independent testing.

Crescent Island’s inference economics have to compete with established accelerators.

Diamond Rapids has to prove its server advantages in real workloads.

Wildcat Lake has to deliver useful local AI in cost-sensitive systems.

The Hot Chips announcement shows direction, not victory.

The Specifications Need to Stay Product-Specific

It is easy to combine the numbers into one misleading picture.

256 cores belongs to Diamond Rapids.

480 GB LPDDR5X and 350 watts belong to Crescent Island.

17 TOPS belongs to Wildcat Lake’s NPU.

They are not one chip.

They are three different product classes.

Keeping those boundaries clear matters because Intel’s whole argument depends on specialization.

The strategy only makes sense if each architecture is solving a different part of the agentic workload.

What Intel Has Actually Confirmed

Intel presented Diamond Rapids, Crescent Island and Wildcat Lake as complementary architectures for agentic AI.

Diamond Rapids is a next-generation Xeon design with up to 256 cores, 1.28 GB LLC, 16 memory channels at up to 12,800 MT/s, and 128 lanes of PCIe Gen 6 with CXL 3.0.

Crescent Island is an inference-oriented datacenter GPU with 32 Xe cores, 256 XMX engines, up to 480 GB LPDDR5X and a 350-watt air-cooled PCIe design.

Wildcat Lake is a Core Series 3 SoC with two performance cores, four efficiency cores, Xe3 graphics with XMX, and an NPU rated at up to 17 TOPS.

Intel also says the portfolio uses Intel 18A technologies, advanced Foveros packaging and UCIe.

What We Should Not Claim

We should not say Diamond Rapids alone runs the whole agentic AI stack.

Intel positions it as the orchestration and general-compute foundation.

We should not say Crescent Island is a training flagship.

Intel is positioning it for inference economics.

We should not compare 480 GB LPDDR5X directly with HBM accelerators without discussing different bandwidth and power trade-offs.

We should not say 17 TOPS means Wildcat Lake can run frontier models locally.

We should not say Intel has proven better performance than competing platforms without independent benchmarks.

And we should not treat Hot Chips specifications as evidence of broad production deployment today.

The Bigger Shift Is That AI Hardware Is Becoming a System, Not a Chip

The first AI hardware race was easy to describe.

Who has the fastest accelerator?

Agentic AI makes the answer more complicated.

Agents need CPU orchestration.

Accelerator inference.

Memory.

Networking.

I/O.

Local execution.

Cloud execution.

Security.

Scheduling.

The winning platform may not be the one with the largest single number.

It may be the one that moves work efficiently across the entire stack.

Intel’s Hot Chips roadmap is a bet on exactly that future.

Diamond Rapids handles the orchestration.

Crescent Island handles large-scale inference.

Wildcat Lake handles the edge.

Three architectures.

One agent workflow.

NVIDIA has agreed to acquire Hugging Face for $12.93 billion, one of the chipmaker’s largest deals and a major move beyond GPUs. Hugging Face is a critical distribution layer for open models, datasets, applications and developer tooling. NVIDIA says the platform will remain open, multi-cloud and multi-accelerator, and that NVIDIA hardware will not be required. The strategic question is whether NVIDIA is buying more than a company: it may be buying a direct route to the developers who decide which models, frameworks and infrastructure become standard.

NVIDIA Is Buying More Than an AI Website

NVIDIA has agreed to acquire Hugging Face for $12.93 billion.

The obvious reading is that the world’s dominant AI-chip company is buying the best-known platform for open AI models.

That is true.

It is also incomplete.

Hugging Face has become one of the places where developers discover models, compare them, download them, fine-tune them, test demos, publish datasets and decide which parts of the AI stack they want to use.

NVIDIA already owns a critical layer below that activity: compute.

Now it is moving toward a layer above it: distribution.

That is what makes this deal strategically important.

The company is not only trying to sell more GPUs.

It is moving closer to the point where developers decide what to run in the first place.

The Deal Is an Agreement — Not a Completed Acquisition Yet

The wording matters.

On September 3, 2026, NVIDIA announced that it had agreed to acquire Hugging Face for exactly $12,930,300,000.

That means the transaction has been announced and agreed.

It does not mean we should casually write as if Hugging Face has already been fully absorbed into NVIDIA’s operations.

Reuters reports that roughly $11.9 billion of the deal is for Hugging Face investors, with an equity-based retention program of up to $1 billion for employees who join NVIDIA.

For TUF, the safest language is simple:

NVIDIA has agreed to acquire Hugging Face.

Until the transaction is fully completed, that distinction should remain.

Why Hugging Face Is Worth So Much to NVIDIA

Hugging Face is not just a model-hosting site anymore.

According to NVIDIA’s announcement, the platform is used by more than 18 million developers, researchers and creators.

It hosts more than 3 million models.

More than 500,000 datasets.

Around 1 million applications.

More than 200,000 companies use the platform to discover, evaluate, customize and deploy AI.

Those numbers explain the logic of the acquisition more clearly than the price tag.

Hugging Face sits where models become usable.

That makes it a distribution layer, a discovery layer and increasingly an application layer for open AI.

The AI Stack Is Becoming More Vertical

Modern AI is often described as a stack.

At the bottom are chips and systems.

Above them are networking, training and inference.

Then come model repositories, developer tooling and applications.

NVIDIA has already expanded far beyond the chip itself.

CUDA tied software development closely to NVIDIA GPUs.

DGX turned the company into a systems vendor.

Networking expanded its role inside AI infrastructure.

Inference software and cloud services moved it further up the stack.

Hugging Face would extend that vertical reach again.

The important point is not that NVIDIA suddenly owns every AI layer.

It does not.

The point is that it is becoming present in more of them.

Hugging Face Is Where Model Choice Happens

Developers do not always begin an AI project by choosing a chip.

They often begin by choosing a model.

Which one is small enough?

Which one has the right license?

Which one works in my language?

Which one has the right benchmark profile?

Which one has a good ecosystem?

Hugging Face is one of the first places many developers go to answer those questions.

That gives the platform influence over downstream infrastructure.

Once a model is selected, the developer begins asking how to fine-tune it, serve it and scale it.

That is where NVIDIA’s infrastructure becomes relevant.

Distribution Can Be More Valuable Than Direct Control

NVIDIA does not need every model on Hugging Face to be its own.

It does not need every workload to run on NVIDIA hardware.

It can still benefit if the platform becomes the default place where AI builders begin.

Distribution creates optionality.

If developers use Hugging Face to discover open models, NVIDIA can surface optimized inference paths.

It can integrate libraries.

It can improve deployment tooling.

It can make NVIDIA-accelerated workflows easier.

The strategic value comes from proximity to developers, not only ownership of content.

NVIDIA Is Promising the Platform Will Stay Open

Jensen Huang addressed the biggest concern directly in NVIDIA’s announcement.

He said Hugging Face will remain an open platform for the entire AI ecosystem.

Developers will still be able to choose the models they want.

The frameworks they want.

The clouds and inference providers they want.

The compute platforms they want.

NVIDIA also says NVIDIA compute will not be required to build on or deploy through Hugging Face.

That commitment is central to the deal.

Hugging Face’s value depends heavily on being useful across the industry rather than being perceived as one vendor’s storefront.

The Neutrality Question Will Not Disappear Because of a Promise

The harder question is not whether NVIDIA allows rival hardware.

The harder question is whether developers continue to see the platform as neutral.

A platform can technically support AMD, Intel, Google TPUs or other accelerators while still giving one ecosystem better optimization, documentation, placement or integration.

Reuters reports that some developers and analysts are already concerned about whether rival hardware could gradually receive less attention.

NVIDIA says the opposite.

The platform will remain multi-accelerator and multi-cloud.

The gap between those positions will be measured by what happens after the transaction, not by launch-day statements.

Open Models Are Strategically Useful to NVIDIA

Open models create a different market structure from closed AI APIs.

A company can download an open-weight model.

Run it locally.

Fine-tune it.

Deploy it on its own infrastructure.

Change the serving stack.

Move between vendors.

That flexibility creates more infrastructure decisions.

And infrastructure decisions create more opportunities for NVIDIA.

Closed API products hide more of the compute layer from the user.

Open models expose it.

That makes open AI strategically compatible with NVIDIA’s business.

NVIDIA Was Already Deep Inside Hugging Face

The acquisition does not come from nowhere.

NVIDIA says it is already the largest contributor of open models and data on Hugging Face.

The company says it has published more than 500 models and more than 250 open datasets on the platform.

Its Nemotron work is part of that strategy.

NVIDIA has increasingly positioned open-weight AI as a way for enterprises and institutions to retain more control over deployment.

Buying Hugging Face takes that strategy from participation to ownership of the platform itself.

CUDA Built a Developer Moat Below the Model Layer

NVIDIA’s greatest software advantage historically was not a model repository.

It was CUDA.

CUDA made NVIDIA GPUs programmable for general-purpose parallel computing and helped build a large software ecosystem around the company’s hardware.

That created switching costs.

Hugging Face could create a different kind of developer relationship.

CUDA sits close to hardware.

Hugging Face sits close to model selection and application development.

If NVIDIA can connect those layers without damaging platform neutrality, the company gains influence at both ends of the developer workflow.

The Deal Could Make Deployment Much Easier

Hugging Face already connects model discovery to deployment.

NVIDIA already provides optimized inference stacks.

The natural integration path is obvious.

A developer finds a model.

Checks the license.

Tests it.

Selects an optimized runtime.

Deploys it.

Monitors it.

Scales it.

That flow could become much smoother under common ownership.

The upside for developers is reduced friction.

The risk is that the easiest path gradually becomes the NVIDIA path even when other options technically remain available.

Convenience can shape ecosystems more strongly than explicit exclusivity.

Inference Is Becoming the Bigger Battlefield

Training frontier models attracts attention because the clusters are enormous.

But deployed AI systems run inference continuously.

Every generated token.

Every image.

Every embedding.

Every agent action.

Every local model call.

As the number of AI applications grows, inference becomes a huge infrastructure market.

Hugging Face gives NVIDIA a closer connection to that layer.

The platform is where millions of builders already move from model discovery toward deployment.

That could make the acquisition strategically useful even if it never produces direct revenue at the scale of NVIDIA’s GPU business.

The Acquisition Also Diversifies NVIDIA’s Customer Access

One risk for NVIDIA is concentration.

The largest AI companies buy enormous amounts of compute.

Some of those same companies are developing custom accelerators to reduce dependence on NVIDIA.

Reuters highlights this directly, citing companies including Meta, OpenAI and Microsoft.

Hugging Face gives NVIDIA a more direct route to a much wider developer base.

Instead of relying only on a relatively small number of hyperscale buyers, NVIDIA can strengthen its relationship with startups, enterprises, researchers and independent developers.

Hugging Face Is Also a Dataset and Application Platform

It would be a mistake to reduce Hugging Face to model downloads.

The platform hosts hundreds of thousands of datasets.

Spaces lets developers publish interactive AI applications and demos.

Its software ecosystem helped standardize how many developers work with modern models.

Hugging Face also expanded into robotics through LeRobot and the acquisition of Pollen Robotics.

That means NVIDIA is buying access to multiple AI workflows at once.

Models.

Data.

Apps.

Libraries.

Evaluation.

Deployment.

Robotics.

The strategic surface is much wider than a simple repository.

The Robotics Angle Is Easy to Miss

Hugging Face has been moving into open robotics.

Its LeRobot ecosystem grew rapidly.

It acquired Pollen Robotics.

It began offering Reachy 2.

That overlaps naturally with NVIDIA’s own robotics strategy around Jetson, Isaac and GR00T.

The acquisition therefore connects not only software AI but potentially physical AI as well.

Open robot models and datasets hosted on Hugging Face can ultimately feed demand for training, simulation and edge inference.

Again, the logic returns to the same pattern.

Distribution creates infrastructure demand.

The Price Reflects Strategic Value, Not Just Current Revenue

Reuters notes that Hugging Face was valued at $4.5 billion in its last disclosed funding round in 2023.

The new agreement is worth $12.93 billion.

That is a large jump.

The explanation is unlikely to be current revenue alone.

NVIDIA is paying for strategic position.

Developer reach.

Open-model distribution.

Community.

Tooling.

Data.

The ability to sit closer to where AI projects begin.

This does not prove the price is cheap or expensive.

That would be an investment judgment.

The useful point is that the acquisition price makes more sense when viewed as control of a strategic layer rather than a conventional SaaS purchase.

The Biggest Risk Is Damaging What Makes Hugging Face Valuable

The acquisition contains an obvious paradox.

NVIDIA gains value from owning Hugging Face because Hugging Face is broadly trusted and broadly used.

If ownership causes developers to leave, the value falls.

If rival hardware vendors stop investing in integrations, the ecosystem narrows.

If model builders decide another repository feels more neutral, distribution fragments.

NVIDIA therefore has a strong economic reason to preserve openness.

That does not remove conflicts of interest.

It makes managing them part of the product strategy.

Open Source Can Be Forked — Platforms Are Harder

Open-source code can often be forked.

A platform ecosystem is harder.

You can copy a repository.

You cannot instantly copy millions of users.

Download counts.

Discussion histories.

Model cards.

Community trust.

Datasets.

Spaces.

Brand recognition.

Network effects.

That is why platform ownership matters even in an open ecosystem.

The underlying models may remain downloadable.

But discovery, reputation and distribution still concentrate value.

NVIDIA is buying those network effects.

Hugging Face Could Become the Front Door to NVIDIA Infrastructure

The strongest strategic interpretation is straightforward.

A developer opens Hugging Face.

Finds a model.

Tests it.

Fine-tunes it.

Deploys it.

Behind the scenes, NVIDIA provides the easiest optimized path for each step.

That does not require lock-in.

It only requires default convenience.

If NVIDIA can make its infrastructure the path of least resistance while preserving credible alternatives, it can gain usage without forcing exclusivity.

That is a more subtle strategy than simply blocking competitors.

But NVIDIA Does Not Automatically Own Open AI

The title deliberately says NVIDIA wants the open AI layer.

It does not say NVIDIA now owns open AI.

Open models come from many organizations.

Meta.

Mistral.

DeepSeek.

Qwen.

Google.

Independent labs.

Universities.

Startups.

Developers can host models elsewhere.

Cloud providers have their own catalogs.

Open-source libraries can move.

The acquisition increases NVIDIA’s influence.

It does not convert an open ecosystem into one company’s property.

The Deal Could Pressure Rival Infrastructure Vendors

AMD, Intel, Google and cloud providers will be watching integration decisions closely.

If Hugging Face remains equally strong across accelerators, the ecosystem may continue normally.

If NVIDIA-specific optimizations move faster, competitors may need to invest more heavily in their own developer tooling and distribution channels.

That could accelerate competition around inference software rather than only raw silicon.

The next AI platform war may be fought as much through model hubs, developer tools and deployment pipelines as through chip benchmark charts.

Open AI Is Becoming an Infrastructure Market

Open models began partly as a research and community movement.

They are increasingly becoming enterprise infrastructure.

Companies want models they can customize.

Governments want control over deployment.

Organizations want to run AI in private environments.

Developers want smaller models that can run locally.

That creates demand for optimized compute at every scale.

NVIDIA’s Hugging Face acquisition is a bet that open AI will not shrink the infrastructure market.

It may expand it.

What NVIDIA Has Actually Confirmed

NVIDIA has confirmed that it agreed to acquire Hugging Face for $12,930,300,000.

The company says Hugging Face has more than 18 million users across developers, researchers and creators.

It says the platform hosts more than 3 million models, 500,000 datasets and 1 million applications and is used by more than 200,000 companies.

NVIDIA says Hugging Face will remain open.

It says developers will retain freedom to choose models, frameworks, clouds, inference providers and compute platforms.

It says NVIDIA hardware will not be required.

NVIDIA also says Hugging Face will continue supporting open-source and open-weight models across the ecosystem and remain multi-cloud and multi-accelerator.

What We Should Not Claim Yet

We should not say the acquisition is already fully completed unless NVIDIA later confirms closing.

We should not say NVIDIA hardware is required on Hugging Face.

NVIDIA explicitly says the opposite.

We should not say Hugging Face will stop supporting rival chips.

There is no current evidence for that.

We should not say NVIDIA now owns open-source AI.

It does not.

We should not claim the transaction guarantees more GPU sales.

That is strategic analysis, not an announced outcome.

And we should not assume the platform’s neutrality will remain unchanged forever.

That is something the ecosystem will have to observe.

The Bigger Story Is Where NVIDIA Wants to Sit in the AI Workflow

For years, NVIDIA’s most important position was obvious.

The GPU.

Then the company expanded into systems, networking, software, cloud services, inference and robotics.

Hugging Face pushes it closer to the developer’s first decision.

Which model should I use?

That is strategically powerful.

If NVIDIA can remain underneath the workload as compute and also stand near the top as the platform where developers discover and deploy models, it gains influence across a much larger portion of the AI stack.

The acquisition is not just about Hugging Face.

It is about NVIDIA deciding that the future of AI infrastructure begins before the first GPU is selected.

GoPro has entered a definitive merger agreement involving Starman Optical, a U.S. optical-photonics company whose Starman New Photonics business is building domestic manufacturing for high-speed optical transceivers. If the transaction closes, the combined company intends to keep GoPro’s consumer products while expanding into AI data-center infrastructure, government, defense, robotics and aerospace. The strategic connection is not that camera lenses and data-center transceivers are the same product. It is that both sit inside a broader optics and photonics stack in which light is used either to capture information or to move it.

The Strange Part Is Not That GoPro Is Merging — It Is What GoPro Wants to Become

GoPro built its identity around one simple idea.

Put a small camera where a normal camera is inconvenient.

On a helmet.

On a bike.

On a surfboard.

On a drone.

That business made GoPro synonymous with action cameras.

Now the company is trying to broaden the definition of what it is.

On September 1, 2026, GoPro announced that it had entered into a definitive merger agreement involving Starman Optical, a privately held U.S. optical-photonics company.

The proposed transaction is not framed only as a financial recapitalization.

GoPro says it is intended to reposition the company across consumer, commercial and defense markets, while adding Starman’s U.S.-made optical transceiver business and opening a route into AI data-center infrastructure.

That creates an unusual strategic arc.

A company famous for capturing light is trying to become part of the infrastructure that moves information with light.

The Deal Is Signed, but It Is Not Closed

The first thing to keep straight is the transaction status.

GoPro and Starman have entered into a definitive merger agreement.

That is more advanced than an exploratory partnership or a non-binding announcement.

But the transaction has not yet closed.

GoPro says the deal is expected to close by the end of 2026, subject to regulatory approvals, customary closing conditions and approval by GoPro stockholders.

Until those conditions are satisfied, Starman’s optical-transceiver business has not simply become a finished GoPro division.

That distinction matters because the most ambitious parts of the strategy are still forward-looking.

The merger creates a plan.

Execution comes later.

GoPro Says the Camera Business Is Staying

The proposed strategy is not to abandon GoPro cameras.

The company says it intends to continue supporting its existing consumer products, subscription business and cloud platform.

That means the transformation is additive rather than a clean exit from consumer electronics.

The proposed combined company would have at least two very different operating stories.

One remains familiar: cameras, imaging, software and consumer products.

The other is infrastructure: optical transceivers, domestic manufacturing and photonics for high-speed computing networks.

Those businesses serve different customers and have different product cycles.

The strategic question is whether the underlying optics, imaging, engineering and intellectual-property base is broad enough to justify putting them under one company.

Starman Is Bringing a Different Kind of Optical Product

Starman Optical describes itself as an optical-photonics company focused on optical transceivers and related technologies through its Starman New Photonics business.

An optical transceiver does not take pictures.

It sits at the edge of a fiber link and converts information between electrical and optical forms.

Inside a server or network switch, data begins as electrical signals.

To send that data efficiently across optical fiber, the transceiver modulates light with the information.

At the other end, another transceiver receives the light and converts it back into electrical signals that the electronics can process.

Cisco describes these modules as the devices at the ends of fiber links that convert electrical signals to optical signals and back again.

That small component is one of the places where electronics and photonics meet.

A Camera and a Transceiver Use Light for Opposite Information Problems

The connection between GoPro and Starman becomes clearer if the word optics is not treated as one product category.

A camera uses light to learn about the outside world.

Light enters through the lens.

The optical system focuses it.

An image sensor converts the incoming photons into electrical information.

An optical transceiver solves almost the reverse information problem.

The data already exists electronically.

The transceiver converts that electrical information into modulated light so it can travel through fiber.

One system turns light into data.

The other turns data into light and back again.

They are not interchangeable technologies.

But both depend on controlling, transmitting, receiving and detecting light with very high precision.

Why AI Data Centers Need So Many Optical Links

An AI data center is not one giant processor.

Large training and inference systems distribute work across many accelerators, servers and switches.

Those devices continuously exchange model parameters, activations and other data.

The faster the compute becomes, the more pressure moves onto the network connecting it.

A GPU that is waiting for data is expensive hardware sitting idle.

That is why the network fabric has become part of AI-system performance.

NVIDIA’s data-center documentation treats optical transceivers as a normal part of high-performance cluster cabling, with modules connecting electrical interfaces to optical fiber.

The physical network therefore becomes part of the scaling problem.

More compute does not help if the surrounding interconnect cannot feed it.

Copper Is Excellent — Until Distance and Signal Loss Become the Problem

Electrical connections remain extremely useful inside data centers.

Short copper links can be cheap, simple and power-efficient.

But higher data rates make electrical signaling harder as distance increases.

Signals lose energy as they travel through traces, connectors and cables.

Compensating for that loss requires more complex electronics and can increase power.

NVIDIA has highlighted this issue in its work on co-packaged optics, noting that high-speed electrical paths between a switch ASIC and an external optical module can accumulate substantial signal loss.

Fiber avoids many of those electrical-distance limitations.

The signal travels as light through glass rather than as a high-speed electrical waveform through copper.

That does not make optics free.

The optical modules themselves consume power and add cost.

But they let high-bandwidth connections extend farther while keeping the cable physically small.

The Transceiver Is the Boundary Between Silicon Electronics and Fiber

A useful way to visualize the system is to follow one direction of a link.

A processor sends electrical data to a network interface.

The switch or network device routes that data toward a port.

At the optical module, electronics drive a light source and encode the information onto that light.

The signal enters fiber.

At the far end, a photodetector converts the received light back into an electrical signal.

The receiving electronics recover the data.

Cisco’s photonics documentation describes laser diodes, photodiodes and optical waveguides as core elements used to generate, guide, modulate and detect the light inside optical transceiver systems.

The transceiver therefore lives exactly at the boundary GoPro is trying to enter.

It is where high-speed digital electronics stop being purely electrical.

AI Infrastructure Is Turning Optics Into a Volume Problem

Optical networking is not new.

Telecommunications networks have used fiber for decades.

What is changing is the density and scale of high-performance computing.

AI clusters can require large numbers of high-bandwidth links across accelerators, switches and racks.

Cisco notes that the growing volume of data-center optical modules has helped drive silicon photonics deeper into high-speed networking.

For a manufacturer, that makes the opportunity very different from selling a small number of specialized scientific optical devices.

The target becomes repeatable production of large numbers of modules with controlled performance, reliability and cost.

Starman’s pitch is therefore not just about inventing an optical component.

It is about manufacturing capacity.

Starman New Photonics Is Building a U.S. Manufacturing Footprint

The manufacturing plan predates the GoPro merger announcement.

In June 2026, the New Jersey Economic Development Authority approved the first award under its Next New Jersey Manufacturing Program for Starman New Photonics.

NJEDA says the company plans to invest $150 million in a 100,000-square-foot facility in Warren, New Jersey.

The project is expected to create 250 jobs.

NJEDA describes the business as building a domestic supply of high-speed optical transceivers that are essential for the AI industry.

That description is important because it shows the merger is being attached to a physical manufacturing project, not only to a slide about future markets.

If the project is executed as planned, the optical side of the proposed GoPro strategy would include actual U.S. production capacity.

This Is Also a Supply-Chain Story

GoPro and Starman repeatedly frame the transaction through domestic manufacturing.

The companies argue that important optical and imaging hardware is still manufactured heavily outside the United States.

The proposed combined company wants to use domestic optical-transceiver production as part of its positioning in AI infrastructure, government and defense.

That is a different strategic logic from the action-camera business.

Consumer electronics companies often optimize around global manufacturing networks and cost.

Strategic infrastructure customers may care much more about where a component is made, how the supply chain is controlled and whether production can meet domestic sourcing requirements.

In that environment, manufacturing location becomes a product attribute.

GoPro Is Bringing More Than a Brand Name

The merger announcement emphasizes GoPro’s imaging experience and intellectual property.

The company says it has built a portfolio of more than 2,500 U.S. patents over roughly 24 years.

That does not mean those patents can simply be applied to an optical transceiver.

Imaging optics and data-communications photonics solve different engineering problems.

But a large imaging organization accumulates capabilities beyond one camera model.

Optical design.

Mechanical packaging.

Thermal constraints.

Sensor integration.

Firmware.

Manufacturing tolerances.

Reliability testing.

Miniaturization.

High-volume consumer production.

The proposed strategy assumes some of those capabilities can support expansion into adjacent optical and imaging markets even when the end product is very different.

The Patent Count Sounds Impressive, but Patent Count Alone Does Not Prove Technical Fit

More than 2,500 U.S. patents is a large portfolio.

It is also easy to overread that number.

A patent count does not tell us how many patents are still strategically important.

It does not tell us which patents apply to transceivers.

It does not tell us whether those patents create a competitive advantage in AI networking.

And it does not automatically turn camera intellectual property into photonics intellectual property.

The useful interpretation is narrower.

GoPro has a long history of building compact optical and imaging products and owns a substantial body of intellectual property around that work.

The merger intends to find more markets where that technical base has value.

Whether the portfolio produces meaningful cross-business advantages will only become clear after the combined strategy is implemented.

The Robotics Connection Is More Direct Than the Data-Center Connection

GoPro also says the combined company intends to expand its optics and imaging capabilities into robotics.

That connection is easier to understand.

Robots need cameras.

They need compact optical systems.

They need image sensors, rugged housings, calibration and low-latency video pipelines.

Some robots also need stereo vision, wide-angle imaging or cameras that can survive motion, vibration and outdoor environments.

Those are closer to the engineering problems GoPro already understands.

The data-center transceiver business is different.

Robotics is an imaging adjacency.

Optical networking is a photonics adjacency.

The proposed company is trying to pursue both under a broader identity built around light.

Defense and Aerospace Add Another Reason to Think in Terms of Optical Systems

GoPro and Starman also identify defense, government and aerospace as target markets.

Those sectors use cameras and imaging systems, but they also buy communications, sensing and optical hardware under different requirements from ordinary consumer products.

Reliability can matter more than retail design.

Domestic manufacturing can matter more.

Qualification cycles can be longer.

Volumes can be lower but unit value higher.

The merger announcement does not yet provide a detailed product roadmap for those markets.

So it would be premature to claim that GoPro is about to produce a specific defense or aerospace system.

What is confirmed is the strategic direction.

The company wants to stop defining its addressable market around consumer cameras alone.

The Merger Is Also a Recapitalization

The technology story sits inside a financial restructuring.

GoPro says the proposed transaction would include a cash payment to existing shareholders, repayment of the company’s outstanding debt at closing and continued ownership for existing shareholders in a smaller portion of the combined public company.

Those mechanics matter because the strategy requires capital.

Expanding a consumer brand into manufacturing-heavy optical infrastructure is not a cheap experiment.

Factories, equipment, qualification, inventory and engineering all consume cash before they produce scale.

The company describes the merger as a way to strengthen the balance sheet and invest in a broader product roadmap.

This article is not evaluating the transaction as an investment.

The relevant point is that the technical expansion and the recapitalization are part of the same plan.

The Optical-Transceiver Market Will Not Behave Like the Camera Market

GoPro knows how to launch a consumer product.

Optical infrastructure has a different rhythm.

A camera can win because of image quality, stabilization, software, usability and brand.

A data-center optical module has to meet electrical, optical, thermal and interoperability requirements inside a larger network architecture.

Enterprise buyers care about qualification and reliability.

A failure can take down a high-value link.

Compatibility matters.

Power per bit matters.

Reach matters.

Module form factor matters.

Standards matter.

The sales channel is different too.

GoPro cannot simply place a new transceiver beside a HERO camera in retail and expect the business to scale.

The proposed company will have to operate as both a consumer brand and an infrastructure supplier.

AI Networking Is Moving Beyond the Traditional Pluggable Module Too

There is another complication.

The optical industry itself is changing.

Traditional pluggable transceivers sit at the front of a switch.

High-speed electrical signals travel from the switch silicon across the board to the pluggable module, where they are converted to light.

At very high data rates, that electrical path becomes increasingly expensive in power and signal integrity.

That is why companies including NVIDIA are pushing co-packaged optics, where the optical engines move much closer to the switching silicon.

This does not make pluggable transceivers obsolete tomorrow.

Pluggables remain widely deployed and operationally convenient.

But it means Starman would be entering a market whose architecture is evolving quickly.

Winning requires following where the optical boundary moves next.

A Successful Strategy Would Turn GoPro Into a Portfolio of Light-Control Technologies

The most coherent version of the merger strategy is not “GoPro starts selling networking gear.”

It is broader.

Consumer cameras use optics to capture scenes.

Robotics uses optics to help machines perceive.

Aerospace and defense use imaging and optical systems under demanding physical conditions.

Data centers use photonics to move information between electronic systems.

Those businesses do not share one customer.

They do not share one product.

But they share a technical theme.

Light is being manipulated to carry information.

If the company can build credible products across those markets, GoPro becomes less of a camera company and more of an optics-and-imaging platform.

That is the identity the merger announcement is trying to create.

What Has Actually Been Confirmed

Several parts of the story are confirmed today.

GoPro entered into a definitive merger agreement on September 1, 2026.

The transaction is expected to close by year-end if the required approvals and conditions are satisfied.

GoPro says it intends to continue supporting its existing consumer products, subscription business and cloud platform.

The proposed combined company intends to add Starman’s U.S.-made optical transceivers and expand into AI data-center infrastructure, government, defense, robotics and aerospace.

GoPro says it has a portfolio of more than 2,500 U.S. patents.

NJEDA says Starman New Photonics plans a $150 million investment in a 100,000-square-foot New Jersey manufacturing facility expected to create 250 jobs.

And authoritative networking documentation confirms the basic technical role of optical transceivers: converting high-speed electrical data into optical signals for fiber and converting received light back into electrical data.

What We Should Not Claim Yet

This article does not claim the merger has closed.

It does not claim Starman’s transceiver business already belongs to GoPro.

It does not claim GoPro camera patents are directly applicable to optical networking.

It does not claim the New Jersey facility is already operating at the planned production scale.

It does not claim GoPro has announced a finished AI-data-center product roadmap.

It does not claim pluggable optical transceivers will remain the dominant architecture indefinitely.

It does not claim the combined company will successfully compete with established optical-networking suppliers.

And it does not treat the transaction as investment advice.

The confirmed story is a strategic repositioning attempt.

The outcome still depends on closing the deal, building the manufacturing capability, shipping products and winning customers.

The Bigger Upgrade Is the Definition of the Company

Companies are often trapped by the product that made them famous.

A successful product becomes a category.

The category becomes the brand.

Then the brand becomes the boundary of what customers, investors and even employees think the company is allowed to build.

GoPro is trying to redraw that boundary.

The action camera is still part of the plan.

But the proposed merger asks whether the deeper asset is not the camera itself.

Maybe it is decades of work around optics, imaging, packaging and light.

Starman brings a second use of light: moving digital information through fiber at high speed.

One business captures the world.

The other connects computers.

If the merger closes and the strategy works, GoPro’s most important upgrade will not be another camera specification.

It will be changing the answer to a much larger question.

What kind of company is GoPro?

Closed-Loop Liquid Cooling Is Becoming Part of the AI Compute Stack

AI infrastructure is usually described through processors, memory and networks.

Cooling is now moving into the same architectural conversation.

Meta’s newest AI-optimized data centers use direct-to-chip closed-loop liquid cooling as part of the facility design. The coolant moves heat away from high-density compute hardware, passes through heat exchangers and then returns to the racks in a continuous loop.

That makes the cooling system part of the compute platform rather than a separate room-level utility.

A GPU rack has a power envelope.

That power becomes heat.

The cold plates, coolant flow, pumps, heat exchangers and external heat-rejection system have to be sized around that heat before the rack can operate at its intended density.

Meta described this shift in detail on August 27, 2026.

The company uses a water-and-glycol coolant mixture in closed loops and says the same coolant can remain in service for up to a decade.

The important change is architectural.

Compute density now influences plumbing design.

Plumbing design influences rack layout.

Rack layout influences the building.

Cooling has become another layer of the AI stack.

Power Density Connects Compute Design Directly to Thermal Design

Every watt consumed by a processor eventually becomes heat that has to leave the system.

That relationship becomes more visible as AI racks concentrate larger amounts of compute into a small physical area.

Meta’s infrastructure team previously described a six-rack pod in which two compute racks contained 72 NVIDIA Blackwell GPUs and consumed about 140 kilowatts.

The surrounding pod used four Air-Assisted Liquid Cooling racks to support that deployment in a traditional data-center environment.

The example shows how thermal architecture follows compute density.

The GPUs define the workload capacity.

Power delivery supplies the electrical energy.

The cooling system removes the resulting heat.

Networking keeps the accelerators connected.

All four systems have to fit into the same physical design.

That is why cooling is becoming a first-order infrastructure parameter.

A data-center team planning a new AI system cannot choose compute density first and treat heat removal as an unrelated decision later.

The two are connected from the beginning.

Direct-to-Chip Cooling Moves the Thermal Path Closer to the Processor

Direct-to-chip liquid cooling shortens the thermal path.

A cold plate sits directly on a high-power component such as a CPU or GPU.

Heat moves from the silicon package into the cold plate.

Liquid flowing through channels inside that plate carries the heat away.

The coolant then moves through the rack-level or facility-level loop toward a heat exchanger.

This is different from cooling only the surrounding room air.

The liquid interacts with the component through a dedicated thermal interface.

The Open Compute Project’s Cold Plate Sub-Project describes direct liquid cooling as an ecosystem extending from the cold plate through the technology cooling system and coolant distribution unit.

That framing is useful because the cold plate is only the first piece.

Tubing, quick disconnects, manifolds, pumps, sensors, filtration, coolant chemistry and heat exchangers all become part of the same thermal path.

The cooling architecture therefore begins on the processor and continues through the rack and the facility.

Meta Uses a Water-and-Glycol Coolant in a Sealed Loop

Meta describes its current closed-loop design as circulating a mixture of water and glycol.

The coolant repeatedly travels through the compute equipment, absorbs heat and then moves toward the heat-exchange stage.

After the heat is transferred away, the cooled liquid returns to the racks.

The same fluid circulates again.

That is the meaning of the closed loop.

The coolant is not continuously consumed as part of the ordinary thermal cycle.

Meta says it expects the coolant used in these loops to remain in service for up to ten years before replacement.

The glycol component gives the fluid system properties suitable for long-term thermal operation, including protection across environmental conditions.

The exact coolant formulation and facility design can vary.

The important architectural feature is recirculation.

The thermal loop carries energy away from the chips while the working fluid remains inside the system.

Heat Exchangers Separate the Compute Loop From Heat Rejection

The coolant has to release the heat it collected from the processors.

That happens through heat exchangers.

A heat exchanger allows thermal energy to move from one fluid or system to another without requiring the two circuits to become one shared loop.

This creates a useful boundary.

The technology cooling system can circulate coolant through the racks.

The facility side can then move that heat toward dry coolers or another site-specific heat-rejection system.

Meta’s typical new data-center design uses direct-to-chip closed-loop cooling with dry coolers where local conditions support that approach.

The coolant therefore handles the internal transport.

The heat exchanger hands the thermal load to the facility.

The dry cooler rejects the heat to the outside environment.

This layered structure keeps the processor-side thermal loop and the larger building system connected while allowing each part to perform a different job.

Dry Coolers Let the Facility Reject Heat Without Consuming the Rack Coolant

Meta says many of its closed-loop sites use dry coolers.

A dry cooler moves outdoor air across heat-exchange surfaces to remove energy from the closed liquid system.

The rack coolant stays inside the loop.

The outside air does not mix with it.

Meta’s water-stewardship documentation says its typical direct-to-chip closed-loop design with dry coolers has no operational water use in the cooling system itself, with site water use limited to other facility needs such as domestic use, cleaning and fire protection.

That statement applies to the specific design Meta describes.

Other data centers can use different cooling architectures depending on climate, site resources and equipment.

The important point for the AI compute stack is that heat rejection can be designed as a separate facility stage.

The chips transfer heat to coolant.

The coolant transfers heat through the exchanger.

The facility rejects that heat outdoors.

Each layer has its own interface.

Cold Plates Are Becoming a Standardized Hardware Interface

As direct-to-chip cooling becomes more common, the industry is also standardizing the components that connect to the processors.

The Open Compute Project’s Cold Plate Sub-Project is developing specifications and guidelines for direct liquid cooling.

Its work covers cold plates, connectors, hose routing, coolant loops and coolant distribution systems.

The project describes its goal as enabling an open ecosystem for direct liquid cooling.

That matters because cooling hardware is beginning to look more like the rest of data-center infrastructure.

Interfaces can be specified.

Flow requirements can be documented.

Connector designs can be standardized.

Coolant compatibility can be tested.

A server vendor, cooling supplier and data-center operator can then design around a shared set of engineering expectations.

Thermal infrastructure becomes a platform with interfaces, not only custom plumbing built independently for each deployment.

Coolant Distribution Units Manage Flow Between the Facility and the Rack

A Coolant Distribution Unit, or CDU, is another important layer.

The CDU manages the liquid flow serving the IT equipment.

Depending on the design, it can contain pumps, heat exchangers, sensors, filtration and control hardware.

The facility side supplies one thermal condition.

The CDU creates the controlled loop needed by the rack.

Rows of server racks inside a data center
High-density compute has to coordinate rack space, power, networking and thermal infrastructure. The pictured racks are a generic data-center environment and are not identified as Meta AI racks.

The rack manifold then distributes coolant to individual servers or cold plates.

This gives operators another boundary in the cooling architecture.

Facility water or facility heat-rejection equipment does not need to connect directly to every processor cold plate.

The CDU can isolate and manage the technology cooling system in between.

Open Compute Project specifications treat the CDU as part of the direct-liquid-cooling ecosystem.

That makes it similar to a power-distribution layer.

Electric power moves through conversion and distribution stages before it reaches the processor.

Cooling can now move through heat-exchange and distribution stages before it reaches the same processor.

Air-Assisted Liquid Cooling Lets Liquid-Cooled Hardware Enter Existing Facilities

Not every data center was originally built with facility liquid loops.

Meta uses Air-Assisted Liquid Cooling, or AALC, for some of those environments.

AALC places pumps and heat exchangers near the racks.

The liquid loop cools the high-density hardware.

The AALC system then transfers the heat into the facility’s existing air-based environment.

This creates a bridge between newer liquid-cooled equipment and buildings designed around a different thermal architecture.

Meta used AALC with its Blackwell deployment in traditional data centers and has described it as a smaller, distributed version of the closed-loop concept.

The core thermal idea remains the same.

Liquid collects heat close to the processors.

A heat exchanger moves that energy into the next cooling stage.

The difference is where that transfer happens.

This lets cooling architecture evolve in stages instead of requiring every facility to use exactly the same building-level loop.

Cooling Capacity Influences How Much Compute Fits Into a Rack

Thermal design also changes physical density.

Meta says cooling the same high-density AI hardware entirely through larger air-handling structures could require substantially more tray space in the example it describes.

Liquid cooling moves much of the thermal transport into cold plates and tubing.

That can leave more of the rack volume available for compute, memory, networking and power hardware.

The relationship is simple.

A rack has finite dimensions.

Every fan, heat sink, duct, manifold and cable consumes space.

The cooling method changes how that space is allocated.

As AI hardware concentrates more electrical power per rack, thermal hardware becomes part of the density equation.

Compute density is therefore not only a semiconductor metric.

It is a mechanical and facility metric too.

The number of accelerators that can operate in one rack depends on whether the rack can receive enough power and remove enough heat.

Cooling Loops Are Designed for Long Service Life

A data-center cooling loop is infrastructure, not a disposable accessory.

It has to operate continuously across long deployment cycles.

Meta says the water-and-glycol mixture in its closed-loop system is expected to remain in service for up to a decade.

Open Compute Project work on direct-to-chip cooling also includes coolant chemistry, corrosion, material compatibility, filtration, connectors and long-term loop requirements.

Those subjects matter because the cooling fluid touches metals, seals, hoses and heat-exchange surfaces across the system.

The thermal performance of the loop has to remain predictable over time.

The mechanical interfaces also have to support maintenance and replacement of IT equipment.

This creates a lifecycle engineering problem.

A GPU generation may change quickly.

The cooling infrastructure around it is expected to support several equipment cycles.

That gives thermal standards another role: help make the rack and facility useful across successive generations of compute hardware.

Cooling Operations Are Becoming Sensor-Driven

A closed-loop system is also a control system.

Operators can measure coolant temperature.

Flow rate.

Pressure.

Rack inlet and outlet conditions.

Heat-exchanger behavior.

Pump state.

Server load.

Outside weather.

Those measurements can be used to change cooling operation dynamically.

The workload is not constant.

Training jobs can move.

Inference traffic can rise and fall.

Outdoor temperature changes through the day and across seasons.

The cooling system therefore has to respond to changing thermal demand.

This makes data-center cooling increasingly similar to other automated infrastructure.

Sensors provide state.

Controllers choose operating points.

The facility changes fan speeds, pump behavior or other control settings.

Cooling becomes software-addressable infrastructure built around a physical thermal loop.

Meta Is Using Reinforcement Learning to Optimize Cooling Control

Meta is also applying AI to the operation of its cooling systems.

The company says its engineering teams built a physics-based simulator that models weather conditions, server load and cooling-equipment behavior.

A reinforcement-learning system can test control decisions inside that simulated environment before policies are applied to the real facility.

Meta says the approach has been scaled to air-cooled data centers in its fleet.

In one pilot, Meta reports that the reinforcement-learning approach reduced supply-fan energy consumption by an average of 20 percent while reducing water usage by 4 percent across varying weather conditions.

Those are Meta’s pilot results for the specific facility and control system it tested.

The architectural point is broader.

Cooling is no longer only a fixed mechanical design.

Its operating policy can also be optimized by software.

The AI compute stack produces heat.

AI can then participate in controlling the infrastructure that removes that heat.

Liquid Cooling Is Becoming an Industry Standards Problem

The move toward higher-density liquid cooling is larger than one data-center operator.

The Open Compute Project and ASHRAE formed an alliance in 2025 focused on liquid-cooling standards and best practices for AI data centers.

Their joint scope spans facility water systems, technology cooling systems, direct-to-chip equipment, immersion cooling and coolant distribution units.

That is a sign that thermal management is becoming an interoperable infrastructure discipline.

Facilities need terminology.

Equipment classes.

Temperature ranges.

Flow requirements.

Connector definitions.

Fluid guidance.

Testing methods.

As those elements become standardized, AI hardware vendors and data-center operators can design around shared thermal interfaces.

Cooling starts to resemble power and networking.

It becomes a layer that needs formal specifications because many independent systems have to connect to it.

New AI Data Centers Are Being Designed Around Liquid Cooling From the Beginning

Meta’s infrastructure plans show the difference between adaptation and native design.

Traditional facilities can use Air-Assisted Liquid Cooling to support newer equipment.

New AI-optimized facilities can include direct-to-chip closed loops in the original building architecture.

That changes planning.

Pipe routing is designed with the data hall.

Heat exchangers are sized with the compute plan.

Dry coolers are selected with the local climate.

Rack density is coordinated with power delivery.

Mechanical systems are designed around the expected thermal load.

Meta says it chooses cooling technologies based on local conditions, including climate, resource availability and technical requirements.

The result is not one identical cooling design for every site.

It is a facility architecture where cooling is chosen alongside compute rather than added after the compute platform is finalized.

The AI Compute Stack Now Extends From Silicon to Heat Rejection

The modern AI stack can be traced through several physical layers.

The model runs on accelerators.

Memory feeds the accelerators.

Networks connect them.

Power systems deliver electrical energy.

Cold plates collect the resulting heat.

Coolant carries that heat through the rack.

CDUs and heat exchangers move it into the facility system.

Dry coolers or other site-specific infrastructure reject it to the environment.

Sensors measure the system.

Control software adjusts operation.

Meta’s closed-loop deployments show how those layers now have to be designed together.

The cooling system is not separate from compute capacity.

It helps define the compute capacity that can operate in a given rack and building.

That is why closed-loop liquid cooling is becoming part of the AI compute stack.

The processor creates the work.

The thermal system makes sustained operation possible.

The two are now designed as one infrastructure problem.

That is the upgrade.

Jalapeño Moves OpenAI From Models and Serving Software Into Silicon

OpenAI has spent years working above the chip.

Models.

Inference kernels.

Serving software.

APIs.

Products such as ChatGPT and Codex.

Jalapeño adds another layer underneath them.

OpenAI and Broadcom unveiled Jalapeño in June 2026 as OpenAI’s first custom inference processor. On August 25, OpenAI published its first measured performance results from engineering hardware running public language models.

The chip is not presented as a general consumer processor.

It was designed around large-language-model inference.

That distinction defines the project.

Training builds or updates model weights.

Inference uses those trained weights to answer requests.

Every ChatGPT response, API completion or agent step becomes an inference workload somewhere in the serving infrastructure.

OpenAI is now designing hardware specifically around that workload.

The company describes Jalapeño as the first generation of a multigenerational compute platform built with Broadcom and other infrastructure partners.

The shift is architectural.

OpenAI is no longer optimizing only the model that runs on the machine.

It is also designing part of the machine around the model-serving process.

Inference Has Several Phases With Different Bottlenecks

One reason to design custom inference hardware is that serving a language model is not one uniform operation.

OpenAI separates the workload into phases.

Prefill processes the user’s prompt and existing context.

Decode generates the response token by token.

Those phases stress the system differently.

OpenAI describes prefill as more compute-intensive.

Decode depends more heavily on memory bandwidth because the system repeatedly accesses model state while producing each next token.

Communication becomes another part of the workload when tensors, model state or cached information need to move between cores or accelerators.

The hardware can therefore spend time computing, moving data or waiting for another part of the system.

Jalapeño was designed around those transitions.

Instead of optimizing one isolated arithmetic peak, OpenAI says it designed compute, memory, networking and software together around the full inference request.

The goal is to keep the useful work moving through the system.

That makes inference performance a systems problem.

The chip matters.

The memory matters.

The network matters.

The serving software deciding where every piece of work goes matters too.

KV Cache Placement Becomes a Hardware Design Problem

The KV cache is one of the clearest examples of software behavior turning into hardware architecture.

During autoregressive generation, a transformer reuses information from earlier tokens instead of recomputing everything from the beginning for every new token.

That reusable state is stored in the key-value cache.

Long conversations and agent sessions can make that state substantial.

Where the cache lives affects how far the data must travel and how quickly it can be reused.

OpenAI says Jalapeño allows model state, including the KV cache, to be explicitly placed and kept local while the system activates the required combination of compute, memory and networking.

That is a hardware-software decision.

The serving layer knows what model state exists.

The hardware exposes a structure that lets the system place that state deliberately.

The network connects the pieces that need to communicate.

The architecture is therefore shaped around a pattern created by language models themselves.

The cache is no longer just an implementation detail inside inference software.

It becomes part of how the accelerator system is organized.

The Network Is Part of the Accelerator Architecture

A custom inference chip does not operate alone.

Large models can span many accelerators.

Requests can move through several devices.

Expert models may need to route work to different parts of the system.

Cached state can be distributed.

Jalapeño treats networking as part of the architecture rather than an external connection added afterward.

OpenAI says the chip was designed to minimize data movement and communication delays and that its network is integral to the system.

Broadcom contributes networking technology, including Tomahawk networking silicon, to the wider platform.

That matters because inference speed depends on more than tensor arithmetic.

A processing unit can only continue when the data it needs is available.

If the system can keep model state closer to the compute that will use it and move information efficiently when communication is necessary, more of the request can stay active.

OpenAI calls the resulting connected area a large domain in which the workload can remain inside one coordinated system.

The chip and network therefore form one serving architecture.

The accelerator does the computation.

The network helps make that computation available across the larger model.

OpenAI Designed Jalapeño Around Both Prefill and Decode

Some inference systems can be organized around separate resources for different phases.

Jalapeño takes another route.

OpenAI describes the accelerator as balanced and fungible across prefill and decode.

The same architecture is intended to handle both phases while adapting to the changing mix of compute, memory and communication requirements.

That becomes relevant for interactive agents.

An agent can receive context.

Generate a short output.

Call a tool.

Receive new information.

Generate again.

Repeat the sequence many times.

The workload moves through prefill and decode repeatedly rather than following one long static pattern.

OpenAI says this changing balance was part of the design target.

The hardware therefore reflects the behavior of the software being served.

A conventional chat completion is one inference pattern.

A long, multi-step agent session creates another.

The processor is designed around the idea that the mix can change during real use.

That is an example of custom silicon being shaped by product workload instead of being designed independently from it.

OpenAI Tested the Chip on Three Public Model Families

The August results were not limited to one OpenAI model.

OpenAI tested Jalapeño with GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.

That matters because a chip designed by a model company could otherwise be interpreted as hardware for one internal model family.

OpenAI says Jalapeño was designed to support current and future language models across the industry.

The three public tests give the company a way to measure that claim on architectures developed outside OpenAI as well as its own open-weight model.

The models also differ in scale and serving behavior.

That gives the benchmark more than one workload.

OpenAI reports that the chip remained on the throughput-per-power and latency frontier across the tested operating range for all three.

Those are OpenAI’s reported results from the InferenceX benchmark environment.

The larger architectural point does not depend on one comparison.

Jalapeño is being programmed as a general language-model inference target.

The hardware is custom.

The model support is intended to remain broader than one model.

The Public Results Measure Throughput and Latency Together

Inference performance can be described in several ways.

Tokens per second.

Tokens per user.

Time between generated tokens.

End-to-end request latency.

Throughput per watt.

Maximum total serving capacity.

A system can look different depending on which measurement is chosen.

OpenAI says it evaluated Jalapeño at matched user experience using InferenceX, a public benchmark from SemiAnalysis.

The company measured how much AI work the system could complete per unit of power while meeting latency requirements.

Across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5, OpenAI reports 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems used in those tests.

For highly interactive operating points, OpenAI reports 2.1 to 4.1 times higher performance.

Those figures belong to the specific benchmark configurations OpenAI published.

They are useful because they show what the company is optimizing.

Not only peak arithmetic.

Not only one-user latency.

The design target is the combination of serving volume, response time and power.

Power Efficiency Is Becoming a Serving Metric

Jalapeño also makes power part of the performance discussion.

OpenAI rates the processor package at 700 watts.

The company says measured sustained power remained at or below 550 watts on the workloads used in its August tests.

OpenAI then normalizes benchmark results using published chip power ratings for the compared accelerators.

That produces throughput-per-kilowatt figures.

The reason is practical.

AI infrastructure is limited by more than the number of chips a company can buy.

A data center has electrical capacity.

Cooling capacity.

Rack limits.

Network capacity.

The same amount of available power can support different amounts of useful inference depending on the complete system.

A custom accelerator can therefore be evaluated by how much model-serving work fits inside a power envelope.

This is especially relevant when the operator is also the company serving the model.

The hardware decision eventually reaches the product as response capacity.

More useful work per kilowatt means the same electrical infrastructure can process more inference requests.

That is the connection between chip architecture and service scale.

The Chip Was Co-Designed With Broadcom and Celestica

Custom silicon does not mean one company manufactures every layer itself.

OpenAI designed Jalapeño’s architecture around its model and serving requirements.

Broadcom provides silicon implementation and networking expertise.

Celestica contributes board, rack and system integration work.

OpenAI describes the collaboration as a multi-generation platform rather than one isolated processor.

Close-up photograph of a microprocessor package mounted on a circuit board
A processor only becomes useful as part of a board, memory, networking, power and software system. The pictured microchip is a generic component and is not Jalapeño.

That division of work is important.

A complete accelerator program requires more than a block diagram.

The design has to become physical silicon.

The chip needs packaging.

Boards.

Power delivery.

Memory.

Networking.

Rack integration.

Production systems.

Software.

Deployment tooling.

The Jalapeño project connects OpenAI’s model and serving knowledge to companies that specialize in turning those requirements into production infrastructure.

This is how a model developer can move into custom hardware without becoming every supplier in the semiconductor chain.

The architecture starts closer to the workload.

Partners industrialize the system around it.

AI Models Were Used During the Chip Development Cycle

AI was also used to build the processor that will run AI.

OpenAI says earlier generations of its models assisted engineers during Jalapeño design and bring-up.

The company reports moving from initial design to manufacturing tapeout in nine months for the chip-development phase it describes.

AI was used to explore implementations, shorten design and verification loops and optimize arithmetic circuits.

That creates a feedback loop.

Models run on hardware.

Model behavior tells engineers where the serving system spends time.

AI tools help engineers design and verify a new accelerator.

The new accelerator is then used to serve later models.

This does not remove semiconductor engineering from the process.

It changes some of the tools inside it.

The same class of software being optimized for deployment becomes part of the engineering workflow used to build the deployment hardware.

Jalapeño therefore represents two forms of hardware-software co-design.

The chip is designed around AI workloads.

AI is also used during parts of the chip-design process.

The Programming Model Was Designed for Humans and AI

OpenAI also designed the software interface with AI-assisted programming in mind.

The company describes Jalapeño as a predictable programming target based on local tensors, explicit communication and predictable synchronization.

Those properties make the hardware mapping problem more structured.

An engineer can describe the work.

An AI system can help decide where that work should be placed, scheduled and coordinated across the accelerator system.

This matters because new model families still require new kernels and model-specific optimization.

The hardware does not automatically run every architecture at maximum efficiency.

The software stack has to adapt.

OpenAI reports using Codex with GPT-Astra to bring GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 to high performance on Jalapeño within two months even though those models were not part of the original production plan.

For selected GPT-OSS attention and mixture-of-experts blocks, OpenAI says AI-generated implementations ran 1.5 to 1.8 times faster than the existing expert-written implementations.

OpenAI explicitly limits those numbers to selected blocks rather than the complete model.

The important architectural point is the programming loop.

Custom hardware and AI-assisted kernel development are being designed together.

Serving Software Becomes Part of the Silicon Advantage

A custom accelerator only becomes useful when the serving stack can keep it busy.

Model weights have to be loaded.

Requests have to be scheduled.

KV cache has to be placed.

Communication has to be coordinated.

Kernels have to match the model architecture.

Batching and interactive traffic have to share the system.

That software layer is one reason OpenAI describes Jalapeño as a full-stack project.

The company operates ChatGPT, Codex and the API.

Those products generate real serving patterns.

OpenAI can observe how the workloads behave and use those observations when designing the chip and its runtime.

Then the software can be adjusted around capabilities added to the hardware.

The information moves both directions.

Product workload informs infrastructure.

Infrastructure changes what the product can serve.

That feedback loop is different from purchasing a processor whose architecture was designed independently from one company’s specific deployment patterns.

Jalapeño makes the inference operator part of the silicon-design process.

Custom Silicon Adds Another Compute Source to OpenAI's Infrastructure

Jalapeño adds another compute source to OpenAI’s infrastructure strategy.

OpenAI says it will continue widely deploying accelerators from NVIDIA and other partners for both training and inference.

The custom processor therefore sits beside external hardware in the wider compute fleet.

That is consistent with the scale of modern AI infrastructure.

Training and inference use different workload mixes.

Individual model families can benefit from different accelerator characteristics.

Capacity has to grow across several suppliers and data-center environments.

Jalapeño gives OpenAI a first-party architecture it can shape around its own serving requirements while partner accelerators continue supplying substantial compute capacity.

This makes the custom-chip strategy easier to understand.

It is another layer of infrastructure optimization.

Some workloads can run on OpenAI-designed silicon.

Others can continue running on external accelerator platforms.

The serving system can grow with both.

The First Generation Is Planned for Deployment by the End of 2026

The August results come from engineering hardware and pre-deployment qualification.

OpenAI says it plans to begin deploying Jalapeño within its compute infrastructure by the end of 2026.

Before that scale-up, the company says it is continuing production qualification, software maturation and validation across additional models.

That status matters.

The published benchmark results show working first-party silicon.

They do not mean the chip is already carrying the full production load of OpenAI’s services.

The project is moving from development into deployment.

Broadcom and OpenAI have described the platform as intended for large-scale data-center rollout across multiple generations.

The infrastructure transition will therefore happen over time.

First silicon.

Benchmarks.

Production qualification.

Software maturity.

Rack integration.

Deployment.

Then later generations.

That sequence is how a custom accelerator becomes part of an operating AI service rather than remaining a laboratory prototype.

Gen 2 and Gen 3 Turn Jalapeño Into a Roadmap

OpenAI already describes Jalapeño as the first generation of a longer processor roadmap.

In the August update, the company said Gen 2 was deep in development and Gen 3 was taking shape.

That changes the meaning of the first chip.

A one-off ASIC can optimize one moment in model architecture.

A multigenerational platform can learn from each deployment cycle.

The first generation reveals how real workloads behave on first-party silicon.

The software team learns which kernels need improvement.

The hardware team sees where data movement occurs.

The serving team sees which workloads map well to the architecture.

Those observations can feed the next design.

At the same time, model architecture continues changing.

Context lengths grow.

Agent workflows become more iterative.

Mixture-of-experts routing changes communication patterns.

New numerical formats alter memory and arithmetic requirements.

A recurring silicon roadmap gives OpenAI a way to incorporate those changes into later hardware.

The processor becomes part of the model roadmap instead of a separate infrastructure purchase.

Jalapeño Shows Why AI Inference Is Moving Toward Full-Stack Design

Jalapeño is useful as a case study because it connects parts of AI infrastructure that are often discussed separately.

The model defines the computation.

Inference software organizes the request.

The KV cache creates persistent state.

Memory bandwidth moves that state.

The network connects accelerators.

The scheduler decides where work goes.

Power limits determine how much of the system fits inside a data center.

The chip architecture has to support all of those layers.

OpenAI designed Jalapeño around that complete serving path.

Broadcom and Celestica help turn the architecture into deployable systems.

OpenAI models help program and optimize the hardware.

The same hardware is tested on several public model families.

The first measured results are now available, while production deployment is planned to begin later in 2026.

That makes custom silicon more than a chip story.

It is an inference-stack story.

As AI products depend on larger volumes of interactive model serving, the companies operating those products have more reason to optimize below the software layer.

Jalapeño is OpenAI’s first step into that layer.

That is the upgrade.

AI Is Adding a New Layer Between the App and the Hardware

For most of computing history, the operating system sat between applications and hardware.

An app asked for memory. The operating system managed memory.

An app wanted a file. The operating system exposed a file system.

An app needed graphics. The operating system provided graphics APIs and drivers.

AI is beginning to add another layer to that relationship.

An application may need a language model, an image model, speech recognition, semantic search or a custom ONNX model. That model may run on the NPU, GPU or CPU. It may already be distributed by the operating-system vendor. It may be downloaded locally. It may run in a private cloud environment. It may come from another model provider.

The app increasingly does not have to manage every one of those pieces by itself.

Windows, Apple platforms and Android are all building system-level AI frameworks that sit between the application and the underlying model execution path.

The exact architecture differs by platform.

The direction is similar.

The operating system is becoming part of how an application reaches AI.

Model Selection Is Becoming an Application Architecture Decision

An AI feature no longer has to begin with one fixed model endpoint.

A developer can start with the task.

Does the feature need short text generation on the device?

Does it need a larger context window?

Does it need a custom model trained for one domain?

Does it need to work offline?

Does it need a cloud model for a larger workload?

Those questions can determine where inference happens and which model is used.

Microsoft documents this directly in its Windows AI guidance. Windows applications can combine Windows AI APIs, Foundry Local, Windows ML and cloud AI services in the same product.

Apple’s Foundation Models framework now exposes a common LanguageModel protocol that can represent Apple Foundation Models, Private Cloud Compute models and other providers that conform to the protocol.

Android’s AI guidance similarly separates on-device Gemini Nano, custom local models and cloud Gemini options, and its Agent Development Kit can combine local and cloud models in the same multi-agent system.

The model becomes one component inside the application architecture.

The operating system provides more of the machinery around that choice.

Windows Now Exposes Several AI Paths Inside One Platform

Windows provides a clear example because Microsoft now documents several AI layers under one Windows AI platform.

Windows AI APIs expose ready-to-use capabilities such as language models, OCR, semantic search, imaging and other built-in features on supported hardware.

Foundry Local provides local language and speech models through a runtime designed for on-device use.

Windows ML gives developers a way to bring their own ONNX models and run them locally.

Cloud APIs remain another path when an application is designed to use remote AI services.

Microsoft explicitly says these options can be combined inside the same application.

That changes the way a Windows AI feature can be designed.

The developer can treat Windows AI as a set of execution layers instead of treating every model as a separate infrastructure project.

One feature might call a built-in Windows AI API.

Another might use a local open model through Foundry Local.

A third might use a custom ONNX model through Windows ML.

Another workflow can connect to cloud AI.

The application can choose the path that matches the task.

Windows AI APIs Can Route Supported Workloads to Local Accelerators

The routing concept becomes more literal when hardware enters the picture.

Modern PCs can contain several compute engines.

The CPU remains the general-purpose processor.

The GPU provides highly parallel compute.

The NPU provides dedicated neural-network acceleration on supported systems.

Microsoft’s current Windows AI guidance says supported Windows AI APIs can route inference through the NPU automatically on Copilot+ PCs, while some APIs can also use GPU or CPU paths on other supported Windows 11 hardware.

That means the application can call a platform API without directly implementing every hardware-specific inference path itself.

The operating system and runtime know more about the machine underneath.

They can expose a higher-level capability to the app.

This is the same pattern operating systems have used for graphics, audio and networking for years.

The application asks for a capability.

The platform handles more of the device-specific execution underneath.

AI is moving into that model.

Windows ML Can Select Execution Providers Across CPU, GPU and NPU

Windows ML goes deeper into hardware-aware inference.

It uses ONNX Runtime and supports execution providers that map model execution to different processors.

Microsoft documents providers for CPU, GPU and NPU acceleration, including hardware-specific providers from AMD, Intel, NVIDIA and Qualcomm.

Some execution providers can be dynamically downloaded through Windows ML and maintained through the Windows platform rather than being bundled independently inside every application.

The framework can also use device policies or explicit developer selection to choose an execution provider.

Diagram showing applications, system calls, kernel, device drivers and hardware in an operating system architecture
A simplified operating-system architecture shows the platform layer between applications and hardware. AI runtimes are adding model and inference services to this same platform role.

That makes the operating system part of the model-to-hardware path.

The model itself can remain an ONNX model.

The execution layer decides which compatible processor and provider will run it.

This separates the AI workload from some of the hardware plumbing beneath it.

For developers, the same model can participate in a Windows execution stack that understands CPU, GPU and NPU options.

For the operating system, AI inference becomes another workload that can be mapped onto the hardware available in the machine.

Foundry Local Adds Model Selection Above the Hardware Layer

Foundry Local adds another level to the Windows stack.

Instead of requiring the application to package one specific local model implementation, the runtime can expose models through aliases and a local API.

Microsoft says Foundry Local detects available hardware and can serve a hardware-optimized model variant for the device.

Its current Windows documentation describes support across Qualcomm NPU paths, DirectX 12 GPUs, NVIDIA CUDA and CPU execution depending on the model and hardware configuration.

The application can therefore ask for a model by the interface provided by the runtime while Foundry Local handles more of the relationship between the model package and the machine.

This is another form of routing.

At one layer, the app chooses a local model family.

At another layer, the runtime selects the hardware-compatible execution path.

The application can then keep the same higher-level code across several hardware configurations.

That is the kind of abstraction operating systems are designed to provide.

Apple Is Building a Common Model Interface Into Its Developer Stack

Apple is approaching the same architectural idea through the Foundation Models framework.

At WWDC26, Apple expanded the framework so applications can work with multiple language-model sources through a common LanguageModel protocol.

Apple’s developer documentation says that can include the on-device Apple Foundation Model, the Apple model running through Private Cloud Compute and other providers such as Claude or Gemini when they conform through the framework.

The important part is the shared interface.

The application can build around a language-model abstraction rather than designing every feature around one provider-specific call shape.

Apple also added Dynamic Profiles that can swap models, tools and instructions during a continuous session.

That moves model choice closer to runtime application behavior.

One task can use one model configuration.

Another task can use another.

The surrounding application can keep the same framework structure.

The model becomes replaceable inside a larger session architecture.

Private Cloud Compute Extends the Same Apple Session Beyond the Device

Apple’s on-device and server-side models show how one application framework can span two compute locations.

The SystemLanguageModel runs on the device.

PrivateCloudComputeLanguageModel runs through Apple’s Private Cloud Compute infrastructure.

Apple documents both through the Foundation Models framework and the LanguageModel protocol.

Its current documentation lists a 4K context size for the on-device model and a 32K context size for the Private Cloud Compute model, with additional reasoning capability on the server-side option.

The application can create a LanguageModelSession with either model type while retaining the same broader session API, tools and instructions.

That is a direct example of model routing at the application-framework level.

The developer decides which execution target fits the feature.

The framework keeps the interaction model consistent.

The location of the model can change without requiring the entire application architecture to change with it.

The model endpoint becomes one parameter inside the session.

Core AI Adds a Bring-Your-Own-Model Path on Apple Silicon

Apple is also adding a lower-level path for developers who want to run their own models locally.

At WWDC26, Apple introduced Core AI as a framework built into the operating system for running AI models on Apple Silicon.

Apple describes Core AI as a way to load, specialize and run models on-device through a native Swift API.

That gives the platform two different model layers.

Foundation Models provides access to Apple models and provider abstractions for language-model sessions.

Core AI provides a path for custom on-device models.

The combination is similar to what is happening on Windows.

There is a high-level model service for common AI capabilities.

There is also a lower-level runtime for custom models.

Both sit inside the operating-system developer stack.

The app can choose how much of the model management it wants the platform to handle.

Android Uses AICore as a System Service for Gemini Nano

Android places the operating-system layer directly between applications and its on-device foundation model.

Gemini Nano runs through AICore, an Android system service.

Google says AICore manages model distribution, future model updates, safety functions and the use of on-device hardware acceleration.

Applications can access Gemini Nano through ML Kit GenAI APIs instead of independently packaging the foundation model and its runtime.

That changes the deployment model for on-device AI.

The application does not have to treat a large model file as ordinary app content.

The operating system can provide the model as a shared system capability.

The same platform layer can manage updates and connect inference to supported hardware.

Google’s current documentation describes Gemini Nano as running through AICore for tasks including summarization, rewriting, image description, speech recognition and custom prompting through ML Kit interfaces.

Android is therefore turning the foundation model into an operating-system service that applications can call.

Android Can Combine On-Device and Cloud Models in One Agent System

Android’s agent framework extends the model-selection idea beyond one model at a time.

Google’s Agent Development Kit for Android supports on-device Gemini Nano through ML Kit and cloud Gemini models through cloud integrations.

Its documentation also describes a hybrid multi-agent pattern where a cloud model can act as the root orchestrator while on-device Gemini Nano sub-agents handle selected tasks locally.

That is a different kind of routing.

The decision can happen at the agent level.

One part of the system can use cloud compute.

Another part can run on the phone.

The application can organize those models as cooperating agents inside one workflow.

This matters because future AI applications may not have one universal model call.

They may contain several model roles.

The operating system and its AI frameworks provide the runtime environment in which those roles can be assigned.

The Router Is Also Becoming a Model-Management Layer

Routing is not only about choosing local or cloud.

It is also about managing the model once it becomes part of the device.

Android AICore manages Gemini Nano distribution and updates.

Windows can manage shared ONNX Runtime components and dynamically acquired execution providers through Windows ML.

Foundry Local can manage model catalogs and hardware-optimized variants.

Apple provides system models directly through Foundation Models and adds Core AI for custom on-device execution.

These are different implementations, but they move the same category of work upward into the platform.

The application can depend on an operating-system AI service instead of independently rebuilding distribution, runtime selection, hardware mapping and update logic for every feature.

That gives AI a more conventional place inside software architecture.

The model starts to look less like a separate product bolted onto an app.

It starts to look like a compute resource exposed through the platform.

Applications Can Start Choosing Models by Task Instead of by Brand

A common model interface changes how developers can think about application design.

The first question can become: what does this task need?

A short offline summarization feature may fit an on-device model.

A long document workflow may use a server model with a larger context window.

A specialized vision feature may use a custom local model.

A background classification task may run through a built-in AI API.

A multi-agent workflow may divide work between local and cloud models.

Windows, Apple platforms and Android now all expose pieces of that architecture.

The details remain platform-specific, and the developer still defines the product logic.

But model identity is becoming easier to separate from feature identity.

The feature can be designed around a capability.

The runtime can then connect that capability to the model and compute path selected for the task.

That is the practical meaning of the operating system becoming a model router.

The Operating System Is Becoming Part of the AI Runtime

The operating system has always decided how applications reach hardware and shared system services.

AI is becoming another part of that responsibility.

Windows can expose built-in models, local open models, custom ONNX models, execution providers and cloud paths inside one developer platform.

Apple can expose an on-device foundation model, a Private Cloud Compute model, third-party language-model providers and custom Core AI models through its developer stack.

Android can expose Gemini Nano through AICore, custom local models through its AI toolchain and cloud models through hybrid application architectures.

The common idea is not that one operating system automatically chooses every model for every application.

The common idea is that model access, execution and hardware mapping are moving into platform APIs that applications can build around.

That is what turns the operating system into an AI model router.

The application defines the task.

The platform provides more ways to connect that task to the model and compute path that will run it.

That is the upgrade.