How Custom XPUs Power Scalable AI Factories

Peter Bubenik · Nvidia Research · · Source
Image for How XPUs Meet a World-Class AI Factory

After studying this material, you should be able to:

Explain how NVLink Fusion enables custom XPUs to integrate into a complete, production-ready AI factory infrastructure — and why this matters for performance, time-to-market, and operational risk management.

Specifically, you will:

  • Define what an AI factory is and how its economics work
  • Explain what XPUs are and the challenges of building them
  • Describe what NVLink Fusion provides and how it solves those challenges
  • Identify the four key pillars of a complete AI factory platform

Step-by-Step Study Material

Step 1: Understand What an AI Factory Is

Before anything else, you need to understand the context — what problem are we solving?

What is an AI Factory?

An AI factory is not a traditional data center. Think of it like a manufacturing plant, but instead of producing physical goods, it produces intelligence — specifically, AI-generated outputs called tokens.

How is its success measured?

MetricWhat it means
Tokens per secondHow fast it produces output
Tokens per wattHow energy-efficient it is
Cost per tokenHow economically viable it is
Utilization & uptimeHow continuously it operates

Key Insight

An AI factory must run continuously. Every second of downtime or inefficiency directly increases cost per token and reduces competitiveness.

Think of it this way: A car factory that stops its assembly line every hour is far less profitable than one that runs 24/7 without interruption. AI factories work the same way.


Step 2: Understand What XPUs Are and Why They're Challenging

What is an XPU?

An XPU (Accelerated Processing Unit) is a custom-designed chip built by hyperscalers (like Amazon, Google, Meta) or AI-native companies to accelerate specific AI workloads.

  • They are alternatives or complements to standard GPUs
  • Companies build them to optimize for their specific workloads
  • Examples include custom ASICs designed for inference, training, or specialized model architectures

The Core Problem with Custom XPUs

Building a chip is only part of the challenge. To actually deploy an XPU in production, a company must also build:

Custom XPU
    ↓
+ High-speed CPU interfaces
+ Scale-up networking solution
+ Compute and switch trays
+ Rack architecture (cooling + power)
+ Security and storage integration
+ Supplier ecosystem management
    ↓
= Complete AI Factory

Most teams underestimate this. The chip might take 2 years to design, but the surrounding infrastructure can take just as long — and cost just as much.

The Result Without a Solution

  • Slow time-to-market — infrastructure delays chip deployment
  • High cost — building everything from scratch is expensive
  • High risk — unproven infrastructure fails in production

Step 3: Understand the Three Dimensions of Scale-Up Networking

Modern AI workloads — like trillion-parameter models, mixture-of-experts architectures, and agentic AI — require chips to communicate with each other at extreme speeds.

This communication is called scale-up networking (connecting chips within a single system or rack).

If scale-up networking is too slow:

Slow Network → Chips sit idle waiting for data
             → Utilization drops
             → Cost per token rises
             → AI factory becomes uneconomical

Three Requirements for Good Scale-Up Networking

1. Delivered Performance

  • End-to-end network speed
  • In-network compute capability
  • Mature software integration

2. Factory Resiliency

  • High uptime
  • Continuous health monitoring and telemetry
  • Ability to service components while the factory keeps running

3. Platform Maturity

  • Proven at large scale
  • Demonstrated return on investment
  • Reduced operational risk

Remember: A networking solution that looks good on paper but has never been deployed at scale is a risk, not an asset.


Step 4: Learn What NVLink Fusion Is and What It Delivers

Definition

NVLink Fusion is NVIDIA's solution that allows custom XPUs to plug into NVIDIA's existing, world-class AI infrastructure — rather than building that infrastructure from scratch.

Think of it as a bridge between a company's custom chip and a proven, production-ready platform.

What NVLink Fusion Provides

A. Scale-Up Networking via 6th-Generation NVLink

ComparisonNVLink FusionOff-the-shelf Ethernet
End-to-end latencyBaseline3x higher (worse)
Packet rateBaseline10x lower (worse)
  • Supports up to 72 XPUs in a single domain today
  • Roadmap extends to 1,152 accelerators with co-packaged optics

B. CPU Connectivity via NVLink-C2C

  • Connects XPUs to NVIDIA Vera CPUs or other ecosystem CPUs
  • Delivers 6x better energy efficiency compared to PCIe interfaces
  • Removes barriers between control (CPU) and compute (XPU) — critical for agentic AI systems

C. Ecosystem for Development and Deployment

NVLink Fusion is supported by partners across:

  • ASIC design
  • CPU architecture
  • IP and optical interconnect

This means companies don't need to find and validate every supplier themselves.

D. MGX Rack-Scale Architecture

  • Shared rack design with NVIDIA's own systems (like Vera Rubin NVL72)
  • 100% liquid cooling — no fans, cables, or hoses
  • Trays can be removed for service while the rest of the rack keeps running
  • Manufacturing partners manage design and integration

Step 5: Understand Risk Management Through Infrastructure Standardization

The Planning Problem

AI factory planning starts before the chips are ready:

Timeline:
[Power procurement] → [Facility design] → [Cooling design] → [Rack layout] → [Chip arrives]
        ↑
   This starts YEARS before silicon is finalized

If your entire data center is locked to one specific chip, and that chip is delayed or changed, your entire facility plan is at risk.

How NVLink Fusion Solves This

NVLink Fusion creates a unified architecture where:

  • XPU-based systems and GPU-based systems share the same rack footprint
  • They share the same networking, cooling, power delivery, and management systems
  • Operators can build out the facility first, then decide the exact chip mix later
  • Capacity can be reprovisioned as workload needs change

Practical Example

A hyperscaler can deploy NVIDIA GPU racks today using NVLink Fusion infrastructure, while their custom XPU is still in development. When the XPU is ready, it slots into the same rack design — no facility rework needed.

As MediaTek's VP explained:

"They can deploy their rack-level solution with the NVIDIA GPU, and then decouple the development of their XPU and put it at a different pace."


Step 6: Understand the Software Layer That Completes the Factory

Hardware alone doesn't make a factory run. Software coordinates everything.

Key Software Components

SoftwarePurpose
NVIDIA NCCLDistributed workload communication between chips
NVIDIA Dynamo & NIXLDisaggregation — splitting workloads across systems
NVIDIA Mission ControlCluster management, telemetry, and debugging
NVIDIA Omniverse DSX BlueprintDigital twin for modeling the entire factory before building it

Why the Digital Twin Matters

Factory buildout is expensive. Mistakes require costly rework.

The DSX AI Factory Blueprint lets partners simulate the entire facility — buildings, power, cooling, compute, and networking — before a single rack is installed. This catches design errors early, when they're cheap to fix.


Step 7: Synthesize — The Complete Picture

Now let's put it all together with a mental model:

PROBLEM:
Custom XPU alone ≠ AI Factory
Building everything from scratch = slow, expensive, risky

SOLUTION: NVLink Fusion

┌─────────────────────────────────────────────────────┐
│                  AI FACTORY                         │
│                                                     │
│  [Custom XPU] ←→ [NVLink Scale-Up Networking]      │
│       ↕                    ↕                        │
│  [NVLink-C2C CPU]    [MGX Rack Architecture]        │
│       ↕                    ↕                        │
│  [Ecosystem Partners] [Factory Software Stack]      │
│                                                     │
│  Result: Tokens/sec ↑  Cost/token ↓  Uptime ↑     │
└─────────────────────────────────────────────────────┘

The Core Value Proposition

Without NVLink FusionWith NVLink Fusion
Build everything from scratchUse proven infrastructure
Slow time-to-marketFaster deployment
High infrastructure riskReduced operational risk
Locked to one chipFlexible, mixed silicon
Separate planning for each chipUnified rack architecture

Quick Review: Key Concepts to Remember

  1. AI factories are measured by tokens/second, tokens/watt, cost/token, and uptime
  2. XPUs are custom chips — but the chip is only part of the challenge
  3. Scale-up networking must deliver performance, resiliency, and platform maturity
  4. NVLink Fusion connects XPUs to NVIDIA's proven infrastructure via 6th-gen NVLink
  5. NVLink-C2C provides 6x better energy efficiency than PCIe for CPU-XPU connections
  6. MGX architecture enables shared rack infrastructure across GPU and XPU systems
  7. Risk is managed by decoupling facility buildout from chip development timelines
  8. Software (NCCL, Dynamo, Mission Control) completes the factory as a coordinated system

Self-Check Questions

  1. Why can't a company simply design a great XPU and call it an AI factory?
  2. What happens to cost-per-token when scale-up networking is too slow?
  3. How does NVLink Fusion reduce time-to-market risk for XPU developers?
  4. Why is the ability to share rack infrastructure between GPUs and XPUs strategically valuable?
  5. What role does the digital twin (Omniverse DSX Blueprint) play in factory planning?

More to study