After studying this material, you should be able to:
Explain how NVLink Fusion enables custom XPUs to integrate into a complete, production-ready AI factory infrastructure — and why this matters for performance, time-to-market, and operational risk management.
Specifically, you will:
Before anything else, you need to understand the context — what problem are we solving?
An AI factory is not a traditional data center. Think of it like a manufacturing plant, but instead of producing physical goods, it produces intelligence — specifically, AI-generated outputs called tokens.
| Metric | What it means |
|---|---|
| Tokens per second | How fast it produces output |
| Tokens per watt | How energy-efficient it is |
| Cost per token | How economically viable it is |
| Utilization & uptime | How continuously it operates |
An AI factory must run continuously. Every second of downtime or inefficiency directly increases cost per token and reduces competitiveness.
Think of it this way: A car factory that stops its assembly line every hour is far less profitable than one that runs 24/7 without interruption. AI factories work the same way.
An XPU (Accelerated Processing Unit) is a custom-designed chip built by hyperscalers (like Amazon, Google, Meta) or AI-native companies to accelerate specific AI workloads.
Building a chip is only part of the challenge. To actually deploy an XPU in production, a company must also build:
Custom XPU
↓
+ High-speed CPU interfaces
+ Scale-up networking solution
+ Compute and switch trays
+ Rack architecture (cooling + power)
+ Security and storage integration
+ Supplier ecosystem management
↓
= Complete AI Factory
Most teams underestimate this. The chip might take 2 years to design, but the surrounding infrastructure can take just as long — and cost just as much.
Modern AI workloads — like trillion-parameter models, mixture-of-experts architectures, and agentic AI — require chips to communicate with each other at extreme speeds.
This communication is called scale-up networking (connecting chips within a single system or rack).
Slow Network → Chips sit idle waiting for data
→ Utilization drops
→ Cost per token rises
→ AI factory becomes uneconomical
Remember: A networking solution that looks good on paper but has never been deployed at scale is a risk, not an asset.
NVLink Fusion is NVIDIA's solution that allows custom XPUs to plug into NVIDIA's existing, world-class AI infrastructure — rather than building that infrastructure from scratch.
Think of it as a bridge between a company's custom chip and a proven, production-ready platform.
| Comparison | NVLink Fusion | Off-the-shelf Ethernet |
|---|---|---|
| End-to-end latency | Baseline | 3x higher (worse) |
| Packet rate | Baseline | 10x lower (worse) |
NVLink Fusion is supported by partners across:
This means companies don't need to find and validate every supplier themselves.
AI factory planning starts before the chips are ready:
Timeline:
[Power procurement] → [Facility design] → [Cooling design] → [Rack layout] → [Chip arrives]
↑
This starts YEARS before silicon is finalized
If your entire data center is locked to one specific chip, and that chip is delayed or changed, your entire facility plan is at risk.
NVLink Fusion creates a unified architecture where:
A hyperscaler can deploy NVIDIA GPU racks today using NVLink Fusion infrastructure, while their custom XPU is still in development. When the XPU is ready, it slots into the same rack design — no facility rework needed.
As MediaTek's VP explained:
"They can deploy their rack-level solution with the NVIDIA GPU, and then decouple the development of their XPU and put it at a different pace."
Hardware alone doesn't make a factory run. Software coordinates everything.
| Software | Purpose |
|---|---|
| NVIDIA NCCL | Distributed workload communication between chips |
| NVIDIA Dynamo & NIXL | Disaggregation — splitting workloads across systems |
| NVIDIA Mission Control | Cluster management, telemetry, and debugging |
| NVIDIA Omniverse DSX Blueprint | Digital twin for modeling the entire factory before building it |
Factory buildout is expensive. Mistakes require costly rework.
The DSX AI Factory Blueprint lets partners simulate the entire facility — buildings, power, cooling, compute, and networking — before a single rack is installed. This catches design errors early, when they're cheap to fix.
Now let's put it all together with a mental model:
PROBLEM:
Custom XPU alone ≠ AI Factory
Building everything from scratch = slow, expensive, risky
SOLUTION: NVLink Fusion
┌─────────────────────────────────────────────────────┐
│ AI FACTORY │
│ │
│ [Custom XPU] ←→ [NVLink Scale-Up Networking] │
│ ↕ ↕ │
│ [NVLink-C2C CPU] [MGX Rack Architecture] │
│ ↕ ↕ │
│ [Ecosystem Partners] [Factory Software Stack] │
│ │
│ Result: Tokens/sec ↑ Cost/token ↓ Uptime ↑ │
└─────────────────────────────────────────────────────┘
| Without NVLink Fusion | With NVLink Fusion |
|---|---|
| Build everything from scratch | Use proven infrastructure |
| Slow time-to-market | Faster deployment |
| High infrastructure risk | Reduced operational risk |
| Locked to one chip | Flexible, mixed silicon |
| Separate planning for each chip | Unified rack architecture |