AI Infrastructure

NVIDIA pairs NVLink Fusion with NVHBM to court custom AI accelerator builders

NVIDIA detailed how NVLink Fusion and NVHBM are meant to help hyperscalers and AI-native companies deploy custom XPUs inside rack-scale AI infrastructure.

Published Updated
NVIDIANVLink FusionAI Hardware

NVIDIA has detailed a new infrastructure push aimed at the companies building custom AI accelerators, pairing NVLink Fusion with a memory technology it calls NVHBM. The August 26 technical post targets a central tension in the AI hardware market. Hyperscalers and AI-native companies increasingly want workload-specific XPUs, but those chips still need high-bandwidth memory, rack-scale communication, software support and operational integration if they are going to compete with mature GPU infrastructure.

NVLink Fusion is NVIDIA’s connective technology for bringing custom XPUs and CPUs into the company’s AI infrastructure platform. The idea is to let semi-custom accelerators use NVIDIA’s scale-up and scale-out technology stack, MGX rack architecture and broader ecosystem rather than forcing each buyer to assemble an isolated infrastructure path. That is strategically important because the AI chip market is fragmenting: major labs and cloud providers want specialized silicon, but they also need those chips to participate in large training, inference and expert-parallel workloads that depend on fast communication across a rack.

NVHBM addresses the package-level side of the problem. NVIDIA describes it as a custom HBM base-die technology designed and validated with leading memory vendors. The claimed benefits are substantial: up to 30 percent more memory bandwidth than standard HBM4e, up to 25 percent more compute die area for additional XPU features and up to 15 percent lower HBM power usage. NVIDIA also says the design reduces PHY and support area by up to 67 percent compared with JEDEC HBM4e, freeing more package area for compute or other accelerator logic.

Those numbers matter because modern AI systems are often limited by data movement, not only arithmetic throughput. Large-model inference repeatedly reads model weights and KV-cache data while serving users with tight latency targets. Training and agentic workloads move activations across devices, especially when mixture-of-experts models use expert parallelism. More memory bandwidth can keep compute engines fed, while lower memory power can improve performance per watt and ease cooling constraints across thousands of accelerators.

NVIDIA’s post also frames NVHBM as a way to give custom silicon teams more design flexibility. Package area is one of the hardest constraints in accelerator design because engineers must divide space among matrix engines, vector units, cache, SRAM, memory interfaces, networking and control logic. By shrinking the memory interface footprint, NVHBM could let designers allocate more area to workload-specific features within the same physical envelope. For a hyperscaler building an inference-optimized XPU, that flexibility could be as valuable as raw bandwidth.

The rack-scale argument is equally important. NVIDIA says NVLink Fusion uses chiplet-based connectivity to bridge custom XPUs into NVLink domains, allowing accelerators to communicate across a shared scale-up fabric and connect upstream to CPUs through NVLink-C2C. In mixture-of-experts serving, where tokens may be routed to different experts on different devices, that kind of fabric can determine whether the system behaves like one coordinated machine or a collection of bottlenecked chips.

The announcement is also defensive. NVIDIA dominates AI training and inference with GPUs, but some of its largest customers are designing their own chips to manage cost, supply and workload specialization. NVLink Fusion and NVHBM give NVIDIA a path to remain the infrastructure layer even when a rack includes more non-GPU silicon. If the approach works in production, the company can benefit from the custom accelerator wave instead of treating it only as a threat.

Customers will still judge the platform on practical criteria: price, availability, memory vendor support, software maturity, integration effort and measured performance under real workloads. But the direction is clear. AI infrastructure is moving toward heterogeneous racks where GPUs, custom XPUs, CPUs, high-bandwidth memory and networking are co-designed. NVIDIA’s message is that the next generation of custom AI chips may still need to plug into NVIDIA’s fabric to become useful at data-center scale.