Hardware

NVIDIA Launches NVLink Fusion to Connect Custom XPUs

NVIDIA has unveiled NVLink Fusion and NVHBM, a new technology stack that allows cloud providers to integrate custom AI processors directly into NVIDIA's rack-scale infrastructure.

NVIDIA Developer Blog10 hrs agoHardware
Image: NVIDIA Developer Blog

NVIDIA has introduced NVLink Fusion, an interconnect technology designed to help hyperscalers and AI-native firms integrate custom XPUs and CPUs into the NVIDIA AI infrastructure platform. By leveraging the MGX rack-scale architecture, developers can deploy custom silicon while utilizing NVIDIA's existing scale-up and scale-out technology stacks. Upstream, these custom XPUs can connect directly to CPUs via NVLink-C2C. This integration aims to simplify development, lower deployment complexity, and accelerate time to market for semi-custom AI factories.

At the package level, the architecture is supported by NVHBM, a custom high-bandwidth memory base-die technology designed alongside leading memory manufacturers. NVHBM delivers up to 30 percent more memory bandwidth per stack compared to the standard JEDEC HBM4e specification. It achieves this by moving the memory controller into the 3D HBM stack and integrating a custom physical interface. This redesign slashes the physical interface and support area by up to 67 percent and simplifies interposer routing to provide up to 80 percent more usable silicon across the layout. Consequently, this frees up to 30 percent more main-die silicon, allowing chip designers to allocate up to 25 percent more compute die area for custom capabilities.

Power efficiency is another critical benefit of the new memory technology. NVHBM reduces memory power consumption by up to 15 percent compared to standard HBM4e. At the scale of a one-gigawatt data center running 2,000-watt XPUs, these power savings can free up enough thermal and electrical headroom to support up to 15,000 additional XPUs. When combined, NVLink Fusion and NVHBM yield a 30 percent overall end-to-end performance increase per XPU by compounding these bandwidth, area, and power improvements.

For hardware engineers and cloud architects, this development bridges the gap between proprietary silicon designs and standardized data center deployments. Instead of building custom networking and cooling systems from scratch, practitioners can use the sixth-generation NVLink fabric to connect their custom chips into a single scale-up domain. This allows custom XPUs to run alongside GPUs in heterogeneous environments, making it easier to handle massive AI workloads like expert parallelism and large-scale inference.

This is our own summary of reporting by NVIDIA Developer Blog

More in Hardware