NVIDIA adds NVHBM to NVLink Fusion for custom AI data center chips

NVIDIA adds NVHBM to NVLink Fusion for custom AI data center chips

NVIDIA added NVHBM to NVLink Fusion, promising more bandwidth, lower HBM power and AWS Trainium collaboration for custom AI chips.

Format News Brief
Read Time 3 min
Category Hardware
Updated Aug 28, 2026

NVIDIA has expanded NVLink Fusion with NVHBM, a custom high-bandwidth memory technology aimed at hyperscalers and AI infrastructure builders designing their own XPUs. The company says Amazon's Annapurna Labs will be the first collaborator on NVHBM as part of a broader NVLink Fusion effort that includes future Trainium chips.

What changed

NVLink Fusion is NVIDIA's path for connecting custom CPUs and accelerators into its rack-scale AI platform. NVHBM adds a memory layer to that strategy. Instead of leaving the memory controller on the accelerator die, NVIDIA says NVHBM moves a custom controller into the 3D HBM stack and uses a redesigned physical interface. That is meant to reduce the silicon and package area consumed by memory plumbing while keeping data closer to compute.

The claimed gains are specific. NVIDIA says NVHBM can deliver up to 30 percent more memory bandwidth than standard HBM4E, cut HBM power use by up to 15 percent, and free up to 25 percent more area on the XPU compute die. In a deeper technical post, NVIDIA also says the PHY and support area can shrink by up to 67 percent, which can leave up to 30 percent more main die silicon for compute or other chip features.

Why it matters

For AI infrastructure buyers, the practical issue is not just peak accelerator speed. Large model inference spends a lot of time moving weights and KV cache data, and agent workloads can make that pressure worse by stretching sessions across more steps and context. If NVIDIA's claims hold in production systems, custom chip teams could get more useful memory bandwidth and power headroom without giving up the rack software and networking stack around NVLink.

The AWS detail is important because it shows NVHBM is not framed only as an NVIDIA GPU feature. Annapurna Labs plans to support NVLink Fusion with next-generation Trainium chips starting with Trainium4, according to NVIDIA. That would let Amazon silicon and NVIDIA GPUs share a common rack architecture, which matters for cloud operators trying to mix internal accelerators with broader ecosystem hardware.

What to watch next

The CyberOGZ read is that NVHBM is less about a single memory spec and more about NVIDIA making itself harder to route around in custom AI data centers. Hyperscalers want their own chips for cost and workload control, but they still need networking, memory validation, software, supply options, and operations tooling. NVLink Fusion gives NVIDIA a way to participate even when the accelerator in the rack is not a GPU. The next useful signal will be partner timelines and independent measurements, especially around real inference throughput, power savings, and how much design flexibility customers gain after qualification.

Sources

Cover photo by Armando Are on Pexels, used under the Pexels License.

Comments (0)

Leave a Comment

Loading comments...