NVIDIA GB300 NVL72 Enhances AI with Ray Placement Groups

James Ding Aug 13, 2026 17:10

NVIDIA GB300 NVL72 leverages Ray's NVLink Domain-Aware Placement Groups for optimized multi-node GPU scheduling, boosting AI performance.

NVIDIA GB300 NVL72 Enhances AI with Ray Placement Groups

NVIDIA's GB300 NVL72, the rack-scale AI platform integrating 72 Blackwell Ultra GPUs and 36 Grace CPUs per rack, is now even more powerful with the introduction of NVLink Domain-Aware Placement Groups in Ray. This new scheduling feature, announced on August 13, 2026, allows users to optimize GPU workloads by colocating processes within a single NVLink domain, unlocking significant performance gains for large-scale AI applications.

The GB300 NVL72 represents NVIDIA’s most advanced rack-scale architecture to date, with each rack offering up to 1.1 exaFLOPS of dense FP4 compute. Using NVLink 5, the rack achieves 1,800 GB/s all-to-all bandwidth per GPU, enabling seamless communication across all 72 GPUs as though they were a unified compute unit. The platform is designed for hyperscale AI workloads, including large language model training, generative AI inference, and agentic AI applications.

Why NVLink Domain-Aware Placement Groups Matter

Prior to this update, Ray placement groups only handled node-level scheduling, which could lead to inefficiencies in multi-rack setups like the GB300. The new placement groups add topology awareness, ensuring that tightly coupled tasks, such as GPU-intensive training or inference workloads, stay within the same NVLink domain. This avoids slower inter-node communication and maximizes the benefits of the GB300’s high-bandwidth NVLink fabric.

For example, NVIDIA’s GEAR research lab tested this feature on large-scale Vision Language Action (VLA) pretraining workloads across a GB300 cluster. When NVLink Domain-Aware Placement Groups were used, the system achieved 1.13x faster iterations per second compared to traditional placement methods. This improvement stems from reduced overhead in GPU-to-GPU communication, as actors remained colocated within the same NVLink domain.

Implications for AI and HPC

The GB300 NVL72 positions itself as a critical tool for hyperscale AI systems. By combining its advanced hardware with Ray’s new scheduling capabilities, enterprises can now better utilize the GB300 for demanding tasks like trillion-parameter model inference and reinforcement learning. NVIDIA’s advancements align with the growing demand for AI supercomputers, as evidenced by Microsoft’s Azure cluster deployment in 2025, which featured over 4,600 GB300 GPUs.

This update also simplifies fault tolerance. When a node within an NVLink domain fails, Ray can automatically reschedule workloads within the same rack, preserving performance and minimizing downtime. Previously, this level of fault-aware scheduling required manual labeling, which was prone to errors and inefficiencies.

What’s Next for NVIDIA and Ray

NVIDIA and Ray developers are already planning enhancements to this feature. Future updates aim to support hierarchical topologies, such as datacenter-level scheduling, and introduce additional placement strategies like STRICT_SPREAD for distributing workloads across multiple racks. These developments could further optimize GB300 deployments, particularly in scenarios involving hybrid AI workloads.

As NVIDIA continues to dominate the AI hardware market, the GB300 NVL72’s integration with Ray highlights the importance of both hardware and software innovation in scaling AI systems. Traders and technology investors may want to keep an eye on NVIDIA's market position, as its shares have recently shown resilience, trading at $225.32 with a 0.55% daily gain as of August 13, 2026.

For developers, Ray’s NVLink Domain-Aware Placement Groups are available now. Documentation and API details can be found on the Ray project site, and NVIDIA encourages community feedback to refine and expand its capabilities.

Image source: Shutterstock