Kubernetes
NVIDIA's NVCRE Boosts AI Cluster Reliability With Real Workload Validation
NVIDIA's Cluster Readiness Engine certifies GPU clusters for production by running real AI workloads. Learn how it identifies bottlenecks and faulty nodes.
NVIDIA Launches NodeWright for Kubernetes Node Management
NVIDIA's NodeWright simplifies Kubernetes node ops, enabling GPU fleet updates without disrupting AI workloads.
NVIDIA Launches Topograph for Topology-Aware GPU Scheduling
NVIDIA's Topograph optimizes GPU workload placement for AI factories, reducing latency and boosting efficiency across cloud and on-premises clusters.
NVIDIA FLARE Expands Federated Learning with Kubernetes, Slurm
NVIDIA FLARE 2.9 adds Slurm support, enabling scalable federated learning across heterogeneous infrastructures like Docker, Kubernetes, and HPC clusters.
KubeRay v1.7 Launches With Upgraded Features for Ray on Kubernetes
KubeRay v1.7 introduces major upgrades like History Server beta, enhanced RayJob management, and stronger security for Ray workloads on Kubernetes.
NVIDIA’s KAI Scheduler and vCluster Enable GPU Sharing in Kubernetes
NVIDIA introduces a solution for isolated Kubernetes clusters and efficient GPU sharing, reducing infrastructure costs for AI/ML teams.
NVIDIA Dynamo Snapshot Tackles Kubernetes AI Cold-Start Problem
NVIDIA's Dynamo Snapshot reduces Kubernetes AI inference cold-start times, leveraging CRIU and GPU Memory Service for sub-5-second deployment speed.
NVIDIA Open-Sources Slinky to Run Slurm GPU Workloads on Kubernetes
NVIDIA's Slinky project enables running Slurm clusters on Kubernetes, already deployed on 8,000+ GPU systems for large-scale AI training infrastructure.
NVIDIA MIG Boosts AI Infrastructure ROI by 33% Over Time-Slicing
New NVIDIA benchmarks show Multi-Instance GPU partitioning achieves 1.00 req/s per GPU versus 0.76 for time-slicing in production AI workloads.
NVIDIA Donates GPU Resource Driver to Kubernetes Open Source Project
NVIDIA transfers critical GPU allocation software to CNCF at KubeCon Europe, marking major shift toward community-governed AI infrastructure.