AI Training
NVIDIA Optimizes JAX LLM Training with Host Offloading
NVIDIA's host offloading for JAX LLM training boosts GPU memory efficiency, enabling larger batch sizes and faster throughput.
NVIDIA Pushes Low-Precision Transformer Training with NVFP4
NVIDIA's NVFP4 enables faster, cheaper transformer training with low-precision techniques. Learn about the latest benchmarks and implications for AI modeling.
NVIDIA Blackwell Dominates MLPerf Training v6.0 Benchmarks
NVIDIA's Blackwell GPUs set new records in MLPerf Training v6.0, showcasing unmatched scale and performance in AI model training.
NVIDIA Blackwell GPUs Dominate MLPerf Training 6.0 Benchmarks
NVIDIA's Blackwell GPUs set new records in AI training performance and scalability in MLPerf 6.0, solidifying its dominance in next-gen infrastructure.
Nvidia's New MoE Kernels Promise 93% Speedup for AI Training
Nvidia unveils advanced MoE training kernels, boosting AI model throughput by up to 93% in GPT pre-training and redefining large-scale efficiency.
NVIDIA's NVFP4 Boosts JAX Model Training on Blackwell GPUs
NVIDIA's NVFP4 enables 4-bit precision training on Blackwell GPUs, delivering up to 73% faster throughput for Llama models without accuracy loss.
SkyRL Adds Vision-Language RL Support for Multimodal Models
SkyRL introduces vision-language reinforcement learning, enabling scalable training for multimodal tasks. Learn how this impacts AI development.
Google's Decoupled DiLoCo Redefines Distributed AI Training
Google's Decoupled DiLoCo architecture enables faster, resilient AI training across data centers, leveraging mixed-generation hardware for efficiency.
NVIDIA Megatron Boosts LLM Training With Muon Optimizer
NVIDIA integrates Muon and advanced optimizers into Megatron to enhance large-scale LLM training with near-parity throughput to AdamW.
NVIDIA NeMo RL Achieves 48% Speedup with End-to-End FP8 Precision Training
NVIDIA's new FP8 recipe for reinforcement learning delivers 48% faster training while matching BF16 accuracy, cutting AI infrastructure costs significantly.