GPU Optimization
AMD ROCm 10 Debuts with ROCm.AI, Promises 3.3x AI Gains
AMD launches ROCm 10, marking a 10-year milestone with new AI-native tools like ROCm.AI, delivering 3.3x inference and 2.4x training boosts.
AMD’s TLX Kernel Boosts GEMM Performance for LLM Training
AMD's TLX optimizations accelerate GEMM, cutting memory bottlenecks and enhancing large language model training on GPUs.
Together AI Kernels Team Achieves 3.6x Performance Gains on NVIDIA Hardware
Together AI's kernel research team delivers major GPU optimization breakthroughs, cutting inference latency from 281ms to 77ms for enterprise AI deployments.
NVIDIA MIG Boosts AI Infrastructure ROI by 33% Over Time-Slicing
New NVIDIA benchmarks show Multi-Instance GPU partitioning achieves 1.00 req/s per GPU versus 0.76 for time-slicing in production AI workloads.
NVIDIA Advances AI Infrastructure With Disaggregated LLM Inference on Kubernetes
NVIDIA details new Kubernetes deployment patterns for disaggregated LLM inference using Dynamo and Grove, promising better GPU utilization for AI workloads.
FlashAttention-4 Hits 71% GPU Utilization on NVIDIA Blackwell B200
Together AI's FlashAttention-4 achieves 1,605 TFLOPs/s on B200 GPUs, up to 2.7x faster than Triton. New pipelining overcomes asymmetric hardware scaling bottlenecks.
NVIDIA Releases Flash Attention Optimization Guide for Blackwell GPUs
NVIDIA's new cuTile framework delivers 1.6x speedups for Flash Attention on B200 GPUs, enabling faster LLM inference critical for AI infrastructure.
NVIDIA Run:ai Delivers 2x GPU Utilization Gains for AI Inference Workloads
NVIDIA benchmarks show Run:ai platform doubles GPU utilization while cutting latency 61x for enterprise AI deployments running NIM inference microservices.
NVIDIA MIG Tech Delivers 2.25x Speedups for Power-Constrained AI Workloads
NVIDIA's Multi-Instance GPU technology shows up to 2.25x performance gains for data center workloads under power limits, with implications for AI infrastructure costs.
NVIDIA Blackwell Delivers 4x Inference Boost for India's Sarvam AI Models
NVIDIA's hardware-software co-design achieves 4x inference speedup for Sarvam AI's 30B parameter sovereign models, showcasing Blackwell's NVFP4 capabilities.