GPU Optimization
NVIDIA Hybrid-EP Slashes MoE AI Training Communication Overhead by 14%
NVIDIA's new Hybrid-EP communication library achieves up to 14% faster training for DeepSeek-V3 and other MoE models on Grace Blackwell hardware.
NVIDIA TensorRT for RTX Brings Self-Optimizing AI to Consumer GPUs
NVIDIA's TensorRT for RTX introduces adaptive inference that automatically optimizes AI workloads at runtime, delivering 1.32x performance gains on RTX 5090.
AI Inference Costs Drop 40% With New GPU Optimization Tactics
Together AI reveals production-tested techniques cutting inference latency by 50-100ms while reducing per-token costs up to 5x through quantization and smart decoding.
Together AI Sets New Benchmark with Fastest Inference for Open-Source Models
Together AI achieves unprecedented speed in open-source model inference, leveraging GPU optimization and quantization techniques to outperform competitors on NVIDIA Blackwell architecture.
Exploring Handwritten PTX Code for GPU Optimization in CUDA
Delve into the potential of handwritten PTX code for enhancing GPU performance in CUDA applications, as outlined by NVIDIA experts.
NVIDIA Unveils Advanced Optimization Techniques for LLM Training on Grace Hopper
NVIDIA introduces advanced strategies for optimizing large language model (LLM) training on the Grace Hopper Superchip, enhancing GPU memory management and computational efficiency.
ThunderKittens Framework Enhanced for NVIDIA Blackwell GPUs
Together AI has optimized its ThunderKittens framework for NVIDIA Blackwell GPUs, enhancing performance with new open-source kernels, according to Together AI.
NVIDIA's GPU Innovations Revolutionize Drug Discovery Simulations
NVIDIA's latest GPU optimization techniques, including CUDA Graphs and C++ coroutines, promise to accelerate pharmaceutical research by enhancing molecular dynamics simulations.