Inference

NVIDIA Dynamo's Shadow Engine Slashes LLM Recovery to 7 Seconds
Inference

NVIDIA Dynamo's Shadow Engine Slashes LLM Recovery to 7 Seconds

NVIDIA's Shadow Engine Recovery in Dynamo cuts LLM downtime from minutes to 7 seconds, transforming AI infrastructure resilience.

NVIDIA Groq 3 LPX Achieves 3,431 TPS Benchmark on Vera Rubin
Inference

NVIDIA Groq 3 LPX Achieves 3,431 TPS Benchmark on Vera Rubin

NVIDIA's Groq 3 LPX sets a new standard in AI inference, delivering 3,431 tokens/second on a 100K context benchmark and redefining high-interactivity workloads.

NVIDIA Groq 3 LPX Hits Full Production, Boosts AI Inference Speed
Inference

NVIDIA Groq 3 LPX Hits Full Production, Boosts AI Inference Speed

NVIDIA's Groq 3 LPX enters full production, delivering record-breaking AI inference speeds for latency-critical workloads. Key to agentic AI development.

AMD ZenDNN 6.0 Boosts AI Inference on EPYC CPUs
Inference

AMD ZenDNN 6.0 Boosts AI Inference on EPYC CPUs

AMD's ZenDNN 6.0 introduces FP16 support, MoE optimizations, and expanded vLLM compatibility, enhancing AI inference capabilities on EPYC processors.

NVIDIA Unveils AI Factory Energy Optimization Tools for Token Efficiency
Inference

NVIDIA Unveils AI Factory Energy Optimization Tools for Token Efficiency

NVIDIA introduces tools like DSX and NVFP4 to improve energy efficiency in AI factories, potentially lowering token production costs by up to 25%.

AI Data Processing Shifts to GPUs: Key Trends and Impacts
Inference

AI Data Processing Shifts to GPUs: Key Trends and Impacts

AI pipelines are increasingly GPU-driven as inference-heavy workloads handle unstructured data, reshaping data processing and infrastructure demands.

NVIDIA Unveils AI Grid Architecture for Distributed Edge Inference at GTC 2026
Inference

NVIDIA Unveils AI Grid Architecture for Distributed Edge Inference at GTC 2026

NVIDIA's AI Grid reference design enables telcos to cut inference costs by 76% and meet sub-500ms latency targets through distributed edge computing.

NVIDIA Blackwell Enhances AI Inference with Superior Performance Gains
Inference

NVIDIA Blackwell Enhances AI Inference with Superior Performance Gains

NVIDIA Blackwell architecture delivers substantial performance improvements for AI inference, utilizing advanced software optimizations and hardware innovations to enhance efficiency and throughput.

NVIDIA's Breakthrough: 4x Faster Inference in Math Problem Solving with Advanced Techniques
Inference

NVIDIA's Breakthrough: 4x Faster Inference in Math Problem Solving with Advanced Techniques

NVIDIA achieves a 4x faster inference in solving complex math problems using NeMo-Skills, TensorRT-LLM, and ReDrafter, optimizing large language models for efficient scaling.

Enhancing LLM Inference with NVIDIA Run:ai and Dynamo Integration
Inference

Enhancing LLM Inference with NVIDIA Run:ai and Dynamo Integration

NVIDIA's Run:ai v2.23 integrates with Dynamo to address large language model inference challenges, offering gang scheduling and topology-aware placement for efficient, scalable deployments.

Trending topics