AI Inference
Nvidia's SWE-Serve Highlights AI Inference Testing Gaps
Nvidia's SWE-Serve benchmark reveals critical gaps between local AI tests and live inference serving, targeting real-world deployment challenges.
NVIDIA Brings Confidential AI Inference to Blackwell GPUs
NVIDIA's confidential computing delivers secure AI inference on Blackwell GPUs, retaining 96-98% performance while safeguarding sensitive data.
NVIDIA Dynamo EPD Boosts Multimodal AI Model Speed by 7x
NVIDIA’s EPD disaggregation in Dynamo accelerates multimodal AI inference by up to 7x, optimizing vision encoding, prefill, and decode stages.
NVIDIA Details GPU Sizing for AI Inference and TCO Optimization
NVIDIA explains how to size GPUs for AI inference workloads, balancing performance and TCO. Key for enterprises scaling generative AI.
OpenAI's Jalapeño Chip Outpaces Rivals in AI Inference Performance
OpenAI's Jalapeño custom AI chip achieves up to 1.9x better throughput per kilowatt and 3.6x lower latency than competing systems, redefining efficiency in AI inference.
OpenAI Unveils Jalapeño Chip with Industry-Leading AI Efficiency
OpenAI's Jalapeño chip outperforms commercial systems in AI inference efficiency, signaling a new era in custom silicon for advanced AI workloads.
NVIDIA Groq 3 LPX Achieves Full Production for Agentic AI
NVIDIA's Groq 3 LPX and Vera Rubin NVL72 enhance AI inference with faster token generation and lower costs, reshaping agentic AI infrastructure.
ThunderAgent Boosts Synthetic Data Generation with 2.5x Speedup
ThunderAgent eliminates inefficiencies in agentic inference, achieving 2.5x throughput and near-linear scalability for synthetic data generation.
AMD Acquires FastFlowLM to Boost AI Inference Performance
AMD acquires FastFlowLM, enhancing AI inference with NPU-first tech. Key step in advancing AI efficiency on Ryzen AI platforms.
NVIDIA's Inference Software Slashes AI Token Costs by 5x
NVIDIA's software stack on Blackwell GPUs reduces token costs by 5x, driving AI inference efficiency for major players like Baseten and Deep Infra.