AI Inference

Nvidia's SWE-Serve Highlights AI Inference Testing Gaps
AI Inference

Nvidia's SWE-Serve Highlights AI Inference Testing Gaps

Nvidia's SWE-Serve benchmark reveals critical gaps between local AI tests and live inference serving, targeting real-world deployment challenges.

NVIDIA Brings Confidential AI Inference to Blackwell GPUs
AI Inference

NVIDIA Brings Confidential AI Inference to Blackwell GPUs

NVIDIA's confidential computing delivers secure AI inference on Blackwell GPUs, retaining 96-98% performance while safeguarding sensitive data.

NVIDIA Dynamo EPD Boosts Multimodal AI Model Speed by 7x
AI Inference

NVIDIA Dynamo EPD Boosts Multimodal AI Model Speed by 7x

NVIDIA’s EPD disaggregation in Dynamo accelerates multimodal AI inference by up to 7x, optimizing vision encoding, prefill, and decode stages.

NVIDIA Details GPU Sizing for AI Inference and TCO Optimization
AI Inference

NVIDIA Details GPU Sizing for AI Inference and TCO Optimization

NVIDIA explains how to size GPUs for AI inference workloads, balancing performance and TCO. Key for enterprises scaling generative AI.

OpenAI's Jalapeño Chip Outpaces Rivals in AI Inference Performance
AI Inference

OpenAI's Jalapeño Chip Outpaces Rivals in AI Inference Performance

OpenAI's Jalapeño custom AI chip achieves up to 1.9x better throughput per kilowatt and 3.6x lower latency than competing systems, redefining efficiency in AI inference.

OpenAI Unveils Jalapeño Chip with Industry-Leading AI Efficiency
AI Inference

OpenAI Unveils Jalapeño Chip with Industry-Leading AI Efficiency

OpenAI's Jalapeño chip outperforms commercial systems in AI inference efficiency, signaling a new era in custom silicon for advanced AI workloads.

NVIDIA Groq 3 LPX Achieves Full Production for Agentic AI
AI Inference

NVIDIA Groq 3 LPX Achieves Full Production for Agentic AI

NVIDIA's Groq 3 LPX and Vera Rubin NVL72 enhance AI inference with faster token generation and lower costs, reshaping agentic AI infrastructure.

ThunderAgent Boosts Synthetic Data Generation with 2.5x Speedup
AI Inference

ThunderAgent Boosts Synthetic Data Generation with 2.5x Speedup

ThunderAgent eliminates inefficiencies in agentic inference, achieving 2.5x throughput and near-linear scalability for synthetic data generation.

AMD Acquires FastFlowLM to Boost AI Inference Performance
AI Inference

AMD Acquires FastFlowLM to Boost AI Inference Performance

AMD acquires FastFlowLM, enhancing AI inference with NPU-first tech. Key step in advancing AI efficiency on Ryzen AI platforms.

NVIDIA's Inference Software Slashes AI Token Costs by 5x
AI Inference

NVIDIA's Inference Software Slashes AI Token Costs by 5x

NVIDIA's software stack on Blackwell GPUs reduces token costs by 5x, driving AI inference efficiency for major players like Baseten and Deep Infra.

Trending topics