AI Inference
NVIDIA Details GPU Sizing for AI Inference and TCO Optimization
NVIDIA explains how to size GPUs for AI inference workloads, balancing performance and TCO. Key for enterprises scaling generative AI.
OpenAI's Jalapeño Chip Outpaces Rivals in AI Inference Performance
OpenAI's Jalapeño custom AI chip achieves up to 1.9x better throughput per kilowatt and 3.6x lower latency than competing systems, redefining efficiency in AI inference.
OpenAI Unveils Jalapeño Chip with Industry-Leading AI Efficiency
OpenAI's Jalapeño chip outperforms commercial systems in AI inference efficiency, signaling a new era in custom silicon for advanced AI workloads.
NVIDIA Groq 3 LPX Achieves Full Production for Agentic AI
NVIDIA's Groq 3 LPX and Vera Rubin NVL72 enhance AI inference with faster token generation and lower costs, reshaping agentic AI infrastructure.
ThunderAgent Boosts Synthetic Data Generation with 2.5x Speedup
ThunderAgent eliminates inefficiencies in agentic inference, achieving 2.5x throughput and near-linear scalability for synthetic data generation.
AMD Acquires FastFlowLM to Boost AI Inference Performance
AMD acquires FastFlowLM, enhancing AI inference with NPU-first tech. Key step in advancing AI efficiency on Ryzen AI platforms.
NVIDIA's Inference Software Slashes AI Token Costs by 5x
NVIDIA's software stack on Blackwell GPUs reduces token costs by 5x, driving AI inference efficiency for major players like Baseten and Deep Infra.
NVIDIA Blackwell GPUs Achieve 15x AI Inference Boost With DFlash
NVIDIA's DFlash speculative decoding delivers 15x faster AI inference on Blackwell GPUs, revolutionizing multiagent workflows and boosting throughput.
BTTInferGrid: Decentralized AI Inference Network Announced
BitTorrent Inc. unveils BTTInferGrid, a decentralized compute network for AI inference aimed at connecting idle GPUs with global AI demand.
NVIDIA TensorRT Brings FP8 Quantization to AI Deployment
NVIDIA TensorRT optimizes AI inference with FP8 quantization, offering faster performance and smaller models for scalable deployment.