AI Inference

NVIDIA Details GPU Sizing for AI Inference and TCO Optimization
AI Inference

NVIDIA Details GPU Sizing for AI Inference and TCO Optimization

NVIDIA explains how to size GPUs for AI inference workloads, balancing performance and TCO. Key for enterprises scaling generative AI.

OpenAI's Jalapeño Chip Outpaces Rivals in AI Inference Performance
AI Inference

OpenAI's Jalapeño Chip Outpaces Rivals in AI Inference Performance

OpenAI's Jalapeño custom AI chip achieves up to 1.9x better throughput per kilowatt and 3.6x lower latency than competing systems, redefining efficiency in AI inference.

OpenAI Unveils Jalapeño Chip with Industry-Leading AI Efficiency
AI Inference

OpenAI Unveils Jalapeño Chip with Industry-Leading AI Efficiency

OpenAI's Jalapeño chip outperforms commercial systems in AI inference efficiency, signaling a new era in custom silicon for advanced AI workloads.

NVIDIA Groq 3 LPX Achieves Full Production for Agentic AI
AI Inference

NVIDIA Groq 3 LPX Achieves Full Production for Agentic AI

NVIDIA's Groq 3 LPX and Vera Rubin NVL72 enhance AI inference with faster token generation and lower costs, reshaping agentic AI infrastructure.

ThunderAgent Boosts Synthetic Data Generation with 2.5x Speedup
AI Inference

ThunderAgent Boosts Synthetic Data Generation with 2.5x Speedup

ThunderAgent eliminates inefficiencies in agentic inference, achieving 2.5x throughput and near-linear scalability for synthetic data generation.

AMD Acquires FastFlowLM to Boost AI Inference Performance
AI Inference

AMD Acquires FastFlowLM to Boost AI Inference Performance

AMD acquires FastFlowLM, enhancing AI inference with NPU-first tech. Key step in advancing AI efficiency on Ryzen AI platforms.

NVIDIA's Inference Software Slashes AI Token Costs by 5x
AI Inference

NVIDIA's Inference Software Slashes AI Token Costs by 5x

NVIDIA's software stack on Blackwell GPUs reduces token costs by 5x, driving AI inference efficiency for major players like Baseten and Deep Infra.

NVIDIA Blackwell GPUs Achieve 15x AI Inference Boost With DFlash
AI Inference

NVIDIA Blackwell GPUs Achieve 15x AI Inference Boost With DFlash

NVIDIA's DFlash speculative decoding delivers 15x faster AI inference on Blackwell GPUs, revolutionizing multiagent workflows and boosting throughput.

BTTInferGrid: Decentralized AI Inference Network Announced
AI Inference

BTTInferGrid: Decentralized AI Inference Network Announced

BitTorrent Inc. unveils BTTInferGrid, a decentralized compute network for AI inference aimed at connecting idle GPUs with global AI demand.

NVIDIA TensorRT Brings FP8 Quantization to AI Deployment
AI Inference

NVIDIA TensorRT Brings FP8 Quantization to AI Deployment

NVIDIA TensorRT optimizes AI inference with FP8 quantization, offering faster performance and smaller models for scalable deployment.

Trending topics