AI Inference

NVIDIA Blackwell GPUs Achieve 15x AI Inference Boost With DFlash
AI Inference

NVIDIA Blackwell GPUs Achieve 15x AI Inference Boost With DFlash

NVIDIA's DFlash speculative decoding delivers 15x faster AI inference on Blackwell GPUs, revolutionizing multiagent workflows and boosting throughput.

BTTInferGrid: Decentralized AI Inference Network Announced
AI Inference

BTTInferGrid: Decentralized AI Inference Network Announced

BitTorrent Inc. unveils BTTInferGrid, a decentralized compute network for AI inference aimed at connecting idle GPUs with global AI demand.

NVIDIA TensorRT Brings FP8 Quantization to AI Deployment
AI Inference

NVIDIA TensorRT Brings FP8 Quantization to AI Deployment

NVIDIA TensorRT optimizes AI inference with FP8 quantization, offering faster performance and smaller models for scalable deployment.

MiniMax-M3 Launches 1M-Token Model With Sparse Attention
AI Inference

MiniMax-M3 Launches 1M-Token Model With Sparse Attention

MiniMax-M3 debuts with 1M-token context and multimodality, leveraging Together AI's optimizations for efficient large-scale inference.

NVIDIA Dynamo Snapshot Tackles Kubernetes AI Cold-Start Problem
AI Inference

NVIDIA Dynamo Snapshot Tackles Kubernetes AI Cold-Start Problem

NVIDIA's Dynamo Snapshot reduces Kubernetes AI inference cold-start times, leveraging CRIU and GPU Memory Service for sub-5-second deployment speed.

Together AI Joins Pearl Labs to Cut AI Inference Costs With Blockchain
AI Inference

Together AI Joins Pearl Labs to Cut AI Inference Costs With Blockchain

Together AI partners with Pearl Research Labs to slash AI inference costs using Proof of Useful Work, generating crypto rewards for GPU workloads.

DeepSeek-V4 Tackles Million-Token Context on NVIDIA HGX B200
AI Inference

DeepSeek-V4 Tackles Million-Token Context on NVIDIA HGX B200

DeepSeek-V4 introduces a 1M-token context window with a hybrid attention architecture, shifting the challenge to inference systems on NVIDIA hardware.

Mamba-3 SSM Drops With Inference-First Design Beating Transformers at Decode
AI Inference

Mamba-3 SSM Drops With Inference-First Design Beating Transformers at Decode

Together.ai releases Mamba-3, an open-source state space model built for inference that outperforms Mamba-2 and matches Transformer decode speeds at 16K sequences.

NVIDIA Unveils Groq 3 LPX Rack System for Ultra-Low Latency AI Inference
AI Inference

NVIDIA Unveils Groq 3 LPX Rack System for Ultra-Low Latency AI Inference

NVIDIA's new Groq 3 LPX delivers 315 PFLOPS and 35x better inference throughput per megawatt, targeting agentic AI workloads on the Vera Rubin platform.

NVIDIA Blackwell Smashes Finance AI Benchmark With 3.2x Speed Gains
AI Inference

NVIDIA Blackwell Smashes Finance AI Benchmark With 3.2x Speed Gains

NVIDIA's GB200 NVL72 sets new STAC-AI record for LLM inference in financial trading, delivering up to 3.2x performance over Hopper architecture.

Trending topics