Search Results for "ai inference"

DeepSeek-V4 Tackles Million-Token Context on NVIDIA HGX B200

DeepSeek-V4 Tackles Million-Token Context on NVIDIA HGX B200

DeepSeek-V4 introduces a 1M-token context window with a hybrid attention architecture, shifting the challenge to inference systems on NVIDIA hardware.

Together AI Joins Pearl Labs to Cut AI Inference Costs With Blockchain

Together AI Joins Pearl Labs to Cut AI Inference Costs With Blockchain

Together AI partners with Pearl Research Labs to slash AI inference costs using Proof of Useful Work, generating crypto rewards for GPU workloads.

NVIDIA Dynamo Snapshot Tackles Kubernetes AI Cold-Start Problem

NVIDIA Dynamo Snapshot Tackles Kubernetes AI Cold-Start Problem

NVIDIA's Dynamo Snapshot reduces Kubernetes AI inference cold-start times, leveraging CRIU and GPU Memory Service for sub-5-second deployment speed.

MiniMax-M3 Launches 1M-Token Model With Sparse Attention

MiniMax-M3 Launches 1M-Token Model With Sparse Attention

MiniMax-M3 debuts with 1M-token context and multimodality, leveraging Together AI's optimizations for efficient large-scale inference.

NVIDIA TensorRT Brings FP8 Quantization to AI Deployment

NVIDIA TensorRT Brings FP8 Quantization to AI Deployment

NVIDIA TensorRT optimizes AI inference with FP8 quantization, offering faster performance and smaller models for scalable deployment.

BTTInferGrid: Decentralized AI Inference Network Announced

BTTInferGrid: Decentralized AI Inference Network Announced

BitTorrent Inc. unveils BTTInferGrid, a decentralized compute network for AI inference aimed at connecting idle GPUs with global AI demand.

NVIDIA Blackwell GPUs Achieve 15x AI Inference Boost With DFlash

NVIDIA Blackwell GPUs Achieve 15x AI Inference Boost With DFlash

NVIDIA's DFlash speculative decoding delivers 15x faster AI inference on Blackwell GPUs, revolutionizing multiagent workflows and boosting throughput.

NVIDIA's Inference Software Slashes AI Token Costs by 5x

NVIDIA's Inference Software Slashes AI Token Costs by 5x

NVIDIA's software stack on Blackwell GPUs reduces token costs by 5x, driving AI inference efficiency for major players like Baseten and Deep Infra.

AMD Acquires FastFlowLM to Boost AI Inference Performance

AMD Acquires FastFlowLM to Boost AI Inference Performance

AMD acquires FastFlowLM, enhancing AI inference with NPU-first tech. Key step in advancing AI efficiency on Ryzen AI platforms.

ThunderAgent Boosts Synthetic Data Generation with 2.5x Speedup

ThunderAgent Boosts Synthetic Data Generation with 2.5x Speedup

ThunderAgent eliminates inefficiencies in agentic inference, achieving 2.5x throughput and near-linear scalability for synthetic data generation.

China Looking into the Application of Blockchain and AI for Cross-Border Financing

China Looking into the Application of Blockchain and AI for Cross-Border Financing

China is researching the application of blockchain technology and artificial intelligence in cross-border financing, focusing on risk management.

Exclusive: Blockchain Beats AI and Big Data for the Highest Average Annual Salary in the UK

Exclusive: Blockchain Beats AI and Big Data for the Highest Average Annual Salary in the UK

The report titled “The Disruption of Disruptive Tech” by Capital on Tap highlighted the state of disruptive tech adoption in early 2020. The U.S. had the most businesses in various disruptive technologies, with 71% dominance in cloud consulting and 53% in cybersecurity.

Trending topics