Benchmarking
Nvidia's SWE-Serve Highlights AI Inference Testing Gaps
Nvidia's SWE-Serve benchmark reveals critical gaps between local AI tests and live inference serving, targeting real-world deployment challenges.
Claude AI Improves Alignment Benchmarks While Preserving Capabilities
Anthropic's Claude achieved significant alignment improvements on 10 benchmarks, outperforming human researchers and maintaining model capabilities.
DeepMind Launches First Double-Blind AI Model Evaluation
Google DeepMind introduces double-blind AI evaluations using cryptographic safeguards to improve trust in model benchmarks.
GLM-5.3 Beats Claude Fable 5 on DeepSWE, Costs 5.4x Less
GLM-5.3 outperforms Claude Fable 5 on DeepSWE's coding benchmark for retries and cost efficiency, at just $3.99 per rollout vs. $21.63.
Harvey Expands Legal Agent Benchmark to Contract Negotiation
Harvey extends its Legal Agent Benchmark (LAB) with 500 new tasks focused on contract drafting, review, and negotiation for in-house legal teams.
Harvey Unveils Open-Source Legal Agent Benchmark LAB
Harvey launches LAB, a benchmark to evaluate AI performance in legal tasks, covering 24 practice areas with over 1,200 tasks.
Together AI Introduces Flexible Benchmarking for LLMs
Together AI unveils Together Evaluations, a framework for benchmarking large language models using open-source models as judges, offering customizable insights into model performance.
Optimizing LLM Inference with TensorRT: A Comprehensive Guide
Explore how TensorRT-LLM enhances large language model inference by optimizing performance through benchmarking and tuning, offering developers a robust toolset for efficient deployment.
Optimizing LLM Inference Costs: A Comprehensive Guide
Explore strategies for benchmarking large language model (LLM) inference costs, enabling smarter scaling and deployment in the AI landscape, as detailed by NVIDIA's latest insights.
Evaluating Multi-Agent Architectures: A Performance Benchmark
LangChain's new study benchmarks various multi-agent architectures, focusing on their performance and scalability using the Tau-bench dataset, highlighting the advantages of modular systems.