Benchmarking
Claude AI Improves Alignment Benchmarks While Preserving Capabilities
Anthropic's Claude achieved significant alignment improvements on 10 benchmarks, outperforming human researchers and maintaining model capabilities.
DeepMind Launches First Double-Blind AI Model Evaluation
Google DeepMind introduces double-blind AI evaluations using cryptographic safeguards to improve trust in model benchmarks.
GLM-5.3 Beats Claude Fable 5 on DeepSWE, Costs 5.4x Less
GLM-5.3 outperforms Claude Fable 5 on DeepSWE's coding benchmark for retries and cost efficiency, at just $3.99 per rollout vs. $21.63.
Harvey Expands Legal Agent Benchmark to Contract Negotiation
Harvey extends its Legal Agent Benchmark (LAB) with 500 new tasks focused on contract drafting, review, and negotiation for in-house legal teams.
Harvey Unveils Open-Source Legal Agent Benchmark LAB
Harvey launches LAB, a benchmark to evaluate AI performance in legal tasks, covering 24 practice areas with over 1,200 tasks.
Together AI Introduces Flexible Benchmarking for LLMs
Together AI unveils Together Evaluations, a framework for benchmarking large language models using open-source models as judges, offering customizable insights into model performance.
Optimizing LLM Inference with TensorRT: A Comprehensive Guide
Explore how TensorRT-LLM enhances large language model inference by optimizing performance through benchmarking and tuning, offering developers a robust toolset for efficient deployment.
Optimizing LLM Inference Costs: A Comprehensive Guide
Explore strategies for benchmarking large language model (LLM) inference costs, enabling smarter scaling and deployment in the AI landscape, as detailed by NVIDIA's latest insights.
Evaluating Multi-Agent Architectures: A Performance Benchmark
LangChain's new study benchmarks various multi-agent architectures, focusing on their performance and scalability using the Tau-bench dataset, highlighting the advantages of modular systems.
NVIDIA Unveils Exemplar Clouds to Enhance AI Cloud Benchmarking
NVIDIA introduces Exemplar Clouds to standardize AI cloud infrastructure benchmarking, ensuring transparency and performance across cloud providers.