List of AI News about benchmarks
| Time | Details |
|---|---|
|
2026-07-25 04:41 |
BenchBenchBench Reveals AI Benchmarking Meta-Analysis
According to emollick, a Codex-prompted project produced BBBBB, an executable benchmark testing AI-authored conformance suites for metrics. |
|
2026-07-24 17:47 |
Claude Opus 5 Breakthrough beats rivals at half cost
According to TheRundownAI, Anthropic’s Claude Opus 5 leaps past 4.8 and outperforms Fable on several benchmarks at half the price. |
|
2026-07-17 09:18 |
Moonshot Kimi K3 Tops Benchmarks, Trails Leaders
According to @CNBC, Moonshot AI says Kimi K3 beats top OpenAI and Anthropic models on select benchmarks while still trailing overall leaders. |
|
2026-07-08 18:37 |
Grok 4.5 powers Cursor launch with flash speed
According to The Rundown AI, Grok 4.5 in SpaceXAI and Cursor delivers flash-level speed, strong benchmarks, and affordable pricing, boosting dev workflows. |
|
2026-07-06 11:07 |
Stanford AI Lab unveils ICML 2026 highlights
According to StanfordAILab, Stanford AI Lab lists ICML 2026 papers on coding agents, LLM reasoning, safety, interpretability, and science. |
|
2026-06-26 22:15 |
Zhipu AI narrows gap with Anthropic, OpenAI
According to CNBC... Zhipu’s GLM models gain on OpenAI and Anthropic as open source and export limits reshape AI competition, per benchmarks and funding data. |
|
2026-06-17 20:41 |
LifeSciBench Launches 750-task Benchmark Analysis
According to OpenAI... LifeSciBench debuts with 750 expert tasks across 7 workflows to assess AI for real-world life science research. |
|
2026-06-16 17:23 |
OpenAI Evals Reform Guides Next Benchmarks
According to OpenAI on X, leaders discuss better evals to forecast model progress as saturated benchmarks get gamed, outlining next judgment areas. |
|
2026-06-15 15:44 |
Claude3.5 Crushes benchmark rankings
According to God of Prompt, Anthropic is crushing a new benchmark, signaling Claude3.5 gains for reasoning and eval leadership. |
|
2026-06-11 14:17 |
Hugging Face revives Papers With Code datasets
According to KyeGomezB, Hugging Face acquired Papers With Code domain and datasets, restoring access researchers used for benchmarking and discovery. |
|
2026-06-10 12:54 |
Project Tapestry Unites Open AI Research
According to @ylecun, Project Tapestry invites researchers to collaborate on open AI benchmarks and tooling, as reported by The Alliance for OpenAI. |
|
2026-06-09 18:10 |
Claude Fable 5 Tops SOTA Benchmarks, Big Leap
According to karpathy, Claude Fable 5 adds safeguards to Mythos and achieves SOTA across benchmarks, excelling at long, complex problem solving. |
|
2026-06-09 18:10 |
Claude Fable 5 Achieves SOTA Benchmarks
According to karpathy, Claude Fable 5 posts SOTA scores and excels at long, difficult problem solving with added safeguards versus Mythos. |
|
2026-05-26 14:57 |
Model Routers Unlock Real-World Wins
According to God of Prompt, routers that pick models by product-specific evals beat chasing generic benchmarks. |
|
2026-05-19 17:59 |
Gemini 3.5 Flash Delivers 4x Speed Breakthrough
According to sundarpichai, Gemini 3.5 Flash is live, 4x faster than frontier models and outperforms 3.1 Pro on most benchmarks, with major coding gains. |
|
2026-05-19 17:53 |
Gemini 3.5 Flash Breakthrough beats 3.1 Pro
According to @OriolVinyalsML, Gemini 3.5 Flash launches with frontier-level intelligence and faster speed, outperforming 3.1 Pro on most benchmarks. |
|
2026-05-09 01:32 |
Claude Mythos Preview hits 16hr eval window
According to @emollick, METR estimated a 50% time horizon of 16hrs for Claude Mythos Preview risk tasks, signaling upper-bound capability growth. |
|
2026-05-05 23:10 |
GPQA Benchmark Shows GPT 5.5 Instant Leap
According to emollick, OpenAI’s free GPT 5.5 Instant matches late-2025 paid model levels on GPQA, signaling rapid capability gains. |
|
2026-05-03 22:10 |
Artificial Analysis index debated in 2026
According to emollick, AA index compares models but lacks trend value; chatgpt21 projects GPT at 90 by 2029 using conservative gains. |
|
2026-04-30 16:14 |
GPT5.5 Tops Benchmarks yet Misfires Often
According to @godofprompt, AA-Omniscience shows GPT-5.5 ranks highest for smarts but is most confidently wrong when penalized for guessing. |