AI News List

List of AI News about benchmarks

Time Details
2026-07-25
04:41
BenchBenchBench Reveals AI Benchmarking Meta-Analysis

According to emollick, a Codex-prompted project produced BBBBB, an executable benchmark testing AI-authored conformance suites for metrics.

Source
2026-07-24
17:47
Claude Opus 5 Breakthrough beats rivals at half cost

According to TheRundownAI, Anthropic’s Claude Opus 5 leaps past 4.8 and outperforms Fable on several benchmarks at half the price.

Source
2026-07-17
09:18
Moonshot Kimi K3 Tops Benchmarks, Trails Leaders

According to @CNBC, Moonshot AI says Kimi K3 beats top OpenAI and Anthropic models on select benchmarks while still trailing overall leaders.

Source
2026-07-08
18:37
Grok 4.5 powers Cursor launch with flash speed

According to The Rundown AI, Grok 4.5 in SpaceXAI and Cursor delivers flash-level speed, strong benchmarks, and affordable pricing, boosting dev workflows.

Source
2026-07-06
11:07
Stanford AI Lab unveils ICML 2026 highlights

According to StanfordAILab, Stanford AI Lab lists ICML 2026 papers on coding agents, LLM reasoning, safety, interpretability, and science.

Source
2026-06-26
22:15
Zhipu AI narrows gap with Anthropic, OpenAI

According to CNBC... Zhipu’s GLM models gain on OpenAI and Anthropic as open source and export limits reshape AI competition, per benchmarks and funding data.

Source
2026-06-17
20:41
LifeSciBench Launches 750-task Benchmark Analysis

According to OpenAI... LifeSciBench debuts with 750 expert tasks across 7 workflows to assess AI for real-world life science research.

Source
2026-06-16
17:23
OpenAI Evals Reform Guides Next Benchmarks

According to OpenAI on X, leaders discuss better evals to forecast model progress as saturated benchmarks get gamed, outlining next judgment areas.

Source
2026-06-15
15:44
Claude3.5 Crushes benchmark rankings

According to God of Prompt, Anthropic is crushing a new benchmark, signaling Claude3.5 gains for reasoning and eval leadership.

Source
2026-06-11
14:17
Hugging Face revives Papers With Code datasets

According to KyeGomezB, Hugging Face acquired Papers With Code domain and datasets, restoring access researchers used for benchmarking and discovery.

Source
2026-06-10
12:54
Project Tapestry Unites Open AI Research

According to @ylecun, Project Tapestry invites researchers to collaborate on open AI benchmarks and tooling, as reported by The Alliance for OpenAI.

Source
2026-06-09
18:10
Claude Fable 5 Tops SOTA Benchmarks, Big Leap

According to karpathy, Claude Fable 5 adds safeguards to Mythos and achieves SOTA across benchmarks, excelling at long, complex problem solving.

Source
2026-06-09
18:10
Claude Fable 5 Achieves SOTA Benchmarks

According to karpathy, Claude Fable 5 posts SOTA scores and excels at long, difficult problem solving with added safeguards versus Mythos.

Source
2026-05-26
14:57
Model Routers Unlock Real-World Wins

According to God of Prompt, routers that pick models by product-specific evals beat chasing generic benchmarks.

Source
2026-05-19
17:59
Gemini 3.5 Flash Delivers 4x Speed Breakthrough

According to sundarpichai, Gemini 3.5 Flash is live, 4x faster than frontier models and outperforms 3.1 Pro on most benchmarks, with major coding gains.

Source
2026-05-19
17:53
Gemini 3.5 Flash Breakthrough beats 3.1 Pro

According to @OriolVinyalsML, Gemini 3.5 Flash launches with frontier-level intelligence and faster speed, outperforming 3.1 Pro on most benchmarks.

Source
2026-05-09
01:32
Claude Mythos Preview hits 16hr eval window

According to @emollick, METR estimated a 50% time horizon of 16hrs for Claude Mythos Preview risk tasks, signaling upper-bound capability growth.

Source
2026-05-05
23:10
GPQA Benchmark Shows GPT 5.5 Instant Leap

According to emollick, OpenAI’s free GPT 5.5 Instant matches late-2025 paid model levels on GPQA, signaling rapid capability gains.

Source
2026-05-03
22:10
Artificial Analysis index debated in 2026

According to emollick, AA index compares models but lacks trend value; chatgpt21 projects GPT at 90 by 2029 using conservative gains.

Source
2026-04-30
16:14
GPT5.5 Tops Benchmarks yet Misfires Often

According to @godofprompt, AA-Omniscience shows GPT-5.5 ranks highest for smarts but is most confidently wrong when penalized for guessing.

Source