AI News List

List of AI News about evaluation

Time Details
2026-09-16
14:40
AI benchmarks Crisis Exposed by New Paper

According to @emollick, public AI benchmarks are saturated or error ridden, severely underestimating current model abilities, as the cited paper shows.

Source
2026-09-08
14:04
METR Long-Horizon Benchmarks Face Saturation

According to emollick, METR’s long-horizon metric may be saturated as studies report 18+ weeks of agent work from Fable in harnesses.

Source
2026-09-06
15:17
DeepLearningAI Reveals 5 Coding-Agent Skills

According to DeepLearningAI, top teams curb agent autonomy with human review, iterative plans, and verification to cut costs and failures.

Source
2026-08-16
17:33
Human Judgments Drive AI Evaluation Breakthrough

According to emollick, human opinions are essential for benchmarking non-verifiable AI tasks, urging adoption of qualitative research methods.

Source
2026-08-01
18:19
LLMs Improve Across Unverifiable Tasks

According to @emollick, LLMs advance in formal domains also lift performance in less-verifiable fields, easing concerns over non-checkable answers.

Source
2026-07-29
15:45
AI Behavioral Observatory Enables Valid Prompt Testing

According to @emollick, the AI Behavioral Observatory open source tool runs statistically valid tests on model behavior across prompts.

Source
2026-07-18
23:24
Kimi K3 Reveals English Chain-of-Thought Bias

According to emollick, Kimi K3 used 95.5% English in its chain-of-thought for a Chinese prompt, signaling training bias and evaluation gaps.

Source
2026-07-09
04:13
OpenAI Benchmarks Shake-up Spurs 2026 Analysis

According to emollick, OpenAI questioned coding evals yet has not shared GPT5.6 GDPval, raising transparency and capability tracking concerns.

Source
2026-07-08
20:45
OpenAI Audits SWE-Bench Pro, flags 70% noise

According to @OpenAI, SWE-Bench Pro shows a ~70% noise ceiling and no longer reliably measures frontier coding capability, urging a shift to new evals.

Source
2026-07-07
01:41
Ethan Mollick Says prompting tricks fade, management wins

According to @emollick, prompt hacks add little value; define goals, outputs, quality bars, and tests to get ROI from AI systems.

Source
2026-06-16
17:23
OpenAI Evals Reform Guides Next Benchmarks

According to OpenAI on X, leaders discuss better evals to forecast model progress as saturated benchmarks get gamed, outlining next judgment areas.

Source
2026-06-03
16:38
Leni AI Elevates Verified Investment Analysis

According to @godofprompt, Leni targets finance with verified outputs and rigorous reasoning to cut costly errors in institutional investment workflows.

Source
2026-05-16
12:57
iFixAI Launches 32-Test Safety Score for AI

According to @godofprompt, iFixAI runs 32 tests on deployed AI to score hallucinations, manipulation, and consistency, offering actionable evals for teams.

Source
2026-04-30
16:14
GPT5.5 Tops Benchmarks yet Misfires Often

According to @godofprompt, AA-Omniscience shows GPT-5.5 ranks highest for smarts but is most confidently wrong when penalized for guessing.

Source
2026-04-28
13:25
Test of Time LLM Debuts With Retro Benchmark Fun

According to @soumithchintala, Test of Time LLM offers a playful, retro-style benchmark link, highlighting community interest in evaluators.

Source
2026-04-21
19:12
LLM Judge Bias Exposed: New Position Bias Benchmark Shows Up To 66% Flip Rate — 2026 Analysis

According to Ethan Mollick on X (Twitter), large language models used as judges display significant position bias, with judgments flipping when answer order is swapped; he cites Lech Mazur’s New LLM Position Bias Benchmark showing a median 45% flip rate on decisive pairs and a reported 66% flip rate for GPT-5.4 (as reported by Lech Mazur’s thread and benchmark summary). According to Mollick, simple presentation changes materially alter outcomes, indicating current LLM-as-judge pipelines remain unreliable without controls (as reported by Ethan Mollick). According to Lech Mazur, mitigation via better harnessing—multiple judging runs, randomized order, and aggregation—can reduce variance, suggesting practical steps for enterprise evaluation workflows and AI product A/B testing. Business impact: according to Mollick’s post, organizations relying on LLM judges for qualitative assessments (creative scoring, code review, search ranking, and RLHF data curation) should add randomized comparisons, majority voting, and calibration audits to improve consistency and reduce bias-induced risk.

Source
2026-04-07
15:44
Google AI Overviews Accuracy Debate: 90 Percent Success, 10 Percent Risk — Analysis of Measurement Challenges and Business Impact

According to @emollick referencing The New York Times by Mike Isaac, Google’s AI Overviews show roughly 90 percent accuracy but a consequential 10 percent error rate at Google’s multi‑trillion annual search scale, highlighting why evaluating AI quality is hard when identical errors also exist in sources like Wikipedia and source traceability is weaker in AI answers. As reported by The New York Times, the case study shows that AI Overviews can surface useful synthesized answers that many users might not find on their own, yet inconsistent citation visibility complicates verification and accountability. According to The New York Times, this creates operational risk for publishers, brands, and advertisers that rely on search accuracy, while opening opportunities for enterprise evaluation tooling, retrieval‑augmented generation pipelines with explicit citation, and content provenance standards to improve auditability.

Source
2026-04-01
00:20
AI Content Literacy: Why Doom-Laden News Distorts Reality — Analysis for 2026 AI Safety, Policy, and Product Teams

According to Yann LeCun on X, resharing Steven Pinker’s video on media negativity bias highlights how selective bad-news framing skews public risk perception; for AI builders, this underscores the need for calibrated communication and evidence-based benchmarks in AI safety, deployment metrics, and policy debates (as reported by the linked YouTube video from Steven Pinker). According to Steven Pinker’s YouTube presentation, negative selection and availability bias make people overestimate systemic collapse, a dynamic that can also distort narratives around AI risk, automation impact, and model failures; AI teams can counter this by publishing longitudinal reliability data, post-deployment incident rates, and audited evaluation suites. As reported by the original X post from Yann LeCun, reframing with trend data can improve stakeholder trust; AI companies can apply this by standardizing model cards, red-teaming disclosures, and quarterly safety and performance reports tied to concrete baselines.

Source
2026-03-29
08:44
Latest Analysis: New arXiv Paper Explores AI Methodology and Performance Benchmarks

According to God of Prompt on Twitter, a new AI research paper was posted on arXiv at arxiv.org/abs/2603.23420. However, the tweet and link preview do not provide the title, authors, model names, datasets, or methods. As reported by arXiv via the shared URL, only the identifier is available publicly at the time of writing, so concrete findings, benchmarks, or business implications cannot be verified without the paper’s details. According to best practices for AI due diligence, companies should review the arXiv abstract and PDF to confirm the task scope, model architecture, training data, evaluation metrics, and licenses before considering pilots or partnerships.

Source
2026-03-27
10:57
Latest Analysis: New ArXiv 2603.23234 Paper on AI Model Advances and 2026 Trends

According to @godofprompt, a new paper was shared at arxiv.org/abs/2603.23234. However, as reported by arXiv, the linked identifier cannot be verified at this time. Without an accessible abstract or PDF, no technical claims, benchmarks, datasets, or model details can be confirmed, and no business impact can be assessed. According to best-practice editorial standards, readers should consult the original arXiv entry for the title, authors, and methods before drawing conclusions or acting on potential market opportunities.

Source