List of AI News about evaluation
| Time | Details |
|---|---|
|
2026-09-16 14:40 |
AI benchmarks Crisis Exposed by New Paper
According to @emollick, public AI benchmarks are saturated or error ridden, severely underestimating current model abilities, as the cited paper shows. |
|
2026-09-08 14:04 |
METR Long-Horizon Benchmarks Face Saturation
According to emollick, METR’s long-horizon metric may be saturated as studies report 18+ weeks of agent work from Fable in harnesses. |
|
2026-09-06 15:17 |
DeepLearningAI Reveals 5 Coding-Agent Skills
According to DeepLearningAI, top teams curb agent autonomy with human review, iterative plans, and verification to cut costs and failures. |
|
2026-08-16 17:33 |
Human Judgments Drive AI Evaluation Breakthrough
According to emollick, human opinions are essential for benchmarking non-verifiable AI tasks, urging adoption of qualitative research methods. |
|
2026-08-01 18:19 |
LLMs Improve Across Unverifiable Tasks
According to @emollick, LLMs advance in formal domains also lift performance in less-verifiable fields, easing concerns over non-checkable answers. |
|
2026-07-29 15:45 |
AI Behavioral Observatory Enables Valid Prompt Testing
According to @emollick, the AI Behavioral Observatory open source tool runs statistically valid tests on model behavior across prompts. |
|
2026-07-18 23:24 |
Kimi K3 Reveals English Chain-of-Thought Bias
According to emollick, Kimi K3 used 95.5% English in its chain-of-thought for a Chinese prompt, signaling training bias and evaluation gaps. |
|
2026-07-09 04:13 |
OpenAI Benchmarks Shake-up Spurs 2026 Analysis
According to emollick, OpenAI questioned coding evals yet has not shared GPT5.6 GDPval, raising transparency and capability tracking concerns. |
|
2026-07-08 20:45 |
OpenAI Audits SWE-Bench Pro, flags 70% noise
According to @OpenAI, SWE-Bench Pro shows a ~70% noise ceiling and no longer reliably measures frontier coding capability, urging a shift to new evals. |
|
2026-07-07 01:41 |
Ethan Mollick Says prompting tricks fade, management wins
According to @emollick, prompt hacks add little value; define goals, outputs, quality bars, and tests to get ROI from AI systems. |
|
2026-06-16 17:23 |
OpenAI Evals Reform Guides Next Benchmarks
According to OpenAI on X, leaders discuss better evals to forecast model progress as saturated benchmarks get gamed, outlining next judgment areas. |
|
2026-06-03 16:38 |
Leni AI Elevates Verified Investment Analysis
According to @godofprompt, Leni targets finance with verified outputs and rigorous reasoning to cut costly errors in institutional investment workflows. |
|
2026-05-16 12:57 |
iFixAI Launches 32-Test Safety Score for AI
According to @godofprompt, iFixAI runs 32 tests on deployed AI to score hallucinations, manipulation, and consistency, offering actionable evals for teams. |
|
2026-04-30 16:14 |
GPT5.5 Tops Benchmarks yet Misfires Often
According to @godofprompt, AA-Omniscience shows GPT-5.5 ranks highest for smarts but is most confidently wrong when penalized for guessing. |
|
2026-04-28 13:25 |
Test of Time LLM Debuts With Retro Benchmark Fun
According to @soumithchintala, Test of Time LLM offers a playful, retro-style benchmark link, highlighting community interest in evaluators. |
|
2026-04-21 19:12 |
LLM Judge Bias Exposed: New Position Bias Benchmark Shows Up To 66% Flip Rate — 2026 Analysis
According to Ethan Mollick on X (Twitter), large language models used as judges display significant position bias, with judgments flipping when answer order is swapped; he cites Lech Mazur’s New LLM Position Bias Benchmark showing a median 45% flip rate on decisive pairs and a reported 66% flip rate for GPT-5.4 (as reported by Lech Mazur’s thread and benchmark summary). According to Mollick, simple presentation changes materially alter outcomes, indicating current LLM-as-judge pipelines remain unreliable without controls (as reported by Ethan Mollick). According to Lech Mazur, mitigation via better harnessing—multiple judging runs, randomized order, and aggregation—can reduce variance, suggesting practical steps for enterprise evaluation workflows and AI product A/B testing. Business impact: according to Mollick’s post, organizations relying on LLM judges for qualitative assessments (creative scoring, code review, search ranking, and RLHF data curation) should add randomized comparisons, majority voting, and calibration audits to improve consistency and reduce bias-induced risk. |
|
2026-04-07 15:44 |
Google AI Overviews Accuracy Debate: 90 Percent Success, 10 Percent Risk — Analysis of Measurement Challenges and Business Impact
According to @emollick referencing The New York Times by Mike Isaac, Google’s AI Overviews show roughly 90 percent accuracy but a consequential 10 percent error rate at Google’s multi‑trillion annual search scale, highlighting why evaluating AI quality is hard when identical errors also exist in sources like Wikipedia and source traceability is weaker in AI answers. As reported by The New York Times, the case study shows that AI Overviews can surface useful synthesized answers that many users might not find on their own, yet inconsistent citation visibility complicates verification and accountability. According to The New York Times, this creates operational risk for publishers, brands, and advertisers that rely on search accuracy, while opening opportunities for enterprise evaluation tooling, retrieval‑augmented generation pipelines with explicit citation, and content provenance standards to improve auditability. |
|
2026-04-01 00:20 |
AI Content Literacy: Why Doom-Laden News Distorts Reality — Analysis for 2026 AI Safety, Policy, and Product Teams
According to Yann LeCun on X, resharing Steven Pinker’s video on media negativity bias highlights how selective bad-news framing skews public risk perception; for AI builders, this underscores the need for calibrated communication and evidence-based benchmarks in AI safety, deployment metrics, and policy debates (as reported by the linked YouTube video from Steven Pinker). According to Steven Pinker’s YouTube presentation, negative selection and availability bias make people overestimate systemic collapse, a dynamic that can also distort narratives around AI risk, automation impact, and model failures; AI teams can counter this by publishing longitudinal reliability data, post-deployment incident rates, and audited evaluation suites. As reported by the original X post from Yann LeCun, reframing with trend data can improve stakeholder trust; AI companies can apply this by standardizing model cards, red-teaming disclosures, and quarterly safety and performance reports tied to concrete baselines. |
|
2026-03-29 08:44 |
Latest Analysis: New arXiv Paper Explores AI Methodology and Performance Benchmarks
According to God of Prompt on Twitter, a new AI research paper was posted on arXiv at arxiv.org/abs/2603.23420. However, the tweet and link preview do not provide the title, authors, model names, datasets, or methods. As reported by arXiv via the shared URL, only the identifier is available publicly at the time of writing, so concrete findings, benchmarks, or business implications cannot be verified without the paper’s details. According to best practices for AI due diligence, companies should review the arXiv abstract and PDF to confirm the task scope, model architecture, training data, evaluation metrics, and licenses before considering pilots or partnerships. |
|
2026-03-27 10:57 |
Latest Analysis: New ArXiv 2603.23234 Paper on AI Model Advances and 2026 Trends
According to @godofprompt, a new paper was shared at arxiv.org/abs/2603.23234. However, as reported by arXiv, the linked identifier cannot be verified at this time. Without an accessible abstract or PDF, no technical claims, benchmarks, datasets, or model details can be confirmed, and no business impact can be assessed. According to best-practice editorial standards, readers should consult the original arXiv entry for the title, authors, and methods before drawing conclusions or acting on potential market opportunities. |