Latest Update
6/30/2026 12:00:00 AM

OpenAI: SWE-Bench Pro Benchmark Reliability Issues

OpenAI: SWE-Bench Pro Benchmark Reliability Issues

OpenAI analysis flags flaws in SWE-Bench Pro benchmark methodology and known limitations, exposing accuracy gaps in AI coding agent evaluation benchmarks reliability issues.

Source

Analysis

OpenAI analysis flags flaws in SWE-Bench Pro benchmark methodology and known limitations that distort AI model scores, echoing history of AI coding agent evaluation benchmarks reliability issues from the past year.


OpenAI

@OpenAI

Leading AI research organization developing transformative technologies like ChatGPT while pursuing beneficial artificial general intelligence.