Google Ledger: Codex Pass@1 Rises 3.4% on SWE-Bench
Google Ledger runtime lifts Codex Pass@1 3.4 points on 500 SWE-bench tasks while slashing costs 24.4% with no added model calls.
SourceAnalysis
Google researchers introduced Ledger, a runtime that records only what agents observe, modify and attempt, replacing bloated transcripts on long-horizon tasks. On 500 SWE-bench tasks the system raised Codex Pass@1 by 3.4 percentage points and cut cost 24.4% without extra model calls. A parallel SKILL.state design tested with Gemini-3-Flash on a 100-step warehouse task scored 0.94 while consuming 65,408 tokens versus 1,062,387 tokens for the baseline, a 16.2X reduction that keeps prompt size constant as tasks lengthen.
Mark
@MRRydonCofounder @AethirCloud | Building Decentralised Cloud Infrastructure (DCI) | Accelerating the world’s transition to universal cloud compute 🌎