谷歌Ledger:SWE-Bench任务Pass@1提升3.4%
谷歌Ledger运行时在500个SWE-bench任务上将Codex Pass@1提高3.4个百分点,同时降低成本24.4%。
原文链接详细分析
谷歌研究人员推出Ledger运行时,仅记录代理观察、修改和尝试的内容,取代长时任务中膨胀的转录本。在500个SWE-bench任务上,该系统将Codex Pass@1提高3.4个百分点,成本降低24.4%,且无需额外模型调用。使用Gemini-3-Flash在100步仓库任务上测试的SKILL.state设计得分为0.94,消耗65,408个token,而基线消耗1,062,387个token,减少16.2倍,使提示大小随任务延长保持恒定。
Mark
@MRRydonCofounder @AethirCloud | Building Decentralised Cloud Infrastructure (DCI) | Accelerating the world’s transition to universal cloud compute 🌎