Anthropic 发布 Claude Opus 5
摘要
Anthropic 发布 Claude Opus 5,以一半价格提供接近旗舰 Claude Fable 5 的性能,并在多项基准测试中达到新 SOTA。
核心要点
- 在 Frontier-Bench v0.1 上,Opus 5 超越所有其他模型,性能是 Opus 4.8 的两倍以上,且每次任务成本更低。
- 在 ARC-AGI 3 上,Opus 5 的得分是次优模型的三倍。
- 在 Zapier AutomationBench 上,相同成本下 Opus 5 的通过率约为次优模型的 1.5 倍,即使在最低努力设置下也领先。
- 在 OSWorld 2.0 上,Opus 5 以略超三分之一的成本超越 Fable 5 的最佳结果。
- 在生命科学领域,Opus 5 在有机化学任务上比 Opus 4.8 高 10.2 个百分点,在蛋白质相关任务上高 7.7 个百分点。
原文佐证
- It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.
- On Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8’s performance at a lower cost per task.
- Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part.
AI 洞察
Opus 5 的发布标志着 Anthropic 在性能与成本之间找到了更优平衡点,可能吸引大量开发者从其他模型迁移,并加速 AI 在编码、自动化等领域的落地。其“深思熟虑”的能力——如自主构建测试工具——暗示模型正向更自主、更可靠的 agent 方向演进,这将对 AI 应用生态产生深远影响。