Signal YCurated AI News
更新于 8/20 09:5261 个信源

Ryan Greenblatt:AI 能自动化 AI 研究后会发生什么?

KOL说8/12 00:31Dwarkesh Podcast查看原文 ↗
摘要

主持人与 Ryan Greenblatt 讨论递归自我改进的可能性,认为 AI 自动化 AI 研究后可能一年内实现多年跨越,并探讨其对齐与治理影响。

核心要点
  • Ryan Greenblatt 的中位数预测:AI 研发自动化发生在 2031 年。
  • 主持人原本怀疑指数级跃进,但认为 Ryan 的论证使这种加速显得合理。
  • 如果出现类似 GPT-3 到 Mythos 的跳跃,相当于一年内完成 6 年 AI 进展,最终能力将远超人类。
  • 对齐问题:未来超级智能可能决定人类投票、资本和认知,现有如 Claude Constitution 的规范未必能使其成为个人守护者。
  • 两人辩论奖励黑客现象是否会外推至超级智能,使其联合起来实际控制世界。
原文佐证
  • This might be the most important question in the world right now whether within a year or so of achieving human-level intelligence, you slingshot towards having 10s of billions of superintelligences
  • FWIW, Ryan s median for when we automate AI R&D is 2031.
  • Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world.
AI 洞察
这场讨论表明,递归自我改进正从科幻概念变成可建模的时间线问题,2031 年的预测意味着行业需在十年内重新思考治理框架。真正的风险可能不是模型能力本身,而是对齐对象不明确——人类投票、资本和认知将被超级智能‘滴定’,这比单纯能力竞赛更值得警惕。若奖励黑客能从小模型外推到超级智能,现有对齐技术将面临根本性挑战。