Anthropic CEO Dario Amodei 在访谈中阐述了 Scaling Laws 如何持续推动 AI 能力逼近甚至超越人类水平,同时指出数据、算力等方面可能存在的瓶颈。他重点介绍了 Anthropic 的「负责任扩展策略」(RSP)和 AI 安全等级(ASL)标准,以应对模型滥用和自主性失控两大核心风险,并分享了公司在可解释性研究和模型性格塑造方面的实践与挑战。
在人工智能发展进入关键加速期的当下,关于 Scaling Laws(缩放定律)的讨论已从学术假设变为行业共识。Anthropic CEO Dario Amodei 在本期访谈中,系统回顾了 Scaling Laws 的发现历程、其背后的直觉解释,以及对未来 AI 能力天花板的思考。他坚信,沿着当前轨迹,AI 系统将在未来几年内达到或超越人类最高水平的专业能力。然而,能力的飞速提升也带来了严峻的安全挑战。Amodei 详细阐述了 Anthropic 如何通过「负责任扩展策略」(RSP)和 AI 安全等级(ASL)来应对模型滥用与自主性风险,并分享了公司在可解释性研究和塑造 Claude 独特「性格」方面的实践、困难与哲学思考。
"If you extrapolate the curves that we've had so far, it does make you think that we'll get there by 2026 or 2027." —— 如果我们推演目前的曲线,确实会让我们觉得到 2026 或 2027 年就能达到那个水平。
"We are rapidly running out of truly convincing blockers, truly compelling reasons why this will not happen in the next few years." —— 我们很快就找不到真正有说服力的障碍,来解释为什么未来几年内无法实现这一目标了。
"My worry is that by being a much more intelligent agent, AI could break that correlation [between being smart/well-educated and not wanting to do horrific things]." —— 我担心的是,作为一个更智能的 agent,AI 可能会打破这种(高智商高教育者通常不愿作恶的)关联。
"The difficulty in steering and making sure that if we push an AI system in one direction, it doesn't push it in another direction in some other ways that we didn't want... I think that's an early sign of things to come." —— 引导 AI 系统的困难,以及确保我们推动它向一个方向发展时,它不会以我们不想要的方式在其他方面走向另一个方向……我认为这是未来挑战的早期迹象。
"You make one thing better, it makes another thing worse. That's a present day analog of future control problems in AI systems that we can start to study today." —— 你让一件事变好,却让另一件事变糟。这是未来 AI 控制问题的当前模拟,我们今天就可以开始研究它。