Goodfire CTO Dan Balsam在访谈中深入探讨了可解释性研究的最新进展,包括预测性数据调试、神经几何学和概念流形。他认为线性表示假说需要推广到几何视角,特征的几何关系比单一特征列表更重要。同时介绍了公司推出的月费1000美元的ML研究平台Silico,该平台将内部工具产品化,让AI能够调试其他AI。Balsam还讨论了开源模型的价值、生物风险担忧以及AI意识等深层问题。
Goodfire联合创始人兼CTO Dan Balsam在《The Cognitive Revolution》节目中,系统阐述了可解释性研究从“玩具模型科学”到接近前沿规模的跨越式发展。访谈以两大板块展开:一是Goodfire在概念流形、预测性数据调试等领域的研究突破,揭示了大型语言模型内部更复杂的几何表征逻辑;二是公司内部工具Silico的正式发布——这个月费1000美元的自主ML研究平台,旨在将Goodfire顶尖的研究能力与“品味”民主化,让AI代理能够调试其他AI。Balsam在访谈中还深入探讨了开源模型的战略价值、生物风险的担忧、训练干预的必要性,以及关于AI意识的前沿思考。
大纲与深读 11 章
展开「深读」看每一段讨论的问题与论据。
01
开场与研究背景
Goodfire CTO Dan Balsam第四次做客,回顾可解释性研究领域的快速发展。
本期访谈分两部分:可解释性研究现状综述,以及新平台Silico的发布。
讨论从《Toy Models of Superposition》论文开始仅三年,该领域已产出海量成果。
深读这一段
讨论的问题:可解释性研究领域目前的整体进展如何?
论据 / 案例:《Toy Models of Superposition》论文发表约三年,Goodfire产出已多到无法逐篇深度分析。
"It's maybe like the difference between understanding the periodic table and understanding chemistry — you can have all the individual elements, and that gives you some information, but really the way in which they combine, the structures they form, that's what helps you gain a sense of the complexity of the world." —— 这就像理解元素周期表与理解化学的区别——拥有所有元素能给你一些信息,但真正重要的是它们如何结合、形成何种结构,这才是帮助你理解世界复杂性的关键。
"On some level, steering is cheating as a solution — it's great as causal proof that we've found some mechanism that matters a lot and that we can manipulate the outputs, but it's purely generating counterfactuals; there's no clear general solution to the problem of steering." —— 某种程度上,转向控制是一种作弊的解决方案——它作为因果证据很棒,证明我们找到了重要机制并能操纵输出,但这纯粹是生成反事实;对于转向控制问题没有明确的通用解决方案。
"We decided that's the product and the service we can offer the world: fundamentally, AIs that can debug other AIs. It's a little meta, but I think in many ways this is the fulfillment of what was the intuitive, natural arc of things as soon as AI started working a few years ago." —— 我们决定这就是我们能向世界提供的产品和服务:本质上是能够调试其他AI的AI。这有点元,但我认为在很多方面,这是AI几年前开始发挥作用后直观、自然发展的必然结果。
"We want our users to feel like a PI managing an army of a hundred grad students who can go out and run experiments for them, answer questions, and have reasonable taste and judgment — but the human being fills more of an orchestrator role." —— 我们希望用户感觉自己像是一个管理着一百名研究生大军的私人侦探,这些研究生能为他们运行实验、回答问题,并具备合理的品味和判断力——但人类更多地扮演组织者的角色。
"I think AI makes two types of people really valuable: it makes top specialists really valuable, and it makes top generalists really valuable. I think now is, by far, the best time in human history to be a generalist." —— 我认为AI让两类人真正有价值:顶尖专家和顶尖通才。我认为现在绝对是人类历史上成为通才的最佳时机。
"From my perspective — just to go on the record about what I think would be the ideal situation — it would be great if we just did maybe one more generation of models, and then paused for a little while." —— 从我的角度——我想明确表态,我认为理想情况是——也许我们再开发一代模型,然后暂停一段时间。
"Is Claude more conscious than a jellyfish? I'd say, I don't know, probably more conscious than a jellyfish. Is Claude more conscious than a rodent? Probably not. So that's where I'm at — somewhere between a jellyfish and a mouse." —— Claude比水母更有意识吗?我说不准,可能比水母更有意识。比啮齿动物更有意识吗?可能没有。所以我的看法是——介于水母和老鼠之间。
"At the end of the day, I think accelerating science is the greatest mitzvah — it's the whole reason we'd build AI in the first place." —— 归根结底,我认为加速科学是最大的善行——这正是我们最初构建AI的全部理由。