唐杰
立场时间线 6 条 · 6 个议题
同一议题按时间排列,说法变化一目了然。
参数量单独讨论已失去意义,必须结合数据量、计算投入和部署运行条件一起衡量
“Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions.(参数规模只有与其他三个因素结合才有意义——你拥有多少数据、打算在哪里投入计算,以及谁在什么条件下运行模型。)”
来源访谈 →GLM-5.3相比前代架构并无根本改变,性能飞跃完全来自长周期高复杂度环境中的强化学习
“在基础模型架构上并无根本改变,其显著性能提升完全源于在长周期、高复杂度环境中的强化学习(RL)”
来源访谈 →训练环境可完全合成、大规模自动化生成,扩展后训练的难度重心已从模型转移到环境本身
“As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.(随着智能体能力的提升,扩展后训练的大部分难度从模型转移到了环境本身。)”
来源访谈 →模型扩展应看五个关键维度,除参数量外还包括数据、计算、部署条件以及MoE稀疏度
“提出了模型扩展的五个关键维度,除了传统的参数量,还包括数据、计算、部署条件等。他特别提到了混合专家模型(MoE)的稀疏度,并引入了新的“XA-YB”表示法来规范描述”
来源访谈 →一旦跨过知识容量阈值,长程推理等高级技能与总参数量关系不大,更多取决于训练质量和推理时的计算分配
“Advanced skills... require carrying long causal chains (20+ inference steps) without losing the thread. This ability does not live in total parameter count once a certain knowledge-holding threshold is reached.(高级技能……需要在不偏离主线的情况下进行长达20步以上的长因果链推理。一旦模型达到某个知识容量阈值,这种能力就不再体现在总参数量中了。)”
来源访谈 →金句墙 3 条
“Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions.”
参数规模只有与其他三个因素结合才有意义——你拥有多少数据、打算在哪里投入计算,以及谁在什么条件下运行模型。
AI模型进入后训练时代:参数规模不再是唯一标尺 · 2026/8/20“Advanced skills... require carrying long causal chains (20+ inference steps) without losing the thread. This ability does not live in total parameter count once a certain knowledge-holding threshold is reached.”
高级技能……需要在不偏离主线的情况下进行长达20步以上的长因果链推理。一旦模型达到某个知识容量阈值,这种能力就不再体现在总参数量中了。
AI模型进入后训练时代:参数规模不再是唯一标尺 · 2026/8/20“As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.”
随着智能体能力的提升,扩展后训练的大部分难度从模型转移到了环境本身。
AI模型进入后训练时代:参数规模不再是唯一标尺 · 2026/8/20