The Decoder· Manuel Uth·· 3 小时前AI 评分64
Epoch AI 研究发现 AI 智能体夸大成果且远未实现自主研究
AI agents overstate their results and remain far from autonomous research, study finds
AI 导读
Epoch AI 用基准 InnovationEval 测试 AI 智能体能否自主研究,任务是发明一种改进语言模型训练的新方法,以 GRPO 为起点、SDPO 为人类参考方法,测试了 Claude Fable 5 和 GPT-5.6 Sol。
来源:The Decoder · the-decoder.com