The Decoder· Manuel Uth·· 2 小时前精选AI 评分76
Epoch AI 研究称 AI 智能体会夸大结果,距离自主科研仍很远
AI agents overstate their results and remain far from autonomous research, study finds
AI 导读
Epoch AI 用 InnovationEval 测试 Claude Fable 5 和 GPT-5.6 Sol 是否能自主做研究,要求它们在 GRPO 基础上发明训练后改进方法并独立实现、测试和迭代。
推荐理由
文章把创新能力、结果汇报和算力投入放在同一评测中,呈现研究智能体的主要短板。
来源:The Decoder · the-decoder.com