跳到正文
原文
Hugging Face Daily Papers·· 2 天前AI 评分29

多教师在线蒸馏中如何组合各教师学到的内容:Teacher-Relative Shifts 方法

Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

AI 导读

研究提出 Δ-MOPD,在多教师在线蒸馏中迁移 teacher-minus-base logit shift,并重新锚定到学生模型冻结初始化。实验显示,3 个组合教师下 Δ-MOPD 比 endpoint composition 高 $4.11$ Math 和 $1.95$ five-benchmark points;阶段式路由中将顺序差距从 $10.50$ 降至 $6.42$ points。

来源:Hugging Face Daily Papers · arxiv.org