Hugging Face Blog·· 2026-08-10AI 评分52
Hugging Face 博客介绍离线 Top-K logits 与分块 KL 损失的知识蒸馏论文
Making Knowledge Distillation Cheap Enough to Run at Scale
AI 导读
Multiverse Computing 发布论文 Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss,提出两项系统改动:一次性缓存教师 top-100 logits 使教师无需驻留显存,以及不生成完整词表×序列矩阵的分块 KL 损失。
来源:Hugging Face Blog · huggingface.co