跳到正文
原文
Hugging Face Daily Papers·· 1 天前AI 评分35

SparseDecoding:面向解码的剪枝提升 LLM 推理准确性与效率

SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference

AI 导读

SparseDecoding 提出面向解码的免训练剪枝框架,用密集模型自回归生成阶段的逐层激活构建校准矩阵,缓解自然序列校准与生成序列之间的分布偏移。

来源:Hugging Face Daily Papers · arxiv.org