Hugging Face Daily Papers·· 2 天前AI 评分58
Galahad 用字节精确记忆让 LLM 阅读成为一次性成本
Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost
AI 导读
Galahad 是面向 vLLM、SGLang 和 this http URL 的记忆层,旨在把 LLM 对相同文本的重复阅读变成一次性成本。论文称,在 7 个真实数据集上,98.7% 的 prompt tokens 是模型已读过的文本;Taliesin 保存文本块的 KV 状态并在相同字节再次出现时加载,Blaise 保存文档并只向模型传入问题所需片段。
来源:Hugging Face Daily Papers · arxiv.org