Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
- 类型:arxiv
- 标识:2607.17524
- 链接:https://arxiv.org/abs/2607.17524
- 主分类:engineering
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, while avoiding the pitfalls that come with directly training the model to generate off-policy tokens. Experiments on document summarization tasks show that TOPL achieves strong out-of-distribution generalization across 11 datasets against a diverse set of sequence-level and token-lev
- 待LLM分类:否
- 标题中文:分布偏移下忠实生成的 token 级离线策略学习
- TLDR中文:本文提出 Token-Level Off-Policy Labeling (TOPL),一种将后训练重构为 token 级正确性预测任务的离线策略训练范式。其核心思路是:通过训练模型区分响应中的好 token 与坏 token,自然引导模型生成好 token,同时避免直接训练模型生成离线策略 token 所带来的缺陷。在文档摘要任务上的实验表明,TOPL 在 11 个数据集上针对多种序列级与 token 级方法实现了强大的分布外泛化能力。
- 来源文件:
- /inbox/tom/_candidates/2026-07-21-agent-rag-longcontext-candidates.json
- [S2 enrich]