onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction
- 类型:arxiv
- 标识:2609.24983
- 链接:https://arxiv.org/abs/2609.24983
- 主分类:agent
- 形态:position
- TLDR:We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model response, the annotator locates the first inappropriate token and either picks a substitute from the model's candidate tokens or types the correct text via free-form editing. The system then truncates everything after that position and continues generation from the corrected prefix, repeating this locate-correct-continue loop until a satisfactory response is obtained. This mechanism lets annotators pre
- 副分类:risk
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-22-agent-rag-longcontext-candidates.json