PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
- 类型:arxiv
- 标识:2608.30597
- 链接:https://arxiv.org/abs/2608.30597
- 主分类:engineering
- 形态:method
- TLDR:Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as ac
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-14-agent-rag-longcontext-candidates.json