Language Chain in Alignment: Cross-lingual Ranking Preference Optimization
- 类型:arxiv
- 标识:2608.23149
- 链接:https://arxiv.org/abs/2608.23149
- 主分类:risk
- 形态:method
- TLDR:The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages. In this paper, we propose Cross-lingual Ranking Preference Optimization~(CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language. We design a hierarchical structure within parallel preference pairs across the target language and English to jointly optimize intra- and inter-lingual preferences, thereby enhancing language adaptation and output qualit
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-08-31-agent-rag-longcontext-candidates.json