Language Chain in Alignment: Cross-lingual Ranking Preference Optimization

  • 类型:arxiv
  • 标识:2608.23149
  • 链接:https://arxiv.org/abs/2608.23149
  • 主分类:risk
  • 形态:method
  • TLDR:The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages. In this paper, we propose Cross-lingual Ranking Preference Optimization~(CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language. We design a hierarchical structure within parallel preference pairs across the target language and English to jointly optimize intra- and inter-lingual preferences, thereby enhancing language adaptation and output qualit
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-08-31-agent-rag-longcontext-candidates.json