Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models

  • 类型:arxiv
  • 标识:2609.21094
  • 链接:https://arxiv.org/abs/2609.21094
  • 主分类:risk
  • 形态:benchmark
  • TLDR:Large Language Models (LLMs) are increasingly deployed in applications that must weigh clashing moral values, yet even strong models exhibit hidden biases and brittle instruction-following across languages. We introduce a 12,000-instance dataset of two-option dilemmas covering pairwise three value conflicts: Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty, along with their translations into Hindi, Arabic, Spanish, and Chinese, to probe cross-lingual behavior. Benchmarking on GPT-5-mini reveals that it consistently favors Honesty over Autonomy across all five languages when
  • 待LLM分类:否
  • 标题中文:价值几何:语言模型伦理偏好对齐中的任务向量组合
  • TLDR中文:大语言模型 (LLM) 越来越多地部署于必须权衡冲突道德价值观的应用中,但即便是强模型也表现出隐藏偏见和跨语言指令遵循能力的脆弱性。我们引入一个包含 12,000 个二选一困境实例的数据集,涵盖三对成对价值冲突:诚实 vs. 正义、正义 vs. 自主、自主 vs. 诚实,并提供印地语、阿拉伯语、西班牙语和中文翻译,用于探测跨语言行为。对 GPT-5-mini 的基准测试显示,在所有五种语言中它始终偏好诚实而非自主,当
  • 来源文件
  • /inbox/tom/_candidates/2026-09-21-agent-rag-longcontext-candidates.json