Collective Bias Mitigation via Model Routing and Collaboration

  • 类型:arxiv
  • 标识:2610.03240
  • 链接:https://arxiv.org/abs/2610.03240
  • 主分类:engineering
  • 形态:application
  • TLDR:Large language models (LLMs) are increasingly deployed in public health, finance, and governance, requiring both accuracy and societal value alignment. Despite recent advances, LLMs often perpetuate or amplify bias embedded in their training data, posing challenges to fairness. While self-debiasing encourages an LLM to identify and correct its own biases, relying on a single model's intrinsic knowledge may be insufficient to address deeply ingrained stereotypes. To address this limitation, we introduce Collective Bias Mitigation (CBM), a framework that alleviates bias by learning fine-grained
  • 待LLM分类:否
  • 标题中文:Collective Bias Mitigation:通过模型路由与协作的集体偏见缓解
  • TLDR中文:大语言模型(LLM)日益部署于公共健康、金融和治理等关键场景,要求准确性与社会价值对齐。尽管近期取得进展,LLM 仍会延续或放大训练数据中嵌入的偏见,对公平性构成挑战。自我去偏鼓励 LLM 识别并纠正自身偏见,但依赖单个模型的内在知识可能不足以消除根深蒂固的刻板印象。为突破这一局限,我们提出 Collective Bias Mitigation(CBM),通过学习细粒度
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-05-agent-rag-longcontext-candidates.json