UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

  • 类型:arxiv
  • 标识:2608.09209
  • 链接:https://arxiv.org/abs/2608.09209
  • 主分类:evaluation
  • 形态:benchmark
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:U N M ASK is presented, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation, and demonstrates that the discovery and validation stages generalize to reward model preference data.
  • 待LLM分类:否
  • 标题中文:UNMASK:文本分类器中虚假捷径的发现与因果验证
  • TLDR中文:提出 UNMASK,一个全自动 pipeline,可在无需额外人工标注的情况下发现、因果验证并缓解文本分类器中的伪相关,并证明其发现与验证阶段可泛化至奖励模型的偏好数据。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-17-agent-rag-longcontext-candidates.json
  • [S2 enrich]