Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
- 类型:arxiv
- 标识:2609.30867
- 链接:https://arxiv.org/abs/2609.30867
- 主分类:evaluation
- 形态:benchmark
- TLDR:Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Acr
- 待LLM分类:否
- 标题中文:[标题中文] 气候政策因果评估中识别假设的循证审计
- TLDR中文:[TLDR中文] 双重差分(DID)研究被广泛用于评估气候政策,但评估支持其识别假设的证据仍具挑战。我们提出 ARGUS,一个结构化的语言模型流水线,针对十一维的假设—含义—证据评估标准对所报告的证据进行审计,并在无法检索到相关证据时选择弃答。我们通过注入缺陷、经济学论文以及一个使用经协调标签的小规模试点评估 ARGUS。在 11 类缺陷基准上,ARGUS 检测出 73% 的植入缺陷,而基于关键词的流水线仅能检测出 18%。
- 来源文件:
- /inbox/tom/_candidates/2026-09-28-agent-rag-longcontext-candidates.json