arXiv:2609.14302 · 评测基准
E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
E2A-Bench:金融图表推理中证据到行动可靠性的基准测试
E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
- 类型:arxiv
- 标识:2609.14302
- 链接:https://arxiv.org/abs/2609.14302
- 主分类:evaluation
- 形态:benchmark
- TLDR:Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucination evaluations are mostly claim-centric; they assess whether generated statements are supported, but not whether evidence remains traceable through rationale, confidence, and final action. We introduce E2A-Bench, a 969-query benchmark for financial chart reasoning, constructed from 323 HS300 constituents under three input modalities with deterministic OHLCV-derived evidence anchors. E2A-Bench evaluates grounding, reasoning-action consistency, evidence-confidence calibration
- 待LLM分类:否
- 标题中文:E2A-Bench:金融图表推理中证据到行动可靠性的基准测试
- TLDR中文:金融视觉语言模型(VLMs)能否将图表证据转化为可靠的操作建议?现有幻觉评估大多以声明为中心,判断生成的陈述是否有证据支持,但并不评估证据能否在推理过程、置信度与最终行动中保持可追溯。本文提出 E2A-Bench,一个包含 969 个查询的金融图表推理基准,基于 323 只 HS300 成份股在三种输入模态下构建,并采用由 OHLCV 衍生的确定性证据锚点。E2A-Bench 评估 grounding、推理-行动一致性、证据-置信度校准...
- 来源文件:
- /inbox/tom/_candidates/2026-09-16-agent-rag-longcontext-candidates.json