IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
- 类型:arxiv
- 标识:2609.29167
- 链接:https://arxiv.org/abs/2609.29167
- 主分类:evaluation
- 形态:benchmark
- TLDR:Banking assistants must use account-specific information to answer requests and, in many cases, take actions through tools. Evaluating only the final response misses important errors. An assistant may ask for information it already has, rely on stale context, select the wrong account, or write an invalid value after stating the correct one. We introduce IndicBankBench, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes. Cases are evaluated at four stages: safety, action and tool use, response adequacy, and advi
- 副分类:risk
- 待LLM分类:否
- 标题中文:[标题中文] IndicBankBench:评估语言模型助手在印度零售银行中的安全性与可靠性
- TLDR中文:[TLDR中文] 银行助手必须使用账户专属信息来回应请求,并在许多情况下通过工具执行操作。仅评估最终回答会遗漏重要错误。助手可能会询问其已知的信息、依赖过时的上下文、选择错误的账户,或在陈述正确值后写入无效值。我们推出 IndicBankBench,一个包含 799 个案例的印度零售银行基准,涵盖五个运营领域、一个能力/拒答领域以及二十个主要评估维度。案例在四个阶段进行评估:安全性、动作与工具使用、回答充分性以及建议
- 来源文件:
- /inbox/tom/_candidates/2026-09-28-agent-rag-longcontext-candidates.json