Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

  • 类型:arxiv
  • 标识:2608.31082
  • 链接:https://arxiv.org/abs/2608.31082
  • 主分类:agent
  • 形态:benchmark
  • TLDR:Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI is deploying LLM agents that reason over this data to answer complex questions for every knowledge worker. Agents can do this today, but at prohibitive cost. Each question repeatedly opens large documents to recover scattered evidence, consuming up to a million tokens. However, if the data were already structured, the same question would reduce to a cheap database lookup. For example, on FanOutQA benchmark, reasoning over an ideal pre-structured
  • 待LLM分类:否
  • 标题中文:通过对非结构化数据的自适应结构化实现 token 高效的数据推理 Agent
  • TLDR中文:大量有价值的数据仍埋藏于非结构化来源之中:网页、报告、合同、申报文件、财报电话会议与 PDF。企业 AI 的大赌注是部署 LLM Agent,跨这些数据进行推理,为每位知识工作者回答复杂问题。Agent 如今可以做到这一点,但成本极高——每个问题都要反复打开大型文档以收集分散的证据,单次消耗可达上百万 token。然而,若数据已被预先结构化,同样的问题将退化为廉价的数据库查询。例如在 FanOutQA 基准上,基于理想预结构化数据的推理
  • 来源文件
  • /inbox/tom/_candidates/2026-09-03-agent-rag-longcontext-candidates.json