LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

  • 类型:arxiv
  • 标识:2608.23200
  • 链接:https://arxiv.org/abs/2608.23200
  • 主分类:evaluation
  • 形态:method
  • TLDR:Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet successful execution experience is typically lost after a single run, forcing subsequent models to rediscover strategies and failure modes from scratch. We study whether such experience can instead be externalized and reused through EvoMap, where verifier-confirmed execution trajectories are consolidated into structured Gene. To evaluate this setting, we introduce the Long-Workflow B
  • 待LLM分类:否
  • 标题中文:LongWoF-Bench:评估 EvoMap Gene 的可验证长工作流任务基准
  • TLDR中文:大语言模型日益被期望执行复杂工作流,其成功依赖于维持相互依赖的约束并产出满足严格端到端验证的产物。然而成功的执行经验通常在单次运行后就丢失了,迫使后续模型从头重新发现策略和失败模式。我们研究能否通过 EvoMap 将此类经验外部化并复用,其中验证器确认的执行轨迹被整合为结构化的 Gene。为评估该设定,我们引入长工作流基准
  • 来源文件
  • /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json