Self-Supervised Scaling of Terminal Environments for Scientific Domains

  • 类型:arxiv
  • 标识:2610.02710
  • 链接:https://arxiv.org/abs/2610.02710
  • 主分类:agent
  • 形态:application
  • TLDR:Terminal agents are increasingly deployed beyond software engineering in science and other specialized domains. Constructing training environments requires executable reference behavior and a domain-specific verifier that distinguishes semantic correctness from superficially plausible artifacts. Authoring these components for each task requires repeated engineering and limits reuse. We introduce software-in-the-loop reconstruction, a self-supervised framework that obtains reference outputs and verification targets from existing software workflows, executable programs mapping structured inputs
  • 待LLM分类:否
  • 标题中文:科学领域终端环境的自监督规模化
  • TLDR中文:Terminal Agent 正日益超越软件工程,部署到科学及其他专业领域。构建训练环境需要可执行参考行为与领域特定的验证器,以区分语义正确性与表面似是而非的产物。为每个任务编写这些组件需要反复工程投入,限制了复用性。本文提出 software-in-the-loop reconstruction,一种自监督框架,从现有软件工作流中获取参考输出与验证目标,即可将结构化输入映射为
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-06-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-10-06-agent-rag-longcontext-candidates.json