SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
- 类型:arxiv
- 标识:2608.19799
- 链接:https://arxiv.org/abs/2608.19799
- 主分类:agent
- 形态:benchmark
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:A paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context shows that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success.
- 副分类:evaluation
- 待LLM分类:否
- 标题中文:SWE-bench Science:编码 Agent 能解决科学领域的工程任务吗?
- TLDR中文:一项配对消融实验在保留仓库与可执行工程上下文的同时移除显式科学指导,表明科学知识并非一律有益:可靠信息能约束修复、提升平均表现与 token 效率,而错位指导则会诱发锚定,不必然提升精确修复成功率。
- 来源文件:
- /inbox/tom/_candidates/2026-08-21-agent-rag-longcontext-candidates.json
- [S2 enrich]