研究库 论文知识库
Papers · organized/paper_cards

论文

63 张论文卡片 · 评测基准 · 方法

开放获取 全部 绿色 · 1640
1. Recursive Agent Harnesses (RAH)
1. Recursive Agent Harnesses(RAH)
arXiv:2606.13643 评测基准 方法 OA · 绿色 被引 5 · S2

本文命名并研究这两条研究脉络之间的模式:其递归单元是配备文件系统工具、代码执行与规划的完整 Agent harness,而非无工具的模型调用,并给出针对长上下文推理的受控评估。This work names and studies the pattern between these two lines of work, where the recursive unit is a full agent harness with filesystem tools, code execution, and planning rather than a model call with no tools, and provides a controlled evaluation on long-context reasoning.

🟡 保留 4:"The Last Harness" — Meta-Evolution 双层循环
arXiv:2604.21003 评测基准 方法 OA · 绿色 被引 5 · S2

一个两级框架将手动 harness 工程转变为自动化 harness 工程,并更进一步——将"自动化本身的设计"也自动化。A two-level framework shifts manual harness engineering into automated harness engineering, and takes one step further --automating the design of the automation itself.

🔴 保留 3:Agentic Harness Engineering (AHE) — arXiv 实证论文
🔴 保留 3:Agentic Harness Engineering(AHE)— arXiv 实证论文
arXiv:2604.25850 评测基准 方法 OA · 绿色 被引 111 · S2

提出 Agentic Harness Engineering(AHE),一个通过三个相互匹配的 observability 支柱应对 harness 工程挑战的闭环,将每一次编辑转化为可证伪的契约,使 harness 演进能够自主进行而不退化为试错。Agentic Harness Engineering (AHE) is introduced, a closed loop that addresses harness engineering challenges through three matched observability pillars that turn every edit into a falsifiable contract, so harness evolution proceeds autonomously without collapsing into trial-and-error.

Agent runtime / security / harness 补充候选
Agent runtime / security / harness 补充候选
arXiv:2603.25723 评测基准 方法 OA · 绿色 被引 53 · S2

本文提出 Natural-Language Agent Harnesses,即可编辑的、描述运行级 harness 策略的文档,以及 Intelligent Harness Runtime(IHR),一个将上述文档解释为 agent 调用、交接、状态更新、验证门控与 artifact 契约的共享运行时。This paper introduces Natural-Language Agent Harnesses, editable documents that describe run-level harness policy, and Intelligent Harness Runtime (IHR), a shared runtime that interprets these documents into agent calls, handoffs, state updates, validation gates, and artifact contracts.

Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds
大语言模型是否在玩"六度分隔"?长上下文流形中的拓扑压缩度量
arXiv:2608.17950 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作数学形式化了 Transformer 如何执行抽象推理,并提出一种新颖的严格几何签名用于评估事实可靠性,证明了深度 LLM 潜空间天然组织为小世界网络。This work mathematically formalizes how transformers execute abstract reasoning and provides a novel, strictly geometric signature for evaluating factual reliability, proving that deep LLM latent spaces natively organize into Small-World networks.

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
Zetta ζ:面向自进化物理智能的高效闭环具身 Harness
arXiv:2608.16590 评测基准 方法 OA · 绿色 被引 9 · S2

本文提出 Zetta,一种闭环具身 harness,在保持基础策略冻结的同时在线演化基于代码的运行时评判器与恢复技能,表明闭环 harness 的自进化为可靠的物理智能开辟了一条可扩展的路径Zetta is presented, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen, and shows that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.

TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity
TinyCast:基于计算周期性的概率性零样本预测
arXiv:2608.15767 评测基准 方法 OA · 绿色 被引 1 · S2

我们提出 TinyCast,一个注意力无关的零样本预测器,仅用 146,505 个参数输出预测分布,其前提是在该规模下,上下文中的周期结构值得通过计算而非学习方式得到。一个零参数谱检测器给出主导周期,上下文按其相位进行折叠,再由一个膨胀卷积编码器和一个分块自回归分位数解码器建模其余部分。它在 GIFT-Eval 榜单上所有可确认参数量的零样本条目中体积最小;在概率准确性方面,它划定了 size-accuracy 前沿。We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 parameters, on the premise that at this size the periodic structure of a context is worth computing rather than learning. A zero-parameter spectral detector supplies the dominant periods, the context is folded on their phase, and a dilated convolutional encoder and a block-autoregressive quantile decoder model the rest. It is smaller than every zero-shot entry on the GIFT-Eval board whose parameter count can be established. On probabilistic accuracy it defines the size-accuracy front

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection
Task-CoEvolve:通过自适应验证任务选择实现 Harness 高效优化
arXiv:2608.20169 评测基准 方法 OA · 绿色 被引 4 · S2

Task-CoEvolve 基于以下观察:相比被一致解决或一致失败的候选任务,候选 harness 之间存在分歧的任务更能提供区分信息;它利用基于历史结果的方差加权采样,将评估聚焦在能力前沿附近的任务上。Task-CoEvolve builds on the observation that tasks on which candidate harnesses disagree are more informative for distinguishing among them than tasks that are consistently solved or failed, and uses variance-weighted sampling based on past outcomes to focus evaluation on tasks near the capability frontier.

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks
LongWoF-Bench:评估 EvoMap Gene 的可验证长工作流任务基准
arXiv:2608.23200 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

EvoMap 的结果表明,经过验证的执行经验可以被保留并共享为可复用的外部资源,使模型能够提升长工作流完成度,而无需反复承担经验探索的全部成本。EvoMap results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
StarHarness:通过分层搜索为企业环境进化 Harness
arXiv:2608.24804 评测基准 方法 OA · 绿色 被引 1 · S2

StarHarness 提供了一种实用方法,通过根据基线失败行为对任务进行分层、将 proposer 可见的搜索任务与 proposer 隐藏的选择任务分离,并为评估泛化能力保留 held-out 任务,从而缓解工具密集型企业任务中持续的 model-environment mismatch。StarHarness offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks by stratifying tasks according to baseline failure behavior, separating proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization.

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
思维链忠实性随偏好线索的传递位置与方式而变化
arXiv:2608.29464 评测基准 方法 OA · 绿色 被引 1 · S2

结果表明,在所测试的单调用、预填工具场景下,当偏好信息通过工具传入或需从原始产物中推断时,CoT 监控的可靠性可能下降。The results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.

Evaluating the Hidden Costs of Personalization in Large Language Models
评估大语言模型个性化中的隐性代价
arXiv:2608.28833 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 PRISK,一个具备自动化数据生成和定制化指标的动态评估框架,用于揭示当前 LLM 个性化中的系统性局限以及个性化信息如何塑造其响应。PRISK is proposed, a dynamic evaluation framework with automated data generation and tailored metrics that uncovers systematic limitations in current LLM personalization and how personalized information shapes its responses.

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
Harness-of-Harness:具备持续改进能力的多日自主软件开发
arXiv:2609.01481 评测基准 方法 OA · 绿色 被引 2 · S2

在持续多天、超过 70 轮迭代的部署中,HoH 自主开发出一款第一人称射击游戏,具备完整的主线剧情、完整实现的核心机制、可供人类游玩的体验、精美的画面与集成的音效。In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio.

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds
模仿学习是否保留灵巧操作的时间鲁棒性?跨任务执行速度的专家-学习者对比
arXiv:2609.01453 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

任务名义成功率相同并不意味着跨执行速度保留专家性能;本文在相同任务条件、初始条件采样与加速倍率下对比了专家与学习者。Equal nominal task success does not imply preservation of expert performance across execution speeds, and an expert and learner under the same task conditions, initial-condition draws, and speedup factors is compared.

WHALE: A Simple Recipe for Joint Harness-Weight Optimization
WHALE:联合 Harness-权重优化的简洁方案
arXiv:2609.00196 评测基准 方法 OA · 绿色 被引 8 · S2

本文提出 Weight-Harness Alternating LEarning (WHALE),一种简单的方法,交替进行两个阶段:先在当前 harness 下更新模型,再在更新后的模型下通过在线拒绝采样微调与 Meta-Harness 搜索更优的 harness。Weight-Harness Alternating LEarning (WHALE), a simple recipe that alternates two phases: updating the model under the current harness, then searching for a better harness under the updated model with online rejection-sampling fine-tuning and Meta-Harness, is proposed.

Small Language Models as Judges for Rubric-Based Reinforcement Learning
小语言模型作为基于评分标准的强化学习评判器
arXiv:2608.30005 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文研究较小的语言模型能否作为高效且可靠的 rubric 评分器,并比较了从小模型中提取逐项判断的三种方式:生成式判定、Yes/No Logprob 边际以及探针判别器。This work studies whether smaller language models can serve as efficient and reliable rubric-based judges, and compares three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges.

When Models Edit Too Much: On the Fidelity of Minimal Code Edits
模型过度编辑时:最小代码编辑的保真度问题
arXiv:2609.04061 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作通过向参考解中注入受控的 AST 级 corruption,从 400 个 BigCodeBench 问题构建了一个评估框架,为每个修复任务赋予已知的最小 patch,并将 edit fidelity 定位为 code-repair 质量的一个独立维度,表明其可被度量与学习。This work constructs an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch, and positions edit fidelity as a distinct axis of code-repair quality and shows that it can be measured and learned.

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
思维链表象之下:LLM 中推理操作的机制化解读
arXiv:2609.04753 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作发现 reasoning operations 在 held-out 表征中是可分的,且 separability 在中间层达到峰值,并验证该结构无法由词汇或位置混淆因素解释。This work finds that reasoning operations are separable in held-out representations, with separability peaking in middle layers, and verify that this structure is not explained by lexical or positional confounds.

Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection
借助 CLIP 与 DINO:面向可泛化深度伪造图像检测的不确定性感知级联融合网络
arXiv:2609.07670 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UCF-Net,一个不确定性感知级联融合网络,利用 CLIP 语言对齐的语义先验与 DINO 自监督视觉结构先验,在域内与跨域评估中均取得最优平均 AUC。UCF-Net is proposed, an uncertainty-aware cascaded fusion network that harnesses CLIP's language-aligned semantic priors and DINO's self-supervised visual-structure priors and achieves the best mean AUC among the evaluated methods in both in-domain and cross-domain evaluations.

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
协同进化的 Harness 与模型:On-policy 修正助力弱模型在模仿失效处迎头赶上
arXiv:2609.09134 评测基准 方法 OA · 绿色 被引 3 · S2

开发一条 on-policy 专家修正流水线,由元层级 MLE Agent 自动化,在弱模型自身的 rollout 中定位失败回合,并请专家仅重写该回合,从而保留模型的规划风格,融合 Harness 进化与模型适配带来的收益。An on-policy expert-correction pipeline is developed, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn, which preserves the model's planning style and combines the gains of harness evolution and model adaptation.

Towards a Deterministic Math Solver for Clinical Language Models
Towards a Deterministic Math Solver for Clinical Language Models
arXiv:2609.10728 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

即便在公式、变量与笔记访问条件一致时,向 Program-Solve 接口加入 executor 对不同开源权重模型的帮助程度各异,且无论哪种情况都无法替代经过验证的公式或可靠的变量抽取。Adding an executor to the Program-Solve interface helps some open-weight models more than others even under matched formula, variable and note access, and is not a substitute for verified formulas or reliable variable extraction either way.

Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
Pick Your Poison:学习选择投毒集合以实现更强的 LLM 后门攻击
arXiv:2609.15029 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

将投毒选择形式化为 oracle 预算下的集合优化问题,提出 SAILS (Set-level Audit-Informed Iterative Learned Selection),通过数百次微调-评估运行学习集合打分器,对百万级候选集合排序,仅审计少量入围集合。This work formalizes poison selection as oracle-budgeted set optimization and introduces SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist.

HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness
HarnessVLN:通过智能体框架统一免训练的具身导航
arXiv:2609.15195 评测基准 方法 OA · 绿色 被引 4 · S2

本工作提出 HarnessVLN,一个零样本、无需训练的框架:通过共享的 Agent Harness 统一指令跟随与物体目标导航,并展示了其在真实世界中两类导航任务上的适用性This work introduces HarnessVLN, a zero-shot, training-free framework that unifies instruction-following and object-goal navigation through a shared Agent Harness and demonstrates its applicability to both navigation tasks in real-world environments.

OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning
OmniHarness:通过符号化策略学习实现可泛化的视觉生成
arXiv:2609.16057 评测基准 方法 OA · 绿色 被引 1 · S2

本文提出 OmniHarness,一个通过符号策略学习实现可泛化视觉生成的框架,将已验证的执行抽象为视觉生成任务族的符号策略,捕获共享流程和适用条件,同时去除实例特定的输入。OmniHarness is introduced, a framework for generalizable visual generation via symbolic policy learning that abstracts verified executions into symbolic policies for visual generation task families, capturing shared procedures and applicability conditions while removing instance-specific inputs.

1️⃣ arXiv · Learning Rate Matters: Vanilla LoRA May Suffice(⭐⭐⭐⭐⭐ 必读)
学习率至关重要:Vanilla LoRA 可能已足够
arXiv:2602.04998 评测基准 方法 Open MIND OA · 绿色 被引 11 · S2

本文通过大规模超参数搜索,系统地重新评估了 Vanilla LoRA 以及九个代表性 LoRA 变体,发现不同 LoRA 方法偏好的学习率区间各异,并将最优学习率区间的差异归因于最大 Hessian 特征值的变化,与经典学习理论相吻合。This work systematically re-evaluate nine representative LoRA variants alongside vanilla LoRA through extensive hyperparameter searches, finding that different LoRA methods favor distinct learning rate ranges and attributes the differing optimal learning rate ranges to variations in the largest Hessian eigenvalue, aligning with classical learning theories.

DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation
DeformSmith:基于物理 Harness 引导的可形变资产分层生成,用于机器人操作
arXiv:2609.18620 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,DeformSmith 生成的资产在视觉质量与物理合理性上均优于 SOTA 基线,包括 PhysGen3D、PhysGM 和 PhysX-Omni,同时支持为可形变物体的机器人操作合成训练数据。Results show that DeformSmith generates assets with better visual quality and physical plausibility than state-of-the-art baselines, including PhysGen3D, PhysGM, and PhysX-Omni, while supporting the synthesis of data for robotic manipulation of deformable objects.

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
大语言模型的测谎测试:读出模型不愿透露的知识
arXiv:2609.21996 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

内部识别探针(Probe of Internal Recognition, PIR)可区分"不愿回答"与"无法回答"的模型,支持装傻审计与遗忘验证,并从选择题扩展至自由生成。Probe of Internal Recognition (PIR) separates a model that will not answer from one that cannot, which supports sandbagging audits and unlearning verification, and extends from multiple-choice questions to free-form generation.

ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning
ALPINE:面向参数与样本高效少样本学习的自适应定位
arXiv:2609.22323 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种面向 few-shot 图像分类的超轻量级空间-关系架构,结合固定的 Gabor 边缘能量引导与窗口化的、内容自适应的 patch locator。实验表明,该架构中虽然存在 pairwise relational computation,但它并非性能的主要驱动因素;真正起决定作用的是 content-adaptive patch locator。An ultra-lightweight spatial-relational architecture for few-shot image classification that combines fixed Gabor edge-energy guidance with a windowed, content-adaptive patch locator is presented, showing that the architecture's pairwise relational computation, while present, is not the primary driver of its performance; the content-adaptive patch locator is.

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
JEV-as-a-Judge:自信则接受,不确定则升级
arXiv:2609.26550 评测基准 方法 OA · 绿色 被引 14 · S2

LLM-as-a-judge 支持跨任务评估,但在大规模场景下推理成本与置信度可靠性成为关键问题。本文研究仅做判断的 judge 能否提供经济高效的首轮判别,并识别何时需要更强的评估。在与十六种生成式与奖励模型 judge 的对比中(采用盲法人类裁定作为参照),我们发现 jev-as-a-judge 在普通偏好与有证据支撑的事实性任务上,与作为最强对照的 SOTA LLM judge 仅相差 3 个百分点,成本仅为后者的 0.36%。在需要核查LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking

StudentBench: AI and human tutoring yield equivalent GRE learning gains
StudentBench:AI 辅导与人类辅导在 GRE 学习成效上相当
arXiv:2609.28470 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

论文证实,AI 辅导在 GRE 学习收益上与专家人类辅导具有统计等效性(p = .015),且在 GRE 的七个领域中,有五个领域的最佳 AI 导师平均超越了人类导师。It is established that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average.

FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation
FLEET:从 logits 熵到文本生成中的增强轨迹
arXiv:2609.27657 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

FLEET 将每次生成表示为穿越状态的稀疏轨迹(状态的熵超过预设阈值),并基于这些轨迹推断逐 token 的效用分数以调整 logits,是一种将 memory 机制融入生成过程的新方法。FLEET represents each generation as a sparse trajectory through states whose entropy exceeds a predefined threshold and uses these trajectories to infer per-token utility scores that adjust the logits, a novel method that integrates a memory mechanism into the generation process.

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models
能力出众却简洁高效:提取并刻画前沿模型中隐藏的思维链
arXiv:2609.26637 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果发现 Astra 表现出 token 高效的有向推理,能更早选择正确轨迹,在内部解决基础步骤,仅外化关键推理,为前沿模型推理提供了超越基准分数的行为视角。It is found that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning, which provides a behavioral lens on frontier-model reasoning beyond benchmark scores.

Learning to Discover Interesting Mathematics
学习发现有趣的数学
arXiv:2609.28603 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

近年来,大语言模型(LLMs)解决高级数学问题的能力日益增强,包括许多悬而未决数十年的难题。这为以空前规模扩展数学知识打开了大门。然而,尽管 LLM 能够猜想并证明越来越多的定理,这些新数学知识是否有趣或有用仍属未知。我们将定理的内在有趣度定义为其证明长度与陈述长度之比。证明该指标与下载量的外在度量高度相关。Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and more theorems, it remains open whether this new mathematical knowledge is interesting or useful. We define intrinsic interestingness of a theorem as the ratio between the length of its proof and the length of its statement. We show that this correlates strongly with an extrinsic measure of the downs

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance
SAGE:通过拓扑引导缓解长程推理偏置
arXiv:2609.30192 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 SAGE(Structural Admissibility-Guided Exploration),一个通过注入结构引导来缓解长程推理中探索偏差与累积偏差的统一框架,在 Andrews-Curtis 问题上取得了最高 8 倍的提升。This work proposes SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning and achieves up to an 8-fold improvement on the Andrews-Curtis problem.

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Mid-Harness:在 Model 与 Harness 之间扩展终端 Agent 的动作规模
arXiv:2609.39982 评测基准 方法 OA · 绿色 被引 1 · S2

本文提出 Mid-Harness,在执行前对候选动作进行采样和验证,同时保持 generator 和 harness 不变,并将 action scaling 识别为终端 Agent 中 test-time compute scaling 的一个有前景的目标。Mid-Harness is introduced, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged, and identifies action scaling as a promising target for test-time compute scaling in terminal agents.

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
MILO:通过编排式多 Agent 进化实现自动化 Harness 发现
arXiv:2609.38349 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MILO(Meta-evolutionary Island Orchestration),一个共同演化 agent harness 及其发现策略的框架,使用前沿模型(Opus 4.8)与开源权重模型(gpt-oss-120b)超越了八个 SOTA harness 与六种搜索方法。This work introduces MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them and outperforms eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models.