研究库 论文知识库
Papers · organized/paper_cards

论文

188 张论文卡片 · 工程化 · OA 绿色

开放获取 全部 绿色 · 1640
Persona Dosing: Calibrated Activation Steering for Graded Trait Control
人设剂量:用于分级特性控制的校准激活引导
arXiv:2609.36388 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果将控制器所学到的行为范围与该范围内请求的准确性分离开来,尽管连贯性下限并非对每种特性都成立。The results separate the behavioral range learned by a controller from the accuracy of requests within that range, although the coherence floor does not hold for every trait there.

Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability
高效推理训练并不总是损害 CoT 忠实性与可监控性
arXiv:2610.03509 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用三种以不同方式施加长度压力的方法(固定生成预算、逐样本长度目标、组相对长度奖励)对多种模型进行微调,发现它们对忠实性和可监控性具有不同的影响。This work fine-tune a variety of models with three methods that apply length pressure differently, namely a fixed generation budget, a per-example length target, and a group-relative length reward, and finds that it affects faithfulness and monitorability differently.

Collective Bias Mitigation via Model Routing and Collaboration
Collective Bias Mitigation:通过模型路由与协作的集体偏见缓解
arXiv:2610.03240 工程化 应用落地 OA · 绿色 被引 1 · S2

本文首次系统地探索了不同 LLM 的有效选择与组织,以培育更公平的 LLM 回答,并展示了 CBM 显著优于独立基线。This work is the first to systematically explore the effective selection and organization of distinct LLMs to cultivate fairer LLM responses and show CBM substantially outperforms standalone baselines.

JEPA-TTT: Persistent Test-Time Training of Latent World Models for Planning under Dynamics Shifts
JEPA-TTT:用于动态漂移下规划的潜在世界模型持续测试时训练
arXiv:2610.00722 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

JEPA-TTT 在测试时间内持续适配预训练动作条件 Joint-Embedding Predictive Architecture 世界模型的潜在动力学预测器,平均将自回归潜在预测误差降低 83%,并将规划性能相较冻结的 JEPA 世界模型提升 153%。JEPA-TTT, which adapts the latent dynamics predictor of a pretrained action-conditioned Joint-Embedding Predictive Architecture world model throughout test time, reduces autoregressive latent prediction error by 83% on average and improves planning performance by 153% over the frozen JEPA world model.

Closing the Context Gap: Activation Alignment for Tabular In-Context Learning
弥合上下文鸿沟:表格上下文学习的激活对齐
arXiv:2610.06679 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

表格基础模型通过以标注训练样本为条件进行上下文学习 (ICL) 来预测。与分离训练与推理的传统模型不同,这些模型必须在每次前向传播中处理所有训练样本,导致每次预测成本高昂。限制训练样本数量可降低成本,但会显著降低性能。我们不丢弃上下文,而是提出激活对齐 (activation alignment),利用完整上下文来教会模型在仅看到子集时如何行为。这是通过训练一个……Tabular foundation models perform in-context learning (ICL) by conditioning predictions on labeled training examples provided as context. Unlike traditional models that separate training from inference, these models must process all training examples in every forward pass, making each prediction expensive. Restricting the number of training examples reduces this cost but substantially degrades performance. Instead of discarding context, we propose activation alignment, a method that leverages the full context to teach a model how to behave when seeing only a subset. This is achieved by trainin

LLM-as-Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
LLM-as-Jev:LLM 已是 Jev 风格决策模型 —— 何时以及如何对其进行微调
arXiv:2610.02076 工程化 方法 OA · 绿色 被引 1 · S2

提出 LLM-as-Jev,一种保留架构的框架,直接从带括号数字标识符的下一 token 概率中提取校准决策,并同时给出无需训练的推理方案与一种微调目标——通过树分解的列表式损失优化候选选择,并使用 KL 散度惩罚将辅助预测锚定到基模型。This work presents LLM-as-Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers, and provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties.

CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training
CONFLUX:用于胸部 CT 三维合成的潜在扩散模型与强化学习后训练
arXiv:2607.02998 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

CONFLUX 是一个面向胸部 CT 的潜在扩散模型:由 3D 变分自编码器压缩每个体数据,整流流 Transformer 在潜在空间中生成,并以分类器从生成体中恢复所请求病征的可靠性作为奖励。CONFLUX, a latent diffusion model for chest computed tomography (CT): a 3D variational autoencoder compresses each volume, and a rectified-flow transformer generates in the latent space, that rewards how reliably a classifier recovers the requested findings from each generated volume.

Transition-Aware best-of-N sampling for Longitudinal Chest X-ray Reports
面向纵向胸部 X 光报告的过渡感知 best-of-N 采样
arXiv:2606.28393 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

首个面向预训练胸部 X 光报告生成器的无需训练 best-of-N 采样方案,显式建模"既往-当前-过渡"纵向先验,全面优于随机选择。This work presents the first training-free best-of-N sampling scheme for pre-trained chest X-ray report generators that is explicitly aware of this longitudinal prior to current transition, and outperforms random selection across the board.

WARP: Weight-Space Analysis for Recovering Training Data Portfolios
WARP:基于权重空间分析的训练数据组合还原
arXiv:2607.01686 工程化 观点 OA · 绿色 被引 1 · S2

WARP:一个直接从已发布权重还原微调模型训练数据混合比例的框架,抽取几何特征并映射至各领域占比,可采用无参数 softmax 读出器,或基于合成混合训练的 MLP 投影器。WARP is introduced, a framework that recovers a fine-tuned model's training mixtures directly from its released weights and extracts geometric features and maps them to domain proportions using either a parameter-free softmax readout or an MLP projector trained on synthetic mixtures.

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination
图原生强化学习通过概念重组实现可追溯的科学假设生成
arXiv:2607.00924 工程化 方法 OA · 绿色 被引 3 · S2

本文提出了 Graph-PRefLexOR,这是一族基于图结构的推理模型,使用 Group Relative Policy Optimization 进行微调,将推理过程组织为显式阶段,分别用于机理探索、图构建、模式提取和假设合成,确立了面向图结构的强化学习作为通向可解释 AI 系统的路径,可应用于材料设计及其他科学领域的科学假设生成。Graph-PRefLexOR is developed, a family of graph-native reasoning models fine-tuned with Group Relative Policy Optimization to organize reasoning into explicit phases for mechanism exploration, graph construction, pattern extraction, and hypothesis synthesis, establishing graph-native reinforcement learning as a pathway toward interpretable AI systems for scientific hypothesis generation in materials design and other scientific applications.

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
训练策略优化的幻象:单调推理策略才是 LLM 强化学习的真正目标
arXiv:2606.29526 工程化 方法 OA · 绿色 被引 1 · S2

本文提出 Monotonic Inference Policy Update (MIPU),一种两步式 LLM 强化学习框架:构建采样器引用的候选更新,并使用推理侧差距代理选择性接受同步候选;实验表明 MIPU 提升了平均推理性能与训练稳定性。Monotonic Inference Policy Update (MIPU) is introduced, a two-step LLM RL framework that constructs sampler-referenced candidate updates and selectively accepts synchronized candidates using an inference-side gap proxy, and experiments show that MIPU improves average reasoning performance and training stability.

One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
一个场景,两种深度:探究单目基础模型中的几何歧义
arXiv:2606.29600 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MultiDepth-3k (MD-3k),一个用于衡量深度层偏好与多层空间关系准确率 (ML-SRA) 的稀疏双层序数基准;领先的深度基础模型在标准 RGB 输入下表现出不同的层偏好,表明同一分层几何可在不同模型中被差异化地解析。MultiDepth-3k (MD-3k), a sparse two-layer ordinal benchmark for measuring depth-layer preference and multi-layer spatial relationship accuracy (ML-SRA), is introduced and leading depth foundation models exhibit diverse layer preferences under standard RGB input, showing that the same layered geometry can be resolved differently across models.

Large-Scale Tunnel Air-Ground Collaboration With FLISP: Fast LiDAR-IMU Synchronized Path Planner
基于 FLISP 的大规模隧道空地协同:Fast LiDAR-IMU Synchronized Path Planner
arXiv:2606.25393 工程化 方法 OA · 绿色 被引 1 · S2

水电隧道检测对基础设施完整性至关重要,但人工方式效率低下且具有危险性。本文提出 FLISP(Fast LiDAR-IMU Synchronized Path Planner),一种面向 UGV-UAV 协同检测的无地图规划框架。不同于传统基于地图的范式,FLISP 具有三项核心贡献:(1) 统一架构,由单套 UGV 搭载的 LiDAR-IMU 驱动两平台的同步路径生成;(2) 平台特定的求解器,采用增强型萤火虫算法用于 UGV 避障,以及动态迭代优化器用于 UAHydropower tunnel inspection is critical for infrastructure integrity yet remains inefficient and hazardous using manual methods. We propose FLISP (Fast LiDAR-IMU Synchronized Path Planner), a mapless planning framework for cooperative UGV-UAV inspection. Unlike traditional map-based paradigms, FLISP features three core contributions: (1) a unified architecture where a single UGV-mounted LiDAR-IMU suite drives synchronized path generation for both platforms; (2) platform-specific solvers utilizing an enhanced Firefly Algorithm for UGV obstacle avoidance and a dynamic iterative optimizer for UA

LISA: Likelihood Score Alignment for Visual-condition Controllable Generation
LISA:面向视觉条件可控生成的似然分数对齐
arXiv:2606.27192 工程化 方法 OA · 绿色 被引 2 · S2

实验表明,LISA 不仅能持续加速训练收敛并提升最终合成结果,还能促使侧网络特征在条件建模中更加解耦,且几乎无额外训练成本,推理成本为零。Experiments demonstrated that LISA can not only consistently accelerate the training convergence and improve final synthetic results, but also encourage the side network's features to be more disentangled for conditional modeling with negligible additional training cost and zero extra inference cost.

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
arXiv:2606.26080 工程化 评测集 OA · 绿色 被引 1 · S2

本文表明强化学习(RL)后训练已具备实现有效 step-level 评分所需的要素,从而完全无需额外的奖励模型训练,并在通用随机 Markov 决策过程下推导出一种隐式 advantage,称为 progress advantage。This work shows that reinforcement learning (RL) post-training already provides the ingredients for effective step-level scoring, eliminating the need for dedicated reward model training altogether, and derives an implicit advantage under a general stochastic Markov decision process, which is term progress advantage.

Randomized YaRN Improves Length Generalization for Long-Context Reasoning
Randomized YaRN 改善长上下文推理的长度泛化能力
arXiv:2606.23687 工程化 观点 OA · 绿色 被引 1 · S2

本文提出 Randomized YaRN,一种通过将基于 YaRN 的位置外推与随机位置编码和长度课程相结合来提升长度泛化能力的训练方法,表明渐进式地将模型暴露于分布外位置分布是实现可泛化长上下文推理的有效方案。Randomized YaRN is proposed, a training method that improves length generalization by combining YaRN-based positional extrapolation with randomized positional encoding and a length curriculum, and suggests that progressively exposing models to OOD positional distributions provides an effective recipe for generalizable long-context reasoning.

The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data
Stanford EDGAR Filings Dataset:将美国企业及金融披露重建为版面保真且 token 高效的预训练数据
arXiv:2606.18192 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Stanford EDGAR Filings Dataset(SEFD),一个将 SEC filings 开放重建为版面保真 MultiMarkdown 的数据集,用于金融语言建模与评估;同时推出两个基于 SEFD 的基准:EDGAR-Forecast,用于评估模型知识截止后基于 filings 的数值预测;EDGAR-OCR,用于评估复杂金融表格的转录质量。The Stanford EDGAR Filings Dataset (SEFD), an open reconstruction of SEC filings into layout-faithful MultiMarkdown for financial language modeling and evaluation, is introduced and two SEFD-derived benchmarks are introduced: EDGAR-Forecast, which evaluates filing-grounded numerical forecasting after model knowledge cutoffs, and EDGAR-OCR, which evaluates transcription of complex financial tables.

Trustworthy Self-Composable Big-Data-as-a-Service: An LLM-Orchestrated Multi-Agent Framework for Automated Data Engineering, AutoML, MLOps Deployment, and Drift-Aware Lifecycle Optimization
可信自组合 Big-Data-as-a-Service:用于自动化数据工程、AutoML、MLOps 部署与漂移感知生命周期优化的 LLM 编排多 Agent 框架
arXiv:2606.17915 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,由 LLM 编排的多 Agent 系统可将传统 AutoML 扩展为可信、自适应且面向生产的 BDaaS 生命周期自动化。The results suggest that LLM-orchestrated multi-agent systems can extend conventional AutoML toward trustworthy, adaptive, and production-oriented BDaaS lifecycle automation.

MedEasy: Designing AI Standardized Patients for Clinical Consultation Training
MedEasy:为临床问诊训练设计 AI 标准化病人
arXiv:2606.17512 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MedEasy,一个多 Agent 系统,通过患者对话、临床操作、决策提交、文档记录与反馈来组织虚拟患者练习,为使用案例特定标准连接情境化实践的 AI 辅助职业训练系统贡献了设计启示。MedEasy is presented, a multi-agent system that organizes virtual-patient practice through patient dialogue, clinical actions, decision submission, documentation, and feedback that contributes design implications for AI-supported professional training systems that use case-specific standards to connect situated practice.

CoRe: A Continuously Reward-Finetuned LLM Query Rewriter for Multi-Stage Context-Aware Relevance in Web-Scale Video Search
CoRe:面向 Web 规模视频搜索中多阶段上下文感知相关性的连续奖励微调 LLM 查询改写器
arXiv:2606.14127 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

CoRe(Context Relevance)被提出。该系统在一个大型短视频搜索引擎中每周重新部署,持续运行超过五个月,使用已部署的多模态相关性模型作为源,并以镜像生产融合代数的乘性比值形式来缩小仿真与生产之间的差距。CoRe (Context Relevance) is presented, such a system, redeployed weekly for over five months in a major short-video search engine, using the deployed multimodal relevance model as its source and a multiplicative ratio form mirroring the production fusion algebra to close the simulation-production gap.

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies
OmniTacTune:面向视觉策略触觉残差适配的策略无关真实世界 RL
arXiv:2607.03723 工程化 方法 OA · 绿色 被引 3 · S2

提出 OmniTacTune,一种策略无关的真实世界 RL 流程,通过残差修正将触觉反馈适配到预训练视觉策略,并在多种接触丰富任务、视觉基础策略与触觉表征间实现泛化。OmniTacTune is introduced, a policy-agnostic real-world RL pipeline that adapts tactile feedback to pretrained visual policies through residual correction and generalizes across diverse contact-rich tasks, visual base policies, and tactile representations.

TESSERA v2: Scaling Pixel-wise Earth Foundation Models
TESSERA v2:扩展像素级地球基础模型
arXiv:2607.03949 工程化 方法 OA · 绿色 被引 5 · S2

给出一个具体且有实证支撑的像素级 EO 基础模型扩展方案:训练大型 encoder,按下游性能筛选,再蒸馏为灵活的学生模型。A concrete, empirically grounded recipe for scaling pixel-wise EO foundation models: train large encoders, select by downstream performance, and distil into flexible student models is given.

Enhancing In-context Panoramic Generation via Geometric-aware Pretraining
基于几何感知预训练增强上下文全景生成
arXiv:2607.08765 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

借助强大的全景先验,Canvas360 构建了一个统一的上下文全景生成框架,通过 token 级拼接支持多样化下游任务,在任务覆盖范围与建模灵活性上均超越已有方法。Empowered by strong panoramic priors, Canvas360 enables a unified in-context panoramic generation framework that supports diverse downstream tasks via token-level concatenation, surpassing prior methods in both task coverage and modeling flexibility.

PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection
PAST-TIDE:基于原型锚定语句调优与主题不变归一化的立场检测方法
arXiv:2607.04690 工程化 方法 OA · 绿色 被引 1 · S2

提出 PAST-TIDE,在官方排行榜上 Subtask A 取得 0.75 的 macro-F1,Subtask B 取得 0.74,表明对预训练模型仅做极少的架构改动即可在低资源场景下保持竞争力。PAST-TIDE is introduced, PAST-TIDE achieves macro-F1 scores of 0.75 for Subtask A and 0.74 for Subtask B on the official leaderboard, indicating that minimal architectural additions to a pre-trained model can remain competitive in low-resource settings.

DrugGen 2: A disease-aware language model for enhancing drug discovery
DrugGen 2:一种用于增强药物发现的疾病感知语言模型
arXiv:2607.08404 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

通过将疾病特异性上下文整合到分子生成中,DrugGen-2 推动了 AI 辅助药物发现,为 de novo 设计和药物再利用提供了强大工具,可同时考虑疾病与分子靶点之间的复杂相互作用。By integrating disease-specific context into molecular generation, DrugGen-2 advances AI-assisted drug discovery, offering a powerful tool for de novo design and drug repurposing that accounts for the complex interplay between diseases and molecular targets.

A Sparse and Truncated State Vector Simulator for Peaked Circuits
一种用于 Peaked Circuits 的稀疏截断态向量模拟器
arXiv:2607.07816 工程化 方法 OA · 绿色 被引 1 · S2

本工作描述了如何在开源实现中满足使用有限项数的截断态向量来模拟 peaked circuits 的要求,并讨论了其性能与局限性。This work describes how the requirements to simulate peaked circuits using a truncated state vector with a limited number of terms were met in an open-source implementation, and discusses its performance and limitations.

PanoWorld: Real-World Panoramic Generation
PanoWorld:真实世界全景图像生成
arXiv:2607.09661 工程化 方法 OA · 绿色 被引 3 · S2

论文提出 PanoWorld,通过固定朝向将相机轨迹简化为平移,并借助 Dense Panoramic Ray-Conditioning 与 Geometry-aware Memory Augmentation 同时支持当前动作建模与长程记忆。PanoWorld is proposed, which simplifies camera trajectories into translations via fixed headings for both current-action modeling and long-range memory through Dense Panoramic Ray-Conditioning and Geometry-aware Memory Augmentation through Dense Panoramic Ray-Conditioning and Geometry-aware Memory Augmentation.

Self-Guided Test-Time Training for Long-Context LLMs
长上下文 LLM 的自引导测试时训练
arXiv:2607.09415 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一种简单方法 Self-Guided TTT (S-TTT),可同时提升 Qwen3-4B-Thinking-2507 与 Llama-3.1-8B-Instruct 的准确率,相对改进最高达 15%。A simple method, Self-Guided TTT (S-TTT), which improves accuracy for both Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct, achieving up to a 15% relative improvement.

Weak-to-Strong Generalization via Direct On-Policy Distillation
通过直接在线策略蒸馏实现弱到强泛化
arXiv:2607.05394 工程化 方法 OA · 绿色 被引 15 · S2

提出 Direct On-Policy Distillation(Direct-OPD),该方法迁移教师模型由 RL 引起的策略偏移,而非在目标模型上运行稀疏奖励 RL,并一致地利用更弱的教师模型来提升更强的目标模型Direct On-Policy Distillation (Direct-OPD) is proposed, which transfers the teacher's RL-induced policy shift instead of running sparse-reward RL on the target model and consistently leverages weaker teachers to improve stronger target models.

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
代理探索与可复用引导:一种通过代理引导更新信号实现的模块化 LLM 后训练范式
arXiv:2607.11505 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Proxy OPD——一种异步后训练框架,迁移奖励驱动的策略改进而非绝对策略分布,将相对策略更新确立为可大规模、按奖励进行后训练的高复用、可调节资产。Proxy OPD is introduced, an asynchronous post-training framework that transfers reward-induced policy improvements rather than absolute policy distributions and establishes relative policy updates as highly reusable, adjustable assets for scalable, reward-based post-training.

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
EgoSteer:基于第一人称视频的可控灵巧操作系统
arXiv:2607.09701 工程化 方法 OA · 绿色 被引 3 · S2

全栈系统:扩展来自第一人称人类视频的灵巧 VLA 预训练,实现数据高效的真实机器人后训练,并在 45 个多样化任务上稳健执行自由形式指令,展示了在多个候选任务中对用户意图的遵循与泛化能力。A full-stack system that scales dexterous VLA pre-training from egocentric human videos and enables data-efficient real-robot post-training, and robustly executes free-form instructions across 45 diverse tasks, demonstrating adherence to user intent amid multiple candidate tasks and generalization.

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
读回:预训练 MLLM 是文本到图像生成的零样本奖励模型
arXiv:2607.11886 工程化 方法 OA · 绿色 被引 1 · S2

提出 SpectraReward,一种无需训练的将预训练 MLLM 转化为即用型奖励模型的奖励函数,用于图像生成强化学习;并引入 Self-SpectraReward,这是统一多模态模型的一种特例,其中策略自身的理解分支充当其生成分支的奖励模型。SpectraReward is proposed, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning, and Self-SpectraReward is introduced, a special case for unified multimodal models where the policy's own understanding branch serves as the reward model for its generation branch.

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations
揭开 On-Policy Distillation 的神秘面纱:作用、病态与调控
arXiv:2607.13399 工程化 方法 OA · 绿色 被引 13 · S2

本文对 OPD 的作用、病态与调控进行系统研究,厘清 OPD 作为探索催化剂的角色,并证实调控良好的信号质量(而非单纯的教师模型规模)才是 OPD 中成功探索的主导因素。A systematic study examining the role, pathologies, and regulations of OPD, clarifying the role of OPD as an exploration catalyst and confirming that well-regulated signal quality, rather than mere teacher scale, governs successful exploration in OPD.

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
LongStraw:固定 GPU 预算下超越 2M token 的长上下文强化学习
arXiv:2607.14952 工程化 应用落地 OA · 绿色 被引 3 · S2

本文提出 LongStraw,一个面向目标、感知架构的系统,用于 resident-state 虚拟化、response replay 和分布式梯度执行,它将实时训练图限制在 response 后缀范围内,同时在完整的 GRPO 组内复用代价高昂的 prompt 计算。This work presents LongStraw, an objective-aware, architecture-aware system for resident-state virtualization, response replay, and distributed-gradient execution that bounds the live training graph by the response suffix while reusing the expensive prompt computation across the complete GRPO group.

Spectral Rewiring for Exploration, Purification, and Model Merging
面向探索、纯化与模型合并的谱重连
arXiv:2607.03065 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

子空间对齐重连(SAR)表明,从参数几何中提取对推理有效的更新,可作为一种无需训练的机制来提升推理与多领域性能。Subspace-Aligned Rewiring (SAR) shows that extracting reasoning-effective updates from parameter geometry can serve as a training-free mechanism to improve reasoning and multi-domain performance.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Xiaomi-Robotics-1:利用超过10万小时真实轨迹数据扩展视觉-语言-动作模型
arXiv:2607.15330 工程化 方法 OA · 绿色 被引 38 · S2

Xiaomi-Robotics-1是一个强大的机器人基础策略,能在复杂灵巧任务上以高数据效率高效微调,并在多个仿真基准上超越SOTA方法。Xiao-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency and across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods.