UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
- 类型:arxiv
- 标识:2610.10164
- 链接:https://arxiv.org/abs/2610.10164
- 主分类:agent
- 形态:method
- TLDR:Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions. Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve. However, as the actor continues learning, rewarding skill proposals through their reuse in subsequent training steps may conflate skill benefits with actor improvement, while directly testing each proposed skill requires costly additional actor rollouts. In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose ski
- 待LLM分类:否
- 标题中文:UniSkill:为持续进化的策略学习与执行者对齐的技能提案
- TLDR中文:大语言模型 Agent 可通过保留从先前交互中提炼的可复用技能来跨任务提升能力。已有研究联合优化任务执行与技能提取,使策略与技能库协同进化。然而,随着执行者持续学习,通过技能在后续训练步骤中的复用对其进行奖励,可能将技能收益与执行者自身的改进混淆;而直接测试每个候选技能又需要代价高昂的额外执行者 rollout。本文提出 UniSkill,利用共享策略与环境交互并提出技能……
- 来源文件:
- /inbox/tom/_candidates/2026-10-08-agent-rag-longcontext-candidates.json