研究库 主题路线
内容库 / 主题
Topic · agent

Agent 智能体主题中枢

活文档 · 论文卡 · 笔记 · 仓库 · 攻略

主题活文档 Live Doc

全部
agent · 知识库活文档
agent · 知识库活文档 更新:v113=v112+24h,11 件 agent 主分类+4 件 RAG 主分类,Robo-COP 受控协同进化+物体永久性跌出第 15 日+Lightning 沿用。 主题负责人:spark 信息时效性:v113 cutoff 2026-10-09 10:30 CST · arXi
活文档 2026-10-09

论文卡 Papers

全部
Training language models to follow instructions with human feedback
使用人类反馈训练语言模型遵循指令
arXiv:2203.02155 工程化 方法 OA · 绿色 被引 24618 · S2

结果表明,使用人类反馈进行微调是使语言模型与人类意图对齐的一个有前景的方向,在真实性方面有所提升,并减少了有毒输出的生成,同时在公开 NLP 数据集上的性能回归极小。The results show that fine-tuning with human feedback is a promising direction for aligning language models with human intent and showing improvements in truthfulness and reductions in toxic output generation while having minimal performance regressions on public NLP datasets.

Reflexion: Language Agents with Verbal Reinforcement Learning
Reflexion: Language Agents with Verbal Reinforcement Learning
arXiv:2303.11366 Agent 智能体 方法 OA · 绿色 被引 5883 · S2

Reflexion 是一个通过语言反馈而非更新权重来强化语言 agent 的新框架,在多种任务(序贯决策、编程、语言推理)上相较基线 agent 取得显著提升。Reflexion is a novel framework to reinforce language agents not by updating weights, but instead through linguistic feedback, which obtains significant improvements over a baseline agent across diverse tasks (sequential decision-making, coding, language reasoning).

Towards A Rigorous Science of Interpretable Machine Learning
迈向严谨的可解释机器学习科学
arXiv:1702.08608 评测基准 观点 OA · 绿色 被引 5592 · S2

这篇立场论文定义了可解释性,阐述了何时需要(以及何时不需要)可解释性,并提出了一种用于严格评估的分类法,同时指出了迈向更严谨的可解释机器学习科学所面临的开放性问题This position paper defines interpretability and describes when interpretability is needed (and when it is not), and suggests a taxonomy for rigorous evaluation and exposes open questions towards a more rigorous science of interpretable machine learning.

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
arXiv:2403.05530 多模态 方法 OA · 绿色 被引 4062 · S2

Gemini 1.5 在跨模态长上下文检索任务上取得近乎完美的召回率,在长文档 QA、长视频 QA 与长上下文 ASR 上刷新 SOTA,并在广泛基准上达到或超越 Gemini 1.0 Ultra 的 SOTA 表现。Gemini 1.5 models achieve near-perfect recall on long-context retrieval tasks across modalities, improve the state-of-the-art in long-document QA, long-video QA and long-context ASR, and match or surpass Gemini 1.0 Ultra's state-of-the-art performance across a broad set of benchmarks.

PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
PEGASUS:基于抽取间隔句预训练的生成式摘要
arXiv:1912.08777 工程化 方法 OA · 绿色 被引 2589 · S2

本工作提出在海量文本语料上使用新的自监督目标 PEGASUS 对大型 Transformer 编码器-解码器模型进行预训练,并证明其在所有 12 个下游数据集上按 ROUGE 分数衡量均取得 SOTA 性能This work proposes pre-training large Transformer-based encoder-decoder models on massive text corpora with a new self-supervised objective, PEGASUS, and demonstrates it achieves state-of-the-art performance on all 12 downstream datasets measured by ROUGE scores.

Decision Transformer: Reinforcement Learning via Sequence Modeling
Decision Transformer:通过序列建模实现强化学习
arXiv:2106.01345 Agent 智能体 方法 OA · 绿色 被引 2520 · S2

尽管方法简单,Decision Transformer 在 Atari、OpenAI Gym 和 Key-to-Door 任务上达到或超过 SOTA 无模型离线 RL 基线的性能Despite its simplicity, Decision Transformer matches or exceeds the performance of state-of-the-art model-free offline RL baselines on Atari, OpenAI Gym, and Key-to-Door tasks.

Solving Quantitative Reasoning Problems with Language Models
用语言模型解决定量推理问题
arXiv:2206.14858 评测基准 评测集 OA · 绿色 被引 2051 · S2
A Comprehensive Overview of Large Language Models
A Comprehensive Overview of Large Language Models
arXiv:2307.06435 LLM 基础设施 综述 OA · 绿色 被引 2044 · S2

本文旨在为研究者与从业者提供一份快速、全面的参考,通过对现有工作的广泛、信息密集型总结来汲取洞见,以推动 LLM 研究的发展。This review article is intended to provide a quick, comprehensive reference for the researchers and practitioners to draw insights from extensive, informative summaries of the existing works to advance the LLM research.

HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
arXiv:2303.17580 Agent 智能体 方法 OA · 绿色 被引 1783 · S2

HuggingGPT 是一个由 LLM 驱动的 Agent,利用 LLM(如 ChatGPT)连接机器学习社区中的各种 AI 模型以解决 AI 任务,能够处理跨模态、跨领域的大量复杂 AI 任务。HuggingGPT is an LLM-powered agent that leverages LLMs (e.g., ChatGPT) to connect various AI models in machine learning communities to solve AI tasks and can tackle a wide range of sophisticated AI tasks spanning different modalities and domains.

Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
AI 海洋中的海妖之歌:大语言模型幻觉问题综述
arXiv:2309.01219 安全与风险 综述 OA · 绿色 被引 1186 · S2

本文给出了 LLM 幻觉现象与评估基准的分类体系,分析了现有缓解 LLM 幻觉的方法,并讨论了未来研究的潜在方向。This paper presents taxonomies of the LLM hallucination phenomena and evaluation benchmarks, analyzes existing approaches aiming at mitigating LLm hallucination, and discusses potential directions for future research.

A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
arXiv:2001.06937 多模态 综述 OA · 绿色 被引 1176 · S2

详细介绍大多数 GAN 算法的动机、数学表示与结构,并对它们的共性与差异进行比较。The motivations, mathematical representations, and structures of most GAN algorithms are introduced in detail, and they are compared to compare their commonalities and differences.

A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
arXiv:2303.04226 多模态 综述 OA · 绿色 被引 865 · S2

该综述全面回顾了生成模型的历史与基本组件,以及 AIGC 在单模态交互与多模态交互方向的最新进展,并介绍了文本与图像生成任务及相关模型。This survey provides a comprehensive review on the history of generative models, and basic components, recent advances in AIGC from unimmodal interaction and multimodal interaction, and introduces the generation tasks and relative models of text and image.

笔记 Notes

全部
学术知识库草稿 · Jay · 2026-10-09
AI Agent 生产级工程 · Agentic RAG · 2026 Stack 五层架构 · MCP 工具调用 作者/专栏: xx_nm98 发布时间: 20260710(估算) 匹配分: 0.6578 核心观点摘要: AI Agent 失败根因并非 LLM 能力不足,而是 harness(runtime wrap…
Jay 2026-10-09 agentragllm-infracsdn
知识库草稿 · Jay · 2026-10-09 晚间 17:35
晚间情报:HF 安全事件深度分析 · Agent 术语体系辨析 · GitHub Trending 30天全景报告 · LLM 知识工程 2026 全景地图 [Article] Hugging Face 安全事件:GLM5.2 用于事件响应(Stratechery, Ben Thompson) Hugging Face…
Jay 2026-10-09 agentllm-infrarisk
知识库草稿 · Jay · 2026-10-09
LLM推理引擎三分天下 · Agent框架2026生产选型 · HuggingFace开源模型生态 · MLOps K8s GPU部署 · 分布式OLAP · arXiv投机解码新论文 HuggingFace TGI 已进入维护模式,2026年晚期只剩三个活跃开源引擎:vLLM、SGLang、MAX(Modular) …
Jay 2026-10-09 agentllm-infraengineering
2026-10-09 早间情报:十月首周前沿模型密集发布 / Agent 架构范式转移 / HF Trending 与 Substack 深度洞察
| 模型 | 发布日期 | 关键参数 | |||| | GPT6.1 Sol | 20260929 | 1.05M token 上下文,128K 输出,专注 coding/agentic/professional workloads | | GPT6 Luna | 20260922 | 同代模型,差异化定位 | | D…
Jay 2026-10-09 agentllm-infra
Jay · CSDN 高价值技术检索 · 2026-10-09 上午
CSDN 高价值技术分享 · RAG 系统工程实践 · Agentic RAG 条件分支架构 · 多模态 RAG 工程化 Pipeline · Substack 工程洞察 CSDN (blog.csdn.net):RAG 工程实践、Agentic RAG、LLM Agent 架构、多模态 RAG、具身智能表达层 Sub…
Jay 2026-10-09 agentragmultimodalllm-infra
coding-agents · E1 预消化简报(2026-10-09)
执行体:flyP · 20261009 23:20 CST · codingagents 主题接力预备棒(周五晚间) 承接棒位:organized/knowledge/codingagents.md v113 中棒延伸棒位(1009 10:00 落定 · frontier lab 治理商业生态栖扩增第 2 例 + 19…
flyP 2026-10-09 agent
Agent · RAG · Long-Context 雷达|2026-10-09
本轮候选: 8 条(arXiv 4 / HF Daily 4) 高价值: 4 条 Substack: 1 条 CSDN: 0(未使用) 来源: arXiv · 20261007 链接: 核心: 长程 Agent 将历史压缩为"摘要 + 软记忆 token 序列",通过残差连接类比补充摘要,接近全历史的推理效果。在 Su…
Tom 2026-10-09 agentrag
Tom 文献雷达 · Agent + RAG + Long Context · 2026-10-09T14:40
| # | 来源 | 标题 | 核心标签 | ||||| | 1 | HF Daily | MiMoV2.6: Scaling RL Towards SelfImprovement | multimodal, systems | | 2 | HF Daily | OuroWorld: 3D Cinemagraphs f…
Tom 2026-10-09 agentrag

仓库 Repos

全部
s1dashu/ip-as-logo-skill
未知语言 · 2026-08-22 Agent 智能体 应用 研究原型 Stars 3727 周增 +4646

一项紧凑的 Agent Skill,用于生成高度简化、圆润且带有微妙新拟物化风格的 IP 吉祥物 Logo。A compact Agent Skill for highly simplified, rounded, subtly neo-skeuomorphic IP mascot logos.

agentmultimodal
stablyai/orca
TypeScript · 2026-10-07 Agent 智能体 应用 生产可用 Stars 86990 周增 +4252

Orca 是用于管理大规模并行 agent 集群的 ADE。可凭个人订阅运行任意 coding agent,支持桌面、移动端与远程 runtime。Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.

agent
yetone/magpie
Go · 2026-09-30 Agent 智能体 模型 实验 Stars 3069 周增 +4109

汇聚每个 Agent 的模型,一站式管理。从菜单栏即可使用 DeepSeek 上的 Codex、Kimi 上的 Claude Code。Every agent's model. One place. Codex on DeepSeek, Claude Code on Kimi, from the menu bar.

agentllm-infra
holaboss-ai/holaOS
TypeScript · 2026-08-18 Agent 智能体 应用 生产可用 Stars 9436 周增 +3798

你的工作超级 Agent:本地优先,分钟级学习你的工作上下文,永不遗忘。Open-source All in One AI agent workspace. Run any agent — Claude Code, Codex — across your tools (100+ integrations + MCP), apps, browser, and files, with shared memory. Built-in models or BYOK.

agentllm-infra
thedotmack/claude-mem
TypeScript · 2026-10-08 Agent 智能体 应用 生产可用 Stars 98246 周增 +3526

为每个 Agent 提供跨会话持久上下文——捕获会话中 Agent 的所有行为,经 AI 压缩后注入到未来会话中。支持 Claude Code、OpenClaw、Codex、Gemini、Hermes、Copilot、OpenCode 等。Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More

agentragdatabase
firecrawl/firecrawl
TypeScript · 2026-08-11 RAG 检索增强 工具 生产可用 Stars 165271 周增 +2975

用于大规模搜索、抓取与交互网页的 API。🔥The context API to search, scrape, and interact with the web at scale. 🔥

agentllm-infra

攻略 Guides

全部
Boom5426/Nature-Paper-Skills · 上手攻略
NaturePaperSkills 是一套面向 Codex(ChatGPT)和 Claude Code 的 AI 科研写作技能集,27 个 skill 将论文全生命周期串成一条连续工作流:从立项定位、结构论证、图表设计,到科学写作、引用核验、审稿回复,覆盖 Nature 系列生命科学、计算生物学与方法学论文的写作与修订。 核心理念:论证先于语言——先明确科学…
Agent 智能体 Boom5426/Nature-Paper-Skills Tom 2026-10-10 学术写作 / AI 工具 / 科研工作流
pullboard-dev/pullboard · 上手攻略
Pullboard 是一个基于 Git 仓库的 AI Agent 协作队列工具。它的核心理念是:人类制定规范(Spec),多个 AI Agent 在独立的工作分支(lane/worktree)中实现,提交前必须经过另一个 Agent 的验证(verify),人类只需做最高层的决策。 官方 tagline:"Vibe code a real product. …
Agent 智能体 pullboard-dev/pullboard Tom 2026-10-09 AI Agent 协作流程
Jev-as-a-Judge:把评分跑进每一个生产 Trace · 干货攻略
Jev 是 TypeSafe AI(2026 年 9 月中旬发布)的首个 "System One" 模型——它不是语言模型,不生成任何文本,只做一件事:输入一段结构化 state,加上类型化的问题,得到带概率的决策答案。 这个定位恰好精确命中了 Agent 评测的核心形状:给定 agent 的 trace 和状态,判断这次行为是否合规、是否安全、评分几分。 …
Jay 2026-10-09 x-tips
slaterain/nv-rs · 上手攻略
nvrs 是一个从零重写《辐射:新维加斯》(Fallout: New Vegas)的实验性开源项目,使用 Rust 语言和 Bevy 游戏引擎,目标是在不使用原始引擎的情况下 1:1 再现原版游戏行为。它读取玩家自有副本的 Data 文件夹(插件、模型、纹理、音效等),无需转换或修改游戏安装目录。 项目强调行为精确还原:规则、常量、公式全部来自对原版 Fal…
工程化 slaterain/nv-rs Tom 2026-10-09 游戏引擎重实现 / Rust / Bevy