研究库 主题路线
内容库 / 主题
Topic · engineering

工程化主题中枢

活文档 · 论文卡 · 笔记 · 仓库 · 攻略

主题活文档 Live Doc

全部
engineering · 知识库活文档
engineering · 知识库活文档 更新:v141精简。JIL攻击+VLA Workload+Behavior-Preserving KV+Agentic-ZTA+Rogue Agent事件+5新立标+1共识149+1争议171+6arXiv+5URL 0. 范围与定调 v141 = 2026-10-08 09:
活文档 2026-10-08

论文卡 Papers

全部
Evaluating Large Language Models Trained on Code
Evaluating Large Language Models Trained on Code
arXiv:2107.03374 评测基准 方法 OA · 绿色 被引 12020 · S2

发现对 GPT 语言模型进行重复采样,是为困难 prompt 生成可用解的一种出乎意料的有效策略,并讨论了部署强大代码生成技术在安全性、安全保障与经济等方面的潜在更广泛影响。It is found that repeated sampling from the GPT language model is a surprisingly effective strategy for producing working solutions to difficult prompts, and the potential broader impacts of deploying powerful code generation technologies, covering safety, security, and economics are discussed.

Scaling Instruction-Finetuned Language Models
指令微调语言模型的规模化
arXiv:2210.11416 工程化 方法 OA · 绿色 被引 4489 · S2

研究发现,在上述多个维度上进行的指令微调可显著提升多种模型类别(PaLM、T5、U-PaLM)、多种提示设定以及多种评测基准(MMLU、BBH、TyDiQA、MGSM、开放式生成)上的表现。It is found that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups, and evaluation benchmarks (MMLU, BBH, TyDiQA, MGSM, open-ended generation).

A Survey on Evaluation of Large Language Models
大语言模型评估综述
arXiv:2307.03109 评测基准 综述 OA · 绿色 被引 3888 · S2

本文对 LLM 的评估方法进行了全面综述,围绕三个关键维度展开:评估什么、在何处评估、如何评估,并为 LLM 评估领域的研究者提供了宝贵洞见。This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate, where to evaluate, and how to evaluate, and offers invaluable insights to researchers in the realm of LLMs evaluation.

Code of "Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs"
《Sirens' Whisper:语音驱动 LLM 的不可听近超声越狱》代码
arXiv:2307.15043 多模态 方法 OA · 绿色 被引 3834 · S2

本文显著推进了针对已对齐语言模型的对抗攻击 SOTA,并提出了关于如何防止此类系统生成不良信息的重要问题。This work significantly advances the state-of-the-art in adversarial attacks against aligned language models, raising important questions about how such systems can be prevented from producing objectionable information.

Code Llama: Open Foundation Models for Code
Code Llama:面向代码的开源基础模型
arXiv:2308.12950 LLM 基础设施 方法 OA · 绿色 被引 3641 · S2
AIKernel Semantic DSL Compiler and Deterministic Agent Execution Architecture
AIKernel 语义 DSL 编译器与确定性 Agent 执行架构
arXiv:2308.08155 Agent 智能体 方法 OA · 绿色 被引 2856 · S2

实证研究表明 AutoGen 框架在多个示例应用中有效,应用领域涵盖数学、编码、问答、运筹学、在线决策、娱乐等。Empirical studies demonstrate the effectiveness of the AutoGen framework in many example applications, with domains ranging from mathematics, coding, question answering, operations research, online decision-making, entertainment, etc.

Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
ChatGPT 生成的代码真的正确吗?面向代码生成的大型语言模型严格评估
arXiv:2305.01210 评测基准 评测集 OA · 绿色 被引 2333 · S2

EvalPlus——一个用于严格基准测试 LLM 生成代码功能正确性的代码合成评估框架,通过 LLM 与基于 mutation 的策略驱动的自动测试输入生成器,为给定评估数据集补充大量新生成的测试用例。EvalPlus -- a code synthesis evaluation framework to rigorously benchmark the functional correctness of LLM-synthesized code and augments a given evaluation dataset with large amounts of test-cases newly produced by an automatic test input generator, powered by both LLM and mutation-based strategies.

The Rise and Potential of Large Language Model Based Agents: A Survey
The Rise and Potential of Large Language Model Based Agents: A Survey
arXiv:2309.07864 Agent 智能体 综述 OA · 绿色 被引 2119 · S2

一篇关于基于 LLM 的 Agent 的全面综述,追溯了 Agent 概念从其哲学起源到在 AI 中的发展历程,解释了为何 LLM 适合作为 Agent 的基础,并提出一个包含三个核心组件的通用框架:大脑、感知与行动。A comprehensive survey on LLM-based agents, tracing the concept of agents from its philosophical origins to its development in AI, and explaining why LLMs are suitable foundations for agents, and presenting a general framework, comprising three main components: brain, perception, and action.

A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
增强 ChatGPT 提示工程的提示模式目录
arXiv:2302.11382 工程化 综述 OA · 绿色 被引 1963 · S2

本文描述了一份以模式形式呈现的 prompt 工程技巧目录,这些技巧已被用于解决与 LLMs 对话时的常见问题,以改进 LLM 对话的输出。A catalog of prompt engineering techniques presented in pattern form that have been applied to solve common problems when conversing with LLMs to improve the outputs of LLM conversations is described.

Large Language Models Are Human-Level Prompt Engineers
大语言模型是人类水平的提示词工程师
arXiv:2211.01910 LLM 基础设施 方法 OA · 绿色 被引 1615 · S2

研究表明,APE 生成的提示词既可引导模型趋向真实性和/或信息量,也可通过将其前置拼接到标准上下文学习提示词之前来提升少样本学习性能。It is shown that APE-engineered prompts can be applied to steer models toward truthfulness and/or informativeness, as well as to improve few-shot learning performance by simply prepending them to standard in-context learning prompts.

PaLM 2 Technical Report
PaLM 2 技术报告
arXiv:2305.10403 LLM 基础设施 方法 OA · 绿色 被引 1564 · S2

PaLM 2 是一个新的 SOTA 语言模型,相比其前身 PaLM 具有更强的多语言和推理能力,并具备更高的计算效率,能够在不增加额外开销或影响其他能力的前提下在推理时控制输出毒性。PaLM 2 is a new state-of-the-art language model that has better multilingual and reasoning capabilities and is more compute-efficient than its predecessor PaLM and enables inference-time control over toxicity without additional overhead or impact on other capabilities.

BloombergGPT: A Large Language Model for Finance
BloombergGPT: A Large Language Model for Finance
arXiv:2303.17564 LLM 基础设施 方法 OA · 绿色 被引 1549 · S2

提出 BloombergGPT,一个 500 亿参数的语言模型,在广泛的金融数据上训练而成,并基于 Bloomberg 丰富的数据源构建了包含 3630 亿 token 的数据集,可能是迄今最大的领域专用数据集。This work presents BloombergGPT, a 50 billion parameter language model that is trained on a wide range of financial data, and constructs a 363 billion token dataset based on Bloomberg's extensive data sources, perhaps the largest domain-specific dataset yet.

笔记 Notes

全部
engineering · E1 预消化简报(2026-10-09)
E1 日间预消化轮 · engineering · 20261009 11:20 CST · Jay 锚定基线:Oct8 e1prep(v141 锚定,JIL/VLA/UndoBench/AgenticZTA/Rogue Agent)+ Oct7 e1prep(v141 综合轮产出) | 来源 | 文件 | 时间 | …
Jay 2026-10-09 engineering
工程筛选报告 · Jay · 2026-10-09 第 3 轮(19:50 UTC+8)
Reasoning Model 安全工程 + MCP 生产实践深度 | # | 条目 | 来源 | 类型 | 工程价值 | |||||| | A | FastMCP 生产 7 大错误(BigDataBoutique) | Medium/Blog | 工程实践 | ⭐⭐⭐⭐ | | B | EchoCoT:隐藏 CoT …
Jay 2026-10-09 engineeringrisk
知识库草稿 · Jay · 2026-10-09
LLM推理引擎三分天下 · Agent框架2026生产选型 · HuggingFace开源模型生态 · MLOps K8s GPU部署 · 分布式OLAP · arXiv投机解码新论文 HuggingFace TGI 已进入维护模式,2026年晚期只剩三个活跃开源引擎:vLLM、SGLang、MAX(Modular) …
Jay 2026-10-09 agentllm-infraengineering
每日简报 · 2026-10-09(周五)
| 渠道 | 工具 | 备注 | |||| | Exa Web Search | exa.web_search_exa | 数据库/后端架构、K8s/Docker、云原生排障、CSDN 工程实践、存储引擎/WAL/MVCC、Substack | | 已有知识库 | /shared/researchkb/inbox/ja…
Jay 2026-10-09 engineeringdatabase
Jay 工程实践筛选 · 2026-10-08 15:05 (UTC+8)
Agent 生产静默失败 · Eval/Observability 差距 · RAG 失败模式 · Substack 高价值工程洞察 arXiv: LLM agent production failures, silent failures taxonomy Substack: theaiengineer, futur…
Jay 2026-10-08 engineering
Jay 工程实践筛选 · 2026-10-08 晚间 (19:50 UTC+8)
Prefill/Decode Disaggregation · Mooncake KV Transfer · vLLM MORIIO · Distributed Inference Networking · PagedAttention 内核设计 Tavily 深度检索:vLLM MORIIO、llmd NIXL、AW…
Jay 2026-10-08 engineering
Jay 工程实践筛选 · 2026-10-08
LLM Inference & RAG 生产工程实践 · vLLM/SGLang 对比 · Prompt 管理 · Chunking 排障 学术:arXiv (LLM inference optimization, production systems) 工程博客:Spheron, Network Bachelor, …
Jay 2026-10-08 engineering
engineering · E1 预消化简报(2026-10-08)
本轮 E1 日间预消化(engineering 主题)检查了以下来源: /shared/researchkb/organized/queue/workqueue.md ✅ 已读 /shared/researchkb/inbox/{jay,tom,flyp,spark,stephen}/ 近 2 天新文件 → 无新增(各…
Jay 2026-10-08 engineering

仓库 Repos

全部
addyosmani/agent-skills
JavaScript · 2026-08-08 Agent 智能体 应用 生产可用 Stars 85866 周增 +1631

面向 AI 编码 Agent 的生产级工程 skillsProduction-grade engineering skills for AI coding agents.

agentengineering
calesthio/OpenMontage
Python · 2026-08-03 Agent 智能体 工具 生产可用 Stars 46849 周增 +1442

全球首个开源、Agentic 视频制作系统,提供 12 条流水线、52 个工具、500+ agent skills,将 AI 编程助手升级为完整视频制作工作室。World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.

agentmultimodalengineering
langgenius/dify
TypeScript · 2026-09-09 Agent 智能体 框架 生产可用 Stars 155117 周增 +755

可直接用于生产环境的 Agent 工作流开发平台。Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.

agentragengineeringllm-infra
MadsLorentzen/ai-job-search
Python · 2026-10-05 评测基准 框架 生产可用 Stars 45021 周增 +673

在你机器上运行的求职工具。基于 Claude Code 构建的 AI 求职框架:评估职位、定制简历、撰写求职信、准备面试。Fork 它并拥有它。The job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.

agentengineering
Imbad0202/academic-research-skills
Python · 2026-10-09 Agent 智能体 工具 生产可用 Stars 51057 周增 +532

Claude Code 的学术研究 Skills:research → write → review → revise → finalizeAcademic Research Skills for Claude Code: research → write → review → revise → finalize

engineering
Vincentwei1021/video-shotcraft
TypeScript · 2026-08-09 多模态 教程 研究原型 Stars 4535 周增 +350

面向 Claude Code 和 Codex 的 AI 视频 skill —— 基于 Remotion 制作电影级产品视频:含 152 张分镜配方卡、209 个动效预览,以及一套开箱即用的模板。AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 152 shot recipe cards, 209 motion previews, a production-ready template

agentmultimodalengineering

攻略 Guides

全部
basementstudio/xmcp · 上手攻略
xmcp 是 basement.studio 出品的 TypeScript 框架,专注于快速构建和部署 MCP(Model Context Protocol)服务器。MCP 是 Anthropic 提出的标准化协议,让 AI 助手(如 Claude、ChatGPT)能够调用外部工具和资源;xmcp 的目标是把这一过程变得极度友好——零配置注册、热重载开发、一…
Agent 智能体 basementstudio/xmcp Tom 2026-10-10 MCP 框架 / TypeScript
Jev-as-a-Judge:把评分跑进每一个生产 Trace · 干货攻略
Jev 是 TypeSafe AI(2026 年 9 月中旬发布)的首个 "System One" 模型——它不是语言模型,不生成任何文本,只做一件事:输入一段结构化 state,加上类型化的问题,得到带概率的决策答案。 这个定位恰好精确命中了 Agent 评测的核心形状:给定 agent 的 trace 和状态,判断这次行为是否合规、是否安全、评分几分。 …
Jay 2026-10-09 x-tips
hjxwz123/Aivory · 上手攻略
Aivory 是一个自部署的 AI 对话与研究平台,将多模型聊天、代码执行、知识库检索、Deep Research 和团队协作整合在一个 Web 界面中。核心卖点是"多工具串联执行"——用户发一条指令,编排器可以在一次对话内自动完成搜索→抓取网页→运行 Python 分析数据→生成文件,最多 48 次工具调用跨越 12 轮模型循环,无需人工介入。⚠️ 公开 …
RAG 检索增强 hjxwz123/Aivory Tom 2026-10-09 AI 平台 / 自部署
Liquid AI d1 决策模型:零输出 Token 跑分类/路由/评分 · 干货攻略
d1 是 Liquid AI 于 2026 年 9 月 29 日发布的决策模型(Decision Model),2026 年 10 月 5 日追加视觉支持后正式对外。 核心设计思路与语言模型相反:不给文字答案,而是对同一个输入直接输出概率分布——每道题一次前向传播,零 Token 生成。 d1 的官方定位是替代那些「本不需要 LLM 做的事情」:分类、路由、…
无(主产品为 Liquid AI 托管 API,非开源权重) Jay 2026-10-08 x-tips