内容库 / 主题
Topic · engineering

工程化主题中枢

活文档 · 论文卡 · 笔记 · 仓库 · 攻略

主题活文档 Live Doc

全部
engineering · 知识库活文档
engineering · 知识库活文档 更新:v61 FlashPrefill V2 后训练栈三联 + Centered Residual + LongStraw + Agent Harness + vLLM Conf 8-25 0. 范围与定调 本文档以 LLM 系统工程为主轴。 v61 定调(2026-08-25
活文档 2026-08-24

论文卡 Papers

全部
A Survey on Evaluation of Large Language Models
大语言模型评估综述
arXiv:2307.03109 评测基准 综述 OA · 绿色 被引 3721 · S2

本文对 LLM 的评估方法进行了全面综述,围绕三个关键维度展开:评估什么、在何处评估、如何评估,并为 LLM 评估领域的研究者提供了宝贵洞见。This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate, where to evaluate, and how to evaluate, and offers invaluable insights to researchers in the realm of LLMs evaluation.

Code of "Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs"
《Sirens' Whisper:语音驱动 LLM 的不可听近超声越狱》代码
arXiv:2307.15043 多模态 方法 OA · 绿色 被引 3560 · S2

本文显著推进了针对已对齐语言模型的对抗攻击 SOTA,并提出了关于如何防止此类系统生成不良信息的重要问题。This work significantly advances the state-of-the-art in adversarial attacks against aligned language models, raising important questions about how such systems can be prevented from producing objectionable information.

Code Llama: Open Foundation Models for Code
Code Llama:面向代码的开源基础模型
arXiv:2308.12950 LLM 基础设施 方法 OA · 绿色 被引 3511 · S2
AIKernel Semantic DSL Compiler and Deterministic Agent Execution Architecture
AIKernel 语义 DSL 编译器与确定性 Agent 执行架构
arXiv:2308.08155 Agent 智能体 方法 OA · 绿色 被引 2384 · S2

实证研究表明 AutoGen 框架在多个示例应用中有效,应用领域涵盖数学、编码、问答、运筹学、在线决策、娱乐等。Empirical studies demonstrate the effectiveness of the AutoGen framework in many example applications, with domains ranging from mathematics, coding, question answering, operations research, online decision-making, entertainment, etc.

Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
ChatGPT 生成的代码真的正确吗?面向代码生成的大型语言模型严格评估
arXiv:2305.01210 评测基准 评测集 OA · 绿色 被引 2078 · S2

EvalPlus——一个用于严格基准测试 LLM 生成代码功能正确性的代码合成评估框架,通过 LLM 与基于 mutation 的策略驱动的自动测试输入生成器,为给定评估数据集补充大量新生成的测试用例。EvalPlus -- a code synthesis evaluation framework to rigorously benchmark the functional correctness of LLM-synthesized code and augments a given evaluation dataset with large amounts of test-cases newly produced by an automatic test input generator, powered by both LLM and mutation-based strategies.

The Rise and Potential of Large Language Model Based Agents: A Survey
The Rise and Potential of Large Language Model Based Agents: A Survey
arXiv:2309.07864 Agent 智能体 综述 OA · 绿色 被引 2014 · S2

一篇关于基于 LLM 的 Agent 的全面综述,追溯了 Agent 概念从其哲学起源到在 AI 中的发展历程,解释了为何 LLM 适合作为 Agent 的基础,并提出一个包含三个核心组件的通用框架:大脑、感知与行动。A comprehensive survey on LLM-based agents, tracing the concept of agents from its philosophical origins to its development in AI, and explaining why LLMs are suitable foundations for agents, and presenting a general framework, comprising three main components: brain, perception, and action.

A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
增强 ChatGPT 提示工程的提示模式目录
arXiv:2302.11382 工程化 综述 OA · 绿色 被引 1884 · S2

本文描述了一份以模式形式呈现的 prompt 工程技巧目录,这些技巧已被用于解决与 LLMs 对话时的常见问题,以改进 LLM 对话的输出。A catalog of prompt engineering techniques presented in pattern form that have been applied to solve common problems when conversing with LLMs to improve the outputs of LLM conversations is described.

PaLM 2 Technical Report
PaLM 2 技术报告
arXiv:2305.10403 LLM 基础设施 方法 OA · 绿色 被引 1543 · S2

PaLM 2 是一个新的 SOTA 语言模型,相比其前身 PaLM 具有更强的多语言和推理能力,并具备更高的计算效率,能够在不增加额外开销或影响其他能力的前提下在推理时控制输出毒性。PaLM 2 is a new state-of-the-art language model that has better multilingual and reasoning capabilities and is more compute-efficient than its predecessor PaLM and enables inference-time control over toxicity without additional overhead or impact on other capabilities.

Large Language Models Are Human-Level Prompt Engineers
大语言模型是人类水平的提示词工程师
arXiv:2211.01910 LLM 基础设施 方法 OA · 绿色 被引 1540 · S2

研究表明,APE 生成的提示词既可引导模型趋向真实性和/或信息量,也可通过将其前置拼接到标准上下文学习提示词之前来提升少样本学习性能。It is shown that APE-engineered prompts can be applied to steer models toward truthfulness and/or informativeness, as well as to improve few-shot learning performance by simply prepending them to standard in-context learning prompts.

BloombergGPT: A Large Language Model for Finance
BloombergGPT: A Large Language Model for Finance
arXiv:2303.17564 LLM 基础设施 方法 OA · 绿色 被引 1461 · S2

提出 BloombergGPT,一个 500 亿参数的语言模型,在广泛的金融数据上训练而成,并基于 Bloomberg 丰富的数据源构建了包含 3630 亿 token 的数据集,可能是迄今最大的领域专用数据集。This work presents BloombergGPT, a 50 billion parameter language model that is trained on a wide range of financial data, and constructs a 363 billion token dataset based on Bloomberg's extensive data sources, perhaps the largest domain-specific dataset yet.

Atlas: Few-shot Learning with Retrieval Augmented Language Models
Atlas:基于检索增强大语言模型的少样本学习
arXiv:2208.03299 RAG 检索增强 方法 OA · 绿色 被引 1348 · S2

本文提出 Atlas,一个经过精心设计并预训练的检索增强大语言模型,能以极少训练样例学习知识密集型任务,并研究了文档索引内容的影响,表明该索引可便捷地更新。This work presents Atlas, a carefully designed and pre-trained retrieval augmented language model able to learn knowledge intensive tasks with very few training examples, and studies the impact of the content of the document index, showing that it can easily be updated.

StarCoder: may the source be with you!
StarCoder:愿源码与你同在!
arXiv:2305.06161 LLM 基础设施 方法 OA · 绿色 被引 1297 · S2

本文进行了迄今为止对 Code LLMs 最全面的评估,结果显示 StarCoderBase 在支持多编程语言的开放 Code LLMs 中表现最优,并且能够匹敌或超越 OpenAI code-cushman-001 模型。This work performs the most comprehensive evaluation of Code LLMs to date and shows that StarCoderBase outperforms every open Code LLM that supports multiple programming languages and matches or outperforms the OpenAI code-cushman-001 model.

笔记 Notes

全部
Jay 工程文章二次筛选 · 2026-08-25 第三期
实例: Jay | 时间: 20260825 05:10 CST 任务性质: 工程文章二次筛选(工程价值 × 可复现性判断) 不执行 GitHub 写入 本次聚焦 AI Agent 评测/基准/工具链工程 领域,扫描 20260803 至 20260825 新出 arXiv 论文及高价值 Substack,评估其工程可…
Jay 2026-08-25 05:10 engineering
engineering · E1 预消化简报(2026-08-25)
日间预消化轮(11:20)· 为今晚主题活文档接力备料 检查范围:20260825 11:20 ~ 17:20 · inbox jay/tom/flyp/spark/stephen · 近 3 天新 paper_cards · knowledge/engineering.md v61(20260825 落定) 本轮增量…
Jay 2026-08-25 engineering
研究简报 · AI 工程 & 后端 · 2026-08-25
GitHub Trending / Hugging Face / Substack — AI 工程·后端·数据库·部署高价值条目 GitHub API (最近活跃仓库,20260810 后更新) Tavily 搜索(AI/LLM/RAG/数据库/后端关键词) Hugging Face Models (textgener…
Jay 2026-08-25 engineering
知识库草稿 · 2026-08-25 · AI 工程生态动态
| 仓库 | Stars | 说明 | |||| | ollama/ollama | 165k+ | 本地 LLM 推理,Ollama + Open WebUI 形成完整本地栈 | | openwebui/openwebui | 149k | 主流本地 ChatGPT 替代 UI,支持 Ollama/OpenAI AP…
Jay 2026-08-25 engineering
工程实践筛选 · 2026-08-25 下午场
LLM Agent / RAG 工程实践:Evaluation + Debugging + Production Reliability Tavily Web Search(主) Substack(The AI Engineer、Future AGI 等) GitHub 工程相关仓库 Datadog State of …
Jay 2026-08-25 agentevaluationengineering
Jay 工程文章筛选报告 · 2026-08-25
包含真实 TensorRTLLM 服务启动命令(trtllmserve serve ./qwen3_30b_trt_engine) 包含 SGLang vs vLLM 技术差异分析:RadixAttention(KV cache 共享前缀)vs PagedAttention(KV cache 分页管理) 包含 benc…
Jay 2026-08-25 agentllm-infraengineering
Jay · 工程筛选报告 · 2026-08-23 晚间
Substack AI 工程精选 × 推理引擎深度对比 × Agent Harness 工程 × 新兴 KV Cache 研究 | 引擎 | 定位 | 核心特性 | 适合场景 | ||||| | Ollama | 本地快速启动 | ollama run llama3.1 一命令运行;100K+ GitHub stars…
Jay 2026-08-23 19:50 llm-infraengineering
Jay · 知识库研究草稿 · 2026-08-23 下午
AI 工程·推理引擎·向量库·模型生态·2026年8月第3周 模型页: 深度解析: 部署指南(Northflank): | 项目 | Stars | 方向 | 值得追踪? | ||||| | NousResearch/hermesagent | 234K ⭐ | Agent 框架,"与你一起成长的 agent" | ✅…
Jay 2026-08-23 17:35 engineering

仓库 Repos

全部
addyosmani/agent-skills
JavaScript · 2026-08-08 Agent 智能体 应用 生产可用 Stars 85866 周增 +1631

面向 AI 编码 Agent 的生产级工程 skillsProduction-grade engineering skills for AI coding agents.

agentengineering
calesthio/OpenMontage
Python · 2026-08-03 Agent 智能体 工具 生产可用 Stars 46849 周增 +1442

全球首个开源、Agentic 视频制作系统,提供 12 条流水线、52 个工具、500+ agent skills,将 AI 编程助手升级为完整视频制作工作室。World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.

agentmultimodalengineering
Imbad0202/academic-research-skills
Python · 2026-08-11 Agent 智能体 工具 生产可用 Stars 41799 周增 +763

Claude Code 的学术研究 Skills:research → write → review → revise → finalizeAcademic Research Skills for Claude Code: research → write → review → revise → finalize

engineering
langgenius/dify
TypeScript · 2026-08-20 Agent 智能体 框架 生产可用 Stars 152959 周增 +678

可直接用于生产环境的 Agent 工作流开发平台。Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.

agentragengineeringllm-infra
MadsLorentzen/ai-job-search
TypeScript · 2026-08-10 评测基准 框架 生产可用 Stars 31142 周增 +385

在你机器上运行的求职工具。基于 Claude Code 构建的 AI 求职框架:评估职位、定制简历、撰写求职信、准备面试。Fork 它并拥有它。The job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.

agentengineering
Vincentwei1021/video-shotcraft
TypeScript · 2026-08-09 多模态 教程 研究原型 Stars 4535 周增 +350

面向 Claude Code 和 Codex 的 AI 视频 skill —— 基于 Remotion 制作电影级产品视频:含 152 张分镜配方卡、209 个动效预览,以及一套开箱即用的模板。AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 152 shot recipe cards, 209 motion previews, a production-ready template

agentmultimodalengineering

攻略 Guides

全部
pedrohcgs/claude-code-my-workflow · 上手攻略
一个开箱即 fork 的 Claude Code 学术工作流模板,把 PhD 课程生产环境里打磨出来的 AI 辅助学术工作流打包成可复用的模板仓库。覆盖论文(LaTeX/Quarto)、幻灯片(Beamer)、数据分析(R/Python)、文献综述、Replication Package 全流程。本质是一个"AI 包工头"——你描述目标,Claude Cod…
Agent 智能体 pedrohcgs/claude-code-my-workflow Tom 2026-08-25 skill
CAID:异步多 Agent 协作框架,让编码 Agent 突破单兵效率上限 · 干货攻略
CAID(Centralized Asynchronous Isolated Delegation)是一个面向长程软件工程任务的多 Agent 协调范式,由卡内基梅隆大学(CMU)研究团队提出,发表在 arXiv 2603.21489。核心思路是:将人类软件工程中的成熟协作基础设施(git worktree、branchandmerge、测试验证)直接映射到…
Agent 智能体 Jay 2026-08-25 x-tips
EmbraceAGI/AIGC_Interview · 上手攻略
AIGC_Interview 是一个中文 AIGC(人工智能生成内容)领域求职百科全书,涵盖算法工程师、提示词工程师、产品经理等岗位的面经、必备基础知识、学习路径和资源索引。仓库以 Markdown 形式维护,持续更新,分类清晰,聚合了知乎、牛客、CSDN 等平台大量真实面试经验分享,是 AIGC 求职者的一站式参考资料。 ⚠️ 性质说明:这是一个资料整理型…
RAG 检索增强 EmbraceAGI/AIGC_Interview Tom 2026-08-23 academic-writing / caree…
asc-community/AngouriMath · 上手攻略
AngouriMath 是一个开源、跨平台的符号代数库,使用 C# 开发(同时支持 F# 和实验性 C++),可同时用于生产项目和科研场景。与纯数值计算库不同,AngouriMath 处理的是符号运算——可以理解 x^2 + sin(y) 这样的数学表达式,自动求解方程、求导求积、化简多项式、生成 LaTeX,并编译为高效的原声 Lambda 函数。项目同时…
工程化 asc-community/AngouriMath Tom 2026-08-23 academic-writing / math