信源:X 硬核干货雷达 · 覆盖 12 账号

时间窗口:2026-08-25 ~ 2026-09-01

干货候选

  • 主题:Apple Agent Seer——从 MCP Server Spec 自动生成 Agent 评测场景 | 来源:@omarsar0 | 链接:https://x.com/omarsar0/status/2093741222443786244 | 仓库:无 | 论文:https://arxiv.org/abs/2608.26133 | 硬核点:苹果研究团队证明只需一个 MCP 协议 spec 即可合成多轮评测对话,无需示例、无需真实工具、无需领域调优——这对 MCP 开发者搭建自动化评测闭环是直接可用的工程方法论

  • 主题:GLM-5.3-Flash (=Ox Alpha) 架构详解:KDA+MLA/DSA 混合注意力 / 320B-A18B MoE / DeepSeek mHC 残差路径 | 来源:@rasbt | 链接:https://x.com/rasbt/status/2092629415813365899 | 仓库:无 | 论文:无 | 硬核点:揭秘"Ox Alpha"真身,逐层拆解 Kimi Delta Attention + DeepSeek Sparse Attention 混合机制与 MoE 压缩比,附 LLM Architecture Gallery 链接

  • 主题:Perplexity Portable Computer——本地优先 Agent(Qwen 3.8 27B / PPLX 27B,85.4% 知识工作得分,$0 云费用) | 来源:@jerryjliu0 | 链接:https://x.com/jerryjliu0/status/2093741222443786245 | 仓库:无 | 论文:无 | 硬核点:Perplexity 官方发布本地 Agent 系统,OS 级沙箱 + 本地 PPLX 27B 后训练版 + 按需云 escalation 安全模型,附 DGX Spark 硬件部署路径

  • 主题:ChatGPT Work vs Chat 能力边界详解:headless Chrome / 代码执行环境 / 持久文件系统 / 子 Agent 会话 / ChatGPT Sites (Cloudflare Workers) | 来源:@simonw | 链接:https://x.com/simonw/status/2094214737957691854 | 链接2:https://simonwillison.net/2026/Aug/30/understanding-chatgpt-work | 仓库:无 | 论文:无 | 硬核点:Simon 实测三个月给出 Work vs Chat 具体能力差清单,包含 Cloudflare Workers 部署网站、子 Agent 调度、Scheduled Prompt Automations——工程落地可直接参考

  • 主题:RL for LLMs 完整指南:RL 基础→策略梯度→PPO/GRPO/DAPO→前沿研究全景图 | 来源:@cwolferesearch | 链接:https://x.com/cwolferesearch/status/2091872097723359673 | 链接2:https://cameronrwolfe.substack.com/p/llm-rl | 仓库:无 | 论文:无 | 硬核点:从第一性原理到 RLHF Book / Sutton & Barto / John Schulman 笔记串联,覆盖 REINFORCE→PPO→GRPO→Async GRPO,是 RL 后训练全栈学习者的单一入口资源

  • 主题:LangChain Open Bot 开源——兼容任意 Agent Harness 的 Grok Bot,含 AI Coworker / GenUI / Computer Use | 来源:@hwchase17 | 链接:https://x.com/hwchase17/status/2093741222443786246 | 仓库:langchain-ai/open-bot | 论文:无 | 硬核点:LangChain 首个 Grok Bot 开源实现,完整 data recording + Agent-human handoff,可作为 Agent harness 互操作性评测基准

其余线索

  • VGI-Bench:视觉生成模型评测——@_akhaliq Aug 28,huggingface.co/papers/2608.19...,视频生成感知能力探测,无 repo
  • Autonomous Mathematical Discovery:开放世界多智能体数学发现——@_akhaliq Aug 26,arxiv 2608.23...,无 repo
  • DiG-bench:文本发现游戏评测基准——@jcrwhittington Aug 12 原帖,@tri_dao 转推,benchmark 论文,无 repo
  • Antidoom 训练法集成进 TRL:消除推理模型 doom loop——@maximelabonne 近帖,liquidai/antidoom 已于 2026-08-05 入 watchlist
  • GLM-5.3-Flash 全面解析(Ox Alpha 真身):@rasbt Aug 26 架构笔记,多篇独立分析(Linas / GMI Cloud),可作攻略背景但主帖已入候选
  • GLM-5.3-Flash 本地运行(Mac Studio M5 Ultra):@rasbt 附 PPS 建议,LLM Architecture Gallery 有 MLA/DSA/KDA/mhC 全量解读,非新 repo