信源:X 硬核干货雷达 · 覆盖 12 账号

干货候选

  • 主题:Meta-Harness: 端到端自动化优化 Agent Harness(COLM 2026) | 来源:@omarsar0 | 链接:https://x.com/omarsar0/status/2086509069762981896 | 仓库:stanford-iris-lab/meta-harness | 论文:https://arxiv.org/abs/2603.28052 | 硬核点:用 coding agent 自动搜索/优化 harness 代码,每次迭代 10M tokens 上下文,击败手工设计 harness,Qwen3-8B ALFWorld 96.9%——自动化 harness 工程化里程碑

  • 主题:Claude 水印原理详解(SynthID-Text,无需重跑 LLM 验证) | 来源:@rasbt | 链接:https://x.com/rasbt/status/2088631263737364818 | 仓库:无 | 论文:无 | 硬核点:Watermarking 在 token 生成时注入水印密钥而非 post-hoc 检测,原理图解 + Aug 18 补充"不需要 rerun LLM"的澄清

  • 主题:Microsoft Agent Lightning v1.0 — 任意 Harness 接入 RL(约 3500 行) | 来源:@omarsar0 | 链接:https://x.com/omarsar0/status/2090078336697733531 | 仓库:无 | 论文:无 | 硬核点:微软正式发布 Agent Framework Harness(附博客 devblogs.microsoft.com/agent-framework),Agent Lightning 将任意 harness 通过 endpoint proxy 接入 GRPO 训练,模型与 harness 解耦

  • 主题:OpenAI Black Hat 披露"Hugging Face 事件"完整时间线(agent privilege escalation) | 来源:@simonw | 链接:https://x.com/simonw/status/2085877951925801274 | 仓库:无 | 论文:无 | 硬核点:17,600 次黑客动作,Linux kernel CVE 提权,Kubernetes service account 误配,完整安全踩坑复盘

  • 主题:DiG-bench — 文字冒险游戏 Discovery 能力评测新基准 | 来源:@tri_dao | 链接:https://x.com/tri_dao/status/2087677140410290302 | 仓库:无 | 论文:无 | 硬核点:Frontier 模型在文字游戏中探查能力的系统性评测基准,Tri Dao 亲推

  • 主题:Cursor Origin — Cursor(SpaceXAI)推出 GitHub 竞品代码托管,IDE 内置 | 来源:@swyx | 链接:https://x.com/swyx/status/2089467492163010836 | 仓库:无(产品/服务) | 论文:无 | 硬核点:Cursor IDE 内置代码托管 + PR + 代码审查,GitHub 同步,agent-native 功能即将推出,GitHub 宕机期间上线引发关注

  • 主题:ExtractBench — LlamaIndex 企业文档信息抽取全面基准(370 docs/67 类) | 来源:@jerryjliu0 | 链接:https://x.com/jerryjliu0/status/2088759243365044618 | 仓库:run-llama/ExtractBench | 论文:无 | 硬核点:复杂企业文档(50+ 页/10k-100k 字段)94%+ 准确率抽取,8 领域系统评测

  • 主题:Jerry Liu — 大规模文档提取 AI agent(50+ 页,94%+ 准确率) | 来源:@jerryjliu0 | 链接:https://x.com/jerryjliu0/status/2089099864554831995 | 仓库:无 | 论文:无 | 硬核点:长文档结构化抽取工程实践,含 10k-100k 字段规模问题的调优经验

其余线索

  • @omarsar0(IBM 研究):弱模型从改写中受益多于损失,强模型则相反——模型规模影响 rephrasing 策略 | https://x.com/omarsar0/status/2088675092238889461
  • @omarsar0(Dyna-2):Meta 新 scaling laws 论文,将模型容量与数据耦合而非独立作用于 loss | https://x.com/omarsar0/status/2086845790983716917
  • @cwolferesearch:rubric-based RL 在安全/对齐研究中的历史,Constitutional AI/Deliberative Alignment/Rules-based Rewards 对比 | https://x.com/cwolferesearch/status/2026151598523625626
  • @maximelabonne:Liquid AI + MacPaw 合作在 macOS 部署端侧 LFM 模型,设备端 agent 加速落地 | https://x.com/maximelabonne/status/2084979489323266176
  • @rasbt:AI 内容检测工具不可靠,Fable 可轻易构造骗过所有检测器的文本 | https://x.com/omarsar0/status/2089895053560864964
  • @swyx:OpenAI Agent Plugins spec 与 Harbor framework 规范对应,AI agent 工具标准化正在收敛 | https://x.com/swyx/status/2085500541707456908