X 硬核干货雷达 · 2026-07-19

信源:X 硬核干货雷达 · 覆盖 12 账号

干货候选

  • 主题:Sakana AI Conductor(ICLR 2026)——7B 模型用 RL 编排多 LLM 协同,SOTA on GPQA-Diamond & LiveCodeBench | 来源:@omarsar0 | 链接:https://x.com/omarsar0/status/2051306659021242635 | 仓库:Sakana-AI-labs/Sakana-Fugu | 论文:ICLR 2026(待补充 arXiv) | 硬核点:RL 调度多模型协同而非单独解决问题的范式革新,含递归拓扑与动态 test-time scaling,值得写实操攻略

  • 主题:Jerry Liu 发布 LlamaParse Retrieval Harness——Agentic Retrieval 的 2026 版 RAG 管道,支持 semantic search/grep/文件导航多工具协作 | 来源:@jerryjliu0 | 链接:https://x.com/jerryjliu0/status/2071729856900215261 | 仓库:run-llama/legacy | 论文:无 | 硬核点:LlamaIndex v2 核心组件,大规模知识库 agent 自主导航的参考实现,可直接复现

  • 主题:IFStruct——Maxime Labonne 开源的 LLM 结构化输出 benchmark,测试 schema following 与 JSON/YAML 有效性 | 来源:@maximelabonne | 链接:https://x.com/maximelabonne/status/2071959400596586702 | 仓库:cleanlab/structured-output-benchmark | 论文:无 | 硬核点:揭示当前结构化输出数据集存在大量标注错误,提供清洗后的可靠 benchmark,实操价值高

  • 主题:Weak-to-Strong Generalization via Direct On-Policy Distillation——AK 分享,Direct-OPD 将小模型 RL 诱导的策略转移蒸馏到大模型,无需在大模型上跑稀疏 reward RL | 来源:@_akhaliq | 链接:https://x.com/_akhaliq/status/2077170804693803266 | 仓库:无 | 论文:https://arxiv.org/abs/2607.05394 | 硬核点:weak-to-strong 新范式,小模型 RL 作为隐式 reward 生成器,复现门槛低于直接 RL

  • 主题:LangChain Fleet Agents 真实降本数据——Harrison Chase 披露 Sonnet($0.19/run)→GLM 5.1($0.07)→Kimi K2.6($0.05),多模型 agent 路由已具成本可行性 | 来源:@hwchase17 | 链接:https://x.com/hwchase17/status/2049552801890771220 | 仓库:无 | 论文:无 | 硬核点:真实生产级成本数字,模型路由降本 3x 以上已可落地,Agent 工程量化参考

其余线索

  • omarsar0:Weak-to-Strong GraphRAG(ICLR 2026),LLM 反馈对齐弱检索器,值得留档追论文 https://x.com/omarsar0/status/1999881513220100336
  • _akhaliq:Read It Back——预训练 MLLM 作为零样本 Reward Model for 文生图评估 https://x.com/_akhaliq/status/2077420925503353037
  • _akhaliq:Fully Open Video MLLM——高效通用视频理解,开源模型卡值得一看 https://x.com/_akhaliq/status/2078179729660629266
  • simonw:LLM 0.32a0 发布,重大向后兼容重构,支持 reasoning 模型,CLI/Python 工具链值得升级 https://x.com/simonw/status/2049567761136058699
  • hwchase17:LangChain Open-Source Software Factory——Dcode/openswe/openswe review/open wiki 全链路 OSS 软件工厂 https://x.com/hwchase17/status/2078576948922708200
  • omarsar0:Autonomous Long-Running Coding Agents 深度长文——从 prompting 到控制系统的范式转变 https://x.com/omarsar0/article/2065880971031834786