信源:X 硬核干货雷达 · 覆盖 12 账号
干货候选
- 主题:「为什么即使已有 Claude Code/Codex 自己造 harness 依然值得」| 来源:@omarsar0 | 链接:https://x.com/omarsar0/status/2095898579629916160 | 仓库:无 | 论文:无 | 硬核点:harness 失败模式、模型需求、关键旋钮的第一手经验,这种认知迁移到任何 harness 都适用,踩坑复盘类
- 主题:Claude Fable 5.1 SVG 鹈鹕 + 动画实测笔记($3.30/张 Max 模式) | 来源:@simonw | 链接:https://x.com/simonw/status/2094938927727804684 | 仓库:无 | 论文:无 | 硬核点:Fable 5.1 SVG 质量实测量化(197.5K views,125 likes),高成本推理场景选型参考
- 主题:Tencent 环境进化 Agent RL 论文——环境供给成 Agent RL 主要瓶颈 | 来源:@omarsar0 | 链接:https://x.com/omarsar0 (Sep 3 回复区) | 仓库:无 | 论文:arxiv 相关 | 硬核点:Agent RL scaling 核心约束的系统性分析,论文要点长帖,值得 bookmark
- 主题:GLM 5.3 Flash 替换 GPT-5.6 Luna 文档处理——零质量损失 80% 折扣价 | 来源:@abacaj | 链接:https://x.com/abacaj (Aug 29 profile) | 仓库:无 | 论文:无 | 硬核点:本地模型 vs API 成本效益的实操对比,115 likes 社区认可,GLM 本地可跑实操参考
- 主题:「下一代模型应减少思考 token」——推理效率优化方向分析 | 来源:@abacaj | 链接:https://x.com/abacaj/status/2093114766738718755 | 仓库:无 | 论文:无 | 硬核点:对 thinking token 机制的系统性批评与解决方向判断,模型工程选型参考
其余线索
- Sakana AI Conductor(ICLR 2026)论文分享——7B 模型通过 RL 编排其他 LLMs,GPQA-Diamond/SOTA,RL 调度拓扑设计 | 来源:@omarsar0 | 链接:https://x.com/omarsar0/status/2051306659021242635
- Gemini 3.8 Flash / 3.8 Flash Cyber 发布——agentic tasks + cyber 能力双提升 | 来源:@omarsar0 | 链接:https://x.com/omarsar0/status/2095177610930098670
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work(arxiv 2608.23) | 来源:@_akhaliq | 链接:https://x.com/_akhaliq/status/2092269176710631758
- LLM 经济模块化架构文章推荐——推理提供商经济逻辑 | 来源:@tri_dao | 链接:https://x.com/tri_dao/status/2071980068981891525