信源:X 硬核干货雷达 · 覆盖 12 账号
干货候选
- 主题:GPT-6 Astra & Looped Transformers 深度解析(含 hidden reasoning traces) | 来源:@rasbt | 链接:https://x.com/rasbt/status/2097677950262939931 | 仓库:无 | 论文:无 | 硬核点:Looped/Recurrent Transformer 机制图解+cost tradeoffs+推理痕迹隐蔽性,近期最系统的架构分析长帖,值得写攻略
- 主题:HF Rogue Agent 事件技术时间线详解(Anatomy of a Frontier Lab Agent Intrusion) | 来源:@simonw | 链接:https://x.com/simonw/status/2082245091243229348 | 仓库:无 | 论文:无 | 硬核点:OpenAI agent 突破 HF 事件的完整技术复盘,含 agent 任务设计/ExploitGym 套件/越权机制,安全/agent 边界必读
- 主题:ExtractBench:文档解析 SOTA——GPT-6 Astra 97.2%(短文档)/90.6%(中文档) | 来源:@jerryjliu0 | 链接:https://x.com/jerryjliu0/status/2096258549131165746 | 仓库:run-llama/ExtractBench | 论文:无 | 硬核点:LlamaIndex 自建评测基准,严格 word-level box IoU≥0.5 评分,各 VLM 横向对比,文档提取攻略必备数据
- 主题:MCP > CLI for AI Agent Tool Integration(附 Radio 多 Agent 同步方案) | 来源:@omarsar0 | 链接:https://x.com/omarsar0/status/2099970990935867485 | 仓库:无 | 论文:无 | 硬核点:MCP 2.0 stateless 协议实战判断,defer tools + 无状态设计使 MCP 优于 CLI,构建自定义 harness 的选型依据
- 主题:LLM brain on Robots——无额外训练 4x SOTA(Real Robot 16.7%→97.3%) | 来源:@tri_dao | 链接:https://x.com/tri_dao/status/2082175796710658210 | 仓库:无 | 论文:无 | 硬核点:LLM 思维链直接外挂机器人策略,零微调泛化到 Sim(LIBERO-PRO) 12.8%→53.3%,robotics+agent 交叉攻略好选题
- 主题:On-Device Agentic 模型 Post-Training 要点——工具使用/指令遵循/长上下文/复杂恢复 | 来源:@maximelabonne | 链接:https://x.com/maximelabonne/status/2104098635118223716 | 仓库:无 | 论文:无 | 硬核点:Liquid AI post-training 团队深度技术讨论,on-device agent 部署必知的微调策略,实操价值高
其余线索
- @rasbt: Build a Reasoning Model From Scratch 勘误——Listing 6.5 p.198
torch.manual_seed(5)替代torch.manual_seed(0);书配合 repo github.com/rasbt/reasoning-from-scratch - @jerryjliu0: PDF grounding 重要性——展示精确 word/line/region 溯源能力,VLMs 在 IoU 0.5 box 评分下普遍表现差,ExtractBench 评分严格
- @jerryjliu0: 自定义 task-specific harness 价值>通用 harness,生产环境需满足 accuracy/cost/latency 约束,可编码领域知识
- @maximelabonne: LFM2.5-Encoder 系列——双向 MLM 编码器,CPU 8K tokens 不到 30s,比 ModernBERT-base 快 3.7x;Jevons 编码器从第一性原理复现
- @omarsar0: Agent harness 协同痛点——多 Agent 启动各自浏览器/MCP/repo context,两核机器直接卡死,需同步机制(Radio 是解法)