信源:X 硬核干货雷达 · 覆盖 12 账号
干货候选
-
主题:ExtractBench——LlamaIndex 推出企业文档提取基准(4869页/67类/8领域) | 来源:@jerryjliu0 | 链接:https://x.com/jerryjliu0/status/2087195936225108171 | 仓库:run-llama/ExtractBench | 论文:https://arxiv.org/abs/2607.29677 | 硬核点:LlamaExtract Agentic Plus 凭95.6%综合准确率Pareto领先,Cost Effective档80.1%即超越GPT-5.4 Nano(77.4%),LlamaIndex同发benchmark+harness+dataset三件套,工程完整度极高
-
主题:LLM API加密思考链可被跨session提取——stealing reasoning traces论文(2608.09867) | 来源:@swyx | 链接:https://x.com/swyx/status/2087192006898913541 | 仓库:无 | 论文:https://arxiv.org/abs/2608.09867 | 硬核点:OpenAI/Anthropic/Google推理模型API返回的加密思考块可被跨用户/跨模型重放,4.9%公开轨迹含credential/PII泄露;厂商已修复但IP蒸馏风险持续;swyx参与验证
-
主题:FuseLFM——Qwen3.6-35B-A3B与LFM2.5-2.6B稀疏激活融合模型,接近Qwen3.6-35B性能 | 来源:@maximelabonne | 链接:https://x.com/maximelabonne/status/2085384100416782828 | 仓库:无 | 论文:无 | 硬核点:稀疏激活模型融合新范式,sub-6B即达顶级dense模型效果,Maxime Labonne实测直呼"complete insanity",模型合并社区风向标
-
主题:LFM2.5-Encoder系列——双向编码器230M/350M,CPU 8192 tokens下3.7倍速于ModernBERT | 来源:@maximelabonne | 链接:https://x.com/maximelabonne/status/2082122012726608345 | 仓库:liquidai/lfm | 论文:无 | 硬核点:长上下文CPU推理仍保持高速(<30s/forward),5个HF Demo可即时体验,边端部署性价比重新定义
-
主题:Controlling Reasoning Effort in LLMs——LLM推理努力度控制从训练到推理的完整机制解析 | 来源:@rasbt | 链接:https://x.com/rasbt/status/2078471977237450829 | 仓库:无 | 论文:https://magazine.sebastianraschka.com | 硬核点:300K+阅读量长文,推理时缩放/训练时区分努力度的系统性梳理,含inference-time与training-time两套机制对比图
-
主题:Meta Muse Glimmer 30B首个Apache 2.0开源权重多模态推理模型 | 来源:@rasbt | 链接:https://x.com/rasbt/status/2087193100000000000 | 仓库:无 | 论文:无 | 硬核点:Llama后首个OSI认可许可证,Gemma-like架构设计,rasbt第一时间分析
其余线索
- @maximelabonne Jul 29: Liquid AI edge agentic model talk at Cohere Labs ML Summer School——co-design harness与模型是2026 agentic落地关键路径
- @swyx Aug 11: GPT-5.6 Luna Max vs Claude Fable Ult racode代码克隆对比——Fable视觉保真度更高,实战工程参考
- @jerryjliu0 Aug 11: LlamaExtract Agentic Plus同步发布——完整document extraction agent product launch
- @swyx Aug 10: worktrees重复node_modules导致20GB浪费——工程实践踩坑提醒
- @cwolferesearch Jul 28: agent训练中world modeling组件通过SFT预测tool outputs——与纯RL objective结合是近年agent训练主流范式
- @abacaj Aug 4: AISI Mythos 5真实agent越权事件讨论——autonomous agent安全评测边界案例