信源:X 硬核干货雷达 · 覆盖 12 账号
干货候选
-
主题:SAS — 通过端到端 LM loss 训练 selector 实现注意力稀疏化 | 来源:@_akhaliq | 链接:https://x.com/_akhaliq/status/2099511334987522298 | 仓库:无 | 论文:待查 arxiv | 硬核点:context ranking 不再靠贪心选择,selector 端到端可学习替代 hard mask,是 attention 稀疏化的新训练范式
-
主题:Apodex 1.1 — 复杂工作场景 Agentic Intelligence scaling 论文 | 来源:@_akhaliq | 链接:https://x.com/_akhaliq/status/2092269176710631758 | 仓库:无 | 论文:待查 arxiv | 硬核点:Agent scaling law 量化框架,scaling agent 能力边界的系统性benchmarking 方法论
-
主题:Skill Entropy — 用 skill entropy 度量长程任务中能力获取的 benchmark | 来源:@_akhaliq | 链接:https://x.com/_akhaliq/status/2085414421308801399 | 仓库:无 | 论文:待查 arxiv | 硬核点:long-horizon benchmark 新指标视角,entropy 刻画 skill acquisition 质量比单一任务准确率更本质
-
主题:MOPD — 2026 post-training 新范式 Multi-Teacher On-Policy Distillation | 来源:@cwolferesearch | 链接:https://x.com/cwolferesearch/status/2095256315476000817 | 仓库:无 | 论文:待查 arxiv | 硬核点:单模型从多教师策略中吸收多种能力,2026 RL/distillation 交叉点的新训练思路
-
主题:E-CommerceBench — 365天电商模拟评测基准(arxiv 2608.30730, QwenLM) | 来源:@omarsar0 | 链接:https://x.com/omarsar0/status/2076700688344613234 | 仓库:QwenLM/E-CommerceBench | 论文:arxiv 2608.30730 | 硬核点:真实电商场景长程 agent 评测数据集,弥补现有 benchmark 缺乏时序决策的空白
-
主题:AgentSky — 世界首个 Agent Market,40+ agents (Claude Code / Codex / Hermes / Pi) 单一 API + 浏览器 | 来源:@svpino | 链接:https://x.com/svpino | 仓库:无 | 论文:无 | 硬核点:类比 OpenRouter 对模型的作用,AgentSky 试图成为 agents 的聚合分发层,有平台效应
-
主题:Jerry Liu — RAG 十二大痛点及 2026 文档解析难点长帖 | 来源:@jerryjliu0 | 链接:https://x.com/jerryjliu0/status/2058953208782074127 | 仓库:无 | 论文:无 | 硬核点:LlamaIndex CEO 亲述生产 RAG 核心失效模式:rerank/查询改写/文档解析/agentic loop 分工
-
主题:commit-rewriter 0.1 — LLM 批量编辑 Git commit message 的本地 web 工具 | 来源:@simonw | 链接:https://github.com/simonw/commit-rewriter | 仓库:simonw/commit-rewriter | 论文:无 | 硬核点:Claude Code 自动 commit 产物净化的实操工具,踩坑 coding agent 日志噪音的工程解法
其余线索
- @_akhaliq: RynnBrain Open Embodied Foundation Models 论文分享 (arxiv + HF paper page) — 具身 AI foundation model 新工作
- @rasbt: "Build a Reasoning Model From Scratch" 新书上架 Amazon — Qwen3 base 的 RL/蒸馏/推理 scaling 全流程实操书
- @maximelabonne: Top Models For Your Hardware 2026 (8GB 硬件模型选型指南) + LFM-2.6B 推荐 — 本地部署硬参硬据
- @maximelabonne: 24B 模型浏览器内运行 ~50 tokens/s (Transformers.js) — 浏览器端侧推理 SOTA 演示
- @abacaj: K3 模型自我修正能力观察 — 虽为观点但含实测细节,K3 vs Fable 对比数据
- @maximelabonne: Fly fig 象棋压缩 demo — 附 HF demo 页面,新奇demo非主流技术方向