Google Deep Search 的实现,支持 1000+ 篇参考文献、本地推理、使用 RAPTOR 与抓取会话对话以及报告生成。An implementation of Google Deep Search with support for 1000+ references, local inference, chatting with your scraping session using RAPTOR, and report generation.
仓库/Skill 库
62 个 · LLM 基础设施 · 应用
Qwen3.8 27B 在 SGLang 上运行于 DGX SparkQwen3.8 27B on SGLang for DGX Spark
OmniStudio 是一个本地大模型一体化桌面工作台,集模型市集下载、llama.cpp/vLLM/SGLang 三引擎推理管理,以及对话、语音合成、ASR语音识别、图片生成、视频生成、OCR 等多种大模型应用于一体,全程本地优先。
2× NVIDIA DGX Spark 上的 GLM-5.3-Flash (NVFP4)——vLLM TP2,262K 上下文,MTP。全球首发部署方案:发现并修复 7 个 day-0 bug,附 sm121 镜像补丁、探针与完整报告。GLM-5.3-Flash (NVFP4) on 2x NVIDIA DGX Spark - vLLM TP2, 262K context, MTP. World-first deploy recipe: 7 day-0 bugs found and fixed, patched sm121 image, probes and full report.
基于开源模型的 Jev 兼容 API 端点(仅 prefill)Jev-compatible API endpoint based on open models (prefill-only)
面向 Qwen3.8(含支持专家卸载的 Qwen3.8-Flash-Next)的 C++/CUDA 单 GPU 推理引擎,运行于 RTX 5090;源自 NInfer。C++/CUDA single-GPU inference engine for Qwen3.8 (incl. Qwen3.8-Flash-Next with offloaded experts) on the RTX 5090; grew from NInfer
面向 Claude Code 的学术研究工作流:包含 20 个 skill,覆盖因果推断(DiD、RDD、IV、合成控制、田野实验、设计分流、预注册)、论文阅读、文献综述、参考文献审计、复现包、LaTeX 与 TikZ,并附带带渲染时质量门禁的 Quarto reveal.js 幻灯片系统。Academic research workflow for Claude Code: 20 skills covering causal inference (DiD, RDD, IV, synthetic control, field experiments, design triage, preregistration), paper reading, lit review, bib auditing, replication packages, LaTeX and TikZ, plus a Quarto reveal.js slide system with render-time quality gates.
为赫尔辛基大学师生打造的 LLM 聊天工具,用于教育与研究。LLM chat built for University of Helsinki staff and students, for education and research.
将 Zotero 阅读器标注评论渲染为 Markdown 和 LaTeX 格式,同时保留存储的原始文本。Render Zotero reader annotation comments as Markdown and LaTex while preserving stored text.
本地 AI 秘书、技术支持与销售一体化方案,基于 XTTS v2 语音克隆、Vosk/Whisper 实时语音识别与 vLLM + Qwen/Llama 等离线 LLM。配备 Vue 3 完整管理面板、Telegram Bot、网站挂件及 fine-tuning pipeline。支持自托管、数据隐私、短信与电话呼叫。📞 Локальный AI-секретарь, тех. поддержка и менеджер по продажам с клонированием голоса XTTS v2, real-time распознаванием речи (Vosk/Whisper) и offline LLM (vLLM + Qwen/Llama и тп). Полноценная админ-панель (Vue 3), Telegram-бот, виджет для сайта, fine-tuning pipeline. Self-hosted, приватность данных, СМС и телефонные звонки .
STEPQuant: Delta 规则循环状态量化中错误在何时何处重要STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
以孟加拉语优先、面向方言的 LLM 研究。开源孟加拉语 tokenizer,性能优于 Sarvam、AI4Bharat 和 GPT-4o(fertility 1.52,几乎零破损 conjuncts)。非商用,保护孟加拉语及其方言。Bengali-first, dialect-aware LLM research. Open-source Bengali tokenizer that outperforms Sarvam, AI4Bharat, and GPT-4o (fertility 1.52, near-zero broken conjuncts). Non-commercial, preserving Bengali and its dialects.
THIS PAPER IS FOR ENTERTAINMENT PURPOSES ONLY.この小説はフィクションであり、実在する団体または個人とは関係ありません。
基于 TurboQuant 优化 FAISS 兼容的向量量化,实现快速、精准的向量检索。Optimize FAISS-compatible vector quantization for fast, accurate vector search with TurboQuant
《大模型推理原理与优化》:面向系统/架构/后端研发工程师的模型原理入门课,目标是通俗易懂的解释推理过程,理解原理有助于系统开发/维护工作
论文 A Survey of On-Policy Distillation for Large Language Models(arXiv:2604.00626)的配套网站。Companion website for A Survey of On-Policy Distillation for Large Language Models (arXiv:2604.00626).
使用纯 C99 MoE 推理引擎在 CPU 上原生运行 DeepSeek-V4-Flash-0731,无需 GPU、CUDA 或 PyTorch。Run native DeepSeek-V4-Flash-0731 on CPU with a pure C99 MoE inference engine — no GPU, CUDA, or PyTorch needed.
通过 MegaQwen CUDA megakernel 加速 Qwen3-0.6B 推理,在 RTX 3090 上达到 531 tok/s decode,较 HuggingFace 提升 3.9×🚀 Achieve faster Qwen3-0.6B inference with the MegaQwen CUDA megakernel, delivering 531 tok/s decode on RTX 3090—3.9x faster than HuggingFace.
完全在浏览器中通过 WebAssembly 实现文档转换、检查、OCR 与翻译。保留版式的翻译,全程本地、离线优先。Convert, inspect, OCR, and translate any document entirely in the browser via WebAssembly. Layout-preserving translation, fully local, offline-first.
metaScreener——基于插件的桌面应用,用于 human-in-the-loop systematic literature screening。在顺序可审计的 pipeline 中结合确定性启发式过滤与 LLM 推理,通过 SHA-256 校验包实现完全可复现。MIT 协议。metaScreener — a plugin-based desktop application for human-in-the-loop systematic literature screening. Combines deterministic heuristic filters with LLM inference in a sequential, auditable pipeline. SHA-256 verified bundles for full reproducibility. MIT licensed.
🛠 通过此硬件插件提升 vLLM 在 Kunlun XPU 上的性能,无缝集成主流 AI 模型并优化执行效率🛠 Enhance vLLM performance on Kunlun XPU with this hardware plugin, offering seamless integration for popular AI models and optimized execution.
面向科学与医学写作的人性化润色的阿英双语 skill,同时保持学术准确性。Bilingual Arabic-English skill for humanizing scientific and medical writing while preserving scholarly accuracy.
语言 case 的范畴论处理,集成 Active Inference 与 CEREBRUM 架构:DisCoPy 弦图横跨类型学、范畴语法、拓扑斯理论及量子扩展,生成 30 个发表级图表与一篇 24 节的手稿。1,207 个测试,覆盖率 95.96%,零 mockCategory-theoretic treatment of linguistic case integrated with Active Inference and the CEREBRUM architecture: DisCoPy string diagrams spanning typology, categorial grammar, topos theory, and quantum extensions, generating 30 publication figures and a 24-section manuscript. 1,207 tests, 95.96% coverage, zero mocks.
基于仿真的 MRI 调度策略分析,结合统计分析与离散事件仿真,以优化资源利用率、等待时间、加班时长与患者吞吐量。Simulation-based analysis of MRI scheduling policies, combining statistical analysis and discrete-event simulation to optimize resource utilization, waiting times, overtime, and patient throughput.
基于 Torch2PC 的预测编码硕士论文可复现项目:在 Ubuntu/ROCm 上对 backpropagation 进行逐层与 compute-matched 对比,实现 PC-CATM/PC-TREF,对 state inference 进行机制诊断,并构建 QWake-PC 以实现自适应 exact inference —— 面向机制可解释的可复现预测编码研究。Воспроизводимый проект магистерской диссертации по predictive coding в Torch2PC: послойное и compute-matched сравнение с backpropagation, PC-CATM/PC-TREF, механизмная диагностика state inference и QWake-PC для адаптивного exact inference в Ubuntu/ROCm — reproducible research on mechanism-aware predictive coding.
SLM-PICO-Screener 是用于自动化系统综述筛选的轻量 NLP 流水线,基于 Phi-3 Mini 配合 LoRA 与 4-bit 量化,将文章分类为 8 个 PICO 标签。体积约 90 MB,可在消费级 GPU/CPU 上运行,加速生物医学研究中的证据综合。SLM-PICO-Screener is a lightweight NLP pipeline automating systematic review screening. Built on Phi-3 Mini with LoRA & 4-bit quantization, it classifies articles into 8 PICO labels. At ~90 MB, it runs on consumer GPUs/CPU, accelerating evidence synthesis in biomedical research.