Codex + DeepSeek 的任务感知推理模式路由(spec/react/weak)Task-aware reasoning-mode router for Codex + DeepSeek | Codex+DeepSeek 思维模式路由(spec/react/weak)
仓库/Skill 库
263 个 · LLM 基础设施
通过一次 wrap() 调用为 LLM 流水线构建运行时可靠性守卫,可配置地防御常见生产故障Build runtime reliability guards for LLM pipelines with one wrap() call and configurable protection against common production failures
WordPress 共享 AI 基础设施:provider 凭证、实时模型、提示词、标准化请求及 AI-Scribe 和兼容插件的使用记录。Shared AI infrastructure for WordPress: provider credentials, live models, prompts, normalised requests and usage records for AI-Scribe and compatible plugins.
论文 A Survey of On-Policy Distillation for Large Language Models(arXiv:2604.00626)的配套网站。Companion website for A Survey of On-Policy Distillation for Large Language Models (arXiv:2604.00626).
使用纯 C 从零构建 LLM 推理引擎,不依赖任何框架。Build an LLM inference engine from scratch in pure C with no frameworks.
论文《Controllable molecular graph generation from natural-language chemical constraints》(基于自然语言化学约束的可控分子图生成)的代码仓库。This is repository for "Controllable molecular graph generation from natural-language chemical constraints"
几何有限主义:语言与思维的几何学Geofinitism: The Geometry of Language and Thought
构建 LongCat-Flash-Prover,利用 LongCat 模型进行快速定理证明与形式化推理。Build LongCat-Flash-Prover for fast theorem proving and formal reasoning with LongCat models
使用纯 C99 MoE 推理引擎在 CPU 上原生运行 DeepSeek-V4-Flash-0731,无需 GPU、CUDA 或 PyTorch。Run native DeepSeek-V4-Flash-0731 on CPU with a pure C99 MoE inference engine — no GPU, CUDA, or PyTorch needed.
构建安全、有治理、可观测、成本可控的云与 AI 平台能力的实用参考框架。A practical reference framework for building secure, governed, observable, and cost-aware Cloud & AI platform capabilities.
本代码库提供个人研究成果《Deep Learning Approach in Time Series Forecasting: Literature Review and Extension on Feature Extraction》的代码与报告存档。This repository provides code and report archive of individual research: Deep Learning Approach in Time Series Forecasting: Literature Review and Extension on Feature Extraction
The Convergence Gap 的代码与 artifact:指令微调模型何时收敛到下一 token 预测Code and artifacts for The Convergence Gap: when instruction-tuned models settle on next-token predictions.
精选并经核验的资源合集:面向学术与科学写作的 LLM 研究论文、数据集、工具与实现。A curated, verified collection of research papers, datasets, tools, and implementations on LLMs for academic and scientific writing.
贝叶斯统计研究语料库:面向贝叶斯统计(推断、计算、先验、模型选择、层次模型、非参数、深度贝叶斯、应用)的数据驱动、自动化校验文献综述Bayesian Statistics Research Corpus: Data-driven, auto-validated literature review for Bayesian statistics (inference, computation, priors, model selection, hierarchical, nonparametric, deep Bayesian, applications)
面向几何感知量子机器学习的可复现研究框架。A reproducible research framework for geometry-aware quantum machine learning.
生产级 LLM 网关:兼容 OpenAI 的 API,支持语义缓存,可在 Claude/OpenAI/Ollama 之间进行智能路由与成本分析。Production LLM gateway: OpenAI-compatible API, semantic caching, intelligent routing across Claude/OpenAI/Ollama, cost analytics.
使用 TurboQuant 和 MLX 在 Windows 上运行 Qwen3.5 语言模型,实现快速的本地推理。Run the Qwen3.5 language model on Windows using TurboQuant and MLX for fast local performance.
Experimental Qwen3.5-derived 752M LLM:基于 Qwen3.5 的实验性 752M 参数 LLM,采用 CPT + SFT 两阶段流程,面向编码、技术推理与指令跟随。Experimental Qwen3.5-derived 752M LLM: a two-phase CPT + SFT pipeline for coding, technical reasoning, and instruction following.
通过 MegaQwen CUDA megakernel 加速 Qwen3-0.6B 推理,在 RTX 3090 上达到 531 tok/s decode,较 HuggingFace 提升 3.9×🚀 Achieve faster Qwen3-0.6B inference with the MegaQwen CUDA megakernel, delivering 531 tok/s decode on RTX 3090—3.9x faster than HuggingFace.
大语言模型(LLM)是一种人工智能(AI)程序,能够识别并生成文本,以及执行其他任务。A large language model (LLM) is a type of artificial intelligence (AI) program that can recognize and generate text, among other tasks.
在双 RTX 3090 GPU 上以 262K 上下文长度并发流运行 Qwen3.6-27B,开发工作已迁移至 club-3090。Run Qwen3.6-27B at 262K context with concurrent streams on dual RTX 3090 GPUs. Development moved to club-3090.
本地优先、隐私至上的 AI 驱动学术写作人性化工具。AI-Powered Academic Writing Humanizer - Local-First, Privacy-Centric Solution
完全在浏览器中通过 WebAssembly 实现文档转换、检查、OCR 与翻译。保留版式的翻译,全程本地、离线优先。Convert, inspect, OCR, and translate any document entirely in the browser via WebAssembly. Layout-preserving translation, fully local, offline-first.
使用支持 macOS 和 Windows 平台的开源工具高效构建与定制机器学习模型Build and customize machine learning models efficiently with an open-source tool that supports macOS and Windows platforms.
MAI-Code 模型的官方仓库,用于发布 MAI-Code 模型版本与更新,并通过 issue 与反馈与开发者社区互动。Official repo for MAI-Code models, where we publish MAI-Code model releases and updates and engage with developer community on issues and feedback.
通过单一严格的 YAML 清单统一管理本地 LLM 运行时——支持状态查看、健康检查、启动、停止,基于原生 systemd 与 Docker。Manage local LLM runtimes from one strict YAML manifest — status, doctor, start, stop over native systemd and Docker.
metaScreener——基于插件的桌面应用,用于 human-in-the-loop systematic literature screening。在顺序可审计的 pipeline 中结合确定性启发式过滤与 LLM 推理,通过 SHA-256 校验包实现完全可复现。MIT 协议。metaScreener — a plugin-based desktop application for human-in-the-loop systematic literature screening. Combines deterministic heuristic filters with LLM inference in a sequential, auditable pipeline. SHA-256 verified bundles for full reproducibility. MIT licensed.
面向 LLM 的阶段感知上下文窗口治理框架。提供不变的上下文长度上限、基于熵的稳定性控制,以及针对降级(碎片化)状态的概率性保证,适用于生产级 LLM 系统。Phase-aware context window governance framework for Large Language Models (LLMs). Provides invariant context length caps, entropy-based stability control, and probabilistic guarantees against degraded (fragmentation) states for production LLM systems.
🛠 通过此硬件插件提升 vLLM 在 Kunlun XPU 上的性能,无缝集成主流 AI 模型并优化执行效率🛠 Enhance vLLM performance on Kunlun XPU with this hardware plugin, offering seamless integration for popular AI models and optimized execution.
SILVA Networks 是一个 Python 包,提供扩展的深度均衡层、定点求解器、隐式微分、结构化算子、诊断工具以及可复现研究 notebook。SILVA Networks is a Python package for extended deep equilibrium layers, fixed-point solvers, implicit differentiation, structured operators, diagnostics, and reproducible research notebooks.
面向科学与医学写作的人性化润色的阿英双语 skill,同时保持学术准确性。Bilingual Arabic-English skill for humanizing scientific and medical writing while preserving scholarly accuracy.
系统综述《Techniques for Adapting Large Language Models to Low-Resource, Morphologically Rich Languages》的补充材料。Supplementary materials for the systematic review: Techniques for Adapting Large Language Models to Low-Resource, Morphologically Rich Languages
Codex skill,面向交通运输与低空出行领域的学术写作(知识蒸馏自 LT)。Codex skill for transportation and low-altitude mobility academic writing (Knowledge distillation from LT)
适用于 DeepSeek、Qwen、GLM 及多种 AI 模型的 OpenAI 兼容 API 示例。OpenAI Compatible API examples for DeepSeek, Qwen, GLM and multiple AI models.
🚀 使用强化学习优化半精度通用矩阵乘法(HGEMM)CUDA kernel,性能超越 cuBLAS 及其他基准。🚀 Optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels using reinforcement learning, surpassing cuBLAS and other benchmarks with superior performance.