研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

263 个 · LLM 基础设施

排序 Stars 周增
Sudharsanselvaraj/Token-Print
TypeScript · 2026-09-24 LLM 基础设施 框架 实验 Stars 198 周增 +46

用于探索 Transformer 架构、张量及实时 LLM 推理的交互式 3D 可视化平台。Interactive 3D visualization platform for exploring transformer architectures, tensors, and real-time LLM inference.

llm-infra
chigwell/llm7.io
TypeScript · 2026-08-11 LLM 基础设施 工具 实验 Stars 193 周增 +0

LLM7.io 提供单一 API 网关,可连接来自多家供应商的众多领先 AI 模型LLM7.io offers a single API gateway that connects you to a wide array of leading AI models from various providers.

llm-infra
0xNatoshi/jev-codex-router
JavaScript · 2026-09-22 LLM 基础设施 工具 实验 Stars 184 周增 +0

Codex 的每轮模型与推理路由,由 Jev(TypeSafe System One)驱动:为每一轮选择模型、思考深度和速度模式。Per-turn model & reasoning routing for Codex, driven by Jev (TypeSafe System One): picks the model, thinking depth and speed mode for every turn.

llm-infra
iblameandrew/deepsearch-academic
Python · 2026-08-14 LLM 基础设施 应用 实验 Stars 181 周增 +0

Google Deep Search 的实现,支持 1000+ 篇参考文献、本地推理、使用 RAPTOR 与抓取会话对话以及报告生成。An implementation of Google Deep Search with support for 1000+ references, local inference, chatting with your scraping session using RAPTOR, and report generation.

llm-infra
MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks
Python · 2026-08-29 LLM 基础设施 模型 实验 Stars 175 周增 +686

GLM-5.3 Flash EXL3,针对 2x DGX Sparks。GLM-5.3 Flash EXL3 for 2x DGX Sparks

maliubiao/dgx-spark-2-deepseek-flash-0731
Shell · 2026-08-11 LLM 基础设施 教程 实验 Stars 169 周增 +0

在两台 dgx-spark 上从零部署 deepseek-flash-0731 的设置指南。setup guide for deepseek-flash-0731 on two dgx-spark from scratch

MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark
Python · 2026-08-19 LLM 基础设施 应用 实验 Stars 168 周增 +0

Qwen3.8 27B 在 SGLang 上运行于 DGX SparkQwen3.8 27B on SGLang for DGX Spark

llm-infra
r-uby-dev/llm
Ruby · 2026-09-14 LLM 基础设施 库 实验 Stars 141 周增 +2

Ruby 的强大 AI 运行时Ruby's capable AI runtime

agentragllm-infra
MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold
Shell · 2026-10-01 LLM 基础设施 工具 研究原型 Stars 129 周增 +0

在单台 DGX Spark 上运行 Qwen3.8 Flash Next(TensorFold)Qwen3.8 Flash Next on one DGX Spark (TensorFold)

bakiraa/qyvaria-hardlogic-kernel-engine
HTML · 2026-09-17 LLM 基础设施 框架 实验 Stars 125 周增 +0

2026 年度顶级开源 AI 工程平台 Qyvaria KernelTop Open-Source AI Engineering Platform 2026 Qyvaria Kernel

agent
ithtelab/workbuddy-manager
Python · 2026-09-16 LLM 基础设施 工具 生产可用 Stars 123 周增 +308

腾讯 CodeBuddy 账号池管理控制台 + OpenAI 兼容反代网关。扫码纳管账号、自动签到、多密钥分发、IP 管控、调用日志与用量统计。UI 对标 linux-do/cdk。

MiaAI-Lab/DeepSeek-v4.1-Flash-DGX-Sparks
Python · 2026-09-15 LLM 基础设施 工具 实验 Stars 118 周增 +122

DeepSeek-V4.1-Flash 部署于 3–4 块 NVIDIA DGX Spark。DeepSeek-V4.1-Flash on 3-4x NVIDIA DGX Sparks

maiphucgiang/codebuddy2api
Python · 2026-09-16 LLM 基础设施 工具 实验 Stars 114 周增 +0

将你的 CodeBuddy 订阅作为本地 OpenAI API 使用。Use your CodeBuddy subscription as local OpenAI APIs.

arikchakma/gpu-time
Python · 2026-09-12 LLM 基础设施 模型 实验 Stars 113 周增 +0

用于解析日期和时间的小型神经模型A small neural model for parsing date & time

kunpengtalk/OmniStudio
TypeScript · 2026-09-16 LLM 基础设施 应用 实验 Stars 108 周增 +114

OmniStudio 是一个本地大模型一体化桌面工作台,集模型市集下载、llama.cpp/vLLM/SGLang 三引擎推理管理,以及对话、语音合成、ASR语音识别、图片生成、视频生成、OCR 等多种大模型应用于一体,全程本地优先。

multimodalllm-infra
Nayjest/ai-microcore
Python · 2026-08-14 LLM 基础设施 库 生产可用 Stars 107 周增 +0

一款便捷的 lib,用于流畅地与大语言模型(LLM)交互及构建 AI 应用。A handy lib for smooth interaction with large language models (LLMs) and crafting AI apps.

ragdatabasellm-infra
allebee/jevk5
Python · 2026-09-25 LLM 基础设施 模型 实验 Stars 107 周增 +0

JevK5:TypeSafe Jev 的开源权重替代方案。单次前向即可输出带概率的类型化决策;权重与代码均采用 Apache-2.0 协议。JevK5: open-weight alternative to TypeSafe Jev. Typed decisions with probabilities in one forward pass; Apache-2.0 weights and code.

446599/ccodex-rotate
Go · 2026-09-22 LLM 基础设施 工具 实验 Stars 93 周增 +0

本地 Codex 反向代理:轮换代理节点池、惰性健康故障转移、每模型 292 turn-state 采集/注入。支持 macOS + Windows。Local Codex reverse proxy: rotating proxy-node pool, lazy health failover, and per-model 292 turn-state collection/injection. macOS + Windows.

penberg/titania
Rust · 2026-09-14 LLM 基础设施 框架 研究原型 Stars 87 周增 +0

Project Titania 是一个完整的大语言模型系统,从 Transformer 到晶体管,简洁到足以让一个人理解。Project Titania is a complete large language model system, from transformer to transistor, simple enough for one person to understand.

llm-infra
trefeon/freebuff-proxy
Go · 2026-08-16 LLM 基础设施 工具 实验 Stars 85 周增 +112

FreeBuff 编码模型的 OpenAI 兼容网关。Token 池、会话生命周期、TLS stealth、嵌入式管理后台。无广告、无 CLI,只有 /v1/chat/completions。OpenAI-compatible gateway for FreeBuff coding models. Token pool, session lifecycle, TLS stealth, embedded admin dashboard. No ads, no CLI, just /v1/chat/completions.

agentllm-infra
kshetrajna12/reflex
Python · 2026-09-20 LLM 基础设施 模型 实验 Stars 85 周增 +77

一个小型开放决策模型:state + 类型化问题 → 校准概率。在 Qwen3.5 上的 Jev / System One 复现。A small open decision model: state + typed questions -> calibrated probabilities. A Jev / System One re-creation on Qwen3.5.

maximpri/mlxtop
Rust · 2026-09-16 LLM 基础设施 工具 实验 Stars 82 周增 +7
0xSero/local-ai-registry
Python · 2026-09-01 LLM 基础设施 工具 实验 Stars 81 周增 +0

本地 AI 注册中心——硬件、模型、recipes、模型实例与价格。Local AI registry — hardware, models, recipes, model instances, and prices

seoan1024/Korean-llm-v3
Python · 2026-08-29 LLM 基础设施 模型 研究原型 Stars 77 周增 +0

基于 PyTorch 的韩语 LLM 实现与训练项目PyTorch 기반 한국어 LLM 구현 및 학습 프로젝트

llm-infra
tomnio/rubric
TypeScript · 2026-09-15 LLM 基础设施 库 实验 Stars 75 周增 +21

以 Schema 优先的 LLM 结构化抽取。Schema-first structured extraction from LLMs.

llm-infra
langfuse/langfuse-java
Java · 2026-08-14 LLM 基础设施 库 生产可用 Stars 72 周增 +0

🪢 Langfuse API 的自动生成 Java 客户端。🪢 Auto-generated Java Client for Langfuse API

llm-infra
zorost/sparkduet
Python · 2026-08-31 LLM 基础设施 评测集 实验 Stars 71 周增 +0

提供 DeepSeek 与 Qwen 服务,运行按需模型库,并在 2 块 NVIDIA DGX Spark 上使用 Unsloth QLoRA 微调。TP=2 的 vLLM 通道汇聚于单一 OpenAI 兼容端点,供 OpenCode、Cursor 与 Hermes 使用。诚实的基准,已提交的工件。Serve DeepSeek and Qwen, run an on-demand model library, and fine-tune with Unsloth QLoRA on 2x NVIDIA DGX Spark. TP=2 vLLM lanes behind one OpenAI-compatible endpoint for OpenCode, Cursor, and Hermes. Honest benchmarks, committed artifacts.

agentllm-infraevaluation
tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark
Python · 2026-08-28 LLM 基础设施 应用 实验 Stars 69 周增 +0

2× NVIDIA DGX Spark 上的 GLM-5.3-Flash (NVFP4)——vLLM TP2,262K 上下文,MTP。全球首发部署方案:发现并修复 7 个 day-0 bug,附 sm121 镜像补丁、探针与完整报告。GLM-5.3-Flash (NVFP4) on 2x NVIDIA DGX Spark - vLLM TP2, 262K context, MTP. World-first deploy recipe: 7 day-0 bugs found and fixed, patched sm121 image, probes and full report.

multimodalllm-infra
NatLabRockies/elm
Python · 2026-08-23 LLM 基础设施 库 实验 Stars 69 周增 +0

ELM 是将 LLM 应用于能源研究的一系列工具集合。ELM is a collection of utilities to apply Large Language Models (LLMs) to energy research.

llm-infra
ekzhang/openjev-sglang
Python · 2026-09-18 LLM 基础设施 应用 实验 Stars 69 周增 +0

基于开源模型的 Jev 兼容 API 端点(仅 prefill)Jev-compatible API endpoint based on open models (prefill-only)

llm-infra
mattn/tensai
Go · 2026-08-26 LLM 基础设施 框架 实验 Stars 67 周增 +70

纯 Go 编写的轻量神经网络框架,使用 AVX2 SIMD 内核(GOEXPERIMENT=simd)。A tiny neural-network framework in pure Go with AVX2 SIMD kernels (GOEXPERIMENT=simd)

Soulmate-Halo/heterogeneous-gpu-pd-lab
Python · 2026-09-07 LLM 基础设施 工具 实验 Stars 66 周增 +0

AI Max+ 395 加速:在 RTX 3060 上实测的异构 GPU PD 与异步融合层 pipeline 实验。AI Max+ 395 acceleration: measured heterogeneous GPU PD and asynchronous fused-layer pipeline experiments with RTX 3060.

engineering
lirantal/tokenu
TypeScript · 2026-09-01 LLM 基础设施 工具 生产可用 Stars 66 周增 +0

类似 unix du 的命令行工具,用于统计每个文件和目录的 token 使用量a unix-like du command line tool to count token usage per files and directories

agentllm-infra
meta-models/meta-oss-cookbook
Python · 2026-08-12 LLM 基础设施 教程 生产可用 Stars 63 周增 +0

Meta Inc. 所有 oss 模型的相关 recipes。All recipes for oss models from Meta Inc.

cyanheads/repo-map
Python · 2026-08-14 LLM 基础设施 工具 实验 Stars 60 周增 +0

repo-map 生成由 LLM 增强的软件仓库摘要与分析,为开发者提供关于项目结构、文件用途以及跨编程语言潜在考量的洞察。repo-map generates LLM-enhanced summaries and analysis of software repositories, providing developers with valuable insights into project structures, file purposes, and potential considerations across various programming languages.

llm-infra
Wallawalla47/Infernix
C++ · 2026-10-09 LLM 基础设施 应用 实验 Stars 54 周增 +0

面向 Qwen3.8(含支持专家卸载的 Qwen3.8-Flash-Next)的 C++/CUDA 单 GPU 推理引擎,运行于 RTX 5090;源自 NInfer。C++/CUDA single-GPU inference engine for Qwen3.8 (incl. Qwen3.8-Flash-Next with offloaded experts) on the RTX 5090; grew from NInfer

llm-infra