用于探索 Transformer 架构、张量及实时 LLM 推理的交互式 3D 可视化平台。Interactive 3D visualization platform for exploring transformer architectures, tensors, and real-time LLM inference.
仓库/Skill 库
263 个 · LLM 基础设施
LLM7.io 提供单一 API 网关,可连接来自多家供应商的众多领先 AI 模型LLM7.io offers a single API gateway that connects you to a wide array of leading AI models from various providers.
Codex 的每轮模型与推理路由,由 Jev(TypeSafe System One)驱动:为每一轮选择模型、思考深度和速度模式。Per-turn model & reasoning routing for Codex, driven by Jev (TypeSafe System One): picks the model, thinking depth and speed mode for every turn.
Google Deep Search 的实现,支持 1000+ 篇参考文献、本地推理、使用 RAPTOR 与抓取会话对话以及报告生成。An implementation of Google Deep Search with support for 1000+ references, local inference, chatting with your scraping session using RAPTOR, and report generation.
GLM-5.3 Flash EXL3,针对 2x DGX Sparks。GLM-5.3 Flash EXL3 for 2x DGX Sparks
在两台 dgx-spark 上从零部署 deepseek-flash-0731 的设置指南。setup guide for deepseek-flash-0731 on two dgx-spark from scratch
Qwen3.8 27B 在 SGLang 上运行于 DGX SparkQwen3.8 27B on SGLang for DGX Spark
在单台 DGX Spark 上运行 Qwen3.8 Flash Next(TensorFold)Qwen3.8 Flash Next on one DGX Spark (TensorFold)
2026 年度顶级开源 AI 工程平台 Qyvaria KernelTop Open-Source AI Engineering Platform 2026 Qyvaria Kernel
腾讯 CodeBuddy 账号池管理控制台 + OpenAI 兼容反代网关。扫码纳管账号、自动签到、多密钥分发、IP 管控、调用日志与用量统计。UI 对标 linux-do/cdk。
DeepSeek-V4.1-Flash 部署于 3–4 块 NVIDIA DGX Spark。DeepSeek-V4.1-Flash on 3-4x NVIDIA DGX Sparks
将你的 CodeBuddy 订阅作为本地 OpenAI API 使用。Use your CodeBuddy subscription as local OpenAI APIs.
OmniStudio 是一个本地大模型一体化桌面工作台,集模型市集下载、llama.cpp/vLLM/SGLang 三引擎推理管理,以及对话、语音合成、ASR语音识别、图片生成、视频生成、OCR 等多种大模型应用于一体,全程本地优先。
一款便捷的 lib,用于流畅地与大语言模型(LLM)交互及构建 AI 应用。A handy lib for smooth interaction with large language models (LLMs) and crafting AI apps.
JevK5:TypeSafe Jev 的开源权重替代方案。单次前向即可输出带概率的类型化决策;权重与代码均采用 Apache-2.0 协议。JevK5: open-weight alternative to TypeSafe Jev. Typed decisions with probabilities in one forward pass; Apache-2.0 weights and code.
本地 Codex 反向代理:轮换代理节点池、惰性健康故障转移、每模型 292 turn-state 采集/注入。支持 macOS + Windows。Local Codex reverse proxy: rotating proxy-node pool, lazy health failover, and per-model 292 turn-state collection/injection. macOS + Windows.
Project Titania 是一个完整的大语言模型系统,从 Transformer 到晶体管,简洁到足以让一个人理解。Project Titania is a complete large language model system, from transformer to transistor, simple enough for one person to understand.
FreeBuff 编码模型的 OpenAI 兼容网关。Token 池、会话生命周期、TLS stealth、嵌入式管理后台。无广告、无 CLI,只有 /v1/chat/completions。OpenAI-compatible gateway for FreeBuff coding models. Token pool, session lifecycle, TLS stealth, embedded admin dashboard. No ads, no CLI, just /v1/chat/completions.
一个小型开放决策模型:state + 类型化问题 → 校准概率。在 Qwen3.5 上的 Jev / System One 复现。A small open decision model: state + typed questions -> calibrated probabilities. A Jev / System One re-creation on Qwen3.5.
本地 AI 注册中心——硬件、模型、recipes、模型实例与价格。Local AI registry — hardware, models, recipes, model instances, and prices
🪢 Langfuse API 的自动生成 Java 客户端。🪢 Auto-generated Java Client for Langfuse API
提供 DeepSeek 与 Qwen 服务,运行按需模型库,并在 2 块 NVIDIA DGX Spark 上使用 Unsloth QLoRA 微调。TP=2 的 vLLM 通道汇聚于单一 OpenAI 兼容端点,供 OpenCode、Cursor 与 Hermes 使用。诚实的基准,已提交的工件。Serve DeepSeek and Qwen, run an on-demand model library, and fine-tune with Unsloth QLoRA on 2x NVIDIA DGX Spark. TP=2 vLLM lanes behind one OpenAI-compatible endpoint for OpenCode, Cursor, and Hermes. Honest benchmarks, committed artifacts.
2× NVIDIA DGX Spark 上的 GLM-5.3-Flash (NVFP4)——vLLM TP2,262K 上下文,MTP。全球首发部署方案:发现并修复 7 个 day-0 bug,附 sm121 镜像补丁、探针与完整报告。GLM-5.3-Flash (NVFP4) on 2x NVIDIA DGX Spark - vLLM TP2, 262K context, MTP. World-first deploy recipe: 7 day-0 bugs found and fixed, patched sm121 image, probes and full report.
ELM 是将 LLM 应用于能源研究的一系列工具集合。ELM is a collection of utilities to apply Large Language Models (LLMs) to energy research.
基于开源模型的 Jev 兼容 API 端点(仅 prefill)Jev-compatible API endpoint based on open models (prefill-only)
纯 Go 编写的轻量神经网络框架,使用 AVX2 SIMD 内核(GOEXPERIMENT=simd)。A tiny neural-network framework in pure Go with AVX2 SIMD kernels (GOEXPERIMENT=simd)
AI Max+ 395 加速:在 RTX 3060 上实测的异构 GPU PD 与异步融合层 pipeline 实验。AI Max+ 395 acceleration: measured heterogeneous GPU PD and asynchronous fused-layer pipeline experiments with RTX 3060.
类似 unix du 的命令行工具,用于统计每个文件和目录的 token 使用量a unix-like du command line tool to count token usage per files and directories
Meta Inc. 所有 oss 模型的相关 recipes。All recipes for oss models from Meta Inc.
repo-map 生成由 LLM 增强的软件仓库摘要与分析,为开发者提供关于项目结构、文件用途以及跨编程语言潜在考量的洞察。repo-map generates LLM-enhanced summaries and analysis of software repositories, providing developers with valuable insights into project structures, file purposes, and potential considerations across various programming languages.
面向 Qwen3.8(含支持专家卸载的 Qwen3.8-Flash-Next)的 C++/CUDA 单 GPU 推理引擎,运行于 RTX 5090;源自 NInfer。C++/CUDA single-GPU inference engine for Qwen3.8 (incl. Qwen3.8-Flash-Next with offloaded experts) on the RTX 5090; grew from NInfer