Qwen3.8-Flash-Next 部署于单台 DGX Spark(TP=1)Qwen3.8-Flash-Next on ONE DGX Spark (TP=1)
仓库/Skill 库
68 个 · LLM 基础设施 · 快速增长
基于 2x DGX Spark 与 TensorFold 的 GLM-5.3-Flash EXL3。GLM-5.3-Flash EXL3 on 2x DGX Spark with TensorFold
在 Mac 上以远低于 104 GB 的内存运行 Qwen3.8-Flash-Next(125B MoE,4-bit 下约 104 GB),通过 SSD 流式加载专家。MLX + Swift,兼容 Ollama 的 API。Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API.
DeepSeek v4 Flash EXL3 运行于单台 DGX SparkDeepSeek v4 Flash EXL3 on one DGX Spark
Codex 的每轮模型与推理路由,由 Jev(TypeSafe System One)驱动:为每一轮选择模型、思考深度和速度模式。Per-turn model & reasoning routing for Codex, driven by Jev (TypeSafe System One): picks the model, thinking depth and speed mode for every turn.
GLM-5.3 Flash EXL3,针对 2x DGX Sparks。GLM-5.3 Flash EXL3 for 2x DGX Sparks
在两台 dgx-spark 上从零部署 deepseek-flash-0731 的设置指南。setup guide for deepseek-flash-0731 on two dgx-spark from scratch
Qwen3.8 27B 在 SGLang 上运行于 DGX SparkQwen3.8 27B on SGLang for DGX Spark
在单台 DGX Spark 上运行 Qwen3.8 Flash Next(TensorFold)Qwen3.8 Flash Next on one DGX Spark (TensorFold)
腾讯 CodeBuddy 账号池管理控制台 + OpenAI 兼容反代网关。扫码纳管账号、自动签到、多密钥分发、IP 管控、调用日志与用量统计。UI 对标 linux-do/cdk。
DeepSeek-V4.1-Flash 部署于 3–4 块 NVIDIA DGX Spark。DeepSeek-V4.1-Flash on 3-4x NVIDIA DGX Sparks
将你的 CodeBuddy 订阅作为本地 OpenAI API 使用。Use your CodeBuddy subscription as local OpenAI APIs.
OmniStudio 是一个本地大模型一体化桌面工作台,集模型市集下载、llama.cpp/vLLM/SGLang 三引擎推理管理,以及对话、语音合成、ASR语音识别、图片生成、视频生成、OCR 等多种大模型应用于一体,全程本地优先。
JevK5:TypeSafe Jev 的开源权重替代方案。单次前向即可输出带概率的类型化决策;权重与代码均采用 Apache-2.0 协议。JevK5: open-weight alternative to TypeSafe Jev. Typed decisions with probabilities in one forward pass; Apache-2.0 weights and code.
本地 Codex 反向代理:轮换代理节点池、惰性健康故障转移、每模型 292 turn-state 采集/注入。支持 macOS + Windows。Local Codex reverse proxy: rotating proxy-node pool, lazy health failover, and per-model 292 turn-state collection/injection. macOS + Windows.
Project Titania 是一个完整的大语言模型系统,从 Transformer 到晶体管,简洁到足以让一个人理解。Project Titania is a complete large language model system, from transformer to transistor, simple enough for one person to understand.
FreeBuff 编码模型的 OpenAI 兼容网关。Token 池、会话生命周期、TLS stealth、嵌入式管理后台。无广告、无 CLI,只有 /v1/chat/completions。OpenAI-compatible gateway for FreeBuff coding models. Token pool, session lifecycle, TLS stealth, embedded admin dashboard. No ads, no CLI, just /v1/chat/completions.
一个小型开放决策模型:state + 类型化问题 → 校准概率。在 Qwen3.5 上的 Jev / System One 复现。A small open decision model: state + typed questions -> calibrated probabilities. A Jev / System One re-creation on Qwen3.5.
本地 AI 注册中心——硬件、模型、recipes、模型实例与价格。Local AI registry — hardware, models, recipes, model instances, and prices
提供 DeepSeek 与 Qwen 服务,运行按需模型库,并在 2 块 NVIDIA DGX Spark 上使用 Unsloth QLoRA 微调。TP=2 的 vLLM 通道汇聚于单一 OpenAI 兼容端点,供 OpenCode、Cursor 与 Hermes 使用。诚实的基准,已提交的工件。Serve DeepSeek and Qwen, run an on-demand model library, and fine-tune with Unsloth QLoRA on 2x NVIDIA DGX Spark. TP=2 vLLM lanes behind one OpenAI-compatible endpoint for OpenCode, Cursor, and Hermes. Honest benchmarks, committed artifacts.
2× NVIDIA DGX Spark 上的 GLM-5.3-Flash (NVFP4)——vLLM TP2,262K 上下文,MTP。全球首发部署方案:发现并修复 7 个 day-0 bug,附 sm121 镜像补丁、探针与完整报告。GLM-5.3-Flash (NVFP4) on 2x NVIDIA DGX Spark - vLLM TP2, 262K context, MTP. World-first deploy recipe: 7 day-0 bugs found and fixed, patched sm121 image, probes and full report.
基于开源模型的 Jev 兼容 API 端点(仅 prefill)Jev-compatible API endpoint based on open models (prefill-only)
纯 Go 编写的轻量神经网络框架,使用 AVX2 SIMD 内核(GOEXPERIMENT=simd)。A tiny neural-network framework in pure Go with AVX2 SIMD kernels (GOEXPERIMENT=simd)
AI Max+ 395 加速:在 RTX 3060 上实测的异构 GPU PD 与异步融合层 pipeline 实验。AI Max+ 395 acceleration: measured heterogeneous GPU PD and asynchronous fused-layer pipeline experiments with RTX 3060.
Meta Inc. 所有 oss 模型的相关 recipes。All recipes for oss models from Meta Inc.
面向 Qwen3.8(含支持专家卸载的 Qwen3.8-Flash-Next)的 C++/CUDA 单 GPU 推理引擎,运行于 RTX 5090;源自 NInfer。C++/CUDA single-GPU inference engine for Qwen3.8 (incl. Qwen3.8-Flash-Next with offloaded experts) on the RTX 5090; grew from NInfer
基于 Jev API 的开放决策模型:通过前向传递给出带概率的类型化答案(yes/no、choice、score、multi),无需生成。基于 Qwen3.5-4B + LoRA,单卡 GPU。Open decisions model with Jev's API: typed answers (yes/no, choice, score, multi) with probabilities from forward passes, no generation. Qwen3.5-4B + LoRA, one GPU.