非自回归的 System 1 决策引擎。单次前向传播即可对任意文本进行类型化选择、打分和 yes/no 决策,支持 100+ 语言,并通过路由按请求选择合适的 checkpoint。Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.
仓库/Skill 库
35 个 · LLM 基础设施 · 模型
FinGPT:开源金融大语言模型!革命性 🔥 已在 HuggingFace 发布训练好的模型。FinGPT: Open-Source Financial Large Language Models! Revolutionize 🔥 We release the trained model on HuggingFace.
中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
复旦大学开源的工具增强对话语言模型An open-source tool-augmented conversational language model from Fudan University
面向本地部署的高速 LLM 服务High-speed Large Language Model Serving for Local Deployment
提供 AI 应用和模型服务的最简方式 —— 构建模型推理 API、任务队列、LLM 应用、多模型 pipeline 等。The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
基于 Qwen3.5 构建的小型 Jev 风格决策模型家族,可自行训练与部署。tiny Jev-like family of decision models built on top of Qwen3.5 you can train and run on your own
面向全模态模型的高效推理框架。A framework for efficient model inference with omni-modality models
用于高性能 AI 模型服务(vLLM、SGLang)和按需 SSH 访问 GPU 实例的 GPU 集群管理器。A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
高质量且极速的预训练深度学习模型与 demo。Pre-trained Deep Learning models and demos (high quality and extremely fast)
Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。
SOTA 低比特 LLM 量化(INT8/FP8/MXFP8/INT4/MXFP4/NVFP4)与稀疏化方案;面向 PyTorch、TensorFlow 与 ONNX Runtime 的领先模型压缩技术SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
一款围绕该模型构建的、面向 Apple silicon 的本地推理引擎。A local inference engine for Apple silicon, built around the model.
基于 2x DGX Spark 与 TensorFold 的 GLM-5.3-Flash EXL3。GLM-5.3-Flash EXL3 on 2x DGX Spark with TensorFold
DeepSeek v4 Flash EXL3 运行于单台 DGX SparkDeepSeek v4 Flash EXL3 on one DGX Spark
GLM-5.3 Flash EXL3,针对 2x DGX Sparks。GLM-5.3 Flash EXL3 for 2x DGX Sparks
JevK5:TypeSafe Jev 的开源权重替代方案。单次前向即可输出带概率的类型化决策;权重与代码均采用 Apache-2.0 协议。JevK5: open-weight alternative to TypeSafe Jev. Typed decisions with probabilities in one forward pass; Apache-2.0 weights and code.
一个小型开放决策模型:state + 类型化问题 → 校准概率。在 Qwen3.5 上的 Jev / System One 复现。A small open decision model: state + typed questions -> calibrated probabilities. A Jev / System One re-creation on Qwen3.5.
基于 Jev API 的开放决策模型:通过前向传递给出带概率的类型化答案(yes/no、choice、score、multi),无需生成。基于 Qwen3.5-4B + LoRA,单卡 GPU。Open decisions model with Jev's API: typed answers (yes/no, choice, score, multi) with probabilities from forward passes, no generation. Qwen3.5-4B + LoRA, one GPU.
运行 Ollama 本地 LLM 服务的 Docker 镜像。默认安全,所有 API 请求需 Bearer token(首次启动时自动生成)。OpenAI 兼容 API。支持首次启动模型预拉取、NVIDIA GPU (CUDA) 加速和持久化模型存储。多架构:amd64、arm64。Docker image to run an Ollama local LLM server. Secure by default, all API requests require a Bearer token (auto-generated on first start). OpenAI-compatible API. Supports first-start model pre-pull, NVIDIA GPU (CUDA) acceleration, and persistent model storage. Multi-arch: amd64, arm64.
Tox21 多终点毒性预测:可复现的研究与推理制品(冻结模型、FastAPI 服务、已审计的安全机制)Tox21 multi-endpoint toxicity prediction: a reproducible research and inference artifact (frozen model, FastAPI service, audited security)
面向生产环境的、有内存约束的跨模型 KV-cache 传输,具备受保护的回退机制与可复现研究工具链。Production-oriented, memory-bounded cross-model KV-cache transfer with guarded fallback and reproducible research tooling.
混合神经符号 AI:Llama 3.2 1B + 精确数学 + 经验证事实,807 MB 即时 CPU 原生回答,无需 GPU。封装 .aef 分发。Hybrid neuro-symbolic AI: Llama 3.2 1B + exact math + verified facts — instant CPU-native answers in 807 MB, no GPU. Sealed .aef distribution
论文《Controllable molecular graph generation from natural-language chemical constraints》(基于自然语言化学约束的可控分子图生成)的代码仓库。This is repository for "Controllable molecular graph generation from natural-language chemical constraints"
Experimental Qwen3.5-derived 752M LLM:基于 Qwen3.5 的实验性 752M 参数 LLM,采用 CPT + SFT 两阶段流程,面向编码、技术推理与指令跟随。Experimental Qwen3.5-derived 752M LLM: a two-phase CPT + SFT pipeline for coding, technical reasoning, and instruction following.
MAI-Code 模型的官方仓库,用于发布 MAI-Code 模型版本与更新,并通过 issue 与反馈与开发者社区互动。Official repo for MAI-Code models, where we publish MAI-Code model releases and updates and engage with developer community on issues and feedback.
🌐 通过 Geo-Llama 利用几何深度学习增强语言理解,结合 conformal manifolds 和递归等距变换提升 AI 模型性能。🌐 Enhance language understanding through geometric deep learning with Geo-Llama, leveraging conformal manifolds and recursive isometries for improved AI models.
RIFT — Race-state Inference From Telemetry。一项可复现研究实现,仅基于位置遥测重建具备事件感知的关联性比赛状态,涵盖路线进度、共享事件标识、检查点通过、物理顺序、前车关系、间距与间隔。RIFT — Race-state Inference From Telemetry. A reproducible research implementation for reconstructing occurrence-aware relational race state from positional telemetry alone, including route progress, shared occurrence identity, checkpoint crossings, physical order, car-ahead relations, gaps and intervals.
目标:将 State Space Models(SSM)/ Mamba 用于医学时间序列分析;总体研究问题:Mamba 模型能否提升基于步态信号检测帕金森病的分类模型性能: To use State Space Models (SSM)/ Mamba for medical time series analysis General research question: Can Mamba models improve the performance of classification models trained to detect Parkinson's disease using gait signals.