研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

281 个

排序 Stars 周增
lukas-blecher/LaTeX-OCR
Python · 2025-01-18 多模态 工具 实验 Stars 16530 周增 +7

pix2tex:使用 ViT 将公式图像转换为 LaTeX 代码pix2tex: Using a ViT to convert images of equations into LaTeX code.

multimodalengineering
memvid/memvid
Rust · 2026-07-14 Agent 智能体 应用 生产可用 Stars 16207 周增 +35

AI Agents 的记忆层。用无服务器、单文件的记忆层替代复杂的 RAG pipeline,为 agents 提供即时检索与长期记忆。Memory layer for AI Agents. Replace complex RAG pipelines with a serverless, single-file memory layer. Give your agents instant retrieval and long-term memory.

agentragmultimodalengineering
HBAI-Ltd/Toonflow-app
TypeScript · 2026-07-28 多模态 框架 生产可用 Stars 13701 周增 +126

Toonflow 是开源一站式 AI 短剧创作工具,将小说、剧本快速转化为动画短剧。集成 AI 编剧、智能分镜、角色与视频生成,跨平台桌面端轻量部署,助力创作者低成本批量产出视觉内容。Toonflow is an open-source AI tool that turns stories and scripts into animated short dramas. Features AI scriptwriting, storyboarding, character and video generation. A cross-platform desktop app for efficient content creation.

multimodalllm-infra
waooAI/waoowaoo
TypeScript · 2026-08-10 Agent 智能体 框架 生产可用 Stars 13554 周增 +21

首家工业级全流程 AI 影视生产平台。行业领先的专业 AI Agent 平台,实现可控的电影与视频制作,涵盖短视频到真人实拍,遵循好莱坞级工作流。首家工业级全流程 AI 影视生产平台。Industry-first professional AI Agent platform for controllable film & video production. From shorts to live-action with Hollywood-standard workflows.

agentmultimodalengineering
CoplayDev/unity-mcp
C# · 2026-08-07 Agent 智能体 应用 生产可用 Stars 13320 周增 +112

Unity MCP 在 AI 助手与 Unity Editor 之间充当桥梁,让 LLM 具备管理资源、控制场景、编辑脚本以及自动化任务的能力。Unity MCP acts as a bridge between AI assistants and your Unity Editor. Give your LLM tools to manage assets, control scenes, edit scripts, and automate tasks within Unity.

multimodalllm-infra
zoicware/RemoveWindowsAI
PowerShell · 2026-08-10 安全与风险 工具 生产可用 Stars 12715 周增 +35

在 Windows 11 中强制移除 Copilot、Recall 等组件Force Remove Copilot, Recall and More in Windows 11

multimodalrisk
bentoml/OpenLLM
Python · 2026-08-10 LLM 基础设施 框架 生产可用 Stars 12473 周增 +77

在云端以 OpenAI 兼容 API 端点形式运行任意开源 LLM,例如 DeepSeek 和 Llama。Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.

multimodalllm-infraengineering
img2threejs/img2threejs
Python · 2026-08-10 多模态 工具 实验 Stars 10708 周增 +1330

将参考图像中的物体重建为纯代码、程序化、质量可控、可直接用于动画的 Three.js 模型。Token 高效的图像转 3D。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
openvinotoolkit/openvino
C++ · 2026-08-11 LLM 基础设施 工具 生产可用 Stars 10637 周增 +21

OpenVINO™ 是用于优化和部署 AI 推理的开源工具包OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

multimodalllm-infra
ultralytics/yolov3
Python · 2026-08-02 LLM 基础设施 应用 生产可用 Stars 10591 周增 +14

基于 PyTorch 的 YOLOv3、YOLOv3-SPP 和 YOLOv3-tiny 实时目标检测实现,支持训练、验证、推理与多格式导出。PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.

multimodalllm-infra
Acly/krita-ai-diffusion
Python · 2026-06-30 多模态 应用 生产可用 Stars 10449 周增 +21

在 Krita 中使用 AI 生成图像的精简界面。支持 Inpaint 与 Outpaint,可选文本提示,无需调参。Streamlined interface for generating images with AI in Krita. Inpaint and outpaint with optional text prompt, no tweaking required.

multimodal
RunanywhereAI/runanywhere-sdks
C++ · 2026-08-11 LLM 基础设施 库 生产可用 Stars 10301 周增 -7

用于本地运行 AI 的生产就绪工具包。Production ready toolkit to run AI locally

multimodalllm-infraengineering
omnigent-ai/omnigent
Python · 2026-09-02 Agent 智能体 框架 生产可用 Stars 9602 周增 +347

Omnigent 是一个开源 AI agent 框架与元 harness:编排 Claude Code、Codex、Cursor、Pi 及自定义 agent——无需重写即可替换 harness,强制执行策略与沙箱化,并支持任意设备实时协作。Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.

agentmultimodalllm-infra
xorbitsai/inference
Python · 2026-08-11 多模态 应用 生产可用 Stars 9486 周增 +0

通过修改一行代码即可将 GPT 替换为任意 LLM。Xinference 让你在云端、本地或笔记本上运行开源、语音和多模态模型,全部通过统一的生产就绪推理 API。Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

multimodalllm-infraengineering
xberg-io/xberg
Rust · 2026-10-02 Agent 智能体 工具 生产可用 Stars 9362 周增 +32

基于 Rust 核心的多语言文档智能:从 106 种格式、140 种文件扩展名中提取文本、元数据、图像、表格与结构化数据,并支持 371 种语言的代码智能。提供十五种语言绑定,配套 CLI、REST API 与 MCP server。Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.

ragmultimodal
activeloopai/deeplake
C++ · 2026-05-21 Agent 智能体 应用 生产可用 Stars 9227 周增 +7

Deeplake 是面向 Agent 的 AI 数据运行时,提供无服务器 postgres 与多模态 datalake,支持可扩展的检索与训练Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.

agentragmultimodalengineering
storytold/photocraft
Rust · 2026-10-07 多模态 库 生产可用 Stars 8993 周增 +27836

纯 Rust 编写的 Adobe Photoshop 开源 clean-room 重新实现An open-source, clean-room reimplementation of Adobe Photoshop in pure Rust

multimodal
bentoml/BentoML
Python · 2026-08-03 LLM 基础设施 模型 生产可用 Stars 8776 周增 +7

提供 AI 应用和模型服务的最简方式 —— 构建模型推理 API、任务队列、LLM 应用、多模型 pipeline 等。The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

multimodalllm-infraengineering
Lightricks/LTX-2
Python · 2026-08-03 多模态 模型 生产可用 Stars 8560 周增 +14

LTX-2 音视频生成模型的官方 Python 推理与 LoRA 训练包。Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.

multimodalllm-infra
helloianneo/ian-xiaohei-illustrations
未知语言 · 2026-06-03 多模态 工具 生产可用 Stars 8494 周增 +119

中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill

agentmultimodal
AI4Finance-Foundation/FinRobot
Python · 2026-09-28 Agent 智能体 框架 生产可用 Stars 8098 周增 +49

FinRobot:基于 LLM 的开源金融分析 AI Agent 平台 🚀 🚀 🚀FinRobot: An Open-Source AI Agent Platform for Financial Applications using Large Language Models

agentmultimodalllm-infra
open-mmlab/mmagic
Jupyter Notebook · 2024-08-06 多模态 模型 实验 Stars 7452 周增 +0

OpenMMLab 多模态高级生成与智能创作工具箱。释放魔力🪄:AIGC、易用 API、丰富模型库、扩散模型,支持文生图、图像/视频修复与增强等任务OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.

multimodal
StarTrail-org/PixelRAG
Python · 2026-07-16 RAG 检索增强 框架 生产可用 Stars 7303 周增 +392

网页解析的终结,可扩展像素原生搜索的开端。链接:https://pixelrag.ai/The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/

agentragmultimodal
teamchong/pxpipe
TypeScript · 2026-07-18 多模态 工具 研究原型 Stars 6425 周增 +140

通过将文本上下文渲染为图像,将 Fable 5 的 token 使用量降低cut Fable 5 token usage by rendering text context as images

multimodal
genkit-ai/genkit
TypeScript · 2026-08-17 Agent 智能体 框架 生产可用 Stars 6342 周增 +0

用于构建 agentic apps 的开源框架,支持 JavaScript、Go、Dart 与 Python,由 Google 在生产环境中构建并使用。Open-source framework for building agentic apps in JavaScript, Go, Dart, and Python, built and used in production by Google

agentragmultimodalengineering
vllm-project/vllm-omni
Python · 2026-08-11 LLM 基础设施 模型 生产可用 Stars 6027 周增 +77

面向全模态模型的高效推理框架。A framework for efficient model inference with omni-modality models

multimodalllm-infra
OpenBMB/UltraRAG
Python · 2026-08-11 Agent 智能体 框架 生产可用 Stars 5668 周增 +0

用于构建复杂创新 RAG 流水线的低代码 MCP 框架A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines

ragmultimodalllm-infraengineering
op7418/guizang-social-card-skill
HTML · 2026-07-01 多模态 工具 研究原型 Stars 5512 周增 +70

🪧 Claude Code / Codex Skill——生成小红书图文卡片与公众号 21:9+1:1 封面配对。编辑 × Swiss 视觉系统,28 套版式,10 种主题,单文件 HTML → PNG。小红书图文 + 公众号封面对🪧 Claude Code / Codex skill — generate Xiaohongshu carousels & WeChat 21:9+1:1 cover pairs. Editorial × Swiss visual systems, 28 layouts, 10 themes, single-file HTML → PNG. 小红书图文 + 公众号封面对

agentmultimodal
LiamGvchi/gc-minimal-zine-poster
未知语言 · 2026-08-09 多模态 工具 实验 Stars 5385 周增 +693

Codex skill,用于生成安静极简的 zine 风格编辑海报提示词与图像。Codex skill for generating quiet minimal zine-style editorial poster prompts and images.

multimodal
WeThinkIn/AIGC-Interview-Book
未知语言 · 2026-09-26 Agent 智能体 应用 研究原型 Stars 4823 周增 +77

【三年面试五年模拟】AIGC/LLM/AI Agent算法工程师面试资源平台。涵盖AIGC、LLM大模型、AI Agent、具身智能、传统深度学习、计算机视觉、自然语言处理、自动驾驶、机器学习、强化学习、大数据挖掘、世界模型、元宇宙、AGI等AI行业面试笔试干货经验与核心跨周期知识。

agentmultimodalllm-infra
Vincentwei1021/video-shotcraft
TypeScript · 2026-08-09 多模态 教程 研究原型 Stars 4535 周增 +350

面向 Claude Code 和 Codex 的 AI 视频 skill —— 基于 Remotion 制作电影级产品视频:含 152 张分镜配方卡、209 个动效预览,以及一套开箱即用的模板。AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 152 shot recipe cards, 209 motion previews, a production-ready template

agentmultimodalengineering
luban-agi/Awesome-AIGC-Tutorials
未知语言 · 2024-03-31 LLM 基础设施 收藏榜 生产可用 Stars 4531 周增 -7

精选的 LLM、AI 绘画等领域的教程和资源。Curated tutorials and resources for Large Language Models, AI Painting, and more.

multimodalllm-infra
EvolvingLMMs-Lab/lmms-eval
Python · 2026-08-29 多模态 工具 研究原型 Stars 4382 周增 +10

一个覆盖文本、图像、视频与音频任务的统一多模态评估工具包。One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

multimodalevaluationllm-infra
wuyoscar/GPT-Image2-Skill
Python · 2026-08-10 Agent 智能体 库 研究原型 Stars 4348 周增 +287

GPT Image 2 的 prompt gallery、image prompt library、agentic skill 以及用于 OpenAI 图像生成/编辑的 CLI。GPT Image 2 prompt gallery, image prompt library, agentic skill, and CLI for OpenAI image generation/editing

agentmultimodal
open-compass/VLMEvalKit
Python · 2026-08-17 评测基准 评测集 研究原型 Stars 4345 周增 +9

大型多模态模型 LMM 的开源评估工具包,支持 220+ LMM 与 80+ 基准测试。Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

multimodalevaluationllm-infra
hoainho/img2threejs
Python · 2026-07-25 多模态 工具 实验 Stars 4176 周增 +2051

将参考图像中的对象重建为纯代码、程序化、带质量门控、可动画化的 Three.js 模型。token 高效的图像到三维转换。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal