仓库/Skill 库
167 个 · 多模态
交互式标注/分割的文献综述Literature Review for Interactive annotation/segmentation
视觉定位的文献综述Literature review of visual localization.
MatlowAI 的 MiniMax-H3 ComfyUI 节点:Contact-Sheet diffusion + Motion Lab(针对快速运动的 test-time 去绳状畸变)MatlowAI's MiniMax-H3 ComfyUI nodes: Contact-Sheet diffusion + Motion Lab (test-time de-roping of fast motion)
[ISPRS 2024] 卫星视频单目标跟踪:系统综述与定向目标跟踪基准[ISPRS 2024] Satellite Video Single Object Tracking: A Systematic Review and An Oriented Object Tracking Benchmark
独立的 Skill 与 Codex 插件,支持 OpenAI 兼容的图像生成、编辑、批处理工作流、QA 以及聚焦画布编辑。Standalone Skill and Codex Plugin for OpenAI-compatible image generation, editing, batch workflows, QA, and focused canvas editing.
DeepSeek Harness (DSH) 论文写作护栏——AI 写作风格检测、证据保留、期刊匹配校准、手稿校对、writing_audit 与自动检查。本地运行,零网络,零 LLM。DeepSeek Harness (DSH) academic writing guard for papers — 论文去AI味 / AI-writing style detection, evidence preservation, journal-fit calibration, manuscript proofreading, writing_audit & automatic checks. Local, zero network, zero LLM.
面向严谨学术论文写作、修订与投稿的 Claude Code skill。跨领域通用,支持按论文设置期刊覆盖规则。Claude Code skill for rigorous academic paper writing, revision, and submission. Field-agnostic with per-paper journal overrides.
Local-first 学术 PDF 阅读器:离线 CPU LLM 翻译与注释、OCR、ECDICT 词典。Tauri + FastAPI。Local-first academic PDF reader: offline CPU LLM translation & glossing, OCR, ECDICT dictionary. Tauri + FastAPI.
证据优先的具身 AI 研究枢纽,覆盖 VLA、世界模型、多模态感知、文献综述、声明审计以及可复用的 Agent skillEvidence-first Embodied AI research hub for VLA, world models, multimodal sensing, literature reviews, claim audits, and reusable agent skills.
[COLMW'26] AdaFlash:通过 On-Policy 蒸馏扩散 Drafters 实现的自适应投机解码[COLMW'26] AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
在 LIBERO 与 RoboTwin2.0 上,基于状态融合解码与冻结视频骨干网络训练面向机器人操作的轻量级世界动作模型。Train lightweight world action models for robot manipulation using state-fusion decoding and frozen video backbones on LIBERO and RoboTwin2.0.
自托管、多用户转录平台:录制或上传音频。支持说话人标记、带时间戳的转录、跨录音识别说话人、摘要、提取行动项,并可使用自有 OpenAI 兼容 LLM 与转录内容对话。你的音频、你的服务器、你的模型。已在笔记本 RTX4070、台式机 RTX3090 与 RTX5090 上测试。Self-hosted, multi-user transcription platform: record or upload audio. Speaker-labeled, timestamped transcripts, Recognize speakers across recordings, Summarize, extract action items and chat over your transcripts with your own OpenAI-compatible LLM. Your Audio, your Server, your Model. Tested on Laptop RTX4070, Desktop RTX3090 and RTX5090
作者感知的科学写作分析:作者空间、目标论文风格、校准的生成证据,以及安全合规的修订Authorship-aware scientific writing analysis: author space, target-paper style, calibrated generation evidence, and integrity-safe revision.
论文《Deep Learning in Dental Image Analysis: A Systematic Review of Datasets, Methodologies, and Emerging Challenges》的官方项目主页。The official project homepage for the paper "Deep Learning in Dental Image Analysis: A Systematic Review of Datasets, Methodologies, and Emerging Challenges"
MCP server,可从 URL 提取任意视频(包括微信视频号)。可选择性地分析并提取场景感知的去重关键帧,并对视频进行转录。完全本地运行MCP server that extracts any video from URL(inc/ WeChat channels). Optionally analyzes and extracts scene-aware deduplicated key frames and transcribes the video. ALL LOCAL
将任意视频链接转化为博士级研究报告。自动化学术流水线:一键完成下载、转写、检索同行评审文献、核查论据并生成带引用的报告。免费且开源。Turn any video url into a PhD-grade research report. Automated academic pipeline: download, transcribe, search peer-reviewed literature, verify claims, generate cited reports — all with one command. Free & open source.
论文《Transformer-Based Tree Extraction from Remote Sensing Imagery: A Systematic Review》的补充代码库。Supplementary repository for: Transformer-Based Tree Extraction from Remote Sensing Imagery: A Systematic Review
AI 教材数字化与互动辅导平台:PDF/扫描教材经 MinerU OCR 结构化出题,文本 + 视觉双模型审校与确定性质量门禁;学生端七种题型互动、分层提示与错题多轮陪练闭环;FastAPI + PostgreSQL JSONB + React 19
TTS、语音转换与口语对话 Agent 的活体系统综述Living systematic review of TTS, voice conversion, and spoken conversational agents
AI 驱动的企业级 Agile OS,使用 Gemini 1.5 将原始客户音频与笔记即时转化为客户 pitch deck、PM epic、UI 规格与后端工程架构。An AI-powered Enterprise Agile OS that instantly translates raw client audio and notes into client pitch decks, PM epics, UI specs, and backend engineering architectures using Gemini 1.5.
面向学术阅读的本地桌面 PDF 阅读器,支持 AI 翻译、OCR、高亮、笔记、翻译历史以及导入/导出。A local desktop PDF reader for academic reading, with AI translation, OCR, highlights, notes, translation history, and import/export support.
《Comparative Analysis of EEG Signal Preprocessing Impact on Generalization of a Transformer-Based EEG Imagined Speech Recognition》研究的可复现研究仓库。Reproducible research repository for 'Comparative Analysis of EEG Signal Preprocessing Impact on Generalization of a Transformer-Based EEG Imagined Speech Recognition' research work
围绕病理 foundation models(尤其是 Prov-GigaPath)的文献综述与可复现性研究Literature review and reproducibility work on pathology foundation models, with a focus on Prov-GigaPath.
基于 Llama 3.2 的本地多模态 RAG AI 助手,可与 PDF、图片和视频对话,完全离线,零数据外泄。Local multimodal RAG AI assistant where you can chat with PDFs, images, and video using Llama 3.2, fully offline, zero data exposure.
开源 Raspberry Pi 5 rover 项目,用于建图、视觉、语音、本地 LLM 及液态神经网络实验。Open-source Raspberry Pi 5 rover project for mapping, vision, voice, local LLMs, and liquid-neural-network experiments.
Embodied Daily:每日具身智能 AI 论文推荐(HF Daily Papers + arXiv)Embodied Daily: daily embodied-AI paper recommendations (HF Daily Papers + arXiv)
数据驱动、自动验证的机器人研究文献综述语料库(manipulation、locomotion、perception、planning、learning、HRI、multi-robot、simulation)。Robotics Research Corpus: Data-driven, auto-validated literature review for robotics research (manipulation, locomotion, perception, planning, learning, HRI, multi-robot, simulation)
本仓库最初是我毕设项目的 PyTorch Lightning 模板,后来演进为包含完整项目实现本身。研究问题是地理定位如何影响遥感视觉-语言模型的训练与推理。注:仓库名称与描述可能变化。This repo originally was a PyTorch Lightning template for my thesis project. It has since evolved to contain the full project implementation itself. The research question is how geolocation influences the training and inference of vision-language models for remote sensing. Note: The repository name and description may change
更简洁的照片编辑体验;开源 RAW 与 JPEG 流水线。Photo editing, simplified. An open-source RAW and JPEG workflow.
Codex skill:将已批准的科研图表蓝图转化为可溯源、使用真实数据的发表级图表。Codex skill for turning approved scientific figure blueprints into traceable real-data publication figures
面向中英双语生物医学研究的术语循证、主张强度与科学写作审计 skill。Evidence-aligned terminology, claim-strength, and scientific-writing audit skill for Chinese and English biomedical research.
SwiftUI 原生 LaTeX 数学渲染,对 LLM 输出具有鲁棒性——支持分隔符归一化、流式渲染与自动主题适配SwiftUI-native LaTeX math rendering robust to LLM output — delimiter normalization, streaming support, and automatic theming