ACL 2026 论文 CAPruner 的官方实现:通过概念邻接场景图剪枝增强大语言模型的 3D 空间推理能力The official implementation for ACL 2026 paper CAPruner: Conceptual-Adjacent Scene Graph Pruner for Enhancing 3D Spatial Reasoning of Large Language Models
仓库/Skill 库
55 个 · 多模态 · AI 核心
[COLMW'26] AdaFlash:通过 On-Policy 蒸馏扩散 Drafters 实现的自适应投机解码[COLMW'26] AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
自托管、多用户转录平台:录制或上传音频。支持说话人标记、带时间戳的转录、跨录音识别说话人、摘要、提取行动项,并可使用自有 OpenAI 兼容 LLM 与转录内容对话。你的音频、你的服务器、你的模型。已在笔记本 RTX4070、台式机 RTX3090 与 RTX5090 上测试。Self-hosted, multi-user transcription platform: record or upload audio. Speaker-labeled, timestamped transcripts, Recognize speakers across recordings, Summarize, extract action items and chat over your transcripts with your own OpenAI-compatible LLM. Your Audio, your Server, your Model. Tested on Laptop RTX4070, Desktop RTX3090 and RTX5090
在 LIBERO 与 RoboTwin2.0 上,基于状态融合解码与冻结视频骨干网络训练面向机器人操作的轻量级世界动作模型。Train lightweight world action models for robot manipulation using state-fusion decoding and frozen video backbones on LIBERO and RoboTwin2.0.
AI 教材数字化与互动辅导平台:PDF/扫描教材经 MinerU OCR 结构化出题,文本 + 视觉双模型审校与确定性质量门禁;学生端七种题型互动、分层提示与错题多轮陪练闭环;FastAPI + PostgreSQL JSONB + React 19
AI 驱动的企业级 Agile OS,使用 Gemini 1.5 将原始客户音频与笔记即时转化为客户 pitch deck、PM epic、UI 规格与后端工程架构。An AI-powered Enterprise Agile OS that instantly translates raw client audio and notes into client pitch decks, PM epics, UI specs, and backend engineering architectures using Gemini 1.5.
基于 Llama 3.2 的本地多模态 RAG AI 助手,可与 PDF、图片和视频对话,完全离线,零数据外泄。Local multimodal RAG AI assistant where you can chat with PDFs, images, and video using Llama 3.2, fully offline, zero data exposure.
开源 Raspberry Pi 5 rover 项目,用于建图、视觉、语音、本地 LLM 及液态神经网络实验。Open-source Raspberry Pi 5 rover project for mapping, vision, voice, local LLMs, and liquid-neural-network experiments.
Neeharika Hemrajani,耶鲁管理学院与耶鲁大学计算机科学系 2026 年秋季独立研究项目。本项目针对光电容积脉搏波 (PPG) 的基础模型与逐人适配技术进行研究,衡量现有方法在成本与准确率之间的权衡。Neeharika Hemrajani, Independent Study, Yale School of Management and Yale University Department of Computer Science, Fall 2026. The following project is an independent study of foundation models for photoplethysmography (PPG) and per-person adaptation techniques to measure the cost versus accuracy gains across existing methods.
本地使用 AI 总结视频,从讲座、会议和教程中提取视觉亮点与文本摘要,基于 RTX 40 系列 GPU 的 CUDA 加速Summarize videos locally with AI, extracting visual highlights and text summaries from lectures, meetings, and tutorials using CUDA on RTX 40-series GPUs
每日提示词,激发 nano banana 生成灵感。Daily prompts to inspire nano banana generation.
🎨 利用 GPT、Gemini 等模型,AI 驱动的学术图表一键生成与定制工具🎨 Generate academic diagrams effortlessly with this AI-driven tool, leveraging models like GPT and Gemini for seamless creation and customization.
使用 Qwen3-TTS 在本地 GPU 上克隆声音并从文本生成语音,提供端到端训练流水线。Clone voices and generate speech from text locally on your GPU using Qwen3-TTS with an end-to-end training pipeline.
保留图表的转换器 —— DOCX/XLSX 转 Markdown,采用原生 OOXML 图表数据提取(无需光栅化/OCR/VLM),并提供零损耗的复合图表标记。Converters where figures survive — DOCX/XLSX to Markdown with native OOXML chart-data extraction (no rasterize/OCR/VLM) and zero-loss composite-figure markers
使用 vision-language model 处理视觉与文本数据,执行多模态推理与图像理解任务。Process visual and textual data with this vision-language model for multimodal reasoning and image understanding tasks.
基于 Whisper、Gemini、Streamlit、yt-dlp 与 FFmpeg 构建的 AI 应用,可即时转录并总结任意 YouTube 视频。Transcribe and summarize any YouTube video instantly with AI-powered app using Whisper, Gemini, Streamlit, yt-dlp & FFmpeg.
🌐 在数据受限条件下重新思考多模态大语言模型的设计与扩展,以 NaViL 通过 Native Training 提升效率与性能🌐 Rethink Multimodal Large Language Models design and scaling under data constraints with NaViL, enhancing efficiency and performance through Native Training.