多阶段气候声明检索与排序:BM25、稠密 ANN、融合、重排序与评估Multi-stage climate claim retrieval and ranking: BM25, dense ANN, fusion, reranking, and evaluation
仓库/Skill 库
1033 个 · AI 核心
由机器强制执行的 AI Agent 前端 UI/UX Skill 包。注册表 + 懒加载:2,018 token 的路由在每次请求中从 334k token 深度内容中按需加载一个 Skill。包含 11 个发布阻塞闸门,其中一项用于校验 Skill 包自身文档与参考资料是否遵循其规则。Machine-enforced frontend UI/UX skill pack for AI agents. Registry + lazy loading: a 2,018-token router loads one skill per request out of 334k tokens of depth. 11 release-blocking gates, including one that checks the pack's own docs and reference material against its own rules.
适用于 macOS 的 local-first 语音+手势 AI 存在。唤醒词 → Whisper → 三层路由 → Claude Agent SDK + MCP,带有 kill switch、设备级 allowlist、声纹门控确认和完整的操作日志。Local-first voice + gesture AI presence for macOS. Wake word → Whisper → three-tier router → Claude Agent SDK + MCP, behind a kill switch, device-scoped allowlist, voiceprint-gated confirmation and a full action log.
Telegram RAG 机器人,以暗黑荒诞的《权力的游戏》谋士风格回答问题;支持英中双语,基于 LLaMA 3.3 70B + Firestore 向量搜索。A Telegram RAG bot that answers your questions as a darkly absurdist Game of Thrones advisor — bilingual (EN/ZH), powered by LLaMA 3.3 70B + Firestore vector search.
💳 使用 XGBoost 检测金融交易欺诈,结合针对不平衡数据集与复杂特征的高级优化技术💳 Detect fraudulent financial transactions using XGBoost. Optimize performance with advanced techniques for imbalanced datasets and complex features.
面向 LLM 的阶段感知上下文窗口治理框架。提供不变的上下文长度上限、基于熵的稳定性控制,以及针对降级(碎片化)状态的概率性保证,适用于生产级 LLM 系统。Phase-aware context window governance framework for Large Language Models (LLMs). Provides invariant context length caps, entropy-based stability control, and probabilistic guarantees against degraded (fragmentation) states for production LLM systems.
这个仓库展示了我的技能、项目,以及从数据工程转型为数据科学家角色的持续学习历程。This repository showcases my skills, projects, and continuous learning journey as I transition from Data Engineering to a Data Scientist role.
硕士论文(KCL):在攻击者-机器留出划分下重新衡量 IoT 入侵检测,并测试基于 LLM 的报文预测是否能提升性能。移除泄露后,macro-F1 从 0.90 降至 0.60。MSc dissertation (KCL): re-measuring IoT intrusion detection under an attacker-machine holdout, and testing whether LLM-based packet prediction improves it. macro-F1 0.90 -> 0.60 once leakage is removed.
🎨 利用 GPT、Gemini 等模型,AI 驱动的学术图表一键生成与定制工具🎨 Generate academic diagrams effortlessly with this AI-driven tool, leveraging models like GPT and Gemini for seamless creation and customization.
自动化 n8n 工作流,抓取 LinkedIn 产品经理职位,通过向量相似度与简历匹配打分,并使用 Claude 为每个职位生成简历修改建议Automated n8n workflow that scrapes LinkedIn PM jobs, scores them against your resume with vector similarity, and uses Claude to suggest per-job resume edits.
Buzz Agent 运行时 harness——监控 stock buzz-acp,可选的 sandbox 前置入口、DNA 完整性校验与 memory checkpoint。Buzz agent runtime harness — supervise stock buzz-acp, optional sandbox front door, DNA integrity, and memory checkpoint.
🛠 通过此硬件插件提升 vLLM 在 Kunlun XPU 上的性能,无缝集成主流 AI 模型并优化执行效率🛠 Enhance vLLM performance on Kunlun XPU with this hardware plugin, offering seamless integration for popular AI models and optimized execution.
🧪 实验性本地优先 RAG,附带桌面 GUI:将文档拖入文件夹即可对话聊天,内置 Ollama / fastembed / OCR / reranker。🧪 Experimental local-first RAG with desktop GUI — drop documents in a folder, chat with them. Ollama / fastembed / OCR / reranker built in.
🔍 Self-Corrective RAG 基于 LangGraph 与 Google Gemini 优化查询并评估文档相关性,提升检索效果🔍 Enhance your searches with Self-Corrective RAG, a system that optimizes queries and evaluates document relevance using LangGraph and Google Gemini.
Semantic Search for Pi 2026:本地知识库与 AI 工具。Semantic Search for Pi 2026: Local Knowledge Base & AI Tool
基于 .NET 的 RAG 文档摄取流水线 — 追踪 git 备份仓库中的文件夹,提取并分块文件,保持向量索引同步。.NET document ingestion pipeline for RAG — tracks a folder of files in a git-backed vault, extracts and chunks them, and keeps a vector index in sync.
为 AI agent 提供持久化、结构化记忆。基于 Go 的本地优先 MCP server:语义召回、去重、provenance、冲突解决。单一静态二进制,零依赖。Persistent, structured memory for AI agents. Local-first MCP server in Go: semantic recall, dedupe, provenance, conflict resolution. One static binary, zero dependencies.
多租户 SaaS 自动化平台,具备事件驱动工作流、AI 驱动的文档处理与实时集成。Multi-tenant SaaS automation platform with event-driven workflows, AI-powered document processing, and real-time integrations.
Vibe-Flow 2.0:通过 Wave Dispatch 实现交付、评审与合并 2026Vibe-Flow 2.0: Ship, Review & Merge via Wave Dispatch 2026
从零开始在 PyTorch 中构建 decoder-only Transformer,涵盖分词、预训练、监督微调,以及基于核心张量操作的对齐。Build decoder-only Transformers from scratch in PyTorch, covering tokenization, pretraining, supervised fine-tuning, and alignment using core tensor operations.
立场论文:编码 Agent 需要对代码库决策的显式建模,而不仅仅是结构建模。Position paper: coding agents need an explicit model of codebase decisions, not just structure"
为仿制药临床法规事务团队打造的法规准备加速器。A regulatory prep accelerator for a generic-drug Clinical Regulatory Affairs team.
使用 Qwen3-TTS 在本地 GPU 上克隆声音并从文本生成语音,提供端到端训练流水线。Clone voices and generate speech from text locally on your GPU using Qwen3-TTS with an end-to-end training pipeline.
自托管共享记忆,面向 AI Agent 团队。支持房间(Rooms)、L0-L3 层级深度与 MCP;读取路径不调用任何语言模型。基于 vectorize-io/hindsight(MIT)的 fork。Self-hosted shared memory for a team of AI agents. Rooms, hierarchical L0-L3 depth, MCP. The read path never invokes a language model. Fork of vectorize-io/hindsight (MIT).
保留图表的转换器 —— DOCX/XLSX 转 Markdown,采用原生 OOXML 图表数据提取(无需光栅化/OCR/VLM),并提供零损耗的复合图表标记。Converters where figures survive — DOCX/XLSX to Markdown with native OOXML chart-data extraction (no rasterize/OCR/VLM) and zero-loss composite-figure markers
3D 记忆岛屿,让 Google ADK agent 将你的记忆保存为有据可查的书:基于 Gemini on Vertex AI、Firestore、Pub/Sub 的 dreaming 机制构建知识图谱与 dark keepers,配合 Cloud TTS/STT 语音。All Things Agentic Hackathon 参赛作品。3D memory island where Google ADK agents keep your memories as grounded books: Gemini on Vertex AI, Firestore, Pub/Sub dreaming that builds a knowledge graph and dark keepers, Cloud TTS/STT voice. All Things Agentic Hackathon entry.
基于证据的构建流水线,编排 10 个 AI 子 Agent,并通过确定性 5 道关卡验证系统证明正确性。Evidence-based build pipeline that orchestrates 10 AI sub-agents and proves correctness through a deterministic 5-gate verification system.
AI 驱动的每周生活规划器 — 由多供应商 Agent 生成饮食、学习、锻炼、习惯和娱乐方案,并附带每周日程追踪。AI-powered weekly life planner - nutrition, study, workouts, habits and fun generated by a multi-provider agent, with a weekly schedule tracker.
🌐 使用 LangExtract 无缝从文本中提取语言,简化语言检测,为项目提供便捷且高准确率的增强。🌐 Extract languages from text seamlessly using LangExtract. Simplify language detection and enhance your projects with ease and accuracy.
基于 Next.js、FastAPI 和 Supabase 的 AI 辅导应用。上传 PDF 与 URL,即可生成结构化摘要、多轮对话和自定义测验。自定义主题通过 Wikipedia、ArXiv 和 DuckDuckGo 进行联网搜索,以生成测验并围绕主题展开对话。Next.js AI tutor with FastAPI and Supabase. Upload PDFs and URLs to generate structured summaries, multi-session chat, and custom quizzes. Custom topics use web search through Wikipedia, ArXiv, and DuckDuckGo to generate quizzes and chat about the topic.
NEXORA 的黑客松团队仓库 - [hackindia-team:ai-first-startup-hackathon-build-a-startup-using-ai-only:nexora]Hackathon team repository for NEXORA - [hackindia-team:ai-first-startup-hackathon-build-a-startup-using-ai-only:nexora]
Codiacs 黑客松团队仓库 - [hackindia-team:ai-first-startup-hackathon-build-a-startup-using-ai-only:codiacs]Hackathon team repository for Codiacs - [hackindia-team:ai-first-startup-hackathon-build-a-startup-using-ai-only:codiacs]
CodeByte 黑客松团队仓库 —— [hackindia-team:ai-first-startup-hackathon-build-a-startup-using-ai-only:codebyte]。Hackathon team repository for CodeByte - [hackindia-team:ai-first-startup-hackathon-build-a-startup-using-ai-only:codebyte]
适用于 DeepSeek、Qwen、GLM 及多种 AI 模型的 OpenAI 兼容 API 示例。OpenAI Compatible API examples for DeepSeek, Qwen, GLM and multiple AI models.
Local-first 可视化任务图,使人类与 AI Agent 在目标、进度、推理与下一步行动上保持对齐。Local-first visual task graph that keeps humans and AI agents aligned on goals, progress, reasoning, and the next action.