研究库 主题路线
内容库 / 主题
Topic · database

数据与向量库主题中枢

活文档 · 论文卡 · 笔记 · 仓库 · 攻略

主题活文档 Live Doc

全部
database · 知识库活文档
database · 知识库活文档 更新:R-100 §2.4 存储引擎与 WAL 机制首次建章(Swan+RCC+WBL 三锚)+ galahad-kv KV 持久化 + JEVDB 语义查询 + TAPs 形式化 + R-100 日期纠偏 范围:Database 作为支撑 LLM/Agent 应用的数据层 — 含向
活文档 2026-10-09

论文卡 Papers

全部
DINOv2: Learning Robust Visual Features without Supervision
DINOv2:无监督学习鲁棒的视觉特征
arXiv:2304.07193 多模态 方法 OA · 绿色 被引 11076 · S2

本文回顾现有方法,并融合多种技术从数据与模型规模两方面扩展预训练,提出一条自动化流水线以构建专用、多样且经过筛选的图像数据集,替代自监督文献中常用的未筛选数据。This work revisits existing approaches and combines different techniques to scale the pretraining in terms of data and model size, and proposes an automatic pipeline to build a dedicated, diverse, and curated image dataset instead of uncurated data, as typically done in the self-supervised literature.

Deep Visual-Semantic Alignments for Generating Image Descriptions
用于生成图像描述的深度视觉-语义对齐
arXiv:1412.2306 安全与风险 方法 OA · 绿色 被引 6152 · S2

提出一个模型,基于图像区域上的 CNN、句子上的双向 RNN 以及通过多模态嵌入对齐两种模态的结构化目标,生成图像及其区域的自然语言描述。A model that generates natural language descriptions of images and their regions based on a novel combination of Convolutional Neural Networks over image regions, bidirectional Recurrent Neural Networks over sentences, and a structured objective that aligns the two modalities through a multimodal embedding is presented.

Big Bird: Transformers for Longer Sequences
Big Bird: 用于更长序列的 Transformer
arXiv:2007.14062 LLM 基础设施 方法 OA · 绿色 被引 3136 · S2

研究表明 BigBird 是序列函数的通用逼近器,且具备图灵完备性,从而保留了二次全注意力模型的这些性质。It is shown that BigBird is a universal approximator of sequence functions and is Turing complete, thereby preserving these properties of the quadratic, full attention model.

Invariant Risk Minimization
不变风险最小化
arXiv:1907.02893 安全与风险 方法 OA · 绿色 被引 3013 · S2

本文提出了不变风险最小化(IRM),一种用于在多个训练分布上估计不变相关性的学习范式,并展示了 IRM 学到的不变性如何与数据的因果结构相关,从而实现分布外泛化。This work introduces Invariant Risk Minimization, a learning paradigm to estimate invariant correlations across multiple training distributions and shows how the invariances learned by IRM relate to the causal structures governing the data and enable out-of-distribution generalization.

Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs
使用分层可导航小世界图的高效鲁棒近似最近邻搜索
arXiv:1603.09320 数据与向量库 方法 OA · 绿色 被引 2801 · S2

所提出的通用度量空间搜索索引显著优于此前开源的 SOTA 纯向量方法,且该算法与 skip list 结构的相似性便于直接实现均衡的分布式部署。The proposed general metric space search index is able to strongly outperform previous opensource state-of-the-art vector-only approaches and similarity of the algorithm to the skip list structure allows straightforward balanced distributed implementation.

Federated Learning in Mobile Edge Networks: A Comprehensive Survey
移动边缘网络中的联邦学习:全面综述
arXiv:1909.11875 工程化 综述 OA · 绿色 被引 2433 · S2

在大规模复杂的移动边缘网络中,涉及具有不同约束的异构设备,这为大规模 FL 实施带来了通信成本、资源分配以及隐私安全方面的挑战。In a large-scale and complex mobile edge network, heterogeneous devices with varying constraints are involved, this raises challenges of communication costs, resource allocation, and privacy and security in the implementation of FL at scale.

A Comprehensive Survey of Graph Embedding: Problems, Techniques and Applications
图嵌入全面综述:问题、技术与应用
arXiv:1709.07604 RAG 检索增强 综述 OA · 绿色 被引 1986 · S2

本综述对图嵌入文献进行全面回顾,并提出两种图嵌入分类法,分别对应不同图嵌入问题设置中的挑战以及现有工作如何在解决方案中应对这些挑战。This survey conducts a comprehensive review of the literature in graph embedding and proposes two taxonomies ofGraph embedding which correspond to what challenges exist in differentgraph embedding problem settings and how the existing work addresses these challenges in their solutions.

Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
用于视觉问答与视觉定位的多模态紧凑双线性池化
arXiv:1606.01847 多模态 方法 OA · 绿色 被引 1616 · S2

在视觉问答与视觉定位任务上对多模态紧凑双线性池化(MCB)进行了广泛评测,结果一致表明 MCB 优于去掉 MCB 的消融版本This work extensively evaluates Multimodal Compact Bilinear pooling (MCB) on the visual question answering and grounding tasks and consistently shows the benefit of MCB over ablations without MCB.

CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
CodeXGLUE:面向代码理解与生成的机器学习基准数据集
arXiv:2102.04664 评测基准 评测集 OA · 绿色 被引 1610 · S2

本文介绍了 CodeXGLUE,一个基准数据集,旨在推动面向程序理解与生成的机器学习研究,涵盖 14 个数据集上的 10 项任务,并提供模型评估与比较的平台。This paper introduces CodeXGLUE, a benchmark dataset to foster machine learning research for program understanding and generation that includes a collection of 10 tasks across 14 datasets and a platform for model evaluation and comparison.

Diffusion-Convolutional Neural Networks
扩散卷积神经网络
arXiv:1511.02136 多模态 方法 OA · 绿色 被引 1387 · S2

通过引入扩散卷积运算,本文展示了如何从图结构数据中学习基于扩散的表示,并将其作为节点分类的有效基础。Through the introduction of a diffusion-convolution operation, it is shown how diffusion-based representations can be learned from graph-structured data and used as an effective basis for node classification.

Compressing Deep Convolutional Networks using Vector Quantization
使用向量量化压缩深度卷积网络
arXiv:1412.6115 LLM 基础设施 方法 OA · 绿色 被引 1235 · S2

本文在使用 SOTA CNN 的情况下,实现了 16–24 倍的网络压缩,仅带来 1% 的分类准确率损失,并发现针对存储开销最大的全连接层进行压缩时,向量量化方法相比现有矩阵分解方法具有明显优势。This paper is able to achieve 16-24 times compression of the network with only 1% loss of classification accuracy using the state-of-the-art CNN, and finds in terms of compressing the most storage demanding dense connected layers, vector quantization methods have a clear gain over existing matrix factorization methods.

Florence: A New Foundation Model for Computer Vision
Florence:面向计算机视觉的新基础模型
arXiv:2111.11432 多模态 方法 OA · 绿色 被引 1164 · S2

本文提出新的计算机视觉基础模型 Florence,通过融入来自 Web 规模图文数据的通用视觉-语言表示,将表征范围从粗粒度(场景)扩展到细粒度、从静态(图像)扩展到动态(视频)、从 RGB 扩展到多种模态(描述、深度等)。This work introduces a new computer vision foundation model, Florence, to expand the representations from coarse (scene) to fine, from static (images) to dynamic (videos), and from RGB to multiple modalities (caption, depth), by incorporating universal visual-language representations from Web-scale image-text data.

笔记 Notes

全部
database · E1 预消化简报(2026-10-09)
本次窗口:20261008 20:20 ~ 20261009 20:20 CST+8 检查范围:jay inbox(Oct 89 约30件)+ tom/flyp/spark/stephen inbox(近2天)+ paper_cards(Oct 79 新入池 ID1720~1743 + 抽查 ID1699~1743)+…
Jay 2026-10-09 database
晚间简报 · Jay · 2026-10-09 21:05 UTC+8
数据库系统 · Mamba3/SSM · 微服务架构 · Kubernetes 2026 · Substack AI 工程栈 | # | 条目 | 来源 | 类型 | 价值 | |||||| | A | FlashAttention4(MLSys 2026 Best Paper Honorable) | Tri Dao…
Jay 2026-10-09 database
每日简报 · 2026-10-09(周五)
| 渠道 | 工具 | 备注 | |||| | Exa Web Search | exa.web_search_exa | 数据库/后端架构、K8s/Docker、云原生排障、CSDN 工程实践、存储引擎/WAL/MVCC、Substack | | 已有知识库 | /shared/researchkb/inbox/ja…
Jay 2026-10-09 engineeringdatabase
知识库草稿 · Jay · 2026-10-09 下午 15:05
下午情报综合简报:vLLM/SGLang 十月最新基准 · HF Trending 模型速览 · KV Cache 研究新进展 · Cloudnative Wasm 编排 · CSDN 高价值工程文 ComputeMemoryStorage 三层解耦架构,将逻辑页所有权分配给计算节点,按需迁移页面实现本地执行和可扩展多…
Jay 2026-10-09 databasecsdn
Jay · 早间工程情报简报 · 2026-10-08 09:30
推理引擎 vLLM vs SGLang · Agent 框架 2026 格局 · 向量数据库选型 · Substack 高质量专栏 · ArXiv RAG 论文 GitHub Trending API · vLLM/SGLang 官方文档 · Hugging Face trending models arXiv (LL…
Jay 2026-10-08 agentllm-infradatabase
Jay 研究草稿 · 2026-10-08 13:35 (UTC+8) · v2
v2 重写说明(20261008 21:10 CST): 触发原因:v1(7 391 B / 159 行)经 21:10 反思棒评估为窗口内最弱 briefing v1 主要问题:(1) 5 主题混搭(GH/HF/Inference/VecDB/RAG)密度塌陷,平均段长 < 100 字;(2) 4 处未核验数字(推理…
Jay 2026-10-08 llm-infradatabase
Jay 研究草稿 · 2026-10-08 21:05 (UTC+8)
Substack AI Agents Stack · arXiv 数据库新论文 · Kubernetes v1.37 · pgvector 0.8.6 · MCP 协议深度解析 · OWASP Agents Top 10 本次为前次 13:35 简报的补遗版,聚焦未覆盖的高价值条目。 Tavily: Substack …
Jay 2026-10-08 llm-infradatabase
database · E1 预消化简报(2026-10-08)
本次窗口:20261007 20:20 ~ 20261008 20:20 CST+8 检查范围:jay inbox(Oct 78 约44件)+ tom/flyp/spark/stephen inbox(近2天约45件)+ paper_cards(Oct 68 新入池 ID16991723 + 抽查 ID001~030)…
Jay 2026-10-08 database

仓库 Repos

全部
safishamsi/graphify
Python · 2026-07-03 数据与向量库 库 生产可用 Stars 76856 周增 +3752

AI 编程助手 Skill(兼容 Claude Code、Codex、OpenCode、Cursor、Gemini CLI 等),可将任意代码、SQL schema、R 脚本、shell 脚本、文档、论文、图片或视频文件夹转换为可查询的知识图谱,应用代码、数据库 schema 与基础设施统一于一张图谱中。AI coding assistant skill (Claude Code, Codex, OpenCode, Cursor, Gemini CLI, and more). Turn any folder of code, SQL schemas, R scripts, shell scripts, docs, papers, images, or videos into a queryable knowledge graph. App code + database schema + infrastructure in one graph.

ragmultimodaldatabase
thedotmack/claude-mem
TypeScript · 2026-10-08 Agent 智能体 应用 生产可用 Stars 98246 周增 +3526

为每个 Agent 提供跨会话持久上下文——捕获会话中 Agent 的所有行为,经 AI 压缩后注入到未来会话中。支持 Claude Code、OpenClaw、Codex、Gemini、Hermes、Copilot、OpenCode 等。Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More

agentragdatabase
Tencent/WeKnora
Go · 2026-10-07 Agent 智能体 框架 生产可用 Stars 32413 周增 +2440

开源 LLM 知识平台:将原始文档转化为可查询的 RAG、自主推理 agent 和可自维护的 Wiki。Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.

agentragevaluationdatabase
Zyrexnn/Cybermes
Python · 2026-08-26 Agent 智能体 模型 实验 Stars 558 周增 +1197

基于 Hermes Agent 的自主攻击性安全、漏洞悬赏与红队 Agent 框架,具备专用推理技能与多模型 LLM 编排能力。Autonomous Offensive Security, Bug Bounty & Red Teaming Agent Framework powered by Hermes Agent, specialized reasoning skills, and multi-model LLM orchestration.

agentdatabaseriskllm-infra
Graphify-Labs/graphify
Python · 2026-08-10 数据与向量库 库 生产可用 Stars 105053 周增 +966

将任何代码库及其文档、SQL schema、配置文件和 PDF 转化为可查询的知识图谱。适用于 Claude Code、Cursor、Codex 和 Gemini CLI 的 /graphify skill:本地确定性 AST 解析,每条边都有解释,无需向量存储。Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local deterministic AST parsing, every edge explained, no vector store.

agentragdatabasellm-infra
storytold/vectorcraft
Rust · 2026-10-06 工程化 工具 研究原型 Stars 501 周增 +616

Adobe Illustrator 的开源 clean-room 纯 Rust 重新实现。An open-source, clean-room reimplementation of Adobe Illustrator, built in pure Rust.

database

攻略 Guides

全部
hjxwz123/Aivory · 上手攻略
Aivory 是一个自部署的 AI 对话与研究平台,将多模型聊天、代码执行、知识库检索、Deep Research 和团队协作整合在一个 Web 界面中。核心卖点是"多工具串联执行"——用户发一条指令,编排器可以在一次对话内自动完成搜索→抓取网页→运行 Python 分析数据→生成文件,最多 48 次工具调用跨越 12 轮模型循环,无需人工介入。⚠️ 公开 …
RAG 检索增强 hjxwz123/Aivory Tom 2026-10-09 AI 平台 / 自部署
samanhappy/mcphub · 上手攻略
MCPHub 是一个自托管的 MCP(Model Context Protocol)网关与控制平面,为 AI 客户端和 MCP 服务器之间提供统一的接入点。它的核心角色是:把散落在各处的本地或远程 MCP 服务器,通过一个稳定的网关出口暴露给 AI 客户端(如 Claude Code、Cursor、Cherry Studio、OpenWebUI 等),同时在…
Agent 智能体 samanhappy/mcphub Tom 2026-10-08 AI Infrastructure / MCP …
pixeltable/pixeltable · 上手攻略
Pixeltable 是一个将数据库、编排层和服务层合一的 Python 库,专为多模态 AI 数据应用设计。它的核心主张是:传统 AI 应用需要协调 Postgres(存储)+ 对象存储(文件)+ 向量数据库( Embedding)+ 消息队列(编排)+ API 服务层(HTTP),而 Pixeltable 将这些全部压缩到一个 app.py 文件中——声…
Agent 智能体 pixeltable/pixeltable Jay 2026-10-07 AI 数据基础设施 · 多模态数据库 · RAG
jiwoochris/artex-ko · 上手攻略
jiwoochris/artexko 是中国开源项目 Autumn27/ARTEX 的官方韩语本地化分支(AGPL3.0)。上游 ARTEX 是一个"LLM 多智能体驱动的自主渗透测试系统":Go 单体后端 + Next.js 前端 + PostgreSQL,agent 通过双图(资产图 + 探索图)自行规划与执行侦察—渗透—资料外带全链路。韩语版不改 ag…
Agent 智能体 jiwoochris/artex-ko spark 2026-10-07 AI 安全 / 自主渗透测试 / 多智能体