█

AI/TLDR

seekdb

A MySQL-compatible search database for agent state — vector, full-text and relational data in one SQL engine, with copy-on-write sandboxes you can fork and merge

Vector DatabasesOpen source
Language
C++
License
Apache-2.0
$pip install -U pyseekdb

Overview

seekdb is an open-source search database from the OceanBase team, built on the OceanBase SQL engine and aimed squarely at agent workloads. It stores vector, full-text, structured and semi-structured data in one engine and queries them in one SQL statement, speaks the MySQL wire protocol, and runs three ways: embedded in-process, as a single-node server, or inside an OceanBase distributed cluster. The Python SDK is `pyseekdb`.

Diagram of the write path: user DML goes to transaction commit, redo log and an immediate return, while an asynchronous Change Stream feeds a fetcher, dispatcher and worker pool that update a delta HNSW index which drains into a snapshot HNSW index; queries read both.
The Change Stream pipeline: commits return before the index is built, and queries read the delta and snapshot indexes together.seekdb README ↗

Its central design problem is the shape of agent traffic: continuous writes followed milliseconds later by a read. seekdb decouples index construction from the write path with an asynchronous index pipeline it calls Change Stream — a DML commit returns without waiting for the index build, while a fetcher, dispatcher and worker pool consume the redo log and update a delta HNSW index. Queries read both the incremental delta index and the snapshot index under fine-grained read locks, so newly written vectors are immediately searchable.

The second distinctive feature is branching. `FORK DATABASE` snapshots an entire database in seconds with no data copy, so an agent can write, query and break things inside a sandbox; `MERGE TABLE … STRATEGY` commits the work back with a FAIL, THEIRS or OURS conflict strategy, and `DROP DATABASE` throws it away. This is kernel-level copy-on-write rather than application-layer save and restore.

Because it keeps the MySQL protocol, the surrounding ecosystem comes for free: SQLAlchemy and any MySQL driver work, as do LangChain, LlamaIndex and Dify. The project publishes the benchmark it quotes — VectorDBBench's StreamingPerformanceCase on the Cohere 10M dataset at 768 dimensions — and the harness to reproduce it.

Bar chart comparing seekdb v1.3.0, Elasticsearch v9.3.3 and Milvus v2.6.9 on VectorDBBench's StreamingPerformanceCase: 1,523 QPS against 470 and 142, 21.7 ms concurrent P99 against 53.6 ms and 153.6 ms, and 1.1x P99 jitter against 10.3x and 9.7x.
The VectorDBBench StreamingPerformanceCase result seekdb publishes, on the Cohere 10M dataset at 768 dimensions.seekdb README ↗

What it does

  • Hybrid search in a single SQL statement: vector distance, full-text `MATCH … AGAINST` and scalar filters pushed into one execution plan
  • Asynchronous Change Stream index pipeline with a two-level HNSW (incremental delta + snapshot), so writes commit without waiting on index construction
  • `FORK DATABASE` / `MERGE TABLE` / `DROP DATABASE` copy-on-write sandboxes for agent exploration and rollback
  • MySQL-compatible with full ACID — works with SQLAlchemy, any MySQL driver, and with LangChain, LlamaIndex and Dify
  • Three deployment shapes from the same engine: embedded library, single-node server, or an OceanBase distributed cluster
  • `pyseekdb`, a Python SDK with a collection API (`get_or_create_collection`, `upsert`, `query`) and built-in embedding functions

Getting started

Embedded mode needs nothing but the Python SDK; the server modes ship as a Docker image, a binary installer and a Homebrew tap.

Install the Python SDK

`pyseekdb` is the Python client. In embedded mode it runs the engine in-process, so there is no server to start.

bashbash
pip install -U pyseekdb
Animated terminal capture of the project's 30-second demo, starting with the pip install of the pyseekdb SDK.
The README's 30-second try: install pyseekdb and run the demo script — no server, no schema, no embedding setup.seekdb README ↗

Write and read agent memory

Point a client at a local file, create a collection, then upsert observations and query them back. Because the index pipeline is asynchronous, a document written on one step is retrievable on the next.

pythonpython
import pyseekdb

client = pyseekdb.Client(path="./agent_state.db")
memory = client.get_or_create_collection(name="episodic")

for step in agent.run():
    # Persist the observation
    memory.upsert(ids=[step.id], documents=[step.observation])

    # Retrieve relevant context — milliseconds after the write
    relevant = memory.query(query_texts=step.next_query, n_results=5)

    agent.act(relevant)

Or run the server with Docker

The server image exposes the MySQL protocol on 2881 and the RPC port on 2886. Mount a volume so the data directory survives a restart.

bashbash
docker run -d \
  --name seekdb \
  -p 2881:2881 \
  -p 2886:2886 \
  -v ./data:/var/lib/oceanbase \
  oceanbase/seekdb:latest

Create a hybrid table and query it

A table can carry a VECTOR column with an HNSW index and a FULLTEXT index side by side; one SELECT then combines similarity, full-text match and scalar predicates.

sqlsql
CREATE TABLE articles (
  id        INT PRIMARY KEY,
  title     TEXT,
  content   TEXT,
  embedding VECTOR(384),
  FULLTEXT INDEX idx_fts (content) WITH PARSER ik,
  VECTOR   INDEX idx_vec (embedding) WITH (DISTANCE=l2, TYPE=hnsw, LIB=vsag)
) ORGANIZATION = HEAP;

SELECT id, title,
       l2_distance(embedding, '[0.12, 0.34, ...]') AS dist
FROM articles
WHERE MATCH(content) AGAINST('quarterly report')
ORDER BY dist APPROXIMATE
LIMIT 10;

Fork a sandbox, then merge or discard it

`FORK DATABASE` is a copy-on-write snapshot taken in seconds. Let the agent work inside the fork, then merge the tables you want to keep — FAIL, THEIRS or OURS decide how conflicts resolve — or drop the whole database.

sqlsql
FORK DATABASE agent_state TO agent_sandbox_42;

USE agent_sandbox_42;
INSERT INTO memory (session_id, embedding, content) VALUES (...);

MERGE TABLE agent_sandbox_42.memory INTO agent_state.memory STRATEGY THEIRS;
-- ...or throw it away:
DROP DATABASE agent_sandbox_42;

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Reach for it when an agent needs durable memory that is writable and searchable in the same loop, without a separate vector store and relational database
  • Reach for it when you want speculative agent work isolated in a branch you can merge or throw away, rather than reconciled by hand
  • Reach for it for RAG where the query needs vector similarity, keyword match and access-control filters resolved together in one plan
  • Reach for it when a MySQL-compatible, ACID engine is easier to operate and embed than another bespoke database

How seekdb compares

seekdb alongside other open-source vector databases tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Supabase★ 111kManaged Postgres backend whose Vector toolkit (pgvector) stores, indexes, and queries embeddings next to transactional data.
Redis Cloud★ 76.6kFully-managed Redis with built-in vector search, offering low-latency similarity and hybrid queries over any embeddings.
Milvus★ 46.3kA distributed vector database for storing and searching billions of embeddings at scale, with multiple index types and Kubernetes-native deployment.
FAISS★ 41kA library from Meta for efficient similarity search and clustering of dense vectors, with both exact and approximate indexes.
Qdrant★ 34.9kA Rust-based vector search engine that stores embeddings with rich payload filtering for semantic search and recommendation systems.
Chroma★ 29.4kA developer-focused vector database designed for quickly building retrieval and RAG features with a simple Python and JavaScript API.
pgvector★ 23.2kA PostgreSQL extension that adds a vector data type and similarity search so you can store and query embeddings inside an existing Postgres database.
seekdb★ 3.1kA MySQL-compatible search database for agent state — vector, full-text and relational data in one SQL engine, with copy-on-write sandboxes you can fork and merge