█

AI/TLDR

Endee

A C++ vector database server with dense and sparse (BM25) search, payload filtering and backups, tuned per CPU instruction set

Vector DatabasesOpen core
Latest
1.3.5
Updated
22 May 2026
Language
C++
License
AGPL-3.0
$docker run \

What's new

1.3.522 May 2026

Bug-fix release: FP16 NEON builds now compile on AArch64 CPUs without the FP16FML extension, and backup upload was moved into the index manager to fix uploading backups.

Overview

Endee is a vector database server written in C++. It stores embeddings in named indexes and answers nearest-neighbour queries over an HTTP API on port 8080, with Python, TypeScript and Java client SDKs on top. The project pitches it as the retrieval layer for RAG pipelines, semantic and hybrid search, recommendations and long-term memory for AI agents, and the docs include integrations for LangChain, LlamaIndex and CrewAI.

Beyond plain dense-vector search, Endee keeps a sparse-vector engine (an inverted index of term weights, used for BM25-style keyword matching) alongside the dense one, so a single index can serve hybrid queries. Metadata filters run first: the filter design doc describes estimating each filter's result size, intersecting the cheapest ones first, and then either brute-forcing a small candidate set or passing a bitmap into a filtered HNSW search. Indexes can be stored at reduced precision such as INT8, and the server has backup and restore endpoints, startup sanity checks and optional token authentication.

The binary is compiled for a specific CPU instruction set — AVX2 or AVX512 on x86, NEON on Apple Silicon and ARM, SVE2 on ARMv9 servers — which is where the project says its speed comes from. Note what the repository is: its README says it tracks an earlier line of Endee and that the actively developed version has diverged; the code is published under the AGPLv3 for the community to fork, study and build on. The company also runs Endee Cloud, a managed serverless version documented separately as "v2".

What it does

  • Dense vector similarity search over named indexes, with a chosen distance space (such as cosine) and configurable precision (for example INT8)
  • Sparse vector search through an inverted index, so dense and BM25-style keyword retrieval can be combined for hybrid search
  • Payload filtering with operators such as $eq and $range, executed as a pre-filter before the vector search
  • Backup, restore and upload of index backups as tar archives through the HTTP API
  • CPU-targeted builds for AVX2, AVX512, NEON and SVE2, plus a prebuilt Docker image and a bundled web dashboard
  • Python, TypeScript and Java SDKs, with LangChain, LlamaIndex and CrewAI integrations in the docs

Getting started

The quickest path is the prebuilt Docker image; the repo also documents a Docker build from source (which it recommends), an install.sh script for Linux and macOS, and a manual CMake build. Docker is the only option on Windows.

Start the server with Docker

Data lands in an endee-data folder in the current directory. Add -e NDD_AUTH_TOKEN=your_token to require a token on every request. Then open http://localhost:8080 to see the Endee dashboard.

bashbash
docker run \
  --ulimit nofile=100000:100000 \
  -p 8080:8080 \
  -v ./endee-data:/data \
  --name endee-server \
  --restart unless-stopped \
  endeeio/endee-server:latest

Or build it for your CPU

On Linux or macOS the install script installs dependencies, compiles the binary and downloads the frontend. Use --avx2 for Intel/AMD, --neon for Apple Silicon, --avx512 for Xeon/EPYC servers or --sve2 for ARMv9.

bashbash
git clone https://github.com/endee-io/endee.git
cd endee
chmod +x ./install.sh ./run.sh
./install.sh --release --avx2
./run.sh

Check it is up

bashbash
curl http://localhost:8080/api/v1/health
curl http://localhost:8080/api/v1/index/list

Create an index and insert vectors from Python

Install the SDK with pip install endee (Python 3.8+). The vectors here are placeholders — use embeddings with the dimension you declared.

pythonpython
from endee import Endee, Precision

client = Endee()

client.create_index(
    name="my_index",
    dimension=384,
    space_type="cosine",
    precision=Precision.INT8
)

index = client.get_index(name="my_index")

index.upsert([
    {
        "id": "doc1",
        "vector": [...],
        "meta": {"title": "First Document"},
        "filter": {"category": "tech"}
    }
])

Run a filtered search

pythonpython
results = index.query(
    vector=[...],
    top_k=5,
    filter=[
        {"category": {"$eq": "tech"}},
        {"score": {"$range": [80, 100]}}
    ]
)

for item in results:
    print(f"ID: {item['id']}, Similarity: {item['similarity']:.3f}")

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Self-host the retrieval layer of a RAG app or chat assistant when you need vector search narrowed by metadata filters
  • Combine dense embeddings with sparse BM25 vectors for hybrid search over documents, products or support content
  • Give LangChain, LlamaIndex or CrewAI agents a long-term memory store they query mid-run
  • Study or fork a compact C++ vector database with documented filter, sparse-index and backup internals

Version history

Every verified update to Endee that AI/TLDR tracked, newest first — each links to our coverage and the official changeset.

  1. 2026-05-221.3.5

    Bug-fix release: FP16 NEON builds now compile on AArch64 CPUs without the FP16FML extension, and backup upload was moved into the index manager to fix uploading backups.

  2. 2026-04-17v1.3.4

    Lowered the minimum RAM requirement to 2 GB, raised the default vector cache from 15% to 50%, and bundled Web UI v1.6.1.

How Endee compares

Endee alongside other open-source vector databases tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Supabase★ 111kManaged Postgres backend whose Vector toolkit (pgvector) stores, indexes, and queries embeddings next to transactional data.
Redis Cloud★ 76.6kFully-managed Redis with built-in vector search, offering low-latency similarity and hybrid queries over any embeddings.
Milvus★ 46.3kA distributed vector database for storing and searching billions of embeddings at scale, with multiple index types and Kubernetes-native deployment.
FAISS★ 41kA library from Meta for efficient similarity search and clustering of dense vectors, with both exact and approximate indexes.
Qdrant★ 34.9kA Rust-based vector search engine that stores embeddings with rich payload filtering for semantic search and recommendation systems.
Chroma★ 29.4kA developer-focused vector database designed for quickly building retrieval and RAG features with a simple Python and JavaScript API.
pgvector★ 23.2kA PostgreSQL extension that adds a vector data type and similarity search so you can store and query embeddings inside an existing Postgres database.
Endee★ 1.3kA C++ vector database server with dense and sparse (BM25) search, payload filtering and backups, tuned per CPU instruction set