█

AI/TLDR

AutoFlow

PingCAP's graph-RAG knowledge base: crawl your docs site, build a knowledge graph in TiDB, and answer questions in a Perplexity-style search page or an embeddable widget

RAG Frameworks & PlatformsOpen source
Language
TypeScript
License
Apache-2.0

Overview

AutoFlow is an open-source knowledge-base tool from PingCAP, built on top of TiDB Vector, LlamaIndex and DSPy. Rather than shipping a library you assemble yourself, it is a deployable application: a crawler that walks a documentation or marketing site through its sitemap, an indexing pipeline that turns the result into both vector embeddings and a knowledge graph, and a front end that answers questions against them. PingCAP runs the same stack publicly at tidb.ai, which doubles as the project's live demo.

The "graph" in graph RAG is the part that distinguishes it from a plain vector store. Alongside chunk embeddings, AutoFlow extracts entities and the relationships between them into a graph held in TiDB, so an answer can follow a chain of related concepts instead of relying on whichever chunks happened to land nearest the query vector. TiDB does the work of holding chat history, vectors, JSON and analytics in one database, which is why the deployment is a single Docker Compose stack rather than a database plus a separate vector service.

The AutoFlow conversational search page at tidb.ai: a thread sidebar on the left, an 'Ask anything about TiDB' prompt box in the centre with suggested questions, and an admin section listing documents, data sources, import tasks and indexes.
The conversational search page, with the admin sections for documents, data sources and indexes in the sidebar.AutoFlow README ↗

Two front ends come with it. The first is a Perplexity-style conversational search page with threads, source citations and a selectable chat engine, backed by an admin area for documents, data sources, import tasks and indexes. The second is an embeddable JavaScript snippet: paste it into your own site and a conversational search window — typically bottom-right — answers product questions in place. The repository carries a warning that AutoFlow is still in early development; the maintainers' stated next step is to package it as a Python library (`autoflow-ai`) in addition to the application.

What it does

  • Graph RAG — entities and relationships extracted into a knowledge graph in TiDB, queried alongside vector similarity rather than instead of it
  • A built-in website crawler that walks official and documentation sites through their sitemap, so a knowledge base can be pointed at a URL rather than hand-fed files
  • A Perplexity-style conversational search page with threads, suggested questions, citations and a selectable chat engine
  • An embeddable JavaScript snippet that drops the same conversational search window into any site
  • An admin area for documents, data sources, import tasks and index tasks, so ingestion is observable rather than a black box
  • One database for everything — TiDB holds chat history, vectors, JSON and analytics, so deployment is a single Compose stack
  • Built on LlamaIndex for retrieval and DSPy for programming the prompts rather than hand-tuning them

Getting started

AutoFlow is deployed as an application, not installed as a package. The project documents Docker Compose as the way to run it, on a machine with at least 4 CPU cores and 8 GB of RAM.

Try the hosted instance first

PingCAP runs AutoFlow publicly at tidb.ai over the TiDB documentation. It is the fastest way to see what a finished knowledge base looks like — the threads, the citations and the graph-backed answers — before you deploy anything.

bashbash
open https://tidb.ai

Deploy with Docker Compose

The project's deployment guide at autoflow.tidb.ai/deploy-with-docker walks through the Compose stack. The two published images are tidbai/backend and tidbai/frontend; both carry semver tags on Docker Hub. Plan for 4 CPU cores and 8 GB of RAM.

bashbash
git clone https://github.com/pingcap/autoflow.git
cd autoflow
# follow https://autoflow.tidb.ai/deploy-with-docker for the
# .env values (TiDB connection string, LLM + embedding keys)
docker compose up -d

Point the crawler at your documentation

In the admin area, add a data source and give it the site you want indexed. The built-in crawler walks the site through its sitemap; import tasks and index tasks then show ingestion progress, so you can see which pages made it into the vector index and the knowledge graph.

Embed the search widget in your own site

Once the index is built, AutoFlow generates a JavaScript snippet for the conversational search window. Copy it into your site's HTML — it is usually placed at the bottom right — and visitors get in-page answers to product questions.

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Reach for it when a documentation site needs an answer box rather than a search box, and the answers must cite the page they came from
  • Reach for it when questions span several concepts at once, so a knowledge graph earns its keep over plain top-k vector search
  • Reach for it when you already run TiDB and would rather not add a separate vector database to the stack
  • Reach for it when you want a deployable RAG application to study end to end — crawler, ingestion, graph, chat — instead of assembling one from libraries

How AutoFlow compares

AutoFlow alongside other open-source rag frameworks & platforms tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Dify★ 158kAn open-source platform with a visual workflow builder for creating LLM and RAG applications without writing much code.
graphify★ 123kTurns a folder of code, docs, PDFs and images into a local knowledge graph with tree-sitter AST parsing and Leiden communities — queryable by agents over MCP, no vector store.
RAGFlow★ 91.6kA RAG engine built around deep document understanding that turns complex files into a grounded, citation-backed question-answering layer.
Context7★ 62.6kContext7 pulls current, version-specific documentation and code examples for any library and feeds them into your LLM, available as a CLI skill or an MCP server.
Pathway★ 62.2kA Python framework with a Rust streaming engine that keeps ETL, real-time analytics and RAG pipelines continuously up to date as source data changes.
LightRAG★ 40kA graph-based RAG system that builds an entity-and-relationship knowledge graph for fast retrieval and easy incremental updates.
Quivr★ 39.6kQuivr is an open-source RAG framework that ingests your documents and answers questions about them, working with any LLM and any file type.
AutoFlow★ 3kPingCAP's graph-RAG knowledge base: crawl your docs site, build a knowledge graph in TiDB, and answer questions in a Perplexity-style search page or an embeddable widget