Overview

SWIRL is an open-source search platform built on the opposite assumption to most RAG stacks: instead of copying documents into a vector store, it sends the user's query out to the systems that already hold them — SharePoint, Confluence, Google Drive, GitHub, Jira, mailboxes, databases and around a hundred other connectors — collects the results, ranks them together and generates an answer with citations back to the originals.
That design mainly buys two things. Permissions stay where they are enforced: each source is queried as the user, so nothing surfaces in SWIRL that the user could not already open. And there is no index to build or refresh, which removes the ETL pipeline, the embedding bill and the staleness window that come with a copied corpus. The trade is latency — every search is a live fan-out — and ranking quality, since SWIRL has to reconcile relevance scores from sources that all compute them differently.
The project is open core. The Apache-2.0 community edition ships the federation engine, the Galaxy web UI, the REST API, the connector framework and cosine-similarity re-ranking with an OpenAI key you supply for RAG. Three-pass ranking (BM25, E5 embeddings and a cross-encoder), canonical answers and the MCP server are reserved for the commercial edition. It is a Django application backed by SQLite or PostgreSQL.
What it does
- Federated query across 100+ connectors, run live at search time rather than against a copied index
- Source-side permission enforcement — each system is queried as the signed-in user
- Galaxy web UI with an AI summary, per-source result columns and clickable citations
- Real-time RAG over the retrieved documents using your own OpenAI key, with answers attributed to the source files
- Re-ranking of heterogeneous result sets by cosine similarity (spaCy and NLTK) in the community edition
- REST API with synchronous and asynchronous federation, plus an extensible connector and processor framework
Getting started
The fastest path is the published docker-compose file, which brings up the Django app, the Galaxy UI and a SQLite store. Note that this container setup does not persist data — use the documented install for anything beyond a trial.
Fetch the compose file
Docker must already be installed and running.
curl https://raw.githubusercontent.com/swirlai/swirl-search/main/docker-compose.yaml -o docker-compose.yamlSet an OpenAI key for RAG (optional)
Without a key SWIRL still federates and ranks; the generated answer is what needs the model.
export MSAL_CB_PORT=8000
export MSAL_HOST=localhost
export OPENAI_API_KEY='<your-key>'Start it
Pull the images and bring the stack up.
docker-compose pull && docker-compose upOpen the Galaxy UI
Browse to http://localhost:8000/galaxy and sign in with the default credentials (admin / password), then connect your first sources.
open http://localhost:8000/galaxyCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it

- Enterprise search across SaaS tools where copying documents into a vector store is blocked by policy or licensing
- Answering questions over systems whose permissions change often, since access is resolved at query time
- Standing up a working RAG answer box in an afternoon, before committing to an embedding pipeline
- Adding a new internal source to search by writing a connector rather than an ingestion job
How SWIRL compares
SWIRL alongside other open-source rerank, search & hybrid tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Elasticsearch | ★ 78.2k | Distributed search and analytics engine with a built-in vector database for dense/sparse embeddings and hybrid keyword-plus-semantic retrieval. |
| Meilisearch Cloud | ★ 59.5k | Managed cloud for the Meilisearch engine, combining fast full-text search with hybrid, semantic, and multimodal vector search. |
| Typesense Cloud | ★ 26.6k | Managed hosting for the Typesense search engine, offering typo-tolerant keyword search plus vector and semantic search via a simple API. |
| Tantivy | ★ 16.2k | A fast full-text search engine library in Rust that provides BM25 keyword search for the lexical half of hybrid retrieval. |
| FlagEmbedding | ★ 12.2k | BAAI's retrieval toolkit that provides the BGE embedding and cross-encoder reranker models used widely in RAG pipelines. |
| Vespa | ★ 7.1k | A search and serving engine that natively combines vector, keyword (BM25), and structured search with built-in ranking for large-scale retrieval. |
| RAGatouille | ★ 4k | A wrapper that makes it easy to train and use ColBERT late-interaction retrieval inside RAG pipelines. |
| SWIRL | ★ 3k | Federated search and RAG that queries your apps live instead of indexing them |