█

AI/TLDR

SWIRL

Federated search and RAG that queries your apps live instead of indexing them

Rerank, Search & HybridOpen core
Language
Python
License
Apache-2.0

Overview

SWIRL's Galaxy UI answering a cybersecurity policy query: an AI summary panel with an attributed source list, a sources sidebar showing counts per connector for Outlook, OneDrive, Google News and SWIRL docs, and ranked results below
One query, five live sources: the Galaxy UI shows the generated summary, every document it cited, and which connector each came fromSWIRL README ↗

SWIRL is an open-source search platform built on the opposite assumption to most RAG stacks: instead of copying documents into a vector store, it sends the user's query out to the systems that already hold them — SharePoint, Confluence, Google Drive, GitHub, Jira, mailboxes, databases and around a hundred other connectors — collects the results, ranks them together and generates an answer with citations back to the originals.

That design mainly buys two things. Permissions stay where they are enforced: each source is queried as the user, so nothing surfaces in SWIRL that the user could not already open. And there is no index to build or refresh, which removes the ETL pipeline, the embedding bill and the staleness window that come with a copied corpus. The trade is latency — every search is a live fan-out — and ranking quality, since SWIRL has to reconcile relevance scores from sources that all compute them differently.

The project is open core. The Apache-2.0 community edition ships the federation engine, the Galaxy web UI, the REST API, the connector framework and cosine-similarity re-ranking with an OpenAI key you supply for RAG. Three-pass ranking (BM25, E5 embeddings and a cross-encoder), canonical answers and the MCP server are reserved for the commercial edition. It is a Django application backed by SQLite or PostgreSQL.

What it does

  • Federated query across 100+ connectors, run live at search time rather than against a copied index
  • Source-side permission enforcement — each system is queried as the signed-in user
  • Galaxy web UI with an AI summary, per-source result columns and clickable citations
  • Real-time RAG over the retrieved documents using your own OpenAI key, with answers attributed to the source files
  • Re-ranking of heterogeneous result sets by cosine similarity (spaCy and NLTK) in the community edition
  • REST API with synchronous and asynchronous federation, plus an extensible connector and processor framework

Getting started

The fastest path is the published docker-compose file, which brings up the Django app, the Galaxy UI and a SQLite store. Note that this container setup does not persist data — use the documented install for anything beyond a trial.

Fetch the compose file

Docker must already be installed and running.

bashbash
curl https://raw.githubusercontent.com/swirlai/swirl-search/main/docker-compose.yaml -o docker-compose.yaml

Set an OpenAI key for RAG (optional)

Without a key SWIRL still federates and ranks; the generated answer is what needs the model.

bashbash
export MSAL_CB_PORT=8000
export MSAL_HOST=localhost
export OPENAI_API_KEY='<your-key>'

Start it

Pull the images and bring the stack up.

bashbash
docker-compose pull && docker-compose up

Open the Galaxy UI

Browse to http://localhost:8000/galaxy and sign in with the default credentials (admin / password), then connect your first sources.

bashbash
open http://localhost:8000/galaxy

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

VideoThe maintainers' short walkthrough of how federated ranking differs from a single web indexSWIRL AI ↗
  • Enterprise search across SaaS tools where copying documents into a vector store is blocked by policy or licensing
  • Answering questions over systems whose permissions change often, since access is resolved at query time
  • Standing up a working RAG answer box in an afternoon, before committing to an embedding pipeline
  • Adding a new internal source to search by writing a connector rather than an ingestion job

How SWIRL compares

SWIRL alongside other open-source rerank, search & hybrid tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Elasticsearch★ 78.2kDistributed search and analytics engine with a built-in vector database for dense/sparse embeddings and hybrid keyword-plus-semantic retrieval.
Meilisearch Cloud★ 59.5kManaged cloud for the Meilisearch engine, combining fast full-text search with hybrid, semantic, and multimodal vector search.
Typesense Cloud★ 26.6kManaged hosting for the Typesense search engine, offering typo-tolerant keyword search plus vector and semantic search via a simple API.
Tantivy★ 16.2kA fast full-text search engine library in Rust that provides BM25 keyword search for the lexical half of hybrid retrieval.
FlagEmbedding★ 12.2kBAAI's retrieval toolkit that provides the BGE embedding and cross-encoder reranker models used widely in RAG pipelines.
Vespa★ 7.1kA search and serving engine that natively combines vector, keyword (BM25), and structured search with built-in ranking for large-scale retrieval.
RAGatouille★ 4kA wrapper that makes it easy to train and use ColBERT late-interaction retrieval inside RAG pipelines.
SWIRL★ 3kFederated search and RAG that queries your apps live instead of indexing them