Overview
fastRAG is a Python research framework from Intel Labs for building efficient, optimized retrieval-augmented generation (RAG) pipelines. It is built on Haystack and Hugging Face, and every fastRAG component is fully Haystack compatible, so you assemble pipelines by mixing fastRAG's retrievers, rankers and generators with standard Haystack parts, either in Python code or in Haystack's YAML pipeline format.
Its focus is compute efficiency. The library ships RAG-oriented components such as ColBERT v2 retrieval with the PLAID indexing engine, Fusion-in-Decoder, and REPLUG, which retrieves several documents, runs each one alongside the prompt in parallel and ensembles the resulting token probabilities, so more documents can be used without being limited by the LLM's context window. It also includes int8-optimized embedders and optimized or sparse cross-encoder rankers.

On the generation side, fastRAG can run LLMs on Intel Xeon processors through Intel Extension for PyTorch, Optimum Intel, ONNX Runtime and OpenVINO, on Intel Gaudi accelerators through Optimum Habana, or on a llama.cpp backend. Intel has archived the repository: it no longer accepts patches and Intel provides no further maintenance, so teams that depend on it are pointed to maintaining their own fork.
What it does

- ColBERT v2 retrieval with the PLAID engine via PLAIDDocumentStore and ColBERTRetriever, plus an IndexUpdater for adding and removing documents in an existing index
- REPLUG ensembling inference (ReplugGenerator) and Fusion-in-Decoder for reading many retrieved documents at once
- Intel-optimized LLM backends: Gaudi accelerators, quantized models on ONNX Runtime or OpenVINO, and a llama.cpp backend
- Optimized int8 bi-encoder embedders and optimized or sparse cross-encoder rerankers
- Haystack-compatible components, pipelines defined in Python or in Haystack's YAML format, and example notebooks covering PLAID, REPLUG, Gaudi, OpenVINO and prompt compression
- Chainlit chat demos for a RAG pipeline chat and a multi-modal ReAct agent that picks between retrievers
Getting started
fastRAG needs Python 3.8+ and PyTorch 2.0+. Install it from PyPI (preferably in a fresh virtual environment), add the extras for the backends you need, then run the Chainlit demos from a clone of the repository.
Install fastRAG
Install the package from PyPI, or clone the repository and run pip install . for the source version.
pip install fastragAdd the extras you need
Optional extras pull in Intel-optimized backends, document stores and the ColBERT/PLAID engine.
pip install fastrag[intel] # Intel optimized backend [Optimum-intel, IPEX]
pip install fastrag[openvino] # Intel optimized backend using OpenVINO
pip install fastrag[qdrant] # Support for Qdrant store
pip install fastrag[colbert] # Support for ColBERT+PLAID; requires FAISS
pip install fastrag[faiss-cpu] # CPU-based Faiss libraryClone the repo for configs and demos
The pipeline configs, example notebooks and Chainlit demos live in the repository, so clone it and install from source.
git clone https://github.com/IntelLabs/fastRAG.git
cd fastRAG
pip install .Start a plain chat
config/regular_chat.yaml loads a Hugging Face LLM for text generation; run it with the no-RAG Chainlit app.
CONFIG=config/regular_chat.yaml chainlit run fastrag/ui/chainlit_no_rag.py
Chat with a RAG pipeline
Launch the RAG chat with a config that lists the retrieval tools the chat model may call, each backed by a Haystack YAML pipeline.
CONFIG=config/rag_pipeline_chat.yaml chainlit run fastrag/ui/chainlit_pipeline.pyTry the multi-modal agent
The multi-modal demo uses ReAct prompting so a LLaVA-based agent can decide which retriever to call. The examples folder has the matching notebook that builds it step by step.
CONFIG=config/visual_chat_agent.yaml chainlit run fastrag/ui/chainlit_multi_modal_agent.pyCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Research and compare efficient RAG techniques such as ColBERT/PLAID retrieval, Fusion-in-Decoder and REPLUG inside one Haystack pipeline
- Run RAG pipelines with quantized LLMs on Intel Xeon CPUs or on Intel Gaudi accelerators
- Drop optimized retrievers, rankers or generators into an existing Haystack pipeline without rewriting it
- Prototype a chat assistant or a multi-modal agent that retrieves documents and images to answer questions

How fastRAG compares
fastRAG alongside other open-source rag frameworks & platforms tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Dify | ★ 158k | An open-source platform with a visual workflow builder for creating LLM and RAG applications without writing much code. |
| graphify | ★ 123k | Turns a folder of code, docs, PDFs and images into a local knowledge graph with tree-sitter AST parsing and Leiden communities — queryable by agents over MCP, no vector store. |
| RAGFlow | ★ 91.6k | A RAG engine built around deep document understanding that turns complex files into a grounded, citation-backed question-answering layer. |
| Context7 | ★ 62.6k | Context7 pulls current, version-specific documentation and code examples for any library and feeds them into your LLM, available as a CLI skill or an MCP server. |
| Pathway | ★ 62.2k | A Python framework with a Rust streaming engine that keeps ETL, real-time analytics and RAG pipelines continuously up to date as source data changes. |
| LightRAG | ★ 40k | A graph-based RAG system that builds an entity-and-relationship knowledge graph for fast retrieval and easy incremental updates. |
| Quivr | ★ 39.6k | Quivr is an open-source RAG framework that ingests your documents and answers questions about them, working with any LLM and any file type. |
| fastRAG | ★ 1.8k | Intel Labs' research framework for efficient RAG pipelines, built as Haystack components and tuned for Intel hardware |