█

AI/TLDR

fastRAG

Intel Labs' research framework for efficient RAG pipelines, built as Haystack components and tuned for Intel hardware

RAG Frameworks & PlatformsOpen source
Language
Python
License
Apache-2.0
$pip install fastrag

Overview

fastRAG is a Python research framework from Intel Labs for building efficient, optimized retrieval-augmented generation (RAG) pipelines. It is built on Haystack and Hugging Face, and every fastRAG component is fully Haystack compatible, so you assemble pipelines by mixing fastRAG's retrievers, rankers and generators with standard Haystack parts, either in Python code or in Haystack's YAML pipeline format.

Its focus is compute efficiency. The library ships RAG-oriented components such as ColBERT v2 retrieval with the PLAID indexing engine, Fusion-in-Decoder, and REPLUG, which retrieves several documents, runs each one alongside the prompt in parallel and ensembles the resulting token probabilities, so more documents can be used without being limited by the LLM's context window. It also includes int8-optimized embedders and optimized or sparse cross-encoder rankers.

Diagram of REPLUG: a retriever fetches several documents for the test context, each document is prepended to the input and run through a black-box LM, and the output distributions are ensembled into one prediction
REPLUG runs each retrieved document with the prompt in parallel and combines the token probabilities.fastRAG components overview ↗

On the generation side, fastRAG can run LLMs on Intel Xeon processors through Intel Extension for PyTorch, Optimum Intel, ONNX Runtime and OpenVINO, on Intel Gaudi accelerators through Optimum Habana, or on a llama.cpp backend. Intel has archived the repository: it no longer accepts patches and Intel provides no further maintenance, so teams that depend on it are pointed to maintaining their own fork.

What it does

Diagram of ColBERT late interaction: a query encoder and a document encoder produce per-token vectors, MaxSim operations match each query token to document tokens, and the results are summed into a score, with document encoding done offline
ColBERT keeps a vector per token and scores documents with MaxSim late interaction.fastRAG components overview ↗
  • ColBERT v2 retrieval with the PLAID engine via PLAIDDocumentStore and ColBERTRetriever, plus an IndexUpdater for adding and removing documents in an existing index
  • REPLUG ensembling inference (ReplugGenerator) and Fusion-in-Decoder for reading many retrieved documents at once
  • Intel-optimized LLM backends: Gaudi accelerators, quantized models on ONNX Runtime or OpenVINO, and a llama.cpp backend
  • Optimized int8 bi-encoder embedders and optimized or sparse cross-encoder rerankers
  • Haystack-compatible components, pipelines defined in Python or in Haystack's YAML format, and example notebooks covering PLAID, REPLUG, Gaudi, OpenVINO and prompt compression
  • Chainlit chat demos for a RAG pipeline chat and a multi-modal ReAct agent that picks between retrievers

Getting started

fastRAG needs Python 3.8+ and PyTorch 2.0+. Install it from PyPI (preferably in a fresh virtual environment), add the extras for the backends you need, then run the Chainlit demos from a clone of the repository.

Install fastRAG

Install the package from PyPI, or clone the repository and run pip install . for the source version.

bashbash
pip install fastrag

Add the extras you need

Optional extras pull in Intel-optimized backends, document stores and the ColBERT/PLAID engine.

bashbash
pip install fastrag[intel]      # Intel optimized backend [Optimum-intel, IPEX]
pip install fastrag[openvino]   # Intel optimized backend using OpenVINO
pip install fastrag[qdrant]     # Support for Qdrant store
pip install fastrag[colbert]    # Support for ColBERT+PLAID; requires FAISS
pip install fastrag[faiss-cpu]  # CPU-based Faiss library

Clone the repo for configs and demos

The pipeline configs, example notebooks and Chainlit demos live in the repository, so clone it and install from source.

bashbash
git clone https://github.com/IntelLabs/fastRAG.git
cd fastRAG
pip install .

Start a plain chat

config/regular_chat.yaml loads a Hugging Face LLM for text generation; run it with the no-RAG Chainlit app.

bashbash
CONFIG=config/regular_chat.yaml chainlit run fastrag/ui/chainlit_no_rag.py
The fastRAG chat demo in Chainlit: the chatbot shows a retrieved Forrest Gump poster, the agent describes the image and answers follow-up questions, and a retrieved text document about Jenny's last name is displayed
The Chainlit chat demo answering follow-up questions from a retrieved image and a retrieved document.fastRAG demo docs ↗

Chat with a RAG pipeline

Launch the RAG chat with a config that lists the retrieval tools the chat model may call, each backed by a Haystack YAML pipeline.

bashbash
CONFIG=config/rag_pipeline_chat.yaml chainlit run fastrag/ui/chainlit_pipeline.py

Try the multi-modal agent

The multi-modal demo uses ReAct prompting so a LLaVA-based agent can decide which retriever to call. The examples folder has the matching notebook that builds it step by step.

bashbash
CONFIG=config/visual_chat_agent.yaml chainlit run fastrag/ui/chainlit_multi_modal_agent.py

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Research and compare efficient RAG techniques such as ColBERT/PLAID retrieval, Fusion-in-Decoder and REPLUG inside one Haystack pipeline
  • Run RAG pipelines with quantized LLMs on Intel Xeon CPUs or on Intel Gaudi accelerators
  • Drop optimized retrievers, rankers or generators into an existing Haystack pipeline without rewriting it
  • Prototype a chat assistant or a multi-modal agent that retrieves documents and images to answer questions
The fastRAG multi-modal ReAct agent in Chainlit reasoning in steps, choosing an imageRetriever tool and then a finishAnswer tool before replying that the image of a man on a bench is from the movie Forrest Gump
The multi-modal agent demo picks a retriever tool, observes the result and composes a final answer.fastRAG demo docs ↗

How fastRAG compares

fastRAG alongside other open-source rag frameworks & platforms tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Dify★ 158kAn open-source platform with a visual workflow builder for creating LLM and RAG applications without writing much code.
graphify★ 123kTurns a folder of code, docs, PDFs and images into a local knowledge graph with tree-sitter AST parsing and Leiden communities — queryable by agents over MCP, no vector store.
RAGFlow★ 91.6kA RAG engine built around deep document understanding that turns complex files into a grounded, citation-backed question-answering layer.
Context7★ 62.6kContext7 pulls current, version-specific documentation and code examples for any library and feeds them into your LLM, available as a CLI skill or an MCP server.
Pathway★ 62.2kA Python framework with a Rust streaming engine that keeps ETL, real-time analytics and RAG pipelines continuously up to date as source data changes.
LightRAG★ 40kA graph-based RAG system that builds an entity-and-relationship knowledge graph for fast retrieval and easy incremental updates.
Quivr★ 39.6kQuivr is an open-source RAG framework that ingests your documents and answers questions about them, working with any LLM and any file type.
fastRAG★ 1.8kIntel Labs' research framework for efficient RAG pipelines, built as Haystack components and tuned for Intel hardware