Overview
GPT-RAG is a solution accelerator from Microsoft's Azure organisation for building enterprise agentic RAG assistants on Microsoft Foundry. Instead of a library you import, it is a set of architecture templates and deployment assets: you run the Azure Developer CLI against the template and get a working chat-over-your-data application provisioned in your own subscription, with the agents answering questions grounded in your enterprise documents and data.
The runtime is split into services: an orchestrator that runs the agentic workflow on the Microsoft Agent Framework, a web UI with streaming chat and custom themes, a data ingestion service that extracts, chunks and indexes content, and an optional MCP server for hosting tools and business logic. Foundry IQ is the default retrieval backend, a single knowledge-base endpoint that fans out to sources such as Blob containers, existing Azure AI Search indexes, SharePoint, OneLake, Fabric, Microsoft 365 through Work IQ, Bing web grounding and MCP servers, with permission trimming. Querying an Azure AI Search index directly stays supported as a rollback path.

The project is built around a Zero-Trust architecture: services run inside a controlled, isolated network, communication follows least-privilege principles, and the whole environment is defined as infrastructure as code. Deployments scale from a basic setup without network isolation, suited to demos, to a fully network-isolated one provisioned in phases through a jumpbox, or an integrated mode that plugs into an existing enterprise landing zone and VNet. The project also covers Responsible AI, end-to-end observability and governance with an audit trail. It is MIT-licensed.
What it does
- Agentic orchestration on the Microsoft Agent Framework, with scenarios such as NL2SQL query generation and tool integration through MCP servers
- Foundry IQ retrieval by default: one knowledge base over Blob, Azure AI Search, SharePoint, OneLake, Fabric, Work IQ, web and MCP sources, with permission trimming
- Azure AI Search direct retrieval kept as a supported rollback path, selected with RETRIEVAL_BACKEND
- Zero-Trust, network-isolated deployment option with private endpoints, a jumpbox workflow and preflight checks, all defined as infrastructure as code
- Separate services for orchestration, web UI, data ingestion and an optional MCP server, each deployable on its own
- Chat runtime topologies: a hosted orchestrator served by Foundry Agent Service, or the classic Container Apps orchestrator
Getting started
You need an Azure subscription where you hold the Contributor and User Access Admin roles and have agreed to the Responsible AI terms for Azure AI Services, plus the Azure Developer CLI, Git, Python 3.12 and, on Windows, PowerShell 7+. The steps below are the Deployment Guide's basic deployment without network isolation.
Initialise the template
Pull the GPT-RAG template into a new azd project.
azd init -t azure/gpt-ragSign in to Azure
Authenticate both the Azure CLI and the Azure Developer CLI.
az login
azd auth login
Provision and deploy a basic environment
Turn off network isolation for a demo environment, create the Azure resources, then deploy the services.
azd env set NETWORK_ISOLATION false
azd provision
azd deployGo network-isolated for production
For a Zero-Trust deployment the guide splits the work in three phases: run azd provision from your workstation, then run scripts/postProvision.ps1 and azd deploy from a private host with VNet or VPN access, such as the jumpbox. To plug into an existing landing zone, switch the deployment mode and point it at your VNet.
azd env set DEPLOYMENT_MODE ailz-integrated
azd env set USE_EXISTING_VNET true
azd env set EXISTING_VNET_RESOURCE_ID "/subscriptions/..."Commands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Reach for it when an enterprise on Azure needs a chat assistant grounded in internal documents and data, and wants a reference deployment rather than building one from parts
- Reach for it when security review demands private endpoints, network isolation and least-privilege access from day one
- Reach for it when answers must draw on several Microsoft sources at once, such as SharePoint, OneLake, Fabric and Blob storage, while respecting user permissions
- Reach for it when agents need to go beyond retrieval, for example generating SQL over business data or calling tools through MCP servers
How GPT-RAG compares
GPT-RAG alongside other open-source rag frameworks & platforms tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Dify | ★ 158k | An open-source platform with a visual workflow builder for creating LLM and RAG applications without writing much code. |
| graphify | ★ 123k | Turns a folder of code, docs, PDFs and images into a local knowledge graph with tree-sitter AST parsing and Leiden communities — queryable by agents over MCP, no vector store. |
| RAGFlow | ★ 91.6k | A RAG engine built around deep document understanding that turns complex files into a grounded, citation-backed question-answering layer. |
| Context7 | ★ 62.6k | Context7 pulls current, version-specific documentation and code examples for any library and feeds them into your LLM, available as a CLI skill or an MCP server. |
| Pathway | ★ 62.2k | A Python framework with a Rust streaming engine that keeps ETL, real-time analytics and RAG pipelines continuously up to date as source data changes. |
| LightRAG | ★ 40k | A graph-based RAG system that builds an entity-and-relationship knowledge graph for fast retrieval and easy incremental updates. |
| Quivr | ★ 39.6k | Quivr is an open-source RAG framework that ingests your documents and answers questions about them, working with any LLM and any file type. |
| GPT-RAG | ★ 1.2k | Microsoft's deployable accelerator for enterprise agentic RAG on Azure, with Zero-Trust networking and infrastructure as code |