Overview
Guardian is a Python command-line framework that automates penetration testing by putting a team of LLM agents in charge of established security tools. A Planner decides what to do next, a Tool Selector picks and runs the scanner, an Analyst interprets the output and a Reporter writes up the findings. The project states plainly that it is built exclusively for authorized security testing and education: you must have explicit written permission before testing any system, and its MIT license carries additional terms saying the same.
Assessments run as YAML workflows. The workflow engine schedules steps as a DAG, runs independent steps in parallel, gates steps on earlier output with `when:` clauses, passes results between steps through sandboxed Jinja2 templates, and checkpoints each session so a run can pick up again with `--resume`. The README lists 50 integrated tool wrappers across network scanning, web reconnaissance, subdomain and DNS enumeration, vulnerability scanning, SSL/TLS testing, cloud and container audits, SAST and secrets, API fuzzing, Active Directory, Android and LLM red-teaming. The external tools are optional: Guardian works with whatever is installed and adapts to what is available.
It works with OpenAI, Anthropic Claude, Google Gemini, OpenRouter, Requesty, local Ollama and any OpenAI-compatible endpoint, and third-party providers and tools can plug in through Python entry points. Guardrails are part of the design: safe mode and a confirmation gate for intrusive tools are on by default, target scope is validated with private address ranges blacklisted, tool output is wrapped in untrusted-content delimiters as a prompt-injection defence, API keys are scrubbed from logs and reports, and every AI decision is written to an audit log.
What it does
- Multi-agent pipeline (Planner, Tool Selector, Analyst, Reporter) plus an optional red/blue/judge debate that triages ambiguous findings
- YAML workflow engine with parallel DAG steps, conditional `when:` clauses, Jinja2 parameter passing and `--resume` from checkpoints
- Wrappers for 50 security tools, including nmap, httpx, subfinder, nuclei, sqlmap, ffuf, trivy, semgrep and garak
- Evidence capture that links every finding to the tool execution, command and raw output that produced it
- Optional knowledge base over CVE, CWE and MITRE ATT&CK data to ground the analyst's references
- Reports in Markdown, HTML or JSON with CVSS v3.1 recomputation, and export to SARIF, DefectDojo or Slack
Getting started
Guardian needs Python 3.11 or higher and one AI provider key. Run it only against systems you own or have explicit written authorization to test. Security tools such as nmap, httpx or nuclei are optional extras that widen what the agents can do.
Clone and install
Install into a virtual environment. On Windows, activate with `.\venv\Scripts\activate`.
git clone https://github.com/zakirkun/guardian-cli.git
cd guardian-cli
python3 -m venv venv
source venv/bin/activate
pip install -e .Configure an AI provider
Set `ai.provider` and the model in `config/guardian.yaml`, or export the key for your provider as an environment variable. The same file holds `safe_mode`, `require_confirmation` and the scope blacklist.
export OPENAI_API_KEY="sk-your-key-here"
# or ANTHROPIC_API_KEY, GOOGLE_API_KEY, OPENROUTER_API_KEY, REQUESTY_API_KEYVerify the install
`models` shows which providers and models Guardian can reach.
python -m cli.main --help
python -m cli.main models
python -m cli.main workflow listRun a workflow against an authorized target
Pick a shipped workflow and a target in scope. `--provider` switches the model backend per run.
python -m cli.main workflow run --name web_pentest --target example.com --provider openai
# local model, no cloud
OLLAMA_HOST=http://localhost:11434 python -m cli.main workflow run --name recon --target scanme.nmap.org --provider ollamaGenerate and export the report
Each run is saved as a session. Render it as a report or push the findings to other systems.
python -m cli.main report --session 20260203_175905 --format html
python -m cli.main report --session 20260203_175905 --export sarifCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Automate the reconnaissance and scanning stages of an authorized web application or network penetration test
- Produce evidence-linked Markdown, HTML or SARIF reports that trace each finding back to the command that found it
- Run repeatable, version-controlled assessment workflows in a security lab or training environment
- Red-team an LLM application with the garak, PyRIT and prompt-fuzzing workflow
How Guardian compares
Guardian alongside other open-source security agents tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| PentAGI | ★ 25.2k | PentAGI is a self-hosted AI security platform that plans and runs penetration tests autonomously using a team of agents and 20+ built-in pentesting tools. |
| PentestGPT | ★ 15.7k | An open-source agent that uses large language models to run penetration tests and solve security challenges, either fully autonomously or with a human in the loop. |
| IDA Pro MCP | ★ 12.4k | An MCP server and IDA Pro plugin that exposes decompilation, cross-references, renaming and type editing to an LLM client, letting an agent read and annotate a binary inside your IDA database. |
| HexStrike AI | ★ 12.3k | An MCP server that gives an AI agent a single interface to 150+ installed security tools — Nmap, Nuclei, SQLMap, Ghidra, Hashcat and more — so it can drive reconnaissance, scanning and binary analysis itself. |
| CAI | ★ 9.8k | CAI (Cybersecurity AI) is an open-source Python framework for building AI agents that automate offensive and defensive security tasks like recon, vulnerability discovery, and exploitation. |
| AI-Infra-Guard | ★ 6.7k | Tencent Zhuque Lab's AI red teaming platform: scans agents, Agent Skills and MCP servers, checks AI infra against a CVE library, fingerprints API relays and runs jailbreak evaluations. |
| T3MP3ST | ★ 6.3k | A multi-agent offensive-security harness for authorised testing that drives an already-installed coding agent, or a local OpenAI-compatible model, through recon, exploitation and reporting from a localhost War Room or the CLI. |
| Guardian | ★ 1.9k | AI-powered penetration testing automation from the command line |