Overview
PentestGPT is an open-source agent that uses large language models to perform penetration testing and capture-the-flag (CTF) challenges. It was published at USENIX Security 2024 and has since grown into an agentic tool that can work through a target on its own.
The current v1.0 release runs an autonomous iteration loop: the agent works continuously, keeps a context file with its progress, and restarts with that prior context when it hits a limit. The loop stops when a flag is captured or the maximum number of iterations is reached.
PentestGPT also keeps the classic human-in-the-loop experience from the original paper as a modernized legacy mode. That mode uses three cooperating LLM sessions for reasoning, generation, and parsing, and lets you drive the session interactively while it maintains a Pentesting Task Tree.
What it does
- Autonomous agent pipeline that runs penetration tests and CTFs with little or no human input
- Iteration loop with a saved context file, so the agent can restart with prior progress and stop on flag capture or a max-iteration limit
- Live walkthrough and real-time feedback that show each step as the agent works through a challenge
- Multi-category coverage for web, crypto, reversing, forensics, PWN, and privilege escalation tasks
- Interactive multi-LLM legacy mode supporting OpenAI, Anthropic, Google Gemini, DeepSeek, xAI Grok, Qwen, Moonshot Kimi, and local Ollama models
- Session persistence to save and resume penetration testing runs
Getting started
PentestGPT needs Python 3.12+, the uv package manager, and an installed and authenticated Claude Code CLI for the autonomous agent. Clone the repository, install dependencies, then run the agent against a target.
Clone and install
Clone the repository and install dependencies with the provided Make target, which runs uv sync under the hood.
git clone https://github.com/GreyDGL/PentestGPT.git
cd PentestGPT
make installRun the autonomous agent against a target
Point PentestGPT at a target host. You can add challenge context with an instruction and cap the number of iterations.
pentestgpt --target 10.10.11.234
pentestgpt --target 10.10.11.50 --instruction "WordPress site, focus on plugin vulnerabilities"
pentestgpt --target 10.10.11.234 --max-iterations 5Use the interactive multi-LLM legacy mode
Set an API key for any provider you want, then run the modernized legacy mode. It can auto-pick the best available models or let you choose a model per session, and it can list or smoke-test every supported model.
pentestgpt-legacy
pentestgpt-legacy --reasoning-model claude-opus-4-8 --parsing-model gemini-3.5-flash
pentestgpt-legacy --list-modelsCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Running autonomous penetration tests against an authorized target to find and exploit vulnerabilities
- Solving capture-the-flag challenges across categories such as web, crypto, reversing, forensics, PWN, and privilege escalation
- Driving a human-in-the-loop pentesting session that maintains a task tree and suggests next steps interactively
- Comparing how different LLM providers and models perform on security reasoning tasks via the legacy mode's model registry and smoke test
How PentestGPT compares
PentestGPT alongside other open-source security agents tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| PentAGI | ★ 25.2k | PentAGI is a self-hosted AI security platform that plans and runs penetration tests autonomously using a team of agents and 20+ built-in pentesting tools. |
| PentestGPT | ★ 15.7k | AI-powered autonomous penetration testing agent, published at USENIX Security 2024 |
| IDA Pro MCP | ★ 12.4k | An MCP server and IDA Pro plugin that exposes decompilation, cross-references, renaming and type editing to an LLM client, letting an agent read and annotate a binary inside your IDA database. |
| HexStrike AI | ★ 12.3k | An MCP server that gives an AI agent a single interface to 150+ installed security tools — Nmap, Nuclei, SQLMap, Ghidra, Hashcat and more — so it can drive reconnaissance, scanning and binary analysis itself. |
| CAI | ★ 9.8k | CAI (Cybersecurity AI) is an open-source Python framework for building AI agents that automate offensive and defensive security tasks like recon, vulnerability discovery, and exploitation. |
| AI-Infra-Guard | ★ 6.7k | Tencent Zhuque Lab's AI red teaming platform: scans agents, Agent Skills and MCP servers, checks AI infra against a CVE library, fingerprints API relays and runs jailbreak evaluations. |
| T3MP3ST | ★ 6.3k | A multi-agent offensive-security harness for authorised testing that drives an already-installed coding agent, or a local OpenAI-compatible model, through recon, exploitation and reporting from a localhost War Room or the CLI. |
| RedAmon | ★ 2.9k | A Docker-deployed offensive-security platform for authorised testing that chains parallel recon, exploitation and post-exploitation into a Neo4j attack graph, then triages the findings and opens remediation pull requests on your repository. |
