█

AI/TLDR

PentAGI

Autonomous AI agent that runs penetration tests in an isolated Docker sandbox

Security AgentsOpen source
Language
Go
License
MIT

Overview

PentAGI (Penetration testing Artificial General Intelligence) is an open-source tool for automated security testing. It is built for security professionals, researchers, and enthusiasts who want a powerful and flexible way to run penetration tests with the help of AI.

The platform runs as a self-hosted stack and is fully autonomous: an AI-driven agent decides which steps to take and then carries them out. Every operation runs inside a sandboxed Docker environment for isolation, and a team of specialized agents handles research, development, and infrastructure tasks. PentAGI works with 10+ LLM providers, including OpenAI, Anthropic, Google Gemini, AWS Bedrock, and local models through Ollama.

What it does

  • Fully autonomous AI agent that determines and executes penetration testing steps with optional execution monitoring and task planning
  • All operations run in a sandboxed, isolated Docker environment for safe execution
  • Built-in suite of 20+ professional security tools, including nmap, metasploit, and sqlmap
  • Team of specialized agents for research, development, and infrastructure tasks, with smart memory and a Graphiti knowledge graph for context
  • Works with 10+ LLM providers (OpenAI, Anthropic, Gemini, AWS Bedrock, Ollama, and more) plus aggregators like OpenRouter and DeepInfra
  • Detailed vulnerability reports with exploitation guides, plus REST and GraphQL APIs with Bearer token authentication

Getting started

PentAGI is deployed with Docker Compose. You need Docker and Docker Compose, at least 2 vCPU, 4GB RAM, and 20GB of free disk space. An interactive installer is recommended, but you can also set it up manually.

Create a working directory

Make a folder for the PentAGI stack and move into it.

bashbash
mkdir pentagi && cd pentagi

Download and fill in the environment file

Copy the example .env file, then add at least one LLM provider key (such as OPEN_AI_KEY, ANTHROPIC_API_KEY, or GEMINI_API_KEY) and update the security-related variables.

bashbash
curl -o .env https://raw.githubusercontent.com/vxcontrol/pentagi/master/.env.example

Run the PentAGI stack

Download the docker-compose file and start all services in the background.

bashbash
curl -O https://raw.githubusercontent.com/vxcontrol/pentagi/master/docker-compose.yml
docker compose up -d

Open the web UI

Visit https://localhost:8443 to reach the PentAGI web interface. The default login is admin@pentagi.com / admin (change it right away).

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Run autonomous penetration tests against a target system, letting the AI agent plan and execute the steps in an isolated sandbox
  • Generate detailed vulnerability reports with exploitation guides for security assessments
  • Automate and integrate security testing into other systems through the REST and GraphQL APIs
  • Self-host a private, controlled AI pentesting platform that keeps all data and execution on your own infrastructure

How PentAGI compares

PentAGI alongside other open-source security agents tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
PentAGI★ 25.2kAutonomous AI agent that runs penetration tests in an isolated Docker sandbox
PentestGPT★ 15.7kAn open-source agent that uses large language models to run penetration tests and solve security challenges, either fully autonomously or with a human in the loop.
IDA Pro MCP★ 12.4kAn MCP server and IDA Pro plugin that exposes decompilation, cross-references, renaming and type editing to an LLM client, letting an agent read and annotate a binary inside your IDA database.
HexStrike AI★ 12.3kAn MCP server that gives an AI agent a single interface to 150+ installed security tools — Nmap, Nuclei, SQLMap, Ghidra, Hashcat and more — so it can drive reconnaissance, scanning and binary analysis itself.
CAI★ 9.8kCAI (Cybersecurity AI) is an open-source Python framework for building AI agents that automate offensive and defensive security tasks like recon, vulnerability discovery, and exploitation.
AI-Infra-Guard★ 6.7kTencent Zhuque Lab's AI red teaming platform: scans agents, Agent Skills and MCP servers, checks AI infra against a CVE library, fingerprints API relays and runs jailbreak evaluations.
T3MP3ST★ 6.3kA multi-agent offensive-security harness for authorised testing that drives an already-installed coding agent, or a local OpenAI-compatible model, through recon, exploitation and reporting from a localhost War Room or the CLI.
RedAmon★ 2.9kA Docker-deployed offensive-security platform for authorised testing that chains parallel recon, exploitation and post-exploitation into a Neo4j attack graph, then triages the findings and opens remediation pull requests on your repository.