Overview
T3MP3ST is a multi-agent offensive-security framework built around a harness rather than a model. It does not ship its own LLM or require a new API key: you connect the coding agent already installed on your machine — Claude Code, Codex, Hermes, OpenCode, Oh My Pi — or point it at a local OpenAI-compatible server such as Ollama, LM Studio or vLLM, and T3MP3ST supplies the tooling, orchestration and reporting around it. Tool calling is driven over text, so models without native function-calling still work.
Operators drive it from a browser "War Room" on localhost or from the CLI: describe an authorised target in plain English and the framework runs recon, exploitation and reporting as a coordinated agent workflow. Coverage differs sharply by domain and the README's status table is explicit about which is which — web-app black-box testing and hint-free CTF solving are marked stable, as is the coordinated-disclosure pipeline for OSS vulnerability hunting and multi-language white-box source analysis; smart-contract work is reproduction rather than novel discovery; and the cloud/IaC, mobile and binary domains are static-detection scaffolding whose dynamic exploitation is not yet benchmarked.
This is an offensive tool, licensed AGPL-3.0, and the project is emphatic that it is for authorised testing, research and education only — point it at systems you own or have explicit written permission to test. Its other distinguishing habit is reproducibility: the project states that every headline number in its README re-derives from data committed under `bench/`, and ships `npm run verify-claims` so a reader can recompute them instead of taking them on trust. It reports 90.1% pass@1 on XBOW's 104-challenge suite by that method; treat that as a self-reported figure you can check yourself.
What it does
- Keyless by default — the coding agent already signed in on your machine is the reasoning backbone, with no second subscription or cloud tenant
- Fully offline operation against Ollama, LM Studio, vLLM or llama.cpp, with text-driven tool calling so native function-calling is not required
- Browser War Room on 127.0.0.1 plus a CLI, both driving the same recon → exploit → report workflow
- Domain coverage marked by maturity in the README: web apps and CTF stable, OSS/robotics coordinated disclosure and white-box source analysis stable, cloud IaC, mobile and binary at static-detection scaffolding
- Re-derivable benchmarks — `npm run verify-claims` recomputes every headline number from committed data under `bench/`
- Docker deployment that binds the API to localhost only, plus documented library/SDK, HTTP API and MCP usage
- Session-reuse and timeout guards, including an opt-in flag for reusing a Claude Code session that is off by default because a resumed session can retain ambient tool authority
Getting started
Before anything else: only point this at systems you own or have explicit, written permission to test. The fastest path is the keyless War Room, which takes a couple of minutes to stand up.
Start the War Room
Install dependencies and start the server, then open the UI on localhost.
npm install
npm run server # War Room → http://127.0.0.1:3333/ui/Connect an agent
In the War Room open Settings and connect a local agent — Claude Code, Codex, Hermes, OpenCode or Oh My Pi. Then describe an authorised target to Op Admiral in plain English and launch. No API key is needed for this path.
Or bring a hosted key
Setting a provider key skips the connect step. Note that a hosted provider receives your prompts, context and generated output — review that provider's terms before pointing it at sensitive target data.
export OPENROUTER_API_KEY=... # or ANTHROPIC_API_KEY / OPENAI_API_KEY / VENICE_API_KEY
export XAI_API_KEY=...
export NOVITA_API_KEY=...Or run it fully offline
Defaults to Ollama; any OpenAI-compatible server works. `npm run build` is only needed from a git clone.
ollama serve && ollama pull llama3
export TEMPEST_LOCAL_BASE_URL=http://localhost:11434/api # LM Studio: http://localhost:1234/v1
export TEMPEST_LOCAL_MODEL=llama3
npm run build
npx tempest # → "Change default provider" → localCheck the numbers yourself
The project's claim is that nothing in the README is a trust-me number. This recomputes them from the committed benchmark data.
npm run verify-claimsRun it in Docker
The compose file binds the container to 127.0.0.1:3333 rather than exposing it to the network.
cp .env.example .env
docker compose up -d
docker compose logs -f
curl http://localhost:3333/api/healthCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Run an authorised black-box web-application assessment end to end, driven by a coding agent you already pay for
- Practise on CTF challenges with a sandbox-jailed, hint-free agent loop rather than a hand-assembled tool chain
- Hunt vulnerabilities in open-source projects through the coordinated-disclosure pipeline, with a refuter step to filter false positives
- Keep an engagement entirely on-premises by pointing the harness at a local model, so target data never leaves the machine
How T3MP3ST compares
T3MP3ST alongside other open-source security agents tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| PentAGI | ★ 25.2k | PentAGI is a self-hosted AI security platform that plans and runs penetration tests autonomously using a team of agents and 20+ built-in pentesting tools. |
| PentestGPT | ★ 15.7k | An open-source agent that uses large language models to run penetration tests and solve security challenges, either fully autonomously or with a human in the loop. |
| IDA Pro MCP | ★ 12.4k | An MCP server and IDA Pro plugin that exposes decompilation, cross-references, renaming and type editing to an LLM client, letting an agent read and annotate a binary inside your IDA database. |
| HexStrike AI | ★ 12.3k | An MCP server that gives an AI agent a single interface to 150+ installed security tools — Nmap, Nuclei, SQLMap, Ghidra, Hashcat and more — so it can drive reconnaissance, scanning and binary analysis itself. |
| CAI | ★ 9.8k | CAI (Cybersecurity AI) is an open-source Python framework for building AI agents that automate offensive and defensive security tasks like recon, vulnerability discovery, and exploitation. |
| AI-Infra-Guard | ★ 6.7k | Tencent Zhuque Lab's AI red teaming platform: scans agents, Agent Skills and MCP servers, checks AI infra against a CVE library, fingerprints API relays and runs jailbreak evaluations. |
| T3MP3ST | ★ 6.3k | A multi-agent offensive-security harness that drives the coding agent you already run through recon, exploitation and reporting against authorised targets |
| RedAmon | ★ 2.9k | A Docker-deployed offensive-security platform for authorised testing that chains parallel recon, exploitation and post-exploitation into a Neo4j attack graph, then triages the findings and opens remediation pull requests on your repository. |