█

AI/TLDR

T3MP3ST

A multi-agent offensive-security harness that drives the coding agent you already run through recon, exploitation and reporting against authorised targets

Security AgentsOpen source
Language
TypeScript
License
AGPL-3.0
$npm install

Overview

T3MP3ST is a multi-agent offensive-security framework built around a harness rather than a model. It does not ship its own LLM or require a new API key: you connect the coding agent already installed on your machine — Claude Code, Codex, Hermes, OpenCode, Oh My Pi — or point it at a local OpenAI-compatible server such as Ollama, LM Studio or vLLM, and T3MP3ST supplies the tooling, orchestration and reporting around it. Tool calling is driven over text, so models without native function-calling still work.

Operators drive it from a browser "War Room" on localhost or from the CLI: describe an authorised target in plain English and the framework runs recon, exploitation and reporting as a coordinated agent workflow. Coverage differs sharply by domain and the README's status table is explicit about which is which — web-app black-box testing and hint-free CTF solving are marked stable, as is the coordinated-disclosure pipeline for OSS vulnerability hunting and multi-language white-box source analysis; smart-contract work is reproduction rather than novel discovery; and the cloud/IaC, mobile and binary domains are static-detection scaffolding whose dynamic exploitation is not yet benchmarked.

This is an offensive tool, licensed AGPL-3.0, and the project is emphatic that it is for authorised testing, research and education only — point it at systems you own or have explicit written permission to test. Its other distinguishing habit is reproducibility: the project states that every headline number in its README re-derives from data committed under `bench/`, and ships `npm run verify-claims` so a reader can recompute them instead of taking them on trust. It reports 90.1% pass@1 on XBOW's 104-challenge suite by that method; treat that as a self-reported figure you can check yourself.

What it does

  • Keyless by default — the coding agent already signed in on your machine is the reasoning backbone, with no second subscription or cloud tenant
  • Fully offline operation against Ollama, LM Studio, vLLM or llama.cpp, with text-driven tool calling so native function-calling is not required
  • Browser War Room on 127.0.0.1 plus a CLI, both driving the same recon → exploit → report workflow
  • Domain coverage marked by maturity in the README: web apps and CTF stable, OSS/robotics coordinated disclosure and white-box source analysis stable, cloud IaC, mobile and binary at static-detection scaffolding
  • Re-derivable benchmarks — `npm run verify-claims` recomputes every headline number from committed data under `bench/`
  • Docker deployment that binds the API to localhost only, plus documented library/SDK, HTTP API and MCP usage
  • Session-reuse and timeout guards, including an opt-in flag for reusing a Claude Code session that is off by default because a resumed session can retain ambient tool authority

Getting started

Before anything else: only point this at systems you own or have explicit, written permission to test. The fastest path is the keyless War Room, which takes a couple of minutes to stand up.

Start the War Room

Install dependencies and start the server, then open the UI on localhost.

bashbash
npm install
npm run server        # War Room → http://127.0.0.1:3333/ui/

Connect an agent

In the War Room open Settings and connect a local agent — Claude Code, Codex, Hermes, OpenCode or Oh My Pi. Then describe an authorised target to Op Admiral in plain English and launch. No API key is needed for this path.

Or bring a hosted key

Setting a provider key skips the connect step. Note that a hosted provider receives your prompts, context and generated output — review that provider's terms before pointing it at sensitive target data.

bashbash
export OPENROUTER_API_KEY=...   # or ANTHROPIC_API_KEY / OPENAI_API_KEY / VENICE_API_KEY
export XAI_API_KEY=...
export NOVITA_API_KEY=...

Or run it fully offline

Defaults to Ollama; any OpenAI-compatible server works. `npm run build` is only needed from a git clone.

bashbash
ollama serve && ollama pull llama3
export TEMPEST_LOCAL_BASE_URL=http://localhost:11434/api   # LM Studio: http://localhost:1234/v1
export TEMPEST_LOCAL_MODEL=llama3
npm run build
npx tempest                 # → "Change default provider" → local

Check the numbers yourself

The project's claim is that nothing in the README is a trust-me number. This recomputes them from the committed benchmark data.

bashbash
npm run verify-claims

Run it in Docker

The compose file binds the container to 127.0.0.1:3333 rather than exposing it to the network.

bashbash
cp .env.example .env
docker compose up -d
docker compose logs -f

curl http://localhost:3333/api/health

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Run an authorised black-box web-application assessment end to end, driven by a coding agent you already pay for
  • Practise on CTF challenges with a sandbox-jailed, hint-free agent loop rather than a hand-assembled tool chain
  • Hunt vulnerabilities in open-source projects through the coordinated-disclosure pipeline, with a refuter step to filter false positives
  • Keep an engagement entirely on-premises by pointing the harness at a local model, so target data never leaves the machine

How T3MP3ST compares

T3MP3ST alongside other open-source security agents tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
PentAGI★ 25.2kPentAGI is a self-hosted AI security platform that plans and runs penetration tests autonomously using a team of agents and 20+ built-in pentesting tools.
PentestGPT★ 15.7kAn open-source agent that uses large language models to run penetration tests and solve security challenges, either fully autonomously or with a human in the loop.
IDA Pro MCP★ 12.4kAn MCP server and IDA Pro plugin that exposes decompilation, cross-references, renaming and type editing to an LLM client, letting an agent read and annotate a binary inside your IDA database.
HexStrike AI★ 12.3kAn MCP server that gives an AI agent a single interface to 150+ installed security tools — Nmap, Nuclei, SQLMap, Ghidra, Hashcat and more — so it can drive reconnaissance, scanning and binary analysis itself.
CAI★ 9.8kCAI (Cybersecurity AI) is an open-source Python framework for building AI agents that automate offensive and defensive security tasks like recon, vulnerability discovery, and exploitation.
AI-Infra-Guard★ 6.7kTencent Zhuque Lab's AI red teaming platform: scans agents, Agent Skills and MCP servers, checks AI infra against a CVE library, fingerprints API relays and runs jailbreak evaluations.
T3MP3ST★ 6.3kA multi-agent offensive-security harness that drives the coding agent you already run through recon, exploitation and reporting against authorised targets
RedAmon★ 2.9kA Docker-deployed offensive-security platform for authorised testing that chains parallel recon, exploitation and post-exploitation into a Neo4j attack graph, then triages the findings and opens remediation pull requests on your repository.