Overview
e2e is an end-to-end testing framework for web and mobile apps, built in TypeScript by TesterArmy and released under Apache-2.0. Instead of scripting every selector and wait, a test describes a goal in natural language — "upgrade the workspace to the Pro plan" — and an agent drives the app until it gets there. The same test then checks the result with ordinary locators and assertions, so agent steps and deterministic checks live side by side.
The part that keeps it cheap is the replay cache. When a later assertion verifies an agent step, e2e records the actions the agent took, and the next run replays them with no model calls until the app changes. Tests with no agent steps need no model at all. Model access is bring-your-own: a subscription such as ChatGPT or GitHub Copilot, an API key through providers like OpenRouter or Vercel AI Gateway, or a local model.
The project is split into packages: the `e2e` SDK, runner and CLI; a web engine that drives Chromium, Firefox and WebKit through Playwright; a mobile engine for iOS simulators and Android emulators; a GitHub reporter that posts results as a pull-request comment; adapters for hosted browsers and simulators; and decision-model executors for bounded actions and assertions. The README marks it as in active development on the way to 1.0, so APIs and config can still change between minor releases.
What it does
- Natural-language goals with `agent.act` and `agent.assert`, mixed with normal locators and `expect` assertions in one test
- Verified agent steps are cached and replayed with no model calls until the app changes
- One API across web (Chromium, Firefox, WebKit via Playwright) and mobile (iOS simulators, Android emulators)
- Bring your own model: ChatGPT or Copilot subscriptions, API keys via AI SDK providers, or local models
- `e2e mcp` sessions let coding agents such as Claude Code and Cursor drive a live app session
- GitHub reporter that posts test results as a pull-request comment
- Docs ship inside the npm package, so coding agents can read them offline in `node_modules/e2e/docs`
Getting started
Setup is one command; it asks for an engine and a model provider and writes a config plus an example test. Commands and code below come from the project's README and product page. Node.js 24.8+ (or 22.22.3+) is required, and Windows users run it inside WSL.

Initialise a project
`init` asks whether you test web or mobile and which model provider to use, then writes a config and an example test.
npx e2e initWrite a test with an agent goal
The agent reaches the goal; the locator assertion checks the outcome deterministically.
// tests/checkout.e2e.ts
import { test, expect } from 'e2e';
test('a member upgrades to Pro', async ({ app, agent, screen }) => {
await app.open('/settings/billing');
await agent.act('upgrade the workspace to the Pro plan');
await agent.assert('the invoice preview shows a prorated amount');
await expect(screen.getByRole('status')).toContainText('Pro');
});Run the suite
Later runs replay verified agent steps from the cache instead of calling the model.
npx e2e runTurn off telemetry if you want
The CLI sends anonymous usage data (commands, engines, failure points — no test or app content).
npx e2e telemetry disableCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Reach for it when brittle selector-based UI tests break on every redesign
- Reach for it to cover web and native mobile flows with one test API
- Reach for it when you want AI-driven tests in CI without paying for model calls on every run
- Reach for it to give a coding agent a live app session it can drive and check through MCP
How e2e compares
e2e alongside other open-source ai sdlc automation tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Multica | ★ 52.1k | Self-hostable workspace where coding agents are assigned issues like teammates: 26 agent CLIs, runtimes you own, a replayable execution log per run, and review gates before anything ships. |
| GPT-Pilot | ★ 33.7k | Autonomous AI developer that breaks down an app description into tasks, writes and runs code incrementally, and asks clarifying questions to produce a working production application. |
| oh-my-codex (OMX) | ★ 33.5k | A workflow layer for the OpenAI Codex CLI that adds agent teams, git-worktree isolation, hooks and HUDs behind a set of canonical plan, code-review and QA commands. |
| Vibe Kanban | ★ 28.3k | Vibe Kanban lets you plan tasks on a kanban board, run coding agents like Claude Code and Codex in isolated workspaces, then review their diffs and ship pull requests. |
| Beads | ★ 27.7k | A Dolt-backed dependency-graph issue tracker for coding agents: hash IDs avoid multi-agent collisions, `bd ready` surfaces unblocked work, and `bd remember` keeps durable project memory. |
| Archon | ★ 23.6k | A workflow engine for AI coding agents: describe plan, implement, validate, review and PR phases as YAML, and every run repeats them in its own git worktree. |
| Keploy | ★ 18.5k | An API testing tool that records real traffic with eBPF — including database and queue calls — and replays it as deterministic tests and data mocks, with no SDK to import and no code changes. |
| e2e | — | End-to-end tests for web and mobile apps where an agent reaches the goal and assertions check it |