Overview
Ponytail is an agent skill and plugin whose whole job is to stop a coding agent from over-building. Its persona is the senior developer who looks at your fifty lines, says nothing, and replaces them with one. In practice it installs a ruleset — plus a handful of slash commands — into whichever coding agent you already use, and that ruleset runs before the agent writes code.
The rule is a seven-rung ladder, and the agent stops at the first rung that holds: does this need to exist at all (YAGNI), is it already in the codebase, does the standard library do it, is there a native platform feature, is it in an installed dependency, can it be one line, and only then the minimum that works. The README is explicit that the ladder runs after the agent understands the problem, not instead of it — lazy about the solution, never about reading the code.
It is also explicit about what is never cut. Trust-boundary validation, data-loss handling, security and accessibility stay in, so "less code" means less unnecessary code rather than golfed code. The repository publishes its own benchmark to back the claim: twelve feature tickets against a real FastAPI + React template with a headless Claude Code session, scored on the resulting git diff. Against the same agent with no skill, it reports -54% lines of code, -22% tokens, -20% cost and -27% time at 100% on its adversarial safety tier — and it documents that an earlier single-shot benchmark overstated the effect, with the agentic numbers as the corrected version.
Installation is per-agent rather than per-project: the same repository ships a Claude Code and Codex plugin marketplace, a Copilot CLI plugin, a Gemini CLI / Antigravity extension, an OpenCode plugin entry, Hermes, Devin, Swival, Pi and OpenClaw paths, and an AGENTS.md fallback that agents such as CodeWhale and Qoder pick up with no setup at all. Levels (lite, full, ultra, off) switch how aggressively the ruleset is applied.
What it does
- A seven-rung decision ladder applied before code is written: YAGNI, reuse, stdlib, native platform feature, installed dependency, one line, then minimum that works
- Safety carve-outs that are never traded away — validation at trust boundaries, data-loss handling, security and accessibility
- Levels (lite / full / ultra / off) plus commands such as /ponytail-review, /ponytail-audit, /ponytail-debt and /ponytail-gain
- Installs into 20+ agents: Claude Code, Codex, Copilot CLI, Gemini CLI / Antigravity, OpenCode, Hermes, Devin, Swival, Pi, OpenClaw and more
- An AGENTS.md fallback so agents that auto-load repo instructions get the ruleset with zero configuration
- A published, reproducible benchmark suite (agentic and single-shot) with per-task tables and stated limitations
Getting started
Ponytail installs as a plugin or skill inside your agent, not as a dependency of your project. The Claude Code and Codex plugins run two small Node.js lifecycle hooks, so node must be on your PATH for always-on activation; without it the skills still work.
Install into Claude Code
Add the marketplace, then install the plugin. The README notes these must be sent as two separate prompts.
/plugin marketplace add DietrichGebert/ponytail
/plugin install ponytail@ponytailOr install into Codex / Copilot CLI
The same marketplace works from the Codex and Copilot CLIs; in Copilot CLI the commands are namespaced as /ponytail:ponytail.
codex plugin marketplace add DietrichGebert/ponytail
codex plugin add ponytail@ponytail
copilot plugin marketplace add DietrichGebert/ponytail
copilot plugin install ponytail@ponytailOr install into Gemini CLI, OpenCode or Hermes
Gemini CLI takes it as an extension, OpenCode as an npm plugin entry in opencode.json, Hermes through its plugin installer.
gemini extensions install https://github.com/DietrichGebert/ponytail
# opencode.json
# { "plugin": ["@dietrichgebert/ponytail"] }
hermes plugins install DietrichGebert/ponytail --enableUse it, and pick a level
Once installed the ruleset is injected every turn. Switch how hard it pushes with the level commands, or run a review pass over existing code.
/ponytail lite
/ponytail ultra
/ponytail-reviewReproduce the benchmark (optional)
The benchmarks directory holds both suites; the single-shot run is a promptfoo config.
npx promptfoo eval -c benchmarks/promptfooconfig.yamlCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Reach for it when your agent answers "add a date picker" with a new dependency, a wrapper component and a stylesheet
- Use it on an existing codebase via /ponytail-review or /ponytail-audit to find code that should not have been written
- Add it to a team's agent setup through AGENTS.md so every harness picks up the same restraint rules
- Pick it when agent token spend and latency matter — the project's own agentic benchmark reports lower cost and time alongside smaller diffs
How Ponytail compares
Ponytail alongside other open-source agent skills & plugins tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Superpowers | ★ 297k | A composable skills plugin that installs a spec-first, TDD, subagent-driven development methodology into Claude Code, Codex, Cursor, Gemini CLI and other coding harnesses. |
| Skills for Real Engineers | ★ 282k | Matt Pocock's everyday agent skills for coding agents, covering alignment grilling, planning, code review and research — small, composable and meant to be edited. |
| ECC | ★ 276k | An installable plugin that adds 68 agents, 286 skills, hooks, rules, memory and an agent-config security scanner to Claude Code, Codex and other coding harnesses. |
| Karpathy Coding Guidelines | ★ 218k | A Claude Code plugin and CLAUDE.md rule set built from Andrej Karpathy's observations on LLM coding pitfalls: think before coding, keep it simple, make surgical changes, and work to verifiable success criteria. |
| Ponytail | ★ 159k | Make your coding agent write the least code that actually works |
| Caveman | ★ 111k | A skill plus local proxy that shortens an AI coding agent's prose while leaving code, commands and paths untouched — the rule file cuts what the agent writes, the proxy compresses what it reads. |
| Addy’s Agent Skills | ★ 104k | A pack of production engineering skills for AI coding agents, with nine lifecycle slash commands — /spec, /plan, /build, /test, /review, /ship — installable into 70+ agents. |
| Taste Skill | ★ 94k | A collection of portable Agent Skills for frontend work: design direction, layout, typography, motion and density dials, plus image-generation skills that produce reference boards for a coding agent to implement. |