AI/TLDR

Quesma · 2026-09-11 · notable

RTK does not cut AI coding costs — Quesma's Terminal-Bench 2.1 run

Quesma benchmarked RTK, which compresses terminal output before a coding agent reads it, over 1,740 Terminal-Bench 2.1 attempts. RTK ran 1% more expensive with Fable 5.0 and 17% more expensive with DeepSeek V4 Pro 0813.

Quesma chart showing RTK's cost and pass-rate verdict on Terminal-Bench 2.1

Quesma's benchmark finds that RTK's reported token savings do not show up as lower cost on Terminal-Bench 2.1.

What is it?

RTK filters and compresses terminal output before an AI coding agent reads it, rewriting git, test, package and file commands so they return terser versions while keeping the essential information. Bartosz Kotrys and Jacek Migdal at Quesma set out to check whether those saved tokens show up on the bill, using Terminal-Bench 2.1 as the test.

How does it work?

The setup paired Claude Code on Fable 5.0 with OpenCode on DeepSeek V4 Pro 0813 through OpenRouter: 85 Fable tasks and 89 DeepSeek tasks, each run five times without RTK and five times with it, for 1,740 attempts and over $1,500 of spend. Quesma reports task-level averages rather than aggregate totals, because the savings turned out to sit in a single task.

Why does it matter?

Quesma's numbers run the other way from RTK's pitch: 1% more expensive on Fable and 17% more on DeepSeek, with pass rates moving from 84% to 83% and from 71% to 69%. The authors conclude that compressing terminal output changes how the agent works through a task, and the extra turns cancel out the tokens saved. They do not recommend RTK as a general cost-saving tool.

Who is it for?

teams paying per token for coding agents

Sources · 4 outlets

Tags

  • rtk
  • coding-agents
  • benchmark
  • terminal-bench
  • token-costs
  • cost-analysis
  • claude-code
  • opencode
  • quesma

← All releases · Learn AI