█

AI/TLDR

EnvHarness

A plug-in layer that reshapes a frozen agent benchmark from the outside — without editing the environment or its verifiers

Benchmark HarnessesOpen source
Updated
20 Aug 2026
Language
Python
License
Apache-2.0
Coverage
1 story

What's new

20 Aug 2026

EnvHarness released with its paper and website: the Setup/Rule/Link component layer, the LLM designer loop, and drivers for ALFWorld, WebArena, SWE-bench Verified, OfficeQA and SpreadsheetBench under Apache-2.0.

Latest news

Overview

EnvHarness applies the agent-harness idea to the other side of the interaction. An agent harness makes a frozen model more capable by plugging in skills, memory and tools without touching its weights; EnvHarness wraps a frozen *environment* with plug-in components so it becomes dynamically controllable without touching the environment's internal code. The problem it targets is that interactive environments are expensive to build and, once built, behave identically no matter which agent meets them or how much that agent has improved — so they can neither target a particular agent's weaknesses nor keep teaching after its tasks are solved.

The layer is assembled from three components that operate strictly at the standard `reset` / `step` interface and stack freely. **Setup** reshapes the initial state. **Rule** reshapes the interaction — which actions are allowed, what they do, and what the agent observes. **Link** composes in another environment's tasks. Crucially, they change only what the agent observes, what it may do and where it starts; the goal predicate that decides success is left untouched, so a reshaped environment keeps the original benchmark's human-built verifiers. Because nothing reaches into environment-specific code, the same system works across domains, and an `EnvHarness` is itself an `ActionableEnv` wrapping another one, so layers compose arbitrarily.

Driving the layer is an LLM designer agent running a diagnostic loop: it reads the policy's trajectories to diagnose a specific weakness, writes components that reshape the environment to target it, tests the policy in the new environment, and revises until the environment can actually teach what the agent lacks. The designer emits real Python — a `_Rules(Rules)` subclass — compiled and executed in an isolated subprocess, so a bad mutation becomes a recorded trace rather than a dead run. The repository ships drivers for ALFWorld, WebArena, SWE-bench Verified, OfficeQA, SpreadsheetBench and a toy environment, plus an RL path that trains a policy with GRPO inside EnvHarness environments via verl-agent. The accompanying paper reports skills induced in EnvHarness environments beating both a no-skill baseline and skills induced in the original environments, with gains up to 9.0 points on held-out ALFWorld tasks and about 9.8% fewer interaction steps.

What it does

  • Three composable components — Setup (initial state), Rule (allowed actions, effects, observations) and Link (compose another environment's tasks)
  • Operates only at the `reset` / `step` interface: the benchmark's task set, dynamics and grading stay exactly as published
  • Goal predicates untouched, so reshaped environments keep the original human-built verifiers
  • An LLM designer agent that diagnoses a policy's weakness from trajectories and writes Python components to target it, revising in a loop
  • Generated components compiled and executed in an isolated subprocess — a bad mutation is a recorded trace, not a crashed run
  • Drivers for ALFWorld, WebArena, SWE-bench Verified, OfficeQA, SpreadsheetBench and a toy environment, plus GRPO training via verl-agent
  • Provider-agnostic model config: one model string (OpenAI, Gemini, or Claude on Vertex AI) selects the backend for every stage

Getting started

Each benchmark has its own environment setup and its own one-command driver, documented in that experiment's README. Configuration is one model string per role, and a `MODEL` environment variable overrides every stage of a run so it never ends up split across providers.

Clone the repository and pick a provider

Set the credentials for the family you want. OpenAI takes an API key, Gemini takes a key, and Claude runs on Vertex AI via application-default credentials.

bashbash
git clone https://github.com/google-research/envharness.git && cd envharness
export OPENAI_API_KEY="your-openai-api-key"
# or: export GEMINI_API_KEY="your-gemini-api-key"
# or: gcloud auth application-default login && export GOOGLE_CLOUD_PROJECT="your-project-id"

Preflight the benchmark you want

Open the README in that experiment folder first — it carries the environment setup, run commands and knobs for that benchmark — then run the environment check.

bashbash
python scripts/check_env.py alfworld

Run a smoke test, then the full protocol

Every experiment folder has the same shape: a smoke script that runs the same stages over fewer tasks, and a full reproduce driver.

bashbash
bash experiments/alfworld/reproduce_smoke.sh
python experiments/alfworld/reproduce.py

Pin one model across every stage

`MODEL` overrides whatever the YAML names for the corpus policy, harness agent, skill induction and evaluation. The same override is available directly on the Stage 1 runner.

bashbash
MODEL=openai/gpt-4.1 python experiments/swebench/reproduce.py
python scripts/run_harness.py --config <corpus.yaml> --model openai/gpt-4.1-mini

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Keep getting signal from a benchmark your agent has already saturated, without writing a new environment
  • Target a diagnosed weakness — reshape observations, allowed actions or the starting state around it and re-test
  • Generate training environments for RL or skill induction while keeping the original benchmark's verifiers as ground truth
  • Compose tasks from one benchmark into another through a single interface instead of per-benchmark glue code

How EnvHarness compares

EnvHarness alongside other open-source benchmark harnesses tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
LM Evaluation Harness★ 14.2kEleutherAI's framework for few-shot evaluation of language models across 60+ academic benchmarks, used as the backend for many leaderboards.
OpenCompass★ 7.5kAn LLM evaluation platform that runs models against 100+ datasets covering reasoning, knowledge, coding, and domain tasks, with leaderboards and multi-model support.
SWE-bench★ 6kA benchmark and containerized harness that tests whether language models can resolve real GitHub issues by generating patches that pass a repository's tests.
simple-evals★ 4.7kOpenAI's lightweight library for running standard zero-shot, chain-of-thought benchmarks like MMLU, MATH, and GPQA to measure model accuracy.
lmms-eval★ 4.4kAn evaluation suite for large multimodal models that runs image, video, and audio benchmarks across many tasks with a unified, reproducible interface.
AgentBench★ 3.8kA benchmark that evaluates LLMs as agents across diverse interactive environments such as operating systems, databases, web browsing, and games.
EvalScope★ 3.5kModelScope evaluation framework that runs standard benchmarks against local or OpenAI-compatible models, with agent-loop evaluation, inference stress testing and a comparison dashboard.
EnvHarness★ 640A plug-in layer that reshapes a frozen agent benchmark from the outside — without editing the environment or its verifiers