AI/TLDR

Hugging Face · 2026-08-13 · major

ICML 2026 Open Reproductions — agents re-ran 2,226 papers, contested 496

Hugging Face ran a 19-day hackathon where 1,221 people pointed coding agents at papers accepted to ICML 2026. Agents judged 35,908 claims across 2,226 papers, and every attempt was published as a public logbook.

Thumbnail for the Hugging Face ICML 2026 Open Reproductions report

Hugging Face turned 1,221 volunteers and their coding agents loose on ICML 2026, then published every reproduction attempt.

Key specs

Papers with a claim verified1,103
Papers with a claim falsified or contested496

Quick facts

OrganizerHugging Face
Hackathon windowJuly 15 – August 2, 2026
Participants1,221 community members
Logbooks published6,816
Papers attempted2,226 (34% of ICML 2026)
Claims judged35,908
Fully reproduced266 papers

What is it?

ICML 2026 Open Reproductions is Hugging Face's report on a public experiment: point coding agents at the papers accepted to ICML 2026 and see which results hold up. Volunteers picked a paper, re-ran its experiments with an agent, and wrote down what they found. Every attempt is public, including the ones that failed.

How does it work?

Each accepted ICML 2026 paper was indexed with its abstract, and its core scientific claims were pulled out so they could be judged one at a time. Participants used Claude Code, Codex, Cursor, Pi or OpenResearch's orx to redo the work. Every run produced a Trackio logbook — a static Hugging Face Space holding the write-up, the code, the artifacts and, optionally, the full agent execution trace.

Why does it matter?

Reviewing ICML 2026 was a scale problem: 23,918 submissions and 6,352 accepted papers, roughly double the year before. Open Reproductions gives a reader an auditable trail per paper instead of trust alone. It also shows how far the checks disagree — 242 papers drew opposing verdicts, where one logbook verified a claim and another contested it.

Who is it for?

ML researchers, reviewers and anyone citing ICML work

Frequently asked questions

How many ICML 2026 papers failed to reproduce?
ICML 2026 Open Reproductions found 496 papers with at least one claim falsified or contested, and 49 papers where every claim was falsified. Hugging Face publishes each verdict next to the logbook that produced it, so a reader can check the reasoning rather than take the count on trust.
Which coding agents did participants use?
Participants in ICML 2026 Open Reproductions used Claude Code, Codex, Cursor, Pi and OpenResearch's orx to re-run paper experiments. Hugging Face did not require one single tool, so the results reflect a mix of agents and human judgment. A logbook can optionally include the full agent execution trace, which shows exactly which commands the agent ran.
Can I read the individual reproduction logbooks?
Yes. Every logbook from ICML 2026 Open Reproductions is public on the challenge Space at huggingface.co/spaces/ICML-2026-agent-repro/challenge, alongside a frozen dataset of the verdicts. Each Trackio logbook is a static Hugging Face Space containing the write-up, the code and the artifacts from that attempt, so anyone can inspect a claim without re-running it.
Did the attempts agree with each other?
Not always. ICML 2026 Open Reproductions recorded 242 papers where different attempts reached opposing verdicts on the same claim — one logbook verified it while another contested it. Hugging Face kept both instead of picking a winner, so the disagreement stays visible and both logbooks remain public on the challenge Space.

Try it

https://huggingface.co/spaces/ICML-2026-agent-repro/challenge

Sources

Tags

  • hugging-face
  • icml
  • reproducibility
  • open-science
  • coding-agents
  • Claude Code
  • Codex
  • Cursor
  • peer-review
  • research
  • evaluation

← All releases · Learn AI