OFICIAL Hugging Face Blog

What We Learned by Reproducing 2,200 papers from ICML

What happened
Based on Hugging Face Blog · Aug 13, 2026

Hugging Face’s ICML 2026 Open Reproductions challenge crowdsourced verification of 2,226 papers using coding agents, finding 51% reproducible, 23% with falsified claims, and highlighting the need for human oversight in AI research evaluation.

What We Learned by Reproducing 2,200 papers from ICML
Hugging Face Blog — Hugging Face
Key points
·
Questions about how reproducible AI research really is are older than the current AI wave.
·
ICML 2026 received 23,918 submissions and accepted 6,352 papers, roughly double the previous year, continuing an exponential trend that is at least partly driven by AI agents making it faster to run experiments and write them up.
·
Reviewers at most conferences are volunteers who may not have the time or expertise to fully review a paper.
·
Here is a review of one accepted ICML 2026 spotlight paper, in the reviewer's own words: Note that this paper got strong scores and a spotlight.
Key numbers
·
Participants received $20 in Hugging Face compute credits and launched 2,962 cloud jobs, verifying 3,978 individual claims through real experiments.
·
Of the examined papers, 51% had at least one claim independently verified, with 266 fully reproduced and 632 partially reproduced without falsification.
·
However, 23% had at least one claim falsified or contested, including 49 papers where all claims failed verification.

Hugging Face organized a 19-day hackathon from July 15 to August 2, 2026, where over 1,200 participants used coding agents to reproduce 2,226 papers from ICML 2026. Participants received $20 in Hugging Face compute credits and launched 2,962 cloud jobs, verifying 3,978 individual claims through real experiments. The effort addressed concerns about reproducibility amid a surge in submissions, driven partly by AI agents accelerating research workflows.

Of the examined papers, 51% had at least one claim independently verified, with 266 fully reproduced and 632 partially reproduced without falsification. However, 23% had at least one claim falsified or contested, including 49 papers where all claims failed verification. Notably, 242 papers saw opposing verdicts from different reproduction teams, underscoring the adversarial nature of reproducibility in AI research.

Several high-profile cases emerged, such as a paging algorithm’s robustness claim being refuted after proof errors were identified, and a transformer paper’s evaluation being skewed by padding tokens. Authors have already responded to findings, with corrections submitted to arXiv and confirmations of errors in multiple papers. The challenge demonstrated that while agents can scale verification, human judgment remains critical for nuanced evaluation.

The hackathon raised questions about the future role of human reviewers in AI research. While agents excel at parallel execution, they struggled with local loops, misinterpreted scale-dependent behavior, and occasionally built falsifications on flawed assumptions. Human oversight proved essential for steering experiments, interpreting results, and making perceptual judgments, as seen in a winning entry where a participant manually reviewed image generation quality despite numerical metrics suggesting stability.

Original source → Deals on Clipraptor.com →