← Back to all posts
News

Harvard-MIT LLM Workflow Flags Discrepancies in 3,460 of 4,452 Top Economics Papers

September 29, 2026 · 15:06 UTC · News
Harvard-MIT LLM Workflow Flags Discrepancies in 3,460 of 4,452 Top Economics Papers

TL;DR

Economists Matthew Schwartz (Harvard), Isaiah Andrews (MIT) and Jesse Shapiro (Harvard) pointed a multi-agent LLM workflow at 4,452 published replication packages from the five top general-interest economics journals and asked it to rerun everything. Their NBER working paper w35782 reports that the workflow flagged and verified discrepancies in 3,460 articles or their appendices, and fully reproduced every attempted calculation in only 11.2% of packages. The same pipeline sped up a calculation by more than 10x in 496 articles and proposed a new, assumption-free extension in 923. The caveats matter: about a third of the numeric mismatches are last-digit rounding, the paper is not peer reviewed, one author contracted for Anthropic, and the code is not public yet.


What they did

Five journals (Econometrica, the Quarterly Journal of Economics, the American Economic Review, the Journal of Political Economy and the Review of Economic Studies) have required authors to deposit data and code for years. The team downloaded every available replication package for articles published in 2000 or later, plus the articles and online appendices. That is 4,452 packages, which the paper says is one to two orders of magnitude larger than earlier reproducibility studies that covered "dozens or hundreds" of articles.

The workflow does three jobs in sequence. First it reproduces: it reads the documentation, skips anything declared data-gated (confidential data), reruns every other calculation and logs any value that differs from the printed one. Then it improves: it profiles slow code and tries either a faster implementation of the same algorithm or a better algorithm. Finally it extends: it asks whether the paper's own assumptions already support a further result the authors did not report.

Under the hood, the team first used a multi-agent LLM setup to build an open-source library that renders the original code (often written in proprietary statistical software) in a form the model can inspect and modify. That is the clever part. Instead of asking a model to squint at someone else's Stata logs, they gave it a transparent copy of the machinery.

The quality gates

A pipeline that yells "error" at 78% of the top of a field would be worthless if it hallucinated its findings, so the discrepancy step has two filters. Every flagged discrepancy goes to a second, adversarial LLM agent whose explicit job is to overturn it and defend the original paper. Where the team could get licensed access to the original software, they also reran the calculation there. Only discrepancies that survive both checks are recorded.

Rerun packagein open library Compare withprinted values Adversarial agenttries to overturn Rerun in orig.software only discrepancies that survive every step are recorded
The reproduction pipeline puts a second agent on the defense's side before anything counts as a discrepancy.

Think of it as a prosecutor and a defense attorney who are both the same model with different job descriptions. The paper does not report how often the defense wins, which would be a useful number for anyone copying the design.

The headline numbers

Of the 4,452 packages, 12.4% declared all calculations data-gated, so the workflow could only compare code and logs against the text. For the other 87.6%, it attempted a full reproduction, and in 87.3% of those it flagged and verified a discrepancy. Only 11.2% of all audited packages reproduced every attempted calculation at the precision the article printed.

outcome, % of all 4,452 audited packages Fully reproduced11.2% Not attempted12.4% 1st miss: abstract3.1% 1st miss: intro12.9% 1st miss: body54.0% 1st miss: appendix6.4% not attempted = all calculations declared data-gated
Most packages fail somewhere, and for more than half the first mismatch sits in the body of the paper.

The workflow catalogs each article by where the earliest fix would be needed, reading from the abstract down. In 3.1% of all audited packages the first problem sits in the abstract itself, and in 12.9% it sits in the introduction. Those are the parts most readers never get past.

How bad is "discrepancy"?

This is where you should resist the lazy headline that "78% of economics is wrong." For the 3,166 articles whose earliest discrepancy is a printed number the workflow recomputed differently, the paper sorts the miss by which significant digit it affects.

earliest mismatch, by digit affected (N = 3,166) Last digit35.6% Beyond 3rd3.9% 3rd digit10.5% 2nd digit25.1% 1st digit24.9% last digit = within 1 unit of the last printed digit (rounding)
A third of the misses are rounding noise; half hit the first or second significant digit.

So 35.6% agree to within one unit of the last printed digit, which is consistent with rounding or truncation. But 24.9% differ at or before the first significant digit, and 25.1% at the second. Neither the paper nor this post can tell you whether any given conclusion flips, and the authors explicitly say the workflow "does not aim to evaluate research."

One observation can move a lot

For the 1,908 articles where the estimator treats rows or clusters as independent, the workflow applied a drop-one-influential-unit test proposed by Giordano and coauthors. Removing a single unit changed a reported point estimate by more than 100% in 963 of them, 512 of which were printed to at least two significant digits.

The part builders should steal: improvement and extension

The reproduction numbers will get the press. The improvement results are the more reusable engineering lesson. Of 1,781 articles with approximate numerical methods (optimization, integration, simulation), the workflow cut a calculation's run time by more than 10x in 496.

  • Reimplementation, 440x: a bootstrap in Lee (2026) ran 1,000 replications one at a time in a loop. The agent noticed each replication only enters through means of quantities that do not vary across replications, collapsed the loop into one matrix product, replayed the original seeded draws, and returned the published intervals exactly, roughly 440 times faster.
  • Reformulation, 24,000x: Akcigit, Pearce and Prato (2025) calibrated seven parameters to seven moments with a stochastic search plus a local solver. The workflow found a closed-form solution, certified it as a unique local root with interval arithmetic, matched every printed digit, and evaluated about 24,000 times faster.

Somewhere, a grad student's six-hour cluster job just became a matrix multiply, which is either delightful or deeply upsetting depending on who paid for the cluster.

Extensions came from 923 of 4,418 screened articles: 143 added a new counterfactual or substantive calculation, 367 quantified uncertainty that was previously unreported, and the rest added sensitivity analysis. Each extension had to survive two adversarial passes, one hunting for unstated extra assumptions and one arguing the extension is useless. One example gives the $18.69 per-person welfare loss from a gym commitment contract in Carrera et al. (2022) a 95% confidence interval of $12.75 to $23.99.

Claude as grader, and its consistency

In a side experiment, fresh sessions of Claude Opus 5 and Claude Fable 5.1 read 100 articles ten times each under three conditions: article only, article plus package, and article plus the team's reproduction. With the reproduction in hand, the models rated findings as reproduced with at most minor differences in 83.1% and 83.6% of gradings, and three random sessions agreed on the reproduction rating for 88.6% and 89.1% of articles. Agreement on whether the methodology was implemented correctly did not improve with the reproduction, a reminder that LLM verdicts on judgment calls stay noisier than verdicts on arithmetic.

Caveats

  • Not peer reviewed. NBER working papers circulate for discussion, and the authors call the reported values "current as of this release," with appendix computing less complete than article computing.
  • Anthropic tie. The paper discloses that Schwartz worked as a contractor for Anthropic on the project and that the results are not endorsed by Anthropic. The only models named are Claude models.
  • Code not out yet. The paper calls the workflow open source and points to the MetaEconomics GitHub organization, but at publication time its public repository holds no workflow code. The authors say a public release is coming.
  • Discrepancy is not error. A mismatch with a rerun can come from software versions, seeds, or undocumented steps, and a third of them are rounding.

Key Takeaways

  • An LLM workflow audited all 4,452 replication packages from five top economics journals since 2000 and verified discrepancies in 3,460 articles or appendices.
  • Only 11.2% of packages fully reproduced; about half of numeric misses hit the first or second significant digit, and 35.6% are last-digit rounding.
  • An adversarial second agent plus reruns in the original software is the design pattern that makes the flags credible.
  • The same pipeline sped up calculations more than 10x in 496 articles, including a 24,000x closed-form reformulation.
  • Treat it as a strong preprint, not a verdict: no peer review, an Anthropic-contracted author, and code still unreleased.

Sources: NBER Working Paper 35782, full paper PDF, MetaEconomics on GitHub, The Neuron daily digest, September 28, 2026.

AIAI AgentsResearchReproducibilityEconomicsClaudeNBER
CONSOLE
$