Someone Finally Checked

2026-08-10 - Adam Murphy

Someone Finally Checked

Checking science was always possible. However, checking has never been affordable.

Someone Finally Checked

Adam Murphy · TheoryOfEverything.ai


For three hundred years, science has been running on a shortcut, and it was a reasonable one. The golden gates of peer review and scientific journals.

What is peer review? Maybe I will summarize. Peer review consists of a read of the story a paper tells. A reviewer asks whether the argument seems sound, whether the method seems reasonable, whether the contribution seems real. Actually checking the science... recomputing the tables, re-deriving the proofs, validating "maths", running the code... costs days of expert time per paper, and the system has never had an abundance of time to spend on that aspect. So we sampled, we trusted reputation, and we moved on. It mostly worked, because there was no alternative.

Over the past year, the alternatives finally started showing up. And they are showing up in numbers. New models, new validation AI services, entire companies and agencies being built to work on cutting-edge science. The numbers they bring are starting to tell a story.

Here is one such example: a team at Together AI and Stanford built a correctness checker on frontier models and pointed it at 2,500 published papers from top AI venues. It found:

  • an average of 4.7 objective errors per paper
  • wrong formulas
  • derivations that fail to hold
  • tables that disagree with their own text

Over 99% of papers contained at least one error, and human experts confirmed the checker was right 83% of the time it flagged something.

A separate group went further and used AI agents to replicate every oral paper from ICML 2026... the most selective tier at one of the most selective venues in computer science. Among the papers with enough verifiable claims to judge, the results were:

  • 8 out of 92 reproduced more than 80% of what they claimed
  • the median cost to reproduce a single paper's experiments came out to about $8,900

What does this tell us? Checking was always possible. Checking was never affordable.

So the errors accumulated in the dark, like the stuff you put in your closet when you do not want to deal with it. This took place in every field, in every era. Publication conferred prestige. Correctness was by necessity handled as a separate property, and we let ourselves conflate the two because nobody could afford to truly look.

Now someone can look, and I get to watch what happens when they do.

There's a pattern on TheoryOfEverything.ai that has taught me more about scientists' insecurities than anything else we've built. I will share the pattern as I see it: someone submits a paper or framework... real work, maybe a proposed solution to a genuinely hard problem.

The panel reviews it and finds something:

  • a derivation that fails to close
  • a citation that says something different than what they claimed

And predictably, within a day, the paper gets yanked. Meanwhile the same paper stays up on Zenodo or arXiv, where nobody has checked it. The version they pull is the one with the review. But why?

Now multiple explanations here are possible, and not all the papers have been pulled for the same reason, I am sure. But I cannot help but think it could be because reviews that show exactly where the work still needs to be done are not as sexy as pages of equations and diagrams.

If the goal is looking like you solved a scientific mystery, then a review with visible holes feels worse than silence. And today, that could be reasonable; AI can make mistakes, but... so do humans.

I understand this feeling from the inside. I have been working on scientific processes inside AI for the last three years. And as my own painful lesson, the very first proof-of-concept run of this platform (then called TOE-Share) caught one of my biggest research mistakes ever, in my own paper. I reworked it, and what's published today reads more like a correction than a completed work. That stung. But that is how science is done; we are not looking for AI to placate us into thinking solutions are present without the chain being complete. In our heart of hearts, science is looking for the closest thread to reality, where predictions and understanding are moved forward.

What is possible today goes beyond an institution and a name attached to some portion of proof. And yet, is that what we've all quietly rewarded for so long? A name on something that looks authoritative, hard to test, behind a closed door where the errors sleep undisturbed.

The researchers getting the most out of this AI shift have made an adjustment. They started treating reviews as a tool instead of a verdict. One of our authors, a licensed engineer with a paper already sitting with journal editors, had the panel flag a single word... one word... that would have made his mathematics invalid. The numbers were right. The word was wrong. He caught it because he asked to be checked before the world checked him. Across the work on our platform that authors revised and re-reviewed, scores improved by an average of half a point (the improvement is actually larger, but some of the most-improved papers were later pulled by their authors... the pattern I described above reaches all the way into our statistics), and most of those who have iterated have seen their work improve. The AI review found the holes, the authors filled them, and the work got stronger. That loop, the process, the science of iterating is the whole idea.

It's also why we built the review as a panel instead of a single model. One AI is one perspective, and single models have documented problems as evaluators: a study out of Washington State University found a lone model correctly identifies false scientific claims about 16% of the time, and the same model can give different judgments on identical input. Our panel runs specialist agents for mathematics, for sources and evidence, and for scientific novelty, drawn from multiple independent AI providers, with a coordinator that synthesizes the reports. When the models disagree, the coordinator digs in, because disagreement is usually where the error lives. During our calibration study, that architecture independently surfaced a real error in a peer-reviewed, published paper. We published the full calibration results, all 1,965 scores, because a review system asking for your trust should show its own work first.

Where does this go next? The trajectory for AI and science looks almost vertical. In one benchmark of papers seeded with known errors, detection rates shot from 21% (the old guard) to 55% with a single sharp model, to 90% when a whole pipeline of agents got involved... all in about a year. You can almost see the slope bending upward, fast. Keep following that line and something odd happens: checking science stops being a single hurdle at publication and turns into background noise. Papers get checked before submission, at publication, after a change, and again as new evidence rolls in. Soon, a claim that's survived checking will weigh more than one that merely survived formatting. The venue might still matter, but what a result has withstood, how it's been tested, will matter more. That is just science, and when that happens, an unchecked paper will read the way an unspellchecked one does today: as unfinished. The researchers who thrive will be the ones who invited the scrutiny early, and were okay with it as a process. They fixed what it found, and left the record up for lessons learned. But I get why there are many reasons to remove or hide works, and I am not trying to throw shade in any way; I still have works that I am afraid to run through the system because I think they would be a blow to my fragile ego, knowing that fundamentally, I still have plenty of work to refine.

That future? I think it is going to reward the people in the trenches, who are doing science, trying, failing, iterating... and that changes the value of guarding the gates, or closing the door on the uncomfortable. I, for one, am looking forward to seeing science break free from the limitations that have confined it for so long.


TheoryOfEverything.ai reviews theoretical frameworks and papers with a multi-model specialist panel. Submit your work, get checked, make it stronger.