
2026-08-24 - Adam Murphy
How Do You Know If Your Science Is Any Good?
For an independent researcher, getting an idea is often easier than getting an honest answer about it. This is the story of why we built TOE-Share — and what we learned trying to make AI review worth trusting.
How Do You Know If Your Science Is Any Good?
Adam Murphy · TheoryOfEverything.ai
Having an idea is the first step in science.
The validation part is finding out whether it actually holds up. Does the mathematics work? Are the assumptions stated clearly enough to check? Has someone already disproved it? Does it make a prediction you could test? Is it genuinely new — or an old idea wearing unfamiliar notation? And finally: is it bigger than one paper can hold? That is what frameworks are for — groups of papers that accumulate into a vision no single paper could carry alone.
If you work inside a university, you have people down the hall who can help answer that. If you work independently, you don't. You can write the paper, check the equations, search the literature, and put it online. But getting someone with the right expertise to seriously examine the work — that can be a challenging wall to climb.
And without that examination, you're stuck in a strange place: you might have found something real, or you might have made one small mistake that quietly invalidates everything after it. How would you know which?
AI helps — and then creates a new problem
AI changed what an independent researcher can do. It can explore an idea with you, explain unfamiliar math, find related work, write code, and poke holes in an argument. That's genuinely extraordinary.
But asking an AI "is my theory correct?" is not scientific review.
AI models are influenced by how a question is framed. They can inherit the author's assumptions, favor ideas that resemble their training data, invent supporting details, or agree with something simply because it has been presented confidently.
Sometimes an AI will find a real problem. Sometimes it will miss one. Sometimes you can ask the same model the same question twice and receive two different answers.
So we began with a different question:
Could we build a process that makes AI less likely to miss what humans normally get wrong?
Building a review, not asking a question
That question became TheoryOfEverything.ai and the TOE-Share review platform.
Instead of asking one AI to judge an entire paper, we separate the work.
Specialist agents examine the mathematics and logic. Others inspect sources and evidence. Others evaluate scientific novelty, falsifiability, clarity, and completeness.
Multiple AI models from different providers perform those reviews independently. They do not begin by seeing one another's answers. A coordinating agent then compares the findings, surfaces disagreements, and produces a combined assessment.
The disagreement matters.
If every model reaches the same conclusion independently, that gives us one kind of signal. If one model thinks an equation is valid and another finds a missing term, that disagreement tells us where the work needs deeper examination.
We are not eliminating AI bias. We are trying to stop any one model's bias from becoming the final answer. The walkthrough walks through what that looks like on a real submission.
Then we had to review the reviewers
Building an AI panel raises the obvious next question: how do you know the panel is any good?
So we began testing it.
We submitted papers of different known quality. We ran the same work through the system repeatedly. We tested different model families. We changed prompts and checked whether the rankings remained stable. We attempted to manipulate the reviewers with fabricated citations, circular arguments, social pressure, and misleading mathematical claims.
The completed calibration study now includes 70 submissions, 1,965 specialist scores, nine models, and four AI providers.
The results were not perfect, which is important.
The testing showed that different model families interpret identical scoring standards differently. Some were substantially stricter about mathematical validity. Others responded differently to novelty.
A single AI review would hide those differences.
A panel makes them visible.
A known-error benchmark
Then we had an opportunity to try something interesting.
I saw a headline reporting that a physicist had found an error in a published scientific paper. Instead of beginning with the critique, I took the original paper and submitted it to TOE-Share to see what the review system would say.
The system was not told what error to look for.
Its reviewers flagged a missing quantum-potential term in the paper's central derivation and questioned the claim that the result represented an exact equivalence rather than a semiclassical approximation.
That pointed in the same direction as the physicist's published criticism.
TOE-Share did not discover the error first, and this was not a blind discovery on my part. I selected the paper because I knew that a problem had been reported.
What mattered was whether the review system could examine the original work and locate the same mathematical weakness without being directed to it.
It did.
That does not make TOE-Share a theorem prover, and a few successful tests do not validate the entire system. But it gives us useful pieces of evidence: the review process could identify a meaningful mathematical problem that aligned with expert human analysis. The full story is here.
Where this could go
Today, TOE-Share examines individual papers and frameworks.
But imagine where this becomes more interesting.
Thousands of researchers submit theories from different disciplines and different parts of the world. Eventually AI systems examine their assumptions, equations, evidence, and predictions. They identify theories that appear unrelated but produce similar mathematical structures. They find contradictions nobody noticed because the papers were published in different fields. They determine which future observations could distinguish one explanation from another.
Perhaps the next great scientific insight will not arrive as one perfect paper.
Perhaps it will emerge when human ideas and effort and AI come together and assist us in our understanding that fragments of the answers were already scattered across hundreds of theories, written by people who did not know one another existed.
We are not there yet.
Right now, we are doing the less glamorous work of testing reviewers, finding failure modes, exposing disagreements, and improving the process.
But that is how new scientific infrastructure begins.
Why I built this
This started with a theory of my own — Quantum Harmonia — and the uncomfortable realization that I had no reliable way to find out whether it held together. TOE-Share is the tool I wished I'd had.
I still won't quietly re-run or hand-tune my own work to make it look better. The public review history and the calibration study are how I keep myself honest about that, too.
Someone has an idea.
Someone asks whether it is any good.
And for the first time, we may be able to build a system that helps almost anyone receive an honest answer.
Your move: Submit one real thing — the one you actually want to know the truth about. Read the review as a map of where the argument is thin, not a verdict on the idea, then patch it and run it again to watch the dimensions move.
→ Open your dashboard · See how review works · Read the calibration study