2026-07-30 - Adam Murphy
GPT-5.6 Sol Proposed Solutions to Six Erdős Problems. We Reviewed One.
A Columbia PhD student used OpenAI GPT-5.6 Sol to propose solutions to six open Erdős problems in five days. We imported Problem 486 as a reference paper and ran a full multi-agent review. Score: 4.2/5. Here's what the report actually says.
In late July 2026, a public claim hit the feed: Columbia PhD student Shouqiao Wang had used OpenAI GPT-5.6 Sol — driven through a long-running Codex session — to propose solutions to six open Erdős problems in five days. Problems 390, 486, 536, 788, 1002, and 1038. Thirteen attempts. Roughly a 46% hit rate on the ones he pursued.
That is exactly the kind of moment TOE-Share was built for.
Not "AI solved math." Not "AI failed at math." The useful question is narrower and harder: if a frontier model proposes a proof of an open problem, what does a structured review of that proof look like?
So we took one of them — Erdős Problem 486 — imported the publicly shared proof materials as a reference paper, and ran our multi-agent panel against it.
Overall score: 4.2 / 5. Recommendation: approve.
That is not a coronation, and it is not a dunk. It is a public report with highlights, lowlights, and specific places an independent mathematician should check before treating the claim as settled.
What the Claim Is
Erdős Problem 486 asks whether a certain sieve survivor set must always have a logarithmic density. The proposed answer is no: a finite probabilistic "gliding hump" construction produces a fixed congruence system whose logarithmic averages fail to converge, with explicit constants
liminf ≤ 177/200 < 49/50 ≤ limsup.
Wang published the proof PDFs, LaTeX sources, and prompts on GitHub. The 486 materials are here. We treated the submission as what it is: an AI-generated proposed proof, submitted so the platform could audit mathematical rigor in public — not as a final human peer review, and not as a Lean formalization.
What the Panel Said
The system classified the submission as pure mathematics, so the usual falsifiability dimension ran under a verifiability rubric: can an independent reader re-derive the claims from the stated steps?
| Dimension | Score |
|---|---|
| Internal Consistency | 5 / 5 |
| Mathematical Validity | 4 / 5 |
| Verifiability | 4 / 5 |
| Clarity | 4 / 5 |
| Novelty | 4 / 5 |
| Completeness | 4 / 5 |
The highlights
- Unanimous 5/5 internal consistency. Four math specialists agreed the activation convention, scale separation, and the two cutoff sequences (recovery vs. deletion) are used coherently. No circular argument or definition drift.
- A real conceptual contribution. The construction separates local harmonic deletion mass from a summably small global periodic footprint — then assembles those blocks with a gliding-hump epoch scheme. The tools are classical; the synthesis for this problem is specific.
- Checkable quantitative claims. Constants like 177/200, 49/50, 15/76, and the entropy margins are stated with enough specificity that independent recomputation is possible.
- Honest scoping. The paper itself says the multi-residue construction does not settle the singleton Erdős Problem 25, and it sits outside the summable-regime positive results it cites.
The lowlights
- Load-bearing numerical margins. The panel's strongest follow-up ask is independent recomputation of the entropy inequalities in the small-footprint lemma (the net
e^{-0.04k}margin). Specialists did not find an outright error; they found tight hand-computed constants that a careful reader should verify before any public claim of resolution. - Citation hygiene. Two references carried problematic identifiers (a 1951 Davenport–Erdős DOI that did not verify cleanly, and a malformed Hoeffding identifier). The underlying papers are real; the bibliographic packaging was not clean. That is a completeness issue, not a proof collapse — but for a claimed open-problem solution it matters.
- Presentation gaps. A referenced definition of the survivor set
Bwas missing from the submitted text, and at least one inequality in Lemma 2 is numerically false as written even though the intended conclusion is clear.
The full specialist write-ups, risk flags, and improvement list are on the permanent review profile.
What 4.2 Actually Means
A 4.2 on this platform means: the argument looks coherent and mostly reproducible, with named places to check, not a rubber stamp that the open problem is closed.
That is the gap the news cycle usually skips. Headlines can say "AI solved six Erdős problems." A review report says: here is the construction, here is what held up under multi-provider scrutiny, here are the equations an independent reader should recompute, and here is what still needs a human mathematician (or a Lean formalization) before anyone should treat the claim as settled science.
AI is getting better at generating candidate proofs. Someone still has to make sure they are good proofs. That is the job.
Read the Report
- Paper page: A Proposed Solution to Erdős Problem 486
- Full review profile: scores, specialist findings, risk flags
- Source materials: Wang's GitHub / 486
If you have a proposed proof — human-written, AI-assisted, or fully model-generated — you can run the same process yourself. That is what TheoryOfEverything.ai is for.