paper Review Profile

Arithmetic Coherence: Two-Phase Structure in the BSD Space of Elliptic Curves — Paper II: The R_BSD Invariant and Geometric Ghost Classification

approvedby TestCreated 8/6/2026Reviewed under Calibration v1.3· 1 review
3.0/ 5
AI Rating

We introduce R_BSD, the ratio of analytic instability C(E) to geometric-local structure α_BSD(E), and show on 368,314 rank-0 elliptic curves that (C, α_BSD) exhibits a robust two-phase structure in which Ghost curves (|Ш|>1) occupy a distinct coherence phase with low instability and high structure. The separation is recoverable by unsupervised clustering (GMM accuracy 99.95%), is not driven by conductor or Sha alone, is strongest after torsion control, and is accompanied by spectral null results indicating the effect is local-analytic rather than due to global zero-spacing.

Read the Full Breakdown

This paper introduces RBSD=C(E)/αBSD(E)R_{\text{BSD}} = C(E)/\alpha_{\text{BSD}}(E), a ratio of analytic instability to geometric-local structure for rank-0 elliptic curves, and argues that rank-0 curves exhibit a robust two-phase structure in (C,αBSD)(C, \alpha_{\text{BSD}}) space corresponding to whether Ш>1|\text{Ш}| > 1 (Ghost) or =1= 1 (Normal). The panel domain was classified as pure_mathematics, meaning the falsifiability dimension was converted to VERIFIABILITY — the independent checkability of the numerical claims, not whether the paper makes laboratory predictions. Under this rubric the paper scores moderately across most dimensions, with completeness (4/5) being the strongest outcome and internal consistency (2/5) being the weakest, reflecting two distinct concerns surfaced by different math specialists.

On the mathematics, there is notable specialist disagreement. Two specialists (DeepSeek-V4-Pro and claude-opus-4-8) found internal consistency strong (4/5 each), noting that definitions are used coherently throughout and the BSD-conditional algebraic rewrite RBSD=L(E,1)Ш/(L(E,1)2N)R_{\text{BSD}} = |L'(E,1)| \cdot |\text{Ш}| / (L(E,1)^2 \sqrt{N}) is correctly derived. Two other specialists (gpt-5.2 and gpt-5.5) scored internal consistency at 2/5, flagging two specific structural tensions that are harder to dismiss. First, the definition of 'Ghost' drifts: the paper opens with Ghost Ш>1\Leftrightarrow |\text{Ш}| > 1, but §8.1 proposes replacing this with the geometric criterion C<104C < 10^{-4} and αBSD>102\alpha_{\text{BSD}} > 10^{-2}, without proving equivalence — and §6.2's own quadrant results imply non-equivalence (the structural quadrant still contains Ghost curves). Second, §2.1 asserts the omitted 2π2\pi normalization factor is irrelevant because it cancels in RBSDR_{\text{BSD}}, yet §8.1 then proposes absolute numerical thresholds in CC and §6 clusters in logC\log C — both of which are NOT invariant under rescaling of CC. This normalization-threshold inconsistency is a concrete problem: the proposed coherence-region boundaries C<104C < 10^{-4}, αBSD>102\alpha_{\text{BSD}} > 10^{-2} are not portable without pinning the normalization convention. The coordinator notes the first specialist pair may have underweighted these concerns; the fixed panel score of 2/5 for internal consistency reflects them.

Seven mathematical risk flags were emitted across the two flagging specialists. The highest-severity (HIGH) flags are at: §6.1 (GMM 99.95% accuracy — covariance structure, initialization, label-assignment rule, class-imbalance handling, and in-sample vs. out-of-sample evaluation are all unspecified, making this headline result non-reproducible from the manuscript); §8.1 (thresholds C<104C < 10^{-4} and αBSD>102\alpha_{\text{BSD}} > 10^{-2} — no derivation or optimization criterion is given for how these were extracted from the GMM); and §7.1–7.3 (the spectral null conclusion — failure to reject equality on five zeros for 2,334 curves does not mathematically establish that the phase structure is 'entirely local-analytic and NOT spectral'; this is an overstated inference from limited power). MEDIUM flags are at §2.1 (normalization portability), §2.2 (whether Ш|\text{Ш}| values in the dataset are independently computed or BSD-derived, which affects claimed independence), §4.1 (near-zero Pearson correlation does not establish non-dependence for a quantity with nonlinear algebraic dependence on Ш|\text{Ш}|), §6.2 (quadrant purity of 100% Ghosts — boundary sensitivity and potential circularity from median splits on the full dataset without held-out validation), and §5.1–§5.2 vs. §10 (an internal numerical inconsistency: §10's conclusion table claims typical Ghost RBSD100R_{\text{BSD}} \sim 10^010210^2, while §5's medians of 0.024 and 0.109 correspond to log10R1.6\log_{10} R \approx -1.6 and 0.96-0.96, a significant discrepancy that should be resolved). These flags are reader warnings, not score components.

On evidence and completeness, specialists consistently found the core empirical separation result (Ghost/Normal separation in RBSDR_{\text{BSD}}, confirmed by Welch's tt-test and KS test at p<10300p < 10^{-300}, across four conductor bands) to be well-supported by the large dataset (368,314 curves from Cremona/LMFDB). The reproducibility gaps are concentrated in the validation layer: the GMM procedure and the coherence-quadrant analysis lack the methodological detail needed for independent reproduction. The companion Paper I ('Ghost Rank I') is cited as foundational but its verification status is unconfirmed; readers cannot independently check claims about S(E)S(E) and the detection framework on which this paper builds. These gaps together produce the evidence/completeness scores of 3–4/5. On novelty (3/5), the consensus view is that RBSDR_{\text{BSD}} is a genuinely fresh combination of analytic and geometric-local data, and the spectral null result is a useful empirical contribution, but much of the apparent two-phase structure follows algebraically from the BSD formula combined with the definition of Ghost by Ш>1|\text{Ш}| > 1 — the novelty is primarily organizational and presentational rather than mechanistic.

Internal Consistency
2/5

The paper is mostly coherent at the level of definitions C(E), alpha_BSD(E), and R_BSD(E) (§2), but there is central consistency strain in how conjecture-conditional equalities are interleaved with unconditional definitions. In §2.2–§2.3, R_BSD is defined as C/alpha with alpha computed from local/geometric data; then the identity alpha = L(1)/|Sha| and the expression R_BSD = |L'(1)|·|Sha|/(L(1)^2·sqrt(N)) are introduced under "assuming BSD". Later interpretive arguments ("paradox" resolution §4.2; claims about algebraic dependence but near-zero correlation §4.1) treat the BSD-substituted form as if it explains the observed geometry, without consistently flagging which parts are conditional on BSD and which are purely definitional/empirical. A second internal-consistency issue is the normalization/threshold interplay: the paper claims the omitted 2π is irrelevant (§2.1) yet proposes absolute numerical phase thresholds in C and alpha (§8.1) and uses clustering in (log C, log alpha) (§6.1). The thresholds for C are not invariant under a global rescaling of C, so either the normalization must be fixed as part of the definition, or thresholds must be reported in a normalization-invariant way (e.g., relative to medians/quantiles). These are not merely stylistic: they affect the portability of the classifier and the meaning of "coherence region" boundaries.

Mathematical Validity
3/5

At the equation level, the algebra is mostly correct: given the rank-0 BSD formula, alpha_BSD = Omega∏c_p/|tors|^2 and L(1)=alpha_BSD·|Sha| implies alpha_BSD = L(1)/|Sha|, and therefore R_BSD = C/alpha_BSD = (|L'(1)|/(|L(1)|·sqrt(N)))/(alpha_BSD) = |L'(1)|·|Sha|/(L(1)^2·sqrt(N)) (§2.2–§2.3). This conditional manipulation is mathematically valid. However, the paper’s main results are empirical/statistical and currently lack enough formal specification to be considered mathematically reproducible from the document. The claims of 99.95% unsupervised recovery (§6.1), 100% ghost purity in a quadrant (§6.2), and the specific numeric threshold classifier (§8.1) are load-bearing, but no explicit mathematical definition of the classifier pipeline is given (feature scaling, covariance constraints, initialization, convergence criteria, how labels are mapped to clusters, whether accuracy is in-sample, treatment of class imbalance, etc.). If any of these choices change, accuracy and purity can change materially; thus the main theorems-as-stated are not yet mathematically secured. Separately, the use of correlation ≈ 0 to conclude "not a proxy" (§4.1) is not a valid mathematical implication without additional assumptions (e.g., monotonicity/linearity); at best it establishes lack of linear correlation. Because the central claims depend on underspecified computational/statistical steps (paper-owned, not merely cited theorems), mathematical_validity cannot exceed 3 under the rubric.

Verifiability (converted from Falsifiability)
3/5

Scored using the VERIFIABILITY rubric (pure mathematics). The computations are in-principle reproducible: dataset is the Cremona/LMFDB database, metrics are explicit formulas, and code paths are listed. However, verification is hindered by (a) redacted definitions in the abstract, (b) no exact reproduction seeds or version pins, (c) statistics like 'p<10^-300' and GMM 99.95% accuracy that are self-referential (the GMM recovers a label that is definitionally correlated with the inputs via BSD), and (d) no printed intermediate cross-checks against known LMFDB values. Claims are specific enough to be checked but require substantial reconstruction, and the central 'independence from Sha' claim is not cleanly falsifiable as stated because the algebraic dependence is acknowledged.

Clarity
3/5

The paper is well organized, with clear sectioning, compact definitions, summary tables, and an understandable narrative arc from definition to dataset to validation to interpretation. A graduate-level reader familiar with elliptic curves can follow the intended message. However, clarity is materially weakened by repeated use of physics-flavored terms such as 'instability,' 'phase,' 'coherence,' 'absorption,' and 'local-analytic' without a precise mathematical criterion distinguishing metaphor from claim. The abstract and discussion sometimes present inferentially stronger conclusions than the body warrants, especially around 'demonstrating' a local-analytic phenomenon and 'replacing' the |Ш|>1 definition. In addition, key methodological details for the GMM and spectral tests are too compressed for a reader to assess exactly what was done without consulting code.

Novelty
3/5

Proposing a two-dimensional (C, alpha_BSD) invariant space and reframing Ghost/large-Sha curves as a 'coherence phase' is a modestly novel repackaging. However, the core observation reduces to well-understood BSD relationships: curves with large |Ш| have elevated L(E,1) relative to geometric factors. The 'phase structure' is a consequence of the BSD formula plus the definition of Ghost by |Ш|>1, so much of the apparent novelty is definitional/presentational rather than a new mechanism. The spectral null result is a mild but genuine empirical contribution. Combination of known ideas with some new framing places this at 3.

Completeness
4/5

The paper is structurally complete on its own stated terms. It defines its principal quantities, specifies the dataset scope (368,314 rank-0 curves from Cremona/LMFDB up to conductor 130,000), states the BSD assumption when using the rank-0 formula involving |Ш|, and includes limitations/future work on threshold calibration, higher rank extension, conductor dependence, theoretical grounding, and twist-family robustness. The narrative follows through on the promised goals: definition of R_BSD, independence tests, phase separation, unsupervised recovery, and spectral null checks. The main reason this is not a 5 is that several support details needed for full reproducibility or boundary scrutiny are missing from the paper text. The clustering section does not specify how '99.95% accuracy' is computed for an unsupervised model, how cluster labels were matched to Ghost/Normal labels, or whether train/test splitting, initialization sensitivity, or class-imbalance effects were examined. The spectral analysis states that 2,334 curves were used but does not explain the sampling protocol, matching criteria, or uncertainty handling in enough detail to judge representativeness. The proposed geometric thresholds C<10^-4 and alpha>10^-2 are described as natural boundaries, but the method for extracting them from the GMM or figures is not formally documented. These are meaningful omissions, but they affect secondary methodological support more than the paper’s core internal completeness.

15 derivation flags— equations with compressed or unverified steps identified by math specialist

Strengths

  • +The quantities C(E), α_BSD(E), and R_BSD(E) are explicitly and consistently defined in §2, and the BSD-conditional rewrite R_BSD = |L'(E,1)|·|Ш|/(L(E,1)²·√N) is correctly derived from the rank-0 formula — the algebraic core is sound and all specialists confirmed no errors in the formula manipulations given stated assumptions.
  • +The large dataset (368,314 rank-0 curves from Cremona/LMFDB) with robustness checks across four conductor bands (§5.4) and a torsion-controlled subset (§5.2) makes the core Ghost/Normal separation result in R_BSD unlikely to be a statistical artifact; the Welch t-test and KS test at p < 10^{-300} are strongly significant.
  • +The spectral null result in §7 — no detectable difference in normalized first-zero heights between Ghost and Normal curves (p = 0.55), with a non-replicating variance signal (p = 0.011 vs. p = 0.51 in a held-out conductor band) honestly reported rather than treated as positive — demonstrates methodological discipline and is a genuine empirical contribution localizing the phenomenon.
  • +The limitations section (§9.4) is candid about empirical threshold calibration, rank restriction, lack of theoretical derivation, and the need for twist-family stability tests — this honesty about the framework's boundaries is a strength.
  • +The paper makes available code paths (Appendix B) and references an established public dataset (Cremona/LMFDB), giving readers a clear starting point for independent reproduction.

Areas for Improvement

  • -Resolve the normalization-threshold inconsistency: §2.1 argues the omitted 2π factor is harmless because it cancels in R_BSD, but §8.1 proposes absolute numerical thresholds in C (C < 10^{-4}) and §6 clusters in log C — both of which are not invariant under global rescaling of C. Either pin the normalization as part of the definition of C(E), or report all phase boundaries in normalization-invariant terms (e.g., quantiles relative to the dataset median).
  • -Prove or retract the claimed equivalence between Ghost and the geometric coherence criterion: §8.1 proposes replacing |Ш| > 1 with C < 10^{-4} and α_BSD > 10^{-2}, but §6.2 already shows that the structural quadrant contains Ghost curves, implying these are not equivalent. Either prove the coherence region is a strict subset of Ghosts (and explain the gap), or reframe §8.1 as a partial geometric characterization rather than a definitional replacement.
  • -Fully specify the GMM clustering pipeline (§6.1): covariance matrix type (full, diagonal, tied), initialization method, number of components tested, convergence criterion, how cluster labels are matched to Ghost/Normal labels, whether accuracy is in-sample or held-out, and how class imbalance (53% Ghost) is handled. The 99.95% headline accuracy is the paper's primary validation result and must be reproducible from the manuscript text alone.
  • -Address the circularity concern in the coherence-quadrant analysis (§6.2): the quadrant boundaries are defined using median splits computed on the full dataset, and Ghost purity is evaluated on the same dataset without held-out validation. This creates a potential selection artifact in the 100% Ghost purity claim. A held-out test or cross-validation is needed.
  • -Replace or supplement Pearson correlation with nonlinear or rank-based dependence measures in §4.1. Near-zero Pearson correlation between R_BSD and |Ш| does not establish non-dependence for a quantity that algebraically contains |Ш| in its numerator under BSD; Spearman correlation, mutual information, or partial dependence conditioned on L(E,1) would be more diagnostic.
  • -Resolve the internal numerical inconsistency in the R_BSD scale: §10's conclusion table claims typical Ghost R_BSD ~ 10^0–10^2 and Normals ~ 10^2–10^3 (log₁₀ R ≈ 0–2 for Ghosts), while §5.1 reports medians of 0.024 and 0.109 (log₁₀ R ≈ −1.6 and −0.96). These cannot simultaneously be correct descriptions of the same data and same normalization; the discrepancy should be explained or corrected.
  • -Provide selection criteria, representativeness justification, and sample-size rationale for the spectral subsample of 2,334 curves (0.63% of the full dataset) in §7. Without this, readers cannot assess whether the null result generalizes. The conclusion that the effect is 'entirely local-analytic and NOT spectral' should be softened to match the limited power of the reported tests.
  • -Maintain a clear and consistent conditional logic boundary throughout: α_BSD is defined without L-values, but later explanatory passages (especially §4.2 and §9.1) invoke the BSD formula to interpret the paradox and the absorption mechanism. Clearly flag at every interpretive step which conclusions are unconditional (from computed quantities) and which are conditional on BSD holding for those curves.

Share this Review

Post your AI review credential to social media, or copy the link to share anywhere.

theoryofeverything.ai/review-profile/paper/8933e20e-a5bf-4dd7-a6d6-bb964d2f46fb

Embed a live badge

Paste into a GitHub README, preprint page, or personal site. It reads the current review, so it updates on its own if the score changes or the work is withdrawn.

Preview of the embeddable TOE-Share review badge
[![TOE-Share AI review: 3.0 / 5](https://theoryofeverything.ai/api/badge/paper/8933e20e-a5bf-4dd7-a6d6-bb964d2f46fb)](https://theoryofeverything.ai/review-profile/paper/8933e20e-a5bf-4dd7-a6d6-bb964d2f46fb?utm_source=toeshare_badge)

How to use it: Markdown goes in a GitHub README or any Markdown page. HTML goes in your own web page, wherever you want the badge to appear. Both show the live badge and link back to this review profile.

Zenodo descriptions strip images; there, link the text TOE-Share 3.0 / 5 to your review profile instead. Show us you shared your work on social media — tag @TOE_Share in your post, or reply to our newsletter with the link. Either works — we’ll add one free review credit to your account.

Share by Email

Email clients cannot render the full review profile page. We send a branded HTML summary plus a link to the live credential.

Open in Email App

Sign in as the submission owner to send a branded HTML email from TOE-Share. Anyone can still copy the text or open their email app.

This review was conducted by TOE-Share's multi-agent AI specialist pipeline. Each dimension is independently evaluated by specialist agents (Math/Logic, Sources/Evidence, Science/Novelty), then synthesized by a coordinator agent. This methodology is aligned with the multi-model AI feedback approach validated in Thakkar et al., Nature Machine Intelligence 2026.

TOE-Share — theoryofeverything.ai