THERMOACOUSTIC COMBUSTION INSTABILITY: EVALUATION PROTOCOL
A frozen evaluation protocol for diagnosing thermoacoustic combustion instability that formalizes a general five-stage methodology for domain and operator admission, comparator verification, separate detection and state-estimation assessment, and commercial interpretation; the thermoacoustics study is the first full instantiation. It introduces a new Stage 0B Operator Structural Audit requiring pre-registered algebraic decompositions of proposed diagnostics into variance-like, coupling-like, and residual contributions, invariance checks, and an a priori falsification criterion that gates…
Approved — Pending Publication
This paper has passed AI review and is awaiting publication by the author.
THERMOACOUSTIC COMBUSTION INSTABILITY: EVALUATION PROTOCOL
Status: FRAMEWORK FROZEN. See governance note below — "frozen" applies to the framework layer only, not to study-specific appendices that remain to be written. Date filed: 2026-08-21 Nature of this document: This is not only a domain-specific ledger entry. It formalizes a general five-stage evaluation methodology for diagnostic operators, of which the thermoacoustics study is the first full instantiation. If this protocol survives Rijke-tube testing unchanged, that is itself evidence — independent of whether η succeeds or fails here — that the methodology resists outcome-driven drift.
Governance: Framework Version vs. Study Appendices
"Frozen" describes two different things in this document, and conflating them was a real clarity problem in earlier drafts — flagged by independent review. Distinguish:
- Framework Version (this document's stage architecture, decision rules, and interpretive safeguards): frozen as of the dates recorded in each section. Changes to the framework itself are versioned additions (the dated ADDENDUM entries below), never silent edits.
- Study Appendices (the Challenge-Generator Specification, the Application Specification, and any future domain-specific case configurations): each is its own separate artifact, frozen independently, on its own schedule, when it is actually written. An appendix not yet existing is not evidence the framework is incomplete or in-progress — the framework can be fully frozen while explicitly requiring appendices that don't exist yet, the same way a building code can be finalized while a specific building's permit application is still pending.
This distinction exists so that no one can later argue the framework itself changed simply because a required appendix was subsequently completed — completing Appendix A satisfies a framework requirement; it does not modify the requirement.
Stage/Level Terminology Reference
This protocol accumulated stage labels across multiple working sessions. The table below is a lookup aid — added for disambiguation, not a renumbering of the underlying architecture, since renumbering an already-cross-referenced document risks introducing broken references for no methodological gain.
| Label | Full name | One-line meaning |
|---|---|---|
| Stage 0A | Indicator Admission | Does the comparator measure the physical quantity it claims to? |
| Stage 0B | Operator Structural Audit | Is the proposed η construction structurally capable of measuring coupling, not just amplitude? |
| Stage 1A | Detector Specification and Calibration | Freeze and calibrate the alarm/threshold logic paired with an admitted indicator. |
| Stage 1B | Comparative Detection Evaluation | Compare lead time between η and admitted comparators, at a matched false-trigger rate. |
| Stage 2 | State-Estimation Product | Does η track combustion state usefully, independent of whether it triggers an alarm? |
| Replication Tier 1 | Numerical repeatability | Different noise realizations, same operating point — confidence intervals only, never counted as independent successes. |
| Replication Tier 2 | Operating-condition replication | Distinct physical operating points (equivalence ratio, Mach number, etc.) — the primary experimental unit. |
| Replication Tier 3 | Geometry replication | Different combustor configurations — tests transportability, where commercial credibility begins. |
| Level B (used elsewhere in this document) | Not a replication tier — this is the commercial/operator-evaluation campaign (sweeping operating conditions to test whether an admitted η shows differentiated value). Named identically to "operating-condition" in casual speech but kept as a distinct term from the Replication Tiers above specifically to avoid the ambiguity an earlier draft had, per independent review. |
Note on the renaming above: an earlier draft used "Level A/B/C" for the replication hierarchy, which collided with "Level B" already in heavy use elsewhere as the name of the commercial-evaluation stage. The replication hierarchy is renamed to Tiers 1/2/3 throughout this document; "Level B" is reserved exclusively for the evaluation-campaign sense.
The general methodology this protocol instantiates
- Domain admission — physics says the question is worth asking (existing Domain-Admission Hypothesis: independently derivable M, M independent of P's measurements, degradation/dynamics expected to reorganize covariance relative to M).
- Operator admission — the chosen M is independently justified for this domain, not fit to produce a desired result. Formalized below as Stage 0B: Operator Structural Audit (added 2026-08-27, after three independent constructions — FEMTO Mode 2, FEMTO sliding-lag, Stanford two-channel commutator — showed a recurring failure mode; see that section for the precise, scope-limited statement of what this pattern does and does not establish).
- Comparator admission (Stage 0 below) — each incumbent implementation actually measures the phenomenon it claims to, verified before it is allowed to serve as a comparator.
- Hypothesis-specific evaluation (Stages 1–2 below) — detection and state-estimation are evaluated as separate products with separate endpoints, not a single bundled claim.
- Commercial interpretation — occurs only after empirical performance is known; mechanistic plausibility governs admission to testing, not commercial ranking.
Mechanistic fit determines what gets tested. Comparator studies determine whether it has practical value. These are not interchangeable, and this document exists to prevent them from being collapsed into each other before any data exists.
Stage 0B — Operator Structural Audit (new, 2026-08-27)
Origin and scope of the finding this responds to. Three independent investigations — FEMTO Mode 2 (frozen-baseline construction, failed domain-admission/transportability), FEMTO sliding-lag (apparent positives undermined by overlap, conditioning, and parameter sensitivity), and the Stanford two-channel commutator (mechanistically dominated by variance imbalance rather than cross-covariance) — share a common thread: the covariance-based commutator constructions tested so far have repeatedly shown a tendency to be dominated by variance-related structure rather than the relational quantity they were intended to represent.
This is a scope-limited statement, not a general claim about η. The evidence supports a statement about these specific constructions — closely related covariance-based commutator formulations — not about η as an abstract diagnostic. A substantially different operator construction could behave differently without requiring optimism about it in advance. Do not restate this finding as "η tends to collapse to amplitude" in any future document; that overstates what three related-construction failures establish.
What this gate requires, before any operating-condition (Level B) results are interpreted:
- Is the operator invariant to simple amplitude scaling, if that is an intended property of the construction?
- Does the operator decompose into identifiable physical contributions (e.g., a variance/amplitude term vs. a cross-covariance/coupling term vs. residual)?
- Are those contributions reported prospectively — i.e., is the decomposition a pre-registered secondary endpoint, computed and reported regardless of the headline result, not assembled after the fact to explain a disappointing outcome?
- Is there an obvious dominant term that would make the interpretation trivial, decidable before any outcome analysis?
If the answer to question 4 is yes before any Level B data is examined, that construction should not advance to the full operating-condition campaign. This is not moving the goalposts after the fact — it is exactly the kind of lesson a pre-registration process exists to absorb once a recurring failure mode has been identified across prior domains.
Required pre-registered decomposition, reported as a Level B secondary endpoint (not a diagnostic invoked only when results disappoint):
η = η_variance + η_coupling + η_other
This three-term form is a general template, not itself covariance-specific — "variance/amplitude-like," "coupling-like," and "other" are roles any operator's decomposition should identify, even though the exact algebraic content of each role differs by construction. The exact algebraic form depends on the specific operator construction (for the Stanford commutator, this took the form c(u−w) + v(b−a), a decomposition derivable directly from the operator's own algebra — see the Stanford closure entry above for the worked example). Whatever the construction, the decomposition must be specified before Level B data is collected, not derived reactively.
Pre-registered falsification criterion, stated in advance: if more than a pre-specified fraction of η is attributable to variance/amplitude terms across the operating-condition sweep, the coupling-state interpretation is not supported for that construction, and this closes as a mechanistic negative — not as "inconclusive" or grounds for retuning the same construction further. (The specific fraction threshold is a construction-specific choice to be frozen alongside the rest of Stage 0B for whatever operator is proposed for Level B; not set generically here.)
How this reframes the commercial question. Prior framing asked "does η work?" — a question that invites indefinite incremental tuning after each negative result. The correct framing after this gate exists: "Is there any operator construction that survives both the structural audit and the comparator study?" A well-designed Level B campaign that returns "no" after testing constructions that clear Stage 0B is a strong, efficient negative result — not a failure of the research program. A construction that clears Stage 0B and then shows differentiated value in Level B has survived a materially higher evidential bar than passing a comparator study alone, since it has also been checked against the specific failure mode three prior domains already exhibited.
Making Stage 0B executable: two parts, not a statement of intent
Naming a gate doesn't make it a gate. Question 4 above ("is there an obvious dominant term, decidable before Level B data exists") is answered through the two-part procedure below — not by inspecting the eventual Level B sweep, which would just relocate the same circularity one step later.
Stage 0B.1 — Algebraic audit (before any simulation)
General principle, applicable to any operator construction, not only covariance-based ones: for the proposed operator, derive its decomposition into identifiable structural sensitivities, and state its expected behavior under the transformations relevant to that construction's algebra — before running anything. What "relevant transformations" means depends on the operator: a covariance-based construction needs the covariance-specific list below; a phase-based, information-theoretic, or other construction needs its own analogous list, derived from its own algebra, not borrowed from this one.
Covariance-based instantiation (this is one example, not the general requirement — flagged by independent review as language that had been written specifically enough to fit only this case): if the operator is built from covariance/commutator structure, identify which terms are sensitive to marginal variances, covariance, phase, normalization, and scale, and state expected behavior under:
- common scaling (x ↦ ax)
- independent channel scaling
- sign changes
- channel permutation
- isotropic covariance
- perfect correlation
- zero correlation
This specific list is what would have caught the Stanford construction's failure mode immediately: the v(b−a) term's capacity to dominate even when the intended interpretation concerns cross-covariance is visible directly from the algebra, without needing real pressure/OH data to reveal it after the fact.
For a non-covariance operator, Stage 0B.1 is not satisfied by checking the list above — a new, analogous set of algebraic invariance checks must be derived from that operator's own mathematical structure before it can proceed to Stage 0B.2. What must be preserved across any construction is the principle (identify structural sensitivities and test them algebraically before any data exists), not the specific covariance-flavored checklist.
Stage 0B.2 — Controlled numerical challenge suite
A small, permanent, methodological calibration set — structurally separate from Level B and permanently excluded from it (see the exclusion rules below). Unlike a normal operating-condition sweep, which may change amplitude and coupling together and reintroduce the same ambiguity, this suite requires counterfactual controls that OSCILOS's own natural parameter variation would not generate:
- Amplitude challenge — multiply simulated channels by known scale factors while preserving their normalized temporal relationship.
- Variance-imbalance challenge — rescale one channel only, leaving coupling timing unchanged.
- Coupling challenge — change FTF gain, delay, or a controlled pressure–heat-release phase relation while renormalizing marginal amplitudes to remain fixed.
- Null coupling challenge — preserve spectra and amplitudes but disrupt cross-channel timing (protected circular shifts, or independently-generated realizations — the same surrogate discipline already established for the Stanford coherence work).
- Combined factorial challenge — cross several amplitude levels with several coupling levels (see table below).
| Challenge axis | Held fixed | Varied | Expected behavior of a coupling-specific operator |
|---|---|---|---|
| Amplitude-only | coupling law, phase relation, normalized covariance geometry | common signal amplitude and channel variance scale | invariant, or changes only within a frozen tolerance |
| Coupling-only | channel amplitudes and marginal variances | cross-channel coupling, phase locking, FTF gain/delay, or normalized covariance orientation | responds monotonically and detectably |
| Mixed | neither | amplitude and coupling varied independently | tracks coupling after conditioning on amplitude |
For every run, report both total η and its predeclared components (per the η = η_variance + η_coupling + η_other decomposition above) — not just the headline value.
Admission criterion — a discrimination ratio, frozen in advance, not an arbitrary "dominant fraction":
S_A = (Δη under amplitude-only variation) / (Δη under coupling-only variation)
A construction intended as a coupling-state coordinate must satisfy a prespecified bound S_A < ε, where ε is justified from numerical uncertainty and expected practical sensitivity — chosen before seeing the challenge results, not fitted to whatever the construction happens to produce.
Deriving ε — not assigning it by intuition. A generic numerical value is not acceptable. Derive a construction-specific bound from two quantities, both fixed before the challenge results are opened:
ε = (largest amplitude-only response considered practically negligible) / (smallest coupling-only response considered practically useful)
Operationally:
- Estimate the numerical/repeatability floor from replicated identical runs — solver precision, injected-noise variation, filtering, and window-estimation effects all contribute.
Resolving an ordering ambiguity flagged by independent review: this step does not require the candidate operator to exist, and is not itself Stage 0B evaluation. The repeatability floor is an operator-independent benchmark measurement — it characterizes the noise floor of the underlying simulation/data-generation pipeline itself (e.g., OSCILOS's own run-to-run variation from wgn's injected noise, solver precision — the same kind of repeatability already characterized for the deterministic-vs-stochastic distinction in the OSCILOS Headless v1.0 package), not the candidate η construction's behavior on that data. This step is permitted and expected before any operator is proposed, using a standard, construction-independent diagnostic already available from the simulation infrastructure (e.g., replicate the frozen OSCILOS case under several noise seeds and measure the spread in pressure RMS or growth-rate fit — not in η itself, which does not yet exist at this point in the sequence). Stage 0B proper begins only once a specific operator is proposed and evaluated against the frozen challenge suite and frozen ε — the repeatability-floor measurement here is a prerequisite input to that gate, not part of the gate, and is therefore not circular.
- Define the minimum useful coupling response from the intended downstream application — not from the candidate operator's own observed performance. For control: the smallest change that can reliably alter a control decision. For state estimation: the smallest separation needed to distinguish adjacent stability-margin states. For detection: the smallest change compatible with the required warning performance.
- Set the tolerated amplitude response as a small, explicitly justified multiple of the repeatability floor from step 1.
- Freeze ε, its derivation, and its uncertainty before evaluating the candidate construction.
- Use uncertainty bounds, not point estimates, for the actual admission decision:
S_A^upper = U₉₅(Δη_amplitude) / L₉₅(Δη_coupling)
Admission requires S_A^upper < ε (not the point estimate) — this prevents a noisy denominator or a favorable random seed from producing a misleading pass.
Non-identifiability handling, required and stated in advance: if L₉₅(Δη_coupling)'s confidence interval includes zero, S_A^upper is undefined (division by an interval spanning zero has no defensible value) — the operator fails admission by non-identifiability, not by exceeding ε. This is a distinct failure mode from "S_A^upper ≥ ε" and must be recorded and reported as such: it means the challenge suite could not establish that the operator responds to coupling at all within the tested range, which is a different (and in some ways more basic) problem than "it responds too much to amplitude relative to coupling."
Remaining specification gaps in S_A itself, flagged by independent review and not yet closed. The ratio as written is not yet reproducible by a second party without three further frozen definitions:
- Δη's exact definition: is this the difference between the two extreme challenge-suite cells, the slope of η across the swept axis, the range across all cells at that axis, or something else? "Δη under amplitude-only variation" names a quantity, not a computation — the specific estimator must be written down and frozen before any challenge-suite data is collected, not chosen after seeing which definition makes a given operator look best.
- Normalization: is η normalized (e.g., by its own baseline value, by a fixed physical scale, by the challenge suite's own dynamic range) before Δη is computed, or is Δη taken on raw η? Different choices are not interchangeable — an unnormalized Δη conflates "large absolute effect" with "an operator that naturally produces large numbers," which is exactly the kind of ambiguity Stage 0B exists to remove.
- Uncertainty multiplier: the U₉₅/L₉₅ notation specifies a 95% bound conceptually but not the actual method (bootstrap over Tier-1 noise realizations? a parametric normal approximation? something else) or the number of realizations needed to estimate it reliably — recall the STATISTICAL_SUMMARY_VALIDATION_RULE.md finding from the FEMTO/Stanford work that extreme-value and tail statistics can require far more samples to stabilize than central-tendency statistics. This must be specified and its own convergence checked, not assumed adequate at whatever sample size is convenient.
None of these three gaps are optional details — until all three are frozen, S_A is not a reproducible quantity, and ε cannot be considered fully specified even once the Application Specification exists. This is a real specification gap, not a flaw in the underlying logic of the discrimination test; it belongs alongside the Challenge-Generator Specification and Application Specification as a third concrete prerequisite artifact.
Stage 0B cannot be completed until a quantitative downstream-use specification exists. If no application-specific minimum useful coupling discrimination has been declared, ε is formally undefined and no operator may be admitted to Level B. This is not a procedural obstacle to route around — it means the commercial application isn't specified precisely enough yet to know what "useful" means, and that has to be resolved before any operator, however constructed, can be meaningfully admitted. Naming a product category (detection, state estimation, control) is not sufficient — the specification needs a number (e.g., "a controller needs to distinguish stability margins 0.1 apart"), not a category label, before ε can be derived.
Incremental-information test, run on the synthetic challenge set itself:
known coupling parameter ~ A_p + A_OH + η
Even on synthetic data with a fully known, controlled coupling axis, η should explain that axis after channel amplitudes are already included. If it does not clear this bar on data built specifically to make the coupling signal easy to find, there is little reason to spend resources on the real, harder Level B campaign.
Minor specification gap, noted for completeness: the model form above (linear? some other functional form?), the fit method, and the specific criterion for "explains the axis" (a significance threshold on η's coefficient? an effect-size bound? an incremental R² requirement?) are not yet fixed. This belongs alongside the S_A Specification gaps above — a small addition to that same prerequisite artifact, not a separate blocking issue on its own.
The full dependency chain, made explicit
- Challenge-Generator Specification (freeze the generator).
- Challenge-Generator Validation (demonstrate each manipulation affects only its intended axis, within tolerance — the acceptance-check column above).
- S_A Specification (freeze Δη's exact estimator, η's normalization convention, and the uncertainty-multiplier method — see the "remaining specification gaps" note above; without this, step 4 below cannot be computed reproducibly even once step 5's inputs exist).
- Application Specification (state what "useful" means numerically for the intended product — a number, not a category label).
- ε derivation, from that application specification plus measured numerical uncertainty.
- Freeze ε.
- Only then may a proposed operator enter Stage 0B.
This ordering — application → engineering requirement → ε → operator evaluation — is deliberately the reverse of operator → observed behavior → convenient ε. The reversal is what makes the gate resistant to unconscious optimization: it is far harder to repeatedly propose operators until one satisfies a convenient ε when the ε itself was fixed before any operator's behavior was known.
Three independent evidential layers — what a result at each layer does and doesn't tell you
The protocol now separates into three genuinely independent questions, and results at one layer must not be read as evidence about another:
- Stage 0A: does the comparator measure what it claims to measure?
- Stage 0B: is the proposed operator structurally capable of measuring the intended relational quantity (as opposed to collapsing onto a lower-order statistic)?
- Level B: does an admitted operator provide practical value on real operating conditions?
A construction can fail Stage 0B without ever reaching Level B; pass Stage 0B but fail commercially at Level B; or pass both. This separation matters because it prevents two specific misreadings: a negative commercial result at Level B being mistaken for evidence the operator was structurally unsound (it may have been structurally fine and simply not commercially differentiated), and a structural failure at Stage 0B being dismissed as "just needs more tuning at Level B" (if the operator collapses onto amplitude by construction, no amount of Level B parameter sweeping fixes that). A well-executed negative result at Stage 0B is informative in its own right — it is not merely a failed precursor to Level B, and should be recorded and closed with the same weight as a Level B closure, not treated as a lesser or preliminary finding.
The challenge generator itself requires an admission record — synthetic data is not automatically neutral
Ground truth being known does not make a challenge generator neutral by default. Before implementing any operator test against it, freeze a Challenge-Generator Specification documenting:
- base OSCILOS case and source hash;
- time step, duration, noise model, and seeds;
- amplitude levels tested;
- independent channel-rescaling levels;
- coupling levels;
- FTF parameters varied;
- phase/delay manipulation method;
- marginal-amplitude renormalization procedure;
- number of Tier 1 (noise-realization) counts per cell;
- null-generation method;
- factorial cells and any exclusions;
- acceptance checks proving each manipulation changed only its intended axis, within tolerance.
That last item is load-bearing: a challenge cell fails validation, and must be rejected before any operator is tested against it, if the generator itself doesn't cleanly isolate the axis it claims to manipulate.
| Challenge | Must change | Must remain equivalent |
|---|---|---|
| Common amplitude scaling | RMS amplitudes | normalized coupling, phase, coherence structure |
| One-channel scaling | variance imbalance | timing and normalized cross-dependence |
| Coupling-only | known coupling coordinate | marginal RMS and spectra |
| Null coupling | cross-channel temporal organization | individual spectra and amplitude distributions |
| Mixed factorial | both assigned axes | all nuisance settings |
Two kinds of challenge data, matched to what each axis actually needs — not all from fresh simulation:
Deterministic transformations of one frozen trajectory (exact control over what changed, no new simulation needed): common amplitude scaling, independent channel scaling, sign changes, permutations, protected circular shifts, marginal-preserving nulls.
Fresh OSCILOS simulations (needed because these require genuinely different physics, not a post-hoc transform): FTF gain changes, time-delay changes, movement toward/away from the stability boundary, physically meaningful coupling changes.
Combining both is stronger than relying on natural OSCILOS parameter sweeps alone, since a natural sweep (varying af, τf, or Mach number directly) may change amplitude and coupling together — reintroducing the exact ambiguity this whole gate exists to separate out.
Critical separation from outcome tuning — how Stage 0B avoids becoming a hidden model-selection loop
The challenge suite must be treated exactly like the OSCILOS engineering fixture (Stage 0A's exception):
- used to test operator structure, not to generate a headline result;
- permanently excluded from Level B — a construction's challenge-suite performance is never itself reported as evidence of Level B success;
- may reject a construction (failing S_A<ε or the incremental-information test is grounds to stop before Level B);
- may not be used to optimize a succession of constructions until one barely passes, unless that iterative process is explicitly labeled as development work, not as validation. Running construction after construction against the same fixed challenge suite until one clears ε, then presenting that survivor as if it had been the only construction tested, would smuggle outcome-tuning back in through Stage 0B exactly where it was removed from Stage 0A and Stage 1. If more than one construction is tried against the challenge suite, that fact — how many, and why prior ones were abandoned — must be part of the record, not quietly dropped.
Disclosure alone is not sufficient — flagged by independent review as a real gap, not a minor addition. Recording that five constructions were tried before one passed is honest, but it does not by itself prevent the sixth construction from having been implicitly shaped by knowledge of how the first five failed against the same fixed challenge suite — informal learning across iterations can leak information into a construction's design even without deliberate outcome-tuning, and mere disclosure doesn't detect or bound that leakage. This protocol requires one of the following, chosen and stated in advance, not selected after seeing how many iterations were needed:
- Split-suite design: freeze two separate challenge suites from the start — a development suite, which may be used freely and repeatedly while designing a construction, and an independent validation suite, generated by the same Challenge-Generator Specification but with different seeds/realizations, held out and touched exactly once, by the final candidate construction only. A construction may iterate against the development suite indefinitely (labeled as development work, per the rule above); only its validation-suite result counts as the Stage 0B admission decision.
- OR periodic regeneration: regenerate the challenge fixtures from the frozen Challenge-Generator Specification (new seeds, same specification) after some declared number of construction attempts (e.g., every 3 attempts), so that no single fixed fixture set can accumulate enough exposure to be implicitly fit to. The regeneration schedule must be declared before the first construction is tried, not decided reactively.
Whichever option is chosen, it must be named in the record alongside the construction count required by the rule above — "five constructions were tried; validation-suite result reported only for the sixth" is a complete, auditable statement; "five constructions were tried" alone is not.
Stage 0A — Indicator Admission
Renamed from "Stage 0" to make an epistemic distinction explicit. Stage 0A asks only: does this comparator indicator measure the physical quantity it claims to measure? It does not involve alarm thresholds, sustained-crossing rules, or false-trigger rates — those are a property of an indicator paired with a detector, not of the indicator alone, and are handled separately in Stage 1A below. Conflating the two let a single boolean (passes) carry two incompatible meanings in an earlier draft of this protocol's supporting code; this section removes that conflation structurally, not just by relabeling.
This is an executable gate, not a set of placeholders to fill in at convenience. The protocol is not executable — Stage 1 does not begin — until every Stage 0 comparator has all five of: complete algorithm, fixed parameters, committed implementation, construct-validity criterion, and failure criterion. There is no partial Stage 1. This sequencing is a direct consequence of the FEMTO experience, where a comparator that looked adequate produced a retracted result; the gate exists specifically so that can't happen here.
Scientific incompleteness vs. procedural incompleteness. This protocol distinguishes the two deliberately. Scientific incompleteness — not yet knowing whether η succeeds or fails in thermoacoustics — is expected and is what Stages 1 and 2 exist to resolve. Procedural incompleteness — unresolved implementation choices (windowing, fitting procedures, which operating conditions constitute the study) — is not acceptable once Stage 1 begins. The gate exists precisely to prevent the second kind from leaking into and contaminating the first.
Engineering-fixture exception (required for the freeze to be achievable at all). The freeze cannot responsibly happen against a purely imagined data format — no comparator, threshold, or validity criterion can be sensibly fixed without having seen at least one real export's actual structure, units, and scaling. This creates a narrow, explicitly bounded exception:
A non-outcome-bearing interface-validation export is permitted before Stage 0 freezes. It may be used only to verify file structure, units, timestamps, signal identities, sample rate, numerical scaling, and code execution. It cannot be used to choose comparator parameters, validity thresholds, operating conditions, or η settings.
The fixture export should be one of OSCILOS's own published validation cases (e.g., the Section 6.3 Rijke tube case, already independently verified against theory in the technical report) — not a case invented for this study. Once used to validate the loader and confirm the pipeline executes correctly, this fixture is permanently excluded from Stage 1 and Stage 2 — it never re-enters as an evaluation condition, regardless of how convenient that might later seem. This exception exists solely to make the freeze achievable; it does not license using the fixture run's outcome to inform any comparator threshold or parameter choice.
Export field classification (checked against real OSCILOS capability, not assumed). The richer export schema is a requested target, not yet a verified interface. Fields are tiered so the fixture isn't blocked by an unavailable but merely desirable channel — any field OSCILOS cannot provide is recorded as unavailable, never silently replaced by an inferred quantity:
- Required for the first executable pipeline: explicit time vector; at least one pressure signal; condition parameters; stationary/ramp designation; theoretical frequency and growth rate (or a documented method to obtain them); solver/export timing; units and probe position.
- Strongly preferred: multiple pressure probes; acoustic velocity; heat-release fluctuation; mode identifier/eigenvector information.
- Optional extension: additional internal states; controller signals; nonlinear flame-model states.
Capability questions for the MATLAB operator, to be answered before the schema is treated as verified:
- Can OSCILOS log pressure at multiple axial locations?
- Can the Simulink model expose acoustic velocity as a logged signal?
- Is unsteady heat release already available as a model signal?
- Can the frequency-domain solver export numerical eigenvalues directly, rather than requiring manual reading from contour plots?
- For ramps, can σ(t) be batch-computed at the same parameter values as the time-domain trajectory?
Until these are answered, the schema is a target contract, not a verified interface.
Imported directly from the FEMTO reconciliation. Before any η comparison is permitted, each incumbent comparator must independently demonstrate it measures the phenomenon it is supposed to measure. A comparator that fails this stage produces a methodological outcome, not a performance outcome — failure here is never counted as a win for η.
Stationary-condition vs. parameter-ramp simulation design (frozen distinction)
These test different things and must not be conflated, as an earlier pipeline sketch did by averaging a fitted growth rate across nearly an entire trajectory against one fixed theoretical value:
- Stationary-condition simulations: fixed
af,τf, Mach number, etc. for the whole trajectory. The theoretical growth rate is one constant eigenvalue. A time-domain fit may be compared against that single value, but only within a prespecified linear-growth interval — before nonlinear saturation and before the signal falls into the noise floor. These test estimator accuracy and stochastic repeatability (Replication Tier 1/Tier 2), and are the design used for Stage 0 comparator admission. - Parameter-ramp simulations: a parameter changes continuously and crosses the Hopf boundary. The theoretical growth rate is σ(t), not one number — the frequency-domain solver must be run along the ramp to produce a stability-margin curve indexed by instantaneous operating point. These test warning and continuous state tracking, and are the design used for Stages 1–2.
Stage 0 admission uses stationary-condition runs. Stages 1 and 2 use parameter ramps. This split is frozen now, not left for whoever runs the first simulation to decide informally.
Stage 0 Freeze Review — required milestone before Stage 1 is authorized
Stage 1 does not begin at "simulations start." It begins only when this checklist is complete:
- Comparator A — RMS pressure
- ☐ Window definition frozen
- ☐ Construct-validity criterion frozen
- ☐ Failure criterion frozen
- Comparator B — Dominant-mode amplitude
- ☐ Mode-identification algorithm frozen
- ☐ Tracking algorithm frozen
- ☐ Construct-validity criterion frozen
- ☐ Failure criterion frozen
- Comparator C — Growth-rate estimation
- ☐ Fitting procedure frozen
- ☐ Smoothing/regularization frozen
- ☐ Construct-validity criterion frozen
- ☐ Failure criterion frozen
- Replication design
- ☐ Tier 2 (operating conditions) selected
- ☐ Number of operating conditions fixed
- ☐ Tier 3 (geometry) scope declared (included or explicitly deferred)
- ☐ Experimental unit explicitly defined
Only when every box above is checked does the protocol authorize Stage 1. The immediate next milestone for this domain is this checklist, not a simulation run.
Stage 1A — Detector Specification and Calibration (new, resolves the stable-regime circularity)
Inserted between indicator admission and comparative evaluation. This is where false-trigger rate actually belongs — it is a property of an indicator-detector pair, not of an indicator alone, so it cannot be part of Stage 0A's construct-validity check without creating a circular dependency (Stage 0A needing a detector that doesn't exist until Stage 1).
Stage 1A freezes, in order:
- the sustained-crossing algorithm;
- baseline-window construction;
- persistence duration;
- threshold-calibration rule;
- target false-trigger rate;
then calibrates using designated stable calibration trajectories, and verifies the achieved false-trigger rate on separate stable validation trajectories (calibration and verification sets must not be the same trajectories).
Governing dependency, now acyclic: indicator admitted (Stage 0A) → detector calibrated (Stage 1A) → false-trigger performance validated (Stage 1A) → lead-time comparison (Stage 1B). The detector's specification may be preregistered early, in parallel with Stage 0A — but it cannot be calibrated or evaluated until its underlying indicator has passed Stage 0A. This is not co-freezing Stage 0 and Stage 1; it is a strict, one-directional sequence with an honest intermediate stage named for what it actually does.
Stage 1B — Comparative Detection Evaluation
Formerly "Stage 1" below. Once Stage 1A's detector is locked, η and admitted comparators are applied to untouched ramp conditions, and lead time is compared at the matched false-trigger criterion established in Stage 1A — not a criterion invented ad hoc at comparison time.
For each comparator, the following must be specified and frozen before any simulation output is examined:
Comparator: RMS pressure
- Physical target: overall acoustic pressure oscillation amplitude in the combustor.
- Implementation: [TO BE FROZEN — exact windowing, filtering band if any, and RMS computation window length, fixed before simulation data is analyzed]
- Construct-validity criterion: RMS should track a monotonic or near-monotonic rise as the system approaches the Hopf bifurcation in a known-unstable simulated case; must not be flat-then-terminal-spike only (the exact FEMTO envelope-spectrum failure mode).
- Failure criterion: if RMS pressure shows no discernible pre-bifurcation trend when the terminal transition region is excluded, RMS is invalid as a comparator for that run.
Comparator: Dominant-mode amplitude
- Physical target: amplitude of the specific acoustic mode expected to go unstable (identified from the combustor's known acoustic mode shapes).
- Implementation: [TO BE FROZEN — mode identification method, whether via FFT peak-tracking or modal decomposition, fixed before simulation data is analyzed]
- Construct-validity criterion: the tracked mode must correspond to the actual growing mode in the simulation (verifiable against the known Rijke-tube model's unstable mode), not an arbitrary spectral peak.
- Failure criterion: if the "dominant mode" identified by the frozen algorithm does not correspond to the simulation's actual unstable mode, this comparator is invalid.
Comparator: Growth-rate estimation
- Physical target: the linear growth rate of the unstable mode's amplitude envelope.
- Implementation: [TO BE FROZEN — fitting window, model form (e.g., exponential fit to envelope), fixed before simulation data is analyzed]
- Construct-validity criterion: the estimated growth rate must be positive and consistent with the simulation's known stability margin in a known-unstable case; must show genuine pre-onset structure, not just a fit to the terminal region.
- Failure criterion: same shape as the FEMTO failure — if growth-rate estimates are flat or noise-dominated until the terminal transition and only appear meaningful once amplitude has already grown large, the comparator is invalid.
Comparator: Recurrence/nonlinear indicators (if included)
- Physical target: [TO BE SPECIFIED — depends on which recurrence-based indicator is chosen, e.g., recurrence rate, determinism]
- Implementation, construct-validity criterion, failure criterion: to be frozen in the same format before inclusion; if not implemented for the first run, explicitly marked as deferred rather than silently omitted.
Universal Stage 0 rule: for every comparator, construct-validity is checked against known physics, conditional on the regime the simulation was set up to produce — not a single universal growth-ratio floor imported literally from FEMTO. Bearings only had one regime (progressive degradation toward failure), so "must show growth when the terminal region is excluded" was the right universal check there. Thermoacoustics deliberately includes stable and subcritical conditions where a valid comparator should remain flat, so the criterion is regime-conditional:
- Known stable condition: comparator must show no sustained physical drift inconsistent with the known stable regime.
- Known unstable fixed condition: comparator must show a measurable linear-growth interval consistent with the independently-computed positive σ.
- Known damped condition: comparator must show decay consistent with the independently-computed negative σ.
- Ramped crossing: comparator must develop pre-onset structure in the expected direction before the nonlinear/post-Hopf region — the one case that does resemble FEMTO's check.
A comparator that fails its regime-appropriate check is invalid for that condition; a comparator that stays flat in a stable condition is behaving correctly and must not be failed by a growth-based criterion built for a different regime.
Outcome labeling for Stage 0: "Comparator invalid; comparative claim not admissible" is the required outcome language when a comparator fails — never reframed as evidence of η's superiority.
Shared simulations, independent analyses
Stage 1 and Stage 2 reuse the same underlying simulation trajectories rather than requiring separate simulation runs solely because there are two hypotheses. Running separate simulations buys nothing scientifically unless the simulation conditions themselves must differ for a specific reason (e.g., Stage 2's robustness-across-operating-conditions endpoint may require a wider parameter sweep than Stage 1's detection endpoint needs — if so, that additional sweep is an extension of Stage 2, not a replacement of shared data).
The data source is shared; the hypotheses are not. This is legitimate provided three conditions hold, in the same spirit as a clinical trial measuring blood pressure, mortality, and adverse events from the same patient cohort:
- Each analysis (Stage 1, Stage 2) is frozen beforehand — its endpoints and decision rules exist before either analysis is run, not adapted after seeing the other's result.
- Each has its own endpoints, per the Stage 1 / Stage 2 sections below — no endpoint is shared or inferred across stages.
- Neither analysis is permitted to influence the other's interpretation. A Stage 1 detection failure does not get reframed as support for (or against) the Stage 2 state-estimation hypothesis, and vice versa.
Units of replication
This is the most important unresolved issue this protocol carries, and it does not have a natural answer the way FEMTO did. FEMTO's experimental unit was supplied by nature — one physical bearing. Thermoacoustics does not hand this study an obvious unit, and it must be supplied explicitly, before any simulation is analyzed, or the resulting evidence will be easy to unintentionally overstate.
Three tiers, not to be collapsed into each other:
Tier 1 — Numerical repeatability. Different noise realizations under identical parameters. Purpose: confidence intervals and stochastic-robustness estimates only. Not counted as independent successes. A method that "succeeds" on 90 of 100 noise realizations at one fixed operating point has established one robustness estimate at one operating point — not 90 independent confirmations.
Tier 2 — Operating-condition replication. Distinct operating points within the same combustor: equivalence ratio, inlet temperature, damping, flame position, time delay, flow velocity. Each represents a genuinely different physical operating point. This is the primary experimental unit for the first study.
Tier 3 — Geometry replication. Different combustor configurations — different geometries, different mode structures, different parameter regimes. This tests transportability rather than mere repeatability, and is where commercial credibility begins. A diagnostic that only works for one geometry is scientifically interesting but commercially limited, and should be described exactly that way if that's as far as the evidence goes.
Explicit prohibition on collapsing tiers: 100 noise realizations × 4 operating points does not equal 400 successful replications. Before any simulation is analyzed, this protocol must specify which tier constitutes the experimental unit for each hypothesis (Stage 1 detection; Stage 2 state-estimation) — and results are reported and counted at that tier, never inflated by multiplying across tiers. If Tier 2 is the declared unit for Stage 1, the Stage 1 outcome table reports one row per operating condition, with Tier 1 realizations collapsed into that row's confidence interval — not listed as additional successes.
Declared unit for this study: [TO BE FROZEN — which specific operating-condition parameters (from the Tier 2 list above) will vary, and how many distinct conditions constitute the study, must be fixed before the first simulation is analyzed, on the same gated basis as Stage 0's comparators.]
Stage 1 — Detection Product
Commercial question: Does η improve instability warning?
Primary endpoint: lead time to instability onset, at matched false-trigger rate, using sustained-crossing logic (not first-crossing — per the FEMTO lesson that first-crossing claims are frequently non-robust), with thresholds set prospectively from a frozen pre-onset reference window (not tuned against the outcome).
Secondary endpoints: none for this stage — detection is scoped narrowly and deliberately, to avoid the bundling problem this protocol exists to prevent.
Decision outcomes (mirrors FEMTO structure):
- η superior
- equivalent
- mixed (heterogeneous across operating conditions/geometries — not resolved to a single average)
- inferior
- comparison inadmissible (Stage 0 comparator failure)
Frozen interpretation rules for this stage:
- A comparator failure cannot be interpreted as η superiority.
- A tie cannot be described as differentiation.
- Heterogeneous outcomes across operating conditions cannot be summarized as superiority from an average effect — the FEMTO bearing-level pattern is reported as-is, not averaged, and the same rule applies here across whichever Tier 2 operating conditions constitute this study's declared unit.
- Results are reported and counted at the declared unit of replication (Tier 2 by default) only. Tier 1 noise-realization repeats collapse into that unit's confidence interval and are never listed or counted as additional successes.
Stage 2 — State-Estimation Product
Completely separate study. Lead time does not appear anywhere in this stage — its absence is deliberate, to prevent the two product concepts from contaminating each other's interpretation.
Commercial question: Does η provide a better state variable for combustion dynamics — one useful for monitoring or control — independent of whether it triggers an alarm earlier than anything else?
Candidate endpoints:
- Smoothness versus noise (does η vary more smoothly than raw pressure or amplitude-based indicators, as a candidate control-loop input?)
- Monotonicity approaching instability (does η change in a consistent direction as the system approaches the Hopf boundary, even while amplitude is still small?)
- Correlation with an independently estimated stability margin (computed from the simulation's known linear stability analysis, not from η itself or from any of the Stage 1 comparators)
- Robustness across operating conditions (equivalence ratio, flow rate, etc.)
- Transfer across combustor geometries without retraining
- Demonstrated usefulness as an input to a controller or digital-twin state estimator (a concrete downstream-use test, not just a statistical property of η in isolation)
Frozen interpretation rules for this stage:
- Improved smoothness without predictive utility is not commercial validation — smoothness is necessary but not sufficient.
- Stronger correlation with stability margin, without independent robustness testing across conditions, is not commercial validation.
- Geometry-specific success is not evidence of generality — a positive result on one combustor geometry is a single data point, reported as such, not extrapolated.
What is deliberately not yet specified
The next executable unit is the Stage 0B Challenge-Generator Specification and Validation — not the Level B sweep, and not the first candidate operator's test. First establish that the challenge generator can independently manipulate amplitude and coupling as claimed (the acceptance-check column in the validation table above). Then derive ε from numerical uncertainty plus a declared minimum practical sensitivity (which requires the downstream application to be specified precisely enough to state a "minimum useful coupling response" — if it isn't, that is itself the blocking issue, not a detail to defer). Only after both the generator's validation and ε are frozen does the first proposed operator receive a legitimate Stage 0B test.
Bracketed [TO BE FROZEN] items above (exact comparator implementations, mode-identification method, fitting windows) must be filled in and frozen before the first outcome-bearing evaluation simulation is run or analyzed — not before any simulation whatsoever, per the engineering-fixture exception above, which permits one non-outcome-bearing interface-validation export prior to freeze. This document is incomplete by design until those specifics are added — but the structure, decision outcomes, and interpretation rules above are frozen now and are not to be altered once evaluation simulation data exists, per the same discipline already applied to the FEMTO ledger entry.
ADDENDUM (2026-08-25b): Stable-reference search within the Stanford dataset — closed
Ten inspected Stanford recordings spanning two combustor lengths (Lc=230, 305mm), multiple equivalence ratios (Φ=0.5, 0.7, 0.8, 1.0), and hydrogen fractions from 0% to 80% all exhibited coherent thermoacoustic oscillation. The most informative cross-geometry comparison — pure methane at Φ=0.8 and Ub=22.5 m/s — showed nearly identical intermittent-bursting statistics at both lengths (CV≈0.52, fraction of time below 20% of mean envelope amplitude≈0.031 in both), differing primarily in amplitude scale (pressure std 71.7 Pa at Lc=305 vs. 41.7 Pa at Lc=230). This is a real cross-geometry replication of the intermittent regime, not a stable-versus-unstable contrast — it argues the bursting behavior is not an idiosyncrasy of one chamber length.
Bounded finding, stated at the scope the evidence supports: the public dataset does not supply a demonstrated stable reference within the inspected operating envelope (ten conditions, two geometries, Φ=0.5–1.0, XH2=0–80%, fixed Ub=22.5 m/s). This is narrower than "the burner has no stable state at this velocity" — an earlier, overreaching phrasing corrected here. Unmeasured operating points, startup transients, different thermal states, or narrower regions between the sampled conditions could still contain a stable case; none of that has been ruled out, only the specific points actually inspected.
Further stable-case searching within this dataset is discontinued.
Protocol consequence:
- Stanford dataset: suitable for transitional-regime and Stage 2 testing; candidate (not completed) Stage 1B evaluation, pending detector calibration from elsewhere.
- Not suitable for: Stage 0A stable-regime admission or Stage 1A false-trigger calibration.
- Next required source for Stage 0A/1A: controlled OSCILOS stable/damped cases, or a separate experimental dataset with independently verified non-oscillatory recordings (the annular combustor and Cambridge Rijke tube candidates remain unverified for public access as of this entry — see prior conversation record; neither should be treated as available until independently confirmed).
ADDENDUM (2026-08-25c): OSCILOS headless integration — provisional Stage 0A milestone, not yet admitted
Status: provisional. Not yet admitted ground truth. Recorded now for traceability; three verification steps (report cross-check, systematic mode search, time-domain cross-validation) must pass before this eigenvalue is treated as Stage 0A ground truth for any comparator admission decision.
Milestone: The packaged OSCILOS_long Rijke-tube case (cases/Case_Rijke_tube.mat, from the real, unmodified MorgansLab/OSCILOS_long GitHub repository) was loaded successfully in Octave 8.4.0 running headlessly in this sandbox (no MATLAB license). The dedicated linear determinant function Fcn_DetEqn_Linear.m — which does not depend on the GUI-entangled FDF/Fcn_Para_initialization chain that blocked the general-purpose Fcn_DetEqn.m — was evaluated directly. Numerical root-finding (Nelder-Mead simplex, fminsearch on |F(s)|², starting from a coarse grid-scan minimum) identified a complex eigenvalue at:
σ = 1.399 s⁻¹, f = 307.03 Hz, determinant residual |F(s)| ≈ 8.2×10⁻⁵
One confirmed source-code patch was required and applied to a working copy only (original backed up as Fcn_DetEqn_original_backup.m): CI.indexFM → CI.FM.indexFM in Fcn_DetEqn.m line 100, justified by cross-referencing the field name's usage elsewhere in the same codebase (Fcn_TD_main_calculation.m uses CI.FM.indexFM(1)) and confirming via isfield() that CI.indexFM does not exist in the loaded case data. This patch was not needed for the successful Fcn_DetEqn_Linear.m path.
This establishes working access to OSCILOS's frequency-domain physics and supplies a candidate independent stability reference. It does not yet open Stage 0A. Pending, in order:
- Cross-check against the technical report's published value for this case (verifies case identity, not just code function) — flagged as urgent: the loaded case shows flame-filter cutoff
fc=15Hz, while the technical report's Section 6.3 text (as read earlier in this work) describedfc=75Hzfor its worked example. This discrepancy must be resolved before anything else proceeds. - Systematic search for other eigenmodes (this root came from one basin of attraction from one starting guess; it is not yet confirmed to be the fundamental or the dominant unstable mode).
- Independent time-domain cross-validation (fit growth rate from a simulated pressure trajectory's initial linear-growth interval only, compare against σ=1.399/s).
Update (2026-08-25d) — case-identity verification: initial case rejected, exact packaged match found.
The originally-loaded Case_Rijke_tube.mat was checked against the report's Fig. 24 documented parameters (radius=50mm, T1=293.15K, T2=586.3K, M1=0.001, af=1, fc=75Hz, τf=3ms) and rejected: it showed fc=15Hz (not 75Hz) and, more fundamentally, T2=351.5K (not 586.3K) and M1=0.0029 (not 0.001) — a genuinely different thermal/flow configuration, not merely a different flame-model parameter. The ~307Hz roots computed from this case (SHA256 465990018...) are retained in this record as evidence the numerical pipeline runs correctly, but are not Stage 0A physical ground truth — they correspond to an undocumented, unverifiable configuration.
A metadata-only screening pass (no eigenvalue calculations) across all 11 packaged cases in cases/ found an exact match: Case_PD_without_HR.mat (SHA256 29bb4670d...) — radius=50mm (exact), T1=293.1K, T2=586.1K, M1=0.0010 (exact), af=1/fc=75Hz/τf=3ms (exact), no HR/Liner, matching Fig. 24(a) "Rijke tube with no dampers." (Case_PD_with_linear_HR.mat shares the same base parameters plus one HR — likely Fig. 25/26(b), a related but different figure, not used here.) No case reconstruction was necessary.
Running the identical root-finding pipeline on this confirmed exact-match case, starting from the published target as an initial guess, converged to:
σ = 47.16 s⁻¹, f = 201.06 Hz (determinant residual |F| ≈ 1.5×10⁻⁶)
against the published target of f=200 Hz, σ=45 s⁻¹ (report text: "the resonator is seen to split the first unstable mode (200 Hz, 45 rad/s)") — a 0.53% frequency match and 4.8% growth-rate match, consistent with the published value being read from a contour-plot minimum rather than stated as an exact analytical figure.
Update (2026-08-25e) — systematic mode search: complete, dominant mode confirmed.
A frozen grid of 55 starting points (frequency: 0–1000Hz in 100Hz steps; growth rate: −100 to +100/s in 50/s steps) was searched using Case_PD_without_HR.mat. All 55 converged (residuals <10⁻⁵). Clustering (tolerance: 1Hz frequency, 1/s growth) found 10 distinct roots.
The 201.06Hz/σ=47.16 mode — the one matching the published (200Hz, 45/s) target — has the highest growth rate of any mode found in the scan. This is an independent confirmation beyond the earlier numerical match: it is unbiased by the search (55 starting points spread evenly across the domain, not concentrated near this mode) and directly corroborates the report's own characterization of it as "the first unstable mode." Other unstable modes found: 377.64Hz/σ=25.47, 790.26Hz/σ=11.07, 966.12Hz/σ=6.32, 1741.69Hz/σ=0.59 — all with lower growth rates. Stable modes: 581.46Hz/σ=−1.26, 1161.33Hz/σ=−5.60, 1372.22Hz/σ=−8.96, 1551.04Hz/σ=−2.81.
One root, at f≈0Hz/σ=−0.687, is flagged as likely a mathematical artifact of the characteristic equation near s=0 rather than a genuine acoustic mode — not discarded, but not counted as a physical mode without further scrutiny.
Scan domain limitation, stated explicitly: 0–1000Hz, ±100/s. Not exhaustive beyond these bounds; treated as a solid characterization within the scanned domain based on 100% convergence and consistent basin structure across widely-spread starting points.
Update (2026-08-25f) — time-domain branch: SUCCESS. Stage 0A frequency-domain ground truth is now admitted.
Fcn_TD_main_calculation_without_GUI (a local function, called via its public entry point Fcn_TD_main_calculation(0)) was run to completion on the same confirmed exact-match case (Case_PD_without_HR.mat). This required five additional fixes, all documented, all mathematically exact substitutions (not approximations) or minimal missing-default supplements — none of which touch OSCILOS's own physics or simulation logic:
Fcn_Pre_calculation_headless.m(new file,shims/) — replicates the computational core of the GUI-onlyGUI_TD_GreenFcn.mdialog (populatingCI.TD.Greenfor inlet BC, outlet BC, and the flame transfer function), skipping only GUI construction code (handles, uipanel, guidata). Verified: both boundary conditions are trivial static gains (num=−1, den=1, consistent with open-open boundaries); only the flame transfer function needed real impulse-response computation.tf.m,impulse.mshims — replace the missing Control-Toolbox functions via exact partial-fraction decomposition (residue(), a base Octave function). Validated to zero numerical error against the closed-form analytical impulse response of the flame's first-order filter.wgn.mshim — replaces the missing Communications-package noise generator, using the exact convention documented in OSCILOS's own source comment ("background noise level = 10^(value/20) Pa"). Validated against expected RMS amplitude.lsim.mshim — replaces the missing Control-Toolbox forced-response simulator, via exact zero-order-hold discretization (state-space conversion + matrix exponential, using baseexpm()). Validated to machine precision (2.6×10⁻¹⁵ error) against the analytical step response. Documented caveat: MATLAB'slsimmay default to first-order-hold (linear) interpolation for continuous systems, not ZOH — this is a genuine, small, unverified difference in interpolation convention, not claimed as bit-identical to MATLAB's internals. Given dt=1e-4s is ~20× finer than the flame's time constant (~2.1ms), the expected impact is small.hilbert.mshim — replaces the missing signal-package Hilbert transform, via the standard FFT-based analytic-signal construction. Validated against a known constant-envelope test sinusoid.- Manual
ExtForceInfodefault —Fcn_TD_main_calculation_without_GUIitself omits this argument when callingFcn_TD_INI, an apparent small incompleteness in the convenience wrapper (not related to Octave/MATLAB compatibility). SuppliedindexForcing=1("without forcing") plus the same default sub-fields documented inFcn_TD_INI's ownnargin==0branch.
Signal morphology (native OSCILOS output, CI.TD.uRatio — the velocity-ratio-at-the-flame variable OSCILOS's own results-plotting function uses): clean exponential growth from the noise floor (~10⁻⁴) to the nonlinear saturation plateau (~1.0) between t≈0 and t≈0.24s, followed by a complex, non-monotonic nonlinear beating pattern in saturation (amplitude oscillating ~1.0–2.6). This is an unusually clean growth signal — five orders of magnitude of monotonic exponential rise before saturation.
Frozen interval-selection rule (defined from signal morphology alone, before fitting or comparing to any target): noise floor = median envelope over t∈[0.001,0.05]s (measured: 1.11×10⁻⁴); fit interval = envelope between 10× noise floor and 20% of saturation amplitude. This produced a fit window of t∈[0.0853, 0.2051]s, 1149 samples — not adjusted after seeing the result.
Result:
- Fitted growth rate: σ_TD = 44.50 s⁻¹ (R²=0.9915) vs. frequency-domain target σ=47.16 s⁻¹ — 5.64% difference
- Fitted frequency (same interval): f_TD = 205.08 Hz vs. frequency-domain target f=201.06 Hz — 2.00% difference
Stage 0A frequency-domain branch: ADMITTED. Two independent computational paths within OSCILOS — frequency-domain determinant root-finding and nonlinear time-domain simulation with a frozen, non-circular fitting rule — agree closely with each other and with the published literature value (200Hz, 45/s). This satisfies the cross-validation this entire multi-session integration effort was built toward. The eigenvalue (f≈201–205Hz, σ≈44.5–47.2 s⁻¹, with the small spread across methods itself an honest characterization of this case's uncertainty) is now usable as genuine Stage 0A ground truth for future comparator admission decisions on this specific case.
What remains before this generalizes beyond one case: this validates one case (Case_PD_without_HR.mat) at one operating point. Extending to the Level B operating-condition sweep the protocol's Stage 0A/1A structure calls for — varying af, τf, or Mach number — has not been attempted and would need its own runs through this now-working pipeline.
ADDENDUM (2026-08-25): First Stage 2 construction — closed, negative
Dataset: Stanford fuel-flexible combustor experimental data (Akoush et al. 2025), Lc=305mm, Φ=0.8, six conditions spanning XH2=0,40,41,43,50,80%. Independently verified real experimental data (SearchWorks/PURL confirmation), not simulation. This sequence itself was validated at length: dominant-mode construct validity confirmed; a genuine 41%→43% intermittent-to-sustained transition confirmed via matched pressure/OH burst statistics and pressure-OH coherence stability; the coherence result confirmed against cross-file and within-file circular-shift surrogate nulls (p≤0.0196 in every condition); the well-known extreme-value non-convergence of surrogate maxima at small N characterized via a dedicated convergence study (N=50→500), yielding the general STATISTICAL_SUMMARY_VALIDATION_RULE.md now filed alongside this protocol.
This sequence remains what it was scoped as: strong Stage 2 experimental material and candidate (not completed) Stage 1B material. It does not contain a genuinely stable/no-instability reference case despite extensive search (9 real files inspected across two geometries, three equivalence ratios, hydrogen content 0-70%, all showing coherent oscillation) — Stage 0A stable-regime admission and Stage 1A false-trigger calibration remain unresolved by this dataset and require either OSCILOS stationary runs or a different experimental source.
First η construction tested against this sequence — CLOSED, NEGATIVE.
Construction: P(t) = rolling 2×2 covariance of z-scored [pressure, OH*] over 20ms windows; M = frozen, trace-normalized, pooled covariance from the XH2=0% condition (the calmest available, not a genuine stable reference); η(t) = ‖[P(t), M]‖_F. Normalization constants frozen from XH2=0% only, applied identically across all six conditions — no per-condition leakage.
Result: mean η correlates strongly with raw channel amplitude (r=0.86 with pressure RMS, r=0.84 with OH RMS). A post-hoc, explanatory-only decomposition of the 2×2 commutator into its two algebraic terms — c(u−w), the cross-covariance-sensitive term, and v(b−a), the variance-imbalance term — showed v(b−a) dominating in every one of the six conditions, by a factor of roughly 2.4× to 8×. This mechanistically explains the amplitude correlation: the construction is structurally more sensitive to how pressure and OH variances separate from the reference than to their actual cross-covariance. η's intermittency signature was smoother and more monotonic than raw pressure's own, but this smoothness came at a cost: it suppressed the independently-documented (Akoush et al. 2025, Fig. 4) "short, higher-amplitude bursts" precursor at XH2=41%, which raw pressure's intermittency statistic preserved as a real, non-monotonic uptick.
Final characterization: The frozen-reference two-channel commutator does not provide differentiated thermoacoustic coupling-state information on this six-condition sequence. Its response is structurally dominated by pressure–OH variance imbalance rather than cross-covariance, is strongly amplitude-associated, and loses a documented transitional feature preserved by the raw pressure signal. Its smoothness and monotonicity are real but do not constitute commercial state-estimation value.
Scope of this closure: this closes only this specific construction (2×2 raw-channel covariance, XH2=0% frozen reference, 20ms window) on this specific sequence. Per the discipline applied throughout this program (CFRP, S_mech, FEMTO Mode-2 and sliding-lag), the construction is not retuned against this outcome — doing so would convert a clean test sequence into a development set. A deliberately cross-covariance-focused operator (e.g., one that isolates or upweights c relative to the variance-imbalance terms structurally, rather than taking the full commutator) remains a distinct, open hypothesis. It requires prospective declaration and evaluation on untouched data — not a retry against this sequence.
Governing principle, restated
Mechanistic fit determines admission to testing, not expected commercial value. Physics says thermoacoustic instability is a coupling phenomenon and that testing η here is a reasonable use of effort — the same kind of argument that was reasonable and wrong for CFRP composites and the continuum mechanistic weighting work. Whether η actually provides differentiated value — as a detector, as a state variable, or not at all — is an empirical question this protocol is built to answer honestly, in either direction, once real data exists.
Current Status (Protocol Freeze, 2026-08-27)
This section is not for today's readers. It exists so that anyone opening this document later — six months from now, a year from now — cannot mistake a sophisticated Stage 0B methodology section for completed Stage 0B testing. The protocol intentionally stops at the line below until the three listed prerequisites exist. Their absence is not an oversight to be filled in casually; each is itself a required, gated deliverable under the dependency chain specified above.
- Stage 0A methodology: Complete (the architecture — comparator definitions, gate structure, decision rules — is fully specified). No individual comparator has completed Stage 0A admission — this is a distinct claim from methodology completeness. As of this entry, zero comparators (RMS pressure, dominant-mode amplitude, growth-rate estimation) have been run through the freeze-review checklist to actual admission; the OSCILOS Headless v1.0 milestone validated the simulation infrastructure those comparators would eventually run against, not the comparators themselves.
- OSCILOS Headless v1.0: Benchmarked, internally reproduced (Octave 8.4.0), frozen. Independent-machine reproduction still outstanding.
- Stage 0B framework methodology (algebraic audit, challenge-suite concept, discrimination metric S_A/ε, anti-model-selection rule, separation from Level B): Complete. Stage 0B execution remains blocked pending the three prerequisite artifacts below.
- Three-layer protocol architecture (Stage 0A / Stage 0B / Level B as independent evidential questions): Complete.
- Challenge-Generator Specification: Not yet written.
- S_A Specification (Δη estimator, normalization, uncertainty-multiplier method): Not yet written. Flagged by independent review as a real specification gap distinct from the Application Specification below — required even once the Application Specification exists, since S_A must be computable and reproducible on its own terms.
- Application Specification (quantitative "useful" discrimination requirement for the intended product): Not yet written.
- ε: Undefined by design, pending the Application Specification. Cannot be assigned by intuition or inferred from any operator's observed performance without violating the ordering this protocol requires (application → requirement → ε → operator evaluation, never the reverse).
- No operator is currently admitted to Stage 0B. No operator may be admitted until all three prerequisite artifacts above are frozen.
What has been demonstrated, stated plainly: the infrastructure to ask the question rigorously now exists. What has not been demonstrated: whether any η construction provides differentiated value in thermoacoustics. These are different milestones, and the distance between them has not shortened by virtue of the first one being real and hard-won.
No related papers linked yet. Connect papers that extend, correct, or build on this work.
You Might Also Find Interesting
Semantically similar papers and frameworks on TOE-Share
No comments yet. Be the first to discuss this work.