Live test record · updated October 5, 2026

The coupled-scalar test — open record

Is a scalar field coupled to matter — the corpus’s central anomaly — allowed by the data? This page follows the test from start to finish: the model declared before any chain, the data and the decision rule, how two agents built it, every withdrawal, every validation result, and what is still pending.

Status: Built · run specification frozen · minimum-step stall repair passed its one-point test · inference precision not yet certified · no chains run yet. Claim status: Hypothesis. The model’s claim stays a hypothesis until the chains read the data. Either outcome will be published here.

Where the test stands

  1. Scope and model agreed — done
  2. Solver built — done
  3. First validation suite passed — done
  4. Pre-pilot checks — in progress
  5. Sampling pilot — next
  6. Production chains — next
  7. Results published here — next

Pending before the pilot

  1. Revised precision design — The screened 7300 design failed its matched validation against the 8000 reference at the two steep corners (λ = 2): Plik moved −0.213 and −0.189, against the 0.05 limit, while points 1–5 passed every criterion. 7300 is not eligible. A revised design must hold at all seven declared points under the unchanged criteria, and must be declared and approved before it runs; none is declared yet. The repair’s test run, which came first in the agreed order, passed on October 5, 2026. The repaired line exists only in that run’s solver copies so far; any later run that uses it must declare it.
  2. Inference-precision certification — Tighter numerical precision must leave predictions and likelihoods unchanged at every declared point. Not yet certified: the first ladder did not converge, the staged check stopped at its first point, the bounded follow-up stopped at its first updated pair on Plik (+0.0660 against 0.05), and the screened 7300 design failed its matched validation at the steep corners. Any revised candidate needs its own declaration and must pass the unchanged suite.

Then a short sampling pilot — proposals, likelihood, solver failures and speed, plus the declared Halofit sensitivity measured by reweighting pilot samples (plan drafted; approved with the pilot) — then production chains. The pilot’s speed sets the chains’ schedule.

What the test asks

A bounded test that lets the chains decide whether a scalar field coupled to all matter is allowed by cosmological data and local gravity tests. It is the corpus refined back to its sharpest testable core, run in the cosmology project where the hi_class solver lives.

A declared candidate, not claimed as uniquely derived by the corpus. The chain measures β and λ only: it cannot separately identify k_E and M_C, so it tests β ≠ 0 and the ratio, never k_E alone.

The declared model

Field, coupling and potential

A(φ) = exp(βφ/M_Pl)
V(φ) = V₀·exp(−λφ/M_Pl)
K = 1

Reduced Planck mass. Universal coupling to baryons, dark matter and massive neutrinos. The potential is the dark energy — no separate Λ. Radiation and massive neutrinos are explicit.

Matter-frame form (u = φ/M_Pl, X = −½(∂φ)²)

G₄ = (M_Pl²/2)·e^(−2βu)
G₂ = (1 − 6β²)·e^(−2βu)·X − V₀·e^(−(λ+4β)u)
G₃ = G₅ = 0

Derived separately by both agents. Observables are computed in the matter (Jordan) frame.

Property functions

α_M = −2β·(dφ/d ln a)/M_Pl
α_B = −α_M
α_T = 0

α_M(z) is derived from the field evolution, not sampled. α_T = 0: gravitational waves travel at the speed of light.

Local strength

fifth force = 2β² × gravity
γ − 1 = −4β²/(1 + 2β²)

The convention check both agents used.

Symmetry

(φ, β, λ) → (−φ, −β, −λ)

Leaves the model unchanged, so the prior can take λ ≥ 0 with β of either sign. Whether β and λ share a sign decides if a chameleon resting point exists (see the principle below).

Colson’s principle: a dynamic scalar field

“it is a dynamic scalar field”

— Colson

Like a chameleon: present everywhere, but camouflaged by its environment until conditions bring it into function.

V_eff(u) = V₀·e^(−λu) + ρ·e^(βu)
βλ > 0: a density-dependent resting point, m² = β(λ + β)ρ/M_Pl²
βλ < 0: no resting point, so no screening
not yet settled: m² ≈ V″ + β²ρ/M_Pl²

Einstein frame, conserved ρ.

Whether the field settles is the thin-shell question, so its local range is checked environment by environment, never inferred from how light it is on cosmological scales. The camouflaged regime — a chameleon, or a symmetron-type threshold switch — is the screened branch: deferred, not excluded. If the data show tension, the decision rule points there.

Scope

  • Unscreened branch only, over an explicit parameter range.
  • Screened branch deferred, not excluded: it needs a self-consistent environmental profile plus laboratory-G calibration.
  • Dark-matter-only coupling and a field-dependent kinetic term K(C) are separate future models.
  • Why the unscreened branch is self-consistent (order-of-magnitude estimates): Cassini gives |β| ≲ 2.5×10⁻³, so N = β(λ + β) ≲ 10⁻² for λ ≲ 4 — about 10⁸ below the estimated thin-shell onsets (≈ 4×10⁶ for the Sun, ~10¹⁰ for the Earth). With N < 0 there is no minimum at all.
  • In range, G_lab = G_*·A²(1 + 2β²); the correction is ≲ 1.3×10⁻⁵.
  • Universal coupling is a metric theory: no separate first-order clock or spectral-line signal (δ ln A = 2β²Φ_N/c² sits inside G_lab).

Data and local tests

  • Planck 2018 — Cosmic microwave background: full high-ℓ Plik TT/TE/EE with its nuisance parameters, low-ℓ TT, low-ℓ EE and lensing
  • Pantheon+ — Type Ia supernovae, without the SH0ES calibration
  • DESI 2024 — Baryon acoustic oscillations (first data release)

Frozen September 30, 2026 with Colson’s approval in the cosmology project’s run specification (SHA-256 as reported by its agent: FROZEN_SPEC.md 368af53429ee…, frozen_inference.json 8404307c8d59…). DESI’s second data release was declared before any result as a separate follow-on run. Every run uses the linear matter treatment, a declared approximation; a separate Halofit check measures the sensitivity to nonlinear structure. The first run uses the ΛCDM comparison’s neutrino prescription (one massive at 0.06 eV, N_ur = 2.0328) in a matched pair: the same likelihoods and numerics for the coupled model and ΛCDM.

Local gravity terms
TestMeasuredSourceIn this modelStatus
Cassiniγ − 1 = (2.1 ± 2.3)×10⁻⁵Bertotti, Iess & Tortora 2003, Nature 425, 374γ − 1 = −4β²/(1 + 2β²) ⇒ |β| ≲ 2.5×10⁻³Labeled “unscreened”. Enters as a Gaussian likelihood term in the cosmology-plus-local run only.
Lunar laser rangingĠ/G₀ = (7.1 ± 7.6)×10⁻¹⁴ yr⁻¹Hofmann & Müller 2018, Class. Quantum Grav. 35, 035015 (LLR data 1970 – January 2015)Ġ/G = −α_M,0·H₀ ⇒ σ(α_M,0) ≈ 1.1×10⁻³Value checked against the published abstract, September 30, 2026; the term is wired into the likelihood code. It matters when β and λ are sampled independently; on the historical-ratio line Cassini binds first.

Priors

  • Local gravity enters as likelihood terms, not prior cuts.
  • β: uniform on [−0.01, 0.01], either sign (frozen) — four times the Cassini edge, so the data, not the prior, set the edge.
  • λ: uniform on [0, 2] (frozen); λ ≥ 0 by the symmetry. Any pile-up at λ = 2 is reported as a prior-edge effect, not a measured cutoff, and widening the range needs a new declaration. The edge is below 4, so the N ≲ 10⁻² scope line stands as written.

Decision rule and expectations

Stated before any chain.

  1. Two runs: cosmology only, and cosmology plus the local-gravity terms, each matched against a ΛCDM run with the same likelihoods and numerics.
  2. If the two runs agree, the result is an allowed region of β and λ.
  3. If they are in tension, the result shows where the unscreened model fails to reconcile the observations — the data’s signal to open the screened branch or declare a model change.
  4. Expected split: cosmology mainly constrains λ; the local terms mainly constrain β (cosmological coupling effects are small: fifth force ~2β² ≲ 10⁻⁵, matter-mass drift ≲ ~10⁻³). This split is expected, not a failure.
  5. No signal receives a σ label unless its statistic is calibrated by null simulations or by a conservative method shown valid for it. Near β = λ = 0 observables depend on the couplings only quadratically, so χ² tables and Fisher forecasts do not apply there.
  6. Either outcome is useful and is published here; no chain is dropped for disagreeing.

Expectations SHA-256: 0a8f3801badd39342b6c383f6dbf672ae6dbb52480fa5dd1ba2e8b228d3bcd68

SHA-256 of the expectations text: the statements above joined by line breaks, UTF-8. Fingerprinted September 30, 2026, before any chain. A fingerprint identifies exact text; it is not an external preregistration. If the expectations ever change, the old fingerprint stays in the status log.

The historical ratio test

Form agreed by both agents. The thresholds below are frozen in a separate declaration that Colson approved on October 1, 2026 (gate snapshot SHA-256 ba1ad4c1653f…). The approval authorizes no mocks, ratio fits or chains.

Historical prediction: β/λ = k_E·x_V ≈ 0.0499 — a direction θ₀ = arctan(0.0499) ≈ 2.9° in the (β, λ) plane, with θ defined mod 180°.

  • D0 — pre-run gate — Mocks through the real pipeline, β and λ varied jointly. Needs a discriminating outcome in ≥ 80% of θ₀-line mocks and ≤ 5% of θ = 0 mocks. If unmet, it is declared in advance that the run tests the general model, not the ratio.
  • D1 — signal gate — Profile-likelihood ratio against β = λ = 0, calibrated by null simulations (≈ 10⁴ fits for a 3σ-level tail) or a conservative method shown valid for this statistic. Otherwise no σ label, and the gate is not passed.
  • D2 — direction — 95% profile set of θ. Incompatible if θ₀ lies outside it. “Compatible with, and precise enough to discriminate, the proposed ratio” only if the set is one connected interval containing θ₀ and narrower than θ₀; otherwise uninformative.

Already known from local gravity: on the θ₀ line, Cassini excludes the whole tracking regime (which starts at λ ≈ 1.69, β ≈ 0.084). The weak segment, λ ≲ 0.05, survives as a valid drifting solution. Inside the frozen priors the θ₀ line leaves the β range at λ ≈ 0.20 — a prior edge, not a data limit.

Outcomes concern the ratio k_E·x_V only, never the full framework; an incompatible result cannot single out k_E. The ratio is never silently imposed or dropped. Under the frozen priors alone the 95th percentile of |β/λ| is 0.050, so a prior-dominated bound can look like 0.0499; such a bound is never quoted as agreement — the gates use profile likelihoods.

Source identified October 1, 2026: k_E·x_V ≈ 0.0499 appears in 32 of the 207 corpus papers (corpus as of August 3, 2026), mainly as a pulsar-timing slope (Papers LXXIII and CLXXVII) and as the Schumann term χ_S (Paper CL). No corpus paper writes it as β/λ: that identification was made in this test’s refinement on September 30 and is labeled a hypothesis. Colson approved testing it as a stated hypothesis on October 1, 2026. No D0 mock campaign is scheduled, so unless a bounded D0 check is separately approved and passes first, the run tests the general model, not the ratio. x_V (the historical x_m) is unsettled, and this page states no value for it.

Validation results so far

Run September 30 – October 5, 2026 in the cosmology project and reported by its agent; failed checks stay in this table. Plik is Planck’s high-ℓ likelihood; a precision step passes only if both the spectra and the likelihoods hold. The checks share machinery: complementary, not fully independent. No chains were started.
CheckWhat it comparesResultOutcome
Two-frame agreementIndependent Einstein-frame background calculation vs the matter-frame solverH(z) agrees to ≈ 1.6×10⁻⁷passed
Brans–Dicke equivalenceBuilt-in model with λ = −4β, ω_BD = 1/(4β²) − 3/2 (β ≠ 0)Agreespassed
Exponential quintessenceβ = 0 comparisonAgreespassed
Exact ΛCDMβ = λ = 0 at restReproduced, with no kinetic floorpassed
Power spectra, ℓ ≤ 100Spectrum comparisonMaximum difference ≈ 1.2×10⁻⁵passed
Normalization robustnessTensor A₀ = 1 (main) vs laboratory A₀²(1 + 2β²) = 1, at β = 0.01, λ = 0.5≈ 0.010% in the background (≈ β², as expected); ΔlnL = −0.00111 on the Planck low-ℓ temperature likelihoodmeasured
Unit tests13 tests, including regression tests for the BAO adapter fixAll passpassed
Coupling-edge backgroundsβ = ±0.01 with λ = 0, 1 and 2, against the independent Einstein-frame referenceAgree to ≈ 1.6×10⁻⁷passed
Initial conditionsInitial-velocity and single-clock initial-condition runsCompleted as specifiedpassed
Local-gravity termsScalar range (Compton length) in Earth-density matter at the Cassini edge and at the β prior edge; lunar-ranging and Cassini terms in the likelihood code≈ 264 AU and ≈ 66 AU — long-range at lunar-ranging scales, so local Ġ/G tracks the cosmic value; both terms wiredpassed
Inference precision, first ladderPlanck high-ℓ (Plik) log-likelihood over three adjacent precision levels, several controls changed per level; declared limit 0.05+0.112, then +0.096 (+0.208 in all) — not convergedfailed
Precision attributionThree solver paths, each at baseline and with five one-group refinements (18 evaluations)17 completed; the coupled point’s integration variant hit the solver’s minimum step. On the ΛCDM paths multipole sampling (+0.0725) and the lensing extension (+0.0308) leadmeasured
Single-control split17 evaluations, plus a fresh full-refined step against the archived value13 completed; the 4 failures all hit the minimum step, at β ≠ 0 with background tolerance 10⁻¹². Archived refined Plik value reproduced to 3×10⁻¹³measured
Precision, stage 1Candidate settings vs one step finer at 7 declared points; each point must keep spectra within 5×10⁻⁴ and |ΔlnL| below 0.05Stopped at point 1 (ΛCDM, action path): EE 9.16×10⁻⁴ and normalized TE 5.14×10⁻⁴ over the limit; Plik +0.0409 and total +0.0352 within itfailed
Single-control runs at the failed pointFresh baseline plus 15 one-control refinements at point 1 (ΛCDM, action path), each checked against the frozen limitsAll 16 completed and the baseline reproduced. Only the Bessel-function sampling exceeded a limit (EE 5.31×10⁻⁴); the two range controls moved Plik most (+0.0495 and +0.0215)measured
Precision follow-upUpdated candidate (Bessel sampling raised from 12 to 24 by the rule stated in advance) vs one step finer at 7 declared points, then a fit-region pointStopped at point 1 (ΛCDM, action path): Plik +0.0660 against the 0.05 limit; spectra, matter spectrum, distances, the other likelihoods and both totals passfailed
Range probe, first attemptEight declared runs at the failed point: fresh baseline, three one-control diagnostics and internal multipole ranges 5900 – 8000No comparison made: the fresh baseline ran past its 115 s limit. The solver had been pinned to one thread, a setting outside the approved plan; re-declared with 8 threads and revised limitsfailed
Range probeRe-declared run at the failed point: fresh baseline, three one-control diagnostics and internal multipole ranges 5900 – 8000, with 8000 a reference only; a range screens in if its next step moves Plik by less than 0.025All 8 runs completed. Plik steps +0.0658, +0.0446, +0.0312, +0.0225, each below the range predicted in advance; 7300 screens in. Baseline reproduced and all three diagnostics passed the frozen metrics. Screening only: nothing certifiedmeasured
Matched validation, first attemptThe screened 7300 design against the 8000 reference at all 7 declared pointsNo runs: a startup check meant for live runs was applied to the finished range-probe archive and stopped the campaign before its first evaluation. Corrected, then re-authorized as a new attempt with a fresh time limitfailed
Matched validationThe screened 7300 design against the 8000 reference at all 7 declared points (14 fresh runs), every frozen criterion at every point; a pass would only make 7300 eligible as a revised candidatePoints 1–5 passed every criterion, and point 1 reproduced +0.0225 exactly. The two steep corners (λ = 2) failed on Plik (−0.213 and −0.189 against 0.05) and on both totals; spectra, matter spectrum, distances and drag scale passed at every point. 7300 is not eligiblefailed
Identical rerunAll 14 matched-validation runs repeated under the pinned setup, one slot per commandAll 14 bitwise identical to the archive; 28 of 28 normal exits in 26.3 min. Normal path only: crash, interruption and restart paths are not yet exercisedpassed
Stall diagnosis (Step I)At β = 0.01, λ = 0.5: the completing control, the earlier start (a = 10⁻¹¹) and the tighter background tolerance (10⁻¹²), each in a plain copy and a logging-only copyValid: 6 of 6 evaluations, all 3 paired checks passed. The stalls come from a cancellation in the scalar’s acceleration line (radiation-sized terms ≈ 1.2×10⁸ times the result) meeting a starting velocity of exactly zero, not from fast physicsmeasured
Repair test (Step R)At β = 0.01, λ = 0.5 with the cancellation-free line: the control, the earlier start (a = 10⁻¹¹) and the tighter background tolerance (10⁻¹²), each in a plain copy and a logging-only copy, against pass marks fixed in advancePassed: 6 of 6 evaluations, all 3 pairs bit-for-bit identical; old-vs-new identity within 5.1–5.3 of 10 rounding units; neutrino trace within 10⁻¹⁵ (limit 10⁻¹²); below redshift 10, H within 1.1×10⁻¹¹ and the field’s velocity within about 10⁻¹⁰ of the archived control (limits 10⁻⁶). One point only: not precision certificationpassed

Two agents, one record

  • Colson — Brings the questions and the direction, relays each step between the agents, and decides.
  • This project’s agent — Worked the mathematics one step at a time in chat: the physical-margin argument, the lunar-ranging term and the with/without comparison; later reviewed each precision report and draft declaration.
  • The cosmology agent — Works in Colson’s separate cosmology project with hi_class: reviewed each step and supplied the engineering, validation and preservation, then built and validated the solver, froze the run specification and runs the precision checks.
  1. A step of mathematics is worked in chat.
  2. Colson relays it to the cosmology agent.
  3. The review comes back: accepted, strong warning, or still open.
  4. Valid gaps are accepted and each withdrawal is named.
  5. Only what survives accumulates into the next step.

“recursive accumulation ♾️ in action between 2 agents. LITERALLY.”

— Colson

“stick to what you know… we don't need to solve the whole universe in one go”

— Colson

“stop trying to prove every screened branch fails before building a working test”

— The cosmology agent

The last two arrived together. Unresolved calculations became open limitations instead of blockers.

The two agents independently reached the same minimum — the same frozen model, the unscreened scope, Cassini labeled unscreened — then supplied complementary halves.

Withdrawals

Withdrawals are part of the product. Each item below was dropped or narrowed in review, and each drop made the surviving test more trustworthy. The test does not depend on any of them.

  • The Cassini window — γ ≈ −4β_Sβ_P is not a Cassini measurement model.
  • “Atoms are never screened” — Too absolute.
  • The 50% redshift figure — Dropped in review.
  • The 10⁻⁸ bound — Dropped in review.
  • The 0.0475 slope — Corrected: the slope in measured density is s̃ = β/(λ + 4β) ≈ 0.0416 (ρ̂ = A³ρ̃).
  • GPS Δln A ≈ 1.5 — Assumed an unsolved environmental profile; 1.5 is not perturbative.
  • The 9.6% solar-line figure — Same unsolved-profile assumption; line shifts also need the matter-frame redshift.
  • “Outside classical treatment” above N ~ 10⁴⁴ — A sub-spacing Compton length calls for a microscopic treatment; M_C is a normalization scale, not the scalar’s mass.
  • An unconditional order-β² mass correction — Narrowed to a field that has not settled (m² ≈ V″ + β²ρ/M_Pl²).
  • “β = 1.92 has no matter era” — Wrong: the steep potential makes it track.
  • “Plain ρ − 3p cancellation fits the stall’s step pattern” — Too loose: it predicts only a sixth to a third of the observed step. The located line cancels terms about four times radiation’s size, and that matches to 8 digits.
  • A few-times-10⁻⁸ spot-check band for the repair — Assumed the saved states meet the Friedmann constraint to about one rounding unit; they carry a residual near −3×10⁻¹⁵, which the old line amplifies. Replaced by two checks with limits fixed before any data.
  • A two-flag compiler guard (-march=, -mfma) — Incomplete: other flags also switch on fused multiply–add (-mavx512*, -mavx10*) or change rounding (-mfpmath=387), and Nix can pass flags through variables the guard does not read. None is in this build, and any effect would fail a check rather than pass one. A future guard should allow only the recipe’s own flags.

The earlier run, kept separate

An earlier joint run used a different setup: the propto_omega parameterization with a separate ΛCDM expansion, independent α_B and α_M, and input c_T = 0.001 — so α_T(a) = 0.001·Ω_smg(a), about 7×10⁻⁴ today (estimate), far above the GW170817 bound under the standard reading. Stability tests were off, so some samples may be inadmissible.

It is a different test. Its chains are kept separately and are never resumed or merged with this one.

Preservation

  • Configuration and scripts frozen; resolved settings and software versions saved.
  • Every chain kept; weighted summaries reported with bulk and tail convergence.
  • Disagreeing chains are never silently dropped.
  • Failed checks and withdrawn steps stay on this page.
  • Settings outside an approved declaration are not authorized. Resource limits change only through a new declaration made before any numbers exist; acceptance criteria never change after results.

Status log

  1. — Result of the repair’s test run: passed. All six evaluations completed, including the earlier start and the tighter background tolerance that used to stall, in 472 s of the 3600 s allowed, and every pass mark fixed in advance held. Each logging copy matched its plain copy bit for bit. In the earliest steps the old line equaled the new one plus the written-out correction to within 5.1–5.3 of the 10 allowed rounding units. The neutrino trace matched the solver’s own ρ − 3p to within 10⁻¹⁵, against 10⁻¹². Below redshift 10 the repaired control matched the archived control to 1.1×10⁻¹¹ in H and about 10⁻¹⁰ in the field’s velocity (scaled), against 10⁻⁶ for each. This project’s agent reviewed the full results package: every file matches its recorded fingerprint, the run used exactly the reviewed declaration and test inputs, all 14 commands exited normally, and none of the flags missed by the incomplete compiler guard appear in the actual build commands. Two estimates stated before the run, scored: the worst identity case, put at about 5 of the 10 units, came in at 5.1–5.3; the tight case’s trace, put at 0.6–0.9 GB, was 0.18 GB, about four times smaller, so its 1 GiB cap was never close. The largest Friedmann-constraint residuals (about 10⁻⁶) all occurred in deliberately perturbed states used to estimate the solver’s Jacobian; they are reported as required, with no pass mark attached. The pass covers β = 0.01, λ = 0.5 only: not precision certification, convergence, other points in the prior box or chain approval. The repaired line exists only in the run’s solver copies so far.
  2. — Colson gave the go for the repair’s test run, and the cosmology agent is running it under the reviewed, hash-bound declaration. It is the bounded background-only run described in the runtime plan below, not chains. Its result will be published here whatever it shows.
  3. — Logging code and execution-ready declaration for the repair’s test run reviewed together by this project’s agent: accepted with no required fixes (declaration SHA-256 673db397…). The physics change is exactly the accepted repair line, applied only to the run’s two fresh solver copies, and every pass mark in the runtime plan is wired in as declared. This project’s agent also checked the logging code’s compensated arithmetic with an exact-arithmetic copy of its own: the working sums stay far below their 255-term limit, the correction is good to about one part in 10¹⁶, and the worst case is estimated at about 5 of the 10 allowed rounding units — an arithmetic check, not a measurement. Two notes did not block the run. The compiler-flag guard this project’s agent asked for is incomplete (see Withdrawals); none of the missed flags is in this build, and any effect would fail a check rather than pass one. The trace for the tight case is estimated at roughly 0.6–0.9 GB against its 1 GiB cap. Stated before the run: a stop at that cap would be a resource stop, not a result about the repair, and changing the cap would take a new declaration, like any other resource limit.
  4. — Public record brought up to date through the repair’s declared test run. The expectations and their fingerprint are unchanged.
  5. — Runtime plan for the repair’s test run accepted by this project’s agent with three additions; not yet run. Six background-only evaluations at β = 0.01, λ = 0.5: the control and the two settings that stalled, each solved in a plain copy and a logging-only copy, both carrying the repair, inside Step I’s unchanged limits (14 one-command slots, 3600 s overall). Pass marks fixed before any data: all six complete; each logging copy matches its plain copy bit for bit; in the earliest steps the old line equals the new one plus the written-out correction to within 10 rounding units of that call; the massive-neutrino trace used by the new line matches the solver’s own ρ − 3p to 10⁻¹² at every saved row below redshift 10; and below redshift 10 the repaired control matches the archived control to 10⁻⁶ in H and in the field’s velocity (scaled to its largest value there) — the stretch of history the chains rely on. The additions were that last mark, a compiler-flag guard that protects the bit-for-bit match, and using the same stored density the old line reads. Any failure ends the run, with its evidence kept. The logging code is written next and reviewed once before Colson’s go. A pass would establish only this repair at this point, not precision certification.
  6. — Repair accepted as source, as the basis for that plan. Only the scalar’s acceleration line changes, and only for this model: it is rewritten in an equivalent form built on the matter trace ρ − 3p, so nothing radiation-sized has to cancel. Radiation drops out analytically, and the massive-neutrino trace comes from its own momentum sum. The two forms agree exactly when the Friedmann constraint holds; off it they differ by a written-out term, 12βa²C/e, where C is the constraint residual and e = A⁻² = exp(−2βφ/M_Pl). This project’s agent checked the derivation two ways — through the Einstein-frame equations and with exact rational arithmetic — and estimated analytically that the change in the residual’s growth rate stays below 0.09 per e-fold anywhere in the frozen prior box (for negative β it can be slight growth rather than damping). No tolerance, minimum step, start time, parameter or grid changes. One spot check failed because of this project’s agent’s own band, which assumed the saved states meet the constraint to about one rounding unit; they carry a residual near −3×10⁻¹⁵, which the old line amplifies. The band was withdrawn, and two checks with limits fixed in advance replace it.
  7. — The cosmology agent confirmed the spot in the solver’s source: the scalar’s acceleration line in hi_class’s background code, solved by Cramer’s rule. There two radiation-sized terms nearly cancel — about 1.2×10⁸ times the surviving matter term at the earlier start (≈ 4ρ_r/ρ_m) — and the step pattern predicted from that matched the traces to 8 digits. The cancellation-free form agrees with this line only on the Friedmann constraint, because the solver evolves H separately; this project’s agent chose it anyway, since keeping the off-constraint term means computing exactly the radiation-sized difference that causes the noise.
  8. — Stall diagnosis (Step I) completed and valid: 6 of 6 evaluations, all 3 paired checks passed, every package hash verified. The stalls at the earlier start (a = 10⁻¹¹) and the tighter background tolerance (10⁻¹²) are a numerical artifact of the first steps, not fast physics; the other variables change at physical rates. The field’s velocity starts at exactly zero, so the solver’s error test weighs it against a fixed 10⁻¹⁵ floor, while its rate of change comes out in coarse jumps — exact whole multiples of one quantum, far above rounding level — the signature of a large cancellation. Those jumps reproduce the tight case’s rejected-step error ratio, 10.02, to four digits. Smaller steps are not available (the minimum step is about 11 rounding units of ln a), so the agreed repair removes the precision loss at its source, with no tolerance, minimum-step or start-time change.
  9. — Two more slips before the valid run: a report reading a value no longer written, caught by the cosmology agent before launch; then a post-calibration replay that passed the solver a setting it rejects, after the control calibration had converged in 14 calls (residual 6.2×10⁻¹²). Neither had been caught in review. Both fixes were reviewed and run under a fresh approval from Colson.
  10. — Stall diagnosis (Step I) approved by Colson: six background-only evaluations at β = 0.01, λ = 0.5 — the control that completes plus the earlier-start and tight-tolerance settings that stall, each in a plain copy and a logging-only copy — in 14 one-command slots under a 3600 s ceiling. Three attempts stopped safely on setup slips before computing anything: a memory-limit setting under the wrong name (after 6 ms, before any build); an inherited compiler setting that the build guard refused; and, after the first full build passed its compiler audit, a mismatch between the compiler’s system library and Python’s that stopped the run at import. Review by this project’s agent had caught none of the three. The fix pins GCC 15.2, whose library is the one Python loads; the archived build had used GCC 14.3. Each stopped attempt is kept, and each fix was reviewed and freshly approved.
  11. — 14-case identical rerun, approved by Colson: all 14 matched-validation runs repeated under the pinned setup, one slot per command inside the agent tool’s 300 s limit. All 14 outputs were bitwise identical to the archive, with 28 of 28 normal exits in 26.3 minutes. This shows repeatability on the normal path only — crash, interruption and restart paths are not yet exercised — and the 7300 design stays failed. Along the way a false stop from the harness’s own 50 ms timer (exit 142 after the work had finished) was found, reproduced and fixed.
  12. — Route chosen with Colson: build the test around checks of what should stay the same — identical reruns, checkpoint and resume, agreement between independent chains — each with a tolerance declared in advance, rather than reconstructions of past snapshots. In Colson’s words: “let’s get the test built just that way.” A records-only reconstruction of the lensing calculation (R0), begun as a yardstick, became optional: the cosmology agent found no material chain risk that only it would catch. Agreed order: the identical rerun; stall diagnosis (I) and repair (R); a revised seven-point precision design and certification; chain software and preflight; then the pilot, which is the first chains. The ceilings for I through certification total about 30 days — limits, not forecasts.
  13. — Matched validation run as authorized: 14 fresh runs and 7 comparisons in 22 min 48 s of the 55-minute limit, with no retries, solver failures or resource stops. Points 1–5 passed every frozen criterion. The two steep corners (λ = 2, β = −0.01 and +0.01) failed: Plik moved −0.213 and −0.189 from 7300 to 8000, against the 0.05 limit, and both totals failed their 0.1 limit. Spectra, matter spectrum, distances and drag scale passed at every point, each at least 17 times inside its limit. So 7300 is not eligible as a revised candidate. Point 1 reproduced the archived +0.0225 exactly. The spectrum change was similar in size at every point — about two-thirds of point 1’s at the steep corners — but the likelihood shift also depends on how the change lines up with each point’s misfit to the data: at the steep corners, which fit Planck poorly at the fixed benchmark parameters, that alignment is 12 – 15 times larger and of opposite sign, so the shifts were 9.4 and 8.4 times point 1’s, both negative. Estimates recorded in advance by this project’s agent: the steep corners had the largest shifts and were the only points over 0.05, as estimated, but their negative sign was not predicted; the size of the spectrum change stayed within a factor of 2 of point 1’s at all seven points; pairs differing only in β agreed within 25% at points 4/5 and 6/7 but not 2/3 (53%); of the numeric ranges, only point 3 fell inside — point 2 below its range, points 4 and 5 above. The saved evidence passed its audit; no further run was started.
  14. — First matched-validation attempt stopped before any evaluation: a startup check meant for live runs was applied to the finished range-probe archive and halted the campaign — an implementation error that failed safe. The cosmology agent corrected it, and Colson authorized a new attempt with a fresh 55-minute limit; the stopped attempt is kept with its evidence. The rules for scoring this project’s agent’s estimates were frozen before the first evaluation.
  15. — Matched validation declared and approved by Colson, with four revisions from this project’s agent; the cosmology agent confirmed nothing else changed. The screened 7300 design runs against the 8000 reference at all seven declared points — 14 fresh runs, 7 comparisons — under the frozen criteria (frozen_inference.json, SHA-256 8404307c8d59…); the intended eighth, fit-region point was never established and is not added. A solver failure fails only its point; timeouts and other technical stops end the campaign. Limits: 190 s per run, 10 GiB per run, 12 GiB shared, 55 minutes overall. A pass would only make 7300 eligible as a revised candidate, needing its own declaration and the unchanged suite; 0.025 is not a pass mark. Advance estimates by this project’s agent, not criteria: point 1 repeats +0.0225; points 2–3 between +0.012 and +0.033; points 4–5 within about 15% of point 1’s value; points 6–7 the largest shifts, furthest from point 1 and likeliest to exceed 0.05, sign not predicted; pairs differing only in β within 25% of each other; the size of the spectrum change within a factor of 2 of point 1’s.
  16. — Range probe run as re-declared: all 8 runs and 11 comparisons completed in 539 s, with no stops, retries or extra runs. Plik moved +0.0658, +0.0446, +0.0312 and +0.0225 over the steps 5200 → 5900 → 6600 → 7300 → 8000 — positive and shrinking — and over ℓ 30–2508 each step kept close to the declared shape (centered R² 0.985 for the first TT step and at least 0.998 after), with no growing departure. Range 7300 is the lowest that screens in: its next step, +0.0225, is under 0.025. That selects a design to test and certifies nothing. All four steps fell below the ranges this project’s agent predicted in advance: the trend held, the numbers did not. The fresh baseline reproduced in 34.7 s, against 32.1 s before, which supports the thread-count explanation of the earlier timeout; all three one-control diagnostics passed the frozen metrics. The added high-multipole tails account for 85–89% of each step’s increase in lensing-deflection variance, and the common range changed too. Peak memory was 7.38 GiB per run (8.12 GiB shared, sampled); live thread counts ranged from 1 to 18 with the pool set to 8. Nothing beyond the approved probe was started.
  17. — Range probe re-declared with revised resources only, and approved by Colson; not yet run. The solver’s built-in thread pool is set explicitly to 8 threads, one per available CPU, and each run’s live thread count is logged with its memory. Time limits come from the measured 32.1 s candidate run: 2.5 × 32.1 s plus a 60 s margin, rounded up, gives 145 s per run (150 s supervised) and 30 minutes overall. Runs, order, memory limits, criteria, predictions and the 0.025 screen are unchanged; the stopped attempt is not resumed.
  18. — Preflight with no solver runs: the compiled solver matches the build behind the 32.1 s candidate run, and the workspace has 8 CPUs. The solver uses its own C++ thread pool, sized by OMP_NUM_THREADS, not OpenMP. Eight threads was the likely earlier default — an inference, not a recorded measurement — so the one-thread setting is the leading suspect for the timeout, unconfirmed.
  19. — First range-probe attempt stopped by its own time limit: the fresh baseline ran past its 115 s limit, so no run completed and the other seven were not attempted. Memory peaked at 0.53 GiB, far below the ≈ 2.45 GiB lensing table, consistent with the solver not yet reaching lensing. The implementation had pinned the solver to one thread, a setting not in the approved text, as the cosmology agent acknowledged. That attempt is closed with its evidence kept; the cause of the timeout was not measured.
  20. — Range probe declared and approved by Colson: eight runs at the failed point — a fresh baseline; three one-control diagnostics (Bessel sampling doubled, the Limber wavenumber coefficient doubled, the lensing quadrature offset raised from 70 to 770); and internal ranges 5900, 6600, 7300 and 8000, with 8000 a reference only. A range screens in only if its next step moves Plik by less than 0.025; the screen selects a design to test and replaces no frozen criterion. Predicted Plik steps, recorded before any run by this project’s agent as conditional estimates, not criteria: 5200 → 5900 +0.066 to +0.067; 5900 → 6600 +0.046 to +0.053; 6600 → 7300 +0.033 to +0.043; 7300 → 8000 +0.025 to +0.035.
  21. — Source reading in the cosmology project, with no solver runs: both range controls set one internal multipole range L — the requested maximum plus its lensing margin — and L drives several calculations at once: the lensing-potential tail, the lensing quadrature and, through a full-Limber calculation this build enables by default, the wavenumber coverage of the potential spectrum. Reading the source cannot rank them, so the probe measures their combined effect. The lensing table alone needs about 2.45 GiB at L = 5200 and 5.77 GiB at L = 8000, so the probe also declares memory limits.
  22. — Review of the saved follow-up outputs by this project’s agent, accepted by the cosmology agent as useful evidence without repeating it: two evaluations with identical settings gave bitwise-identical spectra, and the two range controls shift Plik with aligned shapes — strong suspects, not an established mechanism. Agreed order: understand the source first, then declare a bounded probe, then test any revised candidate against the unchanged suite.
  23. — Bounded precision follow-up run as approved. 16 fresh runs at the failed point — a baseline plus 15 one-control refinements — all completed; only the Bessel-function sampling exceeded a limit (EE 5.31×10⁻⁴), so the rule stated in advance raised it from 12 to 24 in the updated candidate. That pair stopped the campaign at point 1 (ΛCDM, action path): Plik moved +0.0660 one step finer, over the 0.05 limit, while the spectra, matter spectrum, distances, the other likelihoods, both totals and the reconstruction check passed — a numerical-precision failure, not a physics result. One run took 32.1 s with the updated candidate and 145 s one step finer. 18 of 56 slots completed; the rest were not attempted, with no retry.
  24. — Colson approved the bounded precision follow-up exactly as drafted, and the separate 0.0499 declaration with one provenance correction: the page check it cites was this project’s agent reading the page source, relayed by Colson. The declaration freezes the D0–D2 gates (snapshot SHA-256 ba1ad4c1653f…) and authorizes no mocks, ratio fits or chains; with no D0 campaign scheduled, the run is declared in advance to test the general model, not the ratio.
  25. — Public record brought up to date: the frozen specification, the later validation and every precision result so far. The expectations and their fingerprint are unchanged.
  26. — Pilot plan drafted with its nonlinear check: every run stays linear, and the declared Halofit sensitivity is measured by reweighting pilot samples (at most 512 extra evaluations). It is approved together with the pilot.
  27. — Separate 0.0499 declaration drafted for Colson’s approval: the identification β/λ = k_E·x_V becomes a stated hypothesis, with the D0–D2 gates frozen under their own fingerprint. No D0 mock campaign is scheduled, so unless a bounded D0 check is separately approved and passes first, the run is declared in advance to test the general model, not the ratio.
  28. — Corpus source of the 0.0499 prediction counted against the August 3, 2026 corpus: 32 of 207 papers, mainly as a pulsar-timing slope and the Schumann term χ_S. No corpus paper writes it as β/λ.
  29. — Bounded precision follow-up drafted for Colson’s approval: fresh single-control runs at the failed point, a rule stated in advance for updating the candidate settings, the seven declared points again, then a ΛCDM fit-region point. At most 56 evaluations and 186 minutes, in the foreground.
  30. — Precision stage 1, approved by Colson as declared: candidate settings against one step finer at seven points. It stopped at point 1 (ΛCDM, action path): EE 9.16×10⁻⁴ and normalized TE 5.14×10⁻⁴ exceed the 5×10⁻⁴ spectrum limit, while Plik (+0.0409), the total likelihood (+0.0352) and the reconstruction check (1.7×10⁻¹²) pass. One evaluation took 33.5 s with the candidate settings and 166 s one step finer. The other 12 slots were not attempted.
  31. — Single-control split (17 evaluations): 13 completed; the 4 failures all hit the solver’s minimum step — every β ≠ 0 point tried fails once the background tolerance is 10⁻¹². A fresh full-refined step reproduced the archived refined Plik value to 3×10⁻¹³.
  32. — Decomposition of saved outputs, with no new solver runs: the same numerical change can fail the 0.05 likelihood limit at one point and pass at another, depending on how it aligns with that point’s residuals — so every declared point is checked on its own.
  33. — Precision attribution (18 evaluations across three solver paths): 17 completed; the coupled point’s integration variant hit the solver’s minimum step. On the ΛCDM paths, multipole sampling (+0.0725) and the lensing extension (+0.0308) lead the Plik shifts.
  34. — First inference-precision ladder failed its declared limit: Plik moved +0.112, then +0.096 over adjacent precision levels (+0.208 in all, against 0.05) — not converged. Starting the field at an earlier epoch hits the solver’s minimum step; undiagnosed.
  35. — Further validation passed: 13 unit tests, coupling-edge backgrounds, initial-condition runs and the local-gravity terms. A fault in the BAO likelihood adapter was found and fixed, with regression tests; total likelihoods computed before the fix are void, while Plik comparisons stand.
  36. — Run specification frozen with Colson’s approval: Planck 2018 (full Plik TT/TE/EE, low-ℓ TT and EE, lensing), DESI 2024 BAO and Pantheon+ without SH0ES; β uniform on [−0.01, 0.01] and λ uniform on [0, 2]; linear treatment with a separate Halofit check; DESI’s second data release declared as a separate follow-on before any result.
  37. — Public record opened, with the expectations fingerprinted before any chain.
  38. — Lunar-ranging value checked against the published abstract of Hofmann & Müller 2018: Ġ/G₀ = (7.1 ± 7.6)×10⁻¹⁴ yr⁻¹.
  39. — λ prior form set by the model’s symmetry: λ ≥ 0 with β of either sign; the upper edge must be wide enough that the data set the limit.
  40. — First validation suite passed in the cosmology project. No chains started.
  41. — Solver model and independent Einstein-frame reference built in the cosmology project.
  42. — Scope, matter-frame model and validation plan agreed by both agents.

Next: a revised precision design that also holds at the two steep corners, declared and approved before any run and tested at all seven declared points under the unchanged criteria; a design that passes then needs its own declaration and the unchanged precision suite. Then chain software and preflight, and the pilot — the first chains — once both pending items are closed.

Open limitations

  • No chains have run. Nothing here is yet a result about whether the coupled scalar exists.
  • The model is a declared candidate, not claimed as uniquely derived by the corpus.
  • The screened branch is deferred: it needs a self-consistent environmental profile and laboratory-G calibration. Exhaustive exclusion of screened points is still open, and the cosmology agent’s strong warning stands that exponential screening may produce unacceptable environmental effects.
  • The safety margins above (thin-shell onsets, the G_lab correction) are order-of-magnitude estimates, not bounds.
  • Inference precision is not yet certified, and the linear matter treatment is a declared approximation, not a certification that nonlinear effects are negligible.
  • The ratio test’s thresholds are frozen in an approved separate declaration, but no D0 mock campaign is scheduled: unless a bounded D0 check is separately approved and passes first, the run tests the general model, not the ratio.
  • The historical 0.0499 prediction: its corpus source is identified, but no corpus paper writes it as β/λ. That identification is a hypothesis made in this test; Colson approved testing it as a stated hypothesis on October 1, 2026. No x_m value is stated here.
  • The validation checks are complementary, not fully independent.
  • The repair passed its test at one coupled point (β = 0.01, λ = 0.5) in three solver settings. That establishes the repair at that point and that its logging changes nothing, not precision certification, convergence or behavior elsewhere in the prior box, including negative β.