The Ground Truth Engine · Verifiable research

The harder you look, the more it checks out.

Most claims degrade under scrutiny — dig, and you find the overclaim or the number with no denominator. These do the opposite. Every result below is led by an honest claim (a bound, a null, a first measurement — never a hyped headline) and carries a VERIFY THIS panel: the sealed pre-registration that fixed the method before the data, the Bitcoin timestamp that proves when, and — where it applies — a standalone script you run yourself.

01

Sealed pre-registration

A SHA-256 of the frozen method, statistic, and pass criterion — locked before the data was scored.

02

Bitcoin anchor

That hash is timestamped into the Bitcoin blockchain. The block height is proof of when, independent of us.

03

Published nulls

We publish our failures with the same proof as our results. A group that hides them degrades under scrutiny.

04

Run it yourself

Where a result is a record, a stdlib-only script — none of our code — recomputes it. Do not trust us; check.

Results & bounds

What we can stand behind

Each is a first measurement, a bound, or an auditable record — labeled with its metric and sample, and carrying its proof.

Solid-Earth physics · first measurement

The first observational constraint on the l=4 tidal-gravity admixture

Using 21 superconducting-gravimeter stations, we place the first observational bound on the degree-4 (l=4) admixture in the M2 gravity tide predicted by Wahr (1981): epsilon_4 = +0.076% +/- 0.30%, consistent with zero. |epsilon_4| < 0.675% at 95% confidence -- a bound that sits just below Wahr's predicted 0.7%.

epsilon_4 = +0.076% +/- 0.30% | |epsilon_4| < 0.675% (95%) | NULL_BOUND, not a Wahr detection
21 superconducting-gravimeter stations (IGETS network)
Positive control

Signal injection recovers +/-0.7012% for an injected +/-0.7% (unit gain); a negative fixture correctly recovers 0.20%; the shuffle-null is alive. 3 of 3 pre-registered controls pass.

What would have falsified it

Wahr's predicted 0.7% sits at 2.34 sigma from our estimate -- had the admixture been that large, the weighted fit would have shown it. The bound is falsifiable at exactly the value of interest.

Verify this
Sealed pre-reg
92cf0cc31399907f847dcfe85c864bbb337de7a3ce6549d6617977cbee6c10cb
locked 2026-08-29 · sealed before scored
Bitcoin anchor
block 964512block 964513block 964545 confirmed
Anchor file
hypotheses/locked/20260829_l_4_wahr_1981_admixture_in_the_m_2cb6bcb0.canonical-sha256.ots
Run it yourself
python scripts/verify_hypothesis_anchors.py
Data & code
scripts/igets/l4_sealed_scored_fit.py (seed 20260828); IGETS SG network, 21 stations; TPXO10v2a + FES2014b ocean-loading models

Why it matters. A physical constant nobody had measured. It is a finding, not a method -- safe to publish, and the Bitcoin anchor already gives quiet priority. TPXO10v2a ocean loading adjudicated the one outlier station (Onsala) rather than excluding it.

Fundamental physics · pre-registered line search

A dark-matter line search we refused to call a detection

A pre-registered search for a sidereal-frequency magnetic line across 11 INTERMAGNET observatories returned a bar-passing candidate at a 95% upper limit of 1.711 nT. We did NOT call it a detection: the ionospheric K1 tide sits at exactly the sidereal frequency and our resolution cannot separate the two. The sealed pre-registration committed IN ADVANCE that any such line stays a CANDIDATE pending a separate adjudication -- never a detection from this analysis alone.

95% upper limit 1.711 nT | SEALED CANDIDATE -- explicitly not a detection
11 INTERMAGNET observatories (coherent CoSMoS stack)
Positive control

Injection gate passes; split-halves 4 of 4 (stations even/odd, time first/second half, amplitude ratio + phase agreement); empirical p = 0.005 against a 200-point dummy-frequency grid.

What would have falsified it

The pre-registration pre-committed the K1 cap: a bar-passing sidereal line is a candidate, not a detection, until a separately sealed K1 adjudication clears it. A downgrade to a plain null still yields the published 1.711 nT upper limit -- so the result is honest in either direction.

Verify this
Sealed pre-reg
5090ff310f2d1c1dfc3ea0292491e881c6c364a95ebefb583e4c624f2e57e424
locked 2026-08-07 · sealed before scored
Bitcoin anchor
block 961369block 961372block 961402 confirmed
Anchor file
hypotheses/tested/20260807_dm_magnetometer_channel_v1_sider_296ee71c.json.ots
Run it yourself
python scripts/verify_hypothesis_anchors.py
Data & code
scripts/dm_mag_channel_v1.py (sha256 7db981e1...); sealed prereg strategy/dm_mag_channel_prereg_v1_SEALED.md (sha256 3fcad0c7...); INTERMAGNET raw 1-min B-field, 11-station curated subset

Why it matters. This is the anti-crank posture in one result: a signal that passed every pre-registered bar, and we still refused to overclaim it because a mundane ionospheric line sits at the same frequency.

Forecasting · run-it-yourself verifier

A prediction ledger you can verify yourself

Every forecast is written to a SHA-256 hash-chained ledger before the event, so no entry can be edited, inserted, or reordered after the fact. A standalone script that imports ONLY the Python standard library -- none of our code -- recomputes the whole chain against a Bitcoin-anchored root. We publish COVERAGE (how much observed seismicity fell under an alarm), never a skill headline; our measured skill, including nulls, is reported separately.

SHA-256 hash-chained ledger | COVERAGE reported, not skill | chain root Bitcoin-anchored
Public prediction ledger + ComCat ground truth
Live ledger: 11,000+ sealed predictions · 7,600+ matched to real events · 12,000+ seismic stations. Match counts are COVERAGE, not a skill score.
Positive control

The verifier ships a self-test that proves it DETECTS tampering -- an edited field, a deleted entry, and a forged insert carrying its own valid hash (caught by the chain link). A verifier that always says OK would launder exactly what it should catch.

What would have falsified it

Skill is a separate question, scored separately. Our own measurement: no alarm-volume skill (Area Skill Score 0.4877 vs 0.5 for none) and modest out-of-sample probabilistic skill (pooled Brier Skill Score +0.055). Both are published, nulls included -- see /v2/numbers.

Verify this
Bitcoin anchor
chain-root anchored (see verifier)
Run it yourself
python scripts/verify_public_ledger.py deepmap_ledger_YYYY-MM-DD.json
Data & code
scripts/verify_public_ledger.py (stdlib only); scripts/publish_ledger_snapshot.py; ground truth ComCat via common/fdsn_multi.py

Why it matters. Until now, verifying our record meant paying us or asking us. This makes it: don't trust us, run the script. Read it, re-implement it in any language, and you should get the same answer.

The null museum

The failures we publish on purpose

A journal will not carry these. That is exactly why they belong here. Each was sealed and Bitcoin-anchored before it was scored, so the timestamp proves when we knew — including the one we sealed, tested, and then retracted after our own verification found it was wrong.

Null museum · Solid-Earth physics

We searched for the Slichter mode and found nothing -- cleanly

The Slichter mode -- a translational oscillation of Earth's solid inner core, never definitively detected -- searched with a 20-pier coherent superconducting-gravimeter stack over 303 station-years. Null. The sealed pre-registration's primary prediction held and our armed false-alarm flag did not fire; the estimator is confirmed off every one of 38 masked tidal lines.

NULL | methods-clean | estimator off all 38 masked tidal constituents
20 SG piers / 36 segments / 303 station-years (IGETS)
Positive control

The fix is estimator-level and the test was forward-blind: the sealed pre-registration armed a red flag (KSTAR_STILL_ON_A_LINE) that would fire if the search were still sitting on a tidal line. It was graded AFTER the OTS seal and did not fire.

What would have falsified it

A real Slichter triplet would place the detection statistic ON a predicted split frequency; ours sits 31.4 milli-cycles/day off the nearest masked line edge -- exactly where a genuine mode would NOT be.

Verify this
Sealed pre-reg
d882abb27233ef67752ba80c93e721d89f1877b7c1f00348816ec415b3582c73
locked 2026-08-29 · sealed before scored
Bitcoin anchor
block 964581 confirmed
Anchor file
hypotheses/tested/20260829_slichter_1s1_search_v9_line_mask_8579f6b3.json.ots
Run it yourself
python scripts/verify_hypothesis_anchors.py
Data & code
IGETS SG network; cleaned store igets_cleaned_v3sol (20 piers / 36 segments / 303.3 station-yr, 38 constituent lines); trackAB v9 caller, estimator sha b4089bce

Why it matters. Same instrument and program as the l=4 first-measurement above. Here is where it returned nothing -- and we sealed and anchored the null just as carefully as a positive result.

Null museum · Fundamental physics

A clean null on ultralight scalar dark matter

A coherent 21-station superconducting-gravimeter stack searched 50 candidate masses for an oscillating scalar-dark-matter coupling across the band 1 microHz to 8.3 mHz (4.56 million frequency bins, 10,000 permutations). Clean null: no candidate crossed the pre-registered, look-elsewhere-corrected detection threshold.

NO_CANDIDATE_LIMITS | 50 masses | LEE-corrected threshold S >= 4494.87
21 IGETS SG stations · 4.56M frequency bins
Positive control

The unblinding seed was fixed to the hash of a Bitcoin block (961333) that was mined AFTER the analysis was frozen -- so the randomization provably could not have been chosen to flatter the result.

What would have falsified it

Any candidate mass whose coherent amplitude exceeded the sealed look-elsewhere threshold (S >= 4494.87) would have been reported; none did.

Verify this
Sealed pre-reg
972dcbd603607cfa5d96a6627e5365451ad4c5bac4be89cc09f2d1f65b4ee367
locked 2026-08-07 · sealed before scored
Bitcoin anchor
block 961369block 961372block 961402 confirmed
Anchor file
hypotheses/locked/20260807_ultralight_scalar_dm_oscillating_4d41c614.json.ots
Run it yourself
python scripts/verify_hypothesis_anchors.py
Data & code
scripts/nobel_shots/dm_oscillating_g_v2_run.py (commit 276f91e3); IGETS Level-2 CORMIN 1-min SG data, 21 stations >= 5 yr; unblinding seed data/dm_v2_seed.json (BTC block 961333 hash)

Why it matters. A null on a frontier search, done with the discipline that makes a null trustworthy -- a sealed threshold and a randomization seed nobody could game.

Null museum · Earthquake precursors

Solar high-speed streams do not trigger big earthquakes

We tested whether coronal-hole high-speed solar-wind streams elevate the global M6.5+ earthquake rate within 48 hours. Over 1996-2025 (31 stream onsets; 1,310 declustered M6.5+ events) we observed 6 events in the exposure windows against a null expectation of 7.4 -- below chance. Null: no hint of elevation. The bound excludes a rate lift of 1.76x or more at 97.5%.

NULL | observed 6 vs null 7.4 | excludes rate lift >= 1.76x (97.5%)
31 HSS onsets · 1,310 declustered M6.5+ events (1996-2025)
Positive control

The null is a circular-shift surrogate that preserves the earthquake catalogue's own temporal structure; the observed count sits inside its 95% CI [2, 13], below the mean.

What would have falsified it

Had solar streams triggered quakes, the exposure-window count would have exceeded the shuffle null's upper tail; instead it fell below the mean.

Verify this
Sealed pre-reg
06cafff0efa2d6b2e2cef171e54d7fa34190304747f94666b755b3ec42bc3e03
locked 2026-08-15 · sealed before scored
Bitcoin anchor
block 962628block 962629block 962631block 962676 confirmed
Anchor file
hypotheses/locked/20260815_coronal_hole_high_speed_streams__d599a50d.canonical-sha256.ots
Run it yourself
python scripts/verify_hypothesis_anchors.py
Data & code
NASA OMNI2 hourly 1996-2025; ComCat via common/fdsn_multi.py (global M6.5+, declustered 500 km / 30 d)

Why it matters. A popular precursor claim, tested at scale with a proper null and a stated sensitivity floor. It fails -- and the sensitivity bound tells you exactly how large an effect we could have ruled in.

Null museum · Retraction with proof

A precursor we sealed -- then retracted, on the record

Not every sealed hypothesis survives. In one case we sealed a physical precursor hypothesis in our earthquake-forecasting program, forward-tested it, and our OWN later verification found the apparent signal was an artifact -- consistent with the earthquakes recording themselves rather than preceding them. We retracted it. The original seal and the retraction both sit on the Bitcoin-anchored registry, unedited.

STATUS: RETRACTED after our own verification | seal + retraction both anchored, unedited
Sealed pre-registration (identity under commit-reveal)
Positive control

The retraction was driven by an independent internal verification wave that reproduced the original analysis and then found the artifact -- the same discipline that lets a null be trusted, turned on our own positive.

What would have falsified it

The claim was falsifiable by construction; our verification is what falsified it. Because the seal is timestamped, the record proves WHEN we knew we were wrong.

Verify this
Bitcoin anchor
sealed under commit-reveal (identity disclosed on reveal)
Open record
Data & code
Commit-reveal registry entry; the retraction principle is auditable across the public hypothesis registry at /hypotheses

Why it matters. This is the point of the whole page. A group that hides its failures degrades under scrutiny; a group that publishes them -- with a timestamp proving when it knew -- gets stronger. No journal carries a retraction like this. The anchored registry does.

Deliberately generic: this precursor is under a commit-reveal seal. Its identity is disclosed only on reveal.

Look closer

The whole record is open.

Every sealed hypothesis — passed, failed, or withdrawn — lives in the public registry. The forecast ledger recomputes in your browser or from a script. Nothing here asks for your trust; it hands you the means to check.

Open the hypothesis registry →   Verify the ledger →   Ground your AI agent →