Research & profilesZenodoYouTubeORCID
Reflective Time ModelRTM Research Platform
HE
Reflective Time ModelResearch › Statistics
RTM · STATISTICS & NULL TESTS

How do we test whether RTM findings could arise by chance?

RTM findings contain networks linking people, times, dates, places, texts and external records. The statistical question is not simply “is there a matching number?” It is: if we repeatedly generate matched alternatives and give them the same kind of search freedom, can they reproduce the full structure?

No statistics background requiredThe page starts in plain language. Numbers and formulas are kept in optional technical sections.
We do not just count equationsThe tests examine roles, sources, timing, recurring targets and dependency between paths.
This is not all of RTMOnly findings that already have a defined quantitative test are included in the statistical corpus.
What is tested?Whether the full network can be reproduced, not just one match.
What was found?Within the tested reference models, the documented corpus is very difficult to reproduce.
What does this not determine?The statistics alone do not identify the mechanism behind the measured pattern.
In one sentence: this section asks how easily a matched comparison process can reproduce RTM-like network architecture under declared rules.
THE BOTTOM LINE

What did the tests show, in plain language?

The simplest conclusion is that the combined result for the tested corpus is highly unusual within the main comparison models built for it. More search freedom weakens the anomaly, so the site reports sensitivity rather than one isolated headline number.

Latest reported corpus baseline

about 1 in 2.94 quintillionalmost 3 billion billions · within the latest reported comparison modelThis is a corpus-level estimate across 18 quantified finding families, not a single-case probability.

This means a synthetic corpus at least as strong as the tested one is estimated to be extremely rare inside this model. It is not “the probability RTM is true.”

How to read that number

It does meanhow hard it is for a matched alternative world to reach the observed level under the declared rules.
It does not meanthat RTM or any specific physical mechanism has been proven with that probability.
It depends on assumptionsmore allowed wording sets or complete candidate events make the result less unusual.
How to read this without statistics: if matched alternative worlds could easily reproduce the observed architecture, many of them would reach the real case’s score. In many of the tests they remain far below it. The statistics quantify that gap.
How it works

Four steps, no formulas

1

Freeze the real case

Keep the documented people, roles, timing, source and calculation rules.

2

Generate matched alternatives

Replace person with person, time with time, room with room and date with date.

3

Let the alternatives search too

They receive the same allowed operations and representations.

4

Compare the full network

Measure roles, recurring targets, independent paths, loops, chronology and dependency — not just one numerical hit.

RTM LAB · STATISTICS ENGINE

The RTM Lab statistics engine — the same search contract for the observed case and the Null

RTM Lab includes a dedicated Statistics view that connects structural analysis to Null testing. The goal is not merely to take a discovered equation and ask how rare it looks, but to freeze roles and search grammar and give alternative worlds the same search freedom.

01

Enter the event and data roles

Identity, date, time, external source, question, answer, and other anchors can be marked by role. If roles are left automatic, the engine can use a role-blind mode.

02

Resolve dependence before p-values

The engine distinguishes reconstruction groups, Primary units, dependent support, and network components so variants of the same information are not automatically counted as new evidence.

03

Run Multi-Null under the same Search Contract

Alternative worlds receive the same representation and search rules. Multiple Null models can be compared to test sensitivity instead of relying on one number.

04

Role tests and reproducible export

When the role space is finite, exact permutation tests can be used. Session and RTM Audit exports preserve inputs, rules, and results.

Where is it?Open RTM Lab and choose 06 · Statistics. The link here opens that view directly.
Open the RTM Lab statistics engine →
CASE MAP

What are the 18 quantified finding families?

One line of context for each family before the numbers. This is only the currently quantified subset of RTM, not the complete archive.

Nimrod Les

Identity, two life dates, age and birthday converge through several paths.

4338

An original text and a full biblical verse share 4338, followed by a role-sensitive structural test.

6833 — Idan / Ecclesiastes

An Idan-related text is compared with a verse in Ecclesiastes, with structure beyond the initial match.

“One Ordinary Child”

An old physical book sticker and date connect to present-day family ages and the rediscovery date.

Albania / Plaza Tirana

Hotel, room, floor, date and identities form one event network.

Roey Harpaz

Identity, address, place and years are tested as one network.

Sun finding

External scientific-source data enter a network with identity and date anchors.

“How many minutes in an hour?”

A simple fixed-time case tested against additional identity and time anchors.

Ofer / Zephaniah

Verse selection followed by a test of whether each value stays in the correct semantic role.

922

A WhatsApp-group event where content, identities and event roles are tested together.

Jeep

A photo and event time linked to an anchor fixed before the numerical search.

21:33

Text, correction, time and identity form a locked relation skeleton tested in a large alternative space.

Spain

An earlier record is compared with a later event outcome.

Winner form

An external receipt, printed odds and send time create a partially externally controlled event network.

Golden Dragon

A multi-equation network reduced to its true algebraic rank and penalized for search freedom.

104 / Beersheba covenant

A physical stone, verse and dates are linked through external anchors.

Abraham / Gil

Several calculation clusters return to shared nodes, so dependency is part of the test.

Idan’s wedding

Venue, couple, address and dates form one event network under a defined search grammar.

Easy-to-read examples

Four examples of what was actually tested

These examples use different comparison models. They should not be multiplied together as independent p-values.

4338 · text vs external source

Observed case: 11 of 11 conditions

The primary 4338 match was given to the Null for free and dependent families were collapsed.

Across 50 million alternative worlds, the best result was 3 of 11. None reached even 4 of 11.
Titanic · role assignment

All 39,916,800 assignments were checked

This was exhaustive, not a small sample.

The historical assignment was the unique global maximum: 190/190 versus 183/190 for the best alternative.
Emily Dickinson · observer reconstruction

Which coherent profile completes the full structure?

423,360 coherent profiles were checked, plus 100 million identity-replacement worlds.

Only the documented profile completed all tested layers; no joint full reproduction appeared in 100 million replacements.
Ofer / Zephaniah · negative and positive control

Why “many equations” is not enough

A broad arithmetic search was not unusual. The strong separation appeared only when semantic roles had to remain correct.

Broad arithmetic: p=0.44338. Role preservation: 12/12 versus a 3/12 maximum in one million permutations.
CASE-LEVEL BENCHMARK · WEDDING 2021

Wedding network — 6/6 after dependency compression

Six core families: question↔ketubah, 2337 convergence, return to 1229, 810 convergence, the 371 identity chain, and return to the bride. This case-level benchmark is kept separate from corpus-wide estimates.

Observed6/6all core structures
±5% · exact1 / 9.47Qp = 1.056×10⁻¹⁶
Conditional≈ 1:399Bketubah and 1229 fixed
Identity-only≈ 1:8.34Mexternal event data fixed
1 / 9,468,289,223,531,775The exact ±5% space contains one full 6/6 world — the observed case.
MC: max 2/620M in each of ±5%, ±10% and ±25%; zero full 6/6 reconstructions.

Why show conditioned benchmarks too? They show what remains when more external event facts are fixed in advance. The benchmarks are not multiplied together and are not a global search-adjusted p-value.

For readers who want the detail

Full numbers, collapsed by default

You do not need these sections to understand the conclusion. They are here for transparency.

High-N simulation scale and targeted tail runs

18 finding families × 100,000 Null worlds = 1.8 million event-level worlds. Four closer cases were extended to one million runs each.

CaseHits / 1MEmpirical p95% conservative bound
Nimrod Les1,3601.36×10⁻³1.42×10⁻³
Albania949.5×10⁻⁵1.12×10⁻⁴
Winner form1251.26×10⁻⁴1.45×10⁻⁴
6833212.2×10⁻⁵3.02×10⁻⁵
Search-freedom sensitivity — High-N v2.2 table

K = generic wording sets fixed in advance. B = complete candidate events that may be tried and omitted. B is not arithmetic-operation count.

Version note: this table belongs to the High-N v2.2 combination (B=1 baseline ≈6.46×10⁻¹⁹). The later post-tail baseline, 3.4074×10⁻¹⁹, is reported separately at the top of the page.

AllowanceOmnibus estimatePlain language
B=1≈6.46×10⁻¹⁹≈1 in 1.55×10¹⁸
B=3≈1.46×10⁻¹⁶≈1 in 6.84×10¹⁵
B=10≈5.21×10⁻¹⁴≈1 in 1.92×10¹³
B=30≈1.02×10⁻¹¹≈1 in 98.5B
B=100≈2.88×10⁻⁹≈1 in 347M
B=1,000≈6.71×10⁻⁵≈1 in 14,900
How should a global correction avoid double-counting search opportunities?

The trial unit represented by pcorpus is a complete ~18-family research program, not an hour or a thought. A valid external correction would use an empirically defensible Meff of complete equivalent search programs.

p_global = 1 − (1 − p_corpus)^M_eff
Additional case-level results

9/11; Titanic; Emily Dickinson; World Cup + Spain; 14–15 July 2026; Dan Panorama; The Living Subscriber; Ford Capri; Golden Dragon; 21:33. Each belongs to its own comparison model and should not be multiplied into a universal p-value.

Why two nearby corpus baselines appear

High-N v2.2 reported p≈6.46×10⁻¹⁹ (≈1 in 1.55×10¹⁸). After targeted tail updates, the latest reported baseline was pcorpus=3.4074×10⁻¹⁹ (≈1 in 2.94×10¹⁸). They are shown by version rather than silently merged.

PLAIN-LANGUAGE GLOSSARY

The technical words, simply

Null / comparison worldA matched alternative used to see whether the observed structure is genuinely unusual.
Monte CarloGenerate many alternative worlds automatically and count how many succeed.
p-valueThe share of comparison worlds at least as strong as the observed one under the declared model.
DependencyIf two equations reuse the same information, they should not be counted as two independent pieces of evidence.
Search budgetHow many wordings, candidate events or other choices may be tried before selecting the best result.
PermutationKeep the same values but swap which role each one represents.
RESEARCH STATUS

Three layers already completed — one layer still needs independent replication

This separates what has already been done from the next evidential step.

Discovery & documentation

A documented corpus with external sources, timing and role-sensitive network structure.

Nulls & stress tests

Simulation, role permutation, finite-space enumeration and search-freedom sensitivity tests.

Confirmatory tests

New-data tests under rules fixed in advance, beyond the discovery corpus.

Independent external replication

Outside researchers run frozen protocol, code and scoring on new data without involvement in model construction.

What can we say today, in plain language?

The central quantitative result is that the tested corpus is very difficult to reproduce within the comparison models built for it. The signal does not rest on the sheer number of equations; it comes from correct roles, external anchors, timing, recurring targets, multiple paths and explicit dependency control.

1. This is not just “many matches”The stronger tests evaluate whole-network architecture and do not automatically count reused information as independent evidence.
2. “Search enough and you will find something” becomes testableThe Null receives explicit search freedom. When that freedom is expanded, the anomaly weakens, which is why sensitivity curves are shown alongside the headline result.
3. Statistics measure extremity, not mechanismThey measure how hard the architecture is to reproduce under tested reference models. The physical or informational explanation is a separate question.
Where does the research stand, and what comes next? Discovery, severe Null tests, simulations and confirmatory tests on new data have already been performed. The next layer of evidence is independent external replication: researchers not involved in building RTM run a frozen protocol, code and scoring rules on new data. The goal is no longer simply to find additional closures, but to test whether the architecture reproduces independently.