Research, October 2026
Out-of-sample results / 2026
Eighteen fortnights, scored against what happened.
We trained the simulator on data to December 2025, then ran it forward on every two-week window of 2026 for books of 8, 16, 24 and 50 US stocks. These are the results, including the check it fails.
2 January to 25 September 2026
Four book sizes, equal-weight
3 of 5 at 50 stocks; joint crashes fail at every size
The test
What we ran, and on what.
For each window, the simulator saw only the last 382 half-hour closes (about four weeks) of every stock in the book. It drew 32 joint 10-trading-day futures for the whole book. We then compared those futures with what actually happened over the next 10 trading days. Windows do not overlap, and no 2026 data was used in training.
How the books were chosen. The 8-, 16- and 24-stock books are hand-picked large US companies across sectors. The 50-stock book is drawn from S&P 500 names. All books are equal-weight. The 24-stock book, our worked example: AAPL, MSFT, NVDA, AMZN, JPM, XOM, PG, JNJ, GOOGL, NFLX, TSLA, META, V, UNH, HD, AVGO, ABBV, COST, MRK, CVX, PEP, KO, CRM, ADBE.
The five checks
Four pass up to 24 stocks. One fails at every size.
| Check | 8 stocks | 16 stocks | 24 stocks | 50 stocks |
|---|---|---|---|---|
| 1. Do the stocks move together the right amount? (slope, 1.00 = perfect) | 0.93 Pass | 0.94 Pass | 0.95 Pass | 0.95 Pass |
| 2. When one crashes, do others crash too? (share of real joint crashes) | 64% Fail | 58% Fail | 57% Fail | 55% Fail |
| 3. Is the size of each stock’s moves right? (simulated / real) | 0.99 Pass | 0.99 Pass | 0.99 Pass | 0.98 Pass |
| 4. Is the range centred? (0 = centred) | +0.005 Pass | −0.016 Pass | −0.020 Pass | −0.021 Just misses |
| 5. Is the loss estimate reliable? (breaches of 18; 3 or fewer is green) | 1 Pass | 3 Pass | 2 Pass | 1 Pass |
Check 4 at 50 stocks misses its threshold by two thousandths. We report it as a miss. Check 5 uses the binomial expectation for 18 windows: a correct 5% estimate is breached about once in 18, and three or fewer breaches is consistent with that.
The range
A range that corrects its own width.
The raw 90% range was too narrow for larger books. So after each window, the range is widened or narrowed by a multiplier learned only from earlier windows’ misses. The first five windows of the year have no history and stay raw. In 2026 the multiplier settled near 1.4 for the 24-stock book.
| Stocks | Raw range | Self-corrected, all 18 windows | Self-corrected, windows 6–18 |
|---|---|---|---|
| 8 | 88.2% | 89.6% | 88.5% |
| 16 | 81.6% | 87.8% | 90.9% |
| 24 | 80.8% | 87.3% | 91.3% |
| 50 | 79.7% | 87.3% | 93.5% |
How sure can we be? Eighteen windows is a small sample. A hit rate measured on 18 windows carries about ±15 points of uncertainty, so these results support “roughly a 90% range”, not “exactly 90%”. Testing over earlier years is the next step.
The bad fortnight
The loss estimate, against real losses.
| Stocks | Bad-fortnight estimate, average | Estimate, worst window | Worst real fortnight | Breaches of 18 |
|---|---|---|---|---|
| 8 | -3.8% | -5.4% | -3.3% | 1 |
| 16 | -3.7% | -6.5% | -5.9% | 3 |
| 24 | -2.9% | -5.0% | -4.3% | 2 |
| 50 | -2.8% | -4.4% | -4.0% | 1 |
The bad-fortnight estimate is the 10-day loss the book should exceed in one fortnight out of twenty (a 10-day 95% value at risk). Read the raw estimate as a floor: because joint crashes are under-reproduced, the true tail is probably somewhat deeper.
The known weakness
Joint crashes are under-shown.
We count the half-hours in which two stocks both have one of their worst 5% moves. In the 24-stock book that happened at a rate of 0.225 in reality and 0.128 in the simulation: 57% of real. A textbook Gaussian model at the same correlations gives 0.089. So we do better than the textbook, and still fall short of reality. Until that changes, we say so next to every loss number.
The market is alive
Correlations moved more than chance allows.
We split the 24-stock book’s half-hour returns from December 2025 to September 2026 into 7 back-to-back four-week blocks and measured how much each pair’s correlation changed from one block to the next. On average it moved by 0.099. Shuffling the same returns in time 500 times gives the change that chance alone produces: 0.062. Real correlations moved 1.6 times more, and no shuffle came close.
The clearest case was the March 2026 sell-off. The average correlation across the book went from 0.10 to 0.21 within four weeks, then back to 0.08 by May. This is why the simulator re-measures co-movement from the last four weeks on every run.
What the test does not show
Limits.
No direction. The middle of the range carried no information about which way prices went (its correlation with outcomes ranged from −0.02 to +0.13 across book sizes). The product is the shape of the range, not a forecast.
One year. All results come from 2026. Testing over earlier years has not been done yet.
Large US stocks only. Small caps, ADRs, ETFs and other markets are untested in book form.
Prices only. The model does not know earnings dates or news. Gap-prone names are flagged instead: in the 50-stock book the screen flagged NFLX, ORCL, META, UNH, MCD, CSCO, IBM.
Questions about these results?
Read how the simulation works, or talk to us directly.