Phi9 / research & engineering
Better selection, with the losing result beside it
A retained two-sensor calibration pilot, displayed condition by condition rather than as a single success score.
Empirical working draft · 2 October 2026 · synthetic one-step regression
Raising the inherited acquisition margin reduced unnecessary full-model deployment in two null conditions. It did not remove discovery latency: rolling full sensing predicted the planted signal more accurately.
Values come from the retained study summary, not a new website experiment. Final article authorship and venue remain unresolved. This is not peer review, external replication or physical-world validation.
Independent-noise null
Full-model deployment fell from 29/100 to 5/100 for both calibrated selectors.
Online MSE · shared scale 0 to 0.6
| Policy | Online MSE | Charged work | Full selections / 100 |
|---|---|---|---|
| Inherited margin | 0.58790 | 4007 | 29 |
| Fixed-window calibrated | 0.57834 | 10790 | 5 |
| Alarm-conditioned | 0.57809 | 10788 | 5 |
| Rolling compact | 0.58044 | 15368 | Not a selector |
| Rolling full | 0.58450 | 23261 | Not a selector |
Five selections out of 100 have a Wilson 95% interval of about 2.15%–11.18%. A 5% point estimate is not a 5% population guarantee.
Planted sensor signal
All selectors deployed full sensing on 100/100 streams. Yet online MSE was 0.14988 versus 0.08996 for rolling full.
Online MSE · shared scale 0 to 0.6
| Policy | Online MSE | Charged work | Full selections / 100 |
|---|---|---|---|
| Inherited margin | 0.14988 | 4604 | 100 |
| Fixed-window calibrated | 0.14988 | 11618 | 100 |
| Alarm-conditioned | 0.14988 | 11618 | 100 |
| Rolling compact | 0.57176 | 15368 | Not a selector |
| Rolling full | 0.08996 | 23261 | Not a selector |
Alarm-calibrated minus rolling-full MSE: +0.05992, paired 95% interval [+0.05578, +0.06406]. Positive means the selector lost.
Shuffled-sensor null
Full-model deployment fell from 28/100 to 5/100 for both calibrated selectors. The shuffle is an approximate additional null, not an exact conditional-randomization test.
Online MSE · shared scale 0 to 0.6
| Policy | Online MSE | Charged work | Full selections / 100 |
|---|---|---|---|
| Inherited margin | 0.57796 | 3987 | 28 |
| Fixed-window calibrated | 0.56964 | 10790 | 5 |
| Alarm-conditioned | 0.56964 | 10790 | 5 |
| Rolling compact | 0.57176 | 15368 | Not a selector |
| Rolling full | 0.57765 | 23261 | Not a selector |
Five selections out of 100 have a Wilson 95% interval of about 2.15%–11.18%. A 5% point estimate is not a 5% population guarantee.
What did the calibration buy?
Both calibrated policies pay 7,014.11 units per stream for combined calibration work amortized over 300 planned streams. Equal ceilings are not equal actual spend. Rolling full incurs eleven denied refits per stream. Prediction rows are workload, not independent replicates.
In the noise null, the five alarm-calibrated selected candidates show mean validation gain +0.1732 but fresh evaluator-only gain −0.2143. Those tiny selected subsets are descriptive and unstable; the shadow diagnostic never feeds back to the policy and scores the pre-refit candidates.
Protocol and interpretation
96 initial examples, then 512 predictions scored before their own feedback. One residual alarm buys one 48-example probe: fit on 32, select on 16, refit on 48, then deploy. Thresholds 0.1144 and 0.1193 were frozen before evaluation on seeds 44001–44100, disjoint from calibration seeds 33001–33200.
Two independent Gaussian sensors; noise, planted signal and shuffled-sensor conditions. One alarm and two candidate predictors do not test general diagnosis, deep learning or dynamical control. Ordinary paired intervals are unadjusted, not familywise or anytime-valid. The two calibration modes have no demonstrated difference in useful acquisition here.
The retained audit reports nineteen internal engineering checks, including frozen bytes, disjoint seeds, raw loss reaggregation, fifteen deterministic replays and feedback noninterference. This page verifies display values against that summary; it does not rerun or independently certify the study.
Inspect the complete displayed values · Read chapter 13 · Separate discovery evidence