Phi9 / research & engineering

Better selection, with the losing result beside it

A retained two-sensor calibration pilot, displayed condition by condition rather than as a single success score.

Empirical working draft · 2 October 2026 · synthetic one-step regression

Raising the inherited acquisition margin reduced unnecessary full-model deployment in two null conditions. It did not remove discovery latency: rolling full sensing predicted the planted signal more accurately.

Values come from the retained study summary, not a new website experiment. Final article authorship and venue remain unresolved. This is not peer review, external replication or physical-world validation.

100fresh paired seeds / condition
1,500policy runs, not independent seeds
24,000abstract-work ceiling / run

Independent-noise null

Full-model deployment fell from 29/100 to 5/100 for both calibrated selectors.

Online MSE · shared scale 0 to 0.6

Inherited margin0.58790
Fixed-window calibrated0.57834
Alarm-conditioned0.57809
Rolling compact0.58044
Rolling full0.58450
Online MSE and actual charged work. Lower MSE is better; work is not money or wall time.
PolicyOnline MSECharged workFull selections / 100
Inherited margin0.58790400729
Fixed-window calibrated0.57834107905
Alarm-conditioned0.57809107885
Rolling compact0.5804415368Not a selector
Rolling full0.5845023261Not a selector

Five selections out of 100 have a Wilson 95% interval of about 2.15%–11.18%. A 5% point estimate is not a 5% population guarantee.

Planted sensor signal

All selectors deployed full sensing on 100/100 streams. Yet online MSE was 0.14988 versus 0.08996 for rolling full.

Online MSE · shared scale 0 to 0.6

Inherited margin0.14988
Fixed-window calibrated0.14988
Alarm-conditioned0.14988
Rolling compact0.57176
Rolling full0.08996
Online MSE and actual charged work. Lower MSE is better; work is not money or wall time.
PolicyOnline MSECharged workFull selections / 100
Inherited margin0.149884604100
Fixed-window calibrated0.1498811618100
Alarm-conditioned0.1498811618100
Rolling compact0.5717615368Not a selector
Rolling full0.0899623261Not a selector

Alarm-calibrated minus rolling-full MSE: +0.05992, paired 95% interval [+0.05578, +0.06406]. Positive means the selector lost.

Shuffled-sensor null

Full-model deployment fell from 28/100 to 5/100 for both calibrated selectors. The shuffle is an approximate additional null, not an exact conditional-randomization test.

Online MSE · shared scale 0 to 0.6

Inherited margin0.57796
Fixed-window calibrated0.56964
Alarm-conditioned0.56964
Rolling compact0.57176
Rolling full0.57765
Online MSE and actual charged work. Lower MSE is better; work is not money or wall time.
PolicyOnline MSECharged workFull selections / 100
Inherited margin0.57796398728
Fixed-window calibrated0.56964107905
Alarm-conditioned0.56964107905
Rolling compact0.5717615368Not a selector
Rolling full0.5776523261Not a selector

Five selections out of 100 have a Wilson 95% interval of about 2.15%–11.18%. A 5% point estimate is not a 5% population guarantee.

What did the calibration buy?

Both calibrated policies pay 7,014.11 units per stream for combined calibration work amortized over 300 planned streams. Equal ceilings are not equal actual spend. Rolling full incurs eleven denied refits per stream. Prediction rows are workload, not independent replicates.

In the noise null, the five alarm-calibrated selected candidates show mean validation gain +0.1732 but fresh evaluator-only gain −0.2143. Those tiny selected subsets are descriptive and unstable; the shadow diagnostic never feeds back to the policy and scores the pre-refit candidates.

Protocol and interpretation

96 initial examples, then 512 predictions scored before their own feedback. One residual alarm buys one 48-example probe: fit on 32, select on 16, refit on 48, then deploy. Thresholds 0.1144 and 0.1193 were frozen before evaluation on seeds 44001–44100, disjoint from calibration seeds 33001–33200.

Two independent Gaussian sensors; noise, planted signal and shuffled-sensor conditions. One alarm and two candidate predictors do not test general diagnosis, deep learning or dynamical control. Ordinary paired intervals are unadjusted, not familywise or anytime-valid. The two calibration modes have no demonstrated difference in useful acquisition here.

The retained audit reports nineteen internal engineering checks, including frozen bytes, disjoint seeds, raw loss reaggregation, fifteen deterministic replays and feedback noninterference. This page verifies display values against that summary; it does not rerun or independently certify the study.

Inspect the complete displayed values · Read chapter 13 · Separate discovery evidence

Research updates

Notes on the physics of AI — modeling and algorithm design. Sent only when there is something worth reading.

Contact [email protected]