Phi9 / research & engineering

The evidence, not the pitch

What we ran, what failed and what these results do not establish. Recorded 2 October 2026.

SIMULATED DISCOVERY · REPLAYED BASELINE AND REVISION

A single surviving model can still be wrong

The 1,440-run revised study yielded 518 exact singleton predictions, 272 incorrect singleton predictions and 650 nonsingleton or no-model outcomes. These are evaluation categories, not 518 accepted discoveries.

0 accepted / 1,440 abstentions

Under the explicitly conservative full-domain acceptance rule, every run abstained. All eight probes cost 13; the observation budget was 12. Exhaustive checks found a quadratic match for every cubic target on every seven-input subset. Every expandable cubic run returned an incorrect singleton.

This identifies an information obstruction in this protocol. It does not demonstrate general discovery, independent novelty or sample-efficiency superiority.

Original and revised 1,440-run matrices were replayed and checked for completeness and uniqueness. The original 14 tests and revised 18 tests passed. The revision corrected pre-generation budget admission, acceptance reporting and corruption checks; standard-budget trajectories stayed unchanged.

Inspect grouped results · Inspect exhaustive obstruction check · Source and raw runs

BOUNDED PROGRAM · ACTUAL MODEL GENERATION, SEPARATE EXECUTION

An inference-time plan that runs

One real Codex turn generated a task-local JSON program. A separate allowlisted interpreter validated and executed scale → add → dot → readout, returning 6 from the initial vector [1, 2]. Deterministic and command-line replay matched; 15 targeted tests passed.

The expected answer was supplied to the generator. This is a plan-synthesis demonstration, not independent discovery. Tokens encode a program; arithmetic executes in the interpreter, not literally inside tokens. Model weights and global permissions were not changed.

Limits cover modeled work, logical state elements and trace records—not total process memory or provider tokens. Provider token and cost totals were not captured.

Read the execution report · Replay the implementation

CONTROLLED EXTENSION · NOISELESS FINITE ENUMERATION

Evidence and hypotheses are separate limits

The integrated workflow freshly executed and deterministically replayed 5,856 cases. Its 75 tests cover execution and reporting checks; they are not evidence of scientific superiority.

The new matched primary comparison is separate from the original 518/272 study above. Each arm contains 1,440 cases and retains the original search and stopping rules:

Observation capCorrect singletonsWrong singletonsOther outcomesAccepted
125142736530
13653134653293

In a separate exhaustive control, 128 cubic truth tables × three query policies produced 384 cases per condition. Both arms disabled early-singleton stopping. With the complete degree-three hypothesis set and observation cap 13, all 384 cases were exact and accepted. At cap 12, all 384 remained ambiguous and abstained. A quadratic-only set at cap 13 returned no model in all 384 cases.

Finite cubic control: quadratic-only cap 12 produces 384 wrong singleton predictions that abstain; cap 13 detects 384 no-model cases. Complete degree-three cap 12 leaves 384 ambiguous cases; cap 13 accepts 384 exact cases.

The image records the local pre-publication experiment. Actual probe costs were 10 under cap 12 and 13 under cap 13: the last probe cost three units and could not fit the two remaining units. The larger grammar also has more hypotheses and modeled work.

This is an exact noiseless finite-domain result, not open-ended discovery or an equal-cost sample-efficiency comparison. Merely raising the budget does not fix the original search: the primary matched cap-13 experiment still produced 134 incorrect singletons, alongside 653 exact predictions and 653 other outcomes. Its acceptance rule accepted 293 cases and abstained on 1,147.

Inspect the complete summary · Read the recorded protocol · Download the public review bundle

SEPARATE RETAINED AI STUDY · NOT RERUN FOR THIS EDITION

Calibration reduces a failure, not every loss

The two-sensor pilot uses 100 paired seeds per condition and five policies. Null full-model deployment decreases after threshold calibration; rolling full remains more accurate on planted signal. Do not pool these results with the discovery counts above.

Inspect all conditions, costs and uncertainty · Read the chapter

Open work

The eight-page working manuscript (v0.2) reports the actual studies; read the HTML edition. It is not peer reviewed. Outreach is not represented as sent. Physical-workflow data collection and downstream validation remain proposed work.

How to use this record

These are internally checked research materials, not peer-reviewed papers, production guarantees or physical-world validation. Test counts establish engineering checks, not scientific superiority. Send a concrete critique or collaboration enquiry.

Research updates

Notes on the physics of AI — modeling and algorithm design. Sent only when there is something worth reading.

Contact [email protected]