Phi9 · reproducible local research release · version 1
This release connects a validated archived token program, a fresh bounded discovery matrix, a controlled observation/hypothesis extension and a representation-sufficiency workload to pinned results and an existing versioned artifact pipeline. It is a working offline command, not a hosted autonomous research service.
Fresh baseline: 1,440 learner executions plus deterministic audit. Extension: 4,416 learner executions (2,880 primary and 1,536 exhaustive controls). Token interpreter: archived Codex-produced JSON, newly validated and executed; no fresh model inference. Research-core: fresh 64-integer representation workload. Replaying reruns the workloads and compares deterministic results, excluding wall-clock samples.
| Arm | runs | exactSingleton | wrongSingleton | nonSingletonOrNoModel | accepted | abstained | observationCostTotal | modeledWorkTotal |
|---|---|---|---|---|---|---|---|---|
| all | 1440 | 518 | 272 | 650 | 0 | 1440 | 12122 | 910582 |
A single surviving hypothesis is not an accepted answer. Acceptance requires all eight input values observed and exactly one hypothesis. The complete probe domain costs 13, so cap 12 cannot meet that conservative rule.
| Arm | runs | exactSingleton | wrongSingleton | nonSingletonOrNoModel | accepted | abstained | observationCostTotal | modeledWorkTotal |
|---|---|---|---|---|---|---|---|---|
| 12 | 1440 | 514 | 273 | 653 | 0 | 1440 | 12035 | 915702 |
| 13 | 1440 | 653 | 134 | 653 | 293 | 1147 | 12605 | 1485533 |
Only the observation cap changes in this arm. The original learner can stop after two audit probes or be trapped in a misspecified singleton without requesting the decisive probe. Therefore a sufficient available budget need not be used; inspect the actual outcomes rather than equating budget with correctness.
| Arm | runs | exactSingleton | wrongSingleton | nonSingletonOrNoModel | accepted | abstained | observationCostTotal | modeledWorkTotal |
|---|---|---|---|---|---|---|---|---|
| 12 | 2 | active | 128 | 0 | 128 | 0 | 0 | 128 | 1280 | 392960 |
| 12 | 2 | cost-aware | 128 | 0 | 128 | 0 | 0 | 128 | 1280 | 392960 |
| 12 | 2 | random | 128 | 0 | 128 | 0 | 0 | 128 | 1280 | 163584 |
| 12 | 3 | active | 128 | 0 | 0 | 128 | 0 | 128 | 1280 | 785920 |
| 12 | 3 | cost-aware | 128 | 0 | 0 | 128 | 0 | 128 | 1280 | 785920 |
| 12 | 3 | random | 128 | 0 | 0 | 128 | 0 | 128 | 1280 | 327168 |
| 13 | 2 | active | 128 | 0 | 0 | 128 | 0 | 128 | 1664 | 393216 |
| 13 | 2 | cost-aware | 128 | 0 | 0 | 128 | 0 | 128 | 1664 | 393216 |
| 13 | 2 | random | 128 | 0 | 0 | 128 | 0 | 128 | 1664 | 163712 |
| 13 | 3 | active | 128 | 128 | 0 | 0 | 128 | 0 | 1664 | 786432 |
| 13 | 3 | cost-aware | 128 | 128 | 0 | 0 | 128 | 0 | 1664 | 786432 |
| 13 | 3 | random | 128 | 128 | 0 | 0 | 128 | 0 | 1664 | 327424 |
All 128 three-bit cubic coefficients, three query policies, both observation caps and both fixed hypothesis degrees are enumerated. Both caps use the same no-early-singleton stopping rule (audit requirement 8). Degree 2 omits the cubic term; degree 3 contains every Boolean table. These controls change stopping and starting grammar relative to the primary study; comparisons are only within this explicitly separate factorial.
At seven observed inputs each cubic truth table has a quadratic completion. Eight observations distinguish it, but a quadratic-only learner then has no model. A complete hypothesis set plus all eight observations recovers the truth table. This is finite noiseless enumeration, not sample-efficiency superiority, general physical discovery or learned representation revision.
Token readout: 6.0. Modeled budget charges and trace are retained. The target was supplied during the archived Codex generation; replay is not independent solution discovery.
[
{
"aggregateCorrect": 8,
"aggregateTotal": 8,
"encodeNs": 7087,
"maxAbsoluteError": 0.0,
"meanAbsoluteError": 0.0,
"method": "raw-int32",
"payloadSha256": "4b69728f27314651cba62c0e18e0111729ad9cd3b4aa3328bede1b444fb2b5ac",
"pointCorrect": 64,
"pointTotal": 64,
"queryNs": 27156,
"reconstructNs": 28824,
"reconstructedBytes": 256,
"storedBytes": 256,
"totalCorrect": 72,
"totalNs": 63067,
"totalQueries": 72
},
{
"aggregateCorrect": 8,
"aggregateTotal": 8,
"encodeNs": 51217,
"maxAbsoluteError": 0.0,
"meanAbsoluteError": 0.0,
"method": "zlib-lossless",
"payloadSha256": "7e49c57c55b5c554d9fae3190c4893d98707307d318cf15d75b9e4a463f7d341",
"pointCorrect": 64,
"pointTotal": 64,
"queryNs": 8859,
"reconstructNs": 33971,
"reconstructedBytes": 256,
"storedBytes": 86,
"totalCorrect": 72,
"totalNs": 94047,
"totalQueries": 72
},
{
"aggregateCorrect": 8,
"aggregateTotal": 8,
"encodeNs": 9462,
"maxAbsoluteError": 24.375,
"meanAbsoluteError": 5.704861111111111,
"method": "bin-mean-lossy",
"payloadSha256": "34a00a1a23590601202eef57c7113417324623b8449282f3c8e7ba0f27194748",
"pointCorrect": 0,
"pointTotal": 64,
"queryNs": 5452,
"reconstructNs": 4468,
"reconstructedBytes": 512,
"storedBytes": 64,
"totalCorrect": 8,
"totalNs": 19382,
"totalQueries": 72
}
]Bin means preserve aligned sums but lose point-query fidelity on this predetermined workload. Storage payload bytes are not process memory. Timing samples are not a speed ranking.
Record actual observation spend, candidate-generation totals, modeled generation/search/filter work and wall times separately. Modeled work is not FLOPs, dollars or provider tokens. Provider calls and external writes are zero. A full run has finite matrices, a 180-second timeout per subprocess workload and an immutable output directory. No background daemon, live sensing, automatic model-generation loop or publication is installed.
python3 release.py run --output runs/first python3 release.py replay --output runs/first python3 release.py status --output runs/first python3 -m unittest -v test_release
Open research-report.html and prepared/releases/.../index.html locally. The latter is a selected public-safe review bundle, not authorization to publish. Source and output hashes are pinned; changes fail replay. A new run requires a new empty output path; existing runs are never silently overwritten.
This post-review extension targets a known obstruction and is locally protocol-pinned, not external preregistration. Seed blocks are not proof of independence. No noise, held-out task distribution, downstream physical benchmark or novelty claim is established. Research-core supplies only its existing calculation kernel here; its separate Git-source-ingestion system is not rerun or advertised as connected.
Abel et al., state abstractions for lifelong RL and Booker and Majumdar, task-specific memory provide broader context, not empirical support for this finite program. This report is a reproducible technical release, not a peer-reviewed paper.