Inference-time program synthesis and bounded execution

Phi9 · 2 October 2026 · Research demonstration

One real Codex turn generated a task-local JSON program. The expected answer was supplied. A separate Python interpreter validated and executed it.

[1, 2] → scale 2 → [2, 4]
→ add [1, −1] → [3, 3]
→ dot [1, 1] → scalar 6 → readout

15 targeted tests passed. Deterministic and CLI replay matched. Replay makes no new model request. The interpreter allows only scale, add, dot and readout, checks exact fields and finite values, and admits modeled work before operations.

What this establishes

A generated plan can run through a separate bounded execution layer and return an inspectable trace. It does not establish independent discovery, model-weight learning, unrestricted autonomy or physical meaning of the synthetic vector.

Limits

Budgets count admitted coordinate operations, logical state elements and trace records, not total process memory or provider tokens. Provider token/cost totals were not captured. Task-local configuration does not change global tool permissions.

The public source, archived program and tests are supplied for evaluation. This document is an internally checked report, not peer-reviewed research.