Five bounded game runs begin
Each run used a lead-owned state ledger, sequential builders, independent critics, evidence receipts, and promotion gates.
Autonomous production study · August 11–12, 2026
A one-day experiment in bounded autonomy, independent criticism, deterministic evidence, crash recovery, and knowing when not to call unfinished work complete.
The portfolio
The runs produced major playable progress, but none met the portfolio's strict completion definition before the backend budget stop. Accepted gates remain accepted; candidates remain candidates.
Project records
Each record preserves the starting point, accepted boundary, candidate work, evidence, validation facts, known token floor, failure modes, and exact restart action.
A repository skeleton became a deterministic, three-mission tactics slice played by molded-plastic soldiers across a room-scale battlefield.
A deterministic Patio proof became a warm, household-built tower-defense slice with comic framing, original enemies, tactile feedback, and a reviewable evidence trail.
A protected authored campaign gained a separate, deterministic three-stage roguelite mode without replacing or corrupting the original game.
A browser-first voxel architecture became a measured First Excavation playable, while a much larger five-region spell slice was preserved without pretending its missing critic verdict was a pass.
The First Night gray-box route gained a coherent Combat V2 handoff, exact result and rollback contracts, and a readable product shell. A distinct whole-frame fighter pass remains candidate work.
Portfolio timeline
The same operating loop played out differently in each engine. Repository truth and evidence receipts survived the app crash and made the wind-down auditable.
Each run used a lead-owned state ledger, sequential builders, independent critics, evidence receipts, and promotion gates.
Repository state and evidence receipts allowed recovery without restarting accepted work.
Visual overlap, route fallbacks, input leakage, browser lifecycle, evidence naming, and device-truth defects were found outside ordinary unit suites.
Marcus reported approximately 6% remaining. Expansion stopped and truthful checkpoints took priority.
All descendants stopped, Git truth reconciled, postmortems written, and accepted work separated from candidates.
The prepared marble-platformer gauntlet did not consume the freed seat because no run passed the strict completion test before wind-down.
Run architecture
The lead owned state and promotion. Builders changed bounded leaves. Fresh critics could reject them. Reality checks covered what deterministic suites could not.
Cross-project postmortem
The expensive work was not primarily feature code. It was proof: fresh critics, browser matrices, evidence refreshes, recovery, and fixing harnesses that had begun testing the wrong thing.
They found visible overlap, enemies under the floor, hidden automation, false evidence captions, route fallbacks, and input events leaking across scenes. Several defects were invisible to the passing unit suites.
Foxman's final matrix took roughly 15 minutes. Carpet Front repeatedly paid for two-viewport browser proof. The next protocol should require focused gates until a critic says the candidate is materially ready.
Chrome target leaks, key transport, stale audio oracles, mutable evidence provenance, and architecture identity all consumed substantial effort. Once repaired, those harnesses became reusable production infrastructure.
The Codex force-quit did not erase accepted work. Git heads, canonical run-state files, evidence manifests, and critic receipts allowed the portfolio to resume from verified boundaries instead of chat memory.
When the human backend signal reached roughly 6%, expansion stopped. No candidate was promoted by ceremony, no test was weakened, and Afterglow Circuit stayed queued instead of consuming a speculative seat.
Project endpoints exposed lead-side floors, not total descendant or plan consumption. Future runs should record agent, model, gate, retries, wall time, and token deltas in a shared portfolio ledger at every handoff.
What comes next
Carpet needs a corrected oracle and re-critic. Loafer needs real arm64 device truth. Foxman needs a bounded boss leaf. Hex needs provenance reconciliation and a critic. Cock Fight needs a controllable fresh critic. Afterglow remains deliberately queued.