Sixteen planning and orchestration files with no runnable game.
Three.js voxel spell exploration
Hex Quarry
A browser-first voxel architecture became a measured First Excavation playable, while a much larger five-region spell slice was preserved without pretending its missing critic verdict was a pass.
Outcome
First playable accepted; full slice preserved for critique
This report separates accepted product truth from preserved candidate work. A missing critic or device proof is treated as an open gate, not a semantic technicality.
An accepted dense-palette voxel route with Cinder editing and recovery, plus an unaccepted 45-file candidate containing five regions, five spells, secrets, enemy roles, a guardian, results, and replay.
Accepted boundary
- Measured browser-voxel architecture and selected worker-meshing path
- First Excavation playable at accepted commit 2ce1e7b
- Transactional edits, latest-safe Recall, checkpoint recovery, and content-addressed evidence
Held, candidate, or unstarted
- Five-region B2 full slice at a5d48a7 has builder evidence but no independent critic verdict
- Current-head evidence validation is red because provenance pins an earlier control HEAD
- The one-shot evidence writer is consumed and must not be rerun
Gate history
Decision timeline
The useful unit is not the turn. It is the gate: a bounded claim, its evidence, and the verdict that controlled promotion.
Architecture, protocol, cancellation, rollback, and provenance hardened through multiple REVISE loops.
Arrival Shelf, Cinder path opening, recovery, and evidence accepted.
Target law, pose, source binding, clearance, and save integrity repaired.
Five-region additive full slice and one evidence batch completed.
Freshness check stopped before an independent product verdict.
Candidate and evidence preserved; no promotion.
Existing evidence
Frames from the run.
No images were generated for this report. These are existing repository artifacts with shortened SHA-256 anchors. Candidate frames remain visibly labeled.



Testable now
Reproduce the current truth.
Commands are repository-root commands from the project's canonical debrief. They are shown for technical review, not as a public playable deployment claim.
npm ci
npm run typecheck
npm run test:first-playable
npm run build:first-playable
npm run evidence:validate:first-playable
# Candidate verification only
npx vitest run --config full-slice/vitest.config.tsThe First Excavation command set is accepted. The full-slice tests are builder evidence, not an independent acceptance verdict.
Git anchors
- Initial 7c1f4c5
- Accepted first playable 2ce1e7b
- B2 candidate a5d48a7
- Wind-down 423720b
Known token floor
4,002,176
Latest exact platform ledger snapshot. It does not include any unavailable descendant or backend-plan accounting.
Candid postmortem
Failures that changed the run.
The point of the case study is not that agents wrote code. It is that the system found reasons not to trust its own first answer.
Loops, stalls, and defects
- Gate 1 needed several protocol and evidence-path corrections before product work was trustworthy.
- The full-slice critic stopped at freshness preflight and produced no verdict.
- Evidence identity was bound to mutable control HEAD, so later docs-only commits made current-head validation red.
- A one-shot evidence policy prevented cheap regeneration from erasing history.
What carries forward
- Separate accepted and candidate directories made a large unfinished slice safe to preserve.
- Content-addressed evidence builds trust, but provenance should bind product bytes rather than mutable control HEAD.
- Builder evidence is not an independent verdict, regardless of test volume.
Without rerunning the evidence writer, reconcile the provenance validator against the preserved product boundary and run one fresh browser-first critic over the existing B2 evidence. Promote only on PASS.