The full 360-trial campaign was launched, then killed after 4h39m of
CPU time with zero trials completed -- still stuck on the very first
touch of the very first trial.
Root cause, quantified from the workload's own source, not estimated:
RUN-CHAOS5 (workload-5.4th) is 1000000 0 DO CHAOS-FIELD CHAOS-RIPPLE
LOOP, multiplying out to ~2.4 trillion word executions for one call --
~495 days at the observed rate. Never designed to be called to
completion as a single touch in a repeatedly-birthed worker.
A second problem found computing the fix rather than discovering it
mid-run again: workload-1.4th's own birth-time self-execution (20
SQUARE-WAVE + 1500 SQUARE-BURST + 100000x MICRO-BURST, all three run
automatically at capsule load) sums to ~875 million words -- ~4.3
hours just to birth one worker -- and worker index 1 always maps to
it in heterogeneous mode, landing in half the campaign's cells
regardless of concurrency level.
Fix: two new capsules, workload-1-lite.4th and workload-5-lite.4th,
carrying the same word bodies verbatim but a bounded self-execution
tail, comparable in scale to the campaign's other workloads. The
originals are untouched; only multiuser-doe.4th's own WL-CAPSULE/
WL-ENTRY index 1 and 5 mappings were repointed.
Verified live before relaunching a third time: the actual worst case
in isolation (5 0 998 MU-RUN-TRIAL, concurrency=8 heterogeneous,
includes both fixed workloads plus RUN-OMNI) completed in ~20 minutes
wall-clock, Hera stayed healthy throughout. Clean 3-arch qemu boot.
Reps reduced 60 -> 30/cell (Bob's call after seeing the real per-trial
cost) -- still matches the project's "rule of 3's" DoE convention
(same as ACL-RWT's own 30 reps). 6 cfgs x 30 reps = 180 trials.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
Builds the experimental control Bob correctly identified as still
missing after HB-ON/HB-OFF (§XXXIII.5): walk the shuffled matrix of
concurrency-level x workload-mode cells, birth the right worker count
per cell, drive each with two VM-EXEC touches, check VM-ERROR?, kill
them, print a trial marker. Built entirely from existing primitives
(WORKER-BIRTH, VM-EXEC, VM-HEAT, VM-ERROR?, KILL, HB-ON/HB-OFF) plus
doe.4th's own RUN-MATRIX/SHUFFLE-MATRIX pattern -- no new C primitives.
Real capsule-format bug found and fixed: a first draft, chunked purely
by a fixed 16-line count with no regard for word boundaries, split
several CASE...ENDCASE structures and one oversized colon definition
across Block headers. Result was a cascading [CAPSULE][DEFER] failure
from the first split forward -- every subsequent line failed to
compile, and MU-RUN-TRIAL was never actually defined (confirmed:
UNKNOWN WORD when called). Root cause traced to capsule_loader.c
directly: a :...; word and any control structure inside it must fit
entirely within one 16-line block -- the loader's per-block compile
pass has no persistent record of an open CASE's (or an overlong
definition's own) state across a Block boundary. Not previously
documented anywhere in this project's capsule-authoring guidance.
Fixed via manually curated block boundaries and factoring oversized
bodies into smaller helper words.
Reps ratified at 60/cell (not the 10 first drafted), matching this
project's own "rule of 3's" DoE convention (ACL-RWT's 3 seeds/30 reps,
std79's 3x9x3). 6 cfgs x 60 reps = 360 main-block trials.
MU-MAX-REPS raised 20 -> 63 for run-matrix headroom.
Verified live on amd64: one isolated trial (0 0 999 MU-RUN-TRIAL)
produced two real concurrent births, two clean kills, and the exact
expected marker (MU-TRIAL run=999 cfg=0 rep=0 nw=2 mode=0 fail=0).
Hera stayed healthy throughout. Clean 3-arch qemu boot.
Not yet run: the full 360-trial MU-EXEC-CAMPAIGN itself (a genuinely
long-running action under TCG, deliberately not started without
explicit confirmation) or the WIREBIND-automation fixed arm.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K