Commit Graph
100 Commits
Author SHA1 Message Date
Robert Allan JamesandClaude Sonnet 5 9aebbb224b Fix extract_doe_region.py SLOT_FMT drift; regenerate report with real data
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
FABRIC-3.md §XXXV.15. §XXXV.12 added an explicit _align_pad field to
doe_trial_slot_t to fix a struct-alignment bug, but never updated the
Python extractor's SLOT_FMT/SLOT_CRC_SPAN to match -- every field after
isa was read from the wrong byte offset, so slot_crc read wrong bytes
entirely and every legitimate record (including the real, live
campaign's own first two clean trials) was flagged as CRC-corrupt.
Fixed; SLOT_SIZE assertion tightened from <= to == to catch this drift
immediately next time.

Regenerated multiuser_doe_report.pdf against the real campaign's actual
first two trials: run=0 (cfg=0, nw=2, uniform, fail=0) and run=1 (cfg=5,
nw=8, heterogeneous, fail=0), both fleet_conserved=1 at exactly Q48.16
1.0. Real report, real data, first time this has happened for this
campaign. disk/artemis.img itself and the live campaign's log directory
are deliberately left uncommitted -- still being written by the running
campaign.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-17 23:28:50 -04:00
Robert Allan JamesandClaude Sonnet 5 553ca9163d FABRIC-3.md §XXXV.14: flag missing fault-triggered SSM_MODE_C0 fallback
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Raised discussing what the multiuser DoE might reveal: the intended
design is a faulting VM drops to SSM_MODE_C0 and continues as a plain,
non-adaptive VM rather than faulting outright. Confirmed via source:
ssm_l8_update() computes its mode purely from workload metrics, no
fault input anywhere -- a real regression from the hosted StarForth
repo's design ("forgot to carry that over... building the stadium"),
not invented here. ssm_l8_force_config() already exists as the natural
hook a real fix would use. Flagged and logged, not built -- per the
explicit instruction to keep the campaign running and log bugs/gaps as
they surface rather than stopping to fix each one.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-17 21:03:02 -04:00
Robert Allan JamesandClaude Sonnet 5 c02fed140f Fix ISA-tagging: two real bugs found live, campaign quit for a fresh run
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
FABRIC-3.md §XXXV.13. Quit the running amd64 campaign directly on request
once §XXXV.12's build unblocked live testing -- first live check found
the ISA-tagging mechanism completely non-functional, in two independent
ways, both only caught by checking an actual printed value rather than
"it resolves without error."

Bug 1: per-isa-doe.4th's strategy of redefining MU-EMIT-TRIAL-MARKER
never worked, structurally, from the moment it was written. This FORTH
compiles a colon definition's word calls to a direct, early-bound
reference at compile time -- MU-RUN-TRIAL's call to MU-EMIT-TRIAL-MARKER
bound when multiuser-doe.4th loaded, before per-isa-doe.4th's later
redefinition of the same name ever existed. Every trial the running
campaign persisted had an empty isa tag, not "amd64". Fixed by moving
ISA-tag storage (MU-ISA-TAG-BUF/LEN/!) and the one true
MU-EMIT-TRIAL-MARKER into multiuser-doe.4th itself, reading a shared
variable at run time instead. per-isa-doe.4th shrank from 3 blocks to 2
(Block 5118 freed) -- fixing this recovered namespace, didn't cost any.

Bug 2: found immediately while verifying the fix. MU-ISA-TAG!'s CMOVE
call was copied from the original PER-ISA-TAG! it replaced, which had
its own latent stack-order bug -- passed the ISA string's length as the
copy source address instead of the real address, silently copying
blank/garbage bytes with no error. Never caught before because §XXXV.8's
own verification only checked that the entry words resolved, not that
the tag they set was actually correct. Fixed the stack order; the old
buggy word no longer exists anywhere (removed with Bug 1's fix, not
patched in place).

Verified live, together, one boot: MU-ISA-TAG-BUF readback -> "amd64"
(was blank), full MU-EMIT-TRIAL-MARKER with all 6 fields set manually ->
exact match on every field, VM-ERROR? -> 0, DOE-RESUME-COUNT correctly
isolates by ISA (amd64->1, riscv64->0 in the same ring). Clean BYE, zero
leaked processes, synthetic test write reverted before committing.

Preserved logs/CSVs from the killed campaign and both verification boots
per this repo's own convention -- never delete these.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-17 19:28:36 -04:00
Robert Allan JamesandClaude Sonnet 5 51940f49d0 Add per-slot CRC + campaign crash-resume (DOE-RESUME-COUNT)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
FABRIC-3.md §XXXV.12: closes two real gaps found by walking the disk-to-
PDF pipeline end to end. (1) doe_trial_slot_t had no integrity check of
its own -- 64 slots share one devblock via a read-modify-write, so a
crash mid-write to any slot could silently corrupt its siblings with
nothing to catch it. Added an 8-byte CRC (same discipline the control
header already had), computed/verified the same way. (2) The campaign
itself had no crash-resume -- a reboot always restarted the whole
18-trial loop from index 0 even with good persisted data already on
disk. Since the trial matrix is fully deterministic from (seed, n-reps),
resuming needs no persisted "which trials ran" state, just a count:
DOE-RESUME-COUNT (new kernel primitive, same pattern as
DOE-PERSIST-TRIAL) walks the ring counting CRC-valid records for a given
ISA; MU-EXEC-CAMPAIGN's signature changes to accept a start-idx and seeds
MU-RUN-ID/the DO loop from it. Each RUN-*-CAMPAIGN entry word now calls
DOE-RESUME-COUNT before launching. Both block-namespace-checked to fit
existing headroom before committing to the design -- no new blocks
needed.

Real bug caught by the build, not shipped: an implicit compiler-inserted
alignment byte before slot_crc silently broke the hand-computed trailing
pad size (doe_trial_slot_size_check failed to compile). Fixed with an
explicit alignment field, matching log_slot_t's own zero-implicit-
padding discipline.

extract_doe_region.py now verifies both CRCs with the REAL CRC-64/XZ
algorithm from block_subsystem.c's compute_crc64() (ported and verified
bit-for-bit against a native C reference this session), replacing
§XXXV.11's incorrect guessed implementation.

Verified: clean compile on all 3 ISAs (amd64/aarch64/riscv64 -- amd64
verified via isolated build only, not boot, since the real per-ISA
campaign is currently live on the only QEMU instance this project's own
rule allows), both capsules pass mkcapsule --lint, CRC64 bit-exact via
native reference. Live boot-verification explicitly deferred until that
campaign finishes or is interrupted -- not claimed as tested when it
wasn't. Found (not missed) a real compatibility consequence: the running
campaign's data is in the old pre-CRC format, so it reads as "corrupt"
under the new extractor and DOE-RESUME-COUNT will resume it from 0, not
mid-campaign -- no data loss, cheap redundant re-run of ~1-2 trials.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-17 19:17:04 -04:00
Robert Allan JamesandClaude Sonnet 5 c535252907 Build the DoE region reader + automated R/LaTeX report pipeline
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Closes the "still open" reader-tool note from §XXXV.10: scripts/
extract_doe_region.py spools DOE-PERSIST-TRIAL's on-disk binary ring off
disk/artemis.img into a CSV -- direct file read by default, or via a real
loop device (--loop, losetup) per the literal request. Independent
Python reimplementation of doe_region.c's own ring-walk, verified
against a live 3-record write/read round trip.

analyse_multiuser_doe.R (house style, ggplot2/svglite) and
multiuser_doe_report.tex (same preamble convention as
bare_metal_doe_report.tex) auto-discover whatever data exists and
degrade gracefully to a placeholder when a campaign has 0 trials so far
-- the whole point being this can be rebuilt at any point during a
multi-day campaign, including mid-run. build_multiuser_doe_report.sh
chains extraction -> R -> svg_to_pdf.py -> pdflatex into one command.

Found and fixed a real bug while verifying: the R script's SCRIPT_DIR
detection silently wrote charts/tables/report output to the repo root
instead of the analysis directory when invoked from a different cwd --
switched to the robust commandArgs() "--file=" idiom. Verified live both
branches (0 trials and a real 3-record write), reverted synthetic test
data and incidental svg_to_pdf.py churn on ~85 unrelated committed
figures before committing. FABRIC-3.md §XXXV.11 has the full account.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-17 16:59:26 -04:00
Robert Allan JamesandClaude Sonnet 5 aaf349946f Preserve serial logs and DoE CSVs from the last two days' campaign attempts
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Audit artifacts from the original 180-trial single-ISA run (killed after
8 trials, colliding-session incident), the two boot verification runs
around it, and the per-isa-doe.4th amd64 leg (killed deliberately, no
persistence layer yet). Kept per this repo's own convention -- never
delete logs/ or experiments/bare_metal/runs/*.csv.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-17 16:44:10 -04:00
Robert Allan JamesandClaude Sonnet 5 2da746c01c Build DOE-PERSIST-TRIAL: block-backed DoE campaign persistence, live-verified
FABRIC-3.md §XXXV.3/.4 designed this and it was never built -- two live
campaigns got launched relying entirely on a live serial-log tail before
this was caught. doe_region.c/.h mirrors log_region.c's own growable-ring
pattern (own devblock_from_top fence at 98, past log_region's ceiling at
97) to persist one trial-summary record per completed trial straight to
disk/artemis.img, surviving past QEMU exit with no host required.

Verified end to end: booted amd64, called DOE-PERSIST-TRIAL at the
console, clean BYE, then read the raw disk bytes back with QEMU fully
dead -- confirmed the exact values passed in. Wired into both
multiuser-doe.4th and per-isa-doe.4th's trial markers. Clean build, zero
new warnings, all 3 ISAs; mkcapsule --lint passes both edited capsules.

FABRIC-3.md §XXXV.10 documents the fix and the root-cause correction: two
campaigns were launched on an unbuilt persistence design before this was
caught and fixed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-17 16:43:46 -04:00
Robert Allan JamesandClaude Sonnet 5 8b8e8e1405 Build capsules/per-isa-doe.4th: per-ISA campaign driver (N=3, 18 trials/ISA)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Thin wrapper over multiuser-doe.4th's own MU-EXEC-CAMPAIGN, which already
implements the ratified 6-cell factorial generically -- adds only what
FABRIC-3.md §XXXIV.4's per-ISA pivot needed: one fixed seed per ISA,
N=3 reps/cell=18 trials/ISA (ratified over §XXXV.5's own N=2 safety-margin
recommendation, Captain Bob's explicit choice), an ISA tag threaded into
the trial marker so records from all 3 ISAs' sequential campaigns can be
told apart on the shared disk/artemis.img, and three single-word
no-argument entry points (RUN-AMD64-CAMPAIGN/RUN-AARCH64-CAMPAIGN/
RUN-RISCV64-CAMPAIGN).

Blocks 5117-5119 -- the only 3 remaining in the [2048,5120) user block
namespace; mkcapsule --lint caught the first draft going one block past
the ceiling. Verified live, load-only: both capsules load with zero
UNKNOWN WORD/redefinition/error lines, all 3 entry words resolve,
constants compile correctly. No campaign trial has been run yet --
build+verify only, per explicit scope for this pass.

FABRIC-3.md §XXXV.8/.9 records the decision and verification.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:30:43 -04:00
Robert Allan JamesandClaude Sonnet 5 3223fac2a0 Execute FABRIC-3.md §XXXV.6: format new artemis.img, re-mint all 9 identities live
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Boots amd64 with Zuse's bleached thumbdrive, confirms live: fresh 1 GiB
disk/artemis.img formats ("Artemis: blank disk -- formatting"), Zuse
genesis-mints onto it ("Zuse: genesis minted onto attached thumbdrive",
root CA written to the new fence's devblock 0), session authenticates
(zuse@Hera ok>). With Zuse live, mints the other 8 identities (bob, 00-06)
one at a time via QMP-driven xHCI hotplug + MINT at the console, each
confirmed by its own "MINT: identity minted" log line before detaching
and moving to the next. Clean BYE shutdown, zero leaked qemu/socat
processes, all 9 images verified nonzero.

Root cause confirmed live before executing: Zuse's genesis marker lives
on artemis.img's own fence (devblock_from_top=0), not on her thumbdrive
-- with the marker gone, her old drive could not have re-authenticated
regardless of its own content, confirming re-mint was genuinely required.

FABRIC-3.md §XXXV.7/.8, disk/README.md updated with the full account.
Serial log preserved: logs/20260917-114652/amd64/.

Still open: the block-backed DoE persistence writer itself is unbuilt
(design only); ZUSE-ELIGIBILITY-ADD (elevation eligibility, narrower
than baseline auth) not re-populated for any of the 8; old 30 MiB
artemis.img backup disposal undecided.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 11:57:10 -04:00
Robert Allan JamesandClaude Sonnet 5 e56974e0bf Bleach all 9 thumbdrive identity images to blank ahead of re-mint
Prerequisite for FABRIC-3.md §XXXV.6's ratified re-mint: capsule_mint.c's
capsule_mint_identity() refuses to write over an already-recognized drive
(MINT_ERR_ALREADY_MINTED, mirrors WRITE(10)'s refuse-on-non-blank-media
rule), so all 9 existing certs had to be destroyed before any of them
could be re-minted against the new 1 GiB artemis.img. Confirmed root
cause first: Zuse's genesis marker (zuse_pubkey) lives on Artemis's own
disk fence (devblock_from_top=0), not on her thumbdrive -- capsule_zuse_
boot.c's capsule_zuse_boot_try_attach() silently no-ops if her drive
isn't blank and the marker is missing, so the old zuse-thumb-ident.img
could not re-authenticate against the fresh artemis.img regardless.

All 9 images (zuse, bob, 00-06) verified all-zero, 64 MiB size intact.
Re-mint itself (live QEMU boot + MINT) not yet performed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 11:45:12 -04:00
Robert Allan JamesandClaude Sonnet 5 7c5ba7a874 Rebuild disk/artemis.img fresh at 1 GiB, blank; preserve old 30 MiB image
Executes the drive-resize decision ratified in FABRIC-3.md §XXXV.6:
the prior 30 MiB image was too small for the block-backed per-ISA DoE
persistence design (§XXXV.2). Built a new blank 1 GiB image rather than
migrate the old one's top-of-device metadata fence -- re-minting Zuse
and the thumbdrive identities invalidates them either way, so skip the
migration entirely.

- disk/artemis.img: replaced with a fresh blank 1 GiB image (was 30 MiB)
- disk/artemis-30mb-pre-1gib-backup.img: the superseded 30 MiB image,
  preserved rather than deleted; disposal remains open (FABRIC-3.md §XXXV.7)
- disk/README.md: documented both images per this repo's own convention

Not yet formatted or re-minted -- that's a live-boot step, not done here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 11:33:10 -04:00
Robert Allan JamesandClaude Sonnet 5 e4a52c01be FABRIC-3.md §XXXV: bump ratified artemis.img target from 256 MiB to 1 GiB
Captain Bob asked for extra safety margin over the earlier 256 MiB
recommendation. Updated every reference in §XXXV.2/.4/.6 consistently.
Still decision only -- no image built yet.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 11:30:34 -04:00
Robert Allan JamesandClaude Sonnet 5 c19b6febc2 FABRIC-3.md §XXXV.6: ratify fresh 256 MiB artemis.img + full re-mint, no fence migration
Captain Bob resolved the drive-resize fork from §XXXV.2: build a new
blank 256 MiB disk/artemis.img and re-mint Zuse + all 8 thumbdrive
identities fresh, rather than attempt to migrate the top-of-device
fence on the existing 30 MiB image. Skips migration complexity since
a re-mint invalidates prior identities either way. Decision only --
no new image built, no re-mint performed yet.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 11:29:34 -04:00
Robert Allan JamesandClaude Sonnet 5 a83cbe016f FABRIC-3.md §XXXV: block-backed DoE persistence design + N=2 per-ISA shape (design only, no code yet)
Scopes the self-contained (no host serial capture) persistence needed
before the per-ISA campaign driver can run on real bare-metal hardware.
Found disk/artemis.img is only 30 MiB (measured) against a 516 MB/8-trial
raw per-tick baseline -- raw mirroring is ruled out at any realistic
device size. Proposes reusing log_region.c's binary-slot/control-header
pattern for a new trial-summary-only region, flags the fence-addressing
risk in growing the image, and recommends N=2 reps/cell over N=3 given
the added block-IO overhead. Nothing built -- ratified in conversation only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 11:25:57 -04:00
Robert Allan JamesandClaude Sonnet 5 94215e8b47 Fix two catastrophically heavy workloads found running the real campaign
The full 360-trial campaign was launched, then killed after 4h39m of
CPU time with zero trials completed -- still stuck on the very first
touch of the very first trial.

Root cause, quantified from the workload's own source, not estimated:
RUN-CHAOS5 (workload-5.4th) is 1000000 0 DO CHAOS-FIELD CHAOS-RIPPLE
LOOP, multiplying out to ~2.4 trillion word executions for one call --
~495 days at the observed rate. Never designed to be called to
completion as a single touch in a repeatedly-birthed worker.

A second problem found computing the fix rather than discovering it
mid-run again: workload-1.4th's own birth-time self-execution (20
SQUARE-WAVE + 1500 SQUARE-BURST + 100000x MICRO-BURST, all three run
automatically at capsule load) sums to ~875 million words -- ~4.3
hours just to birth one worker -- and worker index 1 always maps to
it in heterogeneous mode, landing in half the campaign's cells
regardless of concurrency level.

Fix: two new capsules, workload-1-lite.4th and workload-5-lite.4th,
carrying the same word bodies verbatim but a bounded self-execution
tail, comparable in scale to the campaign's other workloads. The
originals are untouched; only multiuser-doe.4th's own WL-CAPSULE/
WL-ENTRY index 1 and 5 mappings were repointed.

Verified live before relaunching a third time: the actual worst case
in isolation (5 0 998 MU-RUN-TRIAL, concurrency=8 heterogeneous,
includes both fixed workloads plus RUN-OMNI) completed in ~20 minutes
wall-clock, Hera stayed healthy throughout. Clean 3-arch qemu boot.

Reps reduced 60 -> 30/cell (Bob's call after seeing the real per-trial
cost) -- still matches the project's "rule of 3's" DoE convention
(same as ACL-RWT's own 30 reps). 6 cfgs x 30 reps = 180 trials.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 15:47:30 -04:00
Robert Allan JamesandClaude Sonnet 5 29f91f9531 multiuser-doe.4th: the campaign trial-loop capsule, verified live
Builds the experimental control Bob correctly identified as still
missing after HB-ON/HB-OFF (§XXXIII.5): walk the shuffled matrix of
concurrency-level x workload-mode cells, birth the right worker count
per cell, drive each with two VM-EXEC touches, check VM-ERROR?, kill
them, print a trial marker. Built entirely from existing primitives
(WORKER-BIRTH, VM-EXEC, VM-HEAT, VM-ERROR?, KILL, HB-ON/HB-OFF) plus
doe.4th's own RUN-MATRIX/SHUFFLE-MATRIX pattern -- no new C primitives.

Real capsule-format bug found and fixed: a first draft, chunked purely
by a fixed 16-line count with no regard for word boundaries, split
several CASE...ENDCASE structures and one oversized colon definition
across Block headers. Result was a cascading [CAPSULE][DEFER] failure
from the first split forward -- every subsequent line failed to
compile, and MU-RUN-TRIAL was never actually defined (confirmed:
UNKNOWN WORD when called). Root cause traced to capsule_loader.c
directly: a :...; word and any control structure inside it must fit
entirely within one 16-line block -- the loader's per-block compile
pass has no persistent record of an open CASE's (or an overlong
definition's own) state across a Block boundary. Not previously
documented anywhere in this project's capsule-authoring guidance.
Fixed via manually curated block boundaries and factoring oversized
bodies into smaller helper words.

Reps ratified at 60/cell (not the 10 first drafted), matching this
project's own "rule of 3's" DoE convention (ACL-RWT's 3 seeds/30 reps,
std79's 3x9x3). 6 cfgs x 60 reps = 360 main-block trials.
MU-MAX-REPS raised 20 -> 63 for run-matrix headroom.

Verified live on amd64: one isolated trial (0 0 999 MU-RUN-TRIAL)
produced two real concurrent births, two clean kills, and the exact
expected marker (MU-TRIAL run=999 cfg=0 rep=0 nw=2 mode=0 fail=0).
Hera stayed healthy throughout. Clean 3-arch qemu boot.

Not yet run: the full 360-trial MU-EXEC-CAMPAIGN itself (a genuinely
long-running action under TCG, deliberately not started without
explicit confirmation) or the WIREBIND-automation fixed arm.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 10:27:45 -04:00
Robert Allan JamesandClaude Sonnet 5 41918a28a4 doe_log.c: add vm_name/vm_id_hex identity columns; new calibration workload
Adds the identity columns HB-ON's existing per-tick CSV logger
(doe_log_tick_row(), doe_log.c) was missing. Previously every column
described *a* VM's state each row, but nothing said which VM emitted
it -- concurrent VMs' rows were indistinguishable by source. Two new
first columns, vm_name and vm_id_hex, resolved via a reverse lookup on
vm->stadium_vm_id (capsule_vm_registry_get()).

Real bug found and fixed while building this: the freestanding
snprintf here silently prints the literal format string instead of
substituting for %016llx (width+ll+hex unsupported) -- caught by
reading the actual emitted row, not assumed to work. Replaced with a
hand-rolled hex nibble-table loop, the same idiom vm_uuid_format()/
MINT-SCRATCH-EMIT already use.

Corrects FABRIC-3.md's own prior "still not started: CSV driver"
framing: no bespoke CSV emitter is needed for the multiuser DoE at
all -- this per-tick logger already exists, fires automatically inside
every VM's own execution loop, and just needed HB-ON plus these
identity columns to be usable for concurrent workers. Also corrects
doe_log.h's own stale doc comment claiming g_doe_log_enabled defaults
to 1 -- doe_log.c's own source is authoritative: 0, off by default,
matching the HB-ON/HB-OFF naming.

One real, honest limitation recorded rather than smoothed over:
vm_name reads blank for any tick captured during a WORKER-BIRTH'd
capsule's own self-execution at birth time, since capsule_birth_baby()
runs that work (and its heartbeat ticks) before the registry name can
be set. vm_id_hex is unaffected (reads directly from vm->stadium_vm_id)
and still uniquely disambiguates every row -- confirmed live: a
blank-name row's own vm_id_hex matched its later PARITY:KILL line's
vm_id exactly.

New workload-calib1.4th: Bob confirmed the existing 10 workload-N.4th
files aren't fixed. Sized to cross HEARTBEAT_CHECK_FREQUENCY (256
word-executions/tick) many times over while staying far lighter than
RUN-FIB/RUN-CHAOS5, both too slow under TCG for a bounded verification
run. A real, reusable addition, not a throwaway.

Verified live on amd64 via a self-contained one-shot EXEC'd test
(HB-ON, WORKER-BIRTH two calibration workers, HB-OFF, VM-ERROR? on
both, KILL both, completion banner) rather than interactive polling,
which breaks once HB-ON makes the log grow continuously via Hera/
Hermes/Artemis's own background ticks. Clean 3-arch qemu boot on the
real committed change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 09:53:25 -04:00
Robert Allan JamesandClaude Sonnet 5 e33eb36361 Multiuser DoE punch list: WORKER-BIRTH + VM-ERROR? + vm_physics_init fix
Core mechanism for §XXXIII's main concurrency block, built and
live-verified on amd64. Two new primitives:

- WORKER-BIRTH ( capsule-c capsule-u name-c name-u -- ok? ): births a
  named, VM-EXEC-addressable VM from an arbitrary (p) capsule with no
  identity involved. Corrects the ratified design's own assumption that
  the main block would use UNATTENDED-BIRTH -- that requires a
  committed capsule per identity, impractical for dozens of trial VMs.
  The concurrency block never needed identity at all.
- VM-ERROR? ( c-addr u -- flag ): reads a named VM's error state from
  Hera, mirroring VM-HEAT's silent/always-returns-a-value contract.
  Needed to check a VM-EXEC-driven trial VM's own fault state after
  the fact -- nothing existing let Hera do this.

Real bug found and fixed in both WORKER-BIRTH and UNATTENDED-BIRTH:
capsule_birth_baby() never calls vm_physics_init() either (same shape
as the registry-name gap found building UNATTENDED-BIRTH) -- without
it a born VM is never in the VM Fleet Attractor physics list, so
VM-HEAT returns 0 forever regardless of work done. Fixed by adding
vm_physics_init() alongside the existing registry-name call in both
words.

Traced (not guessed) why heat still read 0 after one VM-EXEC touch
even post-fix: vm_physics_touch()'s transfer logic only fires from a
VM's *second* touch onward -- the first touch just records a baseline
tick. Verified live across three sequential touches: heat 0 -> 7039 ->
11333. This is a real design requirement for the DoE's heat/CV
response variable (each trial must touch a worker at least twice), not
a bug to route around.

Also found live: all 10 existing workload-N.4th capsules self-execute
their full workload at load/birth time (a bare top-level call to their
own RUN-* word at file end) -- missed on an earlier, too-shallow
8-line survey of each file. WORKER-BIRTH alone already runs a worker's
first pass as a side effect of birth.

Verified live on amd64: two concurrent workers (fib + matrix-mul)
birthed, run, measured (heat + error state), and killed cleanly;
VM-HEAT/VM-ERROR? both confirmed silent-0 on an unknown name. Clean
3-arch qemu boot on the real committed change.

Still open: the run-matrix/shuffle/CSV driver capsule itself, the
WIREBIND-automation path for the fixed arm, and the per-VM touch-count
budget's exact value.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 08:51:21 -04:00
Robert Allan JamesandClaude Sonnet 5 bc2e294c50 FABRIC-3.md §XXXIII: multiuser/multitasking DoE design (no code yet)
Scopes the DoE Bob asked for after §XXXII closed -- "a pretty big
experiment... long running multi-factor DoE that sort of puts the OS
through its paces." Design only, per this project's standing
plan-before-code discipline.

Three corrections found and recorded before the design could be
trusted:
- Console sessions are not a concurrency axis and this isn't a defect:
  console.c hardcodes one physical UART as static global state, not an
  instantiable abstraction. One physical terminal, one foreground
  session, by construction. Bob's reaction (eventual multi-seat/
  telnet-ish/pty-multiplexer idea) recorded separately in memory,
  explicitly deferred behind this DoE.
- The real, provable concurrency primitive is VM-EXEC (doe-campaign.4th's
  own THREE-VM-CAMPAIGN already proves the pattern), resolved through
  the full VM registry (unbounded kmalloc list), not messaging.4th's
  16-slot routing table -- checked live, not assumed.
- No FORTH-reachable monotonic clock exists; latency dropped as a
  response variable rather than building a new timing primitive under
  this scope.

Also flagged, not chased: every acceptance boot this session produced
an L8-DoE CSV despite init.4th never calling it and KERNEL_ARGS
defaulting to empty -- likely a stale starforth.cfg survivng `clean`.
Practical decision: the new campaign is its own driver, independent of
whatever causes that.

Ratified design: concurrency level (2/4/8 VM-EXEC-driven VMs) x
workload assignment (uniform/heterogeneous across the existing 10
workload capsules) as the crossed factorial; identity origin
(WIREBIND vs UNATTENDED-BIRTH) as a fixed arm rather than crossed, since
WIREBIND's per-identity USB-attach cost doesn't scale into a full
factorial. Response variables: correctness-proxy (completes without
vm->error), heat/CV stability under contention, completed-trials-per-
interval in place of the dropped latency variable.

Punch list recorded; implementation not started per standing "plan
approval is not a start signal" rule.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 08:01:26 -04:00
Robert Allan JamesandClaude Sonnet 5 af351bff92 Punch list item 4 (final): CONSOLE-ATTACH, full unattended-identity flow verified end-to-end
Closes FABRIC-3.md §XXXII.2's punch list. CONSOLE-ATTACH ( name-c
name-u -- ok? ) pairs a fresh console VM to an already-live VM
registered as "<name>~user". Deliberately a plain, unconditional
primitive with no VMIdentity capability-bit check -- corrects this
session's own first-pass design (§XXXII.2 amended in the same commit):
identity.installed is 0 for Hera/Hermes/Artemis and for every
console-proxy VM, so a capability-bit gate would be unreachable for
every VM a human actually types at, Zuse included. Matches
ZUSE-ELIGIBILITY-ADD's own "no bespoke gate" precedent in this file;
real gating is `' CONSOLE-ATTACH ACL-PIN` in ACL.4th if ever wanted.

Two real bugs found and fixed via live testing, not assumed correct:
- capsule_birth_baby() never sets a VM's registry name (documented
  gap, same one capsule_runcap_birth()'s own history already hit) --
  UNATTENDED-BIRTH gained a second `name` argument and now calls
  capsule_vm_registry_set_name(new_vm_id, "<name>~user") itself.
- CONSOLE-ATTACH's first draft took an independent console name from
  the target's name. sk_repl_dispatch_line()'s pairing check (repl.c)
  reconstructs the target as console_get_vm_name()+"~user" -- a
  mismatched console name silently falls back to direct interpretation
  with no error. Caught live (typed `5 6 + .` at a mismatched console,
  got a direct `11` instead of a relay) and fixed by collapsing to one
  name argument, matching WIREBIND's own by-construction invariant.

Full end-to-end live verification on amd64: minted a test identity,
UNATTENDED-BIRTH'd it as "bob", CONSOLE-ATTACH'd a console named "bob",
USE'd it, typed `5 6 + .` -- no direct output at [zuse@bob] (relay
path taken), then `[zuse@bob~user] 11 ok>` appeared: the relayed
command executed on the target identity VM itself and printed its own
answer back through the shared console. Hera stayed healthy throughout
(2 2 + . -> 4 after switching back). CONSOLE-ATTACH also verified to
refuse cleanly on a nonexistent target with no orphaned VM. Test
capsule reverted after capture per this project's probe convention --
never committed.

FABRIC-3.md §XXXII.2 fully closed: all 4 original questions ratified,
the mid-course drive_uuid and ACL-bit corrections both recorded
plainly rather than silently folded in, and a doc-accuracy note left
for CLAUDE.md's own stale "1024-byte block limit" framing (mkcapsule's
real limits are range [2048,5120) and 16 content lines/block) --
flagged, not fixed, out of this punch list's scope.

Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real committed
C-only change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 07:24:27 -04:00
Robert Allan JamesandClaude Sonnet 5 5879c8b3bc Punch list item 3: UNATTENDED-BIRTH call site, verified live end-to-end
Implements the unattended-birth mechanism FABRIC-3.md §XXXII.2 designed:
UNATTENDED-BIRTH ( name-c name-u -- ok? ) births a VM from a named (p)
capsule via capsule_birth_baby() (completely unmodified, the same
generic build-time-capsule path CAPSULE-BIRTH already uses), then
installs its identity the same way capsule_wirebind.c already does
live for WIREBIND attaches -- vm_identity_from_cert() verification
followed by a plain post-birth struct assignment -- rather than
anything RUNCAP-shaped, since RUNCAP requires a real blkio_dev+
homeblocks_sig_t an unattended identity never has.

The born VM's own capsule payload is expected to lay down two CREATE'd
buffers (UNATTENDED-ID-UUID, UNATTENDED-ID-CERT) via MINT-SCRATCH-
EMIT's own literal format; their addresses are fetched by interpreting
a two-word line inside the *new* VM's own context
(vm_interpret(born_vm, ...)), the same "run inside that VM's own
dictionary" idiom capsule_wirebind.c already uses for VM-NAME-REG.

Explicit invariant preserved: never touches g_wirebind_attached_username
or any console-pairing state, births no console VM -- an unattended
identity stays un-promptable (§VIII.1) until a human pairs a console to
it later via the already-working VM-NAME-REG mechanism.

Verified live end-to-end on amd64: minted a real test identity via
MINT-SCRATCH, captured its MINT-SCRATCH-EMIT output, built a throwaway
test capsule from it (discovered along the way: mkcapsule's real block
constraints are range [2048,5120) and max 16 content lines per block --
neither matches this repo's own doc comment, corrected via ground
truth from the tool itself, not assumed), then ran UNATTENDED-BIRTH
against it: cert verified against Zuse's root pubkey, identity
installed, "no console attached" reported, Hera stayed healthy
afterward (5 6 + . -> 11). Test capsule reverted after capture per
this project's own probe convention -- not a real identity, never
committed. Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real
committed C-only change.

Remaining punch-list item (the ACL cap bit for console attachment) not
started.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 07:06:47 -04:00
Robert Allan JamesandClaude Sonnet 5 5fc709a228 Punch list item 2: MINT-SCRATCH-EMIT, verified live on all 3 arches
Adds the hand-transcription mechanism FABRIC-3.md §XXXII.2's punch list
item 2 calls for: MINT-SCRATCH-EMIT prints the last successful
MINT-SCRATCH's drive_uuid + cert devblock as ready-to-paste FORTH
source (HEX-based CREATE ... C, ... byte sequences), so packaging an
unattended identity into a capsule is a mechanical copy out of the
captured boot log rather than a manual hex-to-FORTH translation an
operator could transpose a digit in.

Refuses (no output) if no MINT-SCRATCH has ever succeeded -- printing
4112 zero bytes as if they were a real identity would be a silent,
misleading success, matching this session's own error-handling audit
discipline rather than adding a new silent-failure primitive right
after finishing one.

Verified live on amd64: MINT-SCRATCH-EMIT correctly refuses before any
mint, then after MINT-SCRATCH succeeds, emits UNATTENDED-ID-UUID and
UNATTENDED-ID-CERT as valid FORTH literals. Cross-checked byte-exact:
the cert's own embedded ASN.1 serialNumber field matches the emitted
UUID bytes exactly, confirming x509_build_user_cert()'s drive_uuid
binding round-trips correctly through the scratch-device path. Clean
3-arch qemu boot (amd64/aarch64/riscv64).

Remaining punch-list items (writing/committing a real capsule file for
an actual named identity, the unattended-birth call site, the ACL cap
bit) not started -- authoring a real committed capsule needs a name/
purpose decision that isn't mine to make.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 05:12:43 -04:00
Robert Allan JamesandClaude Sonnet 5 63864c4b01 Punch list item 1: scratch-device MINT-SCRATCH, verified live on all 3 arches
Implements FABRIC-3.md §XXXII.2's scratch-thumbdrive mint mechanism:
capsule_mint_identity_scratch() (capsule_mint.c/.h) builds a throwaway
RAM-backed blkio_dev via blkio_ram.c's backend and runs
capsule_mint_identity() against it completely unmodified -- same live
Zuse-signing operation, same rng_get_bytes() draw for drive_uuid a real
thumbdrive gets. Reads back only drive_uuid + the cert devblock; the
seed devblock is written into the scratch buffer internally but never
read out (no seed is ever baked into a capsule, per the ratified
no-seed decision).

New FORTH word MINT-SCRATCH (mama_forth_words.c), same stack signature
as MINT, mints into the scratch device instead of any attached drive
and never touches sk_repl_get_attached_blk_dev() or console-pairing
state. Prints the drive_uuid as hex so a live boot log itself proves
each call drew fresh entropy.

Build correction found along the way: blkio_ram.c was excluded from
the kernel build (Makefile.starkernel VM_EXCLUDE) alongside
blkio_factory.c/blkio_file.c. blkio_factory_open() unconditionally
references blkio_file.c's real fopen()/fread() file I/O, which has no
freestanding-kernel equivalent, so the factory function couldn't be
used as-is. blkio_ram.c itself is pure memcpy over a caller buffer --
pulled it alone into the kernel build and wired it directly in
capsule_mint.c, the same way blkio_factory.c's own extern declarations
do internally.

Verified live on amd64: two MINT-SCRATCH calls produced two genuinely
different drive_uuids (b533246d.../ae11b2b2...), confirming fresh
entropy per call rather than stale reuse; VM stayed healthy afterward
(5 6 + . -> 11). Clean 3-arch qemu boot (amd64/aarch64/riscv64),
logs and DoE CSVs committed per standing convention.

Remaining punch-list items (hand-transcription into a .4th block, the
unattended-birth call site, the ACL cap bit) not started.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 04:42:15 -04:00
Robert Allan JamesandClaude Sonnet 5 60bcdc09a7 Stage E amendment: resolve drive_uuid wrinkle via scratch-device mint (FABRIC-3.md §XXXII.2)
Bob's proposal: mint against a "sim thumbdrive" -- a scratch blkio_dev
from the existing blk_subsys_add_raw_device() mechanism (same one the
kernel ramdrive already uses), then hand-transcribe just drive_uuid +
DER cert as FORTH literals into a normal .4th capsule block.

This dissolves the wrinkle the first pass of Q2 left open:
capsule_mint_identity() runs completely unmodified against the scratch
device (same live Zuse-signing op, same rng_get_bytes() draw for
drive_uuid), so vm_identity_from_cert() needs zero changes -- no
origin-aware variant, no content-hash-derived substitute. "Sim"
describes only where the bytes were written, invisible to verify.

Also corrects an error in the first pass: cert production cannot be
"offline" -- capsule_mint_identity() requires Zuse's live in-kernel
signing key, so minting is inescapably a two-boot runtime operation
(mint on boot N, package by hand, birth on boot N+1).

Punch list updated to match. Still design-only, no code.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 04:01:32 -04:00
Robert Allan JamesandClaude Sonnet 5 06e2507ce4 Stage E ratified: unattended identity is cert+personality only, no seed (FABRIC-3.md §XXXII.2)
Settles all 4 open questions on paper, no code:
- No seed baked into capsules for this pass (Bob's call); runtime-minted
  seed documented as a future, separately-scoped possibility
- Birth via capsule_birth_baby() (unmodified, already generic), not
  capsule_runcap_birth() -- that path requires a real blkio_dev*+
  homeblocks_sig_t* an unattended identity can't provide, and its own
  header says build-time-baked content is exactly the wrong case for it
- Identity population is the same post-birth struct-assignment pattern
  capsule_wirebind.c already uses live (vm->identity = identity)
- §XXIV's block-collision machinery already covers a cert-carrying
  capsule -- no new risk, no change needed there
- ACL gate lives on the attaching human's credentials, since an
  unattended instance holds no secret to prove anything about itself
- xHCI/live-table machinery confirmed irrelevant, independent of the
  seed question

One real wrinkle flagged, not resolved: vm_identity_from_cert()'s
serial-number check binds to a physical drive_uuid an unattended
identity doesn't have -- needs Bob's call before implementation.

Punch list recorded; implementation not started per standing "plan
approval is not a start signal" rule.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 03:41:10 -04:00
Robert Allan JamesandClaude Sonnet 5 42d4bf3dad Stage D batch 7 (final): mama_forth_words.c groups 5+6 -- dictionary lookup/test words + Stadium primitives; Stage D closed (FABRIC-3.md §XXXII.6)
NAME>XT/RUNCAP-TEST/PAIR-TEST and the six STADIUM-* physics primitives
(STADIUM-ADMIT/STADIUM-EVICT/STADIUM-RES-PULL/STADIUM-RES-PUSH/
STADIUM-HEAT@/STADIUM-HEAT!), 12 sites -- closes mama_forth_words.c
and the entire kernel-only error-handling audit.

All 105 sites from the §XXXII.3 triage now accounted for: 28 already
correct, 75 silent sites fixed across repl.c/inference_words.c/
log_words.c/vm_core.c/mama_forth_words.c, 2 special cases resolved by
dropping the error per their own documented contract, 1 resolved via
console_println() per its own recursion constraint. Verified with awk:
zero remaining vm->error=1 sites in mama_forth_words.c lack a
diagnostic within the preceding three lines.

Three-arch clean qemu acceptance passed. This closes Stage D and the
USE/logging/audit thread opened in §XXXII; Stage E (human-vs-unattended
identity model) remains open, not started this pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 20:25:30 -04:00
Robert Allan JamesandClaude Sonnet 5 90c86c6006 Stage D batch 6: mama_forth_words.c group 4 -- identity/crypto words (FABRIC-3.md §XXXII.6)
MINT + mint_pop_string()/ZUSE-ELIGIBILITY-ADD/ZUSE-ELIGIBLE?/
ELEVATE-PUBKEY-UNPACK, 10 sites. mint_pop_string() gained a field_name
parameter so its diagnostics name which of MINT's four string
arguments failed (phone/email/username/full_name), rather than a
generic message that would leave the operator guessing. Diagnostic
placement matched each function's own sibling convention where one
exists (ZUSE-ELIGIBILITY-ADD), log_message() default otherwise.

Three-arch clean qemu acceptance passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 19:53:47 -04:00
Robert Allan JamesandClaude Sonnet 5 8b5300fc4f Stage D batch 5: mama_forth_words.c group 3 -- cross-VM execution/dispatch, both flagged special cases resolved (FABRIC-3.md §XXXII.6)
VM-STEP/VM-EXEC/VM-CALL (5 sites) gained console_println() diagnostics
matching their own existing sibling guards.

SWITCH-MARK-WORK and VM-HEAT (3 sites) resolved per §XXXII.3's own
recommendation: dropped vm->error entirely rather than diagnosing it,
matching each function's own doc comment ("must never error or spam
the console" / "does not print/error"). SWITCH-MARK-WORK is the same
function §XXVIII.3 already fixed once for an off-by-one that fired
silently on every MSG-SEND in the system -- the guard now genuinely
cannot repeat that by contract, not just by the threshold being right.
VM-HEAT's guards now push 0 and return, matching its own
always-returns-a-value stack effect.

Three-arch clean qemu acceptance passed. Zuse's own WIREBIND attach
(every boot) drives MSG-SEND -> SWITCH-MARK-WORK, so this batch's most
safety-critical fix is exercised by standard acceptance, not just
compiled.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 19:23:12 -04:00
Robert Allan JamesandClaude Sonnet 5 b42c3b195c Stage D batch 4: mama_forth_words.c groups 1+2 -- capsule/lifecycle words + BIRTH/START/KILL (FABRIC-3.md §XXXII.6)
15 of mama_forth_words.c's 48 silent sites fixed, split by functional
grouping per direct instruction: CAPSULE@/CAPSULE-HASH@/CAPSULE-FLAGS@/
CAPSULE-LEN@/CAPSULE-BIRTH/CAPSULE-RUN/EXEC (9), and BIRTH/START/KILL
(6).

Refined the diagnostic-placement rule: match whichever convention that
same function's other already-correct guards use, rather than
defaulting uniformly. BIRTH/START/KILL/EXEC each already had a
console_println() sibling guard ("name too long or empty") -- their
newly-diagnosed guards now match that, same shape as USE's own fix.
The CAPSULE*@ words have no sibling guard to match, so they keep
log_message() (defer_words.c's gold-standard default).

Three-arch clean qemu acceptance passed; BIRTH itself is exercised by
every boot (Hermes/Artemis birth).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 18:51:54 -04:00
Robert Allan JamesandClaude Sonnet 5 66c2f3e539 Stage D batch 3: fix silent error sites in vm_core.c, incl. the two highest-value primitives; two real NULL-deref bugs found and fixed (FABRIC-3.md §XXXII.6)
All 13 silent vm->error=1 sites in vm_core.c now log a diagnostic
first: vm_enter_compile_mode, vm_compile_word, vm_compile_literal,
vm_compile_call, vm_exit_compile_mode, execute_colon_word (2 sites
each/combined), and the four fundamental memory primitives
vm_load_u8/vm_store_u8/vm_load_cell/vm_store_cell.

Found and fixed two real NULL-pointer-dereference risks while adding
the diagnostics: vm_compile_call() and vm_exit_compile_mode() each had
a combined `if (!vm || <cond>) { vm->error = 1; ... }` guard that
dereferenced vm->error even on the !vm branch of its own condition.
Split both, and applied the same defensive split to the four memory
primitives since vm_ptr()/vm_addr_ok() both tolerate vm==NULL
internally.

Live-verified, amd64: 999999999 @ . recovers correctly at Zuse's own
console (Stage A's mechanism holds), but the new vm_load_cell
diagnostic itself didn't print -- traced to memory_words.c's own
redundant, still-silent vm_addr_ok() pre-check in memory_word_fetch()
(and the same shape in memory_word_store()), which intercepts before
ever reaching vm_load_cell(). memory_words.c is vendored, out of this
initiative's scope, spun off to FABRIC-4.md -- recorded as a concrete
cross-reference for that future work rather than left to be
rediscovered.

Three-arch clean qemu acceptance passed, full POST suite included.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 17:29:37 -04:00
Robert Allan JamesandClaude Sonnet 5 1a8c0e4fcf Stage D batch 2: fix silent error sites in log_words.c, resolve the flagged category-iii special case (FABRIC-3.md §XXXII.6)
log_word_set_level(), log_do_emit(), and log_emit_string() (7 sites
total) now log a diagnostic via log_message(LOG_ERROR, ...) before
setting vm->error, matching this file's own already-correct
log_str_emit() sibling and defer_words.c's gold-standard pattern.

log_word_append_raw() (backing (LOG-APPEND-RAW), 3 sites) resolved
differently per §XXXII.3's own triage note: its doc comment forbids
log_message() here (recursion into the log ring it writes to), but the
word can also be invoked by hand at the console -- console_println()
carries no such recursion risk and matches Stage B's policy for a
manual interactive invocation. Added the console.h include this
required.

Three-arch clean qemu acceptance passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 16:53:18 -04:00
Robert Allan JamesandClaude Sonnet 5 5417ffb2bf Stage D batch 1: fix silent error sites in repl.c and inference_words.c (FABRIC-3.md §XXXII.6)
repl.c's sk_word_blk_attach_ack() and inference_words.c's array_ptr()
helper + infer_word_run()'s allocation guard now log a diagnostic via
log_message(LOG_ERROR, ...) before setting vm->error, matching
defer_words.c's own gold-standard pattern (§XXXII.3) -- these are
internal/background conditions (a malformed message-callback, a bad
array reference or allocation failure), not interactive usage mistakes,
so the fix keeps the fault and reports it rather than dropping it like
USE's own fix did.

Batched together (4 sites total, smaller combined than the next file)
rather than two separate acceptance cycles for negligible size.

Three-arch clean qemu acceptance passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 16:23:14 -04:00
Robert Allan JamesandClaude Sonnet 5 09959c6ca2 Stage C: primitive error-handling audit triage complete, 105 sites classified, no code changed (FABRIC-3.md §XXXII.3)
Read every one of the 105 vm->error=1 sites across the 8 kernel-only
files in context (not sampled): 28 already correct (diagnostic before/
without erroring, e.g. defer_words.c's 15-for-15 gold-standard
pattern), 75 silent (the exact USE-defect shape), 2 special cases
requiring individual handling rather than a generic fix.

Two findings flagged above the rest: SWITCH-MARK-WORK and VM-HEAT
(mama_forth_words.c) each set vm->error on their own guard despite
their own doc comments explicitly saying they must never error --
SWITCH-MARK-WORK is the same function whose off-by-one already caused
a stray silent error to fire on every MSG-SEND once before (§XXVIII.3).
And vm_core.c's four fundamental memory primitives (vm_load_u8/
vm_store_u8/vm_load_cell/vm_store_cell) silently fault on any
out-of-bounds address -- the highest-reach fix candidates in the audit,
hit by far more FORTH words than any single mama_forth_words.c site.

Full per-file site list and per-category breakdown recorded for Stage
D's own reference. Triage only -- no code changed this stage.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 14:36:37 -04:00
Robert Allan JamesandClaude Sonnet 5 c05b70c8d6 Stage B: logging policy documented, level-aware log-ring eviction built, LOG-FLUSH deferred again (FABRIC-3.md §XXXII.4)
Policy decided for the kernel-only audit scope: an interactive
command's direct response stays on console_println/console_puts;
everything else (state transitions, background diagnostics, audit
trails) routes through log_message() at the appropriate level,
matching capsule_mint.c's verify_mint() precedent. Documented, not
code-swept here -- reclassifying individual sites is Stage D's job.

LOG-FLUSH deferred again, explicitly: the per-VM log buffer its own
doc comment presumes (vm_log_buffer.h) doesn't exist anywhere in the
tree -- building it is real feature work needing its own scoped stage.

Level-aware eviction built: log_region_append() now reads the oldest
ring slot's own level before evicting it, protecting ERROR/WARN
records from being pushed out by INFO/DEBUG churn -- drops the
incoming low-priority record instead. Found and fixed an adjacent bug
while making this change: the prior two-valued return contract would
have made a benign "dropped by design" outcome indistinguishable from
a genuine write failure to its one caller, which unconditionally set
vm->error on any nonzero return. Changed to a three-valued contract (0
success, 1 dropped by design, -1 genuine failure).

Three-arch clean qemu acceptance passed. Eviction path itself not
live-exercised (needs 128+ LOG-APPEND calls to fill the ring) --
flagged, matching this project's own precedent for that kind of gap.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 14:01:36 -04:00
Robert Allan JamesandClaude Sonnet 5 5c5896fbc1 Stage A: fix USE's silent stack-underflow/bad-address guards, and the Hera fault-scoping gap they exposed (FABRIC-3.md §XXXII.1)
mama_word_use()'s two silent vm->error=1 guards (dsp<1 stack underflow,
NULL from vm_ptr()) now print a diagnostic and return, matching the
function's other five guards.

Live verification of that fix alone surfaced a bigger problem: with the
guard no longer silent, the REPL proceeds to interpret the leftover
token as an unrecognized word, which independently sets vm->error, and
sk_repl_step()/sk_repl_run()'s Hera-branch still hard-halted on that.
Investigated kernel_main.c's boot/capsule-load paths directly: they
already catch and clear mama->error entirely separately, before
sk_repl_run() is ever entered -- so the "no fallthrough surface" halt
in these two REPL functions was never protecting a boot-time fault, only
an ordinary interactive REPL-turn one. Both functions now recover
unconditionally on any VM's error, Hera included, matching how a
redirected (WIREBIND/USE'd) identity's session already recovered.
Removed the now-fully-unused sk_fault_handler().

Verified live, amd64: USE rajames at Zuse's own console prints the new
diagnostic, then "VM fault -- session recovered, resuming", console
stays interactive afterward. Three-arch clean qemu acceptance passed
before and after the Hera-fault-scoping change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 09:06:33 -04:00
Robert Allan JamesandClaude Sonnet 5 05159c9f9e Scope new initiative: unattended identity model, USE fail-closed-halt root cause, primitive error-handling audit, logging cleanup (FABRIC-3.md §XXXII)
USE crash fully root-caused against current source: mama_word_use() has
seven guards, two of which set vm->error silently (stack underflow,
bad VM address) while the other five print diagnostics; sk_repl_step()
halts the whole kernel only when the faulting VM is Hera; Zuse's
console runs directly on Hera's own VM (confirmed in
capsule_zuse_boot.c), making a benign typo at her prompt the one
deterministic path to the fail-closed halt. Three fix options named,
none applied yet.

Human-vs-unattended identity/console-birth model scoped: console/user
VM pairing is pure name-convention + a per-console-VM VM-NAME-REG
call, so after-the-fact console attach needs no new mechanism. Named
the invariant an unattended birth path must not violate (never touch
the global driving the "no thumbdrive, no prompt" gate).

Error-handling audit scoped kernel-only (mama_forth_words.c + 5
kernel-only word_source files + vm_core.c/repl.c): 107 vm->error=1
sites across 8 files, counted directly. The vendored word_source sweep
is explicitly out of scope here, spun off as future FABRIC-4.md work.

Logging cleanup sequenced ahead of the audit's own fixes, with the two
open §XXVII gaps (LOG-FLUSH doesn't exist; FIFO eviction is
level-blind) named for an explicit in/out-of-scope call.

Staged A-E plan proposed; nothing implemented this pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 07:36:35 -04:00
Robert Allan JamesandClaude Sonnet 5 16cc74243c std79 DoE campaign rerun post-Stage-4: 81/81, §XIV caught live by rerun's own harness bug (FABRIC-3.md §XXXI)
Reran the established 3x9x3 randomized full-factorial campaign
(std79-doe.fth) on all 3 architectures per standing project discipline
(any Stadium-adjacent change reruns the whole DoE from the top).

First amd64 attempt exposed a real test-harness bug that re-triggered
the already-known §XIV concurrent-attach gap: the sequential-attach wait
loop checked for any recent WIREBIND-attached line instead of the
specific identity requested, firing the next device_add before the
kernel finished the current one. Only 2 of 8 identities attached; the
campaign itself completed cleanly with graceful "VM-EXEC: VM not found"
refusals rather than corrupting anything. Fixed the wait loop, discarded
the invalid run's campaign result (its boot log kept for the record),
reran clean.

Corrected reruns: all 8 identities individually confirmed on all 3
architectures, zero VM-not-found errors, zero faults, 81/81 trials
correct against established baseline values. Raw logs archived at
experiments/std79-doe/results-20260915-stage4/.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 04:56:31 -04:00
Robert Allan JamesandClaude Sonnet 5 597f5a6cd8 Stage 4 verified: reap mechanism proven, real WIREBIND multiuser+multitasking confirmed live (FABRIC-3.md §XXX)
Reap mechanism (increments 2+3): temporary probe using Artemis as a safe
stand-in parked identity, 3/3 checks PASS on all 3 architectures (refuse
on SWITCHED_OUT, correct post-reap state, switch-signal slot released).
Probe reverted, all 3 architectures re-verified clean.

Live multiuser verification (increment 4, no code changes): a real
previously-unattached identity thumbdrive attached via QMP on a running
boot on all 3 architectures. VM-EXEC dispatch into her own live VM
computed correctly, tagged with her own name in console output, full
Tripod fleet unaffected. EJECT cleanly tore down both her VMs and
released the switch-signal slot -- confirms increment 1's per-device fix
under a real live attach/detach.

Found and flagged, not fixed: interactive USE on a freshly-attached
identity halts the kernel outright. Confirmed NOT caused by Stage 4 --
reproduced identically on the commit before any Stage 4 work. Direct
VM-EXEC dispatch into the same identity works correctly; this is
specific to the USE/BINDSTEP codepath, plausibly never caught before
since every prior identity campaign used VM-EXEC, never interactive USE.

Stage 4's original scope is complete: multitasking (Tripod) and
multiuser (WIREBIND) are now genuinely composed, verified against real
hardware-driven identity attach on all 3 architectures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 03:39:37 -04:00
Robert Allan JamesandClaude Sonnet 5 d9da82b065 Stage 4 increments 2+3: WIREBIND VMs as switch-signal participants + mark-and-defer tombstone reap (FABRIC-3.md §XXVIII Stage 4)
Increment 2: WIREBIND user VMs (the ones that actually run FORTH work;
console VMs are pure REPL proxies and never participate) register as
Stage 3 switch-signal participants at attach, unregister at teardown.
Slot table bumped 8 -> 16, matching messaging.4th's own VM-MAX -- a real,
already-agreed ceiling, not an invented number. Added
sk_vm_switch_signal_unregister() (compaction-based; Tripod VMs never
needed removal, WIREBIND VMs cycle constantly and would otherwise
exhaust the bounded table).

Increment 3: implements the plan's own ratified option (A) for the
async-detach UAF risk -- mark-and-defer via a new pending_reap flag on
VMRegistryEntry, deliberately not a new VMState (capsule_vm_kill()
already treats VM_STATE_DEAD as idempotent success, which would silently
swallow a reap attempt; SWITCHED_OUT still accurately describes a
tombstoned VM until the moment it's actually freed). unclean_detach()
sets it when capsule_vm_kill() refuses a SWITCHED_OUT target; the Stage 3
checkpoint (vm_core.c) checks it before ever attempting to resume a
pending switch target, and calls the new capsule_vm_force_reap() instead
-- the one caller allowed to bypass capsule_vm_kill()'s own refusal,
because it runs at the exact safe cooperative point the switcher itself
controls. A new idle-tick sweep cleans up the WIREBIND live-table entry
once the reap has actually happened.

Verified clean on all 3 architectures (baseline regression -- no
WIREBIND attach happens in a plain boot). The reap mechanism's own
correctness under a genuinely parked context is verified separately,
next, via a temporary deterministic probe.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 00:58:34 -04:00
Robert Allan JamesandClaude Sonnet 5 9f0f33dfc5 Stage 4 increment 1: per-device WIREBIND tracking, fixing a real multi-identity detach leak (FABRIC-3.md §XXVIII Stage 4)
capsule_wirebind_unclean_detach()/eject() tracked "the attached identity"
as a single global, correct for the console-pairing UX (one physical
console) but wrong for detach safety: since §XV/§XVI proved multiple
identities genuinely live simultaneously via this same attach path, every
attach after the first silently overwrote the singleton, so an unclean
detach of any but the most-recently-attached identity was silently
ignored -- that VM leaked forever, no trace in the log.

Adds a per-device live-identity table, separate from the (unchanged)
console-pairing singleton, so unclean-detach resolves any attached
device to its own identity. Sized off messaging.4th's own VM-MAX (16)
minus Tripod's 3 reserved slots, not an invented number. Corrects the
stale "single-USB-device constraint" doc claim in capsule_wirebind.h,
false since §XV/§XVI.

Groundwork for Stage 4's real deliverable (WIREBIND VMs as switch-signal
participants) -- this increment only fixes detach targeting; switch-
signal registration is next.

Verified clean on all 3 architectures (no WIREBIND attach happens in a
plain boot, so this is a regression check on the existing Tripod-only
path; live multi-identity verification comes with the switch-signal
registration increment).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 00:02:20 -04:00
Robert Allan JamesandClaude Sonnet 5 1c220ad4b4 Close §XXVIII: defer 2 known bugs to Stage 4 (FABRIC-3.md §XXIX)
Bob's call after an honest end-to-end status check: multitasking (fixed
Tripod fleet) and multiuser (Zuse/WIREBIND) each work for their own
tested paths, but two known bugs remain rather than zero -- the §XIV
concurrent WIREBIND attach detection gap, and the never-reconciled
MSG-TICK/Stage-3-switch dual-ownership rough edge. Both deliberately
deferred to be addressed during or at the close of Stage 4, since both
bear directly on WIREBIND VMs joining the switch-signal population.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-14 21:34:58 -04:00
Robert Allan JamesandClaude Sonnet 5 05ae7aa886 Fix SWITCH-MARK-WORK off-by-one: was tripping vm->error on every MSG-SEND (FABRIC-3.md §XXVIII.2)
mama_word_switch_mark_work()'s stack-underflow guard checked dsp < 2,
requiring 3+ items, when it only ever needs the 2 IDX>NAME leaves it
(caddr u). Since dsp is index-based (2 items == dsp 1), this rejected
every normal call. MSG-SEND tail-calls SWITCH-MARK-WORK unconditionally,
so this fired on every message sent anywhere in the system -- visible
only where a caller happened to check the target VM's error flag
afterward (mama_word_vm_exec()'s "VM-EXEC: ERROR in Artemis" report).

Verified clean on all 3 architectures: boot reaches Startup: Artemis
live -> zuse@Hera] ok> with no VM-EXEC: ERROR line at all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-14 20:17:01 -04:00
Robert Allan JamesandClaude Sonnet 5 66beae7fd4 Stage 3 follow-on: message-arrival eligibility hook + trampoline-blind switch-storm fix (FABRIC-3.md §XXVIII.2)
Implements the message-arrival eligibility signal FABRIC-3.md §XXVIII.1 left
open (has_work per-slot flag, set via new SWITCH-MARK-WORK primitive from
MSG-SEND) so an idle VM never becomes a switch target purely by waiting out
the readiness threshold.

Also root-causes and fixes a second, independent switch-storm: the tick's
"who is current" check used vm_log_attributed_vm(), which can't see a VM
parked in switch.c's own raw trampoline. Replaced with a dedicated
g_switch_current_vm tracked by the switch mechanism itself, and moved
target-slot eligibility reset to the switch decision point instead of
relying on ISR polling to observe a window that can be only a few
instructions wide.

Verified live on all 3 architectures: clean boot to zuse@Hera] ok>, live
cross-VM message dispatch, and (since a quiet log looks identical to a
livelocked storm once the DoE probe is gone) confirmed genuine REPL
liveness via QMP send-key + screendump on aarch64/riscv64, not log
inspection alone.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-14 17:39:20 -04:00
Robert Allan JamesandClaude Sonnet 5 862d7d9c48 Stage 3 follow-on: fix stack-ownership corruption + DoE switch columns (FABRIC-3.md §XXVIII.1)
DoE CSV gained 6 switch-signal columns (switch_count_cumulative,
switch_current_slot, switch_*_readiness, switch_ticks_since), and verifying
them with a boot-time HB-ON probe surfaced a real livelock: the preemption
checkpoint could fire inside a VM-EXEC-nested execute_colon_word() call and
switch away from a stack it didn't own, parking a borrowed region of the
caller's stack under the wrong VM's saved-context pointer. The trampoline
bounce was the visible (safe) half of this; the corruption was the quiet
half, live in every prior "clean" Stage 3 boot without ever showing up in
the log.

Fixed by gating the checkpoint on being at the outermost vm_interpret()
call (g_vm_interpret_depth / sk_vm_at_outermost_interpret(), vm_core.c),
per Bob's decision. Also fixed two related bugs found in the same pass:
g_switch_back_to was a single global stale after first entry, now per-VM
state (native_switch_back_to); note_switch_performed() fired on resume
instead of switch-out, now called before the switch.

Verified on all 3 architectures: steady log growth (no freeze), zero
leaked QEMU processes, DoE columns internally consistent, Hermes/Artemis
confirmed genuinely executing (not just trampoline-bouncing). Temporary
HB-ON boot probe reverted after capture.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 23:58:59 -04:00
Robert Allan JamesandClaude Sonnet 5 986d042aa7 Stage 3: timer-driven preemptive switching, live on all 3 arches (FABRIC-3.md §XXVIII)
Fourth stage of the preemptive context-switching plan, and the biggest.
LithosAnanke now genuinely, continuously preempts between Hera, Hermes,
and Artemis -- timer-driven, running live for the entire remainder of
every boot once the Tripod fleet registers, not a bounded probe.

A real design fork was resolved before writing code: the naive approach
(the timer ISR calling Stage 2's sk_vm_context_switch() directly) is
broken -- Stage 0's trap frame lives on whatever stack was active at
interrupt time, and jumping to a different stack via Stage 2's own
independent swap mid-handler would abandon that trap frame unresumed,
guaranteed corruption on the first tick. Chose the safer of two named
options: the ISR only ever sets a flag and returns completely normally
through its own full epilogue; the actual switch happens moments later,
via Stage 2's already-proven mechanism, at a safe cooperative checkpoint
on the mainline (execute_colon_word()'s per-word dispatch loop, checked
on literally every word, not throttled to the existing 256-word
heartbeat-tuning cadence) -- confirmed with the user that word-level
granularity is fine-grained enough given the eventual Zynq FPGA target
where a word is a mnemonic.

New capsule_vm_switch_signal.c/.h: a purpose-built run-readiness signal,
deliberately separate from capsule_vm_physics.c's execution-heat engine
(that one's own header documents itself as never touched from interrupt
context, by design). Slot table sized with headroom (8) rather than
hardcoded to today's 3 participants, so extending participation later is
another register() call, not a redesign -- per direct request to leave
room for swapping the participant set. Simple linear accumulate-then-
threshold for this first cut; a fancier law can replace it later without
touching the mechanism around it. heartbeat_tick() gains its one
deliberate, documented amendment to this file's own top-half/bottom-half
discipline -- the first time this codebase reaches into VM-scheduling
state from real ISR context.

Registration happens only after all three VMs are fully born, right
before the REPL starts -- no critical-section protection yet against
being switched away mid-birth-setup.

Known, flagged rough edge (not reconciled this pass): MSG-TICK's own
idle-pump and this new mechanism can still independently move control
between the same VMs; not observed to interact badly in verification,
but not fully unified either.

Verified interactively at the console on all 3 architectures with
continuous background preemption running throughout -- amd64 computed
`1 1 + .` -> 2, aarch64 computed `1 1 + dup DUP * . CR` -> 4, both
correct, REPL fully responsive, zero fault indicators over sustained
runtime.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 19:32:36 -04:00
Robert Allan JamesandClaude Sonnet 5 f790d0995e Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Third stage of the preemptive context-switching plan. The real
save/restore switch mechanism now exists -- the first time anything has
ever executed on a VM's own native stack (Stage 1 allocated them,
unused).

New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function
call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already
covers every caller-saved register; only the callee-saved set needs
explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV
has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64:
s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for
both halves within one call, since FS is genuine global CPU state, not
part of what's switched). A sibling sk_vm_switch_prime() in the same file
builds the synthetic first-entry frame, kept in assembly so the layout
can never drift out of sync with sk_vm_switch_to() itself.

New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry
priming vs. resuming a parked context, and updates registry state (new
VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no
live frame, this means the opposite). sk_vm_switch_entry() is the minimal
permanent trampoline every freshly-entered VM lands in: no production
behavior defined yet, so it just yields straight back to whoever switched
to it, forever.

Closes the confirmed unguarded-KILL UAF found during planning:
capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama()
all now refuse (or silently leak rather than free, on the cold-restart
path where arch_cold_reset() wipes everything immediately after anyway)
tearing down a switched-out VM. Side effect found, not built on purpose:
the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it
automatically stopped dispatching into a switched-out VM with zero
changes needed there.

Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing
can type interactively into a foreground-only QEMU session) that
round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3
architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe
fully reverted after capture; kernel_main.c shows zero diff.

Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new
files added explicitly (this project's "loader" PE binary is the full
running kernel, not a thin bootstrap stage), and aarch64's switch.S
needed the same #ifndef _WIN32 guard around .hidden that isr.S already
carries (aarch64's loader assembles via clang targeting a PE/COFF
target with no .hidden equivalent) -- caught by a build failure, fixed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 15:25:39 -04:00
Robert Allan JamesandClaude Sonnet 5 57ac3fc304 Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Second stage of the preemptive context-switching plan. Every VM (Hera,
every capsule_birth_baby()-born VM including WIREBIND identities) now
gets its own dedicated 2 MiB native C stack at birth -- but nothing runs
on it yet, that's Stage 2. Pure allocation-machinery proof.

Design correction made before writing code: the plan called for cloning
sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to
be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a
plain kmalloc() block for its dictionary arena, not a real guarded PMM
allocation. Stacks get the real treatment instead (new
sk_vm_native_stack_alloc()/_free() in arena.c): independent
pmm_alloc_contiguous() + guard pages for every VM without exception, no
singleton, no kmalloc fallback -- a stack overflow is exactly the
failure mode guard pages exist for, and a corrupted stack could corrupt
whatever saved context Stage 2 trusts.

2 MiB size matches this project's own established kernel-stack
convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one
shared 2 MiB stack today already carries all VMs' combined nested
VM-EXEC recursion.

Three new VM struct fields, freed in vm_cleanup() alongside the existing
call_stack free. Allocation failure is non-fatal to birth.

All 3 architectures re-verified clean boot to ok>, no native-stack
allocation failures for any Tripod-fleet VM.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 14:30:46 -04:00
Robert Allan JamesandClaude Sonnet 5 15672ce17c Stage 0: trap-frame parity across all 3 arches (FABRIC-3.md §XXVIII)
First stage of the preemptive context-switching plan (see
~/.claude/plans/logical-snuggling-bear.md). Pure foundation work -- every
arch's ISR now saves the full register set on interrupt entry, so a trap
frame is in principle sufficient to resume execution anywhere it was
taken. No FORTH-visible behavior changes.

amd64: added FXSAVE/FXRSTOR, closing a genuine pre-existing correctness
gap (not just future-preemption prep) -- confirmed live double-precision
FP code reachable from ordinary interpreter dispatch (vm_runtime.c Loop
#5/#6), and the ISR previously saved zero FP/SSE state. rbp repurposed as
a fixed anchor so the 16-byte-aligned FXSAVE area can be carved out of an
unpredictably-aligned rsp without disturbing existing argument reads.

aarch64: extended the trap frame 672->800 bytes, adding v8-v15 (AAPCS64
callee-saved, previously excluded on call-site-only reasoning that
doesn't hold for an async trap).

riscv64: extended the trap frame 320->512 bytes, adding s0-s11 and
fs0-fs11 (the latter still correctly gated behind sstatus.FS != Off).

All 3 architectures re-verified clean boot to ok> under the new frames --
amd64 through hundreds of timer ticks with FXSAVE/FXRSTOR live on every
interrupt, aarch64 through 987 ticks, riscv64 clean on the now-larger
FS-conditional block.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 09:23:28 -04:00
Robert Allan JamesandClaude Sonnet 5 2a30212bd3 Real per-VM log persistence: source attribution + ACL pin (FABRIC-3.md §XXVII)
Wires the previously-unused vm_log_attributed_vm() into LOG-APPEND's kernel
primitive so persisted log records carry a trustworthy source (the real
attributed VM's registry name, or "HADES" pseudo-source) instead of a
caller-supplied, trivially forgeable string. Drops src-addr/src-u from
LOG-APPEND's stack signature accordingly. Pins LOG-APPEND via bare ACL-PIN
in Artemis's own init.4th, matching BIRTH/CAPSULE-BIRTH's precedent for a
privileged word that can't reach the shared, host-portable ACL.4th.

Also fixes two console-banner nitpicks: a mis-rendering em dash (U+2014)
in the boot banner, and drops "Emergency" from the CLI banner text.

Doc corrections to artemis_sig.h/zuse_eligibility_list.h reconciling the
three fixed devblock ranges now in play. LOG-FLUSH (the intended normal
entry point) and level-aware log eviction remain open, flagged not fixed.

Re-verified clean boot to ok> on all 3 architectures after every change.
riscv64 showed one new, unrelated virtio_blk write-timeout anomaly during
Artemis's early physics self-test (self-recovered, boot unaffected,
sector doesn't map to the log region) -- flagged, not investigated.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 08:56:33 -04:00
Robert Allan JamesandClaude Sonnet 5 61755fde78 Artemis genesis stamp: fix a BAM-corrupting offset before it ever ran (FABRIC-3.md §XXVI follow-on, Step 3)
Step 3: one-time artemis_sig_t genesis stamp, written once
kernel_main.c's virtio-blk path confirms Artemis's own disk, so the disk
image is later recognizable generically (repl.c's idle-loop USB-MSC scan,
built in the prior commit) regardless of which bus found it.

Correction made before this ever touched the real disk: the signature's
first design (committed in 29b6789) placed it at a fixed bottom-of-device
forth-block (4, devblock 1) -- copying homeblocks_sig_t's own convention,
which is safe for a raw identity thumbdrive but not for Artemis's own
disk. Artemis's disk is block_subsystem.c's own STFR/v2-formatted volume:
devblock 0 holds that format's header and devblock 1 is the FIRST
DEVBLOCK OF THE LIVE BAM (blk_compute_fresh_geometry(): bam_start = 1).
The original design would have overwritten Artemis's live allocation map
on the very first real boot. Caught via direct cross-reference against
block_subsystem.c before the genesis-stamp call site was ever run against
the real image -- no corruption occurred.

Fixed by moving the header to a fixed offset from the END of the device
instead (ARTEMIS_SIG_DEVBLOCK_FROM_TOP=64), the same top-of-device region
block_subsystem.c's own meta_fence_blocks reservation (128 devblocks)
already carves out for system metadata, and where Zuse's genesis marker/
eligibility list already live -- but computed independently via
blkio_info() rather than through blk_meta_zone_*(), since that accessor
needs an already-attached, format-detected slot, which is exactly the
state pre-attach generic discovery doesn't have yet. Picked well clear of
Zuse's two tenants (devblock_from_top 0 and 1+, open-ended) so the two
subsystems' independent math can never collide.

Also reordered kernel_main.c: rng_init() now runs before the Artemis
virtio-blk block (was after) -- the genesis stamp needs rng_get_bytes()
for disk_uuid, and the original order would have failed the stamp on
every boot.

Verified live: booted amd64 against the real disk/artemis.img twice --
first boot logs "Artemis: genesis signature stamped" (confirmed blank at
the target offset beforehand via a host-side read), second boot on the
now-stamped image logs no re-stamp (idempotent, CRC/read-back verified)
-- both boots and aarch64/riscv64 (against the same now-stamped image)
all still report "Artemis: 22998 data blocks" / "PASS: persist-read"
unchanged, confirming the BAM and data pool were never touched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 07:31:21 -04:00
Robert Allan JamesandClaude Sonnet 5 29b6789860 Artemis bus-agnostic discovery: signature format + idle-loop generalization (FABRIC-3.md §XXVI follow-on)
Step 1: new artemis_sig_t header format (magic 'ARTM', sibling to
homeblocks_sig_t, distinct so a generic scan can tell Artemis's own disk
apart from an identity thumbdrive by content alone) -- artemis_sig.h/.c,
wired into Makefile.starkernel.

Step 2: sk_repl_idle()'s existing per-USB-MSC-slot attach handling (the
pattern WIREBIND already uses for identity thumbdrives) now also checks
for the ARTM signature whenever a device's home-blocks check comes back
BLANK. On a match, once Artemis's own storage-attach round-trip
(HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK) confirms success,
capsule_zuse_boot_load_root_pubkey() runs -- the same call kernel_main.c's
synchronous QEMU-only virtio-blk path already makes, now reachable
without a hardcoded PCI vendor/device scan. That function is already
idempotent (no-op once zuse_root_pubkey_known is set), so no boot
restructuring was needed despite the initial concern that deferring
Artemis discovery to the idle loop would require one.

Verified: clean build + QEMU boot to [zuse@Hera] ok> on all three
architectures, zero regression to the existing virtio-blk/Zuse-thumbdrive
attach path.

Steps 3 (genesis-stamping onto disk/artemis.img) and 4 (growable
production log-persistence region) not yet started.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 07:16:48 -04:00
Robert Allan JamesandClaude Sonnet 5 a8b16d41da Unify console prompt to [user@VM]; fix real personality-block truncation; correct §XXV's wrong lockdown conclusion (FABRIC-3.md §XXVI)
Investigating the std79 lockdown finding from FABRIC-3.md §XXV led
to a real discovery: WIREBIND births TWO VMs per identity, a console
proxy under the plain username and the actual restricted identity
under <username>~user (capsule_wirebind.c). Every test in §XXV
targeted the console proxy, which was never locked down at all.
Retested against the correct target (rajames~user): the lockdown
works exactly as designed. §XXV's "lockdown never engages" conclusion
was wrong -- corrected here, not deleted, since the mistake and how
it was caught are worth keeping (see the new feedback memory:
confirm which specific VM a name resolves to before concluding
anything, when a subsystem is known to birth more than one VM per
identity).

Two real, separate things found along the way are kept regardless
of that correction:

- capsule_runcap.c: the reserved personality devblock was read in
  full (mostly zero-padding after a short ~200-byte string) with no
  terminator, producing "WARN: block 4998 exceeds 1KB, truncating"
  on every std79-locked identity's birth, universal, since at least
  2026-09-10. Fixed by trimming to the first NUL byte actually found
  -- real, but harmless to execution (real content sat in the
  truncated block's surviving head); it mattered for capsule_id/
  content_hash being computed over padding instead of real content.

- console.h/console.c/repl.c: unified the prompt from a separately-
  computed "[VMName] (user)" into a single "[user@VMName]" line
  prefix -- exactly the ambiguity that caused the original
  misdiagnosis (the prompt showed only the WIREBIND username,
  identical whether USE had targeted the console proxy or the real
  ~user identity). Implemented as a registered callback
  (console_set_user_prefix_provider()) rather than console.c calling
  into WIREBIND/session logic directly, since console.c is a clean
  HAL module with no prior dependency on capsule-level subsystems.

Verified: clean build on all 3 architectures, zero new warnings,
identical dict_hash/capsule_hash to every prior boot this session
(console/prompt-only change). Full 9-identity messaging campaign
re-run end to end: 202s, zero faults, all 8 identities at 99/99
tokens, zero regression.

Also surfaced, not yet acted on: the full campaign's own console
tags now visibly show which VM each identity's tests actually
reached ([zuse@rajames], not [zuse@rajames~user]) -- messaging.4th's
VM-NAMES-INIT registers identities by plain username, so std79-doe.
fth's turn-attractor has been dispatching to each identity's console
proxy, not the actual locked-down identity, since the messaging
rewrite. Flagged for a deliberate decision, not investigated further.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 06:44:58 -04:00
Robert Allan JamesandClaude Sonnet 5 cb32e6632b Root-cause the std79 lockdown gap: likely never engages for any locked identity (FABRIC-3.md §XXV follow-up)
Following up on the messaging-words-not-denied finding: 'MSG-STATUS
ACL-STD79-ALLOWED? .' sent into rajames came back UNKNOWN WORD, not
"denied" -- acl-std79.4th's own supporting words were never compiled
into the VM's dictionary at all.

Root cause, confirmed via the raw boot log: "WARN: block 4998 exceeds
1KB, truncating" fires during every std79-locked identity's own
birth. capsule_exec_payload() parses "Block N" content as everything
from that header to the next "Block N" header or end-of-payload,
truncating at 1024 bytes for both storage and execution.
MINT_RESTRICTED_PERSONALITY is a short (~200 byte) string written
into a zero-padded 4096-byte devblock at mint time; capsule_runcap_
birth() reads back the entire reserved multi-devblock region with no
second "Block N" header anywhere in it to terminate "block 4998"
early, so the parser treats the whole mostly-padding region as one
oversized block.

Confirmed universal, not one identity's quirk: the warning fires
exactly 8 times in a full 9-identity boot -- once per std79-locked
identity (rajames, 00-06; zuse runs natively on Hera, never through
this path) -- and dates back to at least 2026-09-10 in this repo's
own logs, well before this session. The std79 lockdown has likely
never actually engaged, for any of the 8 locked identities, since it
was built.

Not fully closed to the byte: the real content sits at the start of
the oversized block, inside the surviving truncated slice, so
truncating the tail shouldn't by itself stop the head from executing
-- the exact remaining mechanical step (stale leftover data, a
line-boundary artifact, or something else) isn't nailed down yet.

Explicitly not fixed -- documented per Bob's own "stop here, document
it, and continue" call. Security-relevant (an intended lockdown
restricting nothing) and deserves its own deliberate fix with real
test coverage, not a bolt-on to an unrelated change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 22:59:54 -04:00
Robert Allan JamesandClaude Sonnet 5 3c2daf50d1 Extend BIRTH/CAPSULE-BIRTH to all VMs symmetrically; flag a real std79 lockdown gap (FABRIC-3.md §XXV)
Scoping the workload-into-factorial design's placement-mode factor
led to a real architectural improvement: rather than EXEC-ing a
workload capsule into an already-running, ACL-locked identity's own
persistent dictionary (filesystem-shaped, doesn't dodge the block-
collision exposure just traced in §XXIV), a workload now runs as a
fresh ephemeral child VM, BIRTH'd per trial and reaped after --
matching the project's own stated principle of automanagement over
imposed policy. CAPSULE-BIRTH already passes vm->stadium_vm_id (who
is birthing this VM) as the new child's parent, not a hardcoded Hera
constant, confirmed by reading the C -- so a workload trial genuinely
inherits the specific identity's own lineage when that identity does
the birthing.

Which surfaced a real premise: only Hera could call BIRTH/
CAPSULE-BIRTH at all (registered only in register_mama_forth_words(),
confirmed directly, not part of the earlier §XX messaging-symmetry
fix which deliberately kept this as one of her remaining privileges).
Extended symmetrically now, agreed explicitly before touching code:

- mama_forth_words.c: BIRTH and CAPSULE-BIRTH added to
  register_child_vm_words(), matching §XX's own pattern.
- acl-std79.4th: ' BIRTH , ' CAPSULE-BIRTH , added to ACL-STD79-LIST
  (new block 4048) -- a deliberate, explicit, named exception to the
  lockdown's own "standard words only" guarantee, not a silent one.
  Symmetric registration alone can't weaken any lockdown on its own:
  ACL-LOCKDOWN-STD79 is allowlist-based, deny-by-default, so a newly
  registered word is auto-denied there unless explicitly added.

Verified: clean build on all 3 architectures, zero new warnings.
Hera's own dict_hash unchanged (expected); Hermes/Artemis show the
same new dict_hash on all 3 architectures. Live-tested against a
real attached std79-locked identity: CAPSULE-BIRTH executes
correctly (returns vm_uuid_none() for a deliberately out-of-range
capsule-id, zero fault, zero ACL denial).

Found, and explicitly stopped short of fixing, a separate pre-
existing gap while verifying the above: MSG-STATUS and MSG-K
(messaging.4th words, not on the std79 allowlist) execute for a
locked identity instead of being denied. ACL-LOCKDOWN-STD79 is
confirmed to actually run; something more specific isn't reaching
messaging.4th's dictionary entries. Root cause not traced -- needs
its own investigation into vm_core.c's dictionary-link mechanics and
whichever capsule actually loads messaging for these identities.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 22:31:52 -04:00
Robert Allan JamesandClaude Sonnet 5 cd2fda4351 Add mkcapsule --resolve: build-time claim registry for capsule block collisions (FABRIC-3.md §XXIV)
Traced what a "Block NNNN" collision actually means before designing
a fix for it: capsule_loader.c's block-write path routes through the
generic block-subsystem API, which kernel_main.c registers as two
devices in a fixed order -- the volatile ramdrive first (LBN
2048-3071), then Artemis's real virtio-blk device immediately after
(LBN 3072+, backed by disk/artemis.img). Every capsule this project
has lands in Artemis's persistent range, not the ramdrive, and
blk_update()'s dirty-marking + repl.c's idle-loop flush write that
content through to the real disk file on every boot. A block-number
collision is therefore a silent, persistent overwrite of real disk
content surviving reboots, not a transient RAM mixup.

The existing collision gate (check_block_conflicts(), already a hard
non-interactive build failure) already catches capsule-vs-capsule
collisions across the whole flat range. The real gap: zero visibility
into blocks something other than a capsule owns (Artemis's own
non-capsule persistent data), and no device-boundary/capacity
awareness at all.

Added, scoped step by step before writing any code:

- tools/capsule-claims.txt -- derived, auto-created/regenerated,
  git-ignored. Lets --resolve tell "this capsule's own content
  changed" apart from "genuinely new collision with something else."
- tools/capsule-reserved.txt -- human-authored, git-tracked, seeded
  with nothing yet rather than guessed at. Checked by both the plain
  build gate (new check_reserved_conflicts()) and --resolve.
- tools/patches/ -- git-tracked, one file per accepted interactive
  renumber; a structured old->new block list, not a generic diff,
  since that's the only thing a renumber ever changes.
- mkcapsule --resolve <dir> -- the only interactive mkcapsule mode,
  a deliberate separate invocation from the plain build path (which
  stays non-interactive so CI never blocks on a prompt). Suggests a
  renumbering that preserves a capsule's own existing block spacing,
  prompts y/N, rewrites the .4th source in place on acceptance.

Found and fixed a real bug during verification: the registry's
empty-block-list case (workload-5.4th, zero Block headers) serialized
with a stray trailing space that the reader parsed back as a phantom
block 0, causing spurious re-registration every run -- caught by
testing idempotency directly, not assuming a clean first run meant
it worked.

Verified: isolated collision tests confirm both accept and reject
paths, confirm a resolved collision doesn't re-prompt the other side,
confirm reserved-range collisions are caught by both --resolve and
the plain build gate. Full 3-architecture rebuild via the real
Makefile.starkernel succeeded clean; amd64 boots with an unchanged
dict_hash/capsule_hash from every prior boot this session.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 21:56:37 -04:00
Robert Allan JamesandClaude Sonnet 5 c8ba8832c4 Fix N-RUNS/START-REP hardcoded constants; N-REPS increase blocked on reservoir cost, not a bug (FABRIC-3.md §XXIII)
Two real bugs found and fixed while attempting to raise N-REPS from
3 (both harmless only by coincidence at N-REPS=3, since 3 happened
to equal the hardcoded/literal values):

- N-RUNS was `27 CONSTANT`, not derived -- now `N-ID N-REPS *
  CONSTANT N-RUNS`.
- START-REP's run_id decode used a literal `3 *` where it meant
  `N-REPS *` (confirmed against EXEC-STD79-DOE, the serialized
  baseline, which correctly uses N-REPS for the same decode).

Re-verified at N-REPS=3 (amd64): byte-identical to the already-
verified baseline -- 193s, 0 faults, 99/99 tokens every identity,
K conserved on all 656 rows.

Raising N-REPS to 6 was then tested and found to cause real, silent
data loss at full 8-identity scale (rajames 0/198, 00 66/198, 01/02
99/198, 03-06 fully complete) -- same failure class as §XXI defect 2,
just past the budget again since doubling N-REPS roughly doubles
total MSG-SEND volume (24->48).

Measured the reservoir's replenishment directly rather than assume
from source: 40 sends drained it from 21735 to 1; 300s of pure idle
time (zero further sends) brought it back to 2049 -- real, but only
~6.8 units/s, meaning a full refill would take on the order of 53
minutes against a campaign's few-hundred-second runtime. Raising
N-REPS further needs a cheaper per-send cost or an explicit top-up,
not reliance on ambient decay.

Reverted N-REPS to 3 (fixes kept, they're correctness fixes
independent of the value) rather than commit a silently-lossy result.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 20:03:21 -04:00
Robert Allan JamesandClaude Sonnet 5 5c1b31669e Rename init-0.4th..init-9.4th to workload-0.4th..workload-9.4th; fix stale WL-HI/WL-LO doc
These are alternate boot personality capsules, not files doe.4th
dispatches from -- confirmed by reading doe.4th itself (it generates
its own synthetic workload internally, DOE-WORK) and mkcapsule.c
(only the exact filename "init.4th" is special-cased as the active
MAMA_INIT capsule). "workload-N.4th" names them for what they are
without colliding with that reserved name.

Renamed the 10 files (git mv, preserving history) and their own
self-referential header comments (also fixed a pre-existing typo,
"init-4.th" -> "workload-4.4th"), updated capsules/README.md,
capsules/MANIFEST.md (21 references), .claude/CLAUDE.md, and a
tools/mkcapsule.c comment.

Also fixed experiments/bare_metal/README.md's "Adding a Custom
Workload Capsule" section, discovered stale while doing this rename:
it documented a WL-HI/WL-LO dispatch table and a wl_id CSV column
that don't exist anywhere in the current capsules/ tree or doe.4th's
own CSV header -- corrected to describe what's actually there (no
pluggable workload dispatch; a custom workload is run by substituting
it in as the boot's own init.4th).

Verified: mkcapsule --lint clean (35 files, 0 violations), all 3
architectures build with zero new warnings, amd64 boots to zuse)ok>
with an unchanged dict_hash/capsule_hash from every prior boot this
session (0xc8f4b09e36f4fc4a / 0x1ef4939ed32ec1e6) and zero faults.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 19:43:50 -04:00
Robert Allan JamesandClaude Sonnet 5 71bcb72a59 Fix O(N) idle-loop messaging pump; full 3x9x3 turn-attractor campaign clean on all 3 architectures (FABRIC-3.md §XXII)
sk_repl_idle()'s messaging pump (repl.c) walked the entire live-VM
registry every idle beat (~1Hz) and dispatched a full VM-EXEC
"MSG-TICK" -- a dictionary lookup plus a 32-slot arena scan -- into
every live VM, every tick, unconditionally, forever. Fine at Tripod's
original 3-VM scale; a full 9-identity turn-attractor campaign
exposed it as a genuine wall on riscv64 specifically (its TCG makes
each dispatch cost more): the same campaign that completed in 194s/
339s on amd64/aarch64 never finished on riscv64 at 9 VMs across three
attempts, while 8 VMs there was fine in 159s.

Ruled out capacity explanations before touching anything: bumping
riscv64's QEMU RAM 1024->4096 changed nothing (reverted), and a live
STADIUM-RES@/MSG-STATUS probe with all 9 VMs attached showed no
depletion. Host memory pressure was also ruled out directly (one
background task did get OOM-killed once during the investigation,
but the identical stall reproduced again with 9.2GB free). The real
mistake was three premature kills under 4 minutes with no way to
tell "slow" from "stuck" from outside the guest -- fixed by having
run_doe_batch.sh sample the qemu process's own /proc/<pid>/stat utime
every 60s; with that signal, riscv64 at 9 VMs was unambiguously alive
(climbing utime, no hang), just disproportionately slow going from 8
VMs (159s) to 9 (600s+ and climbing).

This was never really a riscv64-only bug: an O(N) per-second walk
over the full VM population doesn't scale to the hundreds of VMs this
fleet is headed toward, on any architecture -- riscv64 just made it
visible first, at N=9, because its per-dispatch cost is highest.

Fixed by round-robin batching: the pump now dispatches to at most
SK_MSG_PUMP_BATCH (4) live VMs per idle beat via a persistent cursor
that resumes where the previous beat left off, instead of all of them
every time. Bounds both the scan and dispatch cost to O(K) regardless
of total VM count; any single VM's queue now drains roughly every
ceil(N/K) beats instead of every beat, still bounded and still
matching the pump's own existing best-effort contract. No new C
primitives, no messaging/Stadium changes.

Verified: clean build on all 3 architectures, then the full 3x9x3
campaign re-run on all 3 (not just riscv64) per the standing rule
that a defect repair requires a clean re-run everywhere before
anything counts as closed:

  amd64   198s  0 faults  99/99 tokens x8  656/656 K-conserved
  aarch64 339s  0 faults  99/99 tokens x8  659/659 K-conserved
  riscv64 178s  0 faults  99/99 tokens x8  659/659 K-conserved

riscv64 went from "never completes" to faster than aarch64, same
campaign, same seed, same identity set. No regression on amd64/
aarch64.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 19:29:14 -04:00
Robert Allan JamesandClaude Sonnet 5 19916717b2 Turn-attractor rebuilt on real messaging: two live defects found and fixed before the full campaign (FABRIC-3.md §XXI)
Rewrote EXEC-STD79-DOE-CD's dispatch to coordinate via MSG-SEND/
MSG-TICK instead of blocking VM-EXEC, now that Hera can genuinely
message (§XX). RUN-TEST/EXEC-STD79-DOE (the serialized baseline) are
untouched, kept as a byte-for-byte-reproducible historical comparison
point.

A small 2-identity smoke test before any multi-architecture
commitment caught two real defects the design alone didn't predict:

1. Absolute VM-HEAT can't produce fine-grained interleaving -- a
   fresh identity starts at heat=0 against Hera's ~62000+, a gap no
   1..24 divisor closes, so priority locked onto whichever identity
   had executed least, for its entire campaign. Fixed with BASE-HEAT:
   each identity's heat is snapshotted once at campaign start, and
   priority is computed from heat gained *this campaign*, not
   lifetime heat.

2. Single-test-per-message granularity silently drops most of a
   campaign's data: MSG-SEND no-ops on MSG-ALLOC failure, and the
   turn bookkeeping advanced regardless, so rows looked complete
   while missing most of their tests. A live reservoir probe showed
   Hera's own STADIUM-RES@ draining ~725/cycle with no
   replenishment observed -- a 648-send full campaign would exhaust
   it almost immediately. Fixed by dispatching a whole rep (24 tests,
   one concatenated command, measured 442 bytes, well under
   VM-EXEC's 1025-byte cap) per message instead -- 27 sends for a
   full campaign, not 648.

Re-verified after both fixes: 99/99 expected test outputs present,
zero drops, zero faults, genuine rep-level interleaving instead of
either the serialized baseline's fixed order or the first cut's
72-test lock-in.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 16:53:01 -04:00
Robert Allan JamesandClaude Sonnet 5 7edccd2c35 Hera can't message: root-caused and fixed by symmetry (FABRIC-3.md §XX)
Hera never loaded common:messaging.4th, unlike every other VM in the
fleet. The standing belief was this was deliberate, to avoid moving
her dict_hash off baseline. Checked live instead of assumed: loading
messaging.4th into her dictionary silently dropped every colon-
definition referencing one of 8 STADIUM-* primitives that
register_child_vm_words() gives every other VM but
register_mama_forth_words() never gave her -- a missing-primitive gap,
not a designed privilege boundary. Confirmed mama_word_birth is
genuinely VM-agnostic and SPAWN-EVENT is an unwired placeholder before
proposing the fix.

Fix: register the same 8 STADIUM-* primitives for Hera, load
messaging.4th from init.4th the same way Hermes/Artemis/console/mint
already do, and give her own idle-loop context a direct MSG-TICK call
(not VM-EXEC, which would hit the same reentrancy class the existing
per-other-VM pump loop already guards against) so her own queued
messages actually drain. Her dictionary is now a proper superset of
every child VM's, plus her remaining extra privileges -- not
structurally different from any other VM, just additionally
privileged.

Verified: dict_hash identical across amd64/aarch64/riscv64
(0xc8f4b09e36f4fc4a), Hermes/Artemis dict_hashes unchanged and still
cross-arch identical, all three boot clean to zuse)ok> with zero
UNKNOWN WORD faults, mkcapsule --lint clean.

Unblocks rewriting the turn-attractor (FABRIC-3.md §XIX) to coordinate
via real MSG-SEND/MSG-TICK instead of blocking VM-EXEC.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 16:05:11 -04:00
Robert Allan JamesandClaude Sonnet 5 5a9425b91b Add VM-HEAT primitive: groundwork for a compudynamic turn-attractor (FABRIC-3.md §XIX)
Prompted by reading the std79-doe K report: checked whether the campaign's
"real cross-VM dispatch load" was actually concurrent or strictly
serialized. It's serialized at two levels -- VM-EXEC's vm_interpret(target,
...) is a direct synchronous C call (Hera fully blocked until it returns),
and even the background physics tick (vm_tick(), vm_runtime.c) is driven
by each VM's own execution loop, so idle identities accrue zero ticks
between their own turns. K's perfect conservation (FABRIC-3.md §XVIII)
verifies sequential per-VM accounting correctness, not concurrent-access
safety, since there was never concurrent access to test.

Agreed direction: fix this without a scheduler, by reusing the same
least-dense-candidate judgment stadium_admit() already trusts for
eviction, applied to "whose turn is next" instead of "who gets evicted" --
a fleet-level turn-attractor giving the next turn to whichever live
identity currently has the lowest execution_heat_q48, no fixed round-robin,
no priorities, no preemption. Lives beside Stadium in
capsule_vm_physics.c (already the fleet-level consumer of Stadium
primitives, e.g. the K mechanism itself), not inside stadium.c ("the
floor" -- residency/eviction, a different concern from turn order) and
not a new subsystem.

This pass lands only the primitive the mechanism needs: VM-HEAT
( c-addr u -- heat-q48 ), pushing a named VM's current
execution_heat_q48 via vm_physics_heat_of() -- previously C-internal
only (doe_log_heat_by_name(), doe_log.c), never exposed to FORTH. Silent
0 on an unknown/dead name (no print/error), since a turn-attractor
scanning many candidates every turn shouldn't have to filter console
noise for names that simply aren't live. Registered everywhere
VM-EXEC/VM-CALL already are. Builds clean on all three architectures;
live-tested on amd64: Hera -> 65452, Hermes -> 43, unknown name -> 0, no
faults.

The turn-attractor loop itself (std79-doe.fth's trial ordering) is not
yet built -- next step, not done here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 14:53:06 -04:00
Robert Allan JamesandClaude Sonnet 5 238ca95b3c Add closing note to std79 DoE K report
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 14:42:23 -04:00
Robert Allan JamesandClaude Sonnet 5 66ba21adb4 Log fleet_k_q48/fleet_conserved; K holds exactly, 775/775 ticks (FABRIC-3.md §XVIII)
doe_log.c's per-heartbeat-tick CSV gains two columns: fleet_k_q48
(vm_physics_fleet_heat_sum() over ALL live VMs -- the genuine fleet-wide
conservation invariant K, not reconstructable from the 3 named-Tripod-
member heat columns already logged, which omit every identity VM's own
heat) and fleet_conserved (vm_physics_conserved() as 0/1). Requested
explicitly after the first heartbeat-telemetry analysis pass
(analysis-20260912/) omitted K entirely.

Kernel rebuilt on all three architectures, full 3x9x3 campaign rerun
(results-20260912-with-k/). K = 1.0000000000 (Q48.16 raw 65536) on every
one of 775 heartbeat-tick observations, sd(K) = 0, 100% fleet_conserved,
across amd64/aarch64/riscv64, nine identities, three replicates -- zero
deviation. Also a free regression check on both recent Stadium fixes
(§XVI/§XVII): neither disturbed the reservoir-transfer accounting K
depends on.

Found and fixed a tooling wrinkle along the way: fleet_conserved, being
the CSV row's very last field with nothing after it to bound a regex
match, can have a resumed trial digit merge into it with zero separator
on the wire -- combine.py now derives it from fleet_k_q48 directly (same
epsilon vm_physics_conserved() uses) instead of trusting the raw field.
fleet_k_q48 itself is unaffected either way.

Full analysis, discussion, and light/dark SVG->PDF figures written up as
a proper LaTeX report (report-20260912/report/std79_doe_report.pdf),
following experiments/bare_metal/analysis/report/bare_metal_doe_report.tex's
established style -- supersedes analysis-20260912/'s markdown-only first
pass as the primary deliverable for this dataset (kept, not discarded).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 13:25:09 -04:00
Robert Allan JamesandClaude Sonnet 5 6573a6d1d5 Commit per-architecture correlated DoE CSVs, not just the merged one
These are data in their own right (the per-trial-labeled telemetry, one
step before merging into combined.csv), not disposable scratch -- keep
them alongside it rather than only documenting how to regenerate them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 12:47:03 -04:00
Robert Allan JamesandClaude Sonnet 5 934be5a257 Add 3x9 factorial analysis of std79-doe heartbeat/physics telemetry
Correlates every HB-ON/HB-OFF heartbeat-tick CSV row (results-20260912-
with-heartbeat-csv/*-doe-raw.log) back to which trial (run_id, id_idx,
id_label, rep) was active when it printed, despite the async tick
printer splicing rows mid-token -- including mid a DOE-RUN marker itself
-- into the trial loop's own console output on the shared serial line.

Pipeline (analysis-20260912/, see its own README.md):
- correlate_doe.py: two-pass reconstruction per architecture (remove
  atomic CSV-row spans to rebuild the clean trial-output stream, map
  each removed row's offset back to the nearest preceding run_id
  marker); identity/rep looked up from a known-clean prior run's
  run_id mapping rather than re-parsed, since one aarch64 marker
  (trial 12, rajames rep 1) lost its id_idx digit to a zero-separator
  collision with an adjacent CSV field and is unrecoverable from that
  log alone -- its rows fold into trial 11 instead, documented as a
  known limitation.
- combine.py: merges all three architectures into combined.csv (768
  rows), decoding Q48.16 fields to floats and jitter_bits' IEEE754 bit
  pattern to real jitter_ns.
- analysis.R: per-cell (architecture x identity) means/SD and two-way
  ANOVA for each of 12 telemetry metrics, one boxplot SVG per metric,
  written up as ANALYSIS.md.

Key findings: identity significantly affects word-heat/window-sizing
metrics (expected -- different identities execute different word
sets), architecture significantly affects timing metrics (APIC
ticks/tick, timing variance, fleet heat -- expected, different QEMU
targets), zero architecture x identity interaction on any metric.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 12:46:34 -04:00
Robert Allan JamesandClaude Sonnet 5 403a7639e1 Bracket std79 DoE trial loop with HB-ON/HB-OFF, capture real heartbeat CSVs
Every prior std79-doe.fth run's automatically-extracted CSV
(scripts/extract_doe.sh) was silently empty -- g_doe_log_enabled
(doe_log.c) defaults off, and nothing in the campaign ever called HB-ON.
EXEC-STD79-DOE now calls HB-ON right before its trial loop and HB-OFF
right after, so the per-heartbeat-tick physics/timing CSV (18 columns:
tick_number, hot_word_count, avg_word_heat_q48, apic_ticks, per-VM
heat_q48, etc.) finally covers the run's own window on all three
architectures.

Found and fixed two bugs along the way, one in a comment and one in the
ad hoc QEMU orchestration script used to drive these runs (not part of
this repo):

- std79-doe.fth's own explanatory comment accidentally spelled out the
  literal "[HADES][DOE ]" tag string doe_log.c prefixes each row with --
  the FORTH REPL's compile-time echo of that comment then matched
  scripts/extract_doe.sh's own extraction grep, corrupting the first
  extracted CSV row with comment text instead of real telemetry. Fixed by
  never spelling out the literal substring.

- The harness script's completion-detection watched for "STD79-DOE:
  complete" anywhere in the log since before the whole capsule was fed,
  which matches the colon definition's own compile-time echo of that same
  string literal, not just the real end-of-run print. Without HB-ON the
  entire 27-trial run finished in a few seconds -- faster than one poll
  interval -- so the false match and the real one always landed in the
  same window and this never surfaced. HB-ON's added per-tick console I/O
  slowed real execution enough to expose it: the script sent BYE the
  moment compilation finished, truncating every trial after whatever
  point compilation had reached (aarch64 lost 9 of 27 trials this way on
  the first attempt). Fixed by feeding definitions and invocation as two
  genuinely separate connections, with the completion-watch window opened
  only after compilation is confirmed landed.

Verified 27/27 trials correct on every architecture via the campcampaign's
most distinctive result markers (both M*/M/MOD 18-19 digit values and the
2147483648 2/ result, all exactly 27 occurrences, zero faults) rather than
exact substring reconstruction: HB-ON's async per-tick CSV printer and the
trial loop's own console output share the same serial line with no
locking, so CSV rows can splice mid-token into trial output on the wire
(confirmed live -- cosmetic only, the underlying FORTH execution and
values are unaffected). results-20260912-with-heartbeat-csv/ holds both
the raw logs and their paired heartbeat CSVs; the earlier truncated runs'
logs are kept too (never delete logs) as the record of how this was found.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 10:44:00 -04:00
Robert Allan JamesandClaude Sonnet 5 8394e375d1 Reconfirm std79 DoE clean, 81/81, no code changes since §XVII (e51a8d2)
Plain reconfirmation rerun the day after both Stadium fixes landed
(O(ncells) scan fix + donor-floor fix). Same seed, single continuous boot
per architecture, all 9 identities simultaneously live throughout.

81/81 trials correct, 0 mismatches, shuffle sequence md5-identical to
every prior run since e2abc56. Total wall-clock: amd64 ~151s, aarch64
~276s, riscv64 ~165s — consistent with the last verified run, no
regression.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 07:14:45 -04:00
Robert Allan JamesandClaude Sonnet 5 e51a8d229e Fix stadium_grant_quota() donor floor; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVII)
capsule_birth.c hardcoded every new VM's initial Stadium quota grant to split
from Hera specifically. Since a grant always halves whatever the donor
currently has, Hera's own free list converges toward empty after a bounded
number of grants — independent of whether the Stadium as a whole still had
spare capacity, since VMs she'd granted to earlier typically still held
nearly all of their own share untouched. Past that point every subsequent
VM birth's Stadium grant would be silently refused (soft-failed, non-fatal
by existing design), even with plenty of capacity sitting idle elsewhere.

Fixed by adding an O(1)-maintained free_count to StadiumVMQuota (incremented
in stadium_evict(), decremented at both of stadium_admit()'s free-list-pop
sites, set/adjusted in stadium_grant_quota()'s own split — this also let
grant_quota drop its old O(free-list length) counting walk in favor of an
O(1) read) and stadium_best_donor(), an O(live VM count) scan over quota
slots returning whichever in-use VM currently has the most free cells.
capsule_birth.c's birth path now splits from that VM instead of
unconditionally vm_uuid_hera().

Verified with another full rerun of the 3x9x3 std79 DoE campaign from
scratch — same discipline as the prior Stadium fix (any defect repair
reruns the whole DoE from the top) — one continuous boot per architecture,
all 9 identities simultaneously live throughout. 81/81 trials correct, 0
mismatches, DOE-RUN header sequence md5-identical to every prior run.
aarch64 ~280s total (vs ~290s for the O(ncells)-scan fix alone — confirms
no regression). Both known Stadium defects are now closed together on one
clean campaign rerun.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 17:20:07 -04:00
Robert Allan JamesandClaude Sonnet 5 9eff122090 Fix aarch64 Stadium/COOL O(ncells) scan; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVI)
Root cause of the 90+ minute aarch64 VM-birth stall found in §XV: stadium_admit()'s
eviction-fallback scan iterated the entire stadium_ncells array filtered by owner,
not the calling VM's own resident cells as its own doc comment claimed. Combined
with stadium_grant_quota() always splitting from Hera's shrinking free list and
stadium_word_dispatch() calling stadium_admit() per distinct word a VM's capsule
executes, this compounded into a real O(n) blowup — catastrophic specifically on
aarch64 because its -m 4096 (vs 1024 on amd64/riscv64) inflates the kmalloc heap
kmalloc_init() bisects down to, which inflates stadium_ncells 4x (335,544 vs 83,886
cells, measured from boot logs).

Fixed by threading a real per-VM doubly-linked resident-cell list
(StadiumVMQuota.resident_head + stadium_resident_next[]/stadium_resident_prev[])
so the fallback scan is bounded by that VM's own resident count, not the global
cell array size.

Verified with a full rerun of the 3x9x3 std79 DoE campaign from scratch: one
continuous boot per architecture, all 9 identities simultaneously live throughout
(the 3-boot aarch64/riscv64 batching workaround is no longer needed). 81/81 trials
correct, 0 mismatches, DOE-RUN header sequence md5-identical across all three raw
logs. Identity 04's attach on aarch64, which stalled 90+ minutes before, now
completes in ~34s; full boot-to-DoE-complete in ~290s.

Corrects an earlier misreading (carried into §XV, std79-doe.fth's comments, and
the project memory note) that described the symptom as a runaway "335,000+ cycles"
dispatch counter — those were cell array indices, not an event count.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 16:37:44 -04:00
Robert Allan JamesandClaude Sonnet 5 e2abc56306 3(arch) x 9(identity) x 3(rep) randomized full-factorial std79 DoE: 81/81 correct (FABRIC-3.md §XV)
Formal successor to §XII's ad hoc exerciser campaign, requested as a
genuine randomized full-factorial design matching this project's own
DoE methodology (capsules/doe.4th's Fisher-Yates shuffle), and written
entirely in FORTH per explicit request -- not host-orchestrated shell
scripting.

experiments/std79-doe/std79-doe.fth: builds a 27-cell (9 identity x 3
replicate) run matrix, Fisher-Yates shuffles it with a fixed seed
(matching doe.4th's own default), then dispatches each of the 24
exerciser test cases directly into the target identity's own live VM
via VM-EXEC -- no console USE redirection, no per-trial host
interaction. The zuse case runs as directly-compiled native code
(RUN-TEST-NATIVE) rather than VM-EXEC targeting "Hera" herself:
VM-EXEC's own vm_state_push/pop only saves rsp/exit_colon/ecw_nesting,
not input_buffer/input_pos, so a self-targeting call while this
capsule's own vm_interpret call is still mid-line would risk exactly
the class of bug the idle-tick reentrancy guards exist for.

amd64: all 27 trials ran with all 9 identities simultaneously live in
one boot -- clean, zero mismatches.

aarch64: hit a real, uninvestigated bug attaching all 9 simultaneously
-- the 6th live VM's birth stalled for 90+ minutes at 100%+ CPU with a
Stadium COOL-dispatch counter already at 335,000+ cycles, versus tens
of thousands at the same checkpoint for earlier identities. Not
root-caused here (flagged in FABRIC-3.md for later); worked around by
splitting into 3 boots of 3 simultaneously-live identities each, every
boot sharing the same seed so the master 27-slot shuffle is identical,
filtered per boot by a new ACTIVE-LO/ACTIVE-HI range
(EXEC-STD79-DOE's signature: seed lo hi -- ). run_id is always the
slot's true position in the master shuffle, so trial order stays
comparable across boots -- standard DoE blocking.

riscv64: same 3-boot pattern, clean.

Grand total: 81/81 trials correct, 0 mismatches, across all three
architectures, all nine identities, all three replicates.

Also flagged (FABRIC-3.md §XIV, not fixed): attaching several WIREBIND
identities near-simultaneously (whether via rapid hotplug or all
present from boot) causes the kernel to silently detect only some of
them -- confirmed at the host/QMP level that every device was genuinely
present. Worked around throughout this campaign by attaching one
identity at a time with confirmed waits; real hardware hotplug could
hit the same gap, so it's a genuine robustness concern, not just a
test-harness inconvenience.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 13:59:32 -04:00
Robert Allan JamesandClaude Sonnet 5 70db955ac9 Fix M/MOD: hand-rolled 128/64 bit-serial division (FABRIC-3.md §XII.4)
mixed_math_word_m_slash_mod() had the identical bug class already fixed
in M* (commit 9a09949): it reconstructed the dividend from
`(dhigh << 32) | (dlow & 0xFFFFFFFF)`, the same wrong "32-bit halves of a
64-bit value" assumption. Latent for small inputs (fits in 32 bits, so the
reconstruction coincidentally worked), confirmed genuinely broken for a
true wide double -- feeding M*'s own correct 10^24 output into it gave a
quotient/remainder wrong by many orders of magnitude and the wrong sign.

Unlike M*, this one can't just switch to __int128 -- __int128's own `/`/`%`
need libgcc's __udivti3/__umodti3 for 128-bit division, unavailable in
this freestanding, -nostdlib build (the exact constraint
src/starkernel/arch/amd64/timer.c's own doc comment already flagged:
__int128 multiply/shift-by-constant/compare/subtract all compile clean,
only division doesn't). Fixed instead with a hand-rolled unsigned
128-by-64-bit bit-serial (restoring) long division -- 128 iterations of
shift-by-1/compare/subtract only, all in the safe set. Operates on
magnitudes via unsigned negation from 0 (well-defined even for the
extreme negative edge); signs reapplied afterward matching the same C99
truncating-toward-zero convention the previous, narrower implementation
already used, unchanged.

Verified: rebuilt all three architectures, confirmed clean link with no
__udivti3/__umodti3 undefined-symbol errors. Booted and tested all three,
identical results: 1000000000000 1000000000000 M* SWAP 1000000 M/MOD ->
1000000000000000000 remainder 0 (exact division, the case that was wrong
by orders of magnitude before); three sign-combination cases all correct
(-+, +-, --), confirming sign handling survived the magnitude-only
rewrite. T15 (original small-input case) unaffected. Full 24-case
exerciser reran clean on all three, no regressions.

All bugs found by the std79 exerciser campaign, including this one found
while fixing another, are now closed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 10:28:34 -04:00
Robert Allan JamesandClaude Sonnet 5 9a09949c69 Fix M*: use __int128 for a genuine 128-bit double-cell product (FABRIC-3.md §XII.4)
mixed_math_word_m_star() confused "double" (two full cell_t-width cells,
128 bits total on this 64-bit build -- what D+/D-/D./etc. all actually
expect) with "the low/high 32-bit halves of a single 64-bit product" --
code clearly written assuming cell_t is 32-bit. It computed an ordinary
64-bit `long long` product (already wrong for any true product exceeding
64 bits, since long long is the same width as one cell here) and split
that into 32-bit halves via `result & 0xFFFFFFFF` / `result >> 32`. For a
small negative product like -56088, this produced a positive, zero-
extended low cell paired with a correctly-looking dhigh=-1 -- D.'s
overflow check (correctly) rejected the resulting malformed double, on
every architecture, every time (this bug was never architecture-specific,
unlike the D+/D-/DNEGATE/d_compare family already fixed in bea8d74/
1a716c8).

Fixed via __int128 for a genuine 64x64->128-bit multiply, no truncation --
an already-established safe pattern in this codebase for exactly this
operation (src/starkernel/crypto/fe25519.c/scalar25519.c already use it;
src/starkernel/arch/amd64/timer.c's own doc comment confirms __int128
multiply/shift-by-constant compile cleanly with zero undefined symbols on
all three target toolchains -- only division needs unavailable libgcc
support, not used here).

Verified: rebuilt and booted all three architectures clean. T13/T14 both
correct everywhere (56088/-56088). Manual cases confirm genuine 128-bit
precision: -1 -1 M* -> 1; 1000000000000 1000000000000 M* -> dhigh=54210
dlow=2003764205206896640 (10^12 x 10^12 = 10^24, correctly exceeding 64
bits). Full 24-case exerciser reran clean on all three, no regressions.

Found while verifying, NOT fixed (report only, out of scope of this
request): mixed_math_word_m_slash_mod() (M/MOD, the very next function in
the same file) has the identical bug class -- confirmed genuinely broken
for a true wide double (feeding this fix's own correct large-magnitude M*
output into M/MOD produces a result wrong by many orders of magnitude and
the wrong sign). Flagged in FABRIC-3.md for a future fix request.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 08:23:51 -04:00
Robert Allan JamesandClaude Sonnet 5 1a716c8048 Fix d_compare: use ucell_t instead of unsigned long (FABRIC-3.md §XII.4)
d_compare() (double_words.c), the static helper backing DMAX/DMIN/D</D=,
had the same bare-unsigned-long pattern already fixed in D+/D-/DNEGATE
(commit bea8d74) -- 32-bit unsigned long on this aarch64 target silently
truncating the low-cell comparison. Switched to ucell_t.

Verified with four cases designed specifically to expose the old
truncation: two doubles sharing the same high cell but with low cells
differing only in bits 32-63 (invisible to a 32-bit-truncated compare,
e.g. 2^32 vs 0):
  4294967296 0 0 0 D= .            -> 0  (correctly not equal)
  4294967296 0 0 0 DMIN SWAP D.    -> 0  (correctly picks the smaller)
  0 0 4294967296 0 D< .            -> -1 (correct)
  4294967296 0 0 0 D< .            -> 0  (correct)
All four pass identically on amd64, aarch64, and riscv64. These cases
were never run before this fix -- the exerciser campaign never touched
DMAX/DMIN/D</D= at all, so this is the first real evidence d_compare()
was ever exercised on any architecture. Full 24-case exerciser also
reran clean on all three architectures (no regression).

D2*/D2/ use unsigned long long (C-standard-guaranteed >=64-bit) and were
never part of this bug class. All four affected words (D+, D-, DNEGATE,
d_compare) are now fixed; M*'s separate universal DOUBLE-OVERFLOW bug
remains open (unrelated defect, out of scope here).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 07:44:14 -04:00
Robert Allan JamesandClaude Sonnet 5 bea8d7436a Fix D+/D-/DNEGATE: use ucell_t instead of unsigned long (FABRIC-3.md §XII.4)
double_word_d_plus(), double_word_d_minus(), and double_word_dnegate()
(double_words.c) all cast through plain `unsigned long` for their carry/
borrow-detection arithmetic. On this aarch64 bare-metal cross-compile
target, unsigned long is 32-bit (confirmed: sizeof(unsigned long)==4) --
amd64 and riscv64 both happen to have a 64-bit long, so the identical code
only broke on aarch64. The low-cell arithmetic silently truncated to 32
bits, then widened back to cell_t via ordinary (non-sign-extending)
conversion, producing a wrong result whenever the true 64-bit result was
negative -- D. then correctly, faithfully reported DOUBLE-OVERFLOW on the
resulting malformed double.

vm.h already defines ucell_t for exactly this: same conditional as cell_t,
guaranteed width-matched on every target. print_number_formatted()
(format_words.c) already used it correctly; these three words didn't.
Switched all three to ucell_t -- a one-word-class fix, no logic change.

Verified: rebuilt and booted all three architectures clean. T19 (D+) on
aarch64 now correctly prints -2, matching amd64/riscv64; T20 (DNEGATE)
unaffected everywhere. Additional manual cases beyond the original
exerciser, run live on aarch64 to specifically exercise the
>32-bit-magnitude path the old bug depended on: D- (-5-3=-8), DNEGATE on
2^33 (8589934592 -> -8589934592), D+ crossing the same boundary
(3+8589934592=8589934595) -- all correct.

d_compare() (backing DMAX/DMIN/D</D=) has the identical latent pattern but
is out of scope for this fix (not named in the request, never exercised by
the campaign) -- left open, flagged in FABRIC-3.md. M*'s separate,
universal-across-all-three-architectures DOUBLE-OVERFLOW bug is also
untouched -- unrelated defect, not part of this fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 07:37:57 -04:00
Robert Allan JamesandClaude Sonnet 5 91f7b39d3c Root-cause the aarch64-only D+ DOUBLE-OVERFLOW bug (FABRIC-3.md §XII.4)
Confirmed the exact mechanism behind the aarch64-specific D+/D. divergence
found in the std79 exerciser campaign (commit cc81edf). The earlier
writeup's cell_t-width check was real but answered the wrong question --
cell_t is 64-bit everywhere, but double_word_d_plus()'s carry-detection
casts through plain `unsigned long`, and sizeof(unsigned long) is 4 (32-bit)
on this aarch64 bare-metal cross-compile target specifically (amd64 and
riscv64 both happen to have a 64-bit long). The low-cell addition silently
truncates to 32 bits, then widens back to cell_t via ordinary (non-sign-
extending) conversion, producing a wrong positive result_low whenever the
true 64-bit sum is negative -- confirmed live by splitting result_low into
hi/lo 32-bit halves: the real pushed value for `-5 S>D 3 S>D D+` is
+4294967294, not -2. D.'s overflow check is correct and is faithfully
reporting a genuinely malformed double; D+ is the actual defect.

vm.h already defines ucell_t for exactly this class of problem (same
conditional as cell_t, guaranteed width-matched on every target) --
print_number_formatted() (format_words.c) already uses it correctly, D+
doesn't. Also flagged (not exercised by the campaign, same bare-`unsigned
long` pattern, same latent risk): D-, DNEGATE, and the d_compare() helper
behind DMAX/DMIN/D</D=.

Report only, per this project's standing rule (report bugs, don't fix
without being asked) -- no source change in this commit, debug probes
added during the investigation were reverted after use.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 06:41:39 -04:00
Robert Allan JamesandClaude Sonnet 5 cc81edf00a std79 exerciser campaign: 27/27 legs clean, two real D. bugs found (FABRIC-3.md §XII.4)
Completed the FORTH-79 standard-dictionary cross-ISA exerciser campaign: all
9 identities (zuse, rajames/bob, 00-06) x all 3 architectures (amd64,
aarch64, riscv64), 27 legs total. No crashes, no heap corruption across the
full run -- real-world validation that the WIREBIND use-after-free fix
(commit 9142dda, FABRIC-3.md SXIII) holds under genuine multi-cycle load,
not just the synthetic repro used to verify it.

Cross-identity parity is perfect: 0 diffs across all 9 identities on each
architecture. Two real bugs found in double-precision (D.) output, reported
per this project's standing rule (report, don't fix without being asked):

- M* on a negative operand -> D. reports DOUBLE-OVERFLOW, universally
  across all three architectures (engine-level bug, not arch-specific).
- D+ on two negative doubles -> D. reports DOUBLE-OVERFLOW on aarch64
  only; amd64 and riscv64 both correctly print -2. cell_t width confirmed
  64-bit on all three (ruled out as the cause); exact mechanism still open.

Also fixed a test-harness timing issue (run_identities.sh): the retry
budget for USE-after-attach was too tight for boots with several live VMs
already accumulated, causing false "FAILED to USE" verdicts on identities
that actually succeeded a few seconds later. Widened the budget and
switched to smaller per-boot batches (2-3 identities) as the reliable
pattern for this shape of campaign.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-10 23:13:13 -04:00
Robert Allan JamesandClaude Sonnet 5 9142dda2d6 Fix use-after-free in sk_repl_idle()'s idle-tick VM resolution (FABRIC-3.md §XIII)
Root-caused a heap-corruption bug that reliably failed WIREBIND identity
attach on the third attach/detach cycle in one boot. sk_console_readline()
and sk_console_getkey() captured `active_vm` once from their caller and
kept passing that same (possibly long-stale) pointer to sk_repl_idle() on
every idle tick serviced while blocked waiting for input. If the VM it
pointed at was killed (WIREBIND detach) mid-block, the existing bailout
only checked a generic "is anyone attached" boolean -- masked as soon as a
different identity attached next -- so blk_vm_flush_all() kept writing
into a freed VM struct sitting on kmalloc's own free list, corrupting the
free list's linked-list metadata itself.

Both idle branches now re-resolve the live active VM fresh from
g_repl_active_vm on every tick, matching the dispatch-side fix already
made for the sibling bug in §XII.3.

Verified: rebuilt amd64, reran the exact three-cycle repro that reliably
corrupted the heap before the fix -- free-list census stayed stable
through the same idle window that previously collapsed to zero. All three
architectures (amd64/aarch64/riscv64) boot clean to the zuse)ok> prompt.

kmalloc_debug_census()/kmalloc_debug_census_bytes() kept as permanent
diagnostic infrastructure; every other temporary probe added during the
investigation was reverted.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-10 20:30:41 -04:00
Robert Allan JamesandClaude Sonnet 5 70421bdd43 Fix real WIREBIND crash: stale active-VM pointer dispatched after blocking read (FABRIC-3.md §XII.3)
The interpreter_enabled guard added in the previous commit (662ef44) was a
real but incomplete fix -- re-running the exact repro against it still
panicked (this time as a raw #PF page fault), proving something deeper
was wrong.

Root cause, found via targeted console_puts probes (not GDB --
starkernel_kernel.elf's symbols don't correspond to the actual running
starkernel_loader.efi binary for this monolithic build, same gotcha
already on record from the 2026-08-18 aarch64 investigation):

sk_repl_run()'s main loop captures `active` once, before calling
sk_console_readline(), which then blocks for the next full line. If the
identity `active` points at is killed while that read is still blocked,
the bailout meant to catch this (sk_console_identity_present()) only
checks a generic "is anyone attached" boolean, not "is the specific
identity active belonged to still attached" -- a fast detach of one
identity followed by attach of a different one never produces an
observable gap in that boolean, so the bailout never fires. The stale
`active`, now pointing at freed memory, gets dispatched into.

Fix: re-resolve `active` fresh from g_repl_active_vm immediately before
dispatch, right after sk_console_readline() returns. One line, no
registry lookup, no dereference of the stale pointer -- closes the race
regardless of whether the bailout catches it first.

Verified: rebuilt amd64 clean, reproduced the exact same attach/USE/
detach/attach/USE sequence against the fixed build -- clean switch, no
fault, exerciser runs correctly afterward.

FABRIC-3.md §XII.2 also corrected to stop claiming the interpreter_
enabled guard alone closed the crash -- it didn't, per the above.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-10 17:19:01 -04:00
Robert Allan JamesandClaude Sonnet 5 662ef44e59 Fix EXEC/LOAD block-persistence gap and WIREBIND/USE interpreter-race panic (FABRIC-3.md §XII)
Found live while building a cross-ISA FORTH-79 dictionary exerciser:

- capsule_exec_init() zeroed a capsule's block content immediately after
  running it, so LOAD (a genuine FORTH-79 standard word, ACL-allowed even
  for locked identities) could never actually read back what EXEC had
  just written. Removed the clear from capsule_exec_init(); block content
  now persists like any other Standard BLOCK/BUFFER/UPDATE write.
  kernel_main.c's own explicit post-birth clear of Mama's init.4th range
  is untouched.

- USE could redirect the console to a WIREBIND identity's VM before that
  VM's vm_enable_interpreter() step of its own birth sequence had run,
  causing the next typed line to hit vm_assert_interpreter_enabled() and
  panic the entire machine -- not the per-session-recoverable ACL-fault
  path a redirected VM otherwise gets. USE now checks interpreter_enabled
  first and refuses with a retry message instead.

Also includes the amd64/aarch64/riscv64 acceptance boot logs and DoE CSVs
from this session.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-10 16:07:38 -04:00
Robert Allan JamesandClaude Sonnet 5 d5722986b2 Re-mint bob with the restricted personality, closing out the identity set (FABRIC-3.md §XI.6)
bob-thumb-ident.img predated MINT_PERSONALITY_STD79_LOCKDOWN (minted
2026-09-06, f4ded3e landed 2026-09-07) -- the one identity §XI.4's
00-06 conversion missed. Zeroed its homeblocks_sig_t signature block
and re-minted with the same real identity data it originally had
(full_name="Captain Bob", username="rajames",
email="rajames440@gmail.com"), this time with the restricted
personality flag.

Verified standalone (no Zuse, fresh boot): fast attach as "rajames",
USE, standard words computing correctly (11, 49), and VM-EXEC
(ACL-denied) recovering gracefully without halting -- exercising both
the §XI.4 fault-scoping fix and the §XI.5 xHCI/WIREBIND fixes live.

All 9 thumbdrives (zuse, bob/rajames, 00-06) now carry real,
MINT-verified identities; all 8 non-Zuse ones carry the restricted
FORTH-79/83 personality.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-10 11:16:02 -04:00
Robert Allan JamesandClaude Sonnet 5 b301317902 xHCI: drive Port Reset on port reuse; WIREBIND: kill the console VM too (FABRIC-3.md §XI.5)
Two independent bugs that together caused a reliable hotplug wedge:
reusing an xHCI port for a second identity right after an unclean
detach of a first would leave no further hotplug events reaching the
guest at all.

Bug 1 (xhci.c/xhci_driver.h): the xHCI driver never drove PORTSC.PR --
a known, named gap since Milestone 2e (the code's own comment flagged
it, PORTSC_PR/PRC were defined but never referenced). A port's first
connect each boot reads PED already set, so skipping the reset
happened to work; a second device on the same port after a prior
disconnect reads PED clear, and Address Device reliably failed without
an explicit reset cycle. New XHCI_CONN_AWAIT_PORT_RESET state drives
PR and waits for PED to read set before proceeding to Enable Slot.

Bug 2 (capsule_wirebind.c): capsule_wirebind_eject()/unclean_detach()
compared g_repl_active_vm against the *user* VM's pointer
(g_wirebind_attached_vm_id tracks that one, not the console VM) --
never equal, since USE/g_repl_active_vm always points at the console
VM. The guard never fired and the console VM was never killed at all,
only orphaned -- paired to a dead user VM but still the REPL's active
session. New wirebind_teardown_console() helper resolves and tears
down the console VM by its own tracked bare username.

Verified live on amd64: the exact repro (identity 01 on port 2,
unclean detach, identity 02 on the same port immediately after) --
previously wedged with "xhci: address device failed" and no further
hotplug activity; now attaches cleanly and fast, both VMs' KILL
messages appear, console is immediately interactive on the new
identity. Three-arch clean qemu acceptance passed, all clean on the
first attempt.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 22:11:09 -04:00
Robert Allan JamesandClaude Sonnet 5 9ea5580ace FABRIC-3.md §XI.4: record all 7 identities individually verified
00-06 each confirmed standalone (no Zuse): fast attach, USE, a
standard word computing correctly, and an ACL-denied VM-EXEC
recovering gracefully instead of halting -- exercising the fault-
scoping fix from the prior commit across every minted identity, not
just 00/01.

Also notes an out-of-scope hotplug finding: reusing an xHCI port for a
second identity right after an unclean (non-EJECT) detach of a first
wedges that port for further attaches. Worked around (fresh port/boot
per identity) rather than root-caused -- not part of this session's
task.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 20:55:21 -04:00
Robert Allan JamesandClaude Sonnet 5 d6661b5eed Scope VM fault halt to the faulting identity's own session (FABRIC-3.md §XI.4)
A standalone WIREBIND identity (no Zuse, USE'd in directly) hitting an
ACL-denied word halted the entire machine -- Hera, Hermes, Artemis, all
of it -- instead of just that identity's own session. sk_fault_handler()
was being called unconditionally on whichever VM's ->error was set, with
no distinction between Hera's own root session (where "no fallthrough
surface" is the correct, deliberate fail-closed behavior) and a
USE'd-in guest identity (which should recover and resume at its own
prompt instead of taking the fleet down with it).

Both call sites (sk_repl_step, sk_repl_run) now compare the faulting VM
against Hera before deciding: Hera's own session still halts by design;
any other VM prints a recovery message, clears its fault state, and
continues.

Also: mint identities 01-06 with the same FORTH-79/83 restricted
personality identity 00 already had, verified via the fixed fault
scoping above (which this verification pass surfaced).

Verified live on amd64 (both the Hera-halts and identity-recovers
branches); three-arch clean qemu acceptance passed (riscv64's first
attempt hit an unrelated virtio_blk I/O timeout hang, a known QEMU/TCG
flake -- a clean retry booted normally).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 20:41:51 -04:00
Robert Allan JamesandClaude Sonnet 5 9bcc70647b FABRIC-3.md §XI: document the ACL lockdown, messaging migration, and the
two-bug "identity attach without Zuse" investigation

Covers three entangled threads from 2026-09-07/09: the FORTH-79/83 MINT
lockdown personality (commit f4ded3e), the Hera->Artemis storage-attach
message round-trip migration including the ACK-APPEND-NUM stack bug and
the MSG-TICK/ACL collision (commit 63b8b3b), and the extended
investigation into why an identity attaching without Zuse ever attaching
first either took many real minutes or never completed at all -- which
turned out to be two independent, unrelated bugs (a pathological
full-device migration scan self-inflicted earlier the same session,
commit 1a26355; and a trust-chain gate that conflated minting-needs-her-
live-private-seed with verification-needs-only-her-already-public-key,
commit 1839a2b), each one masking clean evidence of the other. Includes
an explicit framing correction: "Zuse must attach first" was never the
actual load-bearing variable, since Artemis's own disk is always
attached first regardless -- worth naming plainly as a trap in this
style of live-forensics-driven investigation, not edited out of the
record.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 16:19:18 -04:00
Robert Allan JamesandClaude Sonnet 5 1839a2b0c3 WIREBIND cert verification: load Zuse's root pubkey independently of her live session
Root cause of the remaining "identity attach doesn't complete when Zuse
never attaches this boot" issue: capsule_wirebind_verify_cert() gated on
mama_vm->zuse_cert_installed, which is only ever set when Zuse's own
drive attaches and authenticates this specific boot
(capsule_zuse_boot_try_attach() -> install_and_activate() ->
vm_zuse_cert_install()). Without her, any other identity's WIREBIND cert
verification silently refused -- correctly, by the old design, but that
design conflated two genuinely different things: "can mint new
identities" (needs Zuse's live private seed, a real privileged
operation) and "can verify an existing identity's cert" (needs nothing
but her already-public key).

That public key was already being persisted independently of her live
session: zuse_genesis_marker_t (zuse_genesis_marker.h) stores it in the
kernel's own top-of-device metadata fence (Artemis's resident storage),
written once at genesis, specifically *not* alongside her private seed
(which stays only on her own removable thumbdrive) -- the type's own doc
comment says as much. It just wasn't being loaded for anything but
confirming which drive is genuinely hers.

Fix: a new capsule_zuse_boot_load_root_pubkey() (capsule_zuse_boot.c)
reads that marker and populates two new VM fields, zuse_root_pubkey_known
/ zuse_root_pubkey (vm.h) -- deliberately separate from
zuse_cert_installed/zuse_cert_seed/zuse_cert_pubkey, which stay
untouched and still gate MINT exactly as before. Called once from
kernel_main.c as soon as Artemis's own storage attaches, unconditionally,
independent of whether Zuse's own drive is ever attached this boot.
capsule_wirebind_verify_cert()/capsule_wirebind_try_attach() now check
zuse_root_pubkey_known instead of zuse_cert_installed.

One identity's attach must not depend on another identity's live
presence -- each identity stands on its own once the fleet's root of
trust has been established once, ever.

Verified live, amd64: identity 00 (disk/thumbdrives/00-thumb-ident.img)
now attaches and completes WIREBIND in 19 seconds with Zuse's own drive
never attached this boot at all (previously: unbounded, many real
minutes or effectively never, before today's other fixes; still slow/
stuck after those, stuck specifically on this silent refusal). Zuse's
own attach flow re-verified unaffected (regression check, amd64).
Three-arch clean qemu acceptance (amd64/aarch64/riscv64) passed with
this change included.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 14:30:14 -04:00
Robert Allan JamesandClaude Sonnet 5 1a263555e2 Fix pathological migration scan that stalled WIREBIND identity attach
Root cause (found by a fresh subagent after an extended live-debugging
investigation into "identity 00 attaches slowly/stalls when Zuse never
attached first this boot"): blk_migration_idle_check() was generalized
earlier today to walk every attached device slot uniformly instead of
hardcoding first_disk_slot() (Artemis's own disk). But its per-slot scan
can only early-exit once it finds a devblock that is BOTH "hot" (claimed
and worn) AND "free" -- and a just-attached, never-claimed USB identity
drive can never satisfy the "hot" half by design (claiming only ever
happens via blk_firsttouch_claim(), which only ever targets
first_disk_slot()). So the scan ran to completion -- the drive's entire
~16,000 devblocks, mostly cache misses over slow emulated USB/BOT --
every single idle tick, forever, blocking sk_repl_idle() (and therefore
the console and the storage-attach message round-trip) each time.

Fix, in src/block_subsystem.c: a new has_ever_claimed flag on
blk_dev_slot_t (set in blk_set_meta(), the single choke point every
BLK_FLAG_CLAIMED transition passes through) skips the scan entirely,
O(1), for any slot nothing has ever claimed -- the common case for a
freshly-attached drive. A new migration_scan_lbn resume cursor bounds
*any* slot's per-tick cost to MIGRATION_SCAN_BUDGET (256) devblocks
examined, picking up where the previous tick left off instead of
restarting from start_lbn every time -- restores this function's own
documented "coarse cadence, cheap early-exit" design intent for every
device, not just the one it used to hardcode.

Also along the way (kept, all real improvements, verified live):
- src/starkernel/usb/xhci.c: xhci_wait_bit()/xhci_bot_wait_for_idle()
  had zero yield hints in their MMIO-polling loops; added arch_relax()
  to both (matches virtio_blk.c below) -- a tight loop of nothing but
  MMIO reads can starve TCG's own host-side timer injection under QEMU.
- src/starkernel/virtio/virtio_blk.c: vblk_io()'s spin bound was 33M
  iterations with zero logging on timeout; a single real (still not
  fully root-caused) timeout cost 31+ minutes of CPU before this was
  caught. Reduced to 1M and added a log line naming the failing sector,
  turning a silent, effectively-unbounded stall into a fast, loud
  failure -- callers already tolerate BLKIO_EIO.
- src/starkernel/repl.c: blk_migration_idle_check() deferred for any
  idle tick where a storage-attach message round-trip is still pending,
  to keep the two block-subsystem-touching paths from interleaving; the
  existing MSG-TICK pump now checks the target VM's own dictionary for
  MSG-TICK (vm_find_word + acl_allow) before VM-EXECing into it, instead
  of spamming "VM-EXEC: ERROR" every tick for a VM that doesn't have --
  or isn't allowed -- the word (e.g. a FORTH-79/83-locked-down identity);
  fixed a real -Wunused-variable build break in the EMERGENCY_CONSOLE_
  ENABLED=1 path (unused today, but needed live to reproduce this bug
  with no identity attached at all).
- src/starkernel/capsule/capsule_mint.c: dropped the dead
  S" common:messaging.4th" EXEC / MSG-CD-INIT lines from
  MINT_RESTRICTED_PERSONALITY -- a FORTH-79/83-locked-down identity has
  no legitimate use for a messaging vocabulary it can never call.

Status: the pathological CPU-climbing scan is confirmed fixed (verified
live: CPU stays flat across an extended run instead of climbing without
bound). The WIREBIND storage-attach message round-trip still does not
complete promptly in the "Zuse never attached, other identity attaches
first" scenario -- a separate, still-open issue in the message-delivery
path itself, not the scan. Tracked as follow-on work.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 12:08:03 -04:00
Robert Allan JamesandClaude Sonnet 5 63b8b3bc29 Migrate Hera->Artemis storage-attach to a real message round-trip
Hera still polls xHCI and sig-checks attached drives, but the storage
registration step (blk_subsys_attach_device(), now wrapped as the
BLK-ATTACH primitive) moves to Artemis's own dictionary, reached via
HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK (VM-EXEC, since Hera can't load her
own messaging.4th -- see the doc comment in repl.c). Identity birth
(Zuse genesis / WIREBIND) is deferred until the ack confirms storage
actually succeeded, instead of running synchronously underneath a
storage call that might fail ("wait for ack, safer for identity data").

Caught and fixed a real bug live during acceptance testing: Artemis's
ACK-APPEND-NUM fed a single-cell value into <# #S #> (which expects a
double-cell pair), causing a stack underflow the first time
HERA-BLK-ATTACH-REQ ran. Fixed with the same `0 SWAP` convention every
other numeric-append helper in this codebase already uses.

Verified booting clean to (zuse) ok> with no VM-EXEC errors on all
three architectures (amd64/aarch64/riscv64).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 21:22:04 -04:00
Robert Allan JamesandClaude Sonnet 5 c4f94c409f FABRIC-2.md: update stale MSGMIGRATE tracking node
Console→user-VM command relay is already real messaging (repl.c:971,
a genuine CONSOLE-CMD-EVENT/MSG-SEND), not hardwired -- this predates
today but the tracking graph still marked the whole node unstarted.
Updated to PARTIAL, with the real remaining gap (attach detection/
verify/birth is still 100% hardwired) and the documented fallback
(relay skips a line containing a double-quote character, repl.c:962-965)
both called out explicitly. Found via a research pass auditing
messaging.4th and the current hardwired-vs-messaged boundary, prompted
by Captain Bob's "get Hermes nailed down" -- Hermes/Artemis themselves
are already equally wired and session-less, not the actual gap.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 19:51:25 -04:00
Robert Allan JamesandClaude Sonnet 5 f4ded3e1a8 MINT: add a FORTH-79/83-standard-words-only lockdown personality
Captain Bob, 2026-09-07: "starting with that 00 user we created, we're
going to give access only to FORTH 79 and 83 standard words. everything
else is locked down."

New capsules/acl-std79.4th (blocks 4023-4047): walks a VM's own
dictionary (>LINK/LINK> traversal, same as ACL-INIT-PRIMITIVES/WORDS
already use) and permanently denies+pins every word not on an explicit
FORTH-79/83 allowlist, extracted from the real registered word set
(stack_words.c through control_words.c), not recited from memory.
Deliberately excludes, beyond plain non-standard words: BYE (100% ACL
bypass to the emergency console -- "needs more discussion, exclude for
now"), COLD/WARM/REBOOT/SAVE-SYSTEM (system lifecycle), the block/screen
editor L/S/SHOW/EDIT/UPDATE/SAVE-BUFFERS (lets a session rewrite
persistent block/capsule content, defeating the lockdown even though
nominally standard), BLK-ACL-*/BLK-OWNER@ (StarForth-specific), and
FORGET/FENCE (flagged as an unrestricted superpower word, 2026-09-03
audit). Keeps WORDS/VLIST/SEE (introspection only -- ACL is enforced
per-target-word at execution time regardless of how an XT was
obtained) and the parenthesized control-flow runtime primitives
((BRANCH) etc. -- IF/DO/LOOP compile calls to these; denying them
breaks ordinary control flow, not security).

MintPersonality enum (capsule_mint.h) lets capsule_mint_identity()
select which personality-source template gets written to a new
identity's devblock -- MINT_PERSONALITY_DEFAULT (unchanged) or
MINT_PERSONALITY_STD79_LOCKDOWN (EXECs acl-std79.4th then
ACL-LOCKDOWN-STD79 as the VM's own last bootstrap step). The actual
restriction logic stays entirely in FORTH per .claude/CLAUDE.md's
Word-Level ACL System rules ("ACL policy belongs in ACL.4th, never in
C") -- capsule_mint.c only picks which few-line bootstrap stub to
write. MINT's own stack signature gains a trailing restrict? flag;
capsule_zuse_boot.c's genesis mint (Zuse herself) explicitly passes
MINT_PERSONALITY_DEFAULT -- the superuser is never restricted.

Two real bugs found and fixed live during testing, both the same class
of self-referential fault: ACL-LOCKDOWN-STD79's own walk loop calls
ACL-STD79-ALLOWED?/ACL-STD79-LIST/ACL-ALLOW!/ACL-PIN on every single
iteration to do its job -- none of those are FORTH-79/83 standard
words, so the walk was denying its own load-bearing infrastructure
partway through and then faulting the next time it tried to call it
("VM fault -- emergency console disabled; halting", reproduced twice
live). Fixed by explicitly protecting all four in the allowlist
(block 4047) -- they must stay allowed for the walk to finish, not
because they belong on a "standard words" list.

Verified live end-to-end: minted a throwaway test identity with the
restrict? flag, confirmed her WIREBIND birth completes cleanly (no
faults, no shadow conflicts) on a single real attach, then USE'd into
her VM and confirmed standard arithmetic and user-defined words work
(1 2 + . -> 3; : X 5 5 * . ; X -> 25) while KILL is entirely unknown to
her dictionary and VM-EXEC is denied. One real, non-fatal side effect
found and left as-is (not asked to fix): the fleet's inter-VM messaging
pump (MSG-ARENA) is also denied by the lockdown, logging a harmless
per-idle-tick warning -- a fully locked-down VM doesn't participate in
message routing.

Not yet applied to the real identity 00 -- this commit is the
mechanism, verified against a disposable test identity only.

Three-arch clean qemu acceptance (single Zuse device, standard
regression case) passed on amd64, aarch64, and riscv64.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 15:41:56 -04:00
Robert Allan JamesandClaude Sonnet 5 8471d529bc FABRIC-3.md §X.4: correct a false claim about kmalloc.c lacking block splitting
The claim ("coalesces but doesn't split") was never actually checked
against kmalloc.c -- it was carried over from alloc_kernel.c's own doc
comment about itself and mis-applied to a different file. Asked to fix
it, re-reading kmalloc.c showed allocate_from_block() already has a
complete, unconditional splitting implementation. Nothing was broken;
correcting the record instead of "fixing" working code.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 14:34:03 -04:00
Robert Allan JamesandClaude Sonnet 5 b106a4b0b5 FABRIC-3.md §X.4: document the allocator resolution — sf_malloc/sf_free now on the real kernel heap
Updates §X.4 (and §X's own title) from "capacity question still OPEN"
to resolved: Captain Bob's framing (unknown VM count in advance, heap
should use whatever memory is actually available once this is a full
OS) led to routing alloc_kernel.c's sf_malloc()/sf_free() through the
kernel's existing kmalloc.c heap (2 GiB floor, PMM-backed, coalescing,
already boot-tested) instead of a separate 4MB arena -- commit 56e19a0.
Verified live: all 9 identities now attach where only 6 did before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 12:47:05 -04:00
Robert Allan JamesandClaude Sonnet 5 56e19a00f9 Route VM-dictionary sf_malloc/sf_free through the kernel's kmalloc heap
Captain Bob's call after seeing the identity-heap-capacity findings
(FABRIC-3.md §X.4): the number of concurrently-running VMs is not known
in advance, and once this is a complete operating system the heap should
be able to use whatever memory is actually available -- not a hardcoded
compile-time ceiling. This was already half-built and just not wired up.

src/starkernel/vm/alloc_kernel.c previously implemented sf_malloc()/
sf_free() (platform_alloc.h's allocator abstraction -- what
vm_create_word() calls for every VM's word dictionary) as its own
isolated static 4MB arena: first-fit free list, no splitting or
coalescing. That's exactly the allocator that topped out around 6
concurrent WIREBIND-born identities, failing from fragmentation before
true capacity exhaustion (§X.4's own measurements).

Sitting right next to it, unused for this purpose: src/starkernel/memory/
kmalloc.c, the kernel's general heap. Already initialized at boot (M6,
kernel_main.c, well before any VM is ever born), reserved from real
PMM-tracked physical memory rather than a fixed array, defaults to a
2 GiB floor explicitly sized "for 256+ baby VMs" per its own comment,
overridable via the --heap= boot flag, and its free list actually
coalesces neighboring blocks on every free.

Change: alloc_kernel.c's sf_malloc()/sf_free() now delegate to
kmalloc_aligned()/kfree() instead of managing a separate arena.
sf_alloc_init() becomes a no-op (kmalloc is already initialized by the
time any VM allocation can happen, and "resetting" a heap now shared by
every kernel subsystem would be actively wrong -- confirmed no external
caller depended on its old reset semantics). sf_alloc_get_stats() reads
kmalloc_get_stats() fresh rather than shadowing byte counts locally;
alloc_count/free_count (which kmalloc.c doesn't track) stay as simple
local counters. sf_calloc()/sf_realloc() are otherwise unchanged. Kernel-
only: the hosted (non-kernel) StarForth build keeps its own separate
alloc_host.c implementation, untouched.

Verified live: replaying the exact hotplug sequence that previously
topped out at 6 identities (Zuse + 8 identities, one at a time via QMP
device_add) now succeeds for all 9, where identity 05 specifically used
to fail. Three-arch clean qemu acceptance (single Zuse device, the
standard regression case) passed on amd64, aarch64, and riscv64 -- one
aarch64 attempt hit an unrelated, already-documented one-off QEMU hiccup
(empty log, boot never progressed past firmware) and passed cleanly on
retry with no rebuild.

Not addressed here: the underlying free-list itself is still first-fit
without splitting (only coalescing changed, inherited from kmalloc.c);
per-VM dictionary sizing (shrinking what each WIREBIND VM's word set
actually needs) is a separate, still-open lever from FABRIC-3.md §X.4's
open architecture question.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 12:46:14 -04:00
Robert Allan JamesandClaude Sonnet 5 36f1d6ae9e FABRIC-3.md §X: document the §IX.5 follow-on — xHCI root-cause/fix, live hotplug walkthrough, identity heap capacity findings
Documents three commits' worth of live-tested work in one continuous
arc, picking up exactly where §IX.5 left off:

- X.1: the "4th-device enumeration failure" was a QEMU test-harness
  port-topology artifact, not a driver bug (commit 30c26ad).
- X.2: the real driver bug it uncovered once ports were fixed --
  Configure Endpoint completions silently dropped under concurrent
  multi-device enumeration, root-caused via temporary reverted probes
  and fixed with the same single-in-flight discipline the rest of the
  driver already uses (commit e10fb76).
- X.3: a live 9-device hotplug walkthrough (Zuse, bob/rajames, 00-06,
  one at a time via QMP) verifying the dynamic attach path and
  per-device identity data integrity, distinct from the boot-time scan
  path X.1/X.2 exercised.
- X.4: exact, live-measured kernel-heap-arena numbers per identity VM
  (457,392 bytes, zero variance across 6 identities), a real leak found
  and fixed (orphaned console VM never torn down on user-VM birth
  failure, commit d8a195b), and the open architecture question this
  surfaces -- identity capacity is fragmentation-limited around 6
  concurrent identities today, not yet decided how to address.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 11:39:18 -04:00
Robert Allan JamesandClaude Sonnet 5 d8a195b8d8 WIREBIND: kill the orphaned console VM when the user VM birth fails
Found live during identity-heap-capacity testing (2026-09-07,
hotplugging Zuse + 8 identities one at a time and measuring the kernel
heap arena via a temporary allocator-stats probe, since reverted): every
WIREBIND identity attach births two VMs in sequence -- a "console" VM,
then the real "user" VM. When the second birth failed (arena
fragmentation under concurrent VM load, a separate, not-yet-fixed
capacity issue), capsule_wirebind_try_attach() logged the failure and
returned, but the console VM that had *already succeeded* was never
torn down. It stays live and registered under the identity's username,
consuming its own ~228KB of the fixed 4MB kernel heap arena forever --
nothing ever points a real user at it, since WIREBIND only ever hands
the caller the user VM's id.

This turns every failed identity attach into a permanent net loss of
heap rather than a neutral retry: confirmed live that a failed attach
left the arena 228,576 bytes worse off than before the attempt, and
every subsequent attempt starts from that worse baseline, compounding.

Fix: call capsule_vm_kill(username) on the now-orphaned console VM
before returning from the failure path -- the same teardown
capsule_wirebind_eject()/capsule_wirebind_unclean_detach() already use
elsewhere in this file (vm_cleanup() + sf_free(), confirmed live to
actually reclaim per-word dictionary allocations, FABRIC-3.md
§IX.2/§IX.3).

Verified live with the same allocator-stats probe (written, captured,
reverted -- not part of this commit): after the fix, a forced user-VM
birth failure now returns the arena to exactly its pre-attempt byte
count (3,526,256, matching the baseline precisely) instead of leaking
228,576 bytes. Three-arch clean qemu acceptance (single Zuse device,
the standard regression case) passed on amd64, aarch64, and riscv64.

The underlying capacity/fragmentation question (why the 7th concurrent
identity's arena allocation fails at all despite technically-sufficient
free bytes) is a separate, open architecture question -- not addressed
here. See project memory for the full measured numbers.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 11:37:06 -04:00
Robert Allan JamesandClaude Sonnet 5 e10fb76fb2 xhci: fix Configure Endpoint completion drop under concurrent multi-device enumeration
Root cause of the FABRIC-3.md §IX.5 follow-on: with 9 devices attached
concurrently at boot (Zuse + 8 identities), only 1 of 9 ever completed
enumeration and reached blkio_usb: MSC device ready -- the other 8 produced
no error and no success, just silence.

xhci_poll_events()'s deferred per-slot dispatch loop submitted a Configure
Endpoint command (a Command Ring op) unconditionally for every slot with
that action pending in a single pass -- unlike every other Command Ring op
in this driver (Enable Slot, Address Device, Disable Slot), which is
correctly gated behind dev->connect_state == XHCI_CONN_IDLE before ever
submitting. With 2+ devices enumerating concurrently, this let multiple
Configure Endpoint commands sit outstanding on the Command Ring at once.
Their completion is correlated purely via the single shared
dev->connect_state field (== XHCI_CONN_AWAIT_CONFIGURE_ENDPOINT), not the
completion event's own Slot ID -- so whichever slot's completion happened
to land while connect_state still read AWAIT_CONFIGURE_ENDPOINT got
correctly chained into SET_CONFIG, and every other slot's completion
arrived after connect_state had already moved on, silently swallowed by
the handler's generic "unrelated command completion" catch-all. No error
path exists for this, which is why it produced total silence rather than
a diagnosable failure.

Root-caused live via temporary WARN-level diagnostic probes (written,
captured, and fully reverted per the project's own probe convention --
this commit contains only the functional fix and its explanatory comment,
no probe code) added at four points: the initial port scan, the connect
handler, the Command Completion Event handler, and the deferred dispatch
loop itself. The probes showed all 9 devices correctly completing Enable
Slot + Address Device (ruling out the connect-state queue as the cause,
the original hypothesis), then all 9 correctly submitting Configure
Endpoint and all 9 commands completing successfully in hardware (code=
SUCCESS, no errors logged) -- but only 1 of 9 ever got its next_action
chained to SET_CONFIG.

Fix: apply the same single-in-flight discipline this driver already uses
for every other Command Ring op. If the Ring isn't free when a slot's
Configure Endpoint action is due, put the action back on that slot instead
of submitting a second command onto a busy Ring -- the next tick's
dispatch pass retries it once the Ring frees up.

Verified live: booting Zuse + all 8 identity drives concurrently (9
devices, one per real xHCI port via the XHCI_PORTS fix from the previous
commit) now produces 9 "MSC device ready" lines and zero xHCI errors,
where it previously produced exactly 1. Three-arch clean qemu acceptance
(single Zuse device, the standard regression case) passed on amd64,
aarch64, and riscv64 -- no change in that baseline behavior.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 09:02:45 -04:00
Robert Allan JamesandClaude Sonnet 5 30c26ade3d Give each xHCI usb-storage device its own port; fix stale ZUSEDISK default
FABRIC-3.md §IX.5's "4th-device enumeration failure" was never a driver
bug: QEMU's default qemu-xhci controller (p2=4,p3=4) exposes only 4 real
dual-role ports, not 8 as the parameter names suggest. Attaching more
devices than that on bus=xhci0.0 without an explicit port= makes QEMU
silently auto-insert a USB2 hub past the 4th slot; the xHCI/BOT driver
correctly reports that hub as "not a Mass Storage/SCSI/BOT device" because
it genuinely isn't one, and every drive behind it is unreachable (no hub
descent in this driver). Confirmed live via QEMU's own `info usb` before
touching any kernel code.

Fix is entirely in the QEMU test harness, not the kernel:
- New XHCI_PORTS Make variable (default 16, overridable) sizes p2/p3 on
  all three arches' qemu-xhci controller with real headroom above the
  current 9-device identity roster, per Bob's standing ruling against
  hardcoding a bound to today's scale (FABRIC-3.md §VII.4).
- ZUSEDISK_QEMU_ARGS now gives Zuse's drive an explicit port=1.
- QEMU_EXTRA's own doc comment shows the port= pattern for additional
  devices.

Also fixed in passing: ZUSEDISK's default path (disk/zuse.img) was stale
-- that file was deleted from git at c3db963, superseded by
disk/thumbdrives/zuse-thumb-ident.img, but the Makefile default was never
updated, so a plain `make qemu` silently failed to attach Zuse at all.
Now defaults to the real minted image.

Not addressed here, flagged for later: scripts/bleach_zuse_img.sh and
disk/README.md still reference the deleted disk/zuse.img path. The
9-device concurrent enumeration itself (the scenario that originally
surfaced §IX.5) remains unverified against the real kernel -- this change
only confirms single-device boots are unaffected on all three arches.

Three-arch clean qemu acceptance passed (single Zuse device, port 1):
amd64, aarch64, riscv64 all reached POST: PASSED / PARITY:OK / "Zuse:
identity confirmed from attached thumbdrive" / (zuse) ok>.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-06 20:24:05 -04:00
Robert Allan JamesandClaude Sonnet 5 2c1b3cd695 Four bugs found live verifying the 8 identity thumbdrives (FABRIC-3.md §IX)
All found by actually running the identity workflow §VII/§VIII made
possible, not by code review:

1. Zuse/WIREBIND cross-contamination on detach: capsule_zuse_boot_logout()
   and capsule_wirebind_unclean_detach() both had no device parameter, so
   an unrelated device detaching (while the real owner's own stayed
   attached) incorrectly tore down the wrong session. Both now compare
   the departing device against their own tracked one, mirroring
   capsule_wirebind.c's pre-existing g_wirebind_attached_dev precedent.

2. Dictionary-entry memory leak: vm_create_word()'s sf_malloc()'d
   DictEntry (plus a second per-entry allocation for transition_metrics)
   was never freed by vm_cleanup(), in both the hosted and kernel
   implementations. Caused a real kernel PANIC after 8-9 repeated VM
   birth/kill cycles in one boot. Fixed by walking vm->latest in both.

3. sf_malloc/sf_free (alloc_kernel.c) was a 4MB bump arena with a
   deliberate no-op free, sized on "VM born once, never killed" -- fix #2
   alone didn't stop the panic because free() itself discarded the
   pointer regardless. Given a real free list (first-fit reuse).

4. Headless-console gate didn't re-engage after a mid-boot logout: the
   original fix (sk_console_mark_login(), one-way sticky) only gated the
   first login of the boot. Replaced with a live check
   (sk_console_identity_present()) re-evaluated continuously, including
   inside sk_console_readline()'s own blocking idle loop -- the console
   is normally sitting blocked there when a hot-unplug logout happens, so
   checking only at the top of the REPL loop wasn't enough.

Also: MINT now verifies its own write (verify_mint(), capsule_mint.c) by
reading back through the same check a real attach performs, rather than
trusting blkio_write()'s BLK_OK alone -- logged via log_message(), not
console_println(), per direct instruction.

Verified live, amd64: the full 8-identity repeated attach/detach cycle
that previously panicked at the same point every time now completes
clean, and a full serial-log sweep found zero bare unauthenticated
prompts anywhere in the run. Three-arch clean-qemu acceptance passed.

Still open, not fixed here: a 3+-simultaneous-device USB enumeration
failure found in a separate live test, not yet root-caused.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-06 01:49:13 -04:00
Robert Allan JamesandClaude Sonnet 5 0bae928aad Mint the 8 identity thumbdrives (bob, 00-06) with real MINT data
zuse-thumb-ident.img was already minted; bob-thumb-ident.img (Captain
Bob / rajames / rajames440@gmail.com) and 00-thumb-ident.img through
06-thumb-ident.img (full_name=username=the number, no email/phone) are
now real minted identities, not blank images.

Done via the actual kernel MINT word at a live console (Zuse
authenticated, drives hot-plugged one at a time via the QEMU HMP
monitor) -- no Python orchestration, per direct instruction: a small
StarForth word, MINT-NUM, wraps MINT for the 7 uniform numeric
identities, driven by plain bash + socat against the monitor/serial
unix sockets.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-05 22:42:41 -04:00
Robert Allan JamesandClaude Sonnet 5 9e81de3f43 xHCI/BOT driver: genuine multi-device support (FABRIC-3.md §VII)
Per-slot registry (xhci_msc_slot_t/dev->msc_slots, sized off the
controller's own reported max_slots) replaces the single-device scalar
fields the driver carried since Milestones 2e-2h. Boot-time port scan no
longer stops at the first connected device; a connect/disconnect that
arrives while the Command Ring is busy is now queued and drained instead
of dropped. blkio_usb.c and repl.c's own single-device state (device
descriptor buffers, blkio_dev_t, attach bookkeeping) became per-slot
registries the same way.

Live multi-device testing (not just compiling) surfaced a second, more
severe bug outside the original plan: transfer_purpose and next_action
were also single scalars shared across the whole controller. Two devices
enumerating concurrently could have one's completion silently overwrite
the other's still-outstanding one, permanently stalling it with no error.
Fixed by moving both per-slot and, critically, reading the Transfer Event
TRB's own real Slot ID field instead of trusting external bookkeeping.

Verified live, all three architectures, mandatory clean-qemu acceptance:
existing single-device path unchanged, and two devices attached
simultaneously (amd64) both progress independently through enumeration
without corrupting or stalling each other.

Also in this pass (implemented and verified in earlier turns this
session, committed together per direct instruction):
- Headless-until-login console policy: no prompt/banner until a real
  identity logs in via an attached thumbdrive (WIREBIND or Zuse, neither
  special), reusing EMERGENCY_CONSOLE_ENABLED as the debug/recovery
  escape hatch (now default-off).
- KILL/g_repl_active_vm dangling-pointer fix: killing the VM the console
  is currently USE'd onto now detaches back to Hera first, matching the
  existing EJECT/UNCLEAN precedent.

FABRIC-3.md §VII/§VIII carry full closure notes for all three.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-05 22:14:14 -04:00
Robert Allan JamesandClaude Sonnet 5 cc6fcb6a0b FABRIC-3.md §VII: rewrite punch list to real code-level detail
Previous pass was too abstract for this series' own bar. Re-traced
xhci.c/repl.c function-by-function: found max_slots is already read and
correctly sizes the DCBAA (so the fix reuses that value, not "start
reading a register"); found xhci_scan_ports_for_already_connected()
deliberately breaks after the first hit, missing a second already-
connected device at boot entirely; found block_subsystem.c's attach
layer is already multi-device-capable, narrowing the real singleton to
xhci_dev_t's own fields plus three repl.c pointers. Punch list items now
name exact functions/fields/line numbers and what's proposed vs. already
true. Also drops the earlier small-fixed-N concurrency bound per
direct correction -- size off dev->max_slots, not a guessed ceiling.

Still plan-only. No driver code touched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-05 21:00:45 -04:00