Commit Graph
235 Commits
Author SHA1 Message Date
Robert Allan JamesandClaude Sonnet 5 41918a28a4 doe_log.c: add vm_name/vm_id_hex identity columns; new calibration workload
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Adds the identity columns HB-ON's existing per-tick CSV logger
(doe_log_tick_row(), doe_log.c) was missing. Previously every column
described *a* VM's state each row, but nothing said which VM emitted
it -- concurrent VMs' rows were indistinguishable by source. Two new
first columns, vm_name and vm_id_hex, resolved via a reverse lookup on
vm->stadium_vm_id (capsule_vm_registry_get()).

Real bug found and fixed while building this: the freestanding
snprintf here silently prints the literal format string instead of
substituting for %016llx (width+ll+hex unsupported) -- caught by
reading the actual emitted row, not assumed to work. Replaced with a
hand-rolled hex nibble-table loop, the same idiom vm_uuid_format()/
MINT-SCRATCH-EMIT already use.

Corrects FABRIC-3.md's own prior "still not started: CSV driver"
framing: no bespoke CSV emitter is needed for the multiuser DoE at
all -- this per-tick logger already exists, fires automatically inside
every VM's own execution loop, and just needed HB-ON plus these
identity columns to be usable for concurrent workers. Also corrects
doe_log.h's own stale doc comment claiming g_doe_log_enabled defaults
to 1 -- doe_log.c's own source is authoritative: 0, off by default,
matching the HB-ON/HB-OFF naming.

One real, honest limitation recorded rather than smoothed over:
vm_name reads blank for any tick captured during a WORKER-BIRTH'd
capsule's own self-execution at birth time, since capsule_birth_baby()
runs that work (and its heartbeat ticks) before the registry name can
be set. vm_id_hex is unaffected (reads directly from vm->stadium_vm_id)
and still uniquely disambiguates every row -- confirmed live: a
blank-name row's own vm_id_hex matched its later PARITY:KILL line's
vm_id exactly.

New workload-calib1.4th: Bob confirmed the existing 10 workload-N.4th
files aren't fixed. Sized to cross HEARTBEAT_CHECK_FREQUENCY (256
word-executions/tick) many times over while staying far lighter than
RUN-FIB/RUN-CHAOS5, both too slow under TCG for a bounded verification
run. A real, reusable addition, not a throwaway.

Verified live on amd64 via a self-contained one-shot EXEC'd test
(HB-ON, WORKER-BIRTH two calibration workers, HB-OFF, VM-ERROR? on
both, KILL both, completion banner) rather than interactive polling,
which breaks once HB-ON makes the log grow continuously via Hera/
Hermes/Artemis's own background ticks. Clean 3-arch qemu boot on the
real committed change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 09:53:25 -04:00
Robert Allan JamesandClaude Sonnet 5 e33eb36361 Multiuser DoE punch list: WORKER-BIRTH + VM-ERROR? + vm_physics_init fix
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Core mechanism for §XXXIII's main concurrency block, built and
live-verified on amd64. Two new primitives:

- WORKER-BIRTH ( capsule-c capsule-u name-c name-u -- ok? ): births a
  named, VM-EXEC-addressable VM from an arbitrary (p) capsule with no
  identity involved. Corrects the ratified design's own assumption that
  the main block would use UNATTENDED-BIRTH -- that requires a
  committed capsule per identity, impractical for dozens of trial VMs.
  The concurrency block never needed identity at all.
- VM-ERROR? ( c-addr u -- flag ): reads a named VM's error state from
  Hera, mirroring VM-HEAT's silent/always-returns-a-value contract.
  Needed to check a VM-EXEC-driven trial VM's own fault state after
  the fact -- nothing existing let Hera do this.

Real bug found and fixed in both WORKER-BIRTH and UNATTENDED-BIRTH:
capsule_birth_baby() never calls vm_physics_init() either (same shape
as the registry-name gap found building UNATTENDED-BIRTH) -- without
it a born VM is never in the VM Fleet Attractor physics list, so
VM-HEAT returns 0 forever regardless of work done. Fixed by adding
vm_physics_init() alongside the existing registry-name call in both
words.

Traced (not guessed) why heat still read 0 after one VM-EXEC touch
even post-fix: vm_physics_touch()'s transfer logic only fires from a
VM's *second* touch onward -- the first touch just records a baseline
tick. Verified live across three sequential touches: heat 0 -> 7039 ->
11333. This is a real design requirement for the DoE's heat/CV
response variable (each trial must touch a worker at least twice), not
a bug to route around.

Also found live: all 10 existing workload-N.4th capsules self-execute
their full workload at load/birth time (a bare top-level call to their
own RUN-* word at file end) -- missed on an earlier, too-shallow
8-line survey of each file. WORKER-BIRTH alone already runs a worker's
first pass as a side effect of birth.

Verified live on amd64: two concurrent workers (fib + matrix-mul)
birthed, run, measured (heat + error state), and killed cleanly;
VM-HEAT/VM-ERROR? both confirmed silent-0 on an unknown name. Clean
3-arch qemu boot on the real committed change.

Still open: the run-matrix/shuffle/CSV driver capsule itself, the
WIREBIND-automation path for the fixed arm, and the per-VM touch-count
budget's exact value.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 08:51:21 -04:00
Robert Allan JamesandClaude Sonnet 5 af351bff92 Punch list item 4 (final): CONSOLE-ATTACH, full unattended-identity flow verified end-to-end
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Closes FABRIC-3.md §XXXII.2's punch list. CONSOLE-ATTACH ( name-c
name-u -- ok? ) pairs a fresh console VM to an already-live VM
registered as "<name>~user". Deliberately a plain, unconditional
primitive with no VMIdentity capability-bit check -- corrects this
session's own first-pass design (§XXXII.2 amended in the same commit):
identity.installed is 0 for Hera/Hermes/Artemis and for every
console-proxy VM, so a capability-bit gate would be unreachable for
every VM a human actually types at, Zuse included. Matches
ZUSE-ELIGIBILITY-ADD's own "no bespoke gate" precedent in this file;
real gating is `' CONSOLE-ATTACH ACL-PIN` in ACL.4th if ever wanted.

Two real bugs found and fixed via live testing, not assumed correct:
- capsule_birth_baby() never sets a VM's registry name (documented
  gap, same one capsule_runcap_birth()'s own history already hit) --
  UNATTENDED-BIRTH gained a second `name` argument and now calls
  capsule_vm_registry_set_name(new_vm_id, "<name>~user") itself.
- CONSOLE-ATTACH's first draft took an independent console name from
  the target's name. sk_repl_dispatch_line()'s pairing check (repl.c)
  reconstructs the target as console_get_vm_name()+"~user" -- a
  mismatched console name silently falls back to direct interpretation
  with no error. Caught live (typed `5 6 + .` at a mismatched console,
  got a direct `11` instead of a relay) and fixed by collapsing to one
  name argument, matching WIREBIND's own by-construction invariant.

Full end-to-end live verification on amd64: minted a test identity,
UNATTENDED-BIRTH'd it as "bob", CONSOLE-ATTACH'd a console named "bob",
USE'd it, typed `5 6 + .` -- no direct output at [zuse@bob] (relay
path taken), then `[zuse@bob~user] 11 ok>` appeared: the relayed
command executed on the target identity VM itself and printed its own
answer back through the shared console. Hera stayed healthy throughout
(2 2 + . -> 4 after switching back). CONSOLE-ATTACH also verified to
refuse cleanly on a nonexistent target with no orphaned VM. Test
capsule reverted after capture per this project's probe convention --
never committed.

FABRIC-3.md §XXXII.2 fully closed: all 4 original questions ratified,
the mid-course drive_uuid and ACL-bit corrections both recorded
plainly rather than silently folded in, and a doc-accuracy note left
for CLAUDE.md's own stale "1024-byte block limit" framing (mkcapsule's
real limits are range [2048,5120) and 16 content lines/block) --
flagged, not fixed, out of this punch list's scope.

Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real committed
C-only change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 07:24:27 -04:00
Robert Allan JamesandClaude Sonnet 5 5879c8b3bc Punch list item 3: UNATTENDED-BIRTH call site, verified live end-to-end
Implements the unattended-birth mechanism FABRIC-3.md §XXXII.2 designed:
UNATTENDED-BIRTH ( name-c name-u -- ok? ) births a VM from a named (p)
capsule via capsule_birth_baby() (completely unmodified, the same
generic build-time-capsule path CAPSULE-BIRTH already uses), then
installs its identity the same way capsule_wirebind.c already does
live for WIREBIND attaches -- vm_identity_from_cert() verification
followed by a plain post-birth struct assignment -- rather than
anything RUNCAP-shaped, since RUNCAP requires a real blkio_dev+
homeblocks_sig_t an unattended identity never has.

The born VM's own capsule payload is expected to lay down two CREATE'd
buffers (UNATTENDED-ID-UUID, UNATTENDED-ID-CERT) via MINT-SCRATCH-
EMIT's own literal format; their addresses are fetched by interpreting
a two-word line inside the *new* VM's own context
(vm_interpret(born_vm, ...)), the same "run inside that VM's own
dictionary" idiom capsule_wirebind.c already uses for VM-NAME-REG.

Explicit invariant preserved: never touches g_wirebind_attached_username
or any console-pairing state, births no console VM -- an unattended
identity stays un-promptable (§VIII.1) until a human pairs a console to
it later via the already-working VM-NAME-REG mechanism.

Verified live end-to-end on amd64: minted a real test identity via
MINT-SCRATCH, captured its MINT-SCRATCH-EMIT output, built a throwaway
test capsule from it (discovered along the way: mkcapsule's real block
constraints are range [2048,5120) and max 16 content lines per block --
neither matches this repo's own doc comment, corrected via ground
truth from the tool itself, not assumed), then ran UNATTENDED-BIRTH
against it: cert verified against Zuse's root pubkey, identity
installed, "no console attached" reported, Hera stayed healthy
afterward (5 6 + . -> 11). Test capsule reverted after capture per
this project's own probe convention -- not a real identity, never
committed. Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real
committed C-only change.

Remaining punch-list item (the ACL cap bit for console attachment) not
started.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 07:06:47 -04:00
Robert Allan JamesandClaude Sonnet 5 5fc709a228 Punch list item 2: MINT-SCRATCH-EMIT, verified live on all 3 arches
Adds the hand-transcription mechanism FABRIC-3.md §XXXII.2's punch list
item 2 calls for: MINT-SCRATCH-EMIT prints the last successful
MINT-SCRATCH's drive_uuid + cert devblock as ready-to-paste FORTH
source (HEX-based CREATE ... C, ... byte sequences), so packaging an
unattended identity into a capsule is a mechanical copy out of the
captured boot log rather than a manual hex-to-FORTH translation an
operator could transpose a digit in.

Refuses (no output) if no MINT-SCRATCH has ever succeeded -- printing
4112 zero bytes as if they were a real identity would be a silent,
misleading success, matching this session's own error-handling audit
discipline rather than adding a new silent-failure primitive right
after finishing one.

Verified live on amd64: MINT-SCRATCH-EMIT correctly refuses before any
mint, then after MINT-SCRATCH succeeds, emits UNATTENDED-ID-UUID and
UNATTENDED-ID-CERT as valid FORTH literals. Cross-checked byte-exact:
the cert's own embedded ASN.1 serialNumber field matches the emitted
UUID bytes exactly, confirming x509_build_user_cert()'s drive_uuid
binding round-trips correctly through the scratch-device path. Clean
3-arch qemu boot (amd64/aarch64/riscv64).

Remaining punch-list items (writing/committing a real capsule file for
an actual named identity, the unattended-birth call site, the ACL cap
bit) not started -- authoring a real committed capsule needs a name/
purpose decision that isn't mine to make.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 05:12:43 -04:00
Robert Allan JamesandClaude Sonnet 5 63864c4b01 Punch list item 1: scratch-device MINT-SCRATCH, verified live on all 3 arches
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Implements FABRIC-3.md §XXXII.2's scratch-thumbdrive mint mechanism:
capsule_mint_identity_scratch() (capsule_mint.c/.h) builds a throwaway
RAM-backed blkio_dev via blkio_ram.c's backend and runs
capsule_mint_identity() against it completely unmodified -- same live
Zuse-signing operation, same rng_get_bytes() draw for drive_uuid a real
thumbdrive gets. Reads back only drive_uuid + the cert devblock; the
seed devblock is written into the scratch buffer internally but never
read out (no seed is ever baked into a capsule, per the ratified
no-seed decision).

New FORTH word MINT-SCRATCH (mama_forth_words.c), same stack signature
as MINT, mints into the scratch device instead of any attached drive
and never touches sk_repl_get_attached_blk_dev() or console-pairing
state. Prints the drive_uuid as hex so a live boot log itself proves
each call drew fresh entropy.

Build correction found along the way: blkio_ram.c was excluded from
the kernel build (Makefile.starkernel VM_EXCLUDE) alongside
blkio_factory.c/blkio_file.c. blkio_factory_open() unconditionally
references blkio_file.c's real fopen()/fread() file I/O, which has no
freestanding-kernel equivalent, so the factory function couldn't be
used as-is. blkio_ram.c itself is pure memcpy over a caller buffer --
pulled it alone into the kernel build and wired it directly in
capsule_mint.c, the same way blkio_factory.c's own extern declarations
do internally.

Verified live on amd64: two MINT-SCRATCH calls produced two genuinely
different drive_uuids (b533246d.../ae11b2b2...), confirming fresh
entropy per call rather than stale reuse; VM stayed healthy afterward
(5 6 + . -> 11). Clean 3-arch qemu boot (amd64/aarch64/riscv64),
logs and DoE CSVs committed per standing convention.

Remaining punch-list items (hand-transcription into a .4th block, the
unattended-birth call site, the ACL cap bit) not started.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 04:42:15 -04:00
Robert Allan JamesandClaude Sonnet 5 42d4bf3dad Stage D batch 7 (final): mama_forth_words.c groups 5+6 -- dictionary lookup/test words + Stadium primitives; Stage D closed (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
NAME>XT/RUNCAP-TEST/PAIR-TEST and the six STADIUM-* physics primitives
(STADIUM-ADMIT/STADIUM-EVICT/STADIUM-RES-PULL/STADIUM-RES-PUSH/
STADIUM-HEAT@/STADIUM-HEAT!), 12 sites -- closes mama_forth_words.c
and the entire kernel-only error-handling audit.

All 105 sites from the §XXXII.3 triage now accounted for: 28 already
correct, 75 silent sites fixed across repl.c/inference_words.c/
log_words.c/vm_core.c/mama_forth_words.c, 2 special cases resolved by
dropping the error per their own documented contract, 1 resolved via
console_println() per its own recursion constraint. Verified with awk:
zero remaining vm->error=1 sites in mama_forth_words.c lack a
diagnostic within the preceding three lines.

Three-arch clean qemu acceptance passed. This closes Stage D and the
USE/logging/audit thread opened in §XXXII; Stage E (human-vs-unattended
identity model) remains open, not started this pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 20:25:30 -04:00
Robert Allan JamesandClaude Sonnet 5 90c86c6006 Stage D batch 6: mama_forth_words.c group 4 -- identity/crypto words (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
MINT + mint_pop_string()/ZUSE-ELIGIBILITY-ADD/ZUSE-ELIGIBLE?/
ELEVATE-PUBKEY-UNPACK, 10 sites. mint_pop_string() gained a field_name
parameter so its diagnostics name which of MINT's four string
arguments failed (phone/email/username/full_name), rather than a
generic message that would leave the operator guessing. Diagnostic
placement matched each function's own sibling convention where one
exists (ZUSE-ELIGIBILITY-ADD), log_message() default otherwise.

Three-arch clean qemu acceptance passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 19:53:47 -04:00
Robert Allan JamesandClaude Sonnet 5 8b5300fc4f Stage D batch 5: mama_forth_words.c group 3 -- cross-VM execution/dispatch, both flagged special cases resolved (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
VM-STEP/VM-EXEC/VM-CALL (5 sites) gained console_println() diagnostics
matching their own existing sibling guards.

SWITCH-MARK-WORK and VM-HEAT (3 sites) resolved per §XXXII.3's own
recommendation: dropped vm->error entirely rather than diagnosing it,
matching each function's own doc comment ("must never error or spam
the console" / "does not print/error"). SWITCH-MARK-WORK is the same
function §XXVIII.3 already fixed once for an off-by-one that fired
silently on every MSG-SEND in the system -- the guard now genuinely
cannot repeat that by contract, not just by the threshold being right.
VM-HEAT's guards now push 0 and return, matching its own
always-returns-a-value stack effect.

Three-arch clean qemu acceptance passed. Zuse's own WIREBIND attach
(every boot) drives MSG-SEND -> SWITCH-MARK-WORK, so this batch's most
safety-critical fix is exercised by standard acceptance, not just
compiled.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 19:23:12 -04:00
Robert Allan JamesandClaude Sonnet 5 b42c3b195c Stage D batch 4: mama_forth_words.c groups 1+2 -- capsule/lifecycle words + BIRTH/START/KILL (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
15 of mama_forth_words.c's 48 silent sites fixed, split by functional
grouping per direct instruction: CAPSULE@/CAPSULE-HASH@/CAPSULE-FLAGS@/
CAPSULE-LEN@/CAPSULE-BIRTH/CAPSULE-RUN/EXEC (9), and BIRTH/START/KILL
(6).

Refined the diagnostic-placement rule: match whichever convention that
same function's other already-correct guards use, rather than
defaulting uniformly. BIRTH/START/KILL/EXEC each already had a
console_println() sibling guard ("name too long or empty") -- their
newly-diagnosed guards now match that, same shape as USE's own fix.
The CAPSULE*@ words have no sibling guard to match, so they keep
log_message() (defer_words.c's gold-standard default).

Three-arch clean qemu acceptance passed; BIRTH itself is exercised by
every boot (Hermes/Artemis birth).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 18:51:54 -04:00
Robert Allan JamesandClaude Sonnet 5 66c2f3e539 Stage D batch 3: fix silent error sites in vm_core.c, incl. the two highest-value primitives; two real NULL-deref bugs found and fixed (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
All 13 silent vm->error=1 sites in vm_core.c now log a diagnostic
first: vm_enter_compile_mode, vm_compile_word, vm_compile_literal,
vm_compile_call, vm_exit_compile_mode, execute_colon_word (2 sites
each/combined), and the four fundamental memory primitives
vm_load_u8/vm_store_u8/vm_load_cell/vm_store_cell.

Found and fixed two real NULL-pointer-dereference risks while adding
the diagnostics: vm_compile_call() and vm_exit_compile_mode() each had
a combined `if (!vm || <cond>) { vm->error = 1; ... }` guard that
dereferenced vm->error even on the !vm branch of its own condition.
Split both, and applied the same defensive split to the four memory
primitives since vm_ptr()/vm_addr_ok() both tolerate vm==NULL
internally.

Live-verified, amd64: 999999999 @ . recovers correctly at Zuse's own
console (Stage A's mechanism holds), but the new vm_load_cell
diagnostic itself didn't print -- traced to memory_words.c's own
redundant, still-silent vm_addr_ok() pre-check in memory_word_fetch()
(and the same shape in memory_word_store()), which intercepts before
ever reaching vm_load_cell(). memory_words.c is vendored, out of this
initiative's scope, spun off to FABRIC-4.md -- recorded as a concrete
cross-reference for that future work rather than left to be
rediscovered.

Three-arch clean qemu acceptance passed, full POST suite included.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 17:29:37 -04:00
Robert Allan JamesandClaude Sonnet 5 1a8c0e4fcf Stage D batch 2: fix silent error sites in log_words.c, resolve the flagged category-iii special case (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
log_word_set_level(), log_do_emit(), and log_emit_string() (7 sites
total) now log a diagnostic via log_message(LOG_ERROR, ...) before
setting vm->error, matching this file's own already-correct
log_str_emit() sibling and defer_words.c's gold-standard pattern.

log_word_append_raw() (backing (LOG-APPEND-RAW), 3 sites) resolved
differently per §XXXII.3's own triage note: its doc comment forbids
log_message() here (recursion into the log ring it writes to), but the
word can also be invoked by hand at the console -- console_println()
carries no such recursion risk and matches Stage B's policy for a
manual interactive invocation. Added the console.h include this
required.

Three-arch clean qemu acceptance passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 16:53:18 -04:00
Robert Allan JamesandClaude Sonnet 5 5417ffb2bf Stage D batch 1: fix silent error sites in repl.c and inference_words.c (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
repl.c's sk_word_blk_attach_ack() and inference_words.c's array_ptr()
helper + infer_word_run()'s allocation guard now log a diagnostic via
log_message(LOG_ERROR, ...) before setting vm->error, matching
defer_words.c's own gold-standard pattern (§XXXII.3) -- these are
internal/background conditions (a malformed message-callback, a bad
array reference or allocation failure), not interactive usage mistakes,
so the fix keeps the fault and reports it rather than dropping it like
USE's own fix did.

Batched together (4 sites total, smaller combined than the next file)
rather than two separate acceptance cycles for negligible size.

Three-arch clean qemu acceptance passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 16:23:14 -04:00
Robert Allan JamesandClaude Sonnet 5 c05b70c8d6 Stage B: logging policy documented, level-aware log-ring eviction built, LOG-FLUSH deferred again (FABRIC-3.md §XXXII.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Policy decided for the kernel-only audit scope: an interactive
command's direct response stays on console_println/console_puts;
everything else (state transitions, background diagnostics, audit
trails) routes through log_message() at the appropriate level,
matching capsule_mint.c's verify_mint() precedent. Documented, not
code-swept here -- reclassifying individual sites is Stage D's job.

LOG-FLUSH deferred again, explicitly: the per-VM log buffer its own
doc comment presumes (vm_log_buffer.h) doesn't exist anywhere in the
tree -- building it is real feature work needing its own scoped stage.

Level-aware eviction built: log_region_append() now reads the oldest
ring slot's own level before evicting it, protecting ERROR/WARN
records from being pushed out by INFO/DEBUG churn -- drops the
incoming low-priority record instead. Found and fixed an adjacent bug
while making this change: the prior two-valued return contract would
have made a benign "dropped by design" outcome indistinguishable from
a genuine write failure to its one caller, which unconditionally set
vm->error on any nonzero return. Changed to a three-valued contract (0
success, 1 dropped by design, -1 genuine failure).

Three-arch clean qemu acceptance passed. Eviction path itself not
live-exercised (needs 128+ LOG-APPEND calls to fill the ring) --
flagged, matching this project's own precedent for that kind of gap.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 14:01:36 -04:00
Robert Allan JamesandClaude Sonnet 5 5c5896fbc1 Stage A: fix USE's silent stack-underflow/bad-address guards, and the Hera fault-scoping gap they exposed (FABRIC-3.md §XXXII.1)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
mama_word_use()'s two silent vm->error=1 guards (dsp<1 stack underflow,
NULL from vm_ptr()) now print a diagnostic and return, matching the
function's other five guards.

Live verification of that fix alone surfaced a bigger problem: with the
guard no longer silent, the REPL proceeds to interpret the leftover
token as an unrecognized word, which independently sets vm->error, and
sk_repl_step()/sk_repl_run()'s Hera-branch still hard-halted on that.
Investigated kernel_main.c's boot/capsule-load paths directly: they
already catch and clear mama->error entirely separately, before
sk_repl_run() is ever entered -- so the "no fallthrough surface" halt
in these two REPL functions was never protecting a boot-time fault, only
an ordinary interactive REPL-turn one. Both functions now recover
unconditionally on any VM's error, Hera included, matching how a
redirected (WIREBIND/USE'd) identity's session already recovered.
Removed the now-fully-unused sk_fault_handler().

Verified live, amd64: USE rajames at Zuse's own console prints the new
diagnostic, then "VM fault -- session recovered, resuming", console
stays interactive afterward. Three-arch clean qemu acceptance passed
before and after the Hera-fault-scoping change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 09:06:33 -04:00
Robert Allan JamesandClaude Sonnet 5 d9da82b065 Stage 4 increments 2+3: WIREBIND VMs as switch-signal participants + mark-and-defer tombstone reap (FABRIC-3.md §XXVIII Stage 4)
Increment 2: WIREBIND user VMs (the ones that actually run FORTH work;
console VMs are pure REPL proxies and never participate) register as
Stage 3 switch-signal participants at attach, unregister at teardown.
Slot table bumped 8 -> 16, matching messaging.4th's own VM-MAX -- a real,
already-agreed ceiling, not an invented number. Added
sk_vm_switch_signal_unregister() (compaction-based; Tripod VMs never
needed removal, WIREBIND VMs cycle constantly and would otherwise
exhaust the bounded table).

Increment 3: implements the plan's own ratified option (A) for the
async-detach UAF risk -- mark-and-defer via a new pending_reap flag on
VMRegistryEntry, deliberately not a new VMState (capsule_vm_kill()
already treats VM_STATE_DEAD as idempotent success, which would silently
swallow a reap attempt; SWITCHED_OUT still accurately describes a
tombstoned VM until the moment it's actually freed). unclean_detach()
sets it when capsule_vm_kill() refuses a SWITCHED_OUT target; the Stage 3
checkpoint (vm_core.c) checks it before ever attempting to resume a
pending switch target, and calls the new capsule_vm_force_reap() instead
-- the one caller allowed to bypass capsule_vm_kill()'s own refusal,
because it runs at the exact safe cooperative point the switcher itself
controls. A new idle-tick sweep cleans up the WIREBIND live-table entry
once the reap has actually happened.

Verified clean on all 3 architectures (baseline regression -- no
WIREBIND attach happens in a plain boot). The reap mechanism's own
correctness under a genuinely parked context is verified separately,
next, via a temporary deterministic probe.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 00:58:34 -04:00
Robert Allan JamesandClaude Sonnet 5 9f0f33dfc5 Stage 4 increment 1: per-device WIREBIND tracking, fixing a real multi-identity detach leak (FABRIC-3.md §XXVIII Stage 4)
capsule_wirebind_unclean_detach()/eject() tracked "the attached identity"
as a single global, correct for the console-pairing UX (one physical
console) but wrong for detach safety: since §XV/§XVI proved multiple
identities genuinely live simultaneously via this same attach path, every
attach after the first silently overwrote the singleton, so an unclean
detach of any but the most-recently-attached identity was silently
ignored -- that VM leaked forever, no trace in the log.

Adds a per-device live-identity table, separate from the (unchanged)
console-pairing singleton, so unclean-detach resolves any attached
device to its own identity. Sized off messaging.4th's own VM-MAX (16)
minus Tripod's 3 reserved slots, not an invented number. Corrects the
stale "single-USB-device constraint" doc claim in capsule_wirebind.h,
false since §XV/§XVI.

Groundwork for Stage 4's real deliverable (WIREBIND VMs as switch-signal
participants) -- this increment only fixes detach targeting; switch-
signal registration is next.

Verified clean on all 3 architectures (no WIREBIND attach happens in a
plain boot, so this is a regression check on the existing Tripod-only
path; live multi-identity verification comes with the switch-signal
registration increment).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 00:02:20 -04:00
Robert Allan JamesandClaude Sonnet 5 05ae7aa886 Fix SWITCH-MARK-WORK off-by-one: was tripping vm->error on every MSG-SEND (FABRIC-3.md §XXVIII.2)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
mama_word_switch_mark_work()'s stack-underflow guard checked dsp < 2,
requiring 3+ items, when it only ever needs the 2 IDX>NAME leaves it
(caddr u). Since dsp is index-based (2 items == dsp 1), this rejected
every normal call. MSG-SEND tail-calls SWITCH-MARK-WORK unconditionally,
so this fired on every message sent anywhere in the system -- visible
only where a caller happened to check the target VM's error flag
afterward (mama_word_vm_exec()'s "VM-EXEC: ERROR in Artemis" report).

Verified clean on all 3 architectures: boot reaches Startup: Artemis
live -> zuse@Hera] ok> with no VM-EXEC: ERROR line at all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-14 20:17:01 -04:00
Robert Allan JamesandClaude Sonnet 5 66beae7fd4 Stage 3 follow-on: message-arrival eligibility hook + trampoline-blind switch-storm fix (FABRIC-3.md §XXVIII.2)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Implements the message-arrival eligibility signal FABRIC-3.md §XXVIII.1 left
open (has_work per-slot flag, set via new SWITCH-MARK-WORK primitive from
MSG-SEND) so an idle VM never becomes a switch target purely by waiting out
the readiness threshold.

Also root-causes and fixes a second, independent switch-storm: the tick's
"who is current" check used vm_log_attributed_vm(), which can't see a VM
parked in switch.c's own raw trampoline. Replaced with a dedicated
g_switch_current_vm tracked by the switch mechanism itself, and moved
target-slot eligibility reset to the switch decision point instead of
relying on ISR polling to observe a window that can be only a few
instructions wide.

Verified live on all 3 architectures: clean boot to zuse@Hera] ok>, live
cross-VM message dispatch, and (since a quiet log looks identical to a
livelocked storm once the DoE probe is gone) confirmed genuine REPL
liveness via QMP send-key + screendump on aarch64/riscv64, not log
inspection alone.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-14 17:39:20 -04:00
Robert Allan JamesandClaude Sonnet 5 862d7d9c48 Stage 3 follow-on: fix stack-ownership corruption + DoE switch columns (FABRIC-3.md §XXVIII.1)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
DoE CSV gained 6 switch-signal columns (switch_count_cumulative,
switch_current_slot, switch_*_readiness, switch_ticks_since), and verifying
them with a boot-time HB-ON probe surfaced a real livelock: the preemption
checkpoint could fire inside a VM-EXEC-nested execute_colon_word() call and
switch away from a stack it didn't own, parking a borrowed region of the
caller's stack under the wrong VM's saved-context pointer. The trampoline
bounce was the visible (safe) half of this; the corruption was the quiet
half, live in every prior "clean" Stage 3 boot without ever showing up in
the log.

Fixed by gating the checkpoint on being at the outermost vm_interpret()
call (g_vm_interpret_depth / sk_vm_at_outermost_interpret(), vm_core.c),
per Bob's decision. Also fixed two related bugs found in the same pass:
g_switch_back_to was a single global stale after first entry, now per-VM
state (native_switch_back_to); note_switch_performed() fired on resume
instead of switch-out, now called before the switch.

Verified on all 3 architectures: steady log growth (no freeze), zero
leaked QEMU processes, DoE columns internally consistent, Hermes/Artemis
confirmed genuinely executing (not just trampoline-bouncing). Temporary
HB-ON boot probe reverted after capture.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 23:58:59 -04:00
Robert Allan JamesandClaude Sonnet 5 986d042aa7 Stage 3: timer-driven preemptive switching, live on all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Fourth stage of the preemptive context-switching plan, and the biggest.
LithosAnanke now genuinely, continuously preempts between Hera, Hermes,
and Artemis -- timer-driven, running live for the entire remainder of
every boot once the Tripod fleet registers, not a bounded probe.

A real design fork was resolved before writing code: the naive approach
(the timer ISR calling Stage 2's sk_vm_context_switch() directly) is
broken -- Stage 0's trap frame lives on whatever stack was active at
interrupt time, and jumping to a different stack via Stage 2's own
independent swap mid-handler would abandon that trap frame unresumed,
guaranteed corruption on the first tick. Chose the safer of two named
options: the ISR only ever sets a flag and returns completely normally
through its own full epilogue; the actual switch happens moments later,
via Stage 2's already-proven mechanism, at a safe cooperative checkpoint
on the mainline (execute_colon_word()'s per-word dispatch loop, checked
on literally every word, not throttled to the existing 256-word
heartbeat-tuning cadence) -- confirmed with the user that word-level
granularity is fine-grained enough given the eventual Zynq FPGA target
where a word is a mnemonic.

New capsule_vm_switch_signal.c/.h: a purpose-built run-readiness signal,
deliberately separate from capsule_vm_physics.c's execution-heat engine
(that one's own header documents itself as never touched from interrupt
context, by design). Slot table sized with headroom (8) rather than
hardcoded to today's 3 participants, so extending participation later is
another register() call, not a redesign -- per direct request to leave
room for swapping the participant set. Simple linear accumulate-then-
threshold for this first cut; a fancier law can replace it later without
touching the mechanism around it. heartbeat_tick() gains its one
deliberate, documented amendment to this file's own top-half/bottom-half
discipline -- the first time this codebase reaches into VM-scheduling
state from real ISR context.

Registration happens only after all three VMs are fully born, right
before the REPL starts -- no critical-section protection yet against
being switched away mid-birth-setup.

Known, flagged rough edge (not reconciled this pass): MSG-TICK's own
idle-pump and this new mechanism can still independently move control
between the same VMs; not observed to interact badly in verification,
but not fully unified either.

Verified interactively at the console on all 3 architectures with
continuous background preemption running throughout -- amd64 computed
`1 1 + .` -> 2, aarch64 computed `1 1 + dup DUP * . CR` -> 4, both
correct, REPL fully responsive, zero fault indicators over sustained
runtime.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 19:32:36 -04:00
Robert Allan JamesandClaude Sonnet 5 f790d0995e Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Third stage of the preemptive context-switching plan. The real
save/restore switch mechanism now exists -- the first time anything has
ever executed on a VM's own native stack (Stage 1 allocated them,
unused).

New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function
call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already
covers every caller-saved register; only the callee-saved set needs
explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV
has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64:
s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for
both halves within one call, since FS is genuine global CPU state, not
part of what's switched). A sibling sk_vm_switch_prime() in the same file
builds the synthetic first-entry frame, kept in assembly so the layout
can never drift out of sync with sk_vm_switch_to() itself.

New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry
priming vs. resuming a parked context, and updates registry state (new
VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no
live frame, this means the opposite). sk_vm_switch_entry() is the minimal
permanent trampoline every freshly-entered VM lands in: no production
behavior defined yet, so it just yields straight back to whoever switched
to it, forever.

Closes the confirmed unguarded-KILL UAF found during planning:
capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama()
all now refuse (or silently leak rather than free, on the cold-restart
path where arch_cold_reset() wipes everything immediately after anyway)
tearing down a switched-out VM. Side effect found, not built on purpose:
the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it
automatically stopped dispatching into a switched-out VM with zero
changes needed there.

Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing
can type interactively into a foreground-only QEMU session) that
round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3
architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe
fully reverted after capture; kernel_main.c shows zero diff.

Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new
files added explicitly (this project's "loader" PE binary is the full
running kernel, not a thin bootstrap stage), and aarch64's switch.S
needed the same #ifndef _WIN32 guard around .hidden that isr.S already
carries (aarch64's loader assembles via clang targeting a PE/COFF
target with no .hidden equivalent) -- caught by a build failure, fixed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 15:25:39 -04:00
Robert Allan JamesandClaude Sonnet 5 57ac3fc304 Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Second stage of the preemptive context-switching plan. Every VM (Hera,
every capsule_birth_baby()-born VM including WIREBIND identities) now
gets its own dedicated 2 MiB native C stack at birth -- but nothing runs
on it yet, that's Stage 2. Pure allocation-machinery proof.

Design correction made before writing code: the plan called for cloning
sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to
be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a
plain kmalloc() block for its dictionary arena, not a real guarded PMM
allocation. Stacks get the real treatment instead (new
sk_vm_native_stack_alloc()/_free() in arena.c): independent
pmm_alloc_contiguous() + guard pages for every VM without exception, no
singleton, no kmalloc fallback -- a stack overflow is exactly the
failure mode guard pages exist for, and a corrupted stack could corrupt
whatever saved context Stage 2 trusts.

2 MiB size matches this project's own established kernel-stack
convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one
shared 2 MiB stack today already carries all VMs' combined nested
VM-EXEC recursion.

Three new VM struct fields, freed in vm_cleanup() alongside the existing
call_stack free. Allocation failure is non-fatal to birth.

All 3 architectures re-verified clean boot to ok>, no native-stack
allocation failures for any Tripod-fleet VM.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 14:30:46 -04:00
Robert Allan JamesandClaude Sonnet 5 15672ce17c Stage 0: trap-frame parity across all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
First stage of the preemptive context-switching plan (see
~/.claude/plans/logical-snuggling-bear.md). Pure foundation work -- every
arch's ISR now saves the full register set on interrupt entry, so a trap
frame is in principle sufficient to resume execution anywhere it was
taken. No FORTH-visible behavior changes.

amd64: added FXSAVE/FXRSTOR, closing a genuine pre-existing correctness
gap (not just future-preemption prep) -- confirmed live double-precision
FP code reachable from ordinary interpreter dispatch (vm_runtime.c Loop
#5/#6), and the ISR previously saved zero FP/SSE state. rbp repurposed as
a fixed anchor so the 16-byte-aligned FXSAVE area can be carved out of an
unpredictably-aligned rsp without disturbing existing argument reads.

aarch64: extended the trap frame 672->800 bytes, adding v8-v15 (AAPCS64
callee-saved, previously excluded on call-site-only reasoning that
doesn't hold for an async trap).

riscv64: extended the trap frame 320->512 bytes, adding s0-s11 and
fs0-fs11 (the latter still correctly gated behind sstatus.FS != Off).

All 3 architectures re-verified clean boot to ok> under the new frames --
amd64 through hundreds of timer ticks with FXSAVE/FXRSTOR live on every
interrupt, aarch64 through 987 ticks, riscv64 clean on the now-larger
FS-conditional block.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 09:23:28 -04:00
Robert Allan JamesandClaude Sonnet 5 2a30212bd3 Real per-VM log persistence: source attribution + ACL pin (FABRIC-3.md §XXVII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Wires the previously-unused vm_log_attributed_vm() into LOG-APPEND's kernel
primitive so persisted log records carry a trustworthy source (the real
attributed VM's registry name, or "HADES" pseudo-source) instead of a
caller-supplied, trivially forgeable string. Drops src-addr/src-u from
LOG-APPEND's stack signature accordingly. Pins LOG-APPEND via bare ACL-PIN
in Artemis's own init.4th, matching BIRTH/CAPSULE-BIRTH's precedent for a
privileged word that can't reach the shared, host-portable ACL.4th.

Also fixes two console-banner nitpicks: a mis-rendering em dash (U+2014)
in the boot banner, and drops "Emergency" from the CLI banner text.

Doc corrections to artemis_sig.h/zuse_eligibility_list.h reconciling the
three fixed devblock ranges now in play. LOG-FLUSH (the intended normal
entry point) and level-aware log eviction remain open, flagged not fixed.

Re-verified clean boot to ok> on all 3 architectures after every change.
riscv64 showed one new, unrelated virtio_blk write-timeout anomaly during
Artemis's early physics self-test (self-recovered, boot unaffected,
sector doesn't map to the log region) -- flagged, not investigated.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 08:56:33 -04:00
Robert Allan JamesandClaude Sonnet 5 61755fde78 Artemis genesis stamp: fix a BAM-corrupting offset before it ever ran (FABRIC-3.md §XXVI follow-on, Step 3)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Step 3: one-time artemis_sig_t genesis stamp, written once
kernel_main.c's virtio-blk path confirms Artemis's own disk, so the disk
image is later recognizable generically (repl.c's idle-loop USB-MSC scan,
built in the prior commit) regardless of which bus found it.

Correction made before this ever touched the real disk: the signature's
first design (committed in 29b6789) placed it at a fixed bottom-of-device
forth-block (4, devblock 1) -- copying homeblocks_sig_t's own convention,
which is safe for a raw identity thumbdrive but not for Artemis's own
disk. Artemis's disk is block_subsystem.c's own STFR/v2-formatted volume:
devblock 0 holds that format's header and devblock 1 is the FIRST
DEVBLOCK OF THE LIVE BAM (blk_compute_fresh_geometry(): bam_start = 1).
The original design would have overwritten Artemis's live allocation map
on the very first real boot. Caught via direct cross-reference against
block_subsystem.c before the genesis-stamp call site was ever run against
the real image -- no corruption occurred.

Fixed by moving the header to a fixed offset from the END of the device
instead (ARTEMIS_SIG_DEVBLOCK_FROM_TOP=64), the same top-of-device region
block_subsystem.c's own meta_fence_blocks reservation (128 devblocks)
already carves out for system metadata, and where Zuse's genesis marker/
eligibility list already live -- but computed independently via
blkio_info() rather than through blk_meta_zone_*(), since that accessor
needs an already-attached, format-detected slot, which is exactly the
state pre-attach generic discovery doesn't have yet. Picked well clear of
Zuse's two tenants (devblock_from_top 0 and 1+, open-ended) so the two
subsystems' independent math can never collide.

Also reordered kernel_main.c: rng_init() now runs before the Artemis
virtio-blk block (was after) -- the genesis stamp needs rng_get_bytes()
for disk_uuid, and the original order would have failed the stamp on
every boot.

Verified live: booted amd64 against the real disk/artemis.img twice --
first boot logs "Artemis: genesis signature stamped" (confirmed blank at
the target offset beforehand via a host-side read), second boot on the
now-stamped image logs no re-stamp (idempotent, CRC/read-back verified)
-- both boots and aarch64/riscv64 (against the same now-stamped image)
all still report "Artemis: 22998 data blocks" / "PASS: persist-read"
unchanged, confirming the BAM and data pool were never touched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 07:31:21 -04:00
Robert Allan JamesandClaude Sonnet 5 29b6789860 Artemis bus-agnostic discovery: signature format + idle-loop generalization (FABRIC-3.md §XXVI follow-on)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Step 1: new artemis_sig_t header format (magic 'ARTM', sibling to
homeblocks_sig_t, distinct so a generic scan can tell Artemis's own disk
apart from an identity thumbdrive by content alone) -- artemis_sig.h/.c,
wired into Makefile.starkernel.

Step 2: sk_repl_idle()'s existing per-USB-MSC-slot attach handling (the
pattern WIREBIND already uses for identity thumbdrives) now also checks
for the ARTM signature whenever a device's home-blocks check comes back
BLANK. On a match, once Artemis's own storage-attach round-trip
(HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK) confirms success,
capsule_zuse_boot_load_root_pubkey() runs -- the same call kernel_main.c's
synchronous QEMU-only virtio-blk path already makes, now reachable
without a hardcoded PCI vendor/device scan. That function is already
idempotent (no-op once zuse_root_pubkey_known is set), so no boot
restructuring was needed despite the initial concern that deferring
Artemis discovery to the idle loop would require one.

Verified: clean build + QEMU boot to [zuse@Hera] ok> on all three
architectures, zero regression to the existing virtio-blk/Zuse-thumbdrive
attach path.

Steps 3 (genesis-stamping onto disk/artemis.img) and 4 (growable
production log-persistence region) not yet started.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 07:16:48 -04:00
Robert Allan JamesandClaude Sonnet 5 a8b16d41da Unify console prompt to [user@VM]; fix real personality-block truncation; correct §XXV's wrong lockdown conclusion (FABRIC-3.md §XXVI)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Investigating the std79 lockdown finding from FABRIC-3.md §XXV led
to a real discovery: WIREBIND births TWO VMs per identity, a console
proxy under the plain username and the actual restricted identity
under <username>~user (capsule_wirebind.c). Every test in §XXV
targeted the console proxy, which was never locked down at all.
Retested against the correct target (rajames~user): the lockdown
works exactly as designed. §XXV's "lockdown never engages" conclusion
was wrong -- corrected here, not deleted, since the mistake and how
it was caught are worth keeping (see the new feedback memory:
confirm which specific VM a name resolves to before concluding
anything, when a subsystem is known to birth more than one VM per
identity).

Two real, separate things found along the way are kept regardless
of that correction:

- capsule_runcap.c: the reserved personality devblock was read in
  full (mostly zero-padding after a short ~200-byte string) with no
  terminator, producing "WARN: block 4998 exceeds 1KB, truncating"
  on every std79-locked identity's birth, universal, since at least
  2026-09-10. Fixed by trimming to the first NUL byte actually found
  -- real, but harmless to execution (real content sat in the
  truncated block's surviving head); it mattered for capsule_id/
  content_hash being computed over padding instead of real content.

- console.h/console.c/repl.c: unified the prompt from a separately-
  computed "[VMName] (user)" into a single "[user@VMName]" line
  prefix -- exactly the ambiguity that caused the original
  misdiagnosis (the prompt showed only the WIREBIND username,
  identical whether USE had targeted the console proxy or the real
  ~user identity). Implemented as a registered callback
  (console_set_user_prefix_provider()) rather than console.c calling
  into WIREBIND/session logic directly, since console.c is a clean
  HAL module with no prior dependency on capsule-level subsystems.

Verified: clean build on all 3 architectures, zero new warnings,
identical dict_hash/capsule_hash to every prior boot this session
(console/prompt-only change). Full 9-identity messaging campaign
re-run end to end: 202s, zero faults, all 8 identities at 99/99
tokens, zero regression.

Also surfaced, not yet acted on: the full campaign's own console
tags now visibly show which VM each identity's tests actually
reached ([zuse@rajames], not [zuse@rajames~user]) -- messaging.4th's
VM-NAMES-INIT registers identities by plain username, so std79-doe.
fth's turn-attractor has been dispatching to each identity's console
proxy, not the actual locked-down identity, since the messaging
rewrite. Flagged for a deliberate decision, not investigated further.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 06:44:58 -04:00
Robert Allan JamesandClaude Sonnet 5 3c2daf50d1 Extend BIRTH/CAPSULE-BIRTH to all VMs symmetrically; flag a real std79 lockdown gap (FABRIC-3.md §XXV)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Scoping the workload-into-factorial design's placement-mode factor
led to a real architectural improvement: rather than EXEC-ing a
workload capsule into an already-running, ACL-locked identity's own
persistent dictionary (filesystem-shaped, doesn't dodge the block-
collision exposure just traced in §XXIV), a workload now runs as a
fresh ephemeral child VM, BIRTH'd per trial and reaped after --
matching the project's own stated principle of automanagement over
imposed policy. CAPSULE-BIRTH already passes vm->stadium_vm_id (who
is birthing this VM) as the new child's parent, not a hardcoded Hera
constant, confirmed by reading the C -- so a workload trial genuinely
inherits the specific identity's own lineage when that identity does
the birthing.

Which surfaced a real premise: only Hera could call BIRTH/
CAPSULE-BIRTH at all (registered only in register_mama_forth_words(),
confirmed directly, not part of the earlier §XX messaging-symmetry
fix which deliberately kept this as one of her remaining privileges).
Extended symmetrically now, agreed explicitly before touching code:

- mama_forth_words.c: BIRTH and CAPSULE-BIRTH added to
  register_child_vm_words(), matching §XX's own pattern.
- acl-std79.4th: ' BIRTH , ' CAPSULE-BIRTH , added to ACL-STD79-LIST
  (new block 4048) -- a deliberate, explicit, named exception to the
  lockdown's own "standard words only" guarantee, not a silent one.
  Symmetric registration alone can't weaken any lockdown on its own:
  ACL-LOCKDOWN-STD79 is allowlist-based, deny-by-default, so a newly
  registered word is auto-denied there unless explicitly added.

Verified: clean build on all 3 architectures, zero new warnings.
Hera's own dict_hash unchanged (expected); Hermes/Artemis show the
same new dict_hash on all 3 architectures. Live-tested against a
real attached std79-locked identity: CAPSULE-BIRTH executes
correctly (returns vm_uuid_none() for a deliberately out-of-range
capsule-id, zero fault, zero ACL denial).

Found, and explicitly stopped short of fixing, a separate pre-
existing gap while verifying the above: MSG-STATUS and MSG-K
(messaging.4th words, not on the std79 allowlist) execute for a
locked identity instead of being denied. ACL-LOCKDOWN-STD79 is
confirmed to actually run; something more specific isn't reaching
messaging.4th's dictionary entries. Root cause not traced -- needs
its own investigation into vm_core.c's dictionary-link mechanics and
whichever capsule actually loads messaging for these identities.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 22:31:52 -04:00
Robert Allan JamesandClaude Sonnet 5 71bcb72a59 Fix O(N) idle-loop messaging pump; full 3x9x3 turn-attractor campaign clean on all 3 architectures (FABRIC-3.md §XXII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
sk_repl_idle()'s messaging pump (repl.c) walked the entire live-VM
registry every idle beat (~1Hz) and dispatched a full VM-EXEC
"MSG-TICK" -- a dictionary lookup plus a 32-slot arena scan -- into
every live VM, every tick, unconditionally, forever. Fine at Tripod's
original 3-VM scale; a full 9-identity turn-attractor campaign
exposed it as a genuine wall on riscv64 specifically (its TCG makes
each dispatch cost more): the same campaign that completed in 194s/
339s on amd64/aarch64 never finished on riscv64 at 9 VMs across three
attempts, while 8 VMs there was fine in 159s.

Ruled out capacity explanations before touching anything: bumping
riscv64's QEMU RAM 1024->4096 changed nothing (reverted), and a live
STADIUM-RES@/MSG-STATUS probe with all 9 VMs attached showed no
depletion. Host memory pressure was also ruled out directly (one
background task did get OOM-killed once during the investigation,
but the identical stall reproduced again with 9.2GB free). The real
mistake was three premature kills under 4 minutes with no way to
tell "slow" from "stuck" from outside the guest -- fixed by having
run_doe_batch.sh sample the qemu process's own /proc/<pid>/stat utime
every 60s; with that signal, riscv64 at 9 VMs was unambiguously alive
(climbing utime, no hang), just disproportionately slow going from 8
VMs (159s) to 9 (600s+ and climbing).

This was never really a riscv64-only bug: an O(N) per-second walk
over the full VM population doesn't scale to the hundreds of VMs this
fleet is headed toward, on any architecture -- riscv64 just made it
visible first, at N=9, because its per-dispatch cost is highest.

Fixed by round-robin batching: the pump now dispatches to at most
SK_MSG_PUMP_BATCH (4) live VMs per idle beat via a persistent cursor
that resumes where the previous beat left off, instead of all of them
every time. Bounds both the scan and dispatch cost to O(K) regardless
of total VM count; any single VM's queue now drains roughly every
ceil(N/K) beats instead of every beat, still bounded and still
matching the pump's own existing best-effort contract. No new C
primitives, no messaging/Stadium changes.

Verified: clean build on all 3 architectures, then the full 3x9x3
campaign re-run on all 3 (not just riscv64) per the standing rule
that a defect repair requires a clean re-run everywhere before
anything counts as closed:

  amd64   198s  0 faults  99/99 tokens x8  656/656 K-conserved
  aarch64 339s  0 faults  99/99 tokens x8  659/659 K-conserved
  riscv64 178s  0 faults  99/99 tokens x8  659/659 K-conserved

riscv64 went from "never completes" to faster than aarch64, same
campaign, same seed, same identity set. No regression on amd64/
aarch64.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 19:29:14 -04:00
Robert Allan JamesandClaude Sonnet 5 7edccd2c35 Hera can't message: root-caused and fixed by symmetry (FABRIC-3.md §XX)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Hera never loaded common:messaging.4th, unlike every other VM in the
fleet. The standing belief was this was deliberate, to avoid moving
her dict_hash off baseline. Checked live instead of assumed: loading
messaging.4th into her dictionary silently dropped every colon-
definition referencing one of 8 STADIUM-* primitives that
register_child_vm_words() gives every other VM but
register_mama_forth_words() never gave her -- a missing-primitive gap,
not a designed privilege boundary. Confirmed mama_word_birth is
genuinely VM-agnostic and SPAWN-EVENT is an unwired placeholder before
proposing the fix.

Fix: register the same 8 STADIUM-* primitives for Hera, load
messaging.4th from init.4th the same way Hermes/Artemis/console/mint
already do, and give her own idle-loop context a direct MSG-TICK call
(not VM-EXEC, which would hit the same reentrancy class the existing
per-other-VM pump loop already guards against) so her own queued
messages actually drain. Her dictionary is now a proper superset of
every child VM's, plus her remaining extra privileges -- not
structurally different from any other VM, just additionally
privileged.

Verified: dict_hash identical across amd64/aarch64/riscv64
(0xc8f4b09e36f4fc4a), Hermes/Artemis dict_hashes unchanged and still
cross-arch identical, all three boot clean to zuse)ok> with zero
UNKNOWN WORD faults, mkcapsule --lint clean.

Unblocks rewriting the turn-attractor (FABRIC-3.md §XIX) to coordinate
via real MSG-SEND/MSG-TICK instead of blocking VM-EXEC.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 16:05:11 -04:00
Robert Allan JamesandClaude Sonnet 5 5a9425b91b Add VM-HEAT primitive: groundwork for a compudynamic turn-attractor (FABRIC-3.md §XIX)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Prompted by reading the std79-doe K report: checked whether the campaign's
"real cross-VM dispatch load" was actually concurrent or strictly
serialized. It's serialized at two levels -- VM-EXEC's vm_interpret(target,
...) is a direct synchronous C call (Hera fully blocked until it returns),
and even the background physics tick (vm_tick(), vm_runtime.c) is driven
by each VM's own execution loop, so idle identities accrue zero ticks
between their own turns. K's perfect conservation (FABRIC-3.md §XVIII)
verifies sequential per-VM accounting correctness, not concurrent-access
safety, since there was never concurrent access to test.

Agreed direction: fix this without a scheduler, by reusing the same
least-dense-candidate judgment stadium_admit() already trusts for
eviction, applied to "whose turn is next" instead of "who gets evicted" --
a fleet-level turn-attractor giving the next turn to whichever live
identity currently has the lowest execution_heat_q48, no fixed round-robin,
no priorities, no preemption. Lives beside Stadium in
capsule_vm_physics.c (already the fleet-level consumer of Stadium
primitives, e.g. the K mechanism itself), not inside stadium.c ("the
floor" -- residency/eviction, a different concern from turn order) and
not a new subsystem.

This pass lands only the primitive the mechanism needs: VM-HEAT
( c-addr u -- heat-q48 ), pushing a named VM's current
execution_heat_q48 via vm_physics_heat_of() -- previously C-internal
only (doe_log_heat_by_name(), doe_log.c), never exposed to FORTH. Silent
0 on an unknown/dead name (no print/error), since a turn-attractor
scanning many candidates every turn shouldn't have to filter console
noise for names that simply aren't live. Registered everywhere
VM-EXEC/VM-CALL already are. Builds clean on all three architectures;
live-tested on amd64: Hera -> 65452, Hermes -> 43, unknown name -> 0, no
faults.

The turn-attractor loop itself (std79-doe.fth's trial ordering) is not
yet built -- next step, not done here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 14:53:06 -04:00
Robert Allan JamesandClaude Sonnet 5 66ba21adb4 Log fleet_k_q48/fleet_conserved; K holds exactly, 775/775 ticks (FABRIC-3.md §XVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
doe_log.c's per-heartbeat-tick CSV gains two columns: fleet_k_q48
(vm_physics_fleet_heat_sum() over ALL live VMs -- the genuine fleet-wide
conservation invariant K, not reconstructable from the 3 named-Tripod-
member heat columns already logged, which omit every identity VM's own
heat) and fleet_conserved (vm_physics_conserved() as 0/1). Requested
explicitly after the first heartbeat-telemetry analysis pass
(analysis-20260912/) omitted K entirely.

Kernel rebuilt on all three architectures, full 3x9x3 campaign rerun
(results-20260912-with-k/). K = 1.0000000000 (Q48.16 raw 65536) on every
one of 775 heartbeat-tick observations, sd(K) = 0, 100% fleet_conserved,
across amd64/aarch64/riscv64, nine identities, three replicates -- zero
deviation. Also a free regression check on both recent Stadium fixes
(§XVI/§XVII): neither disturbed the reservoir-transfer accounting K
depends on.

Found and fixed a tooling wrinkle along the way: fleet_conserved, being
the CSV row's very last field with nothing after it to bound a regex
match, can have a resumed trial digit merge into it with zero separator
on the wire -- combine.py now derives it from fleet_k_q48 directly (same
epsilon vm_physics_conserved() uses) instead of trusting the raw field.
fleet_k_q48 itself is unaffected either way.

Full analysis, discussion, and light/dark SVG->PDF figures written up as
a proper LaTeX report (report-20260912/report/std79_doe_report.pdf),
following experiments/bare_metal/analysis/report/bare_metal_doe_report.tex's
established style -- supersedes analysis-20260912/'s markdown-only first
pass as the primary deliverable for this dataset (kept, not discarded).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 13:25:09 -04:00
Robert Allan JamesandClaude Sonnet 5 e51a8d229e Fix stadium_grant_quota() donor floor; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
capsule_birth.c hardcoded every new VM's initial Stadium quota grant to split
from Hera specifically. Since a grant always halves whatever the donor
currently has, Hera's own free list converges toward empty after a bounded
number of grants — independent of whether the Stadium as a whole still had
spare capacity, since VMs she'd granted to earlier typically still held
nearly all of their own share untouched. Past that point every subsequent
VM birth's Stadium grant would be silently refused (soft-failed, non-fatal
by existing design), even with plenty of capacity sitting idle elsewhere.

Fixed by adding an O(1)-maintained free_count to StadiumVMQuota (incremented
in stadium_evict(), decremented at both of stadium_admit()'s free-list-pop
sites, set/adjusted in stadium_grant_quota()'s own split — this also let
grant_quota drop its old O(free-list length) counting walk in favor of an
O(1) read) and stadium_best_donor(), an O(live VM count) scan over quota
slots returning whichever in-use VM currently has the most free cells.
capsule_birth.c's birth path now splits from that VM instead of
unconditionally vm_uuid_hera().

Verified with another full rerun of the 3x9x3 std79 DoE campaign from
scratch — same discipline as the prior Stadium fix (any defect repair
reruns the whole DoE from the top) — one continuous boot per architecture,
all 9 identities simultaneously live throughout. 81/81 trials correct, 0
mismatches, DOE-RUN header sequence md5-identical to every prior run.
aarch64 ~280s total (vs ~290s for the O(ncells)-scan fix alone — confirms
no regression). Both known Stadium defects are now closed together on one
clean campaign rerun.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 17:20:07 -04:00
Robert Allan JamesandClaude Sonnet 5 9eff122090 Fix aarch64 Stadium/COOL O(ncells) scan; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVI)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Root cause of the 90+ minute aarch64 VM-birth stall found in §XV: stadium_admit()'s
eviction-fallback scan iterated the entire stadium_ncells array filtered by owner,
not the calling VM's own resident cells as its own doc comment claimed. Combined
with stadium_grant_quota() always splitting from Hera's shrinking free list and
stadium_word_dispatch() calling stadium_admit() per distinct word a VM's capsule
executes, this compounded into a real O(n) blowup — catastrophic specifically on
aarch64 because its -m 4096 (vs 1024 on amd64/riscv64) inflates the kmalloc heap
kmalloc_init() bisects down to, which inflates stadium_ncells 4x (335,544 vs 83,886
cells, measured from boot logs).

Fixed by threading a real per-VM doubly-linked resident-cell list
(StadiumVMQuota.resident_head + stadium_resident_next[]/stadium_resident_prev[])
so the fallback scan is bounded by that VM's own resident count, not the global
cell array size.

Verified with a full rerun of the 3x9x3 std79 DoE campaign from scratch: one
continuous boot per architecture, all 9 identities simultaneously live throughout
(the 3-boot aarch64/riscv64 batching workaround is no longer needed). 81/81 trials
correct, 0 mismatches, DOE-RUN header sequence md5-identical across all three raw
logs. Identity 04's attach on aarch64, which stalled 90+ minutes before, now
completes in ~34s; full boot-to-DoE-complete in ~290s.

Corrects an earlier misreading (carried into §XV, std79-doe.fth's comments, and
the project memory note) that described the symptom as a runaway "335,000+ cycles"
dispatch counter — those were cell array indices, not an event count.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 16:37:44 -04:00
Robert Allan JamesandClaude Sonnet 5 70db955ac9 Fix M/MOD: hand-rolled 128/64 bit-serial division (FABRIC-3.md §XII.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
mixed_math_word_m_slash_mod() had the identical bug class already fixed
in M* (commit 9a09949): it reconstructed the dividend from
`(dhigh << 32) | (dlow & 0xFFFFFFFF)`, the same wrong "32-bit halves of a
64-bit value" assumption. Latent for small inputs (fits in 32 bits, so the
reconstruction coincidentally worked), confirmed genuinely broken for a
true wide double -- feeding M*'s own correct 10^24 output into it gave a
quotient/remainder wrong by many orders of magnitude and the wrong sign.

Unlike M*, this one can't just switch to __int128 -- __int128's own `/`/`%`
need libgcc's __udivti3/__umodti3 for 128-bit division, unavailable in
this freestanding, -nostdlib build (the exact constraint
src/starkernel/arch/amd64/timer.c's own doc comment already flagged:
__int128 multiply/shift-by-constant/compare/subtract all compile clean,
only division doesn't). Fixed instead with a hand-rolled unsigned
128-by-64-bit bit-serial (restoring) long division -- 128 iterations of
shift-by-1/compare/subtract only, all in the safe set. Operates on
magnitudes via unsigned negation from 0 (well-defined even for the
extreme negative edge); signs reapplied afterward matching the same C99
truncating-toward-zero convention the previous, narrower implementation
already used, unchanged.

Verified: rebuilt all three architectures, confirmed clean link with no
__udivti3/__umodti3 undefined-symbol errors. Booted and tested all three,
identical results: 1000000000000 1000000000000 M* SWAP 1000000 M/MOD ->
1000000000000000000 remainder 0 (exact division, the case that was wrong
by orders of magnitude before); three sign-combination cases all correct
(-+, +-, --), confirming sign handling survived the magnitude-only
rewrite. T15 (original small-input case) unaffected. Full 24-case
exerciser reran clean on all three, no regressions.

All bugs found by the std79 exerciser campaign, including this one found
while fixing another, are now closed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 10:28:34 -04:00
Robert Allan JamesandClaude Sonnet 5 9a09949c69 Fix M*: use __int128 for a genuine 128-bit double-cell product (FABRIC-3.md §XII.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
mixed_math_word_m_star() confused "double" (two full cell_t-width cells,
128 bits total on this 64-bit build -- what D+/D-/D./etc. all actually
expect) with "the low/high 32-bit halves of a single 64-bit product" --
code clearly written assuming cell_t is 32-bit. It computed an ordinary
64-bit `long long` product (already wrong for any true product exceeding
64 bits, since long long is the same width as one cell here) and split
that into 32-bit halves via `result & 0xFFFFFFFF` / `result >> 32`. For a
small negative product like -56088, this produced a positive, zero-
extended low cell paired with a correctly-looking dhigh=-1 -- D.'s
overflow check (correctly) rejected the resulting malformed double, on
every architecture, every time (this bug was never architecture-specific,
unlike the D+/D-/DNEGATE/d_compare family already fixed in bea8d74/
1a716c8).

Fixed via __int128 for a genuine 64x64->128-bit multiply, no truncation --
an already-established safe pattern in this codebase for exactly this
operation (src/starkernel/crypto/fe25519.c/scalar25519.c already use it;
src/starkernel/arch/amd64/timer.c's own doc comment confirms __int128
multiply/shift-by-constant compile cleanly with zero undefined symbols on
all three target toolchains -- only division needs unavailable libgcc
support, not used here).

Verified: rebuilt and booted all three architectures clean. T13/T14 both
correct everywhere (56088/-56088). Manual cases confirm genuine 128-bit
precision: -1 -1 M* -> 1; 1000000000000 1000000000000 M* -> dhigh=54210
dlow=2003764205206896640 (10^12 x 10^12 = 10^24, correctly exceeding 64
bits). Full 24-case exerciser reran clean on all three, no regressions.

Found while verifying, NOT fixed (report only, out of scope of this
request): mixed_math_word_m_slash_mod() (M/MOD, the very next function in
the same file) has the identical bug class -- confirmed genuinely broken
for a true wide double (feeding this fix's own correct large-magnitude M*
output into M/MOD produces a result wrong by many orders of magnitude and
the wrong sign). Flagged in FABRIC-3.md for a future fix request.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 08:23:51 -04:00
Robert Allan JamesandClaude Sonnet 5 1a716c8048 Fix d_compare: use ucell_t instead of unsigned long (FABRIC-3.md §XII.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
d_compare() (double_words.c), the static helper backing DMAX/DMIN/D</D=,
had the same bare-unsigned-long pattern already fixed in D+/D-/DNEGATE
(commit bea8d74) -- 32-bit unsigned long on this aarch64 target silently
truncating the low-cell comparison. Switched to ucell_t.

Verified with four cases designed specifically to expose the old
truncation: two doubles sharing the same high cell but with low cells
differing only in bits 32-63 (invisible to a 32-bit-truncated compare,
e.g. 2^32 vs 0):
  4294967296 0 0 0 D= .            -> 0  (correctly not equal)
  4294967296 0 0 0 DMIN SWAP D.    -> 0  (correctly picks the smaller)
  0 0 4294967296 0 D< .            -> -1 (correct)
  4294967296 0 0 0 D< .            -> 0  (correct)
All four pass identically on amd64, aarch64, and riscv64. These cases
were never run before this fix -- the exerciser campaign never touched
DMAX/DMIN/D</D= at all, so this is the first real evidence d_compare()
was ever exercised on any architecture. Full 24-case exerciser also
reran clean on all three architectures (no regression).

D2*/D2/ use unsigned long long (C-standard-guaranteed >=64-bit) and were
never part of this bug class. All four affected words (D+, D-, DNEGATE,
d_compare) are now fixed; M*'s separate universal DOUBLE-OVERFLOW bug
remains open (unrelated defect, out of scope here).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 07:44:14 -04:00
Robert Allan JamesandClaude Sonnet 5 bea8d7436a Fix D+/D-/DNEGATE: use ucell_t instead of unsigned long (FABRIC-3.md §XII.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
double_word_d_plus(), double_word_d_minus(), and double_word_dnegate()
(double_words.c) all cast through plain `unsigned long` for their carry/
borrow-detection arithmetic. On this aarch64 bare-metal cross-compile
target, unsigned long is 32-bit (confirmed: sizeof(unsigned long)==4) --
amd64 and riscv64 both happen to have a 64-bit long, so the identical code
only broke on aarch64. The low-cell arithmetic silently truncated to 32
bits, then widened back to cell_t via ordinary (non-sign-extending)
conversion, producing a wrong result whenever the true 64-bit result was
negative -- D. then correctly, faithfully reported DOUBLE-OVERFLOW on the
resulting malformed double.

vm.h already defines ucell_t for exactly this: same conditional as cell_t,
guaranteed width-matched on every target. print_number_formatted()
(format_words.c) already used it correctly; these three words didn't.
Switched all three to ucell_t -- a one-word-class fix, no logic change.

Verified: rebuilt and booted all three architectures clean. T19 (D+) on
aarch64 now correctly prints -2, matching amd64/riscv64; T20 (DNEGATE)
unaffected everywhere. Additional manual cases beyond the original
exerciser, run live on aarch64 to specifically exercise the
>32-bit-magnitude path the old bug depended on: D- (-5-3=-8), DNEGATE on
2^33 (8589934592 -> -8589934592), D+ crossing the same boundary
(3+8589934592=8589934595) -- all correct.

d_compare() (backing DMAX/DMIN/D</D=) has the identical latent pattern but
is out of scope for this fix (not named in the request, never exercised by
the campaign) -- left open, flagged in FABRIC-3.md. M*'s separate,
universal-across-all-three-architectures DOUBLE-OVERFLOW bug is also
untouched -- unrelated defect, not part of this fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 07:37:57 -04:00
Robert Allan JamesandClaude Sonnet 5 9142dda2d6 Fix use-after-free in sk_repl_idle()'s idle-tick VM resolution (FABRIC-3.md §XIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Root-caused a heap-corruption bug that reliably failed WIREBIND identity
attach on the third attach/detach cycle in one boot. sk_console_readline()
and sk_console_getkey() captured `active_vm` once from their caller and
kept passing that same (possibly long-stale) pointer to sk_repl_idle() on
every idle tick serviced while blocked waiting for input. If the VM it
pointed at was killed (WIREBIND detach) mid-block, the existing bailout
only checked a generic "is anyone attached" boolean -- masked as soon as a
different identity attached next -- so blk_vm_flush_all() kept writing
into a freed VM struct sitting on kmalloc's own free list, corrupting the
free list's linked-list metadata itself.

Both idle branches now re-resolve the live active VM fresh from
g_repl_active_vm on every tick, matching the dispatch-side fix already
made for the sibling bug in §XII.3.

Verified: rebuilt amd64, reran the exact three-cycle repro that reliably
corrupted the heap before the fix -- free-list census stayed stable
through the same idle window that previously collapsed to zero. All three
architectures (amd64/aarch64/riscv64) boot clean to the zuse)ok> prompt.

kmalloc_debug_census()/kmalloc_debug_census_bytes() kept as permanent
diagnostic infrastructure; every other temporary probe added during the
investigation was reverted.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-10 20:30:41 -04:00
Robert Allan JamesandClaude Sonnet 5 70421bdd43 Fix real WIREBIND crash: stale active-VM pointer dispatched after blocking read (FABRIC-3.md §XII.3)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
The interpreter_enabled guard added in the previous commit (662ef44) was a
real but incomplete fix -- re-running the exact repro against it still
panicked (this time as a raw #PF page fault), proving something deeper
was wrong.

Root cause, found via targeted console_puts probes (not GDB --
starkernel_kernel.elf's symbols don't correspond to the actual running
starkernel_loader.efi binary for this monolithic build, same gotcha
already on record from the 2026-08-18 aarch64 investigation):

sk_repl_run()'s main loop captures `active` once, before calling
sk_console_readline(), which then blocks for the next full line. If the
identity `active` points at is killed while that read is still blocked,
the bailout meant to catch this (sk_console_identity_present()) only
checks a generic "is anyone attached" boolean, not "is the specific
identity active belonged to still attached" -- a fast detach of one
identity followed by attach of a different one never produces an
observable gap in that boolean, so the bailout never fires. The stale
`active`, now pointing at freed memory, gets dispatched into.

Fix: re-resolve `active` fresh from g_repl_active_vm immediately before
dispatch, right after sk_console_readline() returns. One line, no
registry lookup, no dereference of the stale pointer -- closes the race
regardless of whether the bailout catches it first.

Verified: rebuilt amd64 clean, reproduced the exact same attach/USE/
detach/attach/USE sequence against the fixed build -- clean switch, no
fault, exerciser runs correctly afterward.

FABRIC-3.md §XII.2 also corrected to stop claiming the interpreter_
enabled guard alone closed the crash -- it didn't, per the above.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-10 17:19:01 -04:00
Robert Allan JamesandClaude Sonnet 5 662ef44e59 Fix EXEC/LOAD block-persistence gap and WIREBIND/USE interpreter-race panic (FABRIC-3.md §XII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Found live while building a cross-ISA FORTH-79 dictionary exerciser:

- capsule_exec_init() zeroed a capsule's block content immediately after
  running it, so LOAD (a genuine FORTH-79 standard word, ACL-allowed even
  for locked identities) could never actually read back what EXEC had
  just written. Removed the clear from capsule_exec_init(); block content
  now persists like any other Standard BLOCK/BUFFER/UPDATE write.
  kernel_main.c's own explicit post-birth clear of Mama's init.4th range
  is untouched.

- USE could redirect the console to a WIREBIND identity's VM before that
  VM's vm_enable_interpreter() step of its own birth sequence had run,
  causing the next typed line to hit vm_assert_interpreter_enabled() and
  panic the entire machine -- not the per-session-recoverable ACL-fault
  path a redirected VM otherwise gets. USE now checks interpreter_enabled
  first and refuses with a retry message instead.

Also includes the amd64/aarch64/riscv64 acceptance boot logs and DoE CSVs
from this session.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-10 16:07:38 -04:00
Robert Allan JamesandClaude Sonnet 5 b301317902 xHCI: drive Port Reset on port reuse; WIREBIND: kill the console VM too (FABRIC-3.md §XI.5)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Two independent bugs that together caused a reliable hotplug wedge:
reusing an xHCI port for a second identity right after an unclean
detach of a first would leave no further hotplug events reaching the
guest at all.

Bug 1 (xhci.c/xhci_driver.h): the xHCI driver never drove PORTSC.PR --
a known, named gap since Milestone 2e (the code's own comment flagged
it, PORTSC_PR/PRC were defined but never referenced). A port's first
connect each boot reads PED already set, so skipping the reset
happened to work; a second device on the same port after a prior
disconnect reads PED clear, and Address Device reliably failed without
an explicit reset cycle. New XHCI_CONN_AWAIT_PORT_RESET state drives
PR and waits for PED to read set before proceeding to Enable Slot.

Bug 2 (capsule_wirebind.c): capsule_wirebind_eject()/unclean_detach()
compared g_repl_active_vm against the *user* VM's pointer
(g_wirebind_attached_vm_id tracks that one, not the console VM) --
never equal, since USE/g_repl_active_vm always points at the console
VM. The guard never fired and the console VM was never killed at all,
only orphaned -- paired to a dead user VM but still the REPL's active
session. New wirebind_teardown_console() helper resolves and tears
down the console VM by its own tracked bare username.

Verified live on amd64: the exact repro (identity 01 on port 2,
unclean detach, identity 02 on the same port immediately after) --
previously wedged with "xhci: address device failed" and no further
hotplug activity; now attaches cleanly and fast, both VMs' KILL
messages appear, console is immediately interactive on the new
identity. Three-arch clean qemu acceptance passed, all clean on the
first attempt.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 22:11:09 -04:00
Robert Allan JamesandClaude Sonnet 5 d6661b5eed Scope VM fault halt to the faulting identity's own session (FABRIC-3.md §XI.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
A standalone WIREBIND identity (no Zuse, USE'd in directly) hitting an
ACL-denied word halted the entire machine -- Hera, Hermes, Artemis, all
of it -- instead of just that identity's own session. sk_fault_handler()
was being called unconditionally on whichever VM's ->error was set, with
no distinction between Hera's own root session (where "no fallthrough
surface" is the correct, deliberate fail-closed behavior) and a
USE'd-in guest identity (which should recover and resume at its own
prompt instead of taking the fleet down with it).

Both call sites (sk_repl_step, sk_repl_run) now compare the faulting VM
against Hera before deciding: Hera's own session still halts by design;
any other VM prints a recovery message, clears its fault state, and
continues.

Also: mint identities 01-06 with the same FORTH-79/83 restricted
personality identity 00 already had, verified via the fixed fault
scoping above (which this verification pass surfaced).

Verified live on amd64 (both the Hera-halts and identity-recovers
branches); three-arch clean qemu acceptance passed (riscv64's first
attempt hit an unrelated virtio_blk I/O timeout hang, a known QEMU/TCG
flake -- a clean retry booted normally).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 20:41:51 -04:00
Robert Allan JamesandClaude Sonnet 5 1839a2b0c3 WIREBIND cert verification: load Zuse's root pubkey independently of her live session
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Root cause of the remaining "identity attach doesn't complete when Zuse
never attaches this boot" issue: capsule_wirebind_verify_cert() gated on
mama_vm->zuse_cert_installed, which is only ever set when Zuse's own
drive attaches and authenticates this specific boot
(capsule_zuse_boot_try_attach() -> install_and_activate() ->
vm_zuse_cert_install()). Without her, any other identity's WIREBIND cert
verification silently refused -- correctly, by the old design, but that
design conflated two genuinely different things: "can mint new
identities" (needs Zuse's live private seed, a real privileged
operation) and "can verify an existing identity's cert" (needs nothing
but her already-public key).

That public key was already being persisted independently of her live
session: zuse_genesis_marker_t (zuse_genesis_marker.h) stores it in the
kernel's own top-of-device metadata fence (Artemis's resident storage),
written once at genesis, specifically *not* alongside her private seed
(which stays only on her own removable thumbdrive) -- the type's own doc
comment says as much. It just wasn't being loaded for anything but
confirming which drive is genuinely hers.

Fix: a new capsule_zuse_boot_load_root_pubkey() (capsule_zuse_boot.c)
reads that marker and populates two new VM fields, zuse_root_pubkey_known
/ zuse_root_pubkey (vm.h) -- deliberately separate from
zuse_cert_installed/zuse_cert_seed/zuse_cert_pubkey, which stay
untouched and still gate MINT exactly as before. Called once from
kernel_main.c as soon as Artemis's own storage attaches, unconditionally,
independent of whether Zuse's own drive is ever attached this boot.
capsule_wirebind_verify_cert()/capsule_wirebind_try_attach() now check
zuse_root_pubkey_known instead of zuse_cert_installed.

One identity's attach must not depend on another identity's live
presence -- each identity stands on its own once the fleet's root of
trust has been established once, ever.

Verified live, amd64: identity 00 (disk/thumbdrives/00-thumb-ident.img)
now attaches and completes WIREBIND in 19 seconds with Zuse's own drive
never attached this boot at all (previously: unbounded, many real
minutes or effectively never, before today's other fixes; still slow/
stuck after those, stuck specifically on this silent refusal). Zuse's
own attach flow re-verified unaffected (regression check, amd64).
Three-arch clean qemu acceptance (amd64/aarch64/riscv64) passed with
this change included.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 14:30:14 -04:00
Robert Allan JamesandClaude Sonnet 5 1a263555e2 Fix pathological migration scan that stalled WIREBIND identity attach
Root cause (found by a fresh subagent after an extended live-debugging
investigation into "identity 00 attaches slowly/stalls when Zuse never
attached first this boot"): blk_migration_idle_check() was generalized
earlier today to walk every attached device slot uniformly instead of
hardcoding first_disk_slot() (Artemis's own disk). But its per-slot scan
can only early-exit once it finds a devblock that is BOTH "hot" (claimed
and worn) AND "free" -- and a just-attached, never-claimed USB identity
drive can never satisfy the "hot" half by design (claiming only ever
happens via blk_firsttouch_claim(), which only ever targets
first_disk_slot()). So the scan ran to completion -- the drive's entire
~16,000 devblocks, mostly cache misses over slow emulated USB/BOT --
every single idle tick, forever, blocking sk_repl_idle() (and therefore
the console and the storage-attach message round-trip) each time.

Fix, in src/block_subsystem.c: a new has_ever_claimed flag on
blk_dev_slot_t (set in blk_set_meta(), the single choke point every
BLK_FLAG_CLAIMED transition passes through) skips the scan entirely,
O(1), for any slot nothing has ever claimed -- the common case for a
freshly-attached drive. A new migration_scan_lbn resume cursor bounds
*any* slot's per-tick cost to MIGRATION_SCAN_BUDGET (256) devblocks
examined, picking up where the previous tick left off instead of
restarting from start_lbn every time -- restores this function's own
documented "coarse cadence, cheap early-exit" design intent for every
device, not just the one it used to hardcode.

Also along the way (kept, all real improvements, verified live):
- src/starkernel/usb/xhci.c: xhci_wait_bit()/xhci_bot_wait_for_idle()
  had zero yield hints in their MMIO-polling loops; added arch_relax()
  to both (matches virtio_blk.c below) -- a tight loop of nothing but
  MMIO reads can starve TCG's own host-side timer injection under QEMU.
- src/starkernel/virtio/virtio_blk.c: vblk_io()'s spin bound was 33M
  iterations with zero logging on timeout; a single real (still not
  fully root-caused) timeout cost 31+ minutes of CPU before this was
  caught. Reduced to 1M and added a log line naming the failing sector,
  turning a silent, effectively-unbounded stall into a fast, loud
  failure -- callers already tolerate BLKIO_EIO.
- src/starkernel/repl.c: blk_migration_idle_check() deferred for any
  idle tick where a storage-attach message round-trip is still pending,
  to keep the two block-subsystem-touching paths from interleaving; the
  existing MSG-TICK pump now checks the target VM's own dictionary for
  MSG-TICK (vm_find_word + acl_allow) before VM-EXECing into it, instead
  of spamming "VM-EXEC: ERROR" every tick for a VM that doesn't have --
  or isn't allowed -- the word (e.g. a FORTH-79/83-locked-down identity);
  fixed a real -Wunused-variable build break in the EMERGENCY_CONSOLE_
  ENABLED=1 path (unused today, but needed live to reproduce this bug
  with no identity attached at all).
- src/starkernel/capsule/capsule_mint.c: dropped the dead
  S" common:messaging.4th" EXEC / MSG-CD-INIT lines from
  MINT_RESTRICTED_PERSONALITY -- a FORTH-79/83-locked-down identity has
  no legitimate use for a messaging vocabulary it can never call.

Status: the pathological CPU-climbing scan is confirmed fixed (verified
live: CPU stays flat across an extended run instead of climbing without
bound). The WIREBIND storage-attach message round-trip still does not
complete promptly in the "Zuse never attached, other identity attaches
first" scenario -- a separate, still-open issue in the message-delivery
path itself, not the scan. Tracked as follow-on work.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 12:08:03 -04:00
Robert Allan JamesandClaude Sonnet 5 63b8b3bc29 Migrate Hera->Artemis storage-attach to a real message round-trip
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-amd64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Hera still polls xHCI and sig-checks attached drives, but the storage
registration step (blk_subsys_attach_device(), now wrapped as the
BLK-ATTACH primitive) moves to Artemis's own dictionary, reached via
HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK (VM-EXEC, since Hera can't load her
own messaging.4th -- see the doc comment in repl.c). Identity birth
(Zuse genesis / WIREBIND) is deferred until the ack confirms storage
actually succeeded, instead of running synchronously underneath a
storage call that might fail ("wait for ack, safer for identity data").

Caught and fixed a real bug live during acceptance testing: Artemis's
ACK-APPEND-NUM fed a single-cell value into <# #S #> (which expects a
double-cell pair), causing a stack underflow the first time
HERA-BLK-ATTACH-REQ ran. Fixed with the same `0 SWAP` convention every
other numeric-append helper in this codebase already uses.

Verified booting clean to (zuse) ok> with no VM-EXEC errors on all
three architectures (amd64/aarch64/riscv64).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 21:22:04 -04:00
Robert Allan JamesandClaude Sonnet 5 f4ded3e1a8 MINT: add a FORTH-79/83-standard-words-only lockdown personality
Captain Bob, 2026-09-07: "starting with that 00 user we created, we're
going to give access only to FORTH 79 and 83 standard words. everything
else is locked down."

New capsules/acl-std79.4th (blocks 4023-4047): walks a VM's own
dictionary (>LINK/LINK> traversal, same as ACL-INIT-PRIMITIVES/WORDS
already use) and permanently denies+pins every word not on an explicit
FORTH-79/83 allowlist, extracted from the real registered word set
(stack_words.c through control_words.c), not recited from memory.
Deliberately excludes, beyond plain non-standard words: BYE (100% ACL
bypass to the emergency console -- "needs more discussion, exclude for
now"), COLD/WARM/REBOOT/SAVE-SYSTEM (system lifecycle), the block/screen
editor L/S/SHOW/EDIT/UPDATE/SAVE-BUFFERS (lets a session rewrite
persistent block/capsule content, defeating the lockdown even though
nominally standard), BLK-ACL-*/BLK-OWNER@ (StarForth-specific), and
FORGET/FENCE (flagged as an unrestricted superpower word, 2026-09-03
audit). Keeps WORDS/VLIST/SEE (introspection only -- ACL is enforced
per-target-word at execution time regardless of how an XT was
obtained) and the parenthesized control-flow runtime primitives
((BRANCH) etc. -- IF/DO/LOOP compile calls to these; denying them
breaks ordinary control flow, not security).

MintPersonality enum (capsule_mint.h) lets capsule_mint_identity()
select which personality-source template gets written to a new
identity's devblock -- MINT_PERSONALITY_DEFAULT (unchanged) or
MINT_PERSONALITY_STD79_LOCKDOWN (EXECs acl-std79.4th then
ACL-LOCKDOWN-STD79 as the VM's own last bootstrap step). The actual
restriction logic stays entirely in FORTH per .claude/CLAUDE.md's
Word-Level ACL System rules ("ACL policy belongs in ACL.4th, never in
C") -- capsule_mint.c only picks which few-line bootstrap stub to
write. MINT's own stack signature gains a trailing restrict? flag;
capsule_zuse_boot.c's genesis mint (Zuse herself) explicitly passes
MINT_PERSONALITY_DEFAULT -- the superuser is never restricted.

Two real bugs found and fixed live during testing, both the same class
of self-referential fault: ACL-LOCKDOWN-STD79's own walk loop calls
ACL-STD79-ALLOWED?/ACL-STD79-LIST/ACL-ALLOW!/ACL-PIN on every single
iteration to do its job -- none of those are FORTH-79/83 standard
words, so the walk was denying its own load-bearing infrastructure
partway through and then faulting the next time it tried to call it
("VM fault -- emergency console disabled; halting", reproduced twice
live). Fixed by explicitly protecting all four in the allowlist
(block 4047) -- they must stay allowed for the walk to finish, not
because they belong on a "standard words" list.

Verified live end-to-end: minted a throwaway test identity with the
restrict? flag, confirmed her WIREBIND birth completes cleanly (no
faults, no shadow conflicts) on a single real attach, then USE'd into
her VM and confirmed standard arithmetic and user-defined words work
(1 2 + . -> 3; : X 5 5 * . ; X -> 25) while KILL is entirely unknown to
her dictionary and VM-EXEC is denied. One real, non-fatal side effect
found and left as-is (not asked to fix): the fleet's inter-VM messaging
pump (MSG-ARENA) is also denied by the lockdown, logging a harmless
per-idle-tick warning -- a fully locked-down VM doesn't participate in
message routing.

Not yet applied to the real identity 00 -- this commit is the
mechanism, verified against a disposable test identity only.

Three-arch clean qemu acceptance (single Zuse device, standard
regression case) passed on amd64, aarch64, and riscv64.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 15:41:56 -04:00
Robert Allan JamesandClaude Sonnet 5 56e19a00f9 Route VM-dictionary sf_malloc/sf_free through the kernel's kmalloc heap
Captain Bob's call after seeing the identity-heap-capacity findings
(FABRIC-3.md §X.4): the number of concurrently-running VMs is not known
in advance, and once this is a complete operating system the heap should
be able to use whatever memory is actually available -- not a hardcoded
compile-time ceiling. This was already half-built and just not wired up.

src/starkernel/vm/alloc_kernel.c previously implemented sf_malloc()/
sf_free() (platform_alloc.h's allocator abstraction -- what
vm_create_word() calls for every VM's word dictionary) as its own
isolated static 4MB arena: first-fit free list, no splitting or
coalescing. That's exactly the allocator that topped out around 6
concurrent WIREBIND-born identities, failing from fragmentation before
true capacity exhaustion (§X.4's own measurements).

Sitting right next to it, unused for this purpose: src/starkernel/memory/
kmalloc.c, the kernel's general heap. Already initialized at boot (M6,
kernel_main.c, well before any VM is ever born), reserved from real
PMM-tracked physical memory rather than a fixed array, defaults to a
2 GiB floor explicitly sized "for 256+ baby VMs" per its own comment,
overridable via the --heap= boot flag, and its free list actually
coalesces neighboring blocks on every free.

Change: alloc_kernel.c's sf_malloc()/sf_free() now delegate to
kmalloc_aligned()/kfree() instead of managing a separate arena.
sf_alloc_init() becomes a no-op (kmalloc is already initialized by the
time any VM allocation can happen, and "resetting" a heap now shared by
every kernel subsystem would be actively wrong -- confirmed no external
caller depended on its old reset semantics). sf_alloc_get_stats() reads
kmalloc_get_stats() fresh rather than shadowing byte counts locally;
alloc_count/free_count (which kmalloc.c doesn't track) stay as simple
local counters. sf_calloc()/sf_realloc() are otherwise unchanged. Kernel-
only: the hosted (non-kernel) StarForth build keeps its own separate
alloc_host.c implementation, untouched.

Verified live: replaying the exact hotplug sequence that previously
topped out at 6 identities (Zuse + 8 identities, one at a time via QMP
device_add) now succeeds for all 9, where identity 05 specifically used
to fail. Three-arch clean qemu acceptance (single Zuse device, the
standard regression case) passed on amd64, aarch64, and riscv64 -- one
aarch64 attempt hit an unrelated, already-documented one-off QEMU hiccup
(empty log, boot never progressed past firmware) and passed cleanly on
retry with no rebuild.

Not addressed here: the underlying free-list itself is still first-fit
without splitting (only coalescing changed, inherited from kmalloc.c);
per-VM dictionary sizing (shrinking what each WIREBIND VM's word set
actually needs) is a separate, still-open lever from FABRIC-3.md §X.4's
open architecture question.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 12:46:14 -04:00
Robert Allan JamesandClaude Sonnet 5 d8a195b8d8 WIREBIND: kill the orphaned console VM when the user VM birth fails
Found live during identity-heap-capacity testing (2026-09-07,
hotplugging Zuse + 8 identities one at a time and measuring the kernel
heap arena via a temporary allocator-stats probe, since reverted): every
WIREBIND identity attach births two VMs in sequence -- a "console" VM,
then the real "user" VM. When the second birth failed (arena
fragmentation under concurrent VM load, a separate, not-yet-fixed
capacity issue), capsule_wirebind_try_attach() logged the failure and
returned, but the console VM that had *already succeeded* was never
torn down. It stays live and registered under the identity's username,
consuming its own ~228KB of the fixed 4MB kernel heap arena forever --
nothing ever points a real user at it, since WIREBIND only ever hands
the caller the user VM's id.

This turns every failed identity attach into a permanent net loss of
heap rather than a neutral retry: confirmed live that a failed attach
left the arena 228,576 bytes worse off than before the attempt, and
every subsequent attempt starts from that worse baseline, compounding.

Fix: call capsule_vm_kill(username) on the now-orphaned console VM
before returning from the failure path -- the same teardown
capsule_wirebind_eject()/capsule_wirebind_unclean_detach() already use
elsewhere in this file (vm_cleanup() + sf_free(), confirmed live to
actually reclaim per-word dictionary allocations, FABRIC-3.md
§IX.2/§IX.3).

Verified live with the same allocator-stats probe (written, captured,
reverted -- not part of this commit): after the fix, a forced user-VM
birth failure now returns the arena to exactly its pre-attempt byte
count (3,526,256, matching the baseline precisely) instead of leaking
228,576 bytes. Three-arch clean qemu acceptance (single Zuse device,
the standard regression case) passed on amd64, aarch64, and riscv64.

The underlying capacity/fragmentation question (why the 7th concurrent
identity's arena allocation fails at all despite technically-sufficient
free bytes) is a separate, open architecture question -- not addressed
here. See project memory for the full measured numbers.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 11:37:06 -04:00