Commit Graph
223 Commits
Author SHA1 Message Date
Robert Allan JamesandClaude Sonnet 5 5417ffb2bf Stage D batch 1: fix silent error sites in repl.c and inference_words.c (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
repl.c's sk_word_blk_attach_ack() and inference_words.c's array_ptr()
helper + infer_word_run()'s allocation guard now log a diagnostic via
log_message(LOG_ERROR, ...) before setting vm->error, matching
defer_words.c's own gold-standard pattern (§XXXII.3) -- these are
internal/background conditions (a malformed message-callback, a bad
array reference or allocation failure), not interactive usage mistakes,
so the fix keeps the fault and reports it rather than dropping it like
USE's own fix did.

Batched together (4 sites total, smaller combined than the next file)
rather than two separate acceptance cycles for negligible size.

Three-arch clean qemu acceptance passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 16:23:14 -04:00
Robert Allan JamesandClaude Sonnet 5 c05b70c8d6 Stage B: logging policy documented, level-aware log-ring eviction built, LOG-FLUSH deferred again (FABRIC-3.md §XXXII.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Policy decided for the kernel-only audit scope: an interactive
command's direct response stays on console_println/console_puts;
everything else (state transitions, background diagnostics, audit
trails) routes through log_message() at the appropriate level,
matching capsule_mint.c's verify_mint() precedent. Documented, not
code-swept here -- reclassifying individual sites is Stage D's job.

LOG-FLUSH deferred again, explicitly: the per-VM log buffer its own
doc comment presumes (vm_log_buffer.h) doesn't exist anywhere in the
tree -- building it is real feature work needing its own scoped stage.

Level-aware eviction built: log_region_append() now reads the oldest
ring slot's own level before evicting it, protecting ERROR/WARN
records from being pushed out by INFO/DEBUG churn -- drops the
incoming low-priority record instead. Found and fixed an adjacent bug
while making this change: the prior two-valued return contract would
have made a benign "dropped by design" outcome indistinguishable from
a genuine write failure to its one caller, which unconditionally set
vm->error on any nonzero return. Changed to a three-valued contract (0
success, 1 dropped by design, -1 genuine failure).

Three-arch clean qemu acceptance passed. Eviction path itself not
live-exercised (needs 128+ LOG-APPEND calls to fill the ring) --
flagged, matching this project's own precedent for that kind of gap.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 14:01:36 -04:00
Robert Allan JamesandClaude Sonnet 5 5c5896fbc1 Stage A: fix USE's silent stack-underflow/bad-address guards, and the Hera fault-scoping gap they exposed (FABRIC-3.md §XXXII.1)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
mama_word_use()'s two silent vm->error=1 guards (dsp<1 stack underflow,
NULL from vm_ptr()) now print a diagnostic and return, matching the
function's other five guards.

Live verification of that fix alone surfaced a bigger problem: with the
guard no longer silent, the REPL proceeds to interpret the leftover
token as an unrecognized word, which independently sets vm->error, and
sk_repl_step()/sk_repl_run()'s Hera-branch still hard-halted on that.
Investigated kernel_main.c's boot/capsule-load paths directly: they
already catch and clear mama->error entirely separately, before
sk_repl_run() is ever entered -- so the "no fallthrough surface" halt
in these two REPL functions was never protecting a boot-time fault, only
an ordinary interactive REPL-turn one. Both functions now recover
unconditionally on any VM's error, Hera included, matching how a
redirected (WIREBIND/USE'd) identity's session already recovered.
Removed the now-fully-unused sk_fault_handler().

Verified live, amd64: USE rajames at Zuse's own console prints the new
diagnostic, then "VM fault -- session recovered, resuming", console
stays interactive afterward. Three-arch clean qemu acceptance passed
before and after the Hera-fault-scoping change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 09:06:33 -04:00
Robert Allan JamesandClaude Sonnet 5 d9da82b065 Stage 4 increments 2+3: WIREBIND VMs as switch-signal participants + mark-and-defer tombstone reap (FABRIC-3.md §XXVIII Stage 4)
Increment 2: WIREBIND user VMs (the ones that actually run FORTH work;
console VMs are pure REPL proxies and never participate) register as
Stage 3 switch-signal participants at attach, unregister at teardown.
Slot table bumped 8 -> 16, matching messaging.4th's own VM-MAX -- a real,
already-agreed ceiling, not an invented number. Added
sk_vm_switch_signal_unregister() (compaction-based; Tripod VMs never
needed removal, WIREBIND VMs cycle constantly and would otherwise
exhaust the bounded table).

Increment 3: implements the plan's own ratified option (A) for the
async-detach UAF risk -- mark-and-defer via a new pending_reap flag on
VMRegistryEntry, deliberately not a new VMState (capsule_vm_kill()
already treats VM_STATE_DEAD as idempotent success, which would silently
swallow a reap attempt; SWITCHED_OUT still accurately describes a
tombstoned VM until the moment it's actually freed). unclean_detach()
sets it when capsule_vm_kill() refuses a SWITCHED_OUT target; the Stage 3
checkpoint (vm_core.c) checks it before ever attempting to resume a
pending switch target, and calls the new capsule_vm_force_reap() instead
-- the one caller allowed to bypass capsule_vm_kill()'s own refusal,
because it runs at the exact safe cooperative point the switcher itself
controls. A new idle-tick sweep cleans up the WIREBIND live-table entry
once the reap has actually happened.

Verified clean on all 3 architectures (baseline regression -- no
WIREBIND attach happens in a plain boot). The reap mechanism's own
correctness under a genuinely parked context is verified separately,
next, via a temporary deterministic probe.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 00:58:34 -04:00
Robert Allan JamesandClaude Sonnet 5 9f0f33dfc5 Stage 4 increment 1: per-device WIREBIND tracking, fixing a real multi-identity detach leak (FABRIC-3.md §XXVIII Stage 4)
capsule_wirebind_unclean_detach()/eject() tracked "the attached identity"
as a single global, correct for the console-pairing UX (one physical
console) but wrong for detach safety: since §XV/§XVI proved multiple
identities genuinely live simultaneously via this same attach path, every
attach after the first silently overwrote the singleton, so an unclean
detach of any but the most-recently-attached identity was silently
ignored -- that VM leaked forever, no trace in the log.

Adds a per-device live-identity table, separate from the (unchanged)
console-pairing singleton, so unclean-detach resolves any attached
device to its own identity. Sized off messaging.4th's own VM-MAX (16)
minus Tripod's 3 reserved slots, not an invented number. Corrects the
stale "single-USB-device constraint" doc claim in capsule_wirebind.h,
false since §XV/§XVI.

Groundwork for Stage 4's real deliverable (WIREBIND VMs as switch-signal
participants) -- this increment only fixes detach targeting; switch-
signal registration is next.

Verified clean on all 3 architectures (no WIREBIND attach happens in a
plain boot, so this is a regression check on the existing Tripod-only
path; live multi-identity verification comes with the switch-signal
registration increment).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 00:02:20 -04:00
Robert Allan JamesandClaude Sonnet 5 05ae7aa886 Fix SWITCH-MARK-WORK off-by-one: was tripping vm->error on every MSG-SEND (FABRIC-3.md §XXVIII.2)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
mama_word_switch_mark_work()'s stack-underflow guard checked dsp < 2,
requiring 3+ items, when it only ever needs the 2 IDX>NAME leaves it
(caddr u). Since dsp is index-based (2 items == dsp 1), this rejected
every normal call. MSG-SEND tail-calls SWITCH-MARK-WORK unconditionally,
so this fired on every message sent anywhere in the system -- visible
only where a caller happened to check the target VM's error flag
afterward (mama_word_vm_exec()'s "VM-EXEC: ERROR in Artemis" report).

Verified clean on all 3 architectures: boot reaches Startup: Artemis
live -> zuse@Hera] ok> with no VM-EXEC: ERROR line at all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-14 20:17:01 -04:00
Robert Allan JamesandClaude Sonnet 5 66beae7fd4 Stage 3 follow-on: message-arrival eligibility hook + trampoline-blind switch-storm fix (FABRIC-3.md §XXVIII.2)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Implements the message-arrival eligibility signal FABRIC-3.md §XXVIII.1 left
open (has_work per-slot flag, set via new SWITCH-MARK-WORK primitive from
MSG-SEND) so an idle VM never becomes a switch target purely by waiting out
the readiness threshold.

Also root-causes and fixes a second, independent switch-storm: the tick's
"who is current" check used vm_log_attributed_vm(), which can't see a VM
parked in switch.c's own raw trampoline. Replaced with a dedicated
g_switch_current_vm tracked by the switch mechanism itself, and moved
target-slot eligibility reset to the switch decision point instead of
relying on ISR polling to observe a window that can be only a few
instructions wide.

Verified live on all 3 architectures: clean boot to zuse@Hera] ok>, live
cross-VM message dispatch, and (since a quiet log looks identical to a
livelocked storm once the DoE probe is gone) confirmed genuine REPL
liveness via QMP send-key + screendump on aarch64/riscv64, not log
inspection alone.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-14 17:39:20 -04:00
Robert Allan JamesandClaude Sonnet 5 862d7d9c48 Stage 3 follow-on: fix stack-ownership corruption + DoE switch columns (FABRIC-3.md §XXVIII.1)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
DoE CSV gained 6 switch-signal columns (switch_count_cumulative,
switch_current_slot, switch_*_readiness, switch_ticks_since), and verifying
them with a boot-time HB-ON probe surfaced a real livelock: the preemption
checkpoint could fire inside a VM-EXEC-nested execute_colon_word() call and
switch away from a stack it didn't own, parking a borrowed region of the
caller's stack under the wrong VM's saved-context pointer. The trampoline
bounce was the visible (safe) half of this; the corruption was the quiet
half, live in every prior "clean" Stage 3 boot without ever showing up in
the log.

Fixed by gating the checkpoint on being at the outermost vm_interpret()
call (g_vm_interpret_depth / sk_vm_at_outermost_interpret(), vm_core.c),
per Bob's decision. Also fixed two related bugs found in the same pass:
g_switch_back_to was a single global stale after first entry, now per-VM
state (native_switch_back_to); note_switch_performed() fired on resume
instead of switch-out, now called before the switch.

Verified on all 3 architectures: steady log growth (no freeze), zero
leaked QEMU processes, DoE columns internally consistent, Hermes/Artemis
confirmed genuinely executing (not just trampoline-bouncing). Temporary
HB-ON boot probe reverted after capture.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 23:58:59 -04:00
Robert Allan JamesandClaude Sonnet 5 986d042aa7 Stage 3: timer-driven preemptive switching, live on all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Fourth stage of the preemptive context-switching plan, and the biggest.
LithosAnanke now genuinely, continuously preempts between Hera, Hermes,
and Artemis -- timer-driven, running live for the entire remainder of
every boot once the Tripod fleet registers, not a bounded probe.

A real design fork was resolved before writing code: the naive approach
(the timer ISR calling Stage 2's sk_vm_context_switch() directly) is
broken -- Stage 0's trap frame lives on whatever stack was active at
interrupt time, and jumping to a different stack via Stage 2's own
independent swap mid-handler would abandon that trap frame unresumed,
guaranteed corruption on the first tick. Chose the safer of two named
options: the ISR only ever sets a flag and returns completely normally
through its own full epilogue; the actual switch happens moments later,
via Stage 2's already-proven mechanism, at a safe cooperative checkpoint
on the mainline (execute_colon_word()'s per-word dispatch loop, checked
on literally every word, not throttled to the existing 256-word
heartbeat-tuning cadence) -- confirmed with the user that word-level
granularity is fine-grained enough given the eventual Zynq FPGA target
where a word is a mnemonic.

New capsule_vm_switch_signal.c/.h: a purpose-built run-readiness signal,
deliberately separate from capsule_vm_physics.c's execution-heat engine
(that one's own header documents itself as never touched from interrupt
context, by design). Slot table sized with headroom (8) rather than
hardcoded to today's 3 participants, so extending participation later is
another register() call, not a redesign -- per direct request to leave
room for swapping the participant set. Simple linear accumulate-then-
threshold for this first cut; a fancier law can replace it later without
touching the mechanism around it. heartbeat_tick() gains its one
deliberate, documented amendment to this file's own top-half/bottom-half
discipline -- the first time this codebase reaches into VM-scheduling
state from real ISR context.

Registration happens only after all three VMs are fully born, right
before the REPL starts -- no critical-section protection yet against
being switched away mid-birth-setup.

Known, flagged rough edge (not reconciled this pass): MSG-TICK's own
idle-pump and this new mechanism can still independently move control
between the same VMs; not observed to interact badly in verification,
but not fully unified either.

Verified interactively at the console on all 3 architectures with
continuous background preemption running throughout -- amd64 computed
`1 1 + .` -> 2, aarch64 computed `1 1 + dup DUP * . CR` -> 4, both
correct, REPL fully responsive, zero fault indicators over sustained
runtime.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 19:32:36 -04:00
Robert Allan JamesandClaude Sonnet 5 f790d0995e Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Third stage of the preemptive context-switching plan. The real
save/restore switch mechanism now exists -- the first time anything has
ever executed on a VM's own native stack (Stage 1 allocated them,
unused).

New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function
call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already
covers every caller-saved register; only the callee-saved set needs
explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV
has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64:
s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for
both halves within one call, since FS is genuine global CPU state, not
part of what's switched). A sibling sk_vm_switch_prime() in the same file
builds the synthetic first-entry frame, kept in assembly so the layout
can never drift out of sync with sk_vm_switch_to() itself.

New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry
priming vs. resuming a parked context, and updates registry state (new
VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no
live frame, this means the opposite). sk_vm_switch_entry() is the minimal
permanent trampoline every freshly-entered VM lands in: no production
behavior defined yet, so it just yields straight back to whoever switched
to it, forever.

Closes the confirmed unguarded-KILL UAF found during planning:
capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama()
all now refuse (or silently leak rather than free, on the cold-restart
path where arch_cold_reset() wipes everything immediately after anyway)
tearing down a switched-out VM. Side effect found, not built on purpose:
the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it
automatically stopped dispatching into a switched-out VM with zero
changes needed there.

Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing
can type interactively into a foreground-only QEMU session) that
round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3
architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe
fully reverted after capture; kernel_main.c shows zero diff.

Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new
files added explicitly (this project's "loader" PE binary is the full
running kernel, not a thin bootstrap stage), and aarch64's switch.S
needed the same #ifndef _WIN32 guard around .hidden that isr.S already
carries (aarch64's loader assembles via clang targeting a PE/COFF
target with no .hidden equivalent) -- caught by a build failure, fixed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 15:25:39 -04:00
Robert Allan JamesandClaude Sonnet 5 57ac3fc304 Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Second stage of the preemptive context-switching plan. Every VM (Hera,
every capsule_birth_baby()-born VM including WIREBIND identities) now
gets its own dedicated 2 MiB native C stack at birth -- but nothing runs
on it yet, that's Stage 2. Pure allocation-machinery proof.

Design correction made before writing code: the plan called for cloning
sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to
be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a
plain kmalloc() block for its dictionary arena, not a real guarded PMM
allocation. Stacks get the real treatment instead (new
sk_vm_native_stack_alloc()/_free() in arena.c): independent
pmm_alloc_contiguous() + guard pages for every VM without exception, no
singleton, no kmalloc fallback -- a stack overflow is exactly the
failure mode guard pages exist for, and a corrupted stack could corrupt
whatever saved context Stage 2 trusts.

2 MiB size matches this project's own established kernel-stack
convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one
shared 2 MiB stack today already carries all VMs' combined nested
VM-EXEC recursion.

Three new VM struct fields, freed in vm_cleanup() alongside the existing
call_stack free. Allocation failure is non-fatal to birth.

All 3 architectures re-verified clean boot to ok>, no native-stack
allocation failures for any Tripod-fleet VM.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 14:30:46 -04:00
Robert Allan JamesandClaude Sonnet 5 15672ce17c Stage 0: trap-frame parity across all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
First stage of the preemptive context-switching plan (see
~/.claude/plans/logical-snuggling-bear.md). Pure foundation work -- every
arch's ISR now saves the full register set on interrupt entry, so a trap
frame is in principle sufficient to resume execution anywhere it was
taken. No FORTH-visible behavior changes.

amd64: added FXSAVE/FXRSTOR, closing a genuine pre-existing correctness
gap (not just future-preemption prep) -- confirmed live double-precision
FP code reachable from ordinary interpreter dispatch (vm_runtime.c Loop
#5/#6), and the ISR previously saved zero FP/SSE state. rbp repurposed as
a fixed anchor so the 16-byte-aligned FXSAVE area can be carved out of an
unpredictably-aligned rsp without disturbing existing argument reads.

aarch64: extended the trap frame 672->800 bytes, adding v8-v15 (AAPCS64
callee-saved, previously excluded on call-site-only reasoning that
doesn't hold for an async trap).

riscv64: extended the trap frame 320->512 bytes, adding s0-s11 and
fs0-fs11 (the latter still correctly gated behind sstatus.FS != Off).

All 3 architectures re-verified clean boot to ok> under the new frames --
amd64 through hundreds of timer ticks with FXSAVE/FXRSTOR live on every
interrupt, aarch64 through 987 ticks, riscv64 clean on the now-larger
FS-conditional block.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 09:23:28 -04:00
Robert Allan JamesandClaude Sonnet 5 2a30212bd3 Real per-VM log persistence: source attribution + ACL pin (FABRIC-3.md §XXVII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Wires the previously-unused vm_log_attributed_vm() into LOG-APPEND's kernel
primitive so persisted log records carry a trustworthy source (the real
attributed VM's registry name, or "HADES" pseudo-source) instead of a
caller-supplied, trivially forgeable string. Drops src-addr/src-u from
LOG-APPEND's stack signature accordingly. Pins LOG-APPEND via bare ACL-PIN
in Artemis's own init.4th, matching BIRTH/CAPSULE-BIRTH's precedent for a
privileged word that can't reach the shared, host-portable ACL.4th.

Also fixes two console-banner nitpicks: a mis-rendering em dash (U+2014)
in the boot banner, and drops "Emergency" from the CLI banner text.

Doc corrections to artemis_sig.h/zuse_eligibility_list.h reconciling the
three fixed devblock ranges now in play. LOG-FLUSH (the intended normal
entry point) and level-aware log eviction remain open, flagged not fixed.

Re-verified clean boot to ok> on all 3 architectures after every change.
riscv64 showed one new, unrelated virtio_blk write-timeout anomaly during
Artemis's early physics self-test (self-recovered, boot unaffected,
sector doesn't map to the log region) -- flagged, not investigated.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 08:56:33 -04:00
Robert Allan JamesandClaude Sonnet 5 61755fde78 Artemis genesis stamp: fix a BAM-corrupting offset before it ever ran (FABRIC-3.md §XXVI follow-on, Step 3)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Step 3: one-time artemis_sig_t genesis stamp, written once
kernel_main.c's virtio-blk path confirms Artemis's own disk, so the disk
image is later recognizable generically (repl.c's idle-loop USB-MSC scan,
built in the prior commit) regardless of which bus found it.

Correction made before this ever touched the real disk: the signature's
first design (committed in 29b6789) placed it at a fixed bottom-of-device
forth-block (4, devblock 1) -- copying homeblocks_sig_t's own convention,
which is safe for a raw identity thumbdrive but not for Artemis's own
disk. Artemis's disk is block_subsystem.c's own STFR/v2-formatted volume:
devblock 0 holds that format's header and devblock 1 is the FIRST
DEVBLOCK OF THE LIVE BAM (blk_compute_fresh_geometry(): bam_start = 1).
The original design would have overwritten Artemis's live allocation map
on the very first real boot. Caught via direct cross-reference against
block_subsystem.c before the genesis-stamp call site was ever run against
the real image -- no corruption occurred.

Fixed by moving the header to a fixed offset from the END of the device
instead (ARTEMIS_SIG_DEVBLOCK_FROM_TOP=64), the same top-of-device region
block_subsystem.c's own meta_fence_blocks reservation (128 devblocks)
already carves out for system metadata, and where Zuse's genesis marker/
eligibility list already live -- but computed independently via
blkio_info() rather than through blk_meta_zone_*(), since that accessor
needs an already-attached, format-detected slot, which is exactly the
state pre-attach generic discovery doesn't have yet. Picked well clear of
Zuse's two tenants (devblock_from_top 0 and 1+, open-ended) so the two
subsystems' independent math can never collide.

Also reordered kernel_main.c: rng_init() now runs before the Artemis
virtio-blk block (was after) -- the genesis stamp needs rng_get_bytes()
for disk_uuid, and the original order would have failed the stamp on
every boot.

Verified live: booted amd64 against the real disk/artemis.img twice --
first boot logs "Artemis: genesis signature stamped" (confirmed blank at
the target offset beforehand via a host-side read), second boot on the
now-stamped image logs no re-stamp (idempotent, CRC/read-back verified)
-- both boots and aarch64/riscv64 (against the same now-stamped image)
all still report "Artemis: 22998 data blocks" / "PASS: persist-read"
unchanged, confirming the BAM and data pool were never touched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 07:31:21 -04:00
Robert Allan JamesandClaude Sonnet 5 29b6789860 Artemis bus-agnostic discovery: signature format + idle-loop generalization (FABRIC-3.md §XXVI follow-on)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Step 1: new artemis_sig_t header format (magic 'ARTM', sibling to
homeblocks_sig_t, distinct so a generic scan can tell Artemis's own disk
apart from an identity thumbdrive by content alone) -- artemis_sig.h/.c,
wired into Makefile.starkernel.

Step 2: sk_repl_idle()'s existing per-USB-MSC-slot attach handling (the
pattern WIREBIND already uses for identity thumbdrives) now also checks
for the ARTM signature whenever a device's home-blocks check comes back
BLANK. On a match, once Artemis's own storage-attach round-trip
(HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK) confirms success,
capsule_zuse_boot_load_root_pubkey() runs -- the same call kernel_main.c's
synchronous QEMU-only virtio-blk path already makes, now reachable
without a hardcoded PCI vendor/device scan. That function is already
idempotent (no-op once zuse_root_pubkey_known is set), so no boot
restructuring was needed despite the initial concern that deferring
Artemis discovery to the idle loop would require one.

Verified: clean build + QEMU boot to [zuse@Hera] ok> on all three
architectures, zero regression to the existing virtio-blk/Zuse-thumbdrive
attach path.

Steps 3 (genesis-stamping onto disk/artemis.img) and 4 (growable
production log-persistence region) not yet started.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 07:16:48 -04:00
Robert Allan JamesandClaude Sonnet 5 a8b16d41da Unify console prompt to [user@VM]; fix real personality-block truncation; correct §XXV's wrong lockdown conclusion (FABRIC-3.md §XXVI)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Investigating the std79 lockdown finding from FABRIC-3.md §XXV led
to a real discovery: WIREBIND births TWO VMs per identity, a console
proxy under the plain username and the actual restricted identity
under <username>~user (capsule_wirebind.c). Every test in §XXV
targeted the console proxy, which was never locked down at all.
Retested against the correct target (rajames~user): the lockdown
works exactly as designed. §XXV's "lockdown never engages" conclusion
was wrong -- corrected here, not deleted, since the mistake and how
it was caught are worth keeping (see the new feedback memory:
confirm which specific VM a name resolves to before concluding
anything, when a subsystem is known to birth more than one VM per
identity).

Two real, separate things found along the way are kept regardless
of that correction:

- capsule_runcap.c: the reserved personality devblock was read in
  full (mostly zero-padding after a short ~200-byte string) with no
  terminator, producing "WARN: block 4998 exceeds 1KB, truncating"
  on every std79-locked identity's birth, universal, since at least
  2026-09-10. Fixed by trimming to the first NUL byte actually found
  -- real, but harmless to execution (real content sat in the
  truncated block's surviving head); it mattered for capsule_id/
  content_hash being computed over padding instead of real content.

- console.h/console.c/repl.c: unified the prompt from a separately-
  computed "[VMName] (user)" into a single "[user@VMName]" line
  prefix -- exactly the ambiguity that caused the original
  misdiagnosis (the prompt showed only the WIREBIND username,
  identical whether USE had targeted the console proxy or the real
  ~user identity). Implemented as a registered callback
  (console_set_user_prefix_provider()) rather than console.c calling
  into WIREBIND/session logic directly, since console.c is a clean
  HAL module with no prior dependency on capsule-level subsystems.

Verified: clean build on all 3 architectures, zero new warnings,
identical dict_hash/capsule_hash to every prior boot this session
(console/prompt-only change). Full 9-identity messaging campaign
re-run end to end: 202s, zero faults, all 8 identities at 99/99
tokens, zero regression.

Also surfaced, not yet acted on: the full campaign's own console
tags now visibly show which VM each identity's tests actually
reached ([zuse@rajames], not [zuse@rajames~user]) -- messaging.4th's
VM-NAMES-INIT registers identities by plain username, so std79-doe.
fth's turn-attractor has been dispatching to each identity's console
proxy, not the actual locked-down identity, since the messaging
rewrite. Flagged for a deliberate decision, not investigated further.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 06:44:58 -04:00
Robert Allan JamesandClaude Sonnet 5 3c2daf50d1 Extend BIRTH/CAPSULE-BIRTH to all VMs symmetrically; flag a real std79 lockdown gap (FABRIC-3.md §XXV)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Scoping the workload-into-factorial design's placement-mode factor
led to a real architectural improvement: rather than EXEC-ing a
workload capsule into an already-running, ACL-locked identity's own
persistent dictionary (filesystem-shaped, doesn't dodge the block-
collision exposure just traced in §XXIV), a workload now runs as a
fresh ephemeral child VM, BIRTH'd per trial and reaped after --
matching the project's own stated principle of automanagement over
imposed policy. CAPSULE-BIRTH already passes vm->stadium_vm_id (who
is birthing this VM) as the new child's parent, not a hardcoded Hera
constant, confirmed by reading the C -- so a workload trial genuinely
inherits the specific identity's own lineage when that identity does
the birthing.

Which surfaced a real premise: only Hera could call BIRTH/
CAPSULE-BIRTH at all (registered only in register_mama_forth_words(),
confirmed directly, not part of the earlier §XX messaging-symmetry
fix which deliberately kept this as one of her remaining privileges).
Extended symmetrically now, agreed explicitly before touching code:

- mama_forth_words.c: BIRTH and CAPSULE-BIRTH added to
  register_child_vm_words(), matching §XX's own pattern.
- acl-std79.4th: ' BIRTH , ' CAPSULE-BIRTH , added to ACL-STD79-LIST
  (new block 4048) -- a deliberate, explicit, named exception to the
  lockdown's own "standard words only" guarantee, not a silent one.
  Symmetric registration alone can't weaken any lockdown on its own:
  ACL-LOCKDOWN-STD79 is allowlist-based, deny-by-default, so a newly
  registered word is auto-denied there unless explicitly added.

Verified: clean build on all 3 architectures, zero new warnings.
Hera's own dict_hash unchanged (expected); Hermes/Artemis show the
same new dict_hash on all 3 architectures. Live-tested against a
real attached std79-locked identity: CAPSULE-BIRTH executes
correctly (returns vm_uuid_none() for a deliberately out-of-range
capsule-id, zero fault, zero ACL denial).

Found, and explicitly stopped short of fixing, a separate pre-
existing gap while verifying the above: MSG-STATUS and MSG-K
(messaging.4th words, not on the std79 allowlist) execute for a
locked identity instead of being denied. ACL-LOCKDOWN-STD79 is
confirmed to actually run; something more specific isn't reaching
messaging.4th's dictionary entries. Root cause not traced -- needs
its own investigation into vm_core.c's dictionary-link mechanics and
whichever capsule actually loads messaging for these identities.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 22:31:52 -04:00
Robert Allan JamesandClaude Sonnet 5 71bcb72a59 Fix O(N) idle-loop messaging pump; full 3x9x3 turn-attractor campaign clean on all 3 architectures (FABRIC-3.md §XXII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
sk_repl_idle()'s messaging pump (repl.c) walked the entire live-VM
registry every idle beat (~1Hz) and dispatched a full VM-EXEC
"MSG-TICK" -- a dictionary lookup plus a 32-slot arena scan -- into
every live VM, every tick, unconditionally, forever. Fine at Tripod's
original 3-VM scale; a full 9-identity turn-attractor campaign
exposed it as a genuine wall on riscv64 specifically (its TCG makes
each dispatch cost more): the same campaign that completed in 194s/
339s on amd64/aarch64 never finished on riscv64 at 9 VMs across three
attempts, while 8 VMs there was fine in 159s.

Ruled out capacity explanations before touching anything: bumping
riscv64's QEMU RAM 1024->4096 changed nothing (reverted), and a live
STADIUM-RES@/MSG-STATUS probe with all 9 VMs attached showed no
depletion. Host memory pressure was also ruled out directly (one
background task did get OOM-killed once during the investigation,
but the identical stall reproduced again with 9.2GB free). The real
mistake was three premature kills under 4 minutes with no way to
tell "slow" from "stuck" from outside the guest -- fixed by having
run_doe_batch.sh sample the qemu process's own /proc/<pid>/stat utime
every 60s; with that signal, riscv64 at 9 VMs was unambiguously alive
(climbing utime, no hang), just disproportionately slow going from 8
VMs (159s) to 9 (600s+ and climbing).

This was never really a riscv64-only bug: an O(N) per-second walk
over the full VM population doesn't scale to the hundreds of VMs this
fleet is headed toward, on any architecture -- riscv64 just made it
visible first, at N=9, because its per-dispatch cost is highest.

Fixed by round-robin batching: the pump now dispatches to at most
SK_MSG_PUMP_BATCH (4) live VMs per idle beat via a persistent cursor
that resumes where the previous beat left off, instead of all of them
every time. Bounds both the scan and dispatch cost to O(K) regardless
of total VM count; any single VM's queue now drains roughly every
ceil(N/K) beats instead of every beat, still bounded and still
matching the pump's own existing best-effort contract. No new C
primitives, no messaging/Stadium changes.

Verified: clean build on all 3 architectures, then the full 3x9x3
campaign re-run on all 3 (not just riscv64) per the standing rule
that a defect repair requires a clean re-run everywhere before
anything counts as closed:

  amd64   198s  0 faults  99/99 tokens x8  656/656 K-conserved
  aarch64 339s  0 faults  99/99 tokens x8  659/659 K-conserved
  riscv64 178s  0 faults  99/99 tokens x8  659/659 K-conserved

riscv64 went from "never completes" to faster than aarch64, same
campaign, same seed, same identity set. No regression on amd64/
aarch64.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 19:29:14 -04:00
Robert Allan JamesandClaude Sonnet 5 7edccd2c35 Hera can't message: root-caused and fixed by symmetry (FABRIC-3.md §XX)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Hera never loaded common:messaging.4th, unlike every other VM in the
fleet. The standing belief was this was deliberate, to avoid moving
her dict_hash off baseline. Checked live instead of assumed: loading
messaging.4th into her dictionary silently dropped every colon-
definition referencing one of 8 STADIUM-* primitives that
register_child_vm_words() gives every other VM but
register_mama_forth_words() never gave her -- a missing-primitive gap,
not a designed privilege boundary. Confirmed mama_word_birth is
genuinely VM-agnostic and SPAWN-EVENT is an unwired placeholder before
proposing the fix.

Fix: register the same 8 STADIUM-* primitives for Hera, load
messaging.4th from init.4th the same way Hermes/Artemis/console/mint
already do, and give her own idle-loop context a direct MSG-TICK call
(not VM-EXEC, which would hit the same reentrancy class the existing
per-other-VM pump loop already guards against) so her own queued
messages actually drain. Her dictionary is now a proper superset of
every child VM's, plus her remaining extra privileges -- not
structurally different from any other VM, just additionally
privileged.

Verified: dict_hash identical across amd64/aarch64/riscv64
(0xc8f4b09e36f4fc4a), Hermes/Artemis dict_hashes unchanged and still
cross-arch identical, all three boot clean to zuse)ok> with zero
UNKNOWN WORD faults, mkcapsule --lint clean.

Unblocks rewriting the turn-attractor (FABRIC-3.md §XIX) to coordinate
via real MSG-SEND/MSG-TICK instead of blocking VM-EXEC.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 16:05:11 -04:00
Robert Allan JamesandClaude Sonnet 5 5a9425b91b Add VM-HEAT primitive: groundwork for a compudynamic turn-attractor (FABRIC-3.md §XIX)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Prompted by reading the std79-doe K report: checked whether the campaign's
"real cross-VM dispatch load" was actually concurrent or strictly
serialized. It's serialized at two levels -- VM-EXEC's vm_interpret(target,
...) is a direct synchronous C call (Hera fully blocked until it returns),
and even the background physics tick (vm_tick(), vm_runtime.c) is driven
by each VM's own execution loop, so idle identities accrue zero ticks
between their own turns. K's perfect conservation (FABRIC-3.md §XVIII)
verifies sequential per-VM accounting correctness, not concurrent-access
safety, since there was never concurrent access to test.

Agreed direction: fix this without a scheduler, by reusing the same
least-dense-candidate judgment stadium_admit() already trusts for
eviction, applied to "whose turn is next" instead of "who gets evicted" --
a fleet-level turn-attractor giving the next turn to whichever live
identity currently has the lowest execution_heat_q48, no fixed round-robin,
no priorities, no preemption. Lives beside Stadium in
capsule_vm_physics.c (already the fleet-level consumer of Stadium
primitives, e.g. the K mechanism itself), not inside stadium.c ("the
floor" -- residency/eviction, a different concern from turn order) and
not a new subsystem.

This pass lands only the primitive the mechanism needs: VM-HEAT
( c-addr u -- heat-q48 ), pushing a named VM's current
execution_heat_q48 via vm_physics_heat_of() -- previously C-internal
only (doe_log_heat_by_name(), doe_log.c), never exposed to FORTH. Silent
0 on an unknown/dead name (no print/error), since a turn-attractor
scanning many candidates every turn shouldn't have to filter console
noise for names that simply aren't live. Registered everywhere
VM-EXEC/VM-CALL already are. Builds clean on all three architectures;
live-tested on amd64: Hera -> 65452, Hermes -> 43, unknown name -> 0, no
faults.

The turn-attractor loop itself (std79-doe.fth's trial ordering) is not
yet built -- next step, not done here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 14:53:06 -04:00
Robert Allan JamesandClaude Sonnet 5 66ba21adb4 Log fleet_k_q48/fleet_conserved; K holds exactly, 775/775 ticks (FABRIC-3.md §XVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
doe_log.c's per-heartbeat-tick CSV gains two columns: fleet_k_q48
(vm_physics_fleet_heat_sum() over ALL live VMs -- the genuine fleet-wide
conservation invariant K, not reconstructable from the 3 named-Tripod-
member heat columns already logged, which omit every identity VM's own
heat) and fleet_conserved (vm_physics_conserved() as 0/1). Requested
explicitly after the first heartbeat-telemetry analysis pass
(analysis-20260912/) omitted K entirely.

Kernel rebuilt on all three architectures, full 3x9x3 campaign rerun
(results-20260912-with-k/). K = 1.0000000000 (Q48.16 raw 65536) on every
one of 775 heartbeat-tick observations, sd(K) = 0, 100% fleet_conserved,
across amd64/aarch64/riscv64, nine identities, three replicates -- zero
deviation. Also a free regression check on both recent Stadium fixes
(§XVI/§XVII): neither disturbed the reservoir-transfer accounting K
depends on.

Found and fixed a tooling wrinkle along the way: fleet_conserved, being
the CSV row's very last field with nothing after it to bound a regex
match, can have a resumed trial digit merge into it with zero separator
on the wire -- combine.py now derives it from fleet_k_q48 directly (same
epsilon vm_physics_conserved() uses) instead of trusting the raw field.
fleet_k_q48 itself is unaffected either way.

Full analysis, discussion, and light/dark SVG->PDF figures written up as
a proper LaTeX report (report-20260912/report/std79_doe_report.pdf),
following experiments/bare_metal/analysis/report/bare_metal_doe_report.tex's
established style -- supersedes analysis-20260912/'s markdown-only first
pass as the primary deliverable for this dataset (kept, not discarded).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 13:25:09 -04:00
Robert Allan JamesandClaude Sonnet 5 e51a8d229e Fix stadium_grant_quota() donor floor; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
capsule_birth.c hardcoded every new VM's initial Stadium quota grant to split
from Hera specifically. Since a grant always halves whatever the donor
currently has, Hera's own free list converges toward empty after a bounded
number of grants — independent of whether the Stadium as a whole still had
spare capacity, since VMs she'd granted to earlier typically still held
nearly all of their own share untouched. Past that point every subsequent
VM birth's Stadium grant would be silently refused (soft-failed, non-fatal
by existing design), even with plenty of capacity sitting idle elsewhere.

Fixed by adding an O(1)-maintained free_count to StadiumVMQuota (incremented
in stadium_evict(), decremented at both of stadium_admit()'s free-list-pop
sites, set/adjusted in stadium_grant_quota()'s own split — this also let
grant_quota drop its old O(free-list length) counting walk in favor of an
O(1) read) and stadium_best_donor(), an O(live VM count) scan over quota
slots returning whichever in-use VM currently has the most free cells.
capsule_birth.c's birth path now splits from that VM instead of
unconditionally vm_uuid_hera().

Verified with another full rerun of the 3x9x3 std79 DoE campaign from
scratch — same discipline as the prior Stadium fix (any defect repair
reruns the whole DoE from the top) — one continuous boot per architecture,
all 9 identities simultaneously live throughout. 81/81 trials correct, 0
mismatches, DOE-RUN header sequence md5-identical to every prior run.
aarch64 ~280s total (vs ~290s for the O(ncells)-scan fix alone — confirms
no regression). Both known Stadium defects are now closed together on one
clean campaign rerun.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 17:20:07 -04:00
Robert Allan JamesandClaude Sonnet 5 9eff122090 Fix aarch64 Stadium/COOL O(ncells) scan; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVI)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Root cause of the 90+ minute aarch64 VM-birth stall found in §XV: stadium_admit()'s
eviction-fallback scan iterated the entire stadium_ncells array filtered by owner,
not the calling VM's own resident cells as its own doc comment claimed. Combined
with stadium_grant_quota() always splitting from Hera's shrinking free list and
stadium_word_dispatch() calling stadium_admit() per distinct word a VM's capsule
executes, this compounded into a real O(n) blowup — catastrophic specifically on
aarch64 because its -m 4096 (vs 1024 on amd64/riscv64) inflates the kmalloc heap
kmalloc_init() bisects down to, which inflates stadium_ncells 4x (335,544 vs 83,886
cells, measured from boot logs).

Fixed by threading a real per-VM doubly-linked resident-cell list
(StadiumVMQuota.resident_head + stadium_resident_next[]/stadium_resident_prev[])
so the fallback scan is bounded by that VM's own resident count, not the global
cell array size.

Verified with a full rerun of the 3x9x3 std79 DoE campaign from scratch: one
continuous boot per architecture, all 9 identities simultaneously live throughout
(the 3-boot aarch64/riscv64 batching workaround is no longer needed). 81/81 trials
correct, 0 mismatches, DOE-RUN header sequence md5-identical across all three raw
logs. Identity 04's attach on aarch64, which stalled 90+ minutes before, now
completes in ~34s; full boot-to-DoE-complete in ~290s.

Corrects an earlier misreading (carried into §XV, std79-doe.fth's comments, and
the project memory note) that described the symptom as a runaway "335,000+ cycles"
dispatch counter — those were cell array indices, not an event count.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 16:37:44 -04:00
Robert Allan JamesandClaude Sonnet 5 70db955ac9 Fix M/MOD: hand-rolled 128/64 bit-serial division (FABRIC-3.md §XII.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
mixed_math_word_m_slash_mod() had the identical bug class already fixed
in M* (commit 9a09949): it reconstructed the dividend from
`(dhigh << 32) | (dlow & 0xFFFFFFFF)`, the same wrong "32-bit halves of a
64-bit value" assumption. Latent for small inputs (fits in 32 bits, so the
reconstruction coincidentally worked), confirmed genuinely broken for a
true wide double -- feeding M*'s own correct 10^24 output into it gave a
quotient/remainder wrong by many orders of magnitude and the wrong sign.

Unlike M*, this one can't just switch to __int128 -- __int128's own `/`/`%`
need libgcc's __udivti3/__umodti3 for 128-bit division, unavailable in
this freestanding, -nostdlib build (the exact constraint
src/starkernel/arch/amd64/timer.c's own doc comment already flagged:
__int128 multiply/shift-by-constant/compare/subtract all compile clean,
only division doesn't). Fixed instead with a hand-rolled unsigned
128-by-64-bit bit-serial (restoring) long division -- 128 iterations of
shift-by-1/compare/subtract only, all in the safe set. Operates on
magnitudes via unsigned negation from 0 (well-defined even for the
extreme negative edge); signs reapplied afterward matching the same C99
truncating-toward-zero convention the previous, narrower implementation
already used, unchanged.

Verified: rebuilt all three architectures, confirmed clean link with no
__udivti3/__umodti3 undefined-symbol errors. Booted and tested all three,
identical results: 1000000000000 1000000000000 M* SWAP 1000000 M/MOD ->
1000000000000000000 remainder 0 (exact division, the case that was wrong
by orders of magnitude before); three sign-combination cases all correct
(-+, +-, --), confirming sign handling survived the magnitude-only
rewrite. T15 (original small-input case) unaffected. Full 24-case
exerciser reran clean on all three, no regressions.

All bugs found by the std79 exerciser campaign, including this one found
while fixing another, are now closed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 10:28:34 -04:00
Robert Allan JamesandClaude Sonnet 5 9a09949c69 Fix M*: use __int128 for a genuine 128-bit double-cell product (FABRIC-3.md §XII.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
mixed_math_word_m_star() confused "double" (two full cell_t-width cells,
128 bits total on this 64-bit build -- what D+/D-/D./etc. all actually
expect) with "the low/high 32-bit halves of a single 64-bit product" --
code clearly written assuming cell_t is 32-bit. It computed an ordinary
64-bit `long long` product (already wrong for any true product exceeding
64 bits, since long long is the same width as one cell here) and split
that into 32-bit halves via `result & 0xFFFFFFFF` / `result >> 32`. For a
small negative product like -56088, this produced a positive, zero-
extended low cell paired with a correctly-looking dhigh=-1 -- D.'s
overflow check (correctly) rejected the resulting malformed double, on
every architecture, every time (this bug was never architecture-specific,
unlike the D+/D-/DNEGATE/d_compare family already fixed in bea8d74/
1a716c8).

Fixed via __int128 for a genuine 64x64->128-bit multiply, no truncation --
an already-established safe pattern in this codebase for exactly this
operation (src/starkernel/crypto/fe25519.c/scalar25519.c already use it;
src/starkernel/arch/amd64/timer.c's own doc comment confirms __int128
multiply/shift-by-constant compile cleanly with zero undefined symbols on
all three target toolchains -- only division needs unavailable libgcc
support, not used here).

Verified: rebuilt and booted all three architectures clean. T13/T14 both
correct everywhere (56088/-56088). Manual cases confirm genuine 128-bit
precision: -1 -1 M* -> 1; 1000000000000 1000000000000 M* -> dhigh=54210
dlow=2003764205206896640 (10^12 x 10^12 = 10^24, correctly exceeding 64
bits). Full 24-case exerciser reran clean on all three, no regressions.

Found while verifying, NOT fixed (report only, out of scope of this
request): mixed_math_word_m_slash_mod() (M/MOD, the very next function in
the same file) has the identical bug class -- confirmed genuinely broken
for a true wide double (feeding this fix's own correct large-magnitude M*
output into M/MOD produces a result wrong by many orders of magnitude and
the wrong sign). Flagged in FABRIC-3.md for a future fix request.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 08:23:51 -04:00
Robert Allan JamesandClaude Sonnet 5 1a716c8048 Fix d_compare: use ucell_t instead of unsigned long (FABRIC-3.md §XII.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
d_compare() (double_words.c), the static helper backing DMAX/DMIN/D</D=,
had the same bare-unsigned-long pattern already fixed in D+/D-/DNEGATE
(commit bea8d74) -- 32-bit unsigned long on this aarch64 target silently
truncating the low-cell comparison. Switched to ucell_t.

Verified with four cases designed specifically to expose the old
truncation: two doubles sharing the same high cell but with low cells
differing only in bits 32-63 (invisible to a 32-bit-truncated compare,
e.g. 2^32 vs 0):
  4294967296 0 0 0 D= .            -> 0  (correctly not equal)
  4294967296 0 0 0 DMIN SWAP D.    -> 0  (correctly picks the smaller)
  0 0 4294967296 0 D< .            -> -1 (correct)
  4294967296 0 0 0 D< .            -> 0  (correct)
All four pass identically on amd64, aarch64, and riscv64. These cases
were never run before this fix -- the exerciser campaign never touched
DMAX/DMIN/D</D= at all, so this is the first real evidence d_compare()
was ever exercised on any architecture. Full 24-case exerciser also
reran clean on all three architectures (no regression).

D2*/D2/ use unsigned long long (C-standard-guaranteed >=64-bit) and were
never part of this bug class. All four affected words (D+, D-, DNEGATE,
d_compare) are now fixed; M*'s separate universal DOUBLE-OVERFLOW bug
remains open (unrelated defect, out of scope here).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 07:44:14 -04:00
Robert Allan JamesandClaude Sonnet 5 bea8d7436a Fix D+/D-/DNEGATE: use ucell_t instead of unsigned long (FABRIC-3.md §XII.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
double_word_d_plus(), double_word_d_minus(), and double_word_dnegate()
(double_words.c) all cast through plain `unsigned long` for their carry/
borrow-detection arithmetic. On this aarch64 bare-metal cross-compile
target, unsigned long is 32-bit (confirmed: sizeof(unsigned long)==4) --
amd64 and riscv64 both happen to have a 64-bit long, so the identical code
only broke on aarch64. The low-cell arithmetic silently truncated to 32
bits, then widened back to cell_t via ordinary (non-sign-extending)
conversion, producing a wrong result whenever the true 64-bit result was
negative -- D. then correctly, faithfully reported DOUBLE-OVERFLOW on the
resulting malformed double.

vm.h already defines ucell_t for exactly this: same conditional as cell_t,
guaranteed width-matched on every target. print_number_formatted()
(format_words.c) already used it correctly; these three words didn't.
Switched all three to ucell_t -- a one-word-class fix, no logic change.

Verified: rebuilt and booted all three architectures clean. T19 (D+) on
aarch64 now correctly prints -2, matching amd64/riscv64; T20 (DNEGATE)
unaffected everywhere. Additional manual cases beyond the original
exerciser, run live on aarch64 to specifically exercise the
>32-bit-magnitude path the old bug depended on: D- (-5-3=-8), DNEGATE on
2^33 (8589934592 -> -8589934592), D+ crossing the same boundary
(3+8589934592=8589934595) -- all correct.

d_compare() (backing DMAX/DMIN/D</D=) has the identical latent pattern but
is out of scope for this fix (not named in the request, never exercised by
the campaign) -- left open, flagged in FABRIC-3.md. M*'s separate,
universal-across-all-three-architectures DOUBLE-OVERFLOW bug is also
untouched -- unrelated defect, not part of this fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 07:37:57 -04:00
Robert Allan JamesandClaude Sonnet 5 9142dda2d6 Fix use-after-free in sk_repl_idle()'s idle-tick VM resolution (FABRIC-3.md §XIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Root-caused a heap-corruption bug that reliably failed WIREBIND identity
attach on the third attach/detach cycle in one boot. sk_console_readline()
and sk_console_getkey() captured `active_vm` once from their caller and
kept passing that same (possibly long-stale) pointer to sk_repl_idle() on
every idle tick serviced while blocked waiting for input. If the VM it
pointed at was killed (WIREBIND detach) mid-block, the existing bailout
only checked a generic "is anyone attached" boolean -- masked as soon as a
different identity attached next -- so blk_vm_flush_all() kept writing
into a freed VM struct sitting on kmalloc's own free list, corrupting the
free list's linked-list metadata itself.

Both idle branches now re-resolve the live active VM fresh from
g_repl_active_vm on every tick, matching the dispatch-side fix already
made for the sibling bug in §XII.3.

Verified: rebuilt amd64, reran the exact three-cycle repro that reliably
corrupted the heap before the fix -- free-list census stayed stable
through the same idle window that previously collapsed to zero. All three
architectures (amd64/aarch64/riscv64) boot clean to the zuse)ok> prompt.

kmalloc_debug_census()/kmalloc_debug_census_bytes() kept as permanent
diagnostic infrastructure; every other temporary probe added during the
investigation was reverted.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-10 20:30:41 -04:00
Robert Allan JamesandClaude Sonnet 5 70421bdd43 Fix real WIREBIND crash: stale active-VM pointer dispatched after blocking read (FABRIC-3.md §XII.3)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
The interpreter_enabled guard added in the previous commit (662ef44) was a
real but incomplete fix -- re-running the exact repro against it still
panicked (this time as a raw #PF page fault), proving something deeper
was wrong.

Root cause, found via targeted console_puts probes (not GDB --
starkernel_kernel.elf's symbols don't correspond to the actual running
starkernel_loader.efi binary for this monolithic build, same gotcha
already on record from the 2026-08-18 aarch64 investigation):

sk_repl_run()'s main loop captures `active` once, before calling
sk_console_readline(), which then blocks for the next full line. If the
identity `active` points at is killed while that read is still blocked,
the bailout meant to catch this (sk_console_identity_present()) only
checks a generic "is anyone attached" boolean, not "is the specific
identity active belonged to still attached" -- a fast detach of one
identity followed by attach of a different one never produces an
observable gap in that boolean, so the bailout never fires. The stale
`active`, now pointing at freed memory, gets dispatched into.

Fix: re-resolve `active` fresh from g_repl_active_vm immediately before
dispatch, right after sk_console_readline() returns. One line, no
registry lookup, no dereference of the stale pointer -- closes the race
regardless of whether the bailout catches it first.

Verified: rebuilt amd64 clean, reproduced the exact same attach/USE/
detach/attach/USE sequence against the fixed build -- clean switch, no
fault, exerciser runs correctly afterward.

FABRIC-3.md §XII.2 also corrected to stop claiming the interpreter_
enabled guard alone closed the crash -- it didn't, per the above.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-10 17:19:01 -04:00
Robert Allan JamesandClaude Sonnet 5 662ef44e59 Fix EXEC/LOAD block-persistence gap and WIREBIND/USE interpreter-race panic (FABRIC-3.md §XII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Found live while building a cross-ISA FORTH-79 dictionary exerciser:

- capsule_exec_init() zeroed a capsule's block content immediately after
  running it, so LOAD (a genuine FORTH-79 standard word, ACL-allowed even
  for locked identities) could never actually read back what EXEC had
  just written. Removed the clear from capsule_exec_init(); block content
  now persists like any other Standard BLOCK/BUFFER/UPDATE write.
  kernel_main.c's own explicit post-birth clear of Mama's init.4th range
  is untouched.

- USE could redirect the console to a WIREBIND identity's VM before that
  VM's vm_enable_interpreter() step of its own birth sequence had run,
  causing the next typed line to hit vm_assert_interpreter_enabled() and
  panic the entire machine -- not the per-session-recoverable ACL-fault
  path a redirected VM otherwise gets. USE now checks interpreter_enabled
  first and refuses with a retry message instead.

Also includes the amd64/aarch64/riscv64 acceptance boot logs and DoE CSVs
from this session.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-10 16:07:38 -04:00
Robert Allan JamesandClaude Sonnet 5 b301317902 xHCI: drive Port Reset on port reuse; WIREBIND: kill the console VM too (FABRIC-3.md §XI.5)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Two independent bugs that together caused a reliable hotplug wedge:
reusing an xHCI port for a second identity right after an unclean
detach of a first would leave no further hotplug events reaching the
guest at all.

Bug 1 (xhci.c/xhci_driver.h): the xHCI driver never drove PORTSC.PR --
a known, named gap since Milestone 2e (the code's own comment flagged
it, PORTSC_PR/PRC were defined but never referenced). A port's first
connect each boot reads PED already set, so skipping the reset
happened to work; a second device on the same port after a prior
disconnect reads PED clear, and Address Device reliably failed without
an explicit reset cycle. New XHCI_CONN_AWAIT_PORT_RESET state drives
PR and waits for PED to read set before proceeding to Enable Slot.

Bug 2 (capsule_wirebind.c): capsule_wirebind_eject()/unclean_detach()
compared g_repl_active_vm against the *user* VM's pointer
(g_wirebind_attached_vm_id tracks that one, not the console VM) --
never equal, since USE/g_repl_active_vm always points at the console
VM. The guard never fired and the console VM was never killed at all,
only orphaned -- paired to a dead user VM but still the REPL's active
session. New wirebind_teardown_console() helper resolves and tears
down the console VM by its own tracked bare username.

Verified live on amd64: the exact repro (identity 01 on port 2,
unclean detach, identity 02 on the same port immediately after) --
previously wedged with "xhci: address device failed" and no further
hotplug activity; now attaches cleanly and fast, both VMs' KILL
messages appear, console is immediately interactive on the new
identity. Three-arch clean qemu acceptance passed, all clean on the
first attempt.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 22:11:09 -04:00
Robert Allan JamesandClaude Sonnet 5 d6661b5eed Scope VM fault halt to the faulting identity's own session (FABRIC-3.md §XI.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
A standalone WIREBIND identity (no Zuse, USE'd in directly) hitting an
ACL-denied word halted the entire machine -- Hera, Hermes, Artemis, all
of it -- instead of just that identity's own session. sk_fault_handler()
was being called unconditionally on whichever VM's ->error was set, with
no distinction between Hera's own root session (where "no fallthrough
surface" is the correct, deliberate fail-closed behavior) and a
USE'd-in guest identity (which should recover and resume at its own
prompt instead of taking the fleet down with it).

Both call sites (sk_repl_step, sk_repl_run) now compare the faulting VM
against Hera before deciding: Hera's own session still halts by design;
any other VM prints a recovery message, clears its fault state, and
continues.

Also: mint identities 01-06 with the same FORTH-79/83 restricted
personality identity 00 already had, verified via the fixed fault
scoping above (which this verification pass surfaced).

Verified live on amd64 (both the Hera-halts and identity-recovers
branches); three-arch clean qemu acceptance passed (riscv64's first
attempt hit an unrelated virtio_blk I/O timeout hang, a known QEMU/TCG
flake -- a clean retry booted normally).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 20:41:51 -04:00
Robert Allan JamesandClaude Sonnet 5 1839a2b0c3 WIREBIND cert verification: load Zuse's root pubkey independently of her live session
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Root cause of the remaining "identity attach doesn't complete when Zuse
never attaches this boot" issue: capsule_wirebind_verify_cert() gated on
mama_vm->zuse_cert_installed, which is only ever set when Zuse's own
drive attaches and authenticates this specific boot
(capsule_zuse_boot_try_attach() -> install_and_activate() ->
vm_zuse_cert_install()). Without her, any other identity's WIREBIND cert
verification silently refused -- correctly, by the old design, but that
design conflated two genuinely different things: "can mint new
identities" (needs Zuse's live private seed, a real privileged
operation) and "can verify an existing identity's cert" (needs nothing
but her already-public key).

That public key was already being persisted independently of her live
session: zuse_genesis_marker_t (zuse_genesis_marker.h) stores it in the
kernel's own top-of-device metadata fence (Artemis's resident storage),
written once at genesis, specifically *not* alongside her private seed
(which stays only on her own removable thumbdrive) -- the type's own doc
comment says as much. It just wasn't being loaded for anything but
confirming which drive is genuinely hers.

Fix: a new capsule_zuse_boot_load_root_pubkey() (capsule_zuse_boot.c)
reads that marker and populates two new VM fields, zuse_root_pubkey_known
/ zuse_root_pubkey (vm.h) -- deliberately separate from
zuse_cert_installed/zuse_cert_seed/zuse_cert_pubkey, which stay
untouched and still gate MINT exactly as before. Called once from
kernel_main.c as soon as Artemis's own storage attaches, unconditionally,
independent of whether Zuse's own drive is ever attached this boot.
capsule_wirebind_verify_cert()/capsule_wirebind_try_attach() now check
zuse_root_pubkey_known instead of zuse_cert_installed.

One identity's attach must not depend on another identity's live
presence -- each identity stands on its own once the fleet's root of
trust has been established once, ever.

Verified live, amd64: identity 00 (disk/thumbdrives/00-thumb-ident.img)
now attaches and completes WIREBIND in 19 seconds with Zuse's own drive
never attached this boot at all (previously: unbounded, many real
minutes or effectively never, before today's other fixes; still slow/
stuck after those, stuck specifically on this silent refusal). Zuse's
own attach flow re-verified unaffected (regression check, amd64).
Three-arch clean qemu acceptance (amd64/aarch64/riscv64) passed with
this change included.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 14:30:14 -04:00
Robert Allan JamesandClaude Sonnet 5 1a263555e2 Fix pathological migration scan that stalled WIREBIND identity attach
Root cause (found by a fresh subagent after an extended live-debugging
investigation into "identity 00 attaches slowly/stalls when Zuse never
attached first this boot"): blk_migration_idle_check() was generalized
earlier today to walk every attached device slot uniformly instead of
hardcoding first_disk_slot() (Artemis's own disk). But its per-slot scan
can only early-exit once it finds a devblock that is BOTH "hot" (claimed
and worn) AND "free" -- and a just-attached, never-claimed USB identity
drive can never satisfy the "hot" half by design (claiming only ever
happens via blk_firsttouch_claim(), which only ever targets
first_disk_slot()). So the scan ran to completion -- the drive's entire
~16,000 devblocks, mostly cache misses over slow emulated USB/BOT --
every single idle tick, forever, blocking sk_repl_idle() (and therefore
the console and the storage-attach message round-trip) each time.

Fix, in src/block_subsystem.c: a new has_ever_claimed flag on
blk_dev_slot_t (set in blk_set_meta(), the single choke point every
BLK_FLAG_CLAIMED transition passes through) skips the scan entirely,
O(1), for any slot nothing has ever claimed -- the common case for a
freshly-attached drive. A new migration_scan_lbn resume cursor bounds
*any* slot's per-tick cost to MIGRATION_SCAN_BUDGET (256) devblocks
examined, picking up where the previous tick left off instead of
restarting from start_lbn every time -- restores this function's own
documented "coarse cadence, cheap early-exit" design intent for every
device, not just the one it used to hardcode.

Also along the way (kept, all real improvements, verified live):
- src/starkernel/usb/xhci.c: xhci_wait_bit()/xhci_bot_wait_for_idle()
  had zero yield hints in their MMIO-polling loops; added arch_relax()
  to both (matches virtio_blk.c below) -- a tight loop of nothing but
  MMIO reads can starve TCG's own host-side timer injection under QEMU.
- src/starkernel/virtio/virtio_blk.c: vblk_io()'s spin bound was 33M
  iterations with zero logging on timeout; a single real (still not
  fully root-caused) timeout cost 31+ minutes of CPU before this was
  caught. Reduced to 1M and added a log line naming the failing sector,
  turning a silent, effectively-unbounded stall into a fast, loud
  failure -- callers already tolerate BLKIO_EIO.
- src/starkernel/repl.c: blk_migration_idle_check() deferred for any
  idle tick where a storage-attach message round-trip is still pending,
  to keep the two block-subsystem-touching paths from interleaving; the
  existing MSG-TICK pump now checks the target VM's own dictionary for
  MSG-TICK (vm_find_word + acl_allow) before VM-EXECing into it, instead
  of spamming "VM-EXEC: ERROR" every tick for a VM that doesn't have --
  or isn't allowed -- the word (e.g. a FORTH-79/83-locked-down identity);
  fixed a real -Wunused-variable build break in the EMERGENCY_CONSOLE_
  ENABLED=1 path (unused today, but needed live to reproduce this bug
  with no identity attached at all).
- src/starkernel/capsule/capsule_mint.c: dropped the dead
  S" common:messaging.4th" EXEC / MSG-CD-INIT lines from
  MINT_RESTRICTED_PERSONALITY -- a FORTH-79/83-locked-down identity has
  no legitimate use for a messaging vocabulary it can never call.

Status: the pathological CPU-climbing scan is confirmed fixed (verified
live: CPU stays flat across an extended run instead of climbing without
bound). The WIREBIND storage-attach message round-trip still does not
complete promptly in the "Zuse never attached, other identity attaches
first" scenario -- a separate, still-open issue in the message-delivery
path itself, not the scan. Tracked as follow-on work.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 12:08:03 -04:00
Robert Allan JamesandClaude Sonnet 5 63b8b3bc29 Migrate Hera->Artemis storage-attach to a real message round-trip
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-amd64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Hera still polls xHCI and sig-checks attached drives, but the storage
registration step (blk_subsys_attach_device(), now wrapped as the
BLK-ATTACH primitive) moves to Artemis's own dictionary, reached via
HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK (VM-EXEC, since Hera can't load her
own messaging.4th -- see the doc comment in repl.c). Identity birth
(Zuse genesis / WIREBIND) is deferred until the ack confirms storage
actually succeeded, instead of running synchronously underneath a
storage call that might fail ("wait for ack, safer for identity data").

Caught and fixed a real bug live during acceptance testing: Artemis's
ACK-APPEND-NUM fed a single-cell value into <# #S #> (which expects a
double-cell pair), causing a stack underflow the first time
HERA-BLK-ATTACH-REQ ran. Fixed with the same `0 SWAP` convention every
other numeric-append helper in this codebase already uses.

Verified booting clean to (zuse) ok> with no VM-EXEC errors on all
three architectures (amd64/aarch64/riscv64).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 21:22:04 -04:00
Robert Allan JamesandClaude Sonnet 5 f4ded3e1a8 MINT: add a FORTH-79/83-standard-words-only lockdown personality
Captain Bob, 2026-09-07: "starting with that 00 user we created, we're
going to give access only to FORTH 79 and 83 standard words. everything
else is locked down."

New capsules/acl-std79.4th (blocks 4023-4047): walks a VM's own
dictionary (>LINK/LINK> traversal, same as ACL-INIT-PRIMITIVES/WORDS
already use) and permanently denies+pins every word not on an explicit
FORTH-79/83 allowlist, extracted from the real registered word set
(stack_words.c through control_words.c), not recited from memory.
Deliberately excludes, beyond plain non-standard words: BYE (100% ACL
bypass to the emergency console -- "needs more discussion, exclude for
now"), COLD/WARM/REBOOT/SAVE-SYSTEM (system lifecycle), the block/screen
editor L/S/SHOW/EDIT/UPDATE/SAVE-BUFFERS (lets a session rewrite
persistent block/capsule content, defeating the lockdown even though
nominally standard), BLK-ACL-*/BLK-OWNER@ (StarForth-specific), and
FORGET/FENCE (flagged as an unrestricted superpower word, 2026-09-03
audit). Keeps WORDS/VLIST/SEE (introspection only -- ACL is enforced
per-target-word at execution time regardless of how an XT was
obtained) and the parenthesized control-flow runtime primitives
((BRANCH) etc. -- IF/DO/LOOP compile calls to these; denying them
breaks ordinary control flow, not security).

MintPersonality enum (capsule_mint.h) lets capsule_mint_identity()
select which personality-source template gets written to a new
identity's devblock -- MINT_PERSONALITY_DEFAULT (unchanged) or
MINT_PERSONALITY_STD79_LOCKDOWN (EXECs acl-std79.4th then
ACL-LOCKDOWN-STD79 as the VM's own last bootstrap step). The actual
restriction logic stays entirely in FORTH per .claude/CLAUDE.md's
Word-Level ACL System rules ("ACL policy belongs in ACL.4th, never in
C") -- capsule_mint.c only picks which few-line bootstrap stub to
write. MINT's own stack signature gains a trailing restrict? flag;
capsule_zuse_boot.c's genesis mint (Zuse herself) explicitly passes
MINT_PERSONALITY_DEFAULT -- the superuser is never restricted.

Two real bugs found and fixed live during testing, both the same class
of self-referential fault: ACL-LOCKDOWN-STD79's own walk loop calls
ACL-STD79-ALLOWED?/ACL-STD79-LIST/ACL-ALLOW!/ACL-PIN on every single
iteration to do its job -- none of those are FORTH-79/83 standard
words, so the walk was denying its own load-bearing infrastructure
partway through and then faulting the next time it tried to call it
("VM fault -- emergency console disabled; halting", reproduced twice
live). Fixed by explicitly protecting all four in the allowlist
(block 4047) -- they must stay allowed for the walk to finish, not
because they belong on a "standard words" list.

Verified live end-to-end: minted a throwaway test identity with the
restrict? flag, confirmed her WIREBIND birth completes cleanly (no
faults, no shadow conflicts) on a single real attach, then USE'd into
her VM and confirmed standard arithmetic and user-defined words work
(1 2 + . -> 3; : X 5 5 * . ; X -> 25) while KILL is entirely unknown to
her dictionary and VM-EXEC is denied. One real, non-fatal side effect
found and left as-is (not asked to fix): the fleet's inter-VM messaging
pump (MSG-ARENA) is also denied by the lockdown, logging a harmless
per-idle-tick warning -- a fully locked-down VM doesn't participate in
message routing.

Not yet applied to the real identity 00 -- this commit is the
mechanism, verified against a disposable test identity only.

Three-arch clean qemu acceptance (single Zuse device, standard
regression case) passed on amd64, aarch64, and riscv64.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 15:41:56 -04:00
Robert Allan JamesandClaude Sonnet 5 56e19a00f9 Route VM-dictionary sf_malloc/sf_free through the kernel's kmalloc heap
Captain Bob's call after seeing the identity-heap-capacity findings
(FABRIC-3.md §X.4): the number of concurrently-running VMs is not known
in advance, and once this is a complete operating system the heap should
be able to use whatever memory is actually available -- not a hardcoded
compile-time ceiling. This was already half-built and just not wired up.

src/starkernel/vm/alloc_kernel.c previously implemented sf_malloc()/
sf_free() (platform_alloc.h's allocator abstraction -- what
vm_create_word() calls for every VM's word dictionary) as its own
isolated static 4MB arena: first-fit free list, no splitting or
coalescing. That's exactly the allocator that topped out around 6
concurrent WIREBIND-born identities, failing from fragmentation before
true capacity exhaustion (§X.4's own measurements).

Sitting right next to it, unused for this purpose: src/starkernel/memory/
kmalloc.c, the kernel's general heap. Already initialized at boot (M6,
kernel_main.c, well before any VM is ever born), reserved from real
PMM-tracked physical memory rather than a fixed array, defaults to a
2 GiB floor explicitly sized "for 256+ baby VMs" per its own comment,
overridable via the --heap= boot flag, and its free list actually
coalesces neighboring blocks on every free.

Change: alloc_kernel.c's sf_malloc()/sf_free() now delegate to
kmalloc_aligned()/kfree() instead of managing a separate arena.
sf_alloc_init() becomes a no-op (kmalloc is already initialized by the
time any VM allocation can happen, and "resetting" a heap now shared by
every kernel subsystem would be actively wrong -- confirmed no external
caller depended on its old reset semantics). sf_alloc_get_stats() reads
kmalloc_get_stats() fresh rather than shadowing byte counts locally;
alloc_count/free_count (which kmalloc.c doesn't track) stay as simple
local counters. sf_calloc()/sf_realloc() are otherwise unchanged. Kernel-
only: the hosted (non-kernel) StarForth build keeps its own separate
alloc_host.c implementation, untouched.

Verified live: replaying the exact hotplug sequence that previously
topped out at 6 identities (Zuse + 8 identities, one at a time via QMP
device_add) now succeeds for all 9, where identity 05 specifically used
to fail. Three-arch clean qemu acceptance (single Zuse device, the
standard regression case) passed on amd64, aarch64, and riscv64 -- one
aarch64 attempt hit an unrelated, already-documented one-off QEMU hiccup
(empty log, boot never progressed past firmware) and passed cleanly on
retry with no rebuild.

Not addressed here: the underlying free-list itself is still first-fit
without splitting (only coalescing changed, inherited from kmalloc.c);
per-VM dictionary sizing (shrinking what each WIREBIND VM's word set
actually needs) is a separate, still-open lever from FABRIC-3.md §X.4's
open architecture question.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 12:46:14 -04:00
Robert Allan JamesandClaude Sonnet 5 d8a195b8d8 WIREBIND: kill the orphaned console VM when the user VM birth fails
Found live during identity-heap-capacity testing (2026-09-07,
hotplugging Zuse + 8 identities one at a time and measuring the kernel
heap arena via a temporary allocator-stats probe, since reverted): every
WIREBIND identity attach births two VMs in sequence -- a "console" VM,
then the real "user" VM. When the second birth failed (arena
fragmentation under concurrent VM load, a separate, not-yet-fixed
capacity issue), capsule_wirebind_try_attach() logged the failure and
returned, but the console VM that had *already succeeded* was never
torn down. It stays live and registered under the identity's username,
consuming its own ~228KB of the fixed 4MB kernel heap arena forever --
nothing ever points a real user at it, since WIREBIND only ever hands
the caller the user VM's id.

This turns every failed identity attach into a permanent net loss of
heap rather than a neutral retry: confirmed live that a failed attach
left the arena 228,576 bytes worse off than before the attempt, and
every subsequent attempt starts from that worse baseline, compounding.

Fix: call capsule_vm_kill(username) on the now-orphaned console VM
before returning from the failure path -- the same teardown
capsule_wirebind_eject()/capsule_wirebind_unclean_detach() already use
elsewhere in this file (vm_cleanup() + sf_free(), confirmed live to
actually reclaim per-word dictionary allocations, FABRIC-3.md
§IX.2/§IX.3).

Verified live with the same allocator-stats probe (written, captured,
reverted -- not part of this commit): after the fix, a forced user-VM
birth failure now returns the arena to exactly its pre-attempt byte
count (3,526,256, matching the baseline precisely) instead of leaking
228,576 bytes. Three-arch clean qemu acceptance (single Zuse device,
the standard regression case) passed on amd64, aarch64, and riscv64.

The underlying capacity/fragmentation question (why the 7th concurrent
identity's arena allocation fails at all despite technically-sufficient
free bytes) is a separate, open architecture question -- not addressed
here. See project memory for the full measured numbers.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 11:37:06 -04:00
Robert Allan JamesandClaude Sonnet 5 e10fb76fb2 xhci: fix Configure Endpoint completion drop under concurrent multi-device enumeration
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Root cause of the FABRIC-3.md §IX.5 follow-on: with 9 devices attached
concurrently at boot (Zuse + 8 identities), only 1 of 9 ever completed
enumeration and reached blkio_usb: MSC device ready -- the other 8 produced
no error and no success, just silence.

xhci_poll_events()'s deferred per-slot dispatch loop submitted a Configure
Endpoint command (a Command Ring op) unconditionally for every slot with
that action pending in a single pass -- unlike every other Command Ring op
in this driver (Enable Slot, Address Device, Disable Slot), which is
correctly gated behind dev->connect_state == XHCI_CONN_IDLE before ever
submitting. With 2+ devices enumerating concurrently, this let multiple
Configure Endpoint commands sit outstanding on the Command Ring at once.
Their completion is correlated purely via the single shared
dev->connect_state field (== XHCI_CONN_AWAIT_CONFIGURE_ENDPOINT), not the
completion event's own Slot ID -- so whichever slot's completion happened
to land while connect_state still read AWAIT_CONFIGURE_ENDPOINT got
correctly chained into SET_CONFIG, and every other slot's completion
arrived after connect_state had already moved on, silently swallowed by
the handler's generic "unrelated command completion" catch-all. No error
path exists for this, which is why it produced total silence rather than
a diagnosable failure.

Root-caused live via temporary WARN-level diagnostic probes (written,
captured, and fully reverted per the project's own probe convention --
this commit contains only the functional fix and its explanatory comment,
no probe code) added at four points: the initial port scan, the connect
handler, the Command Completion Event handler, and the deferred dispatch
loop itself. The probes showed all 9 devices correctly completing Enable
Slot + Address Device (ruling out the connect-state queue as the cause,
the original hypothesis), then all 9 correctly submitting Configure
Endpoint and all 9 commands completing successfully in hardware (code=
SUCCESS, no errors logged) -- but only 1 of 9 ever got its next_action
chained to SET_CONFIG.

Fix: apply the same single-in-flight discipline this driver already uses
for every other Command Ring op. If the Ring isn't free when a slot's
Configure Endpoint action is due, put the action back on that slot instead
of submitting a second command onto a busy Ring -- the next tick's
dispatch pass retries it once the Ring frees up.

Verified live: booting Zuse + all 8 identity drives concurrently (9
devices, one per real xHCI port via the XHCI_PORTS fix from the previous
commit) now produces 9 "MSC device ready" lines and zero xHCI errors,
where it previously produced exactly 1. Three-arch clean qemu acceptance
(single Zuse device, the standard regression case) passed on amd64,
aarch64, and riscv64 -- no change in that baseline behavior.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-07 09:02:45 -04:00
Robert Allan JamesandClaude Sonnet 5 2c1b3cd695 Four bugs found live verifying the 8 identity thumbdrives (FABRIC-3.md §IX)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
All found by actually running the identity workflow §VII/§VIII made
possible, not by code review:

1. Zuse/WIREBIND cross-contamination on detach: capsule_zuse_boot_logout()
   and capsule_wirebind_unclean_detach() both had no device parameter, so
   an unrelated device detaching (while the real owner's own stayed
   attached) incorrectly tore down the wrong session. Both now compare
   the departing device against their own tracked one, mirroring
   capsule_wirebind.c's pre-existing g_wirebind_attached_dev precedent.

2. Dictionary-entry memory leak: vm_create_word()'s sf_malloc()'d
   DictEntry (plus a second per-entry allocation for transition_metrics)
   was never freed by vm_cleanup(), in both the hosted and kernel
   implementations. Caused a real kernel PANIC after 8-9 repeated VM
   birth/kill cycles in one boot. Fixed by walking vm->latest in both.

3. sf_malloc/sf_free (alloc_kernel.c) was a 4MB bump arena with a
   deliberate no-op free, sized on "VM born once, never killed" -- fix #2
   alone didn't stop the panic because free() itself discarded the
   pointer regardless. Given a real free list (first-fit reuse).

4. Headless-console gate didn't re-engage after a mid-boot logout: the
   original fix (sk_console_mark_login(), one-way sticky) only gated the
   first login of the boot. Replaced with a live check
   (sk_console_identity_present()) re-evaluated continuously, including
   inside sk_console_readline()'s own blocking idle loop -- the console
   is normally sitting blocked there when a hot-unplug logout happens, so
   checking only at the top of the REPL loop wasn't enough.

Also: MINT now verifies its own write (verify_mint(), capsule_mint.c) by
reading back through the same check a real attach performs, rather than
trusting blkio_write()'s BLK_OK alone -- logged via log_message(), not
console_println(), per direct instruction.

Verified live, amd64: the full 8-identity repeated attach/detach cycle
that previously panicked at the same point every time now completes
clean, and a full serial-log sweep found zero bare unauthenticated
prompts anywhere in the run. Three-arch clean-qemu acceptance passed.

Still open, not fixed here: a 3+-simultaneous-device USB enumeration
failure found in a separate live test, not yet root-caused.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-06 01:49:13 -04:00
Robert Allan JamesandClaude Sonnet 5 9e81de3f43 xHCI/BOT driver: genuine multi-device support (FABRIC-3.md §VII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Per-slot registry (xhci_msc_slot_t/dev->msc_slots, sized off the
controller's own reported max_slots) replaces the single-device scalar
fields the driver carried since Milestones 2e-2h. Boot-time port scan no
longer stops at the first connected device; a connect/disconnect that
arrives while the Command Ring is busy is now queued and drained instead
of dropped. blkio_usb.c and repl.c's own single-device state (device
descriptor buffers, blkio_dev_t, attach bookkeeping) became per-slot
registries the same way.

Live multi-device testing (not just compiling) surfaced a second, more
severe bug outside the original plan: transfer_purpose and next_action
were also single scalars shared across the whole controller. Two devices
enumerating concurrently could have one's completion silently overwrite
the other's still-outstanding one, permanently stalling it with no error.
Fixed by moving both per-slot and, critically, reading the Transfer Event
TRB's own real Slot ID field instead of trusting external bookkeeping.

Verified live, all three architectures, mandatory clean-qemu acceptance:
existing single-device path unchanged, and two devices attached
simultaneously (amd64) both progress independently through enumeration
without corrupting or stalling each other.

Also in this pass (implemented and verified in earlier turns this
session, committed together per direct instruction):
- Headless-until-login console policy: no prompt/banner until a real
  identity logs in via an attached thumbdrive (WIREBIND or Zuse, neither
  special), reusing EMERGENCY_CONSOLE_ENABLED as the debug/recovery
  escape hatch (now default-off).
- KILL/g_repl_active_vm dangling-pointer fix: killing the VM the console
  is currently USE'd onto now detaches back to Hera first, matching the
  existing EJECT/UNCLEAN precedent.

FABRIC-3.md §VII/§VIII carry full closure notes for all three.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-05 22:14:14 -04:00
Robert Allan JamesandClaude Sonnet 5 4d4ab59189 Build the FIRSTTOUCH overflow trigger flagged in FABRIC-2.md §I.2
The overflow-triggered migration path (a WIREBIND-attached identity's own
drive running low on space) was scoped but never built -- only the
trigger-detection call site was missing, per this section's own text.

- capsule_wirebind.c now tracks the attached blkio_dev* alongside the
  already-tracked VM id, set in try_attach() and cleared in both
  EJECT/UNCLEAN paths.
- New capsule_wirebind_overflow_idle_check(), called once per idle tick
  in repl.c right alongside blk_migration_idle_check() (same cadence):
  reads the attached drive's free/total via blk_get_device_free_blocks(),
  and if free space is below a fixed 10% threshold, extends the
  identity's pool with a one-time blk_firsttouch_claim() of 8 additional
  devblocks on Artemis's system-resident device.
- New blk_owner_has_claim(owner_fp) in block_subsystem.c answers the
  debounce question blk_firsttouch_claim()'s own doc comment had left
  open: a disk scan, not a RAM flag, so the already-extended answer
  survives reboot/reattach, matching BMAPFMT's "ownership travels with
  the block" model.

Premise checked before building (does a WIREBIND-attached drive actually
give a real free/total signal, or does it stay PROVISIONAL/raw): traced
repl.c's attach sequence and confirmed blk_subsys_attach_device() runs on
the same dev pointer right after WIREBIND, and a WIREBIND-eligible drive
is always already STFR/v2-formatted, so the signal is real. Premise held,
unlike the BAM item's overstated one.

Verified with the mandatory 3-arch QEMU acceptance (identical dictionary
hashes, no regression) plus a live logic test of blk_owner_has_claim():
a temporary TEST-OWNER-CLAIM word, run once via SK_CMD and reverted,
confirmed it correctly detects the claiming owner and rejects an
unrelated one. The low-disk-space-triggers-a-claim path itself is not
verified end-to-end -- that needs a real minted WIREBIND-user thumbdrive
with deliberately tiny capacity, out of scope for this pass; noted as
such in the FABRIC-2.md §I.2 closure note rather than overclaimed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-05 16:17:19 -04:00
Robert Allan JamesandClaude Sonnet 5 aa33aedca8 Fix blk_meta_t/BAM accounting reconciliation flagged in FABRIC-2.md §I.2
FIRSTTOUCH ownership (blk_meta_t's BLK_FLAG_CLAIMED/owner_fp) and the
generic block-allocation bitmap (BAM, blk_bam_entry_t) were two parallel,
unreconciled accounting systems: blk_firsttouch_claim() never touched the
BAM, and blk_allocate()/devblock_is_free() never checked BLK_FLAG_CLAIMED.
A FIRSTTOUCH claim could be silently overwritten by a later blk_allocate()
call, or could itself steal a devblock already in ordinary use via
BLOCK/UPDATE.

- devblock_is_free() (shared by blk_firsttouch_claim() and
  blk_migration_idle_check()) now also checks the BAM entries of all
  BLK_PACK_RATIO member LBNs, not just blk_meta_t.
- New devblock_claimed_by_lbn() helper wired into blk_allocate()'s
  free-scan, so it skips any LBN whose devblock is BLK_FLAG_CLAIMED.
- blk_firsttouch_claim() now marks the BAM allocated for all 3 member
  LBNs of each devblock it claims, which also fixes vol_meta.free_blocks
  never decrementing for FIRSTTOUCH claims.
- blk_meta_relocate_devblock() traced and confirmed NOT part of the bug —
  it already keeps BAM in sync via blk_subsys_relocate_block()'s own
  blk_mark_free()/blk_update() calls.

Both boundary cases (the reserved/user LBN split at a slot's start_lbn,
and BAM-array bounds) are guarded explicitly.

Verified with the mandatory 3-arch QEMU acceptance (identical dictionary
hashes, clean BYE) plus a live logic test: a temporary TEST-BAM-RECON
word, run once via SK_CMD and fully reverted, confirmed on running code
that ordinary allocation and a FIRSTTOUCH claim land on disjoint LBN
ranges in both directions. FABRIC-2.md §I.2 updated in place with the
closure note, per this project's documentation discipline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-05 15:57:03 -04:00
Claude d6895bdd15 Move vocabulary and control-flow state off file-scope statics onto VM
proof/FINDINGS.md's Isabelle/HOL word-source sweep (§1) found the two
defects severe enough to actively corrupt the live Tripod multi-VM fleet:
file-scope C statics standing in for state that belongs on struct VM.

- vocabulary_words.c (highest severity in the sweep): forth_vocab/
  context_vocab/current_vocab, context_var_addr/current_var_addr, the
  ctx_fc/forth_fc first-char search index, and the `initialized` guard
  were all process-wide statics. Only the first VM to touch any
  vocabulary word ever ran setup; every VM after that silently shared
  VM #1's dictionary-chain pointers and reused VM #1's byte-offset
  addresses as if valid in its own vm->memory. One VM's VOCABULARY/
  DEFINITIONS/FORTH silently changed where every other VM looked up and
  defined words.

- control_words.c: cf_stack/cf_sp/cf_last_mode (IF/THEN/BEGIN/DO/CASE
  compile-time nesting) and the LEAVE/ENDOF patch-site bookkeeping
  (leave_addrs/leave_sp/leave_mark_*, endof_addrs/endof_sp/endof_mark_*)
  were also process-wide statics. Two VMs compiling colon definitions at
  overlapping times would corrupt each other's nesting state.

Both moved onto struct VM, following the existing hold_addr/hold_pos
precedent in include/vm.h ("lives in each VM's own memory... so child
VMs never alias Hera's buffer"):

- New VocabularyState struct (vm->vocab): chain heads, VM-cell addresses,
  first-char index, initialized flag.
- New ControlFlowState struct (vm->cf): cf_stack/cf_sp/cf_last_mode plus
  the LEAVE/ENDOF patch-site stacks. cf_tag_t/cf_item_t/CF_STACK_MAX
  moved from control_words.c into include/vm.h since they're now part of
  the struct VM field's type.
- Sentinel fields (-1/-999, meaning "empty") explicitly initialized in
  both vm_init_with_host() implementations (hosted src/vm_bootstrap.c and
  kernel src/starkernel/vm/vm_bootstrap.c) alongside the existing
  dsp/rsp = -1 initialization, since the preceding zero-init leaves them
  at 0 rather than their empty sentinel.

Every word function in both files already took VM *vm, so no call sites
outside these two files needed to change; cf_push_item/cf_pop_item/
cf_peek_item gained a VM* parameter to reach vm->cf.

Verified: hosted (amd64) and kernel (amd64, __STARKERNEL__) both build
clean with -Wall -Werror after a full clean rebuild (struct VM's layout
changed size, and this Makefile has no header-dependency tracking, so a
stale incremental build would have linked mismatched object layouts).
Hosted POST suite 1012/1012 passing (0 regressions). Manually exercised
VOCABULARY/DEFINITIONS/FORTH/ORDER, and IF/ELSE, DO/LOOP/LEAVE,
BEGIN/WHILE/REPEAT, and CASE/OF/ENDOF/ENDCASE (including nested DO with
I/J) in the REPL -- all correct and unchanged from pre-refactor behavior.

Note: a pre-existing CASE/ENDCASE default-clause bug (the code after the
last OF...ENDOF pair does not correctly become the "default" value once
DROP runs) was found while testing this refactor and confirmed present
on unmodified master too -- not touched here, out of scope for this pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Qf6YcnHgaEtEygq3knx19
2026-09-05 14:07:41 +00:00
Claude c36bd99e1e Fix EXECUTE/?/DUMP/TYPE/DECIMAL-HEX-OCTAL/ALIGN defects from proof sweep
proof/FINDINGS.md's Isabelle/HOL word-source sweep (§4) flagged five real
defects; this fixes all five and records resolution in that doc:

- EXECUTE (system_words.c): cast a popped cell straight to a DictEntry*
  and called through it with only a null check. Now validates via a new
  shared vm_dict_entry_ok(), promoted out of starforth_words.c's
  ENTROPY@/ENTROPY! guard (dictionary_management.c) so EXECUTE gets the
  same live-entry check.

- ? and DUMP (format_words.c): dereferenced the popped cell as a raw host
  pointer, bypassing vm_addr_ok entirely (out-of-VM-bounds read). Both now
  go through VM_ADDR/vm_addr_ok/vm_load_cell/vm_ptr like every other
  memory word (@, `,`, editor_words.c).

- TYPE (io_words.c): bounds check computed addr+count in signed 64-bit
  arithmetic, which can overflow and bypass the check on large operands.
  Replaced with vm_addr_ok(), which is written to avoid that overflow.

- DECIMAL/HEX/OCTAL (format_words.c): wrote only the BASE memory cell,
  never vm->base, the host-mirror field number-output words actually read
  via current_base() -- so these words silently affected number parsing
  but never printing. Now call the existing vm_set_base() (previously
  only used at boot init), which updates both. vm_get_base/vm_set_base
  promoted to public declarations in include/vm.h.

- ALIGN vs ALLOT/,/C,/2, (dictionary_words.c): disagreed on dictionary
  growth ceiling (2MB vs 5MB). Investigated which was correct rather than
  blindly widening: vm_get_block_addr() maps block N to
  vm->memory + N*BLOCK_SIZE across the full 5MB arena, and
  USER_BLOCKS_START (block 2048) lines up exactly with
  DICTIONARY_MEMORY_SIZE -- so ALLOT/,/C,/2, letting `here` grow past 2MB
  could silently corrupt live block/user data sharing that memory.
  Tightened ALLOT/,/C,/2, to DICTIONARY_MEMORY_SIZE to match ALIGN.

Verified: hosted (amd64) and kernel (amd64, __STARKERNEL__) both build
clean with -Wall -Werror; hosted POST suite 1012/1012 passing (0
regressions); manually exercised EXECUTE, ?/DUMP, TYPE, HEX/DECIMAL/OCTAL,
and large-ALLOT rejection in the REPL.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Qf6YcnHgaEtEygq3knx19
2026-09-05 14:07:11 +00:00
Robert Allan JamesandClaude Sonnet 5 70dc8beba4 FABRIC-2.md §I.9 follow-on: blinking | cursor instead of static block
Build / build-riscv64-img (push) Canceled after 0s
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Captain Bob asked for the framebuffer cursor to render as a blinking
vertical bar rather than the previous static solid-block glyph.

vt100_draw_cursor() (hal/vt100.c) now fills a thin bar (cell_w()/8, min
1px, full cell height) at the cursor's left edge instead of the whole
cell -- an I-beam shape. vt100_erase_cursor() is unchanged (clearing the
whole cell already safely covers the narrower bar).

Blinking is new in repl.c: sk_console_readline()'s idle branch toggles the
cursor on/off every SK_CURSOR_BLINK_INTERVAL (50 ticks, 500ms at 100Hz)
via alternating console_fb_draw_cursor()/console_fb_erase_cursor() calls,
independent of the heartbeat/idle-beat mechanism the §I.9 fix just touched
(deliberately not reused, to avoid recoupling to that path). Runs
regardless of n, so it blinks whether sitting at a bare prompt or paused
mid-edit. Every deterministic draw site (initial prompt, prompt reanchor,
backspace, character echo) now goes through a new helper, sk_cursor_show(),
which resets the blink cycle to "on" and redraws -- typing always shows a
solid cursor, never mid-blink.

Verified via the mandatory foreground 3-arch QEMU acceptance boot: amd64
(logs/20260905-021054, extensive live interactive typing including
multi-line : / ; word definitions and error cases, prompts stayed
correctly attached throughout), aarch64 (logs/20260905-021551), riscv64
(logs/20260905-022324) -- all three reached (zuse) ok> and shut down
cleanly via BYE.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var
2026-09-05 02:24:51 -04:00
Robert Allan JamesandClaude Sonnet 5 8edb95b65d FABRIC-2.md §I.9: fix the terminal phantom-linebreak defect
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
console_ensure_line_start() (hal/console.c) used to emit its newline
immediately via a path that deliberately skipped the tx-byte counter, so
that sk_repl_idle()'s unconditional per-beat call to it (repl.c, ~1s idle
heartbeat) could force a real newline the REPL's own prompt-reanchor logic
never noticed -- the prompt was never reprinted, and the next real
keystroke echoed onto the now-blank line, indistinguishable from Enter
having already been pressed at a bare prompt. Root-caused in the previous
commit (704573b); this commit applies the fix per explicit go-ahead.

Fix: defer the newline instead of emitting it eagerly.
console_ensure_line_start() now only sets a flag (g_pending_line_close);
the newline is realized -- for real, and counted by g_console_tx_count
like any other output -- on the next actual console_putc() call, or
silently discarded via the new console_cancel_deferred_line_start() if the
caller decides nothing was actually printed. sk_repl_idle() captures
tx_before_idle right after its console_ensure_line_start() call and cancels
the deferred newline at both of its exit points when console_tx_count()
hasn't moved. A beat with nothing to report now leaves the console
untouched; a beat that does print still closes the dangling prompt line
first, properly counted this time. The pre-existing n > 0 mid-edit gate in
sk_console_readline() is untouched -- independent purpose, not the bug.

Verified via the mandatory foreground 3-arch QEMU acceptance boot
(clean qemu, amd64 -> aarch64 -> riscv64, one at a time): all three reached
(zuse) ok>, all echoed the first real input on the same log line as the
prompt rather than a fresh line, all shut down cleanly via BYE. amd64's log
additionally shows live human backspace-correction still glued to the same
prompt line. Logs: logs/20260905-015854 (amd64), logs/20260905-020248
(aarch64), logs/20260905-020552 (riscv64); logs/20260905-015813 is a
foreground-rule-violation retry killed and redone correctly, kept per this
project's "never delete logs/" convention.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var
2026-09-05 02:08:06 -04:00
Robert Allan JamesandClaude Sonnet 5 dbaead0af2 xhci: route driver chatter through log_message(), silence at default log level
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
xhci.c and repl.c's USB attach/detach path printed every xHCI command
submission, completion, and BOT transfer step unconditionally via
console_println()/console_puts() -- floods the serial log on every boot,
regardless of whether anyone is debugging the USB stack.

Converted every "xhci:"-prefixed line to log_message() with a level
chosen by what it reports, not blanket debug:
- LOG_ERROR: allocation/mapping failures, timeouts, command failures,
  CSW signature/tag mismatches, CSW FAILED/PHASE ERROR, "not implemented"
  refusals, every "deferred ... setup failed" path
- LOG_WARN: dropped/skipped conditions (command ring busy, tracked-port
  range exceeded), unrecognized media (bad version/CRC), TUR retry
- LOG_DEBUG: routine progress (command submitted, succeeded, transfer
  completed, port connected) and expected outcomes (recognized/blank
  media)

Default log level is LOG_INFO, so the LOG_DEBUG chatter that was the
actual complaint is now silent by default and re-enabled with
--log-level=debug; LOG_ERROR/LOG_WARN stay visible so real faults aren't
buried.

xhci_log_hex32() now routes through log_message(LOG_DEBUG, ...) instead
of console_puts()/console_println() directly -- kept its own zero-padded
8-digit hex formatting rather than switching to log_message()'s %x
(which has no width control), since register values lining up in the
log is the reason this helper exists. console.h dropped from xhci.c,
no longer used directly.

Verified functionally unchanged, not just "still boots": all three
architectures reach zuse)ok>, log_message()-instrumented lines are gone
from the default-level log (grep -c xhci == 0 on all three, versus dozens
before), and the USB thumbdrive path still works end to end --
"Zuse: identity confirmed from attached thumbdrive" appears on all three
boots exactly as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var
2026-09-05 01:25:50 -04:00
Robert Allan JamesandClaude Sonnet 5 9b6de5d6c7 riscv64: PLIC base address DTB-discovered, QEMU-virt constant as fallback (§V.3 item 3)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
plic_init() now takes boot_info->dtb (threaded through apic_init()) and
tries fdt_find_node_by_compatible(dtb, "sifive,plic-1.0.0") ->
fdt_find_prop_in_node(..., "reg", ...) before falling back to the
QEMU-virt-specific constant it previously hardcoded unconditionally.
Reuses the node-scoped DTB lookup primitive built for the aarch64 GIC
base fix unchanged. s_plic_base is now a runtime uintptr_t, same shape
as apic.c's s_gicd_base/s_gicc_base.

This system's QEMU/UEFI riscv64 firmware does not forward a DTB to the
guest (timer.c's own timebase-frequency read falls back too, confirmed
in this boot's own log), so only the no-DTB fallback branch is exercised
here -- the success branch (a real DTB with a matching PLIC node) stays
unverified until real Milk-V Mars hardware. FABRIC-3.md's first-drafted
claim that the success branch would run (based on a stale comment in
plic.c's own pre-fix header) was checked against the actual log and
corrected before this commit.

3-arch acceptance: amd64/aarch64 don't compile these files, so their
runs are non-regression on untouched files only. riscv64's own boot log
confirms the fallback path prints exactly as designed and boot reaches
zuse)ok> unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var
2026-09-05 00:56:25 -04:00
Robert Allan JamesandClaude Sonnet 5 90ee8deb6d riscv64: skip satp Bare-mode switch when already Bare (§V.3 item 7 audit fix)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
arch_early_init() unconditionally cleared satp on every boot, justified
only by behavior observed under QEMU/EDK2 firmware (satp.MODE=10/Sv57,
kernel identity-mapped within it). That reasoning never applied to the
native U-Boot+OpenSBI boot path, where satp is conventionally already 0
at S-mode handoff -- the unconditional clear was likely a harmless no-op
there, but on an unverified assumption.

Fix: read satp.MODE first and return early when it's already 0 (nothing
to switch away from, no safety argument needed). The unconditional
csrw/sfence pair still runs unchanged for the confirmed QEMU/EDK2 case.
No Sv39/Sv48/Sv57 page-table walker built -- out of proportion to this
finding's severity.

3-arch acceptance: amd64/aarch64 don't compile this file, so their runs
are non-regression on untouched files only. riscv64's own boot log
confirms satp.MODE = 0xa at entry, so the mode != 0 branch ran and
"satp cleared -- Bare mode, explicit" printed exactly as before the fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var
2026-09-05 00:17:35 -04:00