Commit Graph
4 Commits
Author SHA1 Message Date
Robert Allan JamesandClaude Sonnet 5 51940f49d0 Add per-slot CRC + campaign crash-resume (DOE-RESUME-COUNT)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
FABRIC-3.md §XXXV.12: closes two real gaps found by walking the disk-to-
PDF pipeline end to end. (1) doe_trial_slot_t had no integrity check of
its own -- 64 slots share one devblock via a read-modify-write, so a
crash mid-write to any slot could silently corrupt its siblings with
nothing to catch it. Added an 8-byte CRC (same discipline the control
header already had), computed/verified the same way. (2) The campaign
itself had no crash-resume -- a reboot always restarted the whole
18-trial loop from index 0 even with good persisted data already on
disk. Since the trial matrix is fully deterministic from (seed, n-reps),
resuming needs no persisted "which trials ran" state, just a count:
DOE-RESUME-COUNT (new kernel primitive, same pattern as
DOE-PERSIST-TRIAL) walks the ring counting CRC-valid records for a given
ISA; MU-EXEC-CAMPAIGN's signature changes to accept a start-idx and seeds
MU-RUN-ID/the DO loop from it. Each RUN-*-CAMPAIGN entry word now calls
DOE-RESUME-COUNT before launching. Both block-namespace-checked to fit
existing headroom before committing to the design -- no new blocks
needed.

Real bug caught by the build, not shipped: an implicit compiler-inserted
alignment byte before slot_crc silently broke the hand-computed trailing
pad size (doe_trial_slot_size_check failed to compile). Fixed with an
explicit alignment field, matching log_slot_t's own zero-implicit-
padding discipline.

extract_doe_region.py now verifies both CRCs with the REAL CRC-64/XZ
algorithm from block_subsystem.c's compute_crc64() (ported and verified
bit-for-bit against a native C reference this session), replacing
§XXXV.11's incorrect guessed implementation.

Verified: clean compile on all 3 ISAs (amd64/aarch64/riscv64 -- amd64
verified via isolated build only, not boot, since the real per-ISA
campaign is currently live on the only QEMU instance this project's own
rule allows), both capsules pass mkcapsule --lint, CRC64 bit-exact via
native reference. Live boot-verification explicitly deferred until that
campaign finishes or is interrupted -- not claimed as tested when it
wasn't. Found (not missed) a real compatibility consequence: the running
campaign's data is in the old pre-CRC format, so it reads as "corrupt"
under the new extractor and DOE-RESUME-COUNT will resume it from 0, not
mid-campaign -- no data loss, cheap redundant re-run of ~1-2 trials.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-17 19:17:04 -04:00
Robert Allan JamesandClaude Sonnet 5 cd2fda4351 Add mkcapsule --resolve: build-time claim registry for capsule block collisions (FABRIC-3.md §XXIV)
Traced what a "Block NNNN" collision actually means before designing
a fix for it: capsule_loader.c's block-write path routes through the
generic block-subsystem API, which kernel_main.c registers as two
devices in a fixed order -- the volatile ramdrive first (LBN
2048-3071), then Artemis's real virtio-blk device immediately after
(LBN 3072+, backed by disk/artemis.img). Every capsule this project
has lands in Artemis's persistent range, not the ramdrive, and
blk_update()'s dirty-marking + repl.c's idle-loop flush write that
content through to the real disk file on every boot. A block-number
collision is therefore a silent, persistent overwrite of real disk
content surviving reboots, not a transient RAM mixup.

The existing collision gate (check_block_conflicts(), already a hard
non-interactive build failure) already catches capsule-vs-capsule
collisions across the whole flat range. The real gap: zero visibility
into blocks something other than a capsule owns (Artemis's own
non-capsule persistent data), and no device-boundary/capacity
awareness at all.

Added, scoped step by step before writing any code:

- tools/capsule-claims.txt -- derived, auto-created/regenerated,
  git-ignored. Lets --resolve tell "this capsule's own content
  changed" apart from "genuinely new collision with something else."
- tools/capsule-reserved.txt -- human-authored, git-tracked, seeded
  with nothing yet rather than guessed at. Checked by both the plain
  build gate (new check_reserved_conflicts()) and --resolve.
- tools/patches/ -- git-tracked, one file per accepted interactive
  renumber; a structured old->new block list, not a generic diff,
  since that's the only thing a renumber ever changes.
- mkcapsule --resolve <dir> -- the only interactive mkcapsule mode,
  a deliberate separate invocation from the plain build path (which
  stays non-interactive so CI never blocks on a prompt). Suggests a
  renumbering that preserves a capsule's own existing block spacing,
  prompts y/N, rewrites the .4th source in place on acceptance.

Found and fixed a real bug during verification: the registry's
empty-block-list case (workload-5.4th, zero Block headers) serialized
with a stray trailing space that the reader parsed back as a phantom
block 0, causing spurious re-registration every run -- caught by
testing idempotency directly, not assuming a clean first run meant
it worked.

Verified: isolated collision tests confirm both accept and reject
paths, confirm a resolved collision doesn't re-prompt the other side,
confirm reserved-range collisions are caught by both --resolve and
the plain build gate. Full 3-architecture rebuild via the real
Makefile.starkernel succeeded clean; amd64 boots with an unchanged
dict_hash/capsule_hash from every prior boot this session.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 21:56:37 -04:00
Robert Allan James e58381c121 fb/: untrack and gitignore -- throwaway local verification images
Reversing the earlier decision to track fb/ in git (made when first
setting up the directory for item 4.3.2). Captain Bob: these are
disposable screenshots for eyeballing framebuffer output during Console
work, never meant to be committed. Removed from git tracking (git rm
--cached) and added to .gitignore; files that still exist locally are
untouched, deletions already made locally are left as-is.

Also fixes item 4.3.4's amd64-only blind spot found in the process:
aarch64/riscv64 had no framebuffer device at all (GOP: protocol not
found) -- the "all three architectures boot clean" checks run all session
were REPL/dict_hash parity, a different thing from GOP presence, and
conflating the two was an error. Added -device ramfb (EDK2's
firmware-only GOP framebuffer) to both architectures' qemu targets in
Makefile.starkernel. Both now report GOP: linear framebuffer found at
800x600; cube rendering verified correct on both (screenshots not
committed, per the untrack above -- verified visually this session).
Standard three-arch acceptance boot re-run afterward, all clean,
dict_hash identical and unchanged from before this fix.
2026-08-07 20:13:17 -04:00
Robert Allan James a5ed8c3d87 Initial commit — LithosAnanke kernel 2026-08-01 07:49:56 -04:00