diff --git a/FABRIC-3.5.md b/FABRIC-3.5.md index 3f5b2fce..09f7b4b2 100644 --- a/FABRIC-3.5.md +++ b/FABRIC-3.5.md @@ -3460,3 +3460,126 @@ alloc/free/evict under the new arbiter, the rest is protocol plumbing. **Punch list:** ⬜ **Item 27 — rule on §XXXIII.5**: does kernel-Hermes implement channel negotiation, or one broadcast membership? ⬜ **Item 28 — build and prove the heat-coupled allocator first**, verified against `fleet_conserved`, before any protocol work. + +--- + +## XXXIV. The transition state: how two messaging layers coexist without corrupting K (2026-09-19) + +**The question nobody asked.** §XXII.1 draws the ordering distinction that protects the strip — +Category B (`hermes/init.4th`, most of `messaging.4th`, the routing table, the slot-3 pairing) +is "load-bearing until the replacement boots." That sentence implies a window in which **both +messaging layers are live**, and nothing in this document says how they behave in it. + +### XXXIV.1 — The real hazard is not control. It is heat. + +Item 22 (§XXXII) dealt with two mechanisms moving *control*. The transition creates the same +shape one layer down: **two allocators drawing on the same per-VM Stadium reservoir, both +counted in K.** + +The budget is explicit. `messaging.4th` derives it: + +``` +Q.SLOT = (Q.1 - Q.1/3) / (MSG-MAX + CH-MAX - 1) +``` + +— each VM's reservoir (`Q.1` = `Q48_ONE`, per `stadium.c`'s "item 4.1's reservoir starts each +VM's quota at 1.0"), less `COMMON-CH`'s `Q.1/3` floor, split across **32 messages + 15 channel +slots = 47 items.** That pool is sized for *one* messaging layer. + +**Conservation itself is not at risk** — K holds as long as each layer honestly pulls and +returns, regardless of who holds the heat. **What is at risk is the budget**: two layers each +sized to consume the same reservoir is oversubscription, and the failure mode is allocation +refusals under load, not silent drift. `fleet_conserved` makes it visible rather than quiet, +which is the saving grace. + +**A helpful interaction with §XXXIII:** retiring the channel abstraction removes `CH-MAX - 1` += **15 of those 47 slots**, so the divisor falls from 47 to 32 and per-message heat rises by +~47%. **Deleting the unused generality recovers roughly a third of the per-VM message budget** +— useful headroom precisely during the window when two layers share it. + +### XXXIV.2 — RULING: partition by message type. One owner per message, never shared. + +> **During the transition, every message type is owned by exactly one layer. A given message is +> allocated, held, delivered and released by that layer alone. No message crosses.** + +This is the heat-side analogue of §XXXII's control-side ruling, and it has the same effect: +**two systems coexist without any shared object.** Double-accounting becomes structurally +impossible rather than carefully avoided — there is no message for both layers to charge for. + +### XXXIV.3 — The staging, reusing §XXVIII's own discipline rather than inventing one + +`FABRIC-3.md` §XXVIII faced the structurally identical problem — standing up a second +control-transfer mechanism beside a live one — and solved it with "**explicit go/no-go gates, +not attempted as one pass**," each stage inert until the next enables it: Stage 1 allocated +per-VM stacks with "nothing executes on them yet"; Stage 2 built the switch primitive +"cooperative only, no timer"; Stage 3 turned on the timer. One commit per increment, full +three-architecture acceptance before the next begins. + +**That discipline is the answer here, applied to messaging:** + +| Stage | Content | Live layers | +|---|---|---| +| **A** | Kernel-Hermes structures + heat-coupled allocator. **Wired to nothing, draws no heat.** Boot is byte-identical; `dict_hash` unmoved (no FORTH edit) | FORTH only | +| **B** | **Prove the allocator** (item 28): a test path allocates and frees N messages, `fleet_conserved` holds across it. Heat drawn and returned within the test | FORTH only | +| **C** | **Cut over exactly one message type** — the narrowest live one. `BLK-ATTACH-EVENT` (a single Hera↔Artemis request/ack pair) is the natural candidate | Both, **disjoint** | +| **D** | Remaining live types cut over one at a time | Both, disjoint, shrinking | +| **E** | Category B strip: `hermes/init.4th`, `messaging.4th`, routing table, slot-3 pairing | Kernel only | + +**Stage A is the one that makes this safe**, and it is the stage most likely to feel like +wasted work. It is not: it is what makes every later stage a revert of one commit. + +### XXXIV.4 — A coupling nobody noticed: the Tripod change cannot land before the cutover + +§XIX.6 puts Hestia into `is_fleet_foundation` (`capsule_birth.c:793-796`) and replaces Hermes's +birth at `kernel_main.c:865`. **But FORTH Hermes must stay alive through Stages C and D** — it +is still routing every type not yet cut over. + +**So the two changes are sequenced, not simultaneous, and the intermediate state is explicit:** + +- `is_fleet_foundation` **temporarily holds four names** — Hera, Artemis, Hestia **and** Hermes. + That line is an OR-chain of prefix comparisons, so a fourth is trivial; what matters is that + it is *deliberate and temporary*, not an oversight for a later reader to "clean up." +- **Both births run** at `kernel_main.c`: Hestia's is added, Hermes's is not yet removed. +- Hermes's birth and its `is_fleet_foundation` entry are removed **in Stage E, with the strip**, + not with the Tripod change. + +**This contradicts nothing in §XIX — it sequences it.** Worth writing down because the obvious +reading of §XIX.6 ("swap Hermes for Hestia") would break Stages C and D on the first boot. + +### XXXIV.5 — Cost and rollback, honestly + +**Rollback is one commit per stage**, which is the entire reason for staging. Tags remain the +oracle (`.claude/CLAUDE.md`: "when in doubt about the correct state of any file or branch, look +at the tag first"). + +**The cost is acceptance runs.** Per `.claude/CLAUDE.md` the only valid acceptance is all three +architectures booted in QEMU, one at a time, in the foreground, `clean` before `qemu`. Stages +A–E plus the Category B strips is **at least eight full three-architecture cycles**, each with +logs committed. That is the price of the discipline, and it is the same price §XXVIII paid +across Stages 0–4. **Naming it now so it is a plan rather than a surprise** — amd64 under TCG is +the slow one, and `.claude/CLAUDE.md` warns the session must stay engaged through long runs. + +### XXXIV.6 — What would make me stop and re-plan + +Stated as tripwires rather than hopes, since the point of this pass is confidence: + +- **Stage B fails** — `fleet_conserved` will not hold across a bare alloc/free cycle under the + new allocator. That is the piece §XXXIII named as the only genuinely hard one, and failing it + early is the cheap outcome. **Do not proceed to C.** +- **Stage C shows allocation refusals** that the FORTH-only baseline did not. That is the + oversubscription of §XXXIV.1 arriving, and the answer is to retire channels first (§XXXIII, + recovering ~a third of the budget) rather than to push on. +- **`dict_hash` diverges across architectures** at any stage. Not "changes" — changing is + expected and normal (§III.5). **Diverging between amd64/aarch64/riscv64 is the signal that + something non-deterministic entered the dictionary**, and it is the one failure this + project's acceptance criteria are specifically built to catch. + +### XXXIV.7 — Punch list + +- ⬜ **Item 29 — adopt §XXXIV.2's partition rule and §XXXIV.3's Stage A–E staging** as the + kernel-Hermes build plan (replacing item 6's single line). +- ⬜ **Item 30 — sequence the Tripod change per §XXXIV.4**: Hestia added alongside Hermes, + `is_fleet_foundation` temporarily four names, Hermes removed only at Stage E. +- Item 28 (prove the allocator) becomes **Stage B** and keeps its priority. +- Item 27 (channels: negotiate or one membership) **gains urgency** — §XXXIV.1 shows retiring + channels directly relieves the transition's budget pressure.