Publish path, no dispatch -- FABRIC-3.6.md task 3.3

sk_hermes_publish() allocates one SkHermesMessage per channel member
(heat-cost ruling 2026-09-21: one message per subscriber, funded by
the publisher's own reservoir) and enqueues each onto a new
per-subscriber SkHermesPendingQueue -- found-or-created lazily by
vm_id, sized from stadium_max_vm_count() like the channel/switch
tables. Best-effort across subscribers: a failed allocation or full
queue skips and rolls back just that one subscriber, not the whole
publish -- the natural reading of "ledger and stadium_conserved() hold
across N publishes to M subscribers" (the task's own check), not a
separate ruling.

Dispatches nothing -- sk_hermes_pending_count()/peek()/pop() are the
read/drain primitives task 3.4's real checkpoint-driven drain will
build on; this task's own self-test uses them directly since no
checkpoint hook exists yet.

Verified live on all three architectures: pending-queue table sized
50/202/50 slots (tracking the channel table's own per-arch sizing), a
synthetic publish self-test (2 publishes to 3 subscribers) confirms
exact per-subscriber delivery counts, and ledger/stadium_conserved()
invariants hold both mid-publish and after manually draining every
queue back to baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Robert Allan James
2026-09-22 00:49:22 -04:00
co-authored by Claude Sonnet 5
parent a4afdfa591
commit 1f6343bc03
7 changed files with 28105 additions and 1 deletions
+27 -1
View File
@@ -830,12 +830,38 @@ ruling]** cannot be written precisely until Captain Bob settles the named sub-it
architectures. `dict_hash` for Hermes (`0xc95ef2d92fa0781f`) and Hestia
(`0xfbde9fe105fd3b3d`) identical across all three architectures, unmoved from task 3.1's
values (this task touches no capsule).
- [ ] **3.3** — **Publish path, no dispatch** (§XLIII.3). Publish allocates via
- [x] **3.3** — **Publish path, no dispatch** (§XLIII.3). Publish allocates via
`sk_hermes_alloc()`, enqueues onto each subscriber's own pending queue, and records the
fact; it dispatches nothing. Self-test only. *Check:* ledger audit and
`stadium_conserved()` hold across N publishes to M subscribers (heat cost per subscriber
is a design point to settle **before** writing this: one message per subscriber, or one
shared message with a reference count — **[needs ruling]**).
2026-09-22 · `logs/20260922-002857/amd64/`, `logs/20260922-003404/aarch64/`,
`logs/20260922-003922/riscv64/` — all three reach `[zuse@Hera] ok>`, zero `UNKNOWN WORD`,
`mkcapsule --lint capsules/` clean (38 files, 0 violations, unchanged -- no capsule
touched). Heat cost ruled 2026-09-21 (one message per subscriber); `sk_hermes_publish()`
allocates one `SkHermesMessage` per channel member via `sk_hermes_alloc()` (funded by the
publisher's reservoir) and enqueues each onto a new per-subscriber `SkHermesPendingQueue`
(found-or-created lazily by `vm_id`, same shape as `stadium.c`'s `quota_slot_for_vm()` --
not pre-populated at birth like channel membership is, since not every VM ever receives a
message). Table sized from `stadium_max_vm_count()`, same pattern as tasks 3.1/3.2.
Best-effort, not atomic across subscribers -- a failed allocation or full queue skips just
that one subscriber and rolls back its own allocation; not a separate ruling, the natural
reading of the task's own check (ledger/`stadium_conserved()` invariants hold under partial
delivery too), documented as such in the header. Added `sk_hermes_pending_count()`/
`_peek()`/`_pop()` as the read/drain primitives task 3.4's real checkpoint-driven drain
will build on -- this task's own self-test uses them directly for cleanup since no
checkpoint hook exists yet. New console line confirmed live on all three architectures:
`Kernel-Hermes: N pending-queue slots` (**50 amd64, 202 aarch64, 50 riscv64** -- tracks
the channel/switch-signal tables' own per-arch sizing exactly). Self-test (synthetic
publisher lo=7, three synthetic subscribers lo=8/9/10, N=2 publishes to a 3-member
synthetic channel) confirms `sk_hermes_publish()` returns 3 each time, each subscriber's
queue holds exactly 2 afterward, ledger audit and `stadium_conserved(publisher)` hold
mid-publish, then confirms the same invariants return to baseline after draining every
queue by hand (`pending_peek()`/`release()`/`pop()`) -- `PASS` on all three architectures.
`dict_hash` for Hermes (`0xc95ef2d92fa0781f`) and Hestia (`0xfbde9fe105fd3b3d`) identical
across all three architectures, unmoved from tasks 3.1/3.2's values (this task touches no
capsule).
- [ ] **3.4** — **Drain at the outermost checkpoint** (§XLIII.3–.5). Reuse
`sk_vm_at_outermost_interpret()`; one message per checkpoint **[needs ruling 3.0e]**.
*Check:* a nested interpret does not drain; a queued payload is interpreted exactly once at
+77
View File
@@ -374,6 +374,83 @@ int sk_hermes_channel_member_count(int channel_id);
* DoE/test observability, mirrors sk_vm_switch_signal_slot_capacity(). */
int sk_hermes_channel_capacity(void);
/* Boot-time allocation for the per-subscriber pending-queue table
* (FABRIC-3.6.md task 3.3), kmalloc'd to stadium_max_vm_count() entries --
* same sizing pattern as the channel table (task 3.2) and switch table
* (task 3.1); see kernel_hermes.c's own comment on this call for why. Must
* run after stadium_boot_init(). Soft failure -- returns -1 and leaves the
* table unallocated (every queue lookup then finds nothing) rather than
* halting boot. Idempotent-unsafe: call exactly once. */
int sk_hermes_queues_boot_init(void);
/*
* Publish path, no dispatch (FABRIC-3.6.md task 3.3, SXLIII.3; heat-cost
* ruling 2026-09-21: one message per subscriber, separate heat draw each --
* matches the existing heat-coupled allocator 1:1, no refcount machinery).
*
* sk_hermes_publish() allocates one SkHermesMessage per channel member
* (via sk_hermes_alloc(), funded by the publisher's own reservoir) and
* enqueues each onto that member's own pending queue -- FIFO, one queue
* per subscriber VM, found-or-created lazily on first use (same
* find-or-create-by-vm_id shape stadium.c's quota_slot_for_vm() and
* session.c's session_find() already use). It does not interpret,
* deliver, or otherwise dispatch anything -- draining a queue at a VM's
* own outermost interpret checkpoint is task 3.4's scope, not this one's.
*
* Best-effort, not atomic across subscribers: if a given subscriber's
* allocation or enqueue fails (reservoir exhausted, message arena full,
* or that subscriber's own pending queue full), that one subscriber is
* skipped -- the message already allocated for a failed enqueue is
* released back (rolled back) rather than left orphaned, but delivery to
* every OTHER subscriber already queued is not undone. This was not a
* separate Captain Bob ruling; it is the natural reading of "ledger audit
* and stadium_conserved() hold across N publishes to M subscribers" (task
* 3.3's own check) -- those invariants hold under partial delivery just
* as well as under all-or-nothing, and requiring atomicity across M
* independent reservoir-funded allocations would need a two-phase
* commit/rollback this task's inert scope does not call for.
*
* @param from Publisher, whose reservoir funds every allocation.
* @param channel_id Target channel (SK_HERMES_CHANNEL_COMMON or a
* channel from sk_hermes_channel_create()). Refused
* (-1) if invalid/not in use.
* @param type Message type code, passed through unchanged.
* @param payload_addr Out-of-line payload address, passed through
* unchanged (bound/chunking is task 3.5's scope, not
* this one's -- 3.3 does not enforce a payload size
* limit).
* @param payload_len Payload length in bytes, passed through unchanged.
* @return Count of subscribers successfully enqueued to (0..member count),
* or -1 if channel_id itself was invalid.
*/
int sk_hermes_publish(VMUuid from, int channel_id, uint32_t type,
void *payload_addr, uint32_t payload_len);
/* SK_HERMES_PENDING_MAX - per-subscriber pending-queue depth. The global
* message arena (SK_HERMES_MSG_MAX) is the real ceiling on how many
* messages can ever be in flight system-wide, so sizing each VM's own
* queue to that same bound is a safe, simple upper limit rather than a
* new number to justify. */
#define SK_HERMES_PENDING_MAX SK_HERMES_MSG_MAX
/* sk_hermes_pending_count - number of messages currently queued for
* vm_id (0 if vm_id has no queue yet -- never having received a message
* is not an error). */
int sk_hermes_pending_count(VMUuid vm_id);
/* sk_hermes_pending_peek - the oldest still-queued message for vm_id, or
* NULL if vm_id has no queue or an empty one. Does not remove it -- task
* 3.4's drain logic is expected to peek, interpret, then pop. */
SkHermesMessage *sk_hermes_pending_peek(VMUuid vm_id);
/* sk_hermes_pending_pop - removes (does not release/interpret) the
* oldest queued entry for vm_id. Callers that also want the message's
* heat returned must call sk_hermes_release() on the value
* sk_hermes_pending_peek() returned, themselves, before or after popping
* -- this function only advances the queue. Refused (-1, no effect) if
* vm_id has no queue or an empty one. */
int sk_hermes_pending_pop(VMUuid vm_id);
#endif /* __STARKERNEL__ */
#endif /* STARKERNEL_VM_KERNEL_HERMES_H */
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+81
View File
@@ -544,6 +544,12 @@ static void kernel_main_deep(BootInfo *boot_info) {
* wiring). */
(void)sk_hermes_channels_boot_init();
/* Kernel-Hermes pending-queue table: boot-time allocation (FABRIC-3.6.md
* task 3.3), same sizing/ordering/soft-failure posture as the channel
* table just above. Publish (task 3.3) enqueues here; nothing drains
* it yet (task 3.4). */
(void)sk_hermes_queues_boot_init();
/* item 4.1, FABRIC-0.md item 3.6/§17.7: actually enforce "Hera is patron
* zero" before anything else can land on cell 0 via the free list, then
* bring up the word layer's map. Both must happen before the first word
@@ -1351,6 +1357,81 @@ static void kernel_main_deep(BootInfo *boot_info) {
}
}
/* FABRIC-3.6.md task 3.3 self-test: publish path, no dispatch. A
* synthetic publisher (lo=7, funded via stadium_grant_quota()) sends
* N=2 publishes to a synthetic 3-member channel (lo=8/9/10, message
* targets only -- no reservoir needed to receive) and confirms the
* ruled heat cost (one message per subscriber) lands exactly:
* sk_hermes_publish() returns 3 each time, each subscriber's own
* pending queue holds exactly 2 afterward, and the ledger audit plus
* stadium_conserved(publisher) (SXLIII.3's own check) hold both mid-
* publish and after this test drains every queue back to empty by
* hand (sk_hermes_pending_peek()/release()/pop() directly -- task
* 3.4's real checkpoint-driven drain does not exist yet). Diagnostic
* only, same posture as every other self-test block in this
* function. */
{
VMUuid pub_id, sub_ids[3];
int grant_rc;
int ch;
int i, n;
int pub_ok = 1;
uint64_t held0, pulled0, returned0, consumed0;
uint64_t held1, pulled1, returned1, consumed1;
pub_id.hi = 0;
pub_id.lo = 7;
for (i = 0; i < 3; i++) {
sub_ids[i].hi = 0;
sub_ids[i].lo = (uint64_t)(8 + i);
}
grant_rc = stadium_grant_quota(pub_id, vm_uuid_hera());
console_puts("Kernel-Hermes publish self-test: ");
if (grant_rc != 0) {
console_println("SKIPPED (quota grant failed)");
} else {
ch = sk_hermes_channel_create();
if (ch < 0) pub_ok = 0;
for (i = 0; pub_ok && i < 3; i++) {
if (sk_hermes_channel_subscribe(ch, sub_ids[i]) != 0) pub_ok = 0;
}
sk_hermes_ledger(&held0, &pulled0, &returned0, &consumed0);
for (n = 0; pub_ok && n < 2; n++) {
if (sk_hermes_publish(pub_id, ch, 0, (void *)0, 0) != 3) pub_ok = 0;
}
for (i = 0; pub_ok && i < 3; i++) {
if (sk_hermes_pending_count(sub_ids[i]) != 2) pub_ok = 0;
}
if (!sk_hermes_audit()) pub_ok = 0;
if (!stadium_conserved(pub_id)) pub_ok = 0;
/* Drain every queue by hand -- proves peek/pop/release compose
* correctly, not just that publish enqueued something. */
for (i = 0; pub_ok && i < 3; i++) {
while (sk_hermes_pending_count(sub_ids[i]) > 0) {
SkHermesMessage *msg = sk_hermes_pending_peek(sub_ids[i]);
if (!msg) { pub_ok = 0; break; }
if (sk_hermes_release(msg) != 0) pub_ok = 0;
if (sk_hermes_pending_pop(sub_ids[i]) != 0) pub_ok = 0;
}
}
sk_hermes_ledger(&held1, &pulled1, &returned1, &consumed1);
if (held1 != held0) pub_ok = 0; /* every allocation released, back to baseline */
if (!sk_hermes_audit_values(held1, pulled1, returned1, consumed1)) pub_ok = 0;
if (!stadium_conserved(pub_id)) pub_ok = 0;
if (pub_ok && sk_hermes_channel_destroy(ch) != 0) pub_ok = 0;
console_println(pub_ok ? "PASS" : "FAIL");
}
}
/* Decided 2026-09-05: no console for the running system unless a
* thumbdrive is present -- headless by default (EMERGENCY_CONSOLE_
* ENABLED off), reusing that flag's own existing "does this build
+140
View File
@@ -392,4 +392,144 @@ int sk_hermes_channel_member_count(int channel_id) {
return (int)sk_hermes_channels[channel_id].membership.count;
}
/* Task 3.3: per-subscriber pending queue table. See kernel_hermes.h's own
* doc comments on sk_hermes_queues_boot_init()/SK_HERMES_PENDING_MAX for
* scope and sizing reasoning. A queue is found-or-created lazily by
* vm_id, same shape as stadium.c's quota_slot_for_vm() and session.c's
* session_find() -- not pre-populated at birth like channel membership
* is, since not every VM ever receives a message. */
typedef struct {
VMUuid vm_id;
int in_use;
int slots[SK_HERMES_PENDING_MAX]; /* sk_hermes_msgs[] indices, circular FIFO */
size_t head;
size_t count;
} SkHermesPendingQueue;
static SkHermesPendingQueue *sk_hermes_queues = (SkHermesPendingQueue *)0;
static int sk_hermes_queue_capacity_val = 0;
int sk_hermes_queues_boot_init(void) {
size_t max_vm_count = stadium_max_vm_count();
SkHermesPendingQueue *table;
size_t i;
if (max_vm_count == 0) return -1; /* Stadium not yet initialized */
table = (SkHermesPendingQueue *)kmalloc(max_vm_count * sizeof(SkHermesPendingQueue));
if (!table) return -1;
for (i = 0; i < max_vm_count; i++) {
memset(&table[i], 0, sizeof(SkHermesPendingQueue));
}
sk_hermes_queues = table;
sk_hermes_queue_capacity_val = (int)max_vm_count;
console_puts("Kernel-Hermes: ");
console_put_u64((uint64_t)max_vm_count);
console_println(" pending-queue slots");
return 0;
}
static SkHermesPendingQueue *find_queue(VMUuid vm_id) {
int i;
if (!sk_hermes_queues) return (SkHermesPendingQueue *)0;
for (i = 0; i < sk_hermes_queue_capacity_val; i++) {
if (sk_hermes_queues[i].in_use && vm_uuid_equal(sk_hermes_queues[i].vm_id, vm_id)) {
return &sk_hermes_queues[i];
}
}
return (SkHermesPendingQueue *)0;
}
static SkHermesPendingQueue *find_or_create_queue(VMUuid vm_id) {
int i, free_slot = -1;
if (!sk_hermes_queues) return (SkHermesPendingQueue *)0;
for (i = 0; i < sk_hermes_queue_capacity_val; i++) {
if (sk_hermes_queues[i].in_use && vm_uuid_equal(sk_hermes_queues[i].vm_id, vm_id)) {
return &sk_hermes_queues[i];
}
if (!sk_hermes_queues[i].in_use && free_slot < 0) free_slot = i;
}
if (free_slot < 0) return (SkHermesPendingQueue *)0; /* table full */
memset(&sk_hermes_queues[free_slot], 0, sizeof(SkHermesPendingQueue));
sk_hermes_queues[free_slot].vm_id = vm_id;
sk_hermes_queues[free_slot].in_use = 1;
return &sk_hermes_queues[free_slot];
}
static int queue_push(SkHermesPendingQueue *q, int msg_index) {
size_t tail;
if (q->count >= SK_HERMES_PENDING_MAX) return -1;
tail = (q->head + q->count) % SK_HERMES_PENDING_MAX;
q->slots[tail] = msg_index;
q->count++;
return 0;
}
int sk_hermes_publish(VMUuid from, int channel_id, uint32_t type,
void *payload_addr, uint32_t payload_len) {
SkHermesMembership *m;
size_t i;
int delivered = 0;
if (!channel_valid(channel_id)) return -1;
m = &sk_hermes_channels[channel_id].membership;
for (i = 0; i < m->count; i++) {
VMUuid to = m->members[i];
SkHermesMessage *msg;
SkHermesPendingQueue *q;
int idx;
if (sk_hermes_alloc(from, &msg) != 0) continue; /* this subscriber skipped, not fatal */
msg->type = type;
msg->from = from;
msg->to = to;
msg->payload_addr = payload_addr;
msg->payload_len = payload_len;
msg->channel = (uint32_t)channel_id;
q = find_or_create_queue(to);
if (!q) {
sk_hermes_release(msg); /* roll back this subscriber's allocation */
continue;
}
idx = (int)(msg - sk_hermes_msgs);
if (queue_push(q, idx) != 0) {
sk_hermes_release(msg); /* this subscriber's own queue is full */
continue;
}
delivered++;
}
return delivered;
}
int sk_hermes_pending_count(VMUuid vm_id) {
SkHermesPendingQueue *q = find_queue(vm_id);
return q ? (int)q->count : 0;
}
SkHermesMessage *sk_hermes_pending_peek(VMUuid vm_id) {
SkHermesPendingQueue *q = find_queue(vm_id);
if (!q || q->count == 0) return (SkHermesMessage *)0;
return &sk_hermes_msgs[q->slots[q->head]];
}
int sk_hermes_pending_pop(VMUuid vm_id) {
SkHermesPendingQueue *q = find_queue(vm_id);
if (!q || q->count == 0) return -1;
q->head = (q->head + 1) % SK_HERMES_PENDING_MAX;
q->count--;
return 0;
}
#endif /* __STARKERNEL__ */