diff --git a/docs/v4.0.0/DECOMPOSITION.md b/docs/v4.0.0/DECOMPOSITION.md index abb48215..89170bf4 100644 --- a/docs/v4.0.0/DECOMPOSITION.md +++ b/docs/v4.0.0/DECOMPOSITION.md @@ -567,9 +567,9 @@ width and print a flood of spaces. | `SCAN` `SKIP` | CAP | `( baddr u c -- baddr' u' )`: `SCAN` stops at the first byte equal to the low byte of `c` (or at the end, with `u' = 0`); `SKIP` stops at the first byte not equal to it. See below. A negative count reads as 0, as v3. v3's counted-string auto-detection is not kept. Executed on the golden model (2026-10-03), including a transcript of the v3 binary. Each leaves its caller 5 data cells and 5 return entries. Clobber `A`. | | `(S)` | CAP | Variable, 4 cells: `COMPARE`'s two lengths; `SEARCH`'s string 2 and the whole of string 1. | | `BL` | IN | `32` — executed on the golden model (2026-10-04). | -| `EXPECT` `QUERY` | CC | Built on `KEY` (DEV); source in `v4/capsule/input.v4`. `EXPECT ( baddr u -- )` as v3's `fgets`: at most `u-1` characters, stopping after a new-line (taken, not stored), then a zero; `SPAN` is the count; a line that does not fit leaves its tail to be read next; `u < 0` does nothing and sets `NODE-ERROR`. `QUERY` is `TIB #TIB EXPECT 0 >IN !` with `#TIB` 1025, as v3. Executed on the golden model's host node (2026-10-04), including results recorded from the v3 binary. Both part from FORTH-79, which takes up to `u` characters in `EXPECT` and 80 in `QUERY`: reported 2026-10-04, awaiting a ruling. `EXPECT` leaves its caller 6 data cells and 4 return entries. | +| `EXPECT` `QUERY` | CC | Built on `KEY` (DEV); source in `v4/capsule/input.v4`. FORTH-79 (ruled 2026-10-04). `EXPECT ( baddr n -- )` stores characters from `baddr` upward until a new-line (taken, not stored) or until `n` have been received, then a zero; no action for `n <= 0`; a line longer than `n` leaves the rest to be read next. It also sets `SPAN`, which is not FORTH-79 but is v3's. `QUERY` is `TIB 80 EXPECT 0 >IN !`. v3's `EXPECT` was C's `fgets` and took `n-1` characters, and its `QUERY` took 1024. Executed on the golden model's host node (2026-10-04) against C for every size and line length. `EXPECT` leaves its caller 6 data cells and 4 return entries. | | `SPAN` `TIB` `>IN` `SOURCE` | CC | Interpreter state on the host node. `TIB` is the buffer's byte address and `>IN` and `SPAN` their variables' word addresses, all in line; `SOURCE ( -- baddr u )` is `TIB` and `SPAN @`, a negative `SPAN` read as 0. Executed on the golden model's host node (2026-10-04). | -| `WORD` `ENCLOSE` | CC | Parser; source in `v4/capsule/input.v4`. `WORD ( c -- baddr )`, as v3: the next word of `TIB` delimited by `c`, as a counted string in a 64-byte buffer with a zero after it; delimiters before it are skipped and so are all those after it; a longer word is cut to 62 characters; at the end of the line the count is 0. `ENCLOSE ( baddr c -- baddr n1 n2 n3 )` on a zero-terminated string. Executed on the golden model's host node (2026-10-04) against C and the counts and `>IN` values recorded from the v3 binary (the text in v3's own buffer was corrupt in those runs: `hello` came back as `hlllo`). `WORD` parts from FORTH-79 in skipping every trailing delimiter, storing a zero where the standard stores the delimiter, and cutting at 62: reported 2026-10-04, awaiting a ruling. `WORD` leaves its caller 4 data cells and 3 return entries. | +| `WORD` `ENCLOSE` | CC | Parser; source in `v4/capsule/input.v4`. `WORD ( c -- baddr )` is FORTH-79 (ruled 2026-10-04): characters are taken from `TIB` until the delimiter `c` or the end of the text, leading delimiters ignored, and stored as a counted string; the delimiter met (`c`, or a zero if the text ran out) is stored after them, uncounted; `>IN` is left just past it; with nothing left the count is 0. The count is a byte, so a word over 255 characters is cut to 255; the buffer is 257 bytes. v3's `WORD` skipped every delimiter after the word, always stored a zero, and cut at 62. `ENCLOSE ( baddr c -- baddr n1 n2 n3 )`, not a FORTH-79 word, works on a zero-terminated string as v3's. Executed on the golden model's host node (2026-10-04) against C, and `ENCLOSE` against values recorded from the v3 binary. `WORD` leaves its caller 4 data cells and 3 return entries. | | `NUMBER` `CONVERT` | CC | FORTH-79, not v3 (ruled 2026-10-04); source in `v4/capsule/input.v4`. `CONVERT ( d1 baddr1 -- d2 baddr2 )` takes the characters from `baddr1 + 1` on as digits in the current `BASE`, in either case, accumulating each into the double after multiplying it by `BASE`, and stops at the first that is not one; the double is `( lo hi )`. v3's was base 10 only, started at `baddr1`, and took the double low cell on top. `NUMBER ( baddr -- d )` is the counted string as a signed double in the current `BASE`, with an optional leading minus; anything else gives 0 and sets `NODE-ERROR`. v3's returned `( n flag )` in base 10. Executed on the golden model's host node (2026-10-04) against C in eight bases, including a digit that carries out of the low cell. `NUMBER` leaves its caller 3 data cells and 3 return entries. | | `S"` `(s")` `[']` | CC | Compiler words. | | `LITERAL` `[LITERAL]` (placeholders) | RET | The working `LITERAL` is in §5.17. | diff --git a/v4/capsule/input.v4 b/v4/capsule/input.v4 index 60056019..67a2ed0a 100644 --- a/v4/capsule/input.v4 +++ b/v4/capsule/input.v4 @@ -7,12 +7,11 @@ \ capsule: it runs on the host node. Rests on core.v4. \ \ Constants the loader supplies: -\ TIB byte address of the text input buffer, #TIB bytes long -\ #TIB 1025: 1024 characters and a terminating zero, as v3 +\ TIB byte address of the text input buffer; QUERY fills at most 81 bytes \ >IN word address of the variable: offset in TIB of the next character \ SPAN word address of the variable: how many characters the last \ EXPECT or QUERY read -\ WBUF byte address of WORD's buffer, 64 bytes +\ WBUF byte address of WORD's buffer, 257 bytes \ (P) word address of six cells of scratch for this file \ TIB >IN SPAN and BL are in-line: written in a definition they are literals. @@ -21,25 +20,26 @@ macro BL 32 endmacro \ ( -- baddr u ) the input buffer and how much of it is in use : SOURCE TIB SPAN a! @ -if P drop 0 P: ; -\ ( baddr u -- ) read a line into u bytes at baddr, as v3's fgets did: at -\ most u-1 characters, stopping after a new-line, which is taken but not -\ stored; then a zero. A line of u-1 characters or more leaves what did not -\ fit, its new-line included, to be read next. SPAN is the number of characters stored. With u = 0 -\ nothing is read or stored; u < 0 does nothing and sets NODE-ERROR. +\ ( baddr n -- ) FORTH-79: characters from the terminal are stored from +\ baddr upward until a new-line or until n have been received; the new-line +\ is taken but not stored. A zero is added after the text, so the buffer +\ must hold n + 1 bytes. No action for n <= 0. SPAN is how many characters +\ were stored (SPAN is not FORTH-79; v3 has it). A line longer than n leaves +\ the rest, its new-line included, to be read next. : EXPECT - -if OK drop drop NODE-ERROR b! -1 !b ; - OK: if ZERO + -if NN drop drop ; \ n < 0 + NN: if ZERO over push \ p n R: start - L: -1 + if END \ room for the zero only + L: if END push KEY dup -10 + if NL \ p c x R: start n - drop over C! 1 + pop jump L + drop over C! 1 + pop -1 + jump L NL: drop drop pop END: drop 0 over C! \ p pop - SPAN a! ! ; - ZERO: drop drop 0 SPAN a! ! ; + ZERO: drop drop ; -\ ( -- ) read a line into TIB and start at its beginning -: QUERY TIB #TIB EXPECT 0 >IN a! ! ; +\ ( -- ) FORTH-79: up to 80 characters, or a line, into TIB; >IN to 0. +: QUERY TIB 80 EXPECT 0 >IN a! ! ; \ ---- WORD ------------------------------------------------------------------ \ (P)+0 the delimiter, (P)+1 the length of the text in TIB. @@ -56,21 +56,26 @@ macro (WLEN) (P) 1 + endmacro L: (LEFT) -if E drop (CH) if E drop 1 + jump L E: drop ; -\ ( c -- baddr ) the next word of TIB delimited by c, as a counted string -\ in WBUF with a zero after it. Leading delimiters are skipped; so, as in -\ v3, are all those that follow the word. A word longer than 62 characters -\ is cut to 62. At the end of the text the count is 0. +\ ( c -- baddr ) FORTH-79: characters are taken from TIB until the +\ delimiter c or the end of the text, leading delimiters ignored, and stored +\ in WBUF as a counted string. The delimiter met -- c, or a zero if the text +\ ran out -- is stored after them and is not counted. >IN is left just past +\ that delimiter. With nothing left the count is 0. The count is a byte, +\ so a word longer than 255 characters is cut to 255. : WORD 255 and (WDELIM) a! ! SPAN a! @ -if A drop 0 A: (WLEN) a! ! >IN a! @ -if B drop 0 B: (SKIP) dup push (SCAN) \ end R: start pop over over - \ end start len - dup -63 + -if CAP drop jump FITS CAP: drop drop 62 FITS: + dup -256 + -if CAP drop jump FITS CAP: drop drop 255 FITS: dup WBUF C! dup push push TIB + WBUF 1 + pop CMOVE \ end R: len - 0 pop WBUF + 1 + C! - (SKIP) >IN a! ! WBUF ; + (LEFT) -if RANOUT + drop (WDELIM) a! @ pop WBUF + 1 + C! \ the delimiter after the text + 1 + >IN a! ! WBUF ; + RANOUT: drop 0 pop WBUF + 1 + C! + >IN a! ! WBUF ; \ ---- ENCLOSE --------------------------------------------------------------- \ (P)+0 the delimiter, (P)+1 the string's address. @@ -120,7 +125,8 @@ macro (WLEN) (P) 1 + endmacro \ the current BASE; it may begin with a minus sign. If it is empty, is only \ a sign, or holds anything that is not a digit, the result is 0 and \ NODE-ERROR is set. The character after the string must not be a digit -\ (WORD puts a zero there); if it is, that too is reported as an error. (P)+3 the address of its last character, (P)+4 the sign. +\ (WORD puts the delimiter or a zero there); if it is, that too is reported +\ as an error. (P)+3 the address of its last character, (P)+4 the sign. : NUMBER dup C@ if ERR2 over + (P) 3 + a! ! \ baddr diff --git a/v4/tests/test_host_input.c b/v4/tests/test_host_input.c index a7508285..72630137 100644 --- a/v4/tests/test_host_input.c +++ b/v4/tests/test_host_input.c @@ -5,25 +5,24 @@ * ENCLOSE CONVERT NUMBER and the comment words. The definitions are text, * capsule/core.v4 and capsule/input.v4, assembled here by the text assembler. * - * v3 (v3/src/word_source/string_words.c), kept: - * EXPECT ( baddr u -- ) at most u-1 characters, stopping after a - * new-line; a zero after them; SPAN is how many - * QUERY a line into TIB, >IN to 0 - * WORD ( c -- baddr ) the next word of TIB as a counted string; the - * delimiters before it and all those after it are skipped - * ENCLOSE ( baddr c -- baddr n1 n2 n3 ) - * FORTH-79, where v3 was something else (ruled 2026-10-04): + * These are FORTH-79 words and follow the standard (ruled 2026-10-04) where + * v3 did something else: + * EXPECT ( baddr n -- ) up to n characters, or to a new-line; a zero + * after them; no action for n <= 0. v3 took n-1 (it was fgets). + * QUERY up to 80 characters into TIB, >IN to 0. v3 took 1024. + * WORD ( c -- baddr ) the next word of TIB as a counted string, the + * delimiter met (or a zero at the end of the text) stored after + * it; >IN just past that delimiter. v3 skipped every delimiter + * after the word, stored a zero, and cut the word at 62. * CONVERT ( d1 baddr1 -- d2 baddr2 ) starts at baddr1 + 1, honours BASE, * and the double is ( lo hi ). v3 was base 10 only, started at * baddr1, and took the double low cell on top. * NUMBER ( baddr -- d ) a signed double in the current BASE; anything * that is not a number gives 0 and sets NODE-ERROR. v3 returned * ( n flag ) and read base 10 only. + * SPAN, SOURCE, TIB and ENCLOSE are not in FORTH-79 and behave as v3's. * * The values marked "v3:" were recorded from the v3 binary on 2026-10-04. - * v3's WORD returned the right count and left >IN right, but the text in its - * buffer was corrupt in those runs ("hello" came back as "hlllo"), so only - * its counts and >IN are used. */ #include "v4/text.h" #include "v4/testcode.h" @@ -48,13 +47,13 @@ static int failures = 0, checks = 0; #define TO_IN (TOP - 9) #define SPAN (TOP - 10) #define PVARS (TOP - 16) /* six cells */ -#define WBUF_W (TOP - 40) /* 16 cells = 64 bytes, with a cell free on each side */ -#define TIB_W (TOP - 304) /* 260 cells = 1040 bytes */ -#define SBUF_W (TOP - 384) /* 64 cells = 256 bytes for the tests' own strings */ +#define WBUF_W (TOP - 90) /* 65 cells = 260 bytes, with cells free on each side */ +#define WBUF_CELLS 65 +#define TIB_W (TOP - 360) /* 260 cells = 1040 bytes */ +#define SBUF_W (TOP - 440) /* 64 cells = 256 bytes for the tests' own strings */ #define WBUF (WBUF_W * 4) #define TIB (TIB_W * 4) #define SBUF (SBUF_W * 4) -#define TIB_SIZE 1025 #define GUARD ((v4_cell)0x5EED5EED) static v4_node n; @@ -119,11 +118,13 @@ static v4_cell pop(void) { return v4_dstack_pop(&n.ds); } static int clean(void) { return pop() == CANARY; } static int err(void) { return n.mem[NODE_ERROR] != 0; } -/* The counted string WORD left: its text, and that a zero follows it. */ -static int word_is(v4_cell baddr, const char *want) +/* The counted string WORD left: its text, and after it the delimiter it met + * (`after`), or a zero if the text ran out. */ +static int word_is(v4_cell baddr, const char *want, int after) { unsigned len = (unsigned)strlen(want); - return baddr == WBUF && byte_at(WBUF) == len && bytes_are(WBUF + 1, want, len) && byte_at(WBUF + 1 + (v4_cell)len) == 0; + return baddr == WBUF && byte_at(WBUF) == len && bytes_are(WBUF + 1, want, len) + && byte_at(WBUF + 1 + (v4_cell)len) == (unsigned char)after; } /* Stack headroom, as in test_foundation.c. */ @@ -234,7 +235,6 @@ int main(void) v4_text_constant(&tx, "CONSOLE-STATUS", CONSOLE_ST); v4_text_constant(&tx, "BASE", BASE); v4_text_constant(&tx, "TIB", TIB); - v4_text_constant(&tx, "#TIB", TIB_SIZE); v4_text_constant(&tx, ">IN", TO_IN); v4_text_constant(&tx, "SPAN", SPAN); v4_text_constant(&tx, "WBUF", WBUF); @@ -292,27 +292,26 @@ int main(void) /* ---- EXPECT ---- */ v4_node_console_input_attach(&n, CONSOLE_RX, CONSOLE_ST); - /* v3: B 8 EXPECT on "hello world this is long": SPAN 7, 68 65 6C 6C 6F 20 77 00, and - * the rest of the line ("orld ...") still to be read */ + /* FORTH-79: up to n characters. (v3, which was fgets, stored 7 of + * "hello world this is long" for 8 EXPECT; the standard stores 8.) */ for (i = 0; i < 64; i++) n.mem[SBUF_W + (v4_cell)i] = 0x2A2A2A2A; CHECK(v4_node_console_feed(&n, "hello world this is long\n", 25) == 25, "fed"); - CHECK(run("EXPECT", 2, SBUF, 8, 0) && clean() && !err(), "v3: 8 EXPECT returns"); - CHECK(n.mem[SPAN] == 7 && bytes_are(SBUF, "hello w\0", 8) && byte_at(SBUF + 8) == 0x2A, "v3: 8 EXPECT stores \"hello w\" and a zero, SPAN 7"); - CHECK(run("KEY", 0, 0, 0, 0) && pop() == 'o', "v3: and leaves the rest of the line unread"); + CHECK(run("EXPECT", 2, SBUF, 8, 0) && clean() && !err(), "8 EXPECT returns"); + CHECK(n.mem[SPAN] == 8 && bytes_are(SBUF, "hello wo\0", 9) && byte_at(SBUF + 9) == 0x2A, "8 EXPECT stores eight characters and a zero, SPAN 8"); + CHECK(run("KEY", 0, 0, 0, 0) && pop() == 'r', "and leaves the rest of the line unread"); v4_node_console_input_attach(&n, CONSOLE_RX, CONSOLE_ST); - /* v3: B 8 EXPECT on "hi": SPAN 2. Then B 3 EXPECT on "xyz": SPAN 2, 78 79 00, 'z' unread */ CHECK(v4_node_console_feed(&n, "hi\nxyz\n", 7) == 7, "fed"); - CHECK(run("EXPECT", 2, SBUF, 8, 0) && clean() && n.mem[SPAN] == 2 && bytes_are(SBUF, "hi\0", 3), "v3: a short line: SPAN 2, the new-line taken, not stored"); - CHECK(run("EXPECT", 2, SBUF, 3, 0) && clean() && n.mem[SPAN] == 2 && bytes_are(SBUF, "xy\0", 3), "v3: 3 EXPECT stores two characters"); - CHECK(run("KEY", 0, 0, 0, 0) && pop() == 'z', "v3: and leaves the third unread"); + CHECK(run("EXPECT", 2, SBUF, 8, 0) && clean() && n.mem[SPAN] == 2 && bytes_are(SBUF, "hi\0", 3), "a short line: SPAN 2, the new-line taken, not stored"); + CHECK(run("EXPECT", 2, SBUF, 3, 0) && clean() && n.mem[SPAN] == 3 && bytes_are(SBUF, "xyz\0", 4), "3 EXPECT stores three characters"); + CHECK(run("KEY", 0, 0, 0, 0) && pop() == '\n', "and leaves the new-line unread"); v4_node_console_input_attach(&n, CONSOLE_RX, CONSOLE_ST); /* every size against every line length */ - for (t = 0; t <= 12; t++) + for (t = 1; t <= 12; t++) for (i = 0; i <= 12; i++) { char line[16]; - unsigned want = t == 0 ? 0 : (i < t - 1 ? i : t - 1), k; + unsigned want = i < t ? i : t, k; for (k = 0; k < i; k++) line[k] = (char)('A' + k); line[i] = '\n'; for (k = 0; k < 64; k++) n.mem[SBUF_W + (v4_cell)k] = 0x2A2A2A2A; @@ -321,20 +320,29 @@ int main(void) n.mem[SPAN] = 99; CHECK(run("EXPECT", 2, SBUF + 3, (v4_cell)t, 0) && clean() && !err(), "EXPECT returns [%u,%u]", t, i); CHECK(n.mem[SPAN] == (v4_cell)want, "EXPECT SPAN [%u,%u]: %ld", t, i, (long)n.mem[SPAN]); - CHECK(bytes_are(SBUF + 3, line, want) && (t == 0 || byte_at(SBUF + 3 + (v4_cell)want) == 0), "EXPECT text [%u,%u]", t, i); - CHECK(byte_at(SBUF + 2) == 0x2A && byte_at(SBUF + 3 + (v4_cell)want + (t ? 1 : 0)) == 0x2A, "EXPECT writes nothing else [%u,%u]", t, i); - /* what is left unread: the line's tail if it did not fit, new-line - * included -- and, as with fgets, the new-line alone when the - * text exactly filled the u-1 characters */ + CHECK(bytes_are(SBUF + 3, line, want) && byte_at(SBUF + 3 + (v4_cell)want) == 0, "EXPECT text [%u,%u]", t, i); + CHECK(byte_at(SBUF + 2) == 0x2A && byte_at(SBUF + 4 + (v4_cell)want) == 0x2A, "EXPECT writes nothing else [%u,%u]", t, i); + /* what is left unread: nothing if the new-line was reached, else + * the rest of the line and its new-line */ k = n.input_len - n.input_pos; - CHECK(k == (t == 0 ? i + 1u : (i + 1u < t ? 0u : i + 1u - (t - 1u))), "EXPECT leaves %u unread [%u,%u]", k, t, i); + CHECK(k == (i < t ? 0u : i + 1u - t), "EXPECT leaves %u unread [%u,%u]", k, t, i); } - v4_node_console_input_attach(&n, CONSOLE_RX, CONSOLE_ST); - n.mem[SPAN] = 5; - CHECK(run("EXPECT", 2, SBUF, -1, 0) && clean() && err() && n.mem[SPAN] == 5, "EXPECT of a negative size sets NODE-ERROR and reads nothing"); + /* no action for n <= 0: nothing read, nothing stored, SPAN as it was, no error */ + { + static const v4_cell none[] = { 0, -1, -100, (v4_cell)V4_MSB }; + for (i = 0; i < 4; i++) { + for (t = 0; t < 64; t++) n.mem[SBUF_W + (v4_cell)t] = 0x2A2A2A2A; + v4_node_console_input_attach(&n, CONSOLE_RX, CONSOLE_ST); + (void)v4_node_console_feed(&n, "abc\n", 4); + n.mem[SPAN] = 5; + CHECK(run("EXPECT", 2, SBUF, none[i], 0) && clean() && !err() && n.mem[SPAN] == 5 + && n.input_len - n.input_pos == 4 && byte_at(SBUF) == 0x2A, "EXPECT takes no action for n <= 0 [%u]", i); + } + v4_node_console_input_attach(&n, CONSOLE_RX, CONSOLE_ST); + } /* ---- QUERY and SOURCE ---- */ - n.mem[TIB_W - 1] = GUARD; n.mem[TIB_W + 260] = GUARD; + n.mem[TIB_W - 1] = GUARD; n.mem[TIB_W + 21] = GUARD; n.mem[TO_IN] = 77; CHECK(v4_node_console_feed(&n, " 12 34 +\nnext line\n", 20) == 20, "fed"); CHECK(run("QUERY", 0, 0, 0, 0) && clean() && !err(), "QUERY returns"); @@ -342,13 +350,14 @@ int main(void) CHECK(run("SOURCE", 0, 0, 0, 0) && pop() == 9 && pop() == TIB && clean(), "SOURCE is TIB and SPAN @"); CHECK(run("QUERY", 0, 0, 0, 0) && clean() && n.mem[SPAN] == 9 && bytes_are(TIB, "next line\0", 10), "the next QUERY reads the next line"); { - static char longline[1200]; - for (i = 0; i < 1100; i++) longline[i] = (char)('a' + i % 26); - longline[1100] = '\n'; - CHECK(v4_node_console_feed(&n, longline, 1101) == 1101, "fed"); - CHECK(run("QUERY", 0, 0, 0, 0) && clean() && n.mem[SPAN] == 1024, "a line longer than TIB is cut at 1024 characters"); - CHECK(bytes_are(TIB, longline, 1024) && byte_at(TIB + 1024) == 0, "with a zero after them"); - CHECK(n.mem[TIB_W - 1] == GUARD && n.mem[TIB_W + 260] == GUARD, "and nothing outside TIB written"); + static char longline[200]; + for (i = 0; i < 150; i++) longline[i] = (char)('a' + i % 26); + longline[150] = '\n'; + CHECK(v4_node_console_feed(&n, longline, 151) == 151, "fed"); + CHECK(run("QUERY", 0, 0, 0, 0) && clean() && n.mem[SPAN] == 80, "QUERY takes at most 80 characters"); + CHECK(bytes_are(TIB, longline, 80) && byte_at(TIB + 80) == 0, "with a zero after them"); + CHECK(n.mem[TIB_W - 1] == GUARD && n.mem[TIB_W + 21] == GUARD, "and nothing past the 81st byte written"); + CHECK(run("QUERY", 0, 0, 0, 0) && clean() && n.mem[SPAN] == 70 && bytes_are(TIB, longline + 80, 70), "the rest of the line is the next QUERY's"); v4_node_console_input_attach(&n, CONSOLE_RX, CONSOLE_ST); } n.mem[SPAN] = -5; @@ -356,26 +365,27 @@ int main(void) /* ---- WORD ---- */ - /* v3: " hello world x": BL WORD twice, then >IN is 18 and SPAN 19; a third gives count 1, >IN 19 */ + /* FORTH-79: the delimiter met is stored after the word, and >IN is left + * just past it. (v3 skipped all the delimiters after a word, so its + * >IN after "world" was 18 where the standard's is 17.) */ set_line(" hello world x"); - CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "hello") && clean() && n.mem[TO_IN] == 11, "the first word"); - CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "world") && clean() && n.mem[TO_IN] == 18 && n.mem[SPAN] == 19, "v3: after two words >IN is 18"); - CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "x") && clean() && n.mem[TO_IN] == 19, "v3: the third has count 1 and >IN is 19"); - CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "") && clean() && n.mem[TO_IN] == 19, "at the end of the line the count is 0"); + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "hello", ' ') && clean() && n.mem[TO_IN] == 9, "the first word; >IN is past one blank"); + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "world", ' ') && clean() && n.mem[TO_IN] == 17, "the second"); + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "x", 0) && clean() && n.mem[TO_IN] == 19, "the last word ends the text: a zero after it"); + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "", 0) && clean() && n.mem[TO_IN] == 19, "with nothing left the count is 0"); - /* v3: "a b,,c d,e": SOURCE gives 10; 44 WORD twice, then >IN is 9 */ set_line("a b,,c d,e"); - CHECK(run("WORD", 1, 44, 0, 0) && word_is(pop(), "a b") && clean() && n.mem[TO_IN] == 5, "a comma-delimited word; both commas skipped"); - CHECK(run("WORD", 1, 44, 0, 0) && word_is(pop(), "c d") && clean() && n.mem[TO_IN] == 9, "v3: after two >IN is 9"); - CHECK(run("WORD", 1, 44, 0, 0) && word_is(pop(), "e") && clean() && n.mem[TO_IN] == 10, "the last one"); + CHECK(run("WORD", 1, 44, 0, 0) && word_is(pop(), "a b", ',') && clean() && n.mem[TO_IN] == 4, "a comma-delimited word; one comma taken"); + CHECK(run("WORD", 1, 44, 0, 0) && word_is(pop(), "c d", ',') && clean() && n.mem[TO_IN] == 9, "the leading comma of the next is skipped"); + CHECK(run("WORD", 1, 44, 0, 0) && word_is(pop(), "e", 0) && clean() && n.mem[TO_IN] == 10, "the last one"); - /* v3: "AAAA" with delimiter 65: count 0, >IN 4, count 0 */ + /* v3: "AAAA" with delimiter 65: count 0, >IN 4, count 0 -- the same in the standard */ set_line("AAAA"); - CHECK(run("WORD", 1, 65, 0, 0) && word_is(pop(), "") && clean() && n.mem[TO_IN] == 4, "v3: a line of delimiters: count 0, >IN 4"); - CHECK(run("WORD", 1, 65, 0, 0) && word_is(pop(), "") && clean() && n.mem[TO_IN] == 4, "v3: and again"); + CHECK(run("WORD", 1, 65, 0, 0) && word_is(pop(), "", 0) && clean() && n.mem[TO_IN] == 4, "v3: a line of delimiters: count 0, >IN 4"); + CHECK(run("WORD", 1, 65, 0, 0) && word_is(pop(), "", 0) && clean() && n.mem[TO_IN] == 4, "v3: and again"); /* against C */ - n.mem[WBUF_W - 1] = GUARD; n.mem[WBUF_W + 16] = GUARD; + n.mem[WBUF_W - 1] = GUARD; n.mem[WBUF_W + WBUF_CELLS] = GUARD; for (t = 0; t < 400; t++) { char line[96], tok[96]; unsigned len = rnd() % 70, pos = 0, guard = 0; @@ -385,44 +395,50 @@ int main(void) set_line(line); for (;;) { unsigned start, end, tl; + int after; while (pos < len && line[pos] == delim) pos++; start = pos; while (pos < len && line[pos] != delim) pos++; end = pos; - while (pos < len && line[pos] == delim) pos++; + after = pos < len ? delim : 0; + if (pos < len) pos++; tl = end - start; memcpy(tok, line + start, tl); tok[tl] = 0; - CHECK(run("WORD", 1, (v4_cell)(0x700 + delim), 0, 0) && word_is(pop(), tok) && clean() && !err() + CHECK(run("WORD", 1, (v4_cell)(0x700 + delim), 0, 0) && word_is(pop(), tok, after) && clean() && !err() && n.mem[TO_IN] == (v4_cell)pos, "WORD [%u] \"%s\" at %u", t, tok, pos); - if (tl == 0 || ++guard > 80) break; + if (pos >= len || ++guard > 80) break; } + CHECK(run("WORD", 1, (v4_cell)delim, 0, 0) && word_is(pop(), "", 0) && clean() && n.mem[TO_IN] == (v4_cell)len, "WORD at the end [%u]", t); CHECK(bytes_are(TIB, line, len + 1u) && n.mem[SPAN] == (v4_cell)len, "WORD leaves TIB and SPAN [%u]", t); } - /* a word longer than the buffer is cut to 62, and >IN still passes all of it */ + /* a word of 255 characters fits; a longer one is cut to 255, and >IN still passes all of it */ { - char line[200], cut[64]; - memset(line, 'w', 150); line[150] = ' '; line[151] = 'z'; line[152] = 0; - memset(cut, 'w', 62); cut[62] = 0; + static char line[400], cut[256]; + memset(line, 'w', 255); line[255] = ' '; line[256] = 'z'; line[257] = 0; + memset(cut, 'w', 255); cut[255] = 0; set_line(line); - CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), cut) && clean() && n.mem[TO_IN] == 151, "a 150-character word is cut to 62"); - CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "z") && clean(), "and the next word is the next word"); - CHECK(n.mem[WBUF_W - 1] == GUARD && n.mem[WBUF_W + 16] == GUARD, "nothing outside WORD's buffer written"); + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), cut, ' ') && clean() && n.mem[TO_IN] == 256, "a 255-character word is whole"); + memset(line, 'w', 300); line[300] = ' '; line[301] = 'z'; line[302] = 0; + set_line(line); + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), cut, ' ') && clean() && n.mem[TO_IN] == 301, "a 300-character word is cut to 255"); + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "z", 0) && clean(), "and the next word is the next word"); + CHECK(n.mem[WBUF_W - 1] == GUARD && n.mem[WBUF_W + WBUF_CELLS] == GUARD, "nothing outside WORD's buffer written"); } set_line("abc def"); n.mem[TO_IN] = -3; - CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "abc") && clean(), "a negative >IN reads as 0"); + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "abc", ' ') && clean(), "a negative >IN reads as 0"); n.mem[SPAN] = -3; n.mem[TO_IN] = 0; - CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "") && clean() && n.mem[TO_IN] == 0, "a negative SPAN reads as an empty line"); + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "", 0) && clean() && n.mem[TO_IN] == 0, "a negative SPAN reads as an empty line"); /* ---- the comment words ---- */ set_line("skip this ) kept ( and ) too"); - CHECK(run("PAREN", 0, 0, 0, 0) && clean() && n.mem[TO_IN] == 11, "( skips to after the closing parenthesis"); - CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "kept") && clean(), "and the next word follows"); + CHECK(run("PAREN", 0, 0, 0, 0) && clean() && n.mem[TO_IN] == 11, "( skips to just after the closing parenthesis"); + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "kept", ' ') && clean(), "and the next word follows"); set_line("one two three"); - CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "one") && run("BACKSLASH", 0, 0, 0, 0) && clean() && n.mem[TO_IN] == 13, + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "one", ' ') && run("BACKSLASH", 0, 0, 0, 0) && clean() && n.mem[TO_IN] == 13, "\\ skips the rest of the line"); - CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "") && clean(), "nothing is left after it"); + CHECK(run("WORD", 1, 32, 0, 0) && word_is(pop(), "", 0) && clean(), "nothing is left after it"); /* ---- ENCLOSE ---- */