Repository navigation
train: walk the window — decode context into the cache, train each reply chunk against it - #40
Conversation
…t the cached prefix (WIP) Fable's design (continuum, 2026-10-06): walk her conversation once in one context; context she did not write is a plain decode into the cache; her reply trains in chunks, each attending to everything before it read from the cache as a constant; then the chunk is decoded into the cache under the adapter as it now is. Memory is chunk x window, not window x window, so a 63-73k lived window can train. - build_attn (training): concat(cached prefix K/V, this chunk's K/V); V transposed out of the non-FA cache; the mask sliced to prefix+chunk and made contiguous. - opt_epoch_iter: the walk (decode_span / train_chunk), nothing after the last label. - opt_init: one ubatch per batch, not per context. - graph_max_nodes: the prefix nodes; the training budget follows opt_ctx, not the per-graph flag (a decode's sched_reserve resized it under training=false). - sched_reserve: a training context keeps its scheduler (the walk's first decode re-created it and left ggml-opt on a freed one: GGML_ASSERT(backend)). - server-train: "chunk", the largest multiple of 256 dividing the window. Measured on the 5090, Qwen2.5-Coder-1.5B Q4_K_M, 32 single-reply examples (window 1280), lr 1e-4, 3 epochs: - forward (lr 1e-9): old 0.74658 / eval 0.82223; walk 0.74915 / 0.82235 (correct) - eval: old 0.512 -> 0.247 -> 0.222; walk 0.628 -> 0.452 -> 0.458 - per epoch: old 15-24 s; walk 5-8 s The stop-gradient at the cache costs about half the learning on replies that draw on context. WIP: the WALK fprintf debug lines stay until exact-gradient work lands.
…e-decode stops the run loudly Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…able), so a reply spans chunks on a hybrid The training forward advances a recurrent state in place, and a recurrent memory cannot remove a partial range, so a reply longer than one chunk could not be re-decoded on a hybrid model. Before a training chunk, the recurrent part's state is copied to the scratch sequence 1; after it, the state is restored, the hybrid's attention part drops the chunk's empty cells, and the chunk is decoded for real. The training context gets n_seq_max 2, unified (one attention stream of the window). Forward check on the 5090 (Qwen2.5-Coder-1.5B Q4_K_M, lr 1e-9, 32 single-reply examples, window 1280): chunk=window 0.74915 / eval 0.82235; chunk=256 (replies split across chunks, re-decoded between) 0.74668 / eval 0.82669; old engine 0.74658 / 0.82223. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
With the walk, the training context's cache holds only CONSTANTS (the context before a chunk); the chunk's own K/V carry the gradient in the graph at F32 and are never cached during training. So the cache is F16, cast to the chunk's type where build_attn joins them, which halves the window's cache: about 8 GB to 4 GB at Kimi's 63k on the 27B, the margin beside her live lane. A caller that sends no "chunk" gets 512 until the core passes the lease's S (Fable). The window's length as one chunk is the window x window memory the walk exists to avoid. Forward check, 1.5B, lr 1e-9, chunk 256: 0.74845 / eval 0.82275 (old engine 0.74658 / 0.82223). Learning, chunk 256, 3 epochs: eval 0.626 -> 0.456 -> 0.430, the same truncated-gradient floor as at F32. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…aphs build explicit attention A quantized V cache requires flash attention (llama_init_from_model refuses it otherwise), and FLASH_ATTN_EXT has no backward. So the training context is created WITH flash attention: the walk's context decodes use it, and the cache stores V un-transposed at q8_0, the same as serving. Every training graph builds explicit attention whatever the context's flag (build_attn_mha under cparams.training), so opt_init no longer asserts flash attention off. At Kimi's 63k on the 27B the cache is ~2.3 GB instead of ~4.6 (F16) or ~9 (F32). 1.5B, chunk 256: forward (lr 1e-9) 0.74694 / eval 0.82323 (old engine 0.74658 / 0.82223); learning, 3 epochs: eval 0.621 -> 0.454 -> 0.483 (the truncated floor). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
Measured on the 5090 (1.5B, 3 examples of ~15k, recompute on): one yield per example made a turn arriving mid-walk wait for the whole walk (max 1173-1758 ms). Yielding before every context-decode piece and every training chunk makes the chunk the window: max 265-345 ms, p95 137-178 ms, same loss (0.4415-0.4416). Open: epoch time rose ~80% with no serving load (3.4 -> 6.1 s at chunk 1024), unexplained so far. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…ure path (returns false, like a refused graph) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
|
Rebased onto 2d4d63b (#41 and #42). #42's backend-failure check now also covers the walk's chunk step. Git had merged it in with a bare |
ad2d9c4 to
539c75d
Compare
|
Approve at 539c75d (Fable). This is the walk as designed: unlabelled runs decoded into the cache; her reply trained in chunks against the cached prefix as constants; each chunk then decoded into the cache under the adapter as it now is; nothing past Two notes, both about Joel's 'align bit depth to inference' (tonight). Fine to fix on top in #47 or a follow-up:
|
What it is: the training walk (Fable's design). A window's context is decoded into the KV cache, and each reply chunk is trained against the cached prefix, which it reads as constants. The chunks are then re-decoded, so the next chunk sees them. A training graph therefore holds one chunk, not the whole window. That's what lets Kimi's 63k-token windows train without truncating anything, which Joel's rule forbids.
Contents:
opt_epoch_iter./traintakes achunkparameter (multiple of 256; interim default 512).Measured on the 5090, rebased onto dc1a379:
Why it's a draft: the gradient is truncated. Cached prefixes are stop-gradient constants, so a chunk's loss doesn't reach the earlier context's activations. Before a gene ships from this, Fable wants an exact reverse pass with horizon H. The test plan: H = 0, 1k, 4k and all, on the 1.5B parity set, accepted when H = all matches the full-window loss (0.222). Also open: the qwen35 recompute gap (card 9355b90c). The two questions are separate.
🤖 Generated with Claude Code
https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc