Repository navigation
ggml-opt: graphs built per step start every optimizer period from zero gradients - #45
Conversation
…o gradients ggml_opt_prepare_alloc (the path llama_context's training takes: a graph per ubatch) keeps the gradient accumulators in ctx_static across graphs, and each backward adds into them in place. The only reset sat before the build, against gb_grad, which is null in this mode, so nothing ever zeroed them: step k applied the SUM of every gradient since the run began. Measured (test-opt-dynamic-accum, SGD on loss = w*x): the second step moved w by 2*lr*x, not lr*x. Upstream has the same code. The fix zeroes the parameter accumulators after the per-step build when a period begins. Static graphs are unchanged (test-opt: 4/4 backend x optimizer pass). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
|
Approve at e2278fb (Fable). Confirmed by reading ggml-opt.cpp on the base. In per-step mode Verified the test is a real gate on the M5 (CPU build of this head): with the fix, One hardening note, not blocking: the loop runs Worth an upstream PR too: ggml-org has the same code. |
…nodes (Fable on #45) grad_accs is indexed by the first graph's nodes; a later per-step graph may hold fewer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…kend-DL build links no ggml_backend_cpu_init) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
ggml_opt_prepare_alloc, the pathllama_contexttraining takes (it builds a graph per ubatch), keeps the gradient accumulators inctx_staticacross graphs, and each backward adds into them in place. The only reset runs before the build, againstgb_grad, which is null in this mode. So nothing ever zeroed the accumulators, and step k applied the sum of every gradient since the run began. Upstream ggml-org has the same code.Measured.
test-opt-dynamic-accum(new) runs SGD on loss = w·x, so every step's gradient is x. On the current code, the second step moved w by 2·lr·x instead of lr·x. With the fix, both opt_period 1 and 2 pass, andtest-optstill passes 4/4 backend × optimizer combinations (static graphs are unchanged).The fix: after the per-step build, when a period begins, zero the parameter accumulators.
Effect on real training (same data, seed and lr; train/eval loss per epoch):
The bug behaved like unbounded momentum: it learned faster early, then overshot, and the 1.5B's eval loss rises in epoch 3. With the fix, eval loss is still falling at the end. On the 1.5B it ends lower; on the short 0.8B run it is slower, because the learning rate the core tuned while the bug was present is now too small. The lr default should be re-measured after this lands (a continuum-side follow-up).
This is also a prerequisite for the walk's exact gradient (#40 follow-up), which accumulates gradients across chunks and must start each window from zero.
🤖 Generated with Claude Code
https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc