Skip to content

Whole repo update loop #113

Description

@fabceolin

Summary

codewiki generate --update --update-rung <rung> (the component-level updater in codewiki/src/be/updater/orchestrator.py, introduced in #111) enters an unbounded, non-converging loop when the baseline documentation was itself generated in whole-repository mode (Total modules: 0, meaning the repo or --include scope was too small for LLM clustering, see the else branch at documentation_generator.py:278).

Instead of a small incremental patch, a single --update invocation on a real, small PR (4 files changed, +112/-1) ran for over 2 hours and never terminated on its own (I killed the process manually). It produced 50 markdown files and roughly 258k output tokens: repeated, colliding regenerations of the same ~9-10 modules under names like overview_approval.md, then overview_overview_approval.md, then overview_overview_approval_2.md, then _3, _4, _5 (killed before it reached _6).

Root cause (traced in the actual code, not guessed)

  1. orchestrator.py:78-87 (_load_state): when the on-disk module_tree.json has len(tree) == 0 (our case: the baseline was bootstrapped with a narrow --include that produced only 1 leaf node, so clustering was skipped and the whole-repo branch ran), the updater sets self.whole_repo = True and builds a virtual single-leaf tree (T.virtual_whole_repo_tree) instead of a real module tree.
  2. orchestrator.py:288-295: the fallback-to-full-rebuild check normally aborts cleanly when the diff is large relative to the tracked graph (ratios["fired"]). When self.whole_repo is true, this gets unconditionally overridden to False:
    if self.whole_repo and ratios["fired"]:
        # One virtual leaf: any change is 100% active by construction, and
        # patching that single page is exactly the incremental step.
        ratios["fired"] = False
        ratios["note"] = "whole-repository mode: fallback rule not applied"
    The comment's premise (patching that single page is exactly the incremental step) only holds when the diff is small. The branch never checks the actual size of the diff or graph, so a repo-wide change, or a diff against a baseline that never had a real tree, gets treated the same as a one-line change.
  3. orchestrator.py:309: persisting the repaired tree back to module_tree.json is explicitly skipped in whole-repo mode (if not self.whole_repo: file_manager.save_json(...)). After Step 5, the on-disk module_tree.json is still whatever it was before this --update run (in our case, 0 entries).
  4. orchestrator.py:349-358 (Step 6.1) then unconditionally calls self.doc_generator.generate_module_documentation(new_graph, leaf_nodes), the normal, unbounded bootstrap pipeline, passing the entire current repo graph rather than just the diff.
  5. documentation_generator.py:210-232 loads module_tree.json from disk to decide which branch to take. Because of point 3 it's still empty, so len(module_tree) > 0 is False and it falls into the else branch at documentation_generator.py:278-296 again. That dumps the entire repo's leaf_nodes into a single unconstrained run_module_agent() call, bypassing the bounded, idempotent per-module Phase 1-3 pipeline that real clustering would use.
  6. That single agent call has its own tool (generate_sub_module_documentation_tool) to spin off sub-module pages on the fly. Fed a repo-scale leaf_nodes set instead of the handful of components whole-repo mode was designed for, it seems to have no bounded stopping criterion or duplicate-name detection: it kept re-deriving and regenerating the same ~9-10 sub-modules under incrementally suffixed names for over 2 hours without converging.

Why this didn't show up in our other tests

Any baseline that goes through real LLM clustering (module_tree.json with real modules) never sets self.whole_repo = True, so none of the above applies. The fallback threshold works normally there; we confirmed clean full_fallback behavior on a 63-file diff in a different repo/test. This only triggers when the baseline itself was generated in whole-repo mode and --update is then run against a real diff.

Suggested fix direction

  • Don't blindly override ratios["fired"] = False for whole_repo. Only skip the fallback when the diff itself is actually small, for example by still gating on r_leaf/r_tree, or on the size of new_graph/leaf_nodes against the clustering threshold that would normally trigger real clustering.
  • Alternatively, persist new_tree (or otherwise mark the tree as non-empty) before Step 6.1 runs, so generate_module_documentation doesn't re-enter the unconstrained whole-repo branch a second time in the same invocation.
  • At minimum, the whole-repo agent's sub-module-doc tool loop needs a hard bound (max sub-modules, or dedup by content hash before creating an incrementally-suffixed file) so a misfire like this fails loudly instead of running unbounded.

Repro shape (generalized, not repo-specific)

  1. Bootstrap a repo with codewiki generate --output ./docs --include <a handful of files that parse into ~1 leaf node>. This produces module_tree.json with 0 entries and a single overview.md (whole-repo mode, as intended for tiny scopes).
  2. Without ever bootstrapping the full repo, run codewiki generate --output ./docs --update --update-rung 3b from a worktree/branch with a normal-sized diff elsewhere in the repo (a few files changed, unrelated to the tiny --include scope).
  3. Expected: a bounded incremental update, or a clean, declared full_fallback as documented for the non-whole-repo case. Observed: an unbounded run, a growing set of colliding module files, and no natural termination.

Environment

  • codewiki v2.0.0, installed from this repo (pip install -e .), HEAD at the merge of Feat/component level updater #111 (component-level updater) at the time of testing, no local modifications to codewiki/src/be/updater/ or documentation_generator.py.
  • Main model: anthropic/claude-sonnet-4-6 via an OpenAI-compatible proxy. Cluster model: anthropic/claude-haiku-4-5.

I can share the exact (sanitized) module_tree.json states and the generated .md files if that helps; the source repo itself is private, but nothing in this report depends on its content.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions