Skip to content

Fix #2453: [Bug] sources array duplicates the full original message N times (N = chunk coun - #2454

Open
Memtensor-AI wants to merge 2 commits into
MemTensor:dev-v2.0.36from
Memtensor-AI:bugfix/autodev-2453-20261003090152992
Open

Memtensor-AI wants to merge 2 commits into
MemTensor:dev-v2.0.36from
Memtensor-AI:bugfix/autodev-2453-20261003090152992

Conversation

@Memtensor-AI

Copy link
Copy Markdown
Collaborator

Description

Fixes issue #2453: sources array duplicated per chunk in MultiModalStructMemReader.

Root cause: _split_large_memory_item attached the parent sources list to every chunk (multi_modal_struct.py:169), and _build_window_from_items then concatenated them without dedup, so a message split into N chunks produced memory items whose sources list carried N byte-identical copies of the original message. The reporter measured 27% of production memories affected.

Fix is two-layer: (1) attach sources only to chunk 0 (the "owning" chunk) with an additional internal_info["source_chunk_index"] = 0 marker; (2) dedup all_sources in _build_window_from_items by a stable (type, role, message_id, chat_time, doc_path, content) signature for defense-in-depth. Lineage keys (ingest_batch_id / chunk_index / chunk_total) are preserved on every chunk.

Testing: 8 new regression cases in tests/mem_reader/test_sources_dedup.py (4 reproduced the bug on pristine HEAD; all 8 pass after the fix). Full tests/mem_reader/ suite: 71 pass. Pre-existing failures from a missing markitdown dependency and a torch/transformers DynamicCache compat issue in tests/memories/activation/test_kv.py are unrelated (verified on pristine HEAD). Ruff check + format: clean.

Related Issue (Required): Fixes #2453

Type of change

Please delete options that are not relevant.

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Refactor (does not change functionality, e.g. code style improvements, linting)
  • Documentation update

How Has This Been Tested?

Not run; documentation-only change.

  • Unit Test
  • Test Script Or Test Steps (please provide)
  • Pipeline Automated API Test (please provide)

Checklist

  • I have performed a self-review of my own code
  • I have commented my code in hard-to-understand areas
  • I have added tests that prove my fix is effective or that my feature works
  • I have created related documentation issue/PR in MemOS-Docs (if applicable)
  • I have linked the issue to this PR (if applicable)
  • I have mentioned the person who will review this PR

@WeiminLee please review this PR.

Reviewer Checklist

…2453)

MultiModalStructMemReader._split_large_memory_item was attaching the parent
item's full `sources` list to every chunk, and _build_window_from_items
then concatenated them without dedup. The result: a message split into N
chunks produced memory items whose `sources` list held N byte-identical
copies of the original message. Production data showed 27% of memories
affected, bloating storage, retrieval context (advanced_searcher re-injects
sources into prompts verbatim), and provenance tooling.

Fix is two-layer:
  * In _split_large_memory_item, attach `sources` only to chunk 0
    (the owning chunk) and record source_chunk_index=0 in internal_info.
    Non-owning chunks keep ingest_batch_id + chunk_index + chunk_total
    for lineage discoverability.
  * In _build_window_from_items, dedup `all_sources` by a stable
    (type, role, message_id, chat_time, doc_path, content) signature so
    future regressions or multi-parser overlap also cannot inflate the
    list.

Added 8 regression tests in tests/mem_reader/test_sources_dedup.py
covering the propagation fix, the dedup fix, role-detection preservation,
and an end-to-end _concat_multi_modal_memories invariant.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Memtensor-AI Memtensor-AI added ai:generated Generated or modified by AI | 由 AI 生成或修改 area:memory 记忆存储、检索、更新、召回逻辑 status:in-progress Someone or AI is working on it | 人工或 AI 正在处理 labels Oct 3, 2026
@Memtensor-AI

Memtensor-AI commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator Author

🤖 Open Code Review

Target: PR #2454
Task: c72466b93711d6d9
Base: dev-v2.0.36
Head: bugfix/autodev-2453-20261003090152992
Head SHA: 73621315b028232af3b75179361af75b12f58cfa

✅ OpenCodeReview: Review complete: 0 finding(s) across 2 selected item(s).

Generated by cloud-assistant via Open Code Review.

@Memtensor-AI

Copy link
Copy Markdown
Collaborator Author

🔧 Open Code Review requested Agent fix

Open Code Review found 4 issue(s). I have resumed the development Agent to fix them.

  • Task: c72466b93711d6d9
  • Fix attempt: 1/2
  • Finding delta: 0 repeated / 4 new / 0 likely resolved

The Agent will push a new commit to this PR branch. OCR will recheck after the commit is pushed.

Four follow-ups from the open code review on MemTensor#2454:

1. `_source_signature` now folds in all non-None keys returned by
   `model_dump` instead of only the six known fields. SourceMessage
   declares `extra="allow"`, so two paragraphs from the same document
   that differ only in an extra locator (page, offset, span, …) and
   share `message_id=None` would otherwise collapse incorrectly after
   dedup. Unhashable extras fall back to object identity so we never
   silently merge distinct sources.

2. `test_long_message_produces_windows_with_single_source` additionally
   asserts that at least one window actually carries the owning source.
   The pre-existing `≤ 1` check alone would pass a buggy implementation
   that dropped the source entirely.

3. `test_split_bails_before_marking_when_only_one_chunk` now matches
   the implementation: chunk 0 is the owning chunk even when
   chunk_total==1, so `source_chunk_index=0` is set and the test
   asserts it. Docstring updated.

4. `test_window_role_detection_still_works_after_dedup` additionally
   asserts `len(window.metadata.sources) == 1` so a dedup regression
   that strips every source cannot slip through on role defaulting.

Also adds `test_window_keeps_sources_distinguished_by_extra_fields`
to guard the first fix against future regressions.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ai:generated Generated or modified by AI | 由 AI 生成或修改 area:memory 记忆存储、检索、更新、召回逻辑 status:in-progress Someone or AI is working on it | 人工或 AI 正在处理

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants