Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 14 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,20 @@
Alle wesentlichen Änderungen an diesem Projekt werden hier dokumentiert.
Format basiert auf [Keep a Changelog](https://keepachangelog.com/de/1.1.0/).

## [Unreleased — 2026-10-06]

### Fixed
- GUI workers (`QThread`) are now kept alive until `finished` (`src/gui/worker_utils.py`); a running extraction is never replaced, queued files are no longer dropped, and `closeEvent` stops and waits for all workers before saving.
- HTML export escapes all Markdown text and the document title and only renders `http`/`https`/`mailto`/relative links; the GUI HTML export uses the same `ReportExporter` implementation. TXT export strips Markdown syntax only (keeps `C#`, `#12`, `file_name`).
- Text extraction no longer triggers synchronous embedding on the GUI thread; documents are queued and indexed by the `IndexWorker`. Re-indexing embeds before deleting old chunks; failed indexing resets `is_indexed`; stale `is_indexed` flags are reconciled with the vector index when a project is activated.
- RAG search uses relevance scores (higher = better) instead of raw distances for confidence and thresholds.
- Chat always receives the current project's document manager (no unfiltered queries over the shared collection) and a disabled RAG engine is propagated as `None`.
- New projects inherit provider/model/URL from the saved app config; the LLM client is initialised at startup; projects are opened by ID, the current project is saved first, and project directories get a unique suffix.
- Supported file types come from one shared set (adds `.pptx`, `.html`, `.htm`; drops unsupported `.odt`, `.ods`); HTML is converted to plain text.
- Extractor: Excel `0`/`False` cells kept, RTF `\uN` fallback skipping and multi-byte code pages fixed, BOM/cp1252 detection for plain text, PPTX slides sorted numerically, PDF/Excel/MSG handles closed reliably.
- Ollama availability is re-checked lazily instead of being cached forever after a failed start-up check.
- YAML front matter quotes title/author, GUI PDF export has a timeout, error message and temp-file cleanup, chat messages render as plain text, translator hint matching uses whole words, and the Web Companion stores review notes per project ID (new optional `workspace.id` in the export).

## [Unreleased — 2026-09-22]

### Added
Expand Down
1 change: 1 addition & 0 deletions EXPORTFORMAT.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@ Review-Notizen als Markdown exportiert.
"exported_at": "2026-05-26T00:00:00Z"
},
"workspace": {
"id": "8f0c6a2e-…",
"title": "Projektname",
"question": "Zentrale Fragestellung",
"workflow_type": "analysis",
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,8 @@
[![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/)
[![Platform: Windows](https://img.shields.io/badge/platform-Windows-lightgrey.svg)]()
[![Offline-first](https://img.shields.io/badge/offline--first-yes-green.svg)]()
[![Pytest](https://img.shields.io/badge/Pytest-103%20passed-brightgreen.svg)]()
[![Web Companion Tests](https://img.shields.io/badge/Web%20Companion-58%20passed-brightgreen.svg)]()
[![Pytest](https://img.shields.io/badge/Pytest-147%20passed-brightgreen.svg)]()
[![Web Companion Tests](https://img.shields.io/badge/Web%20Companion-60%20passed-brightgreen.svg)]()
[![Notice: Invariants](https://img.shields.io/badge/Notice-INV--LOCAL--01..10-blue.svg)](NOTICE)
[![SBOM: Level 1](https://img.shields.io/badge/SBOM-Level%201-blue.svg)](THIRD_PARTY_LICENSES.md)
[![Ecosystem: file-bricks](https://img.shields.io/badge/Ecosystem-file--bricks-blue.svg)](https://github.com/file-bricks)
Expand Down
4 changes: 2 additions & 2 deletions README_de.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,8 @@
[![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/)
[![Plattform: Windows](https://img.shields.io/badge/Plattform-Windows-lightgrey.svg)]()
[![Offline-first](https://img.shields.io/badge/offline--first-ja-green.svg)]()
[![Pytest](https://img.shields.io/badge/Pytest-103%20bestanden-brightgreen.svg)]()
[![Web Companion Tests](https://img.shields.io/badge/Web%20Companion-58%20bestanden-brightgreen.svg)]()
[![Pytest](https://img.shields.io/badge/Pytest-147%20bestanden-brightgreen.svg)]()
[![Web Companion Tests](https://img.shields.io/badge/Web%20Companion-60%20bestanden-brightgreen.svg)]()
[![Notice: Invarianten](https://img.shields.io/badge/Notice-INV--LOCAL--01..10-blue.svg)](NOTICE)
[![SBOM: Level 1](https://img.shields.io/badge/SBOM-Level%201-blue.svg)](THIRD_PARTY_LICENSES.md)
[![Ökosystem: file-bricks](https://img.shields.io/badge/%C3%96kosystem-file--bricks-blue.svg)](https://github.com/file-bricks)
Expand Down
2 changes: 1 addition & 1 deletion llms.txt
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

## Last-checked: 2026-09-22

> Local NotebookLM alternative for private document analysis, RAG-assisted research, and multi-format report generation. Verified with 161 passing unit tests (103 Python pytest + 58 Web Companion Node.js tests).
> Local NotebookLM alternative for private document analysis, RAG-assisted research, and multi-format report generation. Verified with 207 passing unit tests (147 Python pytest + 60 Web Companion Node.js tests).

## Description

Expand Down
118 changes: 94 additions & 24 deletions src/core/document_manager.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,8 @@
from typing import Dict, List, Optional, Set, Callable, TYPE_CHECKING
import uuid

from .text_extractor import SUPPORTED_EXTENSIONS as _EXTRACTOR_SUPPORTED_EXTENSIONS

if TYPE_CHECKING:
from ..rag.engine import RAGEngine

Expand Down Expand Up @@ -147,21 +149,8 @@ class DocumentManager:
- Manage sub-query associations
"""

# Supported file types
SUPPORTED_EXTENSIONS = {
# Text
".txt", ".md", ".rst", ".log",
# Documents
".pdf", ".docx", ".doc", ".odt", ".rtf",
# Data
".json", ".xml", ".yaml", ".yml", ".csv",
# Code (for analysis)
".py", ".js", ".ts", ".java", ".cpp", ".c", ".h",
# Spreadsheets
".xlsx", ".xls", ".ods",
# Email
".eml", ".msg"
}
# Supported file types -- abgeleitet aus dem TextExtractor (eine Quelle der Wahrheit)
SUPPORTED_EXTENSIONS = _EXTRACTOR_SUPPORTED_EXTENSIONS

def __init__(self, project_path: Optional[Path] = None, rag_engine: Optional['RAGEngine'] = None):
"""
Expand All @@ -179,6 +168,7 @@ def __init__(self, project_path: Optional[Path] = None, rag_engine: Optional['RA
self._auto_extract: bool = True # Automatisch Text extrahieren bei add_file
self._text_extractor = None # Lazy-loaded
self._pending_extractions: List[str] = [] # doc_ids waiting for async extraction
self._pending_index: List[str] = [] # doc_ids waiting for async (re-)indexing

if project_path:
self._cache_dir = project_path / ".cache"
Expand Down Expand Up @@ -492,15 +482,41 @@ def _try_auto_extract(self, doc: DocumentItem) -> None:
self._notify_change("update", doc)
logger.error(f"Auto-Extraktion Fehler: {doc.name}: {e}")

def set_rag_engine(self, rag_engine: 'RAGEngine') -> None:
def set_rag_engine(self, rag_engine: Optional['RAGEngine']) -> None:
"""
Setzt die RAG Engine für semantische Suche.

Args:
rag_engine: Die RAG Engine Instanz
rag_engine: Die RAG Engine Instanz oder None zum Trennen
"""
self._rag_engine = rag_engine
logger.info("RAG Engine verbunden")
if rag_engine is None:
self._pending_index.clear()
logger.info("RAG Engine getrennt")
else:
logger.info("RAG Engine verbunden")

def sync_index_flags(self) -> None:
"""Gleicht is_indexed mit dem tatsächlichen Inhalt des Vektor-Index ab.

Nötig, weil der Index projektübergreifend geleert werden kann und
documents.json dann veraltete is_indexed-Flags enthält.
"""
if not self._rag_engine:
return
flagged = [d for d in self._documents.values() if d.is_indexed and not d.is_directory]
if not flagged:
return
try:
present = self._rag_engine.get_indexed_document_ids([d.id for d in flagged])
except Exception as e:
logger.warning("Index-Abgleich fehlgeschlagen: %s", e)
return
for doc in flagged:
if doc.id not in present:
doc.is_indexed = False
doc.chunk_count = 0
self._notify_change("deindexed", doc)

def pop_pending_extractions(self) -> List[tuple]:
"""Gibt ausstehende Extraktionen zurück und leert die Queue.
Expand All @@ -516,6 +532,38 @@ def pop_pending_extractions(self) -> List[tuple]:
self._pending_extractions.clear()
return result

def has_pending_extractions(self) -> bool:
"""True, wenn Dokumente auf die asynchrone Extraktion warten."""
return any(
doc_id in self._documents and not self._documents[doc_id].is_directory
for doc_id in self._pending_extractions
)

def discard_pending_extractions(self, doc_ids) -> None:
"""Entfernt Dokumente aus der Extraktions-Queue (z.B. weil sie bereits extrahiert werden)."""
ids = set(doc_ids)
self._pending_extractions = [d for d in self._pending_extractions if d not in ids]

def pop_pending_index(self) -> List[tuple]:
"""Gibt Dokumente zurück, die (neu) indexiert werden sollen, und leert die Queue.

Returns:
Liste von (doc_id, doc_name) Tupeln
"""
result = []
seen = set()
for doc_id in self._pending_index:
doc = self._documents.get(doc_id)
if doc and not doc.is_directory and doc.extracted_text and doc_id not in seen:
seen.add(doc_id)
result.append((doc.id, doc.name))
self._pending_index.clear()
return result

def index_pending_documents(self) -> Dict[str, bool]:
"""Indexiert alle wartenden Dokumente synchron (nur außerhalb des GUI-Threads nutzen)."""
return {doc_id: self.index_document(doc_id) for doc_id, _ in self.pop_pending_index()}

def set_auto_index(self, enabled: bool) -> None:
"""Aktiviert/Deaktiviert automatische Indexierung."""
self._auto_index = enabled
Expand Down Expand Up @@ -562,12 +610,21 @@ def index_document(self, doc_id: str) -> bool:
return True
else:
logger.error(f"Indexierung fehlgeschlagen: {result.error}")
self._mark_not_indexed(doc)
return False

except Exception as e:
logger.error(f"Fehler bei Indexierung von {doc_id}: {e}")
self._mark_not_indexed(doc)
return False

def _mark_not_indexed(self, doc: DocumentItem) -> None:
"""Setzt den Index-Status nach einem Fehlschlag zurück."""
if doc.is_indexed or doc.chunk_count:
doc.is_indexed = False
doc.chunk_count = 0
self._notify_change("deindexed", doc)

def index_all_documents(self) -> Dict[str, bool]:
"""
Indexiert alle Dokumente mit extrahiertem Text.
Expand Down Expand Up @@ -741,17 +798,30 @@ def get_rag_statistics(self) -> Dict:

return stats

# Override update_content to auto-index
def update_content(self, doc_id: str, text: str) -> None:
"""Update extracted text for a document."""
"""Update extracted text for a document.

Indexiert NICHT synchron (Embedding-Aufrufe gehen über HTTP und würden
den GUI-Thread blockieren). Bei aktivem Auto-Index wird das Dokument in
eine Queue gestellt, die die GUI per IndexWorker abarbeitet
(pop_pending_index) bzw. Nicht-GUI-Aufrufer per index_pending_documents().
"""
doc = self._documents.get(doc_id)
if doc:
new_hash = hashlib.md5(text.encode()).hexdigest()
content_changed = new_hash != doc.content_hash
doc.extracted_text = text
doc.text_length = len(text)
doc.content_hash = hashlib.md5(text.encode()).hexdigest()
doc.content_hash = new_hash
doc.status = DocumentStatus.READY
if content_changed and doc.is_indexed:
# Vorhandene Chunks passen nicht mehr zum Text
doc.is_indexed = False
doc.chunk_count = 0
self._notify_change("update", doc)

# Auto-indexieren wenn aktiviert
if self._auto_index and self._rag_engine and text:
self.index_document(doc_id)
# Auto-Indexierung vormerken (asynchron)
needs_index = content_changed or not doc.is_indexed
if (self._auto_index and self._rag_engine and text and needs_index
and doc_id not in self._pending_index):
self._pending_index.append(doc_id)
43 changes: 31 additions & 12 deletions src/core/project.py
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,11 @@ def to_dict(self) -> dict:
"language": self.language
}

@classmethod
def from_app_config(cls) -> "ProjectSettings":
"""Neue Projekt-Settings mit den gespeicherten LLM-Einstellungen der App."""
return cls.from_dict({})

@classmethod
def from_dict(cls, data: dict) -> "ProjectSettings":
from .app_config import get_app_config
Expand Down Expand Up @@ -126,7 +131,9 @@ def create(cls, name: str, main_question: str = "", report_type: str = "analysis
id=str(uuid.uuid4()),
name=name,
main_question=main_question,
report_type=report_type
report_type=report_type,
# LLM-Provider/Modell/URL aus der gespeicherten App-Konfiguration
settings=ProjectSettings.from_app_config(),
)

@property
Expand Down Expand Up @@ -359,23 +366,35 @@ def open_project(self, project_id_or_name: str) -> Optional[Project]:
Returns:
The project if found
"""
# Search by ID or name
for item in self.projects_dir.iterdir():
# ID hat Vorrang vor dem (nicht eindeutigen) Namen
by_name = []
for item in sorted(self.projects_dir.iterdir()):
if item.is_dir():
project_file = item / "project.json"
if project_file.exists():
try:
data = json.loads(project_file.read_text(encoding="utf-8"))
if data["id"] == project_id_or_name or data["name"] == project_id_or_name:
project = Project.load(item)
if project:
self._current_project = project
return project
except Exception:
pass
except (OSError, ValueError):
continue # beschädigte project.json überspringen
if data.get("id") == project_id_or_name:
return self._load_as_current(item)
if data.get("name") == project_id_or_name:
by_name.append((data.get("modified_at", data.get("created_at", "")), item))

if by_name:
# Bei Namensdubletten das zuletzt geänderte Projekt öffnen
by_name.sort(key=lambda entry: entry[0], reverse=True)
return self._load_as_current(by_name[0][1])

return None

def _load_as_current(self, directory: Path) -> Optional[Project]:
"""Load a project directory and make it the current project."""
project = Project.load(directory)
if project:
self._current_project = project
return project

def save_current(self) -> bool:
"""Save the current project."""
if not self._current_project:
Expand Down Expand Up @@ -438,9 +457,9 @@ def _safe_dirname(self, name: str) -> str:
safe = "".join(c for c in name if c.isalnum() or c in " -_").strip()
safe = safe.replace(" ", "_")

# Add timestamp for uniqueness
# Timestamp + Kurz-UUID: zwei Projekte in derselben Sekunde kollidieren nicht
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
return f"{safe}_{timestamp}"
return f"{safe or 'Projekt'}_{timestamp}_{uuid.uuid4().hex[:8]}"

# Output profiles
def get_output_profiles(self) -> List[OutputProfile]:
Expand Down
Loading
Loading