Automatically generate ArchiMate enterprise architecture models from software repositories.
Deriva analyzes code repositories and transforms them into ArchiMate models that can be opened in the Archi modeling tool.
- Clone a Git repository
- Extraction - Build a graph representation:
- Classify phase: Categorize files by type and subtype using registry
- Parse phase: Extract semantic nodes (TypeDefinitions, Methods, BusinessConcepts, etc.)
- Python files use fast AST parsing; other languages use LLM
- Derivation - Generate ArchiMate elements using a hybrid approach:
- Prep phase: Graph enrichment (PageRank, Louvain communities, k-core)
- Generate phase: LLM-based element derivation with graph metrics
- Refine phase: Relationship derivation and quality assurance
- Export to
.xmlfile (ArchiMate format)
- Python 3.14+
- uv (Python package manager)
# Windows (PowerShell)
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
# macOS/Linux
curl -LsSf https://astral.sh/uv/install.sh | shgit clone https://github.com/StevenBtw/Deriva.git
cd Deriva
# Create environment configuration
cp .env.example .env
# Edit .env with your settings (LLM API keys, etc.)uv venv --python 3.14Activate the virtual environment:
# Windows PowerShell
.venv\Scripts\Activate.ps1
# Windows Command Prompt
.venv\Scripts\activate.bat
# macOS/Linux
source .venv/bin/activateuv syncThe business concept step finds candidate terms in the documentation with pinned spaCy pipelines (English, German, French), which uv sync installs with the other dependencies. It gives German and French terms an English form with pinned translation models, which the step downloads into workspace/cache/nlp on its first run and checks by SHA-256. DERIVA_NLP_MODELS_DIR in .env moves that folder.
cd ../../.. # Back to Deriva root
uv run marimo edit deriva/app/app.pyThe marimo notebook opens in your browser at: http://127.0.0.1:2718
When you first open Deriva, you need to seed the configuration database.
Navigate to Column 2: Manage Extraction → File Type Registry
- Click "Seed from JSON"
- This loads default file type mappings from
extraction_config.json - Categories include: Source, Config, Docs, Test, Build, Asset, Data, Exclude
Navigate to Column 2: Manage Extraction → Extraction Step Configuration
Enable the extraction steps you need:
| Step | Purpose | Recommended |
|---|---|---|
| Repository | Creates root node for the repo | Always |
| Directory | Creates directory structure nodes | Always |
| File | Creates file nodes with classification | Always |
| TypeDefinition | Extracts classes, functions (AST for Python) | Yes |
| Method | Extracts methods from type definitions | Optional |
| Edge | Extracts relationships (IMPORTS, USES, CALLS, DECORATED_BY, REFERENCES) | Yes |
| Technology | Finds infrastructure (runtimes, databases, brokers, container platforms) from manifests, build and container files | Optional |
| ExternalDependency | Maps external dependencies | Optional |
| Test | Extracts test definitions | Optional |
If using LLM-assisted extraction, configure your provider in .env:
# Set default model to use
LLM_DEFAULT_MODEL=mistral-devstral
# Configure the model (naming: LLM_{NAME}_*)
LLM_MISTRAL_DEVSTRAL_PROVIDER=mistral
LLM_MISTRAL_DEVSTRAL_MODEL=devstral-2512
LLM_MISTRAL_DEVSTRAL_URL=https://api.mistral.ai/v1/chat/completions
LLM_MISTRAL_DEVSTRAL_KEY=your-key-here
LLM_MISTRAL_DEVSTRAL_STRUCTURED_OUTPUT=trueColumn 1: Configuration → Repositories
- Enter repository URL (e.g.,
https://github.com/user/repo.git) - Optionally specify a target name
- Click "Clone"
Column 0: Run Deriva
- Click "Run Deriva" to run the full pipeline (extraction → derivation)
- Or use individual step buttons: Extraction, Derivation
Results display in a status callout showing nodes/elements created and any errors.
Column 1: Configuration
- Graph Statistics: Node counts by type (Repository, Directory, File, etc.)
- ArchiMate Model: Element and relationship counts by type
Column 1: Configuration → ArchiMate Model
- Set export path (default:
workspace/output/model.xml) - Click "Export Model"
- Open the file with Archi
Via CLI:
deriva export -o workspace/output/model.xmlAll configuration lives in .env. Key settings:
# Graph database (grafeo embedded)
GRAFEO_DB_DIR= # Empty = in-memory; directory = one <repo>.grafeo database per repository
# LLM Provider (mistral, openai, azure, anthropic, ollama, lmstudio)
LLM_MISTRAL_DEVSTRAL_PROVIDER=mistral
LLM_MISTRAL_DEVSTRAL_MODEL=devstral-2512
LLM_MISTRAL_DEVSTRAL_URL=https://api.mistral.ai/v1/chat/completions
LLM_MISTRAL_DEVSTRAL_KEY=your-mistral-api-key
LLM_MISTRAL_DEVSTRAL_STRUCTURED_OUTPUT=true
# Namespaces
GRAPH_NAMESPACE=Graph
ARCHIMATE_NAMESPACE=Model
# NLP translation models for business concepts (default shown)
DERIVA_NLP_MODELS_DIR=workspace/cache/nlpSee .env.example for all available options.
The LLM adapter includes built-in rate limiting to prevent API throttling:
# Requests per minute (0 = use provider default: 60 RPM for cloud, unlimited for local)
LLM_RATE_LIMIT_RPM=0
# Minimum delay between requests in seconds
LLM_RATE_LIMIT_DELAY=0.0
# Max retries on rate limit (429) errors
LLM_RATE_LIMIT_RETRIES=3
# Adaptive throttling (reduces RPM when hitting rate limits)
LLM_THROTTLE_ENABLED=true
LLM_THROTTLE_MIN_FACTOR=0.25 # Minimum 25% of configured RPM
LLM_THROTTLE_RECOVERY_TIME=60 # Seconds before trying to increase RPM
# Circuit breaker (stops requests when provider is failing)
LLM_CIRCUIT_BREAKER_ENABLED=true
LLM_CIRCUIT_FAILURE_THRESHOLD=5 # Consecutive failures to open circuit
LLM_CIRCUIT_RECOVERY_TIME=30 # Seconds before testing recoveryDefault rate limits by provider:
| Provider | Default RPM |
|---|---|
| OpenAI | 30 |
| Anthropic | 30 |
| Mistral | 24 |
| Ollama | Unlimited |
| LM Studio | Unlimited |
The rate limiter automatically:
- Throttles requests to stay within limits
- Applies exponential backoff on rate limit errors (HTTP 429)
- Respects Retry-After headers from providers
- Adaptively reduces RPM when hitting rate limits (recovers over time)
- Opens circuit breaker after consecutive failures to prevent cascading errors
If you encounter undefined extensions during extraction:
Via UI (Marimo):
- Navigate to Column 2 → Undefined Extensions
- Add them to the registry:
- Extension (e.g.,
.tsx,Dockerfile) - Type (source, config, docs, test, build, asset, data, exclude)
- Subtype (e.g.,
typescript,docker)
- Extension (e.g.,
Via CLI:
# List all registered file types
deriva config filetype list
# Add a new file type
deriva config filetype add ".tsx" source typescript
# Delete a file type
deriva config filetype delete ".tsx"
# Show file type statistics by category
deriva config filetype statsNote: Files with unrecognized extensions are automatically classified as
file_type="unknown"with their extension as the subtype. This ensures all files get proper classification even without explicit registry entries.
Excluded directories: dependency and tool directories are skipped by every repository walk (no Directory or File nodes, no LLM calls), matched as whole path segments. The list is the excluded_directories system setting (JSON list); by default .git, __pycache__, node_modules, bower_components, vendor, .venv, venv and site-packages. Changing it triggers re-extraction.
deriva config setting show excluded_directories
deriva config setting set excluded_directories '[".git", "node_modules", "third_party"]'Derivation name patterns: some derivation steps use include and exclude patterns on the names of their code candidates; business concepts are already classified and skip them. Most of these steps reject a name that contains an exclude pattern and keep one that contains an include pattern, and a name that matches neither follows the step's default (rejected by most steps, kept by BusinessFunction, ApplicationInterface, SystemSoftware and TechnologyService). BusinessObject (type definitions) and BusinessEvent (methods) use the patterns to rank candidates instead, and fill their remaining slots with names that did not match. A step that uses the default candidate filter (ApplicationComponent, ApplicationInterface, BusinessFunction, DataObject, Device, Node, SystemSoftware, TechnologyService) can limit the patterns to candidates with given graph labels with its pattern_labels param, and set a k-core threshold with graph_filter. Patterns are stored per step, type and category.
deriva config pattern list Node
deriva config pattern add Node include deployment helm
deriva config pattern delete Node include --category deployment helm # a category left empty is deactivatedDeriva uses a versioning system for configurations. When you update a config, a new version is created while preserving previous versions for rollback.
Correct ways to update configs:
- Via UI (Marimo): Navigate to the config section, edit, and click "Save Config"
- Via CLI: Use the
config updatecommand
# Update extraction config instruction
deriva config update extraction BusinessConcept \
-i "New instruction text..."
# Update extraction config with batch size for multi-file LLM calls
deriva config update extraction BusinessConcept \
--batch-size 5
# Update derivation config from file
deriva config update derivation ApplicationComponent \
--instruction-file prompts/app_component.txt
# View all versions
deriva config versionsDo NOT use JSON import/export for config updates. The db_tool import command is only for backup restoration or migration - it overwrites version history. See BENCHMARKS.md for the optimization workflow.
For LLM-assisted extraction steps:
- Navigate to Column 2 → Extraction Step Configuration
- Expand a node type (e.g., TypeDefinition)
- Edit: Input File Types, Input Graph Elements, Instruction, Example
- Click "Save Config" (this creates a new version)
All prompts follow the Input + Instruction + Example pattern.
Deriva uses a multi-column marimo notebook layout:
| Column | Purpose |
|---|---|
| 0 | Run Deriva: Pipeline execution buttons, status display |
| 1 | Configuration: Runs, repositories, graph database, graph stats, ArchiMate, LLM |
| 2 | Extraction Settings: File type registry, extraction step configuration |
| 3 | Derivation Settings: Element type configuration (13 types across Business/Application/Technology layers), relationship derivation |
The UI is powered by PipelineSession from the services layer, providing a clean separation between presentation and business logic.
- Grafeo (embedded graph database):
- Graph namespace: Intermediate representation (Modules, Files, Dependencies)
- Model namespace: ArchiMate elements and relationships
- DuckDB (
deriva/adapters/database/sql.db): File type registry, extraction configs, settings
Column 0: Run Overview
- Clear Graph: Removes all nodes/edges from Graph namespace
- Clear Model: Removes all ArchiMate elements and relationships
You can query the embedded grafeo graph database using Cypher via the CLI or Marimo notebook:
// See all repositories
MATCH (r:Graph:Repository) RETURN r
// See files in a repo
MATCH (repo:Graph:Repository)-[:Graph:CONTAINS*]->(f:Graph:File)
WHERE repo.name = 'my-repo'
RETURN f.name, f.file_type
// See type definitions
MATCH (td:Graph:TypeDefinition) RETURN td.name, td.type_categoryDeriva includes a full CLI for headless operation and automation:
# Help
deriva --help
# View configuration
deriva config list extraction
deriva config show extraction BusinessConcept
deriva status
# Add a derivation step (created disabled, then enable it); a refine step
# must also be implemented and registered in code under the same name
deriva config add derivation my_refine_step --phase refine --sequence 4 --params '{"dry_run": true}'
deriva config enable derivation my_refine_step
# Manage file types
deriva config filetype list
deriva config filetype add ".lock" dependency lock
deriva config filetype stats
# System settings (e.g. directories skipped during extraction)
deriva config setting show excluded_directories
# Derivation name patterns (include and exclude, per step and category)
deriva config pattern list SystemSoftware
# Run pipeline stages
deriva run extraction --repo flask_invoice_generator -v
deriva run derivation -v
deriva run derivation --phase generate -v # Run specific phase (prep, generate, refine)
deriva run all --repo myrepo
# Export ArchiMate model
deriva export -o workspace/output/model.xmlCLI Options:
| Option | Description |
|---|---|
--repo NAME |
Process specific repository (default: all) |
--phase PHASE |
Run specific derivation phase: prep, generate, or refine |
-v, --verbose |
Print detailed progress |
--no-llm |
Skip LLM-based steps (structural extraction only) |
-o, --output PATH |
Output file path for export |
Deriva includes a multi-model benchmarking system for comparing LLM performance across different providers and models. See BENCHMARKS.md for the full guide and OPTIMIZATION.md for detailed case studies.
# List available benchmark models
deriva benchmark models
# Run a benchmark with specific models
deriva benchmark run \
--repos flask_invoice_generator \
--models openai-gptx,ollama-devstral \
-n 3 \
-d "Comparing gptx with devstral" \
-v
# List benchmark sessions
deriva benchmark list
# Analyze a benchmark session
deriva benchmark analyze bench_20260101_150724Add models to .env using the pattern:
# Azure GPT-4o-mini
LLM_AZURE_GPT4MINI_PROVIDER=azure
LLM_AZURE_GPT4MINI_MODEL=gpt-4
LLM_AZURE_GPT4MINI_URL=https://your-resource.openai.azure.com/...
LLM_AZURE_GPT4MINI_KEY=your-api-key
# Ollama local model
LLM_OLLAMA_LLAMA_PROVIDER=ollama
LLM_OLLAMA_LLAMA_MODEL=devstral
LLM_OLLAMA_LLAMA_URL=http://localhost:11434/api/chatBenchmark runs are logged in OCEL 2.0 (Object-Centric Event Log) format for process mining analysis:
- Events capture pipeline stages, LLM calls, and results
- Object types:
BenchmarkSession,BenchmarkRun,Repository,Model - Logs are saved to
workspace/benchmarks/{session_id}/events.ocel.json
OCEL files can be analyzed with process mining tools like PM4Py, Celonis, or custom analysis scripts.
The business concept step downloads its translation models on the first run. If the download fails, check that the machine can reach the model host and run the step again; a model folder that was not verified is replaced. A SHA-256 mismatch means the downloaded file is not the pinned model (a changed upload or a damaged download), and the step refuses it.
# Check Python version
python --version # Should be 3.14+
# Reinstall dependencies
uv sync --reinstall
# Run without watch mode
uv run marimo edit deriva/app/app.pyFor development setup, architecture details, and contribution guidelines, see CONTRIBUTING.md.
This project is licensed under the GNU Affero General Public License v3.0 (AGPL-3.0).
This means you can freely use, modify, and distribute this software, but if you run a modified version as a network service, you must make the source code available to users of that service.
See LICENSE for the full license text.
- Marimo - Reactive Python notebooks
- Grafeo - Embedded graph database
- ArchiMate - Enterprise architecture standard
- Archi - Open source ArchiMate modeling tool
- Tree-sitter - Multi-language AST parsing
Status: Active Development
