Skip to content

Add reproducible synthetic retail datasets as an optional loader plugin - #469

Open
Ravi Kiran Pagidi (ravikiranpagidi) wants to merge 1 commit into
microsoft:mainfrom
ravikiranpagidi:synthetic-retail-plugin
Open

Ravi Kiran Pagidi (ravikiranpagidi) wants to merge 1 commit into
microsoft:mainfrom
ravikiranpagidi:synthetic-retail-plugin

Conversation

@ravikiranpagidi

Copy link
Copy Markdown

Analysts and workshop instructors can use this optional loader to explore monthly sales by product category, store region, and customer segment without connecting a production database. It provides a configurable, repeatable retail star schema through the existing drop-in plugin interface.

The example uses the published great-generator==0.1.8 package to generate related customer, product, store, date, and sales tables from an embedded SQL contract. It derives gross, discount, and net amounts from product prices, quantities, and customer segments. Table descriptions identify the data as synthetic and explain the join keys.

  • Adds one self-contained example plugin, a setup/analysis walkthrough, and a link from the plugin guide. Users install the optional package in Data Formulator's environment and copy the plugin file into their plugin directory. Core dependencies and built-in loader registration are unchanged.
  • Keeps catalog discovery lightweight and generates the complete dataset on the first fetch. Preview/import filters, sorting, projection, and row limits operate on the same cached tables. Seed and bounded sales-row configuration reproduce the dataset across refreshes and new instances within the same dependency environment.
  • Includes tests for relational integrity, sales arithmetic, concurrency, import options, plugin registration, missing-dependency handling, workspace ingestion, and agent probes. The example tests skip when the optional package is absent.

Validation

  • uv run --no-sync --python .venv/Scripts/python.exe pytest tests/backend/data/test_synthetic_retail_plugin.py tests/backend/data/test_plugin_scanner.py -q — 42 passed, on Windows / Python 3.13 with the published Great Generator 0.1.8 package. DuckDB emits deprecation warnings for the existing fetch_arrow_table() API, also used by the shared probe helper.
  • Generated all five tables at the 100,000-sales-row cap: 200 customers, 50 products, 10 stores, 365 dates, and 100,000 sales; approximately 16.6 seconds and 7.33 MB of final Arrow data in this environment (not a peak-memory measurement).
  • git diff --check passed.

The documentation targets builds with the current plugin interface, explains that complete dimensions should be imported for joins, and limits reproducibility claims to unchanged dependency versions. Browser/model-driven chart creation and the full backend suite were not run.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant