Skip to content

Fixes and improvements - #428

Open
Chenglong Wang (Chenglong-MS) wants to merge 76 commits into
mainfrom
dev
Open

Chenglong Wang (Chenglong-MS) wants to merge 76 commits into
mainfrom
dev

Conversation

@Chenglong-MS

Copy link
Copy Markdown
Collaborator

This pull request introduces several improvements and refinements to the Data Formulator agent and context handling, focusing on clearer terminology, enhanced connector and skill management, and dependency updates. The most important changes are summarized below.

Agent and Context Terminology Improvements:

  • All references to "primary table(s)" and "other available tables" have been replaced with "primary analysis inputs" and "other analysis inputs" throughout agent summaries, context, and system prompts for greater clarity and consistency. The [AVAILABLE TABLES] section is now [ANALYSIS INPUT TABLES]. [1] [2] [3] [4] [5] [6]

Connector Availability and Summary Enhancements:

  • The connector summary block now displays all currently loadable connectors, including those not yet cached, and distinguishes between connected/disconnected sources and catalog availability. Disconnected sources are hidden from the agent, and error handling/logging is improved. [1] [2] [3]

Skill Preloading and System Prompt Improvements:

  • The agent now preloads the "data-loading" skill when no analysis input tables are present, exposing its tools and actions immediately. System prompts include all preloaded skills with clear banners. Skill state is properly rehydrated from both loaded and preloaded markers in the trajectory. [1] [2] [3] [4] [5]

Dependency Updates:

  • Several frontend dependencies have been updated for security and compatibility, including vite, dompurify, echarts, js-yaml, and new resolutions for postcss, esbuild, and tmp. [1] [2] [3]

Documentation Cleanup:

  • The obsolete docs/desktop-portable.md file has been removed.
  • The old model evaluation plan in loops/model-evaluation/plan.md has been deleted.

Comment thread py-src/data_formulator/analyst/agent.py Fixed
Add orcarouter to the built-in providers so Data Formulator users can
connect the OrcaRouter AI gateway from the model picker and via
ORCAROUTER_* environment variables, mirroring the existing ollama/openai
wiring. The client routes through LiteLLM's openai provider against the
OrcaRouter base URL, preserving the orcarouter/ model namespace.
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
…ovider

feat: add OrcaRouter as a built-in model provider
Adds full Hindi translations for all 10 i18n domain files, following
the translation guide's rules (chart type identifiers, encoding channel
keys, and other computation-bound values stay in English). Registers
hi in the locale index/i18next resources and language switcher labels,
and documents AVAILABLE_LANGUAGES=en,zh,hi as opt-in in .env.template.

Relates to #337.
folderPathPlaceholder illustrates filesystem path syntax; translating
the segment words broke the example and was inconsistent with how
every other path/URL placeholder in the app is left untranslated.
Add Hindi (hi) locale for frontend i18n
Comment thread py-src/data_formulator/datalake/workspace.py Fixed
Comment on lines +823 to +826
'<!doctype html><html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width">'
'<title>OpenRouter</title></head><body><p>' + message + '</p>'
'<a href="/">Return to Data Formulator</a>'
f'<script nonce="{nonce}">{script}</script></body></html>',
creator_id = key_info.get("creator_user_id")
connection = {
"creator_user_id": creator_id if isinstance(creator_id, str) else None,
"settings_url": "https://openrouter.ai/keys/" + hashlib.sha256(config["api_key"].encode()).hexdigest(),
The wrapper embeds user-supplied code, so on a system whose preferred
encoding is not UTF-8 (cp1252 on Windows) writing it raised
UnicodeEncodeError for any non-ASCII character, and silently transcoded
the rest -- the container reads run.py as UTF-8 either way.

Matches the explicit encoding already used elsewhere in the package,
e.g. reasoning_log.py and analyst/skills/__init__.py.
The constant table listed DEFAULT_ROW_LIMIT_EPHEMERAL in dfSlice.tsx.
That identifier does not exist anywhere in the repository, and there is
no separate ephemeral row limit in the frontend.
…hemeral-row

docs(dev-guides): drop the nonexistent DEFAULT_ROW_LIMIT_EPHEMERAL row
fix(sandbox): write the docker wrapper script as utf-8
Comment thread py-src/data_formulator/workflows/instances.py Fixed
Translate all 10 frontend locale domains into Bahasa Indonesia (1860 keys, full parity with en). Register the bundle in locales/index.ts and i18n/index.ts, add the 'Bahasa Indonesia' language switcher label, and cover key parity with a unit test.
Add 36 missing Hindi keys and 5 missing Simplified Chinese keys introduced by newer features (semantic model sidebar, connector groups, chart quick actions, virtual sources). All four UI locales now share the same 1860-key set.
Add Indonesian (id) locale + backfill missing hi/zh keys to full parity
Translate all 10 frontend locale domains into Japanese in polite desu/masu form (1860 keys, full parity with en). Register the bundle in locales/index.ts and i18n/index.ts, and cover key parity with a unit test. The language switcher label and backend language support for Japanese already existed, so no changes were needed there.
paths = filesystem['allowWrite']
if (not isinstance(paths, list) or len(paths) > 64
or any(not isinstance(path, str) or not path or len(path) > 2000 or '\0' in path
or not Path(path).expanduser().is_absolute() for path in paths)):
home = Path.home().resolve()
protected = configuration_path().parent.resolve()
for path in paths:
resolved = Path(path).expanduser().resolve()
actor, saved['revision'], ','.join(changed_sections))
return json_ok(snapshot())
except ConfigurationConflict as exc:
return {'status': 'error', 'error': {'code': 'INVALID_REQUEST', 'message': str(exc), 'retry': False}}, 409
Calling ensure_table_keys() at the base call sites does not cover loaders that
override the tree builders. SupersetLoader overrides both list_tables_tree() and
search_catalog() and constructs its nodes by hand, and the clickhouse, mysql and
postgresql search_catalog implementations bypass the base method too.

Move the backfill into _tables_to_catalog_tree(), which every tree builder apart
from the Superset overrides funnels through, and apply the key on the Superset
side separately. That path keys on the dataset uuid in list_tables(), so the
uuid is now carried through ls(), _build_dashboard_group_metadata() and
_dataset_meta() to keep both paths producing the same key.
Comment thread src/app/sessionThunks.ts
* Open a saved session, replacing the current one. Resolves true when it opened;
* failures are reported to the user and resolve false.
*/
export const openSession = (sessionId: string, displayName?: string, options: { saveCurrent?: boolean } = {}) =>
fix(data-loader): backfill table_key in tree and search responses
Add an html_app analyst skill whose write_html_app action saves an
agent-managed .html workspace file with a manifest of the tables it may
read. Revisions go through edit_file.

The canvas renders .html files in a double iframe: a same-origin
container with frame-src 'none' (blocks navigation exfiltration) holding
an opaque-origin sandbox="allow-scripts" app frame with a CSP that
blocks network access. An inlined DF runtime provides DF.query/DF.table
(validated bridge to sample-table, manifest tables only), DF.chart
(Vega-Lite), and theme variables.

Remove the unused Chartifact dialog and its i18n keys.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants