Point it at your own LLM and describe what you want — it builds the agent workflow, and runs it in your terminal.
A TUI-native designer and runner for agentic workflows: flow-chart pipelines with loops, branches, parallel fan-out, human approval gates, and real resilience — defined in plain YAML, generated from natural language, and driven live from the terminal, an HTTP API, or code.
Local-first · no SDK lock-in · your models, your machine, your data.
$ graphx providers --add http://localhost:8000 # point at any OpenAI-compatible server
✔ added openai_local_8000 → qwen3.6-27b
$ graphx generate "fetch a GitHub repo's JSON, summarize the description
with the local model, then save it to a file"
✔ valid workflow generated (oneshot, ~2.8k tokens)
$ graphx run repo_summary.yaml
✔ fetch → ✔ summarize → ✔ save # ran on your own model, start to finishThat's the whole loop: connect a model by URL, describe a pipeline in English, get a working workflow, run it — all offline on your hardware.
Most "agent builders" are either web-based node canvases you can't script, or code-only engines with no UI. graphx is the missing middle, in the terminal:
- 🧠 Build from natural language — a description becomes a real, valid workflow, self-correcting against a schema validator. Two engines (reliable one-shot + agentic tool-driven).
- 🔌 Point at any endpoint — type a URL for vLLM / llama.cpp / Ollama / LM Studio / a gateway; graphx probes it and auto-discovers the model. It also scans your LAN on startup.
- ⚙️ A real engine, not a toy — Pregel-style supersteps, cyclic graphs with loops, shared state with reducers, per-step SQLite checkpointing. Kill a run mid-flight;
resumecontinues from the last step. - 🛡️ Resilience per node — retries with backoff, model fallback chains, validation re-asks, timeouts, token/cost/deadline budgets,
on_erroredges, dead-letters. - 🔐 Credentials done right —
secret://NAMEresolves only at the point of use and never leaks into files, checkpoints, logs, SSE, or the screen. - 🧩 Batteries included — 12 credential-wired connectors (Slack, Discord, Telegram, SendGrid, SMTP, Gmail, GitHub, GitLab, Postgres, S3, webhooks) and OpenAPI auto-scaffolding for anything else.
- ⏰ Runs itself —
triggers:fire a workflow on a cron schedule, an interval, or an inbound webhook; theservedaemon or a systemd timer keeps it going. - 📦 Portable —
graphx exportturns a workflow into a self-contained folder (bundled wheel + Dockerfile) that runs on any machine, no graphx install required. - 📄 YAML is the source of truth — layout-free, diffable, hand-editable. The TUI renders it; it never owns it.
python3 -m venv venv
./venv/bin/pip install -e ".[tui,mcp,server]" # extras: keyring, postgres, s3
./venv/bin/graphx tui # opens a workflow here, or scaffolds a demoPython 3.12+. Nothing leaves your machine unless a node you add reaches out.
version: 1
name: research_review
providers:
local: { base_url: "http://localhost:8000/v1", protocol: openai }
state:
topic: { type: str }
draft: { type: str, default: "" }
entry: [fetch, search] # parallel fan-out
nodes:
- { id: fetch, type: api, url: "https://api/trends?q=<state.topic>",
headers: { Authorization: "Bearer secret://trends_key" } }
- { id: search, type: mcp, server: docs, tool: search, args: { q: "<state.topic>" } }
- { id: gather, type: merge, success_threshold: 1 } # barrier join
- id: write
type: agent
model: "local/qwen3.6-27b"
prompt: "Topic <state.topic>. Trends <fetch.json>. Docs <search.text>. Write a brief."
output_schema: { draft: str } # validated JSON
fallbacks: [ { static: { draft: "unavailable" } } ] # degrade gracefully
updates: { draft: "<self.draft>" }
- id: review
type: human # pause for approval
prompt: "Ship this?"
choices: [approve, revise]
edges:
- { from: fetch, to: gather }
- { from: search, to: gather }
- { from: gather, to: write }
- { from: write, to: review }
- { from: review, to: end, when: "review.choice == 'approve'" }
- { from: review, to: write, when: "review.choice == 'revise'" } # loop backLoops, parallelism, secrets, LLM fallback, and a human gate — in ~25 lines.
graphx providers --add http://192.168.1.50:8000 # probe a URL → discovers the model
graphx generate "poll an RSS feed hourly, summarize new items, post to Slack"
graphx edit myflow.yaml "add a human approval gate before the Slack post"- Grounded, not guessing. The model is handed a compact catalog of every node type + connector + the reference syntax, and its output is checked by the same validator
graphx runuses. Invalid → the exact errors are fed back for a bounded repair loop. The result is always either valid or clearly flagged for a one-line fix. - Two engines.
oneshot(default) emits the whole workflow and repairs it — reliable even on small local models.--agenticdrives builder tools step by step for capable models. - In the TUI, press
g, type your idea, and the generated graph renders in place, ready to run or tweak.
It's an editable first draft, not an oracle — quality scales with your model, and the validator keeps it honest.
graphx new mybot -t review # → mybot.yaml + mybot.eval.yaml, model auto-filledEleven starters in three tiers (n in the TUI). LLM templates wire themselves to a discovered local/LAN inference server and scaffold a paired NAME.eval.yaml, so every workflow starts life with an eval harness.
| template | shape | |
|---|---|---|
| starters | blank agent approval pipeline |
one concept each: shell · LLM w/ tool + schema + fallback · human gate · api loop |
| patterns | review fanout triage |
evaluator-optimizer (fresh-context critic) · parallel map + synthesize · LLM routing w/ escalation |
| real world | inbox digest issueops watchdog |
local-model email triage · cron→fetch→LLM→Slack · GitHub webhook→draft→gate→comment · LLM-free interval health check |
graphx tui examples/hello.yamlA lazygit-style shell: the graph on the left (live node status as it runs), node detail and streaming logs on the right.
| key | action | key | action | |
|---|---|---|---|---|
r |
run live | g |
generate from a description | |
n |
new from template | a |
add node | |
o |
node from an OpenAPI spec | i |
add a service connector | |
c |
connect nodes | k |
manage secrets | |
e |
edit YAML in $EDITOR |
q |
quit |
Every edit writes straight back to the YAML, and the file is watched — edit in vim in another pane and the graph redraws itself.
| group | types |
|---|---|
| work | agent (LLM + tools + validated JSON output + dynamic handoffs) · api (HTTP + $.json.path extraction) · mcp (MCP tool call) · function (Python) · shell (subprocess / CLI agents) |
| flow | condition (branch + loop) · router (LLM picks the path) · critic (independent review → loop on evidence) · map (fan-out over a collection) · merge (barrier join with a success threshold) · subworkflow |
| control | human (approval gate — interrupts, resumable) · wait |
References: <node.field>, <state.key>, <item.x>, secret://NAME. Edges carry when: expressions. LLM providers are per-workflow — anything OpenAI-compatible plus native Anthropic, no SDKs.
Drop-in, credential-wired nodes for popular services. Each declares the secret:// it needs, so you're prompted for it automatically.
graphx add slack notify.yaml message="deploy done"
graphx add github_issue bug.yaml owner=me repo=app title="broken"| category | connectors |
|---|---|
| messaging | slack · discord · telegram · webhook |
sendgrid · smtp · gmail (via MCP) |
|
| dev | github_issue · github_comment · gitlab_issue |
| data | postgres_query · s3_put |
Anything with an OpenAPI spec: graphx scaffold-api wf.yaml http://service builds the api node for you — path params, request body, and response-field extraction included.
Reference secrets as secret://NAME. They resolve only at the point of use — the outbound request, subprocess env, or MCP server — and are never written into the workflow file, checkpoints, the event log, SSE, or the TUI (a redaction net masks any value that slips into output).
graphx secret set slack_webhook_url # hidden prompt (or --value / --stdin)
graphx secret list # names only, never valuesStored 0600 in ~/.graphx/secrets.json (or the OS keyring via the [keyring] extra), with env-var fallback. graphx run refuses to start on a missing secret and tells you exactly how to set it; the TUI prompts inline.
- Retries — per-node exponential backoff + jitter, transient-only (429/5xx/timeout), honoring
Retry-After. - Fallbacks — ordered model chains ending in an optional static degraded output.
- Guards — four independent stops: max steps, token budget, cost budget, wall-clock deadline.
- Checkpoint & resume — full state snapshot to SQLite every superstep.
kill -9a run;graphx resume <thread>continues exactly where it stopped. - Human-in-the-loop — a
humannode interrupts and persists; resume from the CLI, TUI, or API.
Two building blocks for multi-agent workflows, both pure flow/state/logic — no engine changes, no hidden control channels:
-
critic— self-review that can't rubber-stamp. An independent judge scored against explicit criteria in a fresh context: it only ever sees the artifact + the criteria, never the producing agent's conversation, so the same model can't quietly grade its own work. Use a different model (or a deterministichandler:) for a truly independent review. The verdict routes like any other decision — loop back to revise onfail, publish onpass— bounded by the producer'smax_iterations:- { from: review, to: publish, when: "review.verdict == 'pass'" } - { from: review, to: write, when: "review.score < 0.8" }
-
Dynamic handoffs — agent-to-agent transfer at runtime. Give an
agentahandoffs:list and it gets one synthetic tool per target; when it calls one, control and the full conversation transfer to the specialist, which reads the context from<state.handoff.reason>/<state.handoff.messages>. Unlikerouter(routes with no context), a handoff carries the working memory over — and stays bounded by the same step/iteration guards.
See examples/review_loop.yaml (write → critic → loop → publish) and examples/handoff.yaml (triage → specialist).
A read-only observer over the same events + checkpoints every run already writes — it never touches the execution path, state, or routing.
graphx runs # every past run: status · cost · latency · steps
graphx run-show <thread> # per-node tokens, cost, latency, retries, interventions
graphx eval flow.yaml cases.yaml # replay a dataset, assert outcomes (exit 1 on any fail)
graphx eval-compare cases.yaml --a v1.yaml --b v2.yaml # diff two versions: outcomes, metrics, traceEval datasets combine a deterministic backbone with an optional LLM judge (reusing the critic):
cases:
- name: named
input: { name: graphx }
expect:
status: finished
assert: ["shouts == ['HELLO,', 'GRAPHX!']"]
budget: { tokens: 100 }Golden-trace normalization strips volatile noise (timings, tokens, values) so two runs of the same graph diff cleanly — behavioral regressions show up, run-to-run jitter doesn't. Browse past runs in the TUI with b.
graphx run flow.yaml --input topic="graph engines" # stream events in the terminal
graphx resume <thread> --answer approve # answer a human gate
graphx serve examples --port 8420 # REST + SSE server
graphx tui flow.yaml --attach http://localhost:8420 --thread <id> # follow a remote runThe CLI, the TUI, and the HTTP API all consume the same RunEvent stream. The API: GET /workflows · POST /runs · GET /runs/{thread} · POST /runs/{thread}/resume · POST /runs/{thread}/cancel · GET /runs/{thread}/events (SSE, honors Last-Event-ID).
Add a triggers: block and a workflow runs itself — the difference between "a pipeline I run" and "a pipeline that handled 40 emails while I slept."
triggers:
- { type: schedule, cron: "0 7 * * *" } # daily at 07:00
- { type: interval, every: 15m } # every 15 minutes
- { type: webhook, path: "orders", input_from: body } # POST /hooks/orders → runs, body = inputgraphx serve examples --port 8420 # daemon: fires schedules/intervals, receives webhooks
curl -X POST localhost:8420/hooks/orders -d '{"order_id": 4821}' # trigger on demandPrefer the OS to keep it alive? Emit a systemd user timer (or a crontab line) for a schedule-only workflow — no daemon required:
graphx schedule daily.yaml --cron "0 7 * * *" --install # writes + enables a systemd timer
graphx schedule daily.yaml --cron "0 7 * * *" --crontab # or print a crontab lineTurn a workflow into a self-contained, portable program. No graphx install on the target, no internet-to-a-private-repo, no secrets baked in.
graphx export myflow.yaml --docker
# → myflow_export/ · myflow.yaml + run.py + requirements.txt + .env.example + Dockerfile
# + a bundled graphx wheel
cd myflow_export
python3 -m venv venv && ./venv/bin/pip install -r requirements.txt
cp .env.example .env # fill in any secrets (as env vars)
./venv/bin/python run.py # …or: docker build -t myflow . && docker run --env-file .env myflowgraphx installs from the bundled wheel; its dependencies come from PyPI; credentials come from your .env (nothing sensitive is ever written into the export).
| file | what it shows |
|---|---|
examples/hello.yaml |
LLM-free tour — parallel fan-out, merge, map, a counted loop |
examples/gpu_report.yaml |
Real host health report: probe GPUs/disk/inference server → local LLM writes it → human gate → save |
examples/email_triage.yaml |
Classify inbox mail and draft replies on your own model, behind a human gate |
examples/agent_demo.yaml |
Minimal live-LLM demo — agent writes, router branches on the result |
examples/approval.yaml |
Human-in-the-loop gate: draft → approve → publish |
examples/review_loop.yaml |
Self-review: write → independent critic → loop until it passes → publish |
examples/handoff.yaml |
Dynamic handoff — triage agent transfers control + context to a specialist |
examples/hello.eval.yaml |
An eval dataset for hello.yaml — deterministic assertions + budget |
examples/scheduled_report.yaml |
Triggers demo — runs daily on a cron and on a webhook |
generate · edit · new · connectors · add · providers [--scan|--add <url>] · secret set/list/rm · scaffold-api · schedule · export · validate · run · resume · events · history · runs · run-show · eval · eval-compare · tui · serve
v0.9.0 — engine, all node types (incl. critic self-review + agent handoffs), eval/ops layer (replay · assert · compare · per-node cost/latency), natural-language builder (both engines) + point-at-any-endpoint, triggers/scheduling (cron · interval · webhook + systemd/crontab), export-to-portable-program, 12 connectors, secrets, discovery, 11 templates in three tiers with paired eval scaffolds, OpenAPI scaffolding, TUI (designer + runner + editor + run browser), HTTP API + SSE. 298 tests, ruff-clean, Python 3.12+.
Roadmap: richer TUI edge routing, remote run control, provider pricing tables, PyPI publish, more connectors.
Apache-2.0. Free for any use, including commercial — just keep the copyright and NOTICE, and mark any files you change. Patent grant and trademark protection included.