Repository navigation
feat(agents): run queue, one theme per run and the theme toggle - #12112
Merged
Merged
Conversation
Owner decision of 2026-10-09, after spikes S and S2 showed one 4 GiB
instance serves one sandbox at a time safely:
- A two-lane FIFO run queue (agents/anyplot/run_queue.py) in front of
every /messages turn: AGENT_RUN_CONCURRENCY (1) runs in flight,
AGENT_RUNS_PER_MINUTE (1) starts in a sliding 60 s window, at most
AGENT_QUEUE_MAX_WAIT_S (600) of waiting and rate x wait / 60 (10)
entries. A full queue is 503 capacity before the stream, an expired
entry error{code:"capacity"} inside it. The stream sends ready, then
status{step:"queued", position, waiting} at once, on change and every
15 s, then the run; the deadline and its abort start with the run. The
run registry covers queued entries (409 run_active), cancel and purge
withdraw a waiting entry, a gone client withdraws its own, and the stale
sweep uses the queue's clock and maximum wait. /v1/status reports
waiting and in_flight. A premium lane exists but nothing sets it.
- One theme per run: PipelineArgs.theme (light by default), the render
job, the host gates (now judging the job's themes, not THEMES), the
reviewer's images, the artifacts, the padded line and the canvas line
all name that theme. The root's prompt and tool description say when
to pass dark, and its session block names the latest version's theme.
- POST /v1/sessions/{sid}/versions/{version}/render {theme}: renders
another theme of a finished version from its stored run form through
the render backend and gates R1-R3 with the padding fallback, with no
adapter, reviewer or queue; answers ok, needs_attention
(canvas_padded) or failed (render, error) with the version's artifacts,
409 while the session has a run, 404 for an unknown or swept version.
- AGENT_RENDER_CONCURRENCY defaults to 1.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
- POST /debug/agent/sessions/{sid}/versions/{version}/render {theme}
mirrors the agents service's theme toggle: version 0-999 (0 is the
latest), theme light or dark, an unknown field is 422, and upstream
errors keep the agent error shape ({detail, ref}).
- The anyplot/1 relay passes position and waiting on status events, so
status{step:"queued"} reaches the browser.
- Queued time no longer spends the turn budget: every queued status, and
the first event after the wait, restarts AGENT_REQUEST_TIMEOUT_S, as
the agents service starts its own deadline only when the run leaves
the queue. A queue that falls silent still ends with error upstream.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
- Design doc: the bounds paragraph becomes a bounds table with the queue, rate, maximum wait, queue length, per-user and render rows; the request flow, the /v1 table, the render section (one theme, serial renders, the padded line naming the theme), the SSE protocol (queued status, the BFF's budget restart), the Serving flags (--concurrency=20 and --timeout=900 for the queue's open streams), the risks and two open decisions (the anyplot-api 600 s timeout against a 780 s queued turn, and whether "Create plot" carries the site theme). - agents/README.md: the new modules and settings, the adk web bypass of the queue, and AGENT_RUNS_PER_MINUTE=60 for local iteration. - docs/reference/api.md: the toggle route and its answers, the queued status, 503 capacity, waiting and in_flight on /status. - changelog.d/agents-run-queue.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…start before expiry Review fixes for the run queue branch: - SerialRenderer (render/serial.py) wraps every backend in Services.backend: one theme per slot under AGENT_RENDER_CONCURRENCY, a freed slot goes to a waiting pipeline render before a waiting toggle, and a bounded wait raises RenderBusy. The local and sandbox backends drop their own semaphore. - The theme toggle registers in the run registry, so it and a turn refuse each other with 409 run_active and a user has one run or toggle in flight; it waits for a slot at most request deadline minus render timeout (503 capacity after that); a theme that failed the host gates is recorded and rendered at most twice, while an `error` stays unrecorded. - The queue starts an entry whose turn comes at the moment its wait runs out, never one a late pump finds overdue, and documents that the owner's capacity formula assumes runs within the 60 s window. - A user over the daily budget gets the budget refusal before the queue, and a queued turn counts as use of its dataset for the idle sweep. - Reviewer defects name the rendered theme; plot.py's run line names it too. - Tests: serial renderer, toggle in flight, busy slot, budget precheck, dataset touch, the deadline armed only when the run starts, start at the maximum wait, long runs past the capacity promise. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
AGENT_TURN_MAX_S (590 s) bounds a chat turn, queue wait included, so Cloud
Run's 600 s timeout never cuts a stream without error and done. A turn still
queued when its run could no longer finish inside the cap ends with
error{capacity}, and closing the upstream takes it out of the queue before it
spends a token. A queued status restarts the budget with the queue's 15 s
heartbeat on top, because the run may start that long before its first event.
The status route's docstring lists every upstream field.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…s promise The design doc's bounds table, render section, risks and open decisions, the agents README, the API reference and the changelog fragment describe the render slots in front of every backend, the toggle's registry entry and slot wait, the budget check before the queue, the BFF turn cap and what it means for the 600 s maximum wait, and the capacity formula's assumption about run length. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Contributor
There was a problem hiding this comment.
🟡 Changes recommended
A refinement can silently reset a dark plot to light when the pipeline call omits its theme.
3 open findings
What changed in this PR
Adds bounded agent-run scheduling, serial rendering, and single-theme runs with on-demand theme rendering.
Changes:
- Adds FIFO run queue, rate limits, status events, and turn deadlines.
- Serializes rendering and renders one theme per pipeline run.
- Adds theme-toggle APIs, persistence, documentation, and tests.
| File | Description |
|---|---|
.env.example |
Documents the turn cap. |
agents/README.md |
Documents queue and theme behavior. |
agents/anyplot/agent.py |
Exposes the latest theme to the agent. |
agents/anyplot/code/export.py |
Adds theme-aware run instructions. |
agents/anyplot/pipeline.py |
Runs and reviews one theme. |
agents/anyplot/prompts/reviewer.md |
Adapts review instructions for one render. |
agents/anyplot/prompts/root.md |
Guides theme selection and toggling. |
agents/anyplot/render/__init__.py |
Moves concurrency to the wrapper. |
agents/anyplot/render/backends/local.py |
Removes backend-local concurrency. |
agents/anyplot/render/backends/sandbox.py |
Removes backend-local concurrency. |
agents/anyplot/render/contract.py |
Centralizes the theme contract. |
agents/anyplot/render/gates.py |
Evaluates only requested themes. |
agents/anyplot/render/serial.py |
Adds prioritized render slots. |
agents/anyplot/render/store.py |
Supports adding theme images. |
agents/anyplot/run_queue.py |
Implements queued run scheduling. |
agents/anyplot/schemas.py |
Adds theme and artifact contracts. |
agents/anyplot/services.py |
Stores per-theme outcomes and wraps backends. |
agents/anyplot/settings.py |
Adds queue and concurrency settings. |
agents/anyplot/sub_agents/reviewer.py |
Attaches only rendered-theme images. |
agents/anyplot/theme_render.py |
Implements model-free theme rendering. |
agents/anyplot/tools/session.py |
Extends pipeline tool arguments. |
agents/main.py |
Integrates queue and toggle routes. |
agents/stream.py |
Adds queued SSE statuses. |
api/routers/agent.py |
Relays queue events and toggle requests. |
changelog.d/agents-run-queue.md |
Records the feature. |
core/config.py |
Adds the BFF turn limit. |
docs/concepts/agent-network.md |
Updates the architecture design. |
docs/reference/api.md |
Documents new API behavior. |
tests/unit/agents/code/test_export.py |
Tests theme-aware exports. |
tests/unit/agents/runtime/test_pipeline.py |
Tests single-theme results. |
tests/unit/agents/runtime/test_registry.py |
Verifies the expanded tool schema. |
tests/unit/agents/runtime/test_render.py |
Tests theme-scoped gates and storage. |
tests/unit/agents/runtime/test_run_queue.py |
Tests queue semantics. |
tests/unit/agents/runtime/test_run_queue_flow.py |
Tests queue integration. |
tests/unit/agents/runtime/test_serial_render.py |
Tests render serialization. |
tests/unit/agents/runtime/test_service_flow.py |
Tests theme runs and toggles. |
tests/unit/agents/runtime/test_stream.py |
Tests queued stream events. |
tests/unit/agents/test_settings.py |
Tests queue configuration. |
tests/unit/api/test_agent_router.py |
Tests BFF queue and toggle behavior. |
🧠 Review effort: Balanced
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
… change PipelineArgs.theme is optional: a refinement on base='previous' without a theme renders the latest version's theme instead of falling back to light; a new plot stays light. The tool description and the root prompt say so, the pipeline docstring names the theme field, and the API reference heading no longer reads as a theme statement. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
This was referenced Oct 9, 2026
MarkusNeusinger
added a commit
that referenced
this pull request
Oct 10, 2026
## Summary - anyplot-api's Cloud Run request timeout rises from 600 to 900 seconds (`api/cloudbuild.yaml`) and the BFF's turn cap `AGENT_TURN_MAX_S` from 590 to 890 seconds, so a chat turn at the back of a full run queue (600 seconds of waiting plus the 180-second run) ends with its own `error` and `done` instead of being cut at about 385 seconds. - The design document records the owner's decisions of 2026-10-10 on the three open points from #12112: raise the timeout (this PR), keep the queue's capacity formula, keep one run or theme toggle in flight per user. - Docs: `docs/reference/api.md` and `.env.example` follow the new values. ## Plan `docs/concepts/agent-network.md` (bounds table, SSE section, open decisions, risks). ## Test plan - [x] `uv run ruff check .`, `uv run ruff format --check .`, mypy (95 files), `tools.changelog check`, `uv lock --check` pass - [x] `tests/unit/api/test_agent_router.py` and `tests/unit/core` pass (the turn-cap test overrides the setting, so the default change touches no assertion) - [ ] After merge: the deploy-api Cloud Build applies `--timeout=900` to the `anyplot-api` service (watch the build and `gcloud run services describe anyplot-api --format='value(spec.template.spec.timeoutSeconds)'`) --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
MarkusNeusinger
added a commit
that referenced
this pull request
Oct 10, 2026
…12115) ## Summary - Adds `anyplot-renderer` (`agents/renderer/`): a small FastAPI service with `POST /render`, `POST /render/{job_id}/cancel` and `GET /status` behind Cloud Run IAM and the ID-token claims check, that runs each theme of a job in its own Cloud Run sandbox (`sandbox do`, never `--write` or egress, explicit `PATH`), one at a time, with the limits spikes S and S2 measured: run directories on a 512 MiB in-memory volume the service requires on Cloud Run (`503 volume_missing` without it), a watchdog that kills a run past 64 MiB, 10,000 files or a 512 MiB `MemAvailable` floor, a kill that counts only once the launcher has exited, one retry when the launcher fails before the harness starts, bounded output, a cleaned probe, rlimits, a stop on client disconnect or cancel, JSON `500 internal` without the message, and a self-restart when a launcher outlives its kill. It only runs code; the host gates stay in the agents service. The image installs the plotting libraries without ADK; `agents/renderer/cloudbuild.yaml` builds, deploys a `candidate`, smokes and promotes it like the API, and creates the service on the first build only when `describe` reports it missing. - Adds the `remote` render backend, now the default (`AGENT_RENDERER=remote`, `AGENT_RENDER_URL`), so a local `adk web` and the deployed agents service render in the same sandboxes: an ID token from the metadata server when deployed, or from `AGENT_RENDER_TOKEN` or the developer's Application Default Credentials in development; retries over a cold start (1, 3, 8, 15 s within 30 s) for connection errors and front-end 429/5xx only, never for the renderer's own answers; a cancel call for abandoned renders; limit and measured value in the repair feedback. - Hardens the caller check in both services: `X-Serverless-Authorization` wins whenever it is present, because Cloud Run then checks only that header and passes `Authorization` through unverified. The harness prints a start and an end line and sets soft and hard limits alike. CI gains a `renderer-image` job (`ci-image.yml`) that builds the image and renders a seaborn plot through the harness with no network. ## Plan Design: `docs/concepts/agent-network.md` ("Render", "Serving and infrastructure"; status line and service table updated). The renderer split for phase 1 was decided on 2026-10-10 (one renderer for the deployed agents and for `adk web`). Follows #12111 and #12112. ## Test plan - [x] `uv run ruff check . && ruff format --check .`, `mypy api core agents` (101 files), `pytest tests/unit tests/integration` (7935 passed; executor tests run the real harness against a fake launcher: kill path, cancel, disconnect, stuck launcher, retry once and twice, byte, file and memory watchdog, bounded output, refused output files; service tests cover the caller check, the slot, low memory, the volume check and `/status`; the `remote` backend is tested against the in-process app and once end to end through the real executor and host gates), changelog check, `uv lock --check` - [x] Deployed to Cloud Run as a throwaway service (`anyplot-renderer-spike`, a role-less service account, Cloud Build `5b86cc7a` from this head with `_SMOKE_RENDER=false`, 3 min 51 s): the first-build path created the service; `gcloud beta run deploy --sandbox-launcher` and the repeated `--add-volume` flags are accepted by the Cloud Build gcloud image (587.0.0); an anonymous call gets 403; `/status` with the owner's gcloud token reports `sandbox: true`, `volume: true` (the in-memory volume reports its 512 MiB limit), 3,524 MiB available; a real matplotlib render answered in 2.9 s with a 108 KB PNG; spike X (122 cases) renders through it from a laptop with `--gcloud-token` - [x] Configured build on the throwaway with `_SMOKE_RENDER=true` (Cloud Build `a89a46b2` from `09f50b62d`, 3 min 47 s): the `--no-traffic` candidate path, the smoke as the service's own account (impersonated; a build cannot mint an ID token for itself, which builds `da6f53f5` and `a9d3ac7f` showed and the two fix commits address), `/status` with the volume in place, a real render in 3.2 s, promote. The `renderer-image` CI job ran and passed on every push. - [x] `/tmp` fill through the promoted revision (the review's open question): code writing 40 MiB files into the sandbox's private `/tmp` ends after 4.6 s with `OSError: [Errno 28] No space left on device`; the instance stays up (3,451 MiB available before, 3,328 after), the memory watchdog never fires. Decided in the design doc: 4 GiB with one sandbox at a time, no `RENDERER_BIND_TMP`, no 8 GiB instance. The throwaway service, its service account and its images are deleted. - [ ] Not verifiable before merge: the Application Default Credentials token path against Cloud Run IAM; `.github/workflows/` has no pre-merge loop - [ ] After merge nothing deploys: the real `anyplot-renderer` service account, the first build and the `deploy-renderer` trigger are owner tasks (design doc, task 14) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
MarkusNeusinger
added a commit
that referenced
this pull request
Oct 10, 2026
## Summary - Adds the admin-only "Use with my data" chat page at `/debug/agent?spec=&library=&language=` (`app/src/pages/AgentChatPage.tsx`): data panel with a 200 KB counter, 20-row preview, a column dropdown for every spec role (series families get one dropdown per member), the chat thread, a progress timeline that follows the `anyplot/1` stream, a queue-position counter in the chat's corner, `.stop()` through the cancel route, and library pills. Built only with `VITE_ENABLE_AGENT_CHAT=true` (off in `app/cloudbuild.yaml`, so production ships without the chunk) and opens only in local development or for a browser carrying the admin hint. - Adds the result card (`app/src/sections/agent-chat/ResultCard.tsx`) with the plot page's overlay actions (copy image, download PNG, open full size), a light/dark switch through the theme toggle route, the adapted `plot.py` with copy and download, `data.csv` download, the change list and residual notes, a refine composer and the quick-feedback slot; plus the `.adapt()` overlay button on the plot page, gated by build flag, admin hint and the BFF's eligibility route so public visitors never call the debug API. - Two small contract changes on the service side so the page can address versions and roles reliably: the `plot` event carries `version` (agents translator and BFF allowlist), and the dataset answer lists the spec's `roles` with a 422 that keeps the binding check's lines. Analytics events (enum properties only) are documented in `docs/reference/plausible.md`; a loopback mock of the BFF (`app/scripts/agent-bff-mock.mjs`) drives the whole flow without the API, the agents service or a model. ## Plan Design: `docs/concepts/agent-network.md` (sections "Frontend", "Result card", "Analytics"); this is the frontend PR of the phase-1 roadmap, built on the merged runtime core (#12111) and run queue (#12112). ## Test plan - [x] `cd app && yarn lint && yarn fm:check && yarn type-check && yarn test` (81 files, 750 tests) and `yarn build` with the flag on and off; the default build contains no agent client - [x] `uv run --extra dev ruff check . && ruff format --check .`, `mypy api core agents`, `pytest tests/unit tests/integration` (7816 passed), changelog check, `uv lock --check` - [x] Driven in the browser against the mock at 1280 px and 390 px, light and dark: parse, bindings incl. series families and a refused set, queued counter, progress timeline, result card actions (clipboard PNG, downloads named `plot.py`/`data.csv`, open full size, copy code), theme switch, repair round, refusal, error, stop then a new turn without a 409, ineligible pair; no horizontal scroll, 16 px gutter at 390 px - [ ] After merge: `deploy-app` and `deploy-api` Cloud Builds succeed; production stays unchanged (`_VITE_ENABLE_AGENT_CHAT` false, `AGENT_ENABLED` false) - [ ] Once the agents service is deployed: open `/debug/agent?spec=scatter-basic&library=matplotlib` as admin and run one "Create plot" end to end --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


Summary
AGENT_RUN_CONCURRENCY, 1), at mostAGENT_RUNS_PER_MINUTEstarts (1, sliding window), a maximum wait ofAGENT_QUEUE_MAX_WAIT_S(600 s) with the capacity that wait allows (10),503 capacityorerror{capacity}beyond it,status{step:"queued", position, waiting}on the stream,waitingandin_flightonGET /v1/status, queued time outside the request deadline, one queued or running turn per user, apremiumlane that nothing sets yet, and budget refusals before a queue place.SerialRendererwraps whichever backend renders (sandbox, local, fake,adk web),AGENT_RENDER_CONCURRENCYdefaults to 1, one theme per slot, pipeline renders ahead of theme toggles.plot.pyrun line follow, and the padded fallback names the theme.POST /v1/sessions/{sid}/versions/{version}/render {theme}and its BFF mirrorPOST /debug/agent/sessions/{sid}/versions/{version}/renderrender the other theme of a finished version from its stored code and data through the host gates; a toggle and a turn refuse each other, one toggle per user in flight, at most two tries per theme, a 120 s wait for the render slot.AGENT_TURN_MAX_S(590 s, below anyplot-api's 600 s request timeout) so every stream still ends witherroranddone.Plan
docs/concepts/agent-network.md(bounds table, SSE protocol, render section and open decisions updated in this PR). Follow-ups, not in this PR: the real sandbox backend and the renderer service, the agents image and Cloud Build, the frontend.Open decisions for the owner
--timeoutto about 900 s andAGENT_TURN_MAX_Sto about 890 s, or lowerAGENT_QUEUE_MAX_WAIT_Sto about 385 s.Test plan
uv run ruff check .anduv run ruff format --check .cleanuv run --extra typecheck --extra agents mypy api core agents: no issues in 95 filesENVIRONMENT=test uv run pytest tests/unit tests/integration: 7,807 passed, 1 skippeduv run python -m tools.changelog check --base origin/mainanduv lock --checkpass/v1through the queue with two users on a fake clock: the second seesqueuedposition 1 before its run starts; full queue answers 503; 409 for a second turn or toggle of the same user; the 180 s timer is armed only when the run leaves the queue; a client gone while queued leaves the queuecapacityanddone