computer: Add ws:container and pick exec backends with one option - #172
mattzcarey wants to merge 9 commits into
Conversation
馃 Changeset detectedLatest commit: 0e2e978 The changes in this PR will be included in the next version bump. This PR includes changesets to release 4 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
|
Thanks for your interest in Cloudflare Computer. This repository does not accept unsolicited pull requests. Please use one of the accepted contribution paths instead:
If a maintainer asked you to open this pull request, they can add the |
| const result = await handle.result(); | ||
| return { | ||
| exitCode: result.exitCode, | ||
| stdout: truncate(result.stdout, maxOutputBytes), | ||
| stderr: truncate(result.stderr, maxOutputBytes), |
There was a problem hiding this comment.
馃敶 Large output exhausts host memory
A noisy command makes exec accumulate its full output before truncation. drainModuleResult retains every chunk, so maxOutputBytes cannot bound host memory.
Learn more
The module waits for a container command by calling the aggregate result API. drainModuleResult stores every stdout and stderr chunk and then allocates complete concatenated byte arrays and decoded strings. The module truncates only after those allocations, so a 64 KiB output cap does not limit host memory while a verbose command runs.
Example: A command writes hundreds of megabytes to stdout. The module eventually returns at most 64 KiB, but the Workspace retains and joins the full output first.
Recommended fix: Drain the execution handle's event stream directly in createContainerModule, retaining bounded UTF-8 prefixes and counting omitted bytes as chunks arrive. Read through stream closure so the post-command sync bracket still completes.
Was this helpful? React with 馃憤 or 馃憥 to provide feedback.
| const result = await handle.result(); | ||
| return { | ||
| exitCode: result.exitCode, | ||
| stdout: truncate(result.stdout, maxOutputBytes), | ||
| stderr: truncate(result.stderr, maxOutputBytes), | ||
| }; |
There was a problem hiding this comment.
馃敶 Container writes disappear from successful results
When the post-command pull fails or skips files, exec still returns a successful exit code. drainModuleResult exposes the failed sync, but callers cannot tell their Workspace files remain stale.
Learn more
A command backend syncs container writes back to the Workspace after the event stream drains. runPostPull represents pull failures as a pending sync outcome rather than throwing, and the runtime result includes both this outcome and any skipped files. The module drops those fields, allowing an exit code of zero to be returned when filesystem changes have not reached the Workspace.
Example: exec("echo updated > /workspace/output.txt") exits zero, but a failed pull leaves output.txt unchanged in the local Workspace. The isolate receives { exitCode: 0 } and can immediately read the old file.
Recommended fix: Check result.sync.status and result.skipped before treating the command as finished successfully. Surface pending or skipped synchronization to the isolate as an error or documented result fields so callers can detect incomplete file synchronization.
Was this helpful? React with 馃憤 or 馃憥 to provide feedback.
| usedBytes += charBytes; | ||
| endOffset += char.length; | ||
| } | ||
| return `${value.slice(0, endOffset)}\n\n[truncated, ${totalBytes - usedBytes} more bytes]`; |
There was a problem hiding this comment.
馃煛 Truncated output exceeds its configured limit
When output crosses maxOutputBytes, truncate appends a marker beyond that limit. Near the bridge response cap, the resulting call rejects instead of returning output.
Learn more
The truncation helper fills its entire byte budget with command output and then adds a truncation marker. The trusted-call bridge limits the size of the complete encoded response, including both streams, so output near that limit rejects rather than delivering a truncated response. The configuration documentation describes maxOutputBytes as an output cap, but the returned string can exceed it even at small values.
Example: With maxOutputBytes: 5, stdout 123456 becomes 12345\n\n[truncated, 1 more bytes], exceeding five bytes.
Recommended fix: Reserve marker space within the per-stream limit and account for the encoded response overhead and both streams when validating configuration or preparing the response.
Was this helpful? React with 馃憤 or 馃憥 to provide feedback.
commit: |
|
@mattzcarey can you stack this on top of the trustedModules -> modules change. I can merge those now, would like to properly review this one. |
82ad877 to
82eacc5
Compare
| const info = known?.find((backend) => backend.id === id); | ||
| const callable = info?.callable === true; | ||
| const own = info?.description; |
There was a problem hiding this comment.
馃煛 Callable custom backends lose structured input
For custom workspaces exposing isCallable but not backends, createExecTool treats callable backends as shells. Explicit backend selection still works, but the tool omits input and describes module source as shell commands.
Learn more
The exec tool accepts a workspace-like runtime, not just a Workspace client. Its backends method is optional, so callers can provide explicit backend selections without implementing backend discovery. A runtime that implements isCallable(id) but lacks backends() now loses callable status when the tool builds its schema. The tool therefore omits structured input and labels its commands as shell commands even though the backend accepts module source.
Example: A custom runtime with isCallable: (id) => id === "js" and exec(...), passed with backends: { js: {} }, produces an exec tool without an input field. The same runtime accepts structured input directly.
Recommended fix: Preserve isCallable and describe as fallback methods when backends() is unavailable, or require backend metadata explicitly in the tool options and provide a migration path for workspace-like runtimes.
Was this helpful? React with 馃憤 or 馃憥 to provide feedback.
An agent that wants both isolated JavaScript and a full Linux container has so far needed two exec backends, and the model had to pick one per command. This lets JavaScript be the only backend the model sees, with the container as a library it can call. createContainerModule() in @cloudflare/computer/modules/container is a host module factory. Installed as ws:container, its exec(command, options) runs through workspace.runtime.exec on the container backend, so the container shares the Workspace's files through the usual sync bracket. It returns the exit code and bounded output once the command finishes, kills the command when the execution is cancelled, caps the command's timeout at the host call deadline, and refuses to run on a read-only backend, since a container command can write to the Workspace whatever the isolate's access is. Its description tells the model how to call it, and reaches the exec tool through the JavaScript backend's own description.
* computer: Carry backend information through Workspace clients
A WorkspaceClient from getWorkspace() wrapped the runtime with exec,
getExec, killExec, and disposeExec only. The exec tool also asks the
runtime for backendIds, isCallable, and describe, so a client lost all
three: with exec omitted it offered no exec tool, and a callable
backend lost its input argument and its module list. Over RPC those
calls would be asynchronous, while the tool builds its schema
synchronously.
Backends are fixed when the Workspace is constructed, so the runtime
and its RPC stub now expose one backends() call, and a client takes
that snapshot when it is created and answers the three questions from
it, locally and remotely alike.
CloudflareContainerBackend now describes network access from its
egress setting instead of always claiming it, exec takes precedence
over the deprecated shell option so exec: {} always means no tool, and
the mcp README and think prompt stop implying a default backend.
* computer: Answer backend questions with one backends() call
The exec tool learned about backends through three runtime methods,
backendIds, isCallable, and describe, and every way of reaching a
Workspace had to forward all three. It now reads one list from
runtime.backends(), each entry carrying the id, whether the backend is
callable, and its description. backendIds and describe are gone, and
a Workspace client exposes the same backends() from its snapshot.
isCallable stays on the runtime for its own check before running.
* computer: Freeze the client backend snapshot and test input through it
A client returned its backend snapshot array itself, so a caller that
edited it changed what later tool sets saw. The snapshot is now frozen
once when the client is created.
The client tests only checked that the exec schema offered input. They
now send structured input through the exec tool on a local and a
remote client to a callable backend that echoes it, and check the
value comes back as the result. The mcp README no longer calls
worker-shell the default.
The exec tool's self-description for the container landed on the platform-scheduled backend, which is now LegacyContainerBackend. It moves to ContainerBackend, the backend containers should use, with its network line still following the egress mode. The legacy backend goes back to describing nothing and gets the exec tool's one-line default.
ws:container found a missing or wrong container backend only on its first exec. Its factory runs when the JavaScript backend connects and can read the Workspace's backends, so it now fails there: with no such backend, or with one that runs modules instead of shell commands. Its description also stopped promising network access. The model only sees the module's text in this setup, and whether the container can reach the network depends on the container backend's egress setting, which the module cannot know when it is built.
The ws:container docs, the exec tool docs, and the changesets named the old CloudflareContainerBackend. They now show ContainerBackend, the durable-object-scheduled backend containers should use.
Both examples ran their container through LegacyContainerBackend, the platform-scheduled backend, which no longer describes itself to the model. They now use ContainerBackend: the durable object schedules the container, the containers block names an image under images.app with scheduling_policy "durable_object", and the backend asks for standard-2 at launch, the size the old block requested.
1adc1e6 to
0970232
Compare
ws:container also rejected a backend marked callable, taking that as a sign it runs modules. callable means a backend takes structured input, and a shell backend may do both, so the check refused valid backends and let through a module backend that was not callable. It now checks only that the backend exists.
| // The factory runs when the JavaScript backend connects, so a | ||
| // missing container backend fails there, before any code runs, | ||
| // rather than on the first exec. | ||
| if (!host.runtime.backends().some((info) => info.id === backend)) { |
There was a problem hiding this comment.
馃煛 Shell commands run as JavaScript
When backend names a callable backend, createContainerModule accepts it and execOn runs shell commands as module source. Commands fail instead of running in a shell; targeting the same JavaScript backend can start nested executions.
Learn more
The Workspace lists both shell and module backends through runtime.backends(). A module backend is marked callable: true, while the JavaScript backend interprets execution source as a module. The module factory now accepts that backend by ID. An exec call then forwards its shell command through runtime.exec, which routes the text unchanged to the selected backend. This turns a shell call into a JavaScript evaluation; choosing the owning JavaScript backend also allows nested evaluations.
Example: Configure a JavaScript backend with createContainerModule({ backend: "worker-javascript" }). Calling exec("ls") evaluates ls as JavaScript instead of listing files; calling a valid JavaScript source that imports ws:container can start another evaluation.
Recommended fix: Restore the connect-time check for target.callable alongside the existence check, and keep a test with a callable backend ID. If other non-shell backend kinds are supported, validate the backend kind explicitly rather than relying only on callable.
Was this helpful? React with 馃憤 or 馃憥 to provide feedback.
There was a problem hiding this comment.
Fixed in 0e2e978. runtime.backends() now reports each backend's protocol (command or module), and ws:container refuses a module backend when it connects. A callable shell backend is still accepted. Tests cover both.
ws:container needs a backend that runs shell commands. It first judged that by callable, which describes structured input instead: a shell backend may be callable, and a module backend need not be. Without any check, pointing it at the JavaScript backend would run a shell command as JavaScript, or start a nested run. runtime.backends() now reports each backend's protocol, "command" or "module", and ws:container refuses a module backend when it connects. A callable shell backend is accepted.
Stacked on #171. #181 and #182 were reviewed separately and merged into this branch.
An agent that wants both isolated JavaScript and a full Linux container has needed two
execbackends, and the model had to choose one for every command. With this PR, JavaScript can be the only backend the model sees, and the container becomes a module JavaScript imports:sequenceDiagram participant Model participant JS as worker-javascript isolate participant Mod as ws:container (host) participant RT as workspace.runtime participant C as ContainerBackend Model->>JS: exec tool, module source JS->>Mod: exec("npm test", { cwd }) Mod->>RT: exec(command, { backend: "container-shell" }) RT->>C: push Workspace changes, run command C-->>RT: output, exit code, pull changes RT-->>Mod: result Mod-->>JS: { exitCode, stdout, stderr }ws:container.createContainerModule()is a host module factory (#170). It takeshost.runtimewhen the JavaScript backend connects, and fails there if the Workspace has no such backend, or that backend'sprotocolismodule, which would run shell text as JavaScript.runtime.backends()reportsprotocolfor this, becausecallableonly describes structured input.execis a host call. Output arrives when the command finishes, each stream is cut atmaxOutputBytes(64 KiB by default), and the timeout is capped at the host call deadline. Cancelling the JavaScript execution kills the command, arguments are parsed strictly, andexecrefuses to run on a read-only backend.The
exectool (#181).createAIToolsmoves to@cloudflare/computer/tools/ai-sdkand takesexec, the backends the model can use keyed by id. Leavingexecout offers every backend. With several backends the model must name one on every call; there is no default.WorkerJavaScriptBackenddescribes its modules, andWorkerShellBackendandContainerBackenddescribe themselves, so{}is enough per backend.shellstays as a deprecated alias.Workspace clients (#182). A
WorkspaceClientfromgetWorkspace()answersruntime.backends()from a frozen snapshot taken when the client is created, locally and over RPC. Tools built from thewithWorkspacemixin, or from another Worker, then get the sameexectool as tools built from a Workspace the Durable Object owns.On top of the new container runtime. #161 and #162 renamed the platform-scheduled backend to
LegacyContainerBackendand addedContainerBackend, where the Durable Object schedules the container. This branch targets the new one:ContainerBackend, with its network line followingegress: direct, through a gateway, or none.LegacyContainerBackendis unchanged frommain.ws:containerno longer promises network access. In this setup the model only reads the module's text, and network access depends on the container backend'segress.examples/thinkandexamples/mcpmove fromLegacyContainerBackendtoContainerBackend. Their containers blocks usescheduling_policy: "durable_object"withimages.app, and the backend asks forstandard-2at launch, the size the old blocks requested.thinknow needs nothing beyondcreateAITools({ workspace: this.workspace }).docs/09_tool_interface.md,docs/17_isolate_javascript.md, and the changesets nameContainerBackend.Unit tests cover the container module's argument parsing, timeout cap, kill on abort, read-only refusal, the connect-time backend check, and UTF-8 truncation, plus
ContainerBackend's description for each egress mode. The script runner suite runsimport { exec } from "ws:container"in a real Dynamic Worker against a realWorkspace. The client tests send structuredinputthrough local and remote clients.