Skip to content

[shim] Persist task state to clean up tasks after a restart - #4220

Merged
un-def merged 2 commits into
masterfrom
issue_4182_shim_persist_task_state
Aug 27, 2026
Merged

[shim] Persist task state to clean up tasks after a restart#4220
un-def merged 2 commits into
masterfrom
issue_4182_shim_persist_task_state

Conversation

@un-def

@un-def un-def commented Aug 26, 2026

Copy link
Copy Markdown
Collaborator

Since #4203, tasks outlive the shim process, but a restarted shim cannot finish them properly. Restoring a task from its container recovers the container ID, the GPUs, and the ports, but not the task config, so nothing tells the shim which volumes to unmount and which host SSH keys to remove. The reason a task was terminated for is lost as well, as it is not a property of the container.

Each task now has a task.json file in its runner dir, holding what the container cannot tell: the task config, the termination reason and message, and whether the task resources are already released. It is written atomically and flushed to the disk on every task state change, starting right after the runner dir is created, that is, before any resource is acquired -- releasing a resource that was never acquired is a no-op, while the opposite order would leak.

  • Restored tasks get their config back, therefore their volumes are unmounted and their host SSH keys are removed once they finish. The volumes of a container created by an earlier shim version are still recovered from the container mounts; its host SSH keys are not recoverable.
  • A task with a recorded termination reason is restored as terminated, so that the reason reported by the server, e.g., terminated_by_user, is not replaced with the exit code of the container it stopped.
  • GPUs are locked back only for the restored tasks whose resources are not released yet, so that a task cleaned up before the restart does not hold them forever.
  • Task dirs that have a state file but no container are cleaned up on start. Normally, such a dir is left behind when the shim stops running before the container is created, e.g., while pulling the image. Dirs without a state file are left intact, as there is no way to tell whether they belong to a task.
  • The registry credentials are not persisted: the image is already pulled by the time the state is read back.

Also fixes the runner dir rename fallback in remove(), which never worked: the trash name was built from the absolute path of the dir, so the rename target was a relative path under a .trash- dir that does not exist.

Part-of: #4182

un-def and others added 2 commits August 26, 2026 15:01
Since #4203, tasks outlive the shim process, but a restarted shim cannot finish
them properly. Restoring a task from its container recovers the container ID, the
GPUs, and the ports, but not the task config, so nothing tells the shim which
volumes to unmount and which host SSH keys to remove. The reason a task was
terminated for is lost as well, as it is not a property of the container.

Each task now has a `task.json` file in its runner dir, holding what the container
cannot tell: the task config, the termination reason and message, and whether the
task resources are already released. It is written atomically and flushed to the
disk on every task state change, starting right after the runner dir is created,
that is, before any resource is acquired -- releasing a resource that was never
acquired is a no-op, while the opposite order would leak.

* Restored tasks get their config back, therefore their volumes are unmounted and
  their host SSH keys are removed once they finish. The volumes of a container
  created by an earlier shim version are still recovered from the container
  mounts; its host SSH keys are not recoverable.
* A task with a recorded termination reason is restored as terminated, so that the
  reason reported by the server, e.g., `terminated_by_user`, is not replaced with
  the exit code of the container it stopped.
* GPUs are locked back only for the restored tasks whose resources are not
  released yet, so that a task cleaned up before the restart does not hold them
  forever.
* Task dirs that have a state file but no container are cleaned up on start.
  Normally, such a dir is left behind when the shim stops running before the
  container is created, e.g., while pulling the image. Dirs without a state file
  are left intact, as there is no way to tell whether they belong to a task.
* The registry credentials are not persisted: the image is already pulled by the
  time the state is read back.

Also fixes the runner dir rename fallback in remove(), which never worked: the
trash name was built from the absolute path of the dir, so the rename target was a
relative path under a `.trash-` dir that does not exist.

Part-of: #4182
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@un-def
un-def merged commit a70d98b into master Aug 27, 2026
27 checks passed
@un-def
un-def deleted the issue_4182_shim_persist_task_state branch August 27, 2026 08:09
un-def added a commit that referenced this pull request Aug 28, 2026
`~/.dstack/runners/<name>` is mounted into the task container as
`/tmp/runner`. It used to hold runner's files only, but shim now keeps
its own files there as well -- the image pull log since #2903 and the
task state file since #4220 -- sharing them with the runner and the user
workload, which may corrupt or delete them. Nothing sensitive is stored
there today, but the approach is unsafe: nothing stops a contributor
from putting a secret into a file the container can read.

That dir is now the task dir, private to shim, and only its new `runner`
subdir is mounted into the container:

    ~/.dstack/runners/<container-name>/  0700, shim only
      task.json
      pull.log
      runner/                           0755, mounted as /tmp/runner

* `runnerDir`/`runnersDir` are renamed to `taskDir`/`tasksDir`
  throughout, including the `DockerParameters` methods, to signal that
  the dir is managed by shim rather than by runner. The `runners` path
  itself is kept: an upgraded shim must find the dirs of the tasks
  created by the previous version.
* The dir of a restored task is no longer derived from the container
  mounts, which now point at the `runner` subdir, but from the task ID
  in its state file. The tasks dir is scanned once on start, and the
  result is shared by the restore and the orphan sweep, which used to
  scan the dir a second time.
* A task started by a shim version that did not write state files cannot
  be found by its state file, so its dir, named after the container, is
  looked up by name and, as before, removed along with the task.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant