Skip to content

[9.5] [ML] Backport Sandbox2 stack (#3181, #3182, #3185, #3187, #3188, #3215) - #3237

Draft
valeriy42 wants to merge 10 commits into
elastic:9.5from
valeriy42:backport/9.5/sandbox2-stack
Draft

valeriy42 wants to merge 10 commits into
elastic:9.5from
valeriy42:backport/9.5/sandbox2-stack

Conversation

@valeriy42

Copy link
Copy Markdown
Contributor

Backport of the Sandbox2 stack from main to 9.5, cherry-picked in order with -x:

Four follow-ups from main are included because the stack needs them: #3208 (third-party warnings are not errors, needed for Abseil/Sandbox2), #3224 and #3225 (ship controller-protocol.version in the Linux and DRA bundles so Elasticsearch's protocol check accepts them) and #3233 (retry third-party clones, fail configure on clone errors).

All ten picks apply without conflicts. Not built or run locally; relying on PR CI.

valeriy42 and others added 10 commits October 9, 2026 14:42
## Summary

Reconstructs the dormant dependency/build foundation for Sandbox2 from
current `main`, as a clean first slice ahead of the sandbox policy,
spawner, and controller-routing changes that land in follow-up PRs. No
controller or `pytorch_inference` routing changes in this PR — Sandbox2
is not selectable from any production code path yet.

Frozen PR elastic#2873 attempted this feature in one large branch; this PR
takes just its dependency/build layer and replaces its inline
`file(WRITE)`/`string(REGEX REPLACE)` source rewrites with checked-in,
version-pinned patches that fail the configure step loudly instead of
silently no-op'ing on upstream drift.

- `3rd_party/CMakeLists.txt`: FetchContent `sandboxed-api` `v20241008` on
  Linux, applying 4 checked-in patches via `git apply` (fails configure
  loudly on upstream drift, idempotent across reconfigure).
- `3rd_party/patches/sandboxed-api/`: the 4 patches (disable vendored
  gtest, stop `-fno-exceptions` propagating into ml-cpp targets, make
  Python3 optional, link zlib + static libstdc++/libgcc into the
  forkserver binary).
- `3rd_party/licenses/{abseil,sandbox2}-*`: license/attribution files.
- `lib/sandbox/`: dormant `MlSandbox` target (`CMlSandboxAvailability`
  query only — no policy/spawner/diagnostics) plus a Linux-only
  forkserver runtime smoke test that forks/execs/reaps a
  dynamically-linked payload via `PolicyBuilder::AddLibrariesForBinary()`.

Verified locally (Linux x86_64, both a plain configure and
`-DCMAKE_UNITY_BUILD=ON`) before pushing, and green on this repo's own
Linux/macOS/Windows CI, license/security scans, and Java integration
suites.

<sub>Stack created with <a href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>

(cherry picked from commit 4a8b7ae)
…aded seccomp (elastic#3182)

Replace the hand-maintained BPF jump-offset table in `CSystemCallFilter_Linux.cc` with a program builder that derives every jump from the allowlist vector's own size/index. The applied program is generated from `CPytorchInferenceSyscallAllowlist.h`, a single machine-readable declaration, instead of a parallel hardcoded list.

`CSystemCallFilter::installSystemCallFilter()` returns a typed `ESystemCallFilterInstallOutcome` across all three platform implementations instead of `void`, and logs an `ml.seccomp.installed` readiness marker on success. `pytorch_inference/Main.cc` gains a `decideDegradedModeAction()` decision that would terminate before `CIoManager::initIo()` on any degraded-mode seccomp failure; that termination stays behind an internal switch defaulting to false until the controller has a typed way to know a degraded-mode launch was a deliberate operator choice rather than the only option available. Flipping it on today would terminate every launch on a host lacking seccomp BPF, with no operator fallback to select instead. The four non-PyTorch callers now make their unchanged log-and-continue policy explicit instead of silently discarding the result.

Also adds an explicit structured `degradedModeAttestationMarker()` so a controller/Elasticsearch observer can assert seccomp installation directly instead of inferring it from the absence of a fatal log line, and carries forward two pytorch_inference/libtorch compatibility fixes from elastic#2873 into the shared declaration (`clone3` allowed by literal syscall number, and `prlimit64`) with a named regression test, so a from-scratch rewrite doesn't silently drop them.

`CSeccompFilterBuilderTest.cc` decodes the actually-built BPF program to prove it matches the declaration and that jump offsets are derived, not hand-maintained, plus fault-injected coverage of `decideDegradedModeAction()` for every install-failure class.

A future Sandbox2 filesystem/network policy doesn't exist yet in this branch's history; this change establishes the single declaration for that policy to consume once it's written.

(cherry picked from commit a210a30)
…rence (elastic#3185)

Stacks on elastic#3182.

Replaces raw argument-directory inference in `CPytorchInferenceSandboxPolicy` with a typed launch spec: every `input`/`output`/`restore`/`logPipe` path is validated against a pinned child-root contract (`$TMPDIR/ml-child-ipc/<child-id>`) before any policy is built. Rejects relative, root-level, dot-dot, out-of-root, wrong-depth, duplicate, mutable-symlink/alias, and cross-option child-id-mismatch paths - never widens a mount to recover a rejected argument.

Also minimizes the filesystem policy: enumerates and justifies all seven historically bulk-mounted fixed directories, replaces whole `/etc` with five individually justified files, never binds host `/proc`/`/sys` (relies on Sandbox2's own namespaced procfs/sysfs), uses a private bounded tmpfs at `/tmp` instead of the host's, and consumes the syscall allowlist already shared with the legacy BPF filter instead of hand-duplicating it.

Adds a purpose-built allowlisted mechanism probe (`ml_sandbox_probe`) proving allowed IPC access, denied host reads, denied external egress, loopback reachability, and mount enumeration, plus a portable validator unit-test suite and a Linux-only mechanism integration test.

Verified this session: the validator's core logic compiles clean with `-Wall -Wextra -Werror` and passes a standalone driver covering every rejection/acceptance path against real `mkdtemp`/`mkdir`/`symlink` fixtures. The `SANDBOX2_AVAILABLE`/Linux path compiles clean against stub sandbox2/seccomp headers (no vendored Sandbox2 headers available on this host). Not yet verified: an actual Sandbox2 run of the mechanism probe and the real Linux CMake/build integration - needs a Linux CI or devbox pass, in progress. Also fixes a Windows build break this change would otherwise have introduced (POSIX-only realpath/PATH_MAX used unconditionally in a file ml-cpp builds on every platform).

(cherry picked from commit 2ce41a8)
…elastic#3187)

## Summary

Rebuilds `CSandboxedProcessSpawner` around an explicit lifecycle state machine (`Prepared -> Launched -> IdentityCaptured -> Registered -> Monitoring -> TerminationRequested -> CleanupRequired -> Reaped/Failed`), stacked on elastic#3184.

- Non-throwing kill-and-reap guard, armed immediately after launch and PID capture, holding its own owning reference to the sandbox handle so cleanup is order-independent during unwinding.
- Four injectable seams (pidfd acquisition, registry allocation, monitor-thread launch, sandbox completion) for deterministic fault injection.
- Explicit pidfd outcome classification: a kernel lacking pidfd support (`ENOSYS`) is the *only* case that falls back to signalling the sandboxee through its owned monitor handle (SIGKILL, identity-safe, no numeric-PID lookup); every other pidfd error is treated as a resource/identity failure and fails registration outright. Numeric `kill(pid, ...)` does not appear anywhere in this file.
- A CAS-controlled one-shot latch replacing two-boolean coordination for the timeout-vs-completion race.
- `CSandboxedProcessSpawnerLifecycleTest_Linux.cc`: fault-injection coverage for every pidfd class, allocation/resource failures, stale-generation protection, PID-reuse identity binding, the CAS race, descriptor-baseline cleanup, and destructor-latency.

## Notes

- `E_Launched`/`E_CleanupRequired` lifecycle states are declared but intentionally left unassigned (no natural single point without broader restructuring) — not silently dropped.
- The SIGKILL-only assumption for the owned-monitor termination path (no graceful-SIGTERM variant is buildable against the pinned Sandbox2 dependency version) was cross-checked against the vendored monitor source during review.

(cherry picked from commit 12c3ebf)
…c#3188)

## Summary

Stacks on elastic#3187. Adds typed controller-side routing behind two symmetric controller tokens. Landlock fallback when Sandbox2 is unavailable is split to stacked elastic#3215.

- `CCommandProcessor` parses at most one of `--disableSandbox` (operator kill switch, forces legacy) or `--requireSandbox` (operator opt-in, forces Sandbox2 on capable hosts) for the exact configured `pytorch_inference` path. Duplicate occurrences of either token, a mismatched process path, or both tokens present together are all rejected before spawn. The selected token is stripped before the child ever sees it.
- New `CProcessSpawnerRouter` dispatches the decided route to `CSandboxedProcessSpawner` or `CDetachedProcessSpawner` with no automatic fallback: a required Sandbox2 launch that fails returns a failed start response, it never retries through the legacy spawner.
- A `start` command with neither token always takes the legacy route - permanent behaviour for any caller that sends no routing token (support/debug scripts, direct controller invocation, the test harness), not a rollout seam. Elasticsearch is expected to always send exactly one of the two tokens per launch, chosen from its own operator setting's live value.
- Structured once-per-launch observability signal (`sandbox2_launch`: `deployment_id`, `model_id`, `route`, `sandbox2_established`, `mode`, `legacy_reason`, `sandbox2_compiled_in`) so a consumer can distinguish an operator kill switch from the no-token default, and "Sandbox2 supported but no routing token sent" from "built without Sandbox2 support."
- `ML_SANDBOXED` is stripped from every legacy-route child's environment (POSIX and Windows) so an inherited or injected value can never fail-open the mandatory in-process seccomp filter.
- Active Sandbox2 capability probe logged once at controller start (`CSandbox2Diagnostics`) so operators see whether `--requireSandbox` can succeed before a deployment fails closed.
- Staged user-namespace capability probe plus CI wiring: aarch64 runs enforced-mode coverage; x86_64 runs fail-closed-only coverage, since x86_64 Buildkite k8s pods get `EPERM` on `mount("proc", ...)` and there is currently no userns-capable x86_64 CI runner.
- Repaired `test_sandbox2_attack_defense.py`: reached markers, unsandboxed positive controls, per-case cleanup/reap assertions, the real `ml-child-ipc/<child-id>` IPC layout, an assertion that the controller's own `sandbox2_launch` signal reports the expected route before any security-boundary check runs, and explicit `--requireSandbox`/`--disableSandbox` tokens per case instead of a global env-var default.
- Producer-side controller-protocol capability token (`3rd_party/controller-protocol.version`, now version 2) for a future cross-repo compatibility check.
- Build packaging: per-library `install_libs()` guard and unconditional `$ORIGIN` RPATH on bundled ELF libraries so sibling `.so` dependencies resolve when the controller environment is cleared.

## Filesystem-policy fixes found during enforced-mode qualification

End-to-end qualification on the qaf harness with `xpack.ml.trained_models.sandbox_enabled=true` (real `mode:enforced` route, not the legacy fallback) surfaced several launch/filesystem-policy gaps that previously only manifested once a real `pytorch_inference` ran under the enforced sandbox. Fixed here:

- **Per-child IPC directory is created before validation.** The native controller now creates `$TMPDIR/ml-child-ipc/<child-id>` (mode 0700) before `validateChildIpcLaunchSpec()`'s live `realpath()` calls, on both spawn paths. Elasticsearch only ever constructs the IPC path strings; nothing created the directory they name.
- **Per-child IPC root is mounted at the same path inside and outside the sandbox** (no `/run/elastic/ml-ipc` remap), because `pytorch_inference` receives its `--input=`/`--output=`/`--restore=`/`--logPipe=` argv as host paths under that root - a remap left them unresolvable in the sandbox mount namespace.
- **Sandbox2-specific syscall allowlist** ported from the frozen reference (elastic#2873): Sandbox2's namespace/threading setup exercises syscalls the legacy in-process BPF filter never needed, so `legacyBpfAllowedSyscalls()` alone is insufficient.
- **`/proc` is mounted (PID-namespaced) inside the sandbox rootfs.** Sandbox2 mounts a fresh PID-namespaced procfs on the outer root, then `pivot_root`s into the chroot and detaches the old root, so the rootfs had no `/proc`. Without it `readlink(/proc/self/exe)` fails with `ENOENT`, breaking Intel oneMKL's runtime dispatcher (it reads `/proc/self/exe` to self-locate and `dlopen` its CPU-specific `libmkl_*.so.3` kernels) - every enforced `pytorch_inference` aborted with `Intel oneMKL FATAL ERROR: Cannot load <mkl-loader>`. The mount binds the already-namespaced procfs (never the host's): inside the sandbox `/proc` shows only the sandboxee's own PIDs. `/sys` is left unmounted.
- **Regression test:** `ml_sandbox_probe` now checks `/proc/self/exe` is readable inside the real sandbox, asserted by `CPytorchInferenceSandboxPolicyMechanismTest_Linux` - turning the MKL crash into a build-time failure.

Verified: ml-cpp sandbox unit tests 33/33; qaf boot check and `test_scenario_buildly` (8/8) pass under `mode:enforced` with zero MKL crashes and no legacy fallbacks.
(cherry picked from commit 16c5935)
…tic#3215)

## Summary

Stacks on elastic#3188. When Elasticsearch sends `--requireSandbox` but the host cannot run Sandbox2 (ECH allocators with user namespaces disabled, container seccomp blocking `CLONE_NEWUSER`, and similar), the controller steps down to a Landlock filesystem ruleset plus the existing in-process seccomp filter instead of failing every deployment closed.

- `hostConfinement()` probes once at controller start (same cached verdict as the startup self-check) and walks a fixed ladder: Sandbox2, then Landlock, then refusal with an operator-actionable message returned to Elasticsearch.
- New `sandbox2_launch` mode `"landlock"` (`route:"sandbox2"`, `sandbox2_established:false`). Consumers must not treat `route` alone as full isolation.
- Router-only `--restrictFilesystem` token; `CCommandProcessor` rejects caller-supplied copies.
- Documented in `docs/sandbox2_production_failure_modes.md`. Companion: elastic/elasticsearch#159052 must accept `mode:"landlock"`.

A failed Sandbox2 launch on a host that can run Sandbox2 is never retried under Landlock.

(cherry picked from commit f6a9c5b)
ml-cpp applies its strict warning flags (including -Wconversion,
-Wunused-parameter, ...) to every target via
add_compile_options(${ML_CXX_FLAGS}), and the debug Linux CI build enables
CMAKE_COMPILE_WARNING_AS_ERROR=ON (elastic#3198). Once third-party dependencies are
actually compiled from the 3rd_party subtree this breaks the build in two
places:

  1. Compiling the third-party sources themselves, whose own warnings are
     promoted to errors. Fixed by disabling warnings-as-errors for the
     third-party subtree only (matching the existing save/restore pattern),
     so their warnings stay visible but non-fatal.

  2. Compiling our own targets that include Sandbox2/Abseil/protobuf headers
     (MlSandbox and its unit tests). Those headers are pulled in as normal
     (-I) includes - deliberately not -isystem, so they stay ahead of
     PyTorch's bundled protobuf in the search path - and are not warning-clean
     (conversion, unused-parameter, ...). Fixed by turning off
     warnings-as-errors on just those consuming targets, which is order-safe
     and covers every warning class the third-party headers may trip.

ml-cpp's own code keeps warnings-as-errors everywhere else.

Co-authored-by: Cursor <cursoragent@cursor.com>
(cherry picked from commit 19f76ea)
…lastic#3224)

The controller-protocol.version marker was only added to the Gradle
buildZip packaging path, but the Linux artifacts consumed by CI and DRA
are produced by dev-tools/docker/docker_entrypoint.sh (platform zip) and
.buildkite/scripts/steps/create_dra.sh (-deps/-nodeps split), neither of
which staged the marker. As a result the -nodeps bundle lacked the file
and Elasticsearch's new verifyControllerProtocolVersion gate rejected it,
failing the nightly PyTorch build's triggered Java integration tests.

Stage the marker at the bundle root in the Docker packaging step and add
it to the create_dra.sh -nodeps include list, with guarded fast-fail
checks in both paths so a future omission fails loudly at packaging time
rather than as an opaque downstream Elasticsearch build failure.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
(cherry picked from commit 53f3d64)
The marker was added to the -nodeps include list but not the -deps
prune list, so it leaked into both bundles. Elasticsearch's ml plugin
unzips -deps and -nodeps together, so Gradle's Copy fails with a
duplicate controller-protocol.version entry. Prune it from -deps so it
ships only in -nodeps, where the verifyControllerProtocolVersion gate
expects it.

(cherry picked from commit b0c52aa)
…#3233)

The Eigen and Valijson sources are cloned at CMake configure time from
gitlab.com and github.com respectively. Those hosts occasionally return
transient errors (e.g. GitLab "currently unable to handle this request
due to load"), and a single failed clone was enough to break an entire
CI build, requiring a manual rebuild.

Wrap each clone in a bounded retry loop (5 attempts, increasing backoff)
that starts from a clean slate on every attempt, so a brief hosting
outage no longer fails the build.

Also propagate the failure from the outer execute_process() calls that
run these scripts. Previously the FATAL_ERROR raised inside the child
`cmake -P` process was swallowed: configure logged the error but
continued with an empty 3rd_party/eigen, so the failure only surfaced
much later as a cryptic "Eigen/Core: No such file or directory" compile
error. COMMAND_ERROR_IS_FATAL ANY makes configure stop immediately with
the clear message once retries are exhausted, finally delivering the
behaviour elastic#3164 intended.

(cherry picked from commit bcce4aa)
@valeriy42 valeriy42 changed the title t [9.5] [ML] Backport Sandbox2 stack (#3181, #3182, #3185, #3187, #3188, #3215) Oct 9, 2026
@valeriy42 valeriy42 added the v9.5.6 Release version v9.5.6 label Oct 9, 2026
@valeriy42 valeriy42 removed the v9.5.5 label Oct 9, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants