Repository navigation
refactor(asap-tools): decouple remote_monitor lifecycle and client ownership #696
Description
Activity
milindsrivastava1997 commented
on Sep 3, 2026 ContributorAuthorMore actionsRemote monitor decoupling: staged implementation plan
Status: Plan drafted, not started
Design of record:.design_docs/remote-monitor-decoupling-design.md
PR description (draft):.design_docs/remote-monitor-decoupling-PR-description.md
Prerequisite: PR #522 merged tomain.Principles for this refactor
- Tiny commits, one concern each. Each stage below is a commit (or a small
handful) that leavesmaingreen and the experiment scripts runnable. - Additive-then-subtractive. New mechanisms land alongside the old ones and
one call site is migrated first; the old path is deleted only in a later stage,
once nothing calls it. No stage both adds a mechanism and removes its
predecessor. - Manual edits, no scripts. Renames/migrations done by hand, call site by
call site. - Pause after each stage for a manual commit — do not batch.
- Verification is behavioral, not just typecheck. Where a stage changes
runtime wiring, the verify step names the actual experiment run to exercise.
Ordering rationale: earliest stages are pure deletions/internal refactors with
zero behavior change (lowest risk, build confidence), then additive protocol
work, then the one genuinely behavioral change (client ownership), then
subtractive cleanup once the old path is dead.
Stage 0 — Remove dead
profile_query_engine_pidend-to-endWhy first: completely isolated, provably unreachable, zero behavior change —
the safest possible starting commit, and it shrinks the surface every later stage
touches.Files:
remote_monitor.py: delete theprofile_query_engine_pid = Nonevar (lines
~312-314) and its pass-through intoPrometheusClientService.start(~399).experiment_utils/services/prometheus_client_service.py: drop the
profile_query_engine_pidparam fromstart/_start_containerized/
_start_bare_metaland the twoif profile_query_engine_pid is not None
branches.generate_prometheus_client_compose.py: drop the--profile-query-engine-pid
arg and itstemplate_varswiring.asap-tools/queriers/prometheus-client/docker-compose.yml.j2: remove the
--profile_query_engine_pidline.asap-tools/queriers/prometheus-client/main_prometheus_client.py: remove the
--profile_query_engine_pidargparse arg and thestart_query_engine_profiler
thread block (~688-701); deletestart_query_engine_profilerif now unused.
Verify:
grep -rn profile_query_engine_pidreturns nothing;
grep -rn profile_query_engine\bstill shows the live--profile_query_engine
path intact. Bare-metal + containerized prometheus-client still start (dry-run
the generated compose / command string).Closes: issue #23 second bullet.
Stage 1 — Extract
Profilerinterface (internal refactor, same behavior)Why here: pure reorganization of
remote_monitor.pyinternals; no
orchestrator or wire-format change; nothing outsideremote_monitor.pysees it.Files: new
classes/profilers.py(or a section ofremote_monitor.py):class Profiler(ABC): def start(self, pids, output_dir) -> Any: ... def stop(self, handle, store: bool) -> None: ... class AsprofProfiler(Profiler) # from start/stop_profiling_flink_pids class FlamegraphProfiler(Profiler) # from start/stop_profiling_arroyo_pids class PerfProfiler(Profiler) # from start/stop_profiling_query_engine_pids # + convert_query_engine_perf_data folded into stop()
remote_monitor.pykeeps its current CLI andmain()control flow, but the
if args.profile_flink_pids: .../ arroyo / QE blocks now instantiate and call
these classes instead of the free functions. Behavior identical.Verify: an experiment run with
profile_flink/profile_arroyo/
profile_query_engineenabled produces the same profile artifacts
(flink_profiles/,arroyo_profiles/,query_engine_profiles/) as before.
Diff the output tree against a pre-refactor run.Leave a check behind: a
__main__self-check inprofilers.pyasserting
each profiler builds the same command string it did as a free function (guards
the copy-paste consolidation).
Stage 2 — Introduce
MonitorRequest+--requestarg, additiveWhy here: the wire format is the foundation the protocol stages build on, but
it can land without touching lifecycle. Old flags stay;--requestis an
alternative entry point migrated one call site at a time.Files:
- New
classes/monitor_request.py: theMonitorTarget/ProfilerSpec/
WaitStrategy/CostExporterSpec/MonitorRequestdataclasses +to_json
/from_json. (Decision to record: shared module imported by both
remote_monitor.pyandremote_monitor_service.py, vs. duplicated schema —
see design doc "Not yet decided". Recommend shared module; both run from the
sameexperiments/dir.) remote_monitor.py: add--requestarg. When present, build the run from it;
when absent, fall back to today's flag parsing (a thin adapter that constructs
aMonitorRequestfrom the old args, so there is exactly one downstream code
path).execution_modemaps toWaitStrategy:timed→fixed_seconds,
interactive→stdin,prometheus_client/ingest→stop_signal.remote_monitor_service.py: no change yet.
Verify: run the old flag path (unchanged) and a hand-built
--requestJSON
for the same scenario; assert identicalmonitor_output.json. Round-trip test
MonitorRequest.from_json(r.to_json()) == r.
Stage 3 — PID file + command file protocol in
remote_monitor.pyWhy here: adds the node-side half of the lifecycle contract. Still additive —
the orchestrator doesn't use it yet, so nothing breaks.Files:
remote_monitor.py:- On start (once
output_diris known): write<output_dir>/remote_monitor.pid
=os.getpid(), unconditionally (overwrite stale). Register cleanup on clean
exit. - Install a SIGTERM handler that flushes output + removes the PID file.
stop_signalwait strategy: each poll interval, check
<output_dir>/remote_monitor.cmd; on{"action":"stop", "store":...}, break,
stop profilers withstore, write output, remove PID file, exit. (ControlCommand
dataclass inmonitor_request.pyor a sibling.)fixed_seconds/stdinstrategies also remove the PID file on their normal
exit (uniform cleanup).
Verify: launch
remote_monitor.py --requestwithstop_signaldirectly
(no orchestrator); confirm.pidappears,kill(pid,0)succeeds;touch/write
the.cmdfile → it exits within one interval and writes output and removes
.pid. Separately, send it SIGTERM mid-run → it flushes and removes.pid.Leave a check behind: small
test_control_protocol.pythat spawns the
runner in a subprocess against a temp dir and asserts the.pid/.cmd/output
lifecycle for both the command-file and SIGTERM paths.
Stage 4 —
MonitorSessionorchestrator methods, additiveWhy here: node side (Stage 3) exists, so the orchestrator half can be built
and unit-tested against it without yet migrating any experiment script.Files:
remote_monitor_service.py— add to the existing class (rename to
MonitorSessiondeferred to Stage 6 to keep this diff additive):start(request: MonitorRequest, manual_mode=False): serialize to JSON, one
shlex.quote, launch via the existingnohup … &path.wait_for_start(output_dir, timeout=30): poll for.pidexistence.wait_for_finish(output_dir, timeout=600): poll for.pidgone.stop(output_dir, store=True, timeout=60): write.cmd;wait_for_finish;
on timeout read PID from.pid,SIGTERMover SSH, short grace,SIGKILL.
Old methods (
startlegacy,wait_for_remote_monitor_to_finish,
kill_remote_monitor, and PR #522's six) stay untouched and still used by the
scripts. Both APIs coexist.Verify: on a CloudLab node, drive a bare monitor (no client) through
start/wait_for_start/stopvia a throwaway script; confirm.pidlifecycle
and thatstop's SIGTERM escalation fires when the monitor is artificially
wedged (e.g.kill -STOPthe process first).
Stage 5 — Migrate
experiment_run_clickhouse.pyto the new sessionWhy here: ClickHouse is the smaller / newer call site and already owns its
client (ClickHouseDataLoaderService), so it's the lower-risk first migration
and needs no client-ownership change — only swaps the monitor mechanism.Files:
experiment_run_clickhouse.py:- Ingest monitor: replace
start_clickhouse_ingest_monitor+
signal_ingest_monitor_stop+wait_for_remote_monitor_process_exit+
cleanup_ingest_monitor_stop_filewithsession.start(MonitorRequest(..., wait=stop_signal))/wait_for_start/stop. - Query workload: same
stop_signalsession bracket around the existing SQL
query client launch.
Verify: full
experiment_run_clickhouserun on a node; confirm
monitor_output_ingest.jsonandmonitor_output.jsonare both written and
non-empty, matching a pre-migration baseline run's shape.
Stage 6 — Migrate
experiment_run_e2e.py+ client ownership flipWhy here: the one genuinely behavioral change —
experiment_run_e2e.pymust
now launch/awaitQueryClientServiceitself instead ofremote_monitor.pydoing
it. Isolated to its own stage so it can be reviewed and verified alone.Files:
experiment_run_e2e.py: around the currentremote_monitor_service.start(...)wait_for_remote_monitor_to_finish(...)block:session.start(stop_signal)
→wait_for_start→ launch & awaitQueryClientService(mirroring the
ClickHouse data-loader pattern) →session.stop.
remote_monitor.py: delete theprometheus_clientin-process client launch
block (thePrometheusClientServiceimport and its start/health-loop/stop) —
now dead once e2e owns the client.
Verify: full
experiment_run_e2erun (bothsketchdbandbaselinemodes,
container and bare-metal prometheus-client); confirm query results +
monitor_output.jsonmatch a pre-migration baseline. This is the highest-risk
stage — run against a known-good prior experiment output and diff.Closes: issue #53, issue #23 first bullet.
Stage 7 — Delete the old path
Why last: only safe once Stages 5–6 removed every caller.
Files:
remote_monitor.py: remove the legacy flag-parsing adapter and old CLI args;
--requestbecomes the only entry point;main()is the straight-line runner.remote_monitor_service.py: delete legacystart,kill_remote_monitor,
wait_for_remote_monitor_to_finish, and PR feat(asap-tools): separate clickhouse ingest and query cpu/memory monitoring inexperiment run clickhouse#522's six methods
(start_clickhouse_ingest_monitor,is_remote_monitor_running,
wait_for_remote_monitor_start,wait_for_remote_monitor_process_exit,
signal_ingest_monitor_stop,cleanup_ingest_monitor_stop_file,
_remote_monitor_pgrep_pattern); rename class toMonitorSession; remove
constants.INGEST_MONITOR_STOP_FILE.- Remove now-unused keyword-building
if/eliftower fed only by the oldstart.
Verify:
grepconfirms no references to the deleted symbols;
grep -rn pgrep remote_monitorreturns nothing; one more full e2e + clickhouse
run as a final regression check.Closes: issue #91 (hardcoded keyword tower gone), issue #35 (no more
pgreppolling).
Cross-cutting: process discovery (issue #91)
Stages 5–6 pass
MonitorTarget(pid=...)wherever a service already knows its PID
(query engine, ClickHouse, Arroyo — several already expose
get_monitoring_keyword(); extend to hand over the PID).MonitorTarget(keyword=...)- node-side
resolve_targets()(housing today'sget_pids()) remains only as the
fallback for targets nothing owns. Enumerating exactly which current keywords can
become PIDs vs. must stay keyword-based is deferred to Stage 5/6 execution — see
design doc "Not yet decided".
Suggested verification asset
Before Stage 5, capture one known-good
experiment_run_clickhouseand one
experiment_run_e2eoutput tree onmainas golden baselines to diff every
subsequent stage against — the cheapest guard for a refactor this deep into the
experiment critical path.- Tiny commits, one concern each. Each stage below is a commit (or a small
milindsrivastava1997 commented
on Sep 3, 2026 ContributorAuthorMore actionsDecoupling
remote_monitor.py: designStatus: Design agreed, staged implementation plan written (
.design_docs/remote-monitor-decoupling-implementation-plan.md)
Touches:asap-tools/experiments/remote_monitor.py,experiment_utils/services/remote_monitor_service.py,classes/process_monitor.py,experiment_run_e2e.py,experiment_run_clickhouse.py,experiment_utils/services/prometheus_client_service.py(QueryClientService)
Relationship to PR #522: PR #522 (ClickHouse ingest/query monitor split) merges as-is on its own merits. This is a separate follow-up refactor that subsumes and removes the stop-file/pgrep-pattern plumbing PR #522 introduces, replacing it with the general mechanism below.
Background
remote_monitor.pyruns on a CloudLab node (launched via SSH withnohup … &so it survives the launching SSH exec returning) and currently mixes four concerns: process discovery, CPU/memory sampling, ad hoc profiling (Flink/Arroyo/query-engine), and orchestrating the query client (PrometheusClientService/QueryClientService) itself.experiment_run_e2e.pyandexperiment_run_clickhouse.pydrive it throughRemoteMonitorService.The load-bearing constraint: why
nohup … &The reason for
nohup … &is to detach the launched process's lifetime from the single SSH exec that started it — the orchestrator'sprovider.execute_commandfires each command as its own SSH exec that returns immediately, so withoutnohup &the child would get SIGHUP the moment that exec returns. It is not about the orchestrator disconnecting or staying connected. The consequence that shapes this whole design: once detached, the orchestrator holds no process handle and no pipe, so it must learn status and send control out-of-band. This is what rules out live-pipe IPC (e.g. naivemultiprocessingacross the boundary) and drives the file-based protocol below.Where the protocol files physically live, and who touches them across the SSH boundary
remote_monitor.pyruns on the node and usesCloudLabLocalProvider(local exec, no SSH) for its own subprocess needs. The orchestrator (MonitorSession, on the laptop) usesCloudLabProvider(SSH). The.pidand.cmdfiles live on the node's filesystem underexperiment_output_dir:- Orchestrator → files: writes/reads them over SSH via
provider.execute_command(touch,cat >,cat,test -f,kill). PR feat(asap-tools): separate clickhouse ingest and query cpu/memory monitoring inexperiment run clickhouse#522 already does exactly this —signal_ingest_monitor_stopis aprovider.execute_command("touch <stop_file>")— so the mechanic is proven, just generalized here. remote_monitor.py→ files: writes the.pidand reads the.cmdlocally (plain filesystem calls), since it is already on the node.
So "the orchestrator writes the command file" means an SSH
touch/cat >of a node-local path; "the monitor reads it" means a local file read on its poll interval.Problems found
- Process discovery by string-matching.
get_pids()shells out tops aux | grep -E ... | awkordocker inspect. Every new component adds another keyword threaded through anif/eliftower inremote_monitor_service.py(issue Remove hardcoded keywords from remote_monitor_service #91). - No real lifecycle contract across the SSH boundary. The orchestrator fires
remote_monitor.pydetached, then learns it's done by pollingpgrep -f remote_monitor.py(issue Investigate why remote_monitor takes time to shutdown after experiment is over #35 — the slow-shutdown-detection complaint is this). PR feat(asap-tools): separate clickhouse ingest and query cpu/memory monitoring inexperiment run clickhouse#522 needs a way to tell one specific invocation (the ClickHouse ingest monitor) to stop, and since no such primitive exists, it invents one from scratch: a stop-file, apgrepregex pattern-builder withre.escape/shlex.quoteto disambiguate instances, and five new wait/status methods. execution_modeis a growing, non-orthogonal enum.interactive/timed/prometheus_client, plus PR feat(asap-tools): separate clickhouse ingest and query cpu/memory monitoring inexperiment run clickhouse#522'singest. Each mode has bespoke branches in bothmain()andremote_monitor_service.start(), even though the real orthogonal axes are: which PIDs, how long to sample, which profilers, what happens when done.- Profiling is three copy-pasted subsystems (
start/stop_profiling_{flink,arroyo,query_engine}_pids) — same shape (subprocess per PID, track handle, SIGTERM on stop, optionally convert/store output), different tool per pair. - The wire format is hand-built shell strings.
remote_monitor_service.pyconstructs the entire remote invocation via.format()into a shell command; every new flag touches both the format string andremote_monitor.py's argparse in lockstep. - Components orchestrate each other instead of the master orchestrator controlling everything (issue Re-architect code to use Redis for control messages #23, closed unimplemented):
remote_monitor.pyimports and directly starts/blocks onQueryClientServicejust so it knows when to start/stop sampling;main_prometheus_client.pyin turn starts/stops query-engine profiling itself via--profile_query_engine_pid— but that value is hardcoded toNoneat its only live call site (remote_monitor.py:312-314, "unused for Rust QE"), so this second coupling is dead code for the current Rust query engine (which is profiled directly byremote_monitor.pyviaperf record, through the separate, live--profile_query_engineflag).
Issue #23 proposed Redis as the control-message transport for "master orchestrator sends control signals to components." We don't need it: the actual requirement is single-node start/stop/status signaling, which a PID file + a small command file achieves with zero new infrastructure.
Existing precedent for the target shape
ClickHouseDataLoaderService.start()already has the outer script (experiment_run_clickhouse.py) launch and block on a client directly viaprovider.execute_command, independent ofremote_monitor.py. PR #522's ingest-monitor bracket is this pattern applied to monitoring — it just lacked a clean stop primitive, so it built one (the stop-file) ad hoc. The design below generalizes that pattern and gives it the primitive it was missing, instead of it staying a one-off.
Goals
- One lifecycle contract for every remote-monitor invocation, replacing
pgrep-based polling and PR feat(asap-tools): separate clickhouse ingest and query cpu/memory monitoring inexperiment run clickhouse#522's stop-file. - Collapse
execution_modeinto an orthogonalwait strategy+targets+profilersrequest, not a growing enum. remote_monitor.pyno longer imports or knows aboutQueryClientService— the calling experiment script always owns launching/awaiting its client,remote_monitor.pyonly brackets it (resolves issue Explore how to decouple remote_monitor from PrometheusClient #53 and Re-architect code to use Redis for control messages #23's first bullet).- One
Profilerinterface instead of three copy-pasted function pairs. - Delete dead code:
profile_query_engine_pidend-to-end (issue Re-architect code to use Redis for control messages #23's second bullet — already unreachable, remove rather than redesign).
Non-goals
- Not changing
process_monitor.py'sMyMonitor/start_monitor/stop_monitorinternals — that's an in-processmultiprocessing.Pipeboundary betweenremote_monitor.pyand its sampler subprocess, a different and already-clean seam from the SSH-facing protocol below. - Not introducing Redis or any other new runtime dependency (considered per issue Re-architect code to use Redis for control messages #23, rejected — see Background).
- Not changing how the remote process is launched (
nohup … &over SSH stays; only how the orchestrator tracks/signals it afterward changes).
Target design
Control protocol: PID file + command file + SIGTERM escalation
Three plain-filesystem pieces, all scoped under the already-unique
experiment_output_dir, so no regex disambiguation between instances is ever needed:-
PID file (
<output_dir>/remote_monitor.pid): written unconditionally on start (a stale leftover from a crashed prior run is garbage, not a lock, and gets overwritten). Read by the orchestrator to check liveness (kill(pid, 0)) and as the target for the SIGTERM escalation path. -
Command file (
<output_dir>/remote_monitor.cmd): the general control-signal channel, checked byremote_monitor.pyon its existing per-interval poll — no new wake mechanism needed. Generalizes PR feat(asap-tools): separate clickhouse ingest and query cpu/memory monitoring inexperiment run clickhouse#522's boolean stop-file into a small structured command:@dataclass class ControlCommand: action: Literal["stop"] # room to grow: future actions add values here, not new plumbing store: bool = True # consolidates today's per-profiler `store: bool` into one shared field
-
SIGTERM (sent to the PID from the PID file): the escalation/safety-net only, not the primary channel — used when the command file goes unanswered past a timeout, mirroring the existing
terminate()/kill()escalation already inprocess_monitor.stop_monitor.
Request schema (replaces
execution_mode+ ~15 CLI flags)@dataclass class MonitorTarget: pid: Optional[int] = None # preferred: caller's service already knows it keyword: Optional[str] = None # fallback: resolved via ps/docker on the node label: str = "" # replaces today's "keyword" output tag @dataclass class ProfilerSpec: kind: Literal["asprof", "flamegraph", "perf"] pids: List[int] @dataclass class WaitStrategy: kind: Literal["fixed_seconds", "stdin", "stop_signal"] seconds: Optional[int] = None # only for fixed_seconds @dataclass class CostExporterSpec: addr: str port: int monitors_and_models: dict @dataclass class MonitorRequest: targets: List[MonitorTarget] monitors: List[str] # ["memory_info", "cpu_percent"] include_children: bool thread_attribution_keyword: Optional[str] profilers: List[ProfilerSpec] wait: WaitStrategy interval_seconds: float output_dir: str output_file: str cost_exporter: Optional[CostExporterSpec]
Today's four modes become
waitvalues:timed→fixed_seconds,interactive→stdin, bothprometheus_clientand PR #522'singest→stop_signal(same shape once the client is owned by the outer script — no reason for them to be different modes).Transport:
MonitorRequestis serialized as one compact JSON blob passed via a single CLI arg (remote_monitor.py --request '<json>'), replacing today's ~15 hand-formatted flags. No new remote file-push mechanism — strictly less shell-escaping surface than today, not more.remote_monitor.pyrunner shapedef main(request: MonitorRequest): write_pidfile(request.output_dir) # unconditional overwrite install_sigterm_handler(...) # escalation path only pids, labels = resolve_targets(request.targets) # get_pids() fallback lives here, small monitor, ctrl, mpipe = process_monitor.start_monitor(pids, labels, ...) # unchanged internals profilers = [make_profiler(s).start() for s in request.profilers] wait(request.wait) # sleep(N) | input() | poll command file each interval until action=="stop" for p in profilers: p.stop(store=<from command, default True>) monitor_info = process_monitor.stop_monitor(monitor, ctrl, mpipe) # unchanged write_output(monitor_info, request.output_dir, request.output_file) remove_pidfile(request.output_dir)
No
QueryClientService/PrometheusClientServiceimport. Noexecution_modebranches inmain().Profiler interface (replaces 3 copy-pasted function pairs)
class Profiler(ABC): def start(self, pids: List[int], output_dir: str) -> Any: ... def stop(self, handle: Any, store: bool) -> None: ... class AsprofProfiler(Profiler): ... # today's start/stop_profiling_flink_pids class FlamegraphProfiler(Profiler): ... # today's start/stop_profiling_arroyo_pids class PerfProfiler(Profiler): ... # today's start/stop_profiling_query_engine_pids # + convert_query_engine_perf_data folded into stop()
Orchestrator side:
MonitorSession(replacesRemoteMonitorService)class MonitorSession: def __init__(self, provider, node_offset): ... def start(self, request: MonitorRequest, manual_mode: bool = False) -> None: ... # serialize request, one shell-escape, launch via existing nohup ... & def wait_for_start(self, output_dir: str, timeout: int = 30) -> None: ... # poll for <output_dir>/remote_monitor.pid existing def wait_for_finish(self, output_dir: str, timeout: int = 600) -> None: ... # poll for pidfile gone (covers fixed_seconds/stdin self-exit) def stop(self, output_dir: str, store: bool = True, timeout: int = 60) -> None: ... # write command file {action: stop, store}; wait_for_finish; SIGTERM+SIGKILL escalation on timeout
This one class replaces
RemoteMonitorServiceplus all six of PR #522's new methods (start_clickhouse_ingest_monitor,is_remote_monitor_running,wait_for_remote_monitor_start,wait_for_remote_monitor_process_exit,signal_ingest_monitor_stop,cleanup_ingest_monitor_stop_file,_remote_monitor_pgrep_pattern) — the PID-file scoping means there's never a need to pattern-match "which remote_monitor.py instance."Call sites become uniform
Every place the outer script owns a client (ClickHouse ingest today, sketchdb query phase after this change) becomes the same four lines:
session.start(MonitorRequest(..., wait=WaitStrategy("stop_signal"))) session.wait_for_start(output_dir) client_service.start(...) # data_loader today; query_client_service after this change if client_service.use_container: while client_service.is_healthy(): time.sleep(5) client_service.stop() session.stop(output_dir)
experiment_run_e2e.pychanges to launch/awaitQueryClientServiceitself here, the wayexperiment_run_clickhouse.pyalready does forClickHouseDataLoaderService— this is the one real behavioral/control-flow change outsideremote_monitor.pyitself.Dead code to remove alongside this
profile_query_engine_pidend-to-end:remote_monitor.py,prometheus_client_service.py,generate_prometheus_client_compose.py's flag,docker-compose.yml.j2's template line,main_prometheus_client.py'sstart_query_engine_profilerthread. Confirmed unreachable — the only call site in this repo hardcodes it toNone("unused for Rust QE").
Open questions resolved during design
Question Decision Should the outer script always own client launch/await? Yes — matches existing ClickHouseDataLoaderServiceprecedent, resolves issue #53.Request transport: one JSON CLI arg vs. pushed file? One JSON CLI arg — no new remote file-push mechanism needed. Stale PID files from a crashed prior run? Always overwritten on start; not treated as a lock. Redis (issue #23) for control signals? Dropped — PID file + command file + SIGTERM escalation gives the same decoupling with no new dependency, and generalizes to future signal types via ControlCommand.action.Not yet decided
- Staged implementation/commit sequencing (next step).
- Exact
resolve_targets()responsibility split for target types with no owning service object (still falls back to keyword-basedps/dockerdiscovery — scope of that fallback not yet enumerated). - Whether
MonitorRequest/ControlCommandshould be shared Python types imported by both sides, or independently-maintained schemas kept in sync by convention (matters once this is staged into commits).
- Orchestrator → files: writes/reads them over SSH via
milindsrivastava1997 commented
on Sep 3, 2026 ContributorAuthorMore actionsrefactor(asap-tools): decouple remote_monitor lifecycle, replace pgrep/stop-file with PID+command files
Draft — no PR open yet. Paste this into the PR when the refactor lands. See
.design_docs/remote-monitor-decoupling-design.mdfor the full design.Summary
Reworks how the experiment orchestrator launches, tracks, and stops
remote_monitor.pyon a CloudLab node.remote_monitor.pystops orchestrating
the query client,execution_modecollapses into a request object, the three
copy-pasted profilers become oneProfilerinterface, andpgrep-based status
polling (plus PR #522's ad hoc stop-file) is replaced by a small, uniform
control protocol: a PID file, a command file, and SIGTERM
escalation.Closes #53, closes #23, closes #91; resolves the shutdown-detection cost in #35.
Removes the plumbing PR #522 added as a one-off (that PR merged separately on its
own merits; this subsumes it).Motivation
The orchestrator launches
remote_monitor.pydetached over SSH (nohup … &, so
the child survives the launching SSH exec returning). After that it holds no
process handle and no pipe — the child is a detached process on a remote machine.
Today it recovers status bypgrep -f remote_monitor.py, which matches leftover
crashed instances and needs regex disambiguation (exactly what PR #522 had to
build). Andremote_monitor.pyitself imports and blocks on the query client
just to know when to start/stop sampling, so a new backend phase (ClickHouse
ingest) meant a whole newexecution_modeplus a bespoke stop-file.The control protocol (the important part)
Three plain-filesystem pieces give the orchestrator the full lifecycle —
start / status / graceful stop / forced stop — with no new dependency (no Redis;
see #23). All scoped under the already-uniqueexperiment_output_dir, so there
is never any need to pattern-match which monitor instance is meant.1. PID file — "is it alive?"
<output_dir>/remote_monitor.pid, containing one number: the OS PID of the
remote_monitor.pyprocess.remote_monitor.pywrites it (unconditionally — a stale file from a crashed
prior run is garbage to overwrite, not a lock) as nearly the first thing on
startup, and deletes it as the last thing on clean exit.- The orchestrator reads it to check liveness via
kill(pid, 0)(sends no
signal; just succeeds if the PID is alive). File absent → not started, or
already finished and cleaned up. File present butkill(pid,0)fails →
crashed without cleanup.
This is what lets us delete PR #522's
_remote_monitor_pgrep_pattern/
re.escape/shlex.quotemachinery: the path already identifies the instance.2. Command file — "stop now" (extensible to other commands)
<output_dir>/remote_monitor.cmd, written by the orchestrator, read by
remote_monitor.py. Small structured command, e.g.
{"action": "stop", "store": true}.- The monitor's sampling loop already wakes every
interval_secondsto take a
sample; on each wake (in thestop_signalwait strategy) it also cheaply
checks for this file. No new thread, no new timer — it piggybacks the existing
poll cadence. - On seeing
action: stop, it breaks the loop, stops profilers (honoring
store), writes the output JSON, removes the PID file, exits.
Structured
action(vs. a bare marker file) is what makes this extensible:
future commands (pause,rotate, …) are new enum values, not new plumbing.
storehere consolidates the per-profilerstore: boolthe profiling code
already threads around ("keep the flamegraph" vs. "discard"). This is the
generalized form of PR #522's boolean stop-file.3. SIGTERM escalation — the safety net
The command file is cooperative: it only works if the monitor is actually
running its poll loop. If it's wedged (hung profiler subprocess, stuck syscall),
it never reads the file. SIGTERM is the fallback, not the normal stop path.
MonitorSession.stop():- Write
{"action":"stop"}to the command file. - Poll the PID file for up to
timeouts for the monitor to exit on its own. - Still alive → read PID from the PID file, send
SIGTERMover SSH.
remote_monitor.pyinstalls a SIGTERM handler so even this path flushes
output. - Still alive after a short grace period →
SIGKILL.
This is the same graceful→forceful ladder already in
process_monitor.stop_monitor(terminate()→kill()), applied across the
SSH boundary using the PID from the PID file.Need Mechanism Direction Has it started / is it alive / done? PID file + kill(pid,0)monitor → orchestrator Stop cleanly (extensible) command file ( action,store)orchestrator → monitor Stop now, unresponsive SIGTERM → SIGKILL to PID orchestrator → monitor (force) What changes
remote_monitor.py: driven by one--request '<json>'arg
(MonitorRequest) instead of ~15 flags;main()is a straight line with no
execution_modebranches; no longer importsQueryClientService. Writes PID
file, polls command file, escalates on SIGTERM.RemoteMonitorService→MonitorSession:start/wait_for_start/
wait_for_finish/stop, replacingRemoteMonitorServiceplus all six
methods PR feat(asap-tools): separate clickhouse ingest and query cpu/memory monitoring inexperiment run clickhouse#522 added.- Wait strategy replaces
execution_mode:fixed_seconds(wastimed),
stdin(wasinteractive),stop_signal(was bothprometheus_clientand
feat(asap-tools): separate clickhouse ingest and query cpu/memory monitoring inexperiment run clickhouse#522'singest— same shape once the outer script owns the client). - Profilers: one
Profilerinterface withAsprofProfiler/
FlamegraphProfiler/PerfProfilerimplementations, replacing the three
copy-pastedstart/stop_profiling_*_pidspairs. - Client ownership:
experiment_run_e2e.pylaunches/awaits
QueryClientServiceitself (asexperiment_run_clickhouse.pyalready does for
ClickHouseDataLoaderService);remote_monitor.pyonly brackets it.
Removed (dead code)
profile_query_engine_pidend-to-end (remote_monitor.py,
prometheus_client_service.py,generate_prometheus_client_compose.py,
docker-compose.yml.j2,main_prometheus_client.py's
start_query_engine_profilerthread). Its only live call site hardcodes it to
None("unused for Rust QE"); the Rust engine is profiled directly by
remote_monitor.pyviaperf recordthrough the separate--profile_query_engine
flag. (issue Re-architect code to use Redis for control messages #23, second bullet.)
Testing
- TODO once implemented — end-to-end
experiment_run_e2eand
experiment_run_clickhouseruns on a CloudLab node; verify monitor output JSON
is written for each wait strategy, and thatstopworks via both the command
file (normal) and the SIGTERM path (kill the loop, confirm escalation).
Context
The remote-monitor decoupling design in the .design_docs directory remains relevant and is not implemented on main. The current code still uses execution-mode branches, pgrep-based monitor discovery/lifecycle polling, the PR #522 ingest stop-file plumbing, and the legacy profile_query_engine_pid path.
This issue consolidates the remaining work described in the three design documents attached as comments below. It supersedes the need to track the work only through the older exploratory issues #23 and #53, while directly addressing the still-open concerns in #35 and #91.
Related: #23, #35, #53, #91, and merged PR #522.
Proposed outcome
The three source documents are included as comments on this issue.