Skip to content

comm: add coop-recv --wait for blocking inbox reads #56

Description

@ProKil

Follow-up to @akhatua2's suggestion on #51.

Motivation

Today coop-recv does a non-blocking drain (LPOP-until-empty). Agents that want to wait for a reply have to poll, which:

  • Burns step budget. Each poll = one tool call = one step against LimitsExceeded.
  • Relies on the model to remember to poll. Weaker models drift; missing one poll cycle costs a coordination round.
  • Doesn't model the natural collaboration pattern: "I asked you a question, I'll wait for the answer."

A blocking variant turns a polling loop (N steps) into one step.

Proposed change

Add --wait (alias for blocking) and --timeout SECONDS to coop-recv:

coop-send agent2 \"what does foo() return?\" && coop-recv --wait --timeout 60

Implementation: swap LPOP for BLPOP (or BLMPOP if we want to drain multiple at once) with the configured timeout. No changes to coop-send; composition via shell && is enough.

Default timeout: 60s.

No new send flag needed — composes with existing CLI.

Design decisions

1. Timeout

  • Default 60s. BLPOP 0 (forever) is not an option — mutual --wait would deadlock both agents indefinitely.
  • On timeout: return empty JSON list [], same shape as today's coop-recv. The model treats it as "no reply, move on" rather than an error.
  • --timeout accepts seconds (int). Document the cap (e.g. ≤ the adapter's bash-timeout minus a margin) — see risk (4) below.

2. Semantics: "wait for next message", not "wait for reply"

The inbox is FIFO-from-anyone. --wait will return the next message from any sender, which may not be a reply to the specific send that preceded it.

Start without correlation IDs. Document the semantics: "--wait returns the next inbound message from any agent; if you need correlated request/reply, include a tag in your message content and filter client-side." If real request/reply becomes common we add coop-send --id X / coop-recv --wait-for X later — that's a non-trivial envelope change (every recv has to filter, log harvesting changes shape) and not worth doing speculatively.

3. Wake on peer shutdown

If agent B submits and exits while A is BLPOP-waiting on B, A blocks until the 60s timeout for no reason.

Fix: when any agent submits / the runtime tears down, push a sentinel to every other agent's inbox:

{\"from\": \"_runtime\", \"kind\": \"peer_shutdown\", \"agent\": \"agent2\", \"timestamp\": ...}

coop-recv --wait returns it like any other message. The model sees "agent2 submitted" and moves on. This needs a one-line hook in the runner's teardown path.

4. Step accounting

One tool call = one step regardless of wait duration. Document this in the prompt section that describes the comm tools — otherwise the model might infer "waiting is free" and over-use --wait (e.g. --wait immediately after every send even when one-way notification was fine).

Risks

  • Mutual --wait deadlock: mitigated by 60s default timeout. Both agents return empty, log it, continue. Annoying but not catastrophic.
  • Adapter bash-timeout interaction: if any adapter kills bash before 60s, --wait --timeout 60 returns SIGTERM not []. Need to either lower the default or audit each adapter's bash timeout. Action: before merging, check mini_swe_agent_v2, codex, openhands_agent_sdk, claude_code, swe_agent bash timeouts and cap --wait's timeout at min(adapter_bash_timeout - 5s, requested).
  • Race when reply arrives before --wait: BLPOP returns immediately if the list is non-empty. Correct behavior, worth one test.
  • Redis connection holding: a 60s BLPOP holds one connection. Redis default maxclients is 10000; even with 50-agent runs this is fine, but if we ever go higher, switch to BLMPOP with shorter polls or connection pooling.

Scope

Touches src/cooperbench/agents/_coop/coop_msg.py and the prompt sections that describe comm tools in the per-adapter prompts. Backwards compatible — the new flag is opt-in; existing coop-recv invocations stay non-blocking.

Tests:

  • --wait returns immediately when inbox non-empty
  • --wait --timeout 1 returns [] after ~1s when inbox empty
  • Peer-shutdown sentinel wakes a waiter
  • Adapter bash-timeout cap is enforced

Out of scope for this issue

  • Correlation IDs (see (2))
  • A coop-send --then-wait shorthand (composable via && is fine)
  • Changing coop-broadcast semantics

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions