Skip to content

ensemble: route prompts to workers found by gossip - #64

Merged
adiled merged 3 commits into
mainfrom
hive/discover-remote-workers
Oct 4, 2026
Merged

adiled merged 3 commits into
mainfrom
hive/discover-remote-workers

Conversation

@adiled

@adiled adiled commented Oct 4, 2026

Copy link
Copy Markdown
Owner

Closes #63.

What

A prompt naming a model that no local bee advertises got chi:"error", even when a peer humd advertised a worker for it over hum/hives/announce. Now it routes there.

The far side already worked. humd parses from on ensemble-sourced prompts, records sid_origins, and routes worker replies back to the origin — so the missing piece was exactly one hop.

How

  • Ensemble::hive_discover_all() — unfiltered manifests. hive_discover(name) is now a filter over it.
  • humd holds heard manifests keyed by the humd that can route to them (the bee has no ensemble presence; only its humd is dialable), inner-keyed by bee hid / nestler id / name.
  • On a local worker miss: pick a connected peer advertising that model, set to/from, route.
  • A no-worker error now routes back to a remote caller too, instead of dying in the local session and leaving the origin to time out.
  • Re-advertise on PeerAdd. A bee advertises only when it handshakes with its own humd, which a peer reconnect does not re-trigger — so reconnect left peers invisible.
  • PeerRemove drops that humd's manifests; selection ignores any humd absent from ens.peers(), so a manifest cannot outlive its peer.

Proof

sim/tests/remote_discovery.rs — two humds. The laptop has no worker and sends a bare chi:"prompt" naming only a model; the server's worker is the sole advertiser and the laptop has never heard of it.

Fails without the fix:

laptop never got chi:finish from the discovered worker — discovery did
not route the prompt across the mesh
Some("no worker bee advertises model 'claude-opus-4-7'")

Passes with it. Full suite 210 passed / 0 failed; clippy clean under the CI command.

Notes

  • Selection is ens.peers()-gated rather than lease-gated. A manifest has no timestamp and there is no re-advertise heartbeat, so TTL eviction would drop live entries. The peer set is the honest liveness signal, and it is free.
  • A failed forward falls through to the error reply rather than returning, so a dead peer cannot hang the caller.
  • Kept selection to one source: model match against manifests. pick_overflow_peer still selects by caps/free_slots for overflow — deliberately untouched, see below.

Not in scope

pick_overflow_peer and manifest-discovery are now two selection mechanisms that can drift. Worth unifying, but that is a behaviour change to overflow routing and deserves its own issue.

Address-level discovery is also still hand-rolled: IrohTransport::connect requires an iroh: hint and passes no relay URL, so an EndpointId alone cannot dial us. iroh 1.0's endpoint_info (Pkarr + DNS, with UserData for a descriptor) is the substrate for that, and ENS/chain would be AddressLookup impls. Separate crate, separate issue.

A prompt naming a model no local bee advertises got an error, even
when a peer humd advertised one over hum/hives/announce.

hive_discover now has an unfiltered form; humd keeps what it hears
keyed by the humd that can route to it, and on a local worker miss
picks a connected peer advertising that model and routes the prompt
with to/from set. The far side already resolved remote prompts by
model and replied by sid via sid_origins, so one hop was the whole
gap. A no-worker error now also goes back to a remote caller instead
of dying in the local session.

Re-advertise on PeerAdd: a bee only advertises when it handshakes
with its own humd, which a peer reconnect does not re-trigger.
PeerRemove drops that humd's manifests, and selection ignores any
humd absent from the peer set, so a manifest cannot outlive its peer.

A failed forward falls through to the error reply rather than
returning, so a dead peer cannot hang the caller.
adiled added 2 commits October 4, 2026 19:30
Adds prose for sim/tests/remote_discovery.rs and records which tests
have no scenario yet, so the 1:1 claim in this README stops implying
coverage that does not exist.
eviction_only_touches_the_dead_peer waited for Liveness::Dead, which is
only set once a background drain task observes the closed transport.
That is a scheduling assumption, not a guarantee, so the test timed out
on a loaded CI runner while passing locally every time.

A peer nobody is talking to stops being Live once the TTL passes, which
is wall-clock guaranteed. Poll the sweep for the reap instead, and probe
the surviving peer each pass so it is provably still fresh at reap time.

With a 120ms TTL the old assertion only held because it ran within
~40ms of the probe; any correct fix to the wait would have let the live
peer expire too and tripped the next assertion.

5.07s -> 0.40s, and no longer dependent on task scheduling.
@adiled
adiled merged commit 405c75c into main Oct 4, 2026
8 checks passed
@adiled
adiled deleted the hive/discover-remote-workers branch October 4, 2026 15:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

wire hive_discover: route prompts to workers discovered across the mesh

1 participant