Skip to content

Add nonce endpoint - #47

Merged
GCdePaula merged 7 commits into
mainfrom
feature/nonce-endpoint
Sep 28, 2026
Merged

GCdePaula merged 7 commits into
mainfrom
feature/nonce-endpoint

Conversation

@GCdePaula

@GCdePaula GCdePaula commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Summary

Adds two public ingress reads so a wallet can prepare a signature without guessing: GET /nonce?sender= (the nonce to sign next) and GET /domain (the EIP-712 domain /tx verifies against). It also makes the user-nonce rule the nonce endpoint relies on an explicit application contract, and rejects the unreachable-in-practice u32::MAX nonce instead of panicking.

What changed

GET /nonce?sender=0x… → {"sender": "0xAbC…", "next_nonce": 7}

  • next_nonce is one past the sender's latest user op in a valid batch, or 0. It is named next_nonce because nonce in /tx responses and the WS feed is the nonce an op consumed.
  • The value is derived from user_ops rows the inclusion lane already writes in the FULL commit that authorizes each POST /tx ack. There is no new table and no new writer; a (sender, nonce) index keeps the lookup off a full-history scan.
  • A read after a 200 sees the op, because the commit precedes the ack.
  • Recovery lowers the value on its own: the valid_batches filter excludes invalidated rows, which are kept and whose nonces are reused on resubmit.
  • The value counts soft-confirmed ops. It is a hint for one op in flight per sender; concurrent submits from one sender are unsupported.
  • A never-seen sender gets 200 with 0. A lookup that finds no row is an answer, not a storage fault.

GET /domain → {"name", "version", "chainId", "verifyingContract"}

  • Serves the exact domain /tx verifies against, keyed as eth_signTypedData_v4 expects.
  • Clients pin their own domain and assert against this one, like checking eth_chainId. chainId stays a JSON number; the README explains why.

Also

  • /fee and /nonce send Cache-Control: no-store. /fee is otherwise unchanged.
  • There is no nonce or fee pre-filter on POST /tx. A stale op costs the lane an in-memory check and no transaction; a pre-filter would add a SQLite read to every honest submit and is easy to bypass. The api.rs module doc records the reasoning and when to revisit it.
  • The user-nonce rule is now contract: accounts start at 0, validation requires an exact match, each included op advances the nonce by one, and nothing else changes it. It is stated in the application contract, the trait docs, the C header, and AGENTS.md.
  • The shared validate_and_execute_user_op rejects nonce u32::MAX with InvalidReason::NonceExhausted, which the lane and the canonical scheduler both apply. The wallet used to panic in apply there. Reaching it takes 2³²−1 paid ops from one sender.
  • Rust SDK: adds get_nonce and get_domain. GetFeeError is renamed to QueryError, shared by the three read routes.
  • Performance notes are filed in-tree: a dated measurement record and a new "Known optimizations" section in docs/review/register.md.

Why this design

Asking the app for a nonce would mean querying the inclusion lane, which puts reads into the one bounded queue and onto the single app thread. Reading SQLite does not contend with the lane: WAL readers never block the writer. The lane already persists every included op's (sender, nonce), so the nonce is derived from facts that already exist.

We considered and rejected two replicated tables:

  • Written by the API after each ack: it adds a second hot-path writer, loses updates when the process crashes or the client disconnects after the ack, and never rolls back when recovery lowers nonces.
  • Written by the lane: it needs a recovery-time rewind, and recomputing that rewind is the same query.

Known limitation

After an operator rebuild (setup --recovery), a sender with no surviving op since the rebuild reads 0 even when its nonce in the rebuilt baseline is higher. A 422 bad-nonce rejection still names the expected nonce. This is documented in the README and the contract, and the rebuild-round-trip e2e test asserts it. Cockroach recovery is being redesigned as a separate tool; seeding baseline nonces belongs there (Track 6 note).

Performance

Measured on an M5 Max running macOS, with WAL + synchronous=FULL, 64-op chunks, 192k ops. Full method and limits are in docs/review/2026-09-25-sender-index-commit-cost.md.

  • Write side: the sender index raises a chunk commit from 0.27 to 1.0 ms mean, and from 1.1 to 9 ms at p99, because checkpoints run inline. That is about 12 µs per op, against the 500 ms ack target.
  • Why the index stays: any sender-keyed structure pays about the same once the sender population is large, and a derived index needs no recovery maintenance.
  • Read side: about 2 µs per query plus about 0.2 ms to open a connection per request. After recovery, the lookup also walks the sender's invalidated rows above its nonce, about 0.11 µs each, bounded by that sender's own rolled-back volume.
  • The alternatives and their revisit triggers are in the register.

Risk and compatibility

  • Schema: rewrites baseline migration 0001 to add the index. There are no deployed databases, but existing data directories need a fresh setup to get it.
  • Rejection semantics: adds NonceExhausted in the shared boundary, which changes canonical behavior only for an op at u32::MAX, which used to crash the machine.
  • C ABI: appends APPLICATION_ENGINE_NONCE_EXHAUSTED = 3, carried like the max-fee reason. Engines never report it, and the host treats it as unsupported if one does. @edubart: this is additive only, but it touches the shared header.
  • Contract: contiguous, exact-match, +1 per-sender nonces are now required of every engine; libdex should confirm it follows the rule.
  • SDK: GetFeeError is renamed to QueryError.
  • API: additive only. POST /tx and the /fee body are unchanged.

Every account starts at 0, validation accepts only the expected nonce,
each included user op (business failures included) advances its
sender's nonce by exactly one, nothing else changes a nonce, and an op
at u32::MAX is never included. This was the wallet's behavior; the
sequencer now relies on it to derive GET /nonce from persisted user ops
instead of asking the engine, so the rule is stated in the contract, the
trait docs, the C header, and AGENTS.md.
GET /nonce?sender= returns the nonce a sender signs next: one past its
latest user op in a valid batch, or 0. It reads the rows the lane
already writes in the FULL commit that authorizes each POST /tx ack, so
a read after a 200 sees the op, and recovery lowers the value through
the valid-batch filter. No new table or writer: a (sender, nonce) index
on user_ops keeps the lookup off a full-history scan. POST /tx keeps no
nonce or fee pre-filter; the api.rs module doc records why.

GET /domain serves the EIP-712 domain /tx verifies against, keyed as
eth_signTypedData_v4 expects, for clients to assert against the domain
they pin. /fee and /nonce responses carry Cache-Control: no-store.

The Rust SDK gains get_nonce and get_domain; GetFeeError becomes
QueryError, shared by the three read routes. The README documents the
one-op-in-flight client model and the known gap: a sender idle since a
rebuilt baseline reads 0 until its first op. Closing that gap is a
Track 6 follow-up.

Schema: rewrites baseline migration 0001 (no deployed databases), so
existing data directories need a fresh setup to get the index.
The harness wallet gains served_next_nonce. Restart, stale-batch
recovery, Tip cascade, and cold-replica recovery now assert that
GET /nonce agrees with the independent replay wallet, including the
drop when recovery invalidates soft-confirmed ops. The rebuild
round trip pins the documented gap (a sender idle since the rebuild
reads 0) and exactness after its first op in the new era.
A dated record keeps the 2026-09-25 measurement: the sender index raises
the mean 64-op chunk commit from 0.27 to 1.0 ms and p99 from 1.1 to
9 ms on the measured machine, and a lane-written per-sender table
compares favorably only while the sender population fits in a few pages.

The register gains a "Known optimizations" section for headroom with a
known mechanism: the sender index and its alternatives, checkpoints
running inside the lane's commit, and per-request read connections. Each
names its evidence and revisit trigger; the index's schema comment
points there.
u32::MAX has no successor under the nonce rule. The wallet accepted an
op carrying it when the sender's expected nonce was u32::MAX, then
panicked in apply, taking down the lane (or the canonical machine).
validate_and_execute_user_op now rejects such an op with
InvalidReason::NonceExhausted before app validation, next to the
max-fee guard, so the lane and the canonical scheduler agree. Reaching
it takes 2^32-1 paid ops from one sender; the rejection is a courtesy,
not a live threat.

The C header appends APPLICATION_ENGINE_NONCE_EXHAUSTED, carried like
the max-fee reason so the vocabulary stays whole; engines never report
it, and the host treats it as unsupported if one does.
Adds the read side to the sender-index record: a lookup walks the
sender's invalidated index entries above its current nonce (about
0.11 us each, 16 ms at 100,000), bounded by that sender's own
rolled-back volume, while opening a read connection per request costs
about 0.2 ms, a hundred times the query. The register entries now carry
these numbers and the covering-index option.

States that the rebuilt-baseline gap also covers a sender whose
post-rebuild ops a later recovery invalidated, and why /domain keeps
chainId a JSON number: it is exact for every chain browser wallets
accept, and a larger ID fails closed at POST /tx.
@GCdePaula
GCdePaula marked this pull request as ready for review September 25, 2026 13:47
stephenctw
stephenctw previously approved these changes Sep 26, 2026
The exhausted-nonce guard ran before app validation, so any op carrying
u32::MAX was rejected as "nonce 4294967295 has no successor", even when
the sender expected a lower nonce. The guard now runs after the app
accepts the op: an op at the wrong nonce gets the app's "bad nonce:
expected N, got 4294967295", and NonceExhausted means exactly that the
sender's expected nonce is u32::MAX. Both paths reject, and the
canonical scheduler ignores rejection reasons, so consensus is unchanged.

The threat-model row now states what GET /nonce exposes: a live
per-address count of soft-confirmed ops, published before their batches
reach L1. It reveals activity timing, not op contents, and moves only
after the op is sequenced.
@GCdePaula
GCdePaula merged commit acafcdc into main Sep 28, 2026
11 of 12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants