Repository navigation
Conversation
Closed
bgentry
force-pushed
the
bg/conformance-harness
branch
4 times, most recently
from
October 6, 2026 04:26
3a61e95 to
6e5ad81
Compare
bgentry
force-pushed
the
bg/conformance-harness
branch
from
October 6, 2026 12:38
6e5ad81 to
903f1ab
Compare
bgentry
force-pushed
the
bg/conformance-fixtures
branch
from
October 6, 2026 12:38
db624b4 to
864b4af
Compare
bgentry
force-pushed
the
bg/conformance-harness
branch
2 times, most recently
from
October 6, 2026 13:34
274a4b1 to
d948cfc
Compare
bgentry
force-pushed
the
bg/conformance-harness
branch
2 times, most recently
from
October 6, 2026 15:17
d5649aa to
755282d
Compare
brandur
force-pushed
the
bg/conformance-fixtures
branch
from
October 6, 2026 16:56
864b4af to
35949ea
Compare
bgentry
force-pushed
the
bg/conformance-harness
branch
from
October 6, 2026 17:58
755282d to
ad9186a
Compare
The cross-language conformance suite talks to each implementation through an adapter process. Define the contract it speaks as Go types in a dependency-free `protocol` package: fourteen JSON-RPC methods (`handshake`, `migrate`, `insert`, `list`, `cancel`, `retry`, `queue`, `request_resign`, `tx_begin`, `tx_end`, `start`, `stop`, `stats`, and `release`), the job shape adapters report, the built-in worker's behaviors, and the error codes. Operations that may run in a caller's transaction take an optional `tx` name instead of having transactional mirrors, and the harness reads rows and injects faults with SQL itself, so the contract has no raw-row, fault, reset, or deterministic-value methods. Add River Go's adapter, the reference the other implementations are checked against. It's one handler generic over the driver's transaction type, so PostgreSQL (pgx) and SQLite share every method. It inserts `conformance_echo` jobs, runs a worker client whose jobs follow their `behavior` arg, records the events and counters scenarios observe, and can hold a client's first claim on a barrier through a pilot plugin.
bgentry
force-pushed
the
bg/conformance-harness
branch
from
October 6, 2026 18:44
ad9186a to
8d83e41
Compare
Each River implementation tests itself, but nothing checks that a Go process and a Rust or JavaScript process sharing one database agree: that one works the other's jobs, reads its rows, honors its unique keys, wakes on its notifications, and follows its leadership. Add a harness that runs those scenarios between River Go and a candidate implementation's adapter. The harness builds and starts adapters, drives them through a typed client for the contract, and reads and writes the database itself: rows, leaders, queues, migrations, notifications (`LISTEN` on PostgreSQL, the outbox on SQLite), and `pg_stat_activity` by each process's `application_name`. `EachDriver` and `EachDirection` run every scenario on PostgreSQL and SQLite, with each implementation in each role, in a database of its own: a schema reached through the adapters' `search_path`, or a SQLite file. With nothing shared, scenarios run in parallel, and Go against Go finishes in about 25 seconds. The scenarios cover insert-then-work in both directions, golden row comparisons of what each implementation stores when it inserts, works, claims, snoozes, discards, and rescues the same jobs, exact large numbers and IDs, batches, transactions, unique keys and conflicts, list cursors, migrations and custom schemas, notifications and their payloads, remote cancellation, queue control, leadership, claim order and competition, kind handling across a fleet, rescue and scheduling of the other's jobs, resumable cursors, and reserved metadata. Behavior one implementation exhibits alone stays in that implementation's own tests. `RIVER_CONFORMANCE` names the candidate (`go`, `rust`, or `js`); unset, every scenario skips, so `make test` is unaffected. `make test/conformance` runs the suite, against Go itself by default. The harness builds each adapter once per run and runs the built program directly, so killing an adapter kills the adapter itself. Rust's binary is found under `CARGO_TARGET_DIR` when it's set, and JavaScript's adapter is built after the root `riverqueue` package, which pnpm's `@riverqueue/conformance...` filter doesn't select. Each `Implementation` carries an exported `Build` function returning the directory its adapter runs in and the command that starts it, and `UseImplementations` replaces the set the `RIVER_CONFORMANCE` variables select from. Another module can use them, along with `RunBuild`, to run its own scenarios against adapters of its own; River's suite keeps its `go`, `rust`, and `js` implementations unchanged.
Some cross-language checks are too slow or disruptive for every pull request but still matter before a release: implementations must survive faults while sharing a database, stay within reach of River Go's performance, and run together for long periods. Add a nightly tier to the conformance harness. With `RIVER_CONFORMANCE_NIGHTLY` set, it kills processes holding running attempts and leadership so the other implementation rescues them and takes over, replaces every process in turn as a rolling deploy would, terminates listener and pool connections, takes the database away behind a TCP fault proxy, fails and blocks completions with triggers and row locks, holds SQLite's write lock past the busy timeout, hands workers rows they can't decode, and runs both implementations on PostgreSQL disguised as YugabyteDB. It also checks that completions are batched, that connections stay bounded under pool pressure, and that throughput and p95 latency stay within each implementation's bounds relative to Go's. `RIVER_CONFORMANCE_SOAK` runs mixed traffic with periodic leader restarts for the given duration, and `RIVER_CONFORMANCE_PEER` adds a third implementation for fleet scenarios. Pairs of non-Go implementations run the whole suite with `RIVER_CONFORMANCE_REFERENCE`. `make test/conformance/nightly` runs the tier.
River for Go's `InsertMany` and `InsertManyTx` fail an empty batch with "no jobs to insert", but the JavaScript client resolves it to `[]` without touching the database. A caller porting code between the two, or a mixed fleet comparing their behavior, sees a success in one language where the other reports an error. `insertMany` now throws a `ValidationError` with Go's message for an empty batch, in or out of a caller's transaction. The periodic job enqueuer already skips empty batches, so it's unaffected.
River for Go keeps `MaxAttempts` as an `int` and never caps it in the client. SQLite's integer columns hold 64 bits, so Go inserts and works a job with `MaxAttempts: 40000` on SQLite. On PostgreSQL, whose columns are `smallint`, Go's pgx and `database/sql` drivers clamp `max_attempts` and `priority` to 32767 on insert instead of failing. The JavaScript client rejects anything above 32767 on every driver, the SQLite driver rejects such attempt and max attempt counts again on insert, and the PostgreSQL drivers would fail with the database's range error, so a job Go handles fine can't be inserted from JavaScript. Drop the 16-bit ceiling from the client's `maxAttempts` validation and from prepared pilot rows, and let the SQLite driver store any positive safe integer for `attempt`, `max_attempts`, and attempt error numbers. The PostgreSQL and Prisma drivers clamp `max_attempts` and `priority` to 32767 in `jobInsertMany`, like Go's drivers. New tests insert a job with 40,000 maximum attempts on SQLite and read it back unchanged, and insert one through the PostgreSQL and Prisma drivers and read back 32,767.
River for Go's `ResumableStep` keeps the step function's error as the attempt's error, so a job whose step fails with "second failed" records exactly that. The JavaScript `Resumable` wraps it in a `LifecycleError` reading `resumable step "second" failed`, with the original only as its `cause`, so the same failure stores different text depending on which implementation worked the job. `step` and `stepWithCursor` now rethrow the step's own error and keep it as the attempt's failure. Only a non-`Error` throw is still wrapped, since River needs an `Error` to record.
When a leader's renewal finds no term to renew, because it expired or another process replaced the leader row (even under the same client ID), River for Go's elector gives up leadership at once without resigning. The JavaScript client instead keeps reporting itself leader and running maintenance until its local trust deadline passes, the election interval plus 10 seconds less a 2 second margin, so it runs leader-only work for a term another process holds. Lose leadership as soon as a renewal returns no term, without resigning the term this client no longer holds.
`updateQueue` atomically replaces a running queue's capacity and wakes its producer, but a producer whose workers are all busy waits only for one of them to finish. Raising `maxWorkers` on a full queue therefore starts nothing until a running job ends, which never happens for jobs blocked on each other. Wait for either a freed worker slot or the queue's wake-up, so a reconfiguration (or a stop) takes effect immediately.
The conformance harness launches an adapter per implementation, but JavaScript has none, so `RIVER_CONFORMANCE=js` fails at the build. Add the private, never-published `@riverqueue/conformance` workspace package. It implements the contract in `conformance/protocol`: strict JSON-RPC over stdin and stdout, the 14 methods with Go's error codes, the built-in worker's 10 behaviors, named barriers, the claim barrier (a `PilotClient` whose first claim holds its jobs), and stats from River's events and hooks. One `Adapter` class runs over River's driver interface, so PostgreSQL (node-postgres, one driver per schema) and SQLite (`node:sqlite`, one handle per open transaction) differ only in how they connect and begin transactions. `make test/conformance/js` builds the root `riverqueue` package and then the adapter with its workspace dependencies, and runs the pull request tier with JavaScript as the candidate. The package joins the workspace's build, lint, format, and type-check scripts.
The TypeScript conformance suite checks every Go retry case against its delay bounds, every unique key, and that the notification pump dispatches each of Go's notification payloads. It doesn't check that the drivers write those payloads, that PostgreSQL's spaced `json_build_object` form of a payload reads like Go's compact form, or that the retry delay with no jitter is exactly Go's base delay rather than anywhere inside the jitter window. Check the topics and payload fields the SQLite driver writes, and the PostgreSQL driver sends, for each of Go's notification goldens; that every notification payload reader parses the spaced form of each golden payload; and that the default retry policy schedules exactly Go's minimum delay for every retry case when the jitter is zero. The PostgreSQL driver's notification check is an integration test, so `make test/js/integration` now generates the fixtures first, and the integration CI job sets up Go and generates them before testing, as the unit test job does. `make test/js/conformance` also runs the new payload reader test.
The cross-language harness only covers what two implementations do together. Behaviors River for JavaScript shows alone were tested mostly with fake drivers, or not at all, such as cancelling a refetched attempt after a snooze. Test them against the real SQLite driver: bulk and transactional job updates and deletes, worker outcomes and runtime faults, aborted and transactional completions, queue reads, dynamic queue reconfiguration and pause, error handler cancellation, extension ordering, resumable retries and validation, timeouts, shutdown classification, graceful stop, job and queue cleaner retention, the rescuer paging past a full batch, leadership term replacement, and periodic and scheduled jobs. On PostgreSQL, test leader renewal while the job cleaner waits on a row lock and the reindexer skipping missing and artifact indexes.
River Go stores `attempt` and `max_attempts` as Go `int`s. On SQLite, whose columns are native integers, a Go client can insert a job with 40,000 maximum attempts and read it back unchanged. On PostgreSQL, whose columns are `smallint`, Go's drivers clamp `max_attempts` to 32,767 on insert instead of failing. River Rust modeled both as `i16`, PostgreSQL's column width: it couldn't insert such a job at all, and read one Go wrote as 32,767, so the conformance scenario inserting one from Rust failed. Make attempt counts `i32` across the crates: `JobRow::attempt` and `max_attempts`, `AttemptError::attempt` and `AttemptError::new`, `InsertOpts::max_attempts`/`with_max_attempts`, `InsertParams::max_attempts`, `ClientBuilder::default_max_attempts`, `MAX_ATTEMPTS_DEFAULT`, the derive macro's `max_attempts` bound, and `riverqueue-test`'s builders. SQLite rows decode the native integers directly instead of saturating them into `i16`; a value beyond `i32` makes the row undecodable like other columns River can't represent. PostgreSQL rows widen the `smallint` columns, and a PostgreSQL insert clamps `max_attempts` to 32,767 like Go's pgx and `database/sql` drivers. New tests insert and list a job with 40,000 maximum attempts on SQLite, and check PostgreSQL stores and lists it as 32,767.
The conformance rebuild moves behaviors one implementation exhibits alone out of the cross-language harness and into each port's own tests. Its coverage map lists where the Rust suite was missing assertions for them. Add or extend tests so the Rust suite makes every listed assertion: - Worker outcomes, panics, job timeouts, resumable step validation, and remote cancellation of a snoozed and refetched attempt, in a new `worker_outcomes` test. - An error handler cancelling a job with attempts left, and insert middleware wrapping the insert hook before any work hook runs. - Dynamic queues: a removed queue's job stays available, and a queue added at runtime pauses, resumes, and is removed. - Queue get and list reflecting updates, and pause, resume, and update of unknown queue names (with spaces, 129 characters) reporting not found, plus `|` in queue names. - Transactions: commit of a PostgreSQL transaction aborted by a failed statement rolls back, transactional completion records no errors, and update and delete-many in a transaction stay invisible until commit. - Leadership: a same-ID term replacement is followed by a fresh term, with run-on-start periodic jobs inserted once per term and completed. - Maintenance: the queue cleaner keeps recently updated and actively worked queues, the reindexer skips missing indexes, and the rescuer pages past a full default batch of timeout-disabled jobs to discard an unregistered kind. - Go-generated fixtures: the client's default retry policy stays within Go's bounds, and notification topics and payload fields match Go's on PostgreSQL and in SQLite's outbox. `make test/rust/conformance` also runs the SQLite notification check, so a Go change to notification payloads is checked against Rust.
A River Go worker cancels its job with `JobCancel(err)`, and the attempt records `JobCancelError: <err>`. A Rust worker could only return `WorkOutcome::Cancel`, which records the fixed message "job cancelled by worker", so a Rust-worked job lost the reason a Go-worked one keeps. Add `JobCancelError`, an error a worker returns to cancel its job with a reason. Anywhere in the worker error's source chain, it cancels the job whatever attempts it has left, skips the error handler as Go does, and records `JobCancelError: <reason>`, the same text Go writes. `WorkOutcome::Cancel` keeps cancelling without a reason. The worker outcome tests cover a job with attempts left being cancelled and recording its reason.
Add `riverqueue-conformance`, an unpublished workspace binary that serves the 14-method contract in `conformance/protocol` for River's Rust client, so the harness can run its scenarios with Rust as the candidate against River Go. One `Server<B: Backend>` handles every method. `Backend` holds only what differs between PostgreSQL and SQLite: opening the pool, the schema a client uses, beginning a transaction and handing it to River's `.tx` requests, and migrating. Everything else goes through River's database-independent `Client`. The built-in worker implements the ten behaviors, cancelling with a `JobCancelError` reason as Go's adapter does with `JobCancel`. A pilot implements the claim barrier, and a hook, an error handler, and an event subscription record the stats. Jobs are reported with their stored args and metadata as raw JSON, so numbers no float can hold, such as `1e400`, keep their exact text as they do in Go's adapter. `make test/conformance/rust` and `make test/conformance/rust/nightly` build the adapter first so a compile error is reported once. `cargo package --workspace` includes `publish = false` crates, so `check/rust/package` excludes the adapter by name; `cargo publish --workspace` already skips it.
The harness proves River Go, Rust, and JavaScript agree when they share a database, but nothing runs it, so a change to any of them can break the agreement unnoticed. Each pull request runs one job per candidate, only when that candidate's files or the conformance module change: Go against itself in the `Conformance` workflow, and Rust and JavaScript against Go in their own workflows, which already filter on `rust/**`, `js/**`, and `conformance/**`. A job runs every scenario on PostgreSQL 18 and then SQLite, so the pull request tier adds three runners, each a few minutes, and no matrix. A new `Conformance nightly` workflow, on a schedule and by hand, runs what's too slow or noisy for pull requests: - The nightly tier (process kills, database faults, rolling deploys, completion batching, pool pressure) for each candidate on PostgreSQL 14 through 18, with SQLite only on 18. - A three-engine fleet of Go, Rust, and JavaScript, and the ordinary suite with JavaScript against Rust as the reference. - Throughput and p95 latency of Rust and JavaScript relative to Go, reported without failing the run on shared runners. - A one-hour soak through all three engines. Its `release_candidate` input soaks for 5.5 hours instead, the most a GitHub-hosted job's six-hour limit leaves room for, and makes a throughput or latency regression fail the run.
bgentry
force-pushed
the
bg/conformance-harness
branch
from
October 6, 2026 19:38
8d83e41 to
bc0fb3d
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
River now has three implementations, Go, Rust, and JavaScript, and an application can run any mix of them against one database. They only work together if they agree on everything they share there: the rows they write, the unique keys they compute, the notifications they send, who holds leadership, and how a stuck job gets rescued. Each port's own tests can't prove that, because they never run next to another implementation. For example, River Go can insert a job with 40,000 maximum attempts on SQLite, but the Rust client read it back as 32,767 and couldn't insert one at all, and the JavaScript client rejected it too. A JavaScript leader whose term had been taken over kept running maintenance until its own deadline passed. None of that shows up until two implementations share a database.
This adds a harness that puts them in one. River Go is the reference, and a candidate (Go itself, Rust, or JavaScript) runs each scenario with it: one inserts and the other works, one leads and the other takes over, one cancels a job the other is running, and both must leave the database as Go alone would. Every scenario runs on PostgreSQL and SQLite.
Implementations talk to the harness through a 14-method JSON-RPC contract over stdin and stdout (
conformance/protocol), and the harness drives them through a typed Go client with one method per contract method. Each language has one adapter serving both drivers, since River's driver interface already hides the difference: Go's inconformance/cmd, Rust's as the unpublishedriverqueue-conformancecrate, and JavaScript's as the private@riverqueue/conformanceworkspace package. The harness doesn't ask adapters what's in the database. It reads rows, leadership, queues, and notifications itself with SQL, and injects faults the same way (expiring a leader, terminating connections, holding locks). Each scenario gets its own schema on PostgreSQL or its own file on SQLite, so they all run in parallel. Go test names are the scenario list, with no registry or manifest beside them.There are two tiers. On a pull request, one job per language runs it against Go on PostgreSQL 18 and then SQLite, only when that language's code or the conformance module changes: Go against itself in the
Conformanceworkflow, and Rust and JavaScript in their own workflows. The scenarios take about 30 seconds, so each job is a few minutes, mostly building the adapter. A nightly workflow adds process kills, database faults, rolling deploys, and connection and batching checks on PostgreSQL 14 through 18 for each language, a fleet of all three engines, JavaScript against Rust as the reference, throughput and latency against Go, and an hour-long soak. Dispatching it as a release candidate soaks for 5.5 hours, the most GitHub's six-hour job limit allows, and fails on a performance regression.It builds on the nested
conformancemodule from #1451, which holds the fixtures generated from River Go that the ports' own tests read and is kept out of River's Go module zips.It replaces the earlier version of this PR. That one was about 32,000 lines here plus about 14,000 in the ports' adapters, most of it a 68-method contract with separate PostgreSQL and SQLite copies of every adapter, a scenario registry, a feature inventory, JSON schemas with their own validator, and committed generated JSON. Its PR tier took over an hour per language.
Coverage carries over. Each of the old suite's 167 scenarios maps to a harness scenario on the pull request or nightly tier, to the Go-generated fixtures, or, for behavior one implementation shows alone (a worker's outcomes, queue administration, a maintenance service's batching), to that port's own tests. Where a port's tests didn't already check those, this adds tests that do, against real drivers, in their own commits.
Running the ports against Go found these differences, each fixed in its own commit:
insertManyaccepts an empty batch, where Go fails it with "no jobs to insert".updateQueuestarts nothing until a running job finishes.i32, like Go'sint, and a PostgreSQL insert clamps a largermax_attemptsto 32,767 as Go's drivers do.JobCancel(err)does. A newJobCancelErrordoes, and the attempt records the text Go writes.The harness keeps the attempt count fixes in place: on SQLite, a job one implementation inserts with 40,000 maximum attempts is listed and worked by the other, and rows already past 32,767 attempts are worked, failed, retried, and listed with their attempts, maximum attempts, and recorded error attempt numbers unchanged. On PostgreSQL, the same insert succeeds in every implementation, stores and lists 32,767, and is worked by the other.
The commits go in order: the contract and Go's adapter, the pull request tier, and the nightly tier; then for each port its fixes, its adapter, and its new tests (for Rust, the new tests come before the
JobCancelErrorfix, whose test extends them, and the adapter comes last so it usesJobCancelErrorfrom the start); and then CI.