You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Flaky Rust test: notification_poll_failures_do_not_stop_the_client hits database is locked #1455
@brandurnotification_poll_failures_do_not_stop_the_client in rust/riverqueue/tests/resilience_sqlite.rs (added with the Rust port in #1442) fails intermittently in CI with SQLite's database is locked:
thread 'notification_poll_failures_do_not_stop_the_client' panicked at riverqueue/tests/resilience_sqlite.rs:208:10:
called `Result::unwrap()` on an `Err` value: Database(SqliteError { code: 5, message: "database is locked" })
Line 208 is the unwrap in the job_state helper, which the test calls from wait_until while waiting for the job to complete.
It's shown up in about 6 of the ~12 Rust workflow runs since #1442 merged, 8 jobs in all, always with the same panic. It doesn't follow a particular PostgreSQL or Rust version:
The likely cause is the test's own setup. It opens its pool in rollback-journal (Delete) mode with a 20 ms busy timeout, so that the outbox poll fails quickly while a second connection holds BEGIN EXCLUSIVE. The client and the test's job_state reads share that pool. After the exclusive lock is released, the client keeps writing (fetch polls every 20 ms with a 1 ms cooldown, plus the insert and completion), and in rollback-journal mode a committing writer blocks readers. On a loaded runner a commit can take longer than 20 ms, so job_state gets SQLITE_BUSY and its unwrap panics. The client.insert just before the wait uses the same pool and has the same exposure.
Two possible fixes: read job state for the assertions through a separate pool with a normal busy timeout, or have the wait_until condition treat a busy error as "not yet" and keep polling.
@brandur
notification_poll_failures_do_not_stop_the_clientinrust/riverqueue/tests/resilience_sqlite.rs(added with the Rust port in #1442) fails intermittently in CI with SQLite'sdatabase is locked:Line 208 is the
unwrapin thejob_statehelper, which the test calls fromwait_untilwhile waiting for the job to complete.It's shown up in about 6 of the ~12 Rust workflow runs since #1442 merged, 8 jobs in all, always with the same panic. It doesn't follow a particular PostgreSQL or Rust version:
postgres (15): https://github.com/riverqueue/river/actions/runs/37464710415/job/112272738126postgres (17): https://github.com/riverqueue/river/actions/runs/37464713159/job/112272746416postgres (14)andrust_versions (1.96)in run https://github.com/riverqueue/river/actions/runs/37464710877rust_versions (1.97)), 37376861834 (postgres (18)), and 37376860109 (postgres (15),rust_versions (1.97))The likely cause is the test's own setup. It opens its pool in rollback-journal (
Delete) mode with a 20 ms busy timeout, so that the outbox poll fails quickly while a second connection holdsBEGIN EXCLUSIVE. The client and the test'sjob_statereads share that pool. After the exclusive lock is released, the client keeps writing (fetch polls every 20 ms with a 1 ms cooldown, plus the insert and completion), and in rollback-journal mode a committing writer blocks readers. On a loaded runner a commit can take longer than 20 ms, sojob_stategetsSQLITE_BUSYand itsunwrappanics. Theclient.insertjust before the wait uses the same pool and has the same exposure.Two possible fixes: read job state for the assertions through a separate pool with a normal busy timeout, or have the
wait_untilcondition treat a busy error as "not yet" and keep polling.