You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This PR adds a benchmark harness for pg_durable, initially focussed on HTTP.
It supports benchmarking SQL baseline, normal HTTP post, multipart HTTP, and custom SQL workload scripts.
It can handle sequential and concurrent workflows, and the warmups, repetitions, transaction counts, timeouts, and completion polling are all configurable.
When targetting local HTTP, keep-alive, payload sizes, response delay are all configurable too, and request concurrency and TCP connection reuse are measured.
However, the primary measurements are throuput/latency (but it tracks raw samples, and environment/configuration metadata in its JSON report).
Failures and incomplete runs are all detected too, of course.
Dirty check includes newly generated output artifacts
benchmarks/run.py:216
The dirty-state check runs after creating directory, workload.sql, await.sql, and results.json. If --output names a non-ignored directory inside the checkout, those newly generated untracked files make an otherwise clean source tree report dirty: true, corrupting the source metadata. Capture the Git status before creating benchmark artifacts, or explicitly exclude the output directory from this check.
Literal run label prevents timed-out instance cleanup
benchmarks/test_run.py:350
This passes the literal label :run_label to df.start() because the placeholder is inside the SQL string. When the one-second wait times out, cancel_instances() filters by the generated run label and misses this still-running df.sleep(2) instance; assert_cleaned_up() misses it for the same reason. Use the quoted pgbench variable form used by the other timeout test.
The reason will be displayed to describe this comment to others. Learn more.
I ran a ATV (all-the-vibes) starter kit ce-review with 12 specialist agents, GPT-6 Astra with Extra High effort and 272k context window. No actionable findings.
I also skimmed the changes myself, read the benchmark Readme and workloads. Then I ran the SQL workload and interrogated Copilot about how a developer would use the suite for a perf improvement task.
This suite is a good initial tool for systematically approaching performance work in pg_durable.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR adds a benchmark harness for pg_durable, initially focussed on HTTP.
It supports benchmarking SQL baseline, normal HTTP post, multipart HTTP, and custom SQL workload scripts.
It can handle sequential and concurrent workflows, and the warmups, repetitions, transaction counts, timeouts, and completion polling are all configurable.
When targetting local HTTP, keep-alive, payload sizes, response delay are all configurable too, and request concurrency and TCP connection reuse are measured.
However, the primary measurements are throuput/latency (but it tracks raw samples, and environment/configuration metadata in its JSON report).
Failures and incomplete runs are all detected too, of course.
It's exercised lightly in CI.