Skip to content

Initial benchmarks suite - #396

Merged
Pino de Candia (pinodeca) merged 4 commits into
mainfrom
thomcc/benchmarks
Sep 22, 2026
Merged

Pino de Candia (pinodeca) merged 4 commits into
mainfrom
thomcc/benchmarks

Conversation

@thomcc-work

@thomcc-work Thom Chiovoloni (thomcc-work) commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

This PR adds a benchmark harness for pg_durable, initially focussed on HTTP.

It supports benchmarking SQL baseline, normal HTTP post, multipart HTTP, and custom SQL workload scripts.

It can handle sequential and concurrent workflows, and the warmups, repetitions, transaction counts, timeouts, and completion polling are all configurable.

When targetting local HTTP, keep-alive, payload sizes, response delay are all configurable too, and request concurrency and TCP connection reuse are measured.

However, the primary measurements are throuput/latency (but it tracks raw samples, and environment/configuration metadata in its JSON report).

Failures and incomplete runs are all detected too, of course.

It's exercised lightly in CI.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Incorrect pgbench variable quoting causes invalid IDs, URLs, and labels, breaking execution and cleanup.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 5 High severity · 1 Low severity

Open (6)
What changed in this PR

Adds a reusable pgbench suite for measuring completed SQL and HTTP workflows.

Changes:

  • Adds benchmark workloads, runner, polling helper, and HTTP fixture.
  • Adds unit/integration tests and CI checks.
  • Documents setup, execution, results, and cleanup.
File Description
.github/​workflows/​ci.yml Runs harness unit tests.
.gitignore Ignores benchmark Python cache.
README.md Links benchmark documentation.
benchmarks/​README.md Documents the suite.
benchmarks/​await.sql Polls workflow completion.
benchmarks/​http_server.py Provides the HTTP fixture.
benchmarks/​run.py Implements benchmark orchestration and reporting.
benchmarks/​test_run.py Tests runner and fixture behavior.
benchmarks/​workloads/​http-multipart.sql Adds multipart workload.
benchmarks/​workloads/​http.sql Adds HTTP workload.
benchmarks/​workloads/​sql.sql Adds SQL baseline workload.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread benchmarks/run.py
Comment thread benchmarks/test_run.py
Comment thread benchmarks/workloads/http-multipart.sql
Comment thread benchmarks/workloads/http.sql
Comment thread benchmarks/workloads/sql.sql
Comment thread benchmarks/README.md

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

HTTP workloads and two timeout tests use a literal label, preventing scoped cleanup of benchmark instances.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 3 High severity

Open (3)
Resolved since last review (6)

Comment thread benchmarks/test_run.py
Comment thread benchmarks/workloads/http-multipart.sql
Comment thread benchmarks/workloads/http.sql
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The source dirty-state metadata can be incorrect, and one timeout test fails to label and clean up its workflow correctly.

Review effort: Balanced
Findings: None

Resolved since last review (3)
Previously missed (2)

In code that hasn't changed since last review

Medium severity Dirty check includes newly generated output artifacts

benchmarks/​run.py:216

The dirty-state check runs after creating directory, workload.sql, await.sql, and results.json. If --output names a non-ignored directory inside the checkout, those newly generated untracked files make an otherwise clean source tree report dirty: true, corrupting the source metadata. Capture the Git status before creating benchmark artifacts, or explicitly exclude the output directory from this check.

Medium severity Literal run label prevents timed-out instance cleanup

benchmarks/​test_run.py:350

This passes the literal label :run_label to df.start() because the placeholder is inside the SQL string. When the one-second wait times out, cancel_instances() filters by the generated run label and misses this still-running df.sleep(2) instance; assert_cleaned_up() misses it for the same reason. Use the quoted pgbench variable form used by the other timeout test.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The harness is well-scoped, defensively handles failures and cleanup, and has substantial automated coverage.

Review effort: Balanced
Findings: None

@pinodeca Pino de Candia (pinodeca) left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I ran a ATV (all-the-vibes) starter kit ce-review with 12 specialist agents, GPT-6 Astra with Extra High effort and 272k context window. No actionable findings.

I also skimmed the changes myself, read the benchmark Readme and workloads. Then I ran the SQL workload and interrogated Copilot about how a developer would use the suite for a perf improvement task.

This suite is a good initial tool for systematically approaching performance work in pg_durable.

@pinodeca
Pino de Candia (pinodeca) merged commit f22ff73 into main Sep 22, 2026
19 checks passed
@pinodeca
Pino de Candia (pinodeca) deleted the thomcc/benchmarks branch September 22, 2026 20:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants