Skip to content

Wait for the load balancer address instead of failing on the first read - #53

Merged
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address
Aug 19, 2026
Merged

Wait for the load balancer address instead of failing on the first read#53
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address

Conversation

@kierenj

@kierenj kierenj commented Aug 19, 2026

Copy link
Copy Markdown
Member

The problem

create reads the ingress address once, immediately after the rollout it waits for, and gives up if it is not there yet:

getting external IP...
jq: error (at <stdin>:54): Cannot iterate over null (null)
failed

The address is published by the ingress controller on its own cycle, independently of — and later than — the deployment rollout. So that read is usually too early. The job fails, Kubernetes retries it, and a later attempt eventually catches it. The Red Pepper MCP deploy hit this three times in dev and once in live; it only ever looked like flakiness.

How long it actually takes

Measured on both clusters by creating an ingress with a fresh hostname and timing until .status.loadBalancer.ingress[0].ip appeared:

cluster n min median max
NonLive 6 19.1s 51.8s 56.0s
Live 4 38.3s 51.5s 55.4s

The scattered minima against a hard ceiling near 56s is the signature of a single publication cycle — the job arrives at a random point in it, not variable load. Both clusters behave identically.

The fix

A bounded wait, polling until the address appears.

ADDRESS_TIMEOUT defaults to 120s and ADDRESS_POLL_SECONDS to 5s.

120s is deliberately more than 1.5× the measurement (~84s). Because the mechanism is a discrete ~55s cycle rather than a continuous distribution, ~84s covers barely more than one cycle and would fail outright if a single publication were missed — a controller restart or leader-election handover. 120s survives that. The asymmetry favours it: too generous only delays reporting a genuine failure, too tight reverts to spurious deploy failures. Both are overridable.

On failing fast instead

There is nothing to check. IngressStatus is only:

status.loadBalancer.ingress[] → { hostname, ip, ports[{ error, port, protocol }] }

No conditions, no phase. A pending address and a permanently broken one are byte-identical — both an empty loadBalancer. (ports[].error exists but ingress-nginx never populates it.) So the timeout prints the object's events, where a real failure does show up, and exits 1.

Verification

case result
address appears on the 3rd poll waits, then proceeds
address never appears times out, prints events, exit 1

Record-matching regression matrix from #52, all unchanged:

create pepper      exit=0  PATCH .../AAA
create pepper-mcp  exit=0  PATCH .../BBB
delete pepper      exit=0  DELETE .../AAA
create brand-new   exit=0  create path
delete missing     exit=0  no request
create dupe        exit=1  refusing to guess
delete dupe        exit=1  refusing to guess

Also collapses the duplicated ingress/service and ip/hostname branches, which held four copies of the same fetch-and-check.

Rollout

Needs 0.0.25 published, then a chart bump. k8s#126 currently pins 0.0.24 — if that has not merged yet it is cheaper to retarget it at 0.0.25 and ship both in one bump.

🤖 Generated with Claude Code

The create path read the ingress address once, immediately after the rollout it waits for, and gave
up if it was not there yet - "jq: error ... Cannot iterate over null" followed by "failed". The
address is published by the ingress controller on a separate cycle, so that read is usually too
early: the job failed and was retried until a later attempt happened to catch it. The Red Pepper MCP
deploy failed three times this way in dev and once in live.

Measured on both clusters by creating an ingress with a fresh hostname and timing the address:

  NonLive  n=6  min 19.1s  median 51.8s  max 56.0s
  Live     n=4  min 38.3s  median 51.5s  max 55.4s

The hard ceiling near 56s with scattered minima is one publication cycle - the job arrives at a
random point in it. ADDRESS_TIMEOUT therefore defaults to 120s, two cycles, so a single missed one
is survivable; ADDRESS_POLL_SECONDS defaults to 5.

On timeout the object's events are printed and the job exits 1. There is no state to check instead:
IngressStatus carries only loadBalancer.ingress[] and no conditions, so a pending address and a
permanently broken one are indistinguishable except by waiting.

Also collapses the duplicated ingress/service and ip/hostname branches, which had four copies of the
same fetch-and-check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings August 19, 2026 06:49

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR addresses flaky create runs by waiting for Kubernetes to publish the external load balancer address (IP/hostname) instead of reading it once immediately after rollout and failing when the address is not yet present.

Changes:

  • Add a bounded polling loop (with ADDRESS_TIMEOUT and ADDRESS_POLL_SECONDS) to wait for the ingress/service external address before writing the DNS record.
  • Reduce duplicated branches by unifying ingress vs service selection and IP vs hostname selection.
  • Document the new waiting behavior and configuration knobs in the README, and bump the script version to v0.0.25.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

File Description
readme.md Documents the new bounded wait for ingress/service external address and the related environment variables.
k8s-tools.sh Implements polling for the external address with configurable timeout/interval; bumps version string.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread k8s-tools.sh
Comment on lines +99 to +103
resource=$(kubectl --namespace=$namespace get $kind $target --output=json 2>/dev/null)
if [ -n "$resource" ]; then
dns_record_value=$(echo "$resource" | jq -r ".status.loadBalancer.ingress[0].$field // empty")
[ -n "$dns_record_value" ] && break
fi
@kierenj
kierenj merged commit 684bbcf into develop Aug 19, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants