Skip to content

bug: wait_for_completion() never checks the task when interval > timeout, and overshoots the timeout #1311

Description

@BeArchiTek

Component

Python SDK

Infrahub SDK version

1.23.1 (verified on develop @ 9b39ab4)

Current Behavior

wait_for_completion() derives its loop count with integer division, so when interval exceeds timeout it never checks the task at all.

infrahub_sdk/task/manager.py, sync at :587 (async equivalent at :261):

for _ in range(timeout // interval):
    task = self.get(id=id)
    if task.state in FINAL_STATES:
        return task
    time.sleep(interval)
raise TaskNotCompletedError(id=id, message=f"Task {id} did not complete in {timeout} seconds")

Two consequences:

  1. interval > timeout performs zero checks. wait_for_completion(id, interval=90, timeout=60) evaluates range(0), skips the body entirely, and raises TaskNotCompletedError immediately — reporting that the task "did not complete in 60 seconds" without a single request having been made, and without any time having passed. A task that was already finished before the call is reported as incomplete.

  2. The wait overshoots timeout. The sleep happens after the check, so the last iteration sleeps for a further interval before the loop ends. With the defaults (interval=1, timeout=60) the call takes roughly 60 s of sleeping plus 60 round trips, exceeding the nominal timeout. Any interval that does not divide timeout also truncates: interval=7, timeout=60 gives 8 iterations, so it gives up at ~56 s rather than 60.

The message is misleading in every one of these cases, since it quotes timeout rather than the time actually spent.

Expected Behavior

  • A call always checks the task's state at least once, regardless of the interval/timeout relationship. An already-final task is returned rather than reported incomplete.
  • The wait is bounded by wall-clock elapsed time against timeout, not by a precomputed iteration count — for example, loop while elapsed < timeout, checking first and sleeping only if there is remaining budget.
  • The error names the elapsed time (or states the timeout honestly), so the message matches what happened.

Steps to Reproduce

client = InfrahubClientSync(address="http://localhost:8000")

# a task that has already finished
task_id = "<id of a completed task>"

client.task.wait_for_completion(id=task_id, interval=90, timeout=60)
# raises TaskNotCompletedError immediately — returns in milliseconds,
# never issues a request, and the task is in fact complete

For the overshoot, time a default call against a task that never settles: it sleeps ~60 s and issues 60 requests before raising.

Additional Information

Found while auditing which SDK capabilities the opsmill.infrahub Ansible collection does not yet expose. The collection needs wait_for_completion to determine whether triggered server-side work actually succeeded (opsmill/infrahub-ansible#398, and the artifact defect in opsmill/infrahub-ansible#404), so the timeout semantics matter to what an Ansible task can honestly report.

Metadata

Metadata

Assignees

No one assigned

    Labels

    effort/lowThis issue should be completed in a couple of hoursstate/need-triageThis issue needs to be triagedtype/bugSomething isn't working as expected

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions