Skip to content

fix(tracking): reconcile tracking groups on runs that save no nodes - #1278

Open
ogenstad wants to merge 1 commit into
infrahub-developfrom
po-tracking-group-zero-member-reap
Open

ogenstad wants to merge 1 commit into
infrahub-developfrom
po-tracking-group-zero-member-reap

Conversation

@ogenstad

@ogenstad ogenstad commented Aug 25, 2026 •

Copy link
Copy Markdown
Contributor

Why

update_group() returned early whenever a run tracked zero members, so it never diffed the previous membership against the empty set. Any run that saved nothing left every previously tracked node behind as an orphan, still listed in the tracking group. This bites two ways in the field: a generator that legitimately produces nothing (a decommissioning run) never cleans up, and a repository whose last object file is removed leaves its objects stranded.

Removing the early return on its own would have turned that silent no-op into a run-killer. delete_unused() stopped at the first refused delete, and the group was saved before the reap, so a node whose delete was refused had already dropped out of the group and could never be retried.

Closes #572. Also fixes #737 (closed as a duplicate, code never changed) and is the SDK half of opsmill/infrahub#10134.

This is rebased onto the error catalogue (#1266) as a single commit and builds on its typed exceptions rather than on message text.

What changed

Behavioral changes:

  • With delete_unused_nodes=True, which generators and repository imports use, a run that tracks nothing now prunes the members of an existing tracking group instead of doing nothing.
  • A run that tracks nothing and has no existing group still creates no group, and an already-empty group is no longer re-upserted.
  • delete_unused() attempts every unused member instead of stopping at the first refusal. The members the server refused stay in the tracking group so a later run retries them, and update_group() reports them together as a new TrackingGroupCleanupError. A refusal that persists is raised again on every run until whatever blocks the deletion is removed; previously the member dropped out of the group after the first failure and stayed behind unnoticed.
  • The reap deletes members on the tracking context's branch. It used the client's default branch while the group lookup and upsert used the tracking branch, so on any other branch it deleted the wrong node or silently found nothing.
  • Both context-manager exits reset the client mode in a finally block. update_group() raising left the client in TRACKING mode, silently enrolling every later save into the stale context.

What a failed cleanup does

The reap separates failures that are about the member from failures that are about the request, because only the first kind is a fact about that member:

  • About the member - a server refusal such as a mandatory relationship (UNDEFINED_ERROR or no code), PERMISSION_DENIED, or a schema that no longer has the member's kind. It is recorded against that member and the remaining candidates are still attempted.
  • About the request - a timeout, an unreachable server, an expired or missing token, or a branch that is gone, already merged, needs a rebase, or is locked by a merge. It stops the reap and is raised as itself rather than recorded against every remaining member in turn.

After #1266 the GraphQL path raises a coded failure as a GraphQLError even when the code describes the request (TOKEN_EXPIRED, MERGE_IN_PROGRESS, ...), so the class alone cannot separate the two. The reap reads the catalogue code instead, and falls back to the class's declared CODE for a lookup miss the SDK raised without a server behind it.

Either way the group write is attempted before the failure surfaces, listing the nodes this run created plus the members the reap refused or never reached. An interrupted reap used to abort ahead of the upsert: members already in the group self-heal on the next run's diff, but the nodes the run had just created were in no group at all, and no later run could reach them. When the write fails as well, which it usually does for the same reason, the error that stopped the reap is raised with the write's failure as its cause.

The member whose delete was interrupted is kept in the group only when the server answered (an ApiError), since a rejected delete removed nothing. When no answer came back, such as on a timeout, the delete may have gone through, and the server rejects a group write that names a node that no longer exists, which would lose this run's membership. So that member is left out, matching what happened before this PR.

A member already removed by another member's cascade is tolerated through #1266's NodeNotFoundError. The message check for "Unable to find the node" is kept only for servers that predate the catalogue and report no code.

Also changed

  • delete_unused() returns a ReapResult (refused members, members never attempted, and the error that stopped the reap), importable from infrahub_sdk.query_groups, instead of raising. Nothing in the SDK or in Infrahub calls it directly, but it is public, so the change has a changed fragment of its own.
  • A context reused without finding a group clears what an earlier lookup recorded, so it has nothing to reap. The old early return covered this; the reap now runs on every call.
  • TrackingGroupCleanupError is exported from infrahub_sdk.exceptions and added to the public-names snapshot. It is not a GraphQLError, so a caller that caught a refused cleanup with except GraphQLError must catch it explicitly; the changed fragment says so.
  • TrackingGroupCleanupError survives pickle, so it survives the serialization a task orchestrator applies to a failed run.
  • infrahubctl renders TrackingGroupCleanupError as a table of each member and the server's reason, escaped the way feat: typed exceptions for the server's error catalogue #1266 escapes every other error message, instead of falling through to a traceback.
  • The member diff, candidate selection, failure classification and the decisions around the group write live in shared helpers and in ReapResult.failure(), so the async and sync contexts keep only their own calls.

What stayed the same: no change to when tracking is armed, to delete_unused_nodes defaults, or to the rollback-on-exception behavior.

Deliberately out of scope

  • With delete_unused_nodes=False the existing group is not read, so a run that tracks nothing still leaves it unchanged, while a run that tracks one node replaces the whole membership. Reading the group on that path would add a lookup to every default tracking run. The changelog states the condition.
  • The reap issues one sequential delete per unused member, and the zero-member fix makes that path reachable with an entire membership at once: a decommission of several thousand objects is several thousand sequential round trips. InfrahubBatch already has the right shape (return_exceptions=True) and is the natural follow-up.
  • Known edge: a delete that times out without having gone through leaves that one member out of the group, so it stays behind as an orphan, as it did before this PR. Covering it needs a lookup after the timeout, which is the request most likely to fail next.

How to review

Suggested order:

  1. infrahub_sdk/query_groups.py: the module-level classification helpers, then async update_group() and delete_unused(), then confirm the sync twin matches.
  2. tests/unit/sdk/test_group_context.py, then tests/integration/test_tracking_zero_members.py.

Worth extra scrutiny: _REQUEST_FAILURE_CODES, the list of catalogue codes treated as about the request. A code missing from it is recorded against each member; a code wrongly in it stops the cleanup early.

Also deliberate: raising rather than warning on a refused delete. A decommission that quietly fails to decommission seemed worse than a loud one. The consumer decides what that means in context: opsmill/infrahub#10134 catches it per import phase and continues, because the refused objects stay group members and the next import retries them.

Raised by local cubic reviews and declined:

  • PERMISSION_DENIED is counted as about the member. Object permissions are evaluated per kind and branch, so the denial is accurate for each member of that kind, and the members stay in the group either way.
  • A rate limit stops the reap without being in _REQUEST_FAILURE_CODES: it arrives as RateLimitError, which is not a GraphQLError.
  • When a request failure stops the reap after some refusals, the raised error is the request failure alone. The refused members stay in the group, and the next run that gets through reports them; attaching them would need add_note, which requires Python 3.11.
  • ReapResult is not exported from the package root, matching InfrahubGroupContext, which returns it and is not exported either.
  • The async and sync update_group() still repeat their awaited calls and the four lines handling a failed write, whose bare raise has to stay inside the except.

How to test

uv run pytest tests/unit/sdk/test_group_context.py tests/unit/ctl/test_utils.py tests/unit/sdk/test_exceptions_public_names.py
uv run pytest tests/integration/test_tracking_zero_members.py

The unit tests need no Docker daemon. They drive the reap end to end through httpx_mock with #1266's recorded catalogue responses, covering for both clients:

  • a zero-member run with no group, and one with an already-empty group, each issue a single request and no mutation
  • a coded NODE_NOT_FOUND during the reap drops the member without a failure
  • a MERGE_IN_PROGRESS during the reap stops it, keeps every member in the group, and raises MergeInProgressError rather than a cleanup error
  • a timeout during the reap still writes the group, and when the write fails as well, the error that stopped the reap is the one raised
  • the member deleted before an interruption is dropped and the one never reached is kept

Regression guards checked against the pushed code: classifying every GraphQLError as a refusal, the rule before #1266 was taken into account, fails the classification cases, the locked-branch test and the failed-write test for both clients.

The integration tests need a Docker daemon and were not run locally; CI is their first run.

@codecov

codecov Bot commented Aug 25, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.47887% with 5 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
infrahub_sdk/query_groups.py 95.65% 2 Missing and 3 partials ⚠️
@@                 Coverage Diff                  @@
##           infrahub-develop    #1278      +/-   ##
====================================================
- Coverage             87.00%   86.96%   -0.05%     
====================================================
  Files                   153      152       -1     
  Lines                 16293    14805    -1488     
  Branches               2348     1989     -359     
====================================================
- Hits                  14176    12875    -1301     
+ Misses                 1494     1369     -125     
+ Partials                623      561      -62     
Flag Coverage Δ
integration-tests 42.38% <57.74%> (-0.69%) ⬇️
python-3.10 61.53% <72.53%> (-0.31%) ⬇️
python-3.11 61.54% <72.53%> (-0.30%) ⬇️
python-3.12 61.53% <72.53%> (-0.31%) ⬇️
python-3.13 61.54% <72.53%> (-0.30%) ⬇️
python-3.14 61.54% <72.53%> (-0.30%) ⬇️
python-filler-3.12 22.95% <18.30%> (+0.10%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
infrahub_sdk/client.py 85.33% <100.00%> (-1.43%) ⬇️
infrahub_sdk/ctl/utils.py 76.27% <100.00%> (+1.05%) ⬆️
infrahub_sdk/exceptions/__init__.py 100.00% <ø> (ø)
infrahub_sdk/exceptions/base.py 94.88% <100.00%> (+0.17%) ⬆️
infrahub_sdk/query_groups.py 93.86% <95.65%> (+8.61%) ⬆️

... and 13 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 4 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread infrahub_sdk/query_groups.py Outdated
Comment thread infrahub_sdk/query_groups.py
Comment thread infrahub_sdk/query_groups.py Outdated
Comment thread infrahub_sdk/query_groups.py Outdated
@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Aug 27, 2026 •

Copy link
Copy Markdown

Deploying infrahub-sdk-python with  Cloudflare Pages  Cloudflare Pages

Latest commit: 9d001e5
Status: ✅  Deploy successful!
Preview URL: https://76cfc140.infrahub-sdk-python.pages.dev
Branch Preview URL: https://po-tracking-group-zero-membe.infrahub-sdk-python.pages.dev

View logs

@ogenstad
ogenstad force-pushed the po-tracking-group-zero-member-reap branch from ae12445 to 1dc8bad Compare August 31, 2026 12:53

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 1 file (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

Comment thread infrahub_sdk/query_groups.py Outdated
@ogenstad
ogenstad force-pushed the po-tracking-group-zero-member-reap branch from aae78f5 to 1cbc949 Compare September 10, 2026 13:11

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 7 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="infrahub_sdk/ctl/utils.py">

<violation number="1" location="infrahub_sdk/ctl/utils.py:72">
P2: Custom agent: **Flag AI Slop and Fabricated Changes**

Add a CLI regression test for the new `TrackingGroupCleanupError` path, asserting that `handle_exception` renders every failed node and reason in the table and exits with the requested code; neither this branch nor `print_tracking_group_failures()` is currently exercised.</violation>
</file>

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread infrahub_sdk/ctl/utils.py
if isinstance(exc, GraphQLError):
print_graphql_errors(console=console, errors=exc.errors)
raise typer.Exit(code=exit_code)
if isinstance(exc, TrackingGroupCleanupError):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Custom agent: Flag AI Slop and Fabricated Changes

Add a CLI regression test for the new TrackingGroupCleanupError path, asserting that handle_exception renders every failed node and reason in the table and exits with the requested code; neither this branch nor print_tracking_group_failures() is currently exercised.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At infrahub_sdk/ctl/utils.py, line 72:

<comment>Add a CLI regression test for the new `TrackingGroupCleanupError` path, asserting that `handle_exception` renders every failed node and reason in the table and exits with the requested code; neither this branch nor `print_tracking_group_failures()` is currently exercised.</comment>

<file context>
@@ -67,6 +69,9 @@ def handle_exception(exc: Exception, console: Console, exit_code: int) -> NoRetu
     if isinstance(exc, GraphQLError):
         print_graphql_errors(console=console, errors=exc.errors)
         raise typer.Exit(code=exit_code)
+    if isinstance(exc, TrackingGroupCleanupError):
+        print_tracking_group_failures(console=console, failures=exc.failures)
+        raise typer.Exit(code=exit_code)
</file context>

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in d33fdcd: test_a_tracking_group_cleanup_failure_lists_every_member_and_reason in tests/unit/ctl/test_utils.py asserts every node id and reason, the exit code, and that a bracketed reason is not read as markup.

This reply was written by an AI assistant (Claude Code).

Comment thread tests/integration/test_tracking_zero_members.py Outdated
@ogenstad
ogenstad force-pushed the po-tracking-group-zero-member-reap branch from 1cbc949 to d33fdcd Compare October 6, 2026 11:38
@ogenstad
ogenstad force-pushed the po-tracking-group-zero-member-reap branch from d33fdcd to a42e040 Compare October 7, 2026 07:27
update_group() returned early whenever a run tracked zero members, so it never
diffed the previous membership against the empty set. A generator that
legitimately produced nothing, or a repository whose last object file was
removed, left every previously tracked node behind as an orphan. With
delete_unused_nodes=True a run that tracks nothing now prunes an existing
group; one with no group creates none, and an already-empty group is not
re-saved. A context reused without finding a group has nothing to reap.

The reap deletes on the tracking context's branch rather than the client's
default branch, and both context-manager exits reset the client mode in a
finally block so a raising update_group() cannot leave the client tracking.

delete_unused() attempts every candidate and returns a ReapResult instead of
raising. A failure about the member is recorded against it and the rest are
still attempted; the refused members stay in the group so a later run retries
them, and update_group() reports them together as TrackingGroupCleanupError,
again on every run until whatever blocks the deletion is removed. A failure
about the request stops the reap and is re-raised as itself rather than
recorded against each remaining member. That includes a timeout and an
unreachable server, and also the coded failures the GraphQL path raises as a
GraphQLError: an expired or missing token, and a branch that is gone, merged,
needs a rebase or is locked by a merge.

The group write is attempted first, listing this run's nodes plus the refused
members and the ones never reached, so an interrupted reap cannot leave the
nodes the run created in no group at all. The member whose delete was
interrupted is kept only when the server answered: without an answer the
delete may have gone through, and a write naming a node that no longer exists
is rejected. When the write fails too, the error that stopped the reap is
raised with the write's failure as its cause.

A member already removed by another member's cascade is tolerated through the
catalogue's NodeNotFoundError, with the legacy message check kept only for
servers that report no code. The async and sync contexts share the decisions
around the reap and the write, keeping only their own calls.

TrackingGroupCleanupError is exported from infrahub_sdk.exceptions, survives
pickling, and infrahubctl renders it as a table of each member and the
server's reason rather than a traceback. The changed contract of
delete_unused() and the new exception type get a changed fragment of their
own.
@ogenstad
ogenstad force-pushed the po-tracking-group-zero-member-reap branch from a42e040 to 9d001e5 Compare October 7, 2026 15:47
@ogenstad
ogenstad marked this pull request as ready for review October 7, 2026 19:50
@ogenstad
ogenstad requested a review from a team as a code owner October 7, 2026 19:50

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 issues found across 11 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="infrahub_sdk/query_groups.py">

<violation number="1" location="infrahub_sdk/query_groups.py:66">
P2: `RateLimitError` means the server rejected the delete, but it is not an `ApiError`, so this drops the still-existing member from the group and prevents a later retry. Retain the interrupted member for known rejected requests such as HTTP 429.</violation>
</file>

<file name="changelog/572.changed.md">

<violation number="1" location="changelog/572.changed.md:3">
P2: A group-write failure after a member refusal propagates instead of `TrackingGroupCleanupError`, so this guarantee is unconditional where the implementation is not. Qualify it on the group write succeeding.</violation>
</file>

Reply with feedback, questions, or to request a fix.

View guided diff | Turn on auto-fix | Re-trigger cubic

"""
if not isinstance(exc, GraphQLError) or not _about_the_member(exc):
result.error = exc
first_kept = position if isinstance(exc, ApiError) else position + 1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: RateLimitError means the server rejected the delete, but it is not an ApiError, so this drops the still-existing member from the group and prevents a later retry. Retain the interrupted member for known rejected requests such as HTTP 429.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At infrahub_sdk/query_groups.py, line 66:

<comment>`RateLimitError` means the server rejected the delete, but it is not an `ApiError`, so this drops the still-existing member from the group and prevents a later retry. Retain the interrupted member for known rejected requests such as HTTP 429.</comment>

<file context>
@@ -26,6 +41,73 @@ def _node_already_deleted(exc: GraphQLError) -> bool:
+    """
+    if not isinstance(exc, GraphQLError) or not _about_the_member(exc):
+        result.error = exc
+        first_kept = position if isinstance(exc, ApiError) else position + 1
+        result.unattempted = [candidate_id for _, candidate_id in candidates[first_kept:]]
+        return True
</file context>

Comment thread changelog/572.changed.md
@@ -0,0 +1,3 @@
`InfrahubGroupContext.delete_unused()` and its sync counterpart no longer raise. They attempt every unused member and return a `ReapResult`, importable from `infrahub_sdk.query_groups`, that lists the members the server refused to delete, the members the cleanup never reached, and the failure that stopped it, if any. A caller that invoked `delete_unused()` directly and relied on it raising must now inspect that result. `update_group()` and the tracking context manager still raise, as described below.

Leaving a tracking context whose cleanup was refused now raises `TrackingGroupCleanupError` once every member has been attempted, in place of the first `GraphQLError`. `TrackingGroupCleanupError` is not a `GraphQLError`, so a caller that caught the cleanup failure with `except GraphQLError` must catch it explicitly. A cleanup stopped by a failure that is not about a member, such as a timeout, still raises that failure itself.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: A group-write failure after a member refusal propagates instead of TrackingGroupCleanupError, so this guarantee is unconditional where the implementation is not. Qualify it on the group write succeeding.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At changelog/572.changed.md, line 3:

<comment>A group-write failure after a member refusal propagates instead of `TrackingGroupCleanupError`, so this guarantee is unconditional where the implementation is not. Qualify it on the group write succeeding.</comment>

<file context>
@@ -0,0 +1,3 @@
+`InfrahubGroupContext.delete_unused()` and its sync counterpart no longer raise. They attempt every unused member and return a `ReapResult`, importable from `infrahub_sdk.query_groups`, that lists the members the server refused to delete, the members the cleanup never reached, and the failure that stopped it, if any. A caller that invoked `delete_unused()` directly and relied on it raising must now inspect that result. `update_group()` and the tracking context manager still raise, as described below.
+
+Leaving a tracking context whose cleanup was refused now raises `TrackingGroupCleanupError` once every member has been attempted, in place of the first `GraphQLError`. `TrackingGroupCleanupError` is not a `GraphQLError`, so a caller that caught the cleanup failure with `except GraphQLError` must catch it explicitly. A cleanup stopped by a failure that is not about a member, such as a timeout, still raises that failure itself.
</file context>
Suggested change
Leaving a tracking context whose cleanup was refused now raises `TrackingGroupCleanupError` once every member has been attempted, in place of the first `GraphQLError`. `TrackingGroupCleanupError` is not a `GraphQLError`, so a caller that caught the cleanup failure with `except GraphQLError` must catch it explicitly. A cleanup stopped by a failure that is not about a member, such as a timeout, still raises that failure itself.
After a successful group write, leaving a tracking context whose cleanup was refused now raises `TrackingGroupCleanupError` once every member has been attempted, in place of the first `GraphQLError`. `TrackingGroupCleanupError` is not a `GraphQLError`, so a caller that caught the cleanup failure with `except GraphQLError` must catch it explicitly. A cleanup stopped by a failure that is not about a member, such as a timeout, still raises that failure itself.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant