Conversation
getOrCreate has two paths that each wait for a session to become usable, and their budgets had drifted apart. Creating a new session polls the long-running operation for up to 600s. Attaching to a session that already exists cannot use that machinery -- there is no operation handle -- so it hand-rolls a GetSession poll loop waiting for the Spark Connect endpoint to appear, and that loop allowed only 300s. Both are waiting on the same thing: a server finishing its boot. Since get_active_s8s_session_response treats CREATING as reusable, the attach path can land on a session that has not started yet and then give up at half the budget the create path would have allowed. Raise it to 600s so a caller gets the same 10 minutes either way.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
_wait_for_session_availablenow defaults to a 600s budget instead of 300s.Why
getOrCreatehas two paths that each wait for a session to become usable, and their budgets had drifted apart:__create) polls the long-running operation returned bycreate_session, withtimeout=600on the polling retry._get_exiting_active_session) has no operation handle to poll, so it hand-rolls aGetSessionloop waiting for the"Spark Connect Server"key to appear inruntime_info.endpoints— and that loop allowed only 300s.Both are waiting on the same underlying thing: a server finishing its boot. Because
get_active_s8s_session_responsetreatsCREATINGas reusable, the attach path can land on a session that hasn't started yet and then give up at half the budget the create path would have allowed. The most common way to hit this is passing a custom session ID that resolves to a session someone else just started.Raising the default to 600s means a caller gets the same 10 minutes either way.
Not addressed here
The loop only evaluates its deadline between iterations, and
get_sessionis issued with no per-call timeout (the Dataproc GAPIC setsdefault_timeout=Nonefor every method). So the wall-clock can still overshoot the stated budget by the duration of a stalled final call — it will now read "after 600 seconds" while taking longer. Bounding that properly means threading the remaining budget into theget_sessioncall; left for a separate change._wait_for_terminationhas the same shape with a 180s budget.Testing
pytest tests/unit/test_session.py -k wait_for_session_availablepasses (both tests pass an explicit timeout, so neither depended on the default).pyink --checkclean.