CI: raise self-hosted job timeout to 1380m to absorb Phoenix embers queue wait - #1942
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The limited timeout changes match the stated purpose, preserve SLURM walltimes, and have no identified blocking issues.
Review effort: Balanced
Findings: None
What changed in this PR
Extends self-hosted CI timeouts to tolerate long Phoenix SLURM queue waits.
Changes:
- Raises Test Suite, Case Opt, and Benchmark job timeouts from 480 to 1380 minutes.
- Updates related script comments; SLURM walltimes and script behavior remain unchanged.
| File | Description |
|---|---|
.github/workflows/test.yml |
Extends Test Suite and Case Opt timeouts. |
.github/workflows/bench.yml |
Extends Benchmark timeout. |
.github/scripts/submit-slurm-job.sh |
Updates timeout reference in a comment. |
.github/scripts/run_parallel_benchmarks.sh |
Updates timeout reference in a comment. |
.github/scripts/monitor_slurm_job.sh |
Updates queue-wait timeout commentary. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #1942 +/- ##
=======================================
Coverage 62.80% 62.80%
=======================================
Files 86 86
Lines 22385 22385
Branches 3304 3304
=======================================
Hits 14060 14060
Misses 6073 6073
Partials 2252 2252 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Phoenix GPU jobs run under QOS
embers(priority weight 0). With ~475infernojobs pending at ~231k priority, ours (priority ~20) only start on leftover GPUs, and queue waits of ~7h are now routine.timeout-minutescounts that wait, so jobs get killed shortly after they finally start. Example: carbon-surface-v1 bench gpu-acc (SLURM 13821471) sat PENDING 03:22Z→10:13Z, ran ~1h, then the 480m GitHub timeout cancelled it. The runner'sscancelcleanup is why these show up asCANCELLED by <uid>in sacct.This raises the job-level timeout for the self-hosted Test Suite, Case Opt, and Benchmark jobs from 480m to 1380m (23h). SLURM walltime limits are unchanged (3–4h); only the queue wait we tolerate grows. It stays under 24h because GitHub fails jobs that wait >24h for a runner and
GITHUB_TOKENexpires at 24h.Tradeoffs, to revisit if they bite:
Experimental; revert if runner starvation gets worse than the timeouts it removes.
Acknowledgement