Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/scripts/monitor_slurm_job.sh
Original file line number Diff line number Diff line change
Expand Up @@ -97,7 +97,7 @@ is_terminal_state() {

# Optionally bound how long a job may sit un-started in the queue. On the
# preemptible Phoenix 'embers' QOS a job routinely stays PENDING for hours and
# needs most of the job-level `timeout-minutes` (480m) window to backfill onto a
# needs most of the job-level `timeout-minutes` (1380m) window to backfill onto a
# free node; that job timeout is the real backstop. Default to 0 (wait
# indefinitely, up to the job timeout) so ordinary queue pressure does not turn
# otherwise-healthy jobs into red CI. Set SLURM_MAX_QUEUE_SECONDS>0 to opt into
Expand Down
2 changes: 1 addition & 1 deletion .github/scripts/run_parallel_benchmarks.sh
Original file line number Diff line number Diff line change
Expand Up @@ -86,7 +86,7 @@ else
# On Phoenix 'embers' a long benchmark job can be preempted (PreemptMode=CANCEL,
# so it is killed rather than requeued). On preemption (run_monitored exit 76)
# resubmit a fresh job in the same tree and re-monitor, bounded by
# MAX_PREEMPT_RESUBMITS (the 480m job timeout is the real backstop). Note: a
# MAX_PREEMPT_RESUBMITS (the 1380m job timeout is the real backstop). Note: a
# resubmitted job no longer overlaps its counterpart, slightly reducing
# same-load fairness -- still preferable to failing the run on an infra preempt.
: "${MAX_PREEMPT_RESUBMITS:=10}"
Expand Down
2 changes: 1 addition & 1 deletion .github/scripts/submit-slurm-job.sh
Original file line number Diff line number Diff line change
Expand Up @@ -266,7 +266,7 @@ EOT
# preempted job is killed outright and `--requeue` never restarts it. When the
# monitor reports preemption (exit 76), submit a fresh job and monitor again.
# Bounded by MAX_PREEMPT_RESUBMITS as a runaway guard; the job-level
# `timeout-minutes` (480m) remains the real backstop.
# `timeout-minutes` (1380m) remains the real backstop.
: "${MAX_PREEMPT_RESUBMITS:=10}"
# Node faults get a much tighter bound than preemption: preemption is routine on
# 'embers' and says nothing about the node, whereas hitting a second unusable
Expand Down
3 changes: 2 additions & 1 deletion .github/workflows/bench.yml
Original file line number Diff line number Diff line change
Expand Up @@ -103,7 +103,8 @@ jobs:
runs-on:
group: ${{ matrix.group }}
labels: ${{ matrix.labels }}
timeout-minutes: 480
# Includes SLURM queue wait (Phoenix embers is priority 0, often hours). Kept under 24h, when GITHUB_TOKEN expires.
timeout-minutes: 1380
steps:
- name: Clean stale output files
run: rm -f *.out
Expand Down
6 changes: 4 additions & 2 deletions .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -377,7 +377,8 @@ jobs:
# cpe/25.03 introduced an IPA SIGSEGV in CCE 19.0.0). Allow Frontier to
# fail without blocking PR merges; Phoenix remains a hard gate.
continue-on-error: ${{ matrix.runner == 'frontier' }}
timeout-minutes: 480
# Includes SLURM queue wait (Phoenix embers is priority 0, often hours). Kept under 24h, when GITHUB_TOKEN expires.
timeout-minutes: 1380
strategy:
matrix:
include:
Expand Down Expand Up @@ -555,7 +556,8 @@ jobs:
needs: [lint-gate, file-changes]
# Frontier is non-blocking for the same reason as the self job above.
continue-on-error: ${{ matrix.runner == 'frontier' }}
timeout-minutes: 480
# Includes SLURM queue wait (Phoenix embers is priority 0, often hours). Kept under 24h, when GITHUB_TOKEN expires.
timeout-minutes: 1380
strategy:
matrix:
include:
Expand Down
Loading