slurm: make run-parts.sh exclusive detection work with custom prefix and recent Slurm - #1391
slurm: make run-parts.sh exclusive detection work with custom prefix and recent Slurm#1391100milliongold wants to merge 2 commits into
Conversation
…and recent Slurm
The exclusive-job check in run-parts.sh had two independent failures:
1. It called scontrol/squeue through PATH. slurmd's environment does not
include a custom slurm_install_prefix, so both commands produced empty
output, numcpus_sys and numcpus_job were both "", the comparison was
true, and every job ran the *-exclusive-* prolog/epilog scripts. On a
shared node this reset power limits and clocks on all GPUs and dropped
page caches for every job.
2. It parsed "scontrol show job" with grep -Eio "TRES=cpu=[0-9]+". On
recent Slurm the output has both ReqTRES= and AllocTRES= lines, so the
pattern matched twice and numcpus_job became a multi-line value that
never compared equal. With scontrol on PATH, exclusive jobs were
therefore never detected.
Use {{ slurm_install_prefix }}/bin/squeue by absolute path (the file is
already deployed via the template module) and read allocated CPUs and node
count with -o %C / -o %D instead of parsing scontrol. Guard against an
empty result so a lookup failure means "not exclusive" rather than
"exclusive".
Observed on DGX OS 7.5.0, Slurm 26.05.1, slurm_install_prefix=/raid/slurm/usr/local:
- before the binaries were symlinked into /usr/local/bin, every srun
--gres=gpu:1 job logged "Running .../50-exclusive-gpu" and prolog took
6-8 s (srun: Prolog hung on node);
- after symlinking, a bash -x run of the script showed numcpus_job='1<nl>56'.
Signed-off-by: Jea-Eok-Kim <je.kim@xiilab.com>
dholt
left a comment
There was a problem hiding this comment.
run-parts.sh has set -e, so the direct numcpus_job=$(squeue ...) and numnodes_job=$(squeue ...) assignments terminate the script when squeue returns nonzero; redirecting stderr does not suppress that exit status. Please perform both lookups inside a conditional or otherwise explicitly handle failure so exclusive remains 0 and normal non-exclusive scripts still run. Add the required decision-matrix proof for failed squeue as well as empty output.
Automated triage review (agent-generated on the maintainer's behalf; a human maintainer decides merges).
This script runs with "set -e", so a bare `numcpus_job=$(squeue ...)` assignment aborts the entire prolog/epilog run when squeue exits nonzero. Redirecting stderr does not suppress the exit status. The result is that a transient squeue failure skips every part script, not just the exclusive ones, and the job fails. Both allocation fields are now fetched by a single `-o "%C %D"` call inside an `if` condition, so the failure is visible and handled. A failed, empty or non-numeric lookup leaves exclusive=0 and logs a warning. The same reasoning is applied to the last-user-job count, with one difference: piping squeue into `wc -l` hides its exit status, and a failed lookup would be read as "no other jobs" and run the *-lastuserjob-* cleanup scripts while another job of the same user is still on the node. A failed lookup now leaves last_user_job=0. Decision matrix, verified on a DGX B300 (Ubuntu 24.04, bash 5.2, 256 CPUs) by substituting a squeue stub that honours the -o format and the -j flag: case before after ---------- ------------------------------ ------------------------------ exclusive rc=0 all three parts ran rc=0 all three parts ran shared rc=0 exclusive part skipped rc=0 exclusive part skipped otherjobs rc=0 lastuserjob part skipped rc=0 lastuserjob part skipped empty rc=0 silently non-exclusive rc=0 non-exclusive + 1 warning nonnumeric rc=0 silently non-exclusive rc=0 non-exclusive + 1 warning fail rc=1 NO part ran at all rc=0 normal part ran + 2 warnings The three normal cases are unchanged, so there is no regression; the failure cases stop taking the whole run down and stop deciding silently.
|
Thanks — the The What changed (commit
Decision matrix, verified on a DGX B300 (Ubuntu 24.04, bash 5.2, 256 CPUs) with a
The three normal cases are unchanged, so there is no regression; the failure cases stop taking the whole run down and stop deciding silently. One note on the test itself: my first stub ignored |
Problem
The exclusive-job detection in
roles/slurm/templates/etc/slurm/shared/bin/run-parts.shfails in two independent ways:PATH. It calls
scontrol/squeuethrough PATH. slurmd's environment does not include a customslurm_install_prefix, so both commands return nothing,numcpus_sysandnumcpus_jobare both empty,"" == ""is true, and every job runs the*-exclusive-*scripts. On a shared node that resets power limits and application clocks on all GPUs and drops page caches for each job (prolog took 6–8 s;srun: Prolog hung on node).Parsing.
grep -Eio "TRES=cpu=[0-9]+"matches bothReqTRES=andAllocTRES=lines on recent Slurm, sonumcpus_jobbecomes a multi-line value ('1\n56'in abash -xtrace) that never compares equal. Withscontrolon PATH, exclusive jobs are therefore never detected.Reproduced on DGX OS 7.5.0, Slurm 26.05.1,
slurm_install_prefix: /raid/slurm/usr/local.Fix
{{ slurm_install_prefix }}/bin/squeueby absolute path (the script is deployed with thetemplatemodule, so the variable is available).squeue -o %C/-o %Dinstead of parsingscontrol show job.Verification
On the system above, before the fix every
srun --gres=gpu:1job loggedRunning .../50-exclusive-gpu; after symlinking the binaries onto PATH (which exercises failure 2) abash -xrun showednumcpus_job='1\n56'. The patched logic yieldsnumcpus_job=56,numcpus_sys=256,exclusive=0for that job.