NCP-AIO Workload Management • Set 2
NCP-AIO Workload Management Practice Test 2 — 15 questions with explanations. Free, no signup.
A platform engineer is deploying a multi-node large language model training job on a Slurm cluster managed by NVIDIA Base Command Manager. The job repeatedly fails with a 'node not responding' error, and the scheduler log shows that the job was allocated nodes that had been drained for maintenance but were not yet returned to service. Which Slurm configuration should the engineer verify to ensure the scheduler does not assign jobs to nodes in a drained state?
Choose an answer to begin — your selection is scored in the full session.
15 questions · instant feedback and full explanations after every question.