NCP-AIO Troubleshooting and Optimization Practice Question
A system administrator is troubleshooting a 'CUDA error: invalid device ordinal' when launching a job on a multi-GPU system. What is the most likely cause?
⚠ Common exam trap
Candidates frequently assume this error implies a faulty physical GPU or driver crash, overlooking the simpler and more common configuration error where environment variables restrict access to non-existent indices.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
CUDA_VISIBLE_DEVICES is set to an out-of-range index.
The 'invalid device ordinal' error typically indicates that the application is attempting to access a GPU index (e.g., GPU 4) that does not exist or is not visible to the process. This is often caused by environment variables like CUDA_VISIBLE_DEVICES being configured incorrectly, mapping the application to non-existent hardware. Ensuring the logical-to-physical GPU mapping is accurate is vital for correct job execution on shared infrastructure.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The GPU driver is corrupted and requires a reinstall.
Why it's wrong here
Corrupted drivers usually result in 'device not found' or kernel initialization failures rather than an 'invalid ordinal' error. An invalid ordinal refers to a specific integer index that is out of range, pointing to a configuration issue regarding how the application is selecting the hardware, not a driver malfunction.
- ✓
CUDA_VISIBLE_DEVICES is set to an out-of-range index.
Why this is correct
If the environment variable CUDA_VISIBLE_DEVICES specifies an index that does not exist on the host, the CUDA runtime will throw an 'invalid device ordinal' error. This is a configuration error where the system believes it has fewer GPUs than the application is trying to access via the environment variable.
- ✗
The GPU memory is full.
Why it's wrong here
Memory fullness results in OOM (Out of Memory) errors, not an 'invalid device ordinal' error. The system knows the device exists, but the allocation request fails. An invalid ordinal implies that the device selection itself is the point of failure, regardless of the memory state of any physical GPU.
- ✗
The InfiniBand fabric is misconfigured.
Why it's wrong here
InfiniBand issues affect communication latency and connectivity between nodes, not the local indexing of GPUs within the host system. CUDA device indexing is handled by the local driver and runtime, independent of the network fabric. Misconfigurations here would manifest as NCCL errors, not device ordinal errors at startup.
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.