NCA-GENL Core Machine Learning and AI Knowledge Practice Question
Exhibit
Error: CUDA error: device-side assert triggered. Traceback: ... in forward_pass ... loss = criterion(logits, targets) ...
Refer to the exhibit. This error occurs during the training of an LLM. What is the most likely cause for this 'device-side assert' error?
⚠ Common exam trap
Candidates frequently assume the error is a hardware failure or a corrupted model file, overlooking that device-side asserts in PyTorch are almost always caused by index-out-of-bounds errors on GPU kernels.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The target indices exceed the vocabulary size.
A device-side assert error in PyTorch usually indicates an index-out-of-bounds error or a shape mismatch during a kernel operation on the GPU. When the loss function expects a specific range of indices for targets (e.g., in a cross-entropy loss), providing values that exceed the vocab size triggers this assertion. This is a common debugging hurdle that highlights the importance of rigorous input data validation before submitting tensors to the GPU's highly optimized, non-interactive kernels.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The model has run out of VRAM for the current batch.
Why it's wrong here
Out-of-memory errors in PyTorch typically return an explicit 'CUDA out of memory' message. A device-side assert is a specific check triggered by a kernel failing a logical condition, such as an index being out of bounds, rather than an allocation failure caused by insufficient device memory.
- ✓
The target indices exceed the vocabulary size.
Why this is correct
Loss functions like cross-entropy verify that label indices fall within the range [0, vocab_size - 1]. If an index in the target tensor is out of this range, the underlying CUDA kernel triggers an assertion to prevent invalid memory access, resulting in the reported error during the training step.
- ✗
The GPU has overheated and shut down.
Why it's wrong here
GPU overheating leads to system-level stability issues, driver crashes, or kernel resets, not specific 'device-side assert' errors within a PyTorch training loop. Those errors are logical constraints enforced by the GPU kernel code to prevent illegal operations, and they reflect a software logic error rather than a hardware failure.
- ✗
The learning rate is set too high for convergence.
Why it's wrong here
An overly high learning rate typically causes loss values to explode (NaNs) or oscillate wildly, not trigger a device-side assertion. Assertions are structural checks on data boundaries or tensor shapes; they indicate a programmatic violation of kernel requirements, which is distinct from issues related to gradient descent optimization hyperparameters.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.