NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A team is deploying a large language model for real-time inference on an NVIDIA GPU. They observe that the first few inference requests have high latency, but subsequent requests are much faster. What is the most likely explanation for this behavior?
⚠ Common exam trap
The trap here is attributing the latency to hardware warm-up or model loading, which are one-time startup costs, rather than to runtime kernel compilation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
CUDA kernels are being compiled and cached
The initial high latency is due to just-in-time compilation of CUDA kernels, which occurs on the first execution of each operation. Once compiled, the kernels are cached for reuse, making subsequent inferences faster. This is a well-known behavior in deep learning frameworks and explains the warm-up effect.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The model is being quantized on the fly
Why it's wrong here
Quantization is typically a one-time offline process, not something that happens dynamically per request. If quantization were applied on the fly, it would consistently add overhead to every request, not just the first few. Thus, this does not explain the initial high latency followed by faster responses.
- ✗
The GPU is warming up its clock speed
Why it's wrong here
GPU clock speed warm-up can cause slight performance variations, but modern GPUs ramp up quickly, usually within milliseconds. The latency difference here is described as significant and affecting only the first few requests, which is more characteristic of software initialization than hardware warm-up.
- ✓
CUDA kernels are being compiled and cached
Why this is correct
During the first inference, CUDA kernels for operations like matrix multiplications are compiled and cached. This just-in-time compilation adds latency. Subsequent requests reuse the cached kernels, avoiding recompilation and resulting in faster execution. This is a common behavior in frameworks like PyTorch and TensorRT.
- ✗
The model weights are being loaded from disk
Why it's wrong here
Model weights are typically loaded into GPU memory once at startup, before any inference requests. If they were loaded from disk per request, every request would incur the same I/O latency, not just the first few. Thus, this does not match the observed pattern.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.