NCP-GENL GPU Acceleration and Optimization Practice Question
Exhibit
config.pbtxt: { name: 'llm_model', platform: 'tensorrt_plan', instance_group: [{ count: 1, kind: KIND_GPU, gpus: [0] }], dynamic_batching: { preferred_batch_size: [4, 8] } }Refer to the exhibit. The model performance is inconsistent. What is the most likely reason for the performance variability under load?
⚠ Common exam trap
Candidates frequently assume the issue is related to GPU compute capacity or model size, overlooking the configuration-level settings that govern how requests are queued and processed in a production server.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dynamic batching lacks a max_queue_delay_microseconds parameter.
The configuration uses dynamic batching with a preferred batch size range but lacks a 'max_queue_delay_microseconds' setting. Without a delay buffer, the server might dispatch batches as soon as one request arrives, leading to suboptimal under-filled batches and inconsistent latency. Adding a small delay allows the server to collect more requests, increasing throughput and stabilizing inference latency across fluctuating traffic patterns, which is vital for high-performance production deployments.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The instance count is too high for the GPU.
Why it's wrong here
The instance count is set to 1, which is the minimum value. Increasing this count would likely cause OOM errors or resource contention. The bottleneck is not over-subscription of the GPU but rather poor utilization of the batching window, which does not relate to the single instance count configuration provided.
- ✓
Dynamic batching lacks a max_queue_delay_microseconds parameter.
Why this is correct
Without a max_queue_delay_microseconds, the inference server does not wait for additional requests to fill the batch. This results in requests being processed immediately, often with small batch sizes, leading to inconsistent execution times and poor GPU utilization compared to waiting briefly to aggregate requests into larger, more efficient batches.
- ✗
The GPU index is incorrectly defined in the config.
Why it's wrong here
The GPU index 0 is the default and standard reference for the primary GPU in most systems. Unless multiple GPUs are present and the workload requires a different device, this setting is correct. The performance variability is almost certainly due to the batching logic rather than the device hardware mapping.
- ✗
TensorRT plan files are inherently non-deterministic.
Why it's wrong here
TensorRT plan files are highly deterministic once built. The variation in performance is driven by the dynamic workload environment, not the engine itself. If the engine provided consistent output for the same input, it would always perform identically regardless of the external factors, which is not the case here.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.