Courseiva

NCP-GENL GPU Acceleration and Optimization Practice Question

Exhibit

config.pbtxt: { name: 'llm_model', platform: 'tensorrt_plan', instance_group: [{ count: 1, kind: KIND_GPU, gpus: [0] }], dynamic_batching: { preferred_batch_size: [4, 8] } }

Refer to the exhibit. The model performance is inconsistent. What is the most likely reason for the performance variability under load?

⚠ Common exam trap

Candidates frequently assume the issue is related to GPU compute capacity or model size, overlooking the configuration-level settings that govern how requests are queued and processed in a production server.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dynamic batching lacks a max_queue_delay_microseconds parameter.

The configuration uses dynamic batching with a preferred batch size range but lacks a 'max_queue_delay_microseconds' setting. Without a delay buffer, the server might dispatch batches as soon as one request arrives, leading to suboptimal under-filled batches and inconsistent latency. Adding a small delay allows the server to collect more requests, increasing throughput and stabilizing inference latency across fluctuating traffic patterns, which is vital for high-performance production deployments.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The instance count is too high for the GPU.

    Why it's wrong here

    The instance count is set to 1, which is the minimum value. Increasing this count would likely cause OOM errors or resource contention. The bottleneck is not over-subscription of the GPU but rather poor utilization of the batching window, which does not relate to the single instance count configuration provided.

  • ✓

    Dynamic batching lacks a max_queue_delay_microseconds parameter.

    Why this is correct

    Without a max_queue_delay_microseconds, the inference server does not wait for additional requests to fill the batch. This results in requests being processed immediately, often with small batch sizes, leading to inconsistent execution times and poor GPU utilization compared to waiting briefly to aggregate requests into larger, more efficient batches.

  • ✗

    The GPU index is incorrectly defined in the config.

    Why it's wrong here

    The GPU index 0 is the default and standard reference for the primary GPU in most systems. Unless multiple GPUs are present and the workload requires a different device, this setting is correct. The performance variability is almost certainly due to the batching logic rather than the device hardware mapping.

  • ✗

    TensorRT plan files are inherently non-deterministic.

    Why it's wrong here

    TensorRT plan files are highly deterministic once built. The variation in performance is driven by the dynamic workload environment, not the engine itself. If the engine provided consistent output for the same input, it would always perform identically regardless of the external factors, which is not the case here.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.