NCP-GENL Model Optimization Practice Question
Exhibit
config.pbtxt:
instance_group [
{
count: 2
kind: KIND_GPU
gpus: [0]
}
]Refer to the exhibit. An engineer observes that GPU memory utilization is high, but the GPU is frequently idling. How does the provided Triton configuration optimize the inference pipeline?
⚠ Common exam trap
Candidates often misinterpret high VRAM utilization with optimal performance, missing the fact that unutilized compute cycles indicate poor concurrency and request scheduling bottlenecks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
It enables concurrent execution on the GPU
By setting the instance count to 2 on a single GPU, the configuration enables concurrent model execution. This allows Triton to schedule multiple inference requests to be processed simultaneously on the same hardware. This overlap helps to hide memory latency and pipeline stalls, effectively increasing the utilization of the GPU and raising the overall throughput of the deployment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
It forces the model to use INT8 precision
Why it's wrong here
The provided configuration file defines the instance group properties for the model, such as hardware affinity and concurrency. It does not dictate the numerical precision of the model, which is determined by the underlying TensorRT engine file or the model framework configuration, not the Triton instance settings.
- ✓
It enables concurrent execution on the GPU
Why this is correct
Setting 'count: 2' instructs Triton to launch two instances of the model on the specified GPU. This allows the server to process multiple requests in parallel, which is a standard method to improve throughput and keep the GPU busy during periods where one instance might be blocked.
- ✗
It partitions the GPU for dedicated memory
Why it's wrong here
This configuration does not partition the GPU memory at the hardware or driver level. It simply defines how many model instances Triton will manage. Memory partitioning for isolation would typically require Multi-Instance GPU (MIG) settings, which are configured at the hardware or hypervisor layer, not in Triton's model configuration.
- ✗
It enables dynamic batching for the requests
Why it's wrong here
Dynamic batching requires a separate 'dynamic_batching' block in the configuration file to define parameters like 'preferred_batch_size'. The 'instance_group' settings define execution concurrency, which is a distinct concept from batching, where requests are grouped together to maximize the efficiency of single-instance model runs.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.