Courseiva
Model Optimization →mediumMultiple Choice

NCP-GENL Model Optimization Practice Question

Exhibit

config.pbtxt:
instance_group [
  {
    count: 2
    kind: KIND_GPU
    gpus: [0]
  }
]

Refer to the exhibit. An engineer observes that GPU memory utilization is high, but the GPU is frequently idling. How does the provided Triton configuration optimize the inference pipeline?

⚠ Common exam trap

Candidates often misinterpret high VRAM utilization with optimal performance, missing the fact that unutilized compute cycles indicate poor concurrency and request scheduling bottlenecks.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

It enables concurrent execution on the GPU

By setting the instance count to 2 on a single GPU, the configuration enables concurrent model execution. This allows Triton to schedule multiple inference requests to be processed simultaneously on the same hardware. This overlap helps to hide memory latency and pipeline stalls, effectively increasing the utilization of the GPU and raising the overall throughput of the deployment.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    It forces the model to use INT8 precision

    Why it's wrong here

    The provided configuration file defines the instance group properties for the model, such as hardware affinity and concurrency. It does not dictate the numerical precision of the model, which is determined by the underlying TensorRT engine file or the model framework configuration, not the Triton instance settings.

  • ✓

    It enables concurrent execution on the GPU

    Why this is correct

    Setting 'count: 2' instructs Triton to launch two instances of the model on the specified GPU. This allows the server to process multiple requests in parallel, which is a standard method to improve throughput and keep the GPU busy during periods where one instance might be blocked.

  • ✗

    It partitions the GPU for dedicated memory

    Why it's wrong here

    This configuration does not partition the GPU memory at the hardware or driver level. It simply defines how many model instances Triton will manage. Memory partitioning for isolation would typically require Multi-Instance GPU (MIG) settings, which are configured at the hardware or hypervisor layer, not in Triton's model configuration.

  • ✗

    It enables dynamic batching for the requests

    Why it's wrong here

    Dynamic batching requires a separate 'dynamic_batching' block in the configuration file to define parameters like 'preferred_batch_size'. The 'instance_group' settings define execution concurrency, which is a distinct concept from batching, where requests are grouped together to maximize the efficiency of single-instance model runs.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.