Courseiva
Model Deployment →hardMultiple Choice

NCP-GENL Model Deployment Practice Question

Exhibit

model_config.pbtxt:
instance_group [
  {
    count: 2
    kind: KIND_GPU
    gpus: [0]
  }
]

Refer to the exhibit. An engineer observes that memory utilization spikes during inference, causing OOM errors. Given the configuration, what is the most likely cause of the failure?

⚠ Common exam trap

Candidates often blame the model size or the GPU hardware itself, failing to notice that the configuration defines multiple instances, which causes a multiplicative effect on VRAM consumption.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The instance group count of 2 causes redundant loading of weights, exceeding VRAM.

The instance group configuration defines two concurrent instances on the same GPU. Each instance attempts to load a separate copy of the model weights into the GPU memory. If the model size is large, doubling the instances exceeds the available VRAM capacity. This configuration is a common mistake when deploying LLMs where model footprint is significant relative to total available device memory.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The count of 2 causes the GPU to oversubscribe its thermal limits.

    Why it's wrong here

    While higher instance counts increase power draw, OOM errors are specifically related to memory capacity, not thermal limits. Thermal throttling would manifest as increased latency or performance degradation over time rather than an immediate Out of Memory error during the model loading or inference execution phases.

  • ✓

    The instance group count of 2 causes redundant loading of weights, exceeding VRAM.

    Why this is correct

    Setting the instance count to two instructs Triton to create two independent model runners. Each runner requires its own memory allocation for weights and activations. If the model occupies a large portion of the GPU memory, running two instances simultaneously will inevitably exhaust the total available VRAM.

  • ✗

    Triton requires KIND_CPU for concurrent instance execution.

    Why it's wrong here

    Triton supports concurrent execution on both CPU and GPU. Using KIND_GPU is the preferred way to leverage hardware acceleration for LLM inference. The error is not related to the hardware kind selection, but rather the allocation of memory resources assigned to the specific GPU device count.

  • ✗

    The gpus index [0] is invalid for multi-instance deployment.

    Why it's wrong here

    The index [0] is a valid reference to the first physical GPU in the system. Specifying the GPU index is standard practice in Triton configuration. The error arises from attempting to run multiple instances on that specific device without sufficient memory, not from an incorrect index assignment.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.