Courseiva
Model Deployment →hardMultiple Choice

NCP-GENL Model Deployment Practice Question

Exhibit

config.pbtxt: 
instance_group [
  {
    count: 2
    kind: KIND_GPU
    gpus: [0]
  }
]

error_log: [ERROR] 'Failed to load model: CUDA out of memory' during multi-model parallel execution.

Refer to the exhibit. The Triton Inference Server is failing to load a model. The configuration shows two instances assigned to GPU 0. What is the most likely cause and the correct remediation?

⚠ Common exam trap

Candidates frequently assume the error is due to a software version mismatch or driver issue, failing to calculate the cumulative VRAM consumption of multiple model instances relative to the hardware limit.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The total VRAM required for two instances exceeds the GPU capacity; reduce instance count to 1.

The error indicates that the two instances of the model are collectively requesting more VRAM than is available on the physical GPU. By default, Triton attempts to allocate memory for every configured instance upon startup. The remediation requires either reducing the number of instances or implementing a memory-aware model partitioning strategy, such as using Model Analyzer to determine the safe memory footprint per instance before deployment.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The model count exceeds the number of available CUDA cores, requiring a driver update.

    Why it's wrong here

    CUDA core count is independent of VRAM capacity limits. Adding more instances does not change the requirement for CUDA cores but linearly increases the demand for video memory. Updating the driver will not resolve a physical memory capacity issue caused by parallel model loading on a single device.

  • ✓

    The total VRAM required for two instances exceeds the GPU capacity; reduce instance count to 1.

    Why this is correct

    Each model instance requires a dedicated memory buffer for weights and activation tensors. When configured with 'count: 2', Triton attempts to load the model twice. If the sum exceeds the VRAM, an OOM occurs. Reducing the instance count is the most direct way to resolve the startup conflict.

  • ✗

    The model weights are corrupted, preventing the inference engine from initializing the memory space.

    Why it's wrong here

    A corrupted file would trigger a file I/O or model deserialization error, not a 'CUDA out of memory' error. The specific error message points directly to resource exhaustion within the GPU memory space rather than a data integrity issue within the model artifact itself.

  • ✗

    The batch size is set to zero in the configuration, preventing memory allocation.

    Why it's wrong here

    A batch size of zero would typically result in a configuration validation error or a logic error during request processing. It would not cause an OOM event during the initialization phase, as no memory would be allocated for actual batch processing if the batch size were effectively null.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.