NCP-GENL Model Deployment Practice Question
A team is deploying a 70B-parameter LLM using NVIDIA Triton Inference Server with TensorRT-LLM backend on a node with four A100 80GB GPUs. They observe that during inference, only one GPU is utilized while the others remain idle. They have configured the model with tensor parallelism set to 1. What is the most likely cause of this underutilization?
⚠ Common exam trap
Many exam-takers confuse instance groups (which manage multiple model instances) with tensor parallelism (which splits a single model across GPUs).
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Tensor parallelism is set to 1, which means the model is not partitioned across GPUs; it runs entirely on one GPU.
Tensor parallelism determines how a model is sharded across multiple GPUs. Setting it to 1 means no sharding, so the entire model runs on a single GPU, leaving others idle. To utilize all four GPUs for a single model instance, tensor parallelism should be set to 4. Other factors like instance groups or batch size do not override this fundamental partitioning setting.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The Triton model repository is not configured with the correct instance group for multi-GPU execution.
Why it's wrong here
The instance group configuration controls how many model instances run and on which GPUs, but with tensor parallelism set to 1, the model itself is not sharded across GPUs. Even with multiple instances, each instance would run on a single GPU, not utilize all four GPUs for a single model replica. The core issue is the tensor parallelism setting.
- ✓
Tensor parallelism is set to 1, which means the model is not partitioned across GPUs; it runs entirely on one GPU.
Why this is correct
Tensor parallelism splits model layers across multiple GPUs. When set to 1, no partitioning occurs, so the entire model resides on a single GPU, leaving others idle. To utilize all four GPUs for a single model instance, tensor parallelism must be set to 4, matching the number of GPUs. This is the direct cause of underutilization.
- ✗
The model was compiled with a batch size that is too small to trigger multi-GPU execution.
Why it's wrong here
Batch size affects throughput and memory usage but does not determine whether tensor parallelism is used. Even with a small batch, if tensor parallelism is set greater than 1, the model would still be partitioned across GPUs. The observed behavior is due to tensor parallelism being set to 1, not batch size.
- ✗
The KV cache is not enabled, causing the model to fall back to single-GPU execution.
Why it's wrong here
KV cache is used to store key/value tensors for autoregressive generation and does not control multi-GPU execution. Its absence would affect performance and memory but would not cause the model to run on only one GPU. The underutilization is due to tensor parallelism configuration, not KV cache.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.