Courseiva

NCP-GENL · topic practice

Model Deployment practice questions

This domain covers serving and running LLMs with NVIDIA tooling: Triton Inference Server model repositories, instance groups and GPU placement, TensorRT-LLM engines, multi-GPU parallelism, and in-flight batching. Questions are scenario-based, using exhibits of config.pbtxt or memory profiles, asking you to diagnose load failures, OOM errors, and throughput bottlenecks.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Model Deployment

What the exam tests

What to know about Model Deployment

You must read a Triton config or TensorRT-LLM setup and diagnose why a model fails to load or runs out of memory, then choose the right fix. The single most important thing: verify instance-to-GPU mapping and memory budget before tuning throughput.

Triton model repository layout, config.pbtxt fields, and model load/versioning behavior

Triton instance groups, GPU placement, and scheduling across available devices

TensorRT-LLM engine build options, quantization, and KV cache memory sizing

Multi-GPU parallelism in TensorRT-LLM: tensor, pipeline, and expert parallelism

Watch out for

Common Model Deployment exam traps

  • ▸Assuming Triton auto-balances instances across GPUs; it does not, so duplicate GPU assignments cause load failures.
  • ▸Confusing tensor parallelism with pipeline parallelism, then picking the wrong technique for reducing inter-GPU communication.
  • ▸Ignoring KV cache growth with batch size and sequence length, then blaming weights for OOM spikes during inference.

Practice set

Model Deployment questions

20 questions · select your answer, then reveal the explanation

Refer to the exhibit. An engineer configures Triton for dynamic batching. If requests arrive at 2ms intervals and the current queue is empty, what is the expected batching behavior?

Exhibit

{
  "policy": "allow_dynamic_batching",
  "max_queue_delay_microseconds": 5000,
  "batch_sizes": [1, 2, 4, 8],
  "preferred_batch_size": [4]
}

Which THREE techniques are effective for managing memory in high-concurrency LLM deployments on NVIDIA GPUs? (Select THREE)

Question 3mediummultiple choice
Read the full Model Deployment explanation →

When migrating an LLM deployment to a multi-GPU setup, what is the significance of 'Tensor Parallelism'?

An engineering team is deploying a Transformer-based model using NVIDIA TensorRT-LLM on an H100 GPU cluster. Which TWO factors are most critical to consider when configuring the In-flight Batching (IFB) feature for optimal performance?

Refer to the exhibit. An engineer observes that the LLM deployment exhibits high latency and inefficient GPU utilization under moderate load. Based on the provided Triton configuration, what is the most likely cause?

Exhibit

config.pbtxt: 
backend: "tensorrtllm"
instance_group [{ count: 1, kind: KIND_GPU }]
dynamic_batching { preferred_batch_size: [4, 8] }
Question 6mediummultiple choice
Read the full Model Deployment explanation →

A healthcare company runs an NVIDIA Triton Inference Server hosting a clinical-summarization LLM. Compliance requires that every generated summary be fully reproducible for audit: the same prompt must always yield byte-identical output, and token-level probabilities must be retrievable later. The team currently sends requests with default sampling parameters and stores only the final text. Which change to the deployment configuration best satisfies the audit requirement?

A production LLM service on NVIDIA Triton Inference Server must handle bursty traffic while maintaining low latency. The team wants to cap the number of concurrent requests per model instance and queue excess requests. Which Triton configuration parameter should they tune?

Question 8mediummultiple choice
Read the full Model Deployment explanation →

A financial services company is deploying a 13B-parameter LLM using NVIDIA TensorRT-LLM. Their security policy requires that model weights never leave the GPU memory unencrypted and that inference requests are processed with the lowest possible overhead. They have a single NVIDIA H100 GPU with 80 GB memory. Which TensorRT-LLM feature should they enable to meet these requirements while maximizing throughput?

Question 9mediummultiple choice
Read the full Model Deployment explanation →

A team is deploying a Llama-3-8B model with NVIDIA TensorRT-LLM on a single 24 GB A10G GPU. They configure the TensorRT-LLM build with a max batch size of 64 and the default paged KV cache. During a load test with 40 concurrent requests, the server begins returning CUDA out-of-memory errors after several minutes, even though the model weights alone fit comfortably in VRAM. Which configuration change most directly addresses the root cause?

Question 10mediummultiple choice
Read the full Model Deployment explanation →

A financial services company is deploying a Llama-2 13B model on NVIDIA Triton Inference Server. Compliance requires that the model must never be served with unvalidated weights. The MLOps team stores each retrained model in a new version directory under the model repository. They want Triton to automatically load the newest version as soon as it appears, but they also need the ability to immediately roll back to the previous version if a validation check fails. Which Triton model control mode should they configure?

Question 11mediummultiple choice
Read the full Model Deployment explanation →

An e-commerce company deploys a customer support chatbot backed by a TensorRT-LLM model on NVIDIA Triton Inference Server. The model uses a KV cache and the team wants to support multiple concurrent conversations while keeping per-request latency predictable. They configure the Triton TensorRT-LLM backend with a fixed max_batch_size and a limited KV cache. During testing, they observe that some requests are rejected with out-of-memory errors when many long conversations occur simultaneously. Which configuration change best addresses the issue without reducing the maximum batch size?

Question 12mediummultiple choice
Read the full Model Deployment explanation →

A financial services company is deploying a 13B-parameter LLM on a single NVIDIA A100 80GB GPU using NVIDIA TensorRT-LLM. During peak traffic, they observe that throughput plateaus and GPU memory is nearly exhausted, but latency remains acceptable. They want to increase concurrent request handling without adding another GPU. Which configuration change should they make?

Question 13mediummultiple choice
Read the full Model Deployment explanation →

An enterprise is deploying a large language model on NVIDIA Triton Inference Server. Which deployment strategy minimizes latency for requests that require high-throughput batching while maintaining consistent hardware utilization?

Refer to the exhibit. An engineer observes that memory utilization spikes during inference, causing OOM errors. Given the configuration, what is the most likely cause of the failure?

Exhibit

model_config.pbtxt:
instance_group [
  {
    count: 2
    kind: KIND_GPU
    gpus: [0]
  }
]

Which TWO factors should be prioritized when selecting a quantization strategy for deploying a large language model on constrained edge hardware? (Select TWO)

Question 16mediummultiple choice
Read the full Model Deployment explanation →

When deploying an LLM using NVIDIA TensorRT-LLM, what is the primary benefit of pre-compiling the model into a TensorRT engine?

Which component in the NVIDIA AI Enterprise stack is primarily responsible for serving multiple models, managing model versions, and providing metrics for monitoring model health?

Question 18mediummultiple choice
Read the full Model Deployment explanation →

What is the primary role of a Model Repository in the NVIDIA Triton Inference Server architecture?

Question 19mediummultiple choice
Read the full Model Deployment explanation →

In the context of LLM deployment, why is it recommended to use a dedicated inference server like Triton rather than a basic Flask or FastAPI wrapper?

Refer to the exhibit. What is the most likely cause of the failure based on the log entries?

Exhibit

log_output:
[W] [TRT-LLM] [Performance] Request 1042 processing time exceeded 500ms.
[W] [TRT-LLM] [Performance] KV Cache eviction detected for sequence 1042.
[E] [TRT-LLM] [Memory] Out of Memory: Failed to allocate 128MB.

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Model Deployment sessions

Start a Model Deployment only practice session

Every question in these sessions is drawn from the Model Deployment domain — nothing else.

Related practice questions

Related NCP-GENL topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the NCP-GENL exam test about Model Deployment?
You must read a Triton config or TensorRT-LLM setup and diagnose why a model fails to load or runs out of memory, then choose the right fix. The single most important thing: verify instance-to-GPU mapping and memory budget before tuning throughput.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Model Deployment questions in a focused session?
Yes — the session launcher on this page draws every question from the Model Deployment domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other NCP-GENL topics?
Use the topic links above to move to related areas, or go back to the NCP-GENL question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the NCP-GENL exam covers. They are not copied from any real exam or dump site.