Courseiva

NCP-GENL · domain

Model Deployment

This domain covers serving and running LLMs with NVIDIA tooling: Triton Inference Server model repositories, instance groups and GPU placement, TensorRT-LLM engines, multi-GPU parallelism, and in-flight batching. Questions are scenario-based, using exhibits of config.pbtxt or memory profiles, asking you to diagnose load failures, OOM errors, and throughput bottlenecks.

37 questions7 easy16 medium14 hard

Focused practice

Practice Model Deployment questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Model Deployment

You must read a Triton config or TensorRT-LLM setup and diagnose why a model fails to load or runs out of memory, then choose the right fix. The single most important thing: verify instance-to-GPU mapping and memory budget before tuning throughput.

Triton model repository layout, config.pbtxt fields, and model load/versioning behavior

Triton instance groups, GPU placement, and scheduling across available devices

TensorRT-LLM engine build options, quantization, and KV cache memory sizing

Multi-GPU parallelism in TensorRT-LLM: tensor, pipeline, and expert parallelism

Watch out for

Common Model Deployment exam traps

  • ▸Assuming Triton auto-balances instances across GPUs; it does not, so duplicate GPU assignments cause load failures.
  • ▸Confusing tensor parallelism with pipeline parallelism, then picking the wrong technique for reducing inter-GPU communication.
  • ▸Ignoring KV cache growth with batch size and sequence length, then blaming weights for OOM spikes during inference.

Question index

All Model Deployment questions (37)

Click any question to see the full explanation, or start a practice session above.

1

An engineer is preparing an ensemble in NVIDIA Triton Inference Server that chains a Python preprocessing model with a TensorRT-LLM backend model. The preprocessing model must run on CPU while the LLM must run on GPU, and the ensemble must expose a single HTTP endpoint. Which configuration is required to make the ensemble execute correctly?

Easy
2

A startup is deploying a small LLM for a chatbot on a single NVIDIA L4 GPU using NVIDIA Triton Inference Server. They want to ensure the model is automatically loaded when Triton starts and can be updated without restarting the server. Which Triton feature should they configure?

Easy
3

A team is deploying a large language model on NVIDIA Triton Inference Server with TensorRT-LLM backend. They want to reduce GPU memory consumption to fit a larger model on a single GPU without significantly degrading output quality. Which two techniques should they use? (Choose two.)

Medium
4

A healthcare startup is deploying a Mistral 7B model for internal clinical note summarization. They need to serve the model with NVIDIA Triton Inference Server and want to minimize GPU memory footprint during inference. The team plans to use TensorRT-LLM and is choosing a numerical precision for the engine. Which precision should they select to reduce memory usage while maintaining acceptable accuracy for summarization?

Easy
5

A team is deploying a 70B-parameter LLM across four NVIDIA H100 GPUs using NVIDIA TensorRT-LLM with tensor parallelism. They observe that inference works but throughput is lower than expected, and profiling shows significant inter-GPU communication overhead. Which optimization should they apply first to reduce communication overhead?

Hard
6

An enterprise is deploying a large language model on NVIDIA Triton Inference Server. Which deployment strategy minimizes latency for requests that require high-throughput batching while maintaining consistent hardware utilization?

Medium
7

An enterprise deployment team needs to deploy a Large Language Model on NVIDIA Triton Inference Server. They require the lowest possible latency for real-time inference while maximizing GPU memory utilization. Which configuration strategy should the team implement?

Medium
8

A team is deploying a 70B-parameter LLM using NVIDIA Triton Inference Server with TensorRT-LLM backend on a node with four A100 80GB GPUs. They observe that during inference, only one GPU is utilized while the others remain idle. They have configured the model with tensor parallelism set to 1. What is the most likely cause of this underutilization?

Hard
9

Refer to the exhibit. What is the implication of setting the memory_limit to 0.8 in the context of an LLM inference service?

Hard
10

What is the primary benefit of deploying a model with a 'Model Ensemble' configuration in Triton Inference Server?

Medium
11

An engineer is deploying a LLM using NVIDIA Triton Inference Server with the TensorRT-LLM backend. They need to ensure that the model can handle a sudden surge in requests without increasing latency beyond a specified threshold. They have configured the model with a maximum batch size of 32 and dynamic batching with a preferred batch size of 16. However, during peak load, latency spikes are observed. Which Triton configuration parameter should they adjust to control the maximum time a request waits in the dynamic batching queue before being processed?

Hard
12

Which metric is most critical to monitor for identifying 'bottlenecks' in a high-throughput LLM deployment?

Easy
13

A media-analytics firm serves a 13B-parameter summarization model on two A100 GPUs using NVIDIA TensorRT-LLM behind Triton Inference Server. Traffic is bursty: during live events concurrency triples for about ten minutes, then returns to baseline. Operators report that the first requests after each burst begin are slow and sometimes time out, although steady-state latency is acceptable. Which deployment change most directly addresses the cold-start penalty at the beginning of each burst?

Hard
14

Refer to the exhibit. The Triton Inference Server is failing to load a model. The configuration shows two instances assigned to GPU 0. What is the most likely cause and the correct remediation?

Hard
15

A team is deploying a 70B parameter LLM with NVIDIA Triton Inference Server across four NVIDIA H100 GPUs. They are using TensorRT-LLM and need to fit the model within the combined GPU memory while maintaining high throughput. Which two techniques should they use? (Choose two.)

Hard
16

A healthcare company is deploying an LLM for clinical note summarization using NVIDIA Triton Inference Server. They must ensure that only authorized users can access the model and that all inference requests are logged for audit. Which Triton feature should they configure to enforce authentication and authorization?

Medium
17

An engineer needs to expose an LLM served by NVIDIA Triton Inference Server to a web application over HTTP with token streaming. Which Triton feature should they enable?

Easy
18

An inference team runs a TensorRT-LLM model behind NVIDIA Triton Inference Server with the tensorrtllm_backend. They notice that when clients send requests with widely varying prompt lengths, throughput drops and some requests wait much longer than others. They want Triton to group compatible requests dynamically to improve GPU utilization without changing the model build. Which Triton feature should they configure?

Hard
19

When deploying an LLM using NVIDIA TensorRT-LLM, what is the primary benefit of pre-compiling the model into a TensorRT engine?

Medium
20

Refer to the exhibit. An engineer observes that memory utilization spikes during inference, causing OOM errors. Given the configuration, what is the most likely cause of the failure?

Hard
21

An organization is deploying a high-throughput LLM on NVIDIA Triton Inference Server. They observe significant tail latency spikes when serving multiple concurrent requests. Which strategy most effectively optimizes GPU utilization and reduces latency jitter for these concurrent model instances?

Medium
22

A team is deploying a 13B-parameter LLM with NVIDIA TensorRT-LLM on a single A100 80GB GPU. They want to reduce GPU memory usage during inference without retraining the model, while keeping acceptable output quality. Which technique should they apply?

Medium
23

An AI engineer is deploying a large language model using NVIDIA Triton Inference Server. They need to ensure that the server can handle multiple concurrent requests efficiently while maintaining low latency. Which Triton feature allows the server to dynamically batch incoming requests to maximize GPU utilization?

Easy
24

Which TWO factors should be prioritized when selecting a quantization strategy for deploying a large language model on constrained edge hardware? (Select TWO)

Medium
25

A platform team is deploying a 70B-parameter LLM with NVIDIA TensorRT-LLM across four 80 GB H100 GPUs and needs to serve long-context requests efficiently. They are deciding how to combine parallelism and memory techniques in the build and runtime configuration. (Choose two.)

Hard
26

Which component in the NVIDIA AI Enterprise stack is primarily responsible for serving multiple models, managing model versions, and providing metrics for monitoring model health?

Easy
27

In the context of LLM deployment, why is it recommended to use a dedicated inference server like Triton rather than a basic Flask or FastAPI wrapper?

Medium
28

When deploying a model, what is the benefit of using Triton's 'Model Versioning' feature?

Medium
29

A team is deploying a quantized LLM using NVIDIA NIM. To ensure the highest level of security and compliance, they need to verify that the container image has been scanned for vulnerabilities before production use. Which tool is the primary source for certified, production-ready NIM containers?

Medium
30

A team is deploying a large language model on NVIDIA Triton Inference Server with NVIDIA TensorRT-LLM backend. They need to reduce GPU memory usage to fit a larger model on the same hardware while maintaining acceptable latency. Which two techniques should they use? (Choose two.)

Hard
31

Refer to the exhibit. What is the most likely cause of the failure based on the log entries?

Hard
32

A company needs to deploy a generative AI model that will serve prompts containing regulated customer data. Security policy requires that all inference stays on-premises, that the model be quantized to fit existing GPUs, and that no external network calls occur at runtime. Which deployment approach should the engineer choose?

Medium
33

What is the primary role of a Model Repository in the NVIDIA Triton Inference Server architecture?

Medium
34

Which TWO of the following NVIDIA AI Enterprise tools are specifically designed to optimize and accelerate the deployment of LLMs in containerized environments?

Medium
35

A media analytics company runs a TensorRT-LLM optimized GPT-J model on a single NVIDIA A100 80GB GPU using NVIDIA Triton Inference Server. During peak hours, request concurrency rises sharply and the team observes that the GPU is idle for long periods while waiting on host-side tokenization and detokenization. Profiling shows that CPU preprocessing and postprocessing dominate request latency. The team wants to reduce end-to-end latency without changing model weights or adding GPUs. Which Triton feature should they use?

Hard
36

A team is optimizing an NVIDIA TensorRT-LLM deployment of a 70B model on multiple GPUs. They want to reduce inter-GPU communication overhead and improve throughput. Which two techniques should they consider? (Choose two.)

Medium
37

An LLM service on NVIDIA Triton Inference Server experiences high time-to-first-token because the dynamic batcher waits for full batches. The team wants to reduce time-to-first-token while still benefiting from batching. Which adjustment is most appropriate?

Hard

Frequently asked questions

What does the Model Deployment domain cover on the NCP-GENL exam?
You must read a Triton config or TensorRT-LLM setup and diagnose why a model fails to load or runs out of memory, then choose the right fix. The single most important thing: verify instance-to-GPU mapping and memory budget before tuning throughput.
How many questions are in this domain?
This page lists all 37 Model Deployment questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Model Deployment questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
nvidia-ncp-genl NVIDIA-NCP-GENL ncp model deployment Practice Questions