You must read a Triton config or TensorRT-LLM setup and diagnose why a model fails to load or runs out of memory, then choose the right fix. The single most important thing: verify instance-to-GPU mapping and memory budget before tuning throughput.
Start practicing
Model Deployment — choose a session length
Free · No account required
Domain overview
This domain covers serving and running LLMs with NVIDIA tooling: Triton Inference Server model repositories, instance groups and GPU placement, TensorRT-LLM engines, multi-GPU parallelism, and in-flight batching. Questions are scenario-based, using exhibits of config. pbtxt or memory profiles, asking you to diagnose load failures, OOM errors, and throughput bottlenecks.
Exam objectives
Triton model repository layout, config.pbtxt fields, and model load/versioning behavior
Triton instance groups, GPU placement, and scheduling across available devices
TensorRT-LLM engine build options, quantization, and KV cache memory sizing
Multi-GPU parallelism in TensorRT-LLM: tensor, pipeline, and expert parallelism
Assuming Triton auto-balances instances across GPUs; it does not, so duplicate GPU assignments cause load failures.
Confusing tensor parallelism with pipeline parallelism, then picking the wrong technique for reducing inter-GPU communication.
Ignoring KV cache growth with batch size and sequence length, then blaming weights for OOM spikes during inference.
Click any question to see the full explanation and answer options, or start a focused practice session above.
An enterprise is deploying a large language model on NVIDIA Triton Inference Server. Which deployment strategy minimizes latency for requests that require high-throughput batching while maintaining consistent hardware utilization?
2Refer to the exhibit. An engineer observes that memory utilization spikes during inference, causing OOM errors. Given the configuration, what is the most likely cause of the failure?
3Which TWO factors should be prioritized when selecting a quantization strategy for deploying a large language model on constrained edge hardware? (Select TWO)
4When deploying an LLM using NVIDIA TensorRT-LLM, what is the primary benefit of pre-compiling the model into a TensorRT engine?
5Which component in the NVIDIA AI Enterprise stack is primarily responsible for serving multiple models, managing model versions, and providing metrics for monitoring model health?
6What is the primary role of a Model Repository in the NVIDIA Triton Inference Server architecture?
7In the context of LLM deployment, why is it recommended to use a dedicated inference server like Triton rather than a basic Flask or FastAPI wrapper?
8Refer to the exhibit. What is the most likely cause of the failure based on the log entries?
9What is the primary benefit of deploying a model with a 'Model Ensemble' configuration in Triton Inference Server?
10Which metric is most critical to monitor for identifying 'bottlenecks' in a high-throughput LLM deployment?
11Refer to the exhibit. What is the implication of setting the memory_limit to 0.8 in the context of an LLM inference service?
12When deploying a model, what is the benefit of using Triton's 'Model Versioning' feature?
13An organization is deploying a high-throughput LLM on NVIDIA Triton Inference Server. They observe significant tail latency spikes when serving multiple concurrent requests. Which strategy most effectively optimizes GPU utilization and reduces latency jitter for these concurrent model instances?
14Refer to the exhibit. The Triton Inference Server is failing to load a model. The configuration shows two instances assigned to GPU 0. What is the most likely cause and the correct remediation?
15An enterprise deployment team needs to deploy a Large Language Model on NVIDIA Triton Inference Server. They require the lowest possible latency for real-time inference while maximizing GPU memory utilization. Which configuration strategy should the team implement?
16Which TWO of the following NVIDIA AI Enterprise tools are specifically designed to optimize and accelerate the deployment of LLMs in containerized environments?
17A team is deploying a quantized LLM using NVIDIA NIM. To ensure the highest level of security and compliance, they need to verify that the container image has been scanned for vulnerabilities before production use. Which tool is the primary source for certified, production-ready NIM containers?
18A media-analytics firm serves a 13B-parameter summarization model on two A100 GPUs using NVIDIA TensorRT-LLM behind Triton Inference Server. Traffic is bursty: during live events concurrency triples for about ten minutes, then returns to baseline. Operators report that the first requests after each burst begin are slow and sometimes time out, although steady-state latency is acceptable. Which deployment change most directly addresses the cold-start penalty at the beginning of each burst?
19A team is deploying a 13B-parameter LLM with NVIDIA TensorRT-LLM on a single A100 80GB GPU. They want to reduce GPU memory usage during inference without retraining the model, while keeping acceptable output quality. Which technique should they apply?
20An engineer needs to expose an LLM served by NVIDIA Triton Inference Server to a web application over HTTP with token streaming. Which Triton feature should they enable?
21A team is optimizing an NVIDIA TensorRT-LLM deployment of a 70B model on multiple GPUs. They want to reduce inter-GPU communication overhead and improve throughput. Which two techniques should they consider? (Choose two.)
22A team is deploying a 70B-parameter LLM using NVIDIA Triton Inference Server with TensorRT-LLM backend on a node with four A100 80GB GPUs. They observe that during inference, only one GPU is utilized while the others remain idle. They have configured the model with tensor parallelism set to 1. What is the most likely cause of this underutilization?
23An engineer is preparing an ensemble in NVIDIA Triton Inference Server that chains a Python preprocessing model with a TensorRT-LLM backend model. The preprocessing model must run on CPU while the LLM must run on GPU, and the ensemble must expose a single HTTP endpoint. Which configuration is required to make the ensemble execute correctly?
24An LLM service on NVIDIA Triton Inference Server experiences high time-to-first-token because the dynamic batcher waits for full batches. The team wants to reduce time-to-first-token while still benefiting from batching. Which adjustment is most appropriate?
25An AI engineer is deploying a large language model using NVIDIA Triton Inference Server. They need to ensure that the server can handle multiple concurrent requests efficiently while maintaining low latency. Which Triton feature allows the server to dynamically batch incoming requests to maximize GPU utilization?
26A platform team is deploying a 70B-parameter LLM with NVIDIA TensorRT-LLM across four 80 GB H100 GPUs and needs to serve long-context requests efficiently. They are deciding how to combine parallelism and memory techniques in the build and runtime configuration. (Choose two.)
27A team is deploying a large language model on NVIDIA Triton Inference Server with TensorRT-LLM backend. They want to reduce GPU memory consumption to fit a larger model on a single GPU without significantly degrading output quality. Which two techniques should they use? (Choose two.)
28An inference team runs a TensorRT-LLM model behind NVIDIA Triton Inference Server with the tensorrtllm_backend. They notice that when clients send requests with widely varying prompt lengths, throughput drops and some requests wait much longer than others. They want Triton to group compatible requests dynamically to improve GPU utilization without changing the model build. Which Triton feature should they configure?
29A media analytics company runs a TensorRT-LLM optimized GPT-J model on a single NVIDIA A100 80GB GPU using NVIDIA Triton Inference Server. During peak hours, request concurrency rises sharply and the team observes that the GPU is idle for long periods while waiting on host-side tokenization and detokenization. Profiling shows that CPU preprocessing and postprocessing dominate request latency. The team wants to reduce end-to-end latency without changing model weights or adding GPUs. Which Triton feature should they use?
30An engineer is deploying a LLM using NVIDIA Triton Inference Server with the TensorRT-LLM backend. They need to ensure that the model can handle a sudden surge in requests without increasing latency beyond a specified threshold. They have configured the model with a maximum batch size of 32 and dynamic batching with a preferred batch size of 16. However, during peak load, latency spikes are observed. Which Triton configuration parameter should they adjust to control the maximum time a request waits in the dynamic batching queue before being processed?
31A company needs to deploy a generative AI model that will serve prompts containing regulated customer data. Security policy requires that all inference stays on-premises, that the model be quantized to fit existing GPUs, and that no external network calls occur at runtime. Which deployment approach should the engineer choose?
32A healthcare startup is deploying a Mistral 7B model for internal clinical note summarization. They need to serve the model with NVIDIA Triton Inference Server and want to minimize GPU memory footprint during inference. The team plans to use TensorRT-LLM and is choosing a numerical precision for the engine. Which precision should they select to reduce memory usage while maintaining acceptable accuracy for summarization?
33A team is deploying a 70B parameter LLM with NVIDIA Triton Inference Server across four NVIDIA H100 GPUs. They are using TensorRT-LLM and need to fit the model within the combined GPU memory while maintaining high throughput. Which two techniques should they use? (Choose two.)
34A team is deploying a 70B-parameter LLM across four NVIDIA H100 GPUs using NVIDIA TensorRT-LLM with tensor parallelism. They observe that inference works but throughput is lower than expected, and profiling shows significant inter-GPU communication overhead. Which optimization should they apply first to reduce communication overhead?
35A healthcare company is deploying an LLM for clinical note summarization using NVIDIA Triton Inference Server. They must ensure that only authorized users can access the model and that all inference requests are logged for audit. Which Triton feature should they configure to enforce authentication and authorization?
36A team is deploying a large language model on NVIDIA Triton Inference Server with NVIDIA TensorRT-LLM backend. They need to reduce GPU memory usage to fit a larger model on the same hardware while maintaining acceptable latency. Which two techniques should they use? (Choose two.)
37A startup is deploying a small LLM for a chatbot on a single NVIDIA L4 GPU using NVIDIA Triton Inference Server. They want to ensure the model is automatically loaded when Triton starts and can be updated without restarting the server. Which Triton feature should they configure?
You must read a Triton config or TensorRT-LLM setup and diagnose why a model fails to load or runs out of memory, then choose the right fix. The single most important thing: verify instance-to-GPU mapping and memory budget before tuning throughput.
The Courseiva NCP-GENL question bank contains 37 questions in the Model Deployment domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Model Deployment domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included