NCP-GENL · domain
Model Deployment
This domain covers serving and running LLMs with NVIDIA tooling: Triton Inference Server model repositories, instance groups and GPU placement, TensorRT-LLM engines, multi-GPU parallelism, and in-flight batching. Questions are scenario-based, using exhibits of config.pbtxt or memory profiles, asking you to diagnose load failures, OOM errors, and throughput bottlenecks.
Focused practice
Practice Model Deployment questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Model Deployment
You must read a Triton config or TensorRT-LLM setup and diagnose why a model fails to load or runs out of memory, then choose the right fix. The single most important thing: verify instance-to-GPU mapping and memory budget before tuning throughput.
Triton model repository layout, config.pbtxt fields, and model load/versioning behavior
Triton instance groups, GPU placement, and scheduling across available devices
TensorRT-LLM engine build options, quantization, and KV cache memory sizing
Multi-GPU parallelism in TensorRT-LLM: tensor, pipeline, and expert parallelism
Watch out for
Common Model Deployment exam traps
- ▸Assuming Triton auto-balances instances across GPUs; it does not, so duplicate GPU assignments cause load failures.
- ▸Confusing tensor parallelism with pipeline parallelism, then picking the wrong technique for reducing inter-GPU communication.
- ▸Ignoring KV cache growth with batch size and sequence length, then blaming weights for OOM spikes during inference.
Question index
All Model Deployment questions (37)
Click any question to see the full explanation, or start a practice session above.
An engineer is preparing an ensemble in NVIDIA Triton Inference Server that chains a Python preprocessing model with a TensorRT-LLM backend model. The preprocessing model must run on CPU while the LLM must run on GPU, and the ensemble must expose a single HTTP endpoint. Which configuration is required to make the ensemble execute correctly?
Easy2A startup is deploying a small LLM for a chatbot on a single NVIDIA L4 GPU using NVIDIA Triton Inference Server. They want to ensure the model is automatically loaded when Triton starts and can be updated without restarting the server. Which Triton feature should they configure?
Easy3A team is deploying a large language model on NVIDIA Triton Inference Server with TensorRT-LLM backend. They want to reduce GPU memory consumption to fit a larger model on a single GPU without significantly degrading output quality. Which two techniques should they use? (Choose two.)
Medium4A healthcare startup is deploying a Mistral 7B model for internal clinical note summarization. They need to serve the model with NVIDIA Triton Inference Server and want to minimize GPU memory footprint during inference. The team plans to use TensorRT-LLM and is choosing a numerical precision for the engine. Which precision should they select to reduce memory usage while maintaining acceptable accuracy for summarization?
Easy5A team is deploying a 70B-parameter LLM across four NVIDIA H100 GPUs using NVIDIA TensorRT-LLM with tensor parallelism. They observe that inference works but throughput is lower than expected, and profiling shows significant inter-GPU communication overhead. Which optimization should they apply first to reduce communication overhead?
Hard6An enterprise is deploying a large language model on NVIDIA Triton Inference Server. Which deployment strategy minimizes latency for requests that require high-throughput batching while maintaining consistent hardware utilization?
Medium7An enterprise deployment team needs to deploy a Large Language Model on NVIDIA Triton Inference Server. They require the lowest possible latency for real-time inference while maximizing GPU memory utilization. Which configuration strategy should the team implement?
Medium8A team is deploying a 70B-parameter LLM using NVIDIA Triton Inference Server with TensorRT-LLM backend on a node with four A100 80GB GPUs. They observe that during inference, only one GPU is utilized while the others remain idle. They have configured the model with tensor parallelism set to 1. What is the most likely cause of this underutilization?
Hard9Refer to the exhibit. What is the implication of setting the memory_limit to 0.8 in the context of an LLM inference service?
Hard10What is the primary benefit of deploying a model with a 'Model Ensemble' configuration in Triton Inference Server?
Medium11An engineer is deploying a LLM using NVIDIA Triton Inference Server with the TensorRT-LLM backend. They need to ensure that the model can handle a sudden surge in requests without increasing latency beyond a specified threshold. They have configured the model with a maximum batch size of 32 and dynamic batching with a preferred batch size of 16. However, during peak load, latency spikes are observed. Which Triton configuration parameter should they adjust to control the maximum time a request waits in the dynamic batching queue before being processed?
Hard12Which metric is most critical to monitor for identifying 'bottlenecks' in a high-throughput LLM deployment?
Easy13A media-analytics firm serves a 13B-parameter summarization model on two A100 GPUs using NVIDIA TensorRT-LLM behind Triton Inference Server. Traffic is bursty: during live events concurrency triples for about ten minutes, then returns to baseline. Operators report that the first requests after each burst begin are slow and sometimes time out, although steady-state latency is acceptable. Which deployment change most directly addresses the cold-start penalty at the beginning of each burst?
Hard14Refer to the exhibit. The Triton Inference Server is failing to load a model. The configuration shows two instances assigned to GPU 0. What is the most likely cause and the correct remediation?
Hard15A team is deploying a 70B parameter LLM with NVIDIA Triton Inference Server across four NVIDIA H100 GPUs. They are using TensorRT-LLM and need to fit the model within the combined GPU memory while maintaining high throughput. Which two techniques should they use? (Choose two.)
Hard16A healthcare company is deploying an LLM for clinical note summarization using NVIDIA Triton Inference Server. They must ensure that only authorized users can access the model and that all inference requests are logged for audit. Which Triton feature should they configure to enforce authentication and authorization?
Medium17An engineer needs to expose an LLM served by NVIDIA Triton Inference Server to a web application over HTTP with token streaming. Which Triton feature should they enable?
Easy18An inference team runs a TensorRT-LLM model behind NVIDIA Triton Inference Server with the tensorrtllm_backend. They notice that when clients send requests with widely varying prompt lengths, throughput drops and some requests wait much longer than others. They want Triton to group compatible requests dynamically to improve GPU utilization without changing the model build. Which Triton feature should they configure?
Hard19When deploying an LLM using NVIDIA TensorRT-LLM, what is the primary benefit of pre-compiling the model into a TensorRT engine?
Medium20Refer to the exhibit. An engineer observes that memory utilization spikes during inference, causing OOM errors. Given the configuration, what is the most likely cause of the failure?
Hard21An organization is deploying a high-throughput LLM on NVIDIA Triton Inference Server. They observe significant tail latency spikes when serving multiple concurrent requests. Which strategy most effectively optimizes GPU utilization and reduces latency jitter for these concurrent model instances?
Medium22A team is deploying a 13B-parameter LLM with NVIDIA TensorRT-LLM on a single A100 80GB GPU. They want to reduce GPU memory usage during inference without retraining the model, while keeping acceptable output quality. Which technique should they apply?
Medium23An AI engineer is deploying a large language model using NVIDIA Triton Inference Server. They need to ensure that the server can handle multiple concurrent requests efficiently while maintaining low latency. Which Triton feature allows the server to dynamically batch incoming requests to maximize GPU utilization?
Easy24Which TWO factors should be prioritized when selecting a quantization strategy for deploying a large language model on constrained edge hardware? (Select TWO)
Medium25A platform team is deploying a 70B-parameter LLM with NVIDIA TensorRT-LLM across four 80 GB H100 GPUs and needs to serve long-context requests efficiently. They are deciding how to combine parallelism and memory techniques in the build and runtime configuration. (Choose two.)
Hard26Which component in the NVIDIA AI Enterprise stack is primarily responsible for serving multiple models, managing model versions, and providing metrics for monitoring model health?
Easy27In the context of LLM deployment, why is it recommended to use a dedicated inference server like Triton rather than a basic Flask or FastAPI wrapper?
Medium28When deploying a model, what is the benefit of using Triton's 'Model Versioning' feature?
Medium29A team is deploying a quantized LLM using NVIDIA NIM. To ensure the highest level of security and compliance, they need to verify that the container image has been scanned for vulnerabilities before production use. Which tool is the primary source for certified, production-ready NIM containers?
Medium30A team is deploying a large language model on NVIDIA Triton Inference Server with NVIDIA TensorRT-LLM backend. They need to reduce GPU memory usage to fit a larger model on the same hardware while maintaining acceptable latency. Which two techniques should they use? (Choose two.)
Hard31Refer to the exhibit. What is the most likely cause of the failure based on the log entries?
Hard32A company needs to deploy a generative AI model that will serve prompts containing regulated customer data. Security policy requires that all inference stays on-premises, that the model be quantized to fit existing GPUs, and that no external network calls occur at runtime. Which deployment approach should the engineer choose?
Medium33What is the primary role of a Model Repository in the NVIDIA Triton Inference Server architecture?
Medium34Which TWO of the following NVIDIA AI Enterprise tools are specifically designed to optimize and accelerate the deployment of LLMs in containerized environments?
Medium35A media analytics company runs a TensorRT-LLM optimized GPT-J model on a single NVIDIA A100 80GB GPU using NVIDIA Triton Inference Server. During peak hours, request concurrency rises sharply and the team observes that the GPU is idle for long periods while waiting on host-side tokenization and detokenization. Profiling shows that CPU preprocessing and postprocessing dominate request latency. The team wants to reduce end-to-end latency without changing model weights or adding GPUs. Which Triton feature should they use?
Hard36A team is optimizing an NVIDIA TensorRT-LLM deployment of a 70B model on multiple GPUs. They want to reduce inter-GPU communication overhead and improve throughput. Which two techniques should they consider? (Choose two.)
Medium37An LLM service on NVIDIA Triton Inference Server experiences high time-to-first-token because the dynamic batcher waits for full batches. The team wants to reduce time-to-first-token while still benefiting from batching. Which adjustment is most appropriate?
HardOther domains
All NCP-GENL exam domains
Frequently asked questions
- What does the Model Deployment domain cover on the NCP-GENL exam?
- You must read a Triton config or TensorRT-LLM setup and diagnose why a model fails to load or runs out of memory, then choose the right fix. The single most important thing: verify instance-to-GPU mapping and memory budget before tuning throughput.
- How many questions are in this domain?
- This page lists all 37 Model Deployment questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Model Deployment questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.