Be able to explain how TensorRT and TensorRT-LLM optimize generative models: Tensor Cores for FP16/INT8, layer fusion at build time, paged KV cache for memory, and graph surgery for unsupported ops. The key is knowing which optimization happens when and why.
Start practicing
Model Optimization — choose a session length
Free · No account required
Domain overview
This domain covers how NVIDIA TensorRT and TensorRT-LLM turn trained generative models into fast, low-precision inference engines. Questions target quantization (FP16/INT8), kernel selection, KV cache management, layer fusion, and graph surgery, testing whether you know which optimization runs at build time versus runtime and why each exists.
Exam objectives
Tensor Core acceleration of FP16 and INT8 matrix multiplications in TensorRT engines
TensorRT-LLM paged KV cache and block manager reducing memory fragmentation during generation
Layer fusion combining convolution, bias, and activation into single optimized kernels
Graph surgery to remove unsupported ops or replace subgraphs before TensorRT engine build
Assuming INT8 quantization is lossless; it needs calibration and can degrade accuracy without validation.
Confusing build-time optimizations like layer fusion with runtime scheduling such as KV cache paging.
Believing TensorRT can build any graph as-is, ignoring unsupported layers that require graph surgery or plugins.
Click any question to see the full explanation and answer options, or start a focused practice session above.
An enterprise deploying a large language model on an NVIDIA A100 GPU experiences high memory bandwidth bottlenecks during autoregressive token generation. Which optimization technique specifically addresses this memory-bound phase by merging element-wise operations and reducing global memory round-trips?
2When optimizing a model using NVIDIA TensorRT, what is the primary benefit of enabling 'layer fusion' during the optimization process?
3Refer to the exhibit. The TensorRT build process fails with a memory limit error. Which configuration adjustment is most likely to resolve this build-time error?
4Which strategy is most effective for optimizing an LLM that is too large to fit into a single GPU's VRAM?
5What is the primary role of an inference 'calibrator' when converting a model to INT8 precision?
6Which hardware architecture feature is specifically leveraged by TensorRT to accelerate FP16 and INT8 matrix multiplications?
7Refer to the exhibit. The engineer is attempting to deploy on an NVIDIA Orin platform but encounters a runtime error. What is the most likely cause of the failure?
8In the context of model optimization, why is 'graph surgery' sometimes required before building a TensorRT engine?
9What is the primary advantage of using a 'Quantization Aware Training' (QAT) approach over post-training quantization for LLMs?
10Refer to the exhibit. An engineer receives these logs while converting a Transformer model to a TensorRT engine. What is the most appropriate action to resolve this build failure?
11When optimizing a Generative AI model using NVIDIA TensorRT-LLM, which component is primarily responsible for managing the KV cache to minimize memory fragmentation?
12Refer to the exhibit. An engineer observes that GPU memory utilization is high, but the GPU is frequently idling. How does the provided Triton configuration optimize the inference pipeline?
13An engineer is tasked with optimizing a model that performs poorly due to excessive memory access latency. Which TensorRT optimization strategy specifically targets this issue?
14What is the primary function of the 'Triton Model Analyzer' in an optimization workflow?
15Which THREE factors should be considered when choosing an optimal batch size for LLM inference on NVIDIA GPUs?
16Which optimization technique specifically helps to manage the memory bandwidth bottleneck during the autoregressive decoding phase of an LLM?
17Refer to the exhibit. An engineer notices that the TensorRT engine takes an excessively long time to build. What is the most likely cause, and how can it be mitigated?
18Which optimization method should be prioritized when the model's inference performance is bottlenecked by the CPU-to-GPU data transfer overhead?
19Which THREE of the following are primary benefits of using PagedAttention in NVIDIA TensorRT-LLM deployments?
20Refer to the exhibit. The deployment is facing memory allocation errors during peak load. Based on the error log, what is the most effective configuration change to resolve the issue while keeping the model architecture constant?
21A team is serving a 70B-parameter LLM with TensorRT-LLM in a multi-tenant environment where requests arrive with widely varying prompt lengths and generation lengths. During load testing, they observe that throughput collapses when a long-context request is scheduled alongside many short requests, and GPU memory fragmentation causes intermittent out-of-memory errors even though total free memory appears sufficient. Which TensorRT-LLM runtime configuration change most directly addresses both the throughput collapse and the memory fragmentation?
22A team is deploying a Llama 2 13B model with NVIDIA TensorRT-LLM on a single A100 40GB GPU. They need to serve 32 concurrent requests with a maximum sequence length of 4096 tokens. They observe that the GPU runs out of memory during inference. Which configuration parameter should they adjust to control the maximum GPU memory allocated for the KV cache?
23An engineer is optimizing a BERT-like model for inference using NVIDIA TensorRT. They want to reduce latency further by using lower precision without significant accuracy loss. Which TensorRT precision mode should they choose to enable INT8 inference while maintaining accuracy through calibration?
24An engineer is using NVIDIA TensorRT-LLM to optimize an LLM for inference. They want to reduce the memory footprint of the KV cache during long-context generation. Which TWO techniques are supported by TensorRT-LLM to achieve this? (Choose two.)
25A developer is using NVIDIA TensorRT-LLM to optimize a GPT-based model for inference. They want to reduce the model's memory footprint and improve throughput without retraining. Which two techniques can be applied during the TensorRT-LLM build process to achieve these goals? (Choose two.)
26An engineer is using TensorRT-LLM to serve a chatbot model. They observe that the time to first token (TTFT) is high, but subsequent tokens are generated quickly. Which optimization should they prioritize to reduce TTFT?
27An engineer is using NVIDIA TensorRT to optimize a Transformer model for inference on an NVIDIA A100 GPU. They want to maximize throughput while ensuring that the model runs correctly with varying input sequence lengths. Which TensorRT feature should they configure to allow the engine to handle different input shapes at runtime?
28A team is deploying a 13B-parameter chatbot on a single NVIDIA A10G GPU (24 GB VRAM). The model's weights are stored in FP16, and the runtime runs out of memory during KV cache allocation under concurrent user sessions. They must keep answer quality essentially unchanged while maximizing concurrent sessions. Which optimization should they apply first?
29A team is deploying a large language model using NVIDIA TensorRT-LLM on a multi-GPU node with NVLink. They want to minimize inter-GPU communication overhead during inference. Which parallelism strategy should they use to achieve this?
30An engineer is deploying a 13B-parameter LLM with TensorRT-LLM on a single NVIDIA A100 40GB GPU. The FP16 engine requires 26GB for weights, but during generation the KV cache grows beyond remaining memory, causing out-of-memory errors. The team wants to maximize concurrent requests without retraining. Which optimization should they apply first?
31A team is building a TensorRT-LLM engine for a 7B model that must serve both single-turn short prompts and long multi-turn conversations with a shared system prompt. They want to maximize reuse of computation across requests without changing model weights. Which TWO techniques should they enable? (Choose two.)
32A team is preparing a Llama-based chatbot for production and wants to reduce GPU memory and latency without retraining. They decide to apply post-training quantization. Which TensorRT-LLM workflow correctly produces an INT8 or FP8 quantized engine from an existing FP16 checkpoint?
33A developer wants to reduce the disk and memory footprint of a fine-tuned 70B model before serving it with TensorRT-LLM, and is willing to accept a small, measurable quality drop that they will validate with an evaluation harness. Which approach best matches that requirement?
34An engineer is using NVIDIA TensorRT-LLM's in-flight batching to serve a mix of short and very long prompts. They observe that GPU utilization drops and latency for short requests spikes whenever a long prompt is admitted. Which mechanism should they tune to prevent long sequences from monopolizing the batch?
35An engineer is using TensorRT-LLM to serve a model that occasionally receives prompts far longer than the typical 512 tokens, up to 8K tokens. With the default engine settings, requests near 8K fail with a cache capacity error while short requests succeed. Which configuration change most directly resolves this without rebuilding for a single worst-case shape?
36An engineer is optimizing an LLM for inference with NVIDIA TensorRT-LLM and wants to reduce both memory footprint and latency without retraining the model. Which two techniques should they apply? (Choose two.)
37A team has built a TensorRT-LLM engine for a 70B model on four NVIDIA H100 GPUs using tensor parallelism. They now need to serve the same model on a single H100 for a development environment, accepting higher latency. What is the most appropriate approach?
38A developer has a fine-tuned Llama-family model in Hugging Face format and wants to run it with NVIDIA TensorRT-LLM on an H100. Which artifact must be produced before the runtime can execute the model?
39An engineer is tuning a TensorRT-LLM deployment of a 7B model for a latency-sensitive API. Profiling shows that time per output token is higher than expected and that many small kernels run back to back with gaps between them. Which TWO changes are most likely to reduce the per-token latency by cutting kernel launch overhead and redundant memory traffic? (Choose two.)
40A team must serve a 70B model on a single 80 GB GPU for an internal assistant with modest concurrency. Full FP16 weights will not fit alongside the KV cache for the target context length. They want to keep accuracy loss minimal and are willing to spend additional build time. Which approach best fits these constraints?
Be able to explain how TensorRT and TensorRT-LLM optimize generative models: Tensor Cores for FP16/INT8, layer fusion at build time, paged KV cache for memory, and graph surgery for unsupported ops. The key is knowing which optimization happens when and why.
The Courseiva NCP-GENL question bank contains 40 questions in the Model Optimization domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Model Optimization domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included