NCP-GENL · domain
Model Optimization
This domain covers how NVIDIA TensorRT and TensorRT-LLM turn trained generative models into fast, low-precision inference engines. Questions target quantization (FP16/INT8), kernel selection, KV cache management, layer fusion, and graph surgery, testing whether you know which optimization runs at build time versus runtime and why each exists.
Focused practice
Practice Model Optimization questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Model Optimization
Be able to explain how TensorRT and TensorRT-LLM optimize generative models: Tensor Cores for FP16/INT8, layer fusion at build time, paged KV cache for memory, and graph surgery for unsupported ops. The key is knowing which optimization happens when and why.
Tensor Core acceleration of FP16 and INT8 matrix multiplications in TensorRT engines
TensorRT-LLM paged KV cache and block manager reducing memory fragmentation during generation
Layer fusion combining convolution, bias, and activation into single optimized kernels
Graph surgery to remove unsupported ops or replace subgraphs before TensorRT engine build
Watch out for
Common Model Optimization exam traps
- ▸Assuming INT8 quantization is lossless; it needs calibration and can degrade accuracy without validation.
- ▸Confusing build-time optimizations like layer fusion with runtime scheduling such as KV cache paging.
- ▸Believing TensorRT can build any graph as-is, ignoring unsupported layers that require graph surgery or plugins.
Question index
All Model Optimization questions (40)
Click any question to see the full explanation, or start a practice session above.
A team must serve a 70B model on a single 80 GB GPU for an internal assistant with modest concurrency. Full FP16 weights will not fit alongside the KV cache for the target context length. They want to keep accuracy loss minimal and are willing to spend additional build time. Which approach best fits these constraints?
Hard2Refer to the exhibit. The TensorRT build process fails with a memory limit error. Which configuration adjustment is most likely to resolve this build-time error?
Hard3An engineer is tasked with optimizing a model that performs poorly due to excessive memory access latency. Which TensorRT optimization strategy specifically targets this issue?
Medium4A team is preparing a Llama-based chatbot for production and wants to reduce GPU memory and latency without retraining. They decide to apply post-training quantization. Which TensorRT-LLM workflow correctly produces an INT8 or FP8 quantized engine from an existing FP16 checkpoint?
Easy5Refer to the exhibit. An engineer notices that the TensorRT engine takes an excessively long time to build. What is the most likely cause, and how can it be mitigated?
Hard6Which optimization technique specifically helps to manage the memory bandwidth bottleneck during the autoregressive decoding phase of an LLM?
Medium7What is the primary role of an inference 'calibrator' when converting a model to INT8 precision?
Easy8What is the primary function of the 'Triton Model Analyzer' in an optimization workflow?
Easy9A team is building a TensorRT-LLM engine for a 7B model that must serve both single-turn short prompts and long multi-turn conversations with a shared system prompt. They want to maximize reuse of computation across requests without changing model weights. Which TWO techniques should they enable? (Choose two.)
Medium10What is the primary advantage of using a 'Quantization Aware Training' (QAT) approach over post-training quantization for LLMs?
Medium11Which THREE of the following are primary benefits of using PagedAttention in NVIDIA TensorRT-LLM deployments?
Medium12A team is deploying a Llama 2 13B model with NVIDIA TensorRT-LLM on a single A100 40GB GPU. They need to serve 32 concurrent requests with a maximum sequence length of 4096 tokens. They observe that the GPU runs out of memory during inference. Which configuration parameter should they adjust to control the maximum GPU memory allocated for the KV cache?
Medium13Which hardware architecture feature is specifically leveraged by TensorRT to accelerate FP16 and INT8 matrix multiplications?
Medium14A developer is using NVIDIA TensorRT-LLM to optimize a GPT-based model for inference. They want to reduce the model's memory footprint and improve throughput without retraining. Which two techniques can be applied during the TensorRT-LLM build process to achieve these goals? (Choose two.)
Hard15A developer has a fine-tuned Llama-family model in Hugging Face format and wants to run it with NVIDIA TensorRT-LLM on an H100. Which artifact must be produced before the runtime can execute the model?
Easy16A team is deploying a large language model using NVIDIA TensorRT-LLM on a multi-GPU node with NVLink. They want to minimize inter-GPU communication overhead during inference. Which parallelism strategy should they use to achieve this?
Hard17An engineer is using TensorRT-LLM to serve a model that occasionally receives prompts far longer than the typical 512 tokens, up to 8K tokens. With the default engine settings, requests near 8K fail with a cache capacity error while short requests succeed. Which configuration change most directly resolves this without rebuilding for a single worst-case shape?
Hard18Refer to the exhibit. An engineer receives these logs while converting a Transformer model to a TensorRT engine. What is the most appropriate action to resolve this build failure?
Medium19An engineer is using NVIDIA TensorRT-LLM to optimize an LLM for inference. They want to reduce the memory footprint of the KV cache during long-context generation. Which TWO techniques are supported by TensorRT-LLM to achieve this? (Choose two.)
Medium20Refer to the exhibit. An engineer observes that GPU memory utilization is high, but the GPU is frequently idling. How does the provided Triton configuration optimize the inference pipeline?
Medium21An engineer is tuning a TensorRT-LLM deployment of a 7B model for a latency-sensitive API. Profiling shows that time per output token is higher than expected and that many small kernels run back to back with gaps between them. Which TWO changes are most likely to reduce the per-token latency by cutting kernel launch overhead and redundant memory traffic? (Choose two.)
Medium22A team is deploying a 13B-parameter chatbot on a single NVIDIA A10G GPU (24 GB VRAM). The model's weights are stored in FP16, and the runtime runs out of memory during KV cache allocation under concurrent user sessions. They must keep answer quality essentially unchanged while maximizing concurrent sessions. Which optimization should they apply first?
Medium23Refer to the exhibit. The engineer is attempting to deploy on an NVIDIA Orin platform but encounters a runtime error. What is the most likely cause of the failure?
Hard24An engineer is deploying a 13B-parameter LLM with TensorRT-LLM on a single NVIDIA A100 40GB GPU. The FP16 engine requires 26GB for weights, but during generation the KV cache grows beyond remaining memory, causing out-of-memory errors. The team wants to maximize concurrent requests without retraining. Which optimization should they apply first?
Medium25An engineer is using TensorRT-LLM to serve a chatbot model. They observe that the time to first token (TTFT) is high, but subsequent tokens are generated quickly. Which optimization should they prioritize to reduce TTFT?
Medium26When optimizing a Generative AI model using NVIDIA TensorRT-LLM, which component is primarily responsible for managing the KV cache to minimize memory fragmentation?
Hard27An engineer is optimizing an LLM for inference with NVIDIA TensorRT-LLM and wants to reduce both memory footprint and latency without retraining the model. Which two techniques should they apply? (Choose two.)
Medium28When optimizing a model using NVIDIA TensorRT, what is the primary benefit of enabling 'layer fusion' during the optimization process?
Medium29An engineer is optimizing a BERT-like model for inference using NVIDIA TensorRT. They want to reduce latency further by using lower precision without significant accuracy loss. Which TensorRT precision mode should they choose to enable INT8 inference while maintaining accuracy through calibration?
Easy30In the context of model optimization, why is 'graph surgery' sometimes required before building a TensorRT engine?
Medium31Refer to the exhibit. The deployment is facing memory allocation errors during peak load. Based on the error log, what is the most effective configuration change to resolve the issue while keeping the model architecture constant?
Hard32A team has built a TensorRT-LLM engine for a 70B model on four NVIDIA H100 GPUs using tensor parallelism. They now need to serve the same model on a single H100 for a development environment, accepting higher latency. What is the most appropriate approach?
Hard33A developer wants to reduce the disk and memory footprint of a fine-tuned 70B model before serving it with TensorRT-LLM, and is willing to accept a small, measurable quality drop that they will validate with an evaluation harness. Which approach best matches that requirement?
Easy34An enterprise deploying a large language model on an NVIDIA A100 GPU experiences high memory bandwidth bottlenecks during autoregressive token generation. Which optimization technique specifically addresses this memory-bound phase by merging element-wise operations and reducing global memory round-trips?
Medium35Which strategy is most effective for optimizing an LLM that is too large to fit into a single GPU's VRAM?
Medium36An engineer is using NVIDIA TensorRT-LLM's in-flight batching to serve a mix of short and very long prompts. They observe that GPU utilization drops and latency for short requests spikes whenever a long prompt is admitted. Which mechanism should they tune to prevent long sequences from monopolizing the batch?
Hard37A team is serving a 70B-parameter LLM with TensorRT-LLM in a multi-tenant environment where requests arrive with widely varying prompt lengths and generation lengths. During load testing, they observe that throughput collapses when a long-context request is scheduled alongside many short requests, and GPU memory fragmentation causes intermittent out-of-memory errors even though total free memory appears sufficient. Which TensorRT-LLM runtime configuration change most directly addresses both the throughput collapse and the memory fragmentation?
Hard38Which optimization method should be prioritized when the model's inference performance is bottlenecked by the CPU-to-GPU data transfer overhead?
Medium39An engineer is using NVIDIA TensorRT to optimize a Transformer model for inference on an NVIDIA A100 GPU. They want to maximize throughput while ensuring that the model runs correctly with varying input sequence lengths. Which TensorRT feature should they configure to allow the engine to handle different input shapes at runtime?
Medium40Which THREE factors should be considered when choosing an optimal batch size for LLM inference on NVIDIA GPUs?
MediumOther domains
All NCP-GENL exam domains
Frequently asked questions
- What does the Model Optimization domain cover on the NCP-GENL exam?
- Be able to explain how TensorRT and TensorRT-LLM optimize generative models: Tensor Cores for FP16/INT8, layer fusion at build time, paged KV cache for memory, and graph surgery for unsupported ops. The key is knowing which optimization happens when and why.
- How many questions are in this domain?
- This page lists all 40 Model Optimization questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Model Optimization questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.