Courseiva

CCNA Model Optimization Questions

40 questions · Model Optimization · All types, answers revealed

1
MCQhard

A team must serve a 70B model on a single 80 GB GPU for an internal assistant with modest concurrency. Full FP16 weights will not fit alongside the KV cache for the target context length. They want to keep accuracy loss minimal and are willing to spend additional build time. Which approach best fits these constraints?

A.Split the model with tensor parallelism across two GPUs even though only one is available.
B.Apply FP8 quantization to both weights and activations and rebuild the engine.
C.Keep FP16 weights and reduce the KV cache by capping the context length far below the requirement.
D.Use INT4 weight-only quantization with a group size and calibration set chosen for the model.
AnswerD

Weight-only INT4 quantization shrinks the weight footprint roughly fourfold versus FP16, letting a 70B model fit on one 80 GB GPU while leaving the KV cache and activations in higher precision. Group-wise scales limit accuracy loss, and the extra build and calibration time is acceptable given the stated willingness, making this the best fit.

Why this answer

The binding constraint is fitting a 70B model plus KV cache on one 80 GB GPU while preserving accuracy. Weight-only INT4 quantization cuts weight memory roughly fourfold while keeping activations and cache at higher precision, and group-wise scales with a representative calibration set limit degradation. The team's willingness to accept longer build time matches the calibration and engine-build cost this method requires.

Exam trap

The trap here is reaching for a parallelism or context-shortening workaround when the real constraint is weight footprint on a single device that weight-only low-bit quantization directly addresses.

2
MCQhard

Refer to the exhibit. The TensorRT build process fails with a memory limit error. Which configuration adjustment is most likely to resolve this build-time error?

A.Reducing the number of hidden layers
B.Increasing the builder's workspace memory limit
C.Switching to a lower precision inference
D.Disabling the TensorRT engine cache
AnswerB

The builder requires a temporary workspace to allocate memory for different kernel implementations. If this memory limit is too small for a complex model, the builder will fail during the optimization phase. Increasing the workspace size provides the headroom required to compute the optimal execution plan for the model.

Why this answer

The build error indicates that the workspace memory required for the optimization process exceeds the available VRAM or the limit specified by the build configuration. By explicitly increasing the workspace size in the builder configuration, the engineer provides the necessary scratchpad memory for TensorRT to evaluate and test various kernel implementations, allowing the builder to successfully complete the optimization process for all layers of the model.

Exam trap

Candidates often confuse workspace memory with the total GPU VRAM or system RAM, attempting to optimize the model size instead of the builder's temporary scratchpad allocation during the optimization phase.

3
MCQmedium

An engineer is tasked with optimizing a model that performs poorly due to excessive memory access latency. Which TensorRT optimization strategy specifically targets this issue?

A.Precision Calibration
B.Kernel Auto-Tuning
C.Layer Fusion
D.Weight Pruning
AnswerC

Layer fusion combines sequential operations into a single kernel, reducing the overhead of reading and writing intermediate tensors to global memory. This is the most effective approach for mitigating memory access latency because it keeps necessary data within the GPU's fast-access on-chip caches during processing.

Why this answer

Layer fusion is the primary strategy for reducing memory access latency. By combining multiple kernels into one, the engine keeps data in the high-speed cache of the GPU instead of constantly writing to and reading from slow global VRAM. This minimizes the time spent waiting for data movement, which is usually the dominant bottleneck for modern neural networks.

Exam trap

Candidates often confuse layer fusion with model pruning or quantization, failing to realize that fusion is specifically about minimizing the movement of data between GPU registers and global memory.

4
MCQeasy

A team is preparing a Llama-based chatbot for production and wants to reduce GPU memory and latency without retraining. They decide to apply post-training quantization. Which TensorRT-LLM workflow correctly produces an INT8 or FP8 quantized engine from an existing FP16 checkpoint?

A.Build the FP16 engine first, then apply a post-build conversion script that rewrites the engine plan to INT8 weights.
B.Run the quantization toolkit to produce a quantized checkpoint, then build the TensorRT-LLM engine with the appropriate quantization flags.
C.Convert the checkpoint to ONNX with INT8 operators and load it directly into the TensorRT-LLM runtime without an engine build.
D.Enable automatic mixed precision in the runtime and let the inference server choose INT8 kernels dynamically at request time.
AnswerB

TensorRT-LLM expects a quantized checkpoint and quantization metadata before the engine build, so the supported path is to quantize the model with the provided toolkit and then pass the quantization mode during engine construction. This preserves accuracy through calibrated scaling factors and lets the builder select INT8 or FP8 kernels. It requires no retraining and matches the stated goal of reducing memory and latency in production.

Why this answer

Post-training quantization in TensorRT-LLM is a two-stage process: first produce a quantized checkpoint with calibration or scaling data, then build the engine with the matching quantization flags so the builder selects INT8 or FP8 kernels. This avoids retraining and yields memory and latency improvements. The other options either attempt unsupported post-build modification, rely on runtime precision switching that does not exist, or bypass the required engine compilation step.

Exam trap

The trap here is assuming quantization can be applied to an already-built engine or selected dynamically at runtime, when it must be baked into the checkpoint and engine build.

5
MCQhard

Refer to the exhibit. An engineer notices that the TensorRT engine takes an excessively long time to build. What is the most likely cause, and how can it be mitigated?

A.Decrease the workspace size
B.Use a calibration cache
C.Use an explicit tactic sources list
D.Increase the batch size
AnswerC

Specifying the tactic sources allows the builder to skip certain search paths or limit the number of kernels it tests. By narrowing the scope of the tactic search, the build process completes significantly faster while still producing a highly optimized engine that utilizes the target GPU's capabilities effectively.

Why this answer

The long build time is likely due to the large workspace allocation combined with a large search space for the optimization tactics. TensorRT tests a variety of kernel implementations to find the fastest one. To reduce build time, the engineer can limit the 'tactic selection' or use a profile-based build where common shapes are pre-recorded, preventing the engine from exhaustively searching every possible configuration for all input shapes.

Exam trap

Candidates often guess that the model is too large or the GPU is underpowered, missing the fact that TensorRT's exhaustive search for optimization tactics is the primary cause of slow build times.

6
MCQmedium

Which optimization technique specifically helps to manage the memory bandwidth bottleneck during the autoregressive decoding phase of an LLM?

A.Operator Fusion
B.KV Cache Quantization
C.Weight Pruning
D.Dynamic Batching
AnswerB

Quantizing the KV cache reduces the size of the data that must be read from memory for every single token generated. By using 8-bit or 4-bit representations for the cache, the system significantly decreases the memory bandwidth requirement, allowing for faster generation and higher concurrency on the same hardware.

Why this answer

During decoding, the model must read all KV cache values for every generated token. This is memory-bandwidth bound. Techniques like Quantized KV Cache (reducing the precision of cached tokens) and PagedAttention significantly reduce the amount of data moved between the VRAM and the compute units, effectively alleviating the memory bottleneck and allowing for faster token generation rates.

Exam trap

Candidates often select general model quantization, overlooking that the KV cache is a specific, massive memory bottleneck in LLM decoding that requires specialized cache-specific quantization techniques to resolve.

7
MCQeasy

What is the primary role of an inference 'calibrator' when converting a model to INT8 precision?

A.It prunes zero-value weights
B.It calculates the optimal quantization scaling factors
C.It re-trains the model for higher accuracy
D.It optimizes the GPU kernel execution path
AnswerB

The calibrator processes a representative dataset to find the best range for quantization. It calculates the scaling factors that map the FP32 distribution into the INT8 range, minimizing information loss. This is the core function of the calibration step in the post-training quantization pipeline for TensorRT.

Why this answer

The calibrator identifies the optimal range for mapping floating-point activations to 8-bit integers without significant loss of information. It analyzes the distribution of input data to determine the scaling factors. Because INT8 has a limited dynamic range compared to FP32, the calibrator is essential to prevent overflow or underflow and to ensure that the quantized model remains accurate for the target inference tasks.

Exam trap

Test-takers frequently assume the calibrator alters model weights during training or performs the entire quantization process, missing its specific role in generating scaling factors for FP32-to-INT8 conversion.

8
MCQeasy

What is the primary function of the 'Triton Model Analyzer' in an optimization workflow?

A.It converts models to TensorRT
B.It automatically quantizes the model
C.It benchmarks configuration trade-offs
D.It manages model version control
AnswerC

The Model Analyzer benchmarks various deployment configurations (like batch size and instance count) to identify the settings that offer the best performance. This allows engineers to make data-driven decisions when deploying models to ensure they maximize resource utilization while staying within latency and throughput constraints.

Why this answer

The Model Analyzer is a tool designed to explore the trade-offs between throughput, latency, and memory usage for different model configurations. It automatically runs benchmarks with varying batch sizes and instance counts, providing developers with empirical data to find the optimal deployment parameters that meet their specific service level agreements for generative AI applications.

Exam trap

Examinees often mistake the Triton Model Analyzer for a profiling tool that measures training convergence, confusing runtime inference deployment trade-offs with model training metrics.

9
Multi-Selectmedium

A team is building a TensorRT-LLM engine for a 7B model that must serve both single-turn short prompts and long multi-turn conversations with a shared system prompt. They want to maximize reuse of computation across requests without changing model weights. Which TWO techniques should they enable? (Choose two.)

Select 2 answers
A.Enable paged KV cache with block reuse so freed sequence blocks return to the pool for other requests.
B.Apply INT4 weight-only quantization to the attention projection matrices to shrink the cache footprint.
C.Enable KV cache reuse so identical prompt prefixes share cached key/value blocks.
D.Set the builder to use strongly typed engines so the plugin graph is fully deterministic.
E.Increase the beam width so multiple candidate continuations are evaluated per request.
AnswersA, C

The paged KV cache allocates fixed-size blocks and, with block reuse, returns blocks from finished sequences to a shared pool rather than fragmenting memory. Mixed short and long prompts then coexist efficiently, raising the effective batch size and preventing allocation failures without any change to the underlying model weights.

Why this answer

Sharing a system prompt and multi-turn history means many requests share long identical prefixes, and mixed short and long prompts create variable cache pressure. KV cache reuse skips recomputing cached prefixes, while the paged KV cache with block reuse keeps memory defragmented and available to new sequences. Together they cut redundant prefill and raise concurrency without touching model weights.

Exam trap

The trap here is conflating weight memory with KV cache memory, so an engineer reaches for weight quantization when the reuse problem is really about activations and cache block lifecycle.

10
MCQmedium

What is the primary advantage of using a 'Quantization Aware Training' (QAT) approach over post-training quantization for LLMs?

A.It eliminates the need for a calibration dataset
B.It produces significantly smaller model files
C.It results in higher accuracy for quantized models
D.It avoids the use of TensorRT builder
AnswerC

By simulating quantization during training, the weights are optimized to minimize the impact of precision loss. This process allows the model to learn and compensate for the rounding errors inherent in low-precision formats, leading to significantly better accuracy compared to post-training quantization methods for complex language models.

Why this answer

QAT incorporates quantization errors into the training loop, allowing the model to adapt its weights to the loss of precision. Unlike post-training quantization, which can cause significant accuracy degradation for complex LLMs, QAT ensures that the model remains robust despite the restricted dynamic range of the INT8 or FP8 format. This results in superior final inference accuracy, making it the preferred choice for high-stakes generative applications where precision is critical.

Exam trap

Candidates often incorrectly identify 'smaller model size' or 'faster training time' as the primary advantage, whereas QAT is specifically designed to mitigate the accuracy loss inherent in quantization.

11
MCQmedium

Which THREE of the following are primary benefits of using PagedAttention in NVIDIA TensorRT-LLM deployments?

A.Elimination of external memory fragmentation.
B.Reduction in KV cache memory overhead.
C.Support for longer context lengths.
D.Faster calculation of attention scores.
E.Direct hardware support in the GPU scheduler.
AnswerA, B, C

PagedAttention treats the KV cache as a collection of fixed-size blocks, similar to virtual memory in operating systems. This structure prevents external fragmentation because any free block can be allocated to any request, ensuring that memory usage remains highly efficient even when servicing requests of varying lengths and concurrent patterns.

Why this answer

PagedAttention is a critical optimization for LLM inference, solving the problem of memory fragmentation. By managing memory in non-contiguous pages, it allows the system to allocate only what is needed, reducing memory waste significantly. This enables higher batch sizes, better GPU utilization, and the ability to serve longer sequences without running out of memory, which are essential for maintaining high performance in production-grade LLM serving environments.

Exam trap

Candidates often mistake PagedAttention as a compute acceleration technique rather than a memory management optimization, leading them to select incorrect benefits like reduced floating-point operations.

12
MCQmedium

A team is deploying a Llama 2 13B model with NVIDIA TensorRT-LLM on a single A100 40GB GPU. They need to serve 32 concurrent requests with a maximum sequence length of 4096 tokens. They observe that the GPU runs out of memory during inference. Which configuration parameter should they adjust to control the maximum GPU memory allocated for the KV cache?

A.max_batch_size
B.max_num_tokens
C.kv_cache_free_gpu_memory_fraction
D.max_input_len
AnswerC

kv_cache_free_gpu_memory_fraction specifies the fraction of free GPU memory that TensorRT-LLM can use for the KV cache. Lowering this value reduces the KV cache size, preventing out-of-memory errors when serving long sequences with multiple concurrent requests. This parameter directly controls the memory pool dedicated to the KV cache, making it the correct adjustment.

Why this answer

The KV cache in TensorRT-LLM is allocated from a memory pool whose size is controlled by kv_cache_free_gpu_memory_fraction. Lowering this fraction reduces the KV cache footprint, resolving out-of-memory errors when serving multiple long sequences. Other parameters like max_batch_size or max_num_tokens influence scheduling but do not directly bound the KV cache memory pool.

Exam trap

The trap here is assuming that max_batch_size directly limits KV cache memory, when in fact the KV cache pool is separately governed by kv_cache_free_gpu_memory_fraction.

13
MCQmedium

Which hardware architecture feature is specifically leveraged by TensorRT to accelerate FP16 and INT8 matrix multiplications?

A.CPU AVX-512 vector instructions
B.NVIDIA Tensor Cores
C.Shared memory buffers in L1 cache
D.Global memory coalescing hardware
AnswerB

Tensor Cores are hardware circuits designed for high-speed matrix multiplications in FP16, INT8, and other low-precision formats. TensorRT optimizes the execution graph to ensure that large matrix multiplications are dispatched to these units, providing the massive performance gains seen in modern deep learning inference workloads.

Why this answer

NVIDIA Tensor Cores are specialized hardware units designed to perform mixed-precision matrix multiply-accumulate operations in a single cycle. TensorRT automatically detects the presence of these cores and maps compute-intensive layers to them. By utilizing Tensor Cores, the model achieves significantly higher throughput and reduced latency compared to using standard CUDA cores, which perform operations at a lower efficiency per clock cycle for matrix math.

Exam trap

Candidates often confuse Tensor Cores with CUDA cores, assuming that standard CUDA cores are the primary driver for mixed-precision acceleration rather than the specialized hardware units built for matrix math.

14
Multi-Selecthard

A developer is using NVIDIA TensorRT-LLM to optimize a GPT-based model for inference. They want to reduce the model's memory footprint and improve throughput without retraining. Which two techniques can be applied during the TensorRT-LLM build process to achieve these goals? (Choose two.)

Select 2 answers
A.Enabling FP8 precision for both weights and activations on Hopper GPUs
B.Pruning the model by removing entire attention heads
C.Weight-only quantization (e.g., INT8 or INT4) for linear layers
D.Knowledge distillation from a larger teacher model
E.Using a larger batch size to amortize memory overhead
AnswersA, C

FP8 precision on Hopper GPUs (e.g., H100) reduces memory usage and increases throughput by using 8-bit floating point for weights and activations. TensorRT-LLM supports FP8 quantization, which can be applied during build without retraining. This directly addresses memory footprint and throughput, making it a correct choice.

Why this answer

Weight-only quantization and FP8 precision are both build-time optimizations in TensorRT-LLM that reduce memory footprint and improve throughput without retraining. Weight-only quantization lowers weight precision, while FP8 leverages Hopper GPU capabilities for both weights and activations. The other options either require retraining or increase memory usage.

Exam trap

The trap here is assuming that any memory-reduction technique like pruning is suitable, but pruning often requires retraining and is not a standard TensorRT-LLM build option.

15
MCQeasy

A developer has a fine-tuned Llama-family model in Hugging Face format and wants to run it with NVIDIA TensorRT-LLM on an H100. Which artifact must be produced before the runtime can execute the model?

A.A TensorRT engine built from the model's network definition for the target GPU architecture.
B.An ONNX graph exported with dynamic axes and a matching runtime configuration JSON.
C.A PyTorch TorchScript trace of the full forward pass saved as a .pt file.
D.A quantized GGUF file generated with a community conversion script.
AnswerA

TensorRT-LLM compiles the model graph into a serialized TensorRT engine that is specialized for the target GPU compute capability, chosen precision, and parallelism layout. The runtime loads and executes this engine; without it there is nothing for the executor to run. Building the engine is therefore the mandatory step between a Hugging Face checkpoint and inference.

Why this answer

TensorRT-LLM executes a serialized TensorRT engine that is compiled for a specific GPU architecture, precision, and parallelism configuration. After converting the Hugging Face checkpoint into the TensorRT-LLM checkpoint format, the developer must build the engine for the H100. Only that engine can be loaded by the runtime, so it is the required artifact before any inference can occur.

Exam trap

The trap here is confusing an interchange or community format such as ONNX or GGUF with the compiled TensorRT engine that the TensorRT-LLM runtime actually loads and executes.

16
MCQhard

A team is deploying a large language model using NVIDIA TensorRT-LLM on a multi-GPU node with NVLink. They want to minimize inter-GPU communication overhead during inference. Which parallelism strategy should they use to achieve this?

A.Pipeline parallelism (PP)
B.Expert parallelism (EP) in a Mixture-of-Experts model
C.Data parallelism (DP)
D.Tensor parallelism (TP)
AnswerD

Tensor parallelism splits individual layers across GPUs, requiring frequent all-reduce operations for activations. NVLink provides high bandwidth and low latency, making TP efficient. By using TP, the team can minimize communication overhead compared to other strategies that might use slower interconnects or require more synchronization, thus achieving the goal.

Why this answer

Tensor parallelism splits layers across GPUs and uses all-reduce for activations. On a system with NVLink, the high bandwidth and low latency of NVLink make TP efficient, minimizing communication overhead. Pipeline parallelism introduces bubbles, data parallelism replicates the model, and expert parallelism is for MoE models and incurs all-to-all communication.

Exam trap

The trap here is assuming that pipeline parallelism minimizes communication because it reduces frequency, but it introduces pipeline bubbles and does not leverage NVLink as effectively as tensor parallelism.

17
MCQhard

An engineer is using TensorRT-LLM to serve a model that occasionally receives prompts far longer than the typical 512 tokens, up to 8K tokens. With the default engine settings, requests near 8K fail with a cache capacity error while short requests succeed. Which configuration change most directly resolves this without rebuilding for a single worst-case shape?

A.Set the maximum input length and maximum sequence length to accommodate 8K tokens and size the KV cache pool accordingly.
B.Enable INT8 KV cache quantization so each token's cached keys and values consume fewer bytes.
C.Enable chunked context prefill so long prompts are processed in multiple smaller prefill passes.
D.Reduce the maximum batch size to 1 so all KV cache blocks are available to a single long request.
AnswerA

TensorRT-LLM engines have fixed maximum input and sequence length settings, and the KV cache pool must be large enough to hold the longest sequence the engine was built for. Raising these limits and sizing the pool for 8K tokens lets long prompts allocate their blocks, while shorter requests continue to use only the blocks they need.

Why this answer

The failure is a capacity limit tied to the engine's configured maximum input and sequence lengths and the KV cache pool sized for them. Raising those limits to cover 8K tokens and sizing the pool to match lets long prompts reserve the blocks they need, while paged allocation means short requests still consume only a few blocks, so a single worst-case build shape is unnecessary.

Exam trap

The trap here is reaching for memory-saving features like KV cache quantization or prefill chunking when the real constraint is that the engine was built with a maximum sequence length too small for the incoming prompts.

18
MCQmedium

Refer to the exhibit. An engineer receives these logs while converting a Transformer model to a TensorRT engine. What is the most appropriate action to resolve this build failure?

A.Reduce the input batch size
B.Increase the builder workspace size
C.Switch to a smaller base model
D.Disable INT8 calibration
AnswerB

The error specifically mentions workspace memory limit exceeded. By increasing the memory budget provided to the TensorRT builder, the engine can allocate sufficient scratch space to test various optimized kernels for the Attention and MatMul operations, thereby successfully completing the build process for the engine.

Why this answer

The error indicates that the workspace memory allocated to the TensorRT builder is insufficient to explore the search space of kernels for specific operations. Increasing the workspace size allows the builder to allocate larger temporary memory buffers, which is necessary for complex Transformer operations that require high-memory intermediate calculations to optimize effectively on NVIDIA hardware.

Exam trap

Candidates often assume the error is due to a lack of overall GPU VRAM, attempting to reduce model precision or batch size instead of specifically increasing the builder's workspace memory allocation.

19
Multi-Selectmedium

An engineer is using NVIDIA TensorRT-LLM to optimize an LLM for inference. They want to reduce the memory footprint of the KV cache during long-context generation. Which TWO techniques are supported by TensorRT-LLM to achieve this? (Choose two.)

Select 2 answers
A.Use TensorRT's INT8 calibration for the KV cache.
B.Enable paged KV cache with block sharing.
C.Apply structured pruning to remove attention heads.
D.Use FP8 KV cache quantization.
E.Enable sliding window attention in the model architecture.
AnswersB, D

TensorRT-LLM's paged KV cache divides the cache into blocks and allows sharing of identical blocks across sequences, reducing memory when multiple requests share prefixes. This is a core feature for efficient memory management in long-context scenarios and is enabled by default in recent versions.

Why this answer

TensorRT-LLM provides native support for FP8 KV cache quantization and paged KV cache with block sharing. FP8 quantization halves cache memory, while paged cache with block sharing reduces duplication across sequences. Both are configuration-time features that do not require model changes.

The other options involve model modifications or misapply TensorRT features, making them incorrect for this scenario.

Exam trap

The trap here is confusing model-level optimizations like pruning or sliding window attention with runtime KV cache optimizations that TensorRT-LLM directly supports.

20
MCQmedium

Refer to the exhibit. An engineer observes that GPU memory utilization is high, but the GPU is frequently idling. How does the provided Triton configuration optimize the inference pipeline?

A.It forces the model to use INT8 precision
B.It enables concurrent execution on the GPU
C.It partitions the GPU for dedicated memory
D.It enables dynamic batching for the requests
AnswerB

Setting 'count: 2' instructs Triton to launch two instances of the model on the specified GPU. This allows the server to process multiple requests in parallel, which is a standard method to improve throughput and keep the GPU busy during periods where one instance might be blocked.

Why this answer

By setting the instance count to 2 on a single GPU, the configuration enables concurrent model execution. This allows Triton to schedule multiple inference requests to be processed simultaneously on the same hardware. This overlap helps to hide memory latency and pipeline stalls, effectively increasing the utilization of the GPU and raising the overall throughput of the deployment.

Exam trap

Candidates often misinterpret high VRAM utilization with optimal performance, missing the fact that unutilized compute cycles indicate poor concurrency and request scheduling bottlenecks.

21
Multi-Selectmedium

An engineer is tuning a TensorRT-LLM deployment of a 7B model for a latency-sensitive API. Profiling shows that time per output token is higher than expected and that many small kernels run back to back with gaps between them. Which TWO changes are most likely to reduce the per-token latency by cutting kernel launch overhead and redundant memory traffic? (Choose two.)

Select 2 answers
A.Enable a debug synchronization after every kernel so profiling timings are accurate.
B.Enable CUDA Graphs so the decoding iteration is captured and replayed as a single graph launch.
C.Use TensorRT-LLM fused multi-head attention and fused GEMM plugins instead of separate elementwise kernels.
D.Raise the maximum sequence length in the engine build to accommodate the longest possible prompt.
E.Switch the model weights to FP32 to avoid any dequantization kernels in the decode path.
AnswersB, C

CUDA Graphs record the whole sequence of kernels in a decoding step and replay them with one launch, removing per-kernel launch latency and the CPU-side gaps visible in the profile. For a latency-sensitive decode loop with many small kernels, this directly reduces time per output token without changing numerics or requiring retraining.

Why this answer

The profile shows many short kernels with idle gaps, which is the classic signature of launch overhead and unfused elementwise work. CUDA Graphs collapse the decode iteration into a single replayable launch, and fused attention and GEMM plugins merge operations that would otherwise round-trip intermediates through HBM. Together they reduce both launch count and memory traffic, which are the two costs identified in the profile.

Exam trap

The trap here is chasing sequence length or precision settings when the profile clearly shows launch-bound behavior that graph capture and kernel fusion are designed to eliminate.

22
MCQmedium

A team is deploying a 13B-parameter chatbot on a single NVIDIA A10G GPU (24 GB VRAM). The model's weights are stored in FP16, and the runtime runs out of memory during KV cache allocation under concurrent user sessions. They must keep answer quality essentially unchanged while maximizing concurrent sessions. Which optimization should they apply first?

A.Convert the checkpoint to TF32 and rebuild the TensorRT-LLM engine with builder optimization level 5.
B.Apply INT8 weight-only quantization to the linear layers with NVIDIA TensorRT-LLM, keeping activations in FP16.
C.Increase the paged KV cache block size from 16 to 128 tokens per block and disable block reuse.
D.Enable CUDA graph capture for the decoder and set the maximum batch size to the peak observed concurrency.
AnswerB

Weight-only INT8 quantization halves the memory consumed by the model's linear-layer weights (roughly 26 GB FP16 becomes ~13 GB INT8) while leaving activation precision untouched, which preserves output quality closely. The freed VRAM is then available for KV cache blocks, directly lifting the concurrent session ceiling on the 24 GB A10G.

Why this answer

The bottleneck is persistent weight memory (FP16 weights dominate VRAM), which leaves too little room for paged KV cache blocks under concurrency. Weight-only INT8 quantization in TensorRT-LLM roughly halves the linear-layer weight footprint while keeping activations in FP16, preserving quality. That reclaimed VRAM goes directly to KV cache, raising the number of simultaneous sessions without retraining or a new GPU.

Exam trap

The trap here is assuming that any precision reduction degrades quality unacceptably, when weight-only INT8 with FP16 activations is specifically designed to preserve accuracy while cutting the dominant memory consumer.

23
MCQhard

Refer to the exhibit. The engineer is attempting to deploy on an NVIDIA Orin platform but encounters a runtime error. What is the most likely cause of the failure?

A.The workspace memory is too large
B.The DLA hardware is not enabled on this device
C.Unsupported operators are being targeted for the DLA
D.FP16 precision is not supported on DLA
AnswerC

The DLA is a specialized hardware accelerator with limited operator support compared to the GPU. LLMs often use complex or custom operations that the DLA cannot execute. Forcing these operations onto the DLA via the CLI flags will cause the builder to fail because it cannot map the graph.

Why this answer

The inclusion of '--dla 0' forces the engine to run on the Deep Learning Accelerator (DLA) core. Many complex LLM operations, such as specific activation functions or advanced attention mechanisms, are not supported by the DLA's fixed-function logic. If the model graph contains unsupported operators, the build will either fail or generate a non-functional plan, as the DLA has a more restricted operator set than the primary GPU cores.

Exam trap

Candidates often assume the error is a general memory or driver issue, failing to check if the specific operations in the model graph are actually supported by the DLA hardware architecture.

24
MCQmedium

An engineer is deploying a 13B-parameter LLM with TensorRT-LLM on a single NVIDIA A100 40GB GPU. The FP16 engine requires 26GB for weights, but during generation the KV cache grows beyond remaining memory, causing out-of-memory errors. The team wants to maximize concurrent requests without retraining. Which optimization should they apply first?

A.Reduce the KV cache memory footprint by quantizing the KV cache to INT8 while keeping model weights in FP16.
B.Increase the batch size and sequence length limits so the scheduler can pack more requests into the same memory.
C.Enable tensor parallelism across two A100 GPUs to split both weights and KV cache across devices.
D.Convert the model weights from FP16 to FP8 using post-training quantization to free memory for the KV cache.
AnswerA

Quantizing the KV cache to INT8 halves its per-token memory, directly addressing the growth that causes OOM during generation while leaving weights untouched, so no retraining is needed. TensorRT-LLM supports INT8 KV cache with FP16 weights, and it preserves accuracy better than aggressively quantizing weights. This is the least invasive change that increases concurrent request capacity on the constrained A100.

Why this answer

The KV cache grows linearly with sequence length and batch size, and it is the dominant memory consumer during LLM generation after weights are loaded. Quantizing only the KV cache to INT8 halves its footprint without retraining and without changing weight precision, directly relieving the OOM. This keeps the model on the existing A100 while allowing more concurrent requests, which matches the stated goal.

Exam trap

The trap here is assuming that weight quantization is the only lever for memory reduction, when KV cache quantization is often the more targeted fix for generation-time OOM.

25
MCQmedium

An engineer is using TensorRT-LLM to serve a chatbot model. They observe that the time to first token (TTFT) is high, but subsequent tokens are generated quickly. Which optimization should they prioritize to reduce TTFT?

A.Enable in-flight batching to increase GPU utilization.
B.Optimize the prefill phase with kernel fusion and faster attention.
C.Increase the KV cache size to avoid evictions.
D.Use a larger batch size for the prefill phase.
AnswerB

TTFT is dominated by the prefill phase, where the entire prompt is processed. Optimizing this phase with fused kernels, faster attention implementations like FlashAttention, and reducing memory overhead directly cuts TTFT. TensorRT-LLM provides optimized prefill kernels and supports FlashAttention, making this the most effective approach.

Why this answer

Time to first token is primarily determined by the prefill phase, which processes the input prompt. Optimizing this phase with fused kernels, efficient attention mechanisms, and reduced memory latency directly reduces TTFT. TensorRT-LLM includes highly optimized prefill kernels and supports FlashAttention, which significantly speeds up the prefill computation.

Other options target throughput or cache management, not TTFT.

Exam trap

The trap here is conflating throughput optimizations like in-flight batching with latency reductions for the first token, which are distinct performance metrics.

26
MCQhard

When optimizing a Generative AI model using NVIDIA TensorRT-LLM, which component is primarily responsible for managing the KV cache to minimize memory fragmentation?

A.The TensorRT Builder
B.The Paged KV Cache Manager
C.The Model Parallelism Engine
D.The Triton Inference Server
AnswerB

The Paged KV Cache Manager is specifically designed to handle the dynamic growth of sequences in LLMs. By allocating non-contiguous memory blocks, it prevents the fragmentation that would otherwise limit the ability of the system to handle multiple concurrent requests with varying sequence lengths during generation.

Why this answer

TensorRT-LLM utilizes PagedAttention or similar paged memory management techniques to handle the Key-Value (KV) cache. By dividing the cache into fixed-size blocks, the framework avoids the fragmentation issues inherent in contiguous memory allocation. This ensures that the memory is used efficiently during autoregressive generation, allowing for larger context windows and higher concurrent user capacity on the GPU.

Exam trap

Students frequently confuse general GPU memory allocators with specialized LLM components, mistakenly choosing standard cache options instead of paged memory managers designed for attention states.

27
Multi-Selectmedium

An engineer is optimizing an LLM for inference with NVIDIA TensorRT-LLM and wants to reduce both memory footprint and latency without retraining the model. Which two techniques should they apply? (Choose two.)

Select 2 answers
A.Quantize model weights to INT8 or FP8 using post-training quantization with calibration.
B.Enable paged KV cache with block-based memory management.
C.Disable attention and rely solely on feed-forward layers to speed up decoding.
D.Increase the number of transformer layers to improve parallelism across the GPU.
E.Retrain the model with quantization-aware training to learn optimal scaling factors.
AnswersA, B

Post-training weight quantization reduces the bytes per parameter, cutting memory footprint and enabling faster low-precision GEMM kernels on supported GPUs. It requires only calibration data, not retraining, which matches the constraint. This directly lowers both memory and latency, making it a primary optimization for production LLM inference in TensorRT-LLM.

Why this answer

Post-training quantization reduces weight memory and enables faster low-precision kernels, while paged KV cache manages generation-time memory efficiently and improves batching. Together they lower footprint and latency without retraining. The rejected options either require retraining, increase model size, or break the architecture, so they do not fit the no-retraining constraint or the optimization goals.

Exam trap

The trap here is selecting quantization-aware training because it sounds like a precision optimization, when it violates the no-retraining requirement.

28
MCQmedium

When optimizing a model using NVIDIA TensorRT, what is the primary benefit of enabling 'layer fusion' during the optimization process?

A.It increases the overall model weight precision
B.It reduces global memory access overhead
C.It automatically prunes redundant parameters
D.It replaces floating-point math with integer math
AnswerB

By fusing multiple operations into a single kernel, intermediate tensors do not need to be written back to global VRAM. This significantly reduces memory bandwidth consumption, as the fused kernel can pass data directly between operations using fast on-chip memory or registers, leading to improved inference latency.

Why this answer

Layer fusion reduces the overhead of launching multiple small GPU kernels by combining them into a single, optimized kernel. This strategy minimizes the total number of reads and writes to global VRAM, which is often the bottleneck in modern deep learning. By keeping intermediate data in high-speed registers or shared memory, the model spends less time waiting for memory access and more time performing actual floating-point operations.

Exam trap

Candidates often believe layer fusion is primarily about reducing the number of parameters or the total model size, rather than optimizing the execution flow to minimize global memory round-trips.

29
MCQeasy

An engineer is optimizing a BERT-like model for inference using NVIDIA TensorRT. They want to reduce latency further by using lower precision without significant accuracy loss. Which TensorRT precision mode should they choose to enable INT8 inference while maintaining accuracy through calibration?

A.FP16
B.FP32
C.TF32
D.INT8
AnswerD

INT8 precision mode in TensorRT enables 8-bit integer inference, which significantly reduces latency and memory usage. To maintain accuracy, TensorRT uses a calibrator to determine scaling factors from a representative dataset. This matches the engineer's goal of using INT8 with calibration to minimize accuracy loss, making it the correct choice.

Why this answer

INT8 precision mode in TensorRT allows inference using 8-bit integers, reducing latency and memory footprint. To preserve accuracy, a calibrator computes scaling factors from a calibration dataset. The other precisions (FP32, FP16, TF32) do not provide INT8 inference and do not use calibration for quantization.

Exam trap

The trap here is confusing TF32 with INT8, as TF32 is often mentioned for Tensor Core acceleration but is not an 8-bit integer format and does not use calibration.

30
MCQmedium

In the context of model optimization, why is 'graph surgery' sometimes required before building a TensorRT engine?

A.To increase the model's parameter count
B.To resolve unsupported operator compatibility
C.To improve the quality of training data
D.To automatically quantize the model to INT8
AnswerB

Often, models exported from frameworks contain operators that are not directly supported by TensorRT. Graph surgery allows engineers to replace these nodes with equivalent, supported sub-graphs. This ensures the builder can successfully create an engine without encountering errors or falling back to inefficient CPU execution paths.

Why this answer

Graph surgery is the process of modifying the model's computational graph to replace unsupported or inefficient operations with more optimized, TensorRT-compliant versions. Some frameworks export nodes that TensorRT cannot parse or optimize effectively. By manually editing the graph to fuse operations or simplify the structure before building, engineers can ensure that the engine builder produces a high-performance execution plan that effectively utilizes the underlying hardware capabilities.

Exam trap

Candidates frequently believe graph surgery is used for fine-tuning weights or pruning dead neurons, confusing model compression techniques with structural compatibility fixes for the inference engine.

31
MCQhard

Refer to the exhibit. The deployment is facing memory allocation errors during peak load. Based on the error log, what is the most effective configuration change to resolve the issue while keeping the model architecture constant?

A.Switch to a smaller model size.
B.Increase the GPU clock frequency.
C.Enable PagedAttention in the TensorRT-LLM runtime.
D.Reduce the batch size to one.
AnswerC

PagedAttention is designed to solve exactly this type of KV cache allocation failure. By moving from static allocation to paged block allocation, the system can utilize memory more efficiently and pack more requests into the same GPU memory footprint, effectively eliminating the OOM errors caused by inflexible memory management.

Why this answer

The error log indicates that the system is running out of memory because the KV cache is allocated statically, which often leads to fragmentation or over-allocation. Enabling PagedAttention is the standard NVIDIA-recommended solution for this scenario. It allows the system to manage KV cache memory dynamically, reclaiming space from finished requests and efficiently allocating blocks for new tokens, thereby preventing the out-of-memory (OOM) errors caused by static allocation.

Exam trap

Candidates often suggest increasing GPU VRAM or reducing batch sizes, which are reactive measures. They overlook PagedAttention, which is the specific architectural solution for KV cache fragmentation and memory OOM errors.

32
MCQhard

A team has built a TensorRT-LLM engine for a 70B model on four NVIDIA H100 GPUs using tensor parallelism. They now need to serve the same model on a single H100 for a development environment, accepting higher latency. What is the most appropriate approach?

A.Use pipeline parallelism with a single stage to split the model across the GPU's SMs automatically.
B.Load the existing four-GPU engine on a single H100 and set the runtime to use only one device.
C.Convert the engine to a TensorRT plan file with FP32 precision so it can use unified memory on the H100.
D.Rebuild the engine with tensor parallelism set to one and apply weight quantization so the model fits in a single GPU's memory.
AnswerD

Tensor parallelism is fixed at engine build time, so moving from four GPUs to one requires rebuilding the engine with a tensor parallel size of one. A 70B FP16 model needs roughly 140GB, which exceeds a single H100's 80GB, so weight quantization is necessary to fit. This approach produces a valid single-GPU engine while accepting the expected latency increase.

Why this answer

Tensor parallelism is a build-time property, so changing the number of GPUs requires rebuilding the engine with tensor parallel size one. Because a 70B model in FP16 exceeds a single H100's memory, weight quantization is also needed to fit. The other options either attempt to reuse an incompatible engine, misapply parallelism concepts, or increase precision, none of which yield a working single-GPU deployment.

Exam trap

The trap here is assuming a multi-GPU engine can be restricted at runtime to fewer devices, when parallelism is compiled into the plan and requires a rebuild.

33
MCQeasy

A developer wants to reduce the disk and memory footprint of a fine-tuned 70B model before serving it with TensorRT-LLM, and is willing to accept a small, measurable quality drop that they will validate with an evaluation harness. Which approach best matches that requirement?

A.Perform post-training quantization to INT4 or INT8 weights and validate with perplexity and task metrics.
B.Retrain the model from scratch using quantization-aware training with fake-quant nodes in the forward pass.
C.Convert the checkpoint to FP8 and rely on the runtime to upcast to FP16 for every matmul.
D.Increase the tensor parallelism degree so each GPU holds a smaller shard of the FP16 weights.
AnswerA

Post-training quantization compresses an already fine-tuned checkpoint without retraining, cutting weight storage and memory roughly 2x for INT8 and 4x for INT4. Because the developer accepts a small, measurable quality drop and will validate it with an evaluation harness, PTQ is the appropriate, low-effort match for the stated goal.

Why this answer

The requirement is a smaller weight footprint with an accepted, validated quality trade-off and no retraining. Post-training quantization of the fine-tuned checkpoint to INT4 or INT8 delivers that reduction immediately and pairs naturally with an evaluation harness to confirm the quality delta stays within tolerance. Retraining-based methods exceed the stated scope.

Exam trap

The trap here is defaulting to quantization-aware training as the highest-quality option when the scenario explicitly accepts a small, measurable quality drop and does not provide for retraining.

34
MCQmedium

An enterprise deploying a large language model on an NVIDIA A100 GPU experiences high memory bandwidth bottlenecks during autoregressive token generation. Which optimization technique specifically addresses this memory-bound phase by merging element-wise operations and reducing global memory round-trips?

A.Applying static INT8 post-training quantization to all linear layers.
B.Enabling TensorRT-LLM custom kernel fusion for multi-head attention and activation blocks.
C.Increasing the global batch size to maximize arithmetic intensity.
D.Switching from FlashAttention to standard vanilla self-attention mechanisms.
AnswerB

TensorRT-LLM provides highly optimized, fused CUDA kernels specifically designed to eliminate redundant global memory round-trips for operations like multi-head attention, layer normalization, and activations. This directly accelerates memory-bound autoregressive text generation workloads on NVIDIA GPUs.

Why this answer

Kernel fusion combines multiple successive GPU operations, such as bias additions and activations, into a single CUDA kernel. This drastically reduces high-latency global memory read and write operations, directly mitigating the memory bandwidth bottleneck characteristic of autoregressive transformer decoding phases on NVIDIA hardware.

Exam trap

Candidates frequently select generic model pruning or distillation, which reduces parameter count but fails to directly address the specific memory bandwidth bottlenecks caused by repetitive global memory access in autoregressive decoding.

35
MCQmedium

Which strategy is most effective for optimizing an LLM that is too large to fit into a single GPU's VRAM?

A.Increase the inference batch size
B.Apply dynamic quantization to the weights
C.Implement tensor parallelism
D.Convert the model to a CPU-only format
AnswerC

Tensor parallelism involves splitting individual weight matrices of the model across multiple GPUs. This allows the model to reside on multiple devices, effectively pooling their VRAM and compute power. It is the primary method for scaling large models beyond the hardware limits of a single GPU device.

Why this answer

Model parallelism, specifically tensor parallelism, is the standard approach for splitting a large model across multiple GPUs. By partitioning individual layers across different devices, the computation can be distributed, and the total memory requirement is spread proportionally. This allows for the deployment of models that exceed the capacity of a single GPU, enabling high-performance inference for massive parameter models that would otherwise be impossible to load.

Exam trap

Students often select data parallelism or pipeline parallelism incorrectly, failing to recognize that single layers exceeding VRAM require splitting the actual weight matrices across multiple devices.

36
MCQhard

An engineer is using NVIDIA TensorRT-LLM's in-flight batching to serve a mix of short and very long prompts. They observe that GPU utilization drops and latency for short requests spikes whenever a long prompt is admitted. Which mechanism should they tune to prevent long sequences from monopolizing the batch?

A.Adjust the maximum number of tokens processed per iteration and the KV cache block allocation policy to limit how much context a single request can consume.
B.Enable tensor parallelism so the long prompt's attention computation is split across multiple GPUs.
C.Switch from in-flight batching to a static batching scheme that groups requests by similar prompt length.
D.Increase the maximum batch size so more short requests can be admitted alongside the long prompt.
AnswerA

In-flight batching processes a token budget per iteration; if one long prompt consumes most of that budget, short requests stall. Tuning the per-iteration token limit and KV cache block allocation constrains how much context a single sequence can occupy, allowing the scheduler to interleave short and long requests fairly. This directly targets the observed utilization drop and latency spike without changing model precision or hardware.

Why this answer

In-flight batching improves throughput by mixing prefill and decode work, but a very long prompt can consume the entire per-iteration token budget, starving shorter sequences. Limiting tokens per iteration and controlling KV cache block allocation prevents any single request from monopolizing the batch. This restores GPU utilization and keeps short-request latency stable while still allowing long prompts to complete progressively.

Exam trap

The trap here is assuming that more GPUs or larger batches fix latency fairness, when the real constraint is per-iteration token budget allocation within the scheduler.

37
MCQhard

A team is serving a 70B-parameter LLM with TensorRT-LLM in a multi-tenant environment where requests arrive with widely varying prompt lengths and generation lengths. During load testing, they observe that throughput collapses when a long-context request is scheduled alongside many short requests, and GPU memory fragmentation causes intermittent out-of-memory errors even though total free memory appears sufficient. Which TensorRT-LLM runtime configuration change most directly addresses both the throughput collapse and the memory fragmentation?

A.Increase the TensorRT-LLM build max_batch_size and max_seq_len to their maximum possible values so the engine supports every request shape.
B.Switch the engine from FP16 to INT8 precision using a calibration dataset that matches the production prompt distribution.
C.Reduce the number of concurrent client connections at the load balancer so that only one request is processed by the GPU at a time.
D.Enable the paged KV cache with a tuned block size and configure the batch scheduler to use in-flight batching with a per-iteration token budget.
AnswerD

Paged KV cache allocates KV memory in fixed-size blocks from a shared pool, eliminating the external fragmentation that causes OOM. In-flight batching lets the scheduler add new requests and retire finished ones every iteration, so short requests are not blocked behind a long one. Together they directly resolve both the memory fragmentation and the head-of-line blocking that collapses throughput.

Why this answer

The paged KV cache solves external fragmentation by allocating KV memory in uniform blocks from a shared pool, which prevents the OOM that occurs when a long-context request needs a large contiguous region. In-flight batching with a token budget lets the scheduler continuously admit and retire requests each iteration, so short requests no longer wait behind a long generation. These two runtime settings together address both symptoms without sacrificing accuracy or throughput.

Exam trap

The trap here is assuming that increasing max_batch_size and max_seq_len gives the engine more flexibility, when in fact it reserves more memory and does nothing to prevent fragmentation or scheduling head-of-line blocking.

38
MCQmedium

Which optimization method should be prioritized when the model's inference performance is bottlenecked by the CPU-to-GPU data transfer overhead?

A.Quantization
B.GPU-based Pre-processing
C.Model Pruning
D.Operator Fusion
AnswerB

Moving pre-processing to the GPU eliminates the need to transfer intermediate results over the slow PCIe bus. By executing the full pipeline on the GPU device, the system avoids the overhead of host-to-device transfers, which is critical for achieving low-latency inference in real-time generative applications.

Why this answer

When data transfer between the host and GPU is the bottleneck, the most effective strategy is to move the pre-processing logic onto the GPU itself. Using CUDA-accelerated kernels to perform operations like tokenization or normalization directly on the GPU avoids moving data across the PCIe bus, which is a slow operation that stalls the GPU during the inference pipeline.

Exam trap

Candidates often suggest optimizing model weights or using faster interconnects like NVLink, ignoring that moving pre-processing to the GPU is the most direct way to eliminate PCIe bus stalling.

39
MCQmedium

An engineer is using NVIDIA TensorRT to optimize a Transformer model for inference on an NVIDIA A100 GPU. They want to maximize throughput while ensuring that the model runs correctly with varying input sequence lengths. Which TensorRT feature should they configure to allow the engine to handle different input shapes at runtime?

A.Static shapes with fixed dimensions
B.Dynamic shapes with optimization profiles
C.INT8 calibration with a representative dataset
D.Multiple engines for each possible input length
AnswerB

Dynamic shapes with optimization profiles allow a TensorRT engine to accept input tensors of different dimensions at runtime. The profiles define minimum, optimal, and maximum shapes for each input, enabling the engine to select the best kernel for the actual shape. This is essential for handling varying sequence lengths in Transformer models without rebuilding the engine.

Why this answer

Dynamic shapes with optimization profiles enable a single TensorRT engine to handle inputs of different sizes by defining shape ranges. This allows the engine to optimize for the actual input shape at runtime, which is crucial for Transformer models with variable sequence lengths. Static shapes, multiple engines, or INT8 calibration do not provide this flexibility.

Exam trap

The trap here is confusing quantization (INT8 calibration) with shape flexibility, but calibration only affects precision, not the ability to accept different input dimensions.

40
Multi-Selectmedium

Which THREE factors should be considered when choosing an optimal batch size for LLM inference on NVIDIA GPUs?

Select 3 answers
A.Available GPU VRAM
B.Inference Latency SLA
C.Total Compute Throughput
D.Model Training Loss
E.CPU Clock Speed
AnswersA, B, C

The available VRAM is a hard constraint on the batch size. As the batch size increases, the memory required to store intermediate activations and the KV cache also grows. If the batch size is too large, the system will encounter out-of-memory errors, making memory capacity the primary limiting factor.

Why this answer

Selecting an optimal batch size involves a balance between hardware saturation, memory capacity, and latency requirements. Larger batches improve throughput by better utilizing GPU compute cores but increase memory demand and individual request latency. The goal is to reach the highest throughput possible while staying within the VRAM limit and meeting the required latency SLA for the specific application.

Exam trap

Candidates often ignore the Inference Latency SLA, focusing solely on maximizing compute throughput. They fail to realize that an overly large batch size can lead to unacceptable latency for real-time applications.

Ready to test yourself?

Try a timed practice session using only Model Optimization questions.