Courseiva

NCP-GENL · topic practice

Model Optimization practice questions

This domain covers how NVIDIA TensorRT and TensorRT-LLM turn trained generative models into fast, low-precision inference engines. Questions target quantization (FP16/INT8), kernel selection, KV cache management, layer fusion, and graph surgery, testing whether you know which optimization runs at build time versus runtime and why each exists.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Model Optimization

What the exam tests

What to know about Model Optimization

Be able to explain how TensorRT and TensorRT-LLM optimize generative models: Tensor Cores for FP16/INT8, layer fusion at build time, paged KV cache for memory, and graph surgery for unsupported ops. The key is knowing which optimization happens when and why.

Tensor Core acceleration of FP16 and INT8 matrix multiplications in TensorRT engines

TensorRT-LLM paged KV cache and block manager reducing memory fragmentation during generation

Layer fusion combining convolution, bias, and activation into single optimized kernels

Graph surgery to remove unsupported ops or replace subgraphs before TensorRT engine build

Watch out for

Common Model Optimization exam traps

  • ▸Assuming INT8 quantization is lossless; it needs calibration and can degrade accuracy without validation.
  • ▸Confusing build-time optimizations like layer fusion with runtime scheduling such as KV cache paging.
  • ▸Believing TensorRT can build any graph as-is, ignoring unsupported layers that require graph surgery or plugins.

Practice set

Model Optimization questions

20 questions · select your answer, then reveal the explanation

An engineer needs to deploy a large language model on resource-constrained edge hardware. Which optimization technique provides the best balance between memory footprint reduction and inference latency?

Which TWO of the following TensorRT optimization techniques are most effective for improving throughput in LLMs using the Transformer architecture?

Which THREE of the following are benefits of using NVIDIA TensorRT over raw PyTorch for LLM inference in production?

Refer to the exhibit. The builder fails during engine construction. What is the most likely cause related to the TensorRT optimization pipeline?

Exhibit

Error: [TensorRT] Incompatible dimensions. Input shape [1, 1024] cannot be broadcast to [1, 512, 1024].

Which TWO of the following are true regarding the use of 'profiles' in TensorRT for LLMs?

An engineer needs to deploy a large language model on resource-constrained edge hardware. Which optimization technique provides the highest inference speedup by reducing numerical precision while maintaining model structural integrity?

Which TWO of the following NVIDIA TensorRT features are primarily responsible for optimizing Transformer-based model performance during the build phase?

Which optimization strategy is most effective for reducing the inference latency of an LLM when the model exceeds the memory capacity of a single NVIDIA GPU?

Which THREE techniques are commonly used to optimize inference throughput in Generative AI applications deployed on NVIDIA Triton Inference Server?

Which NVIDIA technology provides the most significant boost for LLM inference by utilizing specialized hardware to accelerate non-matrix operations like Softmax and LayerNorm?

What is the primary benefit of using FP8 precision for inference in modern Transformer models on NVIDIA H100 GPUs?

Question 12mediummultiple choice
Read the full Model Optimization explanation →

An engineer needs to deploy a large language model on NVIDIA H100 GPUs while maximizing throughput for a high-concurrency inference service. Which optimization technique is most effective for reducing latency by minimizing memory bandwidth bottlenecks?

Question 13mediummultiple choice
Read the full Model Optimization explanation →

Which NVIDIA TensorRT-LLM feature is primarily used to optimize the parallel execution of the KV cache across multiple GPU instances in a multi-node setup?

When optimizing a Transformer-based model using NVIDIA TensorRT, which TWO steps are essential to enable the Fusion of layer normalization and activation kernels to improve performance?

Question 15mediummultiple choice
Read the full Model Optimization explanation →

A team wants to reduce the cold-start latency for an LLM served with TensorRT-LLM. They've already built an engine with FP16 precision and paged KV cache. Which TensorRT-LLM runtime feature should they enable to overlap the compilation of GPU kernels with the loading of model weights?

Question 16mediummultiple choice
Read the full Model Optimization explanation →

An engineer is optimizing a 13B-parameter LLM for a single NVIDIA L40S GPU using TensorRT-LLM. The model must retain high accuracy on a domain-specific question-answering benchmark, and the team has a small calibration set of 512 representative prompts. Post-training INT8 quantization causes a noticeable drop in answer quality, particularly for numeric reasoning. Which approach best preserves accuracy while still reducing memory and latency?

An engineer is optimizing a 70B parameter LLM for inference on a single NVIDIA H100 GPU with 80GB memory. They want to maximize throughput while keeping latency acceptable. Which combination of TensorRT-LLM settings should they use to fit the model and KV cache in memory?

An engineer profiles a TensorRT-LLM Llama-2-70B engine on 4x H100 GPUs with tensor parallelism and observes that decode-step latency is dominated by inter-GPU communication rather than matmul time. The engine currently uses tensor parallelism degree 4 with no pipeline parallelism. Which change most directly reduces the communication share during decoding?

Question 19mediummultiple choice
Read the full Model Optimization explanation →

A team is serving a 70B-parameter LLM with NVIDIA TensorRT-LLM on two NVIDIA H100 GPUs using tensor parallelism. Under production traffic with many concurrent users, GPU utilization drops to roughly 40% and throughput plateaus well below the target, even though per-token latency stays acceptable. Profiling shows the GPUs idle while waiting on memory transfers between the two GPUs. Which optimization should the team apply first to raise throughput?

An engineer is building a TensorRT-LLM engine for a 13B chat model on a single A100 80 GB GPU. The workload mixes short prompts with long multi-turn conversations, and the team wants to minimize end-to-end latency for interactive use without retraining. After enabling FP8 or INT8 weight-only quantization, which additional optimization most directly attacks the memory-bandwidth bottleneck during autoregressive decoding?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Model Optimization sessions

Start a Model Optimization only practice session

Every question in these sessions is drawn from the Model Optimization domain — nothing else.

Related practice questions

Related NCP-GENL topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the NCP-GENL exam test about Model Optimization?
Be able to explain how TensorRT and TensorRT-LLM optimize generative models: Tensor Cores for FP16/INT8, layer fusion at build time, paged KV cache for memory, and graph surgery for unsupported ops. The key is knowing which optimization happens when and why.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Model Optimization questions in a focused session?
Yes — the session launcher on this page draws every question from the Model Optimization domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other NCP-GENL topics?
Use the topic links above to move to related areas, or go back to the NCP-GENL question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the NCP-GENL exam covers. They are not copied from any real exam or dump site.