Be able to profile a GPU workload in Nsight Systems, name the bottleneck, and match it to the right NVIDIA tool: TensorRT for inference, NeMo for training, NCCL for collectives. The single most important thing is correctly identifying whether PCIe, NVLink, or compute is the real limiter.
Start practicing
GPU Acceleration and Optimization — choose a session length
Free · No account required
Domain overview
This domain covers profiling, scaling, and inference optimization on NVIDIA hardware. Questions test whether you can read Nsight Systems traces, distinguish PCIe from NVLink bottlenecks, choose the right library (TensorRT, NeMo, Triton), and pick techniques that cut LLM latency or all-reduce overhead.
Exam objectives
Reading Nsight Systems traces to separate PCIe transfer stalls from NVLink and SM utilization limits
Selecting TensorRT for inference optimization and deployment, versus NeMo for training and Triton for serving
Diagnosing data-parallel all-reduce overhead with NCCL and choosing overlap, larger batches, or gradient accumulation
Reducing LLM inference latency via quantization, KV-cache tuning, and kernel or batching changes
Blaming PCIe for every stall; NVLink or SM occupancy limits often dominate and Nsight Systems metrics reveal which is actually saturated.
Confusing TensorRT with NeMo or Triton: TensorRT optimizes and deploys engines, NeMo trains, Triton serves, and the exam expects the right tool named.
Assuming quantization or smaller batches always cut latency; KV-cache growth and memory-bandwidth limits can dominate, so measure before optimizing.
Click any question to see the full explanation and answer options, or start a focused practice session above.
When deploying a large language model on NVIDIA H100 GPUs using TensorRT-LLM, which TWO configuration strategies are most effective for improving KV cache efficiency and memory utilization?
2In the context of NVIDIA Tensor Cores, what is the primary benefit of using BF16 (Bfloat16) over FP16 during model training and inference?
3Refer to the exhibit. An engineer receives this timeout error during a CUDA kernel execution. What is the most appropriate first step to diagnose the resource contention?
4Which hardware component of an NVIDIA GPU is most responsible for accelerating matrix-multiply-accumulate (MMA) operations used in transformer layers?
5When profiling an application with NVIDIA Nsight Systems, which TWO metrics are most critical to identify if an application is limited by the PCIe bus?
6Which of the following describes the purpose of 'Kernel Fusion' in the context of optimizing a Deep Learning inference pipeline?
7Why is 'Pinned Memory' (page-locked) essential for high-performance data transfers between host and GPU?
8When using the NVIDIA Collective Communications Library (NCCL), what does 'AllReduce' specifically optimize for in a multi-GPU training configuration?
9What is the primary function of the 'TensorRT' optimization engine in the NVIDIA AI software stack?
10Which TWO of the following techniques are best suited for reducing the latency of LLM inference on NVIDIA GPUs?
11When deploying a model using NVIDIA TensorRT, what is the primary benefit of the 'Engine Building' phase?
12Refer to the exhibit. The model performance is inconsistent. What is the most likely reason for the performance variability under load?
13When profiling an application with NVIDIA Nsight Systems, you notice a long gap between kernel execution blocks on the GPU timeline. What is the most likely cause?
14You are optimizing a ResNet-50 model on an NVIDIA A100. Which precision-based optimization will yield the highest throughput without significant accuracy loss?
15Which NVIDIA library is primarily used for optimizing and deploying deep learning inference models?
16A team is training a large language model on 8 NVIDIA A100 GPUs using PyTorch's DistributedDataParallel (DDP). Profiling shows that all GPUs are frequently idle, waiting for gradient synchronization. The network interconnect between nodes is a 100 Gb Ethernet with TCP/IP, and the model has 13 billion parameters. What is the most effective optimization to reduce the idle time?
17A team is serving a 13B-parameter LLM on a single NVIDIA A100 80GB GPU. During generation, they observe that the GPU compute utilization stays below 20% while memory bandwidth utilization is near saturation. They want to improve throughput without changing the model architecture. Which optimization is most appropriate?
18An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. The model uses multi-head attention with a sequence length of 4096. During inference, the GPU's Tensor Cores are underutilized, and the kernel launch overhead is high due to many small operations. Which optimization should be applied to improve Tensor Core utilization and reduce overhead?
19A team is training a 13B-parameter LLM on 8 NVIDIA A100 GPUs using NVIDIA NeMo. They observe that the all-reduce communication during data-parallel training consumes nearly 40% of each iteration. Which of the following changes is most likely to reduce this communication overhead while preserving convergence?
20A team is training a large language model using NVIDIA DGX A100 nodes with 8 GPUs per node. They observe that GPU utilization is high on all GPUs, but the training throughput scales poorly when adding more nodes. Profiling shows that the communication time during all-reduce operations increases significantly with node count. Which of the following optimizations is most likely to improve scaling efficiency?
21A developer is using NVIDIA TensorRT to optimize a BERT-based model for inference. They notice that the engine performs poorly on variable-length input sequences because it was built with a single optimization profile for a fixed sequence length. What should they do to improve performance across different sequence lengths?
22A developer is optimizing a generative AI model for inference on NVIDIA GPUs. They want to reduce memory footprint and improve throughput without sacrificing accuracy. Which two techniques should they apply? (Choose two.)
23An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. They want to maximize throughput for a chatbot workload with variable-length inputs and outputs. Which of the following techniques should they implement to achieve the highest throughput while maintaining acceptable latency?
24An engineer is optimizing a Transformer-based LLM for inference on an NVIDIA A100 GPU. The model uses FP16 precision, but during generation, the GPU's Tensor Cores are underutilized, and latency is higher than expected. Profiling reveals that many small matrix multiplications are executed sequentially. Which technique is most effective to improve Tensor Core utilization and reduce latency?
25An engineer is training a large language model with pipeline parallelism across four NVIDIA GPUs. They observe that GPU utilization is low and training throughput is limited by idle time during pipeline bubbles. Which technique is most effective to reduce pipeline bubbles and improve utilization?
26A developer is optimizing a BERT-based model for inference on an NVIDIA T4 GPU using TensorRT. The model has a fixed input sequence length of 128. Profiling shows that the kernel execution time is high due to many small operations. Which TensorRT feature should they use to reduce kernel launch overhead and improve latency?
27A team is training a large language model on a single NVIDIA H100 GPU. They observe that training throughput is significantly lower than expected, and profiling with Nsight Systems shows long periods where the GPU is idle waiting for data. The data loading pipeline uses the default PyTorch DataLoader with num_workers=0 and no pinned memory. Which change is most likely to improve GPU utilization?
28An engineer is optimizing a large language model for inference on NVIDIA GPUs and wants to reduce memory usage to fit a larger model or increase batch size. Which two techniques are most effective for reducing GPU memory consumption during inference? (Choose two.)
29An engineer is optimizing a large language model for inference on NVIDIA GPUs using TensorRT-LLM. They want to reduce the memory footprint of the KV cache to support longer context lengths and more concurrent requests. Which two techniques should they implement? (Choose two.)
30An engineer is deploying a large language model for real-time inference on an NVIDIA A100 GPU. The model uses FP16 weights but the inference server must handle variable-length input sequences. Profiling shows that the GPU spends significant time on memory-bound operations and that kernel launches are frequent. Which optimization is most appropriate to reduce latency while maintaining accuracy?
31A developer is using NVIDIA Nsight Systems to profile a PyTorch training loop on an NVIDIA GPU. They notice significant gaps between kernel executions and want to identify whether the bottleneck is CPU-side or GPU-side. Which Nsight Systems feature should they use to visualize the CPU and GPU timelines together?
32An engineer is profiling a CUDA kernel and notices that the achieved occupancy is low, leading to underutilization of the GPU. The kernel uses a large number of registers per thread, limiting the number of resident warps. Which optimization should be attempted first to improve occupancy?
33A company is running a large language model inference service on NVIDIA GPUs. They observe that GPU memory is nearly full, limiting the batch size and thus throughput. The model weights are stored in FP16, and the KV cache consumes a significant portion of memory. Which technique can reduce memory usage while maintaining model accuracy and enabling larger batch sizes?
34A team is optimizing a large language model for inference on NVIDIA GPUs. They want to reduce the memory footprint of the model to fit on a single GPU with limited VRAM. Which two techniques are most appropriate for reducing memory usage during inference? (Choose two.)
Be able to profile a GPU workload in Nsight Systems, name the bottleneck, and match it to the right NVIDIA tool: TensorRT for inference, NeMo for training, NCCL for collectives. The single most important thing is correctly identifying whether PCIe, NVLink, or compute is the real limiter.
The Courseiva NCP-GENL question bank contains 34 questions in the GPU Acceleration and Optimization domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the GPU Acceleration and Optimization domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included