NCP-GENL · domain
GPU Acceleration and Optimization
This domain covers profiling, scaling, and inference optimization on NVIDIA hardware. Questions test whether you can read Nsight Systems traces, distinguish PCIe from NVLink bottlenecks, choose the right library (TensorRT, NeMo, Triton), and pick techniques that cut LLM latency or all-reduce overhead.
Focused practice
Practice GPU Acceleration and Optimization questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about GPU Acceleration and Optimization
Be able to profile a GPU workload in Nsight Systems, name the bottleneck, and match it to the right NVIDIA tool: TensorRT for inference, NeMo for training, NCCL for collectives. The single most important thing is correctly identifying whether PCIe, NVLink, or compute is the real limiter.
Reading Nsight Systems traces to separate PCIe transfer stalls from NVLink and SM utilization limits
Selecting TensorRT for inference optimization and deployment, versus NeMo for training and Triton for serving
Diagnosing data-parallel all-reduce overhead with NCCL and choosing overlap, larger batches, or gradient accumulation
Reducing LLM inference latency via quantization, KV-cache tuning, and kernel or batching changes
Watch out for
Common GPU Acceleration and Optimization exam traps
- ▸Blaming PCIe for every stall; NVLink or SM occupancy limits often dominate and Nsight Systems metrics reveal which is actually saturated.
- ▸Confusing TensorRT with NeMo or Triton: TensorRT optimizes and deploys engines, NeMo trains, Triton serves, and the exam expects the right tool named.
- ▸Assuming quantization or smaller batches always cut latency; KV-cache growth and memory-bandwidth limits can dominate, so measure before optimizing.
Question index
All GPU Acceleration and Optimization questions (34)
Click any question to see the full explanation, or start a practice session above.
A team is training a large language model on 8 NVIDIA A100 GPUs using PyTorch's DistributedDataParallel (DDP). Profiling shows that all GPUs are frequently idle, waiting for gradient synchronization. The network interconnect between nodes is a 100 Gb Ethernet with TCP/IP, and the model has 13 billion parameters. What is the most effective optimization to reduce the idle time?
Medium2A developer is optimizing a BERT-based model for inference on an NVIDIA T4 GPU using TensorRT. The model has a fixed input sequence length of 128. Profiling shows that the kernel execution time is high due to many small operations. Which TensorRT feature should they use to reduce kernel launch overhead and improve latency?
Medium3An engineer is optimizing a large language model for inference on NVIDIA GPUs using TensorRT-LLM. They want to reduce the memory footprint of the KV cache to support longer context lengths and more concurrent requests. Which two techniques should they implement? (Choose two.)
Hard4When profiling an application with NVIDIA Nsight Systems, you notice a long gap between kernel execution blocks on the GPU timeline. What is the most likely cause?
Hard5An engineer is profiling a CUDA kernel and notices that the achieved occupancy is low, leading to underutilization of the GPU. The kernel uses a large number of registers per thread, limiting the number of resident warps. Which optimization should be attempted first to improve occupancy?
Easy6Why is 'Pinned Memory' (page-locked) essential for high-performance data transfers between host and GPU?
Medium7When deploying a model using NVIDIA TensorRT, what is the primary benefit of the 'Engine Building' phase?
Easy8An engineer is optimizing a Transformer-based LLM for inference on an NVIDIA A100 GPU. The model uses FP16 precision, but during generation, the GPU's Tensor Cores are underutilized, and latency is higher than expected. Profiling reveals that many small matrix multiplications are executed sequentially. Which technique is most effective to improve Tensor Core utilization and reduce latency?
Hard9You are optimizing a ResNet-50 model on an NVIDIA A100. Which precision-based optimization will yield the highest throughput without significant accuracy loss?
Medium10When profiling an application with NVIDIA Nsight Systems, which TWO metrics are most critical to identify if an application is limited by the PCIe bus?
Hard11Which NVIDIA library is primarily used for optimizing and deploying deep learning inference models?
Easy12Refer to the exhibit. The model performance is inconsistent. What is the most likely reason for the performance variability under load?
Medium13An engineer is deploying a large language model for real-time inference on an NVIDIA A100 GPU. The model uses FP16 weights but the inference server must handle variable-length input sequences. Profiling shows that the GPU spends significant time on memory-bound operations and that kernel launches are frequent. Which optimization is most appropriate to reduce latency while maintaining accuracy?
Hard14A team is optimizing a large language model for inference on NVIDIA GPUs. They want to reduce the memory footprint of the model to fit on a single GPU with limited VRAM. Which two techniques are most appropriate for reducing memory usage during inference? (Choose two.)
Medium15A team is training a large language model using NVIDIA DGX A100 nodes with 8 GPUs per node. They observe that GPU utilization is high on all GPUs, but the training throughput scales poorly when adding more nodes. Profiling shows that the communication time during all-reduce operations increases significantly with node count. Which of the following optimizations is most likely to improve scaling efficiency?
Medium16A team is serving a 13B-parameter LLM on a single NVIDIA A100 80GB GPU. During generation, they observe that the GPU compute utilization stays below 20% while memory bandwidth utilization is near saturation. They want to improve throughput without changing the model architecture. Which optimization is most appropriate?
Medium17A company is running a large language model inference service on NVIDIA GPUs. They observe that GPU memory is nearly full, limiting the batch size and thus throughput. The model weights are stored in FP16, and the KV cache consumes a significant portion of memory. Which technique can reduce memory usage while maintaining model accuracy and enabling larger batch sizes?
Medium18When deploying a large language model on NVIDIA H100 GPUs using TensorRT-LLM, which TWO configuration strategies are most effective for improving KV cache efficiency and memory utilization?
Hard19An engineer is optimizing a large language model for inference on NVIDIA GPUs and wants to reduce memory usage to fit a larger model or increase batch size. Which two techniques are most effective for reducing GPU memory consumption during inference? (Choose two.)
Hard20A developer is using NVIDIA Nsight Systems to profile a PyTorch training loop on an NVIDIA GPU. They notice significant gaps between kernel executions and want to identify whether the bottleneck is CPU-side or GPU-side. Which Nsight Systems feature should they use to visualize the CPU and GPU timelines together?
Easy21A developer is optimizing a generative AI model for inference on NVIDIA GPUs. They want to reduce memory footprint and improve throughput without sacrificing accuracy. Which two techniques should they apply? (Choose two.)
Medium22A developer is using NVIDIA TensorRT to optimize a BERT-based model for inference. They notice that the engine performs poorly on variable-length input sequences because it was built with a single optimization profile for a fixed sequence length. What should they do to improve performance across different sequence lengths?
Medium23An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. They want to maximize throughput for a chatbot workload with variable-length inputs and outputs. Which of the following techniques should they implement to achieve the highest throughput while maintaining acceptable latency?
Hard24In the context of NVIDIA Tensor Cores, what is the primary benefit of using BF16 (Bfloat16) over FP16 during model training and inference?
Easy25Which hardware component of an NVIDIA GPU is most responsible for accelerating matrix-multiply-accumulate (MMA) operations used in transformer layers?
Medium26When using the NVIDIA Collective Communications Library (NCCL), what does 'AllReduce' specifically optimize for in a multi-GPU training configuration?
Medium27A team is training a large language model on a single NVIDIA H100 GPU. They observe that training throughput is significantly lower than expected, and profiling with Nsight Systems shows long periods where the GPU is idle waiting for data. The data loading pipeline uses the default PyTorch DataLoader with num_workers=0 and no pinned memory. Which change is most likely to improve GPU utilization?
Medium28Which of the following describes the purpose of 'Kernel Fusion' in the context of optimizing a Deep Learning inference pipeline?
Easy29What is the primary function of the 'TensorRT' optimization engine in the NVIDIA AI software stack?
Medium30A team is training a 13B-parameter LLM on 8 NVIDIA A100 GPUs using NVIDIA NeMo. They observe that the all-reduce communication during data-parallel training consumes nearly 40% of each iteration. Which of the following changes is most likely to reduce this communication overhead while preserving convergence?
Medium31Refer to the exhibit. An engineer receives this timeout error during a CUDA kernel execution. What is the most appropriate first step to diagnose the resource contention?
Medium32Which TWO of the following techniques are best suited for reducing the latency of LLM inference on NVIDIA GPUs?
Medium33An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. The model uses multi-head attention with a sequence length of 4096. During inference, the GPU's Tensor Cores are underutilized, and the kernel launch overhead is high due to many small operations. Which optimization should be applied to improve Tensor Core utilization and reduce overhead?
Hard34An engineer is training a large language model with pipeline parallelism across four NVIDIA GPUs. They observe that GPU utilization is low and training throughput is limited by idle time during pipeline bubbles. Which technique is most effective to reduce pipeline bubbles and improve utilization?
HardOther domains
All NCP-GENL exam domains
Frequently asked questions
- What does the GPU Acceleration and Optimization domain cover on the NCP-GENL exam?
- Be able to profile a GPU workload in Nsight Systems, name the bottleneck, and match it to the right NVIDIA tool: TensorRT for inference, NeMo for training, NCCL for collectives. The single most important thing is correctly identifying whether PCIe, NVLink, or compute is the real limiter.
- How many questions are in this domain?
- This page lists all 34 GPU Acceleration and Optimization questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only GPU Acceleration and Optimization questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.