Courseiva

NCP-GENL · domain

GPU Acceleration and Optimization

This domain covers profiling, scaling, and inference optimization on NVIDIA hardware. Questions test whether you can read Nsight Systems traces, distinguish PCIe from NVLink bottlenecks, choose the right library (TensorRT, NeMo, Triton), and pick techniques that cut LLM latency or all-reduce overhead.

34 questions6 easy18 medium10 hard

Focused practice

Practice GPU Acceleration and Optimization questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about GPU Acceleration and Optimization

Be able to profile a GPU workload in Nsight Systems, name the bottleneck, and match it to the right NVIDIA tool: TensorRT for inference, NeMo for training, NCCL for collectives. The single most important thing is correctly identifying whether PCIe, NVLink, or compute is the real limiter.

Reading Nsight Systems traces to separate PCIe transfer stalls from NVLink and SM utilization limits

Selecting TensorRT for inference optimization and deployment, versus NeMo for training and Triton for serving

Diagnosing data-parallel all-reduce overhead with NCCL and choosing overlap, larger batches, or gradient accumulation

Reducing LLM inference latency via quantization, KV-cache tuning, and kernel or batching changes

Watch out for

Common GPU Acceleration and Optimization exam traps

  • ▸Blaming PCIe for every stall; NVLink or SM occupancy limits often dominate and Nsight Systems metrics reveal which is actually saturated.
  • ▸Confusing TensorRT with NeMo or Triton: TensorRT optimizes and deploys engines, NeMo trains, Triton serves, and the exam expects the right tool named.
  • ▸Assuming quantization or smaller batches always cut latency; KV-cache growth and memory-bandwidth limits can dominate, so measure before optimizing.

Question index

All GPU Acceleration and Optimization questions (34)

Click any question to see the full explanation, or start a practice session above.

1

A team is training a large language model on 8 NVIDIA A100 GPUs using PyTorch's DistributedDataParallel (DDP). Profiling shows that all GPUs are frequently idle, waiting for gradient synchronization. The network interconnect between nodes is a 100 Gb Ethernet with TCP/IP, and the model has 13 billion parameters. What is the most effective optimization to reduce the idle time?

Medium
2

A developer is optimizing a BERT-based model for inference on an NVIDIA T4 GPU using TensorRT. The model has a fixed input sequence length of 128. Profiling shows that the kernel execution time is high due to many small operations. Which TensorRT feature should they use to reduce kernel launch overhead and improve latency?

Medium
3

An engineer is optimizing a large language model for inference on NVIDIA GPUs using TensorRT-LLM. They want to reduce the memory footprint of the KV cache to support longer context lengths and more concurrent requests. Which two techniques should they implement? (Choose two.)

Hard
4

When profiling an application with NVIDIA Nsight Systems, you notice a long gap between kernel execution blocks on the GPU timeline. What is the most likely cause?

Hard
5

An engineer is profiling a CUDA kernel and notices that the achieved occupancy is low, leading to underutilization of the GPU. The kernel uses a large number of registers per thread, limiting the number of resident warps. Which optimization should be attempted first to improve occupancy?

Easy
6

Why is 'Pinned Memory' (page-locked) essential for high-performance data transfers between host and GPU?

Medium
7

When deploying a model using NVIDIA TensorRT, what is the primary benefit of the 'Engine Building' phase?

Easy
8

An engineer is optimizing a Transformer-based LLM for inference on an NVIDIA A100 GPU. The model uses FP16 precision, but during generation, the GPU's Tensor Cores are underutilized, and latency is higher than expected. Profiling reveals that many small matrix multiplications are executed sequentially. Which technique is most effective to improve Tensor Core utilization and reduce latency?

Hard
9

You are optimizing a ResNet-50 model on an NVIDIA A100. Which precision-based optimization will yield the highest throughput without significant accuracy loss?

Medium
10

When profiling an application with NVIDIA Nsight Systems, which TWO metrics are most critical to identify if an application is limited by the PCIe bus?

Hard
11

Which NVIDIA library is primarily used for optimizing and deploying deep learning inference models?

Easy
12

Refer to the exhibit. The model performance is inconsistent. What is the most likely reason for the performance variability under load?

Medium
13

An engineer is deploying a large language model for real-time inference on an NVIDIA A100 GPU. The model uses FP16 weights but the inference server must handle variable-length input sequences. Profiling shows that the GPU spends significant time on memory-bound operations and that kernel launches are frequent. Which optimization is most appropriate to reduce latency while maintaining accuracy?

Hard
14

A team is optimizing a large language model for inference on NVIDIA GPUs. They want to reduce the memory footprint of the model to fit on a single GPU with limited VRAM. Which two techniques are most appropriate for reducing memory usage during inference? (Choose two.)

Medium
15

A team is training a large language model using NVIDIA DGX A100 nodes with 8 GPUs per node. They observe that GPU utilization is high on all GPUs, but the training throughput scales poorly when adding more nodes. Profiling shows that the communication time during all-reduce operations increases significantly with node count. Which of the following optimizations is most likely to improve scaling efficiency?

Medium
16

A team is serving a 13B-parameter LLM on a single NVIDIA A100 80GB GPU. During generation, they observe that the GPU compute utilization stays below 20% while memory bandwidth utilization is near saturation. They want to improve throughput without changing the model architecture. Which optimization is most appropriate?

Medium
17

A company is running a large language model inference service on NVIDIA GPUs. They observe that GPU memory is nearly full, limiting the batch size and thus throughput. The model weights are stored in FP16, and the KV cache consumes a significant portion of memory. Which technique can reduce memory usage while maintaining model accuracy and enabling larger batch sizes?

Medium
18

When deploying a large language model on NVIDIA H100 GPUs using TensorRT-LLM, which TWO configuration strategies are most effective for improving KV cache efficiency and memory utilization?

Hard
19

An engineer is optimizing a large language model for inference on NVIDIA GPUs and wants to reduce memory usage to fit a larger model or increase batch size. Which two techniques are most effective for reducing GPU memory consumption during inference? (Choose two.)

Hard
20

A developer is using NVIDIA Nsight Systems to profile a PyTorch training loop on an NVIDIA GPU. They notice significant gaps between kernel executions and want to identify whether the bottleneck is CPU-side or GPU-side. Which Nsight Systems feature should they use to visualize the CPU and GPU timelines together?

Easy
21

A developer is optimizing a generative AI model for inference on NVIDIA GPUs. They want to reduce memory footprint and improve throughput without sacrificing accuracy. Which two techniques should they apply? (Choose two.)

Medium
22

A developer is using NVIDIA TensorRT to optimize a BERT-based model for inference. They notice that the engine performs poorly on variable-length input sequences because it was built with a single optimization profile for a fixed sequence length. What should they do to improve performance across different sequence lengths?

Medium
23

An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. They want to maximize throughput for a chatbot workload with variable-length inputs and outputs. Which of the following techniques should they implement to achieve the highest throughput while maintaining acceptable latency?

Hard
24

In the context of NVIDIA Tensor Cores, what is the primary benefit of using BF16 (Bfloat16) over FP16 during model training and inference?

Easy
25

Which hardware component of an NVIDIA GPU is most responsible for accelerating matrix-multiply-accumulate (MMA) operations used in transformer layers?

Medium
26

When using the NVIDIA Collective Communications Library (NCCL), what does 'AllReduce' specifically optimize for in a multi-GPU training configuration?

Medium
27

A team is training a large language model on a single NVIDIA H100 GPU. They observe that training throughput is significantly lower than expected, and profiling with Nsight Systems shows long periods where the GPU is idle waiting for data. The data loading pipeline uses the default PyTorch DataLoader with num_workers=0 and no pinned memory. Which change is most likely to improve GPU utilization?

Medium
28

Which of the following describes the purpose of 'Kernel Fusion' in the context of optimizing a Deep Learning inference pipeline?

Easy
29

What is the primary function of the 'TensorRT' optimization engine in the NVIDIA AI software stack?

Medium
30

A team is training a 13B-parameter LLM on 8 NVIDIA A100 GPUs using NVIDIA NeMo. They observe that the all-reduce communication during data-parallel training consumes nearly 40% of each iteration. Which of the following changes is most likely to reduce this communication overhead while preserving convergence?

Medium
31

Refer to the exhibit. An engineer receives this timeout error during a CUDA kernel execution. What is the most appropriate first step to diagnose the resource contention?

Medium
32

Which TWO of the following techniques are best suited for reducing the latency of LLM inference on NVIDIA GPUs?

Medium
33

An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. The model uses multi-head attention with a sequence length of 4096. During inference, the GPU's Tensor Cores are underutilized, and the kernel launch overhead is high due to many small operations. Which optimization should be applied to improve Tensor Core utilization and reduce overhead?

Hard
34

An engineer is training a large language model with pipeline parallelism across four NVIDIA GPUs. They observe that GPU utilization is low and training throughput is limited by idle time during pipeline bubbles. Which technique is most effective to reduce pipeline bubbles and improve utilization?

Hard

Frequently asked questions

What does the GPU Acceleration and Optimization domain cover on the NCP-GENL exam?
Be able to profile a GPU workload in Nsight Systems, name the bottleneck, and match it to the right NVIDIA tool: TensorRT for inference, NeMo for training, NCCL for collectives. The single most important thing is correctly identifying whether PCIe, NVLink, or compute is the real limiter.
How many questions are in this domain?
This page lists all 34 GPU Acceleration and Optimization questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only GPU Acceleration and Optimization questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
nvidia-ncp-genl NVIDIA-NCP-GENL gpu acceleration optimization Practice Questions