Courseiva

CCNA Gpu Acceleration Optimization Questions

34 questions · Gpu Acceleration Optimization topic · All types, answers revealed

1
MCQmedium

A team is training a large language model on 8 NVIDIA A100 GPUs using PyTorch's DistributedDataParallel (DDP). Profiling shows that all GPUs are frequently idle, waiting for gradient synchronization. The network interconnect between nodes is a 100 Gb Ethernet with TCP/IP, and the model has 13 billion parameters. What is the most effective optimization to reduce the idle time?

A.Increase the batch size per GPU to keep the GPUs busy during communication.
B.Replace the Ethernet interconnect with NVIDIA Mellanox InfiniBand and enable NCCL over RDMA.
C.Enable gradient compression using FP16 all-reduce to halve communication volume.
D.Use NVIDIA GPUDirect Storage to accelerate data loading from NVMe drives.
AnswerB

InfiniBand with RDMA provides higher bandwidth and lower latency than TCP/IP over Ethernet, reducing the all-reduce time for gradient synchronization. For a 13B parameter model, gradient tensors are large, and the communication overhead dominates. NCCL over RDMA bypasses the CPU and kernel network stack, significantly improving throughput and lowering idle time.

Why this answer

The idle time is caused by slow gradient synchronization over TCP/IP Ethernet. InfiniBand with RDMA and NCCL provides the necessary bandwidth and low latency to reduce all-reduce time, directly addressing the bottleneck. Other options either do not target communication or are secondary optimizations that do not resolve the fundamental interconnect limitation.

Exam trap

The trap here is assuming that increasing batch size or compressing gradients will fully resolve communication stalls, when the underlying high-latency interconnect is the true limiting factor.

2
MCQmedium

A developer is optimizing a BERT-based model for inference on an NVIDIA T4 GPU using TensorRT. The model has a fixed input sequence length of 128. Profiling shows that the kernel execution time is high due to many small operations. Which TensorRT feature should they use to reduce kernel launch overhead and improve latency?

A.Increase the batch size to amortize kernel launch overhead across more samples.
B.Use TensorRT's builder optimization level 5 to enable aggressive layer fusion.
C.Apply CUDA graphs to capture the entire inference graph and replay it with a single launch.
D.Enable FP16 precision and calibrate with a representative dataset.
AnswerC

CUDA graphs capture a sequence of kernels and their dependencies into a single graph, then replay it with one launch, drastically reducing CPU launch overhead. For a fixed-shape model like this BERT with sequence length 128, CUDA graphs are ideal because the graph can be captured once and reused. This directly addresses the many small operations causing high kernel execution time.

Why this answer

CUDA graphs are designed to reduce kernel launch overhead by capturing a static sequence of operations and replaying it as a single unit. For a fixed-shape model like the BERT model with sequence length 128, the graph can be captured once and reused across inferences. This eliminates the CPU-side launch latency for each kernel, directly improving latency.

Exam trap

The trap here is focusing on precision or fusion when the bottleneck is CPU-side kernel launch overhead, which CUDA graphs specifically target.

3
Multi-Selecthard

An engineer is optimizing a large language model for inference on NVIDIA GPUs using TensorRT-LLM. They want to reduce the memory footprint of the KV cache to support longer context lengths and more concurrent requests. Which two techniques should they implement? (Choose two.)

Select 2 answers
A.Use FP32 precision for the KV cache to avoid accuracy loss.
B.Enable multi-head attention with larger head dimension.
C.Enable paged KV cache with block-based memory allocation.
D.Use INT8 quantization for the KV cache.
E.Increase the batch size to improve memory reuse.
AnswersC, D

Paged KV cache divides the cache into fixed-size blocks, eliminating fragmentation and allowing non-contiguous storage. This enables more efficient memory utilization and supports a larger number of concurrent sequences. It is a core feature of TensorRT-LLM for high-throughput serving, directly addressing memory footprint and scalability for long contexts.

Why this answer

INT8 quantization and paged KV cache are both designed to reduce memory footprint. Quantization lowers the bit-width of stored keys/values, while paging eliminates fragmentation and allows more efficient allocation. Together, they enable longer contexts and higher concurrency without increasing GPU memory.

Exam trap

The trap here is confusing batch size increases with memory savings, when larger batches actually increase total KV cache memory.

4
MCQhard

When profiling an application with NVIDIA Nsight Systems, you notice a long gap between kernel execution blocks on the GPU timeline. What is the most likely cause?

A.The GPU is overheating and triggering thermal throttling.
B.Excessive host-side synchronization calls.
C.The PCIe bus is saturated with high-frequency data.
D.The kernel is launching with an invalid thread block size.
AnswerB

Explicit synchronization points force the CPU to wait for the GPU to finish all previous tasks before continuing. These gaps represent the time the GPU spends waiting for the CPU to process logic and issue new work, effectively serializing the pipeline and creating idle time on the GPU execution timeline.

Why this answer

Long gaps in the GPU timeline usually indicate host-side synchronization, such as explicit cudaDeviceSynchronize() calls or CPU-bound code that is stalling the kernel launch queue. In production systems, unnecessary CPU-to-GPU synchronization forces the GPU to remain idle while waiting for the CPU to catch up, directly impacting system-wide latency and violating the principle of asynchronous task pipelining necessary for optimal GPU utilization.

Exam trap

Candidates often attribute these gaps to kernel execution time or network latency, ignoring the role of the host CPU and synchronization primitives in managing the GPU task queue execution flow.

5
MCQeasy

An engineer is profiling a CUDA kernel and notices that the achieved occupancy is low, leading to underutilization of the GPU. The kernel uses a large number of registers per thread, limiting the number of resident warps. Which optimization should be attempted first to improve occupancy?

A.Use the __launch_bounds__ qualifier to limit registers per thread.
B.Enable L1 cache to reduce memory latency.
C.Increase the block size to allow more warps per block.
D.Convert the kernel to use shared memory for data reuse.
AnswerA

The __launch_bounds__ qualifier allows the programmer to specify the minimum number of blocks per multiprocessor, which guides the compiler to limit register usage. By reducing registers per thread, more warps can be resident, improving occupancy. This is a direct way to address register-limited occupancy without changing the algorithm, though it may cause spilling if overused.

Why this answer

The low occupancy is caused by high register usage per thread. Using __launch_bounds__ instructs the compiler to limit registers, allowing more warps to be resident and improving occupancy. This directly addresses the bottleneck, whereas other options do not target register pressure and thus would not effectively improve occupancy.

Exam trap

The trap here is assuming that increasing block size or enabling cache automatically improves occupancy, when the real constraint is the number of registers per thread limiting resident warps.

6
MCQmedium

Why is 'Pinned Memory' (page-locked) essential for high-performance data transfers between host and GPU?

A.It provides larger memory capacity on the GPU.
B.It allows direct DMA access to host RAM.
C.It enables automatic data compression.
D.It eliminates the need for CUDA contexts.
AnswerB

Pinned memory is locked in physical RAM, allowing the GPU to perform direct memory access (DMA) transfers without host CPU involvement. This bypasses the overhead of copying data through temporary buffers, leading to vastly improved bandwidth and lower latency for transfers between the host and GPU.

Why this answer

Pinned memory prevents the operating system from swapping data to disk, allowing the GPU to access host memory directly via DMA (Direct Memory Access). This avoids the need for the driver to copy data into intermediate buffers, which is a major source of latency in standard data pipelines. By using pinned memory, applications can achieve significantly higher transfer speeds, which is vital for real-time generative AI applications.

Exam trap

Candidates frequently assume pinned memory increases GPU compute speed directly, rather than understanding that it primarily optimizes the efficiency of data transfer between host RAM and GPU VRAM.

7
MCQeasy

When deploying a model using NVIDIA TensorRT, what is the primary benefit of the 'Engine Building' phase?

A.It converts the model to a generic portable format.
B.It automatically scales the model across multiple nodes.
C.It performs target-specific kernel selection and layer fusion.
D.It ensures the model can run on any CPU architecture.
AnswerC

The builder phase analyzes the network graph to merge redundant layers and select highly optimized CUDA kernels tailored to the specific GPU architecture. This approach maximizes hardware utilization, minimizes memory access patterns, and optimizes the execution flow to achieve peak performance compared to unoptimized, framework-native model execution.

Why this answer

The TensorRT builder phase analyzes the model graph and hardware topology to select the most efficient kernels for the target GPU. This process includes layer fusion, precision calibration, and kernel selection, which are vital for production-grade inference. By tailoring the model specifically to the underlying hardware architecture, TensorRT achieves significantly higher throughput and lower latency than executing generic framework-native code directly on the GPU.

Exam trap

Candidates often confuse the 'Engine Building' phase with the 'Inference' phase, incorrectly assuming it happens during runtime execution rather than as a pre-processing step to optimize the model graph for hardware.

8
MCQhard

An engineer is optimizing a Transformer-based LLM for inference on an NVIDIA A100 GPU. The model uses FP16 precision, but during generation, the GPU's Tensor Cores are underutilized, and latency is higher than expected. Profiling reveals that many small matrix multiplications are executed sequentially. Which technique is most effective to improve Tensor Core utilization and reduce latency?

A.Enable CUDA graphs to capture and replay the sequence of kernel launches.
B.Use NVIDIA TensorRT to quantize the model to INT8.
C.Increase the batch size to amortize kernel launch overhead.
D.Apply kernel fusion to combine multiple small operations into a single kernel.
AnswerD

Kernel fusion merges multiple small matrix multiplications and element-wise operations into a single kernel, reducing launch overhead and enabling better Tensor Core utilization by processing larger, combined workloads. This is especially effective in Transformer inference where many small GEMMs occur. By fusing operations, the GPU can execute them more efficiently, lowering latency and increasing utilization of Tensor Cores.

Why this answer

The underutilization of Tensor Cores stems from many small, sequential matrix multiplications. Kernel fusion combines these into larger operations, allowing Tensor Cores to process more data per instruction and reducing kernel launch overhead. This directly improves utilization and lowers latency, making it the most effective technique among the options for this specific bottleneck.

Exam trap

The trap here is focusing on quantization or CUDA graphs, which address different bottlenecks, rather than recognizing that small sequential operations require fusion to improve Tensor Core efficiency.

9
MCQmedium

You are optimizing a ResNet-50 model on an NVIDIA A100. Which precision-based optimization will yield the highest throughput without significant accuracy loss?

A.Full FP32 precision.
B.Mixed-precision (FP16/BF16).
C.Aggressive FP8 quantization.
D.Custom boolean quantization.
AnswerB

Mixed precision utilizes the high-speed Tensor Cores to perform arithmetic in FP16 or BF16 while maintaining key parts of the model in FP32. This drastically increases throughput and reduces memory bandwidth requirements, allowing for much faster inference with negligible impacts on the overall accuracy of the model, which is ideal.

Why this answer

NVIDIA A100 GPUs feature Tensor Cores that are specifically optimized for FP16 and BF16 arithmetic. Utilizing mixed-precision training or inference allows the model to leverage these high-speed units, effectively doubling throughput compared to FP32. This is the standard approach for balancing the trade-off between floating-point precision and computational speed, providing near-native accuracy while maximizing the hardware's performance capabilities in deep learning workloads.

Exam trap

Candidates often select FP32 for maximum accuracy, failing to recognize that modern NVIDIA architectures provide hardware-accelerated mixed-precision support that maintains high accuracy while significantly improving throughput.

10
Multi-Selecthard

When profiling an application with NVIDIA Nsight Systems, which TWO metrics are most critical to identify if an application is limited by the PCIe bus?

Select 2 answers
A.Host-to-Device (H2D) throughput.
B.GPU Register usage count.
C.Device-to-Host (D2H) throughput.
D.Shared memory bank conflict count.
E.SM clock frequency.
AnswersA, C

H2D throughput measures the speed at which data is sent from the host CPU to the GPU memory. If this metric hits the theoretical maximum of the PCIe bus, it confirms a bottleneck where the GPU must wait for new data to arrive before processing can begin.

Why this answer

PCIe bus saturation occurs when the transfer rate of data between the host and the GPU becomes a bottleneck. By monitoring the H2D (Host-to-Device) and D2H (Device-to-Host) transfer metrics, developers can see if data movement consumes a disproportionate amount of execution time. Identifying these spikes is crucial because offloading data to the GPU is often the slowest part of a pipeline, and minimizing these transfers is key to scaling high-performance AI.

Exam trap

Candidates often confuse PCIe throughput with GPU compute utilization or memory bandwidth metrics. They fail to realize that PCIe saturation is specifically about the data transfer rate between the host and GPU.

11
MCQeasy

Which NVIDIA library is primarily used for optimizing and deploying deep learning inference models?

A.cuBLAS.
B.TensorRT.
C.NCCL.
D.cuDNN.
AnswerB

TensorRT is the specialized SDK designed for high-performance deep learning inference. It provides essential features such as layer fusion, kernel auto-tuning, and precision reduction (INT8/FP8), which are critical for deploying models in production environments where low latency and high throughput are the primary performance and cost objectives.

Why this answer

TensorRT is NVIDIA's dedicated deep learning inference optimizer and runtime engine. It takes pre-trained models from frameworks like PyTorch or TensorFlow and optimizes them for NVIDIA hardware by fusing layers, selecting best-fit kernels, and performing precision calibration. It is the core tool for any developer needing to transition models from research training to production-level deployment with minimal latency and maximal throughput on NVIDIA infrastructure.

Exam trap

Candidates often select training frameworks like PyTorch or TensorFlow, forgetting that while those are used for development, TensorRT is the specific library designed for production inference optimization.

12
MCQmedium

Refer to the exhibit. The model performance is inconsistent. What is the most likely reason for the performance variability under load?

A.The instance count is too high for the GPU.
B.Dynamic batching lacks a max_queue_delay_microseconds parameter.
C.The GPU index is incorrectly defined in the config.
D.TensorRT plan files are inherently non-deterministic.
AnswerB

Without a max_queue_delay_microseconds, the inference server does not wait for additional requests to fill the batch. This results in requests being processed immediately, often with small batch sizes, leading to inconsistent execution times and poor GPU utilization compared to waiting briefly to aggregate requests into larger, more efficient batches.

Why this answer

The configuration uses dynamic batching with a preferred batch size range but lacks a 'max_queue_delay_microseconds' setting. Without a delay buffer, the server might dispatch batches as soon as one request arrives, leading to suboptimal under-filled batches and inconsistent latency. Adding a small delay allows the server to collect more requests, increasing throughput and stabilizing inference latency across fluctuating traffic patterns, which is vital for high-performance production deployments.

Exam trap

Candidates frequently assume the issue is related to GPU compute capacity or model size, overlooking the configuration-level settings that govern how requests are queued and processed in a production server.

13
MCQhard

An engineer is deploying a large language model for real-time inference on an NVIDIA A100 GPU. The model uses FP16 weights but the inference server must handle variable-length input sequences. Profiling shows that the GPU spends significant time on memory-bound operations and that kernel launches are frequent. Which optimization is most appropriate to reduce latency while maintaining accuracy?

A.Increase the number of CUDA streams to parallelize kernel execution.
B.Enable NVIDIA GPUDirect Storage to accelerate data loading.
C.Use NVIDIA TensorRT with dynamic batching and CUDA graphs to capture the inference step.
D.Convert the model to INT8 using TensorRT with calibration.
AnswerC

TensorRT optimizes the graph and supports dynamic shapes for variable-length sequences. Dynamic batching groups requests to improve GPU utilization, while CUDA graphs capture the sequence of kernel launches into a single graph, drastically reducing launch overhead. This combination directly addresses both memory-bound inefficiencies and frequent launches without sacrificing accuracy.

Why this answer

TensorRT with dynamic batching and CUDA graphs targets the two identified issues: variable-length sequences are handled by dynamic shapes and batching, while CUDA graphs eliminate launch overhead by capturing the entire inference step. This approach maintains FP16 accuracy and improves latency without the risks of quantization.

Exam trap

The trap here is focusing solely on precision reduction as the default latency fix, while overlooking launch overhead and batching strategies that can yield larger gains with no accuracy cost.

14
Multi-Selectmedium

A team is optimizing a large language model for inference on NVIDIA GPUs. They want to reduce the memory footprint of the model to fit on a single GPU with limited VRAM. Which two techniques are most appropriate for reducing memory usage during inference? (Choose two.)

Select 2 answers
A.Quantization of weights to INT8 or FP8.
B.Using NVIDIA TensorRT with layer fusion and kernel auto-tuning.
C.Enabling CUDA graphs to capture the inference step.
D.Pruning or sparsifying the model weights.
E.Increasing the batch size to improve utilization.
AnswersA, D

Quantizing weights to lower precision reduces the memory required to store the model parameters, often by 2x or 4x. This directly lowers VRAM usage and can also speed up inference on supported hardware. It is a standard technique for fitting large models on constrained GPUs, though it may require calibration to maintain accuracy.

Why this answer

Quantization and pruning directly reduce the memory required to store and execute the model. Quantization lowers the precision of weights, while pruning removes unnecessary parameters. Both techniques decrease VRAM usage and can be combined for greater effect.

The other options either increase memory usage or address different bottlenecks like launch overhead.

Exam trap

The trap here is assuming that any optimization that improves inference efficiency also reduces memory; techniques like CUDA graphs and larger batches do not lower VRAM footprint.

15
MCQmedium

A team is training a large language model using NVIDIA DGX A100 nodes with 8 GPUs per node. They observe that GPU utilization is high on all GPUs, but the training throughput scales poorly when adding more nodes. Profiling shows that the communication time during all-reduce operations increases significantly with node count. Which of the following optimizations is most likely to improve scaling efficiency?

A.Use gradient accumulation to increase the effective batch size.
B.Increase the number of CPU threads dedicated to data loading.
C.Enable NCCL's tree algorithm for all-reduce operations.
D.Configure NCCL to use InfiniBand with GPUDirect RDMA and enable adaptive routing.
AnswerD

Using InfiniBand with GPUDirect RDMA allows GPUs to communicate directly with the network adapter, bypassing host memory and reducing latency. Adaptive routing dynamically selects less congested paths, improving bandwidth utilization across multiple nodes. This combination significantly reduces communication overhead in large-scale all-reduce operations, directly addressing the poor scaling observed when adding nodes.

Why this answer

The poor scaling with increasing node count points to inter-node communication overhead, specifically during all-reduce operations. Leveraging InfiniBand with GPUDirect RDMA and adaptive routing optimizes the network path, reducing latency and improving bandwidth. This directly targets the communication bottleneck, enabling better scaling efficiency for distributed LLM training across multiple DGX nodes.

Exam trap

The trap here is assuming that gradient accumulation or CPU thread tuning addresses communication overhead, when the bottleneck is specifically inter-node all-reduce performance.

16
MCQmedium

A team is serving a 13B-parameter LLM on a single NVIDIA A100 80GB GPU. During generation, they observe that the GPU compute utilization stays below 20% while memory bandwidth utilization is near saturation. They want to improve throughput without changing the model architecture. Which optimization is most appropriate?

A.Apply 4-bit weight-only quantization to reduce memory traffic during token generation.
B.Enable multi-GPU tensor parallelism across two A100 GPUs to split the model.
C.Increase the batch size to improve arithmetic intensity and better utilize tensor cores.
D.Increase the number of CPU threads used for token sampling to speed up generation.
AnswerA

The workload is memory-bandwidth bound, as shown by low compute utilization and saturated memory bandwidth. Reducing weight precision to 4-bit lowers the bytes transferred per token, directly alleviating the bottleneck and increasing throughput without changing the architecture. This is a standard optimization for memory-bound LLM inference.

Why this answer

The symptoms indicate a memory-bandwidth-bound workload, common in autoregressive LLM decoding. Reducing weight precision via 4-bit quantization decreases the volume of data read from GPU memory per token, directly increasing throughput. Other options either do not target memory bandwidth or introduce new bottlenecks.

Exam trap

The trap here is assuming that low compute utilization always means the GPU needs more parallel work, when it can instead indicate a memory-bandwidth bottleneck.

17
MCQmedium

A company is running a large language model inference service on NVIDIA GPUs. They observe that GPU memory is nearly full, limiting the batch size and thus throughput. The model weights are stored in FP16, and the KV cache consumes a significant portion of memory. Which technique can reduce memory usage while maintaining model accuracy and enabling larger batch sizes?

A.Quantize the KV cache to INT8.
B.Increase the number of GPUs to distribute the model.
C.Convert the model to FP32 to improve numerical stability.
D.Use gradient checkpointing during inference.
AnswerA

Quantizing the KV cache to INT8 reduces its memory footprint by half compared to FP16, allowing larger batch sizes. Modern techniques like NVIDIA's TensorRT-LLM support INT8 KV cache quantization with minimal accuracy loss. This directly addresses the memory bottleneck while preserving model accuracy, making it an effective solution for increasing throughput.

Why this answer

The memory bottleneck is largely due to the KV cache. Quantizing it to INT8 halves its memory footprint, freeing space for larger batch sizes and improving throughput. This technique is supported in frameworks like TensorRT-LLM and maintains accuracy with proper calibration.

Other options either do not apply to inference or increase memory usage.

Exam trap

The trap here is confusing training-time memory optimizations like gradient checkpointing with inference-time techniques, or assuming that increasing precision helps when the goal is to reduce memory.

18
Multi-Selecthard

When deploying a large language model on NVIDIA H100 GPUs using TensorRT-LLM, which TWO configuration strategies are most effective for improving KV cache efficiency and memory utilization?

Select 2 answers
A.Enable PagedAttention to reduce fragmentation.
B.Increase the static sequence length allocation.
C.Enable continuous inflight batching.
D.Disable multi-head attention optimizations.
E.Switch to FP64 precision for all tensors.
AnswersA, C

PagedAttention treats the KV cache as non-contiguous memory blocks, similar to virtual memory in an OS. This eliminates internal fragmentation caused by over-allocating memory for sequence lengths that never materialize, allowing the GPU to pack more requests into the same VRAM capacity for higher throughput.

Why this answer

Optimizing the KV cache is critical for scaling generative AI, as it occupies significant VRAM during inference. PagedAttention dynamically manages memory blocks to prevent fragmentation, while inflight batching ensures that continuous requests are processed without waiting for the entire batch to finish. Combining these allows for higher concurrency and reduced memory overhead, enabling more efficient deployment of models with large context windows on limited GPU hardware.

Exam trap

Candidates frequently select generic training optimizations like data parallelism instead of focusing specifically on inference-centric KV cache and batching strategies.

19
Multi-Selecthard

An engineer is optimizing a large language model for inference on NVIDIA GPUs and wants to reduce memory usage to fit a larger model or increase batch size. Which two techniques are most effective for reducing GPU memory consumption during inference? (Choose two.)

Select 2 answers
A.Enable activation checkpointing to recompute activations during the backward pass.
B.Increase the number of micro-batches to overlap computation and memory transfers.
C.Quantize model weights to 8-bit integers using TensorRT or similar tools.
D.Use a paged attention mechanism to manage the KV cache efficiently.
E.Store the model weights on the CPU and transfer them to the GPU on demand.
AnswersC, D

Quantizing weights to 8-bit integers reduces the memory footprint by up to 4x compared to FP32. This allows larger models to fit in GPU memory and enables larger batch sizes. TensorRT supports INT8 quantization with calibration to maintain accuracy, making it a standard technique for memory reduction during inference.

Why this answer

Quantizing weights to INT8 reduces memory footprint by storing weights in lower precision, while paged attention optimizes KV cache memory by reducing fragmentation. Both are effective for fitting larger models or increasing batch size during inference. Activation checkpointing and micro-batching are training or throughput techniques, and CPU offloading introduces latency.

Exam trap

The trap here is confusing training-time memory optimizations, like activation checkpointing, with inference-time memory reductions, such as weight quantization and paged attention.

20
MCQeasy

A developer is using NVIDIA Nsight Systems to profile a PyTorch training loop on an NVIDIA GPU. They notice significant gaps between kernel executions and want to identify whether the bottleneck is CPU-side or GPU-side. Which Nsight Systems feature should they use to visualize the CPU and GPU timelines together?

A.The Nsight Compute kernel profiling report.
B.The NVIDIA Management Library (NVML) GPU utilization metrics.
C.The PyTorch autograd profiler output.
D.The CUDA API trace and GPU activity timeline.
AnswerD

Nsight Systems provides a unified timeline that shows CUDA API calls on the CPU and corresponding kernel executions on the GPU. By correlating these, developers can see gaps where the GPU is idle waiting for CPU work, indicating a CPU-side bottleneck. This is the primary feature for identifying such imbalances.

Why this answer

Nsight Systems' unified timeline displays CPU activities (including CUDA API calls) and GPU kernels on the same time axis. This allows developers to see gaps where the GPU is idle while the CPU is busy, indicating a CPU-bound workload. It is the standard tool for this type of system-level analysis.

Exam trap

The trap here is confusing Nsight Compute (kernel-level) with Nsight Systems (system-level timeline), when the latter is needed for CPU-GPU correlation.

21
Multi-Selectmedium

A developer is optimizing a generative AI model for inference on NVIDIA GPUs. They want to reduce memory footprint and improve throughput without sacrificing accuracy. Which two techniques should they apply? (Choose two.)

Select 2 answers
A.Use NVIDIA TensorRT with INT8 quantization and calibration.
B.Use NVIDIA GPUDirect Storage to offload model weights to NVMe.
C.Increase the batch size to amortize memory overhead.
D.Enable NVIDIA Ampere TF32 precision for matrix multiplications.
E.Apply NVIDIA's structured sparsity with 2:4 pattern and TensorRT support.
AnswersA, E

TensorRT with INT8 quantization reduces memory footprint and increases throughput by using 8-bit integer operations, which are faster on Tensor Cores. Calibration ensures minimal accuracy loss by determining optimal scaling factors. This is a standard optimization for inference on NVIDIA GPUs.

Why this answer

INT8 quantization and structured sparsity both reduce memory footprint and improve throughput. INT8 quantization uses 8-bit integers, cutting memory usage by 4x compared to FP32, and Tensor Cores accelerate INT8 operations. Structured sparsity prunes weights in a 2:4 pattern, halving memory for weights and enabling sparse Tensor Core acceleration.

Both are supported by TensorRT and maintain accuracy with proper calibration and fine-tuning.

Exam trap

The trap here is confusing TF32 (which accelerates compute but not memory) with true memory-reducing techniques like INT8 and sparsity.

22
MCQmedium

A developer is using NVIDIA TensorRT to optimize a BERT-based model for inference. They notice that the engine performs poorly on variable-length input sequences because it was built with a single optimization profile for a fixed sequence length. What should they do to improve performance across different sequence lengths?

A.Pad all input sequences to the maximum length supported by the model.
B.Rebuild the engine with multiple optimization profiles covering the range of sequence lengths.
C.Use CUDA graphs to capture and replay the inference execution.
D.Enable FP16 precision to reduce memory usage and speed up kernels.
AnswerB

TensorRT optimization profiles define the min, opt, and max shapes for dynamic input dimensions. Building multiple profiles allows the engine to select the best kernel configurations for different sequence lengths, improving performance across the range. This is the recommended approach for variable-length inputs.

Why this answer

TensorRT optimization profiles allow the engine to tune kernels for specific input shape ranges. When input lengths vary, a single profile cannot cover all cases efficiently. Creating multiple profiles, each with appropriate min/opt/max dimensions, enables the engine to choose the best kernels for each length, improving overall performance.

Exam trap

The trap here is thinking that precision reduction or launch overhead reduction can compensate for a mismatched optimization profile, when the core issue is kernel selection for varying shapes.

23
MCQhard

An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. They want to maximize throughput for a chatbot workload with variable-length inputs and outputs. Which of the following techniques should they implement to achieve the highest throughput while maintaining acceptable latency?

A.Enable in-flight batching (continuous batching) and use paged KV cache with block-based memory management.
B.Use static batching with a fixed batch size of 32 and pad all sequences to the maximum length.
C.Use FP32 precision to ensure accuracy and rely on CUDA graphs to reduce launch overhead.
D.Increase the number of GPUs and use tensor parallelism with a fixed batch size per GPU.
AnswerA

In-flight batching dynamically adds and removes sequences at each iteration, eliminating padding and improving GPU utilization. Paged KV cache stores attention keys and values in non-contiguous blocks, reducing memory fragmentation and allowing more concurrent sequences. Together, they maximize throughput for variable-length chatbot traffic while keeping latency low, as implemented in TensorRT-LLM.

Why this answer

In-flight batching and paged KV cache are key optimizations in TensorRT-LLM for LLM serving. In-flight batching allows sequences to join and leave the batch at each decoding step, avoiding padding and keeping the GPU busy. Paged KV cache manages memory in blocks, enabling more sequences to run concurrently.

This combination delivers the highest throughput for variable-length chatbot workloads.

Exam trap

The trap here is assuming that increasing batch size or GPUs alone will maximize throughput, when dynamic batching and memory management are the real enablers.

24
MCQeasy

In the context of NVIDIA Tensor Cores, what is the primary benefit of using BF16 (Bfloat16) over FP16 during model training and inference?

A.BF16 provides double the precision of FP16.
B.BF16 offers a larger dynamic range for gradients.
C.BF16 requires significantly less memory than FP16.
D.BF16 is faster on non-NVIDIA hardware.
AnswerB

The 8-bit exponent in BF16 matches the dynamic range of FP32, allowing it to represent a much wider range of values than FP16. This prevents underflow and overflow issues during deep learning operations, leading to more stable model convergence without requiring complex loss scaling techniques.

Why this answer

BF16 uses the same exponent range as FP32, which prevents overflow issues commonly encountered in deep learning training when using FP16. This increased dynamic range makes it more robust for gradient calculations and weight updates. By providing a wider range while maintaining the same performance advantages of half-precision, BF16 has become the industry standard for stabilizing training and inference of modern large language models.

Exam trap

Candidates mistakenly believe BF16 provides higher precision than FP16, confusing the mantissa size with the exponent range benefits.

25
MCQmedium

Which hardware component of an NVIDIA GPU is most responsible for accelerating matrix-multiply-accumulate (MMA) operations used in transformer layers?

A.CUDA Cores.
B.Streaming Multiprocessor (SM) Scheduler.
C.Tensor Cores.
D.L2 Cache controller.
AnswerC

Tensor Cores are specialized hardware units optimized for high-performance matrix-multiply-accumulate operations. They are the engine behind modern generative AI, allowing GPUs to process large transformer models with extreme efficiency, significantly outperforming general-purpose cores for the math-heavy tasks required by LLMs and neural networks.

Why this answer

Tensor Cores are specialized hardware units designed to perform high-speed matrix-multiply-accumulate operations in a single clock cycle. By accelerating these core operations, Tensor Cores provide the massive compute throughput needed for deep learning. Understanding the role of Tensor Cores is vital because they define the performance limits for modern LLMs, and optimizing code to utilize them is the single most important task in GPU performance tuning.

Exam trap

Candidates often confuse general-purpose CUDA cores with specialized Tensor Cores when asked about matrix-multiply-accumulate acceleration.

26
MCQmedium

When using the NVIDIA Collective Communications Library (NCCL), what does 'AllReduce' specifically optimize for in a multi-GPU training configuration?

A.Local GPU thread synchronization.
B.Gradient aggregation across all devices.
C.Direct CPU-to-GPU data copy.
D.Increasing GPU clock frequency.
AnswerB

AllReduce performs a summation of gradients across all participating GPUs and returns the result to every device. This ensures all model replicas are synchronized during the weight update phase, making it the fundamental operation for scaling deep learning training across clusters of multiple NVIDIA GPUs.

Why this answer

AllReduce is critical for distributed training because it synchronizes the gradient updates from all GPUs across the cluster. It aggregates data from all devices and distributes the result back to each one, allowing models to train concurrently on massive datasets. By optimizing this collective operation, NCCL minimizes the time GPUs spend waiting for synchronization, which is the primary hurdle in scaling training to hundreds or thousands of GPUs.

Exam trap

Candidates often confuse AllReduce with point-to-point communication or broadcast operations, failing to recognize it as the specific collective operation for synchronizing gradients across distributed nodes.

27
MCQmedium

A team is training a large language model on a single NVIDIA H100 GPU. They observe that training throughput is significantly lower than expected, and profiling with Nsight Systems shows long periods where the GPU is idle waiting for data. The data loading pipeline uses the default PyTorch DataLoader with num_workers=0 and no pinned memory. Which change is most likely to improve GPU utilization?

A.Use CUDA graphs to capture the training step and reduce kernel launch overhead.
B.Switch from FP32 to TF32 precision for matrix multiplications.
C.Increase the batch size to better saturate the GPU.
D.Enable pinned memory and increase num_workers in the DataLoader.
AnswerD

Pinned memory allows asynchronous host-to-device copies, and multiple worker processes prefetch batches in parallel, overlapping data loading with GPU compute. This directly addresses the idle GPU time observed in profiling. With num_workers=0 and no pinning, data loading is serialized and synchronous, starving the GPU.

Why this answer

The profiling evidence points to a data-loading bottleneck: the GPU sits idle waiting for batches. Enabling pinned memory and increasing DataLoader workers allows asynchronous, overlapped data transfer and prefetching, keeping the GPU fed. The other options target compute or launch overhead, which are not the limiting factors in this scenario.

Exam trap

The trap here is assuming that GPU underutilization always means the model or kernels need optimization, rather than checking whether the input pipeline is starving the device.

28
MCQeasy

Which of the following describes the purpose of 'Kernel Fusion' in the context of optimizing a Deep Learning inference pipeline?

A.To increase the number of parallel GPU threads.
B.To reduce redundant global memory read/write cycles.
C.To enable multi-GPU distributed training.
D.To improve model accuracy through extra precision.
AnswerB

Kernel fusion minimizes global memory traffic by keeping intermediate results in registers or shared memory. By avoiding writing intermediate tensors back to VRAM, the pipeline becomes significantly faster, as reading from and writing to high-latency VRAM is the primary bottleneck for many AI inference tasks.

Why this answer

Kernel fusion combines multiple small operations into a single GPU kernel to reduce the overhead of launching kernels and accessing global memory. Every kernel launch involves CPU-side overhead, and global memory accesses are costly in terms of energy and time. Fusion minimizes both, significantly increasing the effective throughput of the GPU by keeping data in high-speed, on-chip storage for as long as possible.

Exam trap

Candidates often think kernel fusion increases parallel thread execution count, confusing instruction-level merging with hardware scaling.

29
MCQmedium

What is the primary function of the 'TensorRT' optimization engine in the NVIDIA AI software stack?

A.It automates the training of complex models.
B.It provides a Python API for model debugging.
C.It performs architecture-specific model optimization.
D.It converts models to run on mobile CPUs.
AnswerC

TensorRT performs deep optimizations like layer fusion, kernel auto-tuning, and precision reduction, all tailored to the specific GPU architecture being used. This allows the model to run at peak throughput and minimal latency by taking advantage of the unique features of the target NVIDIA hardware architecture.

Why this answer

TensorRT optimizes neural network models by performing layer fusion, precision calibration (e.g., to FP8 or INT8), and kernel selection optimized for the specific GPU architecture. By transforming the model into a highly efficient, platform-specific format, TensorRT significantly reduces latency and increases throughput for production inference. It is the core tool for moving from research-grade PyTorch models to production-ready deployments on NVIDIA hardware.

Exam trap

Candidates often mistake TensorRT for a general-purpose library for model training or data preprocessing, ignoring its specific role as an inference-time optimization engine for NVIDIA hardware.

30
MCQmedium

A team is training a 13B-parameter LLM on 8 NVIDIA A100 GPUs using NVIDIA NeMo. They observe that the all-reduce communication during data-parallel training consumes nearly 40% of each iteration. Which of the following changes is most likely to reduce this communication overhead while preserving convergence?

A.Switch from data parallelism to tensor parallelism across all 8 GPUs.
B.Enable gradient accumulation with a larger micro-batch size and use NCCL with tree algorithm.
C.Use NVIDIA NCCL with the ring algorithm and overlap communication with computation via gradient bucketing.
D.Reduce the number of GPUs to 4 and increase the per-GPU batch size.
AnswerC

NCCL's ring algorithm is optimized for large messages and high-bandwidth interconnects like NVLink, making it efficient for all-reduce in data-parallel training. Overlapping communication with computation using gradient bucketing hides latency behind backpropagation. Together, these reduce the effective communication overhead without changing model convergence, directly addressing the observed bottleneck.

Why this answer

The communication bottleneck in data-parallel training is best mitigated by using an efficient collective algorithm and hiding latency. NCCL's ring algorithm excels with large messages on high-bandwidth links, and gradient bucketing enables overlap with compute. This preserves the data-parallel semantics and convergence while reducing wall-clock time spent in all-reduce.

Exam trap

The trap here is assuming that any parallelism change (like tensor parallelism) will reduce communication, when it often increases it due to more frequent synchronization.

31
MCQmedium

Refer to the exhibit. An engineer receives this timeout error during a CUDA kernel execution. What is the most appropriate first step to diagnose the resource contention?

A.Increase the block size to 2048 threads.
B.Use NVIDIA Nsight Compute to profile occupancy.
C.Disable the watchdog timer in the OS.
D.Switch the kernel to run on the CPU.
AnswerB

Nsight Compute provides detailed analysis of occupancy, register usage, and shared memory allocation. It identifies whether the kernel is struggling with resource contention, allowing the developer to adjust thread block configuration or refine memory usage to prevent the execution time from exceeding the watchdog timer.

Why this answer

CUDA kernel timeouts are typically caused by long-running kernels that exceed the GPU's watchdog timer or by excessive resource usage (registers/shared memory) that limits occupancy. Using the NVIDIA Nsight Compute profiler allows the engineer to see the exact occupancy metrics and register pressure, enabling targeted optimizations. This step is crucial for identifying if the kernel is over-provisioned for the specific GPU architecture being targeted.

Exam trap

Candidates frequently recommend restarting the driver or increasing the OS watchdog timeout limit instead of using dedicated profiling tools to inspect kernel resource usage.

32
Multi-Selectmedium

Which TWO of the following techniques are best suited for reducing the latency of LLM inference on NVIDIA GPUs?

Select 2 answers
A.KV cache quantization.
B.Increasing the number of CPU worker threads.
C.Implementing PagedAttention.
D.Disabling ECC memory on the GPU.
E.Switching from Tensor Cores to CUDA Cores.
AnswersA, C

KV cache quantization compresses the storage of intermediate tokens, significantly reducing memory bandwidth consumption. Since LLM inference is often memory-bandwidth bound during the decoding phase, this technique allows more tokens to be processed concurrently and speeds up the transfer of cache data to the compute units.

Why this answer

Reducing LLM latency requires optimizing both the compute throughput and the memory access patterns. Key-Value (KV) cache quantization and PagedAttention are industry standards for LLM acceleration. KV cache quantization reduces memory bandwidth usage, while PagedAttention manages memory dynamically, preventing fragmentation.

These methods significantly improve the token generation rate, which is the primary metric for user-perceived performance in large-scale generative AI deployments.

Exam trap

Candidates often suggest generic performance tweaks like overclocking or batch size adjustments, missing that PagedAttention and KV cache quantization are specific, high-impact techniques for LLM memory management.

33
MCQhard

An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. The model uses multi-head attention with a sequence length of 4096. During inference, the GPU's Tensor Cores are underutilized, and the kernel launch overhead is high due to many small operations. Which optimization should be applied to improve Tensor Core utilization and reduce overhead?

A.Increase the batch size to maximize parallelism and hide latency.
B.Enable kernel fusion by using TensorRT-LLM's fused multi-head attention plugin.
C.Use NVIDIA Triton Inference Server with dynamic batching to improve GPU utilization.
D.Convert the model to FP16 precision to double the Tensor Core throughput.
AnswerB

TensorRT-LLM provides fused multi-head attention plugins that combine multiple operations (e.g., QKV projection, attention, and output projection) into a single kernel. This reduces kernel launch overhead and increases arithmetic intensity, allowing better utilization of Tensor Cores. The fusion also minimizes memory traffic, which is critical for long sequences.

Why this answer

The underutilization of Tensor Cores and high kernel launch overhead stem from many small operations in multi-head attention. TensorRT-LLM's fused multi-head attention plugin combines these operations into a single optimized kernel, reducing overhead and improving Tensor Core utilization. This is a targeted optimization for Transformer models, especially with long sequences.

Exam trap

The trap here is focusing on batch size or precision when the core issue is kernel fragmentation and launch overhead, which require fusion.

34
MCQhard

An engineer is training a large language model with pipeline parallelism across four NVIDIA GPUs. They observe that GPU utilization is low and training throughput is limited by idle time during pipeline bubbles. Which technique is most effective to reduce pipeline bubbles and improve utilization?

A.Increase the number of micro-batches to keep the pipeline stages busy.
B.Reduce the batch size to decrease the memory footprint and allow more concurrent kernels.
C.Enable gradient checkpointing to trade compute for memory and increase batch size.
D.Switch from pipeline parallelism to data parallelism across the four GPUs.
AnswerA

Pipeline bubbles occur when stages wait for data from previous stages. Increasing the number of micro-batches allows more overlapping of forward and backward passes across stages, filling the pipeline and reducing idle time. This is a standard technique in pipeline parallelism to improve utilization.

Why this answer

Pipeline bubbles are idle periods when stages wait for data. Increasing the number of micro-batches allows more fine-grained overlap of computation across stages, keeping all GPUs busy and reducing idle time. This is a direct and effective way to improve pipeline utilization.

Exam trap

The trap here is confusing memory optimization techniques with those that address pipeline scheduling inefficiencies, such as increasing micro-batches.

Ready to test yourself?

Try a timed practice session using only Gpu Acceleration Optimization questions.