Courseiva

NCP-GENL · topic practice

GPU Acceleration and Optimization practice questions

This domain covers profiling, scaling, and inference optimization on NVIDIA hardware. Questions test whether you can read Nsight Systems traces, distinguish PCIe from NVLink bottlenecks, choose the right library (TensorRT, NeMo, Triton), and pick techniques that cut LLM latency or all-reduce overhead.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: GPU Acceleration and Optimization

What the exam tests

What to know about GPU Acceleration and Optimization

Be able to profile a GPU workload in Nsight Systems, name the bottleneck, and match it to the right NVIDIA tool: TensorRT for inference, NeMo for training, NCCL for collectives. The single most important thing is correctly identifying whether PCIe, NVLink, or compute is the real limiter.

Reading Nsight Systems traces to separate PCIe transfer stalls from NVLink and SM utilization limits

Selecting TensorRT for inference optimization and deployment, versus NeMo for training and Triton for serving

Diagnosing data-parallel all-reduce overhead with NCCL and choosing overlap, larger batches, or gradient accumulation

Reducing LLM inference latency via quantization, KV-cache tuning, and kernel or batching changes

Watch out for

Common GPU Acceleration and Optimization exam traps

  • ▸Blaming PCIe for every stall; NVLink or SM occupancy limits often dominate and Nsight Systems metrics reveal which is actually saturated.
  • ▸Confusing TensorRT with NeMo or Triton: TensorRT optimizes and deploys engines, NeMo trains, Triton serves, and the exam expects the right tool named.
  • ▸Assuming quantization or smaller batches always cut latency; KV-cache growth and memory-bandwidth limits can dominate, so measure before optimizing.

Practice set

GPU Acceleration and Optimization questions

20 questions · select your answer, then reveal the explanation

An AI engineer is deploying a large language model on an NVIDIA A100 GPU using TensorRT-LLM. During inference profiling, they notice that token generation latency is higher than expected due to memory bandwidth bottlenecks during the autoregressive decoding phase. Which optimization technique should be applied first to mitigate this bandwidth limitation?

A machine learning operations team is profiling a custom Transformer model trained on NVIDIA H100 GPUs using NVIDIA Nsight Systems. They observe that GPU utilization drops significantly during the data loading and tokenization phases between training steps. Which TWO strategies should the team implement to resolve this host-to-device pipeline bottleneck? (Choose two)

An enterprise developer is optimizing a computer vision pipeline on NVIDIA A100 GPUs using TensorRT. They want to maximize throughput for high-resolution image batches while maintaining acceptable latency bounds. Which execution mode should they configure in the TensorRT builder configuration?

Refer to the exhibit. A deep learning engineer encounters a CUDA out-of-memory error while initializing a deep neural network training job on a partitioned NVIDIA H100 GPU instance. Based on the error log, what is the most appropriate remediation step?

Exhibit

gpu 0: NVIDIA H100 PCIe (UUID: GPU-12345678-abcd-ef01-2345-6789abcdef01)
  MIG 3g.40gb:enabled
  Max Batch Size: 32
  Current Memory Allocation: 98.4%
  Error: CUDA out of memory during workspace allocation for cudnnFindConvolutionForwardAlgorithm

An inference server is experiencing high GPU latency when using FP16 precision. Profiling reveals that the kernel execution time is dominated by memory bandwidth bottlenecks. Which optimization technique is most likely to mitigate this latency while maintaining model throughput?

Refer to the exhibit. An engineer observes the provided output from a GPU currently running an inference task. What does this indicate regarding the current workload state?

Network Topology
query-gpu=utilization.gpunvidia-smiformat=csv

Which CUDA-aware communication primitive is specifically designed to optimize data movement between multiple GPUs in a multi-node cluster by bypassing the CPU host memory?

Which THREE techniques are primary methods for reducing the footprint of an LLM on GPU VRAM to enable larger models on restricted hardware?

In an H100 environment, which factor is most significant when choosing between TensorRT-LLM and standard PyTorch for high-throughput inference deployment?

Which technique provides the most significant performance gain for large-model inference by splitting the model across multiple GPUs while keeping the sequence length per batch constant?

An engineer is profiling a Transformer-based model and observes low GPU utilization despite high batch sizes. The bottleneck analysis identifies frequent host-to-device memory copies. Which optimization strategy most effectively addresses this latency?

Refer to the exhibit. An inference engine is failing during peak load. What is the most likely cause, and which strategy should be prioritized for immediate resolution?

Exhibit

LOG_ENTRY: [TensorRT] Kernel execution error at layer 42: cuBLAS_STATUS_EXECUTION_FAILED. Memory usage at 98%. Profiler output: Memory bandwidth bottleneck observed. GPU VRAM allocated: 23.8GB/24GB.

Which THREE of the following are valid methods to optimize NVIDIA GPU memory bandwidth usage in deep learning?

A team is serving a large language model with TensorRT-LLM on an NVIDIA H100 GPU. They observe that during the generation phase, GPU utilization is high, but the number of requests processed per second is lower than expected. Profiling shows that the GPU spends a significant amount of time waiting for memory operations. Which optimization technique should they apply to improve throughput?

A developer is deploying a large language model on NVIDIA GPUs and wants to reduce inference latency while maintaining model accuracy. The model currently uses FP32 weights and activations. Which two techniques are most appropriate to achieve lower latency without significant accuracy loss? (Choose two.)

A developer is fine-tuning a BERT model on an NVIDIA V100 GPU. They notice that the training step time is much longer than expected, and nvidia-smi shows that GPU memory is nearly full. The model uses a batch size of 64 and the maximum sequence length is 512. Which action is most likely to allow a larger batch size without running out of memory?

A research team is training a Transformer model on a cluster of NVIDIA A100 GPUs using data parallelism with PyTorch DistributedDataParallel (DDP). They observe that scaling efficiency drops significantly when going from 4 to 8 GPUs. Profiling shows high communication overhead in NCCL AllReduce operations. Which optimization is most likely to improve scaling efficiency in this scenario?

When deploying a large language model on NVIDIA H100 GPUs using TensorRT-LLM, which TWO configuration strategies are most effective for improving KV cache efficiency and memory utilization?

In the context of NVIDIA Tensor Cores, what is the primary benefit of using BF16 (Bfloat16) over FP16 during model training and inference?

Refer to the exhibit. An engineer receives this timeout error during a CUDA kernel execution. What is the most appropriate first step to diagnose the resource contention?

Exhibit

Error: Kernel execution timed out after 5000ms.
Potential cause: Excessive occupancy or resource contention.
Action: Adjust block size or shared memory usage.

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused GPU Acceleration and Optimization sessions

Start a GPU Acceleration and Optimization only practice session

Every question in these sessions is drawn from the GPU Acceleration and Optimization domain — nothing else.

Related practice questions

Related NCP-GENL topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the NCP-GENL exam test about GPU Acceleration and Optimization?
Be able to profile a GPU workload in Nsight Systems, name the bottleneck, and match it to the right NVIDIA tool: TensorRT for inference, NeMo for training, NCCL for collectives. The single most important thing is correctly identifying whether PCIe, NVLink, or compute is the real limiter.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just GPU Acceleration and Optimization questions in a focused session?
Yes — the session launcher on this page draws every question from the GPU Acceleration and Optimization domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other NCP-GENL topics?
Use the topic links above to move to related areas, or go back to the NCP-GENL question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the NCP-GENL exam covers. They are not copied from any real exam or dump site.