NCP-GENL GPU Acceleration and Optimization Practice Question
When profiling an application with NVIDIA Nsight Systems, which TWO metrics are most critical to identify if an application is limited by the PCIe bus?
⚠ Common exam trap
Candidates often confuse PCIe throughput with GPU compute utilization or memory bandwidth metrics. They fail to realize that PCIe saturation is specifically about the data transfer rate between the host and GPU.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Host-to-Device (H2D) throughput.
PCIe bus saturation occurs when the transfer rate of data between the host and the GPU becomes a bottleneck. By monitoring the H2D (Host-to-Device) and D2H (Device-to-Host) transfer metrics, developers can see if data movement consumes a disproportionate amount of execution time. Identifying these spikes is crucial because offloading data to the GPU is often the slowest part of a pipeline, and minimizing these transfers is key to scaling high-performance AI.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Host-to-Device (H2D) throughput.
Why this is correct
H2D throughput measures the speed at which data is sent from the host CPU to the GPU memory. If this metric hits the theoretical maximum of the PCIe bus, it confirms a bottleneck where the GPU must wait for new data to arrive before processing can begin.
- ✗
GPU Register usage count.
Why it's wrong here
Register usage is an internal GPU metric related to occupancy and thread scheduling. It has no correlation with the bandwidth limitations of the external PCIe bus. High register usage suggests internal compute constraints, not a data transport bottleneck between the CPU and the GPU.
- ✓
Device-to-Host (D2H) throughput.
Why this is correct
D2H throughput measures the speed at which results are sent back to the host. In applications like generative AI, transferring large generated sequences back to the host can saturate the PCIe bus if not handled efficiently, leading to delays in response times for the end user.
- ✗
Shared memory bank conflict count.
Why it's wrong here
Shared memory bank conflicts are an issue within the streaming multiprocessor's memory subsystem. They cause latency during kernel execution but are unrelated to the PCIe bus bandwidth. Addressing bank conflicts improves local execution speed, not the efficiency of the link between the host and GPU.
- ✗
SM clock frequency.
Why it's wrong here
SM clock frequency is a measure of the GPU's compute speed. While it affects overall performance, it does not provide information about the PCIe bus saturation. A low clock frequency does not necessarily imply PCIe-bound performance, as it could be caused by thermal throttling or power limits.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.