Courseiva

NCP-GENL GPU Acceleration and Optimization Practice Question

A developer is using NVIDIA Nsight Systems to profile a PyTorch training loop on an NVIDIA GPU. They notice significant gaps between kernel executions and want to identify whether the bottleneck is CPU-side or GPU-side. Which Nsight Systems feature should they use to visualize the CPU and GPU timelines together?

⚠ Common exam trap

Many exam-takers confuse Nsight Compute (kernel-level) with Nsight Systems (system-level timeline), when the latter is needed for CPU-GPU correlation.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The CUDA API trace and GPU activity timeline.

Nsight Systems' unified timeline displays CPU activities (including CUDA API calls) and GPU kernels on the same time axis. This allows developers to see gaps where the GPU is idle while the CPU is busy, indicating a CPU-bound workload. It is the standard tool for this type of system-level analysis.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The Nsight Compute kernel profiling report.

    Why it's wrong here

    Nsight Compute is a separate tool for deep-diving into individual kernel performance, providing metrics like occupancy and memory throughput. It does not show the system-wide timeline or CPU-GPU interaction. While useful for kernel optimization, it is not the right tool for identifying CPU-side gaps in the overall timeline.

  • ✗

    The NVIDIA Management Library (NVML) GPU utilization metrics.

    Why it's wrong here

    NVML provides coarse-grained GPU utilization percentages and memory usage, but it does not offer a detailed timeline or correlation with CPU activities. It can indicate overall GPU busyness but cannot pinpoint whether gaps are due to CPU-side delays. It is a monitoring tool, not a profiler for timeline analysis.

  • ✗

    The PyTorch autograd profiler output.

    Why it's wrong here

    The PyTorch autograd profiler records operations and their durations within PyTorch, but it does not provide a system-level view of CUDA API calls or GPU kernels in a unified timeline. It is useful for understanding model-level performance but lacks the granularity to diagnose CPU-GPU synchronization issues.

  • ✓

    The CUDA API trace and GPU activity timeline.

    Why this is correct

    Nsight Systems provides a unified timeline that shows CUDA API calls on the CPU and corresponding kernel executions on the GPU. By correlating these, developers can see gaps where the GPU is idle waiting for CPU work, indicating a CPU-side bottleneck. This is the primary feature for identifying such imbalances.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.