Courseiva
Workload Management →mediumMultiple Choice

NCP-AIO Workload Management Practice Question

A production inference service experiences intermittent latency spikes. The service is deployed on shared infrastructure. Which tool would best help an AI Ops engineer identify if GPU resource contention is the cause?

⚠ Common exam trap

Candidates might suggest looking at CPU logs or standard Kubernetes pod metrics. These metrics are often blind to GPU-specific resource contention, such as memory bus saturation or shared SM usage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

NVIDIA DCGM Exporter with Prometheus/Grafana.

NVIDIA DCGM (Data Center GPU Manager) provides detailed telemetry data, including GPU utilization, memory bandwidth, and power usage per process. By correlating latency spikes with DCGM metrics, engineers can pinpoint if another process on the same GPU is competing for compute resources or memory bandwidth. This visibility is essential for performance tuning and workload placement, allowing engineers to verify if resource isolation policies are functioning as expected.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Standard Kubernetes 'kubectl top' command.

    Why it's wrong here

    'kubectl top' only reports CPU and system RAM usage of pods. It lacks any visibility into GPU-specific hardware metrics, such as GPU utilization, SM occupancy, or VRAM usage. Therefore, it cannot provide the necessary insights to diagnose performance issues rooted in GPU resource contention.

  • ✓

    NVIDIA DCGM Exporter with Prometheus/Grafana.

    Why this is correct

    DCGM collects granular, real-time metrics directly from the GPU hardware. By exporting these to Prometheus and visualizing them in Grafana, engineers can identify spikes in GPU activity or bandwidth saturation that correlate with the application's latency, providing a clear path to identifying the source of resource contention.

  • ✗

    The Linux 'top' command on the worker node.

    Why it's wrong here

    The 'top' utility is restricted to host-level process monitoring. It provides information about CPU load and standard memory consumption but does not interface with GPU hardware metrics. Consequently, it cannot identify if a specific process is hogging GPU compute cycles or causing stalls on the accelerator device.

  • ✗

    A simple network latency test tool like 'ping'.

    Why it's wrong here

    Network latency testing only measures the time taken for packets to travel across the network. It does not provide any visibility into the internal processing time of the GPU or the resource utilization on the compute node. Therefore, it is incapable of identifying GPU-related bottlenecks in inference.

About these practice questions

One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.