Courseiva

NCP-AIO Troubleshooting and Optimization Practice Question

An AI operations engineer notices that a real-time inference service on an NVIDIA T4 GPU has highly variable latency, with occasional spikes to over 100 ms. The service uses TensorRT and runs in a Docker container. Which action should the engineer take to reduce latency variability?

⚠ Common exam trap

The trap here is assuming that persistence mode or CPU allocation affects runtime latency, when the primary cause of variability is often GPU clock throttling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Lock the GPU clocks to a fixed frequency using nvidia-smi -lgc.

Latency variability in GPU inference often results from dynamic clock adjustments as the GPU balances power and thermal constraints. Locking clocks to a stable frequency eliminates this variability, providing consistent execution times. Other options either add latency (dynamic batching), address different issues (persistence mode), or do not target GPU clock behavior.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable dynamic batching in the inference server configuration.

    Why it's wrong here

    Dynamic batching groups multiple inference requests into a single batch to improve throughput, but it can increase latency for individual requests as they wait to be batched. In a latency-sensitive service with spikes, this would likely worsen variability. The goal is to reduce spikes, not add batching delays.

  • ✗

    Increase the number of CPU cores allocated to the container.

    Why it's wrong here

    While CPU resources can affect preprocessing and data transfer, the latency spikes are likely due to GPU-side variability. Adding CPU cores may help if CPU is a bottleneck, but it does not directly control GPU clock behavior. Without evidence of CPU saturation, this is unlikely to resolve the issue and could waste resources.

  • ✗

    Set the GPU to persistence mode using nvidia-smi -pm 1.

    Why it's wrong here

    Persistence mode keeps the NVIDIA driver loaded even when no clients are connected, reducing initialization time for new processes. However, it does not address runtime latency variability caused by GPU clock fluctuations or other factors. It is more relevant for environments where processes frequently start and stop, not for a continuously running service.

  • ✓

    Lock the GPU clocks to a fixed frequency using nvidia-smi -lgc.

    Why this is correct

    Locking GPU clocks prevents the GPU from dynamically adjusting frequencies based on power and thermal headroom, which can cause latency variability. A fixed clock ensures consistent performance, reducing spikes. This is especially effective for latency-sensitive inference where predictable execution time is critical. It trades off some power efficiency for stability.

About these practice questions

Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.