Courseiva

NCP-AIO Troubleshooting and Optimization Practice Question

An inference service runs a 70B parameter model with TensorRT-LLM on a single H100 using in-flight batching. Operators report that time-to-first-token is acceptable, but inter-token latency degrades sharply once concurrent request count exceeds a certain point, and GPU memory utilization sits near 98 percent. Which change most directly addresses the inter-token latency degradation?

⚠ Common exam trap

The trap here is treating higher concurrency as a throughput problem to solve by enlarging the batch, when the observed memory saturation means the batch ceiling is already too high for the available KV cache and workspace.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable FP8 quantization for the KV cache and reduce the maximum batch size in the TensorRT-LLM build configuration.

When GPU memory is nearly exhausted, the runtime cannot hold enough KV cache plus workspace to serve all admitted sequences efficiently, so each decode step stretches and inter-token latency rises with concurrency. Shrinking the KV cache with FP8 quantization and lowering the maximum batch size reduces per-step memory and compute demand, restoring shorter decode iterations. These changes target the resource constraint that actually causes the latency curve to bend upward.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Disable in-flight batching so each request is processed to completion before the next begins.

    Why it's wrong here

    Removing in-flight batching eliminates the mechanism that keeps the GPU busy across sequences and typically raises aggregate latency dramatically under load. While it would avoid contention effects, it sacrifices throughput and does nothing to make decode steps shorter. The reported symptom is degradation under concurrency, so discarding concurrency handling attacks the wrong layer of the problem.

  • ✗

    Increase the maximum batch size so more requests are processed per decode step.

    Why it's wrong here

    Raising the batch ceiling when memory is already at 98 percent worsens the problem: larger batches expand KV cache allocation, trigger more frequent cache eviction or out-of-memory handling, and lengthen each decode step. Inter-token latency grows because every sequence waits longer for its turn in a heavier step. The symptom described, degradation as concurrency rises, calls for reducing pressure rather than adding more concurrent work.

  • ✓

    Enable FP8 quantization for the KV cache and reduce the maximum batch size in the TensorRT-LLM build configuration.

    Why this is correct

    Near-saturated memory with growing concurrency means the KV cache is crowding out the workspace and forcing the scheduler to admit requests it cannot serve efficiently. FP8 KV cache roughly halves cache footprint, and lowering the maximum batch size keeps the runtime from over-admitting sequences. Together they reduce per-step memory pressure and shorten decode iterations, directly improving inter-token latency under high concurrency.

  • ✗

    Switch the service to a round-robin load balancer across two replicas of the same model on one GPU.

    Why it's wrong here

    Two replicas on a single GPU share the same memory and compute resources, so total KV cache capacity does not increase. Splitting requests across replicas adds scheduling overhead and can cause cache thrashing, while each replica still faces the same memory ceiling. This does not relieve the underlying capacity constraint that causes decode steps to slow as concurrency climbs.

About these practice questions

One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.