Courseiva
Model Optimization →hardMultiple Choice

NCP-GENL Model Optimization Practice Question

A team is serving a 70B-parameter LLM with TensorRT-LLM in a multi-tenant environment where requests arrive with widely varying prompt lengths and generation lengths. During load testing, they observe that throughput collapses when a long-context request is scheduled alongside many short requests, and GPU memory fragmentation causes intermittent out-of-memory errors even though total free memory appears sufficient. Which TensorRT-LLM runtime configuration change most directly addresses both the throughput collapse and the memory fragmentation?

⚠ Common exam trap

The trap here is assuming that increasing max_batch_size and max_seq_len gives the engine more flexibility, when in fact it reserves more memory and does nothing to prevent fragmentation or scheduling head-of-line blocking.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable the paged KV cache with a tuned block size and configure the batch scheduler to use in-flight batching with a per-iteration token budget.

The paged KV cache solves external fragmentation by allocating KV memory in uniform blocks from a shared pool, which prevents the OOM that occurs when a long-context request needs a large contiguous region. In-flight batching with a token budget lets the scheduler continuously admit and retire requests each iteration, so short requests no longer wait behind a long generation. These two runtime settings together address both symptoms without sacrificing accuracy or throughput.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the TensorRT-LLM build max_batch_size and max_seq_len to their maximum possible values so the engine supports every request shape.

    Why it's wrong here

    Raising max_batch_size and max_seq_len increases the static memory reservation for activation buffers and worst-case KV cache, which reduces available memory and can worsen, not fix, fragmentation and OOM. It does not implement paged memory management, so the long request still monopolizes contiguous space and short requests still wait, leaving the throughput collapse unaddressed.

  • ✗

    Switch the engine from FP16 to INT8 precision using a calibration dataset that matches the production prompt distribution.

    Why it's wrong here

    INT8 quantization reduces the per-token memory footprint and can improve throughput, but it does not change how KV cache blocks are allocated. Without paged memory management, the allocator still requires large contiguous regions, so fragmentation and OOM persist. It also introduces accuracy risk and does not fix scheduling head-of-line blocking from a long-context request.

  • ✗

    Reduce the number of concurrent client connections at the load balancer so that only one request is processed by the GPU at a time.

    Why it's wrong here

    Serializing requests eliminates the concurrent scheduling that causes contention, but it also removes the batching that gives the GPU its efficiency, so overall throughput drops sharply. It does not address memory fragmentation because the long request still needs a large contiguous KV allocation, and it sacrifices latency for all tenants rather than fixing the runtime configuration.

  • ✓

    Enable the paged KV cache with a tuned block size and configure the batch scheduler to use in-flight batching with a per-iteration token budget.

    Why this is correct

    Paged KV cache allocates KV memory in fixed-size blocks from a shared pool, eliminating the external fragmentation that causes OOM. In-flight batching lets the scheduler add new requests and retire finished ones every iteration, so short requests are not blocked behind a long one. Together they directly resolve both the memory fragmentation and the head-of-line blocking that collapses throughput.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.