Courseiva
Model Deployment →hardMultiple Choice

NCP-GENL Model Deployment Practice Question

A media-analytics firm serves a 13B-parameter summarization model on two A100 GPUs using NVIDIA TensorRT-LLM behind Triton Inference Server. Traffic is bursty: during live events concurrency triples for about ten minutes, then returns to baseline. Operators report that the first requests after each burst begin are slow and sometimes time out, although steady-state latency is acceptable. Which deployment change most directly addresses the cold-start penalty at the beginning of each burst?

⚠ Common exam trap

The trap here is diagnosing burst-boundary latency as a batching or KV-cache capacity problem when it is actually resource initialization that occurs on first use.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure Triton's instance groups to keep a resident model instance and enable the model's warmup configuration so activation buffers and CUDA graphs are exercised before live traffic arrives.

The reported pattern, acceptable steady-state latency but slow and timing-out requests at the start of each burst, is a classic cold-start symptom. Keeping a resident Triton instance and supplying a warmup configuration forces engine loading, workspace allocation, and CUDA graph capture to happen at server startup rather than on the first live request. This removes the initialization stall precisely when traffic spikes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switch the Triton scheduler from dynamic batching to sequence batching so that in-flight conversations are tracked per client.

    Why it's wrong here

    Sequence batching is designed to maintain state across related requests such as multi-turn generation, and it changes how requests are correlated, not when GPU resources are allocated. The reported symptom occurs at burst onset before any sequence state matters. This option adds scheduling complexity without warming the engine, so the initial slow requests would persist.

  • ✓

    Configure Triton's instance groups to keep a resident model instance and enable the model's warmup configuration so activation buffers and CUDA graphs are exercised before live traffic arrives.

    Why this is correct

    Resident instances plus a warmup configuration cause Triton to load the engine and run representative inference requests at startup, allocating workspace, compiling or replaying CUDA graphs, and paging in weights. When the burst begins, those resources are already hot, so the first real requests no longer pay the initialization penalty. This directly targets the cold-start behavior described.

  • ✗

    Enable TensorRT-LLM's in-flight batching and raise the KV cache fraction so more concurrent sequences can share the cache.

    Why it's wrong here

    In-flight batching improves throughput under sustained concurrency by admitting new sequences mid-batch, and KV cache sizing affects how many sequences fit. Neither pre-allocates the execution resources that cause the first requests after idle periods to stall. The problem is a cold-start artifact at burst boundaries, not steady-state capacity, so tuning the cache fraction does not fix it.

  • ✗

    Increase the TensorRT-LLM engine's max_batch_size so larger batches can be formed during the burst peaks.

    Why it's wrong here

    Raising max_batch_size changes the largest batch the engine can process, but it does not pre-allocate or pre-warm anything. The slow first requests stem from resources being acquired or warmed on demand. A larger maximum batch could even increase memory pressure and make initial allocation slower, so it does not resolve the observed timeouts at burst onset.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.