NCA-GENL Core Machine Learning and AI Knowledge Practice Question
An ML engineer is deploying a Transformer-based inference service on an NVIDIA TensorRT-LLM runtime. To maximize inference throughput and reduce latency under heavy concurrent user traffic, the engineer needs to select the optimal decoding batching strategy. Which technique allows multiple incoming dynamic sequence requests to be batched together at the token level rather than waiting for entire sequences to complete?
⚠ Common exam trap
Candidates frequently confuse inflight batching with traditional static request batching or data-parallel batching. Static batching forces early-finishing requests to wait idly for the longest sequence in the batch to complete.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Inflight batching (continuous batching) that schedules individual generation iterations at the token level across requests.
Inflight batching (also known as continuous batching) dynamically groups individual generation steps from different requests at each token generation iteration. This eliminates the idle padding overhead and waiting times associated with traditional static batching, maximizing GPU compute utilization in LLM serving.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Static request batching that groups fixed numbers of incoming inference prompts and pads them to equal sequence lengths.
Why it's wrong here
Static batching pads every prompt to the longest sequence and holds the batch until all sequences finish, so short requests wait on long ones and GPU cycles are wasted on padding. It is the classic throughput approach for uniform workloads, but continuous batching is required for dynamic token-level scheduling.
- ✗
Data-parallel replica sharding that duplicates the model across separate GPU devices to process independent streams.
Why it's wrong here
Data-parallel replicas each process separate request streams, raising aggregate throughput but doing nothing to batch tokens from different sequences within one model instance. It suits scaling independent traffic across devices, whereas continuous batching is what merges dynamic requests at token level.
- ✓
Inflight batching (continuous batching) that schedules individual generation iterations at the token level across requests.
Why this is correct
Inflight batching schedules generation at the token level, admitting new requests into the running batch as others finish rather than waiting for whole sequences. This keeps the GPU saturated under concurrent load, maximising throughput and cutting latency compared with static batching.
- ✗
Model parallelism that splits individual transformer weight matrices across multiple interconnected GPU devices.
Why it's wrong here
Model parallelism partitions weights across GPUs to fit a model that exceeds single-device memory; it addresses capacity, not request scheduling, and adds inter-GPU communication latency. It would be the right choice when the model cannot fit on one GPU, not for batching concurrent dynamic sequences.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.