Courseiva

NCA-GENL Core Machine Learning and AI Knowledge Practice Question

An ML engineer is deploying a Transformer-based inference service on an NVIDIA TensorRT-LLM runtime. To maximize inference throughput and reduce latency under heavy concurrent user traffic, the engineer needs to select the optimal decoding batching strategy. Which technique allows multiple incoming dynamic sequence requests to be batched together at the token level rather than waiting for entire sequences to complete?

⚠ Common exam trap

Candidates frequently confuse inflight batching with traditional static request batching or data-parallel batching. Static batching forces early-finishing requests to wait idly for the longest sequence in the batch to complete.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Inflight batching (continuous batching) that schedules individual generation iterations at the token level across requests.

Inflight batching (also known as continuous batching) dynamically groups individual generation steps from different requests at each token generation iteration. This eliminates the idle padding overhead and waiting times associated with traditional static batching, maximizing GPU compute utilization in LLM serving.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Static request batching that groups fixed numbers of incoming inference prompts and pads them to equal sequence lengths.

    Why it's wrong here

    Static batching pads every prompt to the longest sequence and holds the batch until all sequences finish, so short requests wait on long ones and GPU cycles are wasted on padding. It is the classic throughput approach for uniform workloads, but continuous batching is required for dynamic token-level scheduling.

  • ✗

    Data-parallel replica sharding that duplicates the model across separate GPU devices to process independent streams.

    Why it's wrong here

    Data-parallel replicas each process separate request streams, raising aggregate throughput but doing nothing to batch tokens from different sequences within one model instance. It suits scaling independent traffic across devices, whereas continuous batching is what merges dynamic requests at token level.

  • ✓

    Inflight batching (continuous batching) that schedules individual generation iterations at the token level across requests.

    Why this is correct

    Inflight batching schedules generation at the token level, admitting new requests into the running batch as others finish rather than waiting for whole sequences. This keeps the GPU saturated under concurrent load, maximising throughput and cutting latency compared with static batching.

  • ✗

    Model parallelism that splits individual transformer weight matrices across multiple interconnected GPU devices.

    Why it's wrong here

    Model parallelism partitions weights across GPUs to fit a model that exceeds single-device memory; it addresses capacity, not request scheduling, and adds inter-GPU communication latency. It would be the right choice when the model cannot fit on one GPU, not for batching concurrent dynamic sequences.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.