Courseiva
Software Development →mediumMultiple Choice

NCA-GENL Software Development Practice Question

A developer is building a document summarization service using an NVIDIA NIM microservice for Llama-3. The service must process batches of 20 documents at once to maximize throughput. The NIM container is already running with default settings. Which API parameter should the developer configure to enable efficient batched inference?

⚠ Common exam trap

The trap here is assuming that batching is controlled by a client-side parameter like `batch_size` in the request body, rather than by server-side configuration and sending an array of prompts.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure the `max_batch_size` parameter when starting the NIM container, and send multiple prompts in a single request using the `prompt` field as an array.

To efficiently process a batch of documents with an NVIDIA NIM microservice, the developer must ensure the server is configured with an appropriate `max_batch_size` at launch. Then, the client should send a single request containing an array of prompts. This leverages the NIM's dynamic batching capabilities, which group concurrent requests or batched prompts to maximize GPU utilization and throughput. Other options either misuse parameters or misunderstand the batching mechanism.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set the `batch_size` field in the request payload to 20, allowing the NIM to process all documents in a single forward pass.

    Why it's wrong here

    NVIDIA NIM microservices do not expose a `batch_size` field in the request payload. Batching is controlled via the `batch_size` parameter in the model configuration or the `--max-batch-size` flag when launching the container. Sending a `batch_size` field in the JSON payload will be ignored or cause a validation error, not enable batching.

  • ✗

    Enable the `tensor_parallel` option in the request headers and set it to 20, so the NIM distributes the batch across multiple GPUs.

    Why it's wrong here

    `tensor_parallel` is not a request header parameter; it is a model configuration setting that controls how a single model is split across GPUs. Setting it to 20 is unrealistic and unrelated to batching. Tensor parallelism affects model sharding, not batch processing, and must be configured at deployment time, not per request.

  • ✗

    Use the `stream` parameter set to `true` and send each document as a separate request; the NIM will automatically coalesce them into a batch on the server side.

    Why it's wrong here

    Setting `stream` to `true` enables token-by-token streaming responses for a single request; it does not trigger automatic batching across separate requests. While the server may perform dynamic batching, this is controlled by `max_batch_size`, not by the `stream` flag. Streaming individual requests would increase overhead and not guarantee batching.

  • ✓

    Configure the `max_batch_size` parameter when starting the NIM container, and send multiple prompts in a single request using the `prompt` field as an array.

    Why this is correct

    NVIDIA NIM microservices support dynamic batching at the server level, controlled by the `max_batch_size` parameter set during container launch. To process a batch, the client sends a single request with the `prompt` field as an array of strings. The server then groups these into a batch, improving GPU utilization and throughput for document summarization.

About these practice questions

This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.