Courseiva
Software Development →hardMultiple Select

NCA-GENL Software Development Practice Question

A developer is deploying a large language model using NVIDIA TensorRT-LLM and wants to optimize inference for a production environment with limited GPU memory. Which two techniques can be used to reduce memory footprint while maintaining acceptable performance? (Choose two.)

⚠ Common exam trap

The trap here is thinking that increasing tensor parallelism or batch size reduces memory footprint, when they actually distribute or increase memory usage, respectively.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enabling paged KV cache and tuning the maximum sequence length to a realistic value.

Quantization and paged KV cache with tuned sequence length are effective techniques to reduce GPU memory footprint in TensorRT-LLM. Quantization lowers precision of weights, and paged KV cache optimizes memory allocation for key-value pairs. Both help fit larger models on limited GPUs while preserving performance. Other options either increase memory usage or introduce latency, making them unsuitable for the scenario.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Using a larger batch size to amortize memory overhead across more requests.

    Why it's wrong here

    Increasing batch size typically increases memory usage because more sequences require KV cache storage simultaneously. While it can improve throughput, it does not reduce memory footprint; in fact, it may cause out-of-memory errors on limited GPUs. The goal is to reduce memory, so larger batch size is counterproductive.

  • ✓

    Enabling paged KV cache and tuning the maximum sequence length to a realistic value.

    Why this is correct

    Paged KV cache in TensorRT-LLM manages key-value cache memory in fixed-size blocks, reducing fragmentation and allowing more efficient memory usage. Tuning the maximum sequence length to a realistic value prevents over-allocation of KV cache for sequences that are never reached. Together, these techniques can significantly reduce memory footprint while maintaining performance for typical workloads.

  • ✗

    Increasing the tensor parallelism degree to spread the model across more GPUs.

    Why it's wrong here

    Increasing tensor parallelism spreads the model across multiple GPUs, which reduces per-GPU memory usage but requires more GPUs and introduces communication overhead. The question specifies limited GPU memory, but not necessarily multiple GPUs. Moreover, it does not reduce the total memory footprint; it just distributes it. The goal is to reduce memory footprint, so this is not the best choice.

  • ✓

    Quantization of model weights to INT8 or FP8 using TensorRT-LLM's quantization toolkit.

    Why this is correct

    Quantization reduces the precision of model weights, lowering memory usage and enabling faster computations. TensorRT-LLM supports INT8 and FP8 quantization, which can significantly decrease the memory footprint with minimal accuracy loss. This is a standard technique for deploying large models on GPUs with limited memory, as it allows more parameters to fit while maintaining acceptable performance.

  • ✗

    Storing the model weights on CPU and offloading them to GPU on demand.

    Why it's wrong here

    Offloading model weights to CPU and transferring them to GPU on demand introduces significant latency and PCIe bandwidth bottlenecks, severely degrading performance. While it can reduce GPU memory usage, it is not a technique for maintaining acceptable performance in a production environment. TensorRT-LLM is designed for GPU-resident inference; CPU offloading is generally not recommended for latency-sensitive applications.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.