Courseiva
Model Optimization →mediumMultiple Select

NCP-GENL Model Optimization Practice Question

An engineer is optimizing an LLM for inference with NVIDIA TensorRT-LLM and wants to reduce both memory footprint and latency without retraining the model. Which two techniques should they apply? (Choose two.)

⚠ Common exam trap

The trap here is selecting quantization-aware training because it sounds like a precision optimization, when it violates the no-retraining requirement.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Quantize model weights to INT8 or FP8 using post-training quantization with calibration.

Post-training quantization reduces weight memory and enables faster low-precision kernels, while paged KV cache manages generation-time memory efficiently and improves batching. Together they lower footprint and latency without retraining. The rejected options either require retraining, increase model size, or break the architecture, so they do not fit the no-retraining constraint or the optimization goals.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Quantize model weights to INT8 or FP8 using post-training quantization with calibration.

    Why this is correct

    Post-training weight quantization reduces the bytes per parameter, cutting memory footprint and enabling faster low-precision GEMM kernels on supported GPUs. It requires only calibration data, not retraining, which matches the constraint. This directly lowers both memory and latency, making it a primary optimization for production LLM inference in TensorRT-LLM.

  • ✓

    Enable paged KV cache with block-based memory management.

    Why this is correct

    Paged KV cache allocates memory in fixed-size blocks, reducing fragmentation and allowing more concurrent sequences to share the cache efficiently. It lowers peak memory pressure and improves batching, which in turn raises GPU utilization and reduces latency. It works alongside quantization and does not require retraining, so it is a core TensorRT-LLM optimization.

  • ✗

    Disable attention and rely solely on feed-forward layers to speed up decoding.

    Why it's wrong here

    Attention is essential to transformer language models; removing it would destroy the model's ability to use context and produce coherent output. This is not a supported optimization and would require retraining a fundamentally different architecture. It does not reduce memory or latency in a meaningful, production-valid way for an existing LLM.

  • ✗

    Increase the number of transformer layers to improve parallelism across the GPU.

    Why it's wrong here

    Adding transformer layers increases model size, memory consumption, and compute, which is the opposite of the stated optimization goals. It would also change the model architecture and require retraining. This option confuses model capacity scaling with inference optimization and would degrade both memory and latency rather than improve them.

  • ✗

    Retrain the model with quantization-aware training to learn optimal scaling factors.

    Why it's wrong here

    Quantization-aware training can improve accuracy but explicitly requires retraining, which the scenario forbids. While it is a valid technique in other contexts, it does not satisfy the no-retraining constraint here. Post-training quantization and paged KV cache deliver memory and latency benefits without modifying model weights through training.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.