NCP-GENL Model Optimization Practice Question
A developer wants to reduce the disk and memory footprint of a fine-tuned 70B model before serving it with TensorRT-LLM, and is willing to accept a small, measurable quality drop that they will validate with an evaluation harness. Which approach best matches that requirement?
⚠ Common exam trap
The trap here is defaulting to quantization-aware training as the highest-quality option when the scenario explicitly accepts a small, measurable quality drop and does not provide for retraining.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Perform post-training quantization to INT4 or INT8 weights and validate with perplexity and task metrics.
The requirement is a smaller weight footprint with an accepted, validated quality trade-off and no retraining. Post-training quantization of the fine-tuned checkpoint to INT4 or INT8 delivers that reduction immediately and pairs naturally with an evaluation harness to confirm the quality delta stays within tolerance. Retraining-based methods exceed the stated scope.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Perform post-training quantization to INT4 or INT8 weights and validate with perplexity and task metrics.
Why this is correct
Post-training quantization compresses an already fine-tuned checkpoint without retraining, cutting weight storage and memory roughly 2x for INT8 and 4x for INT4. Because the developer accepts a small, measurable quality drop and will validate it with an evaluation harness, PTQ is the appropriate, low-effort match for the stated goal.
- ✗
Retrain the model from scratch using quantization-aware training with fake-quant nodes in the forward pass.
Why it's wrong here
Quantization-aware training inserts simulated quantization during training to let weights adapt, which generally yields better accuracy than PTQ. However, it requires the full training pipeline, data, and compute for a 70B model, far exceeding the developer's stated willingness to accept a small quality drop for a simpler footprint reduction.
- ✗
Convert the checkpoint to FP8 and rely on the runtime to upcast to FP16 for every matmul.
Why it's wrong here
FP8 storage would halve the footprint relative to FP16, but upcasting every matmul back to FP16 negates the compute benefit and is not how FP8 is used in practice. The scenario calls for a validated, supported footprint reduction, and INT4 or INT8 PTQ with an evaluation harness is the fitting method.
- ✗
Increase the tensor parallelism degree so each GPU holds a smaller shard of the FP16 weights.
Why it's wrong here
Tensor parallelism distributes weights across more GPUs, reducing per-GPU memory, but the aggregate footprint across the deployment is unchanged and disk storage is not reduced at all. It also requires more hardware, which contradicts the goal of shrinking the model's footprint before serving.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.