Courseiva
Model Optimization →hardMultiple Choice

NCP-GENL Model Optimization Practice Question

A team must serve a 70B model on a single 80 GB GPU for an internal assistant with modest concurrency. Full FP16 weights will not fit alongside the KV cache for the target context length. They want to keep accuracy loss minimal and are willing to spend additional build time. Which approach best fits these constraints?

⚠ Common exam trap

The trap here is reaching for a parallelism or context-shortening workaround when the real constraint is weight footprint on a single device that weight-only low-bit quantization directly addresses.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use INT4 weight-only quantization with a group size and calibration set chosen for the model.

The binding constraint is fitting a 70B model plus KV cache on one 80 GB GPU while preserving accuracy. Weight-only INT4 quantization cuts weight memory roughly fourfold while keeping activations and cache at higher precision, and group-wise scales with a representative calibration set limit degradation. The team's willingness to accept longer build time matches the calibration and engine-build cost this method requires.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Split the model with tensor parallelism across two GPUs even though only one is available.

    Why it's wrong here

    Tensor parallelism requires the weights to be partitioned across multiple devices with frequent collective communication. The scenario explicitly states a single 80 GB GPU is available, so this option is not deployable. It also does not reduce the numeric precision or memory requirement in a way that solves the single-device constraint.

  • ✗

    Apply FP8 quantization to both weights and activations and rebuild the engine.

    Why it's wrong here

    FP8 halves the weight footprint relative to FP16, which may still leave a 70B model too large for one 80 GB GPU once the KV cache is included. It also demands calibration to preserve accuracy. While FP8 is attractive on Hopper, INT4 weight-only offers a larger footprint reduction for the single-GPU constraint described.

  • ✗

    Keep FP16 weights and reduce the KV cache by capping the context length far below the requirement.

    Why it's wrong here

    The weights alone for a 70B model in FP16 exceed the available memory once the KV cache and runtime buffers are added. Cutting context length reduces the KV cache but does not make the weights fit, and it violates the stated context requirement. This fails on both the memory and the functional requirement.

  • ✓

    Use INT4 weight-only quantization with a group size and calibration set chosen for the model.

    Why this is correct

    Weight-only INT4 quantization shrinks the weight footprint roughly fourfold versus FP16, letting a 70B model fit on one 80 GB GPU while leaving the KV cache and activations in higher precision. Group-wise scales limit accuracy loss, and the extra build and calibration time is acceptable given the stated willingness, making this the best fit.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.