Courseiva
Model Deployment →easyMultiple Choice

NCP-GENL Model Deployment Practice Question

A healthcare startup is deploying a Mistral 7B model for internal clinical note summarization. They need to serve the model with NVIDIA Triton Inference Server and want to minimize GPU memory footprint during inference. The team plans to use TensorRT-LLM and is choosing a numerical precision for the engine. Which precision should they select to reduce memory usage while maintaining acceptable accuracy for summarization?

⚠ Common exam trap

The trap here is equating TF32 with a memory-compression format, when it is actually a Tensor Core compute mode that does not shrink stored weights.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

INT8 or FP8 quantization, because it reduces weight and activation precision further than FP16 while maintaining acceptable accuracy for text summarization.

INT8 or FP8 quantization reduces the bits per weight and activation below FP16, cutting GPU memory usage while keeping accuracy acceptable for summarization. TensorRT-LLM supports quantized engines with calibration or scaling factors, making it the appropriate precision choice when memory footprint is the primary constraint and the task tolerates small accuracy loss.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    FP16, because it halves memory compared to FP32 and is widely supported on NVIDIA data center GPUs.

    Why it's wrong here

    FP16 does reduce memory relative to FP32 and is well supported on NVIDIA GPUs, but it is not the lowest-memory option available for LLM inference. For a 7B model where memory minimization is the explicit goal, INT8 or FP8 quantization can provide a smaller footprint with acceptable accuracy. FP16 is a reasonable baseline but not the best fit for the stated constraint.

  • ✗

    FP32, because it uses the least GPU memory and is the default for TensorRT-LLM engines.

    Why it's wrong here

    FP32 provides the highest numerical fidelity but consumes the most GPU memory, roughly four bytes per parameter plus activations. It is not the default for TensorRT-LLM LLM engines, which favor reduced precision for efficiency. Choosing FP32 would increase memory footprint, directly conflicting with the goal of minimizing GPU memory usage for the summarization service.

  • ✗

    TF32, because it is a Tensor Core mode that automatically compresses weights to one byte per parameter.

    Why it's wrong here

    TF32 is a Tensor Core compute mode that uses 19 bits for the mantissa and exponent representation in matrix multiplications, not a weight compression format. It does not store weights at one byte per parameter and does not reduce model memory the way INT8 or FP8 quantization does. Selecting TF32 would not satisfy the memory minimization goal for a 7B parameter model.

  • ✓

    INT8 or FP8 quantization, because it reduces weight and activation precision further than FP16 while maintaining acceptable accuracy for text summarization.

    Why this is correct

    INT8 and FP8 quantization reduce the number of bits per weight and activation compared to FP16, lowering GPU memory usage and often improving throughput. For summarization, a tolerant task, the accuracy loss is typically acceptable when calibration is done properly. This directly addresses the requirement to minimize memory footprint while keeping output quality suitable for clinical note summarization.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.