Courseiva
Experimentation →mediumMultiple Select

NCA-GENL Experimentation Practice Question

You are designing an experiment to measure how quantization (FP16 versus INT8) affects inference latency and answer quality for an LLM deployed with NVIDIA TensorRT-LLM. Which two practices are required for the comparison to be valid? (Choose two.)

⚠ Common exam trap

The trap here is assuming that a shared random seed makes FP16 and INT8 outputs directly comparable token by token.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Benchmark both configurations on the same hardware and under the same batch size and concurrency conditions.

A valid quantization comparison must control everything except precision. Using identical prompts and generation parameters ensures the workload is the same, and benchmarking on identical hardware under identical batch and concurrency conditions ensures the environment is the same. Together these practices let the team attribute measured differences in latency and quality to FP16 versus INT8 rather than to confounds.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Benchmark both configurations on the same hardware and under the same batch size and concurrency conditions.

    Why this is correct

    Latency depends heavily on GPU model, batch size, and concurrent request load. Measuring FP16 and INT8 on different hardware or load levels would make the results incomparable. Keeping hardware and serving conditions identical isolates precision as the variable under test.

  • ✗

    Retrain the model after quantization so the weights adapt to the lower precision.

    Why it's wrong here

    Quantization-aware training or fine-tuning is a separate technique, not a requirement for a valid inference comparison. The experiment measures the effect of quantization as applied, so retraining would change the object of study. It also adds cost and time without being necessary for the comparison.

  • ✓

    Use the same input prompts and generation parameters, such as max tokens and temperature, for both precision configurations.

    Why this is correct

    Holding prompts and generation parameters constant ensures that any latency or quality difference comes from precision rather than from different workloads. Without this control, the comparison is confounded and cannot attribute effects to quantization. It is a fundamental requirement for a valid experiment.

  • ✗

    Measure latency only on the first generated token, since subsequent tokens are less affected by precision.

    Why it's wrong here

    First-token latency reflects prefill and prompt processing, while later tokens reflect decoding, and both are affected by precision. Restricting measurement to one phase would miss the effect on overall throughput and quality. A valid comparison covers the full generation path.

  • ✗

    Apply the same random seed to both configurations so token sampling is identical.

    Why it's wrong here

    A shared seed does not make FP16 and INT8 outputs identical because the numerical representations differ. Seeding is useful for reproducibility within a configuration, but it cannot remove precision-induced differences. Requiring it here misunderstands the source of variation being studied.

About these practice questions

Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.