Courseiva

NCA-GENL Core Machine Learning and AI Knowledge Practice Question

Which THREE techniques are commonly used to improve the efficiency of inference for large language models?

⚠ Common exam trap

Candidates often include training-specific techniques like gradient clipping or data augmentation, failing to recognize that the question specifically asks for methods targeting inference efficiency.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Weight Quantization to reduce the bit-width of model parameters.

Inference efficiency is critical for deploying LLMs in real-world production environments. Techniques like quantization reduce model size and memory requirements, KV-caching prevents redundant computation of attention keys and values, and pruning eliminates unnecessary weights. These methods combined allow models to run with significantly lower latency and reduced hardware costs on NVIDIA inference hardware, making high-performance AI more accessible and scalable for diverse enterprise applications.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Weight Quantization to reduce the bit-width of model parameters.

    Why this is correct

    Quantization converts model weights from higher precision (like FP32) to lower precision (INT8 or FP8). This substantially decreases the memory footprint and accelerates inference speed on NVIDIA GPUs, as hardware can process more low-precision operations in parallel with less bandwidth consumption and power usage.

  • ✓

    KV-Caching to store previously computed tokens during generation.

    Why this is correct

    During autoregressive decoding, the model computes key and value vectors for all previous tokens. KV-caching stores these results in GPU memory, avoiding redundant computations for each new token generated. This significantly reduces the time-per-token, which is essential for low-latency inference in chatbots and long-form generation.

  • ✗

    Increasing the number of transformer layers to add depth.

    Why it's wrong here

    Adding more layers increases the computational cost and latency of the model. For inference efficiency, one would typically seek to keep the model as compact as possible while maintaining performance, often through techniques like model distillation rather than increasing depth, which makes the model slower to run.

  • ✓

    Weight Pruning to remove redundant connections in the network.

    Why this is correct

    Pruning removes weights that contribute little to the model's output, resulting in a sparse representation. This reduces the total number of operations required during a forward pass. When implemented with hardware support for sparse matrix multiplication, this can lead to substantial speedups and memory efficiency during inference.

  • ✗

    Replacing all activations with Sigmoid functions for speed.

    Why it's wrong here

    Sigmoid functions are computationally more expensive than ReLU or GeLU variants and are generally unsuitable for deep hidden layers in modern transformers. Replacing efficient activations with sigmoid would degrade both inference quality and speed, making it a counterproductive strategy for improving overall model performance and efficiency.

About these practice questions

Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.