Courseiva
LLM Architecture →hardMultiple Choice

NCP-GENL LLM Architecture Practice Question

An engineer is evaluating a sparse Mixture-of-Experts decoder-only LLM for a latency-sensitive inference service. They notice that although the model has far more total parameters than a dense baseline, throughput per token is only modestly better and sometimes worse. Which factor best explains why sparse MoE does not translate total parameter count into proportional speedup during inference?

⚠ Common exam trap

The trap here is equating parameter count with inference cost, when sparse MoE's bottleneck is memory bandwidth and expert communication rather than activated FLOPs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Only the top-k experts are activated per token, so memory bandwidth and all-to-all communication dominate rather than raw FLOPs.

Sparse MoE activates only a few experts per token, so FLOPs are low, but every expert's weights must remain resident and be routed to, often across devices. Inference becomes bound by memory bandwidth and all-to-all communication rather than compute, so a larger total parameter count does not yield proportional latency gains and can even hurt if routing is imbalanced.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Expert parameters are stored in FP32 while dense models use FP16, doubling memory traffic.

    Why it's wrong here

    There is no general rule that MoE experts are stored in higher precision than dense weights; both are typically held in the same low-precision format such as BF16 or FP8. The real memory-traffic issue is that all experts must be resident and fetched even though only a few are used per token, not a precision mismatch between MoE and dense models.

  • ✗

    The router must compute a full softmax over all experts and backpropagate through every expert for each token.

    Why it's wrong here

    During inference no backpropagation occurs at all, and the router only needs to score experts and select the top-k, not run every expert. The softmax is computed over router logits, which is cheap relative to expert computation. Stating that gradients flow through all experts describes training dynamics, not the inference bottleneck that limits speedup.

  • ✓

    Only the top-k experts are activated per token, so memory bandwidth and all-to-all communication dominate rather than raw FLOPs.

    Why this is correct

    Sparse MoE activates only a small subset of experts per token, so compute FLOPs are low, but the full expert weights must still reside in memory and be gathered, often across GPUs via all-to-all. Inference therefore becomes memory-bandwidth and communication bound rather than compute bound, which is why adding total parameters does not yield proportional latency improvements.

  • ✗

    The attention layers in MoE models are replaced by expert routing, removing the quadratic attention cost.

    Why it's wrong here

    MoE replaces or augments the feed-forward network with multiple expert MLPs; attention remains standard and still scales quadratically with sequence length. Claiming that routing removes attention cost misattributes where sparsity is applied. The attention computation is unchanged, so any inference cost from attention persists regardless of how many experts exist.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.