NCP-GENL LLM Architecture Practice Question
An engineer is evaluating a sparse Mixture-of-Experts decoder-only LLM for a latency-sensitive inference service. They notice that although the model has far more total parameters than a dense baseline, throughput per token is only modestly better and sometimes worse. Which factor best explains why sparse MoE does not translate total parameter count into proportional speedup during inference?
⚠ Common exam trap
The trap here is equating parameter count with inference cost, when sparse MoE's bottleneck is memory bandwidth and expert communication rather than activated FLOPs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Only the top-k experts are activated per token, so memory bandwidth and all-to-all communication dominate rather than raw FLOPs.
Sparse MoE activates only a few experts per token, so FLOPs are low, but every expert's weights must remain resident and be routed to, often across devices. Inference becomes bound by memory bandwidth and all-to-all communication rather than compute, so a larger total parameter count does not yield proportional latency gains and can even hurt if routing is imbalanced.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Expert parameters are stored in FP32 while dense models use FP16, doubling memory traffic.
Why it's wrong here
There is no general rule that MoE experts are stored in higher precision than dense weights; both are typically held in the same low-precision format such as BF16 or FP8. The real memory-traffic issue is that all experts must be resident and fetched even though only a few are used per token, not a precision mismatch between MoE and dense models.
- ✗
The router must compute a full softmax over all experts and backpropagate through every expert for each token.
Why it's wrong here
During inference no backpropagation occurs at all, and the router only needs to score experts and select the top-k, not run every expert. The softmax is computed over router logits, which is cheap relative to expert computation. Stating that gradients flow through all experts describes training dynamics, not the inference bottleneck that limits speedup.
- ✓
Only the top-k experts are activated per token, so memory bandwidth and all-to-all communication dominate rather than raw FLOPs.
Why this is correct
Sparse MoE activates only a small subset of experts per token, so compute FLOPs are low, but the full expert weights must still reside in memory and be gathered, often across GPUs via all-to-all. Inference therefore becomes memory-bandwidth and communication bound rather than compute bound, which is why adding total parameters does not yield proportional latency improvements.
- ✗
The attention layers in MoE models are replaced by expert routing, removing the quadratic attention cost.
Why it's wrong here
MoE replaces or augments the feed-forward network with multiple expert MLPs; attention remains standard and still scales quadratically with sequence length. Claiming that routing removes attention cost misattributes where sparsity is applied. The attention computation is unchanged, so any inference cost from attention persists regardless of how many experts exist.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.