Courseiva
LLM Architecture →hardMultiple Select

NCP-GENL LLM Architecture Practice Question

An ML engineer is evaluating Mixture-of-Experts (MoE) routing for a large decoder-only model to increase capacity without proportionally increasing compute per token. Which TWO statements accurately describe how top-k token routing behaves in such an architecture? (Choose two.)

⚠ Common exam trap

The trap here is assuming that adding more experts automatically increases per-token compute, when in fact the top-k selection keeps activated expert count fixed regardless of total expert count.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Routing decisions are computed per token, allowing different tokens in the same sequence to be processed by different expert subsets

Sparse MoE replaces a single feed-forward block with many expert blocks plus a router that, for each token, selects the top-k experts to run. This yields token-level dynamic routing and keeps per-token FLOPs tied to k and expert size rather than the total expert population. Dense averaging, static assignment, and automatic k growth all contradict the sparse, learned, token-specific routing that defines the architecture.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Routing decisions are computed per token, allowing different tokens in the same sequence to be processed by different expert subsets

    Why this is correct

    The gating network produces a distribution over experts independently for each token, so token routing is dynamic and token-specific. This token-level granularity is what allows an MoE layer to specialize experts across linguistic or semantic patterns while keeping computation sparse for any individual token in the sequence.

  • ✗

    Increasing the number of experts proportionally increases the number of activated experts per token

    Why it's wrong here

    Adding experts expands total parameters and capacity but does not change k, the number of experts activated per token. The router still selects only the top-k experts, so per-token compute stays roughly constant while model capacity grows. This decoupling of capacity from compute is a central property of sparse MoE designs.

  • ✗

    The router assigns tokens to experts once at initialization and keeps the assignment fixed throughout training

    Why it's wrong here

    Router parameters are learned jointly with the rest of the network, and assignments change continuously as gating weights update. A static assignment would remove the adaptive specialization that makes MoE effective. Fixed routing also ignores input context, defeating the purpose of a learned gating function that responds to token representations.

  • ✓

    Only the experts selected by the router for a given token are activated, so FLOPs per token scale with k rather than total expert count

    Why this is correct

    In sparse MoE layers, the router scores each token against all experts, then dispatches the token only to the top-k experts. Because only those experts run their feed-forward computation for that token, per-token compute depends on k and expert size, not on how many experts exist in total, enabling capacity growth without proportional FLOPs.

  • ✗

    All experts receive every token, but their outputs are averaged with learned scalar weights

    Why it's wrong here

    That describes a dense ensemble or soft mixture, not sparse top-k routing. In sparse MoE, unselected experts do not receive or process the token at all. Averaging all expert outputs would eliminate the compute savings that motivate MoE and would contradict the sparse activation pattern the architecture is designed to provide.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.