Diagnose LLM deployment and training failures by reasoning about attention, positional encoding, KV-cache memory, and parallelism strategy. The single most important thing: separate weight memory from KV-cache and activation memory when predicting whether a model fits on given NVIDIA GPUs.
Start practicing
LLM Architecture — choose a session length
Free · No account required
Domain overview
This domain covers the transformer internals that determine how NVIDIA LLMs are built, positioned, and served: attention variants, positional encoding schemes such as RoPE and YaRN, KV-cache behavior, quantization, and parallelism. Questions are scenario-based, asking you to diagnose OOM or scaling failures and pick the architectural feature that fixes them.
Exam objectives
Positional encoding choices: RoPE, YaRN scaling, and learned absolute vs. relative schemes for context extension
KV-cache memory growth and how it drives OOM beyond FP16 weight size on A100/H100 GPUs
Tensor, pipeline, and sequence parallelism for sharding massive decoder-only LLMs across multi-GPU nodes
Grouped-query and multi-query attention plus quantization for efficient large-model inference
Sizing GPU memory from parameter count alone, ignoring KV cache, activations, and framework overhead that cause load-time OOM.
Confusing YaRN context-window extension with fine-tuning or quantization; it rescales rotary frequencies, not weights.
Assuming tensor parallelism alone suffices; massive LLMs also need pipeline or sequence parallelism plus communication overlap.
Click any question to see the full explanation and answer options, or start a focused practice session above.
An engineer is designing a Transformer-based model for long-context document summarization. They decide to replace the standard dense self-attention mechanism with a sliding window attention approach. What is the primary architectural implication of this change?
2Refer to the exhibit. An engineer is fine-tuning an LLM using the provided configuration. What is the primary purpose of applying 'yarn' scaling in this architecture?
3Which TWO of the following statements correctly describe the role of Grouped Query Attention (GQA) in modern LLM architectures?
4In the context of the Transformer architecture, what is the primary function of the Feed-Forward Network (FFN) layers applied after the attention mechanism?
5Why do many modern LLMs use SwiGLU as their activation function in the feed-forward network instead of the traditional ReLU?
6Which THREE architectural features are essential for enabling efficient inference of massive LLMs on multi-GPU systems?
7Refer to the exhibit. The model is encountering an OOM error during long-context processing. Which architectural adjustment is most appropriate to resolve this while maintaining context length?
8In Mixture-of-Experts (MoE) architectures, why does the use of a router mechanism significantly impact performance compared to dense models?
9What is the architectural role of Layer Normalization in a Transformer, and where is it typically placed to ensure stable training?
10Which TWO of the following are benefits of using Rotary Positional Embeddings (RoPE) compared to absolute positional embeddings?
11What is the function of the 'Masked' component in a Decoder-only Transformer's self-attention during training?
12Which architectural component is responsible for projecting the model's hidden states back into the vocabulary space to predict the next token?
13A team is pretraining a decoder-only Transformer LLM on a large corpus of code and natural language. They observe that the model's training loss decreases smoothly, but during generation it sometimes produces degenerate repetition, and attention entropy on long sequences collapses. They suspect the issue is related to the positional encoding scheme. Which architectural change is most likely to mitigate the attention entropy collapse while preserving the model's ability to generalize to sequences longer than those seen during pretraining?
14A developer is building a decoder-only generative model and wants to prevent the model from attending to future tokens during training so that each position can only use information from itself and earlier positions. Which architectural mechanism should they implement in the self-attention layer?
15A team is training a large language model and wants to reduce the memory used by the optimizer without changing the model architecture. They are using Adam and notice that optimizer state consumes more GPU memory than the model weights. Which technique should they apply to reduce optimizer memory while keeping the model architecture unchanged?
16An engineer is deploying a 70B parameter model and needs to serve many concurrent users on a single GPU with limited memory. They want to store the attention keys and values for past tokens efficiently so that generation does not recompute them at every step. Which technique should they implement?
17A team is deploying a 13B-parameter decoder-only LLM on a single NVIDIA A100 40GB GPU for a real-time chatbot. During load, the process runs out of memory even though the model weights in FP16 require roughly 26GB. The team wants to reduce GPU memory usage with minimal impact on output quality and no change to the model architecture. Which technique is most appropriate?
18A team is pretraining a 13B-parameter decoder-only LLM on English text using byte-pair encoding with a 50,000-token vocabulary. They observe that the model produces fluent but repetitive continuations and that the average log-probability assigned to ground-truth tokens plateaus early. The training loss curve shows the model is underfitting rather than overfitting. Which architectural change is most likely to improve the model's capacity to capture long-range dependencies?
19A team is fine-tuning a pretrained decoder-only model on a small domain-specific dataset. They observe that the model quickly overfits and loses general language ability. They want to update only a small number of additional parameters while keeping the base weights frozen. Which approach should they use?
20A team is deploying a 70B-parameter decoder-only LLM on an NVIDIA H100 GPU node. During generation they observe that the KV cache grows linearly with sequence length and is consuming most of the available HBM, forcing them to limit batch size. They want to reduce KV cache memory without retraining the model from scratch. Which architectural change should they apply?
21A developer is inspecting a decoder-only Transformer and notices that during training, the model attends to future tokens in the sequence, causing the loss to drop unrealistically fast but generation to be incoherent. Which architectural mechanism is missing or misconfigured?
22An ML engineer is evaluating Mixture-of-Experts (MoE) routing for a large decoder-only model to increase capacity without proportionally increasing compute per token. Which TWO statements accurately describe how top-k token routing behaves in such an architecture? (Choose two.)
23A team is pre-training a 13B-parameter decoder-only LLM on a cluster of NVIDIA GPUs. They observe that gradient norms spike sharply during the first few hundred steps, destabilizing training. They want to keep the standard post-layer-normalization placement but stabilize early optimization. Which architectural technique should they apply?
24A research team is pretraining a decoder-only LLM and observes that gradient magnitudes in the earliest layers are extremely small while later layers train normally, causing slow convergence. They are using post-layer normalization. Which architectural change is most likely to improve gradient flow to the early layers?
25An engineer is reviewing the attention implementation of a decoder-only LLM used for chat. During inference with a KV cache, generated tokens must not attend to future positions. Which mechanism enforces this constraint inside scaled dot-product attention?
26An engineer is evaluating a sparse Mixture-of-Experts decoder-only LLM for a latency-sensitive inference service. They notice that although the model has far more total parameters than a dense baseline, throughput per token is only modestly better and sometimes worse. Which factor best explains why sparse MoE does not translate total parameter count into proportional speedup during inference?
27An engineer must serve a 70B-parameter LLM for a workload with many concurrent users and long shared system prompts, and wants to maximize throughput without retraining. Which inference-time optimization most directly reduces redundant computation across requests sharing the same prompt prefix?
28A developer is building a retrieval-augmented generation pipeline and needs to choose a component that produces dense vector representations of passages for semantic search. The passages are up to 512 tokens long, and the developer wants a model specifically trained to map semantically similar text to nearby points in embedding space. Which type of model should be selected?
29A research team wants to train a large decoder-only LLM where each token's representation is computed independently of token order, then inject order information afterward. They are choosing between learned absolute positional embeddings and sinusoidal absolute positional embeddings. Which statement accurately characterizes the tradeoff they face?
30An inference engineer is serving a 70B-parameter decoder-only LLM and wants to reduce KV cache memory to fit longer contexts on each GPU. They consider Multi-Query Attention (MQA), Grouped-Query Attention (GQA), and standard Multi-Head Attention (MHA). Which statement correctly describes the memory and quality tradeoff among these attention variants?
31An engineer is reviewing the architecture of a decoder-only LLM that must support very long input contexts for document analysis. They are considering architectural choices that extend effective context length beyond what the model saw during pretraining. Which TWO techniques are designed specifically to extend usable context length without retraining the entire model from scratch? (Choose two.)
32An engineer is analyzing why a decoder-only LLM with 32,000-token context length fails to answer questions that require information from the beginning of a long document when the answer is near the end. The model was trained with standard causal attention. Which two architectural or training factors are most likely contributing to this failure? (Choose two.)
Diagnose LLM deployment and training failures by reasoning about attention, positional encoding, KV-cache memory, and parallelism strategy. The single most important thing: separate weight memory from KV-cache and activation memory when predicting whether a model fits on given NVIDIA GPUs.
The Courseiva NCP-GENL question bank contains 32 questions in the LLM Architecture domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the LLM Architecture domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included