Courseiva

CCNA Llm Architecture Questions

32 questions · Llm Architecture topic · All types, answers revealed

1
MCQmedium

Why do many modern LLMs use SwiGLU as their activation function in the feed-forward network instead of the traditional ReLU?

A.It reduces the number of parameters by half compared to ReLU.
B.It allows the model to perform faster matrix multiplications.
C.It provides better gradient propagation and higher model performance.
D.It makes the model compatible with 4-bit quantization.
AnswerC

SwiGLU's gating mechanism allows for dynamic control over information flow, which leads to superior convergence rates and higher final perplexity scores compared to ReLU. The smoother gradient landscape facilitates training deeper models without encountering the zero-gradient issues that are common with strictly linear activation functions like ReLU.

Why this answer

SwiGLU is a gated linear unit that incorporates the Swish activation, providing a smoother gradient flow and improved representational capacity over ReLU. ReLU's 'dying gradient' problem can hinder training progress, whereas SwiGLU's multiplicative gating allows the model to learn more flexible feature activation patterns. This is fundamental for stabilizing the training of very deep models and achieving state-of-the-art performance in complex linguistic tasks.

Exam trap

Candidates often assume SwiGLU is about reducing compute cost or latency. While efficient, its primary advantage is the improvement of gradient flow and model representational capacity during training.

2
MCQmedium

A team is pre-training a 13B-parameter decoder-only LLM on a cluster of NVIDIA GPUs. They observe that gradient norms spike sharply during the first few hundred steps, destabilizing training. They want to keep the standard post-layer-normalization placement but stabilize early optimization. Which architectural technique should they apply?

A.Initialize all residual output projections with a small scaling factor such as 1/sqrt(2N).
B.Apply learning rate warmup over the first several thousand steps.
C.Insert RMSNorm immediately after the token embedding layer only.
D.Replace GELU activations in the feed-forward network with ReLU.
AnswerA

Scaling residual branch outputs at initialization by a depth-dependent factor keeps the variance of the residual stream roughly constant across layers. This directly tames the growth that causes exploding gradients in post-LN Transformers, allowing stable early training without moving normalization. It is a recognized architectural fix used in deep Transformer initialization schemes such as those in Megatron-style training.

Why this answer

Post-layer-normalization Transformers suffer from residual stream variance that grows with depth, producing large gradients early in training. Scaling residual branch outputs at initialization by a depth-dependent factor keeps variance bounded, directly stabilizing optimization without changing normalization placement or the optimizer schedule. Optimizer tweaks and activation swaps do not address the architectural root cause.

Exam trap

The trap here is assuming that learning rate warmup alone fixes post-LN instability when the underlying issue is depth-dependent residual variance growth.

3
MCQmedium

What is the architectural role of Layer Normalization in a Transformer, and where is it typically placed to ensure stable training?

A.It normalizes the entire batch to reduce training time.
B.It is placed before each sub-layer (Pre-LN) to improve training stability.
C.It is used to perform dimensionality reduction on input embeddings.
D.It replaces the need for residual connections in the model.
AnswerB

Pre-LN configurations move the normalization layer inside the residual branch, which creates a more stable gradient flow. This architectural choice is standard for modern LLMs as it significantly reduces the risk of loss divergence during the early stages of training compared to the older Post-LN approach.

Why this answer

Layer normalization stabilizes the hidden states by ensuring their mean and variance remain within a controlled range throughout the network layers. In modern LLMs, placing it before the attention and FFN blocks (Pre-LN) is preferred over the original Post-LN. This prevents gradient explosion early in training, allowing for higher learning rates and more reliable convergence for deep models during their pre-training phase.

Exam trap

Candidates often confuse Pre-LN with Post-LN. They may remember Layer Normalization exists but fail to recognize that Pre-LN is the modern standard for training stability in deep Transformers.

4
MCQeasy

A developer is inspecting a decoder-only Transformer and notices that during training, the model attends to future tokens in the sequence, causing the loss to drop unrealistically fast but generation to be incoherent. Which architectural mechanism is missing or misconfigured?

A.Pre-layer normalization placement before attention
B.Causal attention mask in the self-attention layers
C.Dropout regularization on attention weights
D.Rotary positional embedding frequency base
AnswerB

In a decoder-only model, self-attention must be masked so each position can only attend to itself and earlier positions. Without this causal mask, the model sees future tokens during training, trivially minimizing next-token loss while learning nothing useful for autoregressive generation, which exactly matches the described symptom.

Why this answer

Autoregressive language models require a causal mask that sets attention scores to negative infinity for positions after the current token. When that mask is absent, the training objective becomes trivial because the target token is visible in the input context, producing the fast loss drop and incoherent generation. Positional encodings, dropout, and normalization placement do not enforce temporal directionality.

Exam trap

The trap here is confusing positional encoding with causal masking, assuming that relative position information alone prevents a decoder from attending to future tokens.

5
MCQhard

An engineer must serve a 70B-parameter LLM for a workload with many concurrent users and long shared system prompts, and wants to maximize throughput without retraining. Which inference-time optimization most directly reduces redundant computation across requests sharing the same prompt prefix?

A.Increasing the beam width during generation
B.Speculative decoding with a smaller draft model
C.Enabling FP8 quantization of the KV cache
D.Prefix caching of key/value tensors for shared prompt prefixes
AnswerD

Prefix caching stores the key/value tensors computed for a shared prefix, such as a long system prompt, and reuses them across requests instead of recomputing attention for those tokens on every call. This directly eliminates the redundant prefill work the scenario describes and is supported in NVIDIA TensorRT-LLM and similar serving stacks.

Why this answer

When many requests share a long system prompt, recomputing attention over that prefix for every request wastes prefill compute. Prefix caching persists the key/value tensors for the shared prefix and reuses them, cutting redundant work and improving throughput under concurrency. Speculative decoding targets decode latency, beam search adds work, and KV cache quantization saves memory but not duplicated prefill computation.

Exam trap

The trap here is conflating memory-footprint optimizations such as KV cache quantization with compute-deduplication techniques like prefix caching, since both involve the KV cache but solve different bottlenecks.

6
MCQeasy

An engineer is reviewing the attention implementation of a decoder-only LLM used for chat. During inference with a KV cache, generated tokens must not attend to future positions. Which mechanism enforces this constraint inside scaled dot-product attention?

A.Dropping the scaling factor 1/sqrt(d_k) so that scores decay for distant tokens.
B.A causal mask that sets attention scores for future positions to negative infinity before the softmax.
C.Applying LayerNorm to the query and key projections before computing attention scores.
D.Shifting the position IDs of cached keys so they appear before the current token.
AnswerB

A causal (lower-triangular) mask adds negative infinity to scores at positions beyond the current token, so after softmax those weights become zero. This guarantees each position only attends to itself and earlier tokens, which is exactly the autoregressive constraint required for decoder-only generation, and it works identically whether or not a KV cache is used.

Why this answer

Causal masking is the standard mechanism that makes decoder-only attention autoregressive. By adding negative infinity to scores of future positions before softmax, those positions receive zero weight, so each token can only attend to itself and prior tokens. Other listed techniques affect numerical stability or positional encoding but do not enforce the temporal constraint.

Exam trap

The trap here is confusing numerical stabilization techniques like score scaling or QK-norm with the masking mechanism that actually enforces causal visibility.

7
MCQhard

In Mixture-of-Experts (MoE) architectures, why does the use of a router mechanism significantly impact performance compared to dense models?

A.It forces all parameters to update at every training step.
B.It allows scaling parameters while keeping compute costs manageable.
C.It forces the model to use all experts simultaneously for every input.
D.It replaces the attention mechanism with a standard linear layer.
AnswerB

MoE allows for a massive total parameter count while maintaining a constant amount of active parameters per token. By routing tokens to specialized experts, the model achieves high capacity and knowledge breadth without the linear compute cost increases that would occur in a fully dense architecture of equivalent size.

Why this answer

The router determines which expert blocks to activate for a given input token, ensuring that only a subset of the model's total parameters are used per forward pass. This decoupling of model size from computation latency allows for the creation of massive, high-capacity models that run with the speed of much smaller dense models. This architectural choice is central to modern scaling laws.

Exam trap

Candidates mistakenly think MoE architectures reduce total parameter count, confusing sparse activation of routing with physical parameter pruning or compression.

8
Multi-Selecthard

An engineer is analyzing why a decoder-only LLM with 32,000-token context length fails to answer questions that require information from the beginning of a long document when the answer is near the end. The model was trained with standard causal attention. Which two architectural or training factors are most likely contributing to this failure? (Choose two.)

Select 2 answers
A.The feed-forward network in each layer compresses the hidden state, discarding information from early tokens.
B.The vocabulary size is too small to represent the document's domain-specific terms, causing tokenization errors.
C.The positional encoding scheme may not generalize well beyond the sequence lengths seen during pretraining.
D.Causal attention masks prevent each token from attending to future tokens, which limits bidirectional reasoning over the document.
E.The model was pretrained primarily on sequences much shorter than 32,000 tokens, so it did not learn to attend across the full context.
AnswersC, E

Many positional encoding schemes, especially learned absolute embeddings or RoPE without scaling, degrade when sequences exceed the lengths seen during training. The model cannot reliably distinguish or weight positions far beyond its training range, so attention to early tokens becomes imprecise. This directly contributes to the failure to retrieve information from the beginning of a long document.

Why this answer

Long-context retrieval failures in decoder-only LLMs typically stem from two sources: pretraining on sequences much shorter than the target context, which prevents the model from learning to attend across the full window, and positional encoding schemes that do not generalize beyond training lengths. Together these cause the model to underuse information at the start of a long document. The causal mask, feed-forward network, and vocabulary size are not primary causes of this behavior.

Exam trap

The trap here is blaming the causal attention mask for limiting long-context reasoning, when causal masking is a defining feature of decoder-only models and does not prevent attending to earlier tokens.

9
MCQeasy

A developer is building a retrieval-augmented generation pipeline and needs to choose a component that produces dense vector representations of passages for semantic search. The passages are up to 512 tokens long, and the developer wants a model specifically trained to map semantically similar text to nearby points in embedding space. Which type of model should be selected?

A.An autoregressive sequence-to-sequence model fine-tuned for summarization.
B.A cross-encoder that jointly encodes the query and passage.
C.A decoder-only generative LLM fine-tuned for instruction following.
D.A bidirectional encoder model trained with a contrastive objective on sentence pairs.
AnswerD

Bidirectional encoders trained with contrastive objectives, such as sentence-transformer style models, are explicitly optimized to place semantically similar passages close together in vector space. They produce a single dense embedding per passage and are efficient for semantic search over 512-token chunks. This matches the developer's requirement for a model specifically trained for embedding-based retrieval.

Why this answer

Dense retrieval requires a model that maps each passage to a single vector such that semantically similar passages are close in the embedding space. Bidirectional encoders trained with contrastive objectives are purpose-built for this, producing high-quality passage embeddings efficiently. Generative LLMs, cross-encoders, and summarization models are designed for other tasks and do not provide the required independent passage vectors.

Exam trap

The trap here is assuming any transformer can serve as an embedding model, when in fact only models trained with a contrastive or similar objective produce reliable dense vectors for semantic search.

10
MCQmedium

A team is training a large language model and wants to reduce the memory used by the optimizer without changing the model architecture. They are using Adam and notice that optimizer state consumes more GPU memory than the model weights. Which technique should they apply to reduce optimizer memory while keeping the model architecture unchanged?

A.Replace the feed-forward activation with a gated variant to reduce activation memory.
B.Switch to a memory-efficient optimizer such as 8-bit Adam or Adafactor that stores reduced-precision or factored optimizer state.
C.Reduce the number of attention heads while keeping the hidden size constant.
D.Increase the batch size proportionally to the number of GPUs so optimizer states are shared across more tokens.
AnswerB

8-bit Adam quantizes the first and second moment estimates to 8-bit, and Adafactor factors the second moment, both cutting optimizer memory substantially. These keep the model architecture and parameter count intact while lowering the dominant memory cost during training.

Why this answer

Adam maintains two full-precision moment estimates for every trainable parameter, so optimizer state can be several times the size of the model weights. Memory-efficient optimizers such as 8-bit Adam or Adafactor reduce this footprint by quantizing or factoring the state, directly lowering the dominant memory cost without changing the model architecture.

Exam trap

The trap here is assuming that larger batches or architectural tweaks solve optimizer memory, when the moments themselves are the cost.

11
Multi-Selectmedium

Which TWO of the following are benefits of using Rotary Positional Embeddings (RoPE) compared to absolute positional embeddings?

Select 2 answers
A.RoPE allows for better extrapolation to unseen sequence lengths.
B.RoPE completely removes the need for attention mechanisms.
C.RoPE improves the model's ability to capture relative token order.
D.RoPE reduces the parameter count of the embedding layer.
E.RoPE requires retraining from scratch if sequence length increases.
AnswersA, C

Because RoPE relies on rotational transformations, the relative distance between tokens is preserved regardless of absolute index. This property makes it easier to extend the context window during inference, as the model can interpret relative relationships between tokens even when they appear at indices beyond the training limit.

Why this answer

RoPE encodes relative position information by rotating vectors in complex space, which allows for better generalization to sequence lengths not seen during training. This relative approach is significantly more effective than absolute embeddings, which assign a fixed position to every index, as it allows the model to interpret the relationships between tokens regardless of their exact position in the sequence.

Exam trap

Candidates mistakenly choose absolute positional embedding benefits, confusing how standard positional encodings behave with RoPE's dynamic complex-space rotation mechanism designed for relative distances.

12
MCQmedium

A team is pretraining a 13B-parameter decoder-only LLM on English text using byte-pair encoding with a 50,000-token vocabulary. They observe that the model produces fluent but repetitive continuations and that the average log-probability assigned to ground-truth tokens plateaus early. The training loss curve shows the model is underfitting rather than overfitting. Which architectural change is most likely to improve the model's capacity to capture long-range dependencies?

A.Reduce the vocabulary size to 32,000 tokens to lower the softmax computation cost.
B.Increase the number of attention heads while keeping the hidden dimension fixed.
C.Increase the number of transformer layers and the hidden dimension proportionally.
D.Replace the learned positional embeddings with sinusoidal absolute positional encodings.
AnswerC

Underfitting at 13B parameters on English text signals insufficient model capacity. Scaling depth (more layers) and width (larger hidden dimension) together increases both the number of sequential transformation steps and the representational bandwidth, allowing the network to model longer-range dependencies and richer token distributions. This directly addresses the plateau in log-probability on ground-truth tokens.

Why this answer

Underfitting in a large decoder-only LLM indicates the model lacks sufficient capacity for the data distribution. Scaling both depth and width increases the parameter count and the number of nonlinear transformations applied to the residual stream, which improves the model's ability to capture long-range dependencies. The other options either rebalance existing capacity, change positional encoding without adding parameters, or reduce computation without improving fit.

Exam trap

The trap here is assuming that any architectural tweak that increases parameter count (such as adding attention heads) will resolve underfitting, when in fact head rebalancing at a fixed hidden dimension does not add representational capacity.

13
Multi-Selecthard

An ML engineer is evaluating Mixture-of-Experts (MoE) routing for a large decoder-only model to increase capacity without proportionally increasing compute per token. Which TWO statements accurately describe how top-k token routing behaves in such an architecture? (Choose two.)

Select 2 answers
A.Routing decisions are computed per token, allowing different tokens in the same sequence to be processed by different expert subsets
B.Increasing the number of experts proportionally increases the number of activated experts per token
C.The router assigns tokens to experts once at initialization and keeps the assignment fixed throughout training
D.Only the experts selected by the router for a given token are activated, so FLOPs per token scale with k rather than total expert count
E.All experts receive every token, but their outputs are averaged with learned scalar weights
AnswersA, D

The gating network produces a distribution over experts independently for each token, so token routing is dynamic and token-specific. This token-level granularity is what allows an MoE layer to specialize experts across linguistic or semantic patterns while keeping computation sparse for any individual token in the sequence.

Why this answer

Sparse MoE replaces a single feed-forward block with many expert blocks plus a router that, for each token, selects the top-k experts to run. This yields token-level dynamic routing and keeps per-token FLOPs tied to k and expert size rather than the total expert population. Dense averaging, static assignment, and automatic k growth all contradict the sparse, learned, token-specific routing that defines the architecture.

Exam trap

The trap here is assuming that adding more experts automatically increases per-token compute, when in fact the top-k selection keeps activated expert count fixed regardless of total expert count.

14
MCQhard

A team is fine-tuning a pretrained decoder-only model on a small domain-specific dataset. They observe that the model quickly overfits and loses general language ability. They want to update only a small number of additional parameters while keeping the base weights frozen. Which approach should they use?

A.Full fine-tuning with a very low learning rate and early stopping.
B.Increasing the dropout rate in the feed-forward layers during fine-tuning.
C.Freezing the embedding layer only and training all transformer blocks.
D.Low-Rank Adaptation (LoRA), which injects trainable low-rank matrices into existing layers while freezing the base weights.
AnswerD

LoRA adds small trainable rank-decomposition matrices to selected weight matrices and leaves the pretrained weights frozen. This updates far fewer parameters, reduces optimizer memory, and is less prone to catastrophic forgetting, which matches the team's constraints.

Why this answer

Parameter-efficient fine-tuning methods such as LoRA freeze the pretrained weights and train only small injected matrices, which reduces optimizer memory and limits drift from the base model. This is well suited to small domain datasets where full fine-tuning would overfit and degrade general language ability.

Exam trap

The trap here is thinking that a low learning rate or extra dropout makes full fine-tuning parameter-efficient, when the base weights are still being updated.

15
MCQmedium

A team is pretraining a decoder-only Transformer LLM on a large corpus of code and natural language. They observe that the model's training loss decreases smoothly, but during generation it sometimes produces degenerate repetition, and attention entropy on long sequences collapses. They suspect the issue is related to the positional encoding scheme. Which architectural change is most likely to mitigate the attention entropy collapse while preserving the model's ability to generalize to sequences longer than those seen during pretraining?

A.Increase the number of attention heads while keeping the total hidden dimension constant, without changing the positional encoding.
B.Apply Layer Normalization before the self-attention sublayer instead of after it, leaving the positional encoding unchanged.
C.Replace learned absolute positional embeddings with Rotary Positional Embeddings (RoPE) applied to queries and keys in every self-attention layer.
D.Add a sinusoidal positional encoding to the input embeddings and freeze those embeddings during training.
AnswerC

RoPE encodes relative position by rotating query and key vectors by an angle proportional to their absolute position. Because attention scores depend only on the relative rotation between queries and keys, the model naturally generalizes to longer contexts and avoids the entropy collapse seen when learned absolute embeddings are extrapolated beyond their trained range. This matches the observed repetition and entropy issues.

Why this answer

The symptoms point to a positional encoding that does not generalize beyond the pretraining sequence length. Rotary Positional Embeddings inject relative position directly into the attention computation by rotating queries and keys, so attention scores depend on relative offsets rather than absolute indices. This preserves attention entropy on longer sequences and reduces degenerate repetition, making it the appropriate architectural change for this scenario.

Exam trap

The trap here is assuming that any positional encoding change, such as switching to sinusoidal or adding normalization, will fix length generalization, when only a relative-position scheme applied inside attention addresses the described entropy collapse.

16
Multi-Selectmedium

Which TWO of the following statements correctly describe the role of Grouped Query Attention (GQA) in modern LLM architectures?

Select 2 answers
A.GQA eliminates the need for separate query, key, and value projection layers.
B.GQA reduces the memory footprint of the KV cache during inference.
C.GQA improves model training speed by increasing the number of compute operations.
D.GQA achieves a compromise between Multi-Head Attention and Multi-Query Attention.
E.GQA prevents the model from using positional embeddings.
AnswersB, D

By reducing the total number of key and value heads, the size of the KV cache stored in GPU memory is significantly decreased. This reduction is highly beneficial for serving models at scale, as it allows for larger batch sizes or longer context lengths without exceeding available memory capacity.

Why this answer

GQA balances the performance of Multi-Head Attention (MHA) and the memory efficiency of Multi-Query Attention (MQA). By sharing keys and values across groups of heads, it reduces the size of the KV cache during inference. This is crucial for high-throughput deployment environments, as it optimizes memory bandwidth utilization without sacrificing the granular representational capacity that individual query heads provide.

Exam trap

Candidates often mistake GQA for a training-only optimization or a method that increases accuracy. GQA is specifically designed for inference-time efficiency by reducing memory bandwidth requirements.

17
MCQhard

Refer to the exhibit. An engineer is fine-tuning an LLM using the provided configuration. What is the primary purpose of applying 'yarn' scaling in this architecture?

A.It increases the number of attention heads to improve feature extraction.
B.It enables the model to process sequences longer than the original pre-trained limit.
C.It compresses the model weights to reduce storage requirements.
D.It replaces the attention mechanism with a recurrent neural network.
AnswerB

The primary goal of YaRN is to allow effective extrapolation and interpolation of positional information. By adjusting the base frequency of RoPE, the model can interpret position indices that fall outside the range seen during initial training, thereby allowing for a significantly larger context window during inference and fine-tuning.

Why this answer

YaRN (Yet another RoPE extension) is designed to extend the context window of pre-trained LLMs beyond their original training length. By modifying the Rotary Positional Embeddings (RoPE) base frequency, it allows the model to interpolate position indices without severe degradation in perplexity. This is essential for enterprise deployments requiring document retrieval or analysis tasks where input lengths exceed the base model's pre-trained constraints.

Exam trap

Candidates often confuse YaRN with quantization or pruning techniques. They assume any acronym related to model optimization is about reducing memory footprint rather than extending the context window.

18
MCQmedium

Refer to the exhibit. The model is encountering an OOM error during long-context processing. Which architectural adjustment is most appropriate to resolve this while maintaining context length?

A.Increase the number of hidden layers.
B.Switch to FlashAttention-2 kernels to optimize memory usage.
C.Reduce the embedding dimension size.
D.Disable the KV cache entirely.
AnswerB

FlashAttention-2 provides a fused kernel that computes attention in blocks, avoiding the storage of the full attention matrix in VRAM. This is the optimal way to handle long sequences like 128k, as it solves the memory bottleneck at the architectural kernel level without sacrificing model capability.

Why this answer

When sequence lengths scale to 128k, the attention matrix size grows quadratically, consuming massive memory. Implementing FlashAttention-2 or similar memory-efficient attention kernels is the industry-standard solution. These kernels optimize the memory layout and tile operations to compute attention without materializing the massive N×N matrix, drastically reducing peak memory usage and enabling the processing of very long sequences within existing GPU capacity.

Exam trap

Candidates often suggest reducing the batch size or model precision. While these help, they do not address the fundamental quadratic memory growth of attention that FlashAttention-2 is designed to solve.

19
Multi-Selecthard

An engineer is reviewing the architecture of a decoder-only LLM that must support very long input contexts for document analysis. They are considering architectural choices that extend effective context length beyond what the model saw during pretraining. Which TWO techniques are designed specifically to extend usable context length without retraining the entire model from scratch? (Choose two.)

Select 2 answers
A.Switching the feed-forward activation from SwiGLU to GeLU to improve long-range dependency modeling.
B.Position interpolation, which rescales positional indices so longer sequences fall within the range seen during pretraining.
C.Increasing the number of attention heads while keeping the head dimension fixed.
D.YaRN scaling, which adjusts rotary positional embedding frequencies to improve extrapolation to longer sequences.
E.Reducing the vocabulary size to shorten the embedding matrix and free memory for longer sequences.
AnswersB, D

Position interpolation compresses the positional index range so that a longer sequence maps into the positions the model encountered during pretraining. This allows the model to generalize to longer contexts with only brief fine-tuning. It directly targets the mismatch between pretraining context length and desired inference length, making it a standard context-extension technique.

Why this answer

Position interpolation and YaRN both operate on positional encoding to reconcile desired sequence lengths with the range seen during pretraining. Interpolation compresses indices, while YaRN adjusts rotary frequencies to improve extrapolation. Both are context-extension methods that require only light fine-tuning, unlike changes to head count, MLP activation, or vocabulary size, which do not address positional generalization.

Exam trap

The trap here is treating any memory-saving or capacity change as context extension, when only techniques that alter positional encoding behavior actually lengthen usable context.

20
MCQhard

An inference engineer is serving a 70B-parameter decoder-only LLM and wants to reduce KV cache memory to fit longer contexts on each GPU. They consider Multi-Query Attention (MQA), Grouped-Query Attention (GQA), and standard Multi-Head Attention (MHA). Which statement correctly describes the memory and quality tradeoff among these attention variants?

A.GQA eliminates the KV cache entirely by recomputing keys and values on the fly during decoding.
B.GQA uses a single key and value head shared by all query heads, matching MQA memory savings while preserving MHA quality.
C.MHA has the smallest KV cache because each head stores its own keys and values independently.
D.MQA shares key and value projections across all query heads, giving the smallest KV cache but often degrading quality relative to MHA.
AnswerD

In MQA, every query head uses the same key and value head, so the KV cache stores only one K and one V vector per layer per token, dramatically reducing memory. This aggressive sharing can hurt model quality and training stability compared with MHA, which is why intermediate GQA designs were introduced. The statement accurately captures both the memory benefit and the quality risk.

Why this answer

MQA minimizes KV cache by sharing one K and V head across all query heads, at some cost to quality. GQA interpolates by grouping query heads with dedicated K and V heads, balancing memory and quality. MHA provides the richest representation but the largest cache.

The other options misstate GQA's structure, reverse MHA's memory ranking, or incorrectly claim GQA removes the cache.

Exam trap

The trap here is conflating GQA's grouped key-value heads with MQA's single shared head, which understates GQA's memory footprint.

21
MCQmedium

Which architectural component is responsible for projecting the model's hidden states back into the vocabulary space to predict the next token?

A.The Feed-Forward Network.
B.The Softmax normalization layer.
C.The LM Head linear layer.
D.The positional embedding layer.
AnswerC

The LM Head is a linear projection layer that maps the final hidden state to the vocabulary size. This is the standard architectural design for generative models, allowing the transformer to map abstract internal features to specific token IDs that correspond to the model's fixed training vocabulary.

Why this answer

The language model head (LM Head) is a final linear projection layer that maps the high-dimensional hidden states from the Transformer blocks into a vector representing the probability distribution over the entire vocabulary. This projection is the final step in the forward pass, converting the learned internal representation into actionable predictions that can be sampled to generate text.

Exam trap

Candidates frequently confuse the LM Head with intermediate Transformer attention blocks or embedding layers, forgetting which component maps hidden states directly to the vocabulary.

22
MCQmedium

A team is deploying a 70B-parameter decoder-only LLM on an NVIDIA H100 GPU node. During generation they observe that the KV cache grows linearly with sequence length and is consuming most of the available HBM, forcing them to limit batch size. They want to reduce KV cache memory without retraining the model from scratch. Which architectural change should they apply?

A.Switch the feed-forward network activation from SwiGLU to ReLU to reduce parameter count.
B.Apply rotary positional embeddings to the key and query vectors instead of learned absolute embeddings.
C.Increase the number of attention heads while keeping the head dimension constant.
D.Replace multi-head attention with Grouped Query Attention (GQA), where multiple query heads share a single key/value head.
AnswerD

GQA reduces the number of distinct key and value projections by having groups of query heads share KV heads, which directly shrinks the KV cache proportionally to the number of KV heads rather than query heads. This lowers HBM pressure during inference and allows larger batch sizes, and it can be introduced through continued pretraining or uptraining rather than a full retrain from scratch.

Why this answer

Grouped Query Attention reduces the number of key/value heads relative to query heads, so the KV cache that must be stored for each token shrinks by the grouping factor. Because the cache dominates HBM during long-context decoding, this architectural change directly relieves memory pressure and permits larger batches, and it can be adopted via uptraining rather than a full retrain.

Exam trap

The trap here is assuming that any attention efficiency change such as RoPE or more heads reduces KV cache size, when only reducing the number of KV heads actually shrinks the cache.

23
MCQhard

An engineer is evaluating a sparse Mixture-of-Experts decoder-only LLM for a latency-sensitive inference service. They notice that although the model has far more total parameters than a dense baseline, throughput per token is only modestly better and sometimes worse. Which factor best explains why sparse MoE does not translate total parameter count into proportional speedup during inference?

A.Expert parameters are stored in FP32 while dense models use FP16, doubling memory traffic.
B.The router must compute a full softmax over all experts and backpropagate through every expert for each token.
C.Only the top-k experts are activated per token, so memory bandwidth and all-to-all communication dominate rather than raw FLOPs.
D.The attention layers in MoE models are replaced by expert routing, removing the quadratic attention cost.
AnswerC

Sparse MoE activates only a small subset of experts per token, so compute FLOPs are low, but the full expert weights must still reside in memory and be gathered, often across GPUs via all-to-all. Inference therefore becomes memory-bandwidth and communication bound rather than compute bound, which is why adding total parameters does not yield proportional latency improvements.

Why this answer

Sparse MoE activates only a few experts per token, so FLOPs are low, but every expert's weights must remain resident and be routed to, often across devices. Inference becomes bound by memory bandwidth and all-to-all communication rather than compute, so a larger total parameter count does not yield proportional latency gains and can even hurt if routing is imbalanced.

Exam trap

The trap here is equating parameter count with inference cost, when sparse MoE's bottleneck is memory bandwidth and expert communication rather than activated FLOPs.

24
MCQmedium

An engineer is deploying a 70B parameter model and needs to serve many concurrent users on a single GPU with limited memory. They want to store the attention keys and values for past tokens efficiently so that generation does not recompute them at every step. Which technique should they implement?

A.Quantization-aware training that reduces weight precision before deployment.
B.A key-value cache that stores the projected keys and values for previously processed tokens.
C.Speculative decoding that uses a smaller draft model to propose tokens.
D.Gradient checkpointing that recomputes activations during the backward pass.
AnswerB

A key-value cache retains the key and value projections for all past tokens so each new step only computes the query and attends to cached entries. This avoids recomputing the full prefix at every decoding step, dramatically reducing compute for autoregressive generation.

Why this answer

Autoregressive decoding generates one token at a time, and without caching the model would recompute keys and values for the entire prefix at every step, wasting compute. A key-value cache stores those projections so each step only processes the new token, which is the standard way to make long-context generation practical.

Exam trap

The trap here is treating training-time memory tricks like gradient checkpointing as if they accelerated autoregressive inference.

25
MCQeasy

In the context of the Transformer architecture, what is the primary function of the Feed-Forward Network (FFN) layers applied after the attention mechanism?

A.To calculate the attention scores between different tokens in the sequence.
B.To process the hidden state representations with non-linear activations.
C.To manage the memory allocation for the KV cache during multi-token generation.
D.To reduce the sequence length of the input tokens to a fixed size.
AnswerB

The FFN typically consists of two linear transformations with a non-linear activation function, like SwiGLU or ReLU, in between. This structure enables the model to learn complex mappings of input features, significantly increasing its capacity to understand nuances that linear attention projections alone might fail to capture effectively.

Why this answer

The FFN layers provide non-linear transformations that allow the model to process information extracted by the attention mechanism. While attention focuses on relationships between tokens, the FFN applies point-wise non-linearities to project these features into higher-dimensional spaces. This is essential for learning complex representations and mappings that ultimately drive the model's predictive accuracy and reasoning capabilities across diverse input patterns.

Exam trap

Candidates often confuse the FFN's role with the attention mechanism's role. They incorrectly attribute the 'capturing of relationships between tokens' to the FFN, rather than the attention mechanism.

26
MCQmedium

An engineer is designing a Transformer-based model for long-context document summarization. They decide to replace the standard dense self-attention mechanism with a sliding window attention approach. What is the primary architectural implication of this change?

A.The model's total parameter count increases due to added positional embedding layers.
B.The model loses the ability to perform parallel training across multiple GPU nodes.
C.The computational complexity of the self-attention mechanism is reduced from quadratic to linear.
D.The model is no longer compatible with standard softmax normalization functions.
AnswerC

By limiting the attention span to a fixed window size, the number of operations per token becomes constant rather than proportional to the sequence length. This shift from O(n²) to O(n*w) complexity is the fundamental architectural advantage for long-context tasks, enabling processing of documents that would otherwise be computationally prohibitive.

Why this answer

Sliding window attention restricts the receptive field of each token to a local neighborhood, drastically reducing the quadratic memory complexity of standard attention to linear. This is critical for scaling LLMs to long contexts, as it prevents the O(n²) memory growth that typically causes GPU out-of-memory errors on large input sequences while maintaining local coherence.

Exam trap

Candidates often focus on the 'loss of accuracy' or 'semantic degradation.' While these are concerns, the question specifically asks for the architectural implication regarding computational complexity.

27
MCQmedium

A team is deploying a 13B-parameter decoder-only LLM on a single NVIDIA A100 40GB GPU for a real-time chatbot. During load, the process runs out of memory even though the model weights in FP16 require roughly 26GB. The team wants to reduce GPU memory usage with minimal impact on output quality and no change to the model architecture. Which technique is most appropriate?

A.Enable tensor parallelism across the single GPU's SMs
B.Switch the attention implementation to FlashAttention-2
C.Apply INT8 weight-only quantization using NVIDIA TensorRT-LLM
D.Increase the KV cache block size to 32 tokens
AnswerC

Weight-only INT8 quantization stores model weights at 8 bits while keeping activations in higher precision, cutting the 26GB FP16 footprint to roughly 13GB with typically small quality loss. TensorRT-LLM supports this natively and is designed for production inference on NVIDIA GPUs, so it directly addresses the memory shortfall without altering the network topology.

Why this answer

The bottleneck is static weight storage on a single 40GB device. Weight-only INT8 quantization halves the parameter footprint while preserving the architecture and offering strong quality retention, and TensorRT-LLM provides a supported path for this on NVIDIA hardware. The other choices target attention compute, KV cache paging, or multi-GPU partitioning, none of which reduce the resident weight memory on one GPU.

Exam trap

The trap here is assuming that attention-kernel optimizations such as FlashAttention reduce total GPU memory enough to fit a model whose weights alone exceed available VRAM.

28
MCQmedium

A research team is pretraining a decoder-only LLM and observes that gradient magnitudes in the earliest layers are extremely small while later layers train normally, causing slow convergence. They are using post-layer normalization. Which architectural change is most likely to improve gradient flow to the early layers?

A.Move layer normalization to before each sublayer (pre-LN) and add a final normalization before the output projection
B.Replace layer normalization with batch normalization across the sequence dimension
C.Remove residual connections to shorten the gradient path
D.Increase the number of attention heads while keeping head dimension constant
AnswerA

Pre-LN applies normalization on the sublayer input, creating a clean residual path that carries gradients directly from the loss to early layers. This is the standard remedy for vanishing gradients in deep Transformers and typically removes the need for learning-rate warmup. Adding a final normalization stabilizes the output scale before the vocabulary projection.

Why this answer

Post-LN places normalization inside the residual branch, so gradients must pass through normalization at every layer, attenuating them in deep stacks. Pre-LN normalizes the sublayer input and leaves the residual stream unnormalized, giving a direct gradient highway from the loss to early parameters. That structural change, plus a final normalization before the output head, is the recognized fix for the described symptom.

Exam trap

The trap here is treating vanishing early-layer gradients as a capacity or width problem and adjusting head counts, when the root cause is where normalization sits relative to the residual branch.

29
MCQmedium

A research team wants to train a large decoder-only LLM where each token's representation is computed independently of token order, then inject order information afterward. They are choosing between learned absolute positional embeddings and sinusoidal absolute positional embeddings. Which statement accurately characterizes the tradeoff they face?

A.Learned embeddings are strictly better because they allow the model to discover relative offsets between tokens.
B.Sinusoidal embeddings are parameter-free and can be computed for any position, while learned embeddings require a fixed maximum length and add parameters.
C.Learned embeddings can extrapolate to sequence lengths longer than those seen in training, whereas sinusoidal embeddings cannot.
D.Sinusoidal embeddings must be recomputed for every batch, which makes them slower than learned embeddings at inference.
AnswerB

Sinusoidal embeddings use fixed sine and cosine functions of position and dimension, so they add no trainable parameters and can be evaluated at arbitrary indices. Learned embeddings allocate a trainable vector per position, which consumes parameters proportional to the maximum length and cannot produce values for unseen indices. This is the classic tradeoff between the two absolute schemes.

Why this answer

Sinusoidal absolute positional embeddings are deterministic functions of position, adding no parameters and allowing evaluation at any index. Learned absolute embeddings allocate a trainable vector per position, consuming parameters and capping usable length at the trained maximum. The other options either reverse the extrapolation behavior or misattribute relative-offset abilities to learned absolute embeddings.

Exam trap

The trap here is assuming learned positional embeddings extrapolate better because they are trainable, when in fact their fixed table size is their key limitation.

30
Multi-Selecthard

Which THREE architectural features are essential for enabling efficient inference of massive LLMs on multi-GPU systems?

Select 3 answers
A.KV Caching to store previous sequence states.
B.Tensor Parallelism to split layers across GPUs.
C.Batch Size reduction to increase memory throughput.
D.Weight Quantization to reduce the memory footprint.
E.Increasing the learning rate during inference.
AnswersA, B, D

KV caching prevents redundant computations by storing previously calculated keys and values, which is critical for reducing inference latency in autoregressive models. Without this, the model would need to recompute the entire attention history for every new token generated, leading to prohibitive performance costs in real-time scenarios.

Why this answer

Efficient inference at scale requires techniques that address memory constraints, compute latency, and communication overhead. Model parallelism, KV caching, and weight quantization are fundamental pillars that allow large models to fit within limited memory pools while maintaining high throughput. These techniques collectively ensure that the model remains responsive and cost-effective when serving large-scale requests in a production environment.

Exam trap

Students often select training-specific optimizations like gradient accumulation or data parallelism instead of focusing on inference-specific architectural requirements like tensor parallelism and KV caching.

31
MCQeasy

What is the function of the 'Masked' component in a Decoder-only Transformer's self-attention during training?

A.To filter out stop words from the input text.
B.To prevent the model from attending to future tokens in the sequence.
C.To increase the randomness of the model's predictions.
D.To reduce the computation of the attention mechanism by 50%.
AnswerB

The causal mask is a triangular matrix applied to the attention scores that sets values for future tokens to negative infinity. This ensures that when the softmax is applied, the weights for these positions become zero, effectively hiding the future context from the model during training and ensuring autoregressive integrity.

Why this answer

The mask ensures that each token can only attend to itself and the tokens that precede it in the sequence. This autoregressive property is essential for training decoder models, as it prevents the model from 'cheating' by looking at future tokens. This ensures that the model learns to predict the next word based solely on the context that would be available during actual inference.

Exam trap

Candidates often think the mask is for ignoring padding tokens. While padding masks exist, the primary 'Masked' component in decoder-only training is strictly for preventing look-ahead at future tokens.

32
MCQeasy

A developer is building a decoder-only generative model and wants to prevent the model from attending to future tokens during training so that each position can only use information from itself and earlier positions. Which architectural mechanism should they implement in the self-attention layer?

A.A padding mask that ignores tokens added to equalize sequence lengths across a batch.
B.A learned positional embedding added to the token embeddings at the input layer.
C.A causal attention mask that sets attention scores for future positions to negative infinity before the softmax.
D.A residual connection that adds the attention output back to the input hidden state.
AnswerC

A causal mask adds negative infinity to the attention logits of positions after the current token, so after softmax those future positions receive zero probability. This enforces left-to-right autoregressive conditioning during training and matches the inference behavior where future tokens are unavailable.

Why this answer

Decoder-only generative models must not see future tokens during training, otherwise next-token prediction becomes trivial and the model will not generalize to autoregressive inference. A causal attention mask accomplishes this by making future positions invisible to the softmax, ensuring each position only conditions on itself and earlier tokens.

Exam trap

The trap here is confusing a padding mask, which hides filler tokens, with a causal mask, which hides future tokens.

Ready to test yourself?

Try a timed practice session using only Llm Architecture questions.