Courseiva

NCP-GENL · domain

LLM Architecture

This domain covers the transformer internals that determine how NVIDIA LLMs are built, positioned, and served: attention variants, positional encoding schemes such as RoPE and YaRN, KV-cache behavior, quantization, and parallelism. Questions are scenario-based, asking you to diagnose OOM or scaling failures and pick the architectural feature that fixes them.

32 questions6 easy16 medium10 hard

Focused practice

Practice LLM Architecture questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about LLM Architecture

Diagnose LLM deployment and training failures by reasoning about attention, positional encoding, KV-cache memory, and parallelism strategy. The single most important thing: separate weight memory from KV-cache and activation memory when predicting whether a model fits on given NVIDIA GPUs.

Positional encoding choices: RoPE, YaRN scaling, and learned absolute vs. relative schemes for context extension

KV-cache memory growth and how it drives OOM beyond FP16 weight size on A100/H100 GPUs

Tensor, pipeline, and sequence parallelism for sharding massive decoder-only LLMs across multi-GPU nodes

Grouped-query and multi-query attention plus quantization for efficient large-model inference

Watch out for

Common LLM Architecture exam traps

  • ▸Sizing GPU memory from parameter count alone, ignoring KV cache, activations, and framework overhead that cause load-time OOM.
  • ▸Confusing YaRN context-window extension with fine-tuning or quantization; it rescales rotary frequencies, not weights.
  • ▸Assuming tensor parallelism alone suffices; massive LLMs also need pipeline or sequence parallelism plus communication overlap.

Question index

All LLM Architecture questions (32)

Click any question to see the full explanation, or start a practice session above.

1

Why do many modern LLMs use SwiGLU as their activation function in the feed-forward network instead of the traditional ReLU?

Medium
2

A team is pre-training a 13B-parameter decoder-only LLM on a cluster of NVIDIA GPUs. They observe that gradient norms spike sharply during the first few hundred steps, destabilizing training. They want to keep the standard post-layer-normalization placement but stabilize early optimization. Which architectural technique should they apply?

Medium
3

What is the architectural role of Layer Normalization in a Transformer, and where is it typically placed to ensure stable training?

Medium
4

A developer is inspecting a decoder-only Transformer and notices that during training, the model attends to future tokens in the sequence, causing the loss to drop unrealistically fast but generation to be incoherent. Which architectural mechanism is missing or misconfigured?

Easy
5

An engineer must serve a 70B-parameter LLM for a workload with many concurrent users and long shared system prompts, and wants to maximize throughput without retraining. Which inference-time optimization most directly reduces redundant computation across requests sharing the same prompt prefix?

Hard
6

An engineer is reviewing the attention implementation of a decoder-only LLM used for chat. During inference with a KV cache, generated tokens must not attend to future positions. Which mechanism enforces this constraint inside scaled dot-product attention?

Easy
7

In Mixture-of-Experts (MoE) architectures, why does the use of a router mechanism significantly impact performance compared to dense models?

Hard
8

An engineer is analyzing why a decoder-only LLM with 32,000-token context length fails to answer questions that require information from the beginning of a long document when the answer is near the end. The model was trained with standard causal attention. Which two architectural or training factors are most likely contributing to this failure? (Choose two.)

Hard
9

A developer is building a retrieval-augmented generation pipeline and needs to choose a component that produces dense vector representations of passages for semantic search. The passages are up to 512 tokens long, and the developer wants a model specifically trained to map semantically similar text to nearby points in embedding space. Which type of model should be selected?

Easy
10

A team is training a large language model and wants to reduce the memory used by the optimizer without changing the model architecture. They are using Adam and notice that optimizer state consumes more GPU memory than the model weights. Which technique should they apply to reduce optimizer memory while keeping the model architecture unchanged?

Medium
11

Which TWO of the following are benefits of using Rotary Positional Embeddings (RoPE) compared to absolute positional embeddings?

Medium
12

A team is pretraining a 13B-parameter decoder-only LLM on English text using byte-pair encoding with a 50,000-token vocabulary. They observe that the model produces fluent but repetitive continuations and that the average log-probability assigned to ground-truth tokens plateaus early. The training loss curve shows the model is underfitting rather than overfitting. Which architectural change is most likely to improve the model's capacity to capture long-range dependencies?

Medium
13

An ML engineer is evaluating Mixture-of-Experts (MoE) routing for a large decoder-only model to increase capacity without proportionally increasing compute per token. Which TWO statements accurately describe how top-k token routing behaves in such an architecture? (Choose two.)

Hard
14

A team is fine-tuning a pretrained decoder-only model on a small domain-specific dataset. They observe that the model quickly overfits and loses general language ability. They want to update only a small number of additional parameters while keeping the base weights frozen. Which approach should they use?

Hard
15

A team is pretraining a decoder-only Transformer LLM on a large corpus of code and natural language. They observe that the model's training loss decreases smoothly, but during generation it sometimes produces degenerate repetition, and attention entropy on long sequences collapses. They suspect the issue is related to the positional encoding scheme. Which architectural change is most likely to mitigate the attention entropy collapse while preserving the model's ability to generalize to sequences longer than those seen during pretraining?

Medium
16

Which TWO of the following statements correctly describe the role of Grouped Query Attention (GQA) in modern LLM architectures?

Medium
17

Refer to the exhibit. An engineer is fine-tuning an LLM using the provided configuration. What is the primary purpose of applying 'yarn' scaling in this architecture?

Hard
18

Refer to the exhibit. The model is encountering an OOM error during long-context processing. Which architectural adjustment is most appropriate to resolve this while maintaining context length?

Medium
19

An engineer is reviewing the architecture of a decoder-only LLM that must support very long input contexts for document analysis. They are considering architectural choices that extend effective context length beyond what the model saw during pretraining. Which TWO techniques are designed specifically to extend usable context length without retraining the entire model from scratch? (Choose two.)

Hard
20

An inference engineer is serving a 70B-parameter decoder-only LLM and wants to reduce KV cache memory to fit longer contexts on each GPU. They consider Multi-Query Attention (MQA), Grouped-Query Attention (GQA), and standard Multi-Head Attention (MHA). Which statement correctly describes the memory and quality tradeoff among these attention variants?

Hard
21

Which architectural component is responsible for projecting the model's hidden states back into the vocabulary space to predict the next token?

Medium
22

A team is deploying a 70B-parameter decoder-only LLM on an NVIDIA H100 GPU node. During generation they observe that the KV cache grows linearly with sequence length and is consuming most of the available HBM, forcing them to limit batch size. They want to reduce KV cache memory without retraining the model from scratch. Which architectural change should they apply?

Medium
23

An engineer is evaluating a sparse Mixture-of-Experts decoder-only LLM for a latency-sensitive inference service. They notice that although the model has far more total parameters than a dense baseline, throughput per token is only modestly better and sometimes worse. Which factor best explains why sparse MoE does not translate total parameter count into proportional speedup during inference?

Hard
24

An engineer is deploying a 70B parameter model and needs to serve many concurrent users on a single GPU with limited memory. They want to store the attention keys and values for past tokens efficiently so that generation does not recompute them at every step. Which technique should they implement?

Medium
25

In the context of the Transformer architecture, what is the primary function of the Feed-Forward Network (FFN) layers applied after the attention mechanism?

Easy
26

An engineer is designing a Transformer-based model for long-context document summarization. They decide to replace the standard dense self-attention mechanism with a sliding window attention approach. What is the primary architectural implication of this change?

Medium
27

A team is deploying a 13B-parameter decoder-only LLM on a single NVIDIA A100 40GB GPU for a real-time chatbot. During load, the process runs out of memory even though the model weights in FP16 require roughly 26GB. The team wants to reduce GPU memory usage with minimal impact on output quality and no change to the model architecture. Which technique is most appropriate?

Medium
28

A research team is pretraining a decoder-only LLM and observes that gradient magnitudes in the earliest layers are extremely small while later layers train normally, causing slow convergence. They are using post-layer normalization. Which architectural change is most likely to improve gradient flow to the early layers?

Medium
29

A research team wants to train a large decoder-only LLM where each token's representation is computed independently of token order, then inject order information afterward. They are choosing between learned absolute positional embeddings and sinusoidal absolute positional embeddings. Which statement accurately characterizes the tradeoff they face?

Medium
30

Which THREE architectural features are essential for enabling efficient inference of massive LLMs on multi-GPU systems?

Hard
31

What is the function of the 'Masked' component in a Decoder-only Transformer's self-attention during training?

Easy
32

A developer is building a decoder-only generative model and wants to prevent the model from attending to future tokens during training so that each position can only use information from itself and earlier positions. Which architectural mechanism should they implement in the self-attention layer?

Easy

Frequently asked questions

What does the LLM Architecture domain cover on the NCP-GENL exam?
Diagnose LLM deployment and training failures by reasoning about attention, positional encoding, KV-cache memory, and parallelism strategy. The single most important thing: separate weight memory from KV-cache and activation memory when predicting whether a model fits on given NVIDIA GPUs.
How many questions are in this domain?
This page lists all 32 LLM Architecture questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only LLM Architecture questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
nvidia-ncp-genl NVIDIA-NCP-GENL llm architecture Practice Questions