Courseiva

NCP-GENL · topic practice

LLM Architecture practice questions

This domain covers the transformer internals that determine how NVIDIA LLMs are built, positioned, and served: attention variants, positional encoding schemes such as RoPE and YaRN, KV-cache behavior, quantization, and parallelism. Questions are scenario-based, asking you to diagnose OOM or scaling failures and pick the architectural feature that fixes them.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: LLM Architecture

What the exam tests

What to know about LLM Architecture

Diagnose LLM deployment and training failures by reasoning about attention, positional encoding, KV-cache memory, and parallelism strategy. The single most important thing: separate weight memory from KV-cache and activation memory when predicting whether a model fits on given NVIDIA GPUs.

Positional encoding choices: RoPE, YaRN scaling, and learned absolute vs. relative schemes for context extension

KV-cache memory growth and how it drives OOM beyond FP16 weight size on A100/H100 GPUs

Tensor, pipeline, and sequence parallelism for sharding massive decoder-only LLMs across multi-GPU nodes

Grouped-query and multi-query attention plus quantization for efficient large-model inference

Watch out for

Common LLM Architecture exam traps

  • ▸Sizing GPU memory from parameter count alone, ignoring KV cache, activations, and framework overhead that cause load-time OOM.
  • ▸Confusing YaRN context-window extension with fine-tuning or quantization; it rescales rotary frequencies, not weights.
  • ▸Assuming tensor parallelism alone suffices; massive LLMs also need pipeline or sequence parallelism plus communication overlap.

Practice set

LLM Architecture questions

20 questions · select your answer, then reveal the explanation

An engineer is analyzing a 70B-parameter decoder-only LLM that uses Grouped Query Attention (GQA) with 8 key/value heads and 64 query heads. During inference with a long prompt, they notice that the KV cache memory is still a bottleneck on a single GPU, and they want to reduce it further without retraining the model from scratch. Which approach best describes a valid architectural modification that reduces KV cache size while preserving the model's learned attention behavior as much as possible?

An engineer is deploying a 70B-parameter LLM for real-time chat. During load testing, the first token latency is acceptable, but subsequent tokens are generated at only 8 tokens per second, far below the target of 30. Profiling shows that the GPU memory bandwidth is saturated during decoding, and the KV cache is consuming over 60 GB. Which optimization is most likely to increase decoding throughput without retraining the model?

An engineer inspects a decoder-only Transformer and notices that each attention sublayer is followed by a residual addition and then a normalization, and the same pattern repeats in the MLP sublayer. They ask why the residual connection is added before the normalization rather than after. What is the primary purpose of the residual connection in this Pre-LN arrangement?

Question 4mediummultiple choice
Read the full LLM Architecture explanation →

A developer is building a retrieval-augmented generation pipeline and needs embeddings whose dot products reflect semantic similarity rather than magnitude. They plan to use a decoder-only LLM's hidden states as embeddings. Which architectural detail should they account for to make the similarity scores meaningful?

Question 5mediummultiple choice
Read the full LLM Architecture explanation →

A research team is training a decoder-only LLM and notices that gradients in the early layers are vanishing, causing slow convergence. They review their architecture and see that the residual connections are placed after the layer normalization in each sublayer, following a pre-norm design. They want to improve gradient flow without changing the model size. Which adjustment is most appropriate?

Question 6mediummultiple choice
Read the full LLM Architecture explanation →

An engineer is designing a Transformer-based model for long-context document summarization. They decide to replace the standard dense self-attention mechanism with a sliding window attention approach. What is the primary architectural implication of this change?

Refer to the exhibit. An engineer is fine-tuning an LLM using the provided configuration. What is the primary purpose of applying 'yarn' scaling in this architecture?

Exhibit

config: { 'model_type': 'decoder-only', 'rope_base': 10000, 'rope_scaling': { 'type': 'yarn', 'factor': 4.0 } }

Which TWO of the following statements correctly describe the role of Grouped Query Attention (GQA) in modern LLM architectures?

In the context of the Transformer architecture, what is the primary function of the Feed-Forward Network (FFN) layers applied after the attention mechanism?

Question 10mediummultiple choice
Read the full LLM Architecture explanation →

Why do many modern LLMs use SwiGLU as their activation function in the feed-forward network instead of the traditional ReLU?

Which THREE architectural features are essential for enabling efficient inference of massive LLMs on multi-GPU systems?

Question 12mediummultiple choice
Read the full LLM Architecture explanation →

Refer to the exhibit. The model is encountering an OOM error during long-context processing. Which architectural adjustment is most appropriate to resolve this while maintaining context length?

Exhibit

Error Log: [CUDA_ERROR_OUT_OF_MEMORY] during attention calculation. Sequence length: 128k. Model: 70B parameter, FP16.
Question 13hardmultiple choice
Review the full routing breakdown →

In Mixture-of-Experts (MoE) architectures, why does the use of a router mechanism significantly impact performance compared to dense models?

Question 14mediummultiple choice
Read the full LLM Architecture explanation →

What is the architectural role of Layer Normalization in a Transformer, and where is it typically placed to ensure stable training?

Which TWO of the following are benefits of using Rotary Positional Embeddings (RoPE) compared to absolute positional embeddings?

What is the function of the 'Masked' component in a Decoder-only Transformer's self-attention during training?

Question 17mediummultiple choice
Read the full LLM Architecture explanation →

Which architectural component is responsible for projecting the model's hidden states back into the vocabulary space to predict the next token?

Question 18mediummultiple choice
Read the full LLM Architecture explanation →

A team is pretraining a decoder-only Transformer LLM on a large corpus of code and natural language. They observe that the model's training loss decreases smoothly, but during generation it sometimes produces degenerate repetition, and attention entropy on long sequences collapses. They suspect the issue is related to the positional encoding scheme. Which architectural change is most likely to mitigate the attention entropy collapse while preserving the model's ability to generalize to sequences longer than those seen during pretraining?

A developer is building a decoder-only generative model and wants to prevent the model from attending to future tokens during training so that each position can only use information from itself and earlier positions. Which architectural mechanism should they implement in the self-attention layer?

Question 20mediummultiple choice
Read the full LLM Architecture explanation →

A team is training a large language model and wants to reduce the memory used by the optimizer without changing the model architecture. They are using Adam and notice that optimizer state consumes more GPU memory than the model weights. Which technique should they apply to reduce optimizer memory while keeping the model architecture unchanged?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused LLM Architecture sessions

Start a LLM Architecture only practice session

Every question in these sessions is drawn from the LLM Architecture domain — nothing else.

Related practice questions

Related NCP-GENL topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the NCP-GENL exam test about LLM Architecture?
Diagnose LLM deployment and training failures by reasoning about attention, positional encoding, KV-cache memory, and parallelism strategy. The single most important thing: separate weight memory from KV-cache and activation memory when predicting whether a model fits on given NVIDIA GPUs.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just LLM Architecture questions in a focused session?
Yes — the session launcher on this page draws every question from the LLM Architecture domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other NCP-GENL topics?
Use the topic links above to move to related areas, or go back to the NCP-GENL question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the NCP-GENL exam covers. They are not copied from any real exam or dump site.