Courseiva
LLM Architecture →hardMultiple Select

NCP-GENL LLM Architecture Practice Question

An engineer is reviewing the architecture of a decoder-only LLM that must support very long input contexts for document analysis. They are considering architectural choices that extend effective context length beyond what the model saw during pretraining. Which TWO techniques are designed specifically to extend usable context length without retraining the entire model from scratch? (Choose two.)

⚠ Common exam trap

The trap here is treating any memory-saving or capacity change as context extension, when only techniques that alter positional encoding behavior actually lengthen usable context.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Position interpolation, which rescales positional indices so longer sequences fall within the range seen during pretraining.

Position interpolation and YaRN both operate on positional encoding to reconcile desired sequence lengths with the range seen during pretraining. Interpolation compresses indices, while YaRN adjusts rotary frequencies to improve extrapolation. Both are context-extension methods that require only light fine-tuning, unlike changes to head count, MLP activation, or vocabulary size, which do not address positional generalization.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switching the feed-forward activation from SwiGLU to GeLU to improve long-range dependency modeling.

    Why it's wrong here

    The feed-forward activation function governs nonlinearity in the MLP blocks and has no direct relationship to how positional information is encoded or extrapolated. Changing from SwiGLU to GeLU would alter training dynamics and quality but would not extend the model's usable context length, which depends on positional encoding and attention behavior.

  • ✓

    Position interpolation, which rescales positional indices so longer sequences fall within the range seen during pretraining.

    Why this is correct

    Position interpolation compresses the positional index range so that a longer sequence maps into the positions the model encountered during pretraining. This allows the model to generalize to longer contexts with only brief fine-tuning. It directly targets the mismatch between pretraining context length and desired inference length, making it a standard context-extension technique.

  • ✗

    Increasing the number of attention heads while keeping the head dimension fixed.

    Why it's wrong here

    Adding attention heads changes model capacity and the size of the KV cache but does nothing to reconcile positional indices with lengths beyond the pretraining range. The model would still see positions it never learned to handle. Context extension is about positional encoding behavior, not about the count of attention heads.

  • ✓

    YaRN scaling, which adjusts rotary positional embedding frequencies to improve extrapolation to longer sequences.

    Why this is correct

    YaRN modifies the frequency schedule of rotary positional embeddings, blending interpolation and extrapolation so attention remains well-behaved beyond the pretraining length. It is specifically designed for context extension and typically requires only modest fine-tuning. This makes it a targeted method for lengthening usable context without a full pretraining run.

  • ✗

    Reducing the vocabulary size to shorten the embedding matrix and free memory for longer sequences.

    Why it's wrong here

    Vocabulary size affects the embedding and output projection matrices, not the positional encoding scheme that determines whether the model can handle longer contexts. Shrinking the vocabulary might free some memory, but it would not teach the model to interpret positions beyond its pretraining range, so it does not extend usable context length.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.