Courseiva
LLM Architecture →mediumMultiple Choice

NCP-GENL LLM Architecture Practice Question

An engineer is designing a Transformer-based model for long-context document summarization. They decide to replace the standard dense self-attention mechanism with a sliding window attention approach. What is the primary architectural implication of this change?

⚠ Common exam trap

Candidates often focus on the 'loss of accuracy' or 'semantic degradation.' While these are concerns, the question specifically asks for the architectural implication regarding computational complexity.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The computational complexity of the self-attention mechanism is reduced from quadratic to linear.

Sliding window attention restricts the receptive field of each token to a local neighborhood, drastically reducing the quadratic memory complexity of standard attention to linear. This is critical for scaling LLMs to long contexts, as it prevents the O(n²) memory growth that typically causes GPU out-of-memory errors on large input sequences while maintaining local coherence.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The model's total parameter count increases due to added positional embedding layers.

    Why it's wrong here

    Sliding window attention does not require additional parameter layers; rather, it modifies the attention computation mask. Parameters are generally independent of the attention pattern, focusing instead on feature projections in the query, key, and value heads, meaning parameter count remains largely unaffected by the windowing mechanism.

  • ✗

    The model loses the ability to perform parallel training across multiple GPU nodes.

    Why it's wrong here

    Parallel training remains fully functional with sliding window attention. Since the attention mask is locally constrained, it can still be computed efficiently across distributed systems. The transformation of the attention matrix does not introduce serial dependencies that would prevent standard data or tensor parallelism during the training process.

  • ✓

    The computational complexity of the self-attention mechanism is reduced from quadratic to linear.

    Why this is correct

    By limiting the attention span to a fixed window size, the number of operations per token becomes constant rather than proportional to the sequence length. This shift from O(n²) to O(n*w) complexity is the fundamental architectural advantage for long-context tasks, enabling processing of documents that would otherwise be computationally prohibitive.

  • ✗

    The model is no longer compatible with standard softmax normalization functions.

    Why it's wrong here

    Softmax remains the standard normalization function for calculating attention weights, regardless of the window size. The windowing approach simply applies a mask to the input of the softmax function to zero out distant tokens. The mathematical properties of softmax and its differentiability are preserved throughout the computation.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.