NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A data scientist is preparing a transformer-based language model for a text summarization task. She notices that the input sequences in her dataset vary widely in length, from a few tokens to several thousand. She decides to set a fixed maximum sequence length and pad shorter sequences with a special token. Which component of the transformer architecture is primarily responsible for handling the positional information of tokens in these sequences?
⚠ Common exam trap
Candidates often confuse self-attention's ability to model relationships with its inability to encode order, leading to the misconception that self-attention alone handles positional information.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Positional encodings
Positional encodings are essential in transformer models because the self-attention mechanism is permutation-invariant. By adding positional encodings to token embeddings, the model gains awareness of token order, enabling it to handle sequences of varying lengths and padding effectively. This is critical for tasks like summarization where the meaning depends on word order.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Self-attention mechanism
Why it's wrong here
Self-attention computes relationships between tokens but does not inherently encode their order. Without positional information, it treats the sequence as a bag of tokens. In this scenario, padding and varying lengths would cause the model to lose the sequential order necessary for summarization. Therefore, self-attention alone does not handle positional information.
- ✗
Layer normalization
Why it's wrong here
Layer normalization stabilizes training by normalizing activations across features, but it does not provide any positional information. It operates independently of token order and sequence length. In this scenario, layer normalization would not help the model differentiate between tokens based on their positions, so it is not the correct component.
- ✗
Feed-forward network
Why it's wrong here
The feed-forward network applies a pointwise transformation to each token independently, without considering order or context. It does not encode positional information. In the summarization task with variable-length sequences, the feed-forward network would process each token in isolation, missing the sequential structure needed for coherent summaries.
- ✓
Positional encodings
Why this is correct
Positional encodings are added to token embeddings to inject information about the order of tokens in the sequence. This allows the transformer to distinguish between tokens at different positions, which is crucial for tasks like summarization where word order affects meaning. With varying sequence lengths and padding, positional encodings ensure the model understands the relative positions of tokens.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.