AI0-001 Machine Learning and Deep Learning Practice Question
A team trains a recurrent neural network to translate sentences averaging 60 words. During evaluation they notice that translations of the final words in long sentences are frequently wrong, while the opening words are translated accurately. Which architectural change best addresses this behavior?
⚠ Common exam trap
The trap here is assuming more layers, larger batches, or longer training can overcome a positional memory problem, when the limitation is architectural rather than an optimization or capacity issue.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Replace the recurrent layers with an attention-based transformer encoder-decoder
Recurrent networks compress the entire source sentence into a fixed-size hidden state and suffer vanishing gradients across many timesteps, which disproportionately harms recall of distant tokens. An attention-based transformer removes the sequential bottleneck, allowing direct token-to-token connections and preserving context for the final words of long sentences.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Lower the learning rate and train for more epochs
Why it's wrong here
A smaller learning rate may stabilize optimization but does not change the architecture's information bottleneck. The recurrent state remains a fixed-size summary that decays over long sequences, so the model continues to lose early-context fidelity for the final words no matter how long it trains.
- ✓
Replace the recurrent layers with an attention-based transformer encoder-decoder
Why this is correct
The described degradation at the end of long sequences is characteristic of vanishing gradients and limited memory in plain recurrent networks. Self-attention lets every output position attend directly to every input position in constant path length, so information from early and late tokens remains accessible. Transformers are the standard architecture for long-sequence translation and directly resolve this failure.
- ✗
Increase the batch size used during training
Why it's wrong here
Batch size affects gradient estimate variance and hardware utilization, not the model's ability to retain information across many timesteps. The recurrent hidden state still must compress the entire sentence, so the positional degradation pattern persists regardless of how many sentences are processed per update.
- ✗
Increase the number of recurrent layers stacked in the encoder
Why it's wrong here
Adding recurrent layers deepens the unrolled computation graph and can amplify vanishing gradients through the extended time dimension. It increases capacity and training cost without providing direct long-range connections, so accuracy on the final words of long sentences is unlikely to improve and may degrade further.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.