Courseiva

AI0-001 Machine Learning and Deep Learning Practice Question

A machine learning engineer is troubleshooting a recurrent neural network that fails to learn long-range dependencies in sequential data. The gradients are computed using backpropagation through time. Which phenomenon is most likely occurring, and what architectural change would best address it?

⚠ Common exam trap

Many exam-takers confuse 'failure to learn long-range dependencies' with exploding gradients, but the correct clue is the inability to capture distant patterns, not training instability or NaN losses.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Vanishing gradients; use LSTM or GRU units

In standard RNNs, backpropagation through time (BPTT) multiplies gradients across many time steps, causing them to shrink exponentially (vanishing gradients). This prevents the network from learning long-range dependencies. LSTM or GRU units introduce gating mechanisms that preserve gradient flow over many time steps, directly solving this problem.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Underfitting; increase the number of time steps

    Why it's wrong here

    Underfitting describes poor fit on training data generally, not the specific decay of gradients across long sequences. Adding time steps lengthens the unrolled graph, worsening vanishing gradients rather than restoring long-range learning. Underfitting is addressed by increasing model capacity or training longer.

  • ✓

    Vanishing gradients; use LSTM or GRU units

    Why this is correct

    Backpropagation through time multiplies many small Jacobian terms, so gradients shrink exponentially across long sequences, preventing the network from learning distant dependencies. LSTM and GRU units introduce gated cell states that carry information along a near-constant error path, preserving gradient magnitude over long ranges.

  • ✗

    Exploding gradients; apply gradient clipping

    Why it's wrong here

    Exploding gradients produce large, oscillating updates, but the stem describes failure to retain information across many steps, which is vanishing gradients. Clipping caps update magnitude and is the remedy for instability during training, not for restoring long-range signal flow; gated architectures such as LSTM address that.

  • ✗

    Overfitting; reduce the number of layers

    Why it's wrong here

    Overfitting concerns generalisation to unseen data, not the vanishing gradients that prevent learning long-range dependencies. It is tempting because reducing layers is a common remedy for overfitting, and would be correct when training accuracy far exceeds validation accuracy.

About these practice questions

Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.