Courseiva

AI0-001 AI Concepts and Techniques Practice Question

Which neural network architecture is specifically designed to handle sequential data and mitigate the vanishing gradient problem?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Long Short-Term Memory (LSTM) network

LSTM (Long Short-Term Memory) is a type of RNN designed with gating mechanisms to prevent vanishing gradients in long sequences. CNNs are for spatial data; vanilla RNNs suffer from vanishing gradients; transformers use attention but are not specifically designed to mitigate vanishing gradients (they use residual connections).

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convolutional Neural Network (CNN)

    Why it's wrong here

    Convolutional neural networks apply local filters with weight sharing over grid-structured inputs such as images, and their fixed receptive fields do not model the temporal dependencies or gating that suppress vanishing gradients in sequential data; recurrent architectures with LSTM or GRU cells, or transformers, address that requirement. CNNs are tempting because they are the default choice for image classification and can be applied to one-dimensional sequences such as audio or text, but they lack the recurrent state needed for long-range order.

  • ✗

    Transformer

    Why it's wrong here

    Transformers process sequences in parallel through self-attention, which removes recurrence but does not itself target the vanishing gradient of long backpropagation chains. The option tempts because Transformers now dominate sequential tasks, yet the architecture explicitly designed to mitigate that gradient problem is the LSTM.

  • ✗

    Vanilla Recurrent Neural Network (RNN)

    Why it's wrong here

    A vanilla RNN suffers the vanishing gradient problem rather than mitigating it, because gradients shrink as they propagate back through many timesteps. The option tempts because RNNs are indeed built for sequential data, but the gated LSTM architecture was created specifically to address that gradient decay.

  • ✓

    Long Short-Term Memory (LSTM) network

    Why this is correct

    LSTM networks use gated cells (input, forget, output) that regulate information flow across time steps, preserving gradients over long sequences. This architecture directly mitigates the vanishing gradient problem while handling sequential data, matching the stem's requirement.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.