AI0-001 Machine Learning and Deep Learning Practice Question
A deep learning model for sentiment analysis uses a softmax output layer. The hidden layers currently use tanh activation. Which activation function should replace tanh to mitigate vanishing gradients in deeper networks?
⚠ Common exam trap
Candidates often mistakenly believe that any non-linear activation works equally well in deep networks, but the trap is that they may choose sigmoid because it is non-linear, ignoring its saturation-induced vanishing gradient problem in deeper architectures.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ReLU
ReLU (Rectified Linear Unit) is correct because it outputs zero for negative inputs and a positive linear slope for positive inputs, which avoids the saturation problem of tanh. In deeper networks, tanh gradients can vanish as activations approach ±1, slowing or halting learning. ReLU's non-saturating nature keeps gradients flowing for positive inputs, mitigating the vanishing gradient problem.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Sigmoid
Why it's wrong here
Sigmoid saturates at both ends, so its derivative peaks at 0.25 and shrinks toward zero, compounding the vanishing-gradient problem tanh already causes across deep stacks. It suits binary output layers, not hidden layers. ReLU-family activations keep a constant gradient for positive inputs, which is what mitigates the decay.
- ✗
Softmax
Why it's wrong here
Softmax normalises outputs into a probability distribution across classes, so applying it in hidden layers destroys the independent feature representations those layers must build. It belongs only at the output layer for multi-class classification. ReLU-family activations keep gradients stable through depth, which is the requirement here.
- ✓
ReLU
Why this is correct
ReLU outputs the input directly for positive values, giving a derivative of one and avoiding the saturation that drives tanh's gradient toward zero in deep stacks. This preserves gradient magnitude during backpropagation, directly mitigating the vanishing gradient problem in the deeper network.
- ✗
Linear
Why it's wrong here
A linear activation collapses the network to an effective single-layer linear transform, so depth adds no representational power and no non-linearity remains for sentiment boundaries. It is used for regression outputs, not hidden layers. ReLU-family functions preserve non-linearity while avoiding the saturation that causes vanishing gradients.
About these practice questions
One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.