NCA-GENL Core Machine Learning and AI Knowledge Practice Question
Which of the following activation functions is most commonly used in hidden layers of deep neural networks to mitigate the vanishing gradient problem?
⚠ Common exam trap
Candidates mistakenly select Sigmoid or Tanh, forgetting that these functions saturate at extreme values, which causes the vanishing gradient problem. They fail to recall that ReLU is specifically designed to prevent this.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ReLU
The Rectified Linear Unit (ReLU) is the standard activation function for deep networks. Unlike Sigmoid or Tanh, which saturate at high and low values, ReLU maintains a constant gradient of 1 for all positive inputs. This effectively prevents the gradient from vanishing during backpropagation across many layers, allowing for the training of much deeper architectures without needing complex initialization or normalization schemes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Sigmoid
Why it's wrong here
Sigmoid functions suffer from vanishing gradients because their derivative is very small for large positive or negative inputs. When chain-ruled across many layers, these small values cause the gradient to shrink toward zero, making it nearly impossible for early layers in a deep network to learn effectively during training.
- ✗
Tanh
Why it's wrong here
Like sigmoid, the Tanh function saturates at the extremes. Its gradient also becomes very small as the input magnitude grows, leading to the vanishing gradient problem in deep networks. While it is zero-centered and often performs better than sigmoid, it does not solve the fundamental vanishing gradient issue as effectively as ReLU.
- ✓
ReLU
Why this is correct
ReLU outputs the input directly if it is positive and zero otherwise. Its derivative is 1 for positive inputs, which allows gradients to flow through the network without being multiplied by small values. This property is key to training deep architectures efficiently, as it drastically reduces the vanishing gradient problem.
- ✗
Linear
Why it's wrong here
A linear activation function just scales the input. If a deep network uses only linear activations, the entire network acts as a single linear transformation regardless of how many layers it has. This makes it impossible to learn complex, non-linear relationships, rendering the depth of the network entirely useless.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.