Courseiva

NCA-GENL Core Machine Learning and AI Knowledge Practice Question

Which of the following activation functions is most commonly used in hidden layers of deep neural networks to mitigate the vanishing gradient problem?

⚠ Common exam trap

Candidates mistakenly select Sigmoid or Tanh, forgetting that these functions saturate at extreme values, which causes the vanishing gradient problem. They fail to recall that ReLU is specifically designed to prevent this.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

ReLU

The Rectified Linear Unit (ReLU) is the standard activation function for deep networks. Unlike Sigmoid or Tanh, which saturate at high and low values, ReLU maintains a constant gradient of 1 for all positive inputs. This effectively prevents the gradient from vanishing during backpropagation across many layers, allowing for the training of much deeper architectures without needing complex initialization or normalization schemes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Sigmoid

    Why it's wrong here

    Sigmoid functions suffer from vanishing gradients because their derivative is very small for large positive or negative inputs. When chain-ruled across many layers, these small values cause the gradient to shrink toward zero, making it nearly impossible for early layers in a deep network to learn effectively during training.

  • ✗

    Tanh

    Why it's wrong here

    Like sigmoid, the Tanh function saturates at the extremes. Its gradient also becomes very small as the input magnitude grows, leading to the vanishing gradient problem in deep networks. While it is zero-centered and often performs better than sigmoid, it does not solve the fundamental vanishing gradient issue as effectively as ReLU.

  • ✓

    ReLU

    Why this is correct

    ReLU outputs the input directly if it is positive and zero otherwise. Its derivative is 1 for positive inputs, which allows gradients to flow through the network without being multiplied by small values. This property is key to training deep architectures efficiently, as it drastically reduces the vanishing gradient problem.

  • ✗

    Linear

    Why it's wrong here

    A linear activation function just scales the input. If a deep network uses only linear activations, the entire network acts as a single linear transformation regardless of how many layers it has. This makes it impossible to learn complex, non-linear relationships, rendering the depth of the network entirely useless.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.