Which TWO of the following statements regarding Gradient Descent variants are correct?
Trap 1: Stochastic Gradient Descent always converges faster than Adam in…
Adam is generally faster to converge because it utilizes adaptive learning rates for each parameter, effectively handling sparse gradients and non-stationary objectives. SGD often requires extensive manual tuning of the learning rate schedule and may struggle with complex loss surfaces compared to adaptive methods like Adam.
Trap 2: Momentum-based optimization has no effect on the speed of…
Momentum is designed to accelerate gradients in directions of consistent descent while dampening oscillations in high-curvature directions. It acts like a ball rolling down a hill, accumulating velocity, which is crucial for overcoming plateaus and noise during the optimization process in deep networks.
Trap 3: Mini-batch SGD is inherently less efficient than using the full…
Mini-batching provides the best trade-off between the stability of Batch Gradient Descent and the speed of pure SGD. It exploits parallel computation on GPUs, allowing for faster iterations and better memory utilization, which is the standard practice for training modern machine learning models.
- A
Batch Gradient Descent calculates the gradient using the entire dataset, ensuring a stable path to the global minimum.
Batch Gradient Descent computes the loss for every training example before updating weights. While this provides a stable, deterministic gradient direction, it is computationally expensive and memory-intensive for large datasets, often making it impractical for modern large-scale LLM training workflows on GPU clusters.
- B
Stochastic Gradient Descent always converges faster than Adam in all neural network architectures.
Why it fails: Adam is generally faster to converge because it utilizes adaptive learning rates for each parameter, effectively handling sparse gradients and non-stationary objectives. SGD often requires extensive manual tuning of the learning rate schedule and may struggle with complex loss surfaces compared to adaptive methods like Adam.
- C
Adam optimizer maintains per-parameter learning rates based on the first and second moments of the gradients.
Adam combines the benefits of momentum (first moment) and RMSProp (second moment). By maintaining an exponentially decaying average of past gradients and their squares, it adapts the update step for every parameter, which is essential for training deep neural networks with varying gradient magnitudes.
- D
Momentum-based optimization has no effect on the speed of convergence in deep learning.
Why it fails: Momentum is designed to accelerate gradients in directions of consistent descent while dampening oscillations in high-curvature directions. It acts like a ball rolling down a hill, accumulating velocity, which is crucial for overcoming plateaus and noise during the optimization process in deep networks.
- E
Mini-batch SGD is inherently less efficient than using the full dataset for every update.
Why it fails: Mini-batching provides the best trade-off between the stability of Batch Gradient Descent and the speed of pure SGD. It exploits parallel computation on GPUs, allowing for faster iterations and better memory utilization, which is the standard practice for training modern machine learning models.