NCA-GENL Core Machine Learning and AI Knowledge Practice Question
When training a model with a very large dataset, which approach provides the best balance between computational efficiency and model convergence?
⚠ Common exam trap
Candidates often pick full-batch gradient descent for stability or single-sample stochastic descent, ignoring how mini-batching optimizes GPU hardware utilization and convergence.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Mini-batch gradient descent.
Mini-batch gradient descent strikes an ideal balance by processing small subsets of the data (batches) at a time. This provides the efficiency of vectorization on GPU hardware while maintaining enough stochasticity to help the optimizer avoid poor local minima. Compared to full batch (too slow) or single-sample stochastic descent (too noisy), mini-batching is the industry standard for scaling training to large datasets while achieving optimal convergence.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Full-batch gradient descent.
Why it's wrong here
Full-batch gradient descent is computationally inefficient for large datasets. Calculating gradients over the entire dataset requires massive memory and is significantly slower per update than mini-batch approaches. It is rarely used in deep learning because the long wait times between weight updates make convergence prohibitively slow for large models.
- ✗
Stochastic gradient descent with batch size 1.
Why it's wrong here
Using a batch size of 1 introduces extreme noise into the gradient estimation. This makes the model training path highly erratic and prevents the effective use of hardware parallelism (e.g., matrix multiplications on GPUs). It is less efficient than mini-batching because it underutilizes hardware and converges slower.
- ✓
Mini-batch gradient descent.
Why this is correct
Mini-batch gradient descent provides the best of both worlds. It uses GPU parallelism to calculate gradients efficiently across a batch of samples while maintaining enough stochasticity to help the model escape poor local minima. It is the gold standard for scaling training on modern large-scale machine learning systems.
- ✗
Gradient descent without any data shuffling.
Why it's wrong here
Failing to shuffle the data can introduce significant biases into the training process, especially if the data has any inherent ordering. This can cause the model to learn patterns specific to the ordering rather than the data itself, which severely hinders the model's ability to generalize to unseen data.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.