NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A team is preparing a dataset to train a generative AI model for text summarization. They want to ensure the model generalizes well and does not simply memorize the training examples. Which TWO practices should they follow? (Choose two.)
⚠ Common exam trap
The trap here is assuming that more data for training or a larger model always improves performance, but without proper validation and augmentation, the model may overfit.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Split the dataset into training, validation, and test sets
To promote generalization and prevent memorization, the team should use a held-out validation set for tuning and a test set for final evaluation, and they should augment training data to increase diversity. Using all data for training, enlarging the model without regularization, or training until training loss is near zero all increase the risk of overfitting.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the model's parameter count to capture more complex patterns
Why it's wrong here
Increasing model size can improve capacity but also raises the risk of overfitting, especially if the dataset is limited. Without proper regularization or more data, a larger model may simply memorize training examples rather than learn generalizable patterns. It does not directly promote generalization.
- ✓
Split the dataset into training, validation, and test sets
Why this is correct
Dividing data into training, validation, and test sets allows the team to tune hyperparameters on the validation set and evaluate final performance on unseen test data. This separation is essential to detect overfitting and estimate how well the model will generalize to new, real-world examples.
- ✗
Use the entire dataset for training to maximize data availability
Why it's wrong here
Using all data for training leaves no independent data to assess generalization. The model could memorize the training set and appear to perform well, but it would likely fail on new inputs. Without a held-out set, there is no reliable way to detect overfitting or tune hyperparameters.
- ✓
Apply data augmentation to increase the diversity of training examples
Why this is correct
Data augmentation creates modified versions of existing examples, such as synonym replacement or sentence shuffling, which exposes the model to more varied inputs. This reduces the chance of memorizing exact phrases and encourages learning of invariant features, thereby improving generalization to unseen text.
- ✗
Train for as many epochs as possible until training loss is near zero
Why it's wrong here
Training until training loss is near zero often leads to overfitting, as the model starts fitting noise and idiosyncrasies of the training data. The goal is to minimize validation loss, not training loss. Early stopping based on validation performance is a better practice to ensure generalization.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.