Courseiva
Experimentation →easyMultiple Choice

NCA-GENL Experimentation Practice Question

An engineer must decide how to split a labeled dataset of 50,000 customer support conversations before fine-tuning a NeMo LLM for intent classification. The goal is an honest estimate of how the tuned model will behave on never-before-seen tickets once deployed. Which splitting approach best supports that goal?

⚠ Common exam trap

The trap here is treating the test set as another tuning signal, which silently converts it into a validation set and inflates the reported accuracy.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Randomly assign 70 percent to training, 15 percent to validation, and 15 percent to a test set that is touched only once at the end.

An untouched holdout test set is the only split that yields an unbiased generalization estimate, because it never influences training or hyperparameter choices. Pairing it with a validation split lets the team tune decisions without contaminating the final measurement, and random assignment keeps class balance comparable across splits.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Randomly assign 70 percent to training, 15 percent to validation, and 15 percent to a test set that is touched only once at the end.

    Why this is correct

    Holding out a test split that is never used for tuning gives an unbiased estimate of generalization, while the validation split supports hyperparameter decisions. Random assignment keeps class proportions approximately stable in each split, so the final test measurement reflects expected production behavior on unseen tickets.

  • ✗

    Use all 50,000 conversations for training and rely on the training loss curve to judge generalization.

    Why it's wrong here

    Training loss measures fit to data the model has already seen, so it will keep improving even when the model memorizes. With no held-out data there is no way to detect overfitting, and the resulting estimate of production accuracy on unseen tickets would be optimistic and unreliable.

  • ✗

    Tune hyperparameters repeatedly against the test split until accuracy is maximized.

    Why it's wrong here

    Repeatedly consulting the test split leaks information into model selection, so the reported accuracy no longer estimates performance on truly unseen data. The number becomes an artifact of tuning against that specific sample rather than an honest predictor of deployed behavior.

  • ✗

    Split the data by conversation length, putting the longest 15 percent in the test set.

    Why it's wrong here

    A length-based split creates distribution shift between splits: the test set contains only long conversations while training contains mostly short ones. Any accuracy drop would confound genuine generalization error with the length difference, so the measurement would not represent performance on typical production tickets.

About these practice questions

This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.