NCP-GENL Evaluation Practice Question
After fine-tuning a code-generation model with NVIDIA NeMo, an engineer notices the model now produces correct domain-specific function calls but has started emitting malformed JSON in about 15 percent of structured-output requests. The fine-tuning dataset contained no structured-output examples. Which evaluation action best explains and catches this regression?
⚠ Common exam trap
The trap here is assuming a successful domain fine-tune cannot break unrelated behaviors, so the JSON failures get explained away instead of being measured against the base model.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Add a format-compliance check plus a general-capability regression suite to the evaluation, and compare against the base model on the same structured-output prompts.
Fine-tuning on a narrow dataset can degrade capabilities absent from that data, and structured output is a classic casualty. The right response is to measure it: add format-compliance scoring to the harness, rerun the same structured prompts against the base and tuned models, and include a general regression suite. Retraining harder, raising temperature, or dismissing the failures all skip the measurement step the situation demands.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Add a format-compliance check plus a general-capability regression suite to the evaluation, and compare against the base model on the same structured-output prompts.
Why this is correct
The symptom points to a capability regression outside the fine-tuning domain, so the evaluation must cover format validity and previously working skills, not just the target task. Running the same structured-output prompts against the base model establishes whether the fine-tune caused the breakage. A format checker converts the vague observation into a measurable pass rate.
- ✗
Increase the temperature of the structured-output requests so the model has more freedom to produce valid JSON.
Why it's wrong here
Higher temperature increases randomness and typically makes format violations more frequent, not less. Structured output benefits from deterministic, low-temperature decoding, and the underlying issue is a shifted capability distribution after fine-tuning. Adjusting sampling does not diagnose the regression and may mask it.
- ✗
Conclude that the fine-tune succeeded on its target task and that the JSON failures are acceptable collateral given the domain gains.
Why it's wrong here
Declaring the regression acceptable without measuring its scope or knowing which downstream systems depend on valid JSON is an unquantified risk. The failures affect a capability the base model handled, so they represent lost functionality. An evaluation should surface the trade-off with numbers so stakeholders can decide deliberately rather than by assertion.
- ✗
Retrain the model with a larger learning rate so the structured-output behavior is reinforced more strongly during fine-tuning.
Why it's wrong here
A larger learning rate would likely worsen the regression by pushing the weights further from the base distribution, and the dataset contains no structured-output examples to reinforce. The problem is a measurement gap, not insufficient training strength. The team needs to detect and quantify the regression before deciding how to remediate it.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.