NCP-GENL Fine-Tuning Practice Question
A company wants to fine-tune a 70B LLM to follow domain-specific instructions. The base model already performs well on general language tasks. They have limited labeled data and limited GPU memory. Which approach is most appropriate?
⚠ Common exam trap
The trap here is assuming that the largest model change yields the best domain adaptation, when memory and data limits make parameter-efficient adaptation the correct engineering choice.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Parameter-Efficient Fine-Tuning with LoRA adapters targeting attention projection layers.
LoRA is the standard parameter-efficient method for adapting very large models under memory constraints, since it trains only small adapter matrices while freezing base weights. With limited labeled data, the reduced trainable parameter count also lowers overfitting risk and helps retain general capabilities. Full fine-tuning, tokenizer retraining, and prompt-only approaches either exceed resource limits or fail to deliver durable domain adaptation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Retrain the tokenizer on the domain corpus and continue pretraining from scratch.
Why it's wrong here
Retraining the tokenizer and pretraining from scratch discards the base model's learned knowledge and requires massive compute and data, which the scenario explicitly lacks. It also does not leverage the base model's strong general language performance. This is far more expensive than necessary and does not fit the limited-resource, instruction-following objective.
- ✓
Parameter-Efficient Fine-Tuning with LoRA adapters targeting attention projection layers.
Why this is correct
LoRA freezes the base weights and trains small low-rank adapter matrices, dramatically reducing memory and optimizer state requirements. It is well suited to limited labeled data because fewer trainable parameters reduce overfitting risk, and it preserves the base model's general abilities. This matches the scenario's constraints on memory and data volume while still adapting the model to domain instructions.
- ✗
Full-parameter fine-tuning of all 70B weights on the domain dataset.
Why it's wrong here
Full-parameter fine-tuning updates every weight and requires optimizer states and gradients for all 70B parameters, which demands enormous memory and compute. With limited labeled data, full fine-tuning also risks catastrophic forgetting of general capabilities. This approach contradicts both the memory constraint and the small-data condition described in the scenario.
- ✗
Prompt engineering with few-shot examples and no weight updates.
Why it's wrong here
Prompt engineering requires no training and can improve outputs, but it does not durably adapt the model to specialized domain behavior. The scenario states the goal is fine-tuning to follow domain-specific instructions, implying persistent weight-level adaptation. Few-shot prompting consumes context length and may not reliably encode domain knowledge, so it does not meet the stated requirement.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.