Courseiva

NCP-GENL · domain

Fine-Tuning

This domain covers parameter-efficient adaptation of large language models: LoRA, QLoRA, and memory-saving techniques for constrained hardware. Questions test your understanding of how low-rank updates, quantized base weights, and gradient checkpointing trade compute, memory, and quality during fine-tuning, and how configuration choices like rank affect trainable parameters and model capacity.

45 questions7 easy21 medium17 hard

Focused practice

Practice Fine-Tuning questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Fine-Tuning

You must be able to explain how LoRA rank controls adapter capacity, which layers LoRA modifies, why QLoRA uses NF4 quantization, and how gradient checkpointing reduces memory. The key is knowing these techniques trade memory or compute for fine-tuning feasibility without changing the base model's frozen weights.

The role of the LoRA rank parameter in controlling low-rank adapter capacity and trainable parameter count

Which architectural components LoRA modifies, namely attention projection weights, while freezing base model weights

How QLoRA's 4-bit NormalFloat (NF4) quantization reduces base model memory during fine-tuning

Why gradient checkpointing trades extra compute for lower activation memory on a single GPU

Watch out for

Common Fine-Tuning exam traps

  • ▸Assuming LoRA retrains all model weights; it freezes the base model and injects trainable low-rank matrices into selected layers.
  • ▸Confusing NF4 quantization with a quality improvement; its purpose is memory reduction, not higher precision or accuracy.
  • ▸Believing gradient checkpointing speeds up training; it recomputes activations during backward passes, adding compute to save memory.

Question index

All Fine-Tuning questions (45)

Click any question to see the full explanation, or start a practice session above.

1

A team is fine-tuning a 13B-parameter Llama model with NVIDIA NeMo on a node of eight A100 80GB GPUs. They want the optimizer state to be partitioned across data-parallel ranks so that per-GPU memory drops, while keeping the model replicas synchronized. Which distributed strategy should they select in the NeMo training configuration?

Medium
2

Which metric is the most reliable indicator that an LLM is overfitting during the fine-tuning phase?

Medium
3

An ML engineer is fine-tuning a 70B model with NVIDIA NeMo Framework across 16 H100 GPUs. Training completes successfully, but when the fine-tuned checkpoint is evaluated, outputs are incoherent and repeat tokens. The engineer confirms the loss decreased smoothly during training and the validation dataset was held out correctly. Which issue is the most likely explanation?

Hard
4

A team is preparing a supervised fine-tuning job in NVIDIA NeMo Framework for a customer-support assistant. They have a large corpus of raw support chat logs with no labels. They want the model to learn to answer customer questions in the company's tone and format. Which data preparation step is most appropriate before training?

Easy
5

A company wants to teach a pretrained LLM to follow a specific output format for customer support replies using supervised fine-tuning on NVIDIA GPUs. Which data preparation approach best matches supervised fine-tuning for instruction following?

Easy
6

A team is fine-tuning a 70B-parameter LLM with NVIDIA NeMo on a multi-node cluster and wants to reduce the memory footprint per GPU without changing the model architecture. They are already using mixed precision and a reasonable micro-batch size. Which two techniques should they apply? (Choose two.)

Hard
7

A financial services company is fine-tuning an LLM to answer questions about internal policies. The base model performs well on general text but frequently invents policy numbers and effective dates. The team has a curated dataset of 5,000 question-answer pairs with correct citations. Which fine-tuning approach best addresses the hallucination of policy numbers and dates?

Hard
8

A developer is preparing a supervised fine-tuning dataset for an instruction-tuned LLM using NVIDIA NeMo. The dataset contains prompts and responses, but the model sometimes learns to generate the prompt text as part of the response. Which dataset formatting practice should be applied to prevent this?

Easy
9

When fine-tuning on a small, domain-specific dataset, why might adding synthetic data generated by a larger model be beneficial?

Hard
10

Why is it important to use a 'warm-up' period in the learning rate schedule when starting a fine-tuning job?

Medium
11

What is the primary role of the 'Learning Rate Scheduler' during LLM fine-tuning?

Medium
12

When fine-tuning a base LLM using Parameter-Efficient Fine-Tuning (PEFT) on an NVIDIA H100 GPU, what is the primary advantage of utilizing LoRA compared to full fine-tuning?

Medium
13

An ML engineer is fine-tuning a 7B-parameter model with LoRA on a single NVIDIA A100 40GB GPU. The training script reports that the adapter weights are not updating after several hundred steps, and the loss remains flat. The base model weights are frozen as intended. Which LoRA configuration issue is the most likely cause?

Medium
14

An engineer fine-tunes a model on a domain corpus with NVIDIA NeMo and observes that training loss falls steadily while validation loss begins rising after the second epoch. The team must produce the most generalizable checkpoint without changing the dataset. Which action should they take?

Hard
15

What is the primary risk of 'catastrophic forgetting' during the fine-tuning process?

Medium
16

Which TWO of the following practices are considered standard procedures for preparing a dataset for Instruction Fine-Tuning (IFT)? (Choose two)

Hard
17

Refer to the exhibit. An engineer observes that a model fine-tuned with this LoRA configuration is failing to converge on a highly complex legal document domain. What is the most likely cause of this issue?

Hard
18

Refer to the exhibit. In the context of a distributed multi-GPU fine-tuning job, what is the most likely cause of this error?

Hard
19

A developer is preparing a dataset for instruction fine-tuning of an LLM using NVIDIA NeMo. The raw data consists of customer support transcripts with speaker labels and timestamps. Which preprocessing step is most important before training?

Easy
20

Which of the following describes the purpose of 'gradient accumulation' in fine-tuning scenarios?

Medium
21

Which of the following describes the purpose of a 'System Prompt' in Instruction Fine-Tuning?

Medium
22

A team fine-tunes a model with NVIDIA NeMo using a packed sequence dataset and notices that some training samples contain several short conversations concatenated. They must ensure the loss is computed only on assistant responses and not on the packed boundaries. Which configuration detail should they verify?

Hard
23

A team is fine-tuning an 8B-parameter LLM with LoRA on a single NVIDIA A100 80GB GPU using NVIDIA NeMo. They observe that training loss decreases, but validation loss starts to rise after epoch 2. They want to keep the same dataset and hyperparameters but mitigate overfitting. Which change is most appropriate?

Medium
24

An engineer is fine-tuning a 13B parameter model with NVIDIA NeMo using tensor parallelism across four GPUs. After resuming from a checkpoint, training loss spikes and then diverges. The checkpoint was saved with a different tensor parallel size than the current run. What is the most likely cause of the divergence?

Hard
25

A team is fine-tuning a Llama 2 7B model with NVIDIA NeMo Framework on a single A100 80GB GPU. They observe that training loss decreases initially but then diverges, and the model outputs become repetitive and incoherent. The team used a learning rate of 5e-5 with AdamW and no warm-up. Which change is most likely to stabilize training and improve convergence?

Medium
26

What is the primary function of the 'rank' parameter in LoRA?

Easy
27

An enterprise is fine-tuning a large language model using NVIDIA NeMo Framework and encounters GPU out-of-memory errors during the backward pass. The training configuration already uses mixed-precision training (FP16). Which architectural intervention should be applied to resolve memory pressure while retaining the optimizer state precision?

Medium
28

An enterprise is fine-tuning a 34B model with NVIDIA NeMo Framework and observes that the validation loss begins rising after the first epoch while training loss continues to fall. The team wants to reduce this divergence and preserve downstream task quality. (Choose two.)

Hard
29

In the context of NVIDIA NeMo, why is it recommended to use FP8 precision during the fine-tuning process on H100 GPUs?

Medium
30

A research team is fine-tuning a model with NVIDIA NeMo and wants to reduce the risk of catastrophic forgetting of general capabilities while still adapting to a specialized domain. They have a small domain dataset and limited compute. Which fine-tuning approach best balances domain adaptation with retention of pretrained knowledge?

Hard
31

Refer to the exhibit. Which adjustment is the most immediate and effective way to resolve this OOM error while maintaining the same training architecture?

Hard
32

When fine-tuning a Large Language Model using Low-Rank Adaptation (LoRA), which architectural component is primarily modified to reduce computational overhead while maintaining performance?

Medium
33

A team is fine-tuning a Llama 3 8B model with NVIDIA NeMo on a single A100 80GB GPU. They observe that validation loss starts to rise while training loss continues to decrease after epoch 2. They want to keep the best generalizing checkpoint without changing the dataset. Which NeMo training configuration strategy should they apply?

Medium
34

A team is fine-tuning a 70B parameter model with NVIDIA NeMo using LoRA on eight H100 GPUs. They want to reduce GPU memory usage during training while preserving the base model's pretrained knowledge. Which two configuration changes should they apply? (Choose two.)

Medium
35

Which component in the NVIDIA NeMo framework is specifically designed to manage the configuration and orchestration of large-scale fine-tuning jobs?

Easy
36

When using QLoRA for fine-tuning, what is the primary purpose of using the 4-bit NormalFloat (NF4) data type?

Medium
37

A data science team is preparing an instruction fine-tuning dataset in NVIDIA NeMo Framework. They notice that after training, the model performs well on the training instructions but poorly on paraphrased versions of the same instructions. They want to improve generalization without increasing dataset size. Which data preparation change is most appropriate?

Medium
38

An ML engineer is fine-tuning a 13B LLM with LoRA on 4 NVIDIA A100 GPUs using NVIDIA NeMo. They notice that the effective batch size is very small and gradients are noisy, but increasing the per-GPU micro batch size triggers out-of-memory errors. Which technique should they apply to increase the effective batch size without increasing memory per step?

Hard
39

An enterprise team is preparing a supervised fine-tuning job in NVIDIA NeMo for a 20B LLM. They want to reduce GPU memory consumption during training without changing the model architecture or the dataset. Which two configuration changes should they apply? (Choose two.)

Hard
40

A developer needs to fine-tune a 7B LLM for a customer-support chatbot using NVIDIA NeMo. The dataset contains paired instructions and desired responses. Which data format should be used to prepare the dataset for supervised fine-tuning?

Easy
41

Why is gradient checkpointing useful when fine-tuning a model on a single GPU?

Medium
42

Refer to the exhibit. What is the effective batch size for this fine-tuning job?

Hard
43

A company wants to fine-tune a 70B LLM to follow domain-specific instructions. The base model already performs well on general language tasks. They have limited labeled data and limited GPU memory. Which approach is most appropriate?

Medium
44

Which THREE factors significantly influence the memory consumption during LLM fine-tuning? (Choose three)

Hard
45

A team is fine-tuning a 13B-parameter Llama model with NVIDIA NeMo Framework on a single A100 80GB GPU. They apply LoRA adapters to the attention projection layers, but the adapters are producing negligible changes to model behavior even after several epochs, and the loss curve stays flat. They confirm the dataset is clean and the tokenizer is correct. Which LoRA configuration issue is the most likely cause?

Medium

Frequently asked questions

What does the Fine-Tuning domain cover on the NCP-GENL exam?
You must be able to explain how LoRA rank controls adapter capacity, which layers LoRA modifies, why QLoRA uses NF4 quantization, and how gradient checkpointing reduces memory. The key is knowing these techniques trade memory or compute for fine-tuning feasibility without changing the base model's frozen weights.
How many questions are in this domain?
This page lists all 45 Fine-Tuning questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Fine-Tuning questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
nvidia-ncp-genl NVIDIA-NCP-GENL ncp fine tuning Practice Questions