Courseiva

CCNA Fine-Tuning Questions

45 questions · Fine-Tuning · All types, answers revealed

1
MCQmedium

A team is fine-tuning a 13B-parameter Llama model with NVIDIA NeMo on a node of eight A100 80GB GPUs. They want the optimizer state to be partitioned across data-parallel ranks so that per-GPU memory drops, while keeping the model replicas synchronized. Which distributed strategy should they select in the NeMo training configuration?

A.Tensor Parallelism (TP) with tensor_model_parallel_size set to 8
B.Pipeline Parallelism (PP) with pipeline_model_parallel_size set to 8
C.Activation recomputation with recompute_granularity set to full
D.ZeRO Stage 1 via the distributed optimizer in Megatron Core
AnswerD

The distributed optimizer in Megatron Core implements ZeRO Stage 1 semantics by sharding the optimizer states, such as Adam first and second moments plus the master weights, across the data-parallel ranks. Each rank keeps a full model replica for the forward and backward pass, so no extra model-parallel communication is introduced, and per-GPU memory falls in proportion to the data-parallel size.

Why this answer

Sharding optimizer state across data-parallel ranks is exactly what the distributed optimizer in Megatron Core does, providing ZeRO Stage 1 behavior with full model replicas per rank. Tensor and pipeline parallelism change how parameters and layers are split, and recomputation targets activations. Only optimizer-state partitioning reduces the Adam moments and master weights each GPU must hold without adding model-parallel communication.

Exam trap

The trap here is assuming that any multi-GPU parallelism mode reduces optimizer memory, when only data-parallel optimizer sharding actually partitions the Adam state.

2
MCQmedium

Which metric is the most reliable indicator that an LLM is overfitting during the fine-tuning phase?

A.A consistent increase in both training and validation loss.
B.A plateau in the training loss while the validation loss remains stagnant.
C.A decrease in training loss with a simultaneous increase in validation loss.
D.The model generates text that matches the training set exactly.
AnswerC

When the training loss drops while the validation loss rises, the model is overfitting by memorizing specific training data. This divergence signifies that the model is no longer generalizing effectively to new, unseen data, which is the primary definition of overfitting in the context of machine learning.

Why this answer

An increasing validation loss alongside a decreasing training loss is the classic indicator of overfitting. The model is essentially memorizing the training samples rather than learning generalized language patterns. Monitoring this divergence is critical in NVIDIA-based training pipelines, as it allows researchers to implement early stopping or adjust regularization strategies before the model's performance on unseen tasks degrades significantly.

Exam trap

Students often monitor training loss alone, falsely believing that a continually dropping training loss indicates successful model generalization.

3
MCQhard

An ML engineer is fine-tuning a 70B model with NVIDIA NeMo Framework across 16 H100 GPUs. Training completes successfully, but when the fine-tuned checkpoint is evaluated, outputs are incoherent and repeat tokens. The engineer confirms the loss decreased smoothly during training and the validation dataset was held out correctly. Which issue is the most likely explanation?

A.The checkpoint was saved and later loaded with a mismatched tokenizer vocabulary, so input token IDs no longer map to the trained embeddings.
B.The evaluation prompts were longer than the maximum sequence length used during fine-tuning, so the model never learned to handle them.
C.The model was trained with a sequence packing strategy that concatenated unrelated samples without proper attention masking between them.
D.The learning rate was too high, causing the optimizer to overshoot and permanently corrupt the base model weights.
AnswerC

Sequence packing can improve throughput, but if attention masking does not prevent tokens from attending across sample boundaries, the model learns spurious cross-sample dependencies. This degrades generation quality, often producing repetition and incoherence, even though the training loss looks healthy because the model is fitting the corrupted attention pattern.

Why this answer

Sequence packing without correct cross-sample attention masking lets tokens attend to unrelated neighboring samples, teaching the model spurious dependencies that surface as repetition and incoherence at inference. The training loss can still fall smoothly because the model is fitting the corrupted attention structure, so loss alone does not reveal the defect.

Exam trap

The trap here is trusting a smooth loss curve as proof of a correct data pipeline, when packing without proper attention masking can degrade generation quality while the loss still looks healthy.

4
MCQeasy

A team is preparing a supervised fine-tuning job in NVIDIA NeMo Framework for a customer-support assistant. They have a large corpus of raw support chat logs with no labels. They want the model to learn to answer customer questions in the company's tone and format. Which data preparation step is most appropriate before training?

A.Convert the raw chat logs into instruction-response pairs that reflect the desired tone and format, then use them for supervised fine-tuning.
B.Use the raw logs directly as a supervised fine-tuning dataset by treating each message as both instruction and response.
C.Apply reinforcement learning from human feedback using the raw logs as the reward model training data without any preference labels.
D.Run continued pretraining on the raw chat logs so the model absorbs the company's vocabulary and style without any labeling.
AnswerA

Supervised fine-tuning requires paired instruction and response examples that demonstrate the target behavior. Converting raw logs into curated instruction-response pairs gives the model explicit examples of the desired tone and format, which is exactly what supervised fine-tuning learns from. This aligns the training data with the task objective.

Why this answer

Supervised fine-tuning learns behavior from explicit instruction-response pairs, so raw logs must be transformed into curated examples that demonstrate the target tone and format. This gives the model clear, aligned supervision and is the most direct way to shape how the assistant responds to customer questions.

Exam trap

The trap here is assuming any domain-specific text can be used directly for fine-tuning, when supervised fine-tuning specifically requires structured instruction-response pairs that demonstrate the target behavior.

5
MCQeasy

A company wants to teach a pretrained LLM to follow a specific output format for customer support replies using supervised fine-tuning on NVIDIA GPUs. Which data preparation approach best matches supervised fine-tuning for instruction following?

A.Score model outputs with a reward model and update the policy using a policy gradient objective
B.Provide only the desired responses without prompts and train the model to reproduce them verbatim
C.Collect pairs of an instruction prompt and the desired response, then train the model to predict the response tokens
D.Gather a large unlabeled corpus and train the model to predict the next token across all documents
AnswerC

Supervised fine-tuning for instruction following uses prompt and response pairs where the loss is computed on the response tokens. This teaches the model the mapping from instruction to desired output format. NeMo's SFT data formats, such as the prompt-completion and chat schemas, are built around exactly this structure, making it the correct data preparation approach.

Why this answer

Instruction-following SFT learns from prompt and response pairs with the loss applied to response tokens, which directly teaches the desired output format. Continued pretraining on unlabeled text, RLHF with a reward model, and response-only training all lack the prompt-conditioned supervised signal that this task requires.

Exam trap

The trap here is confusing any training on domain text with instruction tuning, when only prompt-response pairs teach the model to follow a requested format.

6
Multi-Selecthard

A team is fine-tuning a 70B-parameter LLM with NVIDIA NeMo on a multi-node cluster and wants to reduce the memory footprint per GPU without changing the model architecture. They are already using mixed precision and a reasonable micro-batch size. Which two techniques should they apply? (Choose two.)

Select 2 answers
A.Enable ZeRO-style optimizer state partitioning across data-parallel ranks.
B.Switch the optimizer to SGD with momentum and remove weight decay.
C.Activate activation recomputation so intermediate activations are discarded and recomputed during the backward pass.
D.Reduce the global batch size by a factor of eight and keep the micro-batch size the same.
E.Increase the number of attention heads to spread the computation across more GPUs.
AnswersA, C

ZeRO-style partitioning shards optimizer states, and optionally gradients and parameters, across data-parallel ranks. Each GPU then holds only a fraction of the optimizer memory, which substantially reduces per-GPU footprint for a 70B model. This directly addresses the memory constraint without altering the architecture and is a standard technique in large-scale NeMo training.

Why this answer

Partitioning optimizer states across data-parallel ranks and recomputing activations during the backward pass both reduce per-GPU memory without changing the model architecture. The first shrinks optimizer memory, and the second shrinks activation memory, which together address the dominant contributors for a 70B model. Changing attention heads, switching optimizers, or shrinking the global batch do not achieve the same effect.

Exam trap

The trap here is confusing batch size reduction with memory reduction, when per-GPU memory during forward and backward passes is largely determined by model states and activations rather than the global batch size.

7
MCQhard

A financial services company is fine-tuning an LLM to answer questions about internal policies. The base model performs well on general text but frequently invents policy numbers and effective dates. The team has a curated dataset of 5,000 question-answer pairs with correct citations. Which fine-tuning approach best addresses the hallucination of policy numbers and dates?

A.Increase the model's temperature during inference so it explores a wider range of possible policy numbers and dates.
B.Apply reinforcement learning from human feedback using a reward model that penalizes any answer longer than two sentences.
C.Perform supervised fine-tuning on the curated question-answer pairs so the model learns to produce answers with correct citations.
D.Continue pre-training on a large corpus of internal documents to expand the model's knowledge of policy numbers and dates.
AnswerC

Supervised fine-tuning on curated question-answer pairs with correct citations directly teaches the model the desired input-output behavior. By training on examples that include accurate policy numbers and effective dates, the model learns to associate questions with grounded answers rather than plausible-sounding fabrications, which is the most direct remedy for the observed hallucination pattern.

Why this answer

Supervised fine-tuning on curated question-answer pairs with correct citations is the most direct way to teach the model the desired behavior. The training examples pair each question with grounded policy numbers and effective dates, so the model learns to reproduce accurate citations instead of inventing them. Continued pre-training, length-based RLHF, and higher inference temperature do not target the specific hallucination of policy details.

Exam trap

The trap here is reaching for continued pre-training or RLHF when the scenario already supplies labeled question-answer pairs, which are the natural input for supervised fine-tuning.

8
MCQeasy

A developer is preparing a supervised fine-tuning dataset for an instruction-tuned LLM using NVIDIA NeMo. The dataset contains prompts and responses, but the model sometimes learns to generate the prompt text as part of the response. Which dataset formatting practice should be applied to prevent this?

A.Shuffle the dataset so prompts and responses are randomly paired during training.
B.Increase the maximum sequence length so the prompt and response always fit in a single sample.
C.Duplicate the prompt in both the input and output fields to reinforce the instruction format.
D.Use a loss mask that excludes prompt tokens so the model is trained only on response tokens.
AnswerD

A loss mask sets the loss contribution of prompt tokens to zero, so the model is optimized only on the response tokens. This prevents the model from learning to reproduce the prompt as output. In NeMo, this is handled through the data configuration and tokenizer settings that mark which tokens contribute to the loss.

Why this answer

Prompt echoing occurs when the loss is computed over both prompt and response tokens. Applying a loss mask that excludes prompt tokens ensures gradients are only derived from response tokens, which teaches the model to generate answers rather than repeat instructions. Duplicating prompts, extending sequence length, or shuffling pairs do not address the loss computation.

Exam trap

The trap here is thinking that dataset formatting alone controls what the model learns, when the loss mask determines which tokens actually contribute gradients.

9
MCQhard

When fine-tuning on a small, domain-specific dataset, why might adding synthetic data generated by a larger model be beneficial?

A.It eliminates the need for any real-world human-annotated data entirely.
B.It helps the model generalize better by increasing data diversity.
C.It reduces the VRAM usage of the training process significantly.
D.It guarantees the model will not have any factual inaccuracies.
AnswerB

By providing a larger volume of varied examples, synthetic data helps the model learn to handle diverse inputs within the domain. This diversity is crucial when the primary dataset is limited, as it prevents the model from overfitting to a narrow set of patterns, effectively increasing its overall robustness.

Why this answer

Small datasets often lack the depth needed for a model to generalize effectively. Synthetic data can fill these gaps by providing more examples of the target task, which helps the model learn the nuances of the domain. This technique, when done correctly, reinforces desired behaviors and prevents overfitting on the limited original data, leading to a more robust, versatile, and high-performing model in the intended application area.

Exam trap

Candidates often assume synthetic data is only for increasing volume, missing the key benefit of improving generalization and reducing overfitting when dealing with limited, niche, or domain-specific training sets.

10
MCQmedium

Why is it important to use a 'warm-up' period in the learning rate schedule when starting a fine-tuning job?

A.To increase the total training time for better hardware utilization.
B.To allow the optimizer to adapt to the new data distribution gradually.
C.To force the model to explore a wider range of the loss landscape immediately.
D.To bypass the need for gradient clipping during training.
AnswerB

A gradual warm-up phase prevents abrupt weight changes that could lead to divergence. By starting with a low learning rate, the optimizer stabilizes its states based on the new data distribution, which is critical for maintaining the model's integrity and achieving steady, high-quality convergence during the fine-tuning process.

Why this answer

A warm-up period gradually increases the learning rate from a near-zero value to the target rate. This prevents early, massive gradient updates from destabilizing the pre-trained weights. By slowly introducing the learning rate, the model remains stable during the initial phase of training, ensuring that the optimizer can effectively adapt to the new domain without destroying the fundamental knowledge captured during the original pre-training process.

Exam trap

Candidates incorrectly believe warm-up is to save energy or reduce latency, missing that its primary purpose is preventing gradient spikes from destabilizing pre-trained weights during the initial training phase.

11
MCQmedium

What is the primary role of the 'Learning Rate Scheduler' during LLM fine-tuning?

A.To increase the batch size dynamically
B.To control the weight updates over the training duration
C.To automatically detect the optimal training loss
D.To determine which layers should be frozen
AnswerB

The scheduler dictates how the learning rate changes over time, allowing the model to make large updates initially and smaller, more precise updates as training progresses. This helps the model converge reliably to a stable state, preventing the training process from oscillating or diverging due to a static, high learning rate.

Why this answer

A scheduler adjusts the learning rate throughout the training process, typically starting with a warmup phase to stabilize the weights and then decaying the rate to allow for fine-grained convergence. This prevents the model from overshooting optimal solutions early on and helps it settle into a good local minimum, which is critical for achieving high-quality results without divergence.

Exam trap

Candidates mistakenly believe learning rate schedulers control batch size or regularization penalties, forgetting their direct responsibility for dynamically modifying weight updates over time.

12
MCQmedium

When fine-tuning a base LLM using Parameter-Efficient Fine-Tuning (PEFT) on an NVIDIA H100 GPU, what is the primary advantage of utilizing LoRA compared to full fine-tuning?

A.It increases the number of trainable parameters to improve model convergence speed.
B.It eliminates the need for any gradient calculation during the backpropagation process.
C.It reduces the memory footprint by freezing base weights and training only low-rank adapter matrices.
D.It forces the model to ignore long-range dependencies to accelerate training time.
AnswerC

Freezing the primary model parameters eliminates the need to store optimizer states for the vast majority of the weights. By training only small, low-rank matrices, the total VRAM consumption is drastically reduced, allowing for larger batch sizes and faster training cycles on NVIDIA GPU hardware platforms.

Why this answer

LoRA reduces memory overhead by freezing pre-trained model weights and injecting trainable rank-decomposition matrices into transformer layers. This approach significantly lowers the VRAM requirements, enabling fine-tuning on consumer or enterprise hardware without needing to load the entire parameter set into the optimizer state. This method is critical for deploying domain-specific models efficiently while maintaining performance parity with full fine-tuning, thus optimizing resource utilization across high-performance NVIDIA compute clusters.

Exam trap

Candidates often assume LoRA reduces inference latency or speeds up training time, when its primary benefit is lowering memory usage by freezing base weights.

13
MCQmedium

An ML engineer is fine-tuning a 7B-parameter model with LoRA on a single NVIDIA A100 40GB GPU. The training script reports that the adapter weights are not updating after several hundred steps, and the loss remains flat. The base model weights are frozen as intended. Which LoRA configuration issue is the most likely cause?

A.The base model is loaded in 8-bit precision, which prevents any gradient computation on the adapters.
B.The learning rate is too low, so the adapter weights change by amounts below floating-point precision.
C.The LoRA rank is set too high, causing the adapters to be initialized to zero.
D.The target modules list is empty or points to modules that are not part of the model's forward pass.
AnswerD

LoRA only updates the adapter matrices attached to the specified target modules. If the target modules list is empty or names modules that never execute, no adapter parameters participate in the forward pass, so no gradients reach them and the adapters remain unchanged. This exactly matches the symptom of frozen adapters and flat loss.

Why this answer

If the LoRA target modules list is empty or references modules that are not executed during the forward pass, the adapter parameters receive no gradients and never update. The base model remains frozen as intended, but the adapters are effectively absent from training, which explains both the flat loss and the unchanged adapter weights.

Exam trap

The trap here is assuming that a frozen base model or a low learning rate explains non-updating adapters, when the more likely cause is that the adapters are not attached to any executed module.

14
MCQhard

An engineer fine-tunes a model on a domain corpus with NVIDIA NeMo and observes that training loss falls steadily while validation loss begins rising after the second epoch. The team must produce the most generalizable checkpoint without changing the dataset. Which action should they take?

A.Increase the number of training epochs so the model fully converges on the domain corpus
B.Disable validation during training and rely on the final training loss to select the checkpoint
C.Enable early stopping based on validation loss and restore the best checkpoint
D.Raise the learning rate to escape the local minimum the model has settled into
AnswerC

The described divergence between falling training loss and rising validation loss is the classic overfitting signal. Monitoring validation loss and stopping when it stops improving, then restoring the checkpoint with the lowest validation loss, yields the most generalizable model. NeMo supports validation-interval evaluation and checkpoint selection based on a monitored metric, so this is the direct remedy.

Why this answer

When training loss keeps falling but validation loss turns upward, the model is overfitting and the best generalizing weights occur near the point where validation loss is lowest. Early stopping with best-checkpoint restoration captures that point. More epochs, a higher learning rate, or dropping validation all fail to address the divergence or actively make it worse.

Exam trap

The trap here is treating a falling training loss as evidence of healthy progress when the rising validation loss is the decisive signal.

15
MCQmedium

What is the primary risk of 'catastrophic forgetting' during the fine-tuning process?

A.The model becomes unable to produce responses in the desired language
B.The model loses the general knowledge acquired during pre-training
C.The model's inference speed increases significantly
D.The model begins to output only empty strings
AnswerB

Catastrophic forgetting refers to the phenomenon where a model's weights are modified so drastically that it loses its proficiency in tasks it was previously capable of. This happens when the fine-tuning loss prioritizes the new dataset too heavily, forcing the weights to shift away from the general knowledge base established during pre-training.

Why this answer

Catastrophic forgetting occurs when a model is updated on new, narrow tasks, causing it to lose the broad, generalized knowledge acquired during its extensive pre-training phase. In production, this renders the model useless for its original intended tasks. Mitigating this risk is crucial for businesses that need to maintain multi-purpose model capabilities while still achieving performance gains in specialized domains.

Exam trap

Candidates often confuse catastrophic forgetting with overfitting to the training set, missing that catastrophic forgetting specifically involves losing broad, generalized pre-trained capabilities.

16
Multi-Selecthard

Which TWO of the following practices are considered standard procedures for preparing a dataset for Instruction Fine-Tuning (IFT)? (Choose two)

Select 2 answers
A.Removing duplicate or highly redundant instructional examples
B.Applying aggressive token truncation to all training samples
C.Ensuring consistent schema and prompt formatting
D.Converting all text into numerical vectors before training
E.Increasing the learning rate by a factor of 100
AnswersA, C

Redundant data leads to overfitting and skewed model responses. Ensuring a clean, deduplicated dataset forces the model to learn the underlying logic of the instructions rather than memorizing specific sequences, which is essential for developing a robust model that can generalize to novel user prompts effectively during inference.

Why this answer

Instruction fine-tuning requires high-quality, diverse, and well-formatted data to ensure the model learns to follow specific user prompts. Removing duplicates prevents the model from over-relying on single examples, while consistent formatting (like ChatML or Alpaca format) ensures the model learns the structural cues for turn-based conversation, which is fundamental for effective performance in downstream instruction-following tasks.

Exam trap

Candidates often overlook data quality, assuming that simply increasing the volume of training examples is sufficient, while ignoring the negative impact of duplicates and inconsistent schema formatting on model convergence.

17
MCQhard

Refer to the exhibit. An engineer observes that a model fine-tuned with this LoRA configuration is failing to converge on a highly complex legal document domain. What is the most likely cause of this issue?

A.The dropout value is set too high, causing the model to underfit.
B.The target modules are incorrectly specified, leading to a loss of attention.
C.The rank value is too low to represent the complex patterns of the domain.
D.The alpha parameter is too high, causing gradient explosion.
AnswerC

For complex domains like legal or technical writing, a rank of 8 may be insufficient to capture the necessary parameter updates. Increasing the rank allows the model to capture a richer set of features, providing the additional capacity needed to learn complex domain-specific linguistic relationships and improve convergence.

Why this answer

The rank of 8 is likely too low for capturing the intricate, nuanced patterns required for legal documentation. While low-rank adaptation is efficient, complex domains often necessitate higher rank values to provide enough expressive capacity within the trainable adapter layers. This configuration mismatch limits the model's ability to learn the specific syntactic and semantic structures inherent in specialized legal texts, leading to poor convergence and inadequate performance.

Exam trap

Candidates often assume the failure stems from the learning rate or data quality, overlooking the LoRA rank parameter, which directly dictates the model's capacity to learn complex, domain-specific semantic patterns.

18
MCQhard

Refer to the exhibit. In the context of a distributed multi-GPU fine-tuning job, what is the most likely cause of this error?

A.The model weights are too large for the GPU memory
B.The MASTER_ADDR or MASTER_PORT environment variables are misconfigured
C.The learning rate is too high, causing gradient explosion
D.The dataset is missing required training samples
AnswerB

Distributed training frameworks rely on these variables to establish a communication channel. If the address or port is blocked, inaccessible, or incorrect, the NCCL library cannot handshake between nodes, leading to the connection refused error. Correcting these settings is the standard solution for resolving distributed networking issues in NVIDIA environments.

Why this answer

This error indicates that the NCCL (NVIDIA Collective Communications Library) is failing to communicate between GPU nodes or processes. It is typically caused by a misconfiguration of the environment variables (like MASTER_ADDR or MASTER_PORT) or a firewall blocking the necessary ports. In distributed training, these network configurations are essential for synchronizing gradients across all participating GPUs during the training process.

Exam trap

Candidates frequently mistake NCCL communication timeouts or address errors for insufficient GPU VRAM, overlooking network-level environment variables like MASTER_ADDR and MASTER_PORT required for multi-node synchronization.

19
MCQeasy

A developer is preparing a dataset for instruction fine-tuning of an LLM using NVIDIA NeMo. The raw data consists of customer support transcripts with speaker labels and timestamps. Which preprocessing step is most important before training?

A.Convert each transcript into a structured instruction-response pair with a clear prompt and a target completion.
B.Duplicate each transcript so the model sees every example at least twice per epoch.
C.Remove all punctuation and capitalization so the model learns a consistent lowercase style.
D.Increase the learning rate to compensate for the noisy transcript data.
AnswerA

Instruction fine-tuning expects examples in a prompt-completion or instruction-response format so the loss is computed on the desired assistant output. Raw transcripts with speaker labels and timestamps do not directly teach the model how to respond to instructions. Converting them into structured pairs aligns the data with the training objective and yields useful gradients.

Why this answer

Instruction fine-tuning requires examples that pair a prompt with the desired response so the model learns to follow instructions. Raw support transcripts contain speaker labels and timestamps that are not part of the target behavior. Converting each transcript into a structured instruction-response pair makes the data compatible with the training objective and ensures the loss is computed on the correct tokens.

Exam trap

The trap here is treating preprocessing as a hyperparameter or volume problem, when the real issue is reformatting raw transcripts into instruction-response pairs.

20
MCQmedium

Which of the following describes the purpose of 'gradient accumulation' in fine-tuning scenarios?

A.To increase the speed of the training process by bypassing the GPU's bottleneck.
B.To enable training with larger effective batch sizes on memory-constrained hardware.
C.To automatically optimize the learning rate based on the model's gradient magnitude.
D.To reduce the number of parameters being fine-tuned in the model.
AnswerB

This technique allows engineers to maintain effective batch sizes that would otherwise exceed available VRAM. By accumulating gradients over multiple steps and updating weights only after reaching a target batch size, the training process achieves the stability of large-batch learning while keeping the peak memory usage manageable.

Why this answer

Gradient accumulation allows for simulating larger batch sizes when GPU VRAM is limited. By performing multiple forward and backward passes without updating the weights, the model can aggregate gradients across several smaller steps. This is crucial when working on hardware with restricted memory, as it enables the model to benefit from the statistical stability of larger batches, improving the overall quality of the fine-tuned model's convergence.

Exam trap

Candidates confuse gradient accumulation with model parallelism, failing to realize it is specifically a memory-saving technique that simulates larger batches without requiring additional GPU memory for simultaneous activation storage.

21
MCQmedium

Which of the following describes the purpose of a 'System Prompt' in Instruction Fine-Tuning?

A.To increase the number of tokens processed in the output
B.To act as a persistent instruction defining the model's persona
C.To replace the need for domain-specific fine-tuning entirely
D.To compress the training dataset size
AnswerB

The system prompt acts as a foundational instruction that dictates the model's behavior, tone, and constraints. It provides the necessary context for the assistant to follow during multi-turn conversations, ensuring that the model remains aligned with its intended role and follows specified safety and quality guidelines throughout the interaction.

Why this answer

The system prompt provides a high-level instruction or persona for the model, which guides its overall behavior throughout the conversation. It sets expectations for tone, safety constraints, and task-specific roles. Effectively designing the system prompt is essential for ensuring that the fine-tuned model consistently adheres to the required persona or operational boundaries in real-world production deployments.

Exam trap

Candidates often confuse system prompts with few-shot examples or model weights. Remember that a system prompt is a high-level, persistent instruction that guides overall behavior, tone, and boundaries throughout a conversation.

22
MCQhard

A team fine-tunes a model with NVIDIA NeMo using a packed sequence dataset and notices that some training samples contain several short conversations concatenated. They must ensure the loss is computed only on assistant responses and not on the packed boundaries. Which configuration detail should they verify?

A.That the global batch size is increased to compensate for the additional tokens in each packed sequence
B.That the tokenizer vocabulary is expanded to include special separators between packed samples
C.That the loss mask aligns with each sample's response tokens and that packed sequences are separated by an attention boundary
D.That gradient accumulation steps equal the number of samples packed into each sequence
AnswerC

Packed sequences concatenate multiple samples into one training sequence for efficiency, so the loss mask must mark only assistant response tokens and the attention mechanism must prevent cross-sample attention. In NeMo this is handled through per-token loss masks and sequence boundary handling. Verifying both ensures the model is not trained on padding or on tokens from adjacent samples.

Why this answer

Packed sequences boost training efficiency by filling each context window with multiple samples, but correctness depends on the loss mask covering only assistant response tokens and on attention being blocked across sample boundaries. Batch size, vocabulary changes, and gradient accumulation do not affect these two requirements.

Exam trap

The trap here is treating packed sequences as a pure throughput optimization and overlooking that masks and attention boundaries must be adjusted to keep the loss correct.

23
MCQmedium

A team is fine-tuning an 8B-parameter LLM with LoRA on a single NVIDIA A100 80GB GPU using NVIDIA NeMo. They observe that training loss decreases, but validation loss starts to rise after epoch 2. They want to keep the same dataset and hyperparameters but mitigate overfitting. Which change is most appropriate?

A.Increase the batch size and scale the learning rate proportionally to speed convergence.
B.Switch the optimizer from AdamW to SGD with momentum to regularize the adapter.
C.Increase the LoRA rank from 8 to 64 to give the adapter more capacity.
D.Reduce the number of training epochs and apply early stopping based on validation loss.
AnswerD

The reported pattern, training loss continuing to fall while validation loss rises after epoch 2, indicates the model has begun to overfit. Stopping training at or before the point of minimum validation loss directly counteracts that behavior without changing the data or architecture. Early stopping is a standard, low-risk mitigation that preserves the best generalizing checkpoint.

Why this answer

The divergence between decreasing training loss and increasing validation loss after epoch 2 is a textbook overfitting signal. The most direct, minimal-risk remedy is to stop training earlier using validation loss as the criterion, preserving the checkpoint that generalizes best. Enlarging adapter capacity or altering optimizer and batch settings does not target the generalization gap and may worsen it.

Exam trap

The trap here is assuming that a richer adapter or faster optimizer will improve results, when the observed validation curve already shows the model is memorizing rather than generalizing.

24
MCQhard

An engineer is fine-tuning a 13B parameter model with NVIDIA NeMo using tensor parallelism across four GPUs. After resuming from a checkpoint, training loss spikes and then diverges. The checkpoint was saved with a different tensor parallel size than the current run. What is the most likely cause of the divergence?

A.The optimizer state and model shards were partitioned differently, so the restored weights do not match the current tensor parallel layout.
B.The learning rate scheduler restarted from step zero and applied a large learning rate to all parameters.
C.Gradient accumulation steps were reduced, effectively increasing the global batch size beyond the original configuration.
D.The data loader random seed was not preserved, causing the model to see samples in a different order.
AnswerA

Tensor parallelism splits weight matrices across GPUs, and the sharding layout depends on tensor parallel size. A checkpoint saved with one tensor parallel degree cannot be directly loaded into a run with a different degree without resharding. The mismatched partitioning causes incorrect weight placement, leading to loss spikes and divergence when training resumes.

Why this answer

Tensor parallelism partitions model weights across GPUs, and the partition layout depends on the tensor parallel size. Loading a checkpoint saved with a different tensor parallel degree without resharding places weights incorrectly, causing loss spikes and divergence. Data order, scheduler restart, and gradient accumulation changes do not explain the immediate failure tied to a changed parallel layout.

Exam trap

The trap here is overlooking that checkpoint compatibility in NeMo depends not only on model architecture but also on the parallelism configuration used when the checkpoint was written.

25
MCQmedium

A team is fine-tuning a Llama 2 7B model with NVIDIA NeMo Framework on a single A100 80GB GPU. They observe that training loss decreases initially but then diverges, and the model outputs become repetitive and incoherent. The team used a learning rate of 5e-5 with AdamW and no warm-up. Which change is most likely to stabilize training and improve convergence?

A.Increase the learning rate to 1e-4 to escape local minima.
B.Switch the optimizer to SGD with momentum 0.9 and keep the learning rate at 5e-5.
C.Reduce the learning rate to 1e-5 and add a linear warm-up over the first 10% of training steps.
D.Increase the batch size to 64 and keep all other hyperparameters unchanged.
AnswerC

A lower learning rate (e.g., 1e-5) combined with warm-up prevents large initial updates that can destabilize training. Warm-up gradually increases the learning rate, allowing the model to adapt smoothly. This is a standard practice for fine-tuning large language models, especially when full fine-tuning or using AdamW, and directly addresses the observed divergence and repetitive outputs.

Why this answer

The model diverges due to a learning rate that is too high for fine-tuning without warm-up. Reducing the learning rate to 1e-5 and adding a warm-up phase allows the optimizer to take smaller, more controlled steps initially, preventing large updates that disrupt pretrained weights. This combination is a well-established best practice in NVIDIA NeMo and other frameworks for stable fine-tuning of LLMs.

Exam trap

The trap here is assuming that a higher learning rate always speeds up convergence, when in fine-tuning it often causes divergence and degraded output quality.

26
MCQeasy

What is the primary function of the 'rank' parameter in LoRA?

A.It sets the total number of layers that are trainable
B.It defines the dimension of the low-rank decomposition
C.It dictates the number of epochs the model will train
D.It determines the learning rate for the adapter weights
AnswerB

The rank 'r' specifies the size of the low-rank matrices. For example, if a weight matrix has dimensions (d, d), LoRA decomposes it into (d, r) and (r, d) matrices. 'r' is typically a small integer, which keeps the parameter count very low compared to full fine-tuning.

Why this answer

The rank parameter ('r') determines the dimensionality of the update matrices injected into the model. A lower rank results in a smaller number of trainable parameters, which is more efficient but less expressive. A higher rank allows for more complex adaptations to the data.

Balancing this rank is essential to achieve the desired model performance while staying within hardware memory limits.

Exam trap

Candidates often incorrectly assume the rank parameter controls the number of layers being trained or the learning rate, rather than identifying it as the dimension of the low-rank decomposition matrices.

27
MCQmedium

An enterprise is fine-tuning a large language model using NVIDIA NeMo Framework and encounters GPU out-of-memory errors during the backward pass. The training configuration already uses mixed-precision training (FP16). Which architectural intervention should be applied to resolve memory pressure while retaining the optimizer state precision?

A.Switch the global precision setting from FP16 to INT8 quantization.
B.Implement activation checkpointing to recompute intermediate activations during backward passes.
C.Disable gradient accumulation completely to process each micro-batch independently.
D.Migrate the optimizer states from FP32 to FP8 format.
AnswerB

Activation checkpointing selectively saves specific layer activations and recomputes the discarded ones during the backward pass. This drastically reduces peak GPU memory consumption, enabling successful fine-tuning of larger models on NVIDIA hardware without modifying the core precision configuration.

Why this answer

Applying activation checkpointing trades off compute for memory by recalculating activations during the backward pass instead of storing them all. This directly mitigates out-of-memory errors on NVIDIA GPUs during LLM fine-tuning without requiring a reduction in batch size or model accuracy, making it a standard best practice in NeMo Framework workflows.

Exam trap

Candidates often try to resolve backward pass OOM errors by reducing optimizer precision or model layers, missing activation checkpointing as the standard memory-compute trade-off.

28
Multi-Selecthard

An enterprise is fine-tuning a 34B model with NVIDIA NeMo Framework and observes that the validation loss begins rising after the first epoch while training loss continues to fall. The team wants to reduce this divergence and preserve downstream task quality. (Choose two.)

Select 2 answers
A.Apply weight decay or increase its coefficient to penalize large parameter updates and discourage overfitting.
B.Remove the validation split and train on all available data so the model sees more examples per epoch.
C.Raise the learning rate significantly so the model converges faster and reaches a flatter minimum.
D.Increase the number of training epochs so the optimizer has more time to escape the overfitting region.
E.Enable early stopping based on validation loss so training halts before the model overfits further.
AnswersA, E

Weight decay regularizes the model by penalizing large weights, which reduces the tendency to memorize training-specific patterns. Increasing its coefficient when validation loss diverges is a targeted regularization response that can narrow the gap between training and validation performance while preserving the learned task behavior.

Why this answer

Rising validation loss with falling training loss is classic overfitting. Early stopping on validation loss halts training before further damage, while stronger weight decay regularizes the model and narrows the generalization gap. Together they reduce divergence and protect downstream task quality without discarding useful training signal.

Exam trap

The trap here is interpreting a rising validation loss as a need for more training or a faster learning rate, when the divergence actually calls for regularization and earlier stopping.

29
MCQmedium

In the context of NVIDIA NeMo, why is it recommended to use FP8 precision during the fine-tuning process on H100 GPUs?

A.It guarantees 100% precision parity with full-precision floating-point training.
B.It significantly improves memory throughput and speed via Hopper-specific hardware.
C.It disables the need for gradient scaling during the backward pass.
D.It forces the model to use only CPU-based memory for training weights.
AnswerB

NVIDIA H100 GPUs include specialized hardware support for FP8, which accelerates matrix multiplication and reduces memory footprint. This allows the model to process data much faster than traditional precision formats, providing a competitive advantage in training workflows that require high-performance compute and rapid turnaround times.

Why this answer

FP8 precision leverages the Transformer Engine on NVIDIA Hopper architecture, providing a significant boost in throughput while reducing memory usage. By using FP8, developers can fit larger models or increase batch sizes without sacrificing significant numerical stability. This optimization is essential for modern fine-tuning workflows where computational efficiency directly impacts the speed of iteration and the scalability of training across multi-GPU nodes in an enterprise environment.

Exam trap

Candidates assume FP8 is universally supported across all GPUs, failing to recognize that it requires specific hardware features like the Hopper architecture's Transformer Engine.

30
MCQhard

A research team is fine-tuning a model with NVIDIA NeMo and wants to reduce the risk of catastrophic forgetting of general capabilities while still adapting to a specialized domain. They have a small domain dataset and limited compute. Which fine-tuning approach best balances domain adaptation with retention of pretrained knowledge?

A.Full-parameter fine-tuning with a high learning rate for many epochs on the domain dataset.
B.Freezing all layers and training only the final classification head on the domain dataset.
C.Parameter-efficient fine-tuning with LoRA using a modest rank and a low learning rate.
D.Training from scratch on the domain dataset using the same architecture and tokenizer.
AnswerC

LoRA freezes the pretrained weights and trains small low-rank adapters, which limits drift from the original model and reduces forgetting. A modest rank and low learning rate further constrain updates, making it well suited for small datasets and limited compute. This approach balances domain adaptation with retention of general capabilities.

Why this answer

LoRA with a modest rank and low learning rate freezes the base model and trains small adapters, which constrains updates and preserves pretrained knowledge while still adapting to the domain. Full fine-tuning with high learning rates, training from scratch, or freezing all layers and training only a head either cause forgetting, require excessive resources, or fail to adapt the generative model.

Exam trap

The trap here is equating strong domain adaptation with full fine-tuning, when a small dataset and limited compute make parameter-efficient methods both safer and more practical.

31
MCQhard

Refer to the exhibit. Which adjustment is the most immediate and effective way to resolve this OOM error while maintaining the same training architecture?

A.Increase the sequence length of the inputs.
B.Decrease the batch size per GPU.
C.Switch from mixed precision to full FP32 precision.
D.Disable the optimizer state checkpointing.
AnswerB

Decreasing the batch size is the most effective way to reduce memory consumption immediately. It lowers the VRAM overhead required for activation storage, allowing the model to fit within the existing hardware constraints. This is the standard approach to resolving OOM errors during the fine-tuning training process.

Why this answer

The error indicates that the current batch size is too large for the available VRAM. Reducing the batch size is the most direct solution to free up enough memory for the model to continue training. If a larger effective batch size is required for convergence, the user can subsequently enable gradient accumulation to compensate for the reduced per-step batch size without increasing the memory footprint.

Exam trap

Candidates often try to change model architecture or hardware settings first, ignoring that reducing the batch size is the most immediate and effective way to resolve OOM errors.

32
MCQmedium

When fine-tuning a Large Language Model using Low-Rank Adaptation (LoRA), which architectural component is primarily modified to reduce computational overhead while maintaining performance?

A.The entire embedding layer matrix
B.The feed-forward network activation functions
C.The attention weight projection matrices
D.The entire decoder hidden state output
AnswerC

LoRA injects trainable low-rank matrices into the attention mechanism's query, key, and value projections. By adapting these specific components, the model learns to capture domain-specific patterns without updating the billions of frozen parameters, drastically reducing VRAM consumption and making model fine-tuning feasible on limited NVIDIA hardware resources.

Why this answer

LoRA freezes the pre-trained model weights and injects trainable rank decomposition matrices into the transformer architecture layers. By targeting specifically the attention query, key, and value projection matrices, practitioners can achieve high performance with a fraction of the trainable parameters. This approach is critical for memory-constrained environments, allowing fine-tuning on consumer-grade NVIDIA GPUs while avoiding the massive memory requirements associated with full parameter updates.

Exam trap

Candidates often mistakenly believe LoRA modifies the entire transformer block, failing to identify that it specifically targets the attention weight projection matrices to optimize memory and computational efficiency.

33
MCQmedium

A team is fine-tuning a Llama 3 8B model with NVIDIA NeMo on a single A100 80GB GPU. They observe that validation loss starts to rise while training loss continues to decrease after epoch 2. They want to keep the best generalizing checkpoint without changing the dataset. Which NeMo training configuration strategy should they apply?

A.Reduce the validation split size so that validation loss more closely tracks training loss.
B.Enable early stopping based on validation loss and save the best checkpoint using the validation metric.
C.Increase the number of training epochs and rely on the final checkpoint for deployment.
D.Switch the optimizer from AdamW to SGD with a higher momentum value.
AnswerB

Early stopping monitors validation loss and halts training when it stops improving, preserving the checkpoint with the best validation metric. In NeMo, this is configured through the checkpointing and early stopping callbacks. Since the scenario shows overfitting after epoch 2, stopping at the best validation point keeps generalization without altering data or model architecture.

Why this answer

The divergence between decreasing training loss and increasing validation loss indicates overfitting. The correct response is to use validation-based early stopping and retain the best checkpoint, which NeMo supports through its checkpointing and early stopping mechanisms. Increasing epochs, changing optimizers, or shrinking validation data do not address the core generalization problem.

Exam trap

The trap here is assuming that more training or a different optimizer will fix overfitting when the validation curve already shows the model is past its best generalization point.

34
Multi-Selectmedium

A team is fine-tuning a 70B parameter model with NVIDIA NeMo using LoRA on eight H100 GPUs. They want to reduce GPU memory usage during training while preserving the base model's pretrained knowledge. Which two configuration changes should they apply? (Choose two.)

Select 2 answers
A.Increase the micro batch size to improve GPU utilization and reduce the number of optimizer steps.
B.Enable LoRA by setting adapter dimensions and target modules so only adapter weights receive gradients.
C.Disable tensor parallelism so each GPU holds a full copy of the model weights.
D.Switch from bfloat16 to float32 precision to improve numerical stability during training.
E.Enable gradient checkpointing to recompute activations during the backward pass instead of storing them.
AnswersB, E

LoRA freezes the base model and trains low-rank adapter matrices, which drastically reduces the number of trainable parameters and optimizer state. This lowers memory usage during fine-tuning while preserving the pretrained weights. In NeMo, this is configured through the PEFT section with adapter dimensions and target modules, making it a direct memory-saving measure.

Why this answer

LoRA reduces trainable parameters and optimizer state while freezing the base model, and gradient checkpointing reduces activation memory by recomputing activations. Together they lower peak GPU memory during fine-tuning. Increasing micro batch size, using float32, or disabling tensor parallelism all increase memory usage and work against the stated goal.

Exam trap

The trap here is treating bfloat16 as a stability-only choice and forgetting that float32 doubles memory, which conflicts with the explicit memory-reduction objective.

35
MCQeasy

Which component in the NVIDIA NeMo framework is specifically designed to manage the configuration and orchestration of large-scale fine-tuning jobs?

A.TensorRT-LLM
B.NeMo Framework Launcher
C.NVIDIA Triton Inference Server
D.CUDA Toolkit
AnswerB

The NeMo Framework Launcher provides the necessary abstractions to configure, submit, and manage large-scale fine-tuning tasks. It streamlines the workflow by handling job scheduling and resource allocation, making it the primary tool for orchestrating model training on NVIDIA hardware platforms effectively and reliably at scale.

Why this answer

The NeMo Framework Launcher is the dedicated tool for managing multi-node, large-scale training jobs. It abstracts the complexities of cluster scheduling, environment setup, and hyperparameter management. By using the launcher, engineers ensure that fine-tuning tasks are executed efficiently across NVIDIA compute clusters, allowing for reproducible and scalable experiments that align with enterprise-grade development standards and best practices for large model management.

Exam trap

Candidates often confuse the NeMo Framework Launcher with base PyTorch or general-purpose CI/CD tools like Jenkins, failing to recognize the launcher's specific role in abstracting multi-node cluster scheduling and environment orchestration.

36
MCQmedium

When using QLoRA for fine-tuning, what is the primary purpose of using the 4-bit NormalFloat (NF4) data type?

A.To increase the training speed by using integer-only arithmetic
B.To reduce the memory footprint of the model weights
C.To improve the convergence speed of the optimizer
D.To enable training without the need for gradient accumulation
AnswerB

NF4 quantizes the base model weights to 4-bit precision, which drastically lowers the memory requirement compared to 16-bit or 32-bit representations. This allows users to fine-tune significantly larger models on a single NVIDIA GPU, as the memory bottleneck is primarily the storage of the frozen model weights during the training process.

Why this answer

The NF4 data type is mathematically optimal for weights that follow a normal distribution, which is typical for pre-trained language model weights. By quantizing weights to 4 bits, QLoRA significantly reduces the memory footprint, allowing large models to fit onto GPUs with lower VRAM. This efficiency does not significantly sacrifice performance, provided the weights are dequantized during the forward pass to maintain precision for activations.

Exam trap

Candidates often assume 4-bit quantization (NF4) permanently degrades model accuracy or serves only for inference, ignoring its primary role in reducing training memory footprints via QLoRA.

37
MCQmedium

A data science team is preparing an instruction fine-tuning dataset in NVIDIA NeMo Framework. They notice that after training, the model performs well on the training instructions but poorly on paraphrased versions of the same instructions. They want to improve generalization without increasing dataset size. Which data preparation change is most appropriate?

A.Apply instruction paraphrasing and template diversification so the same intent appears in multiple surface forms during training.
B.Remove all examples where the instruction phrasing is unique, keeping only the most common templates.
C.Increase the number of epochs so the model sees each instruction more times and memorizes the exact phrasing.
D.Lower the learning rate so the model updates more slowly and avoids overfitting to specific instruction phrasings.
AnswerA

Paraphrasing and template diversification expose the model to varied surface forms of the same intent, which teaches it to map meaning rather than exact wording. This directly improves generalization to unseen paraphrases without adding new intents, and it is a standard data augmentation technique for instruction tuning.

Why this answer

Generalization to paraphrased instructions depends on the model seeing the same intent expressed in varied surface forms. Paraphrasing and template diversification teach the model to respond to meaning rather than exact wording, directly improving performance on unseen paraphrases without enlarging the dataset or changing training hyperparameters.

Exam trap

The trap here is treating poor paraphrase generalization as an optimizer or epoch problem, when the real fix is increasing surface-form diversity in the instruction data itself.

38
MCQhard

An ML engineer is fine-tuning a 13B LLM with LoRA on 4 NVIDIA A100 GPUs using NVIDIA NeMo. They notice that the effective batch size is very small and gradients are noisy, but increasing the per-GPU micro batch size triggers out-of-memory errors. Which technique should they apply to increase the effective batch size without increasing memory per step?

A.Switch from bfloat16 to float32 precision to stabilize gradient values.
B.Enable gradient accumulation to sum gradients across multiple micro batches before the optimizer step.
C.Increase the LoRA rank and alpha so the adapter captures more information per step.
D.Reduce the sequence length by truncating all training examples to 128 tokens.
AnswerB

Gradient accumulation executes several forward and backward passes on small micro batches, accumulating gradients before applying a single optimizer update. This raises the effective batch size without increasing peak memory, because only one micro batch resides in memory at a time. It directly addresses the noisy-gradient problem while respecting the memory ceiling that prevents larger micro batches.

Why this answer

Gradient accumulation performs multiple forward and backward passes on small micro batches and sums their gradients before one optimizer update, emulating a larger batch while keeping peak memory at the level of a single micro batch. It is the standard remedy when memory prevents enlarging the micro batch. Changing adapter rank, precision, or sequence length does not increase the number of samples per update in a memory-neutral way.

Exam trap

The trap here is equating adapter capacity or numeric precision with batch size, when only gradient accumulation increases effective batch size without raising peak memory.

39
Multi-Selecthard

An enterprise team is preparing a supervised fine-tuning job in NVIDIA NeMo for a 20B LLM. They want to reduce GPU memory consumption during training without changing the model architecture or the dataset. Which two configuration changes should they apply? (Choose two.)

Select 2 answers
A.Expand the training dataset with additional synthetic examples.
B.Reduce the maximum sequence length by truncating all samples to 64 tokens.
C.Enable gradient checkpointing to recompute activations during the backward pass.
D.Increase the number of attention heads in each transformer block.
E.Use mixed precision with bfloat16 for forward and backward passes.
AnswersC, E

Gradient checkpointing stores only selected activations and recomputes the rest during backpropagation, trading additional compute for substantially lower activation memory. It does not alter the model architecture or dataset, so it fits the stated constraints. This makes it an effective lever for fitting larger models or batch sizes into the same GPU memory budget during NeMo fine-tuning runs.

Why this answer

Gradient checkpointing and bfloat16 mixed precision both reduce memory during training without modifying model architecture or dataset. Checkpointing lowers stored activation memory by recomputation, while bfloat16 halves tensor memory and is supported natively in NeMo. Architectural changes, data expansion, and sequence truncation either violate the constraints or fail to address memory consumption per training step.

Exam trap

The trap here is treating data-level changes like truncation or augmentation as memory optimizations, when the constraints require configuration-level changes that leave architecture and dataset intact.

40
MCQeasy

A developer needs to fine-tune a 7B LLM for a customer-support chatbot using NVIDIA NeMo. The dataset contains paired instructions and desired responses. Which data format should be used to prepare the dataset for supervised fine-tuning?

A.A YAML configuration file listing hyperparameters and dataset paths only.
B.A pickle file containing a Python list of tokenized integer sequences with no text labels.
C.A JSONL file where each line contains an input text field and an output text field representing the instruction-response pair.
D.A CSV file containing only the raw customer queries with no corresponding responses.
AnswerC

Supervised fine-tuning in NeMo expects paired examples that map an input prompt to a target completion. A JSONL file with input and output fields per line matches that structure and can be consumed by the NeMo data preprocessing pipeline. This format directly supports the instruction-response objective used for chatbot behavior, making it the appropriate choice for this scenario.

Why this answer

Supervised fine-tuning trains a model to produce a target response given an input instruction. The dataset must therefore contain aligned input-output pairs, and NeMo's data pipeline accepts JSONL records with distinct input and output fields. Files lacking response labels, pre-tokenized sequences without pairing, or pure configuration files cannot supply the supervision signal the training objective requires.

Exam trap

The trap here is confusing the training configuration file with the training dataset, or assuming tokenized sequences alone are sufficient without input-output pairing.

41
MCQmedium

Why is gradient checkpointing useful when fine-tuning a model on a single GPU?

A.It speeds up the training process by calculating gradients in parallel
B.It saves GPU memory by recomputing activations during the backward pass
C.It automatically adjusts the learning rate for each layer
D.It prevents the model from using the CPU during training
AnswerB

By discarding intermediate activations and recomputing them on-the-fly, the memory footprint is significantly reduced. This allows the model to process larger sequences or larger batches that would normally trigger an out-of-memory error. This trade-off is critical for fine-tuning large models on consumer-grade NVIDIA hardware with limited VRAM capacity.

Why this answer

Gradient checkpointing trades computation for memory by not storing all intermediate activations during the forward pass. Instead, it recomputes them during the backward pass. This is extremely valuable when working on hardware with limited VRAM, as it allows for larger batch sizes or longer sequences that would otherwise cause OOM errors, though it does increase the total training time slightly due to the recomputation overhead.

Exam trap

Candidates often assume gradient checkpointing increases memory efficiency by using compression, rather than correctly identifying it as a trade-off that saves memory by recomputing activations during the backward pass.

42
MCQhard

Refer to the exhibit. What is the effective batch size for this fine-tuning job?

A.1
B.15
C.16
D.17
AnswerC

The effective batch size is the product of the per_device_train_batch_size (1) and the gradient_accumulation_steps (16). This results in an effective batch size of 16, which helps in stabilizing training by providing a more representative gradient estimate over multiple mini-batches before weights are actually updated in the optimizer.

Why this answer

The effective batch size is calculated by multiplying the per-device batch size by the gradient accumulation steps. In this case, 1 * 16 equals 16. Gradient accumulation allows the model to simulate a larger batch size by running multiple forward and backward passes before performing a single optimizer step, which helps in achieving more stable gradient updates without increasing the immediate memory overhead for activations.

Exam trap

Candidates frequently add or subtract batch size parameters instead of multiplying per-device batch size by gradient accumulation steps to find the effective batch size.

43
MCQmedium

A company wants to fine-tune a 70B LLM to follow domain-specific instructions. The base model already performs well on general language tasks. They have limited labeled data and limited GPU memory. Which approach is most appropriate?

A.Retrain the tokenizer on the domain corpus and continue pretraining from scratch.
B.Parameter-Efficient Fine-Tuning with LoRA adapters targeting attention projection layers.
C.Full-parameter fine-tuning of all 70B weights on the domain dataset.
D.Prompt engineering with few-shot examples and no weight updates.
AnswerB

LoRA freezes the base weights and trains small low-rank adapter matrices, dramatically reducing memory and optimizer state requirements. It is well suited to limited labeled data because fewer trainable parameters reduce overfitting risk, and it preserves the base model's general abilities. This matches the scenario's constraints on memory and data volume while still adapting the model to domain instructions.

Why this answer

LoRA is the standard parameter-efficient method for adapting very large models under memory constraints, since it trains only small adapter matrices while freezing base weights. With limited labeled data, the reduced trainable parameter count also lowers overfitting risk and helps retain general capabilities. Full fine-tuning, tokenizer retraining, and prompt-only approaches either exceed resource limits or fail to deliver durable domain adaptation.

Exam trap

The trap here is assuming that the largest model change yields the best domain adaptation, when memory and data limits make parameter-efficient adaptation the correct engineering choice.

44
Multi-Selecthard

Which THREE factors significantly influence the memory consumption during LLM fine-tuning? (Choose three)

Select 3 answers
A.The precision of the optimizer states
B.The number of trainable parameters
C.The activation maps stored for the backward pass
D.The number of CPU cores available for data loading
E.The version of the operating system kernel
AnswersA, B, C

The optimizer (e.g., AdamW) stores states for every trainable parameter. If using 32-bit precision, these states can occupy 8 bytes per parameter. Reducing this precision or utilizing paged techniques is essential for saving VRAM, as optimizer states are one of the largest contributors to memory exhaustion during the training cycle.

Why this answer

Memory consumption is driven by the model parameters, the optimizer states (which are often larger than the model weights), and the activations generated during the forward pass. Efficient management of these components is critical for scaling to larger models. By optimizing these three areas, practitioners can successfully fit large models into limited VRAM, ensuring stability throughout the fine-tuning process.

Exam trap

Candidates mistakenly select dataset size or sequence length as primary hardware memory drivers, overlooking how optimizer states, trainable parameters, and activation maps dictate fine-tuning VRAM consumption.

45
MCQmedium

A team is fine-tuning a 13B-parameter Llama model with NVIDIA NeMo Framework on a single A100 80GB GPU. They apply LoRA adapters to the attention projection layers, but the adapters are producing negligible changes to model behavior even after several epochs, and the loss curve stays flat. They confirm the dataset is clean and the tokenizer is correct. Which LoRA configuration issue is the most likely cause?

A.The optimizer was configured with a momentum value of zero, preventing the adapter weights from accumulating updates.
B.The LoRA target modules were set to the embedding layer rather than the attention projection layers.
C.The base model weights were accidentally left trainable, so gradients are flowing into the full model instead of the adapters.
D.The LoRA alpha scaling value is set extremely low relative to the rank, so the adapter's effective update magnitude is near zero.
AnswerD

The LoRA scaling factor is alpha divided by rank. If alpha is tiny compared to rank (for example alpha=1 with rank=64), the adapter's contribution to the forward pass is scaled to a negligible magnitude, so the frozen base weights dominate and the loss barely moves. Raising alpha or lowering rank restores a meaningful update magnitude.

Why this answer

LoRA updates are scaled by alpha divided by rank, so an alpha that is very small relative to rank shrinks the adapter's effective contribution and can leave the loss nearly unchanged. Correcting the alpha-to-rank ratio restores meaningful adapter influence, allowing the fine-tune to actually shift model behavior on the target task.

Exam trap

The trap here is assuming any LoRA misconfiguration equally explains a flat loss, when the alpha-to-rank scaling ratio specifically controls how much the adapter can influence the forward pass.

Ready to test yourself?

Try a timed practice session using only Fine-Tuning questions.