Courseiva

CCNA Core Ml Ai Knowledge Questions

75 of 82 questions · Page 1/2 · Core Ml Ai Knowledge topic · Answers revealed

1
MCQeasy

A data scientist is preparing a labeled dataset of 50,000 customer support tickets for supervised fine-tuning of an LLM. Each ticket must be assigned exactly one of eight department labels. Which loss function is most appropriate for training this classification head?

A.Categorical cross-entropy loss
B.Mean squared error loss
C.Contrastive loss
D.Binary cross-entropy loss
AnswerA

Categorical cross-entropy compares the predicted probability distribution over the eight mutually exclusive department labels with the one-hot true label, penalizing probability mass placed on incorrect classes. It is the standard objective for single-label multi-class classification and produces well-calibrated softmax outputs, making it the right choice when each ticket belongs to exactly one department.

Why this answer

Because every support ticket carries exactly one of eight mutually exclusive department labels, the task is single-label multi-class classification. Categorical cross-entropy, paired with a softmax output layer, directly maximizes the probability of the correct department while normalizing across all eight classes, giving the strongest and most stable training signal for this scenario.

Exam trap

The trap here is confusing single-label multi-class classification, which uses categorical cross-entropy, with multi-label tagging, which uses binary cross-entropy.

2
MCQmedium

A data scientist is training a transformer model and observes that the training loss is decreasing while the validation loss is increasing. Which technique should be prioritized to address this specific generalization challenge?

A.Increase the depth of the transformer architecture layers.
B.Implement L2 regularization (weight decay) on the model weights.
C.Reduce the batch size to increase stochasticity during training.
D.Decrease the number of training epochs to save time.
AnswerB

L2 regularization penalizes large weights by adding a cost term proportional to the square of the magnitude of coefficients. This forces the model to learn simpler patterns, significantly reducing the variance of the model and effectively mitigating overfitting, which is the primary cause of the divergent loss curves described.

Why this answer

This scenario indicates overfitting, where the model captures noise in the training set rather than the underlying distribution. Regularization techniques like Dropout, weight decay, or early stopping are essential to improve model generalization. By limiting the model's ability to memorize the training data, these methods ensure that the weights remain small and the model learns robust patterns that perform better on unseen, held-out validation datasets.

Exam trap

Candidates often suggest increasing model size or training epochs to fix the loss gap, failing to recognize that these actions typically exacerbate overfitting rather than solving the underlying generalization problem.

3
Multi-Selecthard

A team is evaluating a generative large language model for a customer support chatbot. They need to ensure the model produces factually accurate and contextually appropriate responses while avoiding harmful or biased outputs. Which two techniques are most effective for aligning the model's behavior with these safety and quality requirements? (Choose two.)

Select 2 answers
A.Using a larger batch size during pretraining
B.Supervised fine-tuning on curated, high-quality demonstrations
C.Reducing the model's parameter count
D.Increasing the model's temperature during inference
E.Reinforcement learning from human feedback (RLHF)
AnswersB, E

Supervised fine-tuning on carefully curated examples of desired responses teaches the model to mimic safe, accurate, and contextually appropriate behavior. It provides direct signal on how to handle sensitive topics and maintain factual grounding. Combined with human oversight, it is a core alignment method that shapes the model's output distribution toward the team's requirements.

Why this answer

RLHF and supervised fine-tuning on curated demonstrations are central alignment techniques. RLHF uses human preferences to optimize for safety and helpfulness, while supervised fine-tuning directly teaches desired responses. Temperature, parameter count, and batch size do not specifically improve factual accuracy or reduce harmful outputs.

Exam trap

The trap here is assuming that scaling model size or tuning inference randomness can enforce safety, when alignment requires targeted training with human feedback or curated data.

4
MCQmedium

A team is deploying a large language model for real-time text generation. They observe that the model sometimes produces repetitive and dull outputs, especially when generating longer sequences. They want to encourage more diverse and creative text without significantly degrading coherence. Which decoding strategy should they consider?

A.Temperature scaling with a low temperature
B.Beam search
C.Greedy search
D.Top-k sampling
AnswerD

Top-k sampling restricts the next token choices to the k most likely tokens and samples from them, introducing randomness while avoiding very low-probability tokens. This promotes diversity and reduces repetition. In this scenario, it balances creativity and coherence better than deterministic methods, making it suitable for real-time generation.

Why this answer

Top-k sampling introduces controlled randomness by sampling from the top k tokens, which increases diversity and reduces repetition. It avoids the pitfalls of greedy and beam search, which are deterministic and prone to dull outputs, while preventing the incoherence that can come from sampling the entire distribution.

Exam trap

The trap here is thinking that beam search or low temperature improves creativity, when they actually make outputs more deterministic and repetitive.

5
MCQmedium

Refer to the exhibit. Which technique is most effective for preventing the reported NaN error during model training?

A.Gradient Clipping
B.Using a larger batch size
C.Enabling Layer Normalization
D.Changing the optimizer to SGD
AnswerA

Gradient clipping involves re-scaling gradients if their norm exceeds a pre-defined threshold. By capping the update magnitude, it prevents the weights from exploding into non-finite numbers when the loss surface is steep. This is a critical stability measure for training large transformers and deep networks on high-performance compute clusters like NVIDIA DGX.

Why this answer

The exhibit shows a classic case of exploding gradients, where the gradient norm spikes and leads to non-finite weight values. Gradient clipping is the standard industry practice to mitigate this, as it caps the gradient magnitude during backpropagation. Implementing this ensures stability in deep networks, preventing training collapses that waste expensive GPU compute resources and time in large-scale AI development cycles.

Exam trap

Candidates often try to resolve NaN errors by lowering the learning rate or increasing batch size, which are indirect fixes, instead of applying gradient clipping to explicitly prevent exploding gradients.

6
MCQmedium

A researcher is training a large language model and notices the training loss plateaus early while validation loss increases. What is the most likely cause, and which action should be taken?

A.The model is underfitting; increase the number of hidden layers.
B.The learning rate is too low; increase it to accelerate convergence.
C.The model is overfitting; apply dropout or weight decay.
D.The dataset is too small; reduce the batch size.
AnswerC

Overfitting occurs when the model complexity exceeds the information content in the training set. Dropout randomly disables neurons during training, preventing co-adaptation, while weight decay penalizes large weights. These techniques effectively reduce the variance of the model, forcing it to focus on generalized representations instead of training noise.

Why this answer

The symptoms described clearly indicate overfitting, where the model captures noise in the training set rather than generalizing to unseen data. In the context of large language models, this is a critical challenge. Implementing regularization techniques such as weight decay or dropout helps constrain model complexity, forcing it to learn more robust features rather than memorizing specific patterns, thereby improving overall model performance and generalizability.

Exam trap

Candidates often misdiagnose increasing validation loss alongside a plateauing training loss as underfitting or a need for a higher learning rate, instead of recognizing classic overfitting.

7
MCQmedium

A developer is building a retrieval-augmented generation (RAG) system for an internal knowledge base. The system must answer questions using company documents that are updated frequently. Which component is primarily responsible for retrieving the most relevant document chunks to include in the LLM's context?

A.The fine-tuning dataset used to adapt the LLM.
B.The vector database and its similarity search.
C.The LLM's attention mechanism.
D.The tokenizer used to preprocess the query.
AnswerB

In a RAG system, documents are chunked and embedded into vectors stored in a vector database. When a query arrives, its embedding is compared to stored vectors using similarity search to retrieve the most relevant chunks. This retrieval step is what supplies the LLM with grounded context, making the vector database and its search the primary responsible component.

Why this answer

Retrieval-augmented generation separates knowledge retrieval from generation. The vector database stores embeddings of document chunks, and a similarity search matches the query embedding to the most relevant chunks. Those chunks are then inserted into the LLM's context.

This design allows the knowledge base to be updated independently of the model, which is essential for frequently changing internal documents.

Exam trap

The trap here is attributing retrieval to the LLM's attention or fine-tuning, when in a RAG architecture the vector database and similarity search perform the actual document selection.

8
MCQhard

When utilizing Pipeline Parallelism (PP) in LLM training, what is the 'pipeline bubble' and how is it minimized?

A.Idle GPU time during stage-to-stage communication.
B.Memory overflow caused by large context windows.
C.The overhead of gradient accumulation steps.
D.The delay caused by slow NVLink interconnects.
AnswerA

The pipeline bubble represents the period during the start and end of a forward/backward pass where some GPUs are waiting for work because they are dependent on the output of previous pipeline stages. Micro-batching ensures these stages are filled with more tasks, reducing the duration of this unproductive idle time.

Why this answer

The pipeline bubble is the idle time GPUs spend waiting for activations or gradients to propagate across the pipeline stages. It is minimized using techniques like Micro-batching, which breaks a single global batch into smaller units. By interleaving these units, the GPU stages can remain active more consistently, significantly increasing the pipeline's overall utilization and throughput, which is essential for scaling models that cannot fit on a single GPU's memory.

Exam trap

Candidates often confuse pipeline bubbles with tensor parallelism communication overhead or data parallelism gradient synchronization delays, failing to recognize pipeline-specific idle GPU time.

9
MCQmedium

Which THREE techniques are commonly used to improve the efficiency of inference for large language models?

A.Weight Quantization to reduce the bit-width of model parameters.
B.KV-Caching to store previously computed tokens during generation.
C.Increasing the number of transformer layers to add depth.
D.Weight Pruning to remove redundant connections in the network.
E.Replacing all activations with Sigmoid functions for speed.
AnswerA, B, D

Quantization converts model weights from higher precision (like FP32) to lower precision (INT8 or FP8). This substantially decreases the memory footprint and accelerates inference speed on NVIDIA GPUs, as hardware can process more low-precision operations in parallel with less bandwidth consumption and power usage.

Why this answer

Inference efficiency is critical for deploying LLMs in real-world production environments. Techniques like quantization reduce model size and memory requirements, KV-caching prevents redundant computation of attention keys and values, and pruning eliminates unnecessary weights. These methods combined allow models to run with significantly lower latency and reduced hardware costs on NVIDIA inference hardware, making high-performance AI more accessible and scalable for diverse enterprise applications.

Exam trap

Candidates often include training-specific techniques like gradient clipping or data augmentation, failing to recognize that the question specifically asks for methods targeting inference efficiency.

10
Multi-Selectmedium

A machine learning engineer is evaluating a fine-tuned LLM for a customer-facing summarization task. The model produces fluent summaries, but the team needs to detect when the model generates content that is not supported by the source document. Which TWO evaluation approaches are appropriate for measuring factual consistency between the generated summary and the source? (Choose two.)

Select 2 answers
A.Compute BLEU score between the generated summary and a single human-written reference summary.
B.Calculate ROUGE-L recall against the source document.
C.Have human annotators label each summary sentence as supported or not supported by the source document.
D.Measure perplexity of the generated summary under the fine-tuned model.
E.Use an entailment-based metric such as FactCC or a natural language inference model to check whether the summary is entailed by the source.
AnswersC, E

Human annotation directly assesses whether each claim in the summary is grounded in the source. Annotators can catch subtle fabrications that automated metrics miss. While more expensive, it provides a reliable gold standard for factual consistency and is often used to validate automated metrics in production settings.

Why this answer

Factual consistency requires checking whether the summary's claims are supported by the source document. Entailment-based metrics like FactCC or NLI models directly test this relationship, and human annotation provides a reliable ground truth. Metrics like BLEU, ROUGE, or perplexity measure overlap or fluency and can be fooled by fluent hallucinations, so they are not appropriate for this specific goal.

Exam trap

The trap here is choosing reference-overlap metrics like BLEU or ROUGE because they are familiar, even though they do not verify that generated content is grounded in the source document.

11
MCQmedium

Refer to the exhibit. Which technique is most appropriate to prevent this specific training failure?

A.Increase the learning rate to bypass the threshold.
B.Apply gradient clipping during the backpropagation step.
C.Remove the normalization layers to increase model capacity.
D.Change the optimizer from Adam to SGD without momentum.
AnswerB

Gradient clipping scales the gradient vector if its norm exceeds a defined threshold, ensuring the updates remain within a stable range. This prevents extreme weight changes that cause numerical overflow, effectively stopping the NaN divergence observed in the logs while allowing the training process to continue successfully.

Why this answer

The exhibit shows a classic case of the exploding gradient problem, where large weight updates lead to numerical instability and NaN loss values. This is common in deep architectures, particularly RNNs or deep Transformers. Gradient clipping is the standard solution to constrain the norm of the gradients during backpropagation, ensuring that updates remain within a numerically stable range, thereby preventing the model parameters from reaching infinity or NaN states.

Exam trap

Candidates often confuse exploding gradients with vanishing gradients, incorrectly suggesting that increasing the learning rate or adding deeper layers will resolve numerical instability.

12
MCQhard

Which component of an NVIDIA Transformer Engine is specifically designed to accelerate training on supported GPUs by dynamically adjusting precision?

A.Tensor Cores
B.Transformer Engine
C.CUDA Streams
D.NCCL (NVIDIA Collective Communications Library)
AnswerB

The Transformer Engine is specifically built for NVIDIA H100 and newer GPUs to provide FP8 support. It monitors the distribution of activations and gradients to dynamically scale precision, ensuring that the model maintains high throughput without sacrificing accuracy. It is the key component for optimizing modern transformer training performance.

Why this answer

The NVIDIA Transformer Engine uses FP8 precision and dynamic scaling to drastically accelerate training for transformer-based models. By utilizing hardware-level support for mixed-precision, it maximizes throughput while maintaining accuracy. Understanding how these hardware-software co-optimizations function is critical for modern LLM development, as they allow for training larger models in less time and with lower power usage compared to standard FP32 or mixed-precision training workflows.

Exam trap

Candidates often select general CUDA or TensorRT frameworks, missing that the NVIDIA Transformer Engine is uniquely engineered to dynamically adjust precision for transformer architectures.

13
MCQeasy

Which of the following activation functions is most commonly used in hidden layers of deep neural networks to mitigate the vanishing gradient problem?

A.Sigmoid
B.Tanh
C.ReLU
D.Linear
AnswerC

ReLU outputs the input directly if it is positive and zero otherwise. Its derivative is 1 for positive inputs, which allows gradients to flow through the network without being multiplied by small values. This property is key to training deep architectures efficiently, as it drastically reduces the vanishing gradient problem.

Why this answer

The Rectified Linear Unit (ReLU) is the standard activation function for deep networks. Unlike Sigmoid or Tanh, which saturate at high and low values, ReLU maintains a constant gradient of 1 for all positive inputs. This effectively prevents the gradient from vanishing during backpropagation across many layers, allowing for the training of much deeper architectures without needing complex initialization or normalization schemes.

Exam trap

Candidates mistakenly select Sigmoid or Tanh, forgetting that these functions saturate at extreme values, which causes the vanishing gradient problem. They fail to recall that ReLU is specifically designed to prevent this.

14
MCQeasy

A data science team at a retail company is building a neural network to predict customer churn from tabular data with mixed numerical and categorical features. They want the model to output a probability between 0 and 1, and they are training with a standard gradient descent optimizer. Which loss function is most appropriate for this binary classification task?

A.Binary cross-entropy loss
B.Categorical cross-entropy loss
C.Mean squared error loss
D.Hinge loss
AnswerA

Binary cross-entropy measures the dissimilarity between the predicted probability and the true binary label, providing well-behaved gradients for logistic outputs. It penalizes confident wrong predictions heavily, which suits churn prediction where calibrated probabilities matter. With a sigmoid output unit, its derivative simplifies to the prediction error, making optimization stable and efficient for this scenario.

Why this answer

Binary cross-entropy is the standard loss for binary classification with a sigmoid output. It directly optimizes the log-likelihood of the correct class, yielding strong gradients when predictions are wrong and well-calibrated probabilities. The other losses either assume multiclass targets, produce non-probabilistic outputs, or suffer from vanishing gradients with sigmoid units.

Exam trap

The trap here is assuming mean squared error is universally safe, when its gradient with a sigmoid output becomes vanishingly small for confident mistakes.

15
Multi-Selecthard

A research team is designing a decoder-only transformer LLM for long-document question answering. They want to reduce the quadratic computational cost of self-attention so that training on sequences of 32,000 tokens is feasible on their GPU cluster. Which two techniques are appropriate for this goal? (Choose two.)

Select 2 answers
A.Using a larger vocabulary size for the tokenizer
B.Sparse attention patterns that restrict each token to a subset of positions
C.Replacing layer normalization with batch normalization in the transformer blocks
D.FlashAttention-style IO-aware exact attention kernels
E.Increasing the number of attention heads while keeping the same total hidden dimension
AnswersB, D

Sparse attention reduces the number of query-key interactions from quadratic to near-linear by having each token attend to a limited set of positions, such as a local window plus selected global tokens. This directly lowers compute and memory for long sequences. It is a standard approach for making 32,000-token training feasible while retaining much of the model's ability to capture long-range dependencies.

Why this answer

Sparse attention lowers the number of attention interactions, and IO-aware exact attention kernels like FlashAttention reduce memory traffic and speed up the exact computation. Both directly target the quadratic cost of self-attention for long sequences. Increasing heads, enlarging the vocabulary, or swapping normalization layers does not change the sequence-length-squared scaling and therefore does not make 32,000-token training feasible.

Exam trap

The trap here is confusing general model-size or architecture changes, such as more heads or a larger vocabulary, with techniques that actually reduce the quadratic scaling of attention over sequence length.

16
MCQmedium

Which activation function is most commonly used in the hidden layers of deep neural networks to mitigate the vanishing gradient problem?

A.Sigmoid
B.ReLU
C.Softmax
D.Hyperbolic Tangent (tanh)
AnswerB

ReLU provides a constant gradient of 1 for all positive inputs, preventing the gradient from diminishing during backpropagation. This simple, non-saturating nature makes it highly effective for training very deep networks. It is computationally efficient and has been a cornerstone of deep learning success across various computer vision and language tasks.

Why this answer

The ReLU (Rectified Linear Unit) activation function is the industry standard for hidden layers because it avoids the saturation characteristic of sigmoid or tanh functions. By outputting zero for negative inputs and linear values for positive inputs, it maintains a gradient of one during backpropagation for active neurons. This allows gradients to flow through deep networks without shrinking exponentially, which is essential for training modern, deep architectures effectively.

Exam trap

Candidates confuse hidden-layer activation requirements with output-layer needs, incorrectly selecting sigmoid or softmax for deep hidden layers.

17
MCQmedium

An AI researcher is fine-tuning a large language model and wants to minimize GPU memory consumption during training without altering the model's primary weight representations or introducing quantization error during inference. Which technique provides this capability by decomposing weight matrices into low-rank trainable adaptation matrices?

A.Post-Training Quantization (PTQ) converting weights from FP16 to INT8 integer representations.
B.Low-Rank Adaptation (LoRA) freezing pre-trained weights and training injected rank decomposition matrices.
C.Full-Parameter Fine-Tuning updating all weights across every transformer layer simultaneously using FP32 precision.
D.Knowledge Distillation transferring learned representations from a large teacher model to a smaller student network architecture.
AnswerB

LoRA freezes the pre-trained weights and injects trainable rank decomposition matrices, so only these small adapters update during fine-tuning. This cuts GPU memory use substantially while leaving the original weight representations untouched, and because the base weights stay full precision, no quantization error is introduced at inference.

Why this answer

Low-Rank Adaptation (LoRA) freezes the original pre-trained model weights and injects trainable rank decomposition matrices into the architecture. This drastically reduces the number of trainable parameters and optimizer states in GPU memory during backpropagation, while allowing weights to be merged back into the original matrices for zero-latency deployment.

Exam trap

Candidates frequently confuse LoRA with post-training quantization methods like INT4 or INT8. While quantization compresses weights for inference, LoRA is an efficient parameter-efficient fine-tuning technique that preserves original precision weight matrices.

18
MCQmedium

A machine learning engineer is training a deep neural network for image classification. They notice that the training loss decreases steadily, but the validation loss starts to increase after a few epochs. Which technique is most directly aimed at addressing this issue?

A.Increasing the learning rate
B.Removing batch normalization
C.Reducing the size of the training set
D.Adding dropout layers
AnswerD

Dropout randomly deactivates neurons during training, which forces the network to learn more robust features and reduces co-adaptation. This regularization technique directly combats overfitting, where validation loss rises while training loss falls. By preventing the model from relying too heavily on any single neuron, dropout improves generalization to unseen data.

Why this answer

The described behavior—training loss decreasing while validation loss increases—is a classic sign of overfitting. Dropout is a regularization method that randomly drops units during training, which reduces the network's capacity to memorize noise and encourages more generalizable representations. Other options either do not address overfitting or would exacerbate the problem.

Exam trap

The trap here is thinking that a larger training set or more complex model is always better, but when validation loss rises, regularization like dropout is needed.

19
MCQmedium

In the context of transformer models, what is the purpose of the 'Attention Mask' during the training process?

A.It serves to reduce the number of parameters in the self-attention layer.
B.It prevents the model from attending to future tokens in autoregressive models.
C.It optimizes the data movement between the GPU's L1 and L2 cache.
D.It performs weight pruning to compress the model size after training.
AnswerB

In autoregressive LLMs, the model must predict the next token based only on previous ones. The attention mask sets the attention scores for future positions to negative infinity before softmax, effectively nullifying their influence. This ensures the model learns causal dependencies during its training phase.

Why this answer

The attention mask is essential for transformer architectures to handle sequences of varying lengths and to enforce causal constraints. In tasks like sequence generation, it prevents the model from 'peeking' at future tokens, ensuring that the prediction at each position depends only on preceding tokens. This mechanism is critical for maintaining the autoregressive nature of models like GPT and ensuring correct model training.

Exam trap

Candidates frequently mistake the attention mask for padding management used to handle variable sequence lengths, overlooking its critical causal role in preventing autoregressive models from seeing future tokens.

20
MCQhard

An ML engineer is training a transformer-based language model on a single NVIDIA A100 GPU. They observe that the training loss decreases initially but then becomes NaN after a few hundred steps. The learning rate is 1e-4, and mixed precision with FP16 is enabled. Which action is most likely to stabilize training while preserving the benefits of mixed precision?

A.Switch to FP32 training for all layers.
B.Reduce the batch size to 1 to minimize memory usage.
C.Enable dynamic loss scaling to automatically adjust the loss scale factor during training.
D.Increase the learning rate to 1e-3 to escape the NaN region faster.
AnswerC

Dynamic loss scaling multiplies the loss by a large factor to prevent small gradients from underflowing in FP16, and it automatically reduces the scale when overflows (NaNs) are detected. This directly addresses the NaN issue while keeping FP16 computation, thus preserving mixed precision benefits and stabilizing training.

Why this answer

Dynamic loss scaling is the standard technique to prevent underflow and overflow in FP16 mixed precision training. It scales the loss up to keep gradients in representable range and automatically backs off when NaNs occur. This stabilizes training while retaining the performance and memory advantages of mixed precision, unlike switching to FP32 or altering hyperparameters.

Exam trap

The trap here is assuming that any change to training hyperparameters or precision will fix NaNs, when the specific cause in mixed precision is often gradient underflow that requires loss scaling.

21
MCQhard

Refer to the exhibit. A machine learning engineer reviews the monitoring output from a two-GPU distributed training job running on an NVIDIA DGX system. GPU 0 shows low utilization despite high memory consumption, while GPU 1 shows high utilization and high memory consumption. What is the most likely root cause of this performance imbalance?

A.A PCIe bus bottleneck causing continuous memory swapping between host system RAM and GPU 0 device memory.
B.Imbalanced data distribution or pipeline synchronization stalls causing GPU 0 to wait idly while holding allocated tensors.
C.A hardware fault in the NVIDIA NVLink bridge connecting GPU 0 to the shared system NVSwitch fabric.
D.An incorrect CUDA driver version mismatch preventing GPU 0 from initializing its primary Tensor Core execution units.
AnswerB

GPU memory remains fully allocated by framework tensors and optimizer states even when the compute units are idle. An imbalance in batch sizes or pipeline parallelism stages forces GPU 0 to wait at synchronization barriers, driving down its utilization metric.

Why this answer

Data parallelism with uneven batch distribution or pipeline parallel bubble inefficiency can cause worker synchronization stalls. When one GPU finishes its workload or waits for synchronization barriers while processing uneven tensor sizes or asynchronous communication primitives, its utilization drops while memory remains allocated to tensors.

Exam trap

Candidates often assume a hardware failure or driver mismatch when GPU utilization diverges. However, in distributed training, high memory with low utilization almost always indicates pipeline stalls, gradient synchronization bottlenecks, or uneven data loader chunking.

22
MCQhard

A researcher is fine-tuning a large language model on a downstream task with a small dataset. They notice that the model achieves high training accuracy but poor validation accuracy. Which regularization technique is most appropriate to address this issue?

A.Increase the learning rate to speed up convergence.
B.Remove early stopping and train for more epochs.
C.Apply dropout to the transformer layers during fine-tuning.
D.Reduce the size of the training dataset further.
AnswerC

Dropout randomly deactivates neurons during training, preventing the model from relying too heavily on specific features and reducing overfitting. In transformer fine-tuning, applying dropout to attention and feed-forward layers is a standard regularization method. It improves generalization on small datasets by encouraging more robust representations.

Why this answer

Overfitting on a small dataset is best addressed by regularization methods like dropout, which introduce noise during training and force the model to learn more generalizable patterns. Dropout is particularly effective in transformers and is easy to apply during fine-tuning. Increasing learning rate, reducing data, or removing early stopping would all worsen the problem.

Exam trap

The trap here is confusing overfitting with underfitting; high training accuracy and low validation accuracy clearly indicate overfitting, so techniques that increase model capacity or training time are counterproductive.

23
MCQhard

A machine learning engineer is fine-tuning a pre-trained language model on a small domain-specific dataset. She notices that the model quickly achieves high accuracy on the training set but performs poorly on the validation set. She wants to mitigate this overfitting without collecting more data. Which technique is most appropriate?

A.Increase the learning rate
B.Apply L2 regularization
C.Train for more epochs
D.Use a larger batch size
AnswerB

L2 regularization adds a penalty term to the loss function proportional to the square of the weights, discouraging large weights and reducing overfitting. In fine-tuning on a small dataset, it helps the model generalize better by preventing it from fitting noise. This is a standard and effective technique when additional data is unavailable.

Why this answer

L2 regularization is a classic method to combat overfitting by penalizing large weights, which encourages the model to learn simpler patterns that generalize better. When fine-tuning on a small dataset, it is particularly effective because it constrains the model's capacity to memorize noise, improving validation performance.

Exam trap

The trap here is thinking that more training or a larger batch size can fix overfitting, when they often exacerbate it; regularization techniques like L2 are needed.

24
MCQeasy

A data scientist is preprocessing a text corpus to train a large language model. They want to convert each word into a dense vector representation that captures semantic relationships before feeding it into the transformer. Which technique should they use?

A.TF-IDF vectorization
B.Word embeddings (e.g., Word2Vec, GloVe)
C.Tokenization
D.One-hot encoding
AnswerB

Word embeddings like Word2Vec or GloVe generate dense, low-dimensional vectors where semantically similar words are close in vector space. They capture relationships such as king - man + woman ≈ queen, making them ideal for initializing the embedding layer of a transformer-based language model.

Why this answer

Word embeddings such as Word2Vec or GloVe convert words into dense vectors that encode semantic relationships, which is essential for language models to understand meaning. Tokenization is only a splitting step, while one-hot and TF-IDF yield sparse vectors lacking semantic depth. Thus, embeddings are the correct choice for capturing semantics before transformer processing.

Exam trap

The trap here is confusing tokenization with embedding, assuming that splitting text into tokens automatically provides semantic vectors.

25
MCQmedium

A machine learning engineer is training a large language model on a cluster of NVIDIA GPUs. During training, she observes that the loss occasionally spikes to NaN, causing the training to fail. She suspects that the issue is related to the numerical precision of the computations. Which technique is most appropriate to mitigate this issue while maintaining training stability?

A.Switch to double precision
B.Increase the learning rate
C.Reduce the batch size
D.Use gradient clipping
AnswerD

Gradient clipping limits the magnitude of gradients during backpropagation, preventing explosive gradients that can cause numerical overflow and NaN losses. In training large language models, especially with mixed precision, gradient clipping is a standard technique to maintain stability. It directly addresses the symptom of loss spikes by keeping updates within a manageable range.

Why this answer

Gradient clipping is a widely used technique to prevent exploding gradients, which are a common cause of NaN losses in deep learning, especially when training large models with mixed precision. By capping gradient norms, it ensures that parameter updates remain stable, allowing training to proceed without numerical overflow.

Exam trap

The trap here is assuming that reducing batch size or increasing learning rate can fix NaN losses, when the root cause is often gradient explosion that gradient clipping directly addresses.

26
MCQhard

A team is deploying a large language model for real-time inference on an NVIDIA GPU. They observe that the first few inference requests have high latency, but subsequent requests are much faster. What is the most likely explanation for this behavior?

A.The model is being quantized on the fly
B.The GPU is warming up its clock speed
C.CUDA kernels are being compiled and cached
D.The model weights are being loaded from disk
AnswerC

During the first inference, CUDA kernels for operations like matrix multiplications are compiled and cached. This just-in-time compilation adds latency. Subsequent requests reuse the cached kernels, avoiding recompilation and resulting in faster execution. This is a common behavior in frameworks like PyTorch and TensorRT.

Why this answer

The initial high latency is due to just-in-time compilation of CUDA kernels, which occurs on the first execution of each operation. Once compiled, the kernels are cached for reuse, making subsequent inferences faster. This is a well-known behavior in deep learning frameworks and explains the warm-up effect.

Exam trap

The trap here is attributing the latency to hardware warm-up or model loading, which are one-time startup costs, rather than to runtime kernel compilation.

27
MCQmedium

An enterprise machine learning team is training a large-scale transformer model on a cluster of NVIDIA A100 GPUs using mixed precision (FP16). During the initial training phase, the team notices sudden numerical underflow resulting in vanishing gradients and stalled loss convergence. Which optimization technique must be applied to mitigate this issue without sacrificing the memory-efficiency benefits of FP16?

A.Implementing gradient clipping by global norm to bound large gradient updates before parameter application.
B.Switching the entire training cluster to FP32 precision to guarantee maximum numerical stability and range.
C.Applying dynamic loss scaling to the loss value prior to backpropagation to shift gradients into the representational range.
D.Increasing the mini-batch size exponentially to smooth out noisy gradient estimates across distributed nodes.
AnswerC

Dynamic loss scaling automatically adjusts a multiplier on the loss value to maintain gradient magnitudes within the representational limits of FP16. This prevents underflow during backward passes without requiring full FP32 precision across all network layers.

Why this answer

Loss scaling multiplies the forward pass loss by a scaling factor to shift small gradient magnitudes into the representational range of the FP16 format, preventing underflow. This is essential in NVIDIA mixed-precision training because FP16 has a narrow dynamic range compared to FP32. Proper scaling prevents gradient values from truncating to zero while retaining half-precision throughput and memory footprint advantages on Tensor Cores.

Exam trap

Candidates frequently confuse loss scaling with learning rate warm-up or gradient clipping. While learning rate warm-up stabilizes early training dynamics and gradient clipping prevents exploding gradients, neither addresses the precision representational underflow caused by the limited dynamic range of FP16.

28
MCQeasy

A data science team is building a model to predict whether a customer will churn based on historical account activity. They have a large dataset with labeled outcomes (churned or not churned). Which type of machine learning is most appropriate for this task?

A.Unsupervised learning with clustering
B.Reinforcement learning with a reward function
C.Supervised learning with a classification algorithm
D.Semi-supervised learning with a small labeled set
AnswerC

The dataset contains labeled examples where the outcome (churned or not churned) is known. Supervised learning uses these labels to learn a mapping from input features to the target class. Classification algorithms are specifically designed for discrete outcomes like churn, making this the correct approach for predicting a binary category.

Why this answer

The task is to predict a binary outcome (churn or no churn) using historical data where the outcome is known. This is a classic supervised learning problem, and classification algorithms are designed to learn from labeled examples to predict discrete categories. Unsupervised, reinforcement, and semi-supervised methods do not directly leverage the available labels for prediction.

Exam trap

The trap here is assuming that any large dataset requires unsupervised learning, but the presence of labeled outcomes clearly indicates a supervised classification problem.

29
MCQmedium

When implementing Retrieval-Augmented Generation (RAG), why is the choice of 'Chunk Size' critical for model retrieval performance?

A.It directly limits the number of documents in the vector database.
B.It determines the semantic density and context of retrieved segments.
C.It affects the latency of the embedding model's inference.
D.It is solely determined by the GPU's memory capacity.
AnswerB

Appropriate chunk sizes ensure that retrieved segments contain complete thoughts and sufficient background information. If chunks are too small, they lack context; if too large, they introduce irrelevant information that dilutes the query's focus. Finding the optimal size is essential to help the LLM generate grounded and accurate, high-quality answers.

Why this answer

Chunk size determines how much context is included in a single document segment during the retrieval process. If chunks are too small, the model lacks sufficient context to answer complex queries. If they are too large, the retrieval results contain excessive noise, causing the model to lose focus.

Optimizing this balance is a core task in RAG engineering to ensure that the retrieved information is both relevant and comprehensive enough for the model to generate accurate responses.

Exam trap

Candidates frequently assume that larger chunks are always better because they contain more information, ignoring the trade-off where excessive noise and irrelevant context degrade the model's ability to focus on specific answers.

30
MCQeasy

A developer is building a text summarization assistant that must produce concise, faithful summaries of long support tickets. The team wants to fine-tune a pre-trained large language model on a small labeled dataset of ticket-summary pairs. Which training approach best matches this goal?

A.Supervised fine-tuning (SFT) on the ticket-summary pairs
B.Unsupervised pre-training from scratch on the ticket corpus
C.Reinforcement learning from human feedback (RLHF) using pairwise summary preferences
D.Retrieval-augmented generation (RAG) with no model weight updates
AnswerA

SFT directly optimizes the model on input-output pairs where the ticket is the prompt and the desired summary is the target. Because the team already has labeled summaries, this approach teaches the model the exact task and style required. It is the standard method for adapting a pre-trained LLM to a specific downstream generation task with limited labeled data.

Why this answer

The team has labeled input-output pairs for a specific generation task, which is exactly what supervised fine-tuning is designed for. SFT updates the pre-trained model's weights to map support tickets to concise summaries, leveraging existing language knowledge while learning the task. Other approaches either ignore the labels, require far more data and compute, or do not modify the model to perform the target task.

Exam trap

The trap here is assuming that any method using the ticket data, such as retrieval or pre-training, will teach the model to summarize, when only supervised fine-tuning uses the labeled summary targets directly.

31
Multi-Selectmedium

A machine learning engineer is training a deep neural network and notices that the training loss decreases but the validation loss starts to increase after several epochs. Which two techniques are most appropriate to mitigate this issue? (Choose two.)

Select 2 answers
A.Increase the number of training epochs.
B.Increase the learning rate.
C.Use early stopping based on validation loss.
D.Apply L2 regularization to the model's weights.
E.Add more layers to the neural network.
AnswersC, D

Early stopping monitors validation loss and halts training when it starts to increase, preventing the model from overfitting further. It effectively selects the model at the point of best generalization. This is a simple and widely used technique to combat overfitting without altering the model architecture.

Why this answer

When validation loss increases while training loss decreases, the model is overfitting. L2 regularization penalizes complexity, and early stopping halts training at the optimal point. Both techniques reduce overfitting and improve generalization.

Other options either worsen overfitting or are unrelated to the problem.

Exam trap

The trap here is thinking that more training or a larger model will help, when the issue is actually overfitting.

32
Multi-Selecthard

A team deploys a retrieval-augmented generation pipeline and observes that answers frequently cite facts not present in the retrieved passages. They want to reduce this unsupported generation behavior. (Choose two.)

Select 2 answers
A.Fine-tune the generator on additional unlabeled web text
B.Improve the retriever so that top-ranked passages are more relevant to the query
C.Enlarge the retriever's index to include every document in the enterprise
D.Increase the generator's temperature to diversify its outputs
E.Instruct the generator to answer only from the provided context and to abstain otherwise
AnswersB, E

Unsupported claims often arise when retrieved context is irrelevant or missing the needed evidence, leaving the model to fill gaps from parametric memory. Raising retrieval precision with better embeddings, hybrid search, or reranking ensures the generator receives passages that actually contain the answer, which measurably reduces hallucinated content.

Why this answer

Unsupported generation in retrieval-augmented pipelines stems from two main sources: the generator lacking relevant evidence, and the generator ignoring the evidence it has. Improving retrieval precision supplies the needed facts, while explicit grounding instructions constrain the model to answer only from context and to abstain otherwise. Together they reduce fabricated claims without degrading answer quality.

Exam trap

The trap here is treating hallucination as a creativity problem and raising temperature, when unsupported claims usually come from weak retrieval or an unconstrained generator.

33
MCQmedium

Refer to the exhibit. The training loss is oscillating and failing to converge. What is the most likely immediate adjustment needed?

A.Reduce the learning rate
B.Increase the batch size to 1024
C.Switch to a larger model architecture
D.Switch to FP16 mixed-precision
AnswerA

An oscillating loss is a classic symptom of an overly aggressive learning rate. By reducing the learning rate, the optimizer makes smaller updates to the weights, preventing the model from 'bouncing' around the loss landscape. This allows for finer adjustments and more stable convergence toward a lower loss value.

Why this answer

The oscillating loss and lack of convergence suggest the learning rate is too high. A learning rate of 0.01 is relatively aggressive for many deep learning tasks, causing the model to jump over the optimal minima. Reducing the learning rate is the standard first step to stabilize the training process, allowing the optimizer to settle into the local minimum more consistently and achieve better overall convergence results.

Exam trap

Many candidates assume that an oscillating loss requires increasing the training epochs or batch size, overlooking the primary symptom of an excessively high learning rate.

34
MCQmedium

Refer to the exhibit. The configuration shows the use of FSDP with mixed precision. What is the main benefit of using 'bf16' (Bfloat16) over 'fp16' in this context?

A.It reduces the size of the model weights on the disk.
B.It provides a larger dynamic range, preventing gradient underflow.
C.It is natively supported on older Pascal-based architectures.
D.It doubles the throughput compared to FP32.
AnswerB

BF16 uses an 8-bit exponent field identical to FP32, which allows it to represent a much wider range of values than FP16. This prevents the underflow of small gradient values that often occur in deep learning, significantly simplifying training stability by removing the need for explicit dynamic loss scaling.

Why this answer

Bfloat16 provides the same dynamic range as FP32, preventing the underflow issues common with FP16 when calculating gradients during deep learning training. This stability allows for training without loss-scaling techniques, simplifying the pipeline and improving convergence consistency. In high-performance training, using BF16 is the standard for modern GPUs like the NVIDIA A100/H100, as it ensures stability without sacrificing speed or requiring complex hyperparameter tuning for numeric precision.

Exam trap

Candidates often believe BF16 is chosen for speed gains alone, missing that its primary architectural advantage over FP16 is the larger dynamic range that prevents gradient underflow without complex scaling.

35
MCQeasy

What is the role of 'Temperature' in the context of LLM text generation?

A.It regulates the cooling of the GPU hardware.
B.It adjusts the probability distribution before sampling the next token.
C.It determines the maximum length of the generated sequence.
D.It compresses the model weights to save memory.
AnswerB

Temperature scales the logits before the softmax normalization. Higher values create a softer distribution (more diversity), while lower values create a sharper distribution (more deterministic). This allows the user to control the creativity and predictability of the model's generated text without changing the underlying model weights.

Why this answer

Temperature controls the randomness of the model's output distribution. A low temperature makes the model more confident and deterministic by sharpening the probability distribution, favoring the most likely next token. Conversely, a high temperature flattens the distribution, increasing the likelihood of selecting less probable tokens.

This parameter is crucial for balancing creativity and coherence in generative AI applications, allowing users to fine-tune the model's output behavior.

Exam trap

Candidates often confuse Temperature with repetition penalty or top-k sampling, mistakenly believing it directly alters the model's underlying knowledge base rather than simply adjusting the probability distribution before the final token selection.

36
MCQeasy

Which machine learning paradigm involves an agent learning to make decisions by performing actions in an environment to maximize a cumulative reward?

A.Supervised Learning
B.Unsupervised Learning
C.Reinforcement Learning
D.Transfer Learning
AnswerC

Reinforcement learning is defined by an agent navigating an environment through actions and receiving rewards. The agent aims to learn a policy that maximizes the total long-term reward. This iterative process of exploration and exploitation is a core concept in modern AI, particularly for dynamic and complex decision-making tasks.

Why this answer

Reinforcement Learning (RL) is fundamentally about agents interacting with environments. Unlike supervised learning, which relies on labeled datasets, or unsupervised learning, which finds hidden structures, RL uses a feedback loop of rewards and penalties. This paradigm is crucial for robotics, game playing, and autonomous systems where the optimal sequence of actions is not pre-defined, but must be discovered through trial and error.

Exam trap

Candidates occasionally confuse reinforcement learning with semi-supervised or active learning due to the presence of feedback, missing the core trial-and-error environment interaction loop.

37
MCQmedium

What is the primary function of the 'Attention' mechanism in Transformer models?

A.To reduce the number of parameters in the model
B.To enable the model to weigh the importance of different input tokens
C.To convert text into numerical vectors
D.To perform sequential processing of inputs
AnswerB

Attention allows the model to dynamically compute the importance of each token relative to others in the sequence. By using query, key, and value vectors, the model can focus on specific parts of the context that are relevant for predicting the current token, enabling better semantic understanding of long-range dependencies.

Why this answer

The attention mechanism allows the model to compute weights that signify the relevance of different parts of the input sequence to one another, regardless of their distance. This captures long-range dependencies effectively, which traditional RNNs struggle with due to vanishing gradients over long sequences. Mastering this concept is key to understanding why transformers have become the dominant architecture for nearly all state-of-the-art generative AI and natural language processing applications.

Exam trap

Candidates often confuse attention with memory mechanisms or RNN hidden states, wrongly focusing on sequential processing rather than the mechanism's ability to weigh the importance of tokens across the entire sequence.

38
MCQhard

An ML engineer is deploying a Transformer-based inference service on an NVIDIA TensorRT-LLM runtime. To maximize inference throughput and reduce latency under heavy concurrent user traffic, the engineer needs to select the optimal decoding batching strategy. Which technique allows multiple incoming dynamic sequence requests to be batched together at the token level rather than waiting for entire sequences to complete?

A.Static request batching that groups fixed numbers of incoming inference prompts and pads them to equal sequence lengths.
B.Data-parallel replica sharding that duplicates the model across separate GPU devices to process independent streams.
C.Inflight batching (continuous batching) that schedules individual generation iterations at the token level across requests.
D.Model parallelism that splits individual transformer weight matrices across multiple interconnected GPU devices.
AnswerC

Inflight batching schedules generation at the token level, admitting new requests into the running batch as others finish rather than waiting for whole sequences. This keeps the GPU saturated under concurrent load, maximising throughput and cutting latency compared with static batching.

Why this answer

Inflight batching (also known as continuous batching) dynamically groups individual generation steps from different requests at each token generation iteration. This eliminates the idle padding overhead and waiting times associated with traditional static batching, maximizing GPU compute utilization in LLM serving.

Exam trap

Candidates frequently confuse inflight batching with traditional static request batching or data-parallel batching. Static batching forces early-finishing requests to wait idly for the longest sequence in the batch to complete.

39
MCQeasy

A data scientist is preparing a transformer-based language model for a text summarization task. She notices that the input sequences in her dataset vary widely in length, from a few tokens to several thousand. She decides to set a fixed maximum sequence length and pad shorter sequences with a special token. Which component of the transformer architecture is primarily responsible for handling the positional information of tokens in these sequences?

A.Self-attention mechanism
B.Layer normalization
C.Feed-forward network
D.Positional encodings
AnswerD

Positional encodings are added to token embeddings to inject information about the order of tokens in the sequence. This allows the transformer to distinguish between tokens at different positions, which is crucial for tasks like summarization where word order affects meaning. With varying sequence lengths and padding, positional encodings ensure the model understands the relative positions of tokens.

Why this answer

Positional encodings are essential in transformer models because the self-attention mechanism is permutation-invariant. By adding positional encodings to token embeddings, the model gains awareness of token order, enabling it to handle sequences of varying lengths and padding effectively. This is critical for tasks like summarization where the meaning depends on word order.

Exam trap

The trap here is confusing self-attention's ability to model relationships with its inability to encode order, leading to the misconception that self-attention alone handles positional information.

40
MCQeasy

Which of the following describes the purpose of a validation set in machine learning?

A.To increase the total amount of training data available
B.To tune hyperparameters and prevent overfitting
C.To calculate the final accuracy on unseen production data
D.To perform backpropagation during the model training loop
AnswerB

The validation set allows for the evaluation of model performance on unseen data during training. By observing performance on this set, developers can adjust hyperparameters, such as learning rates or layer sizes, to optimize for generalization. This acts as a guardrail against overfitting, where the model learns the training data too specifically.

Why this answer

The validation set serves as an independent dataset used to tune hyperparameters and monitor model performance during training. It provides an unbiased evaluation of the model's generalization capabilities, allowing engineers to prevent overfitting and select the best model version before final testing. Distinguishing between training, validation, and test data is fundamental to ensuring models perform reliably in production environments, avoiding common pitfalls in AI development.

Exam trap

Candidates frequently confuse the validation set with the test set, mistakenly believing the validation set is used for the final, unbiased performance report rather than for iterative hyperparameter tuning.

41
MCQhard

A machine learning engineer is deploying a transformer-based language model for real-time translation. They observe that inference latency is too high for the required throughput. The model uses standard multi-head self-attention. Which modification is most likely to reduce latency without significantly degrading translation quality?

A.Increase the number of attention heads
B.Use a sparse attention mechanism
C.Increase the model's hidden dimension
D.Add more layers to the transformer
AnswerB

Sparse attention reduces the number of key-value pairs each query attends to, lowering the quadratic complexity of self-attention to near-linear. This directly cuts computation and memory access, decreasing latency. For translation, where local context often suffices, sparse patterns can maintain quality while meeting real-time throughput requirements.

Why this answer

Sparse attention reduces the quadratic cost of self-attention by limiting each token's attention to a subset of positions, which lowers computation and latency. Increasing heads, hidden dimension, or layers all raise computational load and latency. Sparse attention is a targeted optimization for real-time inference while preserving translation quality.

Exam trap

The trap here is thinking more capacity (heads, dimensions, layers) improves performance without considering the latency penalty, when the goal is to reduce latency.

42
MCQmedium

A machine learning engineer is preprocessing a dataset for a generative AI model and wants to ensure that the input features have a similar scale. Which technique is most appropriate?

A.Increase the learning rate during training.
B.Normalize numerical features to have zero mean and unit variance.
C.Remove outliers from the dataset.
D.Apply one-hot encoding to all categorical features.
AnswerB

Normalization (standardization) transforms numerical features to have zero mean and unit variance, ensuring they are on a similar scale. This helps gradient-based optimization converge faster and prevents features with larger magnitudes from dominating. It is a standard preprocessing step for many machine learning and generative models.

Why this answer

Normalizing numerical features to zero mean and unit variance brings them to a similar scale, which aids optimization and model performance. One-hot encoding is for categorical data, outlier removal is for data cleaning, and learning rate is unrelated to feature scaling. Thus, normalization is the correct preprocessing technique.

Exam trap

The trap here is confusing data cleaning (outlier removal) with feature scaling, which are distinct preprocessing steps.

43
MCQeasy

A machine learning engineer is evaluating a language model's performance on a text summarization task. The model achieves a BLEU score of 0.45 and a ROUGE-L score of 0.62 on the test set. The engineer wants to understand how well the model captures the overall meaning of the source documents. Which evaluation metric should they prioritize?

A.Perplexity, because it measures how well the model predicts the next token.
B.ROUGE-L score, because it measures longest common subsequence and recall of reference content.
C.Accuracy, because it measures the percentage of correctly predicted tokens.
D.BLEU score, because it measures n-gram overlap with reference summaries.
AnswerB

ROUGE-L evaluates the longest common subsequence between generated and reference summaries, emphasizing recall of important content. This aligns well with summarization goals, as it rewards including key information from the source. A high ROUGE-L indicates the model captures the essential meaning and structure of the reference summaries.

Why this answer

ROUGE-L is designed for summarization evaluation and measures recall of the longest common subsequence, which correlates with how much essential content from the reference is included. BLEU focuses on precision and is better for translation. Perplexity and accuracy do not assess summary quality against references.

Therefore, ROUGE-L is the most appropriate metric for capturing overall meaning.

Exam trap

The trap here is assuming that any high score indicates good summarization, when different metrics measure different aspects and only ROUGE-L directly evaluates recall of reference content.

44
Multi-Selectmedium

A team is preparing a dataset to train a generative AI model for text summarization. They want to ensure the model generalizes well and does not simply memorize the training examples. Which TWO practices should they follow? (Choose two.)

Select 2 answers
A.Increase the model's parameter count to capture more complex patterns
B.Split the dataset into training, validation, and test sets
C.Use the entire dataset for training to maximize data availability
D.Apply data augmentation to increase the diversity of training examples
E.Train for as many epochs as possible until training loss is near zero
AnswersB, D

Dividing data into training, validation, and test sets allows the team to tune hyperparameters on the validation set and evaluate final performance on unseen test data. This separation is essential to detect overfitting and estimate how well the model will generalize to new, real-world examples.

Why this answer

To promote generalization and prevent memorization, the team should use a held-out validation set for tuning and a test set for final evaluation, and they should augment training data to increase diversity. Using all data for training, enlarging the model without regularization, or training until training loss is near zero all increase the risk of overfitting.

Exam trap

The trap here is assuming that more data for training or a larger model always improves performance, but without proper validation and augmentation, the model may overfit.

45
MCQmedium

When training a model with a very large dataset, which approach provides the best balance between computational efficiency and model convergence?

A.Full-batch gradient descent.
B.Stochastic gradient descent with batch size 1.
C.Mini-batch gradient descent.
D.Gradient descent without any data shuffling.
AnswerC

Mini-batch gradient descent provides the best of both worlds. It uses GPU parallelism to calculate gradients efficiently across a batch of samples while maintaining enough stochasticity to help the model escape poor local minima. It is the gold standard for scaling training on modern large-scale machine learning systems.

Why this answer

Mini-batch gradient descent strikes an ideal balance by processing small subsets of the data (batches) at a time. This provides the efficiency of vectorization on GPU hardware while maintaining enough stochasticity to help the optimizer avoid poor local minima. Compared to full batch (too slow) or single-sample stochastic descent (too noisy), mini-batching is the industry standard for scaling training to large datasets while achieving optimal convergence.

Exam trap

Candidates often pick full-batch gradient descent for stability or single-sample stochastic descent, ignoring how mini-batching optimizes GPU hardware utilization and convergence.

46
MCQmedium

A team is deploying a large language model for real-time text generation and notices that inference latency is too high. They want to reduce latency without retraining the model. Which technique is most appropriate?

A.Increase the batch size for inference requests.
B.Use a larger model with more parameters.
C.Apply quantization to reduce the model's precision from FP32 to FP16 or INT8.
D.Train the model for additional epochs to improve its efficiency.
AnswerC

Quantization reduces the numerical precision of weights and activations, which decreases memory bandwidth and computational requirements, leading to faster inference. It can be applied post-training without retraining. This directly addresses latency by making the model smaller and faster to execute, while often maintaining acceptable accuracy.

Why this answer

Quantization reduces the precision of model parameters, which lowers memory bandwidth and speeds up computation. It can be applied without retraining and is a common technique for reducing inference latency. Other options either increase latency or do not affect it.

Thus, quantization is the most appropriate choice.

Exam trap

The trap here is confusing throughput with latency; increasing batch size improves throughput but can worsen latency.

47
MCQmedium

When evaluating an LLM for factual accuracy, which metric is most effective at detecting hallucinations compared to simple word-overlap metrics?

A.BLEU Score
B.ROUGE-L
C.NLI-based entailment scoring
D.Perplexity
AnswerC

NLI models are trained to determine if one sentence entails another. By checking if the generated output is entailed by the retrieved context, you can mathematically quantify the factual grounding of the model's response. This effectively identifies hallucinations by detecting when the model makes claims unsupported by the provided source material.

Why this answer

Hallucinations in LLMs are semantic errors that word-overlap metrics like BLEU or ROUGE fail to capture, as they only measure token matching. Model-based evaluation, using frameworks like RAGAS or NLI (Natural Language Inference) models, assesses the truthfulness of generated content against reference contexts. This is critical for building trustworthy enterprise AI, where factual accuracy is paramount and simple text similarity does not suffice for assessing the validity of generated information.

Exam trap

Candidates frequently choose traditional lexical overlap metrics like BLEU or ROUGE, failing to recognize that they only measure token matching and miss semantic hallucinations.

48
MCQmedium

Which THREE of the following are valid methods for improving the inference performance of a deployed deep learning model?

A.Quantization
B.Increasing the number of hidden layers
C.Model Pruning
D.Data Augmentation
E.Knowledge Distillation
AnswerA, C, E

Quantization maps high-precision weights (like FP32) to lower-precision formats (like INT8). This significantly reduces the model size and hardware memory bandwidth requirements, leading to faster inference times on NVIDIA GPUs and edge devices without needing a full rebuild of the underlying architecture or complex training cycles.

Why this answer

Optimizing inference is critical for real-time AI applications. Techniques like quantization reduce precision, lowering memory footprint and increasing throughput. Model pruning removes redundant parameters, reducing computational cost without significant loss in accuracy.

Finally, knowledge distillation transfers the capabilities of a large, high-performance 'teacher' model into a smaller, faster 'student' model, providing a highly optimized deployment candidate that maintains the intelligence of its larger predecessor while being significantly more efficient.

Exam trap

Candidates sometimes select training-time methods like data augmentation or fine-tuning when the question specifically asks for inference performance improvements.

49
MCQmedium

An ML engineer trains a sentiment classifier on 10,000 movie reviews but only 300 are negative. The model predicts positive for nearly every review, including obvious negative ones. Which technique best addresses this class imbalance during training?

A.Apply class weighting in the loss function to penalize errors on the minority class more heavily
B.Normalize the input text by lowercasing and removing stopwords
C.Add L2 regularization to all model weights
D.Increase the learning rate and train for more epochs
AnswerA

Weighting the loss inversely to class frequency increases the gradient contribution of the 300 negative examples, so the optimizer can no longer minimize loss by always predicting the majority class. This directly counteracts the imbalance without discarding data or fabricating examples, and it is a standard, low-risk first remedy for skewed binary classification.

Why this answer

With only 300 negative examples against 9,700 positive ones, an unweighted loss is minimized by predicting the majority class. Applying class weights in the loss raises the cost of minority-class errors, forcing the model to learn features that separate negative reviews. This is the most direct and least invasive fix for the skewed decision boundary.

Exam trap

The trap here is assuming that more training or stronger regularization fixes imbalance, when the real issue is that the loss function ignores how rare the negative class is.

50
MCQmedium

Refer to the exhibit. This error occurs during the training of an LLM. What is the most likely cause for this 'device-side assert' error?

A.The model has run out of VRAM for the current batch.
B.The target indices exceed the vocabulary size.
C.The GPU has overheated and shut down.
D.The learning rate is set too high for convergence.
AnswerB

Loss functions like cross-entropy verify that label indices fall within the range [0, vocab_size - 1]. If an index in the target tensor is out of this range, the underlying CUDA kernel triggers an assertion to prevent invalid memory access, resulting in the reported error during the training step.

Why this answer

A device-side assert error in PyTorch usually indicates an index-out-of-bounds error or a shape mismatch during a kernel operation on the GPU. When the loss function expects a specific range of indices for targets (e.g., in a cross-entropy loss), providing values that exceed the vocab size triggers this assertion. This is a common debugging hurdle that highlights the importance of rigorous input data validation before submitting tensors to the GPU's highly optimized, non-interactive kernels.

Exam trap

Candidates frequently assume the error is a hardware failure or a corrupted model file, overlooking that device-side asserts in PyTorch are almost always caused by index-out-of-bounds errors on GPU kernels.

51
MCQmedium

A data scientist is pretraining a 12-layer transformer encoder on a corpus of legal contracts. To prevent the model from simply copying each token to its output during masked language modeling, the team needs a strategy that forces the model to learn bidirectional context. Which masking approach should they apply?

A.Mask the final token in every sequence and train the model to predict only that token.
B.Replace 50% of tokens with [MASK] and leave the rest unchanged.
C.Mask 15% of input tokens at random, replacing 80% with [MASK], 10% with a random token, and 10% with the original token.
D.Randomly shuffle the order of tokens within each sequence before feeding them to the model.
AnswerC

This is the standard BERT-style masking recipe. Replacing 80% with [MASK] forces the model to predict from context, while the 10% random and 10% unchanged tokens reduce pretrain-finetune mismatch and discourage the model from ignoring non-masked tokens. It directly supports bidirectional learning of legal contract language.

Why this answer

The standard masked language modeling objective masks a small percentage of tokens and requires the model to reconstruct them from bidirectional context. Using mostly [MASK] with some random and unchanged tokens balances learning signal and reduces pretraining-finetuning mismatch. This is the correct way to force bidirectional understanding in a transformer encoder.

Exam trap

The trap here is assuming that any token replacement strategy will work, when the specific 80/10/10 split is designed to prevent the model from ignoring context and to align pretraining with fine-tuning.

52
MCQhard

What is the primary motivation for using Position Embeddings in a transformer model?

A.To reduce the computational burden of attention calculations.
B.To enable the self-attention mechanism to recognize the order of tokens.
C.To optimize the model for inference on NVIDIA Jetson devices.
D.To perform dimensionality reduction on input vocabulary.
AnswerB

Self-attention is inherently position-agnostic; it processes inputs as a set. By injecting position embeddings, we provide the model with essential structural information about the sequence order. This allows the model to learn and respect the sequential nature of natural language, which is crucial for grammar and meaning.

Why this answer

Because the self-attention mechanism is permutation-invariant, it treats every token as if it were independent of its position. Without explicit position information, a transformer would struggle to understand syntax and order-dependent relationships. Adding position embeddings to input tokens injects this necessary structural information, allowing the model to distinguish between 'The dog bit the man' and 'The man bit the dog,' which is vital for language understanding.

Exam trap

Candidates often confuse position embeddings with token embeddings, assuming transformers naturally understand word order without extra mechanisms, or incorrectly believe attention mechanisms track sequence positions inherently.

53
MCQeasy

Which of the following describes the 'Warm-up' phase in the context of training deep neural networks?

A.Pre-loading the model into the GPU cache.
B.Gradually increasing the learning rate at the start of training.
C.A method to reduce training time by caching gradients.
D.Reducing the batch size to fit in memory.
AnswerB

The warm-up phase protects the model from unstable updates by keeping the learning rate low initially. This allows the model to stabilize its internal representations before applying the full, higher learning rate, which is critical for preventing divergence and achieving robust convergence during the initial stages of large-scale training.

Why this answer

Warm-up is a technique where the learning rate starts at a very low value and is gradually increased over the first few hundred or thousand steps. This prevents the model from diverging early in the training process when weights are randomly initialized and gradients can be unstable. Proper warm-up is essential for successfully scaling the training of modern LLMs on large compute clusters, ensuring stable convergence from the very first iterations.

Exam trap

Candidates often mistake warm-up for a data augmentation technique or a method to increase model capacity, failing to recognize it as a stability technique for the initial training phase.

54
MCQhard

In the context of generative AI, what is the 'mode collapse' problem in GANs, and why is it a significant challenge?

A.The generator output becomes purely random noise.
B.The generator fails to produce diverse samples.
C.The discriminator becomes too powerful to train.
D.Training loss increases monotonically to infinity.
AnswerB

Mode collapse occurs when the generator produces only a limited subset of the actual data distribution. Because the generator finds one output that satisfies the discriminator, it stops learning the full variety of the data, resulting in highly repetitive outputs that fail to capture the complexity of the training data.

Why this answer

Mode collapse occurs when the generator in a Generative Adversarial Network learns to map several input noise vectors to the same output or a very limited set of outputs. This prevents the generator from capturing the full diversity of the target data distribution. It is a major challenge because it defeats the purpose of generative modeling, which is to produce a wide range of realistic, diverse samples that represent the training distribution.

Exam trap

Candidates often confuse mode collapse with vanishing gradients or training divergence, failing to realize it specifically refers to the generator's inability to produce diverse outputs despite the discriminator's feedback.

55
MCQhard

A financial services company is deploying an LLM-based assistant that must answer questions about internal compliance documents. The documents are updated weekly, and the company cannot retrain the model every week. The assistant must cite the exact source passage for each answer. Which architecture best satisfies these requirements?

A.Use retrieval-augmented generation (RAG) with a vector index over the compliance documents and return retrieved passages with the answer
B.Train a separate classifier to route questions to static FAQ answers written by the compliance team
C.Fine-tune the LLM weekly on the updated documents and rely on its parametric memory to answer
D.Increase the model's context window and paste all compliance documents into every prompt
AnswerA

RAG retrieves relevant passages from an up-to-date vector index at inference time and can include those passages as citations. When documents change, the index is updated without retraining the model. This directly meets the requirements for freshness and source attribution while keeping the LLM's weights fixed, making it the most suitable architecture for this scenario.

Why this answer

RAG separates knowledge from model weights: documents are indexed in a vector store, relevant passages are retrieved at query time, and the LLM generates an answer grounded in those passages with citations. Weekly document updates require only re-indexing, not retraining. Fine-tuning, stuffing all documents into the prompt, or static FAQ routing either conflicts with the update frequency, does not scale, or cannot provide reliable source citations.

Exam trap

The trap here is assuming that fine-tuning is always the best way to add domain knowledge, when the requirements for weekly updates and exact citations point to retrieval-augmented generation instead.

56
MCQmedium

What is the primary role of 'Loss Scaling' when training deep learning models in FP16 precision?

A.To increase the training speed of the GPU
B.To prevent gradient underflow in FP16 training
C.To reduce the required GPU memory footprint
D.To normalize the input data distribution
AnswerB

FP16 has a limited exponent range, which causes very small gradients to become zero. Scaling the loss by a factor (e.g., 1024) pushes these values into the representable range of the FP16 format. This prevents the model from stalling due to vanishing or zeroed-out gradients during the backpropagation process.

Why this answer

Loss scaling is essential because FP16 has a narrower dynamic range than FP32. Small gradient values can underflow to zero, causing the model to stop learning. By scaling the loss up before backpropagation, the gradients are kept within the representable range of FP16, and then scaled back down during the weight update, ensuring stable and effective training while maintaining the speed advantages of half-precision compute.

Exam trap

Candidates often think loss scaling is for speed or memory optimization. They miss that it is specifically a numerical stability technique to prevent small gradients from becoming zero in FP16.

57
MCQhard

A research team is pretraining a transformer on a corpus of 200 billion tokens. They want the model to learn bidirectional context so each token attends to both left and right neighbors during pretraining. Which pretraining objective fits this requirement?

A.Masked language modeling
B.Contrastive next-sentence prediction
C.Sequence-to-sequence denoising with a causal decoder
D.Autoregressive next-token prediction
AnswerA

Masked language modeling randomly replaces a fraction of input tokens with a mask and trains the model to recover them using the full surrounding context on both sides. Because no causal mask is applied, every prediction can attend to tokens to the left and right, which is exactly the bidirectional pretraining signal the team is asking for.

Why this answer

Bidirectional context means each token's representation is informed by both preceding and following tokens. Masked language modeling achieves this by hiding random tokens and requiring the model to reconstruct them from the full unmasked context, so attention spans the entire sequence. Autoregressive and causal-decoder objectives enforce left-to-right masking, which precludes bidirectional attention.

Exam trap

The trap here is equating any language modeling objective with bidirectional context, when causal masking in autoregressive objectives blocks attention to future tokens.

58
MCQmedium

A researcher is fine-tuning a large language model using PEFT (Parameter-Efficient Fine-Tuning) techniques. Which method is specifically designed to inject trainable low-rank matrices into the transformer layers to reduce the number of trainable parameters?

A.Prefix Tuning
B.LoRA
C.Prompt Tuning
D.Full Fine-Tuning
AnswerB

LoRA uses low-rank decomposition to represent weight updates as the product of two smaller matrices. This approach drastically reduces the total number of trainable parameters, allowing for efficient fine-tuning on consumer-grade or limited-memory GPUs without sacrificing the quality of the original pre-trained model weights.

Why this answer

Low-Rank Adaptation (LoRA) is the standard PEFT method for injecting trainable matrices into frozen pre-trained models. By minimizing the number of updated parameters, it significantly reduces VRAM requirements and storage overhead during the fine-tuning process. This technique is vital for deploying custom LLMs on resource-constrained hardware, as it maintains model performance while drastically simplifying the memory demands of the training pipeline.

Exam trap

Candidates often confuse LoRA with full fine-tuning or quantization techniques like QLoRA, mistakenly believing that LoRA changes the original weights of the frozen pre-trained model directly during the update process.

59
MCQhard

A machine learning engineer is evaluating a large language model (LLM) on a text generation task. They observe that the model produces coherent and fluent sentences, but the content is factually incorrect and sometimes contradicts known facts. Which term best describes this phenomenon?

A.Catastrophic forgetting
B.Hallucination
C.Overfitting
D.Mode collapse
AnswerB

Hallucination refers to the generation of plausible-sounding but factually incorrect or nonsensical content by an LLM. In this scenario, the model produces fluent text that contradicts known facts, which is the hallmark of hallucination. It is a known challenge in generative AI, especially when models are not grounded in external knowledge.

Why this answer

The phenomenon described—fluent but factually incorrect text generation—is known as hallucination. It occurs because LLMs are trained to predict likely sequences of words, not to verify facts. Hallucination is a significant challenge for deploying LLMs in applications requiring factual accuracy, and it can be mitigated with retrieval-augmented generation or factual consistency checks.

Exam trap

The trap here is confusing hallucination with overfitting or mode collapse, but hallucination specifically refers to plausible yet false content, not training issues or lack of diversity.

60
MCQmedium

A developer is using a pretrained large language model for a text summarization task. They want to adapt the model to a domain-specific corpus of legal documents but have limited GPU memory and a small labeled dataset. Which fine-tuning approach is most parameter-efficient and suitable for this scenario?

A.Low-Rank Adaptation (LoRA)
B.Training the model from scratch on the legal corpus
C.Freezing all layers and training only the output head
D.Full fine-tuning of all model parameters
AnswerA

LoRA injects trainable low-rank matrices into existing layers while freezing the original weights, dramatically reducing the number of trainable parameters and memory footprint. This makes it ideal for limited GPU memory and small datasets, as only the adapters are updated. It also preserves pretrained knowledge, reducing overfitting risk and enabling efficient domain adaptation for legal summarization.

Why this answer

LoRA is a parameter-efficient fine-tuning method that freezes the base model and trains small rank-decomposition matrices. It drastically cuts memory usage and trainable parameters, making it feasible on limited hardware and small datasets. Full fine-tuning is resource-heavy, training from scratch is impractical, and training only the head under-adapts to the specialized legal domain.

Exam trap

The trap here is equating parameter efficiency with simply freezing most layers, when methods like LoRA adapt internal representations with minimal added parameters.

61
MCQhard

A data scientist is fine-tuning a pretrained large language model on a small domain-specific dataset. The model achieves high accuracy on the training set but poor performance on a held-out validation set. Which technique is most likely to improve the model's generalization?

A.Increase the number of training epochs to further reduce training loss.
B.Reduce the size of the validation set to decrease evaluation variance.
C.Apply regularization techniques such as dropout or weight decay.
D.Increase the learning rate to speed up convergence on the training set.
AnswerC

Regularization methods like dropout and weight decay constrain the model's capacity to memorize training data, encouraging it to learn more generalizable patterns. In this scenario, the large gap between training and validation performance indicates overfitting, so adding regularization is a direct and effective remedy. These techniques are standard in fine-tuning and can be applied without altering the underlying architecture.

Why this answer

The described symptoms—high training accuracy but poor validation accuracy—are classic signs of overfitting. Regularization techniques such as dropout, weight decay, or early stopping are designed to reduce overfitting by penalizing complexity or adding noise during training. Among the options, applying regularization directly targets the cause, whereas the others either exacerbate overfitting or do not address generalization.

Exam trap

The trap here is confusing overfitting with underfitting, leading to choices that increase model capacity or training time instead of applying regularization.

62
MCQmedium

Why is 'Warmup' used for the learning rate schedule during the initial phase of training large language models?

A.To reduce the amount of data needed to reach convergence.
B.To prevent the optimizer from diverging due to large, noisy initial gradients.
C.To automatically detect the optimal batch size for the hardware.
D.To increase the numerical precision of the gradients.
AnswerB

Early in training, gradients can be highly unstable due to the random initialization of weights. A high learning rate would lead to massive updates, potentially pushing weights into unrecoverable states. Warmup keeps the update magnitude small initially, allowing the optimizer to gain stability before using larger steps.

Why this answer

Learning rate warmup is essential for training stability. At the beginning of training, weights are random, and gradients can be volatile. A low initial learning rate prevents the optimizer from making massive, potentially catastrophic updates that could diverge the model.

Gradually increasing the rate allows the model to stabilize and find a more favorable region in the loss landscape, ensuring a smooth and reliable start to the training process.

Exam trap

Candidates mistakenly believe learning rate warmup is used to accelerate the final convergence speed rather than protecting against volatile initial gradients.

63
MCQmedium

In the context of Large Language Models, what is the primary purpose of 'Attention mechanisms' as introduced in the Transformer architecture?

A.To reduce the number of parameters in the model.
B.To process input sequences sequentially for better memory.
C.To compute dynamic weights representing the relevance of tokens.
D.To enforce a fixed context window for all inputs.
AnswerC

Attention allows the model to dynamically compute the importance of each word in a sentence relative to others. By calculating dot products between query and key vectors, the model assigns weights to values, creating context-aware representations that significantly improve natural language understanding and generation capabilities in large models.

Why this answer

Attention mechanisms allow models to weigh the significance of different tokens in an input sequence relative to one another, regardless of their distance. By computing relevance scores, the model can capture long-range dependencies effectively. This capability is crucial for understanding context, syntax, and semantics, which were historically difficult for sequential models like RNNs that struggled with information retention over long sequences during the training process.

Exam trap

Candidates often describe attention as a memory storage mechanism rather than a dynamic weighting system, failing to recognize that it calculates relevance scores between tokens in the current input sequence.

64
MCQmedium

A machine learning engineer is training a convolutional neural network for image classification and notices that the training loss decreases steadily, but the validation loss starts increasing after a few epochs. The training set is large and representative. Which technique is most directly aimed at addressing this phenomenon?

A.Reduce the size of the training set
B.Add dropout layers
C.Use a smaller batch size
D.Increase the learning rate
AnswerB

Dropout randomly deactivates neurons during training, forcing the network to learn redundant representations and reducing its ability to memorize training noise. This regularization directly combats overfitting, which is the cause of rising validation loss despite decreasing training loss. In this image classification scenario, dropout is a standard and effective remedy when the model has sufficient capacity to overfit.

Why this answer

The described behavior—training loss falling while validation loss rises—is classic overfitting. Dropout is a direct regularization technique that prevents co-adaptation of neurons and improves generalization. Increasing learning rate, reducing training data, or changing batch size do not specifically counter overfitting and may worsen the outcome.

Exam trap

The trap here is confusing overfitting with optimization issues, leading to learning rate or batch size tweaks instead of proper regularization.

65
MCQhard

A research team is training a transformer-based language model and wants to reduce the computational cost of the self-attention mechanism for very long input sequences. They are considering replacing the standard scaled dot-product attention with an approximation. Which statement accurately describes a trade-off of using an approximate attention method?

A.It reduces memory and compute complexity but may sacrifice some model quality compared to full attention.
B.It eliminates the need for positional encodings because approximate attention is order-invariant.
C.It increases the number of parameters in the model, which always improves accuracy.
D.It makes the model immune to the vanishing gradient problem in deep transformer stacks.
AnswerA

Approximate attention methods, such as sparse or low-rank attention, lower the quadratic complexity of full attention to near-linear or sub-quadratic, saving memory and compute on long sequences. However, they approximate the full attention matrix and can miss some token interactions, potentially reducing accuracy or requiring more training to recover quality. This trade-off is the key consideration.

Why this answer

Standard self-attention scales quadratically with sequence length, which is costly for long inputs. Approximate attention methods reduce this complexity to near-linear or sub-quadratic by sparsifying or factorizing the attention matrix. The trade-off is that the approximation may not capture all token interactions, potentially lowering model quality or requiring additional training.

This balance between efficiency and accuracy is the central consideration.

Exam trap

The trap here is assuming approximate attention only adds benefits without drawbacks, when in fact it trades some model quality for reduced memory and compute on long sequences.

66
MCQmedium

What is the primary function of Layer Normalization in a transformer architecture?

A.To increase the total parameter count of the model.
B.To stabilize training by normalizing the inputs to each layer.
C.To replace the need for weight initialization techniques.
D.To compress the model to fit on smaller GPU devices.
AnswerB

By normalizing activations to have zero mean and unit variance, layer normalization reduces the dependence of a layer's output on the specific scaling of its input. This enables the use of higher learning rates and helps mitigate the exploding/vanishing gradient issues common in very deep neural networks.

Why this answer

Layer normalization stabilizes the hidden state distributions across layers by normalizing inputs to have zero mean and unit variance. This prevents internal covariate shift and keeps activations within a stable range, allowing for faster convergence and deeper networks. In transformers, it is typically applied to the output of sub-layers, ensuring that subsequent computations remain numerically stable and well-conditioned for backpropagation.

Exam trap

Candidates often confuse Layer Normalization with Batch Normalization, incorrectly assuming that normalization across the batch dimension is the standard approach used in transformer architectures for stabilizing hidden states.

67
MCQmedium

A machine learning engineer is training a large language model and notices that the model performs exceptionally well on the training data but poorly on a held-out test set. Which technique is most appropriate to mitigate this issue?

A.Apply dropout regularization
B.Reduce the size of the training dataset
C.Train for more epochs
D.Increase the model's parameter count
AnswerA

Dropout randomly deactivates neurons during training, forcing the network to learn redundant representations and preventing co-adaptation. This reduces overfitting and improves generalization to unseen data. In transformer models, dropout is commonly applied to attention weights and feed-forward layers, making it a standard and effective remedy for the described problem.

Why this answer

Overfitting is characterized by high training performance and poor test performance. Dropout regularization mitigates this by preventing the network from relying on specific neurons, encouraging more robust features. Increasing model size, reducing data, or training longer would all aggravate the problem.

Therefore, dropout is the appropriate technique.

Exam trap

The trap here is assuming that more training or a larger model will always improve performance, overlooking the need for regularization when overfitting occurs.

68
MCQhard

In the context of NVIDIA's Tensor Core architecture, what is the primary purpose of 'Sparsity' support?

A.To reduce the amount of VRAM consumed by model weights.
B.To double the computational throughput during GEMM operations.
C.To automatically prune the model during the training process.
D.To allow models to run without any normalization layers.
AnswerB

Structured sparsity enables Tensor Cores to effectively halve the number of required arithmetic operations by skipping zeros. When a matrix satisfies the 2:4 sparsity pattern, the hardware skips calculations for the zero weights, enabling a theoretical 2x speedup in matrix multiplication speed without compromising the model's predictive accuracy.

Why this answer

NVIDIA's Fine-Grained Structured Sparsity is a hardware feature that allows the GPU to skip computations for zero-valued weights in neural networks. By enforcing a 2:4 sparsity pattern (where at least 2 out of every 4 consecutive weights are zero), the hardware can double the throughput of matrix multiplication operations. This feature is vital for accelerating LLMs without sacrificing model accuracy, as it effectively doubles compute density on supported modern NVIDIA architectures.

Exam trap

Candidates often assume sparsity is used solely for reducing model storage size or memory footprint on disk, ignoring its primary hardware execution benefit of doubling computational throughput.

69
MCQeasy

A data scientist is preparing a dataset of 50,000 customer support chat transcripts to fine-tune an LLM for a helpdesk assistant. The raw text contains HTML tags, inconsistent whitespace, and occasional personal information such as email addresses. Which preprocessing step should be performed FIRST to prepare the text for tokenization?

A.Apply tokenization using the model's tokenizer.
B.Convert all text to lowercase and remove all punctuation.
C.Split the transcripts into training and validation sets.
D.Clean the text by stripping HTML, normalizing whitespace, and masking personally identifiable information.
AnswerD

Cleaning raw text before tokenization removes noise that would otherwise become tokens and helps protect sensitive data. Stripping HTML and normalizing whitespace standardizes the input, while masking emails prevents the model from memorizing private information. This step must precede tokenization so the tokenizer operates on the final, intended character sequence.

Why this answer

Raw chat transcripts often contain markup and private data that degrade fine-tuning quality and raise privacy concerns. Cleaning the text by removing HTML, normalizing whitespace, and masking emails or phone numbers before tokenization ensures the tokenizer produces meaningful tokens and that sensitive data never reaches the model. This ordering also avoids retokenization work later.

Exam trap

The trap here is assuming tokenization should happen first because it is the most familiar LLM step, when in fact noisy raw text must be cleaned before tokenization.

70
MCQhard

A team is pre-training a 7-billion-parameter LLM on a large text corpus. They observe that the training loss decreases steadily but the validation loss begins to increase after a certain number of steps. The training and validation data come from the same distribution, and the model has not yet reached the compute budget. Which action is most appropriate to address this behavior?

A.Reduce the model size to match the dataset size.
B.Increase the learning rate to escape the local minimum.
C.Collect more training data from the same distribution.
D.Apply regularization techniques such as dropout or weight decay, and consider early stopping.
AnswerD

The described pattern—training loss decreasing while validation loss increases—is classic overfitting. Adding dropout or weight decay constrains the model's capacity to memorize training data, and early stopping halts training at the point of best validation performance. These are standard, effective remedies when validation loss diverges despite ample compute budget remaining.

Why this answer

A widening gap between decreasing training loss and increasing validation loss signals overfitting. The model is memorizing training-specific patterns rather than generalizing. Applying regularization like dropout or weight decay reduces this tendency, and early stopping captures the best validation checkpoint.

These steps are standard and directly target the observed behavior without unnecessary architectural changes.

Exam trap

The trap here is interpreting rising validation loss as a need for more compute or a larger model, when it actually indicates overfitting and calls for regularization or early stopping.

71
MCQmedium

A data scientist is working with a dataset that has a highly skewed distribution, with one class representing only 2% of the samples. They are training a binary classifier and notice that the model predicts the majority class almost exclusively. Which technique is most appropriate to address this issue?

A.Remove the minority class samples
B.Apply class weighting in the loss function
C.Use a simpler model architecture
D.Increase the learning rate
AnswerB

Class weighting assigns a higher penalty to misclassifying the minority class, which encourages the model to pay more attention to it. This directly addresses the imbalance by adjusting the loss contribution of each class. It is a standard and effective method to improve minority class recall without altering the data distribution.

Why this answer

Class imbalance causes models to favor the majority class because the loss function is dominated by those samples. Applying class weights adjusts the loss to penalize minority class errors more heavily, which balances the influence of each class during training. This is a direct and effective technique for improving minority class detection.

Exam trap

The trap here is thinking that changing the model architecture or learning rate will fix class imbalance, but the core issue is the loss function's bias toward the majority class.

72
MCQmedium

Refer to the exhibit. Which strategy is most effective for resolving this memory error without changing the hardware?

A.Gradient Accumulation
B.Increase the hidden layer size
C.Enable more data augmentation
D.Increase the number of training epochs
AnswerA

Gradient accumulation splits a large batch into smaller sub-batches that fit into GPU memory. The model computes the gradients for each sub-batch and sums them up. Only after several steps is the optimizer updated. This effectively achieves the desired large batch size while keeping the peak memory usage low.

Why this answer

The exhibit indicates that the GPU is nearly at capacity. Gradient accumulation is a standard technique that simulates a larger batch size by accumulating gradients over multiple smaller steps before performing a weight update. This allows the user to maintain the intended effective batch size while fitting the training within the physical memory limits of the GPU, ensuring the training process can proceed.

Exam trap

Candidates often suggest reducing the learning rate or model size to fix memory errors, which ignores the primary issue of batch-size-related memory consumption that gradient accumulation is designed to solve.

73
MCQmedium

A machine learning engineer is evaluating a generative language model for a chatbot application. They notice that the model frequently generates repetitive phrases and gets stuck in loops. Which decoding strategy is most likely to reduce this repetition?

A.Top-k sampling with a small k
B.Greedy search
C.Nucleus sampling with a repetition penalty
D.Beam search with a large beam width
AnswerC

Nucleus sampling (top-p) dynamically selects the smallest set of tokens whose cumulative probability exceeds p, promoting diversity. Adding a repetition penalty reduces the likelihood of tokens that have already appeared, directly mitigating loops. This combination effectively balances coherence and novelty, reducing repetitive phrases.

Why this answer

Repetition in generated text often stems from decoding strategies that favor high-probability tokens. Nucleus sampling with a repetition penalty addresses this by dynamically truncating the probability distribution and penalizing repeated tokens, encouraging diverse and non-repetitive outputs. Greedy and beam search lack such penalties, and top-k with small k restricts diversity.

Exam trap

The trap here is assuming that beam search, which is often used for high-quality outputs, will also prevent repetition, when in fact it can reinforce it.

74
MCQhard

Refer to the exhibit. Which hyperparameter configuration in the provided JSON is directly responsible for preventing overfitting through weight penalty?

A.dropout_rate
B.weight_decay
C.learning_rate_scheduler
D.enable_gradient_check
AnswerB

Weight decay is the standard term for L2 regularization in neural networks. It adds a penalty term to the loss function based on the square of the weights, effectively constraining the weights and preventing them from becoming unnecessarily large, which is the primary mechanism for controlling overfitting through penalty methods.

Why this answer

The JSON defines a standard configuration for model training. The 'weight_decay' key is the parameter that implements L2 regularization. By penalizing large weight values in the loss function, it discourages the model from relying too heavily on specific features, thereby preventing overfitting.

This is a common and essential hyperparameter to tune when training deep neural networks to ensure good generalization on unseen test data, as shown by the provided value of 0.05.

Exam trap

Candidates frequently select learning rate or dropout parameters, confusing general training dynamics or activation controls with the specific weight penalty mechanism of L2 regularization.

75
MCQmedium

What is the primary function of the 'Softmax' layer at the output of a multi-class classification model?

A.To squash output values to a range of -1 to 1.
B.To linearize the output for regression tasks.
C.To convert logits into a probability distribution.
D.To enforce sparsity on the output vector.
AnswerC

Softmax normalizes the raw output scores (logits) by exponentiating them and dividing by the sum of all exponentiated values. This ensures that every output represents the probability of that class, and the sum of all probabilities is 1.0, which is necessary for multi-class classification tasks.

Why this answer

The Softmax function converts the raw output scores (logits) of the final layer into a probability distribution. This involves exponentiating each score and normalizing by the sum of all exponentiated scores, ensuring all outputs are in the range (0, 1) and sum to exactly 1.0. This makes the model's output directly interpretable as probabilities for each class, which is vital for decision-making tasks where confidence scores are needed.

Exam trap

Students often mistake Softmax for an activation function applied to hidden layers or confuse it with Sigmoid, failing to recognize its specific role in multi-class probability normalization.

Page 1 of 2 · 82 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Core Ml Ai Knowledge questions.