NCA-GENL · domain
Core Machine Learning and AI Knowledge
This domain checks the foundational ML and AI concepts behind generative AI and LLMs on NVIDIA platforms. Questions cover transformer internals, mixed-precision training behavior on A100 GPUs, retrieval-augmented generation design choices, and decoding controls like temperature. Expect scenario-based items that ask you to diagnose training issues or select the correct technique.
Focused practice
Practice Core Machine Learning and AI Knowledge questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Core Machine Learning and AI Knowledge
Be able to explain transformer attention, diagnose FP16 numerical issues on NVIDIA GPUs, and tune RAG chunking and LLM temperature. The single most important thing: match each technique to the problem it actually solves, not just its definition.
Transformer attention mechanics and why self-attention captures token relationships across sequences
Mixed-precision FP16 training on NVIDIA A100 GPUs, including loss scaling and numerical stability
RAG pipeline design, including chunk size effects on embedding retrieval quality
LLM decoding controls such as temperature and their effect on output randomness
Watch out for
Common Core Machine Learning and AI Knowledge exam traps
- ▸Assuming FP16 underflow is fixed by lowering the learning rate instead of applying dynamic loss scaling.
- ▸Treating larger RAG chunk sizes as always better, ignoring retrieval precision and context window limits.
- ▸Confusing temperature with top-k or top-p, or believing temperature changes model knowledge rather than sampling randomness.
Question index
All Core Machine Learning and AI Knowledge questions (82)
Click any question to see the full explanation, or start a practice session above.
A data scientist is preparing a labeled dataset of 50,000 customer support tickets for supervised fine-tuning of an LLM. Each ticket must be assigned exactly one of eight department labels. Which loss function is most appropriate for training this classification head?
Easy2A data scientist is training a transformer model and observes that the training loss is decreasing while the validation loss is increasing. Which technique should be prioritized to address this specific generalization challenge?
Medium3A team is evaluating a generative large language model for a customer support chatbot. They need to ensure the model produces factually accurate and contextually appropriate responses while avoiding harmful or biased outputs. Which two techniques are most effective for aligning the model's behavior with these safety and quality requirements? (Choose two.)
Hard4A team is deploying a large language model for real-time text generation. They observe that the model sometimes produces repetitive and dull outputs, especially when generating longer sequences. They want to encourage more diverse and creative text without significantly degrading coherence. Which decoding strategy should they consider?
Medium5Refer to the exhibit. Which technique is most effective for preventing the reported NaN error during model training?
Medium6A researcher is training a large language model and notices the training loss plateaus early while validation loss increases. What is the most likely cause, and which action should be taken?
Medium7A developer is building a retrieval-augmented generation (RAG) system for an internal knowledge base. The system must answer questions using company documents that are updated frequently. Which component is primarily responsible for retrieving the most relevant document chunks to include in the LLM's context?
Medium8When utilizing Pipeline Parallelism (PP) in LLM training, what is the 'pipeline bubble' and how is it minimized?
Hard9Which THREE techniques are commonly used to improve the efficiency of inference for large language models?
Medium10A machine learning engineer is evaluating a fine-tuned LLM for a customer-facing summarization task. The model produces fluent summaries, but the team needs to detect when the model generates content that is not supported by the source document. Which TWO evaluation approaches are appropriate for measuring factual consistency between the generated summary and the source? (Choose two.)
Medium11Refer to the exhibit. Which technique is most appropriate to prevent this specific training failure?
Medium12Which component of an NVIDIA Transformer Engine is specifically designed to accelerate training on supported GPUs by dynamically adjusting precision?
Hard13Which of the following activation functions is most commonly used in hidden layers of deep neural networks to mitigate the vanishing gradient problem?
Easy14A data science team at a retail company is building a neural network to predict customer churn from tabular data with mixed numerical and categorical features. They want the model to output a probability between 0 and 1, and they are training with a standard gradient descent optimizer. Which loss function is most appropriate for this binary classification task?
Easy15A research team is designing a decoder-only transformer LLM for long-document question answering. They want to reduce the quadratic computational cost of self-attention so that training on sequences of 32,000 tokens is feasible on their GPU cluster. Which two techniques are appropriate for this goal? (Choose two.)
Hard16Which activation function is most commonly used in the hidden layers of deep neural networks to mitigate the vanishing gradient problem?
Medium17An AI researcher is fine-tuning a large language model and wants to minimize GPU memory consumption during training without altering the model's primary weight representations or introducing quantization error during inference. Which technique provides this capability by decomposing weight matrices into low-rank trainable adaptation matrices?
Medium18A machine learning engineer is training a deep neural network for image classification. They notice that the training loss decreases steadily, but the validation loss starts to increase after a few epochs. Which technique is most directly aimed at addressing this issue?
Medium19In the context of transformer models, what is the purpose of the 'Attention Mask' during the training process?
Medium20An ML engineer is training a transformer-based language model on a single NVIDIA A100 GPU. They observe that the training loss decreases initially but then becomes NaN after a few hundred steps. The learning rate is 1e-4, and mixed precision with FP16 is enabled. Which action is most likely to stabilize training while preserving the benefits of mixed precision?
Hard21Refer to the exhibit. A machine learning engineer reviews the monitoring output from a two-GPU distributed training job running on an NVIDIA DGX system. GPU 0 shows low utilization despite high memory consumption, while GPU 1 shows high utilization and high memory consumption. What is the most likely root cause of this performance imbalance?
Hard22A researcher is fine-tuning a large language model on a downstream task with a small dataset. They notice that the model achieves high training accuracy but poor validation accuracy. Which regularization technique is most appropriate to address this issue?
Hard23A machine learning engineer is fine-tuning a pre-trained language model on a small domain-specific dataset. She notices that the model quickly achieves high accuracy on the training set but performs poorly on the validation set. She wants to mitigate this overfitting without collecting more data. Which technique is most appropriate?
Hard24A data scientist is preprocessing a text corpus to train a large language model. They want to convert each word into a dense vector representation that captures semantic relationships before feeding it into the transformer. Which technique should they use?
Easy25A machine learning engineer is training a large language model on a cluster of NVIDIA GPUs. During training, she observes that the loss occasionally spikes to NaN, causing the training to fail. She suspects that the issue is related to the numerical precision of the computations. Which technique is most appropriate to mitigate this issue while maintaining training stability?
Medium26A team is deploying a large language model for real-time inference on an NVIDIA GPU. They observe that the first few inference requests have high latency, but subsequent requests are much faster. What is the most likely explanation for this behavior?
Hard27An enterprise machine learning team is training a large-scale transformer model on a cluster of NVIDIA A100 GPUs using mixed precision (FP16). During the initial training phase, the team notices sudden numerical underflow resulting in vanishing gradients and stalled loss convergence. Which optimization technique must be applied to mitigate this issue without sacrificing the memory-efficiency benefits of FP16?
Medium28A data science team is building a model to predict whether a customer will churn based on historical account activity. They have a large dataset with labeled outcomes (churned or not churned). Which type of machine learning is most appropriate for this task?
Easy29When implementing Retrieval-Augmented Generation (RAG), why is the choice of 'Chunk Size' critical for model retrieval performance?
Medium30A developer is building a text summarization assistant that must produce concise, faithful summaries of long support tickets. The team wants to fine-tune a pre-trained large language model on a small labeled dataset of ticket-summary pairs. Which training approach best matches this goal?
Easy31A machine learning engineer is training a deep neural network and notices that the training loss decreases but the validation loss starts to increase after several epochs. Which two techniques are most appropriate to mitigate this issue? (Choose two.)
Medium32A team deploys a retrieval-augmented generation pipeline and observes that answers frequently cite facts not present in the retrieved passages. They want to reduce this unsupported generation behavior. (Choose two.)
Hard33Refer to the exhibit. The training loss is oscillating and failing to converge. What is the most likely immediate adjustment needed?
Medium34Refer to the exhibit. The configuration shows the use of FSDP with mixed precision. What is the main benefit of using 'bf16' (Bfloat16) over 'fp16' in this context?
Medium35What is the role of 'Temperature' in the context of LLM text generation?
Easy36Which machine learning paradigm involves an agent learning to make decisions by performing actions in an environment to maximize a cumulative reward?
Easy37What is the primary function of the 'Attention' mechanism in Transformer models?
Medium38An ML engineer is deploying a Transformer-based inference service on an NVIDIA TensorRT-LLM runtime. To maximize inference throughput and reduce latency under heavy concurrent user traffic, the engineer needs to select the optimal decoding batching strategy. Which technique allows multiple incoming dynamic sequence requests to be batched together at the token level rather than waiting for entire sequences to complete?
Hard39A data scientist is preparing a transformer-based language model for a text summarization task. She notices that the input sequences in her dataset vary widely in length, from a few tokens to several thousand. She decides to set a fixed maximum sequence length and pad shorter sequences with a special token. Which component of the transformer architecture is primarily responsible for handling the positional information of tokens in these sequences?
Easy40Which of the following describes the purpose of a validation set in machine learning?
Easy41A machine learning engineer is deploying a transformer-based language model for real-time translation. They observe that inference latency is too high for the required throughput. The model uses standard multi-head self-attention. Which modification is most likely to reduce latency without significantly degrading translation quality?
Hard42A machine learning engineer is preprocessing a dataset for a generative AI model and wants to ensure that the input features have a similar scale. Which technique is most appropriate?
Medium43A machine learning engineer is evaluating a language model's performance on a text summarization task. The model achieves a BLEU score of 0.45 and a ROUGE-L score of 0.62 on the test set. The engineer wants to understand how well the model captures the overall meaning of the source documents. Which evaluation metric should they prioritize?
Easy44A team is preparing a dataset to train a generative AI model for text summarization. They want to ensure the model generalizes well and does not simply memorize the training examples. Which TWO practices should they follow? (Choose two.)
Medium45When training a model with a very large dataset, which approach provides the best balance between computational efficiency and model convergence?
Medium46A team is deploying a large language model for real-time text generation and notices that inference latency is too high. They want to reduce latency without retraining the model. Which technique is most appropriate?
Medium47When evaluating an LLM for factual accuracy, which metric is most effective at detecting hallucinations compared to simple word-overlap metrics?
Medium48Which THREE of the following are valid methods for improving the inference performance of a deployed deep learning model?
Medium49An ML engineer trains a sentiment classifier on 10,000 movie reviews but only 300 are negative. The model predicts positive for nearly every review, including obvious negative ones. Which technique best addresses this class imbalance during training?
Medium50Refer to the exhibit. This error occurs during the training of an LLM. What is the most likely cause for this 'device-side assert' error?
Medium51A data scientist is pretraining a 12-layer transformer encoder on a corpus of legal contracts. To prevent the model from simply copying each token to its output during masked language modeling, the team needs a strategy that forces the model to learn bidirectional context. Which masking approach should they apply?
Medium52What is the primary motivation for using Position Embeddings in a transformer model?
Hard53Which of the following describes the 'Warm-up' phase in the context of training deep neural networks?
Easy54In the context of generative AI, what is the 'mode collapse' problem in GANs, and why is it a significant challenge?
Hard55A financial services company is deploying an LLM-based assistant that must answer questions about internal compliance documents. The documents are updated weekly, and the company cannot retrain the model every week. The assistant must cite the exact source passage for each answer. Which architecture best satisfies these requirements?
Hard56What is the primary role of 'Loss Scaling' when training deep learning models in FP16 precision?
Medium57A research team is pretraining a transformer on a corpus of 200 billion tokens. They want the model to learn bidirectional context so each token attends to both left and right neighbors during pretraining. Which pretraining objective fits this requirement?
Hard58A researcher is fine-tuning a large language model using PEFT (Parameter-Efficient Fine-Tuning) techniques. Which method is specifically designed to inject trainable low-rank matrices into the transformer layers to reduce the number of trainable parameters?
Medium59A machine learning engineer is evaluating a large language model (LLM) on a text generation task. They observe that the model produces coherent and fluent sentences, but the content is factually incorrect and sometimes contradicts known facts. Which term best describes this phenomenon?
Hard60A developer is using a pretrained large language model for a text summarization task. They want to adapt the model to a domain-specific corpus of legal documents but have limited GPU memory and a small labeled dataset. Which fine-tuning approach is most parameter-efficient and suitable for this scenario?
Medium61A data scientist is fine-tuning a pretrained large language model on a small domain-specific dataset. The model achieves high accuracy on the training set but poor performance on a held-out validation set. Which technique is most likely to improve the model's generalization?
Hard62Why is 'Warmup' used for the learning rate schedule during the initial phase of training large language models?
Medium63In the context of Large Language Models, what is the primary purpose of 'Attention mechanisms' as introduced in the Transformer architecture?
Medium64A machine learning engineer is training a convolutional neural network for image classification and notices that the training loss decreases steadily, but the validation loss starts increasing after a few epochs. The training set is large and representative. Which technique is most directly aimed at addressing this phenomenon?
Medium65A research team is training a transformer-based language model and wants to reduce the computational cost of the self-attention mechanism for very long input sequences. They are considering replacing the standard scaled dot-product attention with an approximation. Which statement accurately describes a trade-off of using an approximate attention method?
Hard66What is the primary function of Layer Normalization in a transformer architecture?
Medium67A machine learning engineer is training a large language model and notices that the model performs exceptionally well on the training data but poorly on a held-out test set. Which technique is most appropriate to mitigate this issue?
Medium68In the context of NVIDIA's Tensor Core architecture, what is the primary purpose of 'Sparsity' support?
Hard69A data scientist is preparing a dataset of 50,000 customer support chat transcripts to fine-tune an LLM for a helpdesk assistant. The raw text contains HTML tags, inconsistent whitespace, and occasional personal information such as email addresses. Which preprocessing step should be performed FIRST to prepare the text for tokenization?
Easy70A team is pre-training a 7-billion-parameter LLM on a large text corpus. They observe that the training loss decreases steadily but the validation loss begins to increase after a certain number of steps. The training and validation data come from the same distribution, and the model has not yet reached the compute budget. Which action is most appropriate to address this behavior?
Hard71A data scientist is working with a dataset that has a highly skewed distribution, with one class representing only 2% of the samples. They are training a binary classifier and notice that the model predicts the majority class almost exclusively. Which technique is most appropriate to address this issue?
Medium72Refer to the exhibit. Which strategy is most effective for resolving this memory error without changing the hardware?
Medium73A machine learning engineer is evaluating a generative language model for a chatbot application. They notice that the model frequently generates repetitive phrases and gets stuck in loops. Which decoding strategy is most likely to reduce this repetition?
Medium74Refer to the exhibit. Which hyperparameter configuration in the provided JSON is directly responsible for preventing overfitting through weight penalty?
Hard75What is the primary function of the 'Softmax' layer at the output of a multi-class classification model?
Medium76A team trains a transformer language model on a large corpus but the model achieves very low training loss while performing poorly on held-out text. Which action most directly addresses this outcome?
Medium77Which of the following best defines 'Generalization' in machine learning?
Easy78A machine learning engineer is evaluating a generative LLM for a customer-facing question-answering system. The model produces fluent answers, but during testing it confidently states incorrect facts about company policies. The team wants a metric that specifically measures whether the model's output is supported by the provided source documents. Which evaluation approach is most appropriate?
Medium79A data scientist is preparing a dataset of 50,000 customer support conversations to fine-tune a large language model. The conversations vary widely in length, and many exceed the model's maximum context window. The team wants to preserve conversational coherence while avoiding truncation that removes critical resolution details. Which preprocessing strategy is most appropriate?
Medium80A team is pretraining a large language model on a cluster of NVIDIA GPUs. They observe that the model's training loss decreases steadily for the first few epochs but then suddenly spikes and eventually becomes NaN. They suspect this is due to exploding gradients. Which technique is most appropriate to address this issue?
Medium81Which THREE of the following factors are critical when choosing a foundation model for an enterprise generative AI application?
Hard82Which component of an NVIDIA AI stack is primarily responsible for providing a low-level API for high-performance collective communication primitives across multi-GPU nodes?
EasyOther domains
All NCA-GENL exam domains
Frequently asked questions
- What does the Core Machine Learning and AI Knowledge domain cover on the NCA-GENL exam?
- Be able to explain transformer attention, diagnose FP16 numerical issues on NVIDIA GPUs, and tune RAG chunking and LLM temperature. The single most important thing: match each technique to the problem it actually solves, not just its definition.
- How many questions are in this domain?
- This page lists all 82 Core Machine Learning and AI Knowledge questions in the NCA-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Core Machine Learning and AI Knowledge questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.