NVIDIA · Free Practice Questions · Last reviewed May 2026
30real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
A data scientist is monitoring a fine-tuning job on a DGX system. The training loss graph shows a sharp, localized spike followed by an immediate return to the previous trend. What is the most likely cause?
The model has reached global convergence prematurely.
The learning rate is set significantly too high.
A single corrupted data sample was processed.
A corrupted sample or an outlier that violates the expected data distribution often causes a sudden, momentary spike in the gradient calculation. Once that batch is processed and the optimizer proceeds to the next valid data point, the loss typically returns to its previous trend as the model resumes learning.
The GPU memory buffer has overflowed.
Which TWO of the following visualization techniques are most effective for identifying latent patterns in high-dimensional embedding spaces during LLM evaluation?
T-distributed Stochastic Neighbor Embedding (t-SNE).
t-SNE is highly effective at capturing local structure in high-dimensional data, making it ideal for visualizing clusters of related concepts in embedding space. It excels at revealing intricate semantic relationships that would otherwise remain hidden within thousands of dimensions, providing researchers with actionable insights into model internal representations.
Uniform Manifold Approximation and Projection (UMAP).
UMAP preserves both local and global data structures better than many alternatives while maintaining high computational efficiency for large embedding datasets. By projecting high-dimensional data into low-dimensional space, it allows developers to visually inspect how the model groups different entities, which is crucial for fine-tuning performance validation.
Standard bar charts of token frequency.
Basic line charts showing epoch time.
Histogram of output sequence length.
Refer to the exhibit. A monitoring script outputs this JSON for an LLM inference service. What does the 'p99' metric represent in this context?
The average latency of all requests processed.
The latency of the slowest 1% of requests.
The p99 value indicates that 99% of requests meet this threshold, effectively capturing the upper bound of latency for the vast majority of users. It is a vital metric for identifying performance spikes or infrastructure bottlenecks that impact the worst-case scenario user experiences in a production environment.
The median latency observed during the period.
The total throughput of the inference server.
Which visualization tool is most suitable for tracking the gradient norm evolution during the training of a large language model to detect vanishing or exploding gradients?
Scatter plot matrix.
Line chart.
Line charts provide a clear chronological representation of scalar values, making them the industry standard for monitoring training metrics. They allow for the rapid identification of trends, spikes, and instabilities in gradient norms, providing immediate visual feedback on the health of the model's weight update process over time.
Heat map.
Pie chart.
When evaluating LLM output quality using human-in-the-loop data, which THREE metrics or techniques are most effective for detecting systemic hallucinations?
Consistency check across multiple temperature settings.
Systemic hallucinations often shift when the model's randomness is adjusted. If the model produces different factual assertions at varying temperatures, it signals a lack of grounding in the training data, helping developers isolate parts of the knowledge base that are prone to model fabrication during generative inference tasks.
Natural Language Inference (NLI) scores.
NLI models check whether a generated statement is entailed by a provided source document. By using NLI to score the factual alignment between generated text and source material, developers can automatically identify contradictions that indicate hallucinations, which is a key process for validating large-scale generative model outputs.
Human-labeled factuality score cards.
Expert human evaluation remains the gold standard for ground-truth verification of model outputs. Factuality score cards provide a structured way to quantify hallucination rates, allowing teams to create high-quality datasets for further reinforcement learning or model evaluation, ensuring that human intent aligns with the generated output content.
Word count distribution analysis.
Model training loss convergence tracking.
You are performing a comparative analysis of two different LLM architectures by visualizing their performance on a RAG (Retrieval-Augmented Generation) benchmark. Which visualization is best for comparing the distributions of answer accuracy scores?
Stacked bar chart of token counts.
Box plot comparing accuracy scores.
Box plots are ideal for comparing statistical distributions. They highlight the median performance and the spread of scores, allowing developers to immediately identify which model has a more consistent performance profile and which one is prone to extreme outliers, which is essential for benchmarking different LLM performance architectures.
Radial plot of training time.
Individual data point scatter plot.
Want more Data Analysis and Visualization practice?
Practice this domainAn AI researcher is fine-tuning a Llama-3 model using NeMo Framework and notices high GPU memory usage during training. Which experimentation technique is most effective for reducing memory footprint without sacrificing model quality?
Increase the learning rate significantly
Reduce the batch size to one
Enable gradient checkpointing
Gradient checkpointing saves memory by discarding intermediate activations and recomputing them during the backward pass. This allows for training larger models or using larger batch sizes on constrained hardware. It is the industry-standard experimentation approach for managing memory overhead without compromising the mathematical integrity of the training process.
Switch to a smaller model architecture
When conducting A/B testing on an LLM-based application, which metric is the most reliable indicator of user satisfaction regarding response quality?
Inference latency
Tokens per second
LLM-as-a-Judge score
Using a stronger model to evaluate the outputs of the model under test provides a consistent, scalable quality metric. It mimics human evaluation, allowing for rapid experimentation cycles where qualitative performance can be measured against specific criteria like reasoning, tone, and accuracy in a reproducible and automated manner.
GPU memory utilization
A data scientist is preparing to perform hyperparameter tuning for a Retrieval-Augmented Generation (RAG) system. Which TWO parameters should be prioritized for experimentation to improve retrieval accuracy?
Chunk size
Chunk size directly determines how much context is captured in each vector index entry. Too small, and the model lacks enough information; too large, and the content becomes noisy, leading to irrelevant retrieval. Finding the optimal balance through experimentation is essential for ensuring high-quality context retrieval for the LLM.
Model quantization bit-width
Similarity search top-k
The top-k parameter determines how many documents are passed to the generator. If top-k is too low, critical information might be missed; if too high, the context window might be flooded with irrelevant data, confusing the LLM. Testing different values is vital for balancing retrieval precision and recall.
System prompt length
GPU clock speed
When fine-tuning a model on a new dataset, why is it important to keep a portion of the original pre-training data in the fine-tuning mix?
To increase the total number of training epochs
To prevent catastrophic forgetting
Mixing a small fraction of original pre-training data ensures the model remains anchored to its general knowledge base. Without this 'replay' technique, the model tends to overwrite its pre-trained weights with specific new patterns, resulting in a loss of general-purpose capabilities that were present before the fine-tuning process started.
To reduce the computational time of fine-tuning
To improve the hardware utilization rates
In the context of NVIDIA NeMo, which THREE actions are part of a robust experiment tracking workflow for fine-tuning?
Logging hyperparameters for every run
Logging hyperparameters is necessary to reproduce experiments. Without knowing the exact settings like learning rate, optimizer parameters, and batch size, it is impossible to verify why a specific run achieved its results, making it difficult to improve performance iteratively or justify the configuration to stakeholders in a professional environment.
Deleting logs to conserve disk space
Saving model checkpoints periodically
Periodic checkpoints allow for fault tolerance and comparison of intermediate model states. If a model starts overfitting or diverging, having earlier checkpoints allows the researcher to revert to a better state, saving time and compute resources while ensuring that the best version of the model is ultimately identified.
Automating metric collection via tools like W&B
Automation tools like Weights & Biases provide visual dashboards that simplify the process of comparing dozens of experiments. This reduces human error in data collection and provides clear, actionable insights into how different hyperparameter configurations impact convergence, enabling faster decision-making throughout the experimentation lifecycle for large language models.
Manually calculating gradients during training
Which of the following is the primary goal of the 'Experimentation' phase in an LLM project?
Deploying the model to production
Determining optimal configurations for a task
The objective of experimentation is to systematically test configurations to find those that yield the best performance. Whether it's hyperparameter tuning for fine-tuning or prompt testing for RAG, this phase provides the data-driven evidence needed to select the model setup that best balances quality with resource constraints.
Purchasing new hardware for the team
Writing marketing copy for the product
Want more Experimentation practice?
Practice this domainAn organization is deploying an LLM for customer support. To ensure Trustworthy AI, which approach best mitigates the risk of model hallucination while maintaining factual grounding?
Increase the model's temperature parameter to maximum to encourage diverse reasoning.
Apply fine-tuning on the entire customer support history to memorize expected responses.
Implement Retrieval-Augmented Generation (RAG) using a vector database for source verification.
RAG architecture provides the model with external, high-quality data at inference time. By retrieving relevant document chunks, the LLM generates answers based on existing, verifiable facts rather than internal weights. This process grounds the response and provides a clear mechanism to link outputs back to specific source material.
Restrict the model to only use pre-computed templates for every possible customer query.
Which technique should an organization prioritize to identify and reduce systematic bias in a generative model's training dataset?
Apply post-processing filters to censor all sensitive keywords in model output.
Conduct a comprehensive audit of the training corpus to identify demographic imbalances.
Data auditing allows engineers to quantify the distribution of demographics and concepts within the training corpus. By identifying and balancing these distributions, developers can prevent the model from learning biased correlations. This proactive approach is the industry gold standard for creating fair and ethical generative models from scratch.
Increase the model size to allow for better internal alignment with human values.
Use an adversarial model to guess the sensitive attributes of the primary model's output.
Which of the following best describes the principle of 'Interpretability' in the context of Trustworthy AI?
The ability of the model to perform multiple tasks simultaneously without loss of accuracy.
The capability to explain the internal decision-making process in human-understandable terms.
Interpretability is defined by the transparency of the model's reasoning. By providing insights into which features or input patterns drove a specific output, stakeholders can verify that the model is operating logically and ethically, which is a fundamental requirement for establishing user trust in complex AI systems.
The speed at which a model can process training data during the fine-tuning phase.
The process of removing all personal identifiers from the training dataset.
Which THREE actions are recommended for establishing a robust 'Human-in-the-Loop' (HITL) system for an AI deployment?
Setting specific confidence thresholds that trigger human review.
Confidence thresholds provide a quantitative way to define when an AI is 'unsure.' By routing low-confidence outputs to human experts, organizations can prevent errors from reaching end-users. This mechanism is vital for maintaining high quality and reliability in automated systems, serving as an essential safety gate for deployment.
Automating all processes to eliminate the potential for human error.
Creating intuitive interfaces for humans to edit or approve AI outputs.
Intervention tools allow humans to efficiently correct or validate AI suggestions. Without intuitive interfaces, the review process becomes a bottleneck, leading to user fatigue or errors. Well-designed UI for HITL processes increases the efficiency of the human expert, making the overall system more responsive and reliable in production.
Incorporating expert feedback to refine and improve the model over time.
The value of HITL extends beyond immediate correction; it provides labeled data that can be used for RLHF or fine-tuning. By feeding human-approved outputs back into the training pipeline, the model continuously learns from expert knowledge, reducing future errors and improving the overall quality and reliability of the AI.
Restricting human access to the model's internal weights and architecture.
Refer to the exhibit. Which concept of Trustworthy AI is primarily demonstrated by the actions shown in the CLI output?
Model Explainability.
Data Privacy.
Safety and Robustness.
The system successfully detects harmful content and executes a safety protocol to prevent it from reaching the user. This demonstrates that the model is robust against generating toxic content and adheres to predefined safety standards, which are fundamental components of maintaining a trustworthy and harmless generative AI system.
Algorithmic Efficiency.
Why is 'Data Provenance' considered a crucial component in maintaining Trustworthy AI?
It ensures that the model can be compressed into a smaller size for edge deployment.
It provides a clear audit trail for data lineage, ethics, and legal compliance.
Provenance is essential for verifying that the model was trained on data that is both legally sourced and ethically managed. It allows organizations to demonstrate compliance during audits and proactively address potential issues related to copyright infringement or data contamination, which are vital for long-term AI sustainability.
It speeds up the GPU training process by indexing the data in a vector database.
It automatically corrects grammatical errors in the training corpus.
Want more Trustworthy AI practice?
Practice this domainWhen deploying a Large Language Model using TensorRT-LLM, which TWO configuration factors must be tuned to maximize KV cache efficiency?
KV cache block size.
The KV cache block size determines how memory is partitioned for attention heads. Setting an optimal block size balances memory fragmentation and allocation overhead, preventing wasted VRAM that would otherwise be unusable for storing new tokens, thus maximizing the total sequence capacity available for concurrent users.
Model quantization bit-width.
Maximum number of blocks.
Configuring the maximum number of blocks controls the upper bound of active tokens stored in VRAM. Properly sizing this limit prevents the inference engine from crashing due to memory exhaustion while ensuring that high-concurrency workloads are supported without triggering constant re-allocation, which is computationally expensive.
Input prompt token limit.
GPU clock speed frequency.
When integrating an LLM into an application using NVIDIA API endpoints, what is the primary purpose of the 'System' role in the messages payload?
To provide the user with a history of previous successful interactions.
To define the model's behavioral guidelines and operational constraints.
The System role is explicitly designed to set the stage for the model's persona, functional boundaries, and safety policies. By defining these at the start, developers ensure the model adheres to application requirements regardless of user input, providing a stable foundation for the conversation context.
To act as a buffer for temporary memory storage during inference.
To specify the hardware architecture used for the inference request.
A developer needs to ensure that an LLM application remains deterministic across multiple runs. Which parameter configuration is most effective?
Increase top_p to 1.0 and set temperature to 0.7.
Set temperature to 0.0 and define a fixed seed.
Temperature 0.0 forces the model to choose the most likely token (greedy decoding), while a fixed seed ensures the underlying noise in the sampling process remains constant. Combined, these create a highly deterministic environment where inputs consistently map to the same output tokens, satisfying the requirement.
Disable streaming and use a large batch size for inference.
Apply top_k filtering with a value of 50.
When designing a scalable inference microservice, which factor most significantly impacts the 'Time to First Token' (TTFT) for concurrent users?
The total number of output tokens generated.
The prefill phase efficiency and GPU throughput.
The prefill phase is where the model processes the prompt and calculates the initial KV cache. High GPU throughput and optimized kernels (like those in TensorRT-LLM) directly reduce the duration of this compute-heavy phase, which is the primary contributor to the time elapsed before the first token appears.
The network latency between the user and the load balancer.
The number of active vector databases connected to the service.
Refer to the exhibit. Given the current configuration, what is the primary risk during high-traffic bursts?
FP8 quantization will cause excessive model hallucinations.
Head-of-line blocking due to the FCFS scheduling policy.
FCFS processes requests in the order they arrive regardless of their compute requirements. A single long-generation request will occupy GPU resources, causing all subsequent, potentially short requests to wait. This leads to poor overall system latency and creates a bottleneck during high-traffic bursts, significantly degrading user experience.
The request limit of 128 is too low to saturate the GPU.
The strict policy setting prevents dynamic batching.
When utilizing NVIDIA NIM for deployment, why is it recommended to use a containerized environment?
To bypass the need for CUDA drivers on the host machine.
To ensure dependency consistency and portability.
AI models require specific versions of libraries (e.g., specific CUDA versions). Containerization bundles these dependencies, ensuring that the environment is reproducible and portable. This eliminates version conflicts and configuration drift, allowing the same microservice to run reliably across local workstations, testing clusters, and cloud production environments.
To improve the GPU's clock speed by optimizing kernel distribution.
To automatically optimize the model's weight distribution for multi-GPU setups.
Want more Software Development practice?
Practice this domainA researcher is training a large language model and notices the training loss plateaus early while validation loss increases. What is the most likely cause, and which action should be taken?
The model is underfitting; increase the number of hidden layers.
The learning rate is too low; increase it to accelerate convergence.
The model is overfitting; apply dropout or weight decay.
Overfitting occurs when the model complexity exceeds the information content in the training set. Dropout randomly disables neurons during training, preventing co-adaptation, while weight decay penalizes large weights. These techniques effectively reduce the variance of the model, forcing it to focus on generalized representations instead of training noise.
The dataset is too small; reduce the batch size.
In the context of transformer models, what is the purpose of the 'Attention Mask' during the training process?
It serves to reduce the number of parameters in the self-attention layer.
It prevents the model from attending to future tokens in autoregressive models.
In autoregressive LLMs, the model must predict the next token based only on previous ones. The attention mask sets the attention scores for future positions to negative infinity before softmax, effectively nullifying their influence. This ensures the model learns causal dependencies during its training phase.
It optimizes the data movement between the GPU's L1 and L2 cache.
It performs weight pruning to compress the model size after training.
What is the primary function of Layer Normalization in a transformer architecture?
To increase the total parameter count of the model.
To stabilize training by normalizing the inputs to each layer.
By normalizing activations to have zero mean and unit variance, layer normalization reduces the dependence of a layer's output on the specific scaling of its input. This enables the use of higher learning rates and helps mitigate the exploding/vanishing gradient issues common in very deep neural networks.
To replace the need for weight initialization techniques.
To compress the model to fit on smaller GPU devices.
Why is 'Warmup' used for the learning rate schedule during the initial phase of training large language models?
To reduce the amount of data needed to reach convergence.
To prevent the optimizer from diverging due to large, noisy initial gradients.
Early in training, gradients can be highly unstable due to the random initialization of weights. A high learning rate would lead to massive updates, potentially pushing weights into unrecoverable states. Warmup keeps the update magnitude small initially, allowing the optimizer to gain stability before using larger steps.
To automatically detect the optimal batch size for the hardware.
To increase the numerical precision of the gradients.
What is the primary motivation for using Position Embeddings in a transformer model?
To reduce the computational burden of attention calculations.
To enable the self-attention mechanism to recognize the order of tokens.
Self-attention is inherently position-agnostic; it processes inputs as a set. By injecting position embeddings, we provide the model with essential structural information about the sequence order. This allows the model to learn and respect the sequential nature of natural language, which is crucial for grammar and meaning.
To optimize the model for inference on NVIDIA Jetson devices.
To perform dimensionality reduction on input vocabulary.
Which THREE techniques are commonly used to improve the efficiency of inference for large language models?
Weight Quantization to reduce the bit-width of model parameters.
Quantization converts model weights from higher precision (like FP32) to lower precision (INT8 or FP8). This substantially decreases the memory footprint and accelerates inference speed on NVIDIA GPUs, as hardware can process more low-precision operations in parallel with less bandwidth consumption and power usage.
KV-Caching to store previously computed tokens during generation.
During autoregressive decoding, the model computes key and value vectors for all previous tokens. KV-caching stores these results in GPU memory, avoiding redundant computations for each new token generated. This significantly reduces the time-per-token, which is essential for low-latency inference in chatbots and long-form generation.
Increasing the number of transformer layers to add depth.
Weight Pruning to remove redundant connections in the network.
Pruning removes weights that contribute little to the model's output, resulting in a sparse representation. This reduces the total number of operations required during a forward pass. When implemented with hardware support for sparse matrix multiplication, this can lead to substantial speedups and memory efficiency during inference.
Replacing all activations with Sigmoid functions for speed.
Want more Core Machine Learning and AI Knowledge practice?
Practice this domainThe NCA-GENL exam has 60–90 questions and must be completed in 120 minutes. The passing score is 700/1000.
Scenario-based questions covering exam objectives with detailed answer explanations.
The exam covers 5 domains: Data Analysis and Visualization, Experimentation, Trustworthy AI, Software Development, Core Machine Learning and AI Knowledge. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official NVIDIA NCA-GENL exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.