Courseiva

NVIDIA Certified Professional: Generative AI LLMs (NCP-GENL) — Questions 226–300

352 questions total · 5pages · All types, answers revealed

Page 3

Page 4 of 5

Page 5
226
MCQmedium

When using the NVIDIA Collective Communications Library (NCCL), what does 'AllReduce' specifically optimize for in a multi-GPU training configuration?

A.Local GPU thread synchronization.
B.Gradient aggregation across all devices.
C.Direct CPU-to-GPU data copy.
D.Increasing GPU clock frequency.
AnswerB

AllReduce performs a summation of gradients across all participating GPUs and returns the result to every device. This ensures all model replicas are synchronized during the weight update phase, making it the fundamental operation for scaling deep learning training across clusters of multiple NVIDIA GPUs.

Why this answer

AllReduce is critical for distributed training because it synchronizes the gradient updates from all GPUs across the cluster. It aggregates data from all devices and distributes the result back to each one, allowing models to train concurrently on massive datasets. By optimizing this collective operation, NCCL minimizes the time GPUs spend waiting for synchronization, which is the primary hurdle in scaling training to hundreds or thousands of GPUs.

Exam trap

Candidates often confuse AllReduce with point-to-point communication or broadcast operations, failing to recognize it as the specific collective operation for synchronizing gradients across distributed nodes.

227
MCQmedium

In the context of LLM deployment, why is it recommended to use a dedicated inference server like Triton rather than a basic Flask or FastAPI wrapper?

A.Flask is incompatible with the NVIDIA CUDA drivers.
B.Triton manages concurrent model execution and hardware utilization.
C.Custom wrappers require more memory than Triton servers.
D.Triton is the only way to support REST API endpoints.
AnswerB

Triton is optimized for GPU utilization, providing features such as dynamic batching, concurrent model execution, and shared memory management. These optimizations are difficult to implement correctly in standard web frameworks and are essential for maximizing the performance of expensive GPU resources in large-scale LLM deployments.

Why this answer

While basic web frameworks are easy to implement, they lack the sophisticated features required for production AI. Triton offers hardware-optimized concurrency, automatic dynamic batching, support for multiple backends (TensorRT, PyTorch, ONNX), and advanced metrics. These features are critical for handling the high computational demands and low-latency requirements of LLMs, providing better scalability and reliability compared to custom-built wrappers that do not natively manage GPU streams and queues.

Exam trap

Candidates assume that because Flask/FastAPI can serve models, they are sufficient for production. They overlook the critical need for dynamic batching and hardware-optimized concurrency provided by dedicated inference servers.

228
MCQeasy

Which of the following is considered a best practice for logging in a production LLM environment?

A.Logging only full chat transcripts
B.Using structured JSON logs
C.Logging all data to standard output only
D.Disabling logging to save disk space
AnswerB

JSON-formatted logs are machine-readable and easily parsed by log aggregators like ELK or Splunk. This structure allows for reliable querying of specific fields such as model version, latency, and error codes, which is critical for maintaining production reliability.

Why this answer

Structured logging (typically in JSON format) is the gold standard for production environments because it allows for efficient searching, indexing, and automated analysis. In LLM applications, keeping track of request metadata, model versions, and latency in a structured way is essential for debugging and auditing. This approach enables operations teams to quickly correlate specific issues with model outputs and performance metrics across thousands of requests.

Exam trap

Engineers often rely on unstructured plain-text print statements, assuming human readability is sufficient for large-scale automated pipeline monitoring and analysis.

229
MCQhard

An engineer is fine-tuning a 13B parameter model with NVIDIA NeMo using tensor parallelism across four GPUs. After resuming from a checkpoint, training loss spikes and then diverges. The checkpoint was saved with a different tensor parallel size than the current run. What is the most likely cause of the divergence?

A.The optimizer state and model shards were partitioned differently, so the restored weights do not match the current tensor parallel layout.
B.The learning rate scheduler restarted from step zero and applied a large learning rate to all parameters.
C.Gradient accumulation steps were reduced, effectively increasing the global batch size beyond the original configuration.
D.The data loader random seed was not preserved, causing the model to see samples in a different order.
AnswerA

Tensor parallelism splits weight matrices across GPUs, and the sharding layout depends on tensor parallel size. A checkpoint saved with one tensor parallel degree cannot be directly loaded into a run with a different degree without resharding. The mismatched partitioning causes incorrect weight placement, leading to loss spikes and divergence when training resumes.

Why this answer

Tensor parallelism partitions model weights across GPUs, and the partition layout depends on the tensor parallel size. Loading a checkpoint saved with a different tensor parallel degree without resharding places weights incorrectly, causing loss spikes and divergence. Data order, scheduler restart, and gradient accumulation changes do not explain the immediate failure tied to a changed parallel layout.

Exam trap

The trap here is overlooking that checkpoint compatibility in NeMo depends not only on model architecture but also on the parallelism configuration used when the checkpoint was written.

230
MCQhard

An engineer is evaluating a sparse Mixture-of-Experts decoder-only LLM for a latency-sensitive inference service. They notice that although the model has far more total parameters than a dense baseline, throughput per token is only modestly better and sometimes worse. Which factor best explains why sparse MoE does not translate total parameter count into proportional speedup during inference?

A.Expert parameters are stored in FP32 while dense models use FP16, doubling memory traffic.
B.The router must compute a full softmax over all experts and backpropagate through every expert for each token.
C.Only the top-k experts are activated per token, so memory bandwidth and all-to-all communication dominate rather than raw FLOPs.
D.The attention layers in MoE models are replaced by expert routing, removing the quadratic attention cost.
AnswerC

Sparse MoE activates only a small subset of experts per token, so compute FLOPs are low, but the full expert weights must still reside in memory and be gathered, often across GPUs via all-to-all. Inference therefore becomes memory-bandwidth and communication bound rather than compute bound, which is why adding total parameters does not yield proportional latency improvements.

Why this answer

Sparse MoE activates only a few experts per token, so FLOPs are low, but every expert's weights must remain resident and be routed to, often across devices. Inference becomes bound by memory bandwidth and all-to-all communication rather than compute, so a larger total parameter count does not yield proportional latency gains and can even hurt if routing is imbalanced.

Exam trap

The trap here is equating parameter count with inference cost, when sparse MoE's bottleneck is memory bandwidth and expert communication rather than activated FLOPs.

231
MCQmedium

When deploying a model, what is the benefit of using Triton's 'Model Versioning' feature?

A.It automatically cleans up old model files from the drive.
B.It allows seamless model updates without service restarts.
C.It provides built-in encryption for sensitive model weights.
D.It compiles models into different formats like ONNX and TorchScript.
AnswerB

Versioning enables the server to detect and load new versions of a model dynamically. This allows updates to be pushed to production without interrupting current inference requests, providing the high availability and zero-downtime deployment capabilities required by enterprise-grade AI production environments.

Why this answer

Model versioning allows multiple versions of a model to exist in the repository simultaneously. This enables A/B testing, gradual rollouts, and instant rollbacks. By simply changing the configuration or updating a symbolic link, administrators can shift traffic to a new model version without downtime, ensuring that the service remains available while testing new model iterations or applying hotfixes to production deployments.

Exam trap

Candidates often assume versioning is only for tracking experiments or storage management, failing to realize its critical role in enabling zero-downtime deployments through seamless traffic routing and instant rollbacks.

232
MCQhard

A financial institution is evaluating a fine-tuned GPT-3 model for generating investment advice summaries. They must ensure the model does not produce harmful or biased recommendations. Which evaluation methodology should they implement using NVIDIA NeMo Guardrails and NeMo Evaluator to systematically detect and quantify such issues?

A.Use NeMo Evaluator to compute the F1 score of the model's outputs against a set of correct summaries.
B.Use NeMo Guardrails to define a set of safety policies, then run the model's outputs through the guardrails and measure the percentage of violations flagged.
C.Compute the model's perplexity on a validation set of financial texts.
D.Fine-tune the model further on a dataset of unbiased financial advice and then evaluate with BLEU score.
AnswerB

NeMo Guardrails allows defining programmable rules (e.g., using Colang) to detect and block harmful or biased content. By applying these guardrails to model outputs and quantifying violation rates, the team can systematically evaluate safety and bias. This approach integrates with NeMo Evaluator to log and analyze flagged instances, providing a quantifiable metric for compliance.

Why this answer

To systematically detect and quantify harmful or biased outputs, the team should leverage NeMo Guardrails to encode safety policies and measure violation rates. This provides a direct, quantifiable assessment of safety compliance. Other metrics like perplexity, BLEU, or F1 do not target harmful content and therefore cannot fulfill the evaluation requirement.

Exam trap

The trap here is confusing general performance metrics like perplexity or F1 with safety-specific evaluation, which requires explicit policy checks.

233
MCQmedium

A team is fine-tuning a Llama 2 7B model with NVIDIA NeMo Framework on a single A100 80GB GPU. They observe that training loss decreases initially but then diverges, and the model outputs become repetitive and incoherent. The team used a learning rate of 5e-5 with AdamW and no warm-up. Which change is most likely to stabilize training and improve convergence?

A.Increase the learning rate to 1e-4 to escape local minima.
B.Switch the optimizer to SGD with momentum 0.9 and keep the learning rate at 5e-5.
C.Reduce the learning rate to 1e-5 and add a linear warm-up over the first 10% of training steps.
D.Increase the batch size to 64 and keep all other hyperparameters unchanged.
AnswerC

A lower learning rate (e.g., 1e-5) combined with warm-up prevents large initial updates that can destabilize training. Warm-up gradually increases the learning rate, allowing the model to adapt smoothly. This is a standard practice for fine-tuning large language models, especially when full fine-tuning or using AdamW, and directly addresses the observed divergence and repetitive outputs.

Why this answer

The model diverges due to a learning rate that is too high for fine-tuning without warm-up. Reducing the learning rate to 1e-5 and adding a warm-up phase allows the optimizer to take smaller, more controlled steps initially, preventing large updates that disrupt pretrained weights. This combination is a well-established best practice in NVIDIA NeMo and other frameworks for stable fine-tuning of LLMs.

Exam trap

The trap here is assuming that a higher learning rate always speeds up convergence, when in fine-tuning it often causes divergence and degraded output quality.

234
MCQmedium

A team is deploying a quantized LLM using NVIDIA NIM. To ensure the highest level of security and compliance, they need to verify that the container image has been scanned for vulnerabilities before production use. Which tool is the primary source for certified, production-ready NIM containers?

A.The public NVIDIA Docker Hub repository.
B.The official NVIDIA NGC catalog.
C.A custom build using the standard PyTorch base image.
D.The GitHub repository containing the source code for the model.
AnswerB

NVIDIA NGC is the authoritative source for enterprise-ready containers. Images in this catalog are built, tested, and scanned by NVIDIA, ensuring compatibility with NVIDIA GPUs and adherence to security standards required for deploying LLMs in production environments with strictly defined compliance needs.

Why this answer

NVIDIA NGC (NVIDIA GPU Cloud) is the central repository for certified, secure, and optimized container images. Using NIM containers from the NVIDIA NGC catalog ensures that the images have been scanned by NVIDIA security pipelines. This is critical for enterprise compliance, as it guarantees the software components, including CUDA libraries and inference runtimes, are patched and verified for production deployment.

Exam trap

Candidates often choose public registries like Docker Hub. While accessible, they lack the specific NVIDIA certification and security vetting provided by NGC, which is a requirement for enterprise-grade LLM deployments.

235
MCQmedium

Refer to the exhibit. An engineer observes that GPU memory utilization is high, but the GPU is frequently idling. How does the provided Triton configuration optimize the inference pipeline?

A.It forces the model to use INT8 precision
B.It enables concurrent execution on the GPU
C.It partitions the GPU for dedicated memory
D.It enables dynamic batching for the requests
AnswerB

Setting 'count: 2' instructs Triton to launch two instances of the model on the specified GPU. This allows the server to process multiple requests in parallel, which is a standard method to improve throughput and keep the GPU busy during periods where one instance might be blocked.

Why this answer

By setting the instance count to 2 on a single GPU, the configuration enables concurrent model execution. This allows Triton to schedule multiple inference requests to be processed simultaneously on the same hardware. This overlap helps to hide memory latency and pipeline stalls, effectively increasing the utilization of the GPU and raising the overall throughput of the deployment.

Exam trap

Candidates often misinterpret high VRAM utilization with optimal performance, missing the fact that unutilized compute cycles indicate poor concurrency and request scheduling bottlenecks.

236
MCQeasy

What is the primary function of the 'rank' parameter in LoRA?

A.It sets the total number of layers that are trainable
B.It defines the dimension of the low-rank decomposition
C.It dictates the number of epochs the model will train
D.It determines the learning rate for the adapter weights
AnswerB

The rank 'r' specifies the size of the low-rank matrices. For example, if a weight matrix has dimensions (d, d), LoRA decomposes it into (d, r) and (r, d) matrices. 'r' is typically a small integer, which keeps the parameter count very low compared to full fine-tuning.

Why this answer

The rank parameter ('r') determines the dimensionality of the update matrices injected into the model. A lower rank results in a smaller number of trainable parameters, which is more efficient but less expressive. A higher rank allows for more complex adaptations to the data.

Balancing this rank is essential to achieve the desired model performance while staying within hardware memory limits.

Exam trap

Candidates often incorrectly assume the rank parameter controls the number of layers being trained or the learning rate, rather than identifying it as the dimension of the low-rank decomposition matrices.

237
Multi-Selecthard

A team is deploying a large language model on NVIDIA Triton Inference Server with NVIDIA TensorRT-LLM backend. They need to reduce GPU memory usage to fit a larger model on the same hardware while maintaining acceptable latency. Which two techniques should they use? (Choose two.)

Select 2 answers
A.Set the Triton dynamic batching max_queue_delay_microseconds to 0 to process requests immediately.
B.Quantize the model weights to INT8 or FP8 using TensorRT-LLM's quantization toolkit.
C.Increase the tensor parallel size to 8 across eight GPUs.
D.Enable paged KV cache and in-flight batching in the TensorRT-LLM backend.
E.Use FP32 precision for all layers to avoid quantization errors.
AnswersB, D

Quantizing weights to INT8 or FP8 reduces the memory footprint by 2x to 4x compared to FP16, allowing larger models to fit. TensorRT-LLM supports post-training quantization and quantization-aware training. This directly addresses the memory constraint while maintaining acceptable latency, as quantized kernels are optimized for NVIDIA GPUs. It is a standard technique for memory-constrained deployments.

Why this answer

Quantizing weights to INT8 or FP8 reduces memory footprint significantly, and paged KV cache with in-flight batching optimizes runtime memory usage. Together they allow larger models or longer sequences on the same GPU. Other options either require more hardware, increase memory usage, or affect scheduling rather than memory footprint.

Exam trap

The trap here is thinking that increasing tensor parallel size reduces memory on the same hardware, when it actually requires additional GPUs.

238
MCQmedium

An enterprise is fine-tuning a large language model using NVIDIA NeMo Framework and encounters GPU out-of-memory errors during the backward pass. The training configuration already uses mixed-precision training (FP16). Which architectural intervention should be applied to resolve memory pressure while retaining the optimizer state precision?

A.Switch the global precision setting from FP16 to INT8 quantization.
B.Implement activation checkpointing to recompute intermediate activations during backward passes.
C.Disable gradient accumulation completely to process each micro-batch independently.
D.Migrate the optimizer states from FP32 to FP8 format.
AnswerB

Activation checkpointing selectively saves specific layer activations and recomputes the discarded ones during the backward pass. This drastically reduces peak GPU memory consumption, enabling successful fine-tuning of larger models on NVIDIA hardware without modifying the core precision configuration.

Why this answer

Applying activation checkpointing trades off compute for memory by recalculating activations during the backward pass instead of storing them all. This directly mitigates out-of-memory errors on NVIDIA GPUs during LLM fine-tuning without requiring a reduction in batch size or model accuracy, making it a standard best practice in NeMo Framework workflows.

Exam trap

Candidates often try to resolve backward pass OOM errors by reducing optimizer precision or model layers, missing activation checkpointing as the standard memory-compute trade-off.

239
MCQmedium

A team is training a large language model on a single NVIDIA H100 GPU. They observe that training throughput is significantly lower than expected, and profiling with Nsight Systems shows long periods where the GPU is idle waiting for data. The data loading pipeline uses the default PyTorch DataLoader with num_workers=0 and no pinned memory. Which change is most likely to improve GPU utilization?

A.Use CUDA graphs to capture the training step and reduce kernel launch overhead.
B.Switch from FP32 to TF32 precision for matrix multiplications.
C.Increase the batch size to better saturate the GPU.
D.Enable pinned memory and increase num_workers in the DataLoader.
AnswerD

Pinned memory allows asynchronous host-to-device copies, and multiple worker processes prefetch batches in parallel, overlapping data loading with GPU compute. This directly addresses the idle GPU time observed in profiling. With num_workers=0 and no pinning, data loading is serialized and synchronous, starving the GPU.

Why this answer

The profiling evidence points to a data-loading bottleneck: the GPU sits idle waiting for batches. Enabling pinned memory and increasing DataLoader workers allows asynchronous, overlapped data transfer and prefetching, keeping the GPU fed. The other options target compute or launch overhead, which are not the limiting factors in this scenario.

Exam trap

The trap here is assuming that GPU underutilization always means the model or kernels need optimization, rather than checking whether the input pipeline is starving the device.

240
MCQmedium

Refer to the exhibit. You are reviewing the configuration file for a data preprocessing pipeline. Why is the 'min_words' filter set to 50 in the context of LLM training?

A.It limits the memory footprint of the dataloader.
B.It enables faster tokenization by reducing the number of input files.
C.It removes low-quality snippets to ensure meaningful semantic context.
D.It forces the model to ignore PII-heavy documents.
AnswerC

Removing short sequences filters out noise like headers, footers, or incomplete sentences that provide little linguistic value. This ensures the model spends its training budget on high-quality text, improving its ability to learn complex long-range dependencies and overall coherence within the target language domain.

Why this answer

The 'min_words' filter removes low-information or malformed snippets that lack sufficient context for effective transformer learning. In large-scale training, such as those performed on NVIDIA H100 GPU clusters, including very short strings increases noise and consumes valuable compute cycles without contributing to meaningful semantic representation. Setting a minimum length ensures the model learns from coherent passages rather than disjointed fragments, promoting better structural understanding of the training corpus.

Exam trap

Candidates often assume filtering is purely for privacy, missing the technical reality that very short, low-information snippets introduce noise that degrades the model's ability to learn complex linguistic structures.

241
MCQmedium

You are responsible for the reliability of an LLM inference service running on NVIDIA Triton Inference Server across a fleet of A100 GPUs. The service is deployed with dynamic batching enabled, but during peak hours you observe that end-to-end latency for some requests exceeds the SLO while GPU utilization remains moderate. You suspect that the dynamic batching configuration is causing the issue. Which Triton configuration parameter should you adjust to directly limit the maximum time a request waits in the scheduler queue before being batched?

A.Increase the `instance_group` count to add more model instances.
B.Set the `max_queue_delay_microseconds` parameter to a lower value.
C.Set the `max_batch_size` parameter to a lower value.
D.Enable the `priority_levels` parameter to prioritize certain requests.
AnswerB

The `max_queue_delay_microseconds` parameter in Triton's dynamic batching configuration specifies the maximum time a request can wait in the queue before the scheduler dispatches the batch. By lowering this value, you reduce the maximum latency contributed by batching, which directly addresses the observed SLO violations. This is the intended parameter for controlling batching-induced latency while still allowing some batching for efficiency.

Why this answer

The `max_queue_delay_microseconds` parameter in Triton's dynamic batching configuration sets the maximum time a request can wait in the scheduler queue before being dispatched for inference. Lowering this value reduces the worst-case latency introduced by batching, directly addressing the SLO violations observed during peak hours. While other parameters like `max_batch_size` and `instance_group` affect performance, only `max_queue_delay_microseconds` explicitly bounds the queue wait time, making it the correct adjustment for this scenario.

Exam trap

The trap here is assuming that reducing `max_batch_size` will directly reduce latency, when the real control for queue wait time is `max_queue_delay_microseconds`.

242
Multi-Selectmedium

An engineer is tuning a TensorRT-LLM deployment of a 7B model for a latency-sensitive API. Profiling shows that time per output token is higher than expected and that many small kernels run back to back with gaps between them. Which TWO changes are most likely to reduce the per-token latency by cutting kernel launch overhead and redundant memory traffic? (Choose two.)

Select 2 answers
A.Enable a debug synchronization after every kernel so profiling timings are accurate.
B.Enable CUDA Graphs so the decoding iteration is captured and replayed as a single graph launch.
C.Use TensorRT-LLM fused multi-head attention and fused GEMM plugins instead of separate elementwise kernels.
D.Raise the maximum sequence length in the engine build to accommodate the longest possible prompt.
E.Switch the model weights to FP32 to avoid any dequantization kernels in the decode path.
AnswersB, C

CUDA Graphs record the whole sequence of kernels in a decoding step and replay them with one launch, removing per-kernel launch latency and the CPU-side gaps visible in the profile. For a latency-sensitive decode loop with many small kernels, this directly reduces time per output token without changing numerics or requiring retraining.

Why this answer

The profile shows many short kernels with idle gaps, which is the classic signature of launch overhead and unfused elementwise work. CUDA Graphs collapse the decode iteration into a single replayable launch, and fused attention and GEMM plugins merge operations that would otherwise round-trip intermediates through HBM. Together they reduce both launch count and memory traffic, which are the two costs identified in the profile.

Exam trap

The trap here is chasing sequence length or precision settings when the profile clearly shows launch-bound behavior that graph capture and kernel fusion are designed to eliminate.

243
Multi-Selecthard

An enterprise is fine-tuning a 34B model with NVIDIA NeMo Framework and observes that the validation loss begins rising after the first epoch while training loss continues to fall. The team wants to reduce this divergence and preserve downstream task quality. (Choose two.)

Select 2 answers
A.Apply weight decay or increase its coefficient to penalize large parameter updates and discourage overfitting.
B.Remove the validation split and train on all available data so the model sees more examples per epoch.
C.Raise the learning rate significantly so the model converges faster and reaches a flatter minimum.
D.Increase the number of training epochs so the optimizer has more time to escape the overfitting region.
E.Enable early stopping based on validation loss so training halts before the model overfits further.
AnswersA, E

Weight decay regularizes the model by penalizing large weights, which reduces the tendency to memorize training-specific patterns. Increasing its coefficient when validation loss diverges is a targeted regularization response that can narrow the gap between training and validation performance while preserving the learned task behavior.

Why this answer

Rising validation loss with falling training loss is classic overfitting. Early stopping on validation loss halts training before further damage, while stronger weight decay regularizes the model and narrows the generalization gap. Together they reduce divergence and protect downstream task quality without discarding useful training signal.

Exam trap

The trap here is interpreting a rising validation loss as a need for more training or a faster learning rate, when the divergence actually calls for regularization and earlier stopping.

244
Multi-Selecthard

You are preparing a massive dataset for training a NeMo-based LLM. Which TWO data preprocessing steps are critical to prevent data leakage and ensure model quality?

Select 2 answers
A.Performing temporal or content-based splitting to isolate the validation set.
B.Implementing fuzzy deduplication to remove near-duplicate documents.
C.Increasing the frequency of data augmentation in the training pipeline.
D.Adding metadata tags to all training sequences for classification.
E.Converting all text to lowercase to increase vocabulary efficiency.
AnswersA, B

Isolating data based on timestamps or content clusters prevents the model from seeing future data or overlapping information during training. This ensures the evaluation set remains truly unseen, providing a realistic assessment of how the model will perform on new, unseen data in production environments.

Why this answer

Data leakage occurs when test data is inadvertently included in the training set, leading to inflated performance metrics. Deduplication is equally vital, as repetitive data causes the model to memorize samples rather than generalize. In NVIDIA workflows, these steps are typically performed via distributed scripts on the cluster before tokenization, ensuring that the model learns unique, non-overlapping information across all shards.

Exam trap

Test-takers often confuse basic random splitting with temporal or content-based splitting, missing the fact that standard random splits fail to prevent data leakage in LLM datasets.

245
MCQeasy

Which of the following describes the purpose of 'Kernel Fusion' in the context of optimizing a Deep Learning inference pipeline?

A.To increase the number of parallel GPU threads.
B.To reduce redundant global memory read/write cycles.
C.To enable multi-GPU distributed training.
D.To improve model accuracy through extra precision.
AnswerB

Kernel fusion minimizes global memory traffic by keeping intermediate results in registers or shared memory. By avoiding writing intermediate tensors back to VRAM, the pipeline becomes significantly faster, as reading from and writing to high-latency VRAM is the primary bottleneck for many AI inference tasks.

Why this answer

Kernel fusion combines multiple small operations into a single GPU kernel to reduce the overhead of launching kernels and accessing global memory. Every kernel launch involves CPU-side overhead, and global memory accesses are costly in terms of energy and time. Fusion minimizes both, significantly increasing the effective throughput of the GPU by keeping data in high-speed, on-chip storage for as long as possible.

Exam trap

Candidates often think kernel fusion increases parallel thread execution count, confusing instruction-level merging with hardware scaling.

246
MCQmedium

An engineer is deploying a 70B parameter model and needs to serve many concurrent users on a single GPU with limited memory. They want to store the attention keys and values for past tokens efficiently so that generation does not recompute them at every step. Which technique should they implement?

A.Quantization-aware training that reduces weight precision before deployment.
B.A key-value cache that stores the projected keys and values for previously processed tokens.
C.Speculative decoding that uses a smaller draft model to propose tokens.
D.Gradient checkpointing that recomputes activations during the backward pass.
AnswerB

A key-value cache retains the key and value projections for all past tokens so each new step only computes the query and attends to cached entries. This avoids recomputing the full prefix at every decoding step, dramatically reducing compute for autoregressive generation.

Why this answer

Autoregressive decoding generates one token at a time, and without caching the model would recompute keys and values for the entire prefix at every step, wasting compute. A key-value cache stores those projections so each step only processes the new token, which is the standard way to make long-context generation practical.

Exam trap

The trap here is treating training-time memory tricks like gradient checkpointing as if they accelerated autoregressive inference.

247
MCQmedium

What is the primary function of the 'TensorRT' optimization engine in the NVIDIA AI software stack?

A.It automates the training of complex models.
B.It provides a Python API for model debugging.
C.It performs architecture-specific model optimization.
D.It converts models to run on mobile CPUs.
AnswerC

TensorRT performs deep optimizations like layer fusion, kernel auto-tuning, and precision reduction, all tailored to the specific GPU architecture being used. This allows the model to run at peak throughput and minimal latency by taking advantage of the unique features of the target NVIDIA hardware architecture.

Why this answer

TensorRT optimizes neural network models by performing layer fusion, precision calibration (e.g., to FP8 or INT8), and kernel selection optimized for the specific GPU architecture. By transforming the model into a highly efficient, platform-specific format, TensorRT significantly reduces latency and increases throughput for production inference. It is the core tool for moving from research-grade PyTorch models to production-ready deployments on NVIDIA hardware.

Exam trap

Candidates often mistake TensorRT for a general-purpose library for model training or data preprocessing, ignoring its specific role as an inference-time optimization engine for NVIDIA hardware.

248
MCQmedium

Why is it important to perform 'domain-specific' data cleaning when preparing a corpus for fine-tuning a medical LLM?

A.It helps to anonymize the patient data by replacing all medical terms with generic labels.
B.Generic cleaning often removes or modifies critical domain-specific nomenclature.
C.It ensures that the dataset size is significantly reduced to fit in GPU cache.
D.It automatically corrects all scientific inaccuracies present in the original documents.
AnswerB

Medical nomenclature is highly specialized and often uses abbreviations or symbols that generic cleaning scripts might identify as noise or formatting errors. By using domain-specific cleaning, developers ensure that these critical terms are preserved, which is essential for maintaining the accuracy of the model's domain knowledge.

Why this answer

Medical text contains specific jargon, abbreviations, and relationships that generic cleaning might misinterpret. Generic tools often remove or alter terms that are critical to medical context, such as drug names or procedural codes. By using domain-specific cleaning, you ensure that the model retains the precise terminology necessary for high-stakes, accurate clinical reasoning in medical applications.

Exam trap

Candidates often assume that generic cleaning tools are sufficient for all data types, overlooking the fact that medical terminology is fragile and easily destroyed by standard normalization or stop-word removal.

249
MCQeasy

A team is deploying a large language model on NVIDIA Triton Inference Server in a production environment. They need to ensure that the model server can automatically recover from GPU failures without manual intervention. Which feature of Triton should they configure to achieve this?

A.Triton's health check endpoints and Kubernetes liveness probes
B.Instance groups with multiple GPU instances
C.Model repository polling
D.Model ensembles
AnswerA

Triton exposes health check endpoints that Kubernetes can use with liveness probes to detect when the server is unhealthy due to GPU failure. Kubernetes can then restart the pod, potentially on a healthy GPU node, providing automatic recovery. This combination is a standard practice for achieving high availability in containerized Triton deployments.

Why this answer

To automatically recover from GPU failures, Triton should be deployed in an orchestrated environment like Kubernetes. Triton provides health check endpoints that, when integrated with Kubernetes liveness probes, allow the orchestrator to detect an unhealthy server and restart it, possibly on a different node with a working GPU. This approach ensures minimal downtime without manual intervention.

Exam trap

The trap here is assuming that Triton's internal features like instance groups or ensembles provide automatic failover, when actually orchestration-level health checks are required for recovery.

250
MCQmedium

In the context of NVIDIA NeMo, why is it recommended to use FP8 precision during the fine-tuning process on H100 GPUs?

A.It guarantees 100% precision parity with full-precision floating-point training.
B.It significantly improves memory throughput and speed via Hopper-specific hardware.
C.It disables the need for gradient scaling during the backward pass.
D.It forces the model to use only CPU-based memory for training weights.
AnswerB

NVIDIA H100 GPUs include specialized hardware support for FP8, which accelerates matrix multiplication and reduces memory footprint. This allows the model to process data much faster than traditional precision formats, providing a competitive advantage in training workflows that require high-performance compute and rapid turnaround times.

Why this answer

FP8 precision leverages the Transformer Engine on NVIDIA Hopper architecture, providing a significant boost in throughput while reducing memory usage. By using FP8, developers can fit larger models or increase batch sizes without sacrificing significant numerical stability. This optimization is essential for modern fine-tuning workflows where computational efficiency directly impacts the speed of iteration and the scalability of training across multi-GPU nodes in an enterprise environment.

Exam trap

Candidates assume FP8 is universally supported across all GPUs, failing to recognize that it requires specific hardware features like the Hopper architecture's Transformer Engine.

251
MCQmedium

A team is deploying a 13B-parameter chatbot on a single NVIDIA A10G GPU (24 GB VRAM). The model's weights are stored in FP16, and the runtime runs out of memory during KV cache allocation under concurrent user sessions. They must keep answer quality essentially unchanged while maximizing concurrent sessions. Which optimization should they apply first?

A.Convert the checkpoint to TF32 and rebuild the TensorRT-LLM engine with builder optimization level 5.
B.Apply INT8 weight-only quantization to the linear layers with NVIDIA TensorRT-LLM, keeping activations in FP16.
C.Increase the paged KV cache block size from 16 to 128 tokens per block and disable block reuse.
D.Enable CUDA graph capture for the decoder and set the maximum batch size to the peak observed concurrency.
AnswerB

Weight-only INT8 quantization halves the memory consumed by the model's linear-layer weights (roughly 26 GB FP16 becomes ~13 GB INT8) while leaving activation precision untouched, which preserves output quality closely. The freed VRAM is then available for KV cache blocks, directly lifting the concurrent session ceiling on the 24 GB A10G.

Why this answer

The bottleneck is persistent weight memory (FP16 weights dominate VRAM), which leaves too little room for paged KV cache blocks under concurrency. Weight-only INT8 quantization in TensorRT-LLM roughly halves the linear-layer weight footprint while keeping activations in FP16, preserving quality. That reclaimed VRAM goes directly to KV cache, raising the number of simultaneous sessions without retraining or a new GPU.

Exam trap

The trap here is assuming that any precision reduction degrades quality unacceptably, when weight-only INT8 with FP16 activations is specifically designed to preserve accuracy while cutting the dominant memory consumer.

252
MCQhard

Refer to the exhibit. The engineer is attempting to deploy on an NVIDIA Orin platform but encounters a runtime error. What is the most likely cause of the failure?

A.The workspace memory is too large
B.The DLA hardware is not enabled on this device
C.Unsupported operators are being targeted for the DLA
D.FP16 precision is not supported on DLA
AnswerC

The DLA is a specialized hardware accelerator with limited operator support compared to the GPU. LLMs often use complex or custom operations that the DLA cannot execute. Forcing these operations onto the DLA via the CLI flags will cause the builder to fail because it cannot map the graph.

Why this answer

The inclusion of '--dla 0' forces the engine to run on the Deep Learning Accelerator (DLA) core. Many complex LLM operations, such as specific activation functions or advanced attention mechanisms, are not supported by the DLA's fixed-function logic. If the model graph contains unsupported operators, the build will either fail or generate a non-functional plan, as the DLA has a more restricted operator set than the primary GPU cores.

Exam trap

Candidates often assume the error is a general memory or driver issue, failing to check if the specific operations in the model graph are actually supported by the DLA hardware architecture.

253
MCQeasy

In the context of the Transformer architecture, what is the primary function of the Feed-Forward Network (FFN) layers applied after the attention mechanism?

A.To calculate the attention scores between different tokens in the sequence.
B.To process the hidden state representations with non-linear activations.
C.To manage the memory allocation for the KV cache during multi-token generation.
D.To reduce the sequence length of the input tokens to a fixed size.
AnswerB

The FFN typically consists of two linear transformations with a non-linear activation function, like SwiGLU or ReLU, in between. This structure enables the model to learn complex mappings of input features, significantly increasing its capacity to understand nuances that linear attention projections alone might fail to capture effectively.

Why this answer

The FFN layers provide non-linear transformations that allow the model to process information extracted by the attention mechanism. While attention focuses on relationships between tokens, the FFN applies point-wise non-linearities to project these features into higher-dimensional spaces. This is essential for learning complex representations and mappings that ultimately drive the model's predictive accuracy and reasoning capabilities across diverse input patterns.

Exam trap

Candidates often confuse the FFN's role with the attention mechanism's role. They incorrectly attribute the 'capturing of relationships between tokens' to the FFN, rather than the attention mechanism.

254
MCQmedium

An engineer is designing a Transformer-based model for long-context document summarization. They decide to replace the standard dense self-attention mechanism with a sliding window attention approach. What is the primary architectural implication of this change?

A.The model's total parameter count increases due to added positional embedding layers.
B.The model loses the ability to perform parallel training across multiple GPU nodes.
C.The computational complexity of the self-attention mechanism is reduced from quadratic to linear.
D.The model is no longer compatible with standard softmax normalization functions.
AnswerC

By limiting the attention span to a fixed window size, the number of operations per token becomes constant rather than proportional to the sequence length. This shift from O(n²) to O(n*w) complexity is the fundamental architectural advantage for long-context tasks, enabling processing of documents that would otherwise be computationally prohibitive.

Why this answer

Sliding window attention restricts the receptive field of each token to a local neighborhood, drastically reducing the quadratic memory complexity of standard attention to linear. This is critical for scaling LLMs to long contexts, as it prevents the O(n²) memory growth that typically causes GPU out-of-memory errors on large input sequences while maintaining local coherence.

Exam trap

Candidates often focus on the 'loss of accuracy' or 'semantic degradation.' While these are concerns, the question specifically asks for the architectural implication regarding computational complexity.

255
MCQmedium

An engineer is deploying a 13B-parameter LLM with TensorRT-LLM on a single NVIDIA A100 40GB GPU. The FP16 engine requires 26GB for weights, but during generation the KV cache grows beyond remaining memory, causing out-of-memory errors. The team wants to maximize concurrent requests without retraining. Which optimization should they apply first?

A.Reduce the KV cache memory footprint by quantizing the KV cache to INT8 while keeping model weights in FP16.
B.Increase the batch size and sequence length limits so the scheduler can pack more requests into the same memory.
C.Enable tensor parallelism across two A100 GPUs to split both weights and KV cache across devices.
D.Convert the model weights from FP16 to FP8 using post-training quantization to free memory for the KV cache.
AnswerA

Quantizing the KV cache to INT8 halves its per-token memory, directly addressing the growth that causes OOM during generation while leaving weights untouched, so no retraining is needed. TensorRT-LLM supports INT8 KV cache with FP16 weights, and it preserves accuracy better than aggressively quantizing weights. This is the least invasive change that increases concurrent request capacity on the constrained A100.

Why this answer

The KV cache grows linearly with sequence length and batch size, and it is the dominant memory consumer during LLM generation after weights are loaded. Quantizing only the KV cache to INT8 halves its footprint without retraining and without changing weight precision, directly relieving the OOM. This keeps the model on the existing A100 while allowing more concurrent requests, which matches the stated goal.

Exam trap

The trap here is assuming that weight quantization is the only lever for memory reduction, when KV cache quantization is often the more targeted fix for generation-time OOM.

256
MCQmedium

A team has deployed a Llama-3-8B model with NVIDIA Triton Inference Server for real-time chatbot responses. They need to evaluate the model's generation quality on a held-out test set of 500 prompts, focusing on semantic similarity to reference answers while ensuring low latency. Which evaluation approach is most appropriate to use with NVIDIA NeMo Evaluator?

A.Calculate perplexity on the test set using the model's own logits.
B.Use BERTScore or a similar embedding-based metric via NeMo Evaluator's custom metric interface.
C.Compute BLEU score using the sacreBLEU library integrated into NeMo Evaluator.
D.Run ROUGE-L scoring on the generated outputs and references.
AnswerB

BERTScore leverages contextual embeddings to compute semantic similarity between generated and reference texts, aligning with the requirement to assess meaning rather than surface form. NeMo Evaluator supports custom metrics, allowing integration of BERTScore. This approach handles paraphrases well and is efficient for 500 prompts, meeting the low-latency evaluation goal.

Why this answer

For evaluating semantic similarity in generated text, embedding-based metrics like BERTScore are preferred because they capture meaning beyond exact wording. NeMo Evaluator's extensibility allows incorporating such metrics, providing a more human-aligned assessment for chatbot responses. Perplexity and n-gram overlap metrics like BLEU and ROUGE-L do not adequately measure semantic equivalence, especially when paraphrasing is expected.

Exam trap

The trap here is assuming that any standard metric like BLEU or ROUGE automatically reflects semantic quality, when they actually measure lexical overlap.

257
MCQhard

A generative AI application built on NVIDIA Triton Inference Server is deployed in a Kubernetes cluster with GPU nodes. The operations team wants to detect silent data corruption in model outputs, which could occur due to GPU memory errors. They plan to implement a monitoring solution using NVIDIA Data Center GPU Manager (DCGM). Which DCGM feature should they enable to detect and alert on GPU memory errors that could lead to silent data corruption?

A.DCGM configuration with `EnableAutoBoost` set to true.
B.DCGM health checks with the `memory` watch.
C.DCGM profiling metrics with `DCGM_FI_DEV_GPU_UTIL`.
D.DCGM diagnostics with the `-r` option for a full run.
AnswerB

DCGM health checks include a `memory` watch that monitors GPU memory for errors such as ECC errors (correctable and uncorrectable). Enabling this watch allows DCGM to detect memory errors that could cause silent data corruption. It can trigger alerts or take corrective actions based on policy. This is the appropriate feature to monitor for GPU memory issues that may affect model output integrity, making it the correct choice for detecting silent data corruption.

Why this answer

DCGM health checks with the `memory` watch continuously monitor GPU memory for ECC errors and other faults that can cause silent data corruption. This feature can be configured to raise alerts when errors exceed thresholds, enabling proactive remediation. In contrast, diagnostics are run on-demand, utilization metrics reflect workload, and AutoBoost affects clocks.

For detecting memory errors that could corrupt LLM outputs, the memory watch is the correct DCGM feature to enable.

Exam trap

The trap here is confusing DCGM diagnostics (run manually) with health checks (continuous monitoring) and assuming that any memory-related test will provide real-time alerting.

258
MCQhard

A research team is fine-tuning a model with NVIDIA NeMo and wants to reduce the risk of catastrophic forgetting of general capabilities while still adapting to a specialized domain. They have a small domain dataset and limited compute. Which fine-tuning approach best balances domain adaptation with retention of pretrained knowledge?

A.Full-parameter fine-tuning with a high learning rate for many epochs on the domain dataset.
B.Freezing all layers and training only the final classification head on the domain dataset.
C.Parameter-efficient fine-tuning with LoRA using a modest rank and a low learning rate.
D.Training from scratch on the domain dataset using the same architecture and tokenizer.
AnswerC

LoRA freezes the pretrained weights and trains small low-rank adapters, which limits drift from the original model and reduces forgetting. A modest rank and low learning rate further constrain updates, making it well suited for small datasets and limited compute. This approach balances domain adaptation with retention of general capabilities.

Why this answer

LoRA with a modest rank and low learning rate freezes the base model and trains small adapters, which constrains updates and preserves pretrained knowledge while still adapting to the domain. Full fine-tuning with high learning rates, training from scratch, or freezing all layers and training only a head either cause forgetting, require excessive resources, or fail to adapt the generative model.

Exam trap

The trap here is equating strong domain adaptation with full fine-tuning, when a small dataset and limited compute make parameter-efficient methods both safer and more practical.

259
MCQhard

Refer to the exhibit. Which adjustment is the most immediate and effective way to resolve this OOM error while maintaining the same training architecture?

A.Increase the sequence length of the inputs.
B.Decrease the batch size per GPU.
C.Switch from mixed precision to full FP32 precision.
D.Disable the optimizer state checkpointing.
AnswerB

Decreasing the batch size is the most effective way to reduce memory consumption immediately. It lowers the VRAM overhead required for activation storage, allowing the model to fit within the existing hardware constraints. This is the standard approach to resolving OOM errors during the fine-tuning training process.

Why this answer

The error indicates that the current batch size is too large for the available VRAM. Reducing the batch size is the most direct solution to free up enough memory for the model to continue training. If a larger effective batch size is required for convergence, the user can subsequently enable gradient accumulation to compensate for the reduced per-step batch size without increasing the memory footprint.

Exam trap

Candidates often try to change model architecture or hardware settings first, ignoring that reducing the batch size is the most immediate and effective way to resolve OOM errors.

260
MCQhard

Refer to the exhibit. What is the most likely cause of the failure based on the log entries?

A.The model weights are corrupted in the repository.
B.The KV cache size is exceeding the allocated GPU memory limit.
C.Network latency is causing the client-side timeout.
D.The GPU driver version is incompatible with TensorRT.
AnswerB

LLMs require a significant amount of memory for the KV cache to store key-value pairs of previous tokens. When the sequence length or batch size grows too large for the allocated memory, the system exhausts VRAM, leading to performance degradation, evictions, and eventual OOM termination during inference.

Why this answer

The logs indicate high processing latency followed by KV cache eviction and finally an OOM error. This sequence suggests that the system is running out of VRAM due to the growing KV cache during long-context generation. As the context length increases, the memory required for the KV cache exceeds the available capacity, forcing evictions, slowing down processing, and ultimately triggering an Out of Memory crash.

Exam trap

Candidates often misidentify the error as a general memory leak or a model weight loading issue, failing to correlate the specific symptom of KV cache eviction with long-context generation demands.

261
MCQhard

An engineer is using an NVIDIA NIM for a Mixtral model to extract structured data from invoices. The model occasionally returns fields with the wrong data type, such as a numeric amount as a string. The team wants a prompt engineering fix that does not require changing the model or adding a separate parser. Which approach is most effective?

A.Instruct the model to output the data in a natural language paragraph and then extract the fields manually.
B.Ask the model to double-check its output for type correctness before returning it.
C.Lower the temperature to 0.0 and increase the repetition penalty to discourage type mistakes.
D.Describe the desired output schema in the system prompt, including field names and expected data types, and provide one fully formatted example.
AnswerD

Specifying the schema with explicit data types and showing a complete example gives the model a precise template to follow. The example demonstrates the exact format, including numeric values without quotes. This combination is the most effective prompt-level fix because it removes ambiguity about both field names and types, reducing type errors without external parsing.

Why this answer

An explicit schema with field names and data types, paired with a complete example, gives the model an unambiguous template. The example shows numeric values without quotes, so the model imitates the correct types. Self-checking, sampling parameters, and natural language output do not provide the concrete structural guidance needed to eliminate type errors at the prompt level.

Exam trap

The trap here is relying on self-checking or sampling parameters to fix data type errors, when the prompt never defined the expected types in the first place.

262
MCQmedium

An engineer is using TensorRT-LLM to serve a chatbot model. They observe that the time to first token (TTFT) is high, but subsequent tokens are generated quickly. Which optimization should they prioritize to reduce TTFT?

A.Enable in-flight batching to increase GPU utilization.
B.Optimize the prefill phase with kernel fusion and faster attention.
C.Increase the KV cache size to avoid evictions.
D.Use a larger batch size for the prefill phase.
AnswerB

TTFT is dominated by the prefill phase, where the entire prompt is processed. Optimizing this phase with fused kernels, faster attention implementations like FlashAttention, and reducing memory overhead directly cuts TTFT. TensorRT-LLM provides optimized prefill kernels and supports FlashAttention, making this the most effective approach.

Why this answer

Time to first token is primarily determined by the prefill phase, which processes the input prompt. Optimizing this phase with fused kernels, efficient attention mechanisms, and reduced memory latency directly reduces TTFT. TensorRT-LLM includes highly optimized prefill kernels and supports FlashAttention, which significantly speeds up the prefill computation.

Other options target throughput or cache management, not TTFT.

Exam trap

The trap here is conflating throughput optimizations like in-flight batching with latency reductions for the first token, which are distinct performance metrics.

263
MCQmedium

When fine-tuning a Large Language Model using Low-Rank Adaptation (LoRA), which architectural component is primarily modified to reduce computational overhead while maintaining performance?

A.The entire embedding layer matrix
B.The feed-forward network activation functions
C.The attention weight projection matrices
D.The entire decoder hidden state output
AnswerC

LoRA injects trainable low-rank matrices into the attention mechanism's query, key, and value projections. By adapting these specific components, the model learns to capture domain-specific patterns without updating the billions of frozen parameters, drastically reducing VRAM consumption and making model fine-tuning feasible on limited NVIDIA hardware resources.

Why this answer

LoRA freezes the pre-trained model weights and injects trainable rank decomposition matrices into the transformer architecture layers. By targeting specifically the attention query, key, and value projection matrices, practitioners can achieve high performance with a fraction of the trainable parameters. This approach is critical for memory-constrained environments, allowing fine-tuning on consumer-grade NVIDIA GPUs while avoiding the massive memory requirements associated with full parameter updates.

Exam trap

Candidates often mistakenly believe LoRA modifies the entire transformer block, failing to identify that it specifically targets the attention weight projection matrices to optimize memory and computational efficiency.

264
MCQhard

When optimizing a Generative AI model using NVIDIA TensorRT-LLM, which component is primarily responsible for managing the KV cache to minimize memory fragmentation?

A.The TensorRT Builder
B.The Paged KV Cache Manager
C.The Model Parallelism Engine
D.The Triton Inference Server
AnswerB

The Paged KV Cache Manager is specifically designed to handle the dynamic growth of sequences in LLMs. By allocating non-contiguous memory blocks, it prevents the fragmentation that would otherwise limit the ability of the system to handle multiple concurrent requests with varying sequence lengths during generation.

Why this answer

TensorRT-LLM utilizes PagedAttention or similar paged memory management techniques to handle the Key-Value (KV) cache. By dividing the cache into fixed-size blocks, the framework avoids the fragmentation issues inherent in contiguous memory allocation. This ensures that the memory is used efficiently during autoregressive generation, allowing for larger context windows and higher concurrent user capacity on the GPU.

Exam trap

Students frequently confuse general GPU memory allocators with specialized LLM components, mistakenly choosing standard cache options instead of paged memory managers designed for attention states.

265
MCQhard

Refer to the exhibit. What is the impact of this filter on the training corpus?

A.It limits the vocabulary size to 20% of the original content.
B.It discards documents that are overly repetitive or lack linguistic richness.
C.It forces the model to use 20% more computation for tokenization.
D.It reduces the training corpus size by exactly 20%.
AnswerB

A low lexical diversity score indicates that a document uses a very small set of unique words relative to its total length. This is characteristic of repetitive or low-quality content. Filtering these out ensures the model learns from diverse, high-quality, and informative text, which improves overall model performance.

Why this answer

This filter removes documents with low lexical diversity, which often contain repetitive, low-value, or 'boilerplate' text. Such content provides little signal for the model to learn meaningful language patterns. By enforcing a minimum diversity threshold, you ensure the corpus consists of richer, more informative language, which typically leads to better convergence and higher quality output in the trained model.

Exam trap

Test-takers often misinterpret lexical diversity filters as removing long documents or high-frequency vocabulary, whereas they actually target repetitive, low-value boilerplate text.

266
MCQmedium

A company needs to deploy a generative AI model that will serve prompts containing regulated customer data. Security policy requires that all inference stays on-premises, that the model be quantized to fit existing GPUs, and that no external network calls occur at runtime. Which deployment approach should the engineer choose?

A.Run the model locally but allow it to fetch updated weights from a vendor endpoint at startup
B.Deploy the model in a public cloud region with a private VPC and customer-managed encryption keys
C.Deploy the model with NVIDIA NIM microservices in an air-gapped environment using locally hosted containers from NVIDIA NGC
D.Use a public hosted API for the model and rely on TLS encryption for data in transit
AnswerC

NVIDIA NIM microservices can be pulled as containers and run fully on-premises, including in air-gapped environments, which satisfies the no-external-calls requirement. They support quantized model variants and expose standard APIs, so the team gets a deployable, optimized inference service without relying on any cloud endpoint.

Why this answer

NVIDIA NIM microservices are distributed as containers that can be deployed on-premises or in air-gapped environments, support quantized model variants, and do not require external network calls at runtime. This satisfies the on-premises, quantized, and no-external-calls requirements simultaneously, unlike hosted or cloud-based options.

Exam trap

The trap here is treating transport encryption or VPC isolation as equivalent to on-premises execution, when the policy specifically forbids external inference calls.

267
MCQmedium

A team is training a 13B-parameter LLM on 8 NVIDIA A100 GPUs using NVIDIA NeMo. They observe that the all-reduce communication during data-parallel training consumes nearly 40% of each iteration. Which of the following changes is most likely to reduce this communication overhead while preserving convergence?

A.Switch from data parallelism to tensor parallelism across all 8 GPUs.
B.Enable gradient accumulation with a larger micro-batch size and use NCCL with tree algorithm.
C.Use NVIDIA NCCL with the ring algorithm and overlap communication with computation via gradient bucketing.
D.Reduce the number of GPUs to 4 and increase the per-GPU batch size.
AnswerC

NCCL's ring algorithm is optimized for large messages and high-bandwidth interconnects like NVLink, making it efficient for all-reduce in data-parallel training. Overlapping communication with computation using gradient bucketing hides latency behind backpropagation. Together, these reduce the effective communication overhead without changing model convergence, directly addressing the observed bottleneck.

Why this answer

The communication bottleneck in data-parallel training is best mitigated by using an efficient collective algorithm and hiding latency. NCCL's ring algorithm excels with large messages on high-bandwidth links, and gradient bucketing enables overlap with compute. This preserves the data-parallel semantics and convergence while reducing wall-clock time spent in all-reduce.

Exam trap

The trap here is assuming that any parallelism change (like tensor parallelism) will reduce communication, when it often increases it due to more frequent synchronization.

268
MCQmedium

What is the primary function of the 'Triton Model Control' API in a production environment?

A.To optimize GPU memory frequency
B.To load and unload models dynamically
C.To adjust the model's architecture
D.To capture raw network packets
AnswerB

The primary purpose of this API is to enable runtime control over model loading and unloading. This capability is critical for production systems that need to update models or deploy new versions without causing service interruptions for active end users.

Why this answer

The Model Control API allows for dynamic model management without requiring a full server restart. This is essential for zero-downtime deployments, where new model versions must be loaded, tested, or swapped out while the server continues to handle existing inference traffic. By decoupling model management from server availability, the API significantly increases the flexibility and operational reliability of large-scale, continuously evolving AI model deployments.

Exam trap

Candidates often assume the API is for monitoring model performance or accuracy. They overlook the word 'Control' and fail to associate it with the dynamic lifecycle management of model instances.

269
MCQmedium

A team is fine-tuning a Llama 3 8B model with NVIDIA NeMo on a single A100 80GB GPU. They observe that validation loss starts to rise while training loss continues to decrease after epoch 2. They want to keep the best generalizing checkpoint without changing the dataset. Which NeMo training configuration strategy should they apply?

A.Reduce the validation split size so that validation loss more closely tracks training loss.
B.Enable early stopping based on validation loss and save the best checkpoint using the validation metric.
C.Increase the number of training epochs and rely on the final checkpoint for deployment.
D.Switch the optimizer from AdamW to SGD with a higher momentum value.
AnswerB

Early stopping monitors validation loss and halts training when it stops improving, preserving the checkpoint with the best validation metric. In NeMo, this is configured through the checkpointing and early stopping callbacks. Since the scenario shows overfitting after epoch 2, stopping at the best validation point keeps generalization without altering data or model architecture.

Why this answer

The divergence between decreasing training loss and increasing validation loss indicates overfitting. The correct response is to use validation-based early stopping and retain the best checkpoint, which NeMo supports through its checkpointing and early stopping mechanisms. Increasing epochs, changing optimizers, or shrinking validation data do not address the core generalization problem.

Exam trap

The trap here is assuming that more training or a different optimizer will fix overfitting when the validation curve already shows the model is past its best generalization point.

270
MCQhard

An LLM inference service on NVIDIA Triton Inference Server is experiencing a gradual increase in P99 latency over several hours without a corresponding increase in request rate. GPU utilization remains stable at around 60%. Which action should the team take first to diagnose the root cause?

A.Increase the number of model instances to handle the load.
B.Check for memory leaks in the model or Triton server process.
C.Profile the inference pipeline with NVIDIA Nsight Systems to identify bottlenecks.
D.Review the model's input data for changes in sequence length.
AnswerB

A gradual increase in latency without increased load or GPU utilization suggests a resource leak, such as memory fragmentation or a memory leak in the model or server. Over time, this can cause more frequent garbage collection or swapping, increasing latency. Checking for memory leaks is a logical first step to identify the root cause before applying fixes.

Why this answer

The gradual increase in P99 latency with stable GPU utilization and request rate suggests a resource leak, such as memory fragmentation or a leak in the model or Triton process. Checking for memory leaks is the most direct first step to diagnose the issue. Scaling or profiling may be premature and could mask the root cause.

Reviewing input data is unlikely given the stable utilization.

Exam trap

The trap here is assuming that latency increases always require scaling, without considering resource leaks that don't affect utilization.

271
Multi-Selectmedium

An engineer is optimizing an LLM for inference with NVIDIA TensorRT-LLM and wants to reduce both memory footprint and latency without retraining the model. Which two techniques should they apply? (Choose two.)

Select 2 answers
A.Quantize model weights to INT8 or FP8 using post-training quantization with calibration.
B.Enable paged KV cache with block-based memory management.
C.Disable attention and rely solely on feed-forward layers to speed up decoding.
D.Increase the number of transformer layers to improve parallelism across the GPU.
E.Retrain the model with quantization-aware training to learn optimal scaling factors.
AnswersA, B

Post-training weight quantization reduces the bytes per parameter, cutting memory footprint and enabling faster low-precision GEMM kernels on supported GPUs. It requires only calibration data, not retraining, which matches the constraint. This directly lowers both memory and latency, making it a primary optimization for production LLM inference in TensorRT-LLM.

Why this answer

Post-training quantization reduces weight memory and enables faster low-precision kernels, while paged KV cache manages generation-time memory efficiently and improves batching. Together they lower footprint and latency without retraining. The rejected options either require retraining, increase model size, or break the architecture, so they do not fit the no-retraining constraint or the optimization goals.

Exam trap

The trap here is selecting quantization-aware training because it sounds like a precision optimization, when it violates the no-retraining requirement.

272
MCQhard

An LLM inference service on NVIDIA Triton Inference Server uses TensorRT-LLM as the backend. During peak load, you notice that the first token latency (time to first token) is high, but subsequent tokens are generated quickly. Which metric should you monitor to diagnose this issue?

A.Time to first token (TTFT)
B.GPU utilization
C.Requests per second (RPS)
D.Inter-token latency (ITL)
AnswerA

TTFT directly measures the latency from request submission to the first generated token. In LLM inference, the prefill phase (processing the input prompt) dominates TTFT. Monitoring TTFT helps identify if the prefill is slow due to large batch sizes, long prompts, or insufficient compute, allowing targeted optimization.

Why this answer

Time to first token (TTFT) is the key metric for diagnosing prefill latency in LLM inference. Since the issue is high TTFT but fast subsequent tokens, monitoring TTFT allows you to isolate the prefill phase and optimize it, for example by reducing batch size or using more efficient attention kernels.

Exam trap

The trap here is focusing on GPU utilization or throughput instead of the specific latency metric that reflects the prefill phase.

273
Multi-Selectmedium

A team is fine-tuning a 70B parameter model with NVIDIA NeMo using LoRA on eight H100 GPUs. They want to reduce GPU memory usage during training while preserving the base model's pretrained knowledge. Which two configuration changes should they apply? (Choose two.)

Select 2 answers
A.Increase the micro batch size to improve GPU utilization and reduce the number of optimizer steps.
B.Enable LoRA by setting adapter dimensions and target modules so only adapter weights receive gradients.
C.Disable tensor parallelism so each GPU holds a full copy of the model weights.
D.Switch from bfloat16 to float32 precision to improve numerical stability during training.
E.Enable gradient checkpointing to recompute activations during the backward pass instead of storing them.
AnswersB, E

LoRA freezes the base model and trains low-rank adapter matrices, which drastically reduces the number of trainable parameters and optimizer state. This lowers memory usage during fine-tuning while preserving the pretrained weights. In NeMo, this is configured through the PEFT section with adapter dimensions and target modules, making it a direct memory-saving measure.

Why this answer

LoRA reduces trainable parameters and optimizer state while freezing the base model, and gradient checkpointing reduces activation memory by recomputing activations. Together they lower peak GPU memory during fine-tuning. Increasing micro batch size, using float32, or disabling tensor parallelism all increase memory usage and work against the stated goal.

Exam trap

The trap here is treating bfloat16 as a stability-only choice and forgetting that float32 doubles memory, which conflicts with the explicit memory-reduction objective.

274
MCQmedium

When optimizing a model using NVIDIA TensorRT, what is the primary benefit of enabling 'layer fusion' during the optimization process?

A.It increases the overall model weight precision
B.It reduces global memory access overhead
C.It automatically prunes redundant parameters
D.It replaces floating-point math with integer math
AnswerB

By fusing multiple operations into a single kernel, intermediate tensors do not need to be written back to global VRAM. This significantly reduces memory bandwidth consumption, as the fused kernel can pass data directly between operations using fast on-chip memory or registers, leading to improved inference latency.

Why this answer

Layer fusion reduces the overhead of launching multiple small GPU kernels by combining them into a single, optimized kernel. This strategy minimizes the total number of reads and writes to global VRAM, which is often the bottleneck in modern deep learning. By keeping intermediate data in high-speed registers or shared memory, the model spends less time waiting for memory access and more time performing actual floating-point operations.

Exam trap

Candidates often believe layer fusion is primarily about reducing the number of parameters or the total model size, rather than optimizing the execution flow to minimize global memory round-trips.

275
MCQmedium

Refer to the exhibit. An engineer receives this timeout error during a CUDA kernel execution. What is the most appropriate first step to diagnose the resource contention?

A.Increase the block size to 2048 threads.
B.Use NVIDIA Nsight Compute to profile occupancy.
C.Disable the watchdog timer in the OS.
D.Switch the kernel to run on the CPU.
AnswerB

Nsight Compute provides detailed analysis of occupancy, register usage, and shared memory allocation. It identifies whether the kernel is struggling with resource contention, allowing the developer to adjust thread block configuration or refine memory usage to prevent the execution time from exceeding the watchdog timer.

Why this answer

CUDA kernel timeouts are typically caused by long-running kernels that exceed the GPU's watchdog timer or by excessive resource usage (registers/shared memory) that limits occupancy. Using the NVIDIA Nsight Compute profiler allows the engineer to see the exact occupancy metrics and register pressure, enabling targeted optimizations. This step is crucial for identifying if the kernel is over-provisioned for the specific GPU architecture being targeted.

Exam trap

Candidates frequently recommend restarting the driver or increasing the OS watchdog timeout limit instead of using dedicated profiling tools to inspect kernel resource usage.

276
MCQmedium

A team is deploying a 13B-parameter decoder-only LLM on a single NVIDIA A100 40GB GPU for a real-time chatbot. During load, the process runs out of memory even though the model weights in FP16 require roughly 26GB. The team wants to reduce GPU memory usage with minimal impact on output quality and no change to the model architecture. Which technique is most appropriate?

A.Enable tensor parallelism across the single GPU's SMs
B.Switch the attention implementation to FlashAttention-2
C.Apply INT8 weight-only quantization using NVIDIA TensorRT-LLM
D.Increase the KV cache block size to 32 tokens
AnswerC

Weight-only INT8 quantization stores model weights at 8 bits while keeping activations in higher precision, cutting the 26GB FP16 footprint to roughly 13GB with typically small quality loss. TensorRT-LLM supports this natively and is designed for production inference on NVIDIA GPUs, so it directly addresses the memory shortfall without altering the network topology.

Why this answer

The bottleneck is static weight storage on a single 40GB device. Weight-only INT8 quantization halves the parameter footprint while preserving the architecture and offering strong quality retention, and TensorRT-LLM provides a supported path for this on NVIDIA hardware. The other choices target attention compute, KV cache paging, or multi-GPU partitioning, none of which reduce the resident weight memory on one GPU.

Exam trap

The trap here is assuming that attention-kernel optimizations such as FlashAttention reduce total GPU memory enough to fit a model whose weights alone exceed available VRAM.

277
MCQhard

You are building a pretraining dataset from a large collection of source-code repositories for an NVIDIA NeMo LLM. The data includes many files with licenses, generated code, and minified JavaScript. Which NeMo Curator-based approach best improves code data quality before tokenization?

A.Convert all code to a single programming language using an automated transpiler before tokenization.
B.Increase the model's context window to 32k tokens so that entire repositories fit in one sequence.
C.Tokenize all files with a byte-level BPE tokenizer and skip any further filtering.
D.Use language-specific filters that detect and remove minified files, license headers, and auto-generated code, then apply deduplication.
AnswerD

Minified JavaScript, license headers, and generated code are low-value or repetitive patterns that harm code pretraining. NeMo Curator supports custom filters for line length, ratio of alphanumeric characters, and detection of boilerplate. Deduplication removes repeated generated files. Applying these before tokenization ensures the model learns from meaningful code rather than noise, improving downstream code generation and understanding.

Why this answer

Language-specific filters combined with deduplication directly target the described quality issues. Minified files, license headers, and generated code are common in code corpora and can be detected with heuristics such as average line length, ratio of whitespace, and presence of standard license text. Deduplication removes repeated generated files.

This preprocessing ensures the pretraining corpus contains meaningful code, which improves the model's code capabilities.

Exam trap

The trap here is believing that a more powerful tokenizer or a larger context window can compensate for low-quality or repetitive code data.

278
MCQmedium

An LLM inference service on NVIDIA Triton Inference Server is experiencing occasional out-of-memory (OOM) errors on the GPU. The team wants to monitor GPU memory usage to predict and prevent OOM conditions. Which NVIDIA tool provides real-time GPU memory metrics that can be integrated with Prometheus for alerting?

A.NVIDIA Nsight Systems
B.NVIDIA Triton Inference Server metrics endpoint
C.NVIDIA System Management Interface (nvidia-smi)
D.NVIDIA Data Center GPU Manager (DCGM)
AnswerD

DCGM is a suite of tools for managing and monitoring NVIDIA GPUs in clusters. It provides detailed metrics including GPU memory usage, and DCGM Exporter can expose these metrics in Prometheus format. This integration enables real-time monitoring and alerting on memory usage, helping predict and prevent OOM conditions in production LLM inference.

Why this answer

DCGM is designed for data center GPU monitoring and provides comprehensive metrics, including memory usage. With DCGM Exporter, these metrics can be scraped by Prometheus, enabling real-time alerts when memory usage approaches limits. This proactive monitoring allows the team to adjust batch sizes or resource allocation before OOM errors occur, ensuring reliable LLM inference.

Exam trap

The trap here is confusing profiling tools or basic command-line utilities with a production-grade monitoring solution that integrates with Prometheus for alerting.

279
MCQeasy

You are curating instruction-tuning data for an NVIDIA NIM-deployed LLM. The raw dataset contains many near-duplicate instruction-response pairs that differ only by punctuation and whitespace. Which data preparation step is most appropriate to remove these before fine-tuning?

A.Increase the batch size during fine-tuning to average out duplicate examples.
B.Convert all instructions to lowercase and remove all punctuation globally.
C.Use NVIDIA Triton Inference Server dynamic batching to filter duplicates at serving time.
D.Apply fuzzy deduplication using MinHash LSH over normalized instruction-response text.
AnswerD

Fuzzy deduplication with MinHash LSH is designed to catch near-duplicates that differ by punctuation, whitespace, or minor edits. Normalizing text before hashing ensures that trivial variations map to the same signature. Removing these duplicates prevents the model from overfitting to repeated examples and reduces wasted training compute, which is especially important when fine-tuning an instruction-following model on a curated dataset.

Why this answer

Fuzzy deduplication with MinHash LSH is the standard approach for near-duplicate removal in instruction-tuning datasets. It normalizes text for comparison, generates MinHash signatures, and uses locality-sensitive hashing to find similar pairs efficiently at scale. This directly removes the redundant examples described, improving training efficiency and reducing overfitting without altering the semantic content of the retained data.

Exam trap

The trap here is thinking that deduplication means exact string matching or global text normalization, when near-duplicates require similarity-based methods that preserve the original text.

280
MCQmedium

What is the primary role of a Model Repository in the NVIDIA Triton Inference Server architecture?

A.To store training datasets for real-time model retraining.
B.To act as a centralized filesystem for serving multiple models.
C.To handle network traffic load balancing between server nodes.
D.To compile the model into an optimized executable format.
AnswerB

The repository is the source of truth for the server. It organizes models into a hierarchical structure, enabling versioning and easy configuration management. Triton periodically scans this path, allowing updates to be deployed simply by adding files, which is essential for high-availability production AI systems.

Why this answer

The Model Repository acts as a centralized storage location for all deployed models, their configurations, and their versions. Triton monitors this directory to detect when new versions are added or updated, allowing for seamless model updates without restarting the server. This design supports robust MLOps practices by decoupling the model storage from the inference engine runtime, ensuring that deployments remain manageable and version-controlled.

Exam trap

Candidates often think the Model Repository is a database or a training dataset storage, failing to recognize it as a simple filesystem-based interface for Triton to manage model versions.

281
Multi-Selectmedium

Which TWO of the following techniques are best suited for reducing the latency of LLM inference on NVIDIA GPUs?

Select 2 answers
A.KV cache quantization.
B.Increasing the number of CPU worker threads.
C.Implementing PagedAttention.
D.Disabling ECC memory on the GPU.
E.Switching from Tensor Cores to CUDA Cores.
AnswersA, C

KV cache quantization compresses the storage of intermediate tokens, significantly reducing memory bandwidth consumption. Since LLM inference is often memory-bandwidth bound during the decoding phase, this technique allows more tokens to be processed concurrently and speeds up the transfer of cache data to the compute units.

Why this answer

Reducing LLM latency requires optimizing both the compute throughput and the memory access patterns. Key-Value (KV) cache quantization and PagedAttention are industry standards for LLM acceleration. KV cache quantization reduces memory bandwidth usage, while PagedAttention manages memory dynamically, preventing fragmentation.

These methods significantly improve the token generation rate, which is the primary metric for user-perceived performance in large-scale generative AI deployments.

Exam trap

Candidates often suggest generic performance tweaks like overclocking or batch size adjustments, missing that PagedAttention and KV cache quantization are specific, high-impact techniques for LLM memory management.

282
MCQmedium

You are preparing a large corpus of customer support transcripts for continued pretraining of an NVIDIA NeMo Megatron model. The transcripts contain personally identifiable information such as names, email addresses, and account numbers, and company policy requires that this information be removed before training while preserving as much linguistic context as possible for the model to learn from. Which data preparation approach best satisfies both requirements?

A.Hash every token in the corpus so that the original text cannot be reconstructed, and train on the hashed sequences.
B.Replace all digits and capitalized words with a generic mask token across the entire corpus before training.
C.Detect PII spans with a combination of regular expressions and a named-entity recognition model, then replace each detected span with a consistent placeholder token that preserves the surrounding sentence structure.
D.Drop every transcript that contains any detected PII so that no sensitive information reaches the training pipeline.
AnswerC

Regexes reliably catch structured identifiers such as emails and account numbers, while a named-entity recognition model covers names and locations that patterns miss. Replacing spans with consistent placeholders removes the sensitive values but keeps sentence structure and surrounding context intact, which is exactly what continued pretraining needs to learn language patterns without memorizing private data.

Why this answer

Combining regex detection for structured identifiers with named-entity recognition for names and locations covers the PII surface area, and replacing detected spans with consistent placeholders removes sensitive values while preserving the sentence context that continued pretraining depends on. Dropping records, masking all digits and capitals, or hashing every token either destroys usable language data or fails to anonymize reliably.

Exam trap

The trap here is treating PII removal as a binary choice between deleting records and destroying all structure, when span-level replacement preserves the linguistic signal.

283
Multi-Selectmedium

A team is using NVIDIA NeMo Evaluator to assess a Llama-3-70B model fine-tuned for medical question answering. They want to evaluate both the correctness of answers and the model's ability to avoid hallucinating unsupported facts. Which two evaluation strategies should they implement? (Choose two.)

Select 2 answers
A.Compute BLEU score on the generated answers.
B.Calculate perplexity of the model on a held-out set of medical questions.
C.Use a fact-checking module that cross-references generated statements with a trusted medical knowledge base.
D.Compare model outputs against gold-standard answers using exact match and F1 score.
E.Run the model through NeMo Guardrails with a policy that blocks any output containing numbers.
AnswersC, D

A fact-checking module can verify each claim in the model's output against a curated medical knowledge base, directly detecting hallucinations or unsupported facts. This approach provides a granular, evidence-based assessment of factual accuracy, which is critical in healthcare. Integrating such a module with NeMo Evaluator allows automated flagging of unsupported statements.

Why this answer

To evaluate correctness, exact match and F1 score against gold answers provide a direct comparison. To detect hallucinations, a fact-checking module that verifies statements against a trusted medical knowledge base is essential. Together, these strategies cover both aspects.

BLEU and perplexity do not measure factual accuracy, and a blanket guardrail against numbers is ineffective.

Exam trap

The trap here is assuming that any automated metric like BLEU or perplexity can substitute for factual verification, when hallucination detection requires external knowledge grounding.

284
Multi-Selectmedium

Which TWO of the following NVIDIA AI Enterprise tools are specifically designed to optimize and accelerate the deployment of LLMs in containerized environments?

Select 2 answers
A.NVIDIA Triton Inference Server
B.NVIDIA Modulus
C.TensorRT-LLM
D.NVIDIA NeMo Framework
E.NVIDIA DriveWorks
AnswersA, C

Triton is a production-grade inference serving software that supports multiple frameworks and provides advanced features like dynamic and in-flight batching. It is essential for managing model lifecycle, scaling inference workloads, and ensuring high availability for LLM services in enterprise production environments.

Why this answer

NVIDIA Triton Inference Server and TensorRT-LLM are the cornerstones of the NVIDIA AI Enterprise deployment stack. Triton provides a unified serving platform that abstracts the complexities of hardware management, while TensorRT-LLM provides the compiler technology to fuse kernels and quantize weights. Together, they enable developers to move from research models to production-ready services with minimal code changes, ensuring optimized performance across NVIDIA GPUs.

Exam trap

Candidates often confuse general management tools like NVIDIA AI Workbench with deployment-specific runtime engines like TensorRT-LLM, leading them to select development tools rather than serving technologies.

285
MCQeasy

An engineer is optimizing a BERT-like model for inference using NVIDIA TensorRT. They want to reduce latency further by using lower precision without significant accuracy loss. Which TensorRT precision mode should they choose to enable INT8 inference while maintaining accuracy through calibration?

A.FP16
B.FP32
C.TF32
D.INT8
AnswerD

INT8 precision mode in TensorRT enables 8-bit integer inference, which significantly reduces latency and memory usage. To maintain accuracy, TensorRT uses a calibrator to determine scaling factors from a representative dataset. This matches the engineer's goal of using INT8 with calibration to minimize accuracy loss, making it the correct choice.

Why this answer

INT8 precision mode in TensorRT allows inference using 8-bit integers, reducing latency and memory footprint. To preserve accuracy, a calibrator computes scaling factors from a calibration dataset. The other precisions (FP32, FP16, TF32) do not provide INT8 inference and do not use calibration for quantization.

Exam trap

The trap here is confusing TF32 with INT8, as TF32 is often mentioned for Tensor Core acceleration but is not an 8-bit integer format and does not use calibration.

286
MCQeasy

What is the primary purpose of 'Few-Shot Prompting' in the context of LLM optimization?

A.To reduce the latency of the underlying GPU cluster.
B.To provide in-context learning examples to guide output.
C.To compress the model weights for deployment.
D.To permanently store data in the model's internal memory.
AnswerB

Providing examples allows the model to observe the desired pattern of input and output. This pattern-matching capability enables the model to perform new tasks accurately without formal retraining, making it an ideal strategy for quickly adapting pre-trained models to specific enterprise data formats and business logic requirements.

Why this answer

Few-shot prompting involves providing a few examples of input-output pairs within the prompt to guide the model's performance on a specific task. This approach helps the model learn the desired format, tone, and logic without the need for intensive fine-tuning. It is a highly efficient way to steer model behavior for specific enterprise use cases, ensuring consistent results across multiple interaction sessions.

Exam trap

Candidates often confuse few-shot prompting with fine-tuning, assuming the model's weights are updated during the few-shot process rather than just providing context within the prompt.

287
MCQmedium

A research team is pretraining a decoder-only LLM and observes that gradient magnitudes in the earliest layers are extremely small while later layers train normally, causing slow convergence. They are using post-layer normalization. Which architectural change is most likely to improve gradient flow to the early layers?

A.Move layer normalization to before each sublayer (pre-LN) and add a final normalization before the output projection
B.Replace layer normalization with batch normalization across the sequence dimension
C.Remove residual connections to shorten the gradient path
D.Increase the number of attention heads while keeping head dimension constant
AnswerA

Pre-LN applies normalization on the sublayer input, creating a clean residual path that carries gradients directly from the loss to early layers. This is the standard remedy for vanishing gradients in deep Transformers and typically removes the need for learning-rate warmup. Adding a final normalization stabilizes the output scale before the vocabulary projection.

Why this answer

Post-LN places normalization inside the residual branch, so gradients must pass through normalization at every layer, attenuating them in deep stacks. Pre-LN normalizes the sublayer input and leaves the residual stream unnormalized, giving a direct gradient highway from the loss to early parameters. That structural change, plus a final normalization before the output head, is the recognized fix for the described symptom.

Exam trap

The trap here is treating vanishing early-layer gradients as a capacity or width problem and adjusting head counts, when the root cause is where normalization sits relative to the residual branch.

288
Multi-Selectmedium

Which TWO actions should be part of a robust incident response plan for an LLM deployment failing in production?

Select 2 answers
A.Rolling back to a previous stable model version
B.Deploying the model to more GPUs immediately
C.Reviewing recent logs and telemetry metrics
D.Turning off all security monitoring systems
E.Forcing a reboot of all servers
AnswersA, C

If a new model deployment causes issues, a rollback is the fastest way to restore service stability. This action minimizes downtime and gives the engineering team the necessary time to debug the problematic version in a non-production environment.

Why this answer

A robust incident response plan focuses on rapid mitigation and root cause analysis. Immediately reverting to a known good version reduces the impact on users, while analyzing logs and telemetry provides the necessary data to understand the failure. These steps ensure that service is restored quickly, which is the primary objective of reliability engineering in generative AI systems serving critical user traffic.

Exam trap

Candidates often select 'retraining the model' as a primary incident response action. Retraining is a long-term fix, not an immediate incident response step for a failing production deployment.

289
MCQhard

Refer to the exhibit. What is the intended outcome of this data cleaning configuration for a Large Language Model pre-training corpus?

A.It specifically targets the removal of personally identifiable information (PII).
B.It optimizes for the removal of low-quality or nonsensical text while minimizing redundancy.
C.It enforces a strict length-based chunking strategy for all documents.
D.It converts all text to a vector space representation before filtering.
AnswerB

The configuration uses perplexity filtering to identify incoherent content and length constraints to exclude short, low-information strings. The MinHash algorithm effectively manages the similarity threshold to eliminate near-duplicate documents. This combination ensures that the training dataset is concise, coherent, and free of redundant, low-value information inputs.

Why this answer

This configuration aims to remove low-quality text that fails to meet minimum length requirements or exhibits high perplexity (indicating gibberish or low coherence). Simultaneously, MinHash deduplication identifies and removes near-duplicate documents exceeding a 95% similarity threshold. This cleanup process is vital for pre-training, as it filters out low-value, noisy data that could impede model convergence and general quality during the training cycle.

Exam trap

Candidates often assume cleaning configurations only remove empty strings, failing to recognize the combined role of length/perplexity filters and MinHash deduplication in eliminating noise.

290
MCQmedium

A research team wants to train a large decoder-only LLM where each token's representation is computed independently of token order, then inject order information afterward. They are choosing between learned absolute positional embeddings and sinusoidal absolute positional embeddings. Which statement accurately characterizes the tradeoff they face?

A.Learned embeddings are strictly better because they allow the model to discover relative offsets between tokens.
B.Sinusoidal embeddings are parameter-free and can be computed for any position, while learned embeddings require a fixed maximum length and add parameters.
C.Learned embeddings can extrapolate to sequence lengths longer than those seen in training, whereas sinusoidal embeddings cannot.
D.Sinusoidal embeddings must be recomputed for every batch, which makes them slower than learned embeddings at inference.
AnswerB

Sinusoidal embeddings use fixed sine and cosine functions of position and dimension, so they add no trainable parameters and can be evaluated at arbitrary indices. Learned embeddings allocate a trainable vector per position, which consumes parameters proportional to the maximum length and cannot produce values for unseen indices. This is the classic tradeoff between the two absolute schemes.

Why this answer

Sinusoidal absolute positional embeddings are deterministic functions of position, adding no parameters and allowing evaluation at any index. Learned absolute embeddings allocate a trainable vector per position, consuming parameters and capping usable length at the trained maximum. The other options either reverse the extrapolation behavior or misattribute relative-offset abilities to learned absolute embeddings.

Exam trap

The trap here is assuming learned positional embeddings extrapolate better because they are trainable, when in fact their fixed table size is their key limitation.

291
MCQmedium

A developer is prompting an NVIDIA NIM for a Code Llama model to generate a Python function. The model produces correct logic but frequently omits type hints and docstrings, which the team requires. Which prompting technique best addresses this specific gap?

A.Include a short example of the desired function signature, type hints, and docstring in the prompt, then ask for the new function in the same style.
B.Increase max_tokens so the model has more room to include type hints and docstrings.
C.Add the instruction "write clean, production-quality code" to the system message.
D.Ask the model to first explain its reasoning about the function, then output the code.
AnswerA

Providing a concrete example that exhibits type hints and a docstring demonstrates the exact format expected. The model imitates the pattern it sees, so this one-shot demonstration directly fills the missing elements. It is more reliable than vague style instructions because the required structure is shown rather than described, leaving little room for interpretation.

Why this answer

A concrete example showing the desired function signature, type hints, and docstring teaches the model the exact format by imitation. Vague quality instructions, larger token budgets, and reasoning steps do not specify these structural requirements, so they fail to close the gap. The example-based approach is the most direct and reliable fix.

Exam trap

The trap here is using subjective quality phrases like production-quality code instead of demonstrating the exact structural elements the model must include.

292
MCQhard

A financial institution uses NVIDIA NeMo Guardrails to enforce ethical guidelines in its customer-facing LLM. During testing, the model occasionally generates responses that violate the company's policy against offering investment advice. The guardrails are configured with a set of dialog flows and safety checks. What is the most effective way to address this issue?

A.Fine-tune the base LLM on a dataset of compliant responses to reduce the likelihood of generating investment advice.
B.Implement a post-processing filter that uses a separate LLM to classify responses as advice or non-advice and redacts them accordingly.
C.Enhance the NeMo Guardrails configuration with a custom action that checks the model's output against a compliance rule set and triggers a safe fallback response when a violation is detected.
D.Add a custom guardrail that detects and blocks any mention of specific financial terms like 'invest' or 'stock'.
AnswerC

NeMo Guardrails supports custom actions that can run arbitrary code to validate outputs. By integrating a compliance rule set, the guardrail can detect investment advice and replace the response with a safe fallback, ensuring policy adherence. This approach is flexible and can be updated as policies evolve.

Why this answer

The most effective solution is to enhance NeMo Guardrails with a custom action that evaluates the model's output against compliance rules and triggers a safe fallback. This leverages the extensibility of NeMo Guardrails to enforce policy dynamically and reliably, ensuring that any investment advice is intercepted and replaced.

Exam trap

The trap here is thinking that fine-tuning or simple keyword blocking is sufficient, when in fact a robust, rule-based guardrail with custom actions provides a more reliable and maintainable compliance mechanism.

293
MCQeasy

Which component in the NVIDIA NeMo framework is specifically designed to manage the configuration and orchestration of large-scale fine-tuning jobs?

A.TensorRT-LLM
B.NeMo Framework Launcher
C.NVIDIA Triton Inference Server
D.CUDA Toolkit
AnswerB

The NeMo Framework Launcher provides the necessary abstractions to configure, submit, and manage large-scale fine-tuning tasks. It streamlines the workflow by handling job scheduling and resource allocation, making it the primary tool for orchestrating model training on NVIDIA hardware platforms effectively and reliably at scale.

Why this answer

The NeMo Framework Launcher is the dedicated tool for managing multi-node, large-scale training jobs. It abstracts the complexities of cluster scheduling, environment setup, and hyperparameter management. By using the launcher, engineers ensure that fine-tuning tasks are executed efficiently across NVIDIA compute clusters, allowing for reproducible and scalable experiments that align with enterprise-grade development standards and best practices for large model management.

Exam trap

Candidates often confuse the NeMo Framework Launcher with base PyTorch or general-purpose CI/CD tools like Jenkins, failing to recognize the launcher's specific role in abstracting multi-node cluster scheduling and environment orchestration.

294
MCQhard

Refer to the exhibit. Which prompt engineering technique would best force the model to prioritize technical detail over marketing language?

A.Add a constraint: 'Exclude marketing language and include specific architectural metrics like TDP, memory bandwidth, and interconnect speeds.'
B.Ask the model to 'Write a shorter summary' in the user prompt.
C.Decrease the temperature to 0.0 to make the model more factual.
D.Use few-shot prompting with generic summaries.
AnswerA

This approach provides clear negative constraints (exclude marketing) and positive constraints (include specific metrics). By defining the required output format and content, the model is compelled to ignore its tendency to generate generic, flowery text and focus on the hard data points that define the technical architecture.

Why this answer

The model is failing to adhere to the implicit expectation of 'technical depth.' By explicitly defining a structural constraint—such as requiring specific architectural metrics or removing marketing terminology—the model is forced to prioritize the technical aspects requested. This type of constraint-driven prompting is essential for professional NVIDIA technical writers and engineers who need precise documentation summaries.

Exam trap

Candidates frequently choose vague instructions like 'be more technical' instead of providing specific, actionable constraints, failing to realize that LLMs require explicit structural boundaries to filter out marketing-heavy language.

295
MCQmedium

In the context of model optimization, why is 'graph surgery' sometimes required before building a TensorRT engine?

A.To increase the model's parameter count
B.To resolve unsupported operator compatibility
C.To improve the quality of training data
D.To automatically quantize the model to INT8
AnswerB

Often, models exported from frameworks contain operators that are not directly supported by TensorRT. Graph surgery allows engineers to replace these nodes with equivalent, supported sub-graphs. This ensures the builder can successfully create an engine without encountering errors or falling back to inefficient CPU execution paths.

Why this answer

Graph surgery is the process of modifying the model's computational graph to replace unsupported or inefficient operations with more optimized, TensorRT-compliant versions. Some frameworks export nodes that TensorRT cannot parse or optimize effectively. By manually editing the graph to fuse operations or simplify the structure before building, engineers can ensure that the engine builder produces a high-performance execution plan that effectively utilizes the underlying hardware capabilities.

Exam trap

Candidates frequently believe graph surgery is used for fine-tuning weights or pruning dead neurons, confusing model compression techniques with structural compatibility fixes for the inference engine.

296
Multi-Selecthard

Which THREE architectural features are essential for enabling efficient inference of massive LLMs on multi-GPU systems?

Select 3 answers
A.KV Caching to store previous sequence states.
B.Tensor Parallelism to split layers across GPUs.
C.Batch Size reduction to increase memory throughput.
D.Weight Quantization to reduce the memory footprint.
E.Increasing the learning rate during inference.
AnswersA, B, D

KV caching prevents redundant computations by storing previously calculated keys and values, which is critical for reducing inference latency in autoregressive models. Without this, the model would need to recompute the entire attention history for every new token generated, leading to prohibitive performance costs in real-time scenarios.

Why this answer

Efficient inference at scale requires techniques that address memory constraints, compute latency, and communication overhead. Model parallelism, KV caching, and weight quantization are fundamental pillars that allow large models to fit within limited memory pools while maintaining high throughput. These techniques collectively ensure that the model remains responsive and cost-effective when serving large-scale requests in a production environment.

Exam trap

Students often select training-specific optimizations like gradient accumulation or data parallelism instead of focusing on inference-specific architectural requirements like tensor parallelism and KV caching.

297
MCQmedium

Why is 'tokenization stability' a critical metric when preparing data for NVIDIA-based LLM deployment?

A.It guarantees that the model will always generate the same output for a given prompt.
B.It ensures that the GPU memory usage remains constant during the training process.
C.It prevents unexpected input shifts between training and inference environments.
D.It reduces the total number of parameters required for the embedding layer.
AnswerC

Stability ensures that the mapping between text and tokens remains consistent. If an inference pipeline tokenizes text differently than the training pipeline, the model encounters a distribution shift. This mismatch can result in degraded model performance, incorrect reasoning, or complete failure, making stability a foundational requirement for robust production systems.

Why this answer

Tokenization stability ensures that the same input text consistently maps to the same sequence of tokens across different environments or library versions. If tokenization is inconsistent, the model might receive unexpected inputs compared to what it observed during training, leading to severe performance degradation. For NVIDIA-based deployments, deterministic tokenization is essential for maintaining production-level reliability and predictable model behavior across various inference pipelines.

Exam trap

Test-takers frequently assume tokenization stability only affects processing speed, missing its critical role in preventing unexpected input shifts between training and inference environments.

298
MCQmedium

You are preparing a dataset of customer reviews for fine-tuning an LLM to generate concise summaries. The reviews are in multiple languages, but the target summaries must be in English. You have a limited budget for translation. Which data preparation step is most critical to ensure the fine-tuned model produces high-quality English summaries?

A.Translate all reviews into English and then train the model to summarize English text.
B.Ensure that each training example pairs a review in its original language with a high-quality English summary.
C.Use a multilingual LLM to translate all non-English reviews into English before training.
D.Filter the dataset to include only reviews originally written in English.
AnswerB

This approach directly trains the model to perform cross-lingual summarization: input in any language, output in English. It leverages the original text without translation errors and teaches the model to generate English summaries regardless of source language. This is the most effective strategy for the stated goal.

Why this answer

Pairing original-language reviews with English summaries directly trains the model for cross-lingual summarization, which is the end goal. This avoids translation errors and preserves the original semantics. It also prepares the model for real-world multilingual inputs, ensuring it can generate English summaries without an intermediate translation step.

Exam trap

The trap here is assuming that translating everything to English first is simpler, but it fails to train the model for multilingual input and may introduce translation artifacts.

299
MCQmedium

When using QLoRA for fine-tuning, what is the primary purpose of using the 4-bit NormalFloat (NF4) data type?

A.To increase the training speed by using integer-only arithmetic
B.To reduce the memory footprint of the model weights
C.To improve the convergence speed of the optimizer
D.To enable training without the need for gradient accumulation
AnswerB

NF4 quantizes the base model weights to 4-bit precision, which drastically lowers the memory requirement compared to 16-bit or 32-bit representations. This allows users to fine-tune significantly larger models on a single NVIDIA GPU, as the memory bottleneck is primarily the storage of the frozen model weights during the training process.

Why this answer

The NF4 data type is mathematically optimal for weights that follow a normal distribution, which is typical for pre-trained language model weights. By quantizing weights to 4 bits, QLoRA significantly reduces the memory footprint, allowing large models to fit onto GPUs with lower VRAM. This efficiency does not significantly sacrifice performance, provided the weights are dequantized during the forward pass to maintain precision for activations.

Exam trap

Candidates often assume 4-bit quantization (NF4) permanently degrades model accuracy or serves only for inference, ignoring its primary role in reducing training memory footprints via QLoRA.

300
MCQmedium

When designing an AI application for the public sector, which ethical principle must be prioritized regarding transparency?

A.Maximizing the model's speed to provide instantaneous public service responses.
B.Concealing the use of AI to prevent user bias and ensure natural interactions.
C.Clearly disclosing that the user is interacting with an AI system.
D.Using proprietary, closed-source models to prevent external scrutiny.
AnswerC

Disclosure is the most direct application of transparency. It allows the user to adjust their expectations, knowing they are not speaking to a human. This builds trust and ensures that the user is not misled, which is a foundational ethical requirement for any deployment in the public domain.

Why this answer

Transparency in public sector AI is critical for maintaining democratic accountability and public trust. Citizens have a right to know when they are interacting with an AI rather than a human, and they should understand the logic behind decisions that affect them. By clearly disclosing AI usage and providing explainable outputs, organizations meet their legal obligations and demonstrate ethical responsibility, ensuring that public-facing systems remain fair, accountable, and open to scrutiny.

Exam trap

Candidates focus exclusively on model accuracy or technical latency metrics, ignoring the specific public sector ethical mandate regarding user disclosure and transparency.

Page 3

Page 4 of 5

Page 5

All pages