Courseiva

NVIDIA Certified Professional: Generative AI LLMs (NCP-GENL) — Questions 301–352

352 questions total · 5pages · All types, answers revealed

Page 4

Page 5 of 5

301
MCQmedium

A data science team is preparing an instruction fine-tuning dataset in NVIDIA NeMo Framework. They notice that after training, the model performs well on the training instructions but poorly on paraphrased versions of the same instructions. They want to improve generalization without increasing dataset size. Which data preparation change is most appropriate?

A.Apply instruction paraphrasing and template diversification so the same intent appears in multiple surface forms during training.
B.Remove all examples where the instruction phrasing is unique, keeping only the most common templates.
C.Increase the number of epochs so the model sees each instruction more times and memorizes the exact phrasing.
D.Lower the learning rate so the model updates more slowly and avoids overfitting to specific instruction phrasings.
AnswerA

Paraphrasing and template diversification expose the model to varied surface forms of the same intent, which teaches it to map meaning rather than exact wording. This directly improves generalization to unseen paraphrases without adding new intents, and it is a standard data augmentation technique for instruction tuning.

Why this answer

Generalization to paraphrased instructions depends on the model seeing the same intent expressed in varied surface forms. Paraphrasing and template diversification teach the model to respond to meaning rather than exact wording, directly improving performance on unseen paraphrases without enlarging the dataset or changing training hyperparameters.

Exam trap

The trap here is treating poor paraphrase generalization as an optimizer or epoch problem, when the real fix is increasing surface-form diversity in the instruction data itself.

302
MCQhard

Refer to the exhibit. The deployment is facing memory allocation errors during peak load. Based on the error log, what is the most effective configuration change to resolve the issue while keeping the model architecture constant?

A.Switch to a smaller model size.
B.Increase the GPU clock frequency.
C.Enable PagedAttention in the TensorRT-LLM runtime.
D.Reduce the batch size to one.
AnswerC

PagedAttention is designed to solve exactly this type of KV cache allocation failure. By moving from static allocation to paged block allocation, the system can utilize memory more efficiently and pack more requests into the same GPU memory footprint, effectively eliminating the OOM errors caused by inflexible memory management.

Why this answer

The error log indicates that the system is running out of memory because the KV cache is allocated statically, which often leads to fragmentation or over-allocation. Enabling PagedAttention is the standard NVIDIA-recommended solution for this scenario. It allows the system to manage KV cache memory dynamically, reclaiming space from finished requests and efficiently allocating blocks for new tokens, thereby preventing the out-of-memory (OOM) errors caused by static allocation.

Exam trap

Candidates often suggest increasing GPU VRAM or reducing batch sizes, which are reactive measures. They overlook PagedAttention, which is the specific architectural solution for KV cache fragmentation and memory OOM errors.

303
Multi-Selecthard

You are building an evaluation harness for a retrieval-augmented generative assistant running on NVIDIA NIM microservices. The product owner wants a single trustworthy number for 'answer quality,' but you need to defend the evaluation design. Which two design choices most directly protect the evaluation from producing misleading quality scores? (Choose two.)

Select 2 answers
A.Let the same generative model that powers the assistant also grade its own answers, since it understands the domain best.
B.Hold out a prompt set that was never used during fine-tuning or prompt engineering, and keep it frozen across model versions.
C.Average the outputs of several automatic metrics into one composite score so the product owner receives a single number.
D.Increase decoding temperature during evaluation so the model explores more of the answer space and scores reflect average-case behavior.
E.Include a retrieval ablation that runs the same prompts with and without retrieved context to isolate the contribution of the retriever.
AnswersB, E

A frozen, genuinely unseen set prevents leakage and version-to-version drift in the test data itself. If prompts were recycled during prompt engineering or fine-tuning, scores inflate because the model has effectively seen the answers. Keeping the set immutable also makes run-over-run deltas interpretable, since any change in score can be attributed to the model rather than to a shifting benchmark.

Why this answer

Leakage control and retrieval attribution are the two structural safeguards that make RAG evaluation trustworthy. A frozen, unseen prompt set ensures score changes reflect the model, and the retrieval ablation shows whether gains come from context or from parametric memory. Composite averaging, self-grading, and elevated temperature all add noise or bias rather than removing it, so they undermine the credibility the product owner needs.

Exam trap

The trap here is reaching for a single blended metric or self-grading for convenience, when the real risks are benchmark contamination and unverified attribution of gains to retrieval.

304
MCQhard

An ML engineer is fine-tuning a 13B LLM with LoRA on 4 NVIDIA A100 GPUs using NVIDIA NeMo. They notice that the effective batch size is very small and gradients are noisy, but increasing the per-GPU micro batch size triggers out-of-memory errors. Which technique should they apply to increase the effective batch size without increasing memory per step?

A.Switch from bfloat16 to float32 precision to stabilize gradient values.
B.Enable gradient accumulation to sum gradients across multiple micro batches before the optimizer step.
C.Increase the LoRA rank and alpha so the adapter captures more information per step.
D.Reduce the sequence length by truncating all training examples to 128 tokens.
AnswerB

Gradient accumulation executes several forward and backward passes on small micro batches, accumulating gradients before applying a single optimizer update. This raises the effective batch size without increasing peak memory, because only one micro batch resides in memory at a time. It directly addresses the noisy-gradient problem while respecting the memory ceiling that prevents larger micro batches.

Why this answer

Gradient accumulation performs multiple forward and backward passes on small micro batches and sums their gradients before one optimizer update, emulating a larger batch while keeping peak memory at the level of a single micro batch. It is the standard remedy when memory prevents enlarging the micro batch. Changing adapter rank, precision, or sequence length does not increase the number of samples per update in a memory-neutral way.

Exam trap

The trap here is equating adapter capacity or numeric precision with batch size, when only gradient accumulation increases effective batch size without raising peak memory.

305
MCQhard

A media analytics company runs a TensorRT-LLM optimized GPT-J model on a single NVIDIA A100 80GB GPU using NVIDIA Triton Inference Server. During peak hours, request concurrency rises sharply and the team observes that the GPU is idle for long periods while waiting on host-side tokenization and detokenization. Profiling shows that CPU preprocessing and postprocessing dominate request latency. The team wants to reduce end-to-end latency without changing model weights or adding GPUs. Which Triton feature should they use?

A.Dynamic batching, because it groups requests on the server and reduces the number of model executions.
B.The TensorRT-LLM backend's in-flight batching and built-in tokenizer, which perform tokenization and detokenization inside the backend and keep the GPU busy with continuous batching.
C.Sequence batching with a custom scheduler, because it keeps tokenization on the host and overlaps it with GPU compute across requests.
D.A Triton ensemble that places a Python model for tokenization before the TensorRT-LLM model and a Python model for detokenization after it.
AnswerB

The TensorRT-LLM backend for Triton includes an integrated tokenizer and supports in-flight batching, which moves tokenization into the backend and schedules new requests into ongoing GPU work. This directly targets the host-side bottleneck described in the profile and improves latency by keeping the GPU continuously utilized without adding hardware or changing weights.

Why this answer

The TensorRT-LLM Triton backend provides an integrated tokenizer and in-flight batching, which eliminate the host-side tokenization and detokenization stalls observed in the profile. In-flight batching also continuously admits new requests into the running batch, improving GPU utilization and reducing end-to-end latency without changing model weights or adding GPUs.

Exam trap

The trap here is assuming that generic batching or ensemble orchestration will fix a host-side preprocessing bottleneck, when the actual fix is moving tokenization into the TensorRT-LLM backend and using in-flight batching.

306
MCQhard

When fine-tuning an LLM to follow specific safety protocols, why is the inclusion of 'adversarial' examples in the training data considered a best practice?

A.To artificially inflate the model's perplexity scores during evaluation.
B.To ensure the model learns to refuse requests that violate established safety policies.
C.To increase the model's fluency in generating complex technical instructions.
D.To allow the model to learn and reproduce the adversarial techniques during generation.
AnswerB

Adversarial training exposes the model to edge-case prompts designed to break safety constraints. By learning to handle these, the model becomes more robust against malicious attempts to bypass safety filters. This ensures consistent enforcement of safety protocols in production, making the model more secure and reliable for enterprise use.

Why this answer

Adversarial examples test the model's ability to maintain safety boundaries even when prompted with manipulative or malicious inputs. By training on these examples, the model learns to identify and refuse requests that violate safety protocols. This proactive preparation is essential for deploying LLMs in enterprise environments where maintaining strict safety and compliance standards is required, protecting against sophisticated prompt injection and bypass techniques.

Exam trap

Candidates often assume adversarial examples are only for improving accuracy or general performance, failing to recognize that they are specifically required to enforce safety boundaries and refusal behaviors in high-stakes environments.

307
MCQeasy

What is the function of the 'Masked' component in a Decoder-only Transformer's self-attention during training?

A.To filter out stop words from the input text.
B.To prevent the model from attending to future tokens in the sequence.
C.To increase the randomness of the model's predictions.
D.To reduce the computation of the attention mechanism by 50%.
AnswerB

The causal mask is a triangular matrix applied to the attention scores that sets values for future tokens to negative infinity. This ensures that when the softmax is applied, the weights for these positions become zero, effectively hiding the future context from the model during training and ensuring autoregressive integrity.

Why this answer

The mask ensures that each token can only attend to itself and the tokens that precede it in the sequence. This autoregressive property is essential for training decoder models, as it prevents the model from 'cheating' by looking at future tokens. This ensures that the model learns to predict the next word based solely on the context that would be available during actual inference.

Exam trap

Candidates often think the mask is for ignoring padding tokens. While padding masks exist, the primary 'Masked' component in decoder-only training is strictly for preventing look-ahead at future tokens.

308
MCQmedium

A global bank runs an NVIDIA NeMo Guardrails-protected assistant for its tellers. During a compliance audit, the auditor asks how the bank can prove that the guardrail configuration itself has not been silently altered between releases. Which practice best satisfies this requirement?

A.Ask the model to self-report at startup which guardrails it believes are enabled, and archive that self-report as evidence.
B.Store the Colang guardrail files and their configuration in a version-controlled repository with signed commits, and record the commit hash in the deployment manifest.
C.Rely on the NVIDIA NIM container image digest alone, because the guardrail policy is compiled into the model weights.
D.Enable verbose logging of every user prompt and model response so the auditor can infer the guardrail rules from observed behavior.
AnswerB

Version-controlled, signed guardrail artifacts plus the commit hash in the deployment manifest create an immutable, auditable trail. An auditor can reproduce exactly which guardrail rules were active for any released model and detect unauthorized edits, which meets change-management and integrity expectations for a regulated financial deployment.

Why this answer

Guardrail integrity is a configuration-management problem, not an inference or logging problem. Signed, version-controlled Colang files with commit hashes recorded in the deployment manifest give auditors a verifiable, reproducible record of exactly which safety rules were active. This is the only approach that can detect silent tampering and withstand regulatory scrutiny.

Exam trap

The trap here is assuming that logging runtime behavior or trusting model self-reports can substitute for verifiable configuration artifacts.

309
Multi-Selectmedium

A team is optimizing an NVIDIA TensorRT-LLM deployment of a 70B model on multiple GPUs. They want to reduce inter-GPU communication overhead and improve throughput. Which two techniques should they consider? (Choose two.)

Select 2 answers
A.Enable pipeline parallelism to split layers into stages.
B.Disable in-flight batching to simplify scheduling.
C.Use FP32 precision for all weights to improve numerical stability.
D.Use tensor parallelism with NVLink-connected GPUs.
E.Increase the KV cache block size to 256 tokens.
AnswersA, D

Pipeline parallelism assigns different layer stages to different GPUs, reducing the frequency of cross-GPU communication compared with tensor parallelism. With micro-batching, stages can overlap and keep GPUs busy. This lowers communication overhead per token and can improve throughput when combined with appropriate batch scheduling.

Why this answer

Tensor parallelism benefits from NVLink's high bandwidth to reduce all-reduce overhead, while pipeline parallelism lowers communication frequency by staging layers across GPUs. Together they address the communication bottleneck for large multi-GPU models. KV cache block size, disabling in-flight batching, and FP32 precision do not reduce inter-GPU communication and can hurt throughput.

Exam trap

The trap here is treating memory-management knobs like KV cache block size as if they also reduce inter-GPU communication overhead.

310
MCQhard

A healthcare AI team is evaluating a large language model for clinical note summarization. They want to measure whether the generated summaries contain fabricated information not present in the source notes. Which evaluation approach is most appropriate for detecting hallucinated content?

A.Use a natural language inference (NLI) model to check entailment between the source note and the generated summary.
B.Compute the perplexity of the generated summaries on a held-out set of clinical notes.
C.Calculate the BLEU score between the generated summary and a reference summary written by a clinician.
D.Measure the ROUGE-L score between the generated summary and the source note.
AnswerA

NLI models determine whether a hypothesis (the summary) is entailed by a premise (the source note). If the summary contains fabricated information, the NLI model would predict contradiction or neutral rather than entailment. This approach directly assesses factual consistency and is widely used for hallucination detection in summarization, making it the most appropriate method here.

Why this answer

Natural language inference (NLI) is effective for hallucination detection because it evaluates whether the generated summary is logically entailed by the source document. If the summary introduces unsupported information, the NLI model will not predict entailment. Other metrics like perplexity, BLEU, and ROUGE-L focus on fluency or surface overlap and cannot reliably identify fabricated content.

Therefore, NLI-based checking is the most appropriate approach.

Exam trap

The trap here is confusing fluency metrics like perplexity with factual consistency metrics, when hallucination detection requires comparing the summary against the source for entailment.

311
MCQeasy

A developer is building a decoder-only generative model and wants to prevent the model from attending to future tokens during training so that each position can only use information from itself and earlier positions. Which architectural mechanism should they implement in the self-attention layer?

A.A padding mask that ignores tokens added to equalize sequence lengths across a batch.
B.A learned positional embedding added to the token embeddings at the input layer.
C.A causal attention mask that sets attention scores for future positions to negative infinity before the softmax.
D.A residual connection that adds the attention output back to the input hidden state.
AnswerC

A causal mask adds negative infinity to the attention logits of positions after the current token, so after softmax those future positions receive zero probability. This enforces left-to-right autoregressive conditioning during training and matches the inference behavior where future tokens are unavailable.

Why this answer

Decoder-only generative models must not see future tokens during training, otherwise next-token prediction becomes trivial and the model will not generalize to autoregressive inference. A causal attention mask accomplishes this by making future positions invisible to the softmax, ensuring each position only conditions on itself and earlier tokens.

Exam trap

The trap here is confusing a padding mask, which hides filler tokens, with a causal mask, which hides future tokens.

312
MCQhard

An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. The model uses multi-head attention with a sequence length of 4096. During inference, the GPU's Tensor Cores are underutilized, and the kernel launch overhead is high due to many small operations. Which optimization should be applied to improve Tensor Core utilization and reduce overhead?

A.Increase the batch size to maximize parallelism and hide latency.
B.Enable kernel fusion by using TensorRT-LLM's fused multi-head attention plugin.
C.Use NVIDIA Triton Inference Server with dynamic batching to improve GPU utilization.
D.Convert the model to FP16 precision to double the Tensor Core throughput.
AnswerB

TensorRT-LLM provides fused multi-head attention plugins that combine multiple operations (e.g., QKV projection, attention, and output projection) into a single kernel. This reduces kernel launch overhead and increases arithmetic intensity, allowing better utilization of Tensor Cores. The fusion also minimizes memory traffic, which is critical for long sequences.

Why this answer

The underutilization of Tensor Cores and high kernel launch overhead stem from many small operations in multi-head attention. TensorRT-LLM's fused multi-head attention plugin combines these operations into a single optimized kernel, reducing overhead and improving Tensor Core utilization. This is a targeted optimization for Transformer models, especially with long sequences.

Exam trap

The trap here is focusing on batch size or precision when the core issue is kernel fragmentation and launch overhead, which require fusion.

313
Multi-Selectmedium

You are preparing a dataset for supervised fine-tuning (SFT) of an NVIDIA NeMo LLM to follow instructions. Which TWO data preparation practices are essential to ensure the model learns to generalize rather than memorize? (Choose two.)

Select 2 answers
A.Include a diverse set of instruction types and domains in the training data rather than repeating a few templates.
B.Increase the learning rate significantly to force the model to escape memorization of individual examples.
C.Duplicate the most common instruction-response pairs to reinforce the desired behavior.
D.Use only English instructions to simplify tokenization and avoid multilingual complexity.
E.Split the dataset into training, validation, and test sets with no overlapping instructions or responses.
AnswersA, E

Diversity in instruction types and domains encourages the model to learn the general pattern of following instructions rather than memorizing specific templates. If the training data is dominated by a few templates, the model may overfit to those formats and fail on novel instructions. A broad distribution of tasks improves zero-shot and few-shot generalization.

Why this answer

Disjoint data splits prevent leakage and give honest evaluation, while diverse instruction types and domains teach the model to follow instructions generally rather than memorize templates. Together, these practices reduce overfitting and improve generalization. The other options either fail to address memorization, actively encourage it, or unnecessarily restrict the data distribution.

Exam trap

The trap here is thinking that training hyperparameters like learning rate or duplicating data can substitute for proper data splitting and diversity when the goal is generalization.

314
Multi-Selecthard

An enterprise team is preparing a supervised fine-tuning job in NVIDIA NeMo for a 20B LLM. They want to reduce GPU memory consumption during training without changing the model architecture or the dataset. Which two configuration changes should they apply? (Choose two.)

Select 2 answers
A.Expand the training dataset with additional synthetic examples.
B.Reduce the maximum sequence length by truncating all samples to 64 tokens.
C.Enable gradient checkpointing to recompute activations during the backward pass.
D.Increase the number of attention heads in each transformer block.
E.Use mixed precision with bfloat16 for forward and backward passes.
AnswersC, E

Gradient checkpointing stores only selected activations and recomputes the rest during backpropagation, trading additional compute for substantially lower activation memory. It does not alter the model architecture or dataset, so it fits the stated constraints. This makes it an effective lever for fitting larger models or batch sizes into the same GPU memory budget during NeMo fine-tuning runs.

Why this answer

Gradient checkpointing and bfloat16 mixed precision both reduce memory during training without modifying model architecture or dataset. Checkpointing lowers stored activation memory by recomputation, while bfloat16 halves tensor memory and is supported natively in NeMo. Architectural changes, data expansion, and sequence truncation either violate the constraints or fail to address memory consumption per training step.

Exam trap

The trap here is treating data-level changes like truncation or augmentation as memory optimizations, when the constraints require configuration-level changes that leave architecture and dataset intact.

315
MCQhard

An engineer is training a large language model with pipeline parallelism across four NVIDIA GPUs. They observe that GPU utilization is low and training throughput is limited by idle time during pipeline bubbles. Which technique is most effective to reduce pipeline bubbles and improve utilization?

A.Increase the number of micro-batches to keep the pipeline stages busy.
B.Reduce the batch size to decrease the memory footprint and allow more concurrent kernels.
C.Enable gradient checkpointing to trade compute for memory and increase batch size.
D.Switch from pipeline parallelism to data parallelism across the four GPUs.
AnswerA

Pipeline bubbles occur when stages wait for data from previous stages. Increasing the number of micro-batches allows more overlapping of forward and backward passes across stages, filling the pipeline and reducing idle time. This is a standard technique in pipeline parallelism to improve utilization.

Why this answer

Pipeline bubbles are idle periods when stages wait for data. Increasing the number of micro-batches allows more fine-grained overlap of computation across stages, keeping all GPUs busy and reducing idle time. This is a direct and effective way to improve pipeline utilization.

Exam trap

The trap here is confusing memory optimization techniques with those that address pipeline scheduling inefficiencies, such as increasing micro-batches.

316
MCQeasy

A developer needs to fine-tune a 7B LLM for a customer-support chatbot using NVIDIA NeMo. The dataset contains paired instructions and desired responses. Which data format should be used to prepare the dataset for supervised fine-tuning?

A.A YAML configuration file listing hyperparameters and dataset paths only.
B.A pickle file containing a Python list of tokenized integer sequences with no text labels.
C.A JSONL file where each line contains an input text field and an output text field representing the instruction-response pair.
D.A CSV file containing only the raw customer queries with no corresponding responses.
AnswerC

Supervised fine-tuning in NeMo expects paired examples that map an input prompt to a target completion. A JSONL file with input and output fields per line matches that structure and can be consumed by the NeMo data preprocessing pipeline. This format directly supports the instruction-response objective used for chatbot behavior, making it the appropriate choice for this scenario.

Why this answer

Supervised fine-tuning trains a model to produce a target response given an input instruction. The dataset must therefore contain aligned input-output pairs, and NeMo's data pipeline accepts JSONL records with distinct input and output fields. Files lacking response labels, pre-tokenized sequences without pairing, or pure configuration files cannot supply the supervision signal the training objective requires.

Exam trap

The trap here is confusing the training configuration file with the training dataset, or assuming tokenized sequences alone are sufficient without input-output pairing.

317
MCQhard

A team has built a TensorRT-LLM engine for a 70B model on four NVIDIA H100 GPUs using tensor parallelism. They now need to serve the same model on a single H100 for a development environment, accepting higher latency. What is the most appropriate approach?

A.Use pipeline parallelism with a single stage to split the model across the GPU's SMs automatically.
B.Load the existing four-GPU engine on a single H100 and set the runtime to use only one device.
C.Convert the engine to a TensorRT plan file with FP32 precision so it can use unified memory on the H100.
D.Rebuild the engine with tensor parallelism set to one and apply weight quantization so the model fits in a single GPU's memory.
AnswerD

Tensor parallelism is fixed at engine build time, so moving from four GPUs to one requires rebuilding the engine with a tensor parallel size of one. A 70B FP16 model needs roughly 140GB, which exceeds a single H100's 80GB, so weight quantization is necessary to fit. This approach produces a valid single-GPU engine while accepting the expected latency increase.

Why this answer

Tensor parallelism is a build-time property, so changing the number of GPUs requires rebuilding the engine with tensor parallel size one. Because a 70B model in FP16 exceeds a single H100's memory, weight quantization is also needed to fit. The other options either attempt to reuse an incompatible engine, misapply parallelism concepts, or increase precision, none of which yield a working single-GPU deployment.

Exam trap

The trap here is assuming a multi-GPU engine can be restricted at runtime to fewer devices, when parallelism is compiled into the plan and requires a rebuild.

318
MCQeasy

A developer wants to reduce the disk and memory footprint of a fine-tuned 70B model before serving it with TensorRT-LLM, and is willing to accept a small, measurable quality drop that they will validate with an evaluation harness. Which approach best matches that requirement?

A.Perform post-training quantization to INT4 or INT8 weights and validate with perplexity and task metrics.
B.Retrain the model from scratch using quantization-aware training with fake-quant nodes in the forward pass.
C.Convert the checkpoint to FP8 and rely on the runtime to upcast to FP16 for every matmul.
D.Increase the tensor parallelism degree so each GPU holds a smaller shard of the FP16 weights.
AnswerA

Post-training quantization compresses an already fine-tuned checkpoint without retraining, cutting weight storage and memory roughly 2x for INT8 and 4x for INT4. Because the developer accepts a small, measurable quality drop and will validate it with an evaluation harness, PTQ is the appropriate, low-effort match for the stated goal.

Why this answer

The requirement is a smaller weight footprint with an accepted, validated quality trade-off and no retraining. Post-training quantization of the fine-tuned checkpoint to INT4 or INT8 delivers that reduction immediately and pairs naturally with an evaluation harness to confirm the quality delta stays within tolerance. Retraining-based methods exceed the stated scope.

Exam trap

The trap here is defaulting to quantization-aware training as the highest-quality option when the scenario explicitly accepts a small, measurable quality drop and does not provide for retraining.

319
MCQmedium

You are preparing a large instruction-tuning dataset with NVIDIA NeMo Curator. The dataset contains many near-duplicate instruction-response pairs that differ only in punctuation or minor wording. Which NeMo Curator stage should you apply to remove these near-duplicates before fine-tuning?

A.Fuzzy deduplication using MinHash locality-sensitive hashing.
B.Exact substring deduplication to remove overlapping text spans.
C.Heuristic filtering to remove short or malformed records.
D.Exact duplicate removal using a hash of the full instruction-response pair.
AnswerA

MinHash with locality-sensitive hashing groups documents by estimated Jaccard similarity, so near-duplicates with minor edits are clustered and one representative is kept. In NeMo Curator this is implemented as fuzzy deduplication and directly targets the scenario's near-duplicate instruction pairs. It scales to large datasets and removes redundancy that would otherwise bias fine-tuning.

Why this answer

Near-duplicate instruction-response pairs that differ only slightly require fuzzy matching, not exact hashing or heuristic quality filters. MinHash locality-sensitive hashing estimates Jaccard similarity between records and clusters near-identical ones, letting you retain a single representative. Applying fuzzy deduplication in NeMo Curator reduces redundancy, prevents the model from overfitting repeated patterns, and improves generalization on the instruction-tuning task.

Exam trap

The trap here is assuming that exact duplicate removal is sufficient, when minor punctuation or wording changes cause hashes to differ and near-duplicates to survive.

320
MCQmedium

Why is gradient checkpointing useful when fine-tuning a model on a single GPU?

A.It speeds up the training process by calculating gradients in parallel
B.It saves GPU memory by recomputing activations during the backward pass
C.It automatically adjusts the learning rate for each layer
D.It prevents the model from using the CPU during training
AnswerB

By discarding intermediate activations and recomputing them on-the-fly, the memory footprint is significantly reduced. This allows the model to process larger sequences or larger batches that would normally trigger an out-of-memory error. This trade-off is critical for fine-tuning large models on consumer-grade NVIDIA hardware with limited VRAM capacity.

Why this answer

Gradient checkpointing trades computation for memory by not storing all intermediate activations during the forward pass. Instead, it recomputes them during the backward pass. This is extremely valuable when working on hardware with limited VRAM, as it allows for larger batch sizes or longer sequences that would otherwise cause OOM errors, though it does increase the total training time slightly due to the recomputation overhead.

Exam trap

Candidates often assume gradient checkpointing increases memory efficiency by using compression, rather than correctly identifying it as a trade-off that saves memory by recomputing activations during the backward pass.

321
MCQmedium

An enterprise deploying a large language model on an NVIDIA A100 GPU experiences high memory bandwidth bottlenecks during autoregressive token generation. Which optimization technique specifically addresses this memory-bound phase by merging element-wise operations and reducing global memory round-trips?

A.Applying static INT8 post-training quantization to all linear layers.
B.Enabling TensorRT-LLM custom kernel fusion for multi-head attention and activation blocks.
C.Increasing the global batch size to maximize arithmetic intensity.
D.Switching from FlashAttention to standard vanilla self-attention mechanisms.
AnswerB

TensorRT-LLM provides highly optimized, fused CUDA kernels specifically designed to eliminate redundant global memory round-trips for operations like multi-head attention, layer normalization, and activations. This directly accelerates memory-bound autoregressive text generation workloads on NVIDIA GPUs.

Why this answer

Kernel fusion combines multiple successive GPU operations, such as bias additions and activations, into a single CUDA kernel. This drastically reduces high-latency global memory read and write operations, directly mitigating the memory bandwidth bottleneck characteristic of autoregressive transformer decoding phases on NVIDIA hardware.

Exam trap

Candidates frequently select generic model pruning or distillation, which reduces parameter count but fails to directly address the specific memory bandwidth bottlenecks caused by repetitive global memory access in autoregressive decoding.

322
Multi-Selecthard

A generative AI application deployed on NVIDIA Triton Inference Server with NVIDIA AI Enterprise is experiencing silent data corruption in model outputs. The MLOps team needs to implement monitoring to detect such issues early. Which two actions should be taken to enhance observability for silent data corruption? (Choose two.)

Select 2 answers
A.Enable NVIDIA Data Center GPU Manager (DCGM) to monitor GPU ECC error counts and XID errors.
B.Configure Triton's model warmup to run dummy inferences at startup to stabilize performance.
C.Increase the batch size for all inference requests to improve throughput and reduce per-request overhead.
D.Set up logging of all inference inputs and outputs to a centralized system for later analysis.
E.Implement output validation by comparing inference results against a known-good baseline for a sample of requests.
AnswersA, E

DCGM tracks GPU hardware errors such as ECC memory errors and XID errors, which can cause silent data corruption. By monitoring these metrics, the team can detect hardware-level issues that might lead to incorrect model outputs. This is a proactive measure to identify and alert on potential corruption sources before they affect production results.

Why this answer

Silent data corruption can stem from hardware faults or software bugs. Monitoring GPU ECC and XID errors via DCGM catches hardware-induced corruption, while output validation against a baseline detects incorrect results regardless of cause. Together, they provide both infrastructure and application-level observability, enabling early detection and mitigation of silent data corruption in production LLM inference.

Exam trap

The trap here is focusing on performance tuning or logging instead of active detection methods like hardware error monitoring and output validation, which are essential for catching silent data corruption.

323
MCQmedium

Which strategy is most effective for optimizing an LLM that is too large to fit into a single GPU's VRAM?

A.Increase the inference batch size
B.Apply dynamic quantization to the weights
C.Implement tensor parallelism
D.Convert the model to a CPU-only format
AnswerC

Tensor parallelism involves splitting individual weight matrices of the model across multiple GPUs. This allows the model to reside on multiple devices, effectively pooling their VRAM and compute power. It is the primary method for scaling large models beyond the hardware limits of a single GPU device.

Why this answer

Model parallelism, specifically tensor parallelism, is the standard approach for splitting a large model across multiple GPUs. By partitioning individual layers across different devices, the computation can be distributed, and the total memory requirement is spread proportionally. This allows for the deployment of models that exceed the capacity of a single GPU, enabling high-performance inference for massive parameter models that would otherwise be impossible to load.

Exam trap

Students often select data parallelism or pipeline parallelism incorrectly, failing to recognize that single layers exceeding VRAM require splitting the actual weight matrices across multiple devices.

324
MCQmedium

You are preparing a multilingual corpus for pretraining with NVIDIA NeMo. The dataset contains documents in 40 languages, but the tokenizer was trained primarily on English. Which data preparation action best ensures that non-English text is represented efficiently during tokenization?

A.Increase the model's maximum sequence length to 8192 tokens to accommodate longer tokenized non-English text.
B.Train a new tokenizer on a balanced sample of all 40 languages using SentencePiece or Hugging Face tokenizers before pretraining.
C.Apply Unicode normalization form NFKC to all text and rely on the existing English tokenizer.
D.Translate all non-English documents to English using an NVIDIA NIM translation model before tokenization.
AnswerB

A tokenizer trained predominantly on English will split non-English words into many subword units, increasing sequence length and reducing effective context. Training a new tokenizer on a balanced multilingual sample ensures that each language has adequate vocabulary coverage, which lowers the number of tokens per document and improves training efficiency. This must be done before pretraining because the tokenizer defines the model's embedding space.

Why this answer

Training a new tokenizer on a balanced multilingual sample directly solves the vocabulary coverage problem. It ensures that frequent subwords in each language are included in the vocabulary, reducing token fragmentation and improving training efficiency. The other options either mask the symptom, alter the data inappropriately, or do not address tokenizer vocabulary at all.

Exam trap

The trap here is assuming that sequence length or Unicode normalization can compensate for a tokenizer that lacks vocabulary for the target languages.

325
MCQhard

Refer to the exhibit. What is the effective batch size for this fine-tuning job?

A.1
B.15
C.16
D.17
AnswerC

The effective batch size is the product of the per_device_train_batch_size (1) and the gradient_accumulation_steps (16). This results in an effective batch size of 16, which helps in stabilizing training by providing a more representative gradient estimate over multiple mini-batches before weights are actually updated in the optimizer.

Why this answer

The effective batch size is calculated by multiplying the per-device batch size by the gradient accumulation steps. In this case, 1 * 16 equals 16. Gradient accumulation allows the model to simulate a larger batch size by running multiple forward and backward passes before performing a single optimizer step, which helps in achieving more stable gradient updates without increasing the immediate memory overhead for activations.

Exam trap

Candidates frequently add or subtract batch size parameters instead of multiplying per-device batch size by gradient accumulation steps to find the effective batch size.

326
MCQmedium

A company wants to fine-tune a 70B LLM to follow domain-specific instructions. The base model already performs well on general language tasks. They have limited labeled data and limited GPU memory. Which approach is most appropriate?

A.Retrain the tokenizer on the domain corpus and continue pretraining from scratch.
B.Parameter-Efficient Fine-Tuning with LoRA adapters targeting attention projection layers.
C.Full-parameter fine-tuning of all 70B weights on the domain dataset.
D.Prompt engineering with few-shot examples and no weight updates.
AnswerB

LoRA freezes the base weights and trains small low-rank adapter matrices, dramatically reducing memory and optimizer state requirements. It is well suited to limited labeled data because fewer trainable parameters reduce overfitting risk, and it preserves the base model's general abilities. This matches the scenario's constraints on memory and data volume while still adapting the model to domain instructions.

Why this answer

LoRA is the standard parameter-efficient method for adapting very large models under memory constraints, since it trains only small adapter matrices while freezing base weights. With limited labeled data, the reduced trainable parameter count also lowers overfitting risk and helps retain general capabilities. Full fine-tuning, tokenizer retraining, and prompt-only approaches either exceed resource limits or fail to deliver durable domain adaptation.

Exam trap

The trap here is assuming that the largest model change yields the best domain adaptation, when memory and data limits make parameter-efficient adaptation the correct engineering choice.

327
MCQeasy

A retail company wants to let its customer-support LLM answer questions about order status. The security team insists the model must never be able to invoke a refund or account-modification function, even if a user crafts a clever prompt. Which approach enforces that constraint most reliably?

A.Add a system prompt instructing the model to refuse any request that would modify an account or issue a refund.
B.Expose only a read-only order-status tool to the model and keep refund and account-modification functions outside the tool registry entirely.
C.Log every tool invocation and alert the security team whenever a refund or account-modification call is detected.
D.Fine-tune the model on examples of refund and account-modification refusals so it learns to decline those requests.
AnswerB

If the sensitive functions are never registered as callable tools, no prompt can invoke them, because the model's action space is defined by the registry rather than by its instructions. This is an architectural control that holds even against adversarial input. Read-only status queries remain fully supported, so the customer experience is preserved while the constraint is enforced at the system boundary.

Why this answer

Authorization in tool-using LLM applications must be enforced by what the model can reach, not by what it is told. Keeping refund and account-modification functions out of the tool registry makes them structurally unreachable, so no prompt-injection technique can trigger them. Read-only status access still works.

Prompts, fine-tuning, and post-hoc alerting all leave a live execution path that an adversary could exploit.

Exam trap

The trap here is treating a strong system prompt or refusal fine-tune as a security boundary when the sensitive function remains callable in the execution layer.

328
MCQhard

An engineer is using NVIDIA TensorRT-LLM's in-flight batching to serve a mix of short and very long prompts. They observe that GPU utilization drops and latency for short requests spikes whenever a long prompt is admitted. Which mechanism should they tune to prevent long sequences from monopolizing the batch?

A.Adjust the maximum number of tokens processed per iteration and the KV cache block allocation policy to limit how much context a single request can consume.
B.Enable tensor parallelism so the long prompt's attention computation is split across multiple GPUs.
C.Switch from in-flight batching to a static batching scheme that groups requests by similar prompt length.
D.Increase the maximum batch size so more short requests can be admitted alongside the long prompt.
AnswerA

In-flight batching processes a token budget per iteration; if one long prompt consumes most of that budget, short requests stall. Tuning the per-iteration token limit and KV cache block allocation constrains how much context a single sequence can occupy, allowing the scheduler to interleave short and long requests fairly. This directly targets the observed utilization drop and latency spike without changing model precision or hardware.

Why this answer

In-flight batching improves throughput by mixing prefill and decode work, but a very long prompt can consume the entire per-iteration token budget, starving shorter sequences. Limiting tokens per iteration and controlling KV cache block allocation prevents any single request from monopolizing the batch. This restores GPU utilization and keeps short-request latency stable while still allowing long prompts to complete progressively.

Exam trap

The trap here is assuming that more GPUs or larger batches fix latency fairness, when the real constraint is per-iteration token budget allocation within the scheduler.

329
MCQmedium

A team is using an NVIDIA NIM-hosted Llama model to generate product descriptions from a list of technical specifications. The descriptions sometimes omit key specifications or include invented features. The team wants to improve reliability without changing the model. Which prompt engineering change is most likely to reduce these errors?

A.Add a system prompt that instructs the model to act as a technical writer and to include every specification exactly as given, and provide a few-shot example of a correct description.
B.Increase the max_tokens parameter to allow the model to generate longer descriptions, ensuring all specifications are covered.
C.Use a chain-of-thought prompt that asks the model to list each specification and then write the description.
D.Set the temperature to 0.0 to make the model deterministic and reduce creativity.
AnswerA

A system prompt sets the model's role and constraints, and few-shot examples demonstrate the desired output format and level of detail. By showing a correct description that includes all specifications without additions, the model is more likely to follow that pattern. This directly addresses both omission and invention by providing clear instructions and a concrete example, without retraining the model.

Why this answer

To reduce omissions and inventions without changing the model, the most effective prompt engineering approach is to provide explicit instructions and few-shot examples. A system prompt that defines the role and constraints, combined with an example that demonstrates including all specifications and avoiding fabrication, guides the model to produce more reliable outputs. Other parameter changes or chain-of-thought do not directly address content fidelity.

Exam trap

The trap here is thinking that lowering temperature or increasing max_tokens will fix content errors, when the real solution is to explicitly instruct the model and show it an example of the desired output.

330
MCQmedium

You are preparing a dataset of support tickets for a RAG system using NVIDIA NeMo. Many tickets are short and contain little context, which hurts retrieval quality. Which data preparation technique best improves retrieval by enriching each ticket with related information before embedding?

A.Increase the embedding dimension of the retrieval model.
B.Lower the similarity threshold used during retrieval.
C.Apply deduplication to remove similar tickets from the corpus.
D.Use an LLM to generate contextual summaries or metadata for each ticket and prepend them to the text before embedding.
AnswerD

Generating a summary or metadata adds missing context to each short ticket, so the embedded representation captures more relevant semantics and matches user queries better. Prepending this enriched text before embedding directly addresses the sparsity problem and improves retrieval, making it the appropriate technique for the scenario.

Why this answer

Short tickets embed poorly because they carry little semantic signal. Using an LLM to generate contextual summaries or metadata and prepending that text before embedding enriches each record, so the vector captures more of the ticket's intent and domain. This improves matching with user queries.

Deduplication, larger embeddings, and lower thresholds do not add missing context, so they cannot resolve the sparsity problem in this scenario.

Exam trap

The trap here is assuming that retrieval-time tuning, such as lowering a similarity threshold, can compensate for source documents that simply lack contextual content.

331
MCQhard

A global bank uses NVIDIA NeMo Guardrails to enforce ethical AI policies in its customer-facing LLM application. The compliance team requires that the system automatically logs all instances where the model attempts to generate financial advice, including the prompt, the blocked response, and the rail that triggered. Which NeMo Guardrails feature should the team enable to capture this audit trail?

A.NVIDIA TensorRT-LLM's runtime debugging output
B.NVIDIA NIM's built-in request logging
C.Colang tracing with a custom logging action
D.NVIDIA Triton Inference Server's model ensemble scheduler
AnswerC

Colang tracing allows developers to instrument the guardrails flow and capture events such as rail triggers. By adding a custom logging action within the Colang flow, the team can record the prompt, the blocked response, and the specific rail that fired. This provides the detailed audit trail required for compliance without modifying the LLM itself.

Why this answer

Colang tracing with a custom logging action is the correct approach because it hooks into the guardrails execution flow, allowing the team to capture the exact rail that triggered and the associated prompt and response. Other options are infrastructure components that lack visibility into NeMo Guardrails' policy decisions, so they cannot provide the required audit detail.

Exam trap

The trap here is assuming that any logging in the stack (like NIM or Triton) can serve as an audit trail for guardrail events, when only the guardrails layer knows why a response was blocked.

332
MCQhard

A production LLM service on NVIDIA Triton Inference Server uses dynamic batching. During peak load, the 99th percentile latency increases significantly, but GPU utilization remains at 60%. Which configuration change is most likely to improve latency while maintaining throughput?

A.Decrease the maximum batch size and adjust the batch timeout to reduce waiting
B.Enable model instances to increase parallelism on the same GPU
C.Switch to a larger GPU with more memory bandwidth
D.Increase the maximum batch size in the dynamic batching configuration
AnswerA

With GPU utilization at 60%, the GPU has spare capacity. High P99 latency during peak load suggests requests are waiting too long for batches to fill. Reducing the maximum batch size and lowering the batch timeout (e.g., from 100 microseconds to 50) allows smaller batches to be processed more quickly, reducing queue time. This can improve latency while still maintaining acceptable throughput because the GPU is not fully utilized.

Why this answer

When GPU utilization is low but latency is high, the bottleneck is often the batching mechanism waiting to accumulate requests. Reducing the maximum batch size and batch timeout allows requests to be processed sooner, lowering queue time and P99 latency. Since the GPU has spare capacity, this change can improve latency without sacrificing throughput.

Exam trap

The trap here is assuming that increasing batch size always improves performance, overlooking the impact of batch timeout on latency.

333
MCQmedium

An enterprise deploying a large language model on NVIDIA Triton Inference Server needs to track GPU utilization, memory allocation, and custom inference latency histograms in Prometheus. Which approach should the MLOps engineer implement to ensure robust production observability without overloading the inference execution threads?

A.Write a custom Python wrapper around every generate call to log execution duration directly to a shared text file.
B.Query the NVIDIA Management Library via a cron job every second to poll current GPU statistics and write them to a database.
C.Enable the native Prometheus metrics endpoint in Triton Inference Server configuration and configure Prometheus to scrape it.
D.Disable all internal instrumentation layers to maximize raw throughput, relying solely on external black-box API health probes.
AnswerC

Enabling the native Prometheus metrics endpoint allows the server to expose internal hardware and performance counters directly over an HTTP port. This scraping mechanism operates asynchronously from the main inference pipeline, preventing metric collection overhead from degrading GPU compute efficiency.

Why this answer

Enabling Triton's native Prometheus metrics endpoint allows the server to expose internal hardware and performance counters directly over an HTTP port. This scraping mechanism operates asynchronously from the main inference pipeline, preventing metric collection overhead from degrading GPU compute efficiency or increasing inference request latency.

Exam trap

Candidates often assume custom Python application loggers must manually wrap every inference call, mistakenly believing built-in server metrics lack the required granularity for enterprise LLM latency tracking.

334
MCQeasy

An engineer is building a customer support assistant using an NVIDIA NIM for a Llama 3 70B model. The assistant must always respond in valid JSON containing exactly the keys "issue" and "urgency", and must never include any other text. Which prompt engineering approach most directly enforces this output contract?

A.Raise the temperature to 1.0 so the model explores more possible response formats and eventually produces valid JSON.
B.Append the phrase "Please be concise" to every user message and rely on the model to infer the JSON structure.
C.Set the top_p value to 0.1 and the max_tokens parameter to 50 to force the model into a JSON-only response mode.
D.Add a system message that specifies the JSON schema and instructs the model to output only JSON, then validate responses programmatically.
AnswerD

A system message defining the exact schema and the instruction to emit only JSON constrains the model's behavior at the highest-priority level of the prompt. Because the model sees this before the user turn, it shapes every completion. Programmatic validation then catches any residual drift, making this the most direct and reliable enforcement for a strict output contract.

Why this answer

The most direct way to enforce a strict JSON contract is to state the schema and the output-only requirement in the system message, which has the highest priority in the prompt hierarchy, and then validate the result programmatically. Sampling parameters and stylistic instructions do not constrain structure, so they cannot guarantee the two required keys or the absence of extra text.

Exam trap

The trap here is assuming that sampling parameters such as temperature or top_p can enforce an output format, when they only affect randomness and token selection.

335
MCQeasy

In the context of NVIDIA AI Enterprise, what is the primary purpose of using 'NVIDIA Triton Model Analyzer'?

A.To perform real-time model training
B.To automate performance optimization and configuration
C.To detect GPU hardware failures
D.To encrypt model weights at rest
AnswerB

The tool systematically benchmarks different combinations of runtime parameters to identify the most efficient setup. This automated approach ensures the model meets performance targets reliably without the human error inherent in manual configuration of complex inference engines.

Why this answer

The Model Analyzer is a critical tool for determining the optimal configuration for models deployed on Triton Inference Server. By automating the testing of various batch sizes and concurrency levels, it ensures that models are tuned for maximum throughput and minimum latency. This tuning is essential for maintaining reliability and cost-efficiency in large-scale production, preventing performance degradation caused by suboptimal configuration settings that often go unnoticed in manual deployments.

Exam trap

Test-takers often confuse Model Analyzer with profiling tools used for training or code execution time, missing its specific role in searching optimal Triton configurations.

336
MCQhard

A team is serving a 70B-parameter LLM with TensorRT-LLM in a multi-tenant environment where requests arrive with widely varying prompt lengths and generation lengths. During load testing, they observe that throughput collapses when a long-context request is scheduled alongside many short requests, and GPU memory fragmentation causes intermittent out-of-memory errors even though total free memory appears sufficient. Which TensorRT-LLM runtime configuration change most directly addresses both the throughput collapse and the memory fragmentation?

A.Increase the TensorRT-LLM build max_batch_size and max_seq_len to their maximum possible values so the engine supports every request shape.
B.Switch the engine from FP16 to INT8 precision using a calibration dataset that matches the production prompt distribution.
C.Reduce the number of concurrent client connections at the load balancer so that only one request is processed by the GPU at a time.
D.Enable the paged KV cache with a tuned block size and configure the batch scheduler to use in-flight batching with a per-iteration token budget.
AnswerD

Paged KV cache allocates KV memory in fixed-size blocks from a shared pool, eliminating the external fragmentation that causes OOM. In-flight batching lets the scheduler add new requests and retire finished ones every iteration, so short requests are not blocked behind a long one. Together they directly resolve both the memory fragmentation and the head-of-line blocking that collapses throughput.

Why this answer

The paged KV cache solves external fragmentation by allocating KV memory in uniform blocks from a shared pool, which prevents the OOM that occurs when a long-context request needs a large contiguous region. In-flight batching with a token budget lets the scheduler continuously admit and retire requests each iteration, so short requests no longer wait behind a long generation. These two runtime settings together address both symptoms without sacrificing accuracy or throughput.

Exam trap

The trap here is assuming that increasing max_batch_size and max_seq_len gives the engine more flexibility, when in fact it reserves more memory and does nothing to prevent fragmentation or scheduling head-of-line blocking.

337
MCQhard

A production LLM inference service on NVIDIA Triton Inference Server runs on a multi-GPU node. You observe that one GPU reports ECC XID errors and the model's throughput gradually degrades. Which NVIDIA tool should you use to monitor GPU health and set up alerts for these errors?

A.NVIDIA Data Center GPU Manager (DCGM)
B.NVIDIA TensorRT
C.NVIDIA Nsight Systems
D.NVIDIA Triton Model Analyzer
AnswerA

DCGM is designed for monitoring and managing NVIDIA data center GPUs. It can detect ECC errors, XID errors, and other health metrics, and can be integrated with alerting systems. In this scenario, DCGM would provide the necessary visibility into GPU health and enable proactive alerts when errors occur, helping maintain reliability.

Why this answer

NVIDIA DCGM is the standard tool for monitoring GPU health in data centers. It tracks ECC and XID errors, temperature, power, and other metrics, and can trigger alerts. For a production LLM service experiencing hardware errors, DCGM provides the necessary observability to detect and respond to GPU failures.

Exam trap

The trap here is confusing profiling tools like Nsight Systems or optimization tools like TensorRT with health monitoring tools.

338
Multi-Selecthard

Which THREE factors significantly influence the memory consumption during LLM fine-tuning? (Choose three)

Select 3 answers
A.The precision of the optimizer states
B.The number of trainable parameters
C.The activation maps stored for the backward pass
D.The number of CPU cores available for data loading
E.The version of the operating system kernel
AnswersA, B, C

The optimizer (e.g., AdamW) stores states for every trainable parameter. If using 32-bit precision, these states can occupy 8 bytes per parameter. Reducing this precision or utilizing paged techniques is essential for saving VRAM, as optimizer states are one of the largest contributors to memory exhaustion during the training cycle.

Why this answer

Memory consumption is driven by the model parameters, the optimizer states (which are often larger than the model weights), and the activations generated during the forward pass. Efficient management of these components is critical for scaling to larger models. By optimizing these three areas, practitioners can successfully fit large models into limited VRAM, ensuring stability throughout the fine-tuning process.

Exam trap

Candidates mistakenly select dataset size or sequence length as primary hardware memory drivers, overlooking how optimizer states, trainable parameters, and activation maps dictate fine-tuning VRAM consumption.

339
MCQhard

When evaluating LLM reliability under stress, what is the primary goal of conducting 'Chaos Engineering' on a Triton inference cluster?

A.To increase the total throughput of the model
B.To verify system recovery and failover logic
C.To reduce the physical power usage of GPUs
D.To train the model on noisy data
AnswerB

Validating that the system automatically handles failures is the cornerstone of chaos engineering. By simulating real-world failures, engineers can confirm that the failover mechanisms, such as health checks and load rebalancing, function correctly in an automated and reliable manner.

Why this answer

Chaos Engineering involves deliberately injecting failures, such as network latency or node crashes, to test system resilience. The goal is to verify that the system can automatically recover and maintain service levels despite these disruptions. This is vital for production-grade LLM reliability, as it identifies hidden weaknesses in the failover logic or load balancing configurations before a real-world outage occurs, ensuring the cluster behaves predictably under adverse conditions.

Exam trap

Candidates often confuse Chaos Engineering with performance testing or load testing. The primary distinction is that Chaos Engineering is specifically about verifying system resilience and recovery during failures.

340
MCQhard

An LLM service on NVIDIA Triton Inference Server experiences high time-to-first-token because the dynamic batcher waits for full batches. The team wants to reduce time-to-first-token while still benefiting from batching. Which adjustment is most appropriate?

A.Set preferred_batch_size to a very large value.
B.Disable dynamic batching entirely for the model.
C.Decrease max_queue_delay_microseconds to shorten batching wait time.
D.Increase max_queue_delay_microseconds to allow larger batches.
AnswerC

The dynamic batcher waits up to max_queue_delay_microseconds before dispatching a batch. Reducing this value shortens the wait, so requests begin processing sooner and time-to-first-token drops. The model still batches whatever requests are available, preserving some throughput benefit while improving responsiveness for interactive workloads.

Why this answer

Time-to-first-token is driven by how long the dynamic batcher holds requests before dispatch. Lowering max_queue_delay_microseconds shortens that wait, so requests start sooner while still being batched with any available peers. Raising the delay, inflating preferred batch size, or disabling batching either increases latency or sacrifices throughput unnecessarily.

Exam trap

The trap here is assuming that larger batches or longer queue delays always improve the user experience, when they actually raise time-to-first-token.

341
MCQhard

A bank's model risk committee is reviewing an LLM-based loan-adverse-action notice generator built on NVIDIA NeMo. Regulators require that the system produce a human-readable rationale for each denial and that the rationale be reproducible for any prior decision. Which architectural choice most directly meets both obligations?

A.Increase the model temperature so the generator explores multiple rationales and select the most favorable one for each applicant.
B.Persist the exact prompt, model version, decoding parameters, retrieved context, and random seed for every decision so the notice can be regenerated identically.
C.Cache the generated notice text in a database keyed by applicant ID and serve the cached copy whenever the same applicant is re-scored.
D.Replace the LLM with a logistic regression scorecard so that every denial has a coefficient-based explanation by construction.
AnswerB

Reproducibility in generative systems comes from capturing every input that influences the output: the prompt template, the pinned model version, temperature and top-p settings, any retrieved context, and the seed. With those artifacts stored, the bank can regenerate the identical notice during an examination. Human readability is preserved because the stored prompt and template define the rationale structure the model was instructed to follow.

Why this answer

Auditable generative decisions require capturing the full provenance chain: prompt, model version, decoding configuration, retrieval context, and seed. Storing those artifacts lets the bank reconstruct any prior notice byte-for-byte during an examination while keeping the natural-language rationale the regulation demands. Changing the model class, raising temperature, or merely caching outputs each fails one of the two obligations the committee must satisfy.

Exam trap

The trap here is assuming that saving the generated text is the same as saving the ability to explain it, when reproducibility actually depends on the full input and configuration provenance.

342
MCQmedium

When evaluating an LLM's response to a complex prompt, what is the 'Persona Adoption' technique?

A.A method to verify if the model has memorized personal data.
B.A method to assign an expert identity to improve response quality.
C.A way to force the model to identify the user's persona.
D.A technique to reduce the model's context window usage.
AnswerB

By setting a persona, such as 'Senior NVIDIA GPU Architect,' the model is primed to utilize more relevant technical terminology and adopt a problem-solving approach consistent with that role. This significantly improves the quality and relevance of the response compared to a generic or default conversational persona.

Why this answer

Persona adoption involves assigning a specific role, expertise level, or professional identity to the LLM within the system prompt. This technique helps calibrate the tone, vocabulary, and depth of the response to match the user's expectations. For NVIDIA applications, this ensures that the model speaks with the authority and technical precision required for engineering and developer-facing communications.

Exam trap

Candidates confuse persona adoption with few-shot prompting or fine-tuning, thinking it requires training data rather than simple system prompt instructions.

343
MCQmedium

A hospital network runs an on-premises NVIDIA NIM microservice hosting a clinical-summarization LLM. Compliance requires that every generated summary be attributable to source records and that no protected health information leave the subnet. Which deployment practice best satisfies both requirements at once?

A.Fine-tune the clinical LLM on de-identified records, then deploy the tuned checkpoint to the public cloud region closest to the hospital.
B.Route prompts to a hosted public LLM API for higher quality, then hash the returned summaries before writing them to the clinical record.
C.Enable NVIDIA NeMo Guardrails output rails to redact names and dates, and continue calling the external model endpoint from the clinical application.
D.Keep inference inside the on-premises NIM endpoint and log retrieval-augmented-generation citations that map each summary sentence back to the source record IDs.
AnswerD

Running the NIM microservice on-premises keeps PHI inside the controlled subnet, and citation-based RAG grounding ties every generated sentence to a retrievable source record, satisfying auditability and data-residency simultaneously. Because the retriever and the NIM endpoint are both local, no prompt or completion crosses the trust boundary, so the attributable-evidence trail never depends on an external service.

Why this answer

On-premises NVIDIA NIM inference keeps protected health information inside the controlled network boundary, and RAG citation logging provides the source-record traceability an auditor needs. Redaction, hashing, and de-identified fine-tuning each address only part of the problem and none of them stops live PHI from crossing the trust boundary. Only a local endpoint combined with grounded citations satisfies residency and attribution together.

Exam trap

The trap here is treating output redaction or hashing as equivalent to preventing data egress, when the residency violation already occurs the moment the prompt is transmitted off-subnet.

344
MCQhard

Refer to the exhibit. An audit reveals that 'INTERNAL_STRATEGY' documents are still being generated by the model. Why is this occurring?

A.The log_level is set to DEBUG, which disables the rejection mechanism.
B.The 'mode' is set to PERMISSIVE, preventing the engine from blocking matches.
C.The 'allow_list' contains too many items, causing a conflict with the 'reject_list'.
D.The engine requires a higher GPU clock speed to process the reject_list.
AnswerB

Setting the mode to PERMISSIVE tells the guardrail engine to allow questionable output, likely to avoid false positives. This configuration allows restricted content to leak through. To ensure compliance and stop the generation of forbidden content, the mode must be changed to one that enforces the rejection list strictly.

Why this answer

The 'mode: PERMISSIVE' setting indicates that the guardrail system is allowing the model output even if there is a partial match or a low-confidence detection of forbidden content. In a production environment, permissive modes are dangerous because they prioritize utility over strict security. To block the sensitive strategy data, the configuration must be set to a stricter enforcement mode that prioritizes safety and rejects any output matching the forbidden list.

Exam trap

Candidates assume the model lacks the ability to detect the documents, missing the configuration setting where permissive enforcement allows low-confidence or partial matches to pass through.

345
MCQeasy

A company is using NVIDIA NeMo Guardrails to enforce safety policies in its LLM application. A developer wants to ensure that the model does not generate content that violates the company's policy against discussing competitor products. Which type of guardrail should the developer configure to prevent the model from mentioning competitor names in its responses?

A.Output rail that detects and blocks responses containing competitor names
B.Input rail that filters user queries containing competitor names
C.Dialog rail that redirects the conversation if a competitor is mentioned
D.Retrieval rail that filters the knowledge base for competitor information
AnswerA

An output rail inspects the model's generated response and can block or modify it if it violates a policy. By configuring an output rail to detect competitor names, the developer ensures that any response mentioning them is intercepted before reaching the user. This directly enforces the policy against discussing competitor products.

Why this answer

To prevent the model from generating competitor names in its responses, the developer should use an output rail. Output rails in NeMo Guardrails inspect the LLM's response and can block or modify it based on defined policies. Input, dialog, and retrieval rails address other parts of the pipeline and do not directly control the final output content.

Exam trap

The trap here is assuming that input filtering or retrieval filtering alone can prevent the model from generating specific content, when output inspection is needed.

346
MCQmedium

A production LLM inference service runs on NVIDIA Triton Inference Server across multiple GPUs. The SRE team wants to detect when the service starts returning incorrect or degraded responses compared to a baseline, even when latency and throughput remain normal. Which monitoring approach is most appropriate?

A.Monitor GPU utilization and memory usage via NVIDIA DCGM to detect anomalies.
B.Implement model output quality checks by comparing responses against a baseline using statistical or embedding-based metrics.
C.Use NVIDIA Nsight Systems to profile the inference pipeline and identify bottlenecks.
D.Set up alerting on Triton Inference Server's error rate and request latency percentiles.
AnswerB

Comparing live model outputs to a baseline using metrics such as BLEU, ROUGE, or embedding similarity can detect semantic drift or degradation that doesn't affect latency or throughput. This directly addresses the need to monitor response correctness. It is the most appropriate because it focuses on output quality, which is the core concern here.

Why this answer

To detect degraded or incorrect responses while latency and throughput are normal, monitoring must focus on output quality. Comparing model outputs to a baseline using statistical or embedding-based metrics can reveal semantic drift or accuracy drops. Resource and performance metrics like GPU utilization, error rates, and latency are insufficient because they do not reflect the correctness of generated text.

Exam trap

The trap here is assuming that normal latency and throughput imply correct model behavior, overlooking the need for output quality monitoring.

347
MCQmedium

Which optimization method should be prioritized when the model's inference performance is bottlenecked by the CPU-to-GPU data transfer overhead?

A.Quantization
B.GPU-based Pre-processing
C.Model Pruning
D.Operator Fusion
AnswerB

Moving pre-processing to the GPU eliminates the need to transfer intermediate results over the slow PCIe bus. By executing the full pipeline on the GPU device, the system avoids the overhead of host-to-device transfers, which is critical for achieving low-latency inference in real-time generative applications.

Why this answer

When data transfer between the host and GPU is the bottleneck, the most effective strategy is to move the pre-processing logic onto the GPU itself. Using CUDA-accelerated kernels to perform operations like tokenization or normalization directly on the GPU avoids moving data across the PCIe bus, which is a slow operation that stalls the GPU during the inference pipeline.

Exam trap

Candidates often suggest optimizing model weights or using faster interconnects like NVLink, ignoring that moving pre-processing to the GPU is the most direct way to eliminate PCIe bus stalling.

348
MCQhard

A team is preparing a mixed-language corpus for continued pretraining of an NVIDIA NeMo Megatron model. The corpus contains English, Japanese, and Arabic documents. Tokenizer analysis shows the current English-centric BPE vocabulary produces very long token sequences for Japanese and Arabic, inflating sequence length and compute cost. The team wants to reduce sequence length for non-English text without retraining the tokenizer from scratch and without degrading English performance. Which data preparation action best achieves this?

A.Lowercase and strip diacritics from all Japanese and Arabic text so that the English-centric vocabulary can represent it more efficiently.
B.Transliterate all Japanese and Arabic documents into Latin script using a standard romanization scheme before tokenization.
C.Increase the model's maximum sequence length and positional embedding size so that long token sequences for Japanese and Arabic fit without truncation.
D.Extend the existing BPE vocabulary with additional merges learned from a balanced multilingual sample, then retrain only the embedding and output layers while freezing the rest of the model.
AnswerD

Adding multilingual merges to the existing vocabulary reduces the number of tokens needed to represent Japanese and Arabic text, directly cutting sequence length and compute. Keeping the original English merges preserves English tokenization behavior, and retraining only the embedding and output layers adapts the new vocabulary entries without disturbing the pretrained transformer weights, which limits the risk of degrading English performance.

Why this answer

The high token counts for Japanese and Arabic stem from a vocabulary that lacks subword units for those scripts. Extending the BPE vocabulary with multilingual merges shortens their token sequences while preserving English merges, and retraining only the embedding and output layers adapts the new entries without perturbing the pretrained transformer. This lowers compute cost and keeps English behavior stable.

Exam trap

The trap here is treating long token sequences as a context-length problem to be solved by enlarging the window rather than as a vocabulary coverage problem rooted in the tokenizer.

349
MCQmedium

An engineer is using NVIDIA TensorRT to optimize a Transformer model for inference on an NVIDIA A100 GPU. They want to maximize throughput while ensuring that the model runs correctly with varying input sequence lengths. Which TensorRT feature should they configure to allow the engine to handle different input shapes at runtime?

A.Static shapes with fixed dimensions
B.Dynamic shapes with optimization profiles
C.INT8 calibration with a representative dataset
D.Multiple engines for each possible input length
AnswerB

Dynamic shapes with optimization profiles allow a TensorRT engine to accept input tensors of different dimensions at runtime. The profiles define minimum, optimal, and maximum shapes for each input, enabling the engine to select the best kernel for the actual shape. This is essential for handling varying sequence lengths in Transformer models without rebuilding the engine.

Why this answer

Dynamic shapes with optimization profiles enable a single TensorRT engine to handle inputs of different sizes by defining shape ranges. This allows the engine to optimize for the actual input shape at runtime, which is crucial for Transformer models with variable sequence lengths. Static shapes, multiple engines, or INT8 calibration do not provide this flexibility.

Exam trap

The trap here is confusing quantization (INT8 calibration) with shape flexibility, but calibration only affects precision, not the ability to accept different input dimensions.

350
Multi-Selectmedium

Which THREE factors should be considered when choosing an optimal batch size for LLM inference on NVIDIA GPUs?

Select 3 answers
A.Available GPU VRAM
B.Inference Latency SLA
C.Total Compute Throughput
D.Model Training Loss
E.CPU Clock Speed
AnswersA, B, C

The available VRAM is a hard constraint on the batch size. As the batch size increases, the memory required to store intermediate activations and the KV cache also grows. If the batch size is too large, the system will encounter out-of-memory errors, making memory capacity the primary limiting factor.

Why this answer

Selecting an optimal batch size involves a balance between hardware saturation, memory capacity, and latency requirements. Larger batches improve throughput by better utilizing GPU compute cores but increase memory demand and individual request latency. The goal is to reach the highest throughput possible while staying within the VRAM limit and meeting the required latency SLA for the specific application.

Exam trap

Candidates often ignore the Inference Latency SLA, focusing solely on maximizing compute throughput. They fail to realize that an overly large batch size can lead to unacceptable latency for real-time applications.

351
MCQmedium

A team is fine-tuning a 13B-parameter Llama model with NVIDIA NeMo Framework on a single A100 80GB GPU. They apply LoRA adapters to the attention projection layers, but the adapters are producing negligible changes to model behavior even after several epochs, and the loss curve stays flat. They confirm the dataset is clean and the tokenizer is correct. Which LoRA configuration issue is the most likely cause?

A.The optimizer was configured with a momentum value of zero, preventing the adapter weights from accumulating updates.
B.The LoRA target modules were set to the embedding layer rather than the attention projection layers.
C.The base model weights were accidentally left trainable, so gradients are flowing into the full model instead of the adapters.
D.The LoRA alpha scaling value is set extremely low relative to the rank, so the adapter's effective update magnitude is near zero.
AnswerD

The LoRA scaling factor is alpha divided by rank. If alpha is tiny compared to rank (for example alpha=1 with rank=64), the adapter's contribution to the forward pass is scaled to a negligible magnitude, so the frozen base weights dominate and the loss barely moves. Raising alpha or lowering rank restores a meaningful update magnitude.

Why this answer

LoRA updates are scaled by alpha divided by rank, so an alpha that is very small relative to rank shrinks the adapter's effective contribution and can leave the loss nearly unchanged. Correcting the alpha-to-rank ratio restores meaningful adapter influence, allowing the fine-tune to actually shift model behavior on the target task.

Exam trap

The trap here is assuming any LoRA misconfiguration equally explains a flat loss, when the alpha-to-rank scaling ratio specifically controls how much the adapter can influence the forward pass.

352
MCQmedium

A financial services company has fine-tuned a Llama-3-70B model using NVIDIA NeMo for automated loan risk assessment. During evaluation, they observe that the model achieves high scores on standard accuracy metrics but produces inconsistent responses when the same query is phrased slightly differently. They need a metric that quantifies this inconsistency. Which evaluation metric should they use?

A.Semantic similarity between outputs for paraphrased inputs
B.ROUGE score
C.BLEU score
D.Perplexity
AnswerA

Semantic similarity measures how closely the model's responses align in meaning when the same question is asked with different wording. High similarity indicates consistent understanding and stable generation. In this scenario, computing similarity across paraphrased loan queries directly quantifies the observed inconsistency. This metric is appropriate for detecting and reducing output variance.

Why this answer

The team needs to measure output consistency across semantically equivalent inputs. Semantic similarity between responses to paraphrased queries directly captures this variance, unlike lexical overlap metrics such as BLEU or ROUGE, which compare to references. Perplexity reflects language modeling confidence, not answer stability.

Therefore, semantic similarity is the correct choice for quantifying inconsistency in the fine-tuned model's responses.

Exam trap

The trap here is assuming that high accuracy on standard benchmarks implies consistent behavior across paraphrased inputs, which it does not.

Page 4

Page 5 of 5

All pages