Courseiva

NVIDIA Certified Professional: Generative AI LLMs (NCP-GENL) — Questions 76–150

352 questions total · 5pages · All types, answers revealed

Page 1

Page 2 of 5

Page 3
76
MCQhard

A financial services company is fine-tuning an LLM to answer questions about internal policies. The base model performs well on general text but frequently invents policy numbers and effective dates. The team has a curated dataset of 5,000 question-answer pairs with correct citations. Which fine-tuning approach best addresses the hallucination of policy numbers and dates?

A.Increase the model's temperature during inference so it explores a wider range of possible policy numbers and dates.
B.Apply reinforcement learning from human feedback using a reward model that penalizes any answer longer than two sentences.
C.Perform supervised fine-tuning on the curated question-answer pairs so the model learns to produce answers with correct citations.
D.Continue pre-training on a large corpus of internal documents to expand the model's knowledge of policy numbers and dates.
AnswerC

Supervised fine-tuning on curated question-answer pairs with correct citations directly teaches the model the desired input-output behavior. By training on examples that include accurate policy numbers and effective dates, the model learns to associate questions with grounded answers rather than plausible-sounding fabrications, which is the most direct remedy for the observed hallucination pattern.

Why this answer

Supervised fine-tuning on curated question-answer pairs with correct citations is the most direct way to teach the model the desired behavior. The training examples pair each question with grounded policy numbers and effective dates, so the model learns to reproduce accurate citations instead of inventing them. Continued pre-training, length-based RLHF, and higher inference temperature do not target the specific hallucination of policy details.

Exam trap

The trap here is reaching for continued pre-training or RLHF when the scenario already supplies labeled question-answer pairs, which are the natural input for supervised fine-tuning.

77
MCQhard

A team uses an NVIDIA NIM-hosted model to draft release notes from a changelog. Reviewers report the drafts omit minor fixes and overstate the significance of small changes. Which prompt engineering adjustment BEST addresses both issues?

A.Instruct the model to include every changelog entry, to describe each change at its actual scope without exaggeration, and to organize the output by category such as features, fixes, and known issues.
B.Instruct the model to summarize only the three most impactful changes so the release notes stay concise and readable.
C.Add a single example of a polished release note and ask the model to 'write something similar' for the new changelog.
D.Increase the model's max_tokens and lower top_p so the output can be longer and more focused on the most probable phrasing.
AnswerA

Explicitly requiring complete coverage of all entries addresses omissions, while the accuracy instruction discourages inflating small changes. Categorizing output gives the model a structural checklist that makes it easier to verify that no entry was dropped and keeps minor fixes visible rather than buried.

Why this answer

The reported defects are content-level problems: missing entries and distorted emphasis. Only explicit instructions can fix them, by requiring complete coverage, accurate scoping of each change, and a categorized structure that makes omissions easy to spot. Sampling parameters and single stylistic examples influence form, not editorial completeness or proportionality.

Exam trap

The trap here is reaching for decoding parameters or a style example to fix what are actually content-coverage and accuracy problems that only explicit instructions can resolve.

78
MCQeasy

A developer is preparing a supervised fine-tuning dataset for an instruction-tuned LLM using NVIDIA NeMo. The dataset contains prompts and responses, but the model sometimes learns to generate the prompt text as part of the response. Which dataset formatting practice should be applied to prevent this?

A.Shuffle the dataset so prompts and responses are randomly paired during training.
B.Increase the maximum sequence length so the prompt and response always fit in a single sample.
C.Duplicate the prompt in both the input and output fields to reinforce the instruction format.
D.Use a loss mask that excludes prompt tokens so the model is trained only on response tokens.
AnswerD

A loss mask sets the loss contribution of prompt tokens to zero, so the model is optimized only on the response tokens. This prevents the model from learning to reproduce the prompt as output. In NeMo, this is handled through the data configuration and tokenizer settings that mark which tokens contribute to the loss.

Why this answer

Prompt echoing occurs when the loss is computed over both prompt and response tokens. Applying a loss mask that excludes prompt tokens ensures gradients are only derived from response tokens, which teaches the model to generate answers rather than repeat instructions. Duplicating prompts, extending sequence length, or shuffling pairs do not address the loss computation.

Exam trap

The trap here is thinking that dataset formatting alone controls what the model learns, when the loss mask determines which tokens actually contribute gradients.

79
MCQhard

An engineer is optimizing a Transformer-based LLM for inference on an NVIDIA A100 GPU. The model uses FP16 precision, but during generation, the GPU's Tensor Cores are underutilized, and latency is higher than expected. Profiling reveals that many small matrix multiplications are executed sequentially. Which technique is most effective to improve Tensor Core utilization and reduce latency?

A.Enable CUDA graphs to capture and replay the sequence of kernel launches.
B.Use NVIDIA TensorRT to quantize the model to INT8.
C.Increase the batch size to amortize kernel launch overhead.
D.Apply kernel fusion to combine multiple small operations into a single kernel.
AnswerD

Kernel fusion merges multiple small matrix multiplications and element-wise operations into a single kernel, reducing launch overhead and enabling better Tensor Core utilization by processing larger, combined workloads. This is especially effective in Transformer inference where many small GEMMs occur. By fusing operations, the GPU can execute them more efficiently, lowering latency and increasing utilization of Tensor Cores.

Why this answer

The underutilization of Tensor Cores stems from many small, sequential matrix multiplications. Kernel fusion combines these into larger operations, allowing Tensor Cores to process more data per instruction and reducing kernel launch overhead. This directly improves utilization and lowers latency, making it the most effective technique among the options for this specific bottleneck.

Exam trap

The trap here is focusing on quantization or CUDA graphs, which address different bottlenecks, rather than recognizing that small sequential operations require fusion to improve Tensor Core efficiency.

80
Multi-Selecthard

When building an NVIDIA NeMo LLM application for automated document review, which THREE of the following prompt design choices are critical for ensuring high-quality output? (Select exactly THREE)

Select 3 answers
A.Use standardized delimiters to isolate user-provided documents.
B.Include instructions that prioritize speed over accuracy.
C.Require the output to be in a machine-readable format like JSON.
D.Instruct the model to perform a chain-of-thought validation of its findings.
E.Use a high temperature setting to ensure diverse review perspectives.
AnswersA, C, D

Properly isolating document segments prevents the model from conflating the input data with its own internal knowledge or instructions. This clarity is essential for document review applications where accuracy is paramount, as it ensures the model is specifically analyzing the provided text rather than hallucinating based on external training.

Why this answer

Effective document review requires accuracy, traceability, and consistency. Using clear delimitation prevents data corruption, requiring structured output (like JSON) allows for downstream programmatic integration, and chain-of-thought prompts ensure the model validates its conclusions against the text. These choices together create a robust, production-ready pipeline that minimizes errors and provides the necessary structure for automated workflows in enterprise environments.

Exam trap

Candidates often neglect the importance of machine-readable output formats, focusing only on the content of the response while ignoring the necessity of programmatic integration for automated document review workflows.

81
MCQhard

An engineer is deploying a LLM using NVIDIA Triton Inference Server with the TensorRT-LLM backend. They need to ensure that the model can handle a sudden surge in requests without increasing latency beyond a specified threshold. They have configured the model with a maximum batch size of 32 and dynamic batching with a preferred batch size of 16. However, during peak load, latency spikes are observed. Which Triton configuration parameter should they adjust to control the maximum time a request waits in the dynamic batching queue before being processed?

A.max_batch_size
B.instance_group
C.preferred_batch_size
D.max_queue_delay_microseconds
AnswerD

max_queue_delay_microseconds sets the maximum time a request can wait in the dynamic batching queue before Triton processes it, even if the preferred batch size is not reached. Reducing this value limits latency spikes during peak load by forcing earlier execution, though it may reduce batching efficiency. It directly controls the trade-off between latency and throughput.

Why this answer

The max_queue_delay_microseconds parameter in Triton's dynamic batching configuration specifies the maximum time a request can wait in the queue before being processed. Lowering it reduces latency spikes under surge conditions by preventing requests from waiting too long for a full batch. Other parameters like max_batch_size and preferred_batch_size affect batching but not the wait timeout.

Exam trap

The trap here is assuming that increasing batch size or instance count will solve latency spikes, when the root cause is the queue wait time controlled by max_queue_delay_microseconds.

82
MCQhard

When fine-tuning on a small, domain-specific dataset, why might adding synthetic data generated by a larger model be beneficial?

A.It eliminates the need for any real-world human-annotated data entirely.
B.It helps the model generalize better by increasing data diversity.
C.It reduces the VRAM usage of the training process significantly.
D.It guarantees the model will not have any factual inaccuracies.
AnswerB

By providing a larger volume of varied examples, synthetic data helps the model learn to handle diverse inputs within the domain. This diversity is crucial when the primary dataset is limited, as it prevents the model from overfitting to a narrow set of patterns, effectively increasing its overall robustness.

Why this answer

Small datasets often lack the depth needed for a model to generalize effectively. Synthetic data can fill these gaps by providing more examples of the target task, which helps the model learn the nuances of the domain. This technique, when done correctly, reinforces desired behaviors and prevents overfitting on the limited original data, leading to a more robust, versatile, and high-performing model in the intended application area.

Exam trap

Candidates often assume synthetic data is only for increasing volume, missing the key benefit of improving generalization and reducing overfitting when dealing with limited, niche, or domain-specific training sets.

83
MCQeasy

What is the primary function of the 'Triton Model Analyzer' in an optimization workflow?

A.It converts models to TensorRT
B.It automatically quantizes the model
C.It benchmarks configuration trade-offs
D.It manages model version control
AnswerC

The Model Analyzer benchmarks various deployment configurations (like batch size and instance count) to identify the settings that offer the best performance. This allows engineers to make data-driven decisions when deploying models to ensure they maximize resource utilization while staying within latency and throughput constraints.

Why this answer

The Model Analyzer is a tool designed to explore the trade-offs between throughput, latency, and memory usage for different model configurations. It automatically runs benchmarks with varying batch sizes and instance counts, providing developers with empirical data to find the optimal deployment parameters that meet their specific service level agreements for generative AI applications.

Exam trap

Examinees often mistake the Triton Model Analyzer for a profiling tool that measures training convergence, confusing runtime inference deployment trade-offs with model training metrics.

84
MCQmedium

You are curating a 2 TB corpus of NVIDIA technical documentation and Python code for continued pretraining of a NeMo-based LLM. A colleague proposes filtering out any document containing the token sequence 'CUDA' to reduce hardware-specific bias. What is the most appropriate response?

A.Accept the filter, because removing every occurrence of a hardware-specific token guarantees the model will not overfit to NVIDIA-specific APIs and will generalize better to other vendors.
B.Accept the filter but apply it only to documents where 'CUDA' appears more than ten times, since low-frequency occurrences are harmless and high-frequency ones indicate redundant marketing material.
C.Reject the filter, because removing documents based on a single high-signal domain token destroys the domain-specific signal the model needs and is better handled by deduplication and quality scoring.
D.Accept the filter, because NVIDIA NeMo requires that proprietary product names be excluded from pretraining corpora to comply with the model card's data provenance requirements.
AnswerC

The proposed filter is a crude keyword removal that discards entire documents whose core subject is the target domain. In NeMo Curator pipelines, quality is improved through exact and fuzzy deduplication, heuristic quality filters, and classifier-based scoring, not by excising a single token that defines the domain. Removing these documents would leave the model undertrained on the very terminology it must learn.

Why this answer

Keyword-based deletion of a domain-defining token is a destructive filter, not a quality filter. Effective NeMo data curation relies on deduplication, quality heuristics, and classifier scoring to remove low-value or redundant records while preserving coherent, in-domain technical content. The corpus exists to teach NVIDIA-specific concepts, so removing 'CUDA' would directly undermine the training objective.

Exam trap

The trap here is assuming that removing a vendor-specific keyword reduces bias, when it actually strips high-value domain signal that the continued pretraining corpus was built to provide.

85
Multi-Selectmedium

A team is building a TensorRT-LLM engine for a 7B model that must serve both single-turn short prompts and long multi-turn conversations with a shared system prompt. They want to maximize reuse of computation across requests without changing model weights. Which TWO techniques should they enable? (Choose two.)

Select 2 answers
A.Enable paged KV cache with block reuse so freed sequence blocks return to the pool for other requests.
B.Apply INT4 weight-only quantization to the attention projection matrices to shrink the cache footprint.
C.Enable KV cache reuse so identical prompt prefixes share cached key/value blocks.
D.Set the builder to use strongly typed engines so the plugin graph is fully deterministic.
E.Increase the beam width so multiple candidate continuations are evaluated per request.
AnswersA, C

The paged KV cache allocates fixed-size blocks and, with block reuse, returns blocks from finished sequences to a shared pool rather than fragmenting memory. Mixed short and long prompts then coexist efficiently, raising the effective batch size and preventing allocation failures without any change to the underlying model weights.

Why this answer

Sharing a system prompt and multi-turn history means many requests share long identical prefixes, and mixed short and long prompts create variable cache pressure. KV cache reuse skips recomputing cached prefixes, while the paged KV cache with block reuse keeps memory defragmented and available to new sequences. Together they cut redundant prefill and raise concurrency without touching model weights.

Exam trap

The trap here is conflating weight memory with KV cache memory, so an engineer reaches for weight quantization when the reuse problem is really about activations and cache block lifecycle.

86
MCQmedium

An LLM inference service on NVIDIA Triton Inference Server is experiencing intermittent failures. The operations team wants to set up alerting to detect when the GPU memory utilization exceeds 90% for more than 5 minutes, as this could lead to out-of-memory errors. Which combination of tools should they use to achieve this?

A.NVIDIA DCGM-Exporter and Prometheus Alertmanager
B.nvidia-smi with custom scripting and email notifications
C.NVIDIA Nsight Systems and Prometheus
D.NVIDIA Triton Inference Server metrics and Grafana
AnswerA

DCGM-Exporter exposes GPU metrics, including memory utilization, to Prometheus. Prometheus can then evaluate alerting rules, such as memory utilization >90% for 5 minutes, and trigger alerts via Alertmanager. This combination is standard for GPU monitoring and alerting in production environments, providing reliable and scalable detection of potential OOM conditions.

Why this answer

DCGM-Exporter collects GPU metrics and exposes them to Prometheus. Prometheus evaluates alerting rules based on thresholds and durations, and Alertmanager sends notifications. This setup is ideal for detecting high GPU memory utilization over time, helping prevent OOM errors in LLM inference.

Exam trap

The trap here is assuming that Triton's built-in metrics include GPU memory utilization, when in fact they are inference-specific and require DCGM-Exporter for GPU telemetry.

87
MCQhard

A production LLM service on NVIDIA Triton Inference Server is deployed across multiple GPUs. The team notices that one GPU consistently shows higher latency for inference requests compared to others, despite similar utilization. Which NVIDIA tool should be used to investigate per-GPU performance discrepancies and identify bottlenecks?

A.NVIDIA Nsight Systems
B.NVIDIA Triton Inference Server metrics
C.NVIDIA Data Center GPU Manager (DCGM)
D.NVIDIA System Management Interface (nvidia-smi)
AnswerA

Nsight Systems is a system-wide profiling tool that can capture GPU activities, including kernel execution times and memory transfers, across multiple GPUs. It can highlight per-GPU performance differences by showing timelines and bottlenecks. In this scenario, it can identify why one GPU is slower, such as inefficient kernel usage or thermal throttling, enabling targeted optimization.

Why this answer

Nsight Systems provides detailed profiling of GPU activities, including kernel execution, memory transfers, and synchronization. By capturing a timeline across all GPUs, it can reveal why one GPU lags, such as longer kernel runtimes or thermal issues. This level of detail is essential for diagnosing per-GPU performance discrepancies and optimizing LLM inference in a multi-GPU deployment.

Exam trap

The trap here is relying on monitoring tools like DCGM or nvidia-smi that provide metrics but not the granular profiling needed to identify kernel-level bottlenecks on a specific GPU.

88
MCQeasy

What is the primary function of the 'NeMo Guardrails' toolkit in an enterprise AI pipeline?

A.To increase the GPU memory utilization and throughput of the inference engine.
B.To enforce safety and alignment constraints on LLM interactions.
C.To compress large models into smaller representations for faster deployment.
D.To automate the labeling of training data for supervised fine-tuning.
AnswerB

The primary role of the toolkit is to act as a governance layer that enforces business, safety, and ethical policies. By intercepting inputs and outputs, it ensures that the model operates within predefined constraints, preventing harmful, biased, or unauthorized content from being generated for the end user.

Why this answer

NeMo Guardrails serves as a specialized layer that sits between the user and the LLM, managing the interaction to ensure safety and alignment. It enables developers to define boundaries, prevent specific topics, and ensure the model adheres to enterprise policies. This is vital for mitigating risks like brand damage, legal non-compliance, and the leakage of intellectual property during user interactions with generative AI systems.

Exam trap

Candidates mistakenly select model fine-tuning or training optimization options, confusing safety alignment wrappers with core model pre-training procedures.

89
MCQeasy

A developer is building a retrieval-augmented generation pipeline and needs to choose a component that produces dense vector representations of passages for semantic search. The passages are up to 512 tokens long, and the developer wants a model specifically trained to map semantically similar text to nearby points in embedding space. Which type of model should be selected?

A.An autoregressive sequence-to-sequence model fine-tuned for summarization.
B.A cross-encoder that jointly encodes the query and passage.
C.A decoder-only generative LLM fine-tuned for instruction following.
D.A bidirectional encoder model trained with a contrastive objective on sentence pairs.
AnswerD

Bidirectional encoders trained with contrastive objectives, such as sentence-transformer style models, are explicitly optimized to place semantically similar passages close together in vector space. They produce a single dense embedding per passage and are efficient for semantic search over 512-token chunks. This matches the developer's requirement for a model specifically trained for embedding-based retrieval.

Why this answer

Dense retrieval requires a model that maps each passage to a single vector such that semantically similar passages are close in the embedding space. Bidirectional encoders trained with contrastive objectives are purpose-built for this, producing high-quality passage embeddings efficiently. Generative LLMs, cross-encoders, and summarization models are designed for other tasks and do not provide the required independent passage vectors.

Exam trap

The trap here is assuming any transformer can serve as an embedding model, when in fact only models trained with a contrastive or similar objective produce reliable dense vectors for semantic search.

90
Multi-Selectmedium

You are deploying a large language model on NVIDIA Triton Inference Server in a Kubernetes cluster. To ensure high availability and reliability, which TWO practices should you implement? (Choose two.)

Select 2 answers
A.Enable Triton's model warmup to reduce cold-start latency.
B.Use a single large GPU instance to maximize performance and reduce complexity.
C.Set up health checks and readiness probes for the Triton server pods.
D.Configure multiple model instances per GPU to increase concurrency.
E.Deploy the Triton server as a Kubernetes Deployment with multiple replicas across nodes.
AnswersC, E

Health checks and readiness probes allow Kubernetes to detect unhealthy pods and stop routing traffic to them, and to restart failed pods. This ensures that only healthy replicas serve requests, improving reliability and availability. It is a standard practice for production deployments.

Why this answer

High availability in Kubernetes requires redundancy and health monitoring. Deploying multiple replicas across nodes ensures that a single node failure does not take down the service. Readiness probes ensure traffic is only sent to healthy pods.

Together, these practices minimize downtime and maintain reliable inference serving.

Exam trap

The trap here is focusing on performance optimizations like model warmup or multiple instances, which do not provide redundancy across failures.

91
MCQhard

When training a model for a highly technical domain with a scarcity of high-quality data, which data augmentation strategy is most likely to preserve the model's reliability?

A.Randomly masking 50% of the words in the training documents.
B.Using a smaller, unverified LLM to generate completely new technical scenarios.
C.Generating synthetic examples based on ground-truth technical documentation templates.
D.Translating the dataset into multiple languages using a generic online translator.
AnswerC

Templated generation ensures that the synthetic data adheres to the logical and structural rules of the domain. By basing synthetic examples on verified ground-truth templates, you maximize data variety while maintaining factual accuracy, which is the safest and most effective way to address data scarcity in technical domains.

Why this answer

In technical domains, synthetic data generation must be grounded in existing, verified documents to avoid 'hallucinating' technical facts. Using LLMs to paraphrase or summarize existing high-quality technical content while maintaining strict constraints ensures that the new data follows the same logic and terminology. This method expands the training set while minimizing the risk of introducing incorrect facts, which is essential for high-stakes technical domains.

Exam trap

Candidates often choose unconstrained generative augmentation, which introduces hallucinations. They fail to realize that grounding synthetic data in verified templates is the only way to maintain technical reliability.

92
MCQmedium

You are optimizing a ResNet-50 model on an NVIDIA A100. Which precision-based optimization will yield the highest throughput without significant accuracy loss?

A.Full FP32 precision.
B.Mixed-precision (FP16/BF16).
C.Aggressive FP8 quantization.
D.Custom boolean quantization.
AnswerB

Mixed precision utilizes the high-speed Tensor Cores to perform arithmetic in FP16 or BF16 while maintaining key parts of the model in FP32. This drastically increases throughput and reduces memory bandwidth requirements, allowing for much faster inference with negligible impacts on the overall accuracy of the model, which is ideal.

Why this answer

NVIDIA A100 GPUs feature Tensor Cores that are specifically optimized for FP16 and BF16 arithmetic. Utilizing mixed-precision training or inference allows the model to leverage these high-speed units, effectively doubling throughput compared to FP32. This is the standard approach for balancing the trade-off between floating-point precision and computational speed, providing near-native accuracy while maximizing the hardware's performance capabilities in deep learning workloads.

Exam trap

Candidates often select FP32 for maximum accuracy, failing to recognize that modern NVIDIA architectures provide hardware-accelerated mixed-precision support that maintains high accuracy while significantly improving throughput.

93
MCQmedium

Refer to the exhibit. Given this NeMo configuration, which prompt modification would best improve the reliability of technical support queries?

A.Append 'Be as creative as possible' to the system prompt.
B.Change the stop sequences to be empty.
C.Require the model to state 'I cannot answer this' if the manual is silent.
D.Increase the temperature to 0.9 to ensure varied answers.
AnswerC

This specific instruction provides a clear 'exit path' for the model when the provided context is inadequate. By explicitly defining the behavior for unsupported queries, you prevent the model from guessing or fabricating answers, which is crucial for maintaining the trust and reliability of your technical documentation bot.

Why this answer

The exhibit shows a relatively low temperature, which is good for consistency, but the system prompt lacks specific constraints on how to handle missing data. By modifying the prompt to strictly enforce grounding in the manuals and providing a fallback mechanism, you significantly improve the model's reliability in technical support scenarios. This is a critical step for maintaining quality in production environments utilizing NVIDIA NeMo technology.

Exam trap

Test-takers frequently choose to adjust model temperature or top-p sampling parameters instead of modifying the system prompt to explicitly handle missing data scenarios.

94
MCQmedium

A team is training a large language model and wants to reduce the memory used by the optimizer without changing the model architecture. They are using Adam and notice that optimizer state consumes more GPU memory than the model weights. Which technique should they apply to reduce optimizer memory while keeping the model architecture unchanged?

A.Replace the feed-forward activation with a gated variant to reduce activation memory.
B.Switch to a memory-efficient optimizer such as 8-bit Adam or Adafactor that stores reduced-precision or factored optimizer state.
C.Reduce the number of attention heads while keeping the hidden size constant.
D.Increase the batch size proportionally to the number of GPUs so optimizer states are shared across more tokens.
AnswerB

8-bit Adam quantizes the first and second moment estimates to 8-bit, and Adafactor factors the second moment, both cutting optimizer memory substantially. These keep the model architecture and parameter count intact while lowering the dominant memory cost during training.

Why this answer

Adam maintains two full-precision moment estimates for every trainable parameter, so optimizer state can be several times the size of the model weights. Memory-efficient optimizers such as 8-bit Adam or Adafactor reduce this footprint by quantizing or factoring the state, directly lowering the dominant memory cost without changing the model architecture.

Exam trap

The trap here is assuming that larger batches or architectural tweaks solve optimizer memory, when the moments themselves are the cost.

95
Multi-Selecthard

When profiling an application with NVIDIA Nsight Systems, which TWO metrics are most critical to identify if an application is limited by the PCIe bus?

Select 2 answers
A.Host-to-Device (H2D) throughput.
B.GPU Register usage count.
C.Device-to-Host (D2H) throughput.
D.Shared memory bank conflict count.
E.SM clock frequency.
AnswersA, C

H2D throughput measures the speed at which data is sent from the host CPU to the GPU memory. If this metric hits the theoretical maximum of the PCIe bus, it confirms a bottleneck where the GPU must wait for new data to arrive before processing can begin.

Why this answer

PCIe bus saturation occurs when the transfer rate of data between the host and the GPU becomes a bottleneck. By monitoring the H2D (Host-to-Device) and D2H (Device-to-Host) transfer metrics, developers can see if data movement consumes a disproportionate amount of execution time. Identifying these spikes is crucial because offloading data to the GPU is often the slowest part of a pipeline, and minimizing these transfers is key to scaling high-performance AI.

Exam trap

Candidates often confuse PCIe throughput with GPU compute utilization or memory bandwidth metrics. They fail to realize that PCIe saturation is specifically about the data transfer rate between the host and GPU.

96
MCQmedium

What is the primary advantage of using a 'Quantization Aware Training' (QAT) approach over post-training quantization for LLMs?

A.It eliminates the need for a calibration dataset
B.It produces significantly smaller model files
C.It results in higher accuracy for quantized models
D.It avoids the use of TensorRT builder
AnswerC

By simulating quantization during training, the weights are optimized to minimize the impact of precision loss. This process allows the model to learn and compensate for the rounding errors inherent in low-precision formats, leading to significantly better accuracy compared to post-training quantization methods for complex language models.

Why this answer

QAT incorporates quantization errors into the training loop, allowing the model to adapt its weights to the loss of precision. Unlike post-training quantization, which can cause significant accuracy degradation for complex LLMs, QAT ensures that the model remains robust despite the restricted dynamic range of the INT8 or FP8 format. This results in superior final inference accuracy, making it the preferred choice for high-stakes generative applications where precision is critical.

Exam trap

Candidates often incorrectly identify 'smaller model size' or 'faster training time' as the primary advantage, whereas QAT is specifically designed to mitigate the accuracy loss inherent in quantization.

97
MCQmedium

A financial analyst is using an NVIDIA NIM-hosted Llama 3.1 70B model to extract key financial metrics from quarterly earnings call transcripts. The model inconsistently returns a prose summary instead of the required structured JSON. The analyst needs the output to be reliably parseable by a downstream script that expects a fixed schema with fields "revenue", "eps", and "guidance". Which prompt engineering technique is most appropriate to enforce this output format?

A.Increase the temperature parameter to 0.9 so the model explores more diverse output formats and may eventually produce JSON.
B.Use a chain-of-thought prompt asking the model to 'think step by step' before answering, which will naturally lead to JSON output.
C.Provide a JSON schema and a completed example in the prompt, and instruct the model to return only JSON matching that schema.
D.Add a system prompt that says 'You are a helpful assistant' and rely on the model's instruction-following ability to infer the JSON requirement.
AnswerC

This is correct because providing an explicit schema plus a worked example gives the model a concrete template to imitate, which strongly biases generation toward valid JSON. The instruction to return only JSON reduces the chance of prose leakage. This combination of structured format specification and demonstration is the most reliable way to enforce a fixed output schema without relying on post-processing.

Why this answer

Enforcing a structured output like JSON requires explicit specification of the schema and often a concrete example. Providing a schema and a completed example leverages the model's in-context learning to mimic the exact format, while the instruction to return only JSON reduces extraneous text. Other techniques like temperature adjustment or chain-of-thought do not directly control output structure and may even worsen format compliance.

Exam trap

The trap here is assuming that a generic instruction like 'return JSON' is sufficient without providing a schema or example, when models often need concrete demonstrations to reliably adhere to complex structured formats.

98
MCQhard

A team is evaluating a retrieval-augmented generation (RAG) pipeline using NVIDIA NeMo Evaluation. They notice that the generated answers are fluent but sometimes contradict the retrieved documents. Which evaluation approach best identifies this issue?

A.Compute BLEU score against reference answers.
B.Calculate perplexity of the generated answers.
C.Measure the average length of retrieved documents.
D.Use a natural language inference (NLI) model to check entailment between retrieved documents and generated answers.
AnswerD

NLI models determine whether a hypothesis (generated answer) is entailed by, neutral to, or contradicts a premise (retrieved document). By running an NLI model, you can flag contradictions directly. This approach is robust for RAG evaluation because it assesses factual consistency without requiring reference answers, aligning with the goal of detecting contradictions.

Why this answer

The problem is factual inconsistency between retrieved documents and generated answers. Natural language inference directly evaluates entailment, identifying contradictions. Other metrics like BLEU, perplexity, or document length do not measure this relationship, making NLI the appropriate choice for this RAG evaluation scenario.

Exam trap

The trap here is relying on reference-based metrics like BLEU when the issue is faithfulness to retrieved context, not similarity to a reference answer.

99
MCQeasy

Which metric is most critical to monitor for identifying 'bottlenecks' in a high-throughput LLM deployment?

A.Total number of files in the model repository.
B.GPU Utilization percentage.
C.The name of the backend framework used for training.
D.The number of times the server was restarted.
AnswerB

GPU utilization is a key indicator of hardware efficiency. Low utilization during high-traffic periods indicates that the model is starving for data or that the batching strategy is inefficient. High utilization confirms that the compute resources are being used effectively to process tokens within the inference engine.

Why this answer

GPU utilization is the primary metric for understanding how well the model is saturating the compute resources. If GPU utilization is low while latency is high, it suggests a bottleneck elsewhere, such as CPU preprocessing, data transfer, or synchronization issues. Monitoring this metric allows engineers to determine if they need to increase batch sizes, optimize the pipeline, or scale the infrastructure to maintain performance.

Exam trap

Candidates often select 'latency' or 'throughput' as the primary bottleneck metric, failing to realize that GPU utilization is the fundamental indicator of whether the underlying hardware is actually being saturated.

100
MCQmedium

A media company is using a large language model to generate news article headlines. They want to evaluate the diversity of the generated headlines to ensure they are not overly repetitive. Which metric should they use to quantify the lexical diversity of the generated headlines?

A.Perplexity
B.ROUGE-L
C.BLEU
D.Distinct-n
AnswerD

Distinct-n measures the number of unique n-grams divided by the total number of n-grams in the generated text. It is specifically designed to quantify lexical diversity and repetition. For evaluating headline diversity, Distinct-n provides a direct measure of how varied the word choices are, making it the appropriate metric.

Why this answer

Distinct-n is a metric that calculates the proportion of unique n-grams in the generated text, directly quantifying lexical diversity. It is commonly used to evaluate repetition in generated outputs. BLEU, ROUGE-L, and perplexity assess translation quality, summarization overlap, or fluency, respectively, and do not measure diversity across multiple generations.

Therefore, Distinct-n is the correct choice.

Exam trap

The trap here is confusing fluency metrics like perplexity with diversity metrics, when diversity requires measuring unique n-grams across generations.

101
MCQhard

An evaluation team is comparing two candidate checkpoints of the same fine-tuned model on a 400-prompt open-ended task. They use an LLM-as-judge that returns a pairwise preference for each prompt. Candidate X wins 214 times, candidate Y wins 152 times, and 34 are ties. The judge is the same family as the models being compared. What should the team do before treating candidate X as the winner?

A.Discard the judge and rerun the comparison with ROUGE-L, since automatic overlap metrics are objective and free of style bias.
B.Verify the judge's reliability with a human-labeled subset, control for position bias by swapping answer order, and check that the win margin exceeds judge noise.
C.Accept the result because 214 wins out of 366 non-tie judgments is a clear majority and the sample is large enough to be conclusive.
D.Increase the prompt count to 4,000 and keep the same judging procedure, because a tenfold larger sample will average out any judge bias.
AnswerB

A same-family judge can favor outputs that resemble its own style, and pairwise judges are known to prefer whichever answer appears first. Swapping presentation order removes position bias, while a human-labeled subset quantifies how often the judge agrees with people. Only after those checks does a 62-win margin carry real evidentiary weight.

Why this answer

Pairwise LLM judging is a measurement instrument that needs validation before its verdicts are trusted. Swapping answer order neutralizes position bias, a human-labeled subset estimates judge accuracy, and comparing the win margin to that measured noise tells the team whether the difference is real. Raw majorities, a switch to overlap metrics, and simply enlarging the sample all fail to address systematic judge bias.

Exam trap

The trap here is treating a clear win count from an LLM judge as a verdict, when the judge itself is an unvalidated model that may share stylistic bias with one candidate and prefer the first position.

102
MCQhard

A media-analytics firm serves a 13B-parameter summarization model on two A100 GPUs using NVIDIA TensorRT-LLM behind Triton Inference Server. Traffic is bursty: during live events concurrency triples for about ten minutes, then returns to baseline. Operators report that the first requests after each burst begin are slow and sometimes time out, although steady-state latency is acceptable. Which deployment change most directly addresses the cold-start penalty at the beginning of each burst?

A.Switch the Triton scheduler from dynamic batching to sequence batching so that in-flight conversations are tracked per client.
B.Configure Triton's instance groups to keep a resident model instance and enable the model's warmup configuration so activation buffers and CUDA graphs are exercised before live traffic arrives.
C.Enable TensorRT-LLM's in-flight batching and raise the KV cache fraction so more concurrent sequences can share the cache.
D.Increase the TensorRT-LLM engine's max_batch_size so larger batches can be formed during the burst peaks.
AnswerB

Resident instances plus a warmup configuration cause Triton to load the engine and run representative inference requests at startup, allocating workspace, compiling or replaying CUDA graphs, and paging in weights. When the burst begins, those resources are already hot, so the first real requests no longer pay the initialization penalty. This directly targets the cold-start behavior described.

Why this answer

The reported pattern, acceptable steady-state latency but slow and timing-out requests at the start of each burst, is a classic cold-start symptom. Keeping a resident Triton instance and supplying a warmup configuration forces engine loading, workspace allocation, and CUDA graph capture to happen at server startup rather than on the first live request. This removes the initialization stall precisely when traffic spikes.

Exam trap

The trap here is diagnosing burst-boundary latency as a batching or KV-cache capacity problem when it is actually resource initialization that occurs on first use.

103
Multi-Selectmedium

Which TWO of the following are benefits of using Rotary Positional Embeddings (RoPE) compared to absolute positional embeddings?

Select 2 answers
A.RoPE allows for better extrapolation to unseen sequence lengths.
B.RoPE completely removes the need for attention mechanisms.
C.RoPE improves the model's ability to capture relative token order.
D.RoPE reduces the parameter count of the embedding layer.
E.RoPE requires retraining from scratch if sequence length increases.
AnswersA, C

Because RoPE relies on rotational transformations, the relative distance between tokens is preserved regardless of absolute index. This property makes it easier to extend the context window during inference, as the model can interpret relative relationships between tokens even when they appear at indices beyond the training limit.

Why this answer

RoPE encodes relative position information by rotating vectors in complex space, which allows for better generalization to sequence lengths not seen during training. This relative approach is significantly more effective than absolute embeddings, which assign a fixed position to every index, as it allows the model to interpret the relationships between tokens regardless of their exact position in the sequence.

Exam trap

Candidates mistakenly choose absolute positional embedding benefits, confusing how standard positional encodings behave with RoPE's dynamic complex-space rotation mechanism designed for relative distances.

104
MCQmedium

An enterprise deploys an LLM application using NVIDIA NeMo Guardrails to prevent the generation of toxic content and PII leakage. During red-teaming, testers discover that prompt injection attacks successfully bypass standard input rails by encoding malicious instructions in Base64 format inside conversational context. Which architectural approach provides the most robust mitigation against this evasion technique while maintaining conversational latency requirements?

A.Increase the sensitivity threshold of the self-check input rail to aggressively flag any anomalous conversational patterns
B.Deploy a secondary LLM instance dedicated exclusively to real-time prompt rewriting and normalization before evaluation
C.Implement a custom Python-based input action within NeMo Guardrails to detect and decode Base64 patterns prior to safety model evaluation
D.Rely on NeMo Guardrails built-in output rails to catch toxic generations regardless of whether the input injection was successfully detected
AnswerC

Adding a programmatic pre-processing action directly intercepts incoming payloads, identifies encoding schemes, and decodes strings into readable text so that downstream topical and safety rails can accurately evaluate the actual semantic intent of the user prompt.

Why this answer

Decoding and validating incoming payloads before rail evaluation ensures that obfuscated prompt injections are neutralized. NeMo Guardrails supports custom input actions that can preprocess text payloads before hitting core moderation models, catching encoded bypass attempts early without adding excessive inference overhead to the main generation loop.

Exam trap

Candidates often assume that standard NeMo Guardrails topical rails automatically parse and decode Base64 strings, missing the fact that custom pre-processing actions must be explicitly defined to handle encoded text payloads.

105
MCQhard

A government agency deploys an NVIDIA NIM-based assistant to help citizens understand benefit eligibility. An oversight panel demands that any refusal to answer be explainable and consistent, and that the assistant not improvise policy. Which design best meets these requirements?

A.Allow the model to answer freely but add a disclaimer that responses are not official policy determinations.
B.Fine-tune the model on past benefit determinations so its answers reflect historical agency decisions.
C.Route eligibility questions through a retrieval action over the official policy corpus, and use guardrails to refuse when no authoritative passage is retrieved.
D.Lower the model's temperature to zero and rely on deterministic decoding to guarantee policy accuracy.
AnswerC

Grounding answers in an authoritative corpus and refusing when retrieval returns nothing prevents the model from improvising policy. The refusal is triggered by the absence of a source, which is a consistent, explainable rule that the oversight panel can inspect and audit. This directly satisfies both transparency and consistency demands.

Why this answer

Explainable, consistent refusals come from a rule that is tied to an external authority. Retrieval over the official policy corpus ensures answers are grounded, and a guardrail that refuses when no authoritative passage is found makes the refusal rule transparent and auditable. Decoding settings and disclaimers cannot substitute for grounding and rule-based refusal.

Exam trap

The trap here is believing that low temperature or a disclaimer yields policy accuracy, when neither grounds the model in authoritative sources.

106
MCQeasy

A technical support team is building a chatbot using an NVIDIA NIM microservice. The chatbot must answer questions about a specific product's warranty policy. The team wants to ensure the model's responses are grounded in the official warranty document, which is 50 pages long, and avoid inventing policy details. Which prompt engineering approach is most effective for this scenario?

A.Ask the model to 'think step by step' and reason about the warranty policy from its pre-trained knowledge.
B.Use a retrieval-augmented generation (RAG) pipeline to fetch the most relevant warranty sections and include them in the prompt as context.
C.Fine-tune the model on the warranty document using NVIDIA NeMo, then use the fine-tuned model without any additional context in the prompt.
D.Include the entire warranty document in the system prompt and instruct the model to answer based only on that document.
AnswerB

RAG retrieves only the pertinent sections of the warranty document and places them in the prompt, providing focused context that the model can use to generate accurate answers. This reduces hallucination because the model is conditioned on specific, relevant text. It also scales to large documents and keeps the prompt within token limits. This is the standard approach for grounding LLM responses in proprietary knowledge bases.

Why this answer

Retrieval-augmented generation (RAG) is the most effective way to ground responses in a specific document. It retrieves relevant sections and includes them in the prompt, ensuring the model's answer is based on factual content rather than pre-trained memory. This reduces hallucinations and handles documents that exceed the context window.

Other methods either risk hallucination, are inefficient, or lack flexibility.

Exam trap

The trap here is assuming that fine-tuning is necessary to teach the model a document, when RAG with prompt context is often simpler, more accurate, and easier to update.

107
MCQeasy

A developer is writing a system prompt for an NVIDIA NIM-hosted assistant that must always respond in formal English, never use slang, and never reveal internal system instructions. Where should these persistent behavioral rules be placed for the MOST consistent effect?

A.Omitted entirely, because well-aligned base models already default to formal English and never disclose their instructions.
B.Encoded only in the model's temperature and top_p settings, since decoding parameters control response style.
C.Appended to the end of each user message as a reminder, so the rules stay close to the model's most recent input.
D.In the system prompt, stated as explicit rules that apply to every turn of the conversation.
AnswerD

System prompts establish persistent, high-priority behavioral guidance that applies across all turns. Placing tone, style, and confidentiality rules there gives the model a stable frame of reference, making consistent adherence far more likely than rules scattered in individual user messages.

Why this answer

Persistent behavioral requirements such as tone, style, and confidentiality belong in the system prompt, which the model treats as standing guidance across the conversation. This placement keeps rules stable regardless of what users type and avoids the fragility of repeating them per message or expecting decoding parameters or base alignment to handle policy.

Exam trap

The trap here is assuming that a well-aligned model needs no explicit style or confidentiality instructions, when consistent behavior still depends on system-level guidance.

108
MCQmedium

Why is it important to use a 'warm-up' period in the learning rate schedule when starting a fine-tuning job?

A.To increase the total training time for better hardware utilization.
B.To allow the optimizer to adapt to the new data distribution gradually.
C.To force the model to explore a wider range of the loss landscape immediately.
D.To bypass the need for gradient clipping during training.
AnswerB

A gradual warm-up phase prevents abrupt weight changes that could lead to divergence. By starting with a low learning rate, the optimizer stabilizes its states based on the new data distribution, which is critical for maintaining the model's integrity and achieving steady, high-quality convergence during the fine-tuning process.

Why this answer

A warm-up period gradually increases the learning rate from a near-zero value to the target rate. This prevents early, massive gradient updates from destabilizing the pre-trained weights. By slowly introducing the learning rate, the model remains stable during the initial phase of training, ensuring that the optimizer can effectively adapt to the new domain without destroying the fundamental knowledge captured during the original pre-training process.

Exam trap

Candidates incorrectly believe warm-up is to save energy or reduce latency, missing that its primary purpose is preventing gradient spikes from destabilizing pre-trained weights during the initial training phase.

109
MCQmedium

An LLM application deployed on NVIDIA Triton Inference Server is experiencing intermittent latency spikes. Which metric is most critical to monitor to determine if the issue stems from GPU compute saturation?

A.Total system CPU usage
B.Network interface throughput
C.SM (Streaming Multiprocessor) occupancy
D.Total GPU memory capacity
AnswerC

SM occupancy directly measures the percentage of active warps compared to the maximum supported by the GPU. High occupancy indicates that the streaming multiprocessors are busy executing instructions, serving as the primary metric for identifying compute-bound bottlenecks in neural network inference tasks.

Why this answer

Monitoring GPU utilization alone is insufficient because it does not distinguish between compute saturation and memory bandwidth bottlenecks. NVML-based metrics specifically tracking SM (Streaming Multiprocessor) occupancy provide the granular insight required to identify whether the execution cores are fully utilized. This metric is essential for capacity planning and ensuring that the inference throughput meets strict service level agreements under high concurrent request loads in production environments.

Exam trap

Candidates often monitor general host CPU utilization or basic GPU memory usage, failing to recognize that SM occupancy directly reflects compute execution core saturation.

110
MCQeasy

A retail company is deploying an LLM-based customer service assistant using NVIDIA NIM. The legal team mandates that the model must not generate content that violates copyright, such as reproducing song lyrics or book excerpts. Which NVIDIA offering should the team use to enforce this policy at runtime?

A.NVIDIA Riva with a content moderation model
B.NVIDIA NeMo Guardrails with a custom output rail for copyright detection
C.NVIDIA Triton Inference Server with a custom backend
D.NVIDIA TensorRT-LLM with a copyright filter plugin
AnswerB

NeMo Guardrails allows the creation of output rails that can run checks on the LLM's generated text. A custom rail can invoke a copyright detection model or service to identify and block responses containing protected material. This enforces the policy at runtime without altering the base model.

Why this answer

NeMo Guardrails is the appropriate tool because it is specifically designed to enforce policies on LLM inputs and outputs. An output rail can be configured to call a copyright detection model, blocking responses that infringe. The other options are infrastructure or speech components that do not offer runtime content policy enforcement for text generation.

Exam trap

The trap here is thinking that any NVIDIA component can be repurposed for content filtering, when only NeMo Guardrails provides the programmable policy layer for LLM outputs.

111
MCQeasy

Which NVIDIA library is primarily used for optimizing and deploying deep learning inference models?

A.cuBLAS.
B.TensorRT.
C.NCCL.
D.cuDNN.
AnswerB

TensorRT is the specialized SDK designed for high-performance deep learning inference. It provides essential features such as layer fusion, kernel auto-tuning, and precision reduction (INT8/FP8), which are critical for deploying models in production environments where low latency and high throughput are the primary performance and cost objectives.

Why this answer

TensorRT is NVIDIA's dedicated deep learning inference optimizer and runtime engine. It takes pre-trained models from frameworks like PyTorch or TensorFlow and optimizes them for NVIDIA hardware by fusing layers, selecting best-fit kernels, and performing precision calibration. It is the core tool for any developer needing to transition models from research training to production-level deployment with minimal latency and maximal throughput on NVIDIA infrastructure.

Exam trap

Candidates often select training frameworks like PyTorch or TensorFlow, forgetting that while those are used for development, TensorRT is the specific library designed for production inference optimization.

112
MCQhard

Refer to the exhibit. The Triton Inference Server is failing to load a model. The configuration shows two instances assigned to GPU 0. What is the most likely cause and the correct remediation?

A.The model count exceeds the number of available CUDA cores, requiring a driver update.
B.The total VRAM required for two instances exceeds the GPU capacity; reduce instance count to 1.
C.The model weights are corrupted, preventing the inference engine from initializing the memory space.
D.The batch size is set to zero in the configuration, preventing memory allocation.
AnswerB

Each model instance requires a dedicated memory buffer for weights and activation tensors. When configured with 'count: 2', Triton attempts to load the model twice. If the sum exceeds the VRAM, an OOM occurs. Reducing the instance count is the most direct way to resolve the startup conflict.

Why this answer

The error indicates that the two instances of the model are collectively requesting more VRAM than is available on the physical GPU. By default, Triton attempts to allocate memory for every configured instance upon startup. The remediation requires either reducing the number of instances or implementing a memory-aware model partitioning strategy, such as using Model Analyzer to determine the safe memory footprint per instance before deployment.

Exam trap

Candidates frequently assume the error is due to a software version mismatch or driver issue, failing to calculate the cumulative VRAM consumption of multiple model instances relative to the hardware limit.

113
MCQhard

An LLM inference service running on NVIDIA Triton Inference Server is configured with model ensembles. During production monitoring, the operations team notices that the end-to-end latency reported by the client is significantly higher than the sum of individual model latencies reported by Triton's metrics. Which Triton feature should be investigated to identify the source of the additional latency?

A.Model ensemble scheduler overhead and inter-model data transfer
B.GPU memory utilization causing queuing of inference requests
C.Network latency between the client and the Triton server
D.Dynamic batching configuration causing delays in request aggregation
AnswerA

Triton's ensemble scheduler manages the execution of multiple models and handles data transfer between them. This can introduce overhead not captured in individual model latency metrics. Investigating ensemble scheduling and data transfer times can reveal bottlenecks. The discrepancy between client-side and server-side latency suggests overhead in the ensemble pipeline, such as serialization or scheduling delays.

Why this answer

In Triton model ensembles, the ensemble scheduler coordinates execution and data transfer between models. This adds overhead that is not attributed to any single model's latency. Monitoring ensemble-specific metrics and tracing the pipeline can identify where the extra time is spent, such as in data serialization or scheduling delays.

Exam trap

The trap here is assuming that client-side latency should always match server-side model latency, overlooking the overhead introduced by ensemble scheduling and inter-model communication.

114
MCQmedium

Which THREE of the following are primary benefits of using PagedAttention in NVIDIA TensorRT-LLM deployments?

A.Elimination of external memory fragmentation.
B.Reduction in KV cache memory overhead.
C.Support for longer context lengths.
D.Faster calculation of attention scores.
E.Direct hardware support in the GPU scheduler.
AnswerA, B, C

PagedAttention treats the KV cache as a collection of fixed-size blocks, similar to virtual memory in operating systems. This structure prevents external fragmentation because any free block can be allocated to any request, ensuring that memory usage remains highly efficient even when servicing requests of varying lengths and concurrent patterns.

Why this answer

PagedAttention is a critical optimization for LLM inference, solving the problem of memory fragmentation. By managing memory in non-contiguous pages, it allows the system to allocate only what is needed, reducing memory waste significantly. This enables higher batch sizes, better GPU utilization, and the ability to serve longer sequences without running out of memory, which are essential for maintaining high performance in production-grade LLM serving environments.

Exam trap

Candidates often mistake PagedAttention as a compute acceleration technique rather than a memory management optimization, leading them to select incorrect benefits like reduced floating-point operations.

115
MCQhard

A developer is using an NVIDIA NIM for a Llama 3.1 70B model to build a legal document review assistant. The model must answer questions based on a provided contract, but the contracts are often 50,000 tokens long, exceeding the model's 8,000-token context window. Which prompt engineering strategy is most appropriate to handle this constraint?

A.Use a retrieval-augmented generation (RAG) approach to fetch only the relevant sections of the contract and include them in the prompt.
B.Split the contract into chunks and ask the model to summarize each chunk sequentially, then combine the summaries.
C.Increase the model's context window by fine-tuning it on longer sequences.
D.Use a sliding window approach where the model processes the contract in overlapping segments and maintains a memory of previous segments.
AnswerA

RAG involves retrieving the most relevant passages from the long document and including only those in the prompt. This reduces the context length to fit within the model's window while still providing the necessary information. It is the standard solution for handling documents that exceed the context limit.

Why this answer

RAG is the most appropriate strategy because it retrieves only the relevant sections of the long contract, fitting them into the model's context window. This allows the model to answer questions accurately without needing to process the entire document. It is a widely adopted prompt engineering pattern for handling context length limitations.

Exam trap

The trap here is assuming that fine-tuning or summarization can easily overcome context window limits, when in fact retrieval-based approaches are the practical solution.

116
Multi-Selectmedium

An engineer is designing prompts for an NVIDIA NIM-hosted model that must extract structured fields from unstructured invoices. The extraction accuracy is inconsistent across vendors with different layouts. Which TWO prompt engineering techniques would MOST improve reliability? (Choose two.)

Select 2 answers
A.Instruct the model to output fields in a fixed JSON schema and to use null for any field not present in the document.
B.Ask the model to summarize the invoice in prose first and then extract fields from its own summary in a second pass.
C.Raise the temperature to 0.9 so the model explores multiple interpretations of ambiguous invoice text before committing to values.
D.Provide few-shot examples that cover several distinct invoice layouts, each showing the exact input-to-output field mapping.
E.Remove all formatting and concatenate the entire invoice text into a single lowercase string before sending it to the model.
AnswersA, D

A fixed schema removes ambiguity about output shape and key names, and the null convention gives the model a defined behavior for missing data instead of inventing values. Together these constraints make outputs parseable and reduce hallucinated field values, which is essential for downstream automation.

Why this answer

Reliable structured extraction combines format constraints with demonstrated patterns. A fixed JSON schema with a null convention makes outputs machine-parseable and defines behavior for absent fields, while few-shot examples across varied layouts teach the model to handle layout diversity. Together they reduce both structural errors and hallucinated values, which prose summarization and high-temperature sampling would worsen.

Exam trap

The trap here is treating higher temperature as a way to reason through ambiguity, when in extraction tasks it mainly adds run-to-run inconsistency and invented values.

117
MCQmedium

What is the primary role of the 'Learning Rate Scheduler' during LLM fine-tuning?

A.To increase the batch size dynamically
B.To control the weight updates over the training duration
C.To automatically detect the optimal training loss
D.To determine which layers should be frozen
AnswerB

The scheduler dictates how the learning rate changes over time, allowing the model to make large updates initially and smaller, more precise updates as training progresses. This helps the model converge reliably to a stable state, preventing the training process from oscillating or diverging due to a static, high learning rate.

Why this answer

A scheduler adjusts the learning rate throughout the training process, typically starting with a warmup phase to stabilize the weights and then decaying the rate to allow for fine-grained convergence. This prevents the model from overshooting optimal solutions early on and helps it settle into a good local minimum, which is critical for achieving high-quality results without divergence.

Exam trap

Candidates mistakenly believe learning rate schedulers control batch size or regularization penalties, forgetting their direct responsibility for dynamically modifying weight updates over time.

118
Multi-Selectmedium

Which TWO of the following telemetry types are essential for detecting 'model drift' in a production LLM deployment?

Select 2 answers
A.Output token distribution statistics
B.GPU power consumption levels
C.User feedback and semantic similarity scores
D.Hardware temperature monitoring
E.Network latency between nodes
AnswersA, C

Monitoring shifts in the distribution of generated tokens helps identify if the model is defaulting to repetitive or nonsensical outputs. Significant deviations from established baseline token patterns are often a leading indicator that the input distribution has shifted, triggering a drift condition.

Why this answer

Model drift refers to the degradation of model output quality over time as real-world data deviates from training distributions. Tracking output tokens and response semantic consistency allows engineers to identify when the model starts producing incoherent or irrelevant results. Proactive monitoring of these metrics is critical because drift often occurs silently, potentially damaging user trust before traditional error logs or system performance metrics indicate a failure.

Exam trap

Candidates often select 'latency' or 'throughput' as drift metrics. While these indicate system health, they do not measure the actual semantic quality or output distribution drift of the LLM's generated content.

119
MCQmedium

A healthcare company deploys an LLM-powered clinical documentation assistant using NVIDIA NIM microservices on-premises. During a compliance review, auditors discover that the model occasionally generates patient names and medical record numbers in its output even though these were not present in the input prompt. The team needs to implement a runtime guardrail that detects and blocks such unintended PII leakage without retraining the model. Which NVIDIA component should they configure to add this output-side detection?

A.NVIDIA NeMo Guardrails with an output rail that invokes a PII detection model
B.NVIDIA Riva with automatic speech recognition enabled
C.NVIDIA TensorRT-LLM with quantization-aware training
D.NVIDIA Triton Inference Server with dynamic batching enabled
AnswerA

NeMo Guardrails supports output rails that run after the LLM generates a response, allowing a PII detection model to scan and block or redact sensitive data before it reaches the user. This directly addresses unintended PII leakage without retraining, as the guardrail operates at inference time and can be configured with custom detection logic.

Why this answer

Output rails in NeMo Guardrails are designed to intercept the model's generated text and apply safety checks, such as PII detection, before the response is returned. This is the correct mechanism because it operates at runtime without modifying the underlying model. Triton, TensorRT-LLM, and Riva serve different purposes—serving, optimization, and speech—and lack content-filtering capabilities for PII.

Exam trap

The trap here is confusing inference optimization or serving tools with runtime content moderation, assuming that any NVIDIA component in the pipeline can enforce PII policies.

120
MCQmedium

You are using NVIDIA NeMo Evaluation to assess a summarization model. The model produces summaries that are grammatically correct but omit key information from the source. Which metric should you use to quantify the amount of missing content?

A.Perplexity
B.BLEU score
C.ROUGE-1 recall
D.ROUGE-1 precision
AnswerC

ROUGE-1 recall measures the proportion of unigram overlaps between the generated summary and the reference summary, relative to the reference. Low recall indicates missing content. Since the issue is omission of key information, recall is the appropriate metric to quantify missing content, as it directly reflects how much of the reference is captured.

Why this answer

The problem is omission of key information from summaries. ROUGE-1 recall directly measures how much of the reference unigrams are present in the generated summary. Low recall indicates missing content.

Precision, BLEU, and perplexity focus on relevance, precision, or fluency, not on capturing all reference content, making recall the correct choice.

Exam trap

The trap here is confusing precision and recall: precision penalizes extraneous content, while recall penalizes missing content. The scenario specifically asks about omissions.

121
MCQmedium

A developer is building a customer support assistant using an NVIDIA NIM microservice for a Llama 3.1 8B Instruct model. The assistant must answer questions about an order solely based on a JSON payload containing order details, and it must not use any outside knowledge. Which prompt engineering approach best ensures the model adheres to this constraint?

A.Prefix the prompt with a system message that says 'You are a helpful assistant' and then ask the question without including the JSON.
B.Include the JSON payload in the prompt and instruct the model to answer using only the provided data, adding a fallback phrase for missing information.
C.Fine-tune the model on a dataset of order-related questions and answers before deploying it as a NIM microservice.
D.Use a low temperature setting (e.g., 0.1) to make the model more deterministic and less likely to hallucinate.
AnswerB

This approach explicitly grounds the model in the provided JSON and sets a clear boundary: answer only from the data. Adding a fallback phrase for missing information prevents the model from inventing details, which is critical for factual accuracy in customer support. It leverages the instruction-following capability of Llama 3.1 without requiring additional guardrail systems.

Why this answer

Grounding the model with the exact JSON payload and instructing it to answer only from that data, with a fallback for missing information, ensures responses are based solely on the provided order details. This leverages the instruction-following ability of Llama 3.1 and avoids reliance on external knowledge. Other methods like fine-tuning or temperature adjustment do not enforce strict data adherence at inference time.

Exam trap

The trap here is assuming that a low temperature setting alone can prevent hallucination and enforce data grounding, when in fact explicit instruction and context inclusion are required.

122
MCQmedium

A team is pretraining a 13B-parameter decoder-only LLM on English text using byte-pair encoding with a 50,000-token vocabulary. They observe that the model produces fluent but repetitive continuations and that the average log-probability assigned to ground-truth tokens plateaus early. The training loss curve shows the model is underfitting rather than overfitting. Which architectural change is most likely to improve the model's capacity to capture long-range dependencies?

A.Reduce the vocabulary size to 32,000 tokens to lower the softmax computation cost.
B.Increase the number of attention heads while keeping the hidden dimension fixed.
C.Increase the number of transformer layers and the hidden dimension proportionally.
D.Replace the learned positional embeddings with sinusoidal absolute positional encodings.
AnswerC

Underfitting at 13B parameters on English text signals insufficient model capacity. Scaling depth (more layers) and width (larger hidden dimension) together increases both the number of sequential transformation steps and the representational bandwidth, allowing the network to model longer-range dependencies and richer token distributions. This directly addresses the plateau in log-probability on ground-truth tokens.

Why this answer

Underfitting in a large decoder-only LLM indicates the model lacks sufficient capacity for the data distribution. Scaling both depth and width increases the parameter count and the number of nonlinear transformations applied to the residual stream, which improves the model's ability to capture long-range dependencies. The other options either rebalance existing capacity, change positional encoding without adding parameters, or reduce computation without improving fit.

Exam trap

The trap here is assuming that any architectural tweak that increases parameter count (such as adding attention heads) will resolve underfitting, when in fact head rebalancing at a fixed hidden dimension does not add representational capacity.

123
Multi-Selecthard

Which THREE actions are essential for maintaining a secure and compliant LLM deployment according to the NVIDIA security guidelines?

Select 3 answers
A.Implement strict role-based access control (RBAC) for all API endpoints.
B.Disable all logging of user prompts to maximize data privacy.
C.Apply robust input sanitization to prevent prompt injection attacks.
D.Run model containers as root to ensure full hardware access permissions.
E.Perform regular security scanning and vulnerability assessment of the container images.
AnswersA, C, E

RBAC is a fundamental security requirement that limits exposure to unauthorized users. By ensuring that only authenticated and authorized services can invoke the LLM, you reduce the attack surface and prevent malicious actors from abusing the model's capabilities to generate prohibited content or access restricted data.

Why this answer

Secure LLM deployment requires a defense-in-depth approach. Implementing role-based access control (RBAC) ensures only authorized users interact with models, while input sanitization prevents injection attacks that could lead to data exfiltration. Finally, regular vulnerability scanning of the containerized model environment identifies weaknesses before they can be exploited.

These measures are critical for protecting the model's integrity and ensuring that the AI system does not become a vector for malicious activities.

Exam trap

Candidates often suggest 'model watermarking' as a primary security guideline. While useful for provenance, it is not a core security measure compared to RBAC, input sanitization, and vulnerability scanning.

124
MCQhard

Refer to the exhibit. The inference server reports an OOM error while running multiple LLMs. What configuration change most effectively improves reliability without upgrading the hardware?

A.Increase the batch size
B.Implement KV cache block paging
C.Enable FP32 precision mode
D.Increase the number of threads
AnswerB

KV cache paging divides the cache into smaller, manageable blocks, significantly reducing external fragmentation. This allows the system to utilize memory more effectively, enabling higher concurrency without requiring a physical hardware upgrade, which is the standard strategy for resolving VRAM saturation issues.

Why this answer

The exhibit shows that the VRAM is almost completely exhausted, likely due to excessive KV cache allocation for concurrent models. Implementing PagedAttention or optimizing the KV cache block size allows for more efficient memory management. This approach improves reliability by preventing memory fragmentation and allowing the system to handle higher request concurrency within existing hardware constraints, which is vital for production systems requiring high uptime and predictable performance.

Exam trap

Candidates often suggest reducing the batch size or model precision as the primary fix. While these help, KV cache paging is the specific architectural solution for memory fragmentation in high-concurrency LLM serving.

125
MCQmedium

When fine-tuning a base LLM using Parameter-Efficient Fine-Tuning (PEFT) on an NVIDIA H100 GPU, what is the primary advantage of utilizing LoRA compared to full fine-tuning?

A.It increases the number of trainable parameters to improve model convergence speed.
B.It eliminates the need for any gradient calculation during the backpropagation process.
C.It reduces the memory footprint by freezing base weights and training only low-rank adapter matrices.
D.It forces the model to ignore long-range dependencies to accelerate training time.
AnswerC

Freezing the primary model parameters eliminates the need to store optimizer states for the vast majority of the weights. By training only small, low-rank matrices, the total VRAM consumption is drastically reduced, allowing for larger batch sizes and faster training cycles on NVIDIA GPU hardware platforms.

Why this answer

LoRA reduces memory overhead by freezing pre-trained model weights and injecting trainable rank-decomposition matrices into transformer layers. This approach significantly lowers the VRAM requirements, enabling fine-tuning on consumer or enterprise hardware without needing to load the entire parameter set into the optimizer state. This method is critical for deploying domain-specific models efficiently while maintaining performance parity with full fine-tuning, thus optimizing resource utilization across high-performance NVIDIA compute clusters.

Exam trap

Candidates often assume LoRA reduces inference latency or speeds up training time, when its primary benefit is lowering memory usage by freezing base weights.

126
MCQmedium

An ML engineer is fine-tuning a 7B-parameter model with LoRA on a single NVIDIA A100 40GB GPU. The training script reports that the adapter weights are not updating after several hundred steps, and the loss remains flat. The base model weights are frozen as intended. Which LoRA configuration issue is the most likely cause?

A.The base model is loaded in 8-bit precision, which prevents any gradient computation on the adapters.
B.The learning rate is too low, so the adapter weights change by amounts below floating-point precision.
C.The LoRA rank is set too high, causing the adapters to be initialized to zero.
D.The target modules list is empty or points to modules that are not part of the model's forward pass.
AnswerD

LoRA only updates the adapter matrices attached to the specified target modules. If the target modules list is empty or names modules that never execute, no adapter parameters participate in the forward pass, so no gradients reach them and the adapters remain unchanged. This exactly matches the symptom of frozen adapters and flat loss.

Why this answer

If the LoRA target modules list is empty or references modules that are not executed during the forward pass, the adapter parameters receive no gradients and never update. The base model remains frozen as intended, but the adapters are effectively absent from training, which explains both the flat loss and the unchanged adapter weights.

Exam trap

The trap here is assuming that a frozen base model or a low learning rate explains non-updating adapters, when the more likely cause is that the adapters are not attached to any executed module.

127
Multi-Selecthard

A team is deploying a 70B parameter LLM with NVIDIA Triton Inference Server across four NVIDIA H100 GPUs. They are using TensorRT-LLM and need to fit the model within the combined GPU memory while maintaining high throughput. Which two techniques should they use? (Choose two.)

Select 2 answers
A.Pipeline parallelism, which assigns entire transformer layers to different GPUs and passes activations between stages.
B.Increasing the maximum batch size to the largest value the backend accepts, so memory is fully utilized.
C.Data parallelism, which replicates the full model on every GPU and splits incoming requests across replicas.
D.Enabling FP32 precision for all weights and activations to maximize numerical stability across GPUs.
E.Tensor parallelism, which shards model layers and attention heads across the GPUs so each GPU holds a fraction of the weights.
AnswersA, E

Pipeline parallelism assigns groups of layers to different GPUs, so each GPU stores only a portion of the model. Combined with tensor parallelism, it allows very large models such as 70B parameters to fit across four H100 GPUs. It is commonly used together with tensor parallelism to balance memory and communication overhead in multi-GPU TensorRT-LLM deployments.

Why this answer

Tensor parallelism and pipeline parallelism are the two model-parallel techniques that split a 70B model across multiple GPUs, allowing the weights to fit in aggregate memory on four H100s. Tensor parallelism shards layers and attention heads, while pipeline parallelism assigns layer groups to stages; together they enable large-model deployment with high throughput.

Exam trap

The trap here is confusing data parallelism, which replicates the full model per GPU, with model parallelism, which actually splits weights to fit a large model across devices.

128
MCQeasy

Which technique is most effective for mitigating data leakage during the training of an LLM on time-series-related document data?

A.Random shuffling of all documents regardless of their timestamps.
B.Chronological splitting of the dataset based on a cutoff date.
C.Oversampling the minority class in the training set to improve balance.
D.Applying aggressive data normalization to all numerical values.
AnswerB

Chronological splitting ensures that the model is trained exclusively on data from the past, while validation and testing are performed on subsequent time periods. This mirrors the real-world deployment environment, where the model must predict future events based only on information that has already occurred in history.

Why this answer

Preventing data leakage in temporal datasets requires strict chronological separation between training, validation, and testing sets. If future information is inadvertently included in the training data, the model will 'cheat' by observing future events during training, leading to artificially inflated performance metrics that do not generalize to actual production deployment scenarios where future data is unavailable.

Exam trap

Candidates often rely on random shuffling or standard cross-validation, failing to recognize that time-series text data requires strict chronological splitting to prevent future information leakage.

129
MCQmedium

Refer to the exhibit. The model performance is inconsistent. What is the most likely reason for the performance variability under load?

A.The instance count is too high for the GPU.
B.Dynamic batching lacks a max_queue_delay_microseconds parameter.
C.The GPU index is incorrectly defined in the config.
D.TensorRT plan files are inherently non-deterministic.
AnswerB

Without a max_queue_delay_microseconds, the inference server does not wait for additional requests to fill the batch. This results in requests being processed immediately, often with small batch sizes, leading to inconsistent execution times and poor GPU utilization compared to waiting briefly to aggregate requests into larger, more efficient batches.

Why this answer

The configuration uses dynamic batching with a preferred batch size range but lacks a 'max_queue_delay_microseconds' setting. Without a delay buffer, the server might dispatch batches as soon as one request arrives, leading to suboptimal under-filled batches and inconsistent latency. Adding a small delay allows the server to collect more requests, increasing throughput and stabilizing inference latency across fluctuating traffic patterns, which is vital for high-performance production deployments.

Exam trap

Candidates frequently assume the issue is related to GPU compute capacity or model size, overlooking the configuration-level settings that govern how requests are queued and processed in a production server.

130
MCQmedium

A team is deploying a Llama 2 13B model with NVIDIA TensorRT-LLM on a single A100 40GB GPU. They need to serve 32 concurrent requests with a maximum sequence length of 4096 tokens. They observe that the GPU runs out of memory during inference. Which configuration parameter should they adjust to control the maximum GPU memory allocated for the KV cache?

A.max_batch_size
B.max_num_tokens
C.kv_cache_free_gpu_memory_fraction
D.max_input_len
AnswerC

kv_cache_free_gpu_memory_fraction specifies the fraction of free GPU memory that TensorRT-LLM can use for the KV cache. Lowering this value reduces the KV cache size, preventing out-of-memory errors when serving long sequences with multiple concurrent requests. This parameter directly controls the memory pool dedicated to the KV cache, making it the correct adjustment.

Why this answer

The KV cache in TensorRT-LLM is allocated from a memory pool whose size is controlled by kv_cache_free_gpu_memory_fraction. Lowering this fraction reduces the KV cache footprint, resolving out-of-memory errors when serving multiple long sequences. Other parameters like max_batch_size or max_num_tokens influence scheduling but do not directly bound the KV cache memory pool.

Exam trap

The trap here is assuming that max_batch_size directly limits KV cache memory, when in fact the KV cache pool is separately governed by kv_cache_free_gpu_memory_fraction.

131
MCQhard

An engineer fine-tunes a model on a domain corpus with NVIDIA NeMo and observes that training loss falls steadily while validation loss begins rising after the second epoch. The team must produce the most generalizable checkpoint without changing the dataset. Which action should they take?

A.Increase the number of training epochs so the model fully converges on the domain corpus
B.Disable validation during training and rely on the final training loss to select the checkpoint
C.Enable early stopping based on validation loss and restore the best checkpoint
D.Raise the learning rate to escape the local minimum the model has settled into
AnswerC

The described divergence between falling training loss and rising validation loss is the classic overfitting signal. Monitoring validation loss and stopping when it stops improving, then restoring the checkpoint with the lowest validation loss, yields the most generalizable model. NeMo supports validation-interval evaluation and checkpoint selection based on a monitored metric, so this is the direct remedy.

Why this answer

When training loss keeps falling but validation loss turns upward, the model is overfitting and the best generalizing weights occur near the point where validation loss is lowest. Early stopping with best-checkpoint restoration captures that point. More epochs, a higher learning rate, or dropping validation all fail to address the divergence or actively make it worse.

Exam trap

The trap here is treating a falling training loss as evidence of healthy progress when the rising validation loss is the decisive signal.

132
MCQmedium

A healthcare technology company is deploying an LLM-powered patient triage assistant using NVIDIA NIM microservices on-premises. During an internal audit, the compliance team discovers that the model occasionally outputs patient names and medical record numbers (MRNs) in its responses, even though the training data was scrubbed. The company must implement a runtime safeguard that detects and redacts PII before the response reaches the user. Which NVIDIA component should they integrate into their inference pipeline to achieve this?

A.NVIDIA Triton Inference Server with dynamic batching enabled
B.NVIDIA TensorRT-LLM with quantization-aware training
C.NVIDIA Riva with custom ASR and TTS models
D.NVIDIA NeMo Guardrails with a custom output rail that invokes a PII detection model
AnswerD

NeMo Guardrails allows defining output rails that intercept the LLM response and apply custom actions, such as calling a PII detection model to redact sensitive entities like names and MRNs. This runtime safeguard operates after generation but before user delivery, exactly matching the requirement to detect and redact PII on the fly without retraining.

Why this answer

The requirement is a runtime safeguard that detects and redacts PII in LLM outputs before they reach users. NeMo Guardrails provides a programmable framework for defining output rails that can invoke custom PII detection models and modify responses accordingly. Other NVIDIA components like Triton, TensorRT-LLM, and Riva focus on inference optimization or speech processing, not content moderation, so they cannot enforce the needed redaction policy.

Exam trap

The trap here is assuming that inference optimization tools like TensorRT-LLM or Triton automatically include content safety features, when they do not.

133
MCQmedium

Which hardware architecture feature is specifically leveraged by TensorRT to accelerate FP16 and INT8 matrix multiplications?

A.CPU AVX-512 vector instructions
B.NVIDIA Tensor Cores
C.Shared memory buffers in L1 cache
D.Global memory coalescing hardware
AnswerB

Tensor Cores are hardware circuits designed for high-speed matrix multiplications in FP16, INT8, and other low-precision formats. TensorRT optimizes the execution graph to ensure that large matrix multiplications are dispatched to these units, providing the massive performance gains seen in modern deep learning inference workloads.

Why this answer

NVIDIA Tensor Cores are specialized hardware units designed to perform mixed-precision matrix multiply-accumulate operations in a single cycle. TensorRT automatically detects the presence of these cores and maps compute-intensive layers to them. By utilizing Tensor Cores, the model achieves significantly higher throughput and reduced latency compared to using standard CUDA cores, which perform operations at a lower efficiency per clock cycle for matrix math.

Exam trap

Candidates often confuse Tensor Cores with CUDA cores, assuming that standard CUDA cores are the primary driver for mixed-precision acceleration rather than the specialized hardware units built for matrix math.

134
Multi-Selecthard

An ML engineer is evaluating Mixture-of-Experts (MoE) routing for a large decoder-only model to increase capacity without proportionally increasing compute per token. Which TWO statements accurately describe how top-k token routing behaves in such an architecture? (Choose two.)

Select 2 answers
A.Routing decisions are computed per token, allowing different tokens in the same sequence to be processed by different expert subsets
B.Increasing the number of experts proportionally increases the number of activated experts per token
C.The router assigns tokens to experts once at initialization and keeps the assignment fixed throughout training
D.Only the experts selected by the router for a given token are activated, so FLOPs per token scale with k rather than total expert count
E.All experts receive every token, but their outputs are averaged with learned scalar weights
AnswersA, D

The gating network produces a distribution over experts independently for each token, so token routing is dynamic and token-specific. This token-level granularity is what allows an MoE layer to specialize experts across linguistic or semantic patterns while keeping computation sparse for any individual token in the sequence.

Why this answer

Sparse MoE replaces a single feed-forward block with many expert blocks plus a router that, for each token, selects the top-k experts to run. This yields token-level dynamic routing and keeps per-token FLOPs tied to k and expert size rather than the total expert population. Dense averaging, static assignment, and automatic k growth all contradict the sparse, learned, token-specific routing that defines the architecture.

Exam trap

The trap here is assuming that adding more experts automatically increases per-token compute, when in fact the top-k selection keeps activated expert count fixed regardless of total expert count.

135
MCQmedium

You are evaluating a fine-tuned LLM for a code generation task. The model was trained using NVIDIA NeMo on a dataset of Python functions. You want to measure the percentage of generated functions that pass a set of unit tests. Which evaluation metric is most appropriate?

A.ROUGE-L
B.BLEU score
C.Perplexity
D.pass@k
AnswerD

pass@k measures the percentage of problems for which at least one of k generated samples passes all unit tests. It directly evaluates functional correctness by executing the code against test cases. This metric is standard for code generation tasks and aligns with the scenario's goal of measuring how many generated functions pass unit tests. It accounts for the stochastic nature of LLM outputs by considering multiple samples.

Why this answer

The goal is to measure functional correctness by running unit tests. pass@k does exactly that by executing generated code and checking test outcomes, while also accounting for multiple samples. BLEU, ROUGE-L, and perplexity are text-based metrics that do not verify execution or test passage, so they cannot reliably assess code generation success in this scenario.

Exam trap

The trap here is choosing a text similarity metric like BLEU or ROUGE for code generation, when functional correctness requires execution-based metrics like pass@k.

136
MCQhard

Refer to the exhibit. Why did the system fail to trigger an alert despite high latency?

A.The error rate threshold was too low
B.The alert is suppressed by a duration requirement
C.The Prometheus backend is down
D.The logging level is set too high
AnswerB

Monitoring systems usually require a threshold to be exceeded for a specific time duration to prevent noise from transient spikes. Since the alert did not fire despite exceeding the 500ms limit, a duration-based trigger condition is the most probable cause for the suppression.

Why this answer

Alerting systems often require sustained violation of thresholds to avoid 'flapping' or false positives. The exhibit shows the current latency is 850ms, which exceeds the threshold, but the alerting logic likely requires consecutive samples or a window-based average to trigger. This is a common reliability feature in monitoring stacks to ensure that transient spikes do not disrupt operational teams with unnecessary alerts during non-critical fluctuations.

Exam trap

Candidates often assume the system failed due to a misconfigured threshold or a bug in the monitoring agent, ignoring the common practice of duration-based alert suppression to prevent false positives.

137
MCQhard

An engineer is deploying a large language model for real-time inference on an NVIDIA A100 GPU. The model uses FP16 weights but the inference server must handle variable-length input sequences. Profiling shows that the GPU spends significant time on memory-bound operations and that kernel launches are frequent. Which optimization is most appropriate to reduce latency while maintaining accuracy?

A.Increase the number of CUDA streams to parallelize kernel execution.
B.Enable NVIDIA GPUDirect Storage to accelerate data loading.
C.Use NVIDIA TensorRT with dynamic batching and CUDA graphs to capture the inference step.
D.Convert the model to INT8 using TensorRT with calibration.
AnswerC

TensorRT optimizes the graph and supports dynamic shapes for variable-length sequences. Dynamic batching groups requests to improve GPU utilization, while CUDA graphs capture the sequence of kernel launches into a single graph, drastically reducing launch overhead. This combination directly addresses both memory-bound inefficiencies and frequent launches without sacrificing accuracy.

Why this answer

TensorRT with dynamic batching and CUDA graphs targets the two identified issues: variable-length sequences are handled by dynamic shapes and batching, while CUDA graphs eliminate launch overhead by capturing the entire inference step. This approach maintains FP16 accuracy and improves latency without the risks of quantization.

Exam trap

The trap here is focusing solely on precision reduction as the default latency fix, while overlooking launch overhead and batching strategies that can yield larger gains with no accuracy cost.

138
Multi-Selectmedium

A team is optimizing a large language model for inference on NVIDIA GPUs. They want to reduce the memory footprint of the model to fit on a single GPU with limited VRAM. Which two techniques are most appropriate for reducing memory usage during inference? (Choose two.)

Select 2 answers
A.Quantization of weights to INT8 or FP8.
B.Using NVIDIA TensorRT with layer fusion and kernel auto-tuning.
C.Enabling CUDA graphs to capture the inference step.
D.Pruning or sparsifying the model weights.
E.Increasing the batch size to improve utilization.
AnswersA, D

Quantizing weights to lower precision reduces the memory required to store the model parameters, often by 2x or 4x. This directly lowers VRAM usage and can also speed up inference on supported hardware. It is a standard technique for fitting large models on constrained GPUs, though it may require calibration to maintain accuracy.

Why this answer

Quantization and pruning directly reduce the memory required to store and execute the model. Quantization lowers the precision of weights, while pruning removes unnecessary parameters. Both techniques decrease VRAM usage and can be combined for greater effect.

The other options either increase memory usage or address different bottlenecks like launch overhead.

Exam trap

The trap here is assuming that any optimization that improves inference efficiency also reduces memory; techniques like CUDA graphs and larger batches do not lower VRAM footprint.

139
MCQeasy

A developer is building an interactive assistant using NVIDIA NIM microservices. The assistant must answer questions about a specific set of internal policies. The developer wants to ensure the model's responses are grounded in those policies and not in its general pre-training knowledge. Which prompt engineering technique should be applied?

A.Fine-tune the model on the internal policy documents, then use a simple zero-shot prompt to ask questions.
B.Use a chain-of-thought prompt that asks the model to reason step by step about the policy before answering.
C.Increase the temperature setting to 0.9 to encourage the model to explore a wider range of policy interpretations.
D.Provide the relevant policy excerpts directly within the prompt, instruct the model to answer only from that context, and to state when the answer is not present.
AnswerD

This is retrieval-augmented generation (RAG) at the prompt level. By injecting the policy excerpts into the context and explicitly constraining the model to use only that information, you prevent it from relying on its pre-trained knowledge. Instructing it to say when the answer is missing further reduces hallucination and keeps the response grounded in the provided internal policies.

Why this answer

Grounding the model in a specific set of documents requires placing the relevant excerpts in the prompt and explicitly instructing the model to answer only from that context. This constrains the model to the provided facts and reduces hallucination. Other techniques like increasing temperature, chain-of-thought, or fine-tuning do not achieve this grounding for a dynamic, interactive policy assistant.

Exam trap

The trap here is assuming that fine-tuning or chain-of-thought alone will make the model answer from a specific set of documents, when in fact the documents must be supplied in the prompt context.

140
MCQmedium

You are preparing a large instruction-tuning dataset for an NVIDIA NeMo-based LLM. The raw data consists of user queries and assistant responses collected from a customer support system, stored as JSON lines. During preprocessing, you notice that many responses contain personally identifiable information (PII) such as names, email addresses, and phone numbers. You need to ensure the dataset is safe for training while preserving as much semantic content as possible. Which approach is most appropriate for handling PII in this dataset?

A.Leave the PII intact but exclude the dataset from production use and train only in a sandbox environment.
B.Hash each detected PII string with SHA-256 and substitute the hash into the text.
C.Apply a rule-based regex filter to remove any line containing PII, then train on the remaining data.
D.Use a named entity recognition (NER) model to detect PII and replace each entity with a generic placeholder (e.g., [NAME]).
AnswerD

Replacing PII with placeholders preserves sentence structure and semantic meaning while removing sensitive information. This is a standard de-identification technique that maintains data utility for instruction tuning. It also handles variations better than simple regex and is compatible with NeMo preprocessing pipelines.

Why this answer

Replacing PII with generic placeholders using an NER model is the most balanced approach: it removes sensitive information while preserving the linguistic patterns and context needed for instruction tuning. This method aligns with NVIDIA NeMo's data preprocessing recommendations for responsible AI. It also avoids the pitfalls of discarding data or introducing unintelligible tokens.

Exam trap

The trap here is assuming that regex-based removal is sufficient, but PII can be unstructured and context-dependent, making NER more reliable.

141
MCQmedium

A healthcare analytics company is deploying an LLM-based patient triage assistant using NVIDIA NIM microservices. Compliance requires that every model response be traceable to a specific model version, input prompt, and retrieved context for a minimum of three years. Which approach best satisfies this auditability requirement?

A.Enable NeMo Guardrails output moderation and rely on the guardrail logs to reconstruct decisions.
B.Instrument the NIM inference pipeline to emit structured audit records containing model ID, prompt hash, retrieved context, and response to a write-once log store.
C.Increase the model's temperature logging verbosity and store raw GPU telemetry alongside responses.
D.Configure the NIM container to retain its model weights and prompt templates for three years in cold storage.
AnswerB

Structured audit records emitted at inference time and stored immutably preserve the exact lineage of each response: which model version, which prompt, which retrieved context, and what was returned. This directly satisfies the traceability requirement and supports long-term retention without relying on downstream reconstruction.

Why this answer

Traceability for regulated LLM deployments requires capturing the full inference lineage at the moment of generation. Structured, immutable audit records that bind model version, prompt, retrieved context, and response give auditors a reproducible chain of evidence. Guardrail logs, GPU telemetry, and cold-stored weights each cover only part of the picture and cannot substitute for per-inference provenance.

Exam trap

The trap here is assuming that storing model weights or guardrail logs is equivalent to storing per-inference audit records that bind prompt, context, model version, and response together.

142
MCQmedium

A team is training a large language model using NVIDIA DGX A100 nodes with 8 GPUs per node. They observe that GPU utilization is high on all GPUs, but the training throughput scales poorly when adding more nodes. Profiling shows that the communication time during all-reduce operations increases significantly with node count. Which of the following optimizations is most likely to improve scaling efficiency?

A.Use gradient accumulation to increase the effective batch size.
B.Increase the number of CPU threads dedicated to data loading.
C.Enable NCCL's tree algorithm for all-reduce operations.
D.Configure NCCL to use InfiniBand with GPUDirect RDMA and enable adaptive routing.
AnswerD

Using InfiniBand with GPUDirect RDMA allows GPUs to communicate directly with the network adapter, bypassing host memory and reducing latency. Adaptive routing dynamically selects less congested paths, improving bandwidth utilization across multiple nodes. This combination significantly reduces communication overhead in large-scale all-reduce operations, directly addressing the poor scaling observed when adding nodes.

Why this answer

The poor scaling with increasing node count points to inter-node communication overhead, specifically during all-reduce operations. Leveraging InfiniBand with GPUDirect RDMA and adaptive routing optimizes the network path, reducing latency and improving bandwidth. This directly targets the communication bottleneck, enabling better scaling efficiency for distributed LLM training across multiple DGX nodes.

Exam trap

The trap here is assuming that gradient accumulation or CPU thread tuning addresses communication overhead, when the bottleneck is specifically inter-node all-reduce performance.

143
Multi-Selecthard

A developer is using NVIDIA TensorRT-LLM to optimize a GPT-based model for inference. They want to reduce the model's memory footprint and improve throughput without retraining. Which two techniques can be applied during the TensorRT-LLM build process to achieve these goals? (Choose two.)

Select 2 answers
A.Enabling FP8 precision for both weights and activations on Hopper GPUs
B.Pruning the model by removing entire attention heads
C.Weight-only quantization (e.g., INT8 or INT4) for linear layers
D.Knowledge distillation from a larger teacher model
E.Using a larger batch size to amortize memory overhead
AnswersA, C

FP8 precision on Hopper GPUs (e.g., H100) reduces memory usage and increases throughput by using 8-bit floating point for weights and activations. TensorRT-LLM supports FP8 quantization, which can be applied during build without retraining. This directly addresses memory footprint and throughput, making it a correct choice.

Why this answer

Weight-only quantization and FP8 precision are both build-time optimizations in TensorRT-LLM that reduce memory footprint and improve throughput without retraining. Weight-only quantization lowers weight precision, while FP8 leverages Hopper GPU capabilities for both weights and activations. The other options either require retraining or increase memory usage.

Exam trap

The trap here is assuming that any memory-reduction technique like pruning is suitable, but pruning often requires retraining and is not a standard TensorRT-LLM build option.

144
MCQmedium

What is the primary benefit of using NVIDIA DCGM (Data Center GPU Manager) for monitoring production LLMs?

A.To train LLMs faster
B.To detect hardware-level issues
C.To optimize model weights
D.To generate natural language output
AnswerB

DCGM is specifically designed to provide deep hardware insights, including temperature, power, and memory errors. Identifying these problems early is critical for infrastructure reliability, as it allows for maintenance to be scheduled before a component causes a hard failure.

Why this answer

DCGM provides hardware-level telemetry that is essential for proactive maintenance and reliability. Unlike higher-level metrics that only show software performance, DCGM can identify physical issues such as ECC memory errors, thermal throttling, or failing power supplies. This level of visibility is crucial for anticipating hardware failure before it results in a service outage, allowing operations teams to migrate workloads gracefully and maintain high system reliability.

Exam trap

Candidates often confuse DCGM with software-level monitoring tools like Prometheus or Grafana. They incorrectly assume it monitors model accuracy or inference throughput rather than physical hardware health and GPU-specific telemetry.

145
MCQmedium

A healthcare company is deploying an LLM for clinical note summarization using NVIDIA Triton Inference Server. They must ensure that only authorized users can access the model and that all inference requests are logged for audit. Which Triton feature should they configure to enforce authentication and authorization?

A.Triton's built-in model access control list (ACL) configured via the model repository.
B.Triton's HTTP/REST and gRPC endpoints with a reverse proxy that performs OAuth 2.0 token validation.
C.Triton's dynamic batching configuration with priority levels.
D.Triton's ensemble scheduler with a custom authentication model.
AnswerB

Triton itself does not provide built-in authentication or authorization. The recommended approach is to place a reverse proxy (such as NGINX or Envoy) in front of Triton to handle OAuth 2.0 token validation and access control. This satisfies the requirement for authorized access while Triton focuses on inference. Logging can be handled at the proxy or application level for audit.

Why this answer

Triton Inference Server does not include native authentication or authorization. The standard pattern is to deploy a reverse proxy that validates OAuth 2.0 tokens and forwards authorized requests to Triton. This separates security concerns from inference and allows audit logging at the proxy.

Other options describe scheduling or non-existent features that do not enforce access control.

Exam trap

The trap here is assuming Triton has built-in user authentication, when it actually relies on external components like reverse proxies for security.

146
MCQmedium

When implementing Chain-of-Thought (CoT) prompting for a complex NVIDIA NeMo-based reasoning task, what is the primary benefit of encouraging the model to generate intermediate steps?

A.It forces the model to use more GPU memory per token.
B.It increases the likelihood of the model selecting a random seed.
C.It decomposes complex problems into verifiable logical segments.
D.It eliminates the need for system-level instructions entirely.
AnswerC

Breaking down multi-step problems into smaller, sequential steps allows the model to maintain context and reduces cumulative error rates. By generating intermediate proofs or calculations, the model provides a trace that developers can analyze to identify where the reasoning failed, which is vital for robust application development.

Why this answer

Chain-of-Thought prompting decomposes complex problems into sequential logical steps, which is critical when using LLMs for technical reasoning tasks. By forcing the model to articulate its internal logic, the likelihood of hallucination decreases significantly. This practice is essential for NVIDIA engineers deploying agents that require high-precision output, as it creates an audit trail for the model's reasoning process and allows for better debugging of multi-turn interactions.

Exam trap

Candidates often confuse CoT with simple 'few-shot' prompting, assuming the primary goal is just to provide examples rather than forcing the model to articulate the logical steps required for complex verification.

147
Multi-Selectmedium

Which TWO of the following practices are recommended for ensuring ethical AI development when using NVIDIA NIMs in an enterprise environment?

Select 2 answers
A.Perform periodic bias audits on model responses using diverse, representative evaluation datasets.
B.Bypass local logging to optimize inference latency for high-throughput applications.
C.Maintain comprehensive logs of inputs and outputs for auditability and compliance tracking.
D.Use the model in a closed-loop system where users cannot report offensive content.
E.Exclude all metadata from model outputs to prevent potential privacy leaks.
AnswersA, C

Regular bias auditing is a fundamental component of ethical AI. By evaluating model outputs against diverse datasets, developers can identify and address discriminatory or toxic tendencies. This proactive approach helps ensure the model behaves equitably across different demographics, which is a critical ethical requirement for large-scale enterprise applications.

Why this answer

Ensuring ethical AI involves both technical validation and process transparency. Regular bias auditing helps detect unintended stereotyping, while implementing robust logging ensures accountability for every generated output. These practices allow organizations to monitor for discriminatory patterns and maintain an audit trail for compliance, which is essential for building trust with users and adhering to global ethical standards for AI deployment in sensitive business sectors.

Exam trap

Candidates often select 'automated model retraining' as an ethical practice. Retraining is not a substitute for active bias auditing and logging, which are required for governance and accountability.

148
MCQhard

A team is fine-tuning a pretrained decoder-only model on a small domain-specific dataset. They observe that the model quickly overfits and loses general language ability. They want to update only a small number of additional parameters while keeping the base weights frozen. Which approach should they use?

A.Full fine-tuning with a very low learning rate and early stopping.
B.Increasing the dropout rate in the feed-forward layers during fine-tuning.
C.Freezing the embedding layer only and training all transformer blocks.
D.Low-Rank Adaptation (LoRA), which injects trainable low-rank matrices into existing layers while freezing the base weights.
AnswerD

LoRA adds small trainable rank-decomposition matrices to selected weight matrices and leaves the pretrained weights frozen. This updates far fewer parameters, reduces optimizer memory, and is less prone to catastrophic forgetting, which matches the team's constraints.

Why this answer

Parameter-efficient fine-tuning methods such as LoRA freeze the pretrained weights and train only small injected matrices, which reduces optimizer memory and limits drift from the base model. This is well suited to small domain datasets where full fine-tuning would overfit and degrade general language ability.

Exam trap

The trap here is thinking that a low learning rate or extra dropout makes full fine-tuning parameter-efficient, when the base weights are still being updated.

149
MCQeasy

An engineer needs to expose an LLM served by NVIDIA Triton Inference Server to a web application over HTTP with token streaming. Which Triton feature should they enable?

A.Model ensembles combining preprocessing and postprocessing models.
B.Dynamic batching in the model configuration.
C.Instance groups with multiple GPU instances per model.
D.Decoupled mode with streaming responses in the model backend.
AnswerD

Decoupled mode allows a model backend to return multiple responses for a single request, which is essential for token streaming in LLMs. Triton's HTTP and gRPC endpoints support streaming when the model is configured for decoupled transactions. This lets the web application receive partial outputs as tokens are generated, improving perceived latency.

Why this answer

Triton's decoupled mode lets a backend emit multiple responses for one request, which is the mechanism used for LLM token streaming. Configuring the model for decoupled transactions and using a streaming-capable client over HTTP or gRPC delivers tokens incrementally. Other Triton features like batching, ensembles, or instance groups improve throughput or composition but do not stream partial outputs.

Exam trap

The trap here is confusing throughput optimizations such as dynamic batching with the response-streaming capability required for token-by-token delivery.

150
Multi-Selectmedium

When assessing the quality of a dataset for instruction fine-tuning, which TWO metrics or methods are considered most reliable for measuring dataset diversity?

Select 2 answers
A.Embedding-based clustering to visualize topical coverage across the corpus.
B.Calculating the total number of words in the dataset.
C.Perplexity distribution analysis across segments of the dataset.
D.Measuring the average response length for every instruction.
E.Checking if the dataset is solely in ASCII format.
AnswersA, C

Clustering document embeddings is a standard way to verify that the dataset covers a wide spectrum of topics. If the embeddings form only a few dense clusters, the dataset is likely too narrow. A diverse dataset should exhibit a broader distribution across the vector space, indicating varied content.

Why this answer

Dataset diversity ensures that the model encounters a broad range of topics and linguistic patterns, which is critical for generalization. Using embedding-based clustering allows for the identification of thematic coverage, while perplexity distribution analysis helps assess whether the dataset contains a balance of common and complex structures. These methods together provide a quantitative view of the data's breadth, reducing the risk of bias or overfitting.

Exam trap

Candidates often rely on simple text length or word count metrics, overlooking advanced quantitative methods like embedding-based clustering and perplexity distribution for measuring true dataset diversity.

Page 1

Page 2 of 5

Page 3

All pages