Courseiva

NVIDIA Certified Professional: Generative AI LLMs (NCP-GENL) — Questions 151–225

352 questions total · 5pages · All types, answers revealed

Page 2

Page 3 of 5

Page 4
151
MCQmedium

A team is using an NVIDIA NIM for a Mistral model to classify support tickets into one of five fixed categories. Accuracy is inconsistent, and the model sometimes invents new categories. Which prompt engineering change is most likely to improve reliability without retraining the model?

A.Add a chain-of-thought instruction that asks the model to reason step by step before naming a category.
B.Increase the context window by concatenating the entire ticket history for every request.
C.Provide three labeled examples per category in the prompt and instruct the model to output only the category name.
D.Lower the temperature to 0.0 and remove all instructions so the model relies on its pretrained knowledge.
AnswerC

Few-shot examples define the exact label set and demonstrate the mapping from ticket text to category. Instructing the model to output only the category name removes room for invented labels. This directly addresses the inconsistency and the hallucinated categories without any fine-tuning, making it the most effective change for a fixed-label classification task.

Why this answer

Few-shot examples with explicit labels define the allowed output space, and the instruction to emit only the category name prevents invented labels. This combination directly targets both symptoms, inconsistency and hallucinated categories, and requires no model retraining, unlike sampling or reasoning changes that leave the label set undefined.

Exam trap

The trap here is believing that lowering temperature guarantees correct classification, when the real issue is that the allowed labels were never defined in the prompt.

152
MCQmedium

A team is serving a 13B-parameter LLM on a single NVIDIA A100 80GB GPU. During generation, they observe that the GPU compute utilization stays below 20% while memory bandwidth utilization is near saturation. They want to improve throughput without changing the model architecture. Which optimization is most appropriate?

A.Apply 4-bit weight-only quantization to reduce memory traffic during token generation.
B.Enable multi-GPU tensor parallelism across two A100 GPUs to split the model.
C.Increase the batch size to improve arithmetic intensity and better utilize tensor cores.
D.Increase the number of CPU threads used for token sampling to speed up generation.
AnswerA

The workload is memory-bandwidth bound, as shown by low compute utilization and saturated memory bandwidth. Reducing weight precision to 4-bit lowers the bytes transferred per token, directly alleviating the bottleneck and increasing throughput without changing the architecture. This is a standard optimization for memory-bound LLM inference.

Why this answer

The symptoms indicate a memory-bandwidth-bound workload, common in autoregressive LLM decoding. Reducing weight precision via 4-bit quantization decreases the volume of data read from GPU memory per token, directly increasing throughput. Other options either do not target memory bandwidth or introduce new bottlenecks.

Exam trap

The trap here is assuming that low compute utilization always means the GPU needs more parallel work, when it can instead indicate a memory-bandwidth bottleneck.

153
MCQhard

An inference team runs a TensorRT-LLM model behind NVIDIA Triton Inference Server with the tensorrtllm_backend. They notice that when clients send requests with widely varying prompt lengths, throughput drops and some requests wait much longer than others. They want Triton to group compatible requests dynamically to improve GPU utilization without changing the model build. Which Triton feature should they configure?

A.Model warmup with sample inputs to pre-allocate memory at load time
B.Dynamic batching with a preferred batch size and max queue delay in the model's config.pbtxt
C.Rate limiting with a max queue size to reject excess requests
D.Sequence batching with a sequence ID to maintain state across requests
AnswerB

Dynamic batching lets Triton combine independent requests arriving within a time window into a single batch, improving GPU utilization when request sizes vary. Setting preferred_batch_size and max_queue_delay controls the tradeoff between latency and throughput, and it requires no model rebuild.

Why this answer

Triton's dynamic batching groups independent requests arriving within max_queue_delay into a single batch, using preferred_batch_size to target efficient batch shapes. This improves GPU utilization under varying prompt lengths without rebuilding the TensorRT-LLM engine, directly addressing the throughput and latency variance described.

Exam trap

The trap here is confusing sequence batching, which maintains state for ordered requests, with dynamic batching, which merges independent concurrent requests.

154
MCQmedium

You are preparing a multilingual corpus for pre-training an LLM. The corpus contains documents in English, Spanish, and German, but the English portion is 80% of the data. You want to avoid the model becoming biased toward English. Which data preparation technique should you apply?

A.Remove all English documents to force multilingual learning
B.Translate all documents into English
C.Apply a language-specific tokenizer for each language
D.Oversample the non-English documents or undersample the English documents
AnswerD

Adjusting the sampling ratios by oversampling minority languages or undersampling the dominant language balances the effective training distribution. This reduces English bias and encourages the model to allocate capacity to Spanish and German. It is a standard technique in multilingual data preparation and can be implemented via NeMo Curator or custom sampling logic.

Why this answer

Balancing the language distribution by oversampling minority languages or undersampling the dominant one directly addresses the 80% English skew. This ensures the model sees a more representative mix during training, reducing bias toward English and improving performance on Spanish and German. It is a standard data preparation technique for multilingual corpora.

Exam trap

The trap here is thinking that tokenizer changes or translation solve imbalance, when the core issue is the proportion of training examples per language.

155
MCQmedium

A company is running a large language model inference service on NVIDIA GPUs. They observe that GPU memory is nearly full, limiting the batch size and thus throughput. The model weights are stored in FP16, and the KV cache consumes a significant portion of memory. Which technique can reduce memory usage while maintaining model accuracy and enabling larger batch sizes?

A.Quantize the KV cache to INT8.
B.Increase the number of GPUs to distribute the model.
C.Convert the model to FP32 to improve numerical stability.
D.Use gradient checkpointing during inference.
AnswerA

Quantizing the KV cache to INT8 reduces its memory footprint by half compared to FP16, allowing larger batch sizes. Modern techniques like NVIDIA's TensorRT-LLM support INT8 KV cache quantization with minimal accuracy loss. This directly addresses the memory bottleneck while preserving model accuracy, making it an effective solution for increasing throughput.

Why this answer

The memory bottleneck is largely due to the KV cache. Quantizing it to INT8 halves its memory footprint, freeing space for larger batch sizes and improving throughput. This technique is supported in frameworks like TensorRT-LLM and maintains accuracy with proper calibration.

Other options either do not apply to inference or increase memory usage.

Exam trap

The trap here is confusing training-time memory optimizations like gradient checkpointing with inference-time techniques, or assuming that increasing precision helps when the goal is to reduce memory.

156
MCQmedium

A team is testing a new LLM application. During red-teaming, the model consistently leaks sensitive internal project codenames. How should the team address this systematically?

A.Increase the number of training epochs on the existing dataset.
B.Deploy a guardrail that filters output against a list of sensitive terms.
C.Add a disclaimer at the end of every response stating that the content is confidential.
D.Randomize the model's weights during every inference run.
AnswerB

Implementing an output guardrail is the most effective way to intercept sensitive information before it reaches the user. By explicitly defining a list of restricted terms, the system can block or sanitize the response in real-time, providing a robust safety net for protecting internal corporate confidential data.

Why this answer

Systematic mitigation of data leakage requires a multi-layered approach. Modifying the base model is rarely sufficient; instead, one must implement output-side guardrails that perform pattern matching and dictionary-based filtering. This ensures that even if the model attempts to generate sensitive info, the guardrail intercepts it.

This process protects intellectual property and maintains compliance with corporate confidentiality agreements by ensuring that protected information remains strictly within the secure environment.

Exam trap

Candidates often suggest fine-tuning or retraining the model to remove sensitive data. This is ineffective because models can still hallucinate or reconstruct sensitive information, and it fails to provide a real-time, auditable safety layer.

157
MCQeasy

A production LLM inference service on NVIDIA Triton Inference Server is experiencing intermittent slowdowns. The team wants a single metric that best indicates whether the GPU is the bottleneck during these periods. Which metric should they monitor?

A.GPU utilization (SM utilization)
B.Disk I/O operations per second on the model repository
C.Network throughput between clients and the server
D.CPU utilization of the inference server host
AnswerA

GPU utilization, specifically SM utilization, directly measures the percentage of time the GPU's streaming multiprocessors are active. In LLM inference, high SM utilization sustained near 100% strongly suggests the GPU is the bottleneck. Monitoring this metric during slowdowns helps determine if the GPU is saturated by compute or memory-bound operations, guiding optimization such as batching or model quantization.

Why this answer

GPU utilization (SM utilization) is the most direct indicator of whether the GPU's compute units are saturated. During LLM inference, if slowdowns coincide with near-100% SM utilization, the GPU is the bottleneck. Other metrics like CPU, network, or disk I/O are less likely to explain intermittent latency spikes when the GPU is the primary compute resource.

Exam trap

The trap here is assuming that host CPU or network metrics are sufficient to diagnose GPU-bound LLM inference slowdowns.

158
MCQmedium

A team is pretraining a decoder-only Transformer LLM on a large corpus of code and natural language. They observe that the model's training loss decreases smoothly, but during generation it sometimes produces degenerate repetition, and attention entropy on long sequences collapses. They suspect the issue is related to the positional encoding scheme. Which architectural change is most likely to mitigate the attention entropy collapse while preserving the model's ability to generalize to sequences longer than those seen during pretraining?

A.Increase the number of attention heads while keeping the total hidden dimension constant, without changing the positional encoding.
B.Apply Layer Normalization before the self-attention sublayer instead of after it, leaving the positional encoding unchanged.
C.Replace learned absolute positional embeddings with Rotary Positional Embeddings (RoPE) applied to queries and keys in every self-attention layer.
D.Add a sinusoidal positional encoding to the input embeddings and freeze those embeddings during training.
AnswerC

RoPE encodes relative position by rotating query and key vectors by an angle proportional to their absolute position. Because attention scores depend only on the relative rotation between queries and keys, the model naturally generalizes to longer contexts and avoids the entropy collapse seen when learned absolute embeddings are extrapolated beyond their trained range. This matches the observed repetition and entropy issues.

Why this answer

The symptoms point to a positional encoding that does not generalize beyond the pretraining sequence length. Rotary Positional Embeddings inject relative position directly into the attention computation by rotating queries and keys, so attention scores depend on relative offsets rather than absolute indices. This preserves attention entropy on longer sequences and reduces degenerate repetition, making it the appropriate architectural change for this scenario.

Exam trap

The trap here is assuming that any positional encoding change, such as switching to sinusoidal or adding normalization, will fix length generalization, when only a relative-position scheme applied inside attention addresses the described entropy collapse.

159
Multi-Selectmedium

Which TWO of the following statements correctly describe the role of Grouped Query Attention (GQA) in modern LLM architectures?

Select 2 answers
A.GQA eliminates the need for separate query, key, and value projection layers.
B.GQA reduces the memory footprint of the KV cache during inference.
C.GQA improves model training speed by increasing the number of compute operations.
D.GQA achieves a compromise between Multi-Head Attention and Multi-Query Attention.
E.GQA prevents the model from using positional embeddings.
AnswersB, D

By reducing the total number of key and value heads, the size of the KV cache stored in GPU memory is significantly decreased. This reduction is highly beneficial for serving models at scale, as it allows for larger batch sizes or longer context lengths without exceeding available memory capacity.

Why this answer

GQA balances the performance of Multi-Head Attention (MHA) and the memory efficiency of Multi-Query Attention (MQA). By sharing keys and values across groups of heads, it reduces the size of the KV cache during inference. This is crucial for high-throughput deployment environments, as it optimizes memory bandwidth utilization without sacrificing the granular representational capacity that individual query heads provide.

Exam trap

Candidates often mistake GQA for a training-only optimization or a method that increases accuracy. GQA is specifically designed for inference-time efficiency by reducing memory bandwidth requirements.

160
MCQhard

Refer to the exhibit. An engineer is fine-tuning an LLM using the provided configuration. What is the primary purpose of applying 'yarn' scaling in this architecture?

A.It increases the number of attention heads to improve feature extraction.
B.It enables the model to process sequences longer than the original pre-trained limit.
C.It compresses the model weights to reduce storage requirements.
D.It replaces the attention mechanism with a recurrent neural network.
AnswerB

The primary goal of YaRN is to allow effective extrapolation and interpolation of positional information. By adjusting the base frequency of RoPE, the model can interpret position indices that fall outside the range seen during initial training, thereby allowing for a significantly larger context window during inference and fine-tuning.

Why this answer

YaRN (Yet another RoPE extension) is designed to extend the context window of pre-trained LLMs beyond their original training length. By modifying the Rotary Positional Embeddings (RoPE) base frequency, it allows the model to interpolate position indices without severe degradation in perplexity. This is essential for enterprise deployments requiring document retrieval or analysis tasks where input lengths exceed the base model's pre-trained constraints.

Exam trap

Candidates often confuse YaRN with quantization or pruning techniques. They assume any acronym related to model optimization is about reducing memory footprint rather than extending the context window.

161
MCQeasy

A developer has a fine-tuned Llama-family model in Hugging Face format and wants to run it with NVIDIA TensorRT-LLM on an H100. Which artifact must be produced before the runtime can execute the model?

A.A TensorRT engine built from the model's network definition for the target GPU architecture.
B.An ONNX graph exported with dynamic axes and a matching runtime configuration JSON.
C.A PyTorch TorchScript trace of the full forward pass saved as a .pt file.
D.A quantized GGUF file generated with a community conversion script.
AnswerA

TensorRT-LLM compiles the model graph into a serialized TensorRT engine that is specialized for the target GPU compute capability, chosen precision, and parallelism layout. The runtime loads and executes this engine; without it there is nothing for the executor to run. Building the engine is therefore the mandatory step between a Hugging Face checkpoint and inference.

Why this answer

TensorRT-LLM executes a serialized TensorRT engine that is compiled for a specific GPU architecture, precision, and parallelism configuration. After converting the Hugging Face checkpoint into the TensorRT-LLM checkpoint format, the developer must build the engine for the H100. Only that engine can be loaded by the runtime, so it is the required artifact before any inference can occur.

Exam trap

The trap here is confusing an interchange or community format such as ONNX or GGUF with the compiled TensorRT engine that the TensorRT-LLM runtime actually loads and executes.

162
MCQeasy

A developer is building a customer support assistant using an NVIDIA NIM microservice for a Llama 3 model. The assistant must always respond in valid JSON with keys 'category' and 'urgency'. The model often returns conversational text instead. Which prompt engineering change most directly enforces the required output format?

A.Append 'Return only JSON with keys category and urgency' and set the NIM request parameter 'guided_json' to the target schema.
B.Add 'Do not hallucinate' to the prompt and lower the max_tokens parameter to 50.
C.Add a system prompt that says 'You are a helpful assistant' and increase the temperature to 0.9.
D.Use few-shot examples of JSON outputs and set top_p to 0.1 without any schema constraint.
AnswerA

Combining an explicit instruction with NVIDIA NIM's guided_json parameter constrains decoding to the provided JSON schema, so the model cannot emit conversational text. The instruction aligns the model's intent while guided_json enforces structural validity at generation time. This is the most direct way to guarantee the required keys and format in the response.

Why this answer

Structured output requires both a clear instruction and a decoding-time constraint. NVIDIA NIM supports guided_json, which restricts token generation to a supplied JSON schema, guaranteeing valid keys and syntax. Pairing that with an explicit prompt instruction aligns the model's behavior with the schema.

Few-shot examples or temperature changes alone cannot guarantee strict JSON compliance.

Exam trap

The trap here is assuming that simply asking the model for JSON or providing examples is enough, when only schema-guided decoding like guided_json can guarantee valid structure.

163
MCQmedium

You are using NVIDIA NeMo Curator to filter a 600 GB web-crawl corpus before pre-training. Your team wants to remove exact duplicates and near-duplicates to reduce memorization and speed up training. Which NeMo Curator stage should you apply?

A.HeuristicFilter stage with a quality threshold
B.DocumentDownloader stage followed by DocumentResolver
C.ExactDuplicates and FuzzyDuplicates stages
D.ClassifierFilter stage using a domain classifier
AnswerC

NeMo Curator provides ExactDuplicates and FuzzyDuplicates stages that identify redundant documents; ExactDuplicates catches byte-identical texts, while FuzzyDuplicates uses MinHash/LSH to find near-duplicates. Applying both in sequence on the 600 GB corpus removes redundant content, lowering memorization risk and reducing the effective training token count.

Why this answer

NeMo Curator's ExactDuplicates and FuzzyDuplicates stages are purpose-built for removing redundant documents. ExactDuplicates handles byte-level matches, while FuzzyDuplicates uses MinHash and LSH to catch near-identical texts. Together they reduce corpus size without discarding unique content, which directly lowers memorization risk and training cost on a large web-crawl dataset.

Exam trap

The trap here is assuming that any quality filter removes duplicates, when quality filters score documents individually rather than comparing them to each other.

164
MCQhard

A healthcare AI team is evaluating a fine-tuned GPT-based model for clinical note summarization using NVIDIA NeMo. They need to assess whether the model's summaries contain no fabricated medical facts. Which evaluation approach is most appropriate?

A.Use a factual consistency metric such as SummaC or Q2
B.Measure BERTScore against reference summaries
C.Compute perplexity of the summaries
D.Calculate ROUGE-L scores
AnswerA

Factual consistency metrics like SummaC or Q2 evaluate whether generated summaries are entailed by the source document, detecting hallucinations. They are designed to catch fabricated facts by comparing summary claims against source content. In this healthcare scenario, applying such a metric directly addresses the need to ensure no invented medical facts, making it the most suitable approach.

Why this answer

The critical requirement is detecting fabricated medical facts, which demands checking entailment between the summary and the source clinical note. Factual consistency metrics such as SummaC or Q2 are specifically designed for this purpose, unlike similarity-based metrics (BERTScore, ROUGE-L) that compare to references and may miss hallucinations. Perplexity assesses fluency, not factuality.

Therefore, a factual consistency metric is the correct evaluation approach.

Exam trap

The trap here is confusing semantic similarity to a reference summary with factual consistency against the source document, which are fundamentally different checks.

165
Multi-Selecthard

When deploying a large language model on NVIDIA H100 GPUs using TensorRT-LLM, which TWO configuration strategies are most effective for improving KV cache efficiency and memory utilization?

Select 2 answers
A.Enable PagedAttention to reduce fragmentation.
B.Increase the static sequence length allocation.
C.Enable continuous inflight batching.
D.Disable multi-head attention optimizations.
E.Switch to FP64 precision for all tensors.
AnswersA, C

PagedAttention treats the KV cache as non-contiguous memory blocks, similar to virtual memory in an OS. This eliminates internal fragmentation caused by over-allocating memory for sequence lengths that never materialize, allowing the GPU to pack more requests into the same VRAM capacity for higher throughput.

Why this answer

Optimizing the KV cache is critical for scaling generative AI, as it occupies significant VRAM during inference. PagedAttention dynamically manages memory blocks to prevent fragmentation, while inflight batching ensures that continuous requests are processed without waiting for the entire batch to finish. Combining these allows for higher concurrency and reduced memory overhead, enabling more efficient deployment of models with large context windows on limited GPU hardware.

Exam trap

Candidates frequently select generic training optimizations like data parallelism instead of focusing specifically on inference-centric KV cache and batching strategies.

166
MCQhard

A team is deploying a large language model using NVIDIA TensorRT-LLM on a multi-GPU node with NVLink. They want to minimize inter-GPU communication overhead during inference. Which parallelism strategy should they use to achieve this?

A.Pipeline parallelism (PP)
B.Expert parallelism (EP) in a Mixture-of-Experts model
C.Data parallelism (DP)
D.Tensor parallelism (TP)
AnswerD

Tensor parallelism splits individual layers across GPUs, requiring frequent all-reduce operations for activations. NVLink provides high bandwidth and low latency, making TP efficient. By using TP, the team can minimize communication overhead compared to other strategies that might use slower interconnects or require more synchronization, thus achieving the goal.

Why this answer

Tensor parallelism splits layers across GPUs and uses all-reduce for activations. On a system with NVLink, the high bandwidth and low latency of NVLink make TP efficient, minimizing communication overhead. Pipeline parallelism introduces bubbles, data parallelism replicates the model, and expert parallelism is for MoE models and incurs all-to-all communication.

Exam trap

The trap here is assuming that pipeline parallelism minimizes communication because it reduces frequency, but it introduces pipeline bubbles and does not leverage NVLink as effectively as tensor parallelism.

167
Multi-Selecthard

An engineer is optimizing a large language model for inference on NVIDIA GPUs and wants to reduce memory usage to fit a larger model or increase batch size. Which two techniques are most effective for reducing GPU memory consumption during inference? (Choose two.)

Select 2 answers
A.Enable activation checkpointing to recompute activations during the backward pass.
B.Increase the number of micro-batches to overlap computation and memory transfers.
C.Quantize model weights to 8-bit integers using TensorRT or similar tools.
D.Use a paged attention mechanism to manage the KV cache efficiently.
E.Store the model weights on the CPU and transfer them to the GPU on demand.
AnswersC, D

Quantizing weights to 8-bit integers reduces the memory footprint by up to 4x compared to FP32. This allows larger models to fit in GPU memory and enables larger batch sizes. TensorRT supports INT8 quantization with calibration to maintain accuracy, making it a standard technique for memory reduction during inference.

Why this answer

Quantizing weights to INT8 reduces memory footprint by storing weights in lower precision, while paged attention optimizes KV cache memory by reducing fragmentation. Both are effective for fitting larger models or increasing batch size during inference. Activation checkpointing and micro-batching are training or throughput techniques, and CPU offloading introduces latency.

Exam trap

The trap here is confusing training-time memory optimizations, like activation checkpointing, with inference-time memory reductions, such as weight quantization and paged attention.

168
MCQhard

An engineer is using TensorRT-LLM to serve a model that occasionally receives prompts far longer than the typical 512 tokens, up to 8K tokens. With the default engine settings, requests near 8K fail with a cache capacity error while short requests succeed. Which configuration change most directly resolves this without rebuilding for a single worst-case shape?

A.Set the maximum input length and maximum sequence length to accommodate 8K tokens and size the KV cache pool accordingly.
B.Enable INT8 KV cache quantization so each token's cached keys and values consume fewer bytes.
C.Enable chunked context prefill so long prompts are processed in multiple smaller prefill passes.
D.Reduce the maximum batch size to 1 so all KV cache blocks are available to a single long request.
AnswerA

TensorRT-LLM engines have fixed maximum input and sequence length settings, and the KV cache pool must be large enough to hold the longest sequence the engine was built for. Raising these limits and sizing the pool for 8K tokens lets long prompts allocate their blocks, while shorter requests continue to use only the blocks they need.

Why this answer

The failure is a capacity limit tied to the engine's configured maximum input and sequence lengths and the KV cache pool sized for them. Raising those limits to cover 8K tokens and sizing the pool to match lets long prompts reserve the blocks they need, while paged allocation means short requests still consume only a few blocks, so a single worst-case build shape is unnecessary.

Exam trap

The trap here is reaching for memory-saving features like KV cache quantization or prefill chunking when the real constraint is that the engine was built with a maximum sequence length too small for the incoming prompts.

169
MCQmedium

An LLM inference service on NVIDIA Triton Inference Server uses dynamic batching with a max_queue_delay of 500 microseconds. During a load test, p99 latency exceeds the SLA while GPU utilization remains below 40%. Which change should you make first to improve latency without reducing throughput?

A.Reduce max_queue_delay to a smaller value to allow quicker dispatch of requests.
B.Increase the max_batch_size in the model configuration to allow larger batches.
C.Enable model warmup to reduce first-inference latency.
D.Switch to a larger GPU with more memory to increase batch capacity.
AnswerA

The p99 latency is caused by requests waiting in the dynamic batching queue for up to 500 microseconds. Lowering max_queue_delay reduces this wait, improving tail latency. Because GPU utilization is low, throughput is not constrained by batch size, so reducing the delay will not significantly reduce throughput and directly addresses the latency symptom.

Why this answer

The high p99 latency with low GPU utilization indicates that requests are spending too long in the dynamic batching queue. Reducing max_queue_delay allows requests to be dispatched sooner, cutting tail latency. Since GPU utilization is low, throughput is not limited by batch size, so this change improves latency without harming throughput.

Exam trap

The trap here is assuming that increasing batch size or GPU resources will fix latency, when the real issue is the batching delay.

170
MCQmedium

Which technique is most appropriate for a task requiring an LLM to generate code in a specific enterprise-internal syntax that is not well-represented in its public training data?

A.Zero-shot prompting with broad general coding instructions.
B.Few-shot prompting with multiple code examples.
C.Increasing the model temperature to encourage exploration.
D.Reducing the context window to force brevity.
AnswerB

By providing multiple examples of the target syntax, you enable the model to perform in-context learning of the specific patterns required. This pattern-matching approach allows the model to generalize the internal syntax correctly, providing accurate outputs that conform to enterprise standards without needing to perform full model retraining.

Why this answer

When dealing with proprietary or rare syntax, few-shot prompting with high-quality, representative examples is the most effective way to guide the model. By including these examples within the prompt, you provide the context the model lacks, significantly reducing syntax errors. This is a critical skill for NVIDIA developers building custom coding assistants for internal proprietary frameworks or legacy infrastructure support.

Exam trap

Candidates often mistakenly suggest fine-tuning as the first step, ignoring that few-shot prompting is a faster, more effective way to introduce specific, rare syntax without the overhead of retraining.

171
MCQmedium

An engineer is building a customer-facing FAQ bot using an NVIDIA NIM-hosted Llama 3.1 70B model. The bot must answer ONLY from a supplied product knowledge base and must respond with 'I don't have that information' when the answer is not present. Which prompt engineering approach BEST enforces this constraint?

A.Instruct the model to answer only from the provided knowledge base, explicitly state that it must reply with the refusal phrase when the answer is absent, and include the knowledge base inside clearly delimited sections.
B.Set the top_p value to 0.1 and rely on the model's pretrained knowledge of the product domain instead of supplying the knowledge base in the prompt.
C.Append the entire knowledge base to every user message without any instruction about scope or refusal behavior, letting the model infer the rules from context.
D.Increase the temperature to 1.0 and add 'Be creative and helpful' to the system prompt so the model can improvise when the knowledge base is incomplete.
AnswerA

Combining an explicit grounding instruction, a mandated refusal string, and clearly delimited context gives the model a precise behavioral contract. The delimiters separate trusted context from user input, while the refusal phrase defines the exact fallback behavior. This is the most reliable prompt-level method to constrain an NIM-hosted model to grounded answers without retraining.

Why this answer

Grounded FAQ bots need an explicit behavioral contract: what source to use, what to do when the source lacks the answer, and clear separation of context from user input. Delimiters prevent prompt injection from blending with instructions, and a fixed refusal phrase makes the fallback deterministic and testable. Sampling parameters alone cannot enforce scope constraints.

Exam trap

The trap here is assuming that lowering temperature or top_p will prevent hallucination, when grounding actually depends on explicit instructions and delimited context, not on sampling parameters.

172
MCQeasy

A developer is using NVIDIA Nsight Systems to profile a PyTorch training loop on an NVIDIA GPU. They notice significant gaps between kernel executions and want to identify whether the bottleneck is CPU-side or GPU-side. Which Nsight Systems feature should they use to visualize the CPU and GPU timelines together?

A.The Nsight Compute kernel profiling report.
B.The NVIDIA Management Library (NVML) GPU utilization metrics.
C.The PyTorch autograd profiler output.
D.The CUDA API trace and GPU activity timeline.
AnswerD

Nsight Systems provides a unified timeline that shows CUDA API calls on the CPU and corresponding kernel executions on the GPU. By correlating these, developers can see gaps where the GPU is idle waiting for CPU work, indicating a CPU-side bottleneck. This is the primary feature for identifying such imbalances.

Why this answer

Nsight Systems' unified timeline displays CPU activities (including CUDA API calls) and GPU kernels on the same time axis. This allows developers to see gaps where the GPU is idle while the CPU is busy, indicating a CPU-bound workload. It is the standard tool for this type of system-level analysis.

Exam trap

The trap here is confusing Nsight Compute (kernel-level) with Nsight Systems (system-level timeline), when the latter is needed for CPU-GPU correlation.

173
MCQmedium

Refer to the exhibit. An engineer receives these logs while converting a Transformer model to a TensorRT engine. What is the most appropriate action to resolve this build failure?

A.Reduce the input batch size
B.Increase the builder workspace size
C.Switch to a smaller base model
D.Disable INT8 calibration
AnswerB

The error specifically mentions workspace memory limit exceeded. By increasing the memory budget provided to the TensorRT builder, the engine can allocate sufficient scratch space to test various optimized kernels for the Attention and MatMul operations, thereby successfully completing the build process for the engine.

Why this answer

The error indicates that the workspace memory allocated to the TensorRT builder is insufficient to explore the search space of kernels for specific operations. Increasing the workspace size allows the builder to allocate larger temporary memory buffers, which is necessary for complex Transformer operations that require high-memory intermediate calculations to optimize effectively on NVIDIA hardware.

Exam trap

Candidates often assume the error is due to a lack of overall GPU VRAM, attempting to reduce model precision or batch size instead of specifically increasing the builder's workspace memory allocation.

174
MCQmedium

What is the primary risk of 'catastrophic forgetting' during the fine-tuning process?

A.The model becomes unable to produce responses in the desired language
B.The model loses the general knowledge acquired during pre-training
C.The model's inference speed increases significantly
D.The model begins to output only empty strings
AnswerB

Catastrophic forgetting refers to the phenomenon where a model's weights are modified so drastically that it loses its proficiency in tasks it was previously capable of. This happens when the fine-tuning loss prioritizes the new dataset too heavily, forcing the weights to shift away from the general knowledge base established during pre-training.

Why this answer

Catastrophic forgetting occurs when a model is updated on new, narrow tasks, causing it to lose the broad, generalized knowledge acquired during its extensive pre-training phase. In production, this renders the model useless for its original intended tasks. Mitigating this risk is crucial for businesses that need to maintain multi-purpose model capabilities while still achieving performance gains in specialized domains.

Exam trap

Candidates often confuse catastrophic forgetting with overfitting to the training set, missing that catastrophic forgetting specifically involves losing broad, generalized pre-trained capabilities.

175
Multi-Selecthard

Which TWO of the following practices are considered standard procedures for preparing a dataset for Instruction Fine-Tuning (IFT)? (Choose two)

Select 2 answers
A.Removing duplicate or highly redundant instructional examples
B.Applying aggressive token truncation to all training samples
C.Ensuring consistent schema and prompt formatting
D.Converting all text into numerical vectors before training
E.Increasing the learning rate by a factor of 100
AnswersA, C

Redundant data leads to overfitting and skewed model responses. Ensuring a clean, deduplicated dataset forces the model to learn the underlying logic of the instructions rather than memorizing specific sequences, which is essential for developing a robust model that can generalize to novel user prompts effectively during inference.

Why this answer

Instruction fine-tuning requires high-quality, diverse, and well-formatted data to ensure the model learns to follow specific user prompts. Removing duplicates prevents the model from over-relying on single examples, while consistent formatting (like ChatML or Alpaca format) ensures the model learns the structural cues for turn-based conversation, which is fundamental for effective performance in downstream instruction-following tasks.

Exam trap

Candidates often overlook data quality, assuming that simply increasing the volume of training examples is sufficient, while ignoring the negative impact of duplicates and inconsistent schema formatting on model convergence.

176
Multi-Selectmedium

A media company is preparing an NVIDIA NIM-hosted LLM that summarizes user-submitted articles. Counsel requires evidence that the system respects copyright and attribution obligations. Which two practices should the team implement? (Choose two.)

Select 2 answers
A.Add a system-prompt line that says the model must not reproduce copyrighted text verbatim.
B.Attach source metadata and a retrieval citation to each generated summary so downstream consumers can trace claims to the original article.
C.Log a hash of each source document alongside the prompt and response so the team can later prove which input produced a given summary.
D.Disable all logging so that copyrighted content is never retained on disk after summarization.
E.Increase the model's context window so it can ingest entire copyrighted works and produce longer summaries.
AnswersB, C

Citation and source metadata create a verifiable chain from generated text back to the original work, supporting attribution obligations and enabling downstream review. This is a concrete, auditable practice that directly addresses counsel's requirement for evidence of respect for copyright and attribution.

Why this answer

Copyright and attribution compliance in a summarization pipeline depends on provenance and auditability. Attaching citations and source metadata to each summary, and logging tamper-evident hashes that tie inputs to outputs, together create a defensible record. These practices let the company demonstrate respect for attribution and investigate disputes without retaining unnecessary copies of protected works.

Exam trap

The trap here is choosing logging suppression or prompt instructions as compliance evidence, when both actually weaken auditability or provide no verifiable record.

177
MCQmedium

When deploying an LLM using NVIDIA TensorRT-LLM, what is the primary benefit of pre-compiling the model into a TensorRT engine?

A.It allows the model to run on any generic CPU architecture.
B.It enables layer fusion and kernel selection for the target GPU.
C.It eliminates the need for any GPU memory during inference.
D.It automatically scales the model across multiple distributed nodes.
AnswerB

TensorRT-LLM optimizes the computational graph by fusing layers and selecting the most efficient kernels for the specific GPU architecture. This significantly reduces memory bandwidth consumption and increases computational throughput, which is essential for the high-performance requirements of modern generative AI models in production environments.

Why this answer

Pre-compiling with TensorRT-LLM allows for layer fusion, kernel auto-tuning, and memory optimization tailored specifically to the target GPU architecture. By performing these heavy optimizations offline, the inference engine can execute at peak performance immediately upon loading. This eliminates the runtime overhead associated with graph compilation or dynamic graph execution, ensuring minimal latency and optimal resource utilization from the very first inference request.

Exam trap

Candidates often think TensorRT-LLM just compresses the model, missing the critical benefit of offline graph optimization and kernel fusion that significantly reduces runtime latency during inference.

178
MCQhard

Refer to the exhibit. An engineer observes that memory utilization spikes during inference, causing OOM errors. Given the configuration, what is the most likely cause of the failure?

A.The count of 2 causes the GPU to oversubscribe its thermal limits.
B.The instance group count of 2 causes redundant loading of weights, exceeding VRAM.
C.Triton requires KIND_CPU for concurrent instance execution.
D.The gpus index [0] is invalid for multi-instance deployment.
AnswerB

Setting the instance count to two instructs Triton to create two independent model runners. Each runner requires its own memory allocation for weights and activations. If the model occupies a large portion of the GPU memory, running two instances simultaneously will inevitably exhaust the total available VRAM.

Why this answer

The instance group configuration defines two concurrent instances on the same GPU. Each instance attempts to load a separate copy of the model weights into the GPU memory. If the model size is large, doubling the instances exceeds the available VRAM capacity.

This configuration is a common mistake when deploying LLMs where model footprint is significant relative to total available device memory.

Exam trap

Candidates often blame the model size or the GPU hardware itself, failing to notice that the configuration defines multiple instances, which causes a multiplicative effect on VRAM consumption.

179
MCQhard

An engineer is optimizing prompts for an NVIDIA NIM-hosted model used in a multi-turn technical troubleshooting chat. The model forgets earlier constraints, such as the customer's environment and the product version, as the conversation grows. Which prompt engineering technique best preserves these constraints across turns?

A.Repeat the full conversation history verbatim in every turn and increase max_tokens.
B.Maintain a compact running summary of key constraints and inject it into the system prompt each turn.
C.Use a single-turn prompt for each message and rely on the model's pretrained knowledge of the product.
D.Lower the temperature to 0 and remove the system prompt to avoid conflicting instructions.
AnswerB

A running summary keeps essential facts such as environment and product version in a stable, high-priority position, so they survive across turns without unbounded token growth. Injecting it into the system prompt gives those constraints consistent influence over every response. This directly addresses forgetting while controlling context length in a multi-turn chat.

Why this answer

Multi-turn memory is best preserved by extracting and re-injecting critical constraints rather than replaying all dialogue. A compact running summary placed in the system prompt keeps key facts salient and stable while avoiding context bloat. Temperature and single-turn designs do not address memory, and verbatim history can dilute attention and waste tokens.

Exam trap

The trap here is thinking that more conversation history or deterministic sampling will fix forgetting, when the real fix is summarizing and re-prioritizing the constraints each turn.

180
MCQhard

A team is building a NeMo-based LLM pipeline and must tokenize a corpus that mixes English, Japanese, and Python source code. They plan to train a custom tokenizer with NVIDIA NeMo. Which tokenizer configuration best supports all three content types without excessive sequence length?

A.Byte-level Byte-Pair Encoding with a vocabulary of 64,000 to 128,000 tokens trained on a balanced sample of all three content types.
B.SentencePiece unigram tokenizer trained exclusively on Japanese text.
C.WordPiece tokenization trained only on the English portion of the corpus.
D.Byte-Pair Encoding with a vocabulary limited to 8,000 tokens.
AnswerA

Byte-level BPE avoids unknown tokens by operating on raw bytes, so any script or code symbol is representable. A large vocabulary of 64k-128k tokens captures common multilingual subwords and code identifiers, reducing sequence length. Training on a balanced sample ensures the merge rules reflect English, Japanese, and Python frequencies, which is exactly what this mixed corpus needs.

Why this answer

Byte-level BPE with a large vocabulary trained on a balanced multilingual and code sample handles arbitrary scripts and symbols without unknown tokens. It compresses common English words, Japanese subwords, and Python identifiers into single tokens, controlling sequence length. Training on all three distributions aligns merge statistics with actual usage, unlike English-only or Japanese-only training, and the 64k-128k range provides enough capacity for multilingual plus code coverage.

Exam trap

The trap here is choosing a tokenizer algorithm without checking whether its training data covers all content types, since out-of-vocabulary scripts and code inflate sequence length regardless of algorithm.

181
Multi-Selectmedium

A developer is optimizing a generative AI model for inference on NVIDIA GPUs. They want to reduce memory footprint and improve throughput without sacrificing accuracy. Which two techniques should they apply? (Choose two.)

Select 2 answers
A.Use NVIDIA TensorRT with INT8 quantization and calibration.
B.Use NVIDIA GPUDirect Storage to offload model weights to NVMe.
C.Increase the batch size to amortize memory overhead.
D.Enable NVIDIA Ampere TF32 precision for matrix multiplications.
E.Apply NVIDIA's structured sparsity with 2:4 pattern and TensorRT support.
AnswersA, E

TensorRT with INT8 quantization reduces memory footprint and increases throughput by using 8-bit integer operations, which are faster on Tensor Cores. Calibration ensures minimal accuracy loss by determining optimal scaling factors. This is a standard optimization for inference on NVIDIA GPUs.

Why this answer

INT8 quantization and structured sparsity both reduce memory footprint and improve throughput. INT8 quantization uses 8-bit integers, cutting memory usage by 4x compared to FP32, and Tensor Cores accelerate INT8 operations. Structured sparsity prunes weights in a 2:4 pattern, halving memory for weights and enabling sparse Tensor Core acceleration.

Both are supported by TensorRT and maintain accuracy with proper calibration and fine-tuning.

Exam trap

The trap here is confusing TF32 (which accelerates compute but not memory) with true memory-reducing techniques like INT8 and sparsity.

182
MCQmedium

You are fine-tuning a LLM on a large corpus of technical documentation. You notice the model struggles with domain-specific terminology that frequently appeared in the training data but was inconsistently formatted. Which data preparation technique best addresses this issue?

A.Increasing the batch size during the training phase.
B.Expanding the model's depth with additional transformer layers.
C.Applying consistent text normalization and domain-specific tokenization rules.
D.Reducing the learning rate to prevent overfitting on the noise.
AnswerC

Standardizing formatting and applying custom tokenization ensures that domain-specific terminology is tokenized consistently across the corpus. This alignment allows the model to build stronger semantic associations for technical terms, significantly reducing the probability of errors caused by variations in casing, punctuation, or special character usage.

Why this answer

Consistent normalization, such as lemmatization or standardizing case and special characters, ensures the tokenizer treats synonymous terms identically. In NVIDIA NeMo workflows, data quality directly impacts convergence speed and model accuracy. By reducing vocabulary noise during the preprocessing stage, the model can dedicate more capacity to learning semantic relationships rather than mapping variants of the same technical term to different embedding spaces, ultimately improving downstream performance on domain-specific benchmarks.

Exam trap

Candidates often focus on increasing model size or training epochs to fix terminology issues, ignoring that inconsistent tokenization and raw data noise are the root causes of poor performance.

183
MCQmedium

Refer to the exhibit. The model is encountering an OOM error during long-context processing. Which architectural adjustment is most appropriate to resolve this while maintaining context length?

A.Increase the number of hidden layers.
B.Switch to FlashAttention-2 kernels to optimize memory usage.
C.Reduce the embedding dimension size.
D.Disable the KV cache entirely.
AnswerB

FlashAttention-2 provides a fused kernel that computes attention in blocks, avoiding the storage of the full attention matrix in VRAM. This is the optimal way to handle long sequences like 128k, as it solves the memory bottleneck at the architectural kernel level without sacrificing model capability.

Why this answer

When sequence lengths scale to 128k, the attention matrix size grows quadratically, consuming massive memory. Implementing FlashAttention-2 or similar memory-efficient attention kernels is the industry-standard solution. These kernels optimize the memory layout and tile operations to compute attention without materializing the massive N×N matrix, drastically reducing peak memory usage and enabling the processing of very long sequences within existing GPU capacity.

Exam trap

Candidates often suggest reducing the batch size or model precision. While these help, they do not address the fundamental quadratic memory growth of attention that FlashAttention-2 is designed to solve.

184
MCQmedium

After fine-tuning a code-generation model with NVIDIA NeMo, an engineer notices the model now produces correct domain-specific function calls but has started emitting malformed JSON in about 15 percent of structured-output requests. The fine-tuning dataset contained no structured-output examples. Which evaluation action best explains and catches this regression?

A.Add a format-compliance check plus a general-capability regression suite to the evaluation, and compare against the base model on the same structured-output prompts.
B.Increase the temperature of the structured-output requests so the model has more freedom to produce valid JSON.
C.Conclude that the fine-tune succeeded on its target task and that the JSON failures are acceptable collateral given the domain gains.
D.Retrain the model with a larger learning rate so the structured-output behavior is reinforced more strongly during fine-tuning.
AnswerA

The symptom points to a capability regression outside the fine-tuning domain, so the evaluation must cover format validity and previously working skills, not just the target task. Running the same structured-output prompts against the base model establishes whether the fine-tune caused the breakage. A format checker converts the vague observation into a measurable pass rate.

Why this answer

Fine-tuning on a narrow dataset can degrade capabilities absent from that data, and structured output is a classic casualty. The right response is to measure it: add format-compliance scoring to the harness, rerun the same structured prompts against the base and tuned models, and include a general regression suite. Retraining harder, raising temperature, or dismissing the failures all skip the measurement step the situation demands.

Exam trap

The trap here is assuming a successful domain fine-tune cannot break unrelated behaviors, so the JSON failures get explained away instead of being measured against the base model.

185
Multi-Selectmedium

A team is building a customer service chatbot using NVIDIA NeMo Guardrails and NVIDIA NIM. They need to ensure compliance with ethical AI guidelines, specifically around transparency and user consent. Which two actions should they implement to meet these ethical requirements? (Choose two.)

Select 2 answers
A.Provide an option for users to request deletion of their conversation history and personal data.
B.Implement a NeMo Guardrails input rail that detects and blocks any user queries containing profanity.
C.Use NVIDIA TensorRT-LLM to optimize the model for faster response times.
D.Display a clear notice to users that they are interacting with an AI system and obtain explicit consent before processing personal data.
E.Log all user interactions with timestamps and store them indefinitely for audit purposes.
AnswersA, D

Offering data deletion supports user autonomy and consent, which are ethical AI principles. It allows users to control their personal data and aligns with regulations like GDPR's right to erasure. This action demonstrates transparency and respect for user rights, directly fulfilling the ethical requirement to obtain and honor consent.

Why this answer

Ethical AI guidelines around transparency and consent require that users are informed they are interacting with an AI and that they can control their personal data. Displaying a clear notice and obtaining consent, as well as providing a way to delete data, directly fulfill these principles. Content moderation, indefinite logging, and performance optimization do not address these specific ethical requirements.

Exam trap

The trap here is confusing content moderation or performance optimization with ethical transparency and consent requirements.

186
MCQmedium

A developer is using an NVIDIA NIM-hosted model to classify support tickets into a fixed set of categories. The model occasionally invents new category names. The team wants to guarantee that only allowed categories are returned. Which approach is most appropriate?

A.Provide the full category list in the prompt and use the NIM guided_choice parameter with those categories.
B.Use few-shot examples of each category and set top_k to 100.
C.Ask the model to explain its reasoning before choosing a category.
D.Add 'Choose from the list' to the prompt and set temperature to 1.0.
AnswerA

Listing the categories gives the model the allowed set, and guided_choice constrains decoding to exactly one of those strings. This eliminates invented labels at the sampling level, guaranteeing compliance. It is the most reliable method when the output must be one of a fixed enumeration.

Why this answer

When output must be one of a fixed set, enumerating the categories in the prompt and applying a decoding constraint such as guided_choice ensures only valid labels are emitted. Few-shot examples and reasoning improve quality but do not enforce the boundary. High temperature and broad top_k work against the requirement by increasing variability.

Exam trap

The trap here is relying on prompt wording or examples to restrict categories, when only an explicit enumeration combined with constrained decoding guarantees the model cannot invent a label.

187
MCQhard

Refer to the exhibit. An engineer observes that a model fine-tuned with this LoRA configuration is failing to converge on a highly complex legal document domain. What is the most likely cause of this issue?

A.The dropout value is set too high, causing the model to underfit.
B.The target modules are incorrectly specified, leading to a loss of attention.
C.The rank value is too low to represent the complex patterns of the domain.
D.The alpha parameter is too high, causing gradient explosion.
AnswerC

For complex domains like legal or technical writing, a rank of 8 may be insufficient to capture the necessary parameter updates. Increasing the rank allows the model to capture a richer set of features, providing the additional capacity needed to learn complex domain-specific linguistic relationships and improve convergence.

Why this answer

The rank of 8 is likely too low for capturing the intricate, nuanced patterns required for legal documentation. While low-rank adaptation is efficient, complex domains often necessitate higher rank values to provide enough expressive capacity within the trainable adapter layers. This configuration mismatch limits the model's ability to learn the specific syntactic and semantic structures inherent in specialized legal texts, leading to poor convergence and inadequate performance.

Exam trap

Candidates often assume the failure stems from the learning rate or data quality, overlooking the LoRA rank parameter, which directly dictates the model's capacity to learn complex, domain-specific semantic patterns.

188
MCQhard

Which of the following describes the 'Chain-of-Verification' (CoVe) prompting technique?

A.A method to verify that the GPU is running at full capacity.
B.A process where the model checks its own claims for factual consistency.
C.A technique to optimize the model's weights during training.
D.A protocol for securing the prompt against injection attacks.
AnswerB

CoVe explicitly mandates that the model critiques its own output. By drafting verification questions and answering them, the model can identify and correct errors in its initial draft. This self-correction loop is a powerful tool for improving the truthfulness and reliability of complex LLM-generated reports in enterprise environments.

Why this answer

Chain-of-Verification is a sophisticated technique designed to reduce hallucinations. It prompts the model to generate a response, then draft questions to verify its own claims, answer those questions, and finally revise the original response based on the verification. This is highly effective for NVIDIA developers building high-stakes applications where factual accuracy is non-negotiable and automated auditing is required.

Exam trap

Candidates often conflate CoVe with standard CoT, failing to realize that CoVe is specifically a post-generation verification loop aimed at fact-checking, rather than just a step-by-step reasoning process.

189
Multi-Selecthard

An engineer is reviewing the architecture of a decoder-only LLM that must support very long input contexts for document analysis. They are considering architectural choices that extend effective context length beyond what the model saw during pretraining. Which TWO techniques are designed specifically to extend usable context length without retraining the entire model from scratch? (Choose two.)

Select 2 answers
A.Switching the feed-forward activation from SwiGLU to GeLU to improve long-range dependency modeling.
B.Position interpolation, which rescales positional indices so longer sequences fall within the range seen during pretraining.
C.Increasing the number of attention heads while keeping the head dimension fixed.
D.YaRN scaling, which adjusts rotary positional embedding frequencies to improve extrapolation to longer sequences.
E.Reducing the vocabulary size to shorten the embedding matrix and free memory for longer sequences.
AnswersB, D

Position interpolation compresses the positional index range so that a longer sequence maps into the positions the model encountered during pretraining. This allows the model to generalize to longer contexts with only brief fine-tuning. It directly targets the mismatch between pretraining context length and desired inference length, making it a standard context-extension technique.

Why this answer

Position interpolation and YaRN both operate on positional encoding to reconcile desired sequence lengths with the range seen during pretraining. Interpolation compresses indices, while YaRN adjusts rotary frequencies to improve extrapolation. Both are context-extension methods that require only light fine-tuning, unlike changes to head count, MLP activation, or vocabulary size, which do not address positional generalization.

Exam trap

The trap here is treating any memory-saving or capacity change as context extension, when only techniques that alter positional encoding behavior actually lengthen usable context.

190
MCQmedium

A developer is using NVIDIA TensorRT to optimize a BERT-based model for inference. They notice that the engine performs poorly on variable-length input sequences because it was built with a single optimization profile for a fixed sequence length. What should they do to improve performance across different sequence lengths?

A.Pad all input sequences to the maximum length supported by the model.
B.Rebuild the engine with multiple optimization profiles covering the range of sequence lengths.
C.Use CUDA graphs to capture and replay the inference execution.
D.Enable FP16 precision to reduce memory usage and speed up kernels.
AnswerB

TensorRT optimization profiles define the min, opt, and max shapes for dynamic input dimensions. Building multiple profiles allows the engine to select the best kernel configurations for different sequence lengths, improving performance across the range. This is the recommended approach for variable-length inputs.

Why this answer

TensorRT optimization profiles allow the engine to tune kernels for specific input shape ranges. When input lengths vary, a single profile cannot cover all cases efficiently. Creating multiple profiles, each with appropriate min/opt/max dimensions, enables the engine to choose the best kernels for each length, improving overall performance.

Exam trap

The trap here is thinking that precision reduction or launch overhead reduction can compensate for a mismatched optimization profile, when the core issue is kernel selection for varying shapes.

191
MCQhard

Refer to the exhibit. In the context of a distributed multi-GPU fine-tuning job, what is the most likely cause of this error?

A.The model weights are too large for the GPU memory
B.The MASTER_ADDR or MASTER_PORT environment variables are misconfigured
C.The learning rate is too high, causing gradient explosion
D.The dataset is missing required training samples
AnswerB

Distributed training frameworks rely on these variables to establish a communication channel. If the address or port is blocked, inaccessible, or incorrect, the NCCL library cannot handshake between nodes, leading to the connection refused error. Correcting these settings is the standard solution for resolving distributed networking issues in NVIDIA environments.

Why this answer

This error indicates that the NCCL (NVIDIA Collective Communications Library) is failing to communicate between GPU nodes or processes. It is typically caused by a misconfiguration of the environment variables (like MASTER_ADDR or MASTER_PORT) or a firewall blocking the necessary ports. In distributed training, these network configurations are essential for synchronizing gradients across all participating GPUs during the training process.

Exam trap

Candidates frequently mistake NCCL communication timeouts or address errors for insufficient GPU VRAM, overlooking network-level environment variables like MASTER_ADDR and MASTER_PORT required for multi-node synchronization.

192
MCQeasy

You are evaluating a generative AI model for a chatbot that must adhere to strict safety guidelines. The model occasionally generates toxic or biased responses. You need to automatically evaluate the model's outputs for toxicity and bias before deployment. Which NVIDIA tool or framework is specifically designed for this purpose?

A.NVIDIA Triton Inference Server
B.NVIDIA DALI
C.NVIDIA TensorRT
D.NVIDIA NeMo Guardrails
AnswerD

NVIDIA NeMo Guardrails is a toolkit for adding programmable guardrails to LLM-based conversational systems. It can be configured to detect and block toxic or biased outputs using predefined or custom policies. It integrates with models to enforce safety guidelines in real time, making it the appropriate tool for automatically evaluating and mitigating toxicity and bias in chatbot responses.

Why this answer

NeMo Guardrails is designed to enforce safety and content policies in conversational AI, including toxicity and bias detection. The other options are infrastructure or optimization tools without safety evaluation features. Thus, NeMo Guardrails is the correct choice for automatically evaluating and mitigating toxic or biased responses before deployment.

Exam trap

The trap here is assuming that any NVIDIA tool that handles models can evaluate safety, when only NeMo Guardrails provides programmable guardrails for content moderation.

193
MCQmedium

Which prompt engineering strategy helps the model maintain focus when processing an extremely long document within a single context window?

A.Always increase the model's temperature to 1.0.
B.Instruct the model to analyze the document in sections and summarize each.
C.Remove all system-level instructions to save tokens.
D.Limit the context window size to 1024 tokens.
AnswerB

Breaking down a large document into segments for analysis ensures that the model provides equal attention to all parts. By summarizing each section before synthesis, you force the model to retain the key information from throughout the text, effectively mitigating the risk of ignoring information in the middle.

Why this answer

Long context windows can lead to the 'Lost in the Middle' phenomenon, where the model performs better on the beginning and end of the document but ignores the middle. Using 'attention-focusing' instructions or prompting the model to summarize segments before synthesizing a final answer helps maintain performance across the entire document. This is critical for NVIDIA engineers working with large-scale technical whitepapers and documentation.

Exam trap

Many candidates suggest increasing the context window size further, overlooking the 'Lost in the Middle' phenomenon where long documents are ignored.

194
MCQeasy

A developer is preparing a dataset for instruction fine-tuning of an LLM using NVIDIA NeMo. The raw data consists of customer support transcripts with speaker labels and timestamps. Which preprocessing step is most important before training?

A.Convert each transcript into a structured instruction-response pair with a clear prompt and a target completion.
B.Duplicate each transcript so the model sees every example at least twice per epoch.
C.Remove all punctuation and capitalization so the model learns a consistent lowercase style.
D.Increase the learning rate to compensate for the noisy transcript data.
AnswerA

Instruction fine-tuning expects examples in a prompt-completion or instruction-response format so the loss is computed on the desired assistant output. Raw transcripts with speaker labels and timestamps do not directly teach the model how to respond to instructions. Converting them into structured pairs aligns the data with the training objective and yields useful gradients.

Why this answer

Instruction fine-tuning requires examples that pair a prompt with the desired response so the model learns to follow instructions. Raw support transcripts contain speaker labels and timestamps that are not part of the target behavior. Converting each transcript into a structured instruction-response pair makes the data compatible with the training objective and ensures the loss is computed on the correct tokens.

Exam trap

The trap here is treating preprocessing as a hyperparameter or volume problem, when the real issue is reformatting raw transcripts into instruction-response pairs.

195
MCQhard

A team is evaluating a large language model (LLM) deployed with NVIDIA Triton Inference Server for a real-time question-answering system. They observe that the model's responses are accurate but latency spikes during peak load, causing timeouts. They need to evaluate the model's performance under high concurrency to identify the maximum throughput while maintaining a 95th percentile latency below 200 ms. Which evaluation approach should they use?

A.Measure GPU utilization with NVIDIA Management Library (NVML) during a single long-running inference.
B.Run a one-off inference with a batch size of 1 and measure latency.
C.Compute the model's FLOPS and compare against the GPU's peak theoretical FLOPS.
D.Use NVIDIA Triton's Performance Analyzer with concurrency sweep and latency constraints.
AnswerD

NVIDIA Triton's Performance Analyzer can simulate concurrent requests by sweeping concurrency levels and measuring latency percentiles. It supports setting a latency constraint (e.g., 95th percentile < 200 ms) and automatically finds the maximum throughput that satisfies it. This directly addresses the need to evaluate performance under high concurrency and identify the optimal operating point.

Why this answer

Evaluating performance under high concurrency requires simulating multiple simultaneous requests and measuring latency percentiles. NVIDIA Triton's Performance Analyzer is designed for this, allowing concurrency sweeps and latency constraints to find the maximum throughput meeting the SLA. Other options either lack concurrency or measure irrelevant metrics, so they cannot identify the operating point that satisfies the 95th percentile latency requirement.

Exam trap

The trap here is confusing theoretical compute limits or single-request latency with actual concurrent performance, which requires load testing with multiple simultaneous requests.

196
MCQmedium

An organization is deploying a high-throughput LLM on NVIDIA Triton Inference Server. They observe significant tail latency spikes when serving multiple concurrent requests. Which strategy most effectively optimizes GPU utilization and reduces latency jitter for these concurrent model instances?

A.Increase the number of CPU threads per model instance in the configuration file.
B.Disable all model instances to ensure requests are processed in serial order.
C.Enable dynamic batching with an optimized max_queue_delay_microseconds setting.
D.Switch the model to a lower precision format like FP8 without model recalibration.
AnswerC

Dynamic batching allows Triton to accumulate requests over a specified window to create larger, more efficient tensor operations. Fine-tuning the queue delay ensures that the server waits just long enough to fill a batch without unnecessarily delaying individual requests, directly mitigating tail latency spikes during peak traffic.

Why this answer

Dynamic Batching is the primary mechanism in Triton to combine individual inference requests into a single batch, significantly improving throughput while minimizing latency. By configuring the 'max_queue_delay_microseconds' parameter, the system balances wait times with compute efficiency. This is critical for enterprise deployments where maximizing GPU hardware investment while maintaining strict service-level agreements is the standard requirement for production-grade generative AI applications.

Exam trap

Candidates often recommend simply scaling up GPU count or adjusting thread counts, overlooking Triton's native queue management mechanisms designed explicitly to control latency jitter.

197
Multi-Selectmedium

You are preparing a large instruction-tuning dataset for an LLM using NVIDIA NeMo. You need to ensure the dataset supports efficient training and evaluation. Which TWO steps are essential in the data preparation pipeline? (Choose two.)

Select 2 answers
A.Convert the dataset into a format compatible with NeMo, such as JSONL with input and output fields
B.Train a tokenizer from scratch on the dataset
C.Apply data augmentation using back-translation
D.Remove all punctuation from the text
E.Split the dataset into training, validation, and test sets
AnswersA, E

NeMo's instruction-tuning pipeline expects data in a structured format, commonly JSONL with fields like input and output or prompt and completion. Converting the raw dataset into this format ensures the data loader can parse and batch examples correctly. Without this step, training would fail or require custom parsing code, so it is essential.

Why this answer

Formatting the dataset into a NeMo-compatible structure such as JSONL with input and output fields ensures the data loader can consume it, while splitting into training, validation, and test sets enables proper model evaluation and hyperparameter tuning. These two steps are foundational; other activities like tokenizer training or augmentation are optional and context-dependent.

Exam trap

The trap here is treating optional enhancements like tokenizer retraining or augmentation as mandatory, when the truly essential steps are formatting and splitting.

198
MCQmedium

Which of the following describes the purpose of 'gradient accumulation' in fine-tuning scenarios?

A.To increase the speed of the training process by bypassing the GPU's bottleneck.
B.To enable training with larger effective batch sizes on memory-constrained hardware.
C.To automatically optimize the learning rate based on the model's gradient magnitude.
D.To reduce the number of parameters being fine-tuned in the model.
AnswerB

This technique allows engineers to maintain effective batch sizes that would otherwise exceed available VRAM. By accumulating gradients over multiple steps and updating weights only after reaching a target batch size, the training process achieves the stability of large-batch learning while keeping the peak memory usage manageable.

Why this answer

Gradient accumulation allows for simulating larger batch sizes when GPU VRAM is limited. By performing multiple forward and backward passes without updating the weights, the model can aggregate gradients across several smaller steps. This is crucial when working on hardware with restricted memory, as it enables the model to benefit from the statistical stability of larger batches, improving the overall quality of the fine-tuned model's convergence.

Exam trap

Candidates confuse gradient accumulation with model parallelism, failing to realize it is specifically a memory-saving technique that simulates larger batches without requiring additional GPU memory for simultaneous activation storage.

199
MCQmedium

A team is deploying a 13B-parameter LLM with NVIDIA TensorRT-LLM on a single A100 80GB GPU. They want to reduce GPU memory usage during inference without retraining the model, while keeping acceptable output quality. Which technique should they apply?

A.Enable INT8 weight-only quantization using TensorRT-LLM's quantization toolkit.
B.Reduce the maximum sequence length to 512 tokens to lower activation memory.
C.Increase the KV cache block size to 128 tokens to reduce memory fragmentation.
D.Enable Tensor Parallelism across two GPUs to split the model.
AnswerA

INT8 weight-only quantization compresses model weights to 8-bit integers while keeping activations in higher precision, cutting weight memory roughly in half with minimal quality loss. TensorRT-LLM supports this via its quantization toolkit and calibration workflow, making it a practical post-training approach for a 13B model on a single 80GB GPU without retraining.

Why this answer

INT8 weight-only quantization is a post-training technique supported by TensorRT-LLM that reduces weight memory by roughly half while preserving output quality. It directly addresses the goal of lowering GPU memory usage for a large model on a single GPU without retraining. Other listed changes affect scheduling or activation memory rather than the core weight footprint.

Exam trap

The trap here is assuming that KV cache tuning or sequence-length reduction will solve weight-dominated memory pressure for a large model.

200
MCQeasy

An AI engineer is deploying a large language model using NVIDIA Triton Inference Server. They need to ensure that the server can handle multiple concurrent requests efficiently while maintaining low latency. Which Triton feature allows the server to dynamically batch incoming requests to maximize GPU utilization?

A.Dynamic batching
B.Instance groups
C.Model versioning
D.Ensemble scheduling
AnswerA

Dynamic batching in Triton automatically groups individual inference requests into batches on the server side, within a configurable time window. This increases throughput and GPU utilization while keeping latency low for high-concurrency scenarios. It is specifically designed to handle multiple concurrent requests efficiently without client-side batching.

Why this answer

Dynamic batching is a Triton feature that aggregates concurrent inference requests into larger batches on the server side, improving GPU utilization and throughput while managing latency. It is the standard mechanism for handling high concurrency efficiently. Model versioning, instance groups, and ensemble scheduling serve different purposes and do not provide dynamic request batching.

Exam trap

The trap here is confusing dynamic batching with instance groups, which also affect concurrency but through multiple model instances rather than request aggregation.

201
Multi-Selectmedium

Which TWO factors should be prioritized when selecting a quantization strategy for deploying a large language model on constrained edge hardware? (Select TWO)

Select 2 answers
A.Maximizing the model training loss to improve generalization.
B.Reducing the memory footprint of the model weights.
C.Increasing the number of neural network layers in the architecture.
D.Improving inference throughput via reduced bit-precision arithmetic.
E.Replacing the Transformer architecture with a linear regression model.
AnswersB, D

On edge devices, VRAM is severely limited. Quantization reduces the bit-depth of weights, directly lowering the memory requirement. This allows larger models to fit into the limited VRAM of edge hardware, which is critical for enabling complex LLM inference tasks that would otherwise fail to load.

Why this answer

Selecting a quantization strategy requires balancing precision loss against hardware performance gains. When deploying to edge devices, memory bandwidth and storage capacity are typically the primary bottlenecks. By reducing precision from FP16 or FP32 to INT8 or FP8, you directly decrease the model's memory footprint and increase the number of operations per clock cycle, which is essential for maintaining acceptable real-time inference speeds on resource-limited hardware.

Exam trap

Candidates often confuse hardware training constraints with edge deployment limitations, incorrectly prioritizing compute scaling factors instead of focusing strictly on memory footprint and memory bandwidth bottlenecks inherent to edge devices.

202
MCQhard

A hospital's AI governance board is reviewing an LLM triage assistant built on NVIDIA NIM. They want an ongoing, automated mechanism that flags when model outputs drift toward unsafe clinical recommendations across thousands of daily conversations, without reviewing every transcript manually. Which approach best fits this need?

A.Deploy NVIDIA NeMo Guardrails with output rails that classify responses against a clinical-safety policy, and emit metrics to a monitoring dashboard for threshold alerts.
B.Fine-tune the base model weekly on the most recent transcripts so unsafe recommendations naturally decrease over time.
C.Increase the model's temperature setting so the assistant produces more varied responses that are easier to spot during spot checks.
D.Require clinicians to sign off on every AI-generated recommendation before it reaches a patient, eliminating the need for automated drift detection.
AnswerA

Output rails evaluate generated responses against defined policies in real time and can produce structured signals. Routing those signals into monitoring and alerting gives the board continuous, automated visibility into unsafe clinical drift without manual transcript review, which directly matches the requirement for scalable oversight.

Why this answer

Continuous safety oversight at scale requires machine-evaluable signals rather than manual review or model retraining. Output rails in NeMo Guardrails can classify responses against a clinical policy and emit telemetry, which feeds dashboards and alerts. This gives the governance board an automated, auditable early-warning system for unsafe drift across large volumes of conversations.

Exam trap

The trap here is conflating a mitigation control such as human sign-off or fine-tuning with an automated monitoring and alerting mechanism.

203
MCQmedium

Which of the following describes the purpose of a 'System Prompt' in Instruction Fine-Tuning?

A.To increase the number of tokens processed in the output
B.To act as a persistent instruction defining the model's persona
C.To replace the need for domain-specific fine-tuning entirely
D.To compress the training dataset size
AnswerB

The system prompt acts as a foundational instruction that dictates the model's behavior, tone, and constraints. It provides the necessary context for the assistant to follow during multi-turn conversations, ensuring that the model remains aligned with its intended role and follows specified safety and quality guidelines throughout the interaction.

Why this answer

The system prompt provides a high-level instruction or persona for the model, which guides its overall behavior throughout the conversation. It sets expectations for tone, safety constraints, and task-specific roles. Effectively designing the system prompt is essential for ensuring that the fine-tuned model consistently adheres to the required persona or operational boundaries in real-world production deployments.

Exam trap

Candidates often confuse system prompts with few-shot examples or model weights. Remember that a system prompt is a high-level, persistent instruction that guides overall behavior, tone, and boundaries throughout a conversation.

204
MCQmedium

You are preparing a dataset for instruction fine-tuning an LLM using NVIDIA NeMo. The dataset contains pairs of instructions and responses, but you notice that some responses are significantly longer than others, and a few are extremely short (e.g., 'Yes' or 'No'). You want to ensure the model learns to generate appropriate-length responses. Which data preparation technique is most effective?

A.Augment the dataset by duplicating short responses to increase their frequency.
B.Filter out all responses shorter than 10 tokens to remove trivial answers.
C.Ensure the dataset has a balanced distribution of response lengths by sampling or resampling.
D.Include a length token or metadata in the input to indicate the desired response length.
AnswerC

Balancing response lengths helps the model learn to generate appropriate lengths based on the instruction. By avoiding overrepresentation of very short or very long responses, the model can better generalize. This can be done by stratified sampling or by capping extreme lengths, but maintaining a natural distribution is key.

Why this answer

A balanced distribution of response lengths allows the model to learn when to generate concise versus detailed answers. Overrepresentation of short responses may lead to truncation, while too many long responses can cause verbosity. Resampling to achieve a more uniform distribution helps the model internalize the relationship between instruction complexity and response length.

Exam trap

The trap here is thinking that removing short responses solves the issue, but it actually biases the model and removes valid training signals.

205
MCQhard

A team fine-tunes a model with NVIDIA NeMo using a packed sequence dataset and notices that some training samples contain several short conversations concatenated. They must ensure the loss is computed only on assistant responses and not on the packed boundaries. Which configuration detail should they verify?

A.That the global batch size is increased to compensate for the additional tokens in each packed sequence
B.That the tokenizer vocabulary is expanded to include special separators between packed samples
C.That the loss mask aligns with each sample's response tokens and that packed sequences are separated by an attention boundary
D.That gradient accumulation steps equal the number of samples packed into each sequence
AnswerC

Packed sequences concatenate multiple samples into one training sequence for efficiency, so the loss mask must mark only assistant response tokens and the attention mechanism must prevent cross-sample attention. In NeMo this is handled through per-token loss masks and sequence boundary handling. Verifying both ensures the model is not trained on padding or on tokens from adjacent samples.

Why this answer

Packed sequences boost training efficiency by filling each context window with multiple samples, but correctness depends on the loss mask covering only assistant response tokens and on attention being blocked across sample boundaries. Batch size, vocabulary changes, and gradient accumulation do not affect these two requirements.

Exam trap

The trap here is treating packed sequences as a pure throughput optimization and overlooking that masks and attention boundaries must be adjusted to keep the loss correct.

206
Multi-Selecthard

A platform team is deploying a 70B-parameter LLM with NVIDIA TensorRT-LLM across four 80 GB H100 GPUs and needs to serve long-context requests efficiently. They are deciding how to combine parallelism and memory techniques in the build and runtime configuration. (Choose two.)

Select 2 answers
A.Use tensor parallelism across the four GPUs so each GPU holds a shard of every layer's weights
B.Configure the paged KV cache with a block size and max sequence length sized for the target context window
C.Reduce the number of GPUs to one and enable FP8 quantization to fit the model on a single device
D.Rely on CUDA Unified Memory to transparently page weights and KV cache between host and device
E.Enable pipeline parallelism with a single micro-batch to minimize inter-GPU communication
AnswersA, B

Tensor parallelism splits each layer's weight matrices across GPUs, allowing a 70B model that would not fit on one 80 GB H100 to be served across four. With NVLink-connected H100s, the all-reduce communication overhead is manageable, and it is the standard way to serve very large models with low latency.

Why this answer

Tensor parallelism shards each layer across the four H100s so a 70B model fits and serves with low latency over NVLink, while a properly sized paged KV cache prevents memory over-reservation for long contexts. Together they address both weight distribution and the dominant dynamic memory consumer in long-context LLM serving.

Exam trap

The trap here is treating pipeline parallelism or Unified Memory as easy wins, when pipeline bubbles and host-device paging both hurt latency-sensitive long-context serving.

207
Multi-Selecthard

An LLM inference service deployed on NVIDIA Triton Inference Server is experiencing occasional failures under high load. The team wants to implement proactive monitoring to predict and prevent these failures. Which two metrics should be prioritized for early detection of potential issues? (Choose two.)

Select 2 answers
A.Disk I/O throughput
B.CPU utilization of the Triton server process
C.Number of active connections
D.GPU memory utilization per instance
E.Request queue time
AnswersD, E

GPU memory utilization is critical because LLMs are memory-intensive. Rising memory usage can indicate memory leaks or increased batch sizes, leading to out-of-memory errors. Monitoring this metric allows proactive scaling or optimization before failures occur. It directly relates to resource exhaustion, a common cause of inference failures under load.

Why this answer

Under high load, LLM inference failures often stem from resource exhaustion or overload. GPU memory utilization is critical because insufficient memory leads to out-of-memory errors. Request queue time indicates if the system is falling behind, predicting timeouts.

Together, these metrics provide early warning of impending failures, enabling proactive measures.

Exam trap

The trap here is focusing on general system metrics like CPU or disk I/O, which are less relevant for GPU-bound LLM inference under load.

208
MCQhard

An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. They want to maximize throughput for a chatbot workload with variable-length inputs and outputs. Which of the following techniques should they implement to achieve the highest throughput while maintaining acceptable latency?

A.Enable in-flight batching (continuous batching) and use paged KV cache with block-based memory management.
B.Use static batching with a fixed batch size of 32 and pad all sequences to the maximum length.
C.Use FP32 precision to ensure accuracy and rely on CUDA graphs to reduce launch overhead.
D.Increase the number of GPUs and use tensor parallelism with a fixed batch size per GPU.
AnswerA

In-flight batching dynamically adds and removes sequences at each iteration, eliminating padding and improving GPU utilization. Paged KV cache stores attention keys and values in non-contiguous blocks, reducing memory fragmentation and allowing more concurrent sequences. Together, they maximize throughput for variable-length chatbot traffic while keeping latency low, as implemented in TensorRT-LLM.

Why this answer

In-flight batching and paged KV cache are key optimizations in TensorRT-LLM for LLM serving. In-flight batching allows sequences to join and leave the batch at each decoding step, avoiding padding and keeping the GPU busy. Paged KV cache manages memory in blocks, enabling more sequences to run concurrently.

This combination delivers the highest throughput for variable-length chatbot workloads.

Exam trap

The trap here is assuming that increasing batch size or GPUs alone will maximize throughput, when dynamic batching and memory management are the real enablers.

209
MCQhard

You are preparing a dataset for pretraining an LLM using NVIDIA NeMo's Megatron-LM. The dataset consists of JSONL files where each line contains a 'text' field. You need to convert these files into the binary format required by Megatron for efficient training. Which tool or method should you use to perform this conversion?

A.Use Apache Arrow to serialize the JSONL data into a columnar format and load it with a custom data loader.
B.Use the `jsonl_to_bin` utility provided by the Hugging Face Transformers library.
C.Use the `nemo_curator` library to convert JSONL to TFRecord format, then train directly on TFRecords.
D.Use the `preprocess_data.py` script from the Megatron repository to tokenize and create indexed binary files.
AnswerD

The `preprocess_data.py` script in Megatron-LM is specifically designed to tokenize JSONL text data and produce binary files with index mapping, which are required for efficient data loading during pretraining. It handles tokenization, concatenation, and sharding, making it the standard tool for this conversion.

Why this answer

Megatron-LM provides the `preprocess_data.py` script specifically for converting JSONL text data into indexed binary files that its data loader can read efficiently. This script handles tokenization, document concatenation, and sharding, ensuring compatibility with Megatron's training pipeline. It is the recommended and standard method for preparing data for Megatron-based pretraining.

Exam trap

The trap here is assuming that any serialization format like TFRecord or Arrow will work, but Megatron requires its own binary format produced by its preprocessing script.

210
MCQmedium

An enterprise deployment of NeMo Guardrails is experiencing hallucinations where the model provides medical advice despite strict system prompts. What is the most effective approach to mitigate this risk?

A.Increase the temperature parameter of the LLM to provide more creative, diverse outputs.
B.Implement NeMo Guardrails 'dialogue rails' to detect and redirect queries related to medical diagnosis.
C.Retrain the entire foundation model on a curated dataset of medical textbooks.
D.Change the model architecture to a smaller parameter size to reduce knowledge density.
AnswerB

Dialogue rails act as a middleware layer that inspects the interaction flow. By explicitly identifying medical diagnosis intents, the system can trigger a predefined flow that refuses to answer or directs the user to a qualified human professional, ensuring the model remains within its safe operational domain.

Why this answer

NeMo Guardrails provides a structured way to intercept model input and output to enforce safety boundaries. By defining specific 'rails' that detect non-compliant topics, the system can pivot the conversation or block the response entirely. This mechanism is critical in high-stakes environments where adherence to strict safety protocols is mandatory to prevent liability and ensure that generative AI tools do not cross into domains requiring human expertise.

Exam trap

Candidates often suggest fine-tuning the model to 'fix' hallucinations. Fine-tuning is expensive and unreliable for enforcing safety boundaries compared to the deterministic control provided by dialogue rails.

211
MCQeasy

In the context of NVIDIA Tensor Cores, what is the primary benefit of using BF16 (Bfloat16) over FP16 during model training and inference?

A.BF16 provides double the precision of FP16.
B.BF16 offers a larger dynamic range for gradients.
C.BF16 requires significantly less memory than FP16.
D.BF16 is faster on non-NVIDIA hardware.
AnswerB

The 8-bit exponent in BF16 matches the dynamic range of FP32, allowing it to represent a much wider range of values than FP16. This prevents underflow and overflow issues during deep learning operations, leading to more stable model convergence without requiring complex loss scaling techniques.

Why this answer

BF16 uses the same exponent range as FP32, which prevents overflow issues commonly encountered in deep learning training when using FP16. This increased dynamic range makes it more robust for gradient calculations and weight updates. By providing a wider range while maintaining the same performance advantages of half-precision, BF16 has become the industry standard for stabilizing training and inference of modern large language models.

Exam trap

Candidates mistakenly believe BF16 provides higher precision than FP16, confusing the mantissa size with the exponent range benefits.

212
MCQmedium

An enterprise is preparing a massive multi-terabyte corpus of specialized PDF documents containing technical schematics and dense tables for training a domain-specific LLM using NeMo Curator. During the document extraction pipeline, you notice that tabular data layouts are scrambled into unstructured string tokens, destroying spatial relationships. Which data preparation strategy should you implement first within NeMo Curator to preserve tabular integrity before tokenization?

A.Apply aggressive regex-based whitespace removal across all extracted string buffers to normalize sentence boundaries prior to embedding generation.
B.Deploy NeMo Curator's layout-aware PDF extraction modules featuring visual bounding box detection to parse tables into structured Markdown formatting.
C.Increase the chunk size parameter in the tokenizer configuration so that entire table blocks fit inside a single oversized context window.
D.Convert all PDF pages into low-resolution JPEG images to bypass text extraction errors and feed raw pixels directly into a text-only causal language model.
AnswerB

Layout-aware extraction uses visual bounding box detection to identify table cells and their spatial relationships, emitting structured Markdown rather than scrambled token strings. This preserves row and column integrity before tokenisation, directly addressing the destroyed tabular layouts described in the stem.

Why this answer

NeMo Curator provides specialized PDF extraction utilities that leverage advanced computer vision and layout parsing models to accurately identify bounding boxes for tables and figures. Preserving tabular markdown structures prevents spatial collapse, ensuring downstream tokenizers capture relational data accurately. This step is critical in domain-specific LLM training because scrambled tables introduce severe noise that degrades reasoning capabilities.

Exam trap

Candidates often assume standard text extraction tools are sufficient for all PDFs, overlooking that raw text extraction strips structural bounding box metadata essential for tabular layout preservation.

213
MCQhard

A global bank uses NVIDIA NeMo Guardrails in front of an LLM assistant that answers employee HR questions. Legal requires that the assistant refuse any request that could constitute unauthorized legal advice, even when the request is phrased indirectly. During testing, a prompt such as 'My manager wants to know if we can terminate someone for discussing pay' bypasses the existing rail. What is the most effective configuration change to close this gap?

A.Enable output moderation to filter responses that contain legal terminology after generation.
B.Lower the model temperature to reduce creative paraphrasing of restricted topics.
C.Add a semantic intent-matching rail with representative indirect phrasings and a canonical refusal response for unauthorized legal advice.
D.Increase the context window so the assistant can see more of the conversation history before deciding.
AnswerC

Semantic intent matching generalizes beyond exact keywords, so representative indirect phrasings teach the rail to recognize the underlying request for legal advice. Pairing that with a canonical refusal ensures the assistant consistently declines regardless of surface wording, which is exactly what closing this bypass requires.

Why this answer

Indirect requests evade keyword-based rails because the restricted intent is expressed through context rather than explicit terms. A semantic intent-matching rail trained on representative indirect phrasings, combined with a canonical refusal, generalizes to new paraphrases and enforces the policy at input time. Temperature, context size, and output filtering do not address the underlying intent-classification gap.

Exam trap

The trap here is treating a bypass as a model-behavior problem solvable with temperature or context tuning, when it is actually an input intent-classification gap.

214
MCQeasy

A team is preparing a customer-support chat dataset for instruction fine-tuning of an LLM. The raw data contains HTML tags, inconsistent date formats, and emoji. Which data preparation step should be performed first to make the text usable for tokenization?

A.Normalize and clean the text
B.Tokenize the text with the model's tokenizer
C.Convert the text to lowercase
D.Split the dataset into train and validation sets
AnswerA

Normalization and cleaning remove HTML tags, standardize date formats, and handle emoji so the text is consistent before tokenization. This step ensures the tokenizer sees clean input, reducing noise in the training data and improving the quality of the instruction-tuning examples. It is the logical first step in the preparation pipeline.

Why this answer

Cleaning and normalizing the raw chat data first removes HTML, standardizes dates, and handles emoji, producing consistent text for tokenization. This order prevents noisy artifacts from becoming tokens and ensures both training and validation splits receive the same treatment. It is the foundational step before tokenization and dataset splitting.

Exam trap

The trap here is treating tokenization as a cleaning step, when tokenizers faithfully encode whatever noise is present in the input text.

215
MCQmedium

A team has fine-tuned a Llama-3-70B model with NVIDIA NeMo and now must decide whether the tuned checkpoint actually improves on the base model for their domain. They run both models over the same 500-prompt held-out set and compute ROUGE-L against reference answers. The fine-tuned model scores 0.41 and the base model scores 0.44. What is the most technically sound conclusion the evaluation lead should draw?

A.The base model should be kept because a fine-tune that does not raise ROUGE-L proves catastrophic forgetting of the pre-training corpus.
B.The fine-tuned model is worse and should be discarded, since a lower ROUGE-L on 500 held-out prompts is a statistically significant regression.
C.ROUGE-L is insensitive to instruction-following style, so a 0.03 gap cannot be used to judge the fine-tune and the team should switch to a pairwise LLM-as-judge with human spot-checking.
D.The evaluation is valid but the held-out set is too small; expanding it to 5,000 prompts would make ROUGE-L reliable enough to select the better checkpoint.
AnswerC

ROUGE-L measures n-gram overlap with a single reference, so it penalizes valid paraphrases and rewards surface copying. A 0.03 deficit is well inside that noise floor, meaning the metric cannot resolve the question. Because the goal is domain behavior rather than lexical matching, pairwise preference judging with human verification is the appropriate instrument here.

Why this answer

ROUGE-L compares token overlap to a single reference, so it cannot distinguish a correct paraphrase from an incorrect answer, and a 0.03 gap is within that measurement noise. Since the team's goal is domain-appropriate behavior rather than lexical imitation, a pairwise preference evaluation with human adjudication gives a defensible signal. Discarding the checkpoint, enlarging the sample, or claiming forgetting all misread what the metric actually measures.

Exam trap

The trap here is assuming any numeric difference in a reference-based overlap metric reflects a real quality difference, when ROUGE-L largely measures lexical similarity rather than task correctness.

216
MCQmedium

Refer to the exhibit. What prompt engineering strategy ensures the model consistently maintains its persona and technical expertise throughout this multi-turn dialogue?

A.Append the persona instruction to every user turn.
B.Use a persistent system prompt that defines the persona and constraints.
C.Increase the temperature to 1.0 to keep the model 'active'.
D.Force the model to provide a summary of its own persona at the end of each turn.
AnswerB

A well-defined system prompt acts as the 'source of truth' that the model refers back to in every turn. By embedding the persona and constraints (like 'CUDA expert') in the system prompt, you ensure that the model stays within its defined boundaries, regardless of the complexity of the interaction.

Why this answer

In multi-turn conversations, the model can 'drift' away from its original instructions. Re-affirming constraints or using a 'System Prompt' that is explicitly included in the context of every turn is critical. For an NVIDIA code optimizer, the model must maintain its technical persona, specifically focusing on CUDA performance, regardless of how complex the dialogue becomes, ensuring the expertise level remains consistent throughout.

Exam trap

Candidates assume a single initial prompt persists automatically across long multi-turn conversations without requiring reinforcement in subsequent turns.

217
MCQmedium

Which hardware component of an NVIDIA GPU is most responsible for accelerating matrix-multiply-accumulate (MMA) operations used in transformer layers?

A.CUDA Cores.
B.Streaming Multiprocessor (SM) Scheduler.
C.Tensor Cores.
D.L2 Cache controller.
AnswerC

Tensor Cores are specialized hardware units optimized for high-performance matrix-multiply-accumulate operations. They are the engine behind modern generative AI, allowing GPUs to process large transformer models with extreme efficiency, significantly outperforming general-purpose cores for the math-heavy tasks required by LLMs and neural networks.

Why this answer

Tensor Cores are specialized hardware units designed to perform high-speed matrix-multiply-accumulate operations in a single clock cycle. By accelerating these core operations, Tensor Cores provide the massive compute throughput needed for deep learning. Understanding the role of Tensor Cores is vital because they define the performance limits for modern LLMs, and optimizing code to utilize them is the single most important task in GPU performance tuning.

Exam trap

Candidates often confuse general-purpose CUDA cores with specialized Tensor Cores when asked about matrix-multiply-accumulate acceleration.

218
MCQhard

An inference engineer is serving a 70B-parameter decoder-only LLM and wants to reduce KV cache memory to fit longer contexts on each GPU. They consider Multi-Query Attention (MQA), Grouped-Query Attention (GQA), and standard Multi-Head Attention (MHA). Which statement correctly describes the memory and quality tradeoff among these attention variants?

A.GQA eliminates the KV cache entirely by recomputing keys and values on the fly during decoding.
B.GQA uses a single key and value head shared by all query heads, matching MQA memory savings while preserving MHA quality.
C.MHA has the smallest KV cache because each head stores its own keys and values independently.
D.MQA shares key and value projections across all query heads, giving the smallest KV cache but often degrading quality relative to MHA.
AnswerD

In MQA, every query head uses the same key and value head, so the KV cache stores only one K and one V vector per layer per token, dramatically reducing memory. This aggressive sharing can hurt model quality and training stability compared with MHA, which is why intermediate GQA designs were introduced. The statement accurately captures both the memory benefit and the quality risk.

Why this answer

MQA minimizes KV cache by sharing one K and V head across all query heads, at some cost to quality. GQA interpolates by grouping query heads with dedicated K and V heads, balancing memory and quality. MHA provides the richest representation but the largest cache.

The other options misstate GQA's structure, reverse MHA's memory ranking, or incorrectly claim GQA removes the cache.

Exam trap

The trap here is conflating GQA's grouped key-value heads with MQA's single shared head, which understates GQA's memory footprint.

219
Multi-Selectmedium

An engineer is using NVIDIA TensorRT-LLM to optimize an LLM for inference. They want to reduce the memory footprint of the KV cache during long-context generation. Which TWO techniques are supported by TensorRT-LLM to achieve this? (Choose two.)

Select 2 answers
A.Use TensorRT's INT8 calibration for the KV cache.
B.Enable paged KV cache with block sharing.
C.Apply structured pruning to remove attention heads.
D.Use FP8 KV cache quantization.
E.Enable sliding window attention in the model architecture.
AnswersB, D

TensorRT-LLM's paged KV cache divides the cache into blocks and allows sharing of identical blocks across sequences, reducing memory when multiple requests share prefixes. This is a core feature for efficient memory management in long-context scenarios and is enabled by default in recent versions.

Why this answer

TensorRT-LLM provides native support for FP8 KV cache quantization and paged KV cache with block sharing. FP8 quantization halves cache memory, while paged cache with block sharing reduces duplication across sequences. Both are configuration-time features that do not require model changes.

The other options involve model modifications or misapply TensorRT features, making them incorrect for this scenario.

Exam trap

The trap here is confusing model-level optimizations like pruning or sliding window attention with runtime KV cache optimizations that TensorRT-LLM directly supports.

220
MCQeasy

Which component in the NVIDIA AI Enterprise stack is primarily responsible for serving multiple models, managing model versions, and providing metrics for monitoring model health?

A.CUDA Toolkit
B.NVIDIA Triton Inference Server
C.NVIDIA NeMo
D.NVIDIA DALI
AnswerB

Triton is the purpose-built inference server that manages model lifecycles, supports various frameworks, and exposes comprehensive metrics via endpoints like Prometheus. It is designed to optimize serving across different hardware configurations and ensures that models are served efficiently and reliably within large-scale enterprise production environments.

Why this answer

NVIDIA Triton Inference Server is the standard tool for model serving in production. It provides a unified API for various frameworks, supports concurrent model execution, and performs health checks. It is designed to handle model versioning and provide detailed telemetry data, which is essential for maintaining reliable and scalable AI deployments in enterprise production pipelines.

Exam trap

Candidates often confuse the model serving layer with the training framework or the orchestration layer, failing to identify Triton as the specific tool for serving and monitoring models.

221
MCQmedium

Which architectural component is responsible for projecting the model's hidden states back into the vocabulary space to predict the next token?

A.The Feed-Forward Network.
B.The Softmax normalization layer.
C.The LM Head linear layer.
D.The positional embedding layer.
AnswerC

The LM Head is a linear projection layer that maps the final hidden state to the vocabulary size. This is the standard architectural design for generative models, allowing the transformer to map abstract internal features to specific token IDs that correspond to the model's fixed training vocabulary.

Why this answer

The language model head (LM Head) is a final linear projection layer that maps the high-dimensional hidden states from the Transformer blocks into a vector representing the probability distribution over the entire vocabulary. This projection is the final step in the forward pass, converting the learned internal representation into actionable predictions that can be sampled to generate text.

Exam trap

Candidates frequently confuse the LM Head with intermediate Transformer attention blocks or embedding layers, forgetting which component maps hidden states directly to the vocabulary.

222
MCQeasy

A team is deploying an LLM for real-time inference using NVIDIA Triton Inference Server. They need to monitor the health and performance of the GPU to ensure reliability. Which NVIDIA tool provides comprehensive GPU telemetry, including utilization, memory, temperature, and power, and integrates with Prometheus for monitoring?

A.NVIDIA Triton Inference Server metrics endpoint
B.NVIDIA Nsight Systems
C.NVIDIA System Management Interface (nvidia-smi)
D.NVIDIA Data Center GPU Manager (DCGM)
AnswerD

DCGM is designed for data center GPU monitoring and management. It provides detailed telemetry on GPU utilization, memory, temperature, power, and more. It integrates with Prometheus via DCGM-Exporter, enabling robust monitoring and alerting for production LLM deployments. This makes it the ideal tool for ensuring GPU health and performance.

Why this answer

DCGM is the comprehensive GPU telemetry tool from NVIDIA, offering metrics on utilization, memory, temperature, power, and more. It integrates with Prometheus through DCGM-Exporter, enabling scalable monitoring and alerting. For production LLM deployments, DCGM is essential for ensuring GPU reliability and performance.

Exam trap

The trap here is confusing basic GPU monitoring utilities like nvidia-smi with full-featured telemetry tools like DCGM, which are required for production-grade monitoring.

223
MCQmedium

A team is using an NVIDIA NeMo-based LLM to answer questions over a product manual. The model sometimes answers using general knowledge instead of the provided manual excerpts. They want to force the model to rely only on the supplied context. Which prompt engineering approach best addresses this?

A.Instruct the model to answer only from the provided context and to reply 'Not in the provided context' when the answer is absent.
B.Increase the model's temperature so it explores more diverse answers from its pretrained knowledge.
C.Add more few-shot examples of correct answers without changing the instruction about using the context.
D.Shorten the prompt by removing the manual excerpts and rely on the model's product knowledge.
AnswerA

An explicit grounding instruction with a fallback phrase constrains the model to the supplied excerpts and gives it a safe response when the context lacks the answer. This reduces reliance on pretrained knowledge and makes unsupported answers visible. It is the most direct prompt-level fix for context adherence in a retrieval-augmented setup.

Why this answer

Grounding a model in retrieved context requires an explicit instruction that restricts answers to that context and defines what to do when the answer is missing. A fallback phrase such as 'Not in the provided context' prevents the model from filling gaps with pretrained knowledge. Few-shot examples help style but do not enforce the boundary; temperature and context removal work against the goal.

Exam trap

The trap here is believing that adding examples or tweaking sampling parameters will stop the model from using pretrained knowledge, when only an explicit grounding instruction with a refusal fallback reliably restricts it.

224
MCQmedium

A team is fine-tuning an 8B-parameter LLM with LoRA on a single NVIDIA A100 80GB GPU using NVIDIA NeMo. They observe that training loss decreases, but validation loss starts to rise after epoch 2. They want to keep the same dataset and hyperparameters but mitigate overfitting. Which change is most appropriate?

A.Increase the batch size and scale the learning rate proportionally to speed convergence.
B.Switch the optimizer from AdamW to SGD with momentum to regularize the adapter.
C.Increase the LoRA rank from 8 to 64 to give the adapter more capacity.
D.Reduce the number of training epochs and apply early stopping based on validation loss.
AnswerD

The reported pattern, training loss continuing to fall while validation loss rises after epoch 2, indicates the model has begun to overfit. Stopping training at or before the point of minimum validation loss directly counteracts that behavior without changing the data or architecture. Early stopping is a standard, low-risk mitigation that preserves the best generalizing checkpoint.

Why this answer

The divergence between decreasing training loss and increasing validation loss after epoch 2 is a textbook overfitting signal. The most direct, minimal-risk remedy is to stop training earlier using validation loss as the criterion, preserving the checkpoint that generalizes best. Enlarging adapter capacity or altering optimizer and batch settings does not target the generalization gap and may worsen it.

Exam trap

The trap here is assuming that a richer adapter or faster optimizer will improve results, when the observed validation curve already shows the model is memorizing rather than generalizing.

225
MCQmedium

A team is deploying a 70B-parameter decoder-only LLM on an NVIDIA H100 GPU node. During generation they observe that the KV cache grows linearly with sequence length and is consuming most of the available HBM, forcing them to limit batch size. They want to reduce KV cache memory without retraining the model from scratch. Which architectural change should they apply?

A.Switch the feed-forward network activation from SwiGLU to ReLU to reduce parameter count.
B.Apply rotary positional embeddings to the key and query vectors instead of learned absolute embeddings.
C.Increase the number of attention heads while keeping the head dimension constant.
D.Replace multi-head attention with Grouped Query Attention (GQA), where multiple query heads share a single key/value head.
AnswerD

GQA reduces the number of distinct key and value projections by having groups of query heads share KV heads, which directly shrinks the KV cache proportionally to the number of KV heads rather than query heads. This lowers HBM pressure during inference and allows larger batch sizes, and it can be introduced through continued pretraining or uptraining rather than a full retrain from scratch.

Why this answer

Grouped Query Attention reduces the number of key/value heads relative to query heads, so the KV cache that must be stored for each token shrinks by the grouping factor. Because the cache dominates HBM during long-context decoding, this architectural change directly relieves memory pressure and permits larger batches, and it can be adopted via uptraining rather than a full retrain.

Exam trap

The trap here is assuming that any attention efficiency change such as RoPE or more heads reduces KV cache size, when only reducing the number of KV heads actually shrinks the cache.

Page 2

Page 3 of 5

Page 4

All pages