Courseiva

NVIDIA Certified Professional: Generative AI LLMs (NCP-GENL) — Questions 1–75

352 questions total · 5pages · All types, answers revealed

Page 1 of 5

Page 2
1
MCQmedium

An enterprise fine-tunes a Llama-3-70B model using NVIDIA NeMo for automated technical support ticketing. The development team needs an automated evaluation pipeline that measures semantic similarity against human-curated reference answers without relying on costly human annotators. Which metric provides the most robust embedding-based semantic similarity assessment for this scenario?

A.ROUGE-1 measures unigram overlap but fails to capture semantic synonyms and contextual nuances common in enterprise technical support documentation.
B.BLEU evaluates n-gram precision with a brevity penalty, heavily penalizing valid creative paraphrasing typically found in conversational AI support responses.
C.BERTScore computes token similarity matrices using contextual embeddings to capture deep semantic meaning regardless of surface-level phrasing variations.
D.Perplexity measures how well a probability distribution predicts a sample, reflecting language fluency rather than semantic alignment with a specific reference answer.
AnswerC

BERTScore aligns tokens between candidate and reference using contextual embeddings, computing precision, recall and F1 over cosine similarity. This satisfies the stem's requirement for embedding-based semantic assessment that captures meaning despite surface-level phrasing variation, without human annotators.

Why this answer

BERTScore leverages contextual embeddings from transformer models to evaluate token-level semantic overlap rather than exact string matching, making it ideal for technical support text where phrasing varies. This automated metric correlates strongly with human judgment, significantly accelerating iteration cycles during enterprise model development workflows on NVIDIA infrastructure.

Exam trap

Candidates frequently choose exact-match string metrics like BLEU or ROUGE, failing to account for semantic synonyms and variations in technical support phrasing.

2
MCQeasy

An engineer is preparing an ensemble in NVIDIA Triton Inference Server that chains a Python preprocessing model with a TensorRT-LLM backend model. The preprocessing model must run on CPU while the LLM must run on GPU, and the ensemble must expose a single HTTP endpoint. Which configuration is required to make the ensemble execute correctly?

A.Write a custom backend that loads both models and orchestrates them in C++
B.Merge both models into a single TensorRT-LLM engine using a custom plugin
C.Use Triton's 'instance_group' to pin both models to the same GPU and rely on default routing
D.Define each step as a separate Triton model and reference them in a platform 'ensemble' model's ensemble_scheduling block
AnswerD

Triton ensembles are defined by a model with platform 'ensemble', which contains an ensemble_scheduling section listing each step, its model name, version, and input/output mappings. This lets Triton route tensors between the CPU Python model and the GPU TensorRT-LLM model while exposing one endpoint.

Why this answer

Triton ensembles use a model with platform 'ensemble' and an ensemble_scheduling block that lists each step with model name, version, and tensor mappings. This enables chaining a CPU Python preprocessing model with a GPU TensorRT-LLM model behind a single endpoint, which is exactly the requirement.

Exam trap

The trap here is reaching for custom code or instance_group tuning when Triton's built-in ensemble scheduling already solves multi-model chaining.

3
MCQeasy

You are evaluating a text generation model using NVIDIA NeMo Evaluation and want to measure how well the generated text matches a reference translation. Which metric is specifically designed for this purpose?

A.BLEU
B.METEOR
C.Perplexity
D.ROUGE
AnswerA

BLEU (Bilingual Evaluation Understudy) is designed for machine translation evaluation. It computes n-gram precision between generated and reference translations, with a brevity penalty. It is the standard metric for translation quality and is directly applicable when a reference translation exists, making it the correct choice for this scenario.

Why this answer

BLEU is the standard metric for machine translation, measuring n-gram precision with a brevity penalty. It directly compares generated text to reference translations. Perplexity measures fluency without references, ROUGE is for summarization, and METEOR, while translation-oriented, is less commonly used in NeMo Evaluation.

Thus, BLEU is the correct metric for this scenario.

Exam trap

The trap here is selecting METEOR because it is also a translation metric, but BLEU is the primary and most widely supported metric in NeMo Evaluation for this purpose.

4
MCQeasy

A startup is deploying a small LLM for a chatbot on a single NVIDIA L4 GPU using NVIDIA Triton Inference Server. They want to ensure the model is automatically loaded when Triton starts and can be updated without restarting the server. Which Triton feature should they configure?

A.Instance groups with multiple model instances per GPU.
B.Model repository with version policy and model control API.
C.Dynamic batching with preferred batch sizes.
D.Ensemble scheduler with a custom preprocessing model.
AnswerB

Triton's model repository stores model files and versions. By configuring the version policy (e.g., latest), Triton can automatically load the latest version at startup. The model control API allows loading and unloading models at runtime without restarting the server. This meets both requirements: automatic loading and dynamic updates.

Why this answer

The model repository with version policy allows Triton to automatically load models at startup, and the model control API enables loading and unloading models at runtime. This combination provides both automatic initialization and dynamic updates without server restarts. Other options address inference scheduling or parallelism, not model lifecycle management.

Exam trap

The trap here is confusing the ensemble scheduler with model management, when ensembles are for inference pipelines, not loading or versioning.

5
MCQmedium

You are curating a 200 GB instruction-tuning corpus for an NVIDIA NeMo fine-tuning job on a Llama-based model. Post-training evaluation reveals the model regurgitates exact validation-set passages verbatim. An audit shows that near-duplicate instruction/response pairs were split randomly at the record level across train and validation partitions. Which data preparation change most directly eliminates this leakage while preserving the maximum amount of usable training data?

A.Increase the validation split ratio from 5% to 20% so that fewer training records overlap with the validation partition.
B.Perform MinHash-based near-duplicate detection across the full corpus, cluster records by similarity, and assign whole clusters to either the train or validation split using a deterministic hash of the cluster ID.
C.Apply aggressive token-level cleaning to strip boilerplate phrases and punctuation from every record before splitting the dataset randomly.
D.Deduplicate only the validation partition by removing any record whose exact text appears in the training partition, then keep the original random split.
AnswerB

Grouping near-duplicates into clusters and assigning entire clusters to one partition prevents the same or paraphrased content from appearing in both train and validation, which is exactly what caused the verbatim regurgitation. Because only duplicate clusters are collapsed rather than all similar-looking records, the maximum amount of unique training data is retained, and the deterministic hash keeps splits reproducible across reruns.

Why this answer

The leakage arises because near-duplicate instruction/response pairs were split at the record level, allowing paraphrased twins to appear in both partitions. Clustering near-duplicates and assigning entire clusters to one split closes that pathway while retaining all unique content for training. Deterministic cluster-to-split hashing also makes the partition reproducible, which matters for auditing and for comparing fine-tuning runs fairly.

Exam trap

The trap here is assuming that deduplication must be exact or that adjusting the split ratio can fix leakage caused by near-duplicate records spanning partitions.

6
MCQhard

A team must serve a 70B model on a single 80 GB GPU for an internal assistant with modest concurrency. Full FP16 weights will not fit alongside the KV cache for the target context length. They want to keep accuracy loss minimal and are willing to spend additional build time. Which approach best fits these constraints?

A.Split the model with tensor parallelism across two GPUs even though only one is available.
B.Apply FP8 quantization to both weights and activations and rebuild the engine.
C.Keep FP16 weights and reduce the KV cache by capping the context length far below the requirement.
D.Use INT4 weight-only quantization with a group size and calibration set chosen for the model.
AnswerD

Weight-only INT4 quantization shrinks the weight footprint roughly fourfold versus FP16, letting a 70B model fit on one 80 GB GPU while leaving the KV cache and activations in higher precision. Group-wise scales limit accuracy loss, and the extra build and calibration time is acceptable given the stated willingness, making this the best fit.

Why this answer

The binding constraint is fitting a 70B model plus KV cache on one 80 GB GPU while preserving accuracy. Weight-only INT4 quantization cuts weight memory roughly fourfold while keeping activations and cache at higher precision, and group-wise scales with a representative calibration set limit degradation. The team's willingness to accept longer build time matches the calibration and engine-build cost this method requires.

Exam trap

The trap here is reaching for a parallelism or context-shortening workaround when the real constraint is weight footprint on a single device that weight-only low-bit quantization directly addresses.

7
MCQhard

A research team is evaluating a large language model's robustness to adversarial attacks. They want to use NVIDIA NeMo Evaluator to measure how often the model's output changes when small, semantically preserving perturbations are applied to input prompts. Which evaluation metric or method should they implement?

A.Use exact match to compare the outputs from original and perturbed prompts.
B.Measure the perplexity of the model on the perturbed prompts.
C.Calculate the semantic similarity between the two outputs using an embedding-based metric like BERTScore.
D.Compute the BLEU score between outputs from original and perturbed prompts.
AnswerC

Robustness to semantically preserving perturbations can be assessed by measuring how similar the model's outputs are. BERTScore captures semantic equivalence, so a high score indicates the model produced consistent meaning despite input changes. This directly quantifies robustness. NeMo Evaluator can integrate BERTScore as a custom metric to automate this comparison.

Why this answer

To measure robustness to semantically preserving perturbations, the team should compare the semantic similarity of outputs from original and perturbed inputs. BERTScore provides a semantic similarity score, making it suitable. BLEU and exact match are lexical and too brittle, while perplexity does not assess output consistency.

Exam trap

The trap here is using lexical metrics like BLEU or exact match to compare outputs, which fail to capture semantic equivalence and thus misrepresent robustness.

8
MCQmedium

When preparing a proprietary technical manual dataset for a RAG pipeline, which data preprocessing step is most critical to ensure the LLM avoids hallucinations regarding specific product configurations?

A.Converting all text to lowercase to ensure uniformity in vector embeddings.
B.Performing aggressive stop-word removal to reduce the dimensionality of the vector space.
C.Implementing document-aware recursive character splitting with overlapping segments and metadata tagging.
D.Applying basic sentence tokenization based strictly on periods to create uniform chunks.
AnswerC

Maintaining document hierarchy via metadata and using recursive splitting preserves logical boundaries within technical manuals. The overlap ensures that context isn't lost at chunk edges, while metadata allows the system to filter by product version, ensuring the LLM only consumes data relevant to the specific hardware revision being queried.

Why this answer

Chunking strategies and metadata tagging ensure that context retrieval is precise. By preserving technical hierarchy and associating data with specific product versions, the LLM retrieves ground-truth documentation rather than generic information. This reduces hallucinations by constraining the search space to relevant, version-controlled text blocks, directly impacting the accuracy and reliability of downstream inference tasks in enterprise NVIDIA-based AI deployments.

Exam trap

Candidates frequently choose basic paragraph splitting over recursive character splitting with metadata, which fails to respect document hierarchy and product versioning, leading to hallucinations.

9
MCQmedium

An engineer is deploying an NVIDIA NeMo Guardrails system to moderate a chatbot's responses. The chatbot must refuse to answer questions about politics but should answer questions about weather. Which prompt engineering strategy in NeMo Guardrails is most appropriate to enforce this behavior?

A.Add a Python function that checks the user input for political keywords and returns a refusal, bypassing the model entirely.
B.Define a dialogue flow with a canonical form for political questions that triggers a refusal response, and a separate flow for weather questions that allows the answer.
C.Use a system prompt that instructs the model to refuse political questions and answer weather questions.
D.Fine-tune the underlying LLM to refuse political questions and answer weather questions.
AnswerB

NeMo Guardrails uses dialogue flows defined in Colang to specify how the bot should respond to different user intents. By creating a flow that detects political questions and triggers a refusal, and another flow for weather that permits answering, you enforce the desired behavior. This is the intended use of the guardrails framework.

Why this answer

NeMo Guardrails uses Colang dialogue flows to define bot behavior based on user intent. Creating separate flows for political and weather questions ensures deterministic refusal or answering. This is the core prompt engineering approach within the guardrails framework and provides reliable, auditable control.

Exam trap

The trap here is confusing runtime guardrail flows with model-level instructions or fine-tuning, which do not provide the same deterministic enforcement.

10
MCQmedium

A healthcare company is deploying an LLM-based patient triage assistant using NVIDIA NIM microservices on-premises. To comply with HIPAA, they need to ensure that no protected health information (PHI) is transmitted to external services. Which deployment approach best meets this requirement?

A.Use a hybrid approach where sensitive data is processed locally but model updates are pulled from NVIDIA's cloud.
B.Use NVIDIA's cloud-based NIM API with a business associate agreement (BAA) to handle PHI.
C.Deploy the NIM microservice on a public cloud instance with encryption in transit and at rest.
D.Host the NIM microservice locally within the company's secure network and configure it to use only local model weights.
AnswerD

Hosting the NIM microservice locally keeps all data within the company's network, preventing PHI from leaving the premises. Using local model weights ensures no external API calls are made for inference, thereby complying with HIPAA's data residency and privacy requirements.

Why this answer

The correct approach is to host the NIM microservice locally and use local model weights. This ensures that all data, including PHI, remains within the company's secure network and is never transmitted to external services, satisfying HIPAA's strict privacy and data residency requirements.

Exam trap

The trap here is assuming that a business associate agreement or encryption alone is sufficient to comply with HIPAA when the requirement explicitly forbids any external transmission of PHI.

11
Multi-Selectmedium

A developer is creating prompts for an NVIDIA NIM-hosted LLM to summarize financial reports. The reports are lengthy and contain many tables and figures. The developer wants to ensure the summaries are accurate and include key numerical data. Which TWO prompt engineering techniques should be applied? (Choose two.)

Select 2 answers
A.Use a chain-of-thought prompt that asks the model to first list all tables and figures, then write the summary.
B.Set the temperature to a high value like 1.0 to encourage the model to creatively interpret the financial data and provide insightful analysis.
C.Instruct the model to extract and include all monetary values and percentages exactly as they appear in the report, and to avoid rounding or paraphrasing numbers.
D.Instruct the model to ignore any tables and figures and focus only on the narrative text to avoid confusion.
E.Provide a few-shot example of a summary that correctly includes key numbers from a sample report, demonstrating the desired format and level of detail.
AnswersC, E

This instruction directly addresses the need for accurate numerical data. By explicitly telling the model to include all monetary values and percentages exactly as they appear, and to avoid rounding or paraphrasing, you reduce the risk of the model altering numbers. It sets a clear constraint that the model must adhere to, which is crucial for financial summaries where precision is paramount.

Why this answer

To ensure accurate financial summaries with key numerical data, the developer should explicitly instruct the model to include all monetary values and percentages exactly as they appear, and provide a few-shot example demonstrating correct inclusion of numbers. These two techniques directly guide the model to preserve numerical accuracy and follow the desired format. High temperature, chain-of-thought, or ignoring tables would not achieve the goal.

Exam trap

The trap here is thinking that chain-of-thought or high temperature will improve numerical accuracy, when the real need is explicit instructions and examples that enforce exact inclusion of figures.

12
MCQmedium

Why do many modern LLMs use SwiGLU as their activation function in the feed-forward network instead of the traditional ReLU?

A.It reduces the number of parameters by half compared to ReLU.
B.It allows the model to perform faster matrix multiplications.
C.It provides better gradient propagation and higher model performance.
D.It makes the model compatible with 4-bit quantization.
AnswerC

SwiGLU's gating mechanism allows for dynamic control over information flow, which leads to superior convergence rates and higher final perplexity scores compared to ReLU. The smoother gradient landscape facilitates training deeper models without encountering the zero-gradient issues that are common with strictly linear activation functions like ReLU.

Why this answer

SwiGLU is a gated linear unit that incorporates the Swish activation, providing a smoother gradient flow and improved representational capacity over ReLU. ReLU's 'dying gradient' problem can hinder training progress, whereas SwiGLU's multiplicative gating allows the model to learn more flexible feature activation patterns. This is fundamental for stabilizing the training of very deep models and achieving state-of-the-art performance in complex linguistic tasks.

Exam trap

Candidates often assume SwiGLU is about reducing compute cost or latency. While efficient, its primary advantage is the improvement of gradient flow and model representational capacity during training.

13
MCQhard

Refer to the exhibit. The TensorRT build process fails with a memory limit error. Which configuration adjustment is most likely to resolve this build-time error?

A.Reducing the number of hidden layers
B.Increasing the builder's workspace memory limit
C.Switching to a lower precision inference
D.Disabling the TensorRT engine cache
AnswerB

The builder requires a temporary workspace to allocate memory for different kernel implementations. If this memory limit is too small for a complex model, the builder will fail during the optimization phase. Increasing the workspace size provides the headroom required to compute the optimal execution plan for the model.

Why this answer

The build error indicates that the workspace memory required for the optimization process exceeds the available VRAM or the limit specified by the build configuration. By explicitly increasing the workspace size in the builder configuration, the engineer provides the necessary scratchpad memory for TensorRT to evaluate and test various kernel implementations, allowing the builder to successfully complete the optimization process for all layers of the model.

Exam trap

Candidates often confuse workspace memory with the total GPU VRAM or system RAM, attempting to optimize the model size instead of the builder's temporary scratchpad allocation during the optimization phase.

14
MCQeasy

A data scientist is using an NVIDIA NeMo LLM to generate Python code from natural language descriptions. The model often produces code that works but does not follow the team's style guide, such as using single quotes instead of double quotes and missing type hints. Which prompt engineering technique should the data scientist use to improve adherence to the style guide?

A.Chain-of-thought prompting to encourage the model to reason about the style guide before writing code.
B.Few-shot prompting with examples that demonstrate the desired coding style.
C.Zero-shot prompting with a detailed instruction describing the style guide rules.
D.Increasing the temperature to allow more creative code generation.
AnswerB

Few-shot prompting provides the model with concrete examples of the desired output style, such as double quotes and type hints. By including a few input-output pairs in the prompt, the model can infer the pattern and apply it to new queries. This is a standard prompt engineering technique to guide formatting and style without retraining.

Why this answer

Few-shot prompting with examples that embody the desired style guide is the most effective way to teach the model the specific formatting rules. By showing the model correct examples, it can mimic the style in new generations. This technique is particularly useful for coding tasks where precise syntax and style matter.

Exam trap

The trap here is assuming that a detailed zero-shot instruction is sufficient to enforce style, when in fact models often require concrete examples to reliably follow formatting rules.

15
MCQmedium

A team is pre-training a 13B-parameter decoder-only LLM on a cluster of NVIDIA GPUs. They observe that gradient norms spike sharply during the first few hundred steps, destabilizing training. They want to keep the standard post-layer-normalization placement but stabilize early optimization. Which architectural technique should they apply?

A.Initialize all residual output projections with a small scaling factor such as 1/sqrt(2N).
B.Apply learning rate warmup over the first several thousand steps.
C.Insert RMSNorm immediately after the token embedding layer only.
D.Replace GELU activations in the feed-forward network with ReLU.
AnswerA

Scaling residual branch outputs at initialization by a depth-dependent factor keeps the variance of the residual stream roughly constant across layers. This directly tames the growth that causes exploding gradients in post-LN Transformers, allowing stable early training without moving normalization. It is a recognized architectural fix used in deep Transformer initialization schemes such as those in Megatron-style training.

Why this answer

Post-layer-normalization Transformers suffer from residual stream variance that grows with depth, producing large gradients early in training. Scaling residual branch outputs at initialization by a depth-dependent factor keeps variance bounded, directly stabilizing optimization without changing normalization placement or the optimizer schedule. Optimizer tweaks and activation swaps do not address the architectural root cause.

Exam trap

The trap here is assuming that learning rate warmup alone fixes post-LN instability when the underlying issue is depth-dependent residual variance growth.

16
MCQmedium

A team is training a large language model on 8 NVIDIA A100 GPUs using PyTorch's DistributedDataParallel (DDP). Profiling shows that all GPUs are frequently idle, waiting for gradient synchronization. The network interconnect between nodes is a 100 Gb Ethernet with TCP/IP, and the model has 13 billion parameters. What is the most effective optimization to reduce the idle time?

A.Increase the batch size per GPU to keep the GPUs busy during communication.
B.Replace the Ethernet interconnect with NVIDIA Mellanox InfiniBand and enable NCCL over RDMA.
C.Enable gradient compression using FP16 all-reduce to halve communication volume.
D.Use NVIDIA GPUDirect Storage to accelerate data loading from NVMe drives.
AnswerB

InfiniBand with RDMA provides higher bandwidth and lower latency than TCP/IP over Ethernet, reducing the all-reduce time for gradient synchronization. For a 13B parameter model, gradient tensors are large, and the communication overhead dominates. NCCL over RDMA bypasses the CPU and kernel network stack, significantly improving throughput and lowering idle time.

Why this answer

The idle time is caused by slow gradient synchronization over TCP/IP Ethernet. InfiniBand with RDMA and NCCL provides the necessary bandwidth and low latency to reduce all-reduce time, directly addressing the bottleneck. Other options either do not target communication or are secondary optimizations that do not resolve the fundamental interconnect limitation.

Exam trap

The trap here is assuming that increasing batch size or compressing gradients will fully resolve communication stalls, when the underlying high-latency interconnect is the true limiting factor.

17
MCQmedium

What is the architectural role of Layer Normalization in a Transformer, and where is it typically placed to ensure stable training?

A.It normalizes the entire batch to reduce training time.
B.It is placed before each sub-layer (Pre-LN) to improve training stability.
C.It is used to perform dimensionality reduction on input embeddings.
D.It replaces the need for residual connections in the model.
AnswerB

Pre-LN configurations move the normalization layer inside the residual branch, which creates a more stable gradient flow. This architectural choice is standard for modern LLMs as it significantly reduces the risk of loss divergence during the early stages of training compared to the older Post-LN approach.

Why this answer

Layer normalization stabilizes the hidden states by ensuring their mean and variance remain within a controlled range throughout the network layers. In modern LLMs, placing it before the attention and FFN blocks (Pre-LN) is preferred over the original Post-LN. This prevents gradient explosion early in training, allowing for higher learning rates and more reliable convergence for deep models during their pre-training phase.

Exam trap

Candidates often confuse Pre-LN with Post-LN. They may remember Layer Normalization exists but fail to recognize that Pre-LN is the modern standard for training stability in deep Transformers.

18
MCQmedium

When auditing an NVIDIA Triton deployment for security, which action is most critical to protect sensitive inference data?

A.Disabling the GPU logging
B.Enabling TLS for client-server communication
C.Increasing the memory buffer
D.Switching to an unsecured port
AnswerB

TLS ensures that the data transmitted between the client application and the Triton server is encrypted. This is essential for preventing unauthorized eavesdropping on sensitive prompts or responses, fulfilling basic data security requirements for production LLM deployments.

Why this answer

Securing the communication channel is the first line of defense in protecting sensitive user input data sent to an LLM. Utilizing TLS/SSL encryption prevents man-in-the-middle attacks where data could be intercepted during transit. In professional production environments, this is a standard requirement for compliance and data privacy, ensuring that prompt data and generated completions remain confidential between the client and the inference server.

Exam trap

Candidates frequently select internal model weight encryption or file system permissions, forgetting that data in transit across network boundaries remains exposed without transport-layer security.

19
MCQeasy

A developer is preparing to fine-tune a model with NVIDIA NeMo and wants a quantitative baseline before training begins. They plan to score the base model on a 300-question multiple-choice reasoning set and report accuracy. Which evaluation setup gives the most defensible baseline number?

A.Run the base model repeatedly at high temperature and report the highest accuracy observed across runs as the baseline.
B.Score the base model on the same held-out set with a fixed prompt template and deterministic decoding, then record the exact configuration alongside the accuracy.
C.Use the training split as the baseline evaluation set so the number reflects how well the model already covers the fine-tuning material.
D.Ask the model to self-report a confidence percentage for each question and average those values as the baseline accuracy.
AnswerB

A baseline is only useful if it is reproducible and comparable to later runs. Fixing the prompt template, decoding parameters, and dataset split means any future change in accuracy can be attributed to training rather than to evaluation drift. Recording the full configuration lets reviewers reproduce the number and detect accidental leakage into the fine-tuning data.

Why this answer

A trustworthy baseline requires a held-out set, a fixed prompt template, deterministic decoding, and a recorded configuration. Those choices make the number reproducible and ensure later deltas reflect training rather than evaluation noise. Best-of-runs reporting, training-set scoring, and self-reported confidence all produce numbers that either inflate performance or cannot be compared to ground-truth labels.

Exam trap

The trap here is treating a convenient or flattering number, such as best-of-runs accuracy or self-reported confidence, as a baseline when it cannot be reproduced or compared to gold labels.

20
Multi-Selecthard

You are building the data preparation stage for an NVIDIA NeMo retrieval-augmented generation pipeline that will ingest millions of internal wiki pages. The ingestion team reports that the same policy text appears in dozens of pages with minor edits, and that some pages contain copied tables from external sources. You need to produce a clean, deduplicated chunk store that supports accurate citation and avoids returning redundant passages. Which TWO actions best address these requirements? (Choose two.)

Select 2 answers
A.Apply a fixed-size token window with no overlap to every page so that chunk boundaries are uniform across the corpus.
B.Run near-duplicate detection at the chunk level using MinHash or SimHash, and retain one canonical chunk per duplicate group while preserving a mapping from the canonical chunk to all source page identifiers.
C.Attach source metadata such as page identifier, section heading, and last-modified timestamp to each chunk, and propagate that metadata through embedding and indexing so retrieved chunks can be cited.
D.Tag chunks that contain tables copied from external sources as low-quality and exclude them from the index entirely.
E.Embed every page as a single vector and store the full page text as the retrieval unit, avoiding chunking entirely.
AnswersB, C

Chunk-level near-duplicate detection removes redundant policy passages that differ only by minor edits, which directly reduces duplicate retrieval results. Keeping a mapping from the canonical chunk back to every source page identifier preserves citation integrity, so a retrieved passage can still be attributed to all originating pages even though only one copy is embedded and indexed.

Why this answer

Redundant retrieval from repeated policy text is best solved by chunk-level near-duplicate detection that keeps one canonical copy while mapping it to all source pages, and accurate citation requires source metadata to travel with each chunk through the embedding and indexing stages. Together these actions reduce duplicate results and preserve provenance, which fixed-size chunking, page-level embedding, or blanket exclusion of external tables cannot achieve.

Exam trap

The trap here is assuming that deduplication alone is sufficient and overlooking that citation integrity depends on metadata being carried alongside the retained canonical chunk.

21
MCQmedium

An engineer is tasked with optimizing a model that performs poorly due to excessive memory access latency. Which TensorRT optimization strategy specifically targets this issue?

A.Precision Calibration
B.Kernel Auto-Tuning
C.Layer Fusion
D.Weight Pruning
AnswerC

Layer fusion combines sequential operations into a single kernel, reducing the overhead of reading and writing intermediate tensors to global memory. This is the most effective approach for mitigating memory access latency because it keeps necessary data within the GPU's fast-access on-chip caches during processing.

Why this answer

Layer fusion is the primary strategy for reducing memory access latency. By combining multiple kernels into one, the engine keeps data in the high-speed cache of the GPU instead of constantly writing to and reading from slow global VRAM. This minimizes the time spent waiting for data movement, which is usually the dominant bottleneck for modern neural networks.

Exam trap

Candidates often confuse layer fusion with model pruning or quantization, failing to realize that fusion is specifically about minimizing the movement of data between GPU registers and global memory.

22
MCQmedium

A developer is optimizing a BERT-based model for inference on an NVIDIA T4 GPU using TensorRT. The model has a fixed input sequence length of 128. Profiling shows that the kernel execution time is high due to many small operations. Which TensorRT feature should they use to reduce kernel launch overhead and improve latency?

A.Increase the batch size to amortize kernel launch overhead across more samples.
B.Use TensorRT's builder optimization level 5 to enable aggressive layer fusion.
C.Apply CUDA graphs to capture the entire inference graph and replay it with a single launch.
D.Enable FP16 precision and calibrate with a representative dataset.
AnswerC

CUDA graphs capture a sequence of kernels and their dependencies into a single graph, then replay it with one launch, drastically reducing CPU launch overhead. For a fixed-shape model like this BERT with sequence length 128, CUDA graphs are ideal because the graph can be captured once and reused. This directly addresses the many small operations causing high kernel execution time.

Why this answer

CUDA graphs are designed to reduce kernel launch overhead by capturing a static sequence of operations and replaying it as a single unit. For a fixed-shape model like the BERT model with sequence length 128, the graph can be captured once and reused across inferences. This eliminates the CPU-side launch latency for each kernel, directly improving latency.

Exam trap

The trap here is focusing on precision or fusion when the bottleneck is CPU-side kernel launch overhead, which CUDA graphs specifically target.

23
MCQeasy

A developer is inspecting a decoder-only Transformer and notices that during training, the model attends to future tokens in the sequence, causing the loss to drop unrealistically fast but generation to be incoherent. Which architectural mechanism is missing or misconfigured?

A.Pre-layer normalization placement before attention
B.Causal attention mask in the self-attention layers
C.Dropout regularization on attention weights
D.Rotary positional embedding frequency base
AnswerB

In a decoder-only model, self-attention must be masked so each position can only attend to itself and earlier positions. Without this causal mask, the model sees future tokens during training, trivially minimizing next-token loss while learning nothing useful for autoregressive generation, which exactly matches the described symptom.

Why this answer

Autoregressive language models require a causal mask that sets attention scores to negative infinity for positions after the current token. When that mask is absent, the training objective becomes trivial because the target token is visible in the input context, producing the fast loss drop and incoherent generation. Positional encodings, dropout, and normalization placement do not enforce temporal directionality.

Exam trap

The trap here is confusing positional encoding with causal masking, assuming that relative position information alone prevents a decoder from attending to future tokens.

24
MCQmedium

A team is fine-tuning a 13B-parameter Llama model with NVIDIA NeMo on a node of eight A100 80GB GPUs. They want the optimizer state to be partitioned across data-parallel ranks so that per-GPU memory drops, while keeping the model replicas synchronized. Which distributed strategy should they select in the NeMo training configuration?

A.Tensor Parallelism (TP) with tensor_model_parallel_size set to 8
B.Pipeline Parallelism (PP) with pipeline_model_parallel_size set to 8
C.Activation recomputation with recompute_granularity set to full
D.ZeRO Stage 1 via the distributed optimizer in Megatron Core
AnswerD

The distributed optimizer in Megatron Core implements ZeRO Stage 1 semantics by sharding the optimizer states, such as Adam first and second moments plus the master weights, across the data-parallel ranks. Each rank keeps a full model replica for the forward and backward pass, so no extra model-parallel communication is introduced, and per-GPU memory falls in proportion to the data-parallel size.

Why this answer

Sharding optimizer state across data-parallel ranks is exactly what the distributed optimizer in Megatron Core does, providing ZeRO Stage 1 behavior with full model replicas per rank. Tensor and pipeline parallelism change how parameters and layers are split, and recomputation targets activations. Only optimizer-state partitioning reduces the Adam moments and master weights each GPU must hold without adding model-parallel communication.

Exam trap

The trap here is assuming that any multi-GPU parallelism mode reduces optimizer memory, when only data-parallel optimizer sharding actually partitions the Adam state.

25
MCQhard

A multinational insurer deploys an NVIDIA NIM-based claims triage assistant across the EU and Brazil. The compliance team must demonstrate that the system honors data-subject deletion requests and that personal data is not transferred outside approved regions. Which design decision addresses both obligations most directly?

A.Anonymize prompts before inference and retain only aggregate statistics, discarding the original records immediately.
B.Deploy regional NIM endpoints and per-region data stores so personal data stays in its jurisdiction, and implement deletion by cascading erasure across the prompt logs, vector index, and cached responses for that subject.
C.Encrypt all personal data with a single key managed by the home-region security team and replicate the encrypted stores globally for resilience.
D.Store all prompts and completions in a single centralized log bucket in the insurer's home region for simplified auditing.
AnswerB

Regional endpoints and region-scoped data stores keep personal data within its approved jurisdiction, satisfying the transfer restriction by construction. Cascading deletion across every derived store, including prompt logs, retrieval indexes, and caches, is what makes a data-subject erasure request actually complete, since copies in secondary stores are a common audit finding. Together these two design choices map one-to-one onto the two stated obligations.

Why this answer

Residency is a placement problem and erasure is a propagation problem, so the design must solve both: region-local NIM endpoints and stores for placement, and cascading deletion across logs, indexes, and caches for propagation. Centralized logging, global replication of encrypted data, and edge anonymization each fail at least one obligation, and none of them produces a defensible deletion trail. Regional architecture with explicit erasure workflows is what an auditor can verify.

Exam trap

The trap here is assuming that encrypting personal data or anonymizing it early removes the need to keep it in-region and to delete it on request.

26
Multi-Selecthard

An engineer is optimizing a large language model for inference on NVIDIA GPUs using TensorRT-LLM. They want to reduce the memory footprint of the KV cache to support longer context lengths and more concurrent requests. Which two techniques should they implement? (Choose two.)

Select 2 answers
A.Use FP32 precision for the KV cache to avoid accuracy loss.
B.Enable multi-head attention with larger head dimension.
C.Enable paged KV cache with block-based memory allocation.
D.Use INT8 quantization for the KV cache.
E.Increase the batch size to improve memory reuse.
AnswersC, D

Paged KV cache divides the cache into fixed-size blocks, eliminating fragmentation and allowing non-contiguous storage. This enables more efficient memory utilization and supports a larger number of concurrent sequences. It is a core feature of TensorRT-LLM for high-throughput serving, directly addressing memory footprint and scalability for long contexts.

Why this answer

INT8 quantization and paged KV cache are both designed to reduce memory footprint. Quantization lowers the bit-width of stored keys/values, while paging eliminates fragmentation and allows more efficient allocation. Together, they enable longer contexts and higher concurrency without increasing GPU memory.

Exam trap

The trap here is confusing batch size increases with memory savings, when larger batches actually increase total KV cache memory.

27
MCQeasy

A retail company wants to release an LLM-powered shopping assistant on NVIDIA NIM. Legal requires that the assistant never provide personalized financial advice, even if a user asks for it. Which control most directly enforces this boundary at runtime?

A.Include a sentence in the system prompt instructing the model to avoid financial advice.
B.Reduce the model's maximum output tokens so responses are too short to contain financial advice.
C.Add a NeMo Guardrails dialog rail that detects financial-advice intent and returns a fixed refusal message before the LLM is invoked.
D.Publish a terms-of-service page stating that the assistant does not provide financial advice.
AnswerC

A dialog rail intercepts the user turn and can short-circuit the request with a canned refusal, preventing the LLM from ever generating financial advice. This is a deterministic, runtime boundary that directly enforces the legal restriction regardless of model behavior, making it the most direct control.

Why this answer

Deterministic runtime boundaries are established by guardrails that intercept intent before generation. A dialog rail in NeMo Guardrails can recognize financial-advice requests and return a fixed refusal, ensuring the model never produces the prohibited content. System prompts and token limits are probabilistic or unrelated controls, and policy documents do not enforce behavior.

Exam trap

The trap here is treating a system-prompt instruction as a hard enforcement boundary when it can be bypassed by adversarial phrasing.

28
Multi-Selectmedium

A team is deploying a large language model on NVIDIA Triton Inference Server with TensorRT-LLM backend. They want to reduce GPU memory consumption to fit a larger model on a single GPU without significantly degrading output quality. Which two techniques should they use? (Choose two.)

Select 2 answers
A.Increase the maximum batch size
B.Use FP32 precision for all computations
C.Use tensor parallelism across multiple GPUs
D.Enable INT8 weight-only quantization
E.Enable paged KV cache
AnswersD, E

INT8 weight-only quantization reduces the precision of model weights from FP16 to INT8, cutting memory usage by roughly half. It often preserves output quality well because activations remain in higher precision. This allows larger models to fit on a single GPU with minimal accuracy loss, making it a suitable technique for the scenario.

Why this answer

INT8 weight-only quantization reduces weight memory footprint with minimal quality loss, and paged KV cache optimizes memory usage during generation by reducing fragmentation. Together, they enable larger models to fit on a single GPU. FP32 increases memory, larger batch sizes demand more memory, and tensor parallelism requires multiple GPUs, so they do not meet the goal.

Exam trap

The trap here is assuming that increasing batch size or using FP32 improves memory efficiency, when they actually increase memory usage.

29
MCQeasy

A healthcare startup is deploying a Mistral 7B model for internal clinical note summarization. They need to serve the model with NVIDIA Triton Inference Server and want to minimize GPU memory footprint during inference. The team plans to use TensorRT-LLM and is choosing a numerical precision for the engine. Which precision should they select to reduce memory usage while maintaining acceptable accuracy for summarization?

A.FP16, because it halves memory compared to FP32 and is widely supported on NVIDIA data center GPUs.
B.FP32, because it uses the least GPU memory and is the default for TensorRT-LLM engines.
C.TF32, because it is a Tensor Core mode that automatically compresses weights to one byte per parameter.
D.INT8 or FP8 quantization, because it reduces weight and activation precision further than FP16 while maintaining acceptable accuracy for text summarization.
AnswerD

INT8 and FP8 quantization reduce the number of bits per weight and activation compared to FP16, lowering GPU memory usage and often improving throughput. For summarization, a tolerant task, the accuracy loss is typically acceptable when calibration is done properly. This directly addresses the requirement to minimize memory footprint while keeping output quality suitable for clinical note summarization.

Why this answer

INT8 or FP8 quantization reduces the bits per weight and activation below FP16, cutting GPU memory usage while keeping accuracy acceptable for summarization. TensorRT-LLM supports quantized engines with calibration or scaling factors, making it the appropriate precision choice when memory footprint is the primary constraint and the task tolerates small accuracy loss.

Exam trap

The trap here is equating TF32 with a memory-compression format, when it is actually a Tensor Core compute mode that does not shrink stored weights.

30
MCQeasy

A team is preparing a Llama-based chatbot for production and wants to reduce GPU memory and latency without retraining. They decide to apply post-training quantization. Which TensorRT-LLM workflow correctly produces an INT8 or FP8 quantized engine from an existing FP16 checkpoint?

A.Build the FP16 engine first, then apply a post-build conversion script that rewrites the engine plan to INT8 weights.
B.Run the quantization toolkit to produce a quantized checkpoint, then build the TensorRT-LLM engine with the appropriate quantization flags.
C.Convert the checkpoint to ONNX with INT8 operators and load it directly into the TensorRT-LLM runtime without an engine build.
D.Enable automatic mixed precision in the runtime and let the inference server choose INT8 kernels dynamically at request time.
AnswerB

TensorRT-LLM expects a quantized checkpoint and quantization metadata before the engine build, so the supported path is to quantize the model with the provided toolkit and then pass the quantization mode during engine construction. This preserves accuracy through calibrated scaling factors and lets the builder select INT8 or FP8 kernels. It requires no retraining and matches the stated goal of reducing memory and latency in production.

Why this answer

Post-training quantization in TensorRT-LLM is a two-stage process: first produce a quantized checkpoint with calibration or scaling data, then build the engine with the matching quantization flags so the builder selects INT8 or FP8 kernels. This avoids retraining and yields memory and latency improvements. The other options either attempt unsupported post-build modification, rely on runtime precision switching that does not exist, or bypass the required engine compilation step.

Exam trap

The trap here is assuming quantization can be applied to an already-built engine or selected dynamically at runtime, when it must be baked into the checkpoint and engine build.

31
MCQmedium

Which metric is the most reliable indicator that an LLM is overfitting during the fine-tuning phase?

A.A consistent increase in both training and validation loss.
B.A plateau in the training loss while the validation loss remains stagnant.
C.A decrease in training loss with a simultaneous increase in validation loss.
D.The model generates text that matches the training set exactly.
AnswerC

When the training loss drops while the validation loss rises, the model is overfitting by memorizing specific training data. This divergence signifies that the model is no longer generalizing effectively to new, unseen data, which is the primary definition of overfitting in the context of machine learning.

Why this answer

An increasing validation loss alongside a decreasing training loss is the classic indicator of overfitting. The model is essentially memorizing the training samples rather than learning generalized language patterns. Monitoring this divergence is critical in NVIDIA-based training pipelines, as it allows researchers to implement early stopping or adjust regularization strategies before the model's performance on unseen tasks degrades significantly.

Exam trap

Students often monitor training loss alone, falsely believing that a continually dropping training loss indicates successful model generalization.

32
MCQhard

An engineer must serve a 70B-parameter LLM for a workload with many concurrent users and long shared system prompts, and wants to maximize throughput without retraining. Which inference-time optimization most directly reduces redundant computation across requests sharing the same prompt prefix?

A.Increasing the beam width during generation
B.Speculative decoding with a smaller draft model
C.Enabling FP8 quantization of the KV cache
D.Prefix caching of key/value tensors for shared prompt prefixes
AnswerD

Prefix caching stores the key/value tensors computed for a shared prefix, such as a long system prompt, and reuses them across requests instead of recomputing attention for those tokens on every call. This directly eliminates the redundant prefill work the scenario describes and is supported in NVIDIA TensorRT-LLM and similar serving stacks.

Why this answer

When many requests share a long system prompt, recomputing attention over that prefix for every request wastes prefill compute. Prefix caching persists the key/value tensors for the shared prefix and reuses them, cutting redundant work and improving throughput under concurrency. Speculative decoding targets decode latency, beam search adds work, and KV cache quantization saves memory but not duplicated prefill computation.

Exam trap

The trap here is conflating memory-footprint optimizations such as KV cache quantization with compute-deduplication techniques like prefix caching, since both involve the KV cache but solve different bottlenecks.

33
MCQmedium

A team is using NVIDIA Triton Inference Server to serve multiple LLMs. They want to automatically detect when a model's inference latency degrades beyond acceptable thresholds and trigger an alert. Which Triton feature should they configure?

A.Model Analyzer
B.Triton's built-in model warmup
C.Triton Metrics API with Prometheus and Alertmanager
D.Triton's dynamic batching configuration
AnswerC

Triton exposes a Metrics API endpoint that provides detailed inference metrics, including latency percentiles, request counts, and queue times. Integrating this with Prometheus allows scraping and storing these metrics, while Alertmanager can evaluate rules and trigger alerts when latency exceeds defined thresholds. This combination provides real-time monitoring and automated alerting, exactly what the team needs for production reliability.

Why this answer

Triton's Metrics API, when paired with Prometheus and Alertmanager, provides the necessary observability to monitor inference latency in real time and automatically alert when thresholds are exceeded. This integration is a standard practice for production monitoring, enabling proactive incident response. Other options are either static optimization tools or lack alerting functionality.

Exam trap

The trap here is confusing performance tuning tools like Model Analyzer or dynamic batching with runtime monitoring and alerting features.

34
MCQhard

A team is deploying a 70B-parameter LLM across four NVIDIA H100 GPUs using NVIDIA TensorRT-LLM with tensor parallelism. They observe that inference works but throughput is lower than expected, and profiling shows significant inter-GPU communication overhead. Which optimization should they apply first to reduce communication overhead?

A.Enable NVLink and ensure the GPUs are connected via NVSwitch for peer-to-peer communication.
B.Reduce the tensor parallel size to 2 and run two independent replicas.
C.Increase the number of attention heads to improve parallelism.
D.Switch from tensor parallelism to pipeline parallelism to eliminate inter-GPU communication.
AnswerA

Tensor parallelism requires frequent all-reduce operations between GPUs. NVLink with NVSwitch provides high-bandwidth, low-latency peer-to-peer communication, which is essential to reduce the overhead. Without NVLink, communication over PCIe becomes a bottleneck. Ensuring NVLink is enabled and the topology uses NVSwitch is the first and most impactful optimization for multi-GPU tensor parallelism.

Why this answer

Tensor parallelism relies on frequent all-reduce operations, so the interconnect bandwidth is critical. NVLink with NVSwitch provides the necessary high-speed peer-to-peer communication to minimize overhead. Without it, PCIe becomes the bottleneck.

Other options either do not address the communication pattern or are infeasible for a 70B model on four GPUs.

Exam trap

The trap here is thinking that pipeline parallelism eliminates inter-GPU communication, when it actually introduces different communication and pipeline bubbles.

35
MCQmedium

A financial firm is deploying a generative AI chatbot using NVIDIA NIM. To comply with strict data residency regulations, where must the inference and data processing occur?

A.On any public cloud infrastructure that supports the specific model architecture.
B.Within the firm's controlled, compliant infrastructure or local sovereign cloud region.
C.On an edge device located at the user's home or mobile office.
D.Through a distributed network of global nodes to minimize latency.
AnswerB

Keeping the data processing within controlled, geographically specified infrastructure is the only way to guarantee residency compliance. By using a private environment or a sovereign cloud, the organization enforces strict physical and logical boundaries that prevent sensitive data from exiting the jurisdiction, satisfying both legal and security obligations.

Why this answer

Data residency regulations require that sensitive data remains within specific geographic or sovereign borders. By deploying models on-premises or within a verified sovereign cloud region using NVIDIA NIM, the organization retains complete control over the data lifecycle. This ensures that personal or financial information is not processed, logged, or stored in unauthorized regions, which is a critical requirement for regulatory compliance in the financial sector.

Exam trap

Candidates often select cloud-agnostic SaaS or public multi-tenant APIs, forgetting that financial regulations mandate dedicated physical control or localized sovereign infrastructure to prevent cross-border data leakage.

36
MCQhard

Refer to the exhibit. A team is preparing log data for a RAG-based troubleshooting assistant. Given the configuration, what is the most significant risk during the retrieval phase?

A.The use of 'nv-embed-v1' will cause OOM errors during the indexing phase.
B.The chunk size of 4096 tokens is too small for modern log analysis.
C.Setting the overlap to 0 will break contextual continuity between consecutive log entries.
D.The 'regex_mask_all' policy will cause the embedding model to fail during vectorization.
AnswerC

Logs are inherently sequential. A zero-overlap configuration ensures that each chunk is treated as an isolated entity, potentially splitting related log events. This prevents the LLM from seeing the full narrative of a system failure, significantly reducing the diagnostic utility of the retrieval-augmented generation output for complex errors.

Why this answer

The lack of overlap between chunks is critical. In log data, sequential events are often related; by having zero overlap, the system loses the transition between log lines that might contain a causal link. This fragmentation forces the model to interpret isolated snapshots, hindering its ability to reconstruct the sequence of errors or system states, which is vital for effective root-cause analysis.

Exam trap

Candidates often focus on chunk size or retrieval speed, missing that zero overlap in sequential data like logs destroys the context necessary for the model to understand causal relationships.

37
MCQeasy

An engineer is reviewing the attention implementation of a decoder-only LLM used for chat. During inference with a KV cache, generated tokens must not attend to future positions. Which mechanism enforces this constraint inside scaled dot-product attention?

A.Dropping the scaling factor 1/sqrt(d_k) so that scores decay for distant tokens.
B.A causal mask that sets attention scores for future positions to negative infinity before the softmax.
C.Applying LayerNorm to the query and key projections before computing attention scores.
D.Shifting the position IDs of cached keys so they appear before the current token.
AnswerB

A causal (lower-triangular) mask adds negative infinity to scores at positions beyond the current token, so after softmax those weights become zero. This guarantees each position only attends to itself and earlier tokens, which is exactly the autoregressive constraint required for decoder-only generation, and it works identically whether or not a KV cache is used.

Why this answer

Causal masking is the standard mechanism that makes decoder-only attention autoregressive. By adding negative infinity to scores of future positions before softmax, those positions receive zero weight, so each token can only attend to itself and prior tokens. Other listed techniques affect numerical stability or positional encoding but do not enforce the temporal constraint.

Exam trap

The trap here is confusing numerical stabilization techniques like score scaling or QK-norm with the masking mechanism that actually enforces causal visibility.

38
MCQhard

An ML engineer is fine-tuning a 70B model with NVIDIA NeMo Framework across 16 H100 GPUs. Training completes successfully, but when the fine-tuned checkpoint is evaluated, outputs are incoherent and repeat tokens. The engineer confirms the loss decreased smoothly during training and the validation dataset was held out correctly. Which issue is the most likely explanation?

A.The checkpoint was saved and later loaded with a mismatched tokenizer vocabulary, so input token IDs no longer map to the trained embeddings.
B.The evaluation prompts were longer than the maximum sequence length used during fine-tuning, so the model never learned to handle them.
C.The model was trained with a sequence packing strategy that concatenated unrelated samples without proper attention masking between them.
D.The learning rate was too high, causing the optimizer to overshoot and permanently corrupt the base model weights.
AnswerC

Sequence packing can improve throughput, but if attention masking does not prevent tokens from attending across sample boundaries, the model learns spurious cross-sample dependencies. This degrades generation quality, often producing repetition and incoherence, even though the training loss looks healthy because the model is fitting the corrupted attention pattern.

Why this answer

Sequence packing without correct cross-sample attention masking lets tokens attend to unrelated neighboring samples, teaching the model spurious dependencies that surface as repetition and incoherence at inference. The training loss can still fall smoothly because the model is fitting the corrupted attention structure, so loss alone does not reveal the defect.

Exam trap

The trap here is trusting a smooth loss curve as proof of a correct data pipeline, when packing without proper attention masking can degrade generation quality while the loss still looks healthy.

39
MCQmedium

When preparing unstructured documentation for a high-performance retrieval system, which approach best balances index size and retrieval relevance?

A.Using extremely large chunks to ensure that every document is a single vector.
B.Implementing sliding window chunking with semantic overlap based on document structure.
C.Storing every sentence as an individual chunk in the vector database.
D.Removing all overlaps to keep the index size at the absolute minimum possible.
AnswerB

Sliding window chunking with semantic overlap allows the retrieval system to maintain context across chunk boundaries. By respecting document structure, the system ensures that chunks are meaningful and logically coherent. This approach provides the best balance between retrieval granularity, context preservation, and overall index size efficiency for high-performance systems.

Why this answer

Utilizing a sliding window approach with semantic overlap ensures that retrieved chunks maintain context. By carefully selecting chunk size and overlap, engineers can optimize the index size to avoid redundant storage while ensuring that the semantic units of the text are not cut off. This balance is vital for maximizing the accuracy of RAG systems running on NVIDIA infrastructure, where memory efficiency is paramount.

Exam trap

Test-takers frequently choose fixed-size character chunking without overlap, mistakenly assuming it preserves context, when semantic overlap is essential to prevent cutting off critical context.

40
MCQhard

When profiling an application with NVIDIA Nsight Systems, you notice a long gap between kernel execution blocks on the GPU timeline. What is the most likely cause?

A.The GPU is overheating and triggering thermal throttling.
B.Excessive host-side synchronization calls.
C.The PCIe bus is saturated with high-frequency data.
D.The kernel is launching with an invalid thread block size.
AnswerB

Explicit synchronization points force the CPU to wait for the GPU to finish all previous tasks before continuing. These gaps represent the time the GPU spends waiting for the CPU to process logic and issue new work, effectively serializing the pipeline and creating idle time on the GPU execution timeline.

Why this answer

Long gaps in the GPU timeline usually indicate host-side synchronization, such as explicit cudaDeviceSynchronize() calls or CPU-bound code that is stalling the kernel launch queue. In production systems, unnecessary CPU-to-GPU synchronization forces the GPU to remain idle while waiting for the CPU to catch up, directly impacting system-wide latency and violating the principle of asynchronous task pipelining necessary for optimal GPU utilization.

Exam trap

Candidates often attribute these gaps to kernel execution time or network latency, ignoring the role of the host CPU and synchronization primitives in managing the GPU task queue execution flow.

41
MCQmedium

A developer is using an NVIDIA NIM for a customer support chatbot. The chatbot must handle multi-turn conversations and maintain context about the user's issue. The developer notices that after several turns, the bot starts giving generic responses and forgets earlier details. Which prompt engineering approach is most effective to maintain context?

A.Use a summarization prompt at each turn to condense the conversation so far, and include only the summary in the next prompt.
B.Add a system prompt that instructs the model to always ask the user to repeat their issue if it is unsure.
C.Use a higher temperature setting to make the model more creative in interpreting the user's issue.
D.Include the entire conversation history in the prompt for each turn, ensuring the model has access to all previous messages.
AnswerD

Including the full conversation history in the prompt allows the model to attend to all previous turns, maintaining context. This is a standard approach for multi-turn dialogues. However, it is limited by the context window size; if the conversation exceeds the window, older messages may be truncated. Despite this, for many support conversations, it is the most straightforward and effective method to preserve context without additional infrastructure.

Why this answer

Including the full conversation history in the prompt is the most direct way to maintain context across turns. The model can attend to all previous messages, ensuring it remembers details from earlier in the conversation. While summarization can reduce token usage, it risks losing critical information.

Temperature and system prompts do not address context retention.

Exam trap

The trap here is assuming that summarizing the conversation will preserve all necessary details, when in fact summarization can omit specifics that are crucial for support interactions.

42
MCQhard

An enterprise is running a mission-critical generative AI application on an NVIDIA DGX cluster. The MLOps team notices occasional silent GPU memory corruption during long-running inference jobs that do not trigger hard crashes. Which monitoring tool and strategy should be utilized for early detection?

A.Deploy NVIDIA Data Center GPU Manager (DCGM) with Prometheus exporter to monitor XID errors and hardware ECC events continuously.
B.Increase the frequency of client-side HTTP ping probes sent from the load balancer to the API gateway.
C.Write a custom bash script that runs nvidia-smi every five minutes and parses plain text output for warning strings.
D.Rely on standard Kubernetes node liveness probes to automatically restart pods when system memory usage exceeds 90%.
AnswerA

DCGM provides specialized diagnostic tests and continuous health monitoring, including ECC error tracking and XID error detection. Configuring DCGM to raise alerts on uncorrectable memory errors or specific XID failure codes allows operators to isolate failing GPUs before they impact production workloads.

Why this answer

DCGM (Data Center GPU Manager) provides specialized diagnostic tests and continuous health monitoring, including ECC error tracking and XID error detection. Configuring DCGM to raise alerts on uncorrectable memory errors or specific XID failure codes allows operators to isolate failing GPUs before they impact production workloads.

Exam trap

Engineers often rely solely on standard Kubernetes pod health checks, completely missing hardware-level silent errors that occur beneath the container runtime layer.

43
MCQeasy

An engineer is profiling a CUDA kernel and notices that the achieved occupancy is low, leading to underutilization of the GPU. The kernel uses a large number of registers per thread, limiting the number of resident warps. Which optimization should be attempted first to improve occupancy?

A.Use the __launch_bounds__ qualifier to limit registers per thread.
B.Enable L1 cache to reduce memory latency.
C.Increase the block size to allow more warps per block.
D.Convert the kernel to use shared memory for data reuse.
AnswerA

The __launch_bounds__ qualifier allows the programmer to specify the minimum number of blocks per multiprocessor, which guides the compiler to limit register usage. By reducing registers per thread, more warps can be resident, improving occupancy. This is a direct way to address register-limited occupancy without changing the algorithm, though it may cause spilling if overused.

Why this answer

The low occupancy is caused by high register usage per thread. Using __launch_bounds__ instructs the compiler to limit registers, allowing more warps to be resident and improving occupancy. This directly addresses the bottleneck, whereas other options do not target register pressure and thus would not effectively improve occupancy.

Exam trap

The trap here is assuming that increasing block size or enabling cache automatically improves occupancy, when the real constraint is the number of registers per thread limiting resident warps.

44
Multi-Selecthard

A production LLM inference service on NVIDIA Triton Inference Server is being monitored for reliability. The team wants to implement effective logging to diagnose issues such as high latency and errors. Which TWO logging practices are recommended for a production LLM environment? (Choose two.)

Select 2 answers
A.Log every inference request and response payload for full traceability.
B.Use structured logging (e.g., JSON) to enable efficient parsing and querying.
C.Disable logging in production to maximize inference throughput.
D.Log GPU temperature and power metrics at debug level for every inference.
E.Include a unique request ID in logs to correlate events across distributed components.
AnswersB, E

Structured logging formats like JSON make it easier to parse and query logs programmatically. This is crucial for automated monitoring and alerting systems. It allows filtering by fields such as request ID, model name, and latency, enabling quick identification of issues in production LLM deployments without manual log inspection.

Why this answer

Including a unique request ID enables end-to-end tracing, which is vital for diagnosing latency and errors in distributed LLM inference pipelines. Structured logging facilitates automated parsing and querying, allowing efficient monitoring and alerting. Together, they provide robust observability without excessive overhead or security risks.

Exam trap

The trap here is assuming that more logging is always better, leading to practices like logging full payloads or per-inference GPU metrics, which can harm performance and privacy.

45
MCQhard

You are preparing a dataset for continued pre-training of an LLM on internal engineering documents. The corpus contains many documents with boilerplate headers, footers, and legal disclaimers repeated across files. Which NeMo Curator approach best reduces this boilerplate while preserving unique technical content?

A.Use the n-gram or sentence-level deduplication stage to remove repeated spans
B.Apply a heuristic filter that drops documents below a word-count threshold
C.Lowercase all text and remove punctuation
D.Split documents into fixed-size chunks and keep only the first chunk
AnswerA

NeMo Curator supports span-level deduplication, which can identify and remove repeated n-grams or sentences that appear across many documents. This directly targets boilerplate headers, footers, and disclaimers while leaving unique technical passages intact. It is the most precise way to reduce redundancy without discarding entire documents.

Why this answer

Span-level deduplication in NeMo Curator detects repeated n-grams or sentences across documents and removes them, which precisely targets boilerplate headers, footers, and legal disclaimers. Unlike document-level filtering, it preserves the unique technical content within each file, making it the best choice for a corpus where redundancy is localized rather than whole-document.

Exam trap

The trap here is assuming that document-level deduplication or filtering will remove boilerplate, when the repeated text is only a small portion of otherwise unique documents.

46
Multi-Selecthard

A team is deploying an NVIDIA NIM for a Llama 3 model as a retrieval-augmented generation (RAG) assistant over internal documentation. Users report that the assistant sometimes answers from its pretrained knowledge instead of the retrieved passages, and occasionally cites a passage that does not support its claim. Which TWO prompt engineering changes best reduce these behaviors? (Choose two.)

Select 2 answers
A.Add a chain-of-thought instruction asking the model to reason about why the user's question is interesting before answering.
B.Raise the temperature to 0.8 so the model produces more varied answers and avoids memorized responses.
C.Instruct the model to answer only from the provided context and to respond with a fixed phrase when the context is insufficient.
D.Require the model to quote the exact sentence from the context that supports each claim before stating the answer.
E.Remove the retrieved passages from the prompt and rely on the model's pretrained knowledge to keep the prompt short.
AnswersC, D

An explicit grounding instruction tells the model that the retrieved passages are the sole source of truth. A fallback phrase for insufficient context prevents the model from filling gaps with pretrained knowledge. This directly addresses both symptoms: unsupported answers and reliance on internal memory, by defining what the model may use and what it must do when the context is inadequate.

Why this answer

Grounding the model with an explicit instruction to use only the provided context, plus a fallback for insufficient information, prevents reliance on pretrained knowledge. Requiring verbatim supporting quotes makes each claim auditable and discourages fabricated citations. Together these changes enforce source adherence and citation accuracy, while sampling changes or removing context would undermine the RAG design.

Exam trap

The trap here is assuming that increasing temperature or adding reasoning steps will improve faithfulness, when the real fix is constraining the model to the retrieved context and requiring verifiable quotes.

47
MCQeasy

A team is preparing a supervised fine-tuning job in NVIDIA NeMo Framework for a customer-support assistant. They have a large corpus of raw support chat logs with no labels. They want the model to learn to answer customer questions in the company's tone and format. Which data preparation step is most appropriate before training?

A.Convert the raw chat logs into instruction-response pairs that reflect the desired tone and format, then use them for supervised fine-tuning.
B.Use the raw logs directly as a supervised fine-tuning dataset by treating each message as both instruction and response.
C.Apply reinforcement learning from human feedback using the raw logs as the reward model training data without any preference labels.
D.Run continued pretraining on the raw chat logs so the model absorbs the company's vocabulary and style without any labeling.
AnswerA

Supervised fine-tuning requires paired instruction and response examples that demonstrate the target behavior. Converting raw logs into curated instruction-response pairs gives the model explicit examples of the desired tone and format, which is exactly what supervised fine-tuning learns from. This aligns the training data with the task objective.

Why this answer

Supervised fine-tuning learns behavior from explicit instruction-response pairs, so raw logs must be transformed into curated examples that demonstrate the target tone and format. This gives the model clear, aligned supervision and is the most direct way to shape how the assistant responds to customer questions.

Exam trap

The trap here is assuming any domain-specific text can be used directly for fine-tuning, when supervised fine-tuning specifically requires structured instruction-response pairs that demonstrate the target behavior.

48
MCQhard

Refer to the exhibit. What is the most likely reason for the high P99 latency despite low GPU utilization?

A.The GPU compute units are failing
B.Inappropriate concurrency settings
C.The model is too large for memory
D.Network bandwidth is saturated
AnswerB

Low GPU utilization combined with high queue depth confirms that the system is not processing enough concurrent requests to utilize the hardware fully. The bottleneck is the configuration limiting the number of parallel inferences, causing requests to queue unnecessarily.

Why this answer

When GPU utilization is low but P99 latency is high, the bottleneck is usually not the compute capability but rather the request queuing or orchestration layer. High queue depth indicates that requests are waiting for a slot, likely due to a restrictive concurrent request setting or inefficient backend task scheduling. This scenario highlights the importance of matching inference concurrency settings with the hardware's throughput capacity to avoid artificial bottlenecks.

Exam trap

Candidates often blame the model size or GPU compute capacity. If utilization is low, the hardware is idle, meaning the bottleneck is clearly in the request handling or queuing configuration.

49
MCQmedium

An enterprise is deploying a large language model on NVIDIA Triton Inference Server. Which deployment strategy minimizes latency for requests that require high-throughput batching while maintaining consistent hardware utilization?

A.Implement static batching with a fixed size of 1.
B.Enable Dynamic Batching in the Triton model configuration file.
C.Disable all batching features to process requests serially.
D.Offload all batching logic to the client-side application layer.
AnswerB

Dynamic Batching aggregates individual requests into batches based on defined delay windows, maximizing GPU compute cycles. By adjusting batching parameters, administrators can balance throughput and latency effectively. This is the industry-standard method for optimizing NVIDIA hardware utilization when serving LLMs in real-world, high-concurrency production environments.

Why this answer

Dynamic Batching is the optimal strategy for Triton Inference Server in production environments. It groups individual inference requests arriving within a short time window into a single batch, allowing the GPU to process them in parallel. This maximizes throughput by fully saturating CUDA cores, reducing the overhead of kernel launches, and ensuring that hardware utilization remains high even under variable traffic loads, effectively balancing latency and overall system capacity.

Exam trap

Candidates often confuse dynamic batching with model parallelism or caching, failing to recognize that Triton's dynamic batching is specifically designed to maximize GPU utilization by grouping requests at runtime.

50
MCQhard

Refer to the exhibit. An engineer notices that the TensorRT engine takes an excessively long time to build. What is the most likely cause, and how can it be mitigated?

A.Decrease the workspace size
B.Use a calibration cache
C.Use an explicit tactic sources list
D.Increase the batch size
AnswerC

Specifying the tactic sources allows the builder to skip certain search paths or limit the number of kernels it tests. By narrowing the scope of the tactic search, the build process completes significantly faster while still producing a highly optimized engine that utilizes the target GPU's capabilities effectively.

Why this answer

The long build time is likely due to the large workspace allocation combined with a large search space for the optimization tactics. TensorRT tests a variety of kernel implementations to find the fastest one. To reduce build time, the engineer can limit the 'tactic selection' or use a profile-based build where common shapes are pre-recorded, preventing the engine from exhaustively searching every possible configuration for all input shapes.

Exam trap

Candidates often guess that the model is too large or the GPU is underpowered, missing the fact that TensorRT's exhaustive search for optimization tactics is the primary cause of slow build times.

51
MCQhard

In Mixture-of-Experts (MoE) architectures, why does the use of a router mechanism significantly impact performance compared to dense models?

A.It forces all parameters to update at every training step.
B.It allows scaling parameters while keeping compute costs manageable.
C.It forces the model to use all experts simultaneously for every input.
D.It replaces the attention mechanism with a standard linear layer.
AnswerB

MoE allows for a massive total parameter count while maintaining a constant amount of active parameters per token. By routing tokens to specialized experts, the model achieves high capacity and knowledge breadth without the linear compute cost increases that would occur in a fully dense architecture of equivalent size.

Why this answer

The router determines which expert blocks to activate for a given input token, ensuring that only a subset of the model's total parameters are used per forward pass. This decoupling of model size from computation latency allows for the creation of massive, high-capacity models that run with the speed of much smaller dense models. This architectural choice is central to modern scaling laws.

Exam trap

Candidates mistakenly think MoE architectures reduce total parameter count, confusing sparse activation of routing with physical parameter pruning or compression.

52
MCQeasy

A company wants to teach a pretrained LLM to follow a specific output format for customer support replies using supervised fine-tuning on NVIDIA GPUs. Which data preparation approach best matches supervised fine-tuning for instruction following?

A.Score model outputs with a reward model and update the policy using a policy gradient objective
B.Provide only the desired responses without prompts and train the model to reproduce them verbatim
C.Collect pairs of an instruction prompt and the desired response, then train the model to predict the response tokens
D.Gather a large unlabeled corpus and train the model to predict the next token across all documents
AnswerC

Supervised fine-tuning for instruction following uses prompt and response pairs where the loss is computed on the response tokens. This teaches the model the mapping from instruction to desired output format. NeMo's SFT data formats, such as the prompt-completion and chat schemas, are built around exactly this structure, making it the correct data preparation approach.

Why this answer

Instruction-following SFT learns from prompt and response pairs with the loss applied to response tokens, which directly teaches the desired output format. Continued pretraining on unlabeled text, RLHF with a reward model, and response-only training all lack the prompt-conditioned supervised signal that this task requires.

Exam trap

The trap here is confusing any training on domain text with instruction tuning, when only prompt-response pairs teach the model to follow a requested format.

53
MCQeasy

You are preparing a customer-support dataset for fine-tuning an LLM with NVIDIA NeMo. The raw data includes personally identifiable information such as names, email addresses, and phone numbers. Which data preparation step must be performed before training to comply with privacy requirements?

A.Apply PII detection and redaction to replace sensitive entities with placeholders.
B.Shuffle the dataset to break associations between PII and responses.
C.Increase the batch size during fine-tuning to average out PII exposure.
D.Tokenize the dataset with a custom vocabulary that includes PII patterns.
AnswerA

Detecting and redacting PII replaces names, emails, and phone numbers with generic placeholders, removing sensitive information while preserving the conversational structure needed for fine-tuning. This directly addresses the privacy requirement and prevents the model from memorizing or emitting real customer data, making it the correct preparation step.

Why this answer

Privacy compliance requires removing or masking personally identifiable information before training. PII detection and redaction replaces sensitive entities with placeholders, preserving the dataset's instructional value while preventing the model from memorizing real customer details. Tokenization, batch size, and shuffling do not alter the presence of PII and therefore cannot satisfy the requirement to protect sensitive data during fine-tuning.

Exam trap

The trap here is confusing training-time hyperparameters like batch size or shuffling with data-preparation steps that actually remove sensitive content.

54
MCQmedium

An enterprise deployment team needs to deploy a Large Language Model on NVIDIA Triton Inference Server. They require the lowest possible latency for real-time inference while maximizing GPU memory utilization. Which configuration strategy should the team implement?

A.Enable dynamic batching with a maximum delay of 500ms to gather requests.
B.Deploy the model as a standard PyTorch model using the default model repository.
C.Implement TensorRT-LLM with In-flight Batching and CUDA Graph capture.
D.Use the Triton Model Analyzer to force a single-request batch size of 1.
AnswerC

This configuration maximizes GPU efficiency by decoupling request scheduling from token generation. In-flight batching allows for concurrent processing of tokens, while CUDA graphs minimize driver launch overhead. This combination is the industry-standard path for achieving high-performance LLM serving on NVIDIA hardware.

Why this answer

To achieve optimal performance on Triton, using TensorRT-LLM backends with In-flight Batching (IFB) is essential. IFB allows the server to process new requests before existing ones finish, significantly increasing throughput and reducing latency. Configuring CUDA Graph capture further reduces CPU overhead by recording GPU operations.

This approach is critical for LLMs because traditional static batching often leads to underutilized GPU cycles during the autoregressive decoding phase.

Exam trap

Candidates often confuse static batching with In-flight Batching, incorrectly assuming that standard model ensemble configurations provide the same latency benefits as native IFB support offered by the TensorRT-LLM backend.

55
MCQhard

An engineer is using an NVIDIA NIM for a code generation model to produce Python functions from natural language descriptions. The model frequently generates code that uses deprecated libraries or incorrect function signatures. The engineer wants to improve the accuracy of the generated code by providing examples. Which prompting strategy is most appropriate?

A.Zero-shot prompting with a detailed description of the desired function and its parameters.
B.Few-shot prompting with examples that demonstrate the correct use of the desired libraries and function signatures.
C.Self-consistency prompting where the model generates multiple code solutions and the most common one is selected.
D.Chain-of-thought prompting that asks the model to first explain its reasoning about which libraries to use, then write the code.
AnswerB

Few-shot prompting provides the model with concrete examples of the desired output, including correct library usage and function signatures. By showing several input-output pairs where the output uses the correct libraries and signatures, the model can learn the pattern and apply it to new inputs. This is especially effective for code generation where precise syntax and API usage matter, and it directly addresses the deprecated library and signature issues.

Why this answer

Few-shot prompting with examples that demonstrate correct library usage and function signatures is the most direct way to guide the model to produce accurate code. By showing the desired pattern, the model can imitate it, reducing the likelihood of deprecated libraries or incorrect signatures. Other strategies like zero-shot, chain-of-thought, or self-consistency do not provide the concrete examples needed to correct systematic API errors.

Exam trap

The trap here is assuming that chain-of-thought or self-consistency will fix incorrect API usage, when the model needs concrete examples of correct code to learn the pattern.

56
MCQmedium

Which metric provides the best indication of 'inference queue saturation' in a Triton deployment?

A.Total GPU memory usage
B.Request queue duration
C.GPU temperature
D.System clock speed
AnswerB

Queue duration measures the time a request waits before being executed. An increasing trend in this metric is the definitive sign that the inference server is unable to process incoming requests as quickly as they arrive, indicating clear saturation.

Why this answer

Queue duration is the most direct indicator of saturation. It measures how long an inference request spends waiting in the server's input buffer before it is processed by the GPU. In a production environment, monitoring this metric is crucial to identify when the server's request capacity has been exceeded, allowing for auto-scaling triggers to provision more instances and maintain performance standards during traffic spikes.

Exam trap

Candidates often choose GPU utilization or throughput. These metrics indicate how well the hardware is working, but they do not reveal if requests are being delayed in the server's input buffer.

57
Multi-Selecthard

A team is designing prompts for an NVIDIA NIM-hosted LLM that must produce concise, citation-backed answers from retrieved documents. They want to improve factual grounding and reduce unsupported claims. Which two prompt engineering practices best support this goal? (Choose two.)

Select 2 answers
A.Remove document identifiers from the context to simplify the prompt and reduce tokens.
B.Instruct the model to cite the specific document ID or snippet for each claim it makes.
C.Increase temperature to encourage the model to synthesize multiple documents creatively.
D.Instruct the model to say 'Insufficient evidence' when the retrieved documents do not contain the answer.
E.Allow the model to answer from general knowledge when retrieved documents are incomplete.
AnswersB, D

Requiring citations forces the model to tie statements to retrieved sources, making unsupported claims easier to detect and reducing free-form invention. It also gives reviewers a way to verify answers. This practice directly supports factual grounding because the model must reference evidence rather than rely on parametric memory.

Why this answer

Factual grounding in retrieval-augmented generation improves when the model must cite sources and when it is allowed to refuse when evidence is missing. Citations create traceability, and a refusal fallback prevents guessing. Allowing general knowledge, raising temperature, or stripping identifiers weakens provenance and increases unsupported claims.

Exam trap

The trap here is treating creativity or general knowledge as helpful, when citation-backed grounding requires restricting answers to retrieved evidence and permitting refusal.

58
MCQhard

A team is deploying a 70B-parameter LLM using NVIDIA Triton Inference Server with TensorRT-LLM backend on a node with four A100 80GB GPUs. They observe that during inference, only one GPU is utilized while the others remain idle. They have configured the model with tensor parallelism set to 1. What is the most likely cause of this underutilization?

A.The Triton model repository is not configured with the correct instance group for multi-GPU execution.
B.Tensor parallelism is set to 1, which means the model is not partitioned across GPUs; it runs entirely on one GPU.
C.The model was compiled with a batch size that is too small to trigger multi-GPU execution.
D.The KV cache is not enabled, causing the model to fall back to single-GPU execution.
AnswerB

Tensor parallelism splits model layers across multiple GPUs. When set to 1, no partitioning occurs, so the entire model resides on a single GPU, leaving others idle. To utilize all four GPUs for a single model instance, tensor parallelism must be set to 4, matching the number of GPUs. This is the direct cause of underutilization.

Why this answer

Tensor parallelism determines how a model is sharded across multiple GPUs. Setting it to 1 means no sharding, so the entire model runs on a single GPU, leaving others idle. To utilize all four GPUs for a single model instance, tensor parallelism should be set to 4.

Other factors like instance groups or batch size do not override this fundamental partitioning setting.

Exam trap

The trap here is confusing instance groups (which manage multiple model instances) with tensor parallelism (which splits a single model across GPUs).

59
MCQeasy

A data engineer is preparing a JSONL instruction dataset for an NVIDIA NeMo supervised fine-tuning run. Each line currently contains a free-form 'text' field with the instruction, context, and response concatenated. The training configuration expects the standard NeMo instruction-tuning schema with separate fields for the task instruction, optional context, and the expected response. What is the most appropriate data preparation step?

A.Keep the free-form text field and modify the NeMo training configuration to treat the entire line as the response.
B.Convert the JSONL file to plain text with one example per paragraph and let the tokenizer infer the instruction and response boundaries.
C.Duplicate each record and label one copy as instruction and the other as response so the loader sees two fields per example.
D.Parse each record and rewrite it into the expected fields, such as instruction, input, and output, while preserving the original text content and escaping any embedded quotes.
AnswerD

NeMo's supervised fine-tuning data loader expects distinct fields for the instruction, optional context, and response, so parsing the concatenated text into those fields makes the dataset consumable without custom loader code. Preserving content and escaping quotes maintains data fidelity and prevents malformed JSONL lines that would break parsing during training.

Why this answer

NeMo's supervised fine-tuning loader expects separate instruction, context, and response fields, so the free-form text must be parsed into that schema with content preserved and quotes escaped. This produces a dataset the standard training configuration can consume directly, keeping the run reproducible and allowing loss to be computed on the response portion alone.

Exam trap

The trap here is assuming the training configuration can be bent to accept free-form text instead of reshaping the data to match the loader's expected schema.

60
Multi-Selecthard

A team is evaluating a large language model for a question-answering system using NVIDIA NeMo Evaluation. They need to assess both the relevance of the answer to the question and its factual correctness. (Choose two.)

Select 2 answers
A.F1 score over tokens
B.BLEU score
C.Exact Match (EM)
D.Factual consistency score using an NLI model
E.Answer relevance score from a QA evaluation model
AnswersD, E

A factual consistency score uses an NLI model to check whether the generated answer is entailed by a trusted knowledge source (e.g., a reference document). This directly measures factual correctness, the second required criterion. It is robust to paraphrasing and can be automated within NeMo Evaluation to flag hallucinations or contradictions.

Why this answer

The team must evaluate relevance and factual correctness. Answer relevance score directly measures how well the answer addresses the question, while factual consistency score via NLI checks whether the answer is factually supported. Exact Match, token F1, and BLEU focus on string overlap and do not separately assess relevance and factual correctness, making them less suitable for this dual requirement.

Exam trap

The trap here is assuming that overlap-based metrics like F1 or BLEU capture both relevance and factual correctness, when they primarily measure lexical similarity and can be misled by paraphrasing or fluent hallucinations.

61
MCQhard

Refer to the exhibit. What is the implication of setting the memory_limit to 0.8 in the context of an LLM inference service?

A.It guarantees that the system will never crash due to memory.
B.It will evict old KV cache entries when 80% limit is reached.
C.It forces the GPU to run at 80% of its clock speed.
D.It expands the memory capacity by 20% using swap space.
AnswerB

The 0.8 setting acts as a cap on the memory footprint of the KV cache. When usage reaches 80% of the assigned memory, the LRU policy triggers the eviction of the least recently used entries, allowing the system to continue operation without crashing, albeit at the cost of losing older sequence data.

Why this answer

A memory limit of 0.8 indicates that the system will reserve up to 80% of the allocated memory for the KV cache. Once this limit is reached, the Least Recently Used (LRU) policy will begin evicting older sequences to make room for new ones. This helps prevent hard OOM crashes, but users of the evicted sequences will experience errors or forced re-computations when trying to continue their generation tasks.

Exam trap

Candidates often mistakenly believe the 0.8 memory limit triggers a system-wide shutdown or error, failing to recognize it as a threshold for the LRU eviction policy used in KV cache management.

62
Multi-Selecthard

An engineer is analyzing why a decoder-only LLM with 32,000-token context length fails to answer questions that require information from the beginning of a long document when the answer is near the end. The model was trained with standard causal attention. Which two architectural or training factors are most likely contributing to this failure? (Choose two.)

Select 2 answers
A.The feed-forward network in each layer compresses the hidden state, discarding information from early tokens.
B.The vocabulary size is too small to represent the document's domain-specific terms, causing tokenization errors.
C.The positional encoding scheme may not generalize well beyond the sequence lengths seen during pretraining.
D.Causal attention masks prevent each token from attending to future tokens, which limits bidirectional reasoning over the document.
E.The model was pretrained primarily on sequences much shorter than 32,000 tokens, so it did not learn to attend across the full context.
AnswersC, E

Many positional encoding schemes, especially learned absolute embeddings or RoPE without scaling, degrade when sequences exceed the lengths seen during training. The model cannot reliably distinguish or weight positions far beyond its training range, so attention to early tokens becomes imprecise. This directly contributes to the failure to retrieve information from the beginning of a long document.

Why this answer

Long-context retrieval failures in decoder-only LLMs typically stem from two sources: pretraining on sequences much shorter than the target context, which prevents the model from learning to attend across the full window, and positional encoding schemes that do not generalize beyond training lengths. Together these cause the model to underuse information at the start of a long document. The causal mask, feed-forward network, and vocabulary size are not primary causes of this behavior.

Exam trap

The trap here is blaming the causal attention mask for limiting long-context reasoning, when causal masking is a defining feature of decoder-only models and does not prevent attending to earlier tokens.

63
MCQmedium

You are preparing a 2 TB corpus of English and German web text for continued pretraining of a NeMo-based LLM. The German portion includes many pages with unescaped HTML entities and mixed-language sentences. Which NeMo Curator stage should you apply to remove boilerplate, fix HTML artifacts, and filter low-quality documents before tokenization?

A.NeMo Curator's text cleaning and heuristic filtering pipeline, including HTML unescaping, language identification, and quality classifier stages.
B.NeMo's Megatron-LM pretraining script with an increased dropout rate on the embedding layer.
C.NeMo Guardrails with a custom Colang flow that blocks documents containing HTML tags.
D.NVIDIA Triton Inference Server with a Python backend that preprocesses each document at inference time.
AnswerA

NeMo Curator provides modular stages for exactly this: HTML unescaping, boilerplate removal, language identification using fastText, and quality filtering with a classifier or heuristic scores. Running these before tokenization ensures the German and English subsets are clean and consistently language-tagged, which reduces noise during continued pretraining and prevents mixed-language documents from degrading the model's language modeling.

Why this answer

NeMo Curator is the correct tool because it supplies production-grade stages for HTML unescaping, language identification, and quality filtering that operate on large-scale text before tokenization. These stages directly address the German/English mixed-language and malformed HTML issues. The other options are training, runtime safety, or serving components, none of which clean a pretraining corpus.

Exam trap

The trap here is assuming that any NVIDIA component that processes text, such as Guardrails or Triton, can substitute for a dedicated data curation pipeline when preparing a training corpus.

64
MCQmedium

You are evaluating a fine-tuned Llama-3-70B model for a customer service chatbot using NVIDIA NeMo Evaluation. The model's outputs are factually correct but often verbose, exceeding the desired response length. Which metric should you prioritize to quantify this issue?

A.Average response length
B.BLEU
C.Perplexity
D.ROUGE-L
AnswerA

Average response length directly measures the mean number of tokens or words in generated responses. Since the issue is verbosity, this metric quantifies the problem precisely. It can be computed easily within NeMo Evaluation and compared against a target threshold to guide fine-tuning adjustments.

Why this answer

The core issue is that responses are too long, so the evaluation must quantify length. Average response length provides a direct, interpretable measure of verbosity. Other metrics like ROUGE-L, BLEU, and perplexity assess content overlap or fluency, not length, and thus cannot reliably diagnose the problem.

Exam trap

The trap here is assuming that standard text generation metrics like BLEU or ROUGE automatically penalize verbosity, when they primarily measure n-gram overlap or recall.

65
Multi-Selecthard

A team is fine-tuning a 70B-parameter LLM with NVIDIA NeMo on a multi-node cluster and wants to reduce the memory footprint per GPU without changing the model architecture. They are already using mixed precision and a reasonable micro-batch size. Which two techniques should they apply? (Choose two.)

Select 2 answers
A.Enable ZeRO-style optimizer state partitioning across data-parallel ranks.
B.Switch the optimizer to SGD with momentum and remove weight decay.
C.Activate activation recomputation so intermediate activations are discarded and recomputed during the backward pass.
D.Reduce the global batch size by a factor of eight and keep the micro-batch size the same.
E.Increase the number of attention heads to spread the computation across more GPUs.
AnswersA, C

ZeRO-style partitioning shards optimizer states, and optionally gradients and parameters, across data-parallel ranks. Each GPU then holds only a fraction of the optimizer memory, which substantially reduces per-GPU footprint for a 70B model. This directly addresses the memory constraint without altering the architecture and is a standard technique in large-scale NeMo training.

Why this answer

Partitioning optimizer states across data-parallel ranks and recomputing activations during the backward pass both reduce per-GPU memory without changing the model architecture. The first shrinks optimizer memory, and the second shrinks activation memory, which together address the dominant contributors for a 70B model. Changing attention heads, switching optimizers, or shrinking the global batch do not achieve the same effect.

Exam trap

The trap here is confusing batch size reduction with memory reduction, when per-GPU memory during forward and backward passes is largely determined by model states and activations rather than the global batch size.

66
MCQhard

A team fine-tunes an NVIDIA NeMo model to classify support tickets into five categories. In production, the model sometimes outputs free-form explanations instead of a single category label, breaking the downstream parser. Which prompt engineering change MOST reliably constrains the output format?

A.Add a polite request asking the model to 'try to keep answers short and to the point' while leaving the output format unspecified.
B.Increase max_tokens so the model has more room to explain its reasoning before stating the final category label.
C.Lower the temperature to zero and trust that deterministic sampling will force the model to emit only a category label.
D.Add three few-shot examples that each end with the exact category label, and instruct the model to output only the label with no additional text.
AnswerD

Few-shot examples demonstrate the precise output pattern, and the explicit instruction forbids extra text. Because the model conditions on the demonstrated format, it strongly biases toward emitting only a label. This combination is the most reliable prompt-level method to enforce a strict output contract without changing decoding or adding post-processing.

Why this answer

Strict output contracts are best enforced by showing the exact desired format through few-shot examples and explicitly prohibiting any additional text. Demonstrations act as in-context conditioning that shapes the model's continuation pattern, while the instruction closes the loophole of adding commentary. Sampling parameters control randomness, not structure.

Exam trap

The trap here is believing that temperature zero guarantees a clean label-only output, when format compliance is determined by prompt instructions and examples rather than by decoding settings.

67
MCQeasy

What is the primary function of data 'normalization' in the context of preparing inputs for a Transformer model?

A.Increasing the complexity of the input text to help the model learn more features.
B.Standardizing the text to ensure consistent interpretation of tokens.
C.Encrypting the dataset to protect sensitive information during training.
D.Compressing the dataset to reduce storage space on the GPU disk.
AnswerB

Normalization removes variations in formatting, such as whitespace or character encoding, that do not carry semantic weight. This consistency is vital, as it ensures that the tokenizer maps the same concepts to the same token IDs, preventing unnecessary ambiguity and ensuring the model learns stable, reliable relationships between tokens.

Why this answer

Normalization, such as standardizing whitespace, handling special tokens, and ensuring consistent character encodings, ensures that the model interprets input text in a predictable manner. By removing noise that doesn't contribute to semantic meaning, the model can focus its capacity on learning complex linguistic patterns. This is a standard and essential step in any high-performance AI data pipeline using NVIDIA accelerated computing.

Exam trap

Candidates often confuse text normalization with tokenization or embedding generation, failing to realize that normalization strictly standardizes raw text characters and whitespace before token processing begins.

68
MCQhard

A team is building an instruction-tuning dataset in NeMo from 40,000 internal support tickets. Each ticket contains a customer problem and a resolved answer, but the resolution text sometimes includes the customer's name, account number, and internal case IDs. The team plans to use NeMo Curator to produce training-ready JSONL. Which approach best prepares this data for instruction tuning while limiting personally identifiable information exposure?

A.Hash the entire ticket text with a cryptographic digest and train on the hashes, since the model can learn the mapping between hashed problems and hashed answers without ever seeing the original identifiers.
B.Run a PII redaction stage that detects and replaces names, account numbers, and case IDs with placeholders, then convert problem-resolution pairs into the instruction, input, and output schema before export.
C.Convert the tickets into instruction, input, and output JSONL first, then fine-tune the model and rely on a post-training output filter to block any generated response that contains an account number or customer name.
D.Drop every ticket whose resolution contains any digit, since account numbers and case IDs always contain numeric characters, and train only on the remaining text-only resolutions.
AnswerB

Redacting identifiers before schema conversion prevents the model from memorizing sensitive strings and keeps the pipeline deterministic. NeMo Curator supports PII detection and replacement as a distinct stage, and converting cleaned records into the instruction, input, and output fields yields the JSONL format the fine-tuning configuration expects. Doing redaction first also avoids leaking identifiers into derived fields.

Why this answer

Sensitive identifiers must be removed from the corpus before training, not mitigated after the fact. Running a PII detection and replacement stage in NeMo Curator, then mapping problem-resolution pairs into the instruction, input, and output schema, produces compliant JSONL that the fine-tuning job can consume directly. This ordering keeps identifiers out of both the training data and the resulting model weights.

Exam trap

The trap here is treating PII handling as a post-training output filter, when the identifiers must be removed from the training corpus itself before the model ever sees them.

69
MCQmedium

Which optimization technique specifically helps to manage the memory bandwidth bottleneck during the autoregressive decoding phase of an LLM?

A.Operator Fusion
B.KV Cache Quantization
C.Weight Pruning
D.Dynamic Batching
AnswerB

Quantizing the KV cache reduces the size of the data that must be read from memory for every single token generated. By using 8-bit or 4-bit representations for the cache, the system significantly decreases the memory bandwidth requirement, allowing for faster generation and higher concurrency on the same hardware.

Why this answer

During decoding, the model must read all KV cache values for every generated token. This is memory-bandwidth bound. Techniques like Quantized KV Cache (reducing the precision of cached tokens) and PagedAttention significantly reduce the amount of data moved between the VRAM and the compute units, effectively alleviating the memory bottleneck and allowing for faster token generation rates.

Exam trap

Candidates often select general model quantization, overlooking that the KV cache is a specific, massive memory bottleneck in LLM decoding that requires specialized cache-specific quantization techniques to resolve.

70
MCQmedium

What is the primary benefit of deploying a model with a 'Model Ensemble' configuration in Triton Inference Server?

A.It increases the number of GPUs available to the model.
B.It enables multi-model pipelining with reduced network overhead.
C.It automatically converts models into TensorRT engines.
D.It provides a load-balancing mechanism across server clusters.
AnswerB

Ensembles allow multiple inference tasks, such as text pre-processing and LLM generation, to occur sequentially on the server. By chaining these models together, the client only needs to send one request, avoiding the latency and security risks associated with multiple network round-trips for each part of the pipeline.

Why this answer

Model ensembles allow developers to chain multiple models together into a single pipeline (e.g., preprocessing, inference, and post-processing). This keeps the logic within the server, reducing the overhead of multiple network round-trips between the client and the server. By managing the entire pipeline as a single unit, you ensure consistent data flow and lower overall latency for the end-to-end application.

Exam trap

Candidates frequently assume ensembles are primarily for predictive accuracy or model voting, missing the architectural benefit of reducing network latency through server-side pipelining of preprocessing and post-processing steps.

71
MCQmedium

Why is 'Pinned Memory' (page-locked) essential for high-performance data transfers between host and GPU?

A.It provides larger memory capacity on the GPU.
B.It allows direct DMA access to host RAM.
C.It enables automatic data compression.
D.It eliminates the need for CUDA contexts.
AnswerB

Pinned memory is locked in physical RAM, allowing the GPU to perform direct memory access (DMA) transfers without host CPU involvement. This bypasses the overhead of copying data through temporary buffers, leading to vastly improved bandwidth and lower latency for transfers between the host and GPU.

Why this answer

Pinned memory prevents the operating system from swapping data to disk, allowing the GPU to access host memory directly via DMA (Direct Memory Access). This avoids the need for the driver to copy data into intermediate buffers, which is a major source of latency in standard data pipelines. By using pinned memory, applications can achieve significantly higher transfer speeds, which is vital for real-time generative AI applications.

Exam trap

Candidates frequently assume pinned memory increases GPU compute speed directly, rather than understanding that it primarily optimizes the efficiency of data transfer between host RAM and GPU VRAM.

72
MCQeasy

What is the primary role of an inference 'calibrator' when converting a model to INT8 precision?

A.It prunes zero-value weights
B.It calculates the optimal quantization scaling factors
C.It re-trains the model for higher accuracy
D.It optimizes the GPU kernel execution path
AnswerB

The calibrator processes a representative dataset to find the best range for quantization. It calculates the scaling factors that map the FP32 distribution into the INT8 range, minimizing information loss. This is the core function of the calibration step in the post-training quantization pipeline for TensorRT.

Why this answer

The calibrator identifies the optimal range for mapping floating-point activations to 8-bit integers without significant loss of information. It analyzes the distribution of input data to determine the scaling factors. Because INT8 has a limited dynamic range compared to FP32, the calibrator is essential to prevent overflow or underflow and to ensure that the quantized model remains accurate for the target inference tasks.

Exam trap

Test-takers frequently assume the calibrator alters model weights during training or performs the entire quantization process, missing its specific role in generating scaling factors for FP32-to-INT8 conversion.

73
Multi-Selectmedium

A media company runs an NVIDIA NIM-hosted content assistant that drafts articles from user prompts. Legal has flagged two risks: the model reproducing long verbatim passages from copyrighted training sources, and the model generating defamatory statements about named private individuals. Which two controls best address these specific risks? (Choose two.)

Select 2 answers
A.Enable request-level rate limiting per user account to reduce the volume of generated articles entering the editorial pipeline.
B.Raise the model's temperature and top-p values so generations vary more and are less likely to match any single source.
C.Deploy a retrieval-based similarity check that compares generated output against a licensed corpus index and blocks or rewrites spans that exceed a verbatim-overlap threshold.
D.Add a named-entity and defamation classifier in the output rail that flags assertions about private individuals and routes them to human review before publication.
E.Restrict the assistant to a smaller parameter model fine-tuned only on the company's own published archive.
AnswersC, D

Copyright regurgitation is a measurable overlap problem, so comparing generations against an indexed corpus of known works and blocking high-similarity spans directly targets the first risk. This is the mechanism behind output-side copyright filters and it produces a concrete, auditable threshold rather than a vague policy statement. Because the check runs on the generated text, it catches memorized passages regardless of how the prompt elicited them.

Why this answer

The two flagged risks require output-side detection tailored to each: a similarity index against licensed works catches verbatim copyright reproduction, and a named-entity plus defamation classifier with human escalation catches risky claims about private individuals. Sampling changes, model swaps, and rate limiting all operate on the wrong layer and leave one or both risks unmitigated. Layered output controls matched to the specific harm are what the legal review requires.

Exam trap

The trap here is reaching for a generation-side setting such as temperature or a smaller model when the flagged harms are detectable only after the text is produced.

74
MCQhard

A research team is evaluating a large language model's ability to follow instructions. They have a dataset of prompts with corresponding reference outputs. They want to use an automated metric that correlates well with human judgments of instruction-following quality. Which evaluation method is most suitable?

A.BLEU score against reference outputs
B.GPT-4-based evaluation with a detailed rubric
C.Perplexity of the model on the reference outputs
D.ROUGE-L score against reference outputs
AnswerB

Using a strong LLM like GPT-4 as a judge with a detailed rubric has been shown to correlate well with human judgments for instruction-following tasks. The rubric can specify criteria such as adherence to format, constraints, and correctness. This method captures nuanced aspects that n-gram metrics miss, making it the most suitable for this scenario.

Why this answer

LLM-based evaluation with a detailed rubric, such as using GPT-4 as a judge, has been demonstrated to align closely with human judgments for instruction-following. It can assess adherence to constraints, format, and correctness in a way that n-gram metrics like BLEU and ROUGE-L cannot. Perplexity measures fluency, not instruction adherence.

Therefore, the LLM-as-judge approach is the most suitable.

Exam trap

The trap here is assuming that reference-based n-gram metrics like BLEU or ROUGE-L can evaluate instruction-following, when they primarily measure surface overlap and miss nuanced adherence.

75
MCQeasy

When deploying a model using NVIDIA TensorRT, what is the primary benefit of the 'Engine Building' phase?

A.It converts the model to a generic portable format.
B.It automatically scales the model across multiple nodes.
C.It performs target-specific kernel selection and layer fusion.
D.It ensures the model can run on any CPU architecture.
AnswerC

The builder phase analyzes the network graph to merge redundant layers and select highly optimized CUDA kernels tailored to the specific GPU architecture. This approach maximizes hardware utilization, minimizes memory access patterns, and optimizes the execution flow to achieve peak performance compared to unoptimized, framework-native model execution.

Why this answer

The TensorRT builder phase analyzes the model graph and hardware topology to select the most efficient kernels for the target GPU. This process includes layer fusion, precision calibration, and kernel selection, which are vital for production-grade inference. By tailoring the model specifically to the underlying hardware architecture, TensorRT achieves significantly higher throughput and lower latency than executing generic framework-native code directly on the GPU.

Exam trap

Candidates often confuse the 'Engine Building' phase with the 'Inference' phase, incorrectly assuming it happens during runtime execution rather than as a pre-processing step to optimize the model graph for hardware.

Page 1 of 5

Page 2

All pages