Courseiva

NCP-GENL · domain

scenario questions

Practise NVIDIA Certified Professional: Generative AI LLMs scenario questions practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.

352 questions58 easy173 medium121 hard

Focused practice

Practice scenario questions questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about scenario questions

scenario questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Watch out for

Common scenario questions exam traps

  • ▸Answering from memory before reading the full scenario.
  • ▸Missing a constraint such as cost, availability, security, scope or command context.
  • ▸Choosing a broad answer when the question asks for the most specific fix.
  • ▸Ignoring why the wrong options are tempting.

Question index

All scenario questions questions (352)

Click any question to see the full explanation, or start a practice session above.

1

An enterprise fine-tunes a Llama-3-70B model using NVIDIA NeMo for automated technical support ticketing. The development team needs an automated evaluation pipeline that measures semantic similarity against human-curated reference answers without relying on costly human annotators. Which metric provides the most robust embedding-based semantic similarity assessment for this scenario?

Medium
2

An engineer is preparing an ensemble in NVIDIA Triton Inference Server that chains a Python preprocessing model with a TensorRT-LLM backend model. The preprocessing model must run on CPU while the LLM must run on GPU, and the ensemble must expose a single HTTP endpoint. Which configuration is required to make the ensemble execute correctly?

Easy
3

You are evaluating a text generation model using NVIDIA NeMo Evaluation and want to measure how well the generated text matches a reference translation. Which metric is specifically designed for this purpose?

Easy
4

A startup is deploying a small LLM for a chatbot on a single NVIDIA L4 GPU using NVIDIA Triton Inference Server. They want to ensure the model is automatically loaded when Triton starts and can be updated without restarting the server. Which Triton feature should they configure?

Easy
5

You are curating a 200 GB instruction-tuning corpus for an NVIDIA NeMo fine-tuning job on a Llama-based model. Post-training evaluation reveals the model regurgitates exact validation-set passages verbatim. An audit shows that near-duplicate instruction/response pairs were split randomly at the record level across train and validation partitions. Which data preparation change most directly eliminates this leakage while preserving the maximum amount of usable training data?

Medium
6

A team must serve a 70B model on a single 80 GB GPU for an internal assistant with modest concurrency. Full FP16 weights will not fit alongside the KV cache for the target context length. They want to keep accuracy loss minimal and are willing to spend additional build time. Which approach best fits these constraints?

Hard
7

A research team is evaluating a large language model's robustness to adversarial attacks. They want to use NVIDIA NeMo Evaluator to measure how often the model's output changes when small, semantically preserving perturbations are applied to input prompts. Which evaluation metric or method should they implement?

Hard
8

When preparing a proprietary technical manual dataset for a RAG pipeline, which data preprocessing step is most critical to ensure the LLM avoids hallucinations regarding specific product configurations?

Medium
9

An engineer is deploying an NVIDIA NeMo Guardrails system to moderate a chatbot's responses. The chatbot must refuse to answer questions about politics but should answer questions about weather. Which prompt engineering strategy in NeMo Guardrails is most appropriate to enforce this behavior?

Medium
10

A healthcare company is deploying an LLM-based patient triage assistant using NVIDIA NIM microservices on-premises. To comply with HIPAA, they need to ensure that no protected health information (PHI) is transmitted to external services. Which deployment approach best meets this requirement?

Medium
11

A developer is creating prompts for an NVIDIA NIM-hosted LLM to summarize financial reports. The reports are lengthy and contain many tables and figures. The developer wants to ensure the summaries are accurate and include key numerical data. Which TWO prompt engineering techniques should be applied? (Choose two.)

Medium
12

Why do many modern LLMs use SwiGLU as their activation function in the feed-forward network instead of the traditional ReLU?

Medium
13

Refer to the exhibit. The TensorRT build process fails with a memory limit error. Which configuration adjustment is most likely to resolve this build-time error?

Hard
14

A data scientist is using an NVIDIA NeMo LLM to generate Python code from natural language descriptions. The model often produces code that works but does not follow the team's style guide, such as using single quotes instead of double quotes and missing type hints. Which prompt engineering technique should the data scientist use to improve adherence to the style guide?

Easy
15

A team is pre-training a 13B-parameter decoder-only LLM on a cluster of NVIDIA GPUs. They observe that gradient norms spike sharply during the first few hundred steps, destabilizing training. They want to keep the standard post-layer-normalization placement but stabilize early optimization. Which architectural technique should they apply?

Medium
16

A team is training a large language model on 8 NVIDIA A100 GPUs using PyTorch's DistributedDataParallel (DDP). Profiling shows that all GPUs are frequently idle, waiting for gradient synchronization. The network interconnect between nodes is a 100 Gb Ethernet with TCP/IP, and the model has 13 billion parameters. What is the most effective optimization to reduce the idle time?

Medium
17

What is the architectural role of Layer Normalization in a Transformer, and where is it typically placed to ensure stable training?

Medium
18

When auditing an NVIDIA Triton deployment for security, which action is most critical to protect sensitive inference data?

Medium
19

A developer is preparing to fine-tune a model with NVIDIA NeMo and wants a quantitative baseline before training begins. They plan to score the base model on a 300-question multiple-choice reasoning set and report accuracy. Which evaluation setup gives the most defensible baseline number?

Easy
20

You are building the data preparation stage for an NVIDIA NeMo retrieval-augmented generation pipeline that will ingest millions of internal wiki pages. The ingestion team reports that the same policy text appears in dozens of pages with minor edits, and that some pages contain copied tables from external sources. You need to produce a clean, deduplicated chunk store that supports accurate citation and avoids returning redundant passages. Which TWO actions best address these requirements? (Choose two.)

Hard
21

An engineer is tasked with optimizing a model that performs poorly due to excessive memory access latency. Which TensorRT optimization strategy specifically targets this issue?

Medium
22

A developer is optimizing a BERT-based model for inference on an NVIDIA T4 GPU using TensorRT. The model has a fixed input sequence length of 128. Profiling shows that the kernel execution time is high due to many small operations. Which TensorRT feature should they use to reduce kernel launch overhead and improve latency?

Medium
23

A developer is inspecting a decoder-only Transformer and notices that during training, the model attends to future tokens in the sequence, causing the loss to drop unrealistically fast but generation to be incoherent. Which architectural mechanism is missing or misconfigured?

Easy
24

A team is fine-tuning a 13B-parameter Llama model with NVIDIA NeMo on a node of eight A100 80GB GPUs. They want the optimizer state to be partitioned across data-parallel ranks so that per-GPU memory drops, while keeping the model replicas synchronized. Which distributed strategy should they select in the NeMo training configuration?

Medium
25

A multinational insurer deploys an NVIDIA NIM-based claims triage assistant across the EU and Brazil. The compliance team must demonstrate that the system honors data-subject deletion requests and that personal data is not transferred outside approved regions. Which design decision addresses both obligations most directly?

Hard
26

An engineer is optimizing a large language model for inference on NVIDIA GPUs using TensorRT-LLM. They want to reduce the memory footprint of the KV cache to support longer context lengths and more concurrent requests. Which two techniques should they implement? (Choose two.)

Hard
27

A retail company wants to release an LLM-powered shopping assistant on NVIDIA NIM. Legal requires that the assistant never provide personalized financial advice, even if a user asks for it. Which control most directly enforces this boundary at runtime?

Easy
28

A team is deploying a large language model on NVIDIA Triton Inference Server with TensorRT-LLM backend. They want to reduce GPU memory consumption to fit a larger model on a single GPU without significantly degrading output quality. Which two techniques should they use? (Choose two.)

Medium
29

A healthcare startup is deploying a Mistral 7B model for internal clinical note summarization. They need to serve the model with NVIDIA Triton Inference Server and want to minimize GPU memory footprint during inference. The team plans to use TensorRT-LLM and is choosing a numerical precision for the engine. Which precision should they select to reduce memory usage while maintaining acceptable accuracy for summarization?

Easy
30

A team is preparing a Llama-based chatbot for production and wants to reduce GPU memory and latency without retraining. They decide to apply post-training quantization. Which TensorRT-LLM workflow correctly produces an INT8 or FP8 quantized engine from an existing FP16 checkpoint?

Easy
31

Which metric is the most reliable indicator that an LLM is overfitting during the fine-tuning phase?

Medium
32

An engineer must serve a 70B-parameter LLM for a workload with many concurrent users and long shared system prompts, and wants to maximize throughput without retraining. Which inference-time optimization most directly reduces redundant computation across requests sharing the same prompt prefix?

Hard
33

A team is using NVIDIA Triton Inference Server to serve multiple LLMs. They want to automatically detect when a model's inference latency degrades beyond acceptable thresholds and trigger an alert. Which Triton feature should they configure?

Medium
34

A team is deploying a 70B-parameter LLM across four NVIDIA H100 GPUs using NVIDIA TensorRT-LLM with tensor parallelism. They observe that inference works but throughput is lower than expected, and profiling shows significant inter-GPU communication overhead. Which optimization should they apply first to reduce communication overhead?

Hard
35

A financial firm is deploying a generative AI chatbot using NVIDIA NIM. To comply with strict data residency regulations, where must the inference and data processing occur?

Medium
36

Refer to the exhibit. A team is preparing log data for a RAG-based troubleshooting assistant. Given the configuration, what is the most significant risk during the retrieval phase?

Hard
37

An engineer is reviewing the attention implementation of a decoder-only LLM used for chat. During inference with a KV cache, generated tokens must not attend to future positions. Which mechanism enforces this constraint inside scaled dot-product attention?

Easy
38

An ML engineer is fine-tuning a 70B model with NVIDIA NeMo Framework across 16 H100 GPUs. Training completes successfully, but when the fine-tuned checkpoint is evaluated, outputs are incoherent and repeat tokens. The engineer confirms the loss decreased smoothly during training and the validation dataset was held out correctly. Which issue is the most likely explanation?

Hard
39

When preparing unstructured documentation for a high-performance retrieval system, which approach best balances index size and retrieval relevance?

Medium
40

When profiling an application with NVIDIA Nsight Systems, you notice a long gap between kernel execution blocks on the GPU timeline. What is the most likely cause?

Hard
41

A developer is using an NVIDIA NIM for a customer support chatbot. The chatbot must handle multi-turn conversations and maintain context about the user's issue. The developer notices that after several turns, the bot starts giving generic responses and forgets earlier details. Which prompt engineering approach is most effective to maintain context?

Medium
42

An enterprise is running a mission-critical generative AI application on an NVIDIA DGX cluster. The MLOps team notices occasional silent GPU memory corruption during long-running inference jobs that do not trigger hard crashes. Which monitoring tool and strategy should be utilized for early detection?

Hard
43

An engineer is profiling a CUDA kernel and notices that the achieved occupancy is low, leading to underutilization of the GPU. The kernel uses a large number of registers per thread, limiting the number of resident warps. Which optimization should be attempted first to improve occupancy?

Easy
44

A production LLM inference service on NVIDIA Triton Inference Server is being monitored for reliability. The team wants to implement effective logging to diagnose issues such as high latency and errors. Which TWO logging practices are recommended for a production LLM environment? (Choose two.)

Hard
45

You are preparing a dataset for continued pre-training of an LLM on internal engineering documents. The corpus contains many documents with boilerplate headers, footers, and legal disclaimers repeated across files. Which NeMo Curator approach best reduces this boilerplate while preserving unique technical content?

Hard
46

A team is deploying an NVIDIA NIM for a Llama 3 model as a retrieval-augmented generation (RAG) assistant over internal documentation. Users report that the assistant sometimes answers from its pretrained knowledge instead of the retrieved passages, and occasionally cites a passage that does not support its claim. Which TWO prompt engineering changes best reduce these behaviors? (Choose two.)

Hard
47

A team is preparing a supervised fine-tuning job in NVIDIA NeMo Framework for a customer-support assistant. They have a large corpus of raw support chat logs with no labels. They want the model to learn to answer customer questions in the company's tone and format. Which data preparation step is most appropriate before training?

Easy
48

Refer to the exhibit. What is the most likely reason for the high P99 latency despite low GPU utilization?

Hard
49

An enterprise is deploying a large language model on NVIDIA Triton Inference Server. Which deployment strategy minimizes latency for requests that require high-throughput batching while maintaining consistent hardware utilization?

Medium
50

Refer to the exhibit. An engineer notices that the TensorRT engine takes an excessively long time to build. What is the most likely cause, and how can it be mitigated?

Hard
51

In Mixture-of-Experts (MoE) architectures, why does the use of a router mechanism significantly impact performance compared to dense models?

Hard
52

A company wants to teach a pretrained LLM to follow a specific output format for customer support replies using supervised fine-tuning on NVIDIA GPUs. Which data preparation approach best matches supervised fine-tuning for instruction following?

Easy
53

You are preparing a customer-support dataset for fine-tuning an LLM with NVIDIA NeMo. The raw data includes personally identifiable information such as names, email addresses, and phone numbers. Which data preparation step must be performed before training to comply with privacy requirements?

Easy
54

An enterprise deployment team needs to deploy a Large Language Model on NVIDIA Triton Inference Server. They require the lowest possible latency for real-time inference while maximizing GPU memory utilization. Which configuration strategy should the team implement?

Medium
55

An engineer is using an NVIDIA NIM for a code generation model to produce Python functions from natural language descriptions. The model frequently generates code that uses deprecated libraries or incorrect function signatures. The engineer wants to improve the accuracy of the generated code by providing examples. Which prompting strategy is most appropriate?

Hard
56

Which metric provides the best indication of 'inference queue saturation' in a Triton deployment?

Medium
57

A team is designing prompts for an NVIDIA NIM-hosted LLM that must produce concise, citation-backed answers from retrieved documents. They want to improve factual grounding and reduce unsupported claims. Which two prompt engineering practices best support this goal? (Choose two.)

Hard
58

A team is deploying a 70B-parameter LLM using NVIDIA Triton Inference Server with TensorRT-LLM backend on a node with four A100 80GB GPUs. They observe that during inference, only one GPU is utilized while the others remain idle. They have configured the model with tensor parallelism set to 1. What is the most likely cause of this underutilization?

Hard
59

A data engineer is preparing a JSONL instruction dataset for an NVIDIA NeMo supervised fine-tuning run. Each line currently contains a free-form 'text' field with the instruction, context, and response concatenated. The training configuration expects the standard NeMo instruction-tuning schema with separate fields for the task instruction, optional context, and the expected response. What is the most appropriate data preparation step?

Easy
60

A team is evaluating a large language model for a question-answering system using NVIDIA NeMo Evaluation. They need to assess both the relevance of the answer to the question and its factual correctness. (Choose two.)

Hard
61

Refer to the exhibit. What is the implication of setting the memory_limit to 0.8 in the context of an LLM inference service?

Hard
62

An engineer is analyzing why a decoder-only LLM with 32,000-token context length fails to answer questions that require information from the beginning of a long document when the answer is near the end. The model was trained with standard causal attention. Which two architectural or training factors are most likely contributing to this failure? (Choose two.)

Hard
63

You are preparing a 2 TB corpus of English and German web text for continued pretraining of a NeMo-based LLM. The German portion includes many pages with unescaped HTML entities and mixed-language sentences. Which NeMo Curator stage should you apply to remove boilerplate, fix HTML artifacts, and filter low-quality documents before tokenization?

Medium
64

You are evaluating a fine-tuned Llama-3-70B model for a customer service chatbot using NVIDIA NeMo Evaluation. The model's outputs are factually correct but often verbose, exceeding the desired response length. Which metric should you prioritize to quantify this issue?

Medium
65

A team is fine-tuning a 70B-parameter LLM with NVIDIA NeMo on a multi-node cluster and wants to reduce the memory footprint per GPU without changing the model architecture. They are already using mixed precision and a reasonable micro-batch size. Which two techniques should they apply? (Choose two.)

Hard
66

A team fine-tunes an NVIDIA NeMo model to classify support tickets into five categories. In production, the model sometimes outputs free-form explanations instead of a single category label, breaking the downstream parser. Which prompt engineering change MOST reliably constrains the output format?

Hard
67

What is the primary function of data 'normalization' in the context of preparing inputs for a Transformer model?

Easy
68

A team is building an instruction-tuning dataset in NeMo from 40,000 internal support tickets. Each ticket contains a customer problem and a resolved answer, but the resolution text sometimes includes the customer's name, account number, and internal case IDs. The team plans to use NeMo Curator to produce training-ready JSONL. Which approach best prepares this data for instruction tuning while limiting personally identifiable information exposure?

Hard
69

Which optimization technique specifically helps to manage the memory bandwidth bottleneck during the autoregressive decoding phase of an LLM?

Medium
70

What is the primary benefit of deploying a model with a 'Model Ensemble' configuration in Triton Inference Server?

Medium
71

Why is 'Pinned Memory' (page-locked) essential for high-performance data transfers between host and GPU?

Medium
72

What is the primary role of an inference 'calibrator' when converting a model to INT8 precision?

Easy
73

A media company runs an NVIDIA NIM-hosted content assistant that drafts articles from user prompts. Legal has flagged two risks: the model reproducing long verbatim passages from copyrighted training sources, and the model generating defamatory statements about named private individuals. Which two controls best address these specific risks? (Choose two.)

Medium
74

A research team is evaluating a large language model's ability to follow instructions. They have a dataset of prompts with corresponding reference outputs. They want to use an automated metric that correlates well with human judgments of instruction-following quality. Which evaluation method is most suitable?

Hard
75

When deploying a model using NVIDIA TensorRT, what is the primary benefit of the 'Engine Building' phase?

Easy
76

A financial services company is fine-tuning an LLM to answer questions about internal policies. The base model performs well on general text but frequently invents policy numbers and effective dates. The team has a curated dataset of 5,000 question-answer pairs with correct citations. Which fine-tuning approach best addresses the hallucination of policy numbers and dates?

Hard
77

A team uses an NVIDIA NIM-hosted model to draft release notes from a changelog. Reviewers report the drafts omit minor fixes and overstate the significance of small changes. Which prompt engineering adjustment BEST addresses both issues?

Hard
78

A developer is preparing a supervised fine-tuning dataset for an instruction-tuned LLM using NVIDIA NeMo. The dataset contains prompts and responses, but the model sometimes learns to generate the prompt text as part of the response. Which dataset formatting practice should be applied to prevent this?

Easy
79

An engineer is optimizing a Transformer-based LLM for inference on an NVIDIA A100 GPU. The model uses FP16 precision, but during generation, the GPU's Tensor Cores are underutilized, and latency is higher than expected. Profiling reveals that many small matrix multiplications are executed sequentially. Which technique is most effective to improve Tensor Core utilization and reduce latency?

Hard
80

When building an NVIDIA NeMo LLM application for automated document review, which THREE of the following prompt design choices are critical for ensuring high-quality output? (Select exactly THREE)

Hard
81

An engineer is deploying a LLM using NVIDIA Triton Inference Server with the TensorRT-LLM backend. They need to ensure that the model can handle a sudden surge in requests without increasing latency beyond a specified threshold. They have configured the model with a maximum batch size of 32 and dynamic batching with a preferred batch size of 16. However, during peak load, latency spikes are observed. Which Triton configuration parameter should they adjust to control the maximum time a request waits in the dynamic batching queue before being processed?

Hard
82

When fine-tuning on a small, domain-specific dataset, why might adding synthetic data generated by a larger model be beneficial?

Hard
83

What is the primary function of the 'Triton Model Analyzer' in an optimization workflow?

Easy
84

You are curating a 2 TB corpus of NVIDIA technical documentation and Python code for continued pretraining of a NeMo-based LLM. A colleague proposes filtering out any document containing the token sequence 'CUDA' to reduce hardware-specific bias. What is the most appropriate response?

Medium
85

A team is building a TensorRT-LLM engine for a 7B model that must serve both single-turn short prompts and long multi-turn conversations with a shared system prompt. They want to maximize reuse of computation across requests without changing model weights. Which TWO techniques should they enable? (Choose two.)

Medium
86

An LLM inference service on NVIDIA Triton Inference Server is experiencing intermittent failures. The operations team wants to set up alerting to detect when the GPU memory utilization exceeds 90% for more than 5 minutes, as this could lead to out-of-memory errors. Which combination of tools should they use to achieve this?

Medium
87

A production LLM service on NVIDIA Triton Inference Server is deployed across multiple GPUs. The team notices that one GPU consistently shows higher latency for inference requests compared to others, despite similar utilization. Which NVIDIA tool should be used to investigate per-GPU performance discrepancies and identify bottlenecks?

Hard
88

What is the primary function of the 'NeMo Guardrails' toolkit in an enterprise AI pipeline?

Easy
89

A developer is building a retrieval-augmented generation pipeline and needs to choose a component that produces dense vector representations of passages for semantic search. The passages are up to 512 tokens long, and the developer wants a model specifically trained to map semantically similar text to nearby points in embedding space. Which type of model should be selected?

Easy
90

You are deploying a large language model on NVIDIA Triton Inference Server in a Kubernetes cluster. To ensure high availability and reliability, which TWO practices should you implement? (Choose two.)

Medium
91

When training a model for a highly technical domain with a scarcity of high-quality data, which data augmentation strategy is most likely to preserve the model's reliability?

Hard
92

You are optimizing a ResNet-50 model on an NVIDIA A100. Which precision-based optimization will yield the highest throughput without significant accuracy loss?

Medium
93

Refer to the exhibit. Given this NeMo configuration, which prompt modification would best improve the reliability of technical support queries?

Medium
94

A team is training a large language model and wants to reduce the memory used by the optimizer without changing the model architecture. They are using Adam and notice that optimizer state consumes more GPU memory than the model weights. Which technique should they apply to reduce optimizer memory while keeping the model architecture unchanged?

Medium
95

When profiling an application with NVIDIA Nsight Systems, which TWO metrics are most critical to identify if an application is limited by the PCIe bus?

Hard
96

What is the primary advantage of using a 'Quantization Aware Training' (QAT) approach over post-training quantization for LLMs?

Medium
97

A financial analyst is using an NVIDIA NIM-hosted Llama 3.1 70B model to extract key financial metrics from quarterly earnings call transcripts. The model inconsistently returns a prose summary instead of the required structured JSON. The analyst needs the output to be reliably parseable by a downstream script that expects a fixed schema with fields "revenue", "eps", and "guidance". Which prompt engineering technique is most appropriate to enforce this output format?

Medium
98

A team is evaluating a retrieval-augmented generation (RAG) pipeline using NVIDIA NeMo Evaluation. They notice that the generated answers are fluent but sometimes contradict the retrieved documents. Which evaluation approach best identifies this issue?

Hard
99

Which metric is most critical to monitor for identifying 'bottlenecks' in a high-throughput LLM deployment?

Easy
100

A media company is using a large language model to generate news article headlines. They want to evaluate the diversity of the generated headlines to ensure they are not overly repetitive. Which metric should they use to quantify the lexical diversity of the generated headlines?

Medium
101

An evaluation team is comparing two candidate checkpoints of the same fine-tuned model on a 400-prompt open-ended task. They use an LLM-as-judge that returns a pairwise preference for each prompt. Candidate X wins 214 times, candidate Y wins 152 times, and 34 are ties. The judge is the same family as the models being compared. What should the team do before treating candidate X as the winner?

Hard
102

A media-analytics firm serves a 13B-parameter summarization model on two A100 GPUs using NVIDIA TensorRT-LLM behind Triton Inference Server. Traffic is bursty: during live events concurrency triples for about ten minutes, then returns to baseline. Operators report that the first requests after each burst begin are slow and sometimes time out, although steady-state latency is acceptable. Which deployment change most directly addresses the cold-start penalty at the beginning of each burst?

Hard
103

Which TWO of the following are benefits of using Rotary Positional Embeddings (RoPE) compared to absolute positional embeddings?

Medium
104

An enterprise deploys an LLM application using NVIDIA NeMo Guardrails to prevent the generation of toxic content and PII leakage. During red-teaming, testers discover that prompt injection attacks successfully bypass standard input rails by encoding malicious instructions in Base64 format inside conversational context. Which architectural approach provides the most robust mitigation against this evasion technique while maintaining conversational latency requirements?

Medium
105

A government agency deploys an NVIDIA NIM-based assistant to help citizens understand benefit eligibility. An oversight panel demands that any refusal to answer be explainable and consistent, and that the assistant not improvise policy. Which design best meets these requirements?

Hard
106

A technical support team is building a chatbot using an NVIDIA NIM microservice. The chatbot must answer questions about a specific product's warranty policy. The team wants to ensure the model's responses are grounded in the official warranty document, which is 50 pages long, and avoid inventing policy details. Which prompt engineering approach is most effective for this scenario?

Easy
107

A developer is writing a system prompt for an NVIDIA NIM-hosted assistant that must always respond in formal English, never use slang, and never reveal internal system instructions. Where should these persistent behavioral rules be placed for the MOST consistent effect?

Easy
108

Why is it important to use a 'warm-up' period in the learning rate schedule when starting a fine-tuning job?

Medium
109

An LLM application deployed on NVIDIA Triton Inference Server is experiencing intermittent latency spikes. Which metric is most critical to monitor to determine if the issue stems from GPU compute saturation?

Medium
110

A retail company is deploying an LLM-based customer service assistant using NVIDIA NIM. The legal team mandates that the model must not generate content that violates copyright, such as reproducing song lyrics or book excerpts. Which NVIDIA offering should the team use to enforce this policy at runtime?

Easy
111

Which NVIDIA library is primarily used for optimizing and deploying deep learning inference models?

Easy
112

Refer to the exhibit. The Triton Inference Server is failing to load a model. The configuration shows two instances assigned to GPU 0. What is the most likely cause and the correct remediation?

Hard
113

An LLM inference service running on NVIDIA Triton Inference Server is configured with model ensembles. During production monitoring, the operations team notices that the end-to-end latency reported by the client is significantly higher than the sum of individual model latencies reported by Triton's metrics. Which Triton feature should be investigated to identify the source of the additional latency?

Hard
114

Which THREE of the following are primary benefits of using PagedAttention in NVIDIA TensorRT-LLM deployments?

Medium
115

A developer is using an NVIDIA NIM for a Llama 3.1 70B model to build a legal document review assistant. The model must answer questions based on a provided contract, but the contracts are often 50,000 tokens long, exceeding the model's 8,000-token context window. Which prompt engineering strategy is most appropriate to handle this constraint?

Hard
116

An engineer is designing prompts for an NVIDIA NIM-hosted model that must extract structured fields from unstructured invoices. The extraction accuracy is inconsistent across vendors with different layouts. Which TWO prompt engineering techniques would MOST improve reliability? (Choose two.)

Medium
117

What is the primary role of the 'Learning Rate Scheduler' during LLM fine-tuning?

Medium
118

Which TWO of the following telemetry types are essential for detecting 'model drift' in a production LLM deployment?

Medium
119

A healthcare company deploys an LLM-powered clinical documentation assistant using NVIDIA NIM microservices on-premises. During a compliance review, auditors discover that the model occasionally generates patient names and medical record numbers in its output even though these were not present in the input prompt. The team needs to implement a runtime guardrail that detects and blocks such unintended PII leakage without retraining the model. Which NVIDIA component should they configure to add this output-side detection?

Medium
120

You are using NVIDIA NeMo Evaluation to assess a summarization model. The model produces summaries that are grammatically correct but omit key information from the source. Which metric should you use to quantify the amount of missing content?

Medium
121

A developer is building a customer support assistant using an NVIDIA NIM microservice for a Llama 3.1 8B Instruct model. The assistant must answer questions about an order solely based on a JSON payload containing order details, and it must not use any outside knowledge. Which prompt engineering approach best ensures the model adheres to this constraint?

Medium
122

A team is pretraining a 13B-parameter decoder-only LLM on English text using byte-pair encoding with a 50,000-token vocabulary. They observe that the model produces fluent but repetitive continuations and that the average log-probability assigned to ground-truth tokens plateaus early. The training loss curve shows the model is underfitting rather than overfitting. Which architectural change is most likely to improve the model's capacity to capture long-range dependencies?

Medium
123

Which THREE actions are essential for maintaining a secure and compliant LLM deployment according to the NVIDIA security guidelines?

Hard
124

Refer to the exhibit. The inference server reports an OOM error while running multiple LLMs. What configuration change most effectively improves reliability without upgrading the hardware?

Hard
125

When fine-tuning a base LLM using Parameter-Efficient Fine-Tuning (PEFT) on an NVIDIA H100 GPU, what is the primary advantage of utilizing LoRA compared to full fine-tuning?

Medium
126

An ML engineer is fine-tuning a 7B-parameter model with LoRA on a single NVIDIA A100 40GB GPU. The training script reports that the adapter weights are not updating after several hundred steps, and the loss remains flat. The base model weights are frozen as intended. Which LoRA configuration issue is the most likely cause?

Medium
127

A team is deploying a 70B parameter LLM with NVIDIA Triton Inference Server across four NVIDIA H100 GPUs. They are using TensorRT-LLM and need to fit the model within the combined GPU memory while maintaining high throughput. Which two techniques should they use? (Choose two.)

Hard
128

Which technique is most effective for mitigating data leakage during the training of an LLM on time-series-related document data?

Easy
129

Refer to the exhibit. The model performance is inconsistent. What is the most likely reason for the performance variability under load?

Medium
130

A team is deploying a Llama 2 13B model with NVIDIA TensorRT-LLM on a single A100 40GB GPU. They need to serve 32 concurrent requests with a maximum sequence length of 4096 tokens. They observe that the GPU runs out of memory during inference. Which configuration parameter should they adjust to control the maximum GPU memory allocated for the KV cache?

Medium
131

An engineer fine-tunes a model on a domain corpus with NVIDIA NeMo and observes that training loss falls steadily while validation loss begins rising after the second epoch. The team must produce the most generalizable checkpoint without changing the dataset. Which action should they take?

Hard
132

A healthcare technology company is deploying an LLM-powered patient triage assistant using NVIDIA NIM microservices on-premises. During an internal audit, the compliance team discovers that the model occasionally outputs patient names and medical record numbers (MRNs) in its responses, even though the training data was scrubbed. The company must implement a runtime safeguard that detects and redacts PII before the response reaches the user. Which NVIDIA component should they integrate into their inference pipeline to achieve this?

Medium
133

Which hardware architecture feature is specifically leveraged by TensorRT to accelerate FP16 and INT8 matrix multiplications?

Medium
134

An ML engineer is evaluating Mixture-of-Experts (MoE) routing for a large decoder-only model to increase capacity without proportionally increasing compute per token. Which TWO statements accurately describe how top-k token routing behaves in such an architecture? (Choose two.)

Hard
135

You are evaluating a fine-tuned LLM for a code generation task. The model was trained using NVIDIA NeMo on a dataset of Python functions. You want to measure the percentage of generated functions that pass a set of unit tests. Which evaluation metric is most appropriate?

Medium
136

Refer to the exhibit. Why did the system fail to trigger an alert despite high latency?

Hard
137

An engineer is deploying a large language model for real-time inference on an NVIDIA A100 GPU. The model uses FP16 weights but the inference server must handle variable-length input sequences. Profiling shows that the GPU spends significant time on memory-bound operations and that kernel launches are frequent. Which optimization is most appropriate to reduce latency while maintaining accuracy?

Hard
138

A team is optimizing a large language model for inference on NVIDIA GPUs. They want to reduce the memory footprint of the model to fit on a single GPU with limited VRAM. Which two techniques are most appropriate for reducing memory usage during inference? (Choose two.)

Medium
139

A developer is building an interactive assistant using NVIDIA NIM microservices. The assistant must answer questions about a specific set of internal policies. The developer wants to ensure the model's responses are grounded in those policies and not in its general pre-training knowledge. Which prompt engineering technique should be applied?

Easy
140

You are preparing a large instruction-tuning dataset for an NVIDIA NeMo-based LLM. The raw data consists of user queries and assistant responses collected from a customer support system, stored as JSON lines. During preprocessing, you notice that many responses contain personally identifiable information (PII) such as names, email addresses, and phone numbers. You need to ensure the dataset is safe for training while preserving as much semantic content as possible. Which approach is most appropriate for handling PII in this dataset?

Medium
141

A healthcare analytics company is deploying an LLM-based patient triage assistant using NVIDIA NIM microservices. Compliance requires that every model response be traceable to a specific model version, input prompt, and retrieved context for a minimum of three years. Which approach best satisfies this auditability requirement?

Medium
142

A team is training a large language model using NVIDIA DGX A100 nodes with 8 GPUs per node. They observe that GPU utilization is high on all GPUs, but the training throughput scales poorly when adding more nodes. Profiling shows that the communication time during all-reduce operations increases significantly with node count. Which of the following optimizations is most likely to improve scaling efficiency?

Medium
143

A developer is using NVIDIA TensorRT-LLM to optimize a GPT-based model for inference. They want to reduce the model's memory footprint and improve throughput without retraining. Which two techniques can be applied during the TensorRT-LLM build process to achieve these goals? (Choose two.)

Hard
144

What is the primary benefit of using NVIDIA DCGM (Data Center GPU Manager) for monitoring production LLMs?

Medium
145

A healthcare company is deploying an LLM for clinical note summarization using NVIDIA Triton Inference Server. They must ensure that only authorized users can access the model and that all inference requests are logged for audit. Which Triton feature should they configure to enforce authentication and authorization?

Medium
146

When implementing Chain-of-Thought (CoT) prompting for a complex NVIDIA NeMo-based reasoning task, what is the primary benefit of encouraging the model to generate intermediate steps?

Medium
147

Which TWO of the following practices are recommended for ensuring ethical AI development when using NVIDIA NIMs in an enterprise environment?

Medium
148

A team is fine-tuning a pretrained decoder-only model on a small domain-specific dataset. They observe that the model quickly overfits and loses general language ability. They want to update only a small number of additional parameters while keeping the base weights frozen. Which approach should they use?

Hard
149

An engineer needs to expose an LLM served by NVIDIA Triton Inference Server to a web application over HTTP with token streaming. Which Triton feature should they enable?

Easy
150

When assessing the quality of a dataset for instruction fine-tuning, which TWO metrics or methods are considered most reliable for measuring dataset diversity?

Medium
151

A team is using an NVIDIA NIM for a Mistral model to classify support tickets into one of five fixed categories. Accuracy is inconsistent, and the model sometimes invents new categories. Which prompt engineering change is most likely to improve reliability without retraining the model?

Medium
152

A team is serving a 13B-parameter LLM on a single NVIDIA A100 80GB GPU. During generation, they observe that the GPU compute utilization stays below 20% while memory bandwidth utilization is near saturation. They want to improve throughput without changing the model architecture. Which optimization is most appropriate?

Medium
153

An inference team runs a TensorRT-LLM model behind NVIDIA Triton Inference Server with the tensorrtllm_backend. They notice that when clients send requests with widely varying prompt lengths, throughput drops and some requests wait much longer than others. They want Triton to group compatible requests dynamically to improve GPU utilization without changing the model build. Which Triton feature should they configure?

Hard
154

You are preparing a multilingual corpus for pre-training an LLM. The corpus contains documents in English, Spanish, and German, but the English portion is 80% of the data. You want to avoid the model becoming biased toward English. Which data preparation technique should you apply?

Medium
155

A company is running a large language model inference service on NVIDIA GPUs. They observe that GPU memory is nearly full, limiting the batch size and thus throughput. The model weights are stored in FP16, and the KV cache consumes a significant portion of memory. Which technique can reduce memory usage while maintaining model accuracy and enabling larger batch sizes?

Medium
156

A team is testing a new LLM application. During red-teaming, the model consistently leaks sensitive internal project codenames. How should the team address this systematically?

Medium
157

A production LLM inference service on NVIDIA Triton Inference Server is experiencing intermittent slowdowns. The team wants a single metric that best indicates whether the GPU is the bottleneck during these periods. Which metric should they monitor?

Easy
158

A team is pretraining a decoder-only Transformer LLM on a large corpus of code and natural language. They observe that the model's training loss decreases smoothly, but during generation it sometimes produces degenerate repetition, and attention entropy on long sequences collapses. They suspect the issue is related to the positional encoding scheme. Which architectural change is most likely to mitigate the attention entropy collapse while preserving the model's ability to generalize to sequences longer than those seen during pretraining?

Medium
159

Which TWO of the following statements correctly describe the role of Grouped Query Attention (GQA) in modern LLM architectures?

Medium
160

Refer to the exhibit. An engineer is fine-tuning an LLM using the provided configuration. What is the primary purpose of applying 'yarn' scaling in this architecture?

Hard
161

A developer has a fine-tuned Llama-family model in Hugging Face format and wants to run it with NVIDIA TensorRT-LLM on an H100. Which artifact must be produced before the runtime can execute the model?

Easy
162

A developer is building a customer support assistant using an NVIDIA NIM microservice for a Llama 3 model. The assistant must always respond in valid JSON with keys 'category' and 'urgency'. The model often returns conversational text instead. Which prompt engineering change most directly enforces the required output format?

Easy
163

You are using NVIDIA NeMo Curator to filter a 600 GB web-crawl corpus before pre-training. Your team wants to remove exact duplicates and near-duplicates to reduce memorization and speed up training. Which NeMo Curator stage should you apply?

Medium
164

A healthcare AI team is evaluating a fine-tuned GPT-based model for clinical note summarization using NVIDIA NeMo. They need to assess whether the model's summaries contain no fabricated medical facts. Which evaluation approach is most appropriate?

Hard
165

When deploying a large language model on NVIDIA H100 GPUs using TensorRT-LLM, which TWO configuration strategies are most effective for improving KV cache efficiency and memory utilization?

Hard
166

A team is deploying a large language model using NVIDIA TensorRT-LLM on a multi-GPU node with NVLink. They want to minimize inter-GPU communication overhead during inference. Which parallelism strategy should they use to achieve this?

Hard
167

An engineer is optimizing a large language model for inference on NVIDIA GPUs and wants to reduce memory usage to fit a larger model or increase batch size. Which two techniques are most effective for reducing GPU memory consumption during inference? (Choose two.)

Hard
168

An engineer is using TensorRT-LLM to serve a model that occasionally receives prompts far longer than the typical 512 tokens, up to 8K tokens. With the default engine settings, requests near 8K fail with a cache capacity error while short requests succeed. Which configuration change most directly resolves this without rebuilding for a single worst-case shape?

Hard
169

An LLM inference service on NVIDIA Triton Inference Server uses dynamic batching with a max_queue_delay of 500 microseconds. During a load test, p99 latency exceeds the SLA while GPU utilization remains below 40%. Which change should you make first to improve latency without reducing throughput?

Medium
170

Which technique is most appropriate for a task requiring an LLM to generate code in a specific enterprise-internal syntax that is not well-represented in its public training data?

Medium
171

An engineer is building a customer-facing FAQ bot using an NVIDIA NIM-hosted Llama 3.1 70B model. The bot must answer ONLY from a supplied product knowledge base and must respond with 'I don't have that information' when the answer is not present. Which prompt engineering approach BEST enforces this constraint?

Medium
172

A developer is using NVIDIA Nsight Systems to profile a PyTorch training loop on an NVIDIA GPU. They notice significant gaps between kernel executions and want to identify whether the bottleneck is CPU-side or GPU-side. Which Nsight Systems feature should they use to visualize the CPU and GPU timelines together?

Easy
173

Refer to the exhibit. An engineer receives these logs while converting a Transformer model to a TensorRT engine. What is the most appropriate action to resolve this build failure?

Medium
174

What is the primary risk of 'catastrophic forgetting' during the fine-tuning process?

Medium
175

Which TWO of the following practices are considered standard procedures for preparing a dataset for Instruction Fine-Tuning (IFT)? (Choose two)

Hard
176

A media company is preparing an NVIDIA NIM-hosted LLM that summarizes user-submitted articles. Counsel requires evidence that the system respects copyright and attribution obligations. Which two practices should the team implement? (Choose two.)

Medium
177

When deploying an LLM using NVIDIA TensorRT-LLM, what is the primary benefit of pre-compiling the model into a TensorRT engine?

Medium
178

Refer to the exhibit. An engineer observes that memory utilization spikes during inference, causing OOM errors. Given the configuration, what is the most likely cause of the failure?

Hard
179

An engineer is optimizing prompts for an NVIDIA NIM-hosted model used in a multi-turn technical troubleshooting chat. The model forgets earlier constraints, such as the customer's environment and the product version, as the conversation grows. Which prompt engineering technique best preserves these constraints across turns?

Hard
180

A team is building a NeMo-based LLM pipeline and must tokenize a corpus that mixes English, Japanese, and Python source code. They plan to train a custom tokenizer with NVIDIA NeMo. Which tokenizer configuration best supports all three content types without excessive sequence length?

Hard
181

A developer is optimizing a generative AI model for inference on NVIDIA GPUs. They want to reduce memory footprint and improve throughput without sacrificing accuracy. Which two techniques should they apply? (Choose two.)

Medium
182

You are fine-tuning a LLM on a large corpus of technical documentation. You notice the model struggles with domain-specific terminology that frequently appeared in the training data but was inconsistently formatted. Which data preparation technique best addresses this issue?

Medium
183

Refer to the exhibit. The model is encountering an OOM error during long-context processing. Which architectural adjustment is most appropriate to resolve this while maintaining context length?

Medium
184

After fine-tuning a code-generation model with NVIDIA NeMo, an engineer notices the model now produces correct domain-specific function calls but has started emitting malformed JSON in about 15 percent of structured-output requests. The fine-tuning dataset contained no structured-output examples. Which evaluation action best explains and catches this regression?

Medium
185

A team is building a customer service chatbot using NVIDIA NeMo Guardrails and NVIDIA NIM. They need to ensure compliance with ethical AI guidelines, specifically around transparency and user consent. Which two actions should they implement to meet these ethical requirements? (Choose two.)

Medium
186

A developer is using an NVIDIA NIM-hosted model to classify support tickets into a fixed set of categories. The model occasionally invents new category names. The team wants to guarantee that only allowed categories are returned. Which approach is most appropriate?

Medium
187

Refer to the exhibit. An engineer observes that a model fine-tuned with this LoRA configuration is failing to converge on a highly complex legal document domain. What is the most likely cause of this issue?

Hard
188

Which of the following describes the 'Chain-of-Verification' (CoVe) prompting technique?

Hard
189

An engineer is reviewing the architecture of a decoder-only LLM that must support very long input contexts for document analysis. They are considering architectural choices that extend effective context length beyond what the model saw during pretraining. Which TWO techniques are designed specifically to extend usable context length without retraining the entire model from scratch? (Choose two.)

Hard
190

A developer is using NVIDIA TensorRT to optimize a BERT-based model for inference. They notice that the engine performs poorly on variable-length input sequences because it was built with a single optimization profile for a fixed sequence length. What should they do to improve performance across different sequence lengths?

Medium
191

Refer to the exhibit. In the context of a distributed multi-GPU fine-tuning job, what is the most likely cause of this error?

Hard
192

You are evaluating a generative AI model for a chatbot that must adhere to strict safety guidelines. The model occasionally generates toxic or biased responses. You need to automatically evaluate the model's outputs for toxicity and bias before deployment. Which NVIDIA tool or framework is specifically designed for this purpose?

Easy
193

Which prompt engineering strategy helps the model maintain focus when processing an extremely long document within a single context window?

Medium
194

A developer is preparing a dataset for instruction fine-tuning of an LLM using NVIDIA NeMo. The raw data consists of customer support transcripts with speaker labels and timestamps. Which preprocessing step is most important before training?

Easy
195

A team is evaluating a large language model (LLM) deployed with NVIDIA Triton Inference Server for a real-time question-answering system. They observe that the model's responses are accurate but latency spikes during peak load, causing timeouts. They need to evaluate the model's performance under high concurrency to identify the maximum throughput while maintaining a 95th percentile latency below 200 ms. Which evaluation approach should they use?

Hard
196

An organization is deploying a high-throughput LLM on NVIDIA Triton Inference Server. They observe significant tail latency spikes when serving multiple concurrent requests. Which strategy most effectively optimizes GPU utilization and reduces latency jitter for these concurrent model instances?

Medium
197

You are preparing a large instruction-tuning dataset for an LLM using NVIDIA NeMo. You need to ensure the dataset supports efficient training and evaluation. Which TWO steps are essential in the data preparation pipeline? (Choose two.)

Medium
198

Which of the following describes the purpose of 'gradient accumulation' in fine-tuning scenarios?

Medium
199

A team is deploying a 13B-parameter LLM with NVIDIA TensorRT-LLM on a single A100 80GB GPU. They want to reduce GPU memory usage during inference without retraining the model, while keeping acceptable output quality. Which technique should they apply?

Medium
200

An AI engineer is deploying a large language model using NVIDIA Triton Inference Server. They need to ensure that the server can handle multiple concurrent requests efficiently while maintaining low latency. Which Triton feature allows the server to dynamically batch incoming requests to maximize GPU utilization?

Easy
201

Which TWO factors should be prioritized when selecting a quantization strategy for deploying a large language model on constrained edge hardware? (Select TWO)

Medium
202

A hospital's AI governance board is reviewing an LLM triage assistant built on NVIDIA NIM. They want an ongoing, automated mechanism that flags when model outputs drift toward unsafe clinical recommendations across thousands of daily conversations, without reviewing every transcript manually. Which approach best fits this need?

Hard
203

Which of the following describes the purpose of a 'System Prompt' in Instruction Fine-Tuning?

Medium
204

You are preparing a dataset for instruction fine-tuning an LLM using NVIDIA NeMo. The dataset contains pairs of instructions and responses, but you notice that some responses are significantly longer than others, and a few are extremely short (e.g., 'Yes' or 'No'). You want to ensure the model learns to generate appropriate-length responses. Which data preparation technique is most effective?

Medium
205

A team fine-tunes a model with NVIDIA NeMo using a packed sequence dataset and notices that some training samples contain several short conversations concatenated. They must ensure the loss is computed only on assistant responses and not on the packed boundaries. Which configuration detail should they verify?

Hard
206

A platform team is deploying a 70B-parameter LLM with NVIDIA TensorRT-LLM across four 80 GB H100 GPUs and needs to serve long-context requests efficiently. They are deciding how to combine parallelism and memory techniques in the build and runtime configuration. (Choose two.)

Hard
207

An LLM inference service deployed on NVIDIA Triton Inference Server is experiencing occasional failures under high load. The team wants to implement proactive monitoring to predict and prevent these failures. Which two metrics should be prioritized for early detection of potential issues? (Choose two.)

Hard
208

An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. They want to maximize throughput for a chatbot workload with variable-length inputs and outputs. Which of the following techniques should they implement to achieve the highest throughput while maintaining acceptable latency?

Hard
209

You are preparing a dataset for pretraining an LLM using NVIDIA NeMo's Megatron-LM. The dataset consists of JSONL files where each line contains a 'text' field. You need to convert these files into the binary format required by Megatron for efficient training. Which tool or method should you use to perform this conversion?

Hard
210

An enterprise deployment of NeMo Guardrails is experiencing hallucinations where the model provides medical advice despite strict system prompts. What is the most effective approach to mitigate this risk?

Medium
211

In the context of NVIDIA Tensor Cores, what is the primary benefit of using BF16 (Bfloat16) over FP16 during model training and inference?

Easy
212

An enterprise is preparing a massive multi-terabyte corpus of specialized PDF documents containing technical schematics and dense tables for training a domain-specific LLM using NeMo Curator. During the document extraction pipeline, you notice that tabular data layouts are scrambled into unstructured string tokens, destroying spatial relationships. Which data preparation strategy should you implement first within NeMo Curator to preserve tabular integrity before tokenization?

Medium
213

A global bank uses NVIDIA NeMo Guardrails in front of an LLM assistant that answers employee HR questions. Legal requires that the assistant refuse any request that could constitute unauthorized legal advice, even when the request is phrased indirectly. During testing, a prompt such as 'My manager wants to know if we can terminate someone for discussing pay' bypasses the existing rail. What is the most effective configuration change to close this gap?

Hard
214

A team is preparing a customer-support chat dataset for instruction fine-tuning of an LLM. The raw data contains HTML tags, inconsistent date formats, and emoji. Which data preparation step should be performed first to make the text usable for tokenization?

Easy
215

A team has fine-tuned a Llama-3-70B model with NVIDIA NeMo and now must decide whether the tuned checkpoint actually improves on the base model for their domain. They run both models over the same 500-prompt held-out set and compute ROUGE-L against reference answers. The fine-tuned model scores 0.41 and the base model scores 0.44. What is the most technically sound conclusion the evaluation lead should draw?

Medium
216

Refer to the exhibit. What prompt engineering strategy ensures the model consistently maintains its persona and technical expertise throughout this multi-turn dialogue?

Medium
217

Which hardware component of an NVIDIA GPU is most responsible for accelerating matrix-multiply-accumulate (MMA) operations used in transformer layers?

Medium
218

An inference engineer is serving a 70B-parameter decoder-only LLM and wants to reduce KV cache memory to fit longer contexts on each GPU. They consider Multi-Query Attention (MQA), Grouped-Query Attention (GQA), and standard Multi-Head Attention (MHA). Which statement correctly describes the memory and quality tradeoff among these attention variants?

Hard
219

An engineer is using NVIDIA TensorRT-LLM to optimize an LLM for inference. They want to reduce the memory footprint of the KV cache during long-context generation. Which TWO techniques are supported by TensorRT-LLM to achieve this? (Choose two.)

Medium
220

Which component in the NVIDIA AI Enterprise stack is primarily responsible for serving multiple models, managing model versions, and providing metrics for monitoring model health?

Easy
221

Which architectural component is responsible for projecting the model's hidden states back into the vocabulary space to predict the next token?

Medium
222

A team is deploying an LLM for real-time inference using NVIDIA Triton Inference Server. They need to monitor the health and performance of the GPU to ensure reliability. Which NVIDIA tool provides comprehensive GPU telemetry, including utilization, memory, temperature, and power, and integrates with Prometheus for monitoring?

Easy
223

A team is using an NVIDIA NeMo-based LLM to answer questions over a product manual. The model sometimes answers using general knowledge instead of the provided manual excerpts. They want to force the model to rely only on the supplied context. Which prompt engineering approach best addresses this?

Medium
224

A team is fine-tuning an 8B-parameter LLM with LoRA on a single NVIDIA A100 80GB GPU using NVIDIA NeMo. They observe that training loss decreases, but validation loss starts to rise after epoch 2. They want to keep the same dataset and hyperparameters but mitigate overfitting. Which change is most appropriate?

Medium
225

A team is deploying a 70B-parameter decoder-only LLM on an NVIDIA H100 GPU node. During generation they observe that the KV cache grows linearly with sequence length and is consuming most of the available HBM, forcing them to limit batch size. They want to reduce KV cache memory without retraining the model from scratch. Which architectural change should they apply?

Medium
226

When using the NVIDIA Collective Communications Library (NCCL), what does 'AllReduce' specifically optimize for in a multi-GPU training configuration?

Medium
227

In the context of LLM deployment, why is it recommended to use a dedicated inference server like Triton rather than a basic Flask or FastAPI wrapper?

Medium
228

Which of the following is considered a best practice for logging in a production LLM environment?

Easy
229

An engineer is fine-tuning a 13B parameter model with NVIDIA NeMo using tensor parallelism across four GPUs. After resuming from a checkpoint, training loss spikes and then diverges. The checkpoint was saved with a different tensor parallel size than the current run. What is the most likely cause of the divergence?

Hard
230

An engineer is evaluating a sparse Mixture-of-Experts decoder-only LLM for a latency-sensitive inference service. They notice that although the model has far more total parameters than a dense baseline, throughput per token is only modestly better and sometimes worse. Which factor best explains why sparse MoE does not translate total parameter count into proportional speedup during inference?

Hard
231

When deploying a model, what is the benefit of using Triton's 'Model Versioning' feature?

Medium
232

A financial institution is evaluating a fine-tuned GPT-3 model for generating investment advice summaries. They must ensure the model does not produce harmful or biased recommendations. Which evaluation methodology should they implement using NVIDIA NeMo Guardrails and NeMo Evaluator to systematically detect and quantify such issues?

Hard
233

A team is fine-tuning a Llama 2 7B model with NVIDIA NeMo Framework on a single A100 80GB GPU. They observe that training loss decreases initially but then diverges, and the model outputs become repetitive and incoherent. The team used a learning rate of 5e-5 with AdamW and no warm-up. Which change is most likely to stabilize training and improve convergence?

Medium
234

A team is deploying a quantized LLM using NVIDIA NIM. To ensure the highest level of security and compliance, they need to verify that the container image has been scanned for vulnerabilities before production use. Which tool is the primary source for certified, production-ready NIM containers?

Medium
235

Refer to the exhibit. An engineer observes that GPU memory utilization is high, but the GPU is frequently idling. How does the provided Triton configuration optimize the inference pipeline?

Medium
236

What is the primary function of the 'rank' parameter in LoRA?

Easy
237

A team is deploying a large language model on NVIDIA Triton Inference Server with NVIDIA TensorRT-LLM backend. They need to reduce GPU memory usage to fit a larger model on the same hardware while maintaining acceptable latency. Which two techniques should they use? (Choose two.)

Hard
238

An enterprise is fine-tuning a large language model using NVIDIA NeMo Framework and encounters GPU out-of-memory errors during the backward pass. The training configuration already uses mixed-precision training (FP16). Which architectural intervention should be applied to resolve memory pressure while retaining the optimizer state precision?

Medium
239

A team is training a large language model on a single NVIDIA H100 GPU. They observe that training throughput is significantly lower than expected, and profiling with Nsight Systems shows long periods where the GPU is idle waiting for data. The data loading pipeline uses the default PyTorch DataLoader with num_workers=0 and no pinned memory. Which change is most likely to improve GPU utilization?

Medium
240

Refer to the exhibit. You are reviewing the configuration file for a data preprocessing pipeline. Why is the 'min_words' filter set to 50 in the context of LLM training?

Medium
241

You are responsible for the reliability of an LLM inference service running on NVIDIA Triton Inference Server across a fleet of A100 GPUs. The service is deployed with dynamic batching enabled, but during peak hours you observe that end-to-end latency for some requests exceeds the SLO while GPU utilization remains moderate. You suspect that the dynamic batching configuration is causing the issue. Which Triton configuration parameter should you adjust to directly limit the maximum time a request waits in the scheduler queue before being batched?

Medium
242

An engineer is tuning a TensorRT-LLM deployment of a 7B model for a latency-sensitive API. Profiling shows that time per output token is higher than expected and that many small kernels run back to back with gaps between them. Which TWO changes are most likely to reduce the per-token latency by cutting kernel launch overhead and redundant memory traffic? (Choose two.)

Medium
243

An enterprise is fine-tuning a 34B model with NVIDIA NeMo Framework and observes that the validation loss begins rising after the first epoch while training loss continues to fall. The team wants to reduce this divergence and preserve downstream task quality. (Choose two.)

Hard
244

You are preparing a massive dataset for training a NeMo-based LLM. Which TWO data preprocessing steps are critical to prevent data leakage and ensure model quality?

Hard
245

Which of the following describes the purpose of 'Kernel Fusion' in the context of optimizing a Deep Learning inference pipeline?

Easy
246

An engineer is deploying a 70B parameter model and needs to serve many concurrent users on a single GPU with limited memory. They want to store the attention keys and values for past tokens efficiently so that generation does not recompute them at every step. Which technique should they implement?

Medium
247

What is the primary function of the 'TensorRT' optimization engine in the NVIDIA AI software stack?

Medium
248

Why is it important to perform 'domain-specific' data cleaning when preparing a corpus for fine-tuning a medical LLM?

Medium
249

A team is deploying a large language model on NVIDIA Triton Inference Server in a production environment. They need to ensure that the model server can automatically recover from GPU failures without manual intervention. Which feature of Triton should they configure to achieve this?

Easy
250

In the context of NVIDIA NeMo, why is it recommended to use FP8 precision during the fine-tuning process on H100 GPUs?

Medium
251

A team is deploying a 13B-parameter chatbot on a single NVIDIA A10G GPU (24 GB VRAM). The model's weights are stored in FP16, and the runtime runs out of memory during KV cache allocation under concurrent user sessions. They must keep answer quality essentially unchanged while maximizing concurrent sessions. Which optimization should they apply first?

Medium
252

Refer to the exhibit. The engineer is attempting to deploy on an NVIDIA Orin platform but encounters a runtime error. What is the most likely cause of the failure?

Hard
253

In the context of the Transformer architecture, what is the primary function of the Feed-Forward Network (FFN) layers applied after the attention mechanism?

Easy
254

An engineer is designing a Transformer-based model for long-context document summarization. They decide to replace the standard dense self-attention mechanism with a sliding window attention approach. What is the primary architectural implication of this change?

Medium
255

An engineer is deploying a 13B-parameter LLM with TensorRT-LLM on a single NVIDIA A100 40GB GPU. The FP16 engine requires 26GB for weights, but during generation the KV cache grows beyond remaining memory, causing out-of-memory errors. The team wants to maximize concurrent requests without retraining. Which optimization should they apply first?

Medium
256

A team has deployed a Llama-3-8B model with NVIDIA Triton Inference Server for real-time chatbot responses. They need to evaluate the model's generation quality on a held-out test set of 500 prompts, focusing on semantic similarity to reference answers while ensuring low latency. Which evaluation approach is most appropriate to use with NVIDIA NeMo Evaluator?

Medium
257

A generative AI application built on NVIDIA Triton Inference Server is deployed in a Kubernetes cluster with GPU nodes. The operations team wants to detect silent data corruption in model outputs, which could occur due to GPU memory errors. They plan to implement a monitoring solution using NVIDIA Data Center GPU Manager (DCGM). Which DCGM feature should they enable to detect and alert on GPU memory errors that could lead to silent data corruption?

Hard
258

A research team is fine-tuning a model with NVIDIA NeMo and wants to reduce the risk of catastrophic forgetting of general capabilities while still adapting to a specialized domain. They have a small domain dataset and limited compute. Which fine-tuning approach best balances domain adaptation with retention of pretrained knowledge?

Hard
259

Refer to the exhibit. Which adjustment is the most immediate and effective way to resolve this OOM error while maintaining the same training architecture?

Hard
260

Refer to the exhibit. What is the most likely cause of the failure based on the log entries?

Hard
261

An engineer is using an NVIDIA NIM for a Mixtral model to extract structured data from invoices. The model occasionally returns fields with the wrong data type, such as a numeric amount as a string. The team wants a prompt engineering fix that does not require changing the model or adding a separate parser. Which approach is most effective?

Hard
262

An engineer is using TensorRT-LLM to serve a chatbot model. They observe that the time to first token (TTFT) is high, but subsequent tokens are generated quickly. Which optimization should they prioritize to reduce TTFT?

Medium
263

When fine-tuning a Large Language Model using Low-Rank Adaptation (LoRA), which architectural component is primarily modified to reduce computational overhead while maintaining performance?

Medium
264

When optimizing a Generative AI model using NVIDIA TensorRT-LLM, which component is primarily responsible for managing the KV cache to minimize memory fragmentation?

Hard
265

Refer to the exhibit. What is the impact of this filter on the training corpus?

Hard
266

A company needs to deploy a generative AI model that will serve prompts containing regulated customer data. Security policy requires that all inference stays on-premises, that the model be quantized to fit existing GPUs, and that no external network calls occur at runtime. Which deployment approach should the engineer choose?

Medium
267

A team is training a 13B-parameter LLM on 8 NVIDIA A100 GPUs using NVIDIA NeMo. They observe that the all-reduce communication during data-parallel training consumes nearly 40% of each iteration. Which of the following changes is most likely to reduce this communication overhead while preserving convergence?

Medium
268

What is the primary function of the 'Triton Model Control' API in a production environment?

Medium
269

A team is fine-tuning a Llama 3 8B model with NVIDIA NeMo on a single A100 80GB GPU. They observe that validation loss starts to rise while training loss continues to decrease after epoch 2. They want to keep the best generalizing checkpoint without changing the dataset. Which NeMo training configuration strategy should they apply?

Medium
270

An LLM inference service on NVIDIA Triton Inference Server is experiencing a gradual increase in P99 latency over several hours without a corresponding increase in request rate. GPU utilization remains stable at around 60%. Which action should the team take first to diagnose the root cause?

Hard
271

An engineer is optimizing an LLM for inference with NVIDIA TensorRT-LLM and wants to reduce both memory footprint and latency without retraining the model. Which two techniques should they apply? (Choose two.)

Medium
272

An LLM inference service on NVIDIA Triton Inference Server uses TensorRT-LLM as the backend. During peak load, you notice that the first token latency (time to first token) is high, but subsequent tokens are generated quickly. Which metric should you monitor to diagnose this issue?

Hard
273

A team is fine-tuning a 70B parameter model with NVIDIA NeMo using LoRA on eight H100 GPUs. They want to reduce GPU memory usage during training while preserving the base model's pretrained knowledge. Which two configuration changes should they apply? (Choose two.)

Medium
274

When optimizing a model using NVIDIA TensorRT, what is the primary benefit of enabling 'layer fusion' during the optimization process?

Medium
275

Refer to the exhibit. An engineer receives this timeout error during a CUDA kernel execution. What is the most appropriate first step to diagnose the resource contention?

Medium
276

A team is deploying a 13B-parameter decoder-only LLM on a single NVIDIA A100 40GB GPU for a real-time chatbot. During load, the process runs out of memory even though the model weights in FP16 require roughly 26GB. The team wants to reduce GPU memory usage with minimal impact on output quality and no change to the model architecture. Which technique is most appropriate?

Medium
277

You are building a pretraining dataset from a large collection of source-code repositories for an NVIDIA NeMo LLM. The data includes many files with licenses, generated code, and minified JavaScript. Which NeMo Curator-based approach best improves code data quality before tokenization?

Hard
278

An LLM inference service on NVIDIA Triton Inference Server is experiencing occasional out-of-memory (OOM) errors on the GPU. The team wants to monitor GPU memory usage to predict and prevent OOM conditions. Which NVIDIA tool provides real-time GPU memory metrics that can be integrated with Prometheus for alerting?

Medium
279

You are curating instruction-tuning data for an NVIDIA NIM-deployed LLM. The raw dataset contains many near-duplicate instruction-response pairs that differ only by punctuation and whitespace. Which data preparation step is most appropriate to remove these before fine-tuning?

Easy
280

What is the primary role of a Model Repository in the NVIDIA Triton Inference Server architecture?

Medium
281

Which TWO of the following techniques are best suited for reducing the latency of LLM inference on NVIDIA GPUs?

Medium
282

You are preparing a large corpus of customer support transcripts for continued pretraining of an NVIDIA NeMo Megatron model. The transcripts contain personally identifiable information such as names, email addresses, and account numbers, and company policy requires that this information be removed before training while preserving as much linguistic context as possible for the model to learn from. Which data preparation approach best satisfies both requirements?

Medium
283

A team is using NVIDIA NeMo Evaluator to assess a Llama-3-70B model fine-tuned for medical question answering. They want to evaluate both the correctness of answers and the model's ability to avoid hallucinating unsupported facts. Which two evaluation strategies should they implement? (Choose two.)

Medium
284

Which TWO of the following NVIDIA AI Enterprise tools are specifically designed to optimize and accelerate the deployment of LLMs in containerized environments?

Medium
285

An engineer is optimizing a BERT-like model for inference using NVIDIA TensorRT. They want to reduce latency further by using lower precision without significant accuracy loss. Which TensorRT precision mode should they choose to enable INT8 inference while maintaining accuracy through calibration?

Easy
286

What is the primary purpose of 'Few-Shot Prompting' in the context of LLM optimization?

Easy
287

A research team is pretraining a decoder-only LLM and observes that gradient magnitudes in the earliest layers are extremely small while later layers train normally, causing slow convergence. They are using post-layer normalization. Which architectural change is most likely to improve gradient flow to the early layers?

Medium
288

Which TWO actions should be part of a robust incident response plan for an LLM deployment failing in production?

Medium
289

Refer to the exhibit. What is the intended outcome of this data cleaning configuration for a Large Language Model pre-training corpus?

Hard
290

A research team wants to train a large decoder-only LLM where each token's representation is computed independently of token order, then inject order information afterward. They are choosing between learned absolute positional embeddings and sinusoidal absolute positional embeddings. Which statement accurately characterizes the tradeoff they face?

Medium
291

A developer is prompting an NVIDIA NIM for a Code Llama model to generate a Python function. The model produces correct logic but frequently omits type hints and docstrings, which the team requires. Which prompting technique best addresses this specific gap?

Medium
292

A financial institution uses NVIDIA NeMo Guardrails to enforce ethical guidelines in its customer-facing LLM. During testing, the model occasionally generates responses that violate the company's policy against offering investment advice. The guardrails are configured with a set of dialog flows and safety checks. What is the most effective way to address this issue?

Hard
293

Which component in the NVIDIA NeMo framework is specifically designed to manage the configuration and orchestration of large-scale fine-tuning jobs?

Easy
294

Refer to the exhibit. Which prompt engineering technique would best force the model to prioritize technical detail over marketing language?

Hard
295

In the context of model optimization, why is 'graph surgery' sometimes required before building a TensorRT engine?

Medium
296

Which THREE architectural features are essential for enabling efficient inference of massive LLMs on multi-GPU systems?

Hard
297

Why is 'tokenization stability' a critical metric when preparing data for NVIDIA-based LLM deployment?

Medium
298

You are preparing a dataset of customer reviews for fine-tuning an LLM to generate concise summaries. The reviews are in multiple languages, but the target summaries must be in English. You have a limited budget for translation. Which data preparation step is most critical to ensure the fine-tuned model produces high-quality English summaries?

Medium
299

When using QLoRA for fine-tuning, what is the primary purpose of using the 4-bit NormalFloat (NF4) data type?

Medium
300

When designing an AI application for the public sector, which ethical principle must be prioritized regarding transparency?

Medium
301

A data science team is preparing an instruction fine-tuning dataset in NVIDIA NeMo Framework. They notice that after training, the model performs well on the training instructions but poorly on paraphrased versions of the same instructions. They want to improve generalization without increasing dataset size. Which data preparation change is most appropriate?

Medium
302

Refer to the exhibit. The deployment is facing memory allocation errors during peak load. Based on the error log, what is the most effective configuration change to resolve the issue while keeping the model architecture constant?

Hard
303

You are building an evaluation harness for a retrieval-augmented generative assistant running on NVIDIA NIM microservices. The product owner wants a single trustworthy number for 'answer quality,' but you need to defend the evaluation design. Which two design choices most directly protect the evaluation from producing misleading quality scores? (Choose two.)

Hard
304

An ML engineer is fine-tuning a 13B LLM with LoRA on 4 NVIDIA A100 GPUs using NVIDIA NeMo. They notice that the effective batch size is very small and gradients are noisy, but increasing the per-GPU micro batch size triggers out-of-memory errors. Which technique should they apply to increase the effective batch size without increasing memory per step?

Hard
305

A media analytics company runs a TensorRT-LLM optimized GPT-J model on a single NVIDIA A100 80GB GPU using NVIDIA Triton Inference Server. During peak hours, request concurrency rises sharply and the team observes that the GPU is idle for long periods while waiting on host-side tokenization and detokenization. Profiling shows that CPU preprocessing and postprocessing dominate request latency. The team wants to reduce end-to-end latency without changing model weights or adding GPUs. Which Triton feature should they use?

Hard
306

When fine-tuning an LLM to follow specific safety protocols, why is the inclusion of 'adversarial' examples in the training data considered a best practice?

Hard
307

What is the function of the 'Masked' component in a Decoder-only Transformer's self-attention during training?

Easy
308

A global bank runs an NVIDIA NeMo Guardrails-protected assistant for its tellers. During a compliance audit, the auditor asks how the bank can prove that the guardrail configuration itself has not been silently altered between releases. Which practice best satisfies this requirement?

Medium
309

A team is optimizing an NVIDIA TensorRT-LLM deployment of a 70B model on multiple GPUs. They want to reduce inter-GPU communication overhead and improve throughput. Which two techniques should they consider? (Choose two.)

Medium
310

A healthcare AI team is evaluating a large language model for clinical note summarization. They want to measure whether the generated summaries contain fabricated information not present in the source notes. Which evaluation approach is most appropriate for detecting hallucinated content?

Hard
311

A developer is building a decoder-only generative model and wants to prevent the model from attending to future tokens during training so that each position can only use information from itself and earlier positions. Which architectural mechanism should they implement in the self-attention layer?

Easy
312

An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. The model uses multi-head attention with a sequence length of 4096. During inference, the GPU's Tensor Cores are underutilized, and the kernel launch overhead is high due to many small operations. Which optimization should be applied to improve Tensor Core utilization and reduce overhead?

Hard
313

You are preparing a dataset for supervised fine-tuning (SFT) of an NVIDIA NeMo LLM to follow instructions. Which TWO data preparation practices are essential to ensure the model learns to generalize rather than memorize? (Choose two.)

Medium
314

An enterprise team is preparing a supervised fine-tuning job in NVIDIA NeMo for a 20B LLM. They want to reduce GPU memory consumption during training without changing the model architecture or the dataset. Which two configuration changes should they apply? (Choose two.)

Hard
315

An engineer is training a large language model with pipeline parallelism across four NVIDIA GPUs. They observe that GPU utilization is low and training throughput is limited by idle time during pipeline bubbles. Which technique is most effective to reduce pipeline bubbles and improve utilization?

Hard
316

A developer needs to fine-tune a 7B LLM for a customer-support chatbot using NVIDIA NeMo. The dataset contains paired instructions and desired responses. Which data format should be used to prepare the dataset for supervised fine-tuning?

Easy
317

A team has built a TensorRT-LLM engine for a 70B model on four NVIDIA H100 GPUs using tensor parallelism. They now need to serve the same model on a single H100 for a development environment, accepting higher latency. What is the most appropriate approach?

Hard
318

A developer wants to reduce the disk and memory footprint of a fine-tuned 70B model before serving it with TensorRT-LLM, and is willing to accept a small, measurable quality drop that they will validate with an evaluation harness. Which approach best matches that requirement?

Easy
319

You are preparing a large instruction-tuning dataset with NVIDIA NeMo Curator. The dataset contains many near-duplicate instruction-response pairs that differ only in punctuation or minor wording. Which NeMo Curator stage should you apply to remove these near-duplicates before fine-tuning?

Medium
320

Why is gradient checkpointing useful when fine-tuning a model on a single GPU?

Medium
321

An enterprise deploying a large language model on an NVIDIA A100 GPU experiences high memory bandwidth bottlenecks during autoregressive token generation. Which optimization technique specifically addresses this memory-bound phase by merging element-wise operations and reducing global memory round-trips?

Medium
322

A generative AI application deployed on NVIDIA Triton Inference Server with NVIDIA AI Enterprise is experiencing silent data corruption in model outputs. The MLOps team needs to implement monitoring to detect such issues early. Which two actions should be taken to enhance observability for silent data corruption? (Choose two.)

Hard
323

Which strategy is most effective for optimizing an LLM that is too large to fit into a single GPU's VRAM?

Medium
324

You are preparing a multilingual corpus for pretraining with NVIDIA NeMo. The dataset contains documents in 40 languages, but the tokenizer was trained primarily on English. Which data preparation action best ensures that non-English text is represented efficiently during tokenization?

Medium
325

Refer to the exhibit. What is the effective batch size for this fine-tuning job?

Hard
326

A company wants to fine-tune a 70B LLM to follow domain-specific instructions. The base model already performs well on general language tasks. They have limited labeled data and limited GPU memory. Which approach is most appropriate?

Medium
327

A retail company wants to let its customer-support LLM answer questions about order status. The security team insists the model must never be able to invoke a refund or account-modification function, even if a user crafts a clever prompt. Which approach enforces that constraint most reliably?

Easy
328

An engineer is using NVIDIA TensorRT-LLM's in-flight batching to serve a mix of short and very long prompts. They observe that GPU utilization drops and latency for short requests spikes whenever a long prompt is admitted. Which mechanism should they tune to prevent long sequences from monopolizing the batch?

Hard
329

A team is using an NVIDIA NIM-hosted Llama model to generate product descriptions from a list of technical specifications. The descriptions sometimes omit key specifications or include invented features. The team wants to improve reliability without changing the model. Which prompt engineering change is most likely to reduce these errors?

Medium
330

You are preparing a dataset of support tickets for a RAG system using NVIDIA NeMo. Many tickets are short and contain little context, which hurts retrieval quality. Which data preparation technique best improves retrieval by enriching each ticket with related information before embedding?

Medium
331

A global bank uses NVIDIA NeMo Guardrails to enforce ethical AI policies in its customer-facing LLM application. The compliance team requires that the system automatically logs all instances where the model attempts to generate financial advice, including the prompt, the blocked response, and the rail that triggered. Which NeMo Guardrails feature should the team enable to capture this audit trail?

Hard
332

A production LLM service on NVIDIA Triton Inference Server uses dynamic batching. During peak load, the 99th percentile latency increases significantly, but GPU utilization remains at 60%. Which configuration change is most likely to improve latency while maintaining throughput?

Hard
333

An enterprise deploying a large language model on NVIDIA Triton Inference Server needs to track GPU utilization, memory allocation, and custom inference latency histograms in Prometheus. Which approach should the MLOps engineer implement to ensure robust production observability without overloading the inference execution threads?

Medium
334

An engineer is building a customer support assistant using an NVIDIA NIM for a Llama 3 70B model. The assistant must always respond in valid JSON containing exactly the keys "issue" and "urgency", and must never include any other text. Which prompt engineering approach most directly enforces this output contract?

Easy
335

In the context of NVIDIA AI Enterprise, what is the primary purpose of using 'NVIDIA Triton Model Analyzer'?

Easy
336

A team is serving a 70B-parameter LLM with TensorRT-LLM in a multi-tenant environment where requests arrive with widely varying prompt lengths and generation lengths. During load testing, they observe that throughput collapses when a long-context request is scheduled alongside many short requests, and GPU memory fragmentation causes intermittent out-of-memory errors even though total free memory appears sufficient. Which TensorRT-LLM runtime configuration change most directly addresses both the throughput collapse and the memory fragmentation?

Hard
337

A production LLM inference service on NVIDIA Triton Inference Server runs on a multi-GPU node. You observe that one GPU reports ECC XID errors and the model's throughput gradually degrades. Which NVIDIA tool should you use to monitor GPU health and set up alerts for these errors?

Hard
338

Which THREE factors significantly influence the memory consumption during LLM fine-tuning? (Choose three)

Hard
339

When evaluating LLM reliability under stress, what is the primary goal of conducting 'Chaos Engineering' on a Triton inference cluster?

Hard
340

An LLM service on NVIDIA Triton Inference Server experiences high time-to-first-token because the dynamic batcher waits for full batches. The team wants to reduce time-to-first-token while still benefiting from batching. Which adjustment is most appropriate?

Hard
341

A bank's model risk committee is reviewing an LLM-based loan-adverse-action notice generator built on NVIDIA NeMo. Regulators require that the system produce a human-readable rationale for each denial and that the rationale be reproducible for any prior decision. Which architectural choice most directly meets both obligations?

Hard
342

When evaluating an LLM's response to a complex prompt, what is the 'Persona Adoption' technique?

Medium
343

A hospital network runs an on-premises NVIDIA NIM microservice hosting a clinical-summarization LLM. Compliance requires that every generated summary be attributable to source records and that no protected health information leave the subnet. Which deployment practice best satisfies both requirements at once?

Medium
344

Refer to the exhibit. An audit reveals that 'INTERNAL_STRATEGY' documents are still being generated by the model. Why is this occurring?

Hard
345

A company is using NVIDIA NeMo Guardrails to enforce safety policies in its LLM application. A developer wants to ensure that the model does not generate content that violates the company's policy against discussing competitor products. Which type of guardrail should the developer configure to prevent the model from mentioning competitor names in its responses?

Easy
346

A production LLM inference service runs on NVIDIA Triton Inference Server across multiple GPUs. The SRE team wants to detect when the service starts returning incorrect or degraded responses compared to a baseline, even when latency and throughput remain normal. Which monitoring approach is most appropriate?

Medium
347

Which optimization method should be prioritized when the model's inference performance is bottlenecked by the CPU-to-GPU data transfer overhead?

Medium
348

A team is preparing a mixed-language corpus for continued pretraining of an NVIDIA NeMo Megatron model. The corpus contains English, Japanese, and Arabic documents. Tokenizer analysis shows the current English-centric BPE vocabulary produces very long token sequences for Japanese and Arabic, inflating sequence length and compute cost. The team wants to reduce sequence length for non-English text without retraining the tokenizer from scratch and without degrading English performance. Which data preparation action best achieves this?

Hard
349

An engineer is using NVIDIA TensorRT to optimize a Transformer model for inference on an NVIDIA A100 GPU. They want to maximize throughput while ensuring that the model runs correctly with varying input sequence lengths. Which TensorRT feature should they configure to allow the engine to handle different input shapes at runtime?

Medium
350

Which THREE factors should be considered when choosing an optimal batch size for LLM inference on NVIDIA GPUs?

Medium
351

A team is fine-tuning a 13B-parameter Llama model with NVIDIA NeMo Framework on a single A100 80GB GPU. They apply LoRA adapters to the attention projection layers, but the adapters are producing negligible changes to model behavior even after several epochs, and the loss curve stays flat. They confirm the dataset is clean and the tokenizer is correct. Which LoRA configuration issue is the most likely cause?

Medium
352

A financial services company has fine-tuned a Llama-3-70B model using NVIDIA NeMo for automated loan risk assessment. During evaluation, they observe that the model achieves high scores on standard accuracy metrics but produces inconsistent responses when the same query is phrased slightly differently. They need a metric that quantifies this inconsistency. Which evaluation metric should they use?

Medium

Frequently asked questions

What does the scenario questions domain cover on the NCP-GENL exam?
scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
How many questions are in this domain?
This page lists all 352 scenario questions questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only scenario questions questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.