Courseiva

NCP-GENL · topic practice

Scenario practice questions

Practise NVIDIA Certified Professional: Generative AI LLMs Scenario practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
15 questionsDomain: Scenario

What the exam tests

What to know about Scenario

Scenario questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Watch out for

Common Scenario exam traps

  • ▸Answering from memory before reading the full scenario.
  • ▸Missing a constraint such as cost, availability, security, scope or command context.
  • ▸Choosing a broad answer when the question asks for the most specific fix.
  • ▸Ignoring why the wrong options are tempting.

Practice set

Scenario questions

15 questions · select your answer, then reveal the explanation

Question 1mediummultiple choice
Read the full Scenario explanation →

Refer to the exhibit. Given this NeMo configuration, which prompt modification would best improve the reliability of technical support queries?

Exhibit

{"model_config": {"temperature": 0.2, "max_tokens": 512, "stop_sequences": ["Human:", "User:"], "system_prompt": "You are a helpful assistant. You must answer based on the provided technical manuals."}}
Question 2easymultiple choice
Read the full Scenario explanation →

Which technique is most effective for mitigating data leakage during the training of an LLM on time-series-related document data?

Question 3mediummultiple choice
Read the full Scenario explanation →

An enterprise fine-tunes a Llama-3-70B model using NVIDIA NeMo for automated technical support ticketing. The development team needs an automated evaluation pipeline that measures semantic similarity against human-curated reference answers without relying on costly human annotators. Which metric provides the most robust embedding-based semantic similarity assessment for this scenario?

Question 4hardmultiple choice
Read the full Scenario explanation →

Refer to the exhibit. The deployment is facing memory allocation errors during peak load. Based on the error log, what is the most effective configuration change to resolve the issue while keeping the model architecture constant?

Exhibit

Error: [TRT-LLM] KV Cache block allocation failed. Current allocation: 80% capacity. Request rejected to prevent OOM. Optimization status: KV_CACHE_ENABLED=True, PAGED_ATTENTION=False.
Question 5easymultiple choice
Read the full Scenario explanation →

A technical support team is building a chatbot using an NVIDIA NIM microservice. The chatbot must answer questions about a specific product's warranty policy. The team wants to ensure the model's responses are grounded in the official warranty document, which is 50 pages long, and avoid inventing policy details. Which prompt engineering approach is most effective for this scenario?

Question 6mediummulti select
Read the full Scenario explanation →

An engineer is using NVIDIA TensorRT-LLM to optimize an LLM for inference. They want to reduce the memory footprint of the KV cache during long-context generation. Which TWO techniques are supported by TensorRT-LLM to achieve this? (Choose two.)

Question 7mediummultiple choice
Read the full Scenario explanation →

A team is training a large language model on a single NVIDIA H100 GPU. They observe that training throughput is significantly lower than expected, and profiling with Nsight Systems shows long periods where the GPU is idle waiting for data. The data loading pipeline uses the default PyTorch DataLoader with num_workers=0 and no pinned memory. Which change is most likely to improve GPU utilization?

Question 8mediummultiple choice
Read the full Scenario explanation →

You are responsible for the reliability of an LLM inference service running on NVIDIA Triton Inference Server across a fleet of A100 GPUs. The service is deployed with dynamic batching enabled, but during peak hours you observe that end-to-end latency for some requests exceeds the SLO while GPU utilization remains moderate. You suspect that the dynamic batching configuration is causing the issue. Which Triton configuration parameter should you adjust to directly limit the maximum time a request waits in the scheduler queue before being batched?

Question 9mediummultiple choice
Read the full Scenario explanation →

You are preparing a dataset of support tickets for a RAG system using NVIDIA NeMo. Many tickets are short and contain little context, which hurts retrieval quality. Which data preparation technique best improves retrieval by enriching each ticket with related information before embedding?

Question 10mediummultiple choice
Study the full Python automation breakdown →

You are evaluating a fine-tuned LLM for a code generation task. The model was trained using NVIDIA NeMo on a dataset of Python functions. You want to measure the percentage of generated functions that pass a set of unit tests. Which evaluation metric is most appropriate?

Question 11hardmultiple choice
Read the full Scenario explanation →

Refer to the exhibit. What is the most likely reason for the high P99 latency despite low GPU utilization?

Exhibit

{
  "model_name": "llama-3",
  "latency_percentile_99": 1200,
  "target_p99": 500,
  "queue_depth": 50,
  "gpu_util": "40%"
}
Question 12easymultiple choice
Read the full Scenario explanation →

You are evaluating a text generation model using NVIDIA NeMo Evaluation and want to measure how well the generated text matches a reference translation. Which metric is specifically designed for this purpose?

Question 13hardmultiple choice
Read the full Scenario explanation →

A team is evaluating a retrieval-augmented generation (RAG) pipeline using NVIDIA NeMo Evaluation. They notice that the generated answers are fluent but sometimes contradict the retrieved documents. Which evaluation approach best identifies this issue?

Question 14mediummultiple choice
Read the full Scenario explanation →

A team is pretraining a decoder-only Transformer LLM on a large corpus of code and natural language. They observe that the model's training loss decreases smoothly, but during generation it sometimes produces degenerate repetition, and attention entropy on long sequences collapses. They suspect the issue is related to the positional encoding scheme. Which architectural change is most likely to mitigate the attention entropy collapse while preserving the model's ability to generalize to sequences longer than those seen during pretraining?

Question 15mediummultiple choice
Read the full Scenario explanation →

Which of the following describes the purpose of 'gradient accumulation' in fine-tuning scenarios?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Scenario sessions

Start a Scenario only practice session

Every question in these sessions is drawn from the Scenario domain — nothing else.

Related practice questions

Related NCP-GENL topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the NCP-GENL exam test about Scenario?
Scenario questions test whether you can apply the concept in context, not just recognise a definition.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Scenario questions in a focused session?
Yes — the session launcher on this page draws every question from the Scenario domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other NCP-GENL topics?
Use the topic links above to move to related areas, or go back to the NCP-GENL question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the NCP-GENL exam covers. They are not copied from any real exam or dump site.