Courseiva
← Back to NVIDIA Certified Professional: Generative AI LLMs questions

Scenario-based practice

Refer to the Exhibit Practice Questions

Practise NVIDIA Certified Professional: Generative AI LLMs practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

15
scenario questions
NCP-GENL
exam code
NVIDIA
vendor

Scenario guide

How to approach refer to the exhibit practice questions

Practise exhibit-style questions that ask you to read a topology, table, command output or diagram before choosing the best answer.

Quick answer

Exhibit-style questions test whether you can read a topology, command output, diagram or table before choosing the best answer.

How to extract the relevant detail from an exhibit.

How topology, command output or routing information affects the answer.

How to avoid answering from memory before reading the evidence.

How to map the exhibit back to the exam objective.

Related practice questions

Related NCP-GENL topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1mediummultiple choice
Full question →

Refer to the exhibit. Given this NeMo configuration, which prompt modification would best improve the reliability of technical support queries?

Exhibit

{"model_config": {"temperature": 0.2, "max_tokens": 512, "stop_sequences": ["Human:", "User:"], "system_prompt": "You are a helpful assistant. You must answer based on the provided technical manuals."}}
Question 2hardmultiple choice
Full question →

Refer to the exhibit. The Triton Inference Server is failing to load a model. The configuration shows two instances assigned to GPU 0. What is the most likely cause and the correct remediation?

Exhibit

config.pbtxt: 
instance_group [
  {
    count: 2
    kind: KIND_GPU
    gpus: [0]
  }
]

error_log: [ERROR] 'Failed to load model: CUDA out of memory' during multi-model parallel execution.
Question 3hardmultiple choice
Full question →

Refer to the exhibit. An engineer is fine-tuning an LLM using the provided configuration. What is the primary purpose of applying 'yarn' scaling in this architecture?

Exhibit

config: { 'model_type': 'decoder-only', 'rope_base': 10000, 'rope_scaling': { 'type': 'yarn', 'factor': 4.0 } }
Question 4hardmultiple choice
Full question →

Refer to the exhibit. Which prompt engineering technique would best force the model to prioritize technical detail over marketing language?

Exhibit

System: You are an expert at summarizing NVIDIA technical whitepapers.
User: Summarize the latest Grace Hopper Superchip architecture.
Assistant: [Model generates a 500-word response]
Issue: The response includes generic marketing fluff and lacks the requested technical depth.
Question 5mediummultiple choice
Full question →

Refer to the exhibit. An engineer receives these logs while converting a Transformer model to a TensorRT engine. What is the most appropriate action to resolve this build failure?

Exhibit

TRT_LOG: [E] Error: Failed to find valid tactics for layer 'Attention_Softmax_0'.
TRT_LOG: [E] Error: Workspace memory limit exceeded for layer 'MatMul_QKV_1'.
TRT_LOG: [I] INFO: Attempting to optimize graph with reduced precision...
Question 6hardmultiple choice
Full question →

Refer to the exhibit. An engineer observes that memory utilization spikes during inference, causing OOM errors. Given the configuration, what is the most likely cause of the failure?

Exhibit

model_config.pbtxt:
instance_group [
  {
    count: 2
    kind: KIND_GPU
    gpus: [0]
  }
]
Question 7hardmultiple choice
Full question →

Refer to the exhibit. The engineer is attempting to deploy on an NVIDIA Orin platform but encounters a runtime error. What is the most likely cause of the failure?

Network Topology
trt-builderonnx model.onnxworkspace 4096fp16dla 0
Question 8hardmultiple choice
Full question →

Refer to the exhibit. Which adjustment is the most immediate and effective way to resolve this OOM error while maintaining the same training architecture?

Exhibit

Error: CUDA out of memory. Tried to allocate 512.00 MiB. GPU 0 has 23.90 GiB total capacity. Memory used at last allocation: 23.50 GiB.
Question 9hardmultiple choice
Full question →

Refer to the exhibit. What is the most likely cause of the failure based on the log entries?

Exhibit

log_output:
[W] [TRT-LLM] [Performance] Request 1042 processing time exceeded 500ms.
[W] [TRT-LLM] [Performance] KV Cache eviction detected for sequence 1042.
[E] [TRT-LLM] [Memory] Out of Memory: Failed to allocate 128MB.
Question 10mediummultiple choice
Full question →

Refer to the exhibit. An engineer receives this timeout error during a CUDA kernel execution. What is the most appropriate first step to diagnose the resource contention?

Exhibit

Error: Kernel execution timed out after 5000ms.
Potential cause: Excessive occupancy or resource contention.
Action: Adjust block size or shared memory usage.
Question 11hardmultiple choice
Full question →

Refer to the exhibit. The deployment is facing memory allocation errors during peak load. Based on the error log, what is the most effective configuration change to resolve the issue while keeping the model architecture constant?

Exhibit

Error: [TRT-LLM] KV Cache block allocation failed. Current allocation: 80% capacity. Request rejected to prevent OOM. Optimization status: KV_CACHE_ENABLED=True, PAGED_ATTENTION=False.
Question 12mediummultiple choice
Full question →

Refer to the exhibit. An engineer observes that GPU memory utilization is high, but the GPU is frequently idling. How does the provided Triton configuration optimize the inference pipeline?

Exhibit

config.pbtxt:
instance_group [
  {
    count: 2
    kind: KIND_GPU
    gpus: [0]
  }
]
Question 13hardmultiple choice
Full question →

Refer to the exhibit. What is the effective batch size for this fine-tuning job?

Exhibit

Config: { "gradient_accumulation_steps": 16, "per_device_train_batch_size": 1 }
Question 14hardmultiple choice
Full question →

Refer to the exhibit. What is the implication of setting the memory_limit to 0.8 in the context of an LLM inference service?

Exhibit

config_policy:
- name: "kv_cache_management"
  type: "dynamic"
  memory_limit: 0.8
  eviction_policy: "lru"
Question 15hardmultiple choice
Full question →

Refer to the exhibit. An engineer observes that a model fine-tuned with this LoRA configuration is failing to converge on a highly complex legal document domain. What is the most likely cause of this issue?

Exhibit

config: { "rank": 8, "alpha": 16, "target_modules": ["query_key_value", "dense"], "dropout": 0.05, "bias": "none" }

These NCP-GENL practice questions are part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style NCP-GENL questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.