Courseiva
← Back to NVIDIA Certified Professional: Generative AI LLMs questions

Scenario-based practice

Hard Difficulty Questions

Practise NVIDIA Certified Professional: Generative AI LLMs practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

20
scenario questions
NCP-GENL
exam code
NVIDIA
vendor

Scenario guide

How to approach hard difficulty questions

These are the questions most candidates get wrong. They require connecting multiple concepts, reading tricky output, or knowing edge-case behaviour that isn't on most study cards. Practising them trains you to operate under uncertainty — a necessary skill on the real exam.

Quick answer

Hard Difficulty Questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Related practice questions

Related NCP-GENL topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmulti select
Full question →

When profiling an application with NVIDIA Nsight Systems, which TWO metrics are most critical to identify if an application is limited by the PCIe bus?

Question 2hardmultiple choice
Full question →

Refer to the exhibit. The Triton Inference Server is failing to load a model. The configuration shows two instances assigned to GPU 0. What is the most likely cause and the correct remediation?

Exhibit

config.pbtxt: 
instance_group [
  {
    count: 2
    kind: KIND_GPU
    gpus: [0]
  }
]

error_log: [ERROR] 'Failed to load model: CUDA out of memory' during multi-model parallel execution.
Question 3hardmultiple choice
Full question →

Refer to the exhibit. An engineer is fine-tuning an LLM using the provided configuration. What is the primary purpose of applying 'yarn' scaling in this architecture?

Exhibit

config: { 'model_type': 'decoder-only', 'rope_base': 10000, 'rope_scaling': { 'type': 'yarn', 'factor': 4.0 } }
Question 4hardmultiple choice
Full question →

When optimizing a Generative AI model using NVIDIA TensorRT-LLM, which component is primarily responsible for managing the KV cache to minimize memory fragmentation?

Question 5hardmultiple choice
Full question →

A financial institution uses NVIDIA NeMo Guardrails to enforce ethical guidelines in its customer-facing LLM. During testing, the model occasionally generates responses that violate the company's policy against offering investment advice. The guardrails are configured with a set of dialog flows and safety checks. What is the most effective way to address this issue?

Question 6hardmultiple choice
Full question →

Refer to the exhibit. Which prompt engineering technique would best force the model to prioritize technical detail over marketing language?

Exhibit

System: You are an expert at summarizing NVIDIA technical whitepapers.
User: Summarize the latest Grace Hopper Superchip architecture.
Assistant: [Model generates a 500-word response]
Issue: The response includes generic marketing fluff and lacks the requested technical depth.
Question 7hardmulti select
Full question →

When building an NVIDIA NeMo LLM application for automated document review, which THREE of the following prompt design choices are critical for ensuring high-quality output? (Select exactly THREE)

Question 8hardmulti select
Full question →

Which THREE architectural features are essential for enabling efficient inference of massive LLMs on multi-GPU systems?

Question 9hardmultiple choice
Full question →

When fine-tuning an LLM to follow specific safety protocols, why is the inclusion of 'adversarial' examples in the training data considered a best practice?

Question 10hardmultiple choice
Full question →

Refer to the exhibit. An engineer observes that memory utilization spikes during inference, causing OOM errors. Given the configuration, what is the most likely cause of the failure?

Exhibit

model_config.pbtxt:
instance_group [
  {
    count: 2
    kind: KIND_GPU
    gpus: [0]
  }
]
Question 11hardmultiple choice
Full question →

A production LLM service on NVIDIA Triton Inference Server uses dynamic batching. During peak load, the 99th percentile latency increases significantly, but GPU utilization remains at 60%. Which configuration change is most likely to improve latency while maintaining throughput?

Question 12hardmultiple choice
Full question →

A production LLM inference service on NVIDIA Triton Inference Server runs on a multi-GPU node. You observe that one GPU reports ECC XID errors and the model's throughput gradually degrades. Which NVIDIA tool should you use to monitor GPU health and set up alerts for these errors?

Question 13hardmultiple choice
Full question →

A team is preparing a mixed-language corpus for continued pretraining of an NVIDIA NeMo Megatron model. The corpus contains English, Japanese, and Arabic documents. Tokenizer analysis shows the current English-centric BPE vocabulary produces very long token sequences for Japanese and Arabic, inflating sequence length and compute cost. The team wants to reduce sequence length for non-English text without retraining the tokenizer from scratch and without degrading English performance. Which data preparation action best achieves this?

Question 14hardmulti select
Full question →

Which THREE factors significantly influence the memory consumption during LLM fine-tuning? (Choose three)

Question 15hardmultiple choice
Full question →

A bank's model risk committee is reviewing an LLM-based loan-adverse-action notice generator built on NVIDIA NeMo. Regulators require that the system produce a human-readable rationale for each denial and that the rationale be reproducible for any prior decision. Which architectural choice most directly meets both obligations?

Question 16hardmultiple choice
Full question →

Refer to the exhibit. The engineer is attempting to deploy on an NVIDIA Orin platform but encounters a runtime error. What is the most likely cause of the failure?

Network Topology
trt-builderonnx model.onnxworkspace 4096fp16dla 0
Question 17hardmultiple choice
Full question →

Refer to the exhibit. Which adjustment is the most immediate and effective way to resolve this OOM error while maintaining the same training architecture?

Exhibit

Error: CUDA out of memory. Tried to allocate 512.00 MiB. GPU 0 has 23.90 GiB total capacity. Memory used at last allocation: 23.50 GiB.
Question 18hardmultiple choice
Full question →

When optimizing a Transformer-based model using NVIDIA TensorRT, which TWO steps are essential to enable the Fusion of layer normalization and activation kernels to improve performance?

Question 19hardmultiple choice
Full question →

Which technique provides the most significant performance gain for large-model inference by splitting the model across multiple GPUs while keeping the sequence length per batch constant?

Question 20hardmulti select
Full question →

Which THREE of the following are valid methods to optimize NVIDIA GPU memory bandwidth usage in deep learning?

These NCP-GENL practice questions are part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style NCP-GENL questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.