Courseiva
← Back to NVIDIA Certified Professional: Generative AI LLMs questions

Scenario-based practice

Select Two (Multi-Select) Questions

Practise NVIDIA Certified Professional: Generative AI LLMs practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

20
scenario questions
NCP-GENL
exam code
NVIDIA
vendor

Scenario guide

How to approach select two (multi-select) questions

Multi-select questions tell you to 'Choose TWO' or 'Choose THREE'. Getting partial credit is not a thing — you must select all correct answers with no incorrect ones. The stem always states how many to choose, so trust it. These questions require precision, not best-guess elimination.

Quick answer

Select Two (Multi-Select) Questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Related practice questions

Related NCP-GENL topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmulti select
Full question →

When profiling an application with NVIDIA Nsight Systems, which TWO metrics are most critical to identify if an application is limited by the PCIe bus?

Question 2mediummulti select
Full question →

Which TWO of the following techniques are best suited for reducing the latency of LLM inference on NVIDIA GPUs?

Question 3mediummulti select
Full question →

Which TWO actions should be part of a robust incident response plan for an LLM deployment failing in production?

Question 4hardmulti select
Full question →

When building an NVIDIA NeMo LLM application for automated document review, which THREE of the following prompt design choices are critical for ensuring high-quality output? (Select exactly THREE)

Question 5hardmulti select
Full question →

Which THREE architectural features are essential for enabling efficient inference of massive LLMs on multi-GPU systems?

Question 6mediummulti select
Full question →

A team is optimizing an NVIDIA TensorRT-LLM deployment of a 70B model on multiple GPUs. They want to reduce inter-GPU communication overhead and improve throughput. Which two techniques should they consider? (Choose two.)

Question 7mediummulti select
Full question →

You are preparing a dataset for supervised fine-tuning (SFT) of an NVIDIA NeMo LLM to follow instructions. Which TWO data preparation practices are essential to ensure the model learns to generalize rather than memorize? (Choose two.)

Question 8hardmulti select
Full question →

Which THREE factors significantly influence the memory consumption during LLM fine-tuning? (Choose three)

Question 9mediummulti select
Full question →

When assessing the quality of a dataset for instruction fine-tuning, which TWO metrics or methods are considered most reliable for measuring dataset diversity?

Question 10mediummulti select
Full question →

Which TWO of the following practices are recommended for ensuring ethical AI development when using NVIDIA NIMs in an enterprise environment?

Question 11hardmulti select
Full question →

You are preparing a massive dataset for training a NeMo-based LLM. Which TWO data preprocessing steps are critical to prevent data leakage and ensure model quality?

Question 12hardmulti select
Full question →

Which TWO of the following practices are considered standard procedures for preparing a dataset for Instruction Fine-Tuning (IFT)? (Choose two)

Question 13mediummulti select
Full question →

Which TWO factors should be prioritized when selecting a quantization strategy for deploying a large language model on constrained edge hardware? (Select TWO)

Question 14hardmulti select
Full question →

You are building an evaluation harness for a retrieval-augmented generative assistant running on NVIDIA NIM microservices. The product owner wants a single trustworthy number for 'answer quality,' but you need to defend the evaluation design. Which two design choices most directly protect the evaluation from producing misleading quality scores? (Choose two.)

Question 15hardmulti select
Full question →

An engineer is reviewing the architecture of a decoder-only LLM that must support very long input contexts for document analysis. They are considering architectural choices that extend effective context length beyond what the model saw during pretraining. Which TWO techniques are designed specifically to extend usable context length without retraining the entire model from scratch? (Choose two.)

Question 16hardmulti select
Full question →

An engineer is analyzing why a decoder-only LLM with 32,000-token context length fails to answer questions that require information from the beginning of a long document when the answer is near the end. The model was trained with standard causal attention. Which two architectural or training factors are most likely contributing to this failure? (Choose two.)

Question 17mediummulti select
Full question →

Which TWO of the following NVIDIA AI Enterprise tools are specifically designed to optimize and accelerate the deployment of LLMs in containerized environments?

Question 18mediummulti select
Full question →

A developer is optimizing a generative AI model for inference on NVIDIA GPUs. They want to reduce memory footprint and improve throughput without sacrificing accuracy. Which two techniques should they apply? (Choose two.)

Question 19mediummulti select
Full question →

An engineer is using NVIDIA TensorRT-LLM to optimize an LLM for inference. They want to reduce the memory footprint of the KV cache during long-context generation. Which TWO techniques are supported by TensorRT-LLM to achieve this? (Choose two.)

Question 20hardmulti select
Full question →

A developer is using NVIDIA TensorRT-LLM to optimize a GPT-based model for inference. They want to reduce the model's memory footprint and improve throughput without retraining. Which two techniques can be applied during the TensorRT-LLM build process to achieve these goals? (Choose two.)

These NCP-GENL practice questions are part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style NCP-GENL questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.