Courseiva
← Back to NVIDIA Certified Associate: Generative AI LLMs questions

Scenario-based practice

Select Two (Multi-Select) Questions

Practise NVIDIA Certified Associate: Generative AI LLMs practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

20
scenario questions
NCA-GENL
exam code
NVIDIA
vendor

Scenario guide

How to approach select two (multi-select) questions

Multi-select questions tell you to 'Choose TWO' or 'Choose THREE'. Getting partial credit is not a thing — you must select all correct answers with no incorrect ones. The stem always states how many to choose, so trust it. These questions require precision, not best-guess elimination.

Quick answer

Select Two (Multi-Select) Questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Related practice questions

Related NCA-GENL topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmulti select
Full question →

A team is fine-tuning an LLM and wants to detect whether individual training examples are causing unusually large gradient updates. They plan to visualize per-example gradient norms alongside other diagnostics. Which two visualizations are most appropriate for identifying these influential examples? (Choose two.)

Question 2hardmulti select
Full question →

Which THREE factors influence the reproducibility of an LLM experiment?

Question 3hardmulti select
Full question →

When designing an experiment to evaluate the performance of an LLM on a downstream classification task, which THREE factors should be controlled to ensure the results are comparable across different model sizes?

Question 4mediummulti select
Full question →

Which TWO factors should be considered when evaluating the cost-benefit of an LLM experimentation strategy?

Question 5hardmulti select
Full question →

Which THREE of the following factors are critical when choosing a foundation model for an enterprise generative AI application?

Question 6mediummulti select
Full question →

An engineer is evaluating different prompting strategies (Zero-shot, Few-shot, Chain-of-Thought) for an RAG pipeline. Which TWO metrics are most effective for quantifying the quality of the generative output during this experimentation?

Question 7mediummulti select
Full question →

A data scientist is preparing to perform hyperparameter tuning for a Retrieval-Augmented Generation (RAG) system. Which TWO parameters should be prioritized for experimentation to improve retrieval accuracy?

Question 8hardmulti select
Full question →

In the context of NVIDIA NeMo, which THREE actions are part of a robust experiment tracking workflow for fine-tuning?

Question 9hardmulti select
Full question →

Which THREE of the following strategies are recommended by NVIDIA to mitigate data leakage in enterprise-grade LLM applications?

Question 10hardmulti select
Full question →

An enterprise AI researcher is conducting ablation studies on a large language model using NVIDIA NeMo. To ensure the experimentation results are scientifically valid and statistically sound, which THREE practices must be enforced during the study?

Question 11mediummulti select
Full question →

Which THREE actions are recommended for establishing a robust 'Human-in-the-Loop' (HITL) system for an AI deployment?

Question 12mediummulti select
Full question →

When evaluating LLM output quality using human-in-the-loop data, which THREE metrics or techniques are most effective for detecting systemic hallucinations?

Question 13mediummulti select
Full question →

An ML engineer is setting up an experiment log for a fine-tuning run and wants to record the metadata necessary to reproduce the resulting model later. Which two items are most essential to capture for reproducibility? (Choose two.)

Question 14mediummulti select
Full question →

You are designing an experiment to measure how quantization (FP16 versus INT8) affects inference latency and answer quality for an LLM deployed with NVIDIA TensorRT-LLM. Which two practices are required for the comparison to be valid? (Choose two.)

Question 15mediummulti select
Full question →

A research team is running a series of controlled LLM fine-tuning experiments on NVIDIA DGX systems using NeMo Framework to compare two learning-rate schedules. They want the comparison to be scientifically valid and repeatable by another engineer next quarter. Which two practices are required to make the experiments reproducible? (Choose two.)

Question 16mediummulti select
Full question →

A research team is running a hyperparameter sweep over learning rate and warmup steps for a NeMo fine-tuning job. They notice that runs with identical hyperparameters produce different final validation losses across repeated executions. Which two changes would most directly improve the reproducibility of these experiments? (Choose two.)

Question 17mediummulti select
Full question →

An engineer is setting up an automated experiment sweep over temperature and top-p for a NeMo-served LLM, and wants the results to be comparable and reproducible. Which two practices should be applied? (Choose two.)

Question 18hardmulti select
Full question →

A research team is comparing three fine-tuning recipes for a NeMo LLM and wants the comparison to be defensible in a later review. Which two practices most improve the credibility of the reported comparison? (Choose two.)

Question 19mediummulti select
Full question →

An AI researcher is designing an experiment to compare two prompt templates for a customer-support LLM using NVIDIA NeMo. To ensure the comparison is fair and reproducible, which two practices should they follow? (Choose two.)

Question 20hardmulti select
Full question →

A research team is designing an experiment to measure how prompt phrasing affects the factuality of an LLM in a retrieval-augmented question-answering pipeline. Which two design choices are necessary to attribute observed factuality differences to the prompt rather than to other pipeline components? (Choose two.)

These NCA-GENL practice questions are part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style NCA-GENL questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.