20+ practice questions focused on Evaluation — one of the most tested topics on the NVIDIA Certified Professional: Generative AI LLMs exam. Each question includes a detailed explanation so you learn why the right answer is correct.
Start Evaluation PracticeA financial services company has fine-tuned a Llama-3-70B model using NVIDIA NeMo to generate investment summaries. They need to evaluate the model's output for factual consistency against a set of gold-standard analyst reports. The team wants a metric that measures the longest common subsequence between the generated summary and the reference summary, ignoring word order. Which evaluation metric should they use?
Explanation: ROUGE-L is designed for summarization evaluation and uses the longest common subsequence to capture in-order matching without requiring contiguous matches. It balances precision and recall via an F-measure, making it suitable for comparing generated summaries against gold references. The other metrics either focus on n-gram precision, semantic similarity, or unigram alignment, and none directly compute LCS as required.
A team is evaluating a fine-tuned GPT-style model using NVIDIA NeMo. They want to assess the model's performance on a question-answering task. Which two metrics are most appropriate for measuring exact match and semantic similarity between the model's answers and reference answers? (Choose two.)
Explanation: Exact Match and F1 score are the standard metrics for question answering. Exact Match checks if the predicted answer exactly matches the reference after normalization, while F1 score computes token-level overlap, capturing partial correctness and semantic similarity. Other metrics like BLEU, ROUGE-L, and perplexity are not designed for QA and do not provide the required measurements.
You are evaluating a fine-tuned NVIDIA Nemotron-4 15B model for a customer-facing summarization task. The model was fine-tuned using NVIDIA NeMo on a domain-specific dataset. During evaluation, you notice that the model produces summaries that are fluent but frequently omit critical numerical data present in the source documents. You need to quantify this omission issue. Which evaluation metric should you prioritize?
Explanation: The model's fluent but incomplete summaries omit numerical data, so the evaluation must specifically detect missing numbers. Standard metrics like ROUGE-L or BERTScore may not penalize numerical omissions sufficiently. FactCC checks entailment but does not isolate numbers. Masking numbers as entities and using ROUGE-1 recall ensures each critical number contributes to the score, making this the most appropriate metric.
You are evaluating a large language model (LLM) deployed with NVIDIA Triton Inference Server for a customer service chatbot. The model's responses are sometimes irrelevant to the user's query, and you need to set up an automated evaluation pipeline to detect this. Which two metrics are most appropriate for this task? (Choose two.)
Explanation: Detecting irrelevance requires metrics that assess semantic similarity between the response and the expected answer. BERTScore uses contextual embeddings to capture meaning, and METEOR aligns synonyms and paraphrases, both making them robust to lexical variations. Perplexity, Exact Match, and ROUGE-L are less effective: perplexity ignores query-response relation, Exact Match is too strict, and ROUGE-L relies on surface overlap.
An enterprise fine-tunes a Llama-3-70B model using NVIDIA NeMo for automated technical support ticketing. The development team needs an automated evaluation pipeline that measures semantic similarity against human-curated reference answers without relying on costly human annotators. Which metric provides the most robust embedding-based semantic similarity assessment for this scenario?
Explanation: BERTScore leverages contextual embeddings from transformer models to evaluate token-level semantic overlap rather than exact string matching, making it ideal for technical support text where phrasing varies. This automated metric correlates strongly with human judgment, significantly accelerating iteration cycles during enterprise model development workflows on NVIDIA infrastructure.
+15 more Evaluation questions available
Practice all Evaluation questions1. Baseline your knowledge
Start with 10 questions to gauge your current understanding of Evaluation. This tells you whether you need a concept refresher or just practice.
2. Review every explanation
For each question — right or wrong — read the full explanation. Understanding why an answer is correct is more valuable than knowing the answer itself.
3. Focus on exam traps
Evaluation questions on the NCP-GENL frequently use trap wording. Look for subtle differences in answers that test your precision, not just general knowledge.
4. Reach 80% consistently
Do repeated sessions until you score 80%+ three times in a row. Then move to mixed-mode practice to test cross-topic recall under realistic conditions.
The exact number varies per candidate. Evaluation is tested as part of the NVIDIA Certified Professional: Generative AI LLMs blueprint. Practicing with targeted Evaluation questions ensures you can handle any format or difficulty that appears.
Yes. Courseiva provides free NCP-GENL practice questions across all exam topics and domains. The platform includes topic-based practice, mock exams, missed-question review, bookmarked questions, and readiness tracking — no account required.
Difficulty is subjective, but Evaluation is a high-priority exam concept tested in multiple ways — direct recall, scenario analysis, and command-output interpretation. Consistent practice is the best way to build confidence.
Launch a full Evaluation practice session with instant scoring and detailed explanations.
Start Evaluation Practice →