Courseiva

NCP-GENL · topic practice

Evaluation practice questions

This domain covers evaluating generative AI LLMs within the NVIDIA ecosystem, including NeMo, NIM, and Triton. It tests designing automated pipelines for semantic similarity, detecting format errors, building RAG evaluation harnesses, and computing quality metrics on held-out sets. Questions present real-world scenarios requiring selection of appropriate tools and metrics.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Evaluation

What the exam tests

What to know about Evaluation

Candidates must design and implement evaluation pipelines using NVIDIA tools like NeMo Evaluator, NIM, and Triton. The most important thing is to select appropriate metrics (e.g., semantic similarity, format validation) and avoid relying on a single number for answer quality.

Using NeMo Evaluator for semantic similarity against human-curated references

Detecting malformed JSON in structured outputs from fine-tuned models

Building evaluation harnesses for RAG with NVIDIA NIM microservices

Computing generation quality metrics on held-out test sets with Triton

Watch out for

Common Evaluation exam traps

  • ▸Assuming a single metric like BLEU or ROUGE fully captures answer quality; instead, combine semantic similarity, format validation, and task-specific checks.
  • ▸Overlooking that fine-tuning can degrade structured output formatting; always validate JSON and other schemas separately from semantic correctness.
  • ▸Neglecting to use held-out data or human-curated references, leading to overfitting and unreliable evaluation results.

Practice set

Evaluation questions

20 questions · select your answer, then reveal the explanation

Question 1mediummultiple choice
Read the full Evaluation explanation →

A financial services company has fine-tuned a Llama-3-70B model using NVIDIA NeMo to generate investment summaries. They need to evaluate the model's output for factual consistency against a set of gold-standard analyst reports. The team wants a metric that measures the longest common subsequence between the generated summary and the reference summary, ignoring word order. Which evaluation metric should they use?

Question 2mediummulti select
Read the full Evaluation explanation →

A team is evaluating a fine-tuned GPT-style model using NVIDIA NeMo. They want to assess the model's performance on a question-answering task. Which two metrics are most appropriate for measuring exact match and semantic similarity between the model's answers and reference answers? (Choose two.)

Question 3mediummultiple choice
Read the full Evaluation explanation →

You are evaluating a fine-tuned NVIDIA Nemotron-4 15B model for a customer-facing summarization task. The model was fine-tuned using NVIDIA NeMo on a domain-specific dataset. During evaluation, you notice that the model produces summaries that are fluent but frequently omit critical numerical data present in the source documents. You need to quantify this omission issue. Which evaluation metric should you prioritize?

You are evaluating a large language model (LLM) deployed with NVIDIA Triton Inference Server for a customer service chatbot. The model's responses are sometimes irrelevant to the user's query, and you need to set up an automated evaluation pipeline to detect this. Which two metrics are most appropriate for this task? (Choose two.)

Question 5mediummultiple choice
Read the full Evaluation explanation →

An enterprise fine-tunes a Llama-3-70B model using NVIDIA NeMo for automated technical support ticketing. The development team needs an automated evaluation pipeline that measures semantic similarity against human-curated reference answers without relying on costly human annotators. Which metric provides the most robust embedding-based semantic similarity assessment for this scenario?

Question 6mediummultiple choice
Read the full Evaluation explanation →

A financial services company has fine-tuned a Llama-3-70B model using NVIDIA NeMo for automated loan risk assessment. During evaluation, they observe that the model achieves high scores on standard accuracy metrics but produces inconsistent responses when the same query is phrased slightly differently. They need a metric that quantifies this inconsistency. Which evaluation metric should they use?

Question 7hardmultiple choice
Read the full Evaluation explanation →

A healthcare AI team is evaluating a fine-tuned GPT-based model for clinical note summarization using NVIDIA NeMo. They need to assess whether the model's summaries contain no fabricated medical facts. Which evaluation approach is most appropriate?

Question 8mediummultiple choice
Read the full Evaluation explanation →

You are evaluating a fine-tuned Llama-3-70B model for a customer service chatbot using NVIDIA NeMo Evaluation. The model's outputs are factually correct but often verbose, exceeding the desired response length. Which metric should you prioritize to quantify this issue?

Question 9hardmultiple choice
Read the full Evaluation explanation →

A team is evaluating a retrieval-augmented generation (RAG) pipeline using NVIDIA NeMo Evaluation. They notice that the generated answers are fluent but sometimes contradict the retrieved documents. Which evaluation approach best identifies this issue?

Question 10mediummultiple choice
Read the full Evaluation explanation →

You are using NVIDIA NeMo Evaluation to assess a summarization model. The model produces summaries that are grammatically correct but omit key information from the source. Which metric should you use to quantify the amount of missing content?

Question 11hardmultiple choice
Read the full Evaluation explanation →

A healthcare AI team is evaluating a large language model for clinical note summarization. They want to measure whether the generated summaries contain fabricated information not present in the source notes. Which evaluation approach is most appropriate for detecting hallucinated content?

Question 12hardmulti select
Read the full Evaluation explanation →

A team is evaluating a large language model for a question-answering system using NVIDIA NeMo Evaluation. They need to assess both the relevance of the answer to the question and its factual correctness. (Choose two.)

Question 13mediummultiple choice
Read the full Evaluation explanation →

A team has deployed a Llama-3-8B model with NVIDIA Triton Inference Server for real-time chatbot responses. They need to evaluate the model's generation quality on a held-out test set of 500 prompts, focusing on semantic similarity to reference answers while ensuring low latency. Which evaluation approach is most appropriate to use with NVIDIA NeMo Evaluator?

Question 14easymultiple choice
Read the full Evaluation explanation →

You are evaluating a text generation model using NVIDIA NeMo Evaluation and want to measure how well the generated text matches a reference translation. Which metric is specifically designed for this purpose?

Question 15hardmultiple choice
Read the full Evaluation explanation →

A financial institution is evaluating a fine-tuned GPT-3 model for generating investment advice summaries. They must ensure the model does not produce harmful or biased recommendations. Which evaluation methodology should they implement using NVIDIA NeMo Guardrails and NeMo Evaluator to systematically detect and quantify such issues?

Question 16mediummultiple choice
Read the full Evaluation explanation →

A media company is using a large language model to generate news article headlines. They want to evaluate the diversity of the generated headlines to ensure they are not overly repetitive. Which metric should they use to quantify the lexical diversity of the generated headlines?

Question 17mediummulti select
Read the full Evaluation explanation →

A team is using NVIDIA NeMo Evaluator to assess a Llama-3-70B model fine-tuned for medical question answering. They want to evaluate both the correctness of answers and the model's ability to avoid hallucinating unsupported facts. Which two evaluation strategies should they implement? (Choose two.)

Question 18hardmultiple choice
Read the full Evaluation explanation →

A research team is evaluating a large language model's ability to follow instructions. They have a dataset of prompts with corresponding reference outputs. They want to use an automated metric that correlates well with human judgments of instruction-following quality. Which evaluation method is most suitable?

Question 19hardmultiple choice
Read the full Evaluation explanation →

A research team is evaluating a large language model's robustness to adversarial attacks. They want to use NVIDIA NeMo Evaluator to measure how often the model's output changes when small, semantically preserving perturbations are applied to input prompts. Which evaluation metric or method should they implement?

Question 20hardmultiple choice
Read the full Evaluation explanation →

A team is evaluating a large language model (LLM) deployed with NVIDIA Triton Inference Server for a real-time question-answering system. They observe that the model's responses are accurate but latency spikes during peak load, causing timeouts. They need to evaluate the model's performance under high concurrency to identify the maximum throughput while maintaining a 95th percentile latency below 200 ms. Which evaluation approach should they use?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Evaluation sessions

Start a Evaluation only practice session

Every question in these sessions is drawn from the Evaluation domain — nothing else.

Related practice questions

Related NCP-GENL topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the NCP-GENL exam test about Evaluation?
Candidates must design and implement evaluation pipelines using NVIDIA tools like NeMo Evaluator, NIM, and Triton. The most important thing is to select appropriate metrics (e.g., semantic similarity, format validation) and avoid relying on a single number for answer quality.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Evaluation questions in a focused session?
Yes — the session launcher on this page draws every question from the Evaluation domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other NCP-GENL topics?
Use the topic links above to move to related areas, or go back to the NCP-GENL question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the NCP-GENL exam covers. They are not copied from any real exam or dump site.