Courseiva

NCP-GENL · domain

Evaluation

This domain covers evaluating generative AI LLMs within the NVIDIA ecosystem, including NeMo, NIM, and Triton. It tests designing automated pipelines for semantic similarity, detecting format errors, building RAG evaluation harnesses, and computing quality metrics on held-out sets. Questions present real-world scenarios requiring selection of appropriate tools and metrics.

23 questions3 easy10 medium10 hard

Focused practice

Practice Evaluation questions

Scored sessions drawing only from this domain — pick a length below.

What this domain covers

What to know about Evaluation

Candidates must design and implement evaluation pipelines using NVIDIA tools like NeMo Evaluator, NIM, and Triton. The most important thing is to select appropriate metrics (e.g., semantic similarity, format validation) and avoid relying on a single number for answer quality.

Using NeMo Evaluator for semantic similarity against human-curated references

Detecting malformed JSON in structured outputs from fine-tuned models

Building evaluation harnesses for RAG with NVIDIA NIM microservices

Computing generation quality metrics on held-out test sets with Triton

Watch out for

Common Evaluation exam traps

  • ▸Assuming a single metric like BLEU or ROUGE fully captures answer quality; instead, combine semantic similarity, format validation, and task-specific checks.
  • ▸Overlooking that fine-tuning can degrade structured output formatting; always validate JSON and other schemas separately from semantic correctness.
  • ▸Neglecting to use held-out data or human-curated references, leading to overfitting and unreliable evaluation results.

Question index

All Evaluation questions (23)

Click any question to see the full explanation, or start a practice session above.

1

An enterprise fine-tunes a Llama-3-70B model using NVIDIA NeMo for automated technical support ticketing. The development team needs an automated evaluation pipeline that measures semantic similarity against human-curated reference answers without relying on costly human annotators. Which metric provides the most robust embedding-based semantic similarity assessment for this scenario?

Medium
2

You are evaluating a text generation model using NVIDIA NeMo Evaluation and want to measure how well the generated text matches a reference translation. Which metric is specifically designed for this purpose?

Easy
3

A research team is evaluating a large language model's robustness to adversarial attacks. They want to use NVIDIA NeMo Evaluator to measure how often the model's output changes when small, semantically preserving perturbations are applied to input prompts. Which evaluation metric or method should they implement?

Hard
4

A developer is preparing to fine-tune a model with NVIDIA NeMo and wants a quantitative baseline before training begins. They plan to score the base model on a 300-question multiple-choice reasoning set and report accuracy. Which evaluation setup gives the most defensible baseline number?

Easy
5

A team is evaluating a large language model for a question-answering system using NVIDIA NeMo Evaluation. They need to assess both the relevance of the answer to the question and its factual correctness. (Choose two.)

Hard
6

You are evaluating a fine-tuned Llama-3-70B model for a customer service chatbot using NVIDIA NeMo Evaluation. The model's outputs are factually correct but often verbose, exceeding the desired response length. Which metric should you prioritize to quantify this issue?

Medium
7

A research team is evaluating a large language model's ability to follow instructions. They have a dataset of prompts with corresponding reference outputs. They want to use an automated metric that correlates well with human judgments of instruction-following quality. Which evaluation method is most suitable?

Hard
8

A team is evaluating a retrieval-augmented generation (RAG) pipeline using NVIDIA NeMo Evaluation. They notice that the generated answers are fluent but sometimes contradict the retrieved documents. Which evaluation approach best identifies this issue?

Hard
9

A media company is using a large language model to generate news article headlines. They want to evaluate the diversity of the generated headlines to ensure they are not overly repetitive. Which metric should they use to quantify the lexical diversity of the generated headlines?

Medium
10

An evaluation team is comparing two candidate checkpoints of the same fine-tuned model on a 400-prompt open-ended task. They use an LLM-as-judge that returns a pairwise preference for each prompt. Candidate X wins 214 times, candidate Y wins 152 times, and 34 are ties. The judge is the same family as the models being compared. What should the team do before treating candidate X as the winner?

Hard
11

You are using NVIDIA NeMo Evaluation to assess a summarization model. The model produces summaries that are grammatically correct but omit key information from the source. Which metric should you use to quantify the amount of missing content?

Medium
12

You are evaluating a fine-tuned LLM for a code generation task. The model was trained using NVIDIA NeMo on a dataset of Python functions. You want to measure the percentage of generated functions that pass a set of unit tests. Which evaluation metric is most appropriate?

Medium
13

A healthcare AI team is evaluating a fine-tuned GPT-based model for clinical note summarization using NVIDIA NeMo. They need to assess whether the model's summaries contain no fabricated medical facts. Which evaluation approach is most appropriate?

Hard
14

After fine-tuning a code-generation model with NVIDIA NeMo, an engineer notices the model now produces correct domain-specific function calls but has started emitting malformed JSON in about 15 percent of structured-output requests. The fine-tuning dataset contained no structured-output examples. Which evaluation action best explains and catches this regression?

Medium
15

You are evaluating a generative AI model for a chatbot that must adhere to strict safety guidelines. The model occasionally generates toxic or biased responses. You need to automatically evaluate the model's outputs for toxicity and bias before deployment. Which NVIDIA tool or framework is specifically designed for this purpose?

Easy
16

A team is evaluating a large language model (LLM) deployed with NVIDIA Triton Inference Server for a real-time question-answering system. They observe that the model's responses are accurate but latency spikes during peak load, causing timeouts. They need to evaluate the model's performance under high concurrency to identify the maximum throughput while maintaining a 95th percentile latency below 200 ms. Which evaluation approach should they use?

Hard
17

A team has fine-tuned a Llama-3-70B model with NVIDIA NeMo and now must decide whether the tuned checkpoint actually improves on the base model for their domain. They run both models over the same 500-prompt held-out set and compute ROUGE-L against reference answers. The fine-tuned model scores 0.41 and the base model scores 0.44. What is the most technically sound conclusion the evaluation lead should draw?

Medium
18

A financial institution is evaluating a fine-tuned GPT-3 model for generating investment advice summaries. They must ensure the model does not produce harmful or biased recommendations. Which evaluation methodology should they implement using NVIDIA NeMo Guardrails and NeMo Evaluator to systematically detect and quantify such issues?

Hard
19

A team has deployed a Llama-3-8B model with NVIDIA Triton Inference Server for real-time chatbot responses. They need to evaluate the model's generation quality on a held-out test set of 500 prompts, focusing on semantic similarity to reference answers while ensuring low latency. Which evaluation approach is most appropriate to use with NVIDIA NeMo Evaluator?

Medium
20

A team is using NVIDIA NeMo Evaluator to assess a Llama-3-70B model fine-tuned for medical question answering. They want to evaluate both the correctness of answers and the model's ability to avoid hallucinating unsupported facts. Which two evaluation strategies should they implement? (Choose two.)

Medium
21

You are building an evaluation harness for a retrieval-augmented generative assistant running on NVIDIA NIM microservices. The product owner wants a single trustworthy number for 'answer quality,' but you need to defend the evaluation design. Which two design choices most directly protect the evaluation from producing misleading quality scores? (Choose two.)

Hard
22

A healthcare AI team is evaluating a large language model for clinical note summarization. They want to measure whether the generated summaries contain fabricated information not present in the source notes. Which evaluation approach is most appropriate for detecting hallucinated content?

Hard
23

A financial services company has fine-tuned a Llama-3-70B model using NVIDIA NeMo for automated loan risk assessment. During evaluation, they observe that the model achieves high scores on standard accuracy metrics but produces inconsistent responses when the same query is phrased slightly differently. They need a metric that quantifies this inconsistency. Which evaluation metric should they use?

Medium

Frequently asked questions

What does the Evaluation domain cover on the NCP-GENL exam?
Candidates must design and implement evaluation pipelines using NVIDIA tools like NeMo Evaluator, NIM, and Triton. The most important thing is to select appropriate metrics (e.g., semantic similarity, format validation) and avoid relying on a single number for answer quality.
How many questions are in this domain?
This page lists all 23 Evaluation questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Evaluation questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
nvidia-ncp-genl NVIDIA-NCP-GENL ncp evaluation Practice Questions