Courseiva

CCNA Ncp Evaluation Questions

23 questions · Ncp Evaluation topic · All types, answers revealed

1
MCQmedium

An enterprise fine-tunes a Llama-3-70B model using NVIDIA NeMo for automated technical support ticketing. The development team needs an automated evaluation pipeline that measures semantic similarity against human-curated reference answers without relying on costly human annotators. Which metric provides the most robust embedding-based semantic similarity assessment for this scenario?

A.ROUGE-1 measures unigram overlap but fails to capture semantic synonyms and contextual nuances common in enterprise technical support documentation.
B.BLEU evaluates n-gram precision with a brevity penalty, heavily penalizing valid creative paraphrasing typically found in conversational AI support responses.
C.BERTScore computes token similarity matrices using contextual embeddings to capture deep semantic meaning regardless of surface-level phrasing variations.
D.Perplexity measures how well a probability distribution predicts a sample, reflecting language fluency rather than semantic alignment with a specific reference answer.
AnswerC

BERTScore aligns tokens between candidate and reference using contextual embeddings, computing precision, recall and F1 over cosine similarity. This satisfies the stem's requirement for embedding-based semantic assessment that captures meaning despite surface-level phrasing variation, without human annotators.

Why this answer

BERTScore leverages contextual embeddings from transformer models to evaluate token-level semantic overlap rather than exact string matching, making it ideal for technical support text where phrasing varies. This automated metric correlates strongly with human judgment, significantly accelerating iteration cycles during enterprise model development workflows on NVIDIA infrastructure.

Exam trap

Candidates frequently choose exact-match string metrics like BLEU or ROUGE, failing to account for semantic synonyms and variations in technical support phrasing.

2
MCQeasy

You are evaluating a text generation model using NVIDIA NeMo Evaluation and want to measure how well the generated text matches a reference translation. Which metric is specifically designed for this purpose?

A.BLEU
B.METEOR
C.Perplexity
D.ROUGE
AnswerA

BLEU (Bilingual Evaluation Understudy) is designed for machine translation evaluation. It computes n-gram precision between generated and reference translations, with a brevity penalty. It is the standard metric for translation quality and is directly applicable when a reference translation exists, making it the correct choice for this scenario.

Why this answer

BLEU is the standard metric for machine translation, measuring n-gram precision with a brevity penalty. It directly compares generated text to reference translations. Perplexity measures fluency without references, ROUGE is for summarization, and METEOR, while translation-oriented, is less commonly used in NeMo Evaluation.

Thus, BLEU is the correct metric for this scenario.

Exam trap

The trap here is selecting METEOR because it is also a translation metric, but BLEU is the primary and most widely supported metric in NeMo Evaluation for this purpose.

3
MCQhard

A research team is evaluating a large language model's robustness to adversarial attacks. They want to use NVIDIA NeMo Evaluator to measure how often the model's output changes when small, semantically preserving perturbations are applied to input prompts. Which evaluation metric or method should they implement?

A.Use exact match to compare the outputs from original and perturbed prompts.
B.Measure the perplexity of the model on the perturbed prompts.
C.Calculate the semantic similarity between the two outputs using an embedding-based metric like BERTScore.
D.Compute the BLEU score between outputs from original and perturbed prompts.
AnswerC

Robustness to semantically preserving perturbations can be assessed by measuring how similar the model's outputs are. BERTScore captures semantic equivalence, so a high score indicates the model produced consistent meaning despite input changes. This directly quantifies robustness. NeMo Evaluator can integrate BERTScore as a custom metric to automate this comparison.

Why this answer

To measure robustness to semantically preserving perturbations, the team should compare the semantic similarity of outputs from original and perturbed inputs. BERTScore provides a semantic similarity score, making it suitable. BLEU and exact match are lexical and too brittle, while perplexity does not assess output consistency.

Exam trap

The trap here is using lexical metrics like BLEU or exact match to compare outputs, which fail to capture semantic equivalence and thus misrepresent robustness.

4
MCQeasy

A developer is preparing to fine-tune a model with NVIDIA NeMo and wants a quantitative baseline before training begins. They plan to score the base model on a 300-question multiple-choice reasoning set and report accuracy. Which evaluation setup gives the most defensible baseline number?

A.Run the base model repeatedly at high temperature and report the highest accuracy observed across runs as the baseline.
B.Score the base model on the same held-out set with a fixed prompt template and deterministic decoding, then record the exact configuration alongside the accuracy.
C.Use the training split as the baseline evaluation set so the number reflects how well the model already covers the fine-tuning material.
D.Ask the model to self-report a confidence percentage for each question and average those values as the baseline accuracy.
AnswerB

A baseline is only useful if it is reproducible and comparable to later runs. Fixing the prompt template, decoding parameters, and dataset split means any future change in accuracy can be attributed to training rather than to evaluation drift. Recording the full configuration lets reviewers reproduce the number and detect accidental leakage into the fine-tuning data.

Why this answer

A trustworthy baseline requires a held-out set, a fixed prompt template, deterministic decoding, and a recorded configuration. Those choices make the number reproducible and ensure later deltas reflect training rather than evaluation noise. Best-of-runs reporting, training-set scoring, and self-reported confidence all produce numbers that either inflate performance or cannot be compared to ground-truth labels.

Exam trap

The trap here is treating a convenient or flattering number, such as best-of-runs accuracy or self-reported confidence, as a baseline when it cannot be reproduced or compared to gold labels.

5
Multi-Selecthard

A team is evaluating a large language model for a question-answering system using NVIDIA NeMo Evaluation. They need to assess both the relevance of the answer to the question and its factual correctness. (Choose two.)

Select 2 answers
A.F1 score over tokens
B.BLEU score
C.Exact Match (EM)
D.Factual consistency score using an NLI model
E.Answer relevance score from a QA evaluation model
AnswersD, E

A factual consistency score uses an NLI model to check whether the generated answer is entailed by a trusted knowledge source (e.g., a reference document). This directly measures factual correctness, the second required criterion. It is robust to paraphrasing and can be automated within NeMo Evaluation to flag hallucinations or contradictions.

Why this answer

The team must evaluate relevance and factual correctness. Answer relevance score directly measures how well the answer addresses the question, while factual consistency score via NLI checks whether the answer is factually supported. Exact Match, token F1, and BLEU focus on string overlap and do not separately assess relevance and factual correctness, making them less suitable for this dual requirement.

Exam trap

The trap here is assuming that overlap-based metrics like F1 or BLEU capture both relevance and factual correctness, when they primarily measure lexical similarity and can be misled by paraphrasing or fluent hallucinations.

6
MCQmedium

You are evaluating a fine-tuned Llama-3-70B model for a customer service chatbot using NVIDIA NeMo Evaluation. The model's outputs are factually correct but often verbose, exceeding the desired response length. Which metric should you prioritize to quantify this issue?

A.Average response length
B.BLEU
C.Perplexity
D.ROUGE-L
AnswerA

Average response length directly measures the mean number of tokens or words in generated responses. Since the issue is verbosity, this metric quantifies the problem precisely. It can be computed easily within NeMo Evaluation and compared against a target threshold to guide fine-tuning adjustments.

Why this answer

The core issue is that responses are too long, so the evaluation must quantify length. Average response length provides a direct, interpretable measure of verbosity. Other metrics like ROUGE-L, BLEU, and perplexity assess content overlap or fluency, not length, and thus cannot reliably diagnose the problem.

Exam trap

The trap here is assuming that standard text generation metrics like BLEU or ROUGE automatically penalize verbosity, when they primarily measure n-gram overlap or recall.

7
MCQhard

A research team is evaluating a large language model's ability to follow instructions. They have a dataset of prompts with corresponding reference outputs. They want to use an automated metric that correlates well with human judgments of instruction-following quality. Which evaluation method is most suitable?

A.BLEU score against reference outputs
B.GPT-4-based evaluation with a detailed rubric
C.Perplexity of the model on the reference outputs
D.ROUGE-L score against reference outputs
AnswerB

Using a strong LLM like GPT-4 as a judge with a detailed rubric has been shown to correlate well with human judgments for instruction-following tasks. The rubric can specify criteria such as adherence to format, constraints, and correctness. This method captures nuanced aspects that n-gram metrics miss, making it the most suitable for this scenario.

Why this answer

LLM-based evaluation with a detailed rubric, such as using GPT-4 as a judge, has been demonstrated to align closely with human judgments for instruction-following. It can assess adherence to constraints, format, and correctness in a way that n-gram metrics like BLEU and ROUGE-L cannot. Perplexity measures fluency, not instruction adherence.

Therefore, the LLM-as-judge approach is the most suitable.

Exam trap

The trap here is assuming that reference-based n-gram metrics like BLEU or ROUGE-L can evaluate instruction-following, when they primarily measure surface overlap and miss nuanced adherence.

8
MCQhard

A team is evaluating a retrieval-augmented generation (RAG) pipeline using NVIDIA NeMo Evaluation. They notice that the generated answers are fluent but sometimes contradict the retrieved documents. Which evaluation approach best identifies this issue?

A.Compute BLEU score against reference answers.
B.Calculate perplexity of the generated answers.
C.Measure the average length of retrieved documents.
D.Use a natural language inference (NLI) model to check entailment between retrieved documents and generated answers.
AnswerD

NLI models determine whether a hypothesis (generated answer) is entailed by, neutral to, or contradicts a premise (retrieved document). By running an NLI model, you can flag contradictions directly. This approach is robust for RAG evaluation because it assesses factual consistency without requiring reference answers, aligning with the goal of detecting contradictions.

Why this answer

The problem is factual inconsistency between retrieved documents and generated answers. Natural language inference directly evaluates entailment, identifying contradictions. Other metrics like BLEU, perplexity, or document length do not measure this relationship, making NLI the appropriate choice for this RAG evaluation scenario.

Exam trap

The trap here is relying on reference-based metrics like BLEU when the issue is faithfulness to retrieved context, not similarity to a reference answer.

9
MCQmedium

A media company is using a large language model to generate news article headlines. They want to evaluate the diversity of the generated headlines to ensure they are not overly repetitive. Which metric should they use to quantify the lexical diversity of the generated headlines?

A.Perplexity
B.ROUGE-L
C.BLEU
D.Distinct-n
AnswerD

Distinct-n measures the number of unique n-grams divided by the total number of n-grams in the generated text. It is specifically designed to quantify lexical diversity and repetition. For evaluating headline diversity, Distinct-n provides a direct measure of how varied the word choices are, making it the appropriate metric.

Why this answer

Distinct-n is a metric that calculates the proportion of unique n-grams in the generated text, directly quantifying lexical diversity. It is commonly used to evaluate repetition in generated outputs. BLEU, ROUGE-L, and perplexity assess translation quality, summarization overlap, or fluency, respectively, and do not measure diversity across multiple generations.

Therefore, Distinct-n is the correct choice.

Exam trap

The trap here is confusing fluency metrics like perplexity with diversity metrics, when diversity requires measuring unique n-grams across generations.

10
MCQhard

An evaluation team is comparing two candidate checkpoints of the same fine-tuned model on a 400-prompt open-ended task. They use an LLM-as-judge that returns a pairwise preference for each prompt. Candidate X wins 214 times, candidate Y wins 152 times, and 34 are ties. The judge is the same family as the models being compared. What should the team do before treating candidate X as the winner?

A.Discard the judge and rerun the comparison with ROUGE-L, since automatic overlap metrics are objective and free of style bias.
B.Verify the judge's reliability with a human-labeled subset, control for position bias by swapping answer order, and check that the win margin exceeds judge noise.
C.Accept the result because 214 wins out of 366 non-tie judgments is a clear majority and the sample is large enough to be conclusive.
D.Increase the prompt count to 4,000 and keep the same judging procedure, because a tenfold larger sample will average out any judge bias.
AnswerB

A same-family judge can favor outputs that resemble its own style, and pairwise judges are known to prefer whichever answer appears first. Swapping presentation order removes position bias, while a human-labeled subset quantifies how often the judge agrees with people. Only after those checks does a 62-win margin carry real evidentiary weight.

Why this answer

Pairwise LLM judging is a measurement instrument that needs validation before its verdicts are trusted. Swapping answer order neutralizes position bias, a human-labeled subset estimates judge accuracy, and comparing the win margin to that measured noise tells the team whether the difference is real. Raw majorities, a switch to overlap metrics, and simply enlarging the sample all fail to address systematic judge bias.

Exam trap

The trap here is treating a clear win count from an LLM judge as a verdict, when the judge itself is an unvalidated model that may share stylistic bias with one candidate and prefer the first position.

11
MCQmedium

You are using NVIDIA NeMo Evaluation to assess a summarization model. The model produces summaries that are grammatically correct but omit key information from the source. Which metric should you use to quantify the amount of missing content?

A.Perplexity
B.BLEU score
C.ROUGE-1 recall
D.ROUGE-1 precision
AnswerC

ROUGE-1 recall measures the proportion of unigram overlaps between the generated summary and the reference summary, relative to the reference. Low recall indicates missing content. Since the issue is omission of key information, recall is the appropriate metric to quantify missing content, as it directly reflects how much of the reference is captured.

Why this answer

The problem is omission of key information from summaries. ROUGE-1 recall directly measures how much of the reference unigrams are present in the generated summary. Low recall indicates missing content.

Precision, BLEU, and perplexity focus on relevance, precision, or fluency, not on capturing all reference content, making recall the correct choice.

Exam trap

The trap here is confusing precision and recall: precision penalizes extraneous content, while recall penalizes missing content. The scenario specifically asks about omissions.

12
MCQmedium

You are evaluating a fine-tuned LLM for a code generation task. The model was trained using NVIDIA NeMo on a dataset of Python functions. You want to measure the percentage of generated functions that pass a set of unit tests. Which evaluation metric is most appropriate?

A.ROUGE-L
B.BLEU score
C.Perplexity
D.pass@k
AnswerD

pass@k measures the percentage of problems for which at least one of k generated samples passes all unit tests. It directly evaluates functional correctness by executing the code against test cases. This metric is standard for code generation tasks and aligns with the scenario's goal of measuring how many generated functions pass unit tests. It accounts for the stochastic nature of LLM outputs by considering multiple samples.

Why this answer

The goal is to measure functional correctness by running unit tests. pass@k does exactly that by executing generated code and checking test outcomes, while also accounting for multiple samples. BLEU, ROUGE-L, and perplexity are text-based metrics that do not verify execution or test passage, so they cannot reliably assess code generation success in this scenario.

Exam trap

The trap here is choosing a text similarity metric like BLEU or ROUGE for code generation, when functional correctness requires execution-based metrics like pass@k.

13
MCQhard

A healthcare AI team is evaluating a fine-tuned GPT-based model for clinical note summarization using NVIDIA NeMo. They need to assess whether the model's summaries contain no fabricated medical facts. Which evaluation approach is most appropriate?

A.Use a factual consistency metric such as SummaC or Q2
B.Measure BERTScore against reference summaries
C.Compute perplexity of the summaries
D.Calculate ROUGE-L scores
AnswerA

Factual consistency metrics like SummaC or Q2 evaluate whether generated summaries are entailed by the source document, detecting hallucinations. They are designed to catch fabricated facts by comparing summary claims against source content. In this healthcare scenario, applying such a metric directly addresses the need to ensure no invented medical facts, making it the most suitable approach.

Why this answer

The critical requirement is detecting fabricated medical facts, which demands checking entailment between the summary and the source clinical note. Factual consistency metrics such as SummaC or Q2 are specifically designed for this purpose, unlike similarity-based metrics (BERTScore, ROUGE-L) that compare to references and may miss hallucinations. Perplexity assesses fluency, not factuality.

Therefore, a factual consistency metric is the correct evaluation approach.

Exam trap

The trap here is confusing semantic similarity to a reference summary with factual consistency against the source document, which are fundamentally different checks.

14
MCQmedium

After fine-tuning a code-generation model with NVIDIA NeMo, an engineer notices the model now produces correct domain-specific function calls but has started emitting malformed JSON in about 15 percent of structured-output requests. The fine-tuning dataset contained no structured-output examples. Which evaluation action best explains and catches this regression?

A.Add a format-compliance check plus a general-capability regression suite to the evaluation, and compare against the base model on the same structured-output prompts.
B.Increase the temperature of the structured-output requests so the model has more freedom to produce valid JSON.
C.Conclude that the fine-tune succeeded on its target task and that the JSON failures are acceptable collateral given the domain gains.
D.Retrain the model with a larger learning rate so the structured-output behavior is reinforced more strongly during fine-tuning.
AnswerA

The symptom points to a capability regression outside the fine-tuning domain, so the evaluation must cover format validity and previously working skills, not just the target task. Running the same structured-output prompts against the base model establishes whether the fine-tune caused the breakage. A format checker converts the vague observation into a measurable pass rate.

Why this answer

Fine-tuning on a narrow dataset can degrade capabilities absent from that data, and structured output is a classic casualty. The right response is to measure it: add format-compliance scoring to the harness, rerun the same structured prompts against the base and tuned models, and include a general regression suite. Retraining harder, raising temperature, or dismissing the failures all skip the measurement step the situation demands.

Exam trap

The trap here is assuming a successful domain fine-tune cannot break unrelated behaviors, so the JSON failures get explained away instead of being measured against the base model.

15
MCQeasy

You are evaluating a generative AI model for a chatbot that must adhere to strict safety guidelines. The model occasionally generates toxic or biased responses. You need to automatically evaluate the model's outputs for toxicity and bias before deployment. Which NVIDIA tool or framework is specifically designed for this purpose?

A.NVIDIA Triton Inference Server
B.NVIDIA DALI
C.NVIDIA TensorRT
D.NVIDIA NeMo Guardrails
AnswerD

NVIDIA NeMo Guardrails is a toolkit for adding programmable guardrails to LLM-based conversational systems. It can be configured to detect and block toxic or biased outputs using predefined or custom policies. It integrates with models to enforce safety guidelines in real time, making it the appropriate tool for automatically evaluating and mitigating toxicity and bias in chatbot responses.

Why this answer

NeMo Guardrails is designed to enforce safety and content policies in conversational AI, including toxicity and bias detection. The other options are infrastructure or optimization tools without safety evaluation features. Thus, NeMo Guardrails is the correct choice for automatically evaluating and mitigating toxic or biased responses before deployment.

Exam trap

The trap here is assuming that any NVIDIA tool that handles models can evaluate safety, when only NeMo Guardrails provides programmable guardrails for content moderation.

16
MCQhard

A team is evaluating a large language model (LLM) deployed with NVIDIA Triton Inference Server for a real-time question-answering system. They observe that the model's responses are accurate but latency spikes during peak load, causing timeouts. They need to evaluate the model's performance under high concurrency to identify the maximum throughput while maintaining a 95th percentile latency below 200 ms. Which evaluation approach should they use?

A.Measure GPU utilization with NVIDIA Management Library (NVML) during a single long-running inference.
B.Run a one-off inference with a batch size of 1 and measure latency.
C.Compute the model's FLOPS and compare against the GPU's peak theoretical FLOPS.
D.Use NVIDIA Triton's Performance Analyzer with concurrency sweep and latency constraints.
AnswerD

NVIDIA Triton's Performance Analyzer can simulate concurrent requests by sweeping concurrency levels and measuring latency percentiles. It supports setting a latency constraint (e.g., 95th percentile < 200 ms) and automatically finds the maximum throughput that satisfies it. This directly addresses the need to evaluate performance under high concurrency and identify the optimal operating point.

Why this answer

Evaluating performance under high concurrency requires simulating multiple simultaneous requests and measuring latency percentiles. NVIDIA Triton's Performance Analyzer is designed for this, allowing concurrency sweeps and latency constraints to find the maximum throughput meeting the SLA. Other options either lack concurrency or measure irrelevant metrics, so they cannot identify the operating point that satisfies the 95th percentile latency requirement.

Exam trap

The trap here is confusing theoretical compute limits or single-request latency with actual concurrent performance, which requires load testing with multiple simultaneous requests.

17
MCQmedium

A team has fine-tuned a Llama-3-70B model with NVIDIA NeMo and now must decide whether the tuned checkpoint actually improves on the base model for their domain. They run both models over the same 500-prompt held-out set and compute ROUGE-L against reference answers. The fine-tuned model scores 0.41 and the base model scores 0.44. What is the most technically sound conclusion the evaluation lead should draw?

A.The base model should be kept because a fine-tune that does not raise ROUGE-L proves catastrophic forgetting of the pre-training corpus.
B.The fine-tuned model is worse and should be discarded, since a lower ROUGE-L on 500 held-out prompts is a statistically significant regression.
C.ROUGE-L is insensitive to instruction-following style, so a 0.03 gap cannot be used to judge the fine-tune and the team should switch to a pairwise LLM-as-judge with human spot-checking.
D.The evaluation is valid but the held-out set is too small; expanding it to 5,000 prompts would make ROUGE-L reliable enough to select the better checkpoint.
AnswerC

ROUGE-L measures n-gram overlap with a single reference, so it penalizes valid paraphrases and rewards surface copying. A 0.03 deficit is well inside that noise floor, meaning the metric cannot resolve the question. Because the goal is domain behavior rather than lexical matching, pairwise preference judging with human verification is the appropriate instrument here.

Why this answer

ROUGE-L compares token overlap to a single reference, so it cannot distinguish a correct paraphrase from an incorrect answer, and a 0.03 gap is within that measurement noise. Since the team's goal is domain-appropriate behavior rather than lexical imitation, a pairwise preference evaluation with human adjudication gives a defensible signal. Discarding the checkpoint, enlarging the sample, or claiming forgetting all misread what the metric actually measures.

Exam trap

The trap here is assuming any numeric difference in a reference-based overlap metric reflects a real quality difference, when ROUGE-L largely measures lexical similarity rather than task correctness.

18
MCQhard

A financial institution is evaluating a fine-tuned GPT-3 model for generating investment advice summaries. They must ensure the model does not produce harmful or biased recommendations. Which evaluation methodology should they implement using NVIDIA NeMo Guardrails and NeMo Evaluator to systematically detect and quantify such issues?

A.Use NeMo Evaluator to compute the F1 score of the model's outputs against a set of correct summaries.
B.Use NeMo Guardrails to define a set of safety policies, then run the model's outputs through the guardrails and measure the percentage of violations flagged.
C.Compute the model's perplexity on a validation set of financial texts.
D.Fine-tune the model further on a dataset of unbiased financial advice and then evaluate with BLEU score.
AnswerB

NeMo Guardrails allows defining programmable rules (e.g., using Colang) to detect and block harmful or biased content. By applying these guardrails to model outputs and quantifying violation rates, the team can systematically evaluate safety and bias. This approach integrates with NeMo Evaluator to log and analyze flagged instances, providing a quantifiable metric for compliance.

Why this answer

To systematically detect and quantify harmful or biased outputs, the team should leverage NeMo Guardrails to encode safety policies and measure violation rates. This provides a direct, quantifiable assessment of safety compliance. Other metrics like perplexity, BLEU, or F1 do not target harmful content and therefore cannot fulfill the evaluation requirement.

Exam trap

The trap here is confusing general performance metrics like perplexity or F1 with safety-specific evaluation, which requires explicit policy checks.

19
MCQmedium

A team has deployed a Llama-3-8B model with NVIDIA Triton Inference Server for real-time chatbot responses. They need to evaluate the model's generation quality on a held-out test set of 500 prompts, focusing on semantic similarity to reference answers while ensuring low latency. Which evaluation approach is most appropriate to use with NVIDIA NeMo Evaluator?

A.Calculate perplexity on the test set using the model's own logits.
B.Use BERTScore or a similar embedding-based metric via NeMo Evaluator's custom metric interface.
C.Compute BLEU score using the sacreBLEU library integrated into NeMo Evaluator.
D.Run ROUGE-L scoring on the generated outputs and references.
AnswerB

BERTScore leverages contextual embeddings to compute semantic similarity between generated and reference texts, aligning with the requirement to assess meaning rather than surface form. NeMo Evaluator supports custom metrics, allowing integration of BERTScore. This approach handles paraphrases well and is efficient for 500 prompts, meeting the low-latency evaluation goal.

Why this answer

For evaluating semantic similarity in generated text, embedding-based metrics like BERTScore are preferred because they capture meaning beyond exact wording. NeMo Evaluator's extensibility allows incorporating such metrics, providing a more human-aligned assessment for chatbot responses. Perplexity and n-gram overlap metrics like BLEU and ROUGE-L do not adequately measure semantic equivalence, especially when paraphrasing is expected.

Exam trap

The trap here is assuming that any standard metric like BLEU or ROUGE automatically reflects semantic quality, when they actually measure lexical overlap.

20
Multi-Selectmedium

A team is using NVIDIA NeMo Evaluator to assess a Llama-3-70B model fine-tuned for medical question answering. They want to evaluate both the correctness of answers and the model's ability to avoid hallucinating unsupported facts. Which two evaluation strategies should they implement? (Choose two.)

Select 2 answers
A.Compute BLEU score on the generated answers.
B.Calculate perplexity of the model on a held-out set of medical questions.
C.Use a fact-checking module that cross-references generated statements with a trusted medical knowledge base.
D.Compare model outputs against gold-standard answers using exact match and F1 score.
E.Run the model through NeMo Guardrails with a policy that blocks any output containing numbers.
AnswersC, D

A fact-checking module can verify each claim in the model's output against a curated medical knowledge base, directly detecting hallucinations or unsupported facts. This approach provides a granular, evidence-based assessment of factual accuracy, which is critical in healthcare. Integrating such a module with NeMo Evaluator allows automated flagging of unsupported statements.

Why this answer

To evaluate correctness, exact match and F1 score against gold answers provide a direct comparison. To detect hallucinations, a fact-checking module that verifies statements against a trusted medical knowledge base is essential. Together, these strategies cover both aspects.

BLEU and perplexity do not measure factual accuracy, and a blanket guardrail against numbers is ineffective.

Exam trap

The trap here is assuming that any automated metric like BLEU or perplexity can substitute for factual verification, when hallucination detection requires external knowledge grounding.

21
Multi-Selecthard

You are building an evaluation harness for a retrieval-augmented generative assistant running on NVIDIA NIM microservices. The product owner wants a single trustworthy number for 'answer quality,' but you need to defend the evaluation design. Which two design choices most directly protect the evaluation from producing misleading quality scores? (Choose two.)

Select 2 answers
A.Let the same generative model that powers the assistant also grade its own answers, since it understands the domain best.
B.Hold out a prompt set that was never used during fine-tuning or prompt engineering, and keep it frozen across model versions.
C.Average the outputs of several automatic metrics into one composite score so the product owner receives a single number.
D.Increase decoding temperature during evaluation so the model explores more of the answer space and scores reflect average-case behavior.
E.Include a retrieval ablation that runs the same prompts with and without retrieved context to isolate the contribution of the retriever.
AnswersB, E

A frozen, genuinely unseen set prevents leakage and version-to-version drift in the test data itself. If prompts were recycled during prompt engineering or fine-tuning, scores inflate because the model has effectively seen the answers. Keeping the set immutable also makes run-over-run deltas interpretable, since any change in score can be attributed to the model rather than to a shifting benchmark.

Why this answer

Leakage control and retrieval attribution are the two structural safeguards that make RAG evaluation trustworthy. A frozen, unseen prompt set ensures score changes reflect the model, and the retrieval ablation shows whether gains come from context or from parametric memory. Composite averaging, self-grading, and elevated temperature all add noise or bias rather than removing it, so they undermine the credibility the product owner needs.

Exam trap

The trap here is reaching for a single blended metric or self-grading for convenience, when the real risks are benchmark contamination and unverified attribution of gains to retrieval.

22
MCQhard

A healthcare AI team is evaluating a large language model for clinical note summarization. They want to measure whether the generated summaries contain fabricated information not present in the source notes. Which evaluation approach is most appropriate for detecting hallucinated content?

A.Use a natural language inference (NLI) model to check entailment between the source note and the generated summary.
B.Compute the perplexity of the generated summaries on a held-out set of clinical notes.
C.Calculate the BLEU score between the generated summary and a reference summary written by a clinician.
D.Measure the ROUGE-L score between the generated summary and the source note.
AnswerA

NLI models determine whether a hypothesis (the summary) is entailed by a premise (the source note). If the summary contains fabricated information, the NLI model would predict contradiction or neutral rather than entailment. This approach directly assesses factual consistency and is widely used for hallucination detection in summarization, making it the most appropriate method here.

Why this answer

Natural language inference (NLI) is effective for hallucination detection because it evaluates whether the generated summary is logically entailed by the source document. If the summary introduces unsupported information, the NLI model will not predict entailment. Other metrics like perplexity, BLEU, and ROUGE-L focus on fluency or surface overlap and cannot reliably identify fabricated content.

Therefore, NLI-based checking is the most appropriate approach.

Exam trap

The trap here is confusing fluency metrics like perplexity with factual consistency metrics, when hallucination detection requires comparing the summary against the source for entailment.

23
MCQmedium

A financial services company has fine-tuned a Llama-3-70B model using NVIDIA NeMo for automated loan risk assessment. During evaluation, they observe that the model achieves high scores on standard accuracy metrics but produces inconsistent responses when the same query is phrased slightly differently. They need a metric that quantifies this inconsistency. Which evaluation metric should they use?

A.Semantic similarity between outputs for paraphrased inputs
B.ROUGE score
C.BLEU score
D.Perplexity
AnswerA

Semantic similarity measures how closely the model's responses align in meaning when the same question is asked with different wording. High similarity indicates consistent understanding and stable generation. In this scenario, computing similarity across paraphrased loan queries directly quantifies the observed inconsistency. This metric is appropriate for detecting and reducing output variance.

Why this answer

The team needs to measure output consistency across semantically equivalent inputs. Semantic similarity between responses to paraphrased queries directly captures this variance, unlike lexical overlap metrics such as BLEU or ROUGE, which compare to references. Perplexity reflects language modeling confidence, not answer stability.

Therefore, semantic similarity is the correct choice for quantifying inconsistency in the fine-tuned model's responses.

Exam trap

The trap here is assuming that high accuracy on standard benchmarks implies consistent behavior across paraphrased inputs, which it does not.

Ready to test yourself?

Try a timed practice session using only Ncp Evaluation questions.