NCP-GENL Evaluation Practice Question
A research team is evaluating a large language model's robustness to adversarial attacks. They want to use NVIDIA NeMo Evaluator to measure how often the model's output changes when small, semantically preserving perturbations are applied to input prompts. Which evaluation metric or method should they implement?
⚠ Common exam trap
The trap here is using lexical metrics like BLEU or exact match to compare outputs, which fail to capture semantic equivalence and thus misrepresent robustness.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Calculate the semantic similarity between the two outputs using an embedding-based metric like BERTScore.
To measure robustness to semantically preserving perturbations, the team should compare the semantic similarity of outputs from original and perturbed inputs. BERTScore provides a semantic similarity score, making it suitable. BLEU and exact match are lexical and too brittle, while perplexity does not assess output consistency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use exact match to compare the outputs from original and perturbed prompts.
Why it's wrong here
Exact match would only count identical strings, which is too strict; semantically equivalent outputs with different wording would be marked as inconsistent. This would underestimate robustness. A semantic similarity metric is needed to properly assess meaning preservation.
- ✗
Measure the perplexity of the model on the perturbed prompts.
Why it's wrong here
Perplexity measures how well the model predicts the perturbed input itself, not the consistency of its generated outputs. A model could have low perplexity on perturbed prompts yet produce different answers. Therefore, perplexity does not measure robustness as defined here.
- ✓
Calculate the semantic similarity between the two outputs using an embedding-based metric like BERTScore.
Why this is correct
Robustness to semantically preserving perturbations can be assessed by measuring how similar the model's outputs are. BERTScore captures semantic equivalence, so a high score indicates the model produced consistent meaning despite input changes. This directly quantifies robustness. NeMo Evaluator can integrate BERTScore as a custom metric to automate this comparison.
- ✗
Compute the BLEU score between outputs from original and perturbed prompts.
Why it's wrong here
BLEU measures n-gram overlap between a candidate and reference, not the consistency between two model outputs. Using BLEU to compare outputs from original and perturbed inputs would not directly quantify robustness; it would only show lexical similarity, which may be high even if the meaning changes. Thus, it is not the appropriate metric.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.