Courseiva
Evaluation →mediumMultiple Choice

NCP-GENL Evaluation Practice Question

You are evaluating a fine-tuned LLM for a code generation task. The model was trained using NVIDIA NeMo on a dataset of Python functions. You want to measure the percentage of generated functions that pass a set of unit tests. Which evaluation metric is most appropriate?

⚠ Common exam trap

The trap here is choosing a text similarity metric like BLEU or ROUGE for code generation, when functional correctness requires execution-based metrics like pass@k.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

pass@k

The goal is to measure functional correctness by running unit tests. pass@k does exactly that by executing generated code and checking test outcomes, while also accounting for multiple samples. BLEU, ROUGE-L, and perplexity are text-based metrics that do not verify execution or test passage, so they cannot reliably assess code generation success in this scenario.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    ROUGE-L

    Why it's wrong here

    ROUGE-L computes longest common subsequence overlap between generated and reference code, focusing on textual similarity. Like BLEU, it does not execute code or verify correctness. Two functions with similar text may behave differently, and correct functions may have low overlap. Thus, ROUGE-L cannot measure the percentage of functions that pass unit tests, making it inappropriate for this evaluation.

  • ✗

    BLEU score

    Why it's wrong here

    BLEU score measures n-gram overlap between generated and reference code, which does not reflect functional correctness. Two functionally equivalent functions can have low BLEU if they use different variable names or structures. Since the scenario requires measuring whether generated functions pass unit tests, BLEU is inadequate because it ignores execution behavior and only compares surface text.

  • ✗

    Perplexity

    Why it's wrong here

    Perplexity measures how well a language model predicts a sequence of tokens, reflecting confidence in the training distribution. It does not indicate whether generated code is functionally correct or passes tests. A low perplexity model can still produce code with syntax errors or logical flaws. Therefore, perplexity is not suitable for evaluating code generation success against unit tests.

  • ✓

    pass@k

    Why this is correct

    pass@k measures the percentage of problems for which at least one of k generated samples passes all unit tests. It directly evaluates functional correctness by executing the code against test cases. This metric is standard for code generation tasks and aligns with the scenario's goal of measuring how many generated functions pass unit tests. It accounts for the stochastic nature of LLM outputs by considering multiple samples.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.