Courseiva

CCDV-F Prompt and Context Engineering Practice Question

A developer is iterating on a classification prompt for support tickets. Each test run uses a different random sample of 200 tickets, and accuracy swings by 8 points between runs. The prompt itself is unchanged. What is the best first step to get trustworthy signal about whether a prompt edit actually helped?

⚠ Common exam trap

The trap here is attributing accuracy swings to model randomness and reaching for temperature, when the dominant source of variance is the changing evaluation sample.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Freeze a held-out evaluation set and run every prompt version against the same fixed examples.

A fixed held-out evaluation set is the foundation of reliable prompt iteration. By scoring every prompt version against identical examples, the developer isolates the effect of the prompt change from the effect of sampling. Temperature and example count affect generation, but neither controls which inputs are being compared.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase max_tokens so the model has more room to reason before classifying.

    Why it's wrong here

    Output length is not the source of the variance here; the input sample is. Raising max_tokens may change behavior slightly, but it does not control which tickets are evaluated, so run-to-run swings would persist. The developer needs a controlled comparison, not a different generation budget.

  • ✓

    Freeze a held-out evaluation set and run every prompt version against the same fixed examples.

    Why this is correct

    Holding the evaluation set constant removes sampling noise as a confounder. Any accuracy difference between prompt versions then reflects the prompt change rather than which tickets happened to be drawn. This is the standard way to make prompt iteration measurable instead of anecdotal.

  • ✗

    Add more few-shot examples until accuracy stops changing between runs.

    Why it's wrong here

    More examples may raise absolute accuracy, but they do not stabilize measurement across changing samples. The run-to-run swing is an evaluation design problem, not an example-count problem. Without a fixed test set, the developer cannot tell whether the added examples helped or whether a lucky sample did.

  • ✗

    Lower the temperature to 0 and rerun the same random sample twice.

    Why it's wrong here

    Temperature 0 reduces generation variance but does not address the changing input sample. Two runs over different random draws can still differ by several points even at temperature 0, because the model is being asked about different tickets. The developer would still be comparing apples to oranges.

About these practice questions

Courseiva writes every CCDV-F question from scratch — 257 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCDV-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCDV-F exam.