Courseiva

CCDV-F Model Selection and Cost Management Practice Question

A team is estimating costs for a new feature that will send 1 million requests per month. Each request has a 2,000-token input and generates a 500-token output. The team wants to reduce the output token cost, which dominates the bill. Which strategy is MOST effective for reducing output token costs?

⚠ Common exam trap

The trap here is applying a familiar optimization like Prompt Caching reflexively, even when the scenario explicitly states that output tokens, not input tokens, are the dominant cost.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Instruct the model to be concise and set a lower max_tokens limit to cap response length.

Because output tokens dominate the monthly bill, the most effective lever is limiting how many tokens the model generates. Lowering max_tokens and prompting for conciseness directly caps output length and thus cost. Optimizing input tokens or context capacity addresses smaller or irrelevant components of the bill.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable Prompt Caching on the input to reduce the input token cost.

    Why it's wrong here

    Prompt Caching reduces input token costs, but the scenario states that output tokens dominate the bill. Optimizing the smaller cost component yields limited total savings. This is a valid optimization in general but not the most effective one for reducing a bill driven primarily by output tokens.

  • ✗

    Increase the max_tokens parameter to allow longer responses.

    Why it's wrong here

    Raising max_tokens permits longer outputs, which would increase output token costs rather than reduce them. Output tokens are billed per token generated, so allowing more tokens directly raises the bill. This strategy works against the team's stated goal of reducing the dominant cost component.

  • ✗

    Switch to a model with a larger context window to accommodate longer inputs.

    Why it's wrong here

    A larger context window addresses input capacity, not output token pricing. Output tokens are billed at their own rate regardless of context window size. This change does not reduce output costs and may even increase per-token pricing if the larger-context model is more expensive.

  • ✓

    Instruct the model to be concise and set a lower max_tokens limit to cap response length.

    Why this is correct

    Output tokens are billed per generated token, so constraining response length directly reduces the dominant cost. Instructing conciseness and lowering max_tokens caps the worst-case output size. This approach targets the largest cost driver without sacrificing the feature's purpose, since the responses remain useful but shorter.

About these practice questions

This CCDV-F question is part of Courseiva's 257-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCDV-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCDV-F exam.