CCDV-F Model Selection and Cost Management Practice Question
An AI engineering team is auditing an expensive Claude API integration processing millions of customer support queries daily. Which TWO strategies should they implement to effectively reduce token expenditure without sacrificing core model intelligence? (Choose two)
⚠ Common exam trap
Candidates often try to reduce costs by shortening model output lengths or lowering generation parameters, overlooking the greater impact of optimizing input context via caching and concise prompting.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement Anthropic prompt caching for large, recurring system prompts and reference knowledge bases.
Optimizing context management and leveraging prompt caching are cornerstone strategies for cost reduction in high-volume Anthropic integrations. Caching static system prompts drastically lowers input token charges across repeated requests, while concise prompts eliminate redundant verbiage. Together, these practices protect infrastructure budgets while sustaining the semantic fidelity required by enterprise-grade support workflows.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Implement Anthropic prompt caching for large, recurring system prompts and reference knowledge bases.
Why this is correct
Prompt caching allows developers to reuse previously processed prompt blocks, offering significant cost discounts and reduced time-to-first-token for recurring system instructions. This directly targets high-volume input token costs in applications with static reference materials or lengthy guidelines.
- ✗
Switch every production workload immediately from Claude 3.5 Sonnet to Claude 3 Opus to maximize parallel processing.
Why it's wrong here
Claude 3 Opus consumes more tokens per response than Sonnet, so switching every workload raises expenditure rather than reducing it. Opus suits complex reasoning tasks where Sonnet's quality falls short, but applying it universally to millions of routine support queries ignores the cost driver the team is auditing.
- ✓
Refactor prompts to remove verbose formatting instructions, conversational filler, and redundant examples.
Why this is correct
Trimming unnecessary tokens from every incoming request reduces cumulative input token consumption. Because API pricing scales linearly with token volume, even small reductions across millions of daily support queries yield substantial financial savings over billing cycles.
- ✗
Disable response streaming to batch full payloads and reduce HTTP connection overhead.
Why it's wrong here
Disabling streaming does not reduce token consumption, since input and output tokens are billed identically regardless of delivery method. Streaming affects perceived latency, not cost. It would be correct when a client cannot handle server-sent events or requires complete payloads for downstream processing.
- ✗
Hardcode maximum output tokens to five thousand across all API calls regardless of task complexity.
Why it's wrong here
Setting excessively high maximum output token limits creates financial exposure if the model hallucinates or loops. It does not reduce costs; instead, it risks unexpected billing spikes if response lengths are unconstrained by prompt design or application logic.
About these practice questions
One of 257 original CCDV-F practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCDV-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCDV-F exam.