CCDV-F Model Selection and Cost Management Practice Question
A developer is building an application that uses the Claude Messages API to generate product descriptions. The application sends a system prompt of 1,500 tokens, a user prompt of 200 tokens, and receives a response of 300 tokens. The developer wants to reduce costs. Which two strategies would directly reduce the cost per API call? (Choose two.)
⚠ Common exam trap
Test-takers frequently confuse features that improve performance or reduce call volume, like streaming or caching, with those that directly lower the token-based cost of a single API call.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a smaller model like Claude 3 Haiku instead of Claude 3 Opus.
The cost per API call is determined by the number of input and output tokens and the model's pricing. Shortening the system prompt reduces input tokens, and using a smaller model like Claude 3 Haiku lowers the per-token cost. Both directly decrease the cost of each call. Increasing max_tokens, enabling streaming, or caching responses do not reduce the token-based cost of an individual call.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a smaller model like Claude 3 Haiku instead of Claude 3 Opus.
Why this is correct
Claude 3 Haiku has a lower cost per token than Claude 3 Opus. Switching to Haiku for a task like generating product descriptions, which does not require the advanced reasoning of Opus, reduces the cost per API call. This is a direct cost-saving measure as long as the smaller model meets quality requirements.
- ✗
Enable streaming to receive tokens as they are generated.
Why it's wrong here
Streaming affects how tokens are delivered to the client, not how they are billed. The total number of input and output tokens remains the same, so the cost per API call is unchanged. Streaming improves perceived latency but does not reduce token costs.
- ✓
Shorten the system prompt by removing redundant instructions.
Why this is correct
Reducing the system prompt length directly decreases the number of input tokens sent to the model. Since input tokens are billed per token, a shorter system prompt lowers the cost per call. This is a straightforward and effective cost-reduction strategy, provided the essential instructions remain to maintain output quality.
- ✗
Increase the max_tokens parameter to allow longer responses.
Why it's wrong here
Increasing max_tokens allows the model to generate more tokens, which would increase output token costs if the model uses them. It does not reduce cost; it could raise it. The parameter only sets an upper limit, so it does not directly affect cost unless the model generates more tokens as a result.
- ✗
Cache the model's responses for identical prompts.
Why it's wrong here
Caching responses on the client side can reduce the number of API calls if identical prompts are repeated, but it does not reduce the cost of a single API call. The question asks for strategies that directly reduce the cost per call, and caching addresses call volume, not per-call cost.
About these practice questions
One of 257 original CCDV-F practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCDV-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCDV-F exam.