CCDV-F Model Selection and Cost Management Practice Question
A developer is using the Claude API to generate creative writing pieces. The application often sends the same 2,000-token instruction set as a prefix to every request. Which feature can reduce the cost of processing this repeated prefix?
⚠ Common exam trap
The trap here is assuming that any cost-saving feature, such as using a smaller model or the Batch API, will address the specific issue of a repeated prefix, when Prompt Caching is the only feature that directly caches and reuses that prefix.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Prompt Caching
Prompt Caching is specifically designed to reduce costs when the same prefix is reused across multiple API calls. By caching the 2,000-token instruction set, subsequent calls avoid reprocessing those tokens at full input rates, directly lowering the cost. Other features like streaming, Batch API, or model selection do not target the redundant processing of a repeated prefix.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Batch API
Why it's wrong here
The Batch API is designed for asynchronous processing of large volumes of requests at a discount, but it does not cache repeated prefixes within individual calls. It reduces cost through volume discounts, not through prefix reuse. For a repeated prefix, Prompt Caching is more direct.
- ✓
Prompt Caching
Why this is correct
Prompt Caching allows developers to cache a static prefix, such as a long instruction set, so that it is not re-processed and re-billed at full input token rates on subsequent calls. This directly reduces the cost of the repeated 2,000-token prefix, making it the ideal solution for this scenario.
- ✗
Streaming responses
Why it's wrong here
Streaming responses only changes how the output is delivered to the client, not how input tokens are billed. The repeated prefix would still be processed and charged at full rate on each call. Streaming does not address the cost of redundant input tokens.
- ✗
Using a smaller model
Why it's wrong here
Switching to a smaller model like Claude 3 Haiku can lower per-token costs, but it does not specifically address the repeated prefix. The entire prefix would still be processed and billed each time. Prompt Caching is the feature designed to avoid re-processing and re-billing the same prefix.
About these practice questions
Courseiva writes every CCDV-F question from scratch — 257 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCDV-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCDV-F exam.