CCDV-F Model Selection and Cost Management Practice Question
A developer is building a long-context application that processes 150,000 tokens per request. To manage costs and maintain performance, which TWO techniques should be prioritized?
⚠ Common exam trap
Candidates often suggest summarizing the entire context before sending it, which ignores the efficiency of prompt caching for static data and the precision of RAG for dynamic data retrieval.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implementing Prompt Caching for the static background context.
Managing long-context applications requires a combination of architectural efficiency and cost-saving features. Developers must minimize redundant data processing and ensure the model remains focused on relevant information. Utilizing prompt caching for static data and implementing RAG to limit the context sent to the model are the two most effective strategies for scaling long-context applications while keeping costs under control.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Implementing Prompt Caching for the static background context.
Why this is correct
Prompt caching allows the developer to store the 150,000-token context on Anthropic's servers after the first request. Subsequent requests that use the same context only pay a small 'cache hit' fee rather than the full input token price. This is the single most impactful feature for reducing costs in applications that repeatedly reference large documents or datasets.
- ✓
Using Retrieval-Augmented Generation (RAG) to only send relevant snippets.
Why this is correct
Instead of sending the full 150,000 tokens every time, RAG identifies the most relevant parts of the document and only includes those in the prompt. This drastically reduces the input token count for each request, leading to significant cost savings and faster response times, as the model has less information to process before generating an answer.
- ✗
Converting the text to a more compact binary format before sending to the API.
Why it's wrong here
The Anthropic API accepts text or specifically formatted image data, not arbitrary binary formats. Attempting to send binary data or non-standard encodings will not reduce token counts; in fact, it often increases them because the tokenizer may treat the binary characters as individual, rare tokens, leading to higher costs and potentially incomprehensible model outputs.
- ✗
Hard-coding the model to Claude 3 Opus for better token compression.
Why it's wrong here
Claude 3 Opus is the most expensive model and does not 'compress' tokens more efficiently than others. While it has a larger context window and better reasoning, it will still charge for every one of the 150,000 tokens. Using Opus for such large prompts without caching or RAG would result in the highest possible cost per request.
- ✗
Increasing the 'temperature' setting to reduce the length of the generated output.
Why it's wrong here
Temperature controls the randomness and creativity of the output, not its length. While a very low temperature might lead to more repetitive and potentially shorter responses in some cases, it is not a reliable or recommended method for cost management. To control output length and cost, developers should use the 'max_tokens' parameter instead of adjusting temperature.
About these practice questions
This CCDV-F question is part of Courseiva's 257-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCDV-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCDV-F exam.