CCDV-F · domain
Model Selection and Cost Management
This domain covers how Claude API costs are driven by token usage and how to control them through model choice, prompt size, and caching. Questions present realistic scenarios—support pipelines, repeated report analysis, knowledge bases—and ask you to pick cost-reduction strategies, interpret prompt caching economics, or identify factors that change total token consumption.
Focused practice
Practice Model Selection and Cost Management questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Model Selection and Cost Management
Be able to compute or reason about token-based cost: identify which factors (input size, output size, request count, model tier, cache hits) drive spend, and choose concrete Claude API features—model selection, context trimming, prompt caching—that lower it without losing required quality.
Selecting an appropriate Claude model tier (e.g., Haiku vs Sonnet vs Opus) for a workload's cost/quality tradeoff
Using prompt caching with cache_control to avoid re-charging full input tokens on repeated large context
Reducing token expenditure by trimming prompt context and limiting output length
Estimating monthly cost from input tokens, output tokens, and request volume
Watch out for
Common Model Selection and Cost Management exam traps
- ▸Assuming prompt caching eliminates all cost—cache writes and reads still incur charges, and only repeated prefixes benefit.
- ▸Choosing the largest model by default instead of matching model tier to task complexity and budget.
- ▸Ignoring output tokens and max_tokens when estimating cost, focusing only on input prompt size.
Question index
All Model Selection and Cost Management questions (40)
Click any question to see the full explanation, or start a practice session above.
Refer to the exhibit. This JSON object represents a single line in a JSONL file intended for the Anthropic Batch API. What is the primary advantage of submitting requests in this format rather than using the standard synchronous Messages API?
Medium2An AI engineering team is auditing an expensive Claude API integration processing millions of customer support queries daily. Which TWO strategies should they implement to effectively reduce token expenditure without sacrificing core model intelligence? (Choose two)
Hard3A developer is building a high-traffic legal research tool. Which TWO conditions must be met for a content block to successfully return a 'cache hit' and reduce the cost of an API call?
Hard4A legal tech company is using Claude 3.5 Sonnet to generate 5,000-word contract drafts. They notice that the output costs are significantly higher than the input costs. What is the most effective way to reduce these specific output costs without switching to a lower-quality model?
Hard5Which THREE criteria are most important when deciding whether to upgrade from Claude 3 Haiku to Claude 3.5 Sonnet for a production application?
Medium6A developer is using the Claude API to generate creative writing pieces. The application often sends the same 2,000-token instruction set as a prefix to every request. Which feature can reduce the cost of processing this repeated prefix?
Easy7A developer is optimizing a Claude-powered chatbot that handles 50,000 daily conversations. Each conversation starts with a 1,200-token system prompt that includes detailed instructions and examples. The system prompt is identical across all conversations. The developer wants to reduce input token costs. Which technique should be applied?
Hard8Which THREE factors directly determine the total cost of a single API call to the Claude Messages endpoint?
Medium9A developer is using the Claude API to generate marketing copy for 50 different product lines. Each request includes a unique set of product details and a unique creative brief. The developer wants to minimize cost while ensuring high-quality output. Which model should they choose?
Easy10A developer is building a long-context application that processes 150,000 tokens per request. To manage costs and maintain performance, which TWO techniques should be prioritized?
Hard11A developer is concerned about the 'token overhead' when using system prompts. How does Anthropic charge for tokens included in the system prompt compared to tokens in the user message?
Easy12A developer runs a nightly batch job that sends 8,000 independent product-description summarization requests to the Anthropic API. Each request shares an identical 12,000-token instruction block and a unique 300-token product description. The team wants to reduce input-token spend without changing output quality or model. Which approach best meets this goal?
Hard13A company needs to summarize 500 academic papers, each approximately 20,000 tokens long. The summaries are needed for a weekly report due in three days. Which approach provides the best balance of cost-efficiency and model capability for this specific task?
Medium14A developer notices that a Claude integration's monthly bill is far higher than expected. Logs show that each request includes a 15,000-token system prompt containing detailed style guidelines, plus a 200-token user question, and the model generates a 50-token answer. The system prompt is identical across all requests. Which change would MOST reduce cost without altering output quality?
Hard15A developer needs to monitor the costs of different departments using a single Anthropic API key. Which API feature should be utilized to categorize and track usage without creating multiple accounts or keys?
Medium16A financial services firm is developing a customer support bot that requires low latency and high throughput for processing simple account inquiries. The firm expects over 500,000 requests per day and prioritizes cost-efficiency above the highest possible reasoning capabilities. Which model should the developer select to meet these specific business requirements?
Medium17A developer is building a customer-facing chatbot using the Claude API. The bot must respond within 1.5 seconds on average. The team initially selected claude-3-opus-20240229 for its high quality, but latency is consistently above 3 seconds. They need to reduce latency while maintaining acceptable response quality. Which action should the developer take?
Medium18A developer is reaching the Rate Limits for their account tier while using Claude 3.5 Sonnet. Which TWO actions would help manage these limits while also potentially reducing costs?
Hard19A developer needs to select a model for a code generation tool that must handle complex multi-file refactoring and advanced algorithmic logic. Which Claude 3.5 model currently provides the highest level of intelligence and coding capability for this task?
Easy20An application uses Claude 3.5 Sonnet and implements Prompt Caching for a 10,000-token system prompt. The application processes 1,000 requests per hour. If the cache is refreshed every request and never expires, how does the billing for the input tokens change after the very first request?
Hard21A developer is using the Claude Messages API to process a 50,000-token document for a summarization task. The summarization prompt and instructions add another 1,000 tokens. The developer wants to minimize output token costs while ensuring a comprehensive summary. Which strategy is most effective?
Hard22A team is estimating costs for a new feature that will send 1 million requests per month. Each request has a 2,000-token input and generates a 500-token output. The team wants to reduce the output token cost, which dominates the bill. Which strategy is MOST effective for reducing output token costs?
Medium23A developer is building a real-time translation service that must respond within 1 second for short phrases. The service will handle thousands of requests per hour. Quality is important, but latency and cost are critical. Which Claude model is the most appropriate?
Medium24When calculating the estimated cost of a project using Claude, which metric is used by Anthropic to measure the volume of data processed and generated?
Easy25Which of the following scenarios describes the most effective use of Prompt Caching for cost management?
Easy26An enterprise is migrating a document processing pipeline that handles 50,000 PDFs daily. Each PDF is converted to text (approx. 2,000 tokens) and requires a summary. The project has a strict budget. Which approach provides the most significant cost reduction while utilizing Claude 3.5 Sonnet?
Medium27A developer wants to implement a 'Summary' feature for a long conversation history. As the conversation grows, the cost of sending the entire history with every new message increases. What is the most cost-effective architectural pattern to handle this?
Medium28A developer is comparing the cost of using Claude 3 Opus versus Claude 3.5 Sonnet for a task that requires complex reasoning. The task involves processing 1,000 requests, each with 500 input tokens and 200 output tokens. Which statement accurately reflects the cost consideration?
Medium29A company needs to process 10 million short customer feedback snippets to identify 'bug reports' vs 'feature requests'. Speed and budget are the primary constraints, while the classification logic is straightforward. Which model provides the best throughput-to-cost ratio?
Medium30A developer is building a real-time chat application where users expect responses in under two seconds. The prompts are short (under 200 tokens) and the responses are typically one or two sentences. Which Claude model should the developer choose to optimize for latency and cost?
Medium31You are building a summarization feature that processes 20,000 support tickets nightly. Each ticket is under 4,000 tokens and the output summary is around 300 tokens. The job must complete within a 6-hour window and you want to minimize cost. Which Claude model should you choose?
Medium32A developer is building an application that uses the Claude Messages API to generate product descriptions. The application sends a system prompt of 1,500 tokens, a user prompt of 200 tokens, and receives a response of 300 tokens. The developer wants to reduce costs. Which two strategies would directly reduce the cost per API call? (Choose two.)
Medium33A developer is using the Claude Messages API to generate a 500-token response. The input consists of a 200-token user message and a 100-token system prompt. Which factor directly determines the output token cost of this API call?
Easy34Refer to the exhibit. A developer is implementing the provided JSON structure to optimize an application that repeatedly analyzes the same large report. What is the primary financial implication of using the 'cache_control' block in this specific API request?
Medium35A developer is optimizing a Claude-powered document processing pipeline that sends large, mostly identical legal templates followed by short variable fields. They want to reduce input token costs while preserving output fidelity. Which TWO strategies are appropriate? (Choose two.)
Hard36A startup is prototyping a chatbot that handles simple FAQ responses for a small user base. The team wants the lowest possible cost per request and does not need advanced reasoning. Which Claude model selection strategy is MOST appropriate?
Easy37In the context of Anthropic's pricing model, why is it generally recommended to provide only the necessary context rather than the entire available dataset in a single prompt?
Easy38A developer needs to estimate the monthly budget for a new internal knowledge base application powered by Claude. Which TWO factors directly influence the total token consumption and subsequent cost of the API requests?
Easy39You are designing an AI agent that performs multi-step reasoning. The first step involves basic data extraction, while the second step requires complex logical deduction based on the extracted data. How should you select models to optimize for both performance and cost?
Medium40A developer is building a Claude-powered coding assistant that must answer questions about a 60,000-token proprietary codebase on every user request. The codebase is static and updated only weekly. The developer wants to minimize per-request input token costs while keeping latency low. Which approach is MOST cost-effective?
MediumOther domains
All CCDV-F exam domains
Frequently asked questions
- What does the Model Selection and Cost Management domain cover on the CCDV-F exam?
- Be able to compute or reason about token-based cost: identify which factors (input size, output size, request count, model tier, cache hits) drive spend, and choose concrete Claude API features—model selection, context trimming, prompt caching—that lower it without losing required quality.
- How many questions are in this domain?
- This page lists all 40 Model Selection and Cost Management questions in the CCDV-F question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Model Selection and Cost Management questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.