Be able to compute or reason about token-based cost: identify which factors (input size, output size, request count, model tier, cache hits) drive spend, and choose concrete Claude API features—model selection, context trimming, prompt caching—that lower it without losing required quality.
Start practicing
Model Selection and Cost Management — choose a session length
Free · No account required
Domain overview
This domain covers how Claude API costs are driven by token usage and how to control them through model choice, prompt size, and caching. Questions present realistic scenarios—support pipelines, repeated report analysis, knowledge bases—and ask you to pick cost-reduction strategies, interpret prompt caching economics, or identify factors that change total token consumption.
Exam objectives
Selecting an appropriate Claude model tier (e.g., Haiku vs Sonnet vs Opus) for a workload's cost/quality tradeoff
Using prompt caching with cache_control to avoid re-charging full input tokens on repeated large context
Reducing token expenditure by trimming prompt context and limiting output length
Estimating monthly cost from input tokens, output tokens, and request volume
Assuming prompt caching eliminates all cost—cache writes and reads still incur charges, and only repeated prefixes benefit.
Choosing the largest model by default instead of matching model tier to task complexity and budget.
Ignoring output tokens and max_tokens when estimating cost, focusing only on input prompt size.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A developer needs to estimate the monthly budget for a new internal knowledge base application powered by Claude. Which TWO factors directly influence the total token consumption and subsequent cost of the API requests?
2An enterprise is migrating a document processing pipeline that handles 50,000 PDFs daily. Each PDF is converted to text (approx. 2,000 tokens) and requires a summary. The project has a strict budget. Which approach provides the most significant cost reduction while utilizing Claude 3.5 Sonnet?
3You are designing an AI agent that performs multi-step reasoning. The first step involves basic data extraction, while the second step requires complex logical deduction based on the extracted data. How should you select models to optimize for both performance and cost?
4A developer is building a long-context application that processes 150,000 tokens per request. To manage costs and maintain performance, which TWO techniques should be prioritized?
5Which of the following scenarios describes the most effective use of Prompt Caching for cost management?
6A company needs to process 10 million short customer feedback snippets to identify 'bug reports' vs 'feature requests'. Speed and budget are the primary constraints, while the classification logic is straightforward. Which model provides the best throughput-to-cost ratio?
7A developer is concerned about the 'token overhead' when using system prompts. How does Anthropic charge for tokens included in the system prompt compared to tokens in the user message?
8A legal tech company is using Claude 3.5 Sonnet to generate 5,000-word contract drafts. They notice that the output costs are significantly higher than the input costs. What is the most effective way to reduce these specific output costs without switching to a lower-quality model?
9In the context of Anthropic's pricing model, why is it generally recommended to provide only the necessary context rather than the entire available dataset in a single prompt?
10A financial services firm is developing a customer support bot that requires low latency and high throughput for processing simple account inquiries. The firm expects over 500,000 requests per day and prioritizes cost-efficiency above the highest possible reasoning capabilities. Which model should the developer select to meet these specific business requirements?
11Refer to the exhibit. A developer is implementing the provided JSON structure to optimize an application that repeatedly analyzes the same large report. What is the primary financial implication of using the 'cache_control' block in this specific API request?
12A developer needs to select a model for a code generation tool that must handle complex multi-file refactoring and advanced algorithmic logic. Which Claude 3.5 model currently provides the highest level of intelligence and coding capability for this task?
13Which THREE factors directly determine the total cost of a single API call to the Claude Messages endpoint?
14A company needs to summarize 500 academic papers, each approximately 20,000 tokens long. The summaries are needed for a weekly report due in three days. Which approach provides the best balance of cost-efficiency and model capability for this specific task?
15A developer is building a high-traffic legal research tool. Which TWO conditions must be met for a content block to successfully return a 'cache hit' and reduce the cost of an API call?
16Refer to the exhibit. This JSON object represents a single line in a JSONL file intended for the Anthropic Batch API. What is the primary advantage of submitting requests in this format rather than using the standard synchronous Messages API?
17A developer needs to monitor the costs of different departments using a single Anthropic API key. Which API feature should be utilized to categorize and track usage without creating multiple accounts or keys?
18Which THREE criteria are most important when deciding whether to upgrade from Claude 3 Haiku to Claude 3.5 Sonnet for a production application?
19An application uses Claude 3.5 Sonnet and implements Prompt Caching for a 10,000-token system prompt. The application processes 1,000 requests per hour. If the cache is refreshed every request and never expires, how does the billing for the input tokens change after the very first request?
20When calculating the estimated cost of a project using Claude, which metric is used by Anthropic to measure the volume of data processed and generated?
21A developer is reaching the Rate Limits for their account tier while using Claude 3.5 Sonnet. Which TWO actions would help manage these limits while also potentially reducing costs?
22A developer wants to implement a 'Summary' feature for a long conversation history. As the conversation grows, the cost of sending the entire history with every new message increases. What is the most cost-effective architectural pattern to handle this?
23An AI engineering team is auditing an expensive Claude API integration processing millions of customer support queries daily. Which TWO strategies should they implement to effectively reduce token expenditure without sacrificing core model intelligence? (Choose two)
24A developer runs a nightly batch job that sends 8,000 independent product-description summarization requests to the Anthropic API. Each request shares an identical 12,000-token instruction block and a unique 300-token product description. The team wants to reduce input-token spend without changing output quality or model. Which approach best meets this goal?
25You are building a summarization feature that processes 20,000 support tickets nightly. Each ticket is under 4,000 tokens and the output summary is around 300 tokens. The job must complete within a 6-hour window and you want to minimize cost. Which Claude model should you choose?
26A developer is optimizing a Claude-powered chatbot that handles 50,000 daily conversations. Each conversation starts with a 1,200-token system prompt that includes detailed instructions and examples. The system prompt is identical across all conversations. The developer wants to reduce input token costs. Which technique should be applied?
27A developer is using the Claude Messages API to generate a 500-token response. The input consists of a 200-token user message and a 100-token system prompt. Which factor directly determines the output token cost of this API call?
28A developer is building a Claude-powered coding assistant that must answer questions about a 60,000-token proprietary codebase on every user request. The codebase is static and updated only weekly. The developer wants to minimize per-request input token costs while keeping latency low. Which approach is MOST cost-effective?
29A developer is building a real-time translation service that must respond within 1 second for short phrases. The service will handle thousands of requests per hour. Quality is important, but latency and cost are critical. Which Claude model is the most appropriate?
30A developer is building a customer-facing chatbot using the Claude API. The bot must respond within 1.5 seconds on average. The team initially selected claude-3-opus-20240229 for its high quality, but latency is consistently above 3 seconds. They need to reduce latency while maintaining acceptable response quality. Which action should the developer take?
31A startup is prototyping a chatbot that handles simple FAQ responses for a small user base. The team wants the lowest possible cost per request and does not need advanced reasoning. Which Claude model selection strategy is MOST appropriate?
32A developer notices that a Claude integration's monthly bill is far higher than expected. Logs show that each request includes a 15,000-token system prompt containing detailed style guidelines, plus a 200-token user question, and the model generates a 50-token answer. The system prompt is identical across all requests. Which change would MOST reduce cost without altering output quality?
33A developer is building a real-time chat application where users expect responses in under two seconds. The prompts are short (under 200 tokens) and the responses are typically one or two sentences. Which Claude model should the developer choose to optimize for latency and cost?
34A team is estimating costs for a new feature that will send 1 million requests per month. Each request has a 2,000-token input and generates a 500-token output. The team wants to reduce the output token cost, which dominates the bill. Which strategy is MOST effective for reducing output token costs?
35A developer is using the Claude Messages API to process a 50,000-token document for a summarization task. The summarization prompt and instructions add another 1,000 tokens. The developer wants to minimize output token costs while ensuring a comprehensive summary. Which strategy is most effective?
36A developer is using the Claude API to generate marketing copy for 50 different product lines. Each request includes a unique set of product details and a unique creative brief. The developer wants to minimize cost while ensuring high-quality output. Which model should they choose?
37A developer is optimizing a Claude-powered document processing pipeline that sends large, mostly identical legal templates followed by short variable fields. They want to reduce input token costs while preserving output fidelity. Which TWO strategies are appropriate? (Choose two.)
38A developer is building an application that uses the Claude Messages API to generate product descriptions. The application sends a system prompt of 1,500 tokens, a user prompt of 200 tokens, and receives a response of 300 tokens. The developer wants to reduce costs. Which two strategies would directly reduce the cost per API call? (Choose two.)
39A developer is using the Claude API to generate creative writing pieces. The application often sends the same 2,000-token instruction set as a prefix to every request. Which feature can reduce the cost of processing this repeated prefix?
40A developer is comparing the cost of using Claude 3 Opus versus Claude 3.5 Sonnet for a task that requires complex reasoning. The task involves processing 1,000 requests, each with 500 input tokens and 200 output tokens. Which statement accurately reflects the cost consideration?
Be able to compute or reason about token-based cost: identify which factors (input size, output size, request count, model tier, cache hits) drive spend, and choose concrete Claude API features—model selection, context trimming, prompt caching—that lower it without losing required quality.
The Courseiva CCDV-F question bank contains 40 questions in the Model Selection and Cost Management domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Model Selection and Cost Management domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included