Courseiva

CCDV-F · domain

Model Selection and Cost Management

This domain covers how Claude API costs are driven by token usage and how to control them through model choice, prompt size, and caching. Questions present realistic scenarios—support pipelines, repeated report analysis, knowledge bases—and ask you to pick cost-reduction strategies, interpret prompt caching economics, or identify factors that change total token consumption.

40 questions10 easy19 medium11 hard

Focused practice

Practice Model Selection and Cost Management questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Model Selection and Cost Management

Be able to compute or reason about token-based cost: identify which factors (input size, output size, request count, model tier, cache hits) drive spend, and choose concrete Claude API features—model selection, context trimming, prompt caching—that lower it without losing required quality.

Selecting an appropriate Claude model tier (e.g., Haiku vs Sonnet vs Opus) for a workload's cost/quality tradeoff

Using prompt caching with cache_control to avoid re-charging full input tokens on repeated large context

Reducing token expenditure by trimming prompt context and limiting output length

Estimating monthly cost from input tokens, output tokens, and request volume

Watch out for

Common Model Selection and Cost Management exam traps

  • ▸Assuming prompt caching eliminates all cost—cache writes and reads still incur charges, and only repeated prefixes benefit.
  • ▸Choosing the largest model by default instead of matching model tier to task complexity and budget.
  • ▸Ignoring output tokens and max_tokens when estimating cost, focusing only on input prompt size.

Question index

All Model Selection and Cost Management questions (40)

Click any question to see the full explanation, or start a practice session above.

1

Refer to the exhibit. This JSON object represents a single line in a JSONL file intended for the Anthropic Batch API. What is the primary advantage of submitting requests in this format rather than using the standard synchronous Messages API?

Medium
2

An AI engineering team is auditing an expensive Claude API integration processing millions of customer support queries daily. Which TWO strategies should they implement to effectively reduce token expenditure without sacrificing core model intelligence? (Choose two)

Hard
3

A developer is building a high-traffic legal research tool. Which TWO conditions must be met for a content block to successfully return a 'cache hit' and reduce the cost of an API call?

Hard
4

A legal tech company is using Claude 3.5 Sonnet to generate 5,000-word contract drafts. They notice that the output costs are significantly higher than the input costs. What is the most effective way to reduce these specific output costs without switching to a lower-quality model?

Hard
5

Which THREE criteria are most important when deciding whether to upgrade from Claude 3 Haiku to Claude 3.5 Sonnet for a production application?

Medium
6

A developer is using the Claude API to generate creative writing pieces. The application often sends the same 2,000-token instruction set as a prefix to every request. Which feature can reduce the cost of processing this repeated prefix?

Easy
7

A developer is optimizing a Claude-powered chatbot that handles 50,000 daily conversations. Each conversation starts with a 1,200-token system prompt that includes detailed instructions and examples. The system prompt is identical across all conversations. The developer wants to reduce input token costs. Which technique should be applied?

Hard
8

Which THREE factors directly determine the total cost of a single API call to the Claude Messages endpoint?

Medium
9

A developer is using the Claude API to generate marketing copy for 50 different product lines. Each request includes a unique set of product details and a unique creative brief. The developer wants to minimize cost while ensuring high-quality output. Which model should they choose?

Easy
10

A developer is building a long-context application that processes 150,000 tokens per request. To manage costs and maintain performance, which TWO techniques should be prioritized?

Hard
11

A developer is concerned about the 'token overhead' when using system prompts. How does Anthropic charge for tokens included in the system prompt compared to tokens in the user message?

Easy
12

A developer runs a nightly batch job that sends 8,000 independent product-description summarization requests to the Anthropic API. Each request shares an identical 12,000-token instruction block and a unique 300-token product description. The team wants to reduce input-token spend without changing output quality or model. Which approach best meets this goal?

Hard
13

A company needs to summarize 500 academic papers, each approximately 20,000 tokens long. The summaries are needed for a weekly report due in three days. Which approach provides the best balance of cost-efficiency and model capability for this specific task?

Medium
14

A developer notices that a Claude integration's monthly bill is far higher than expected. Logs show that each request includes a 15,000-token system prompt containing detailed style guidelines, plus a 200-token user question, and the model generates a 50-token answer. The system prompt is identical across all requests. Which change would MOST reduce cost without altering output quality?

Hard
15

A developer needs to monitor the costs of different departments using a single Anthropic API key. Which API feature should be utilized to categorize and track usage without creating multiple accounts or keys?

Medium
16

A financial services firm is developing a customer support bot that requires low latency and high throughput for processing simple account inquiries. The firm expects over 500,000 requests per day and prioritizes cost-efficiency above the highest possible reasoning capabilities. Which model should the developer select to meet these specific business requirements?

Medium
17

A developer is building a customer-facing chatbot using the Claude API. The bot must respond within 1.5 seconds on average. The team initially selected claude-3-opus-20240229 for its high quality, but latency is consistently above 3 seconds. They need to reduce latency while maintaining acceptable response quality. Which action should the developer take?

Medium
18

A developer is reaching the Rate Limits for their account tier while using Claude 3.5 Sonnet. Which TWO actions would help manage these limits while also potentially reducing costs?

Hard
19

A developer needs to select a model for a code generation tool that must handle complex multi-file refactoring and advanced algorithmic logic. Which Claude 3.5 model currently provides the highest level of intelligence and coding capability for this task?

Easy
20

An application uses Claude 3.5 Sonnet and implements Prompt Caching for a 10,000-token system prompt. The application processes 1,000 requests per hour. If the cache is refreshed every request and never expires, how does the billing for the input tokens change after the very first request?

Hard
21

A developer is using the Claude Messages API to process a 50,000-token document for a summarization task. The summarization prompt and instructions add another 1,000 tokens. The developer wants to minimize output token costs while ensuring a comprehensive summary. Which strategy is most effective?

Hard
22

A team is estimating costs for a new feature that will send 1 million requests per month. Each request has a 2,000-token input and generates a 500-token output. The team wants to reduce the output token cost, which dominates the bill. Which strategy is MOST effective for reducing output token costs?

Medium
23

A developer is building a real-time translation service that must respond within 1 second for short phrases. The service will handle thousands of requests per hour. Quality is important, but latency and cost are critical. Which Claude model is the most appropriate?

Medium
24

When calculating the estimated cost of a project using Claude, which metric is used by Anthropic to measure the volume of data processed and generated?

Easy
25

Which of the following scenarios describes the most effective use of Prompt Caching for cost management?

Easy
26

An enterprise is migrating a document processing pipeline that handles 50,000 PDFs daily. Each PDF is converted to text (approx. 2,000 tokens) and requires a summary. The project has a strict budget. Which approach provides the most significant cost reduction while utilizing Claude 3.5 Sonnet?

Medium
27

A developer wants to implement a 'Summary' feature for a long conversation history. As the conversation grows, the cost of sending the entire history with every new message increases. What is the most cost-effective architectural pattern to handle this?

Medium
28

A developer is comparing the cost of using Claude 3 Opus versus Claude 3.5 Sonnet for a task that requires complex reasoning. The task involves processing 1,000 requests, each with 500 input tokens and 200 output tokens. Which statement accurately reflects the cost consideration?

Medium
29

A company needs to process 10 million short customer feedback snippets to identify 'bug reports' vs 'feature requests'. Speed and budget are the primary constraints, while the classification logic is straightforward. Which model provides the best throughput-to-cost ratio?

Medium
30

A developer is building a real-time chat application where users expect responses in under two seconds. The prompts are short (under 200 tokens) and the responses are typically one or two sentences. Which Claude model should the developer choose to optimize for latency and cost?

Medium
31

You are building a summarization feature that processes 20,000 support tickets nightly. Each ticket is under 4,000 tokens and the output summary is around 300 tokens. The job must complete within a 6-hour window and you want to minimize cost. Which Claude model should you choose?

Medium
32

A developer is building an application that uses the Claude Messages API to generate product descriptions. The application sends a system prompt of 1,500 tokens, a user prompt of 200 tokens, and receives a response of 300 tokens. The developer wants to reduce costs. Which two strategies would directly reduce the cost per API call? (Choose two.)

Medium
33

A developer is using the Claude Messages API to generate a 500-token response. The input consists of a 200-token user message and a 100-token system prompt. Which factor directly determines the output token cost of this API call?

Easy
34

Refer to the exhibit. A developer is implementing the provided JSON structure to optimize an application that repeatedly analyzes the same large report. What is the primary financial implication of using the 'cache_control' block in this specific API request?

Medium
35

A developer is optimizing a Claude-powered document processing pipeline that sends large, mostly identical legal templates followed by short variable fields. They want to reduce input token costs while preserving output fidelity. Which TWO strategies are appropriate? (Choose two.)

Hard
36

A startup is prototyping a chatbot that handles simple FAQ responses for a small user base. The team wants the lowest possible cost per request and does not need advanced reasoning. Which Claude model selection strategy is MOST appropriate?

Easy
37

In the context of Anthropic's pricing model, why is it generally recommended to provide only the necessary context rather than the entire available dataset in a single prompt?

Easy
38

A developer needs to estimate the monthly budget for a new internal knowledge base application powered by Claude. Which TWO factors directly influence the total token consumption and subsequent cost of the API requests?

Easy
39

You are designing an AI agent that performs multi-step reasoning. The first step involves basic data extraction, while the second step requires complex logical deduction based on the extracted data. How should you select models to optimize for both performance and cost?

Medium
40

A developer is building a Claude-powered coding assistant that must answer questions about a 60,000-token proprietary codebase on every user request. The codebase is static and updated only weekly. The developer wants to minimize per-request input token costs while keeping latency low. Which approach is MOST cost-effective?

Medium

Frequently asked questions

What does the Model Selection and Cost Management domain cover on the CCDV-F exam?
Be able to compute or reason about token-based cost: identify which factors (input size, output size, request count, model tier, cache hits) drive spend, and choose concrete Claude API features—model selection, context trimming, prompt caching—that lower it without losing required quality.
How many questions are in this domain?
This page lists all 40 Model Selection and Cost Management questions in the CCDV-F question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Model Selection and Cost Management questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
anthropic-claude-developer ANTHROPIC-CLAUDE-DEVELOPER model selection and cost management Practice Questions