Courseiva

CCDV-F · topic practice

Model Selection and Cost Management practice questions

This domain covers how Claude API costs are driven by token usage and how to control them through model choice, prompt size, and caching. Questions present realistic scenarios—support pipelines, repeated report analysis, knowledge bases—and ask you to pick cost-reduction strategies, interpret prompt caching economics, or identify factors that change total token consumption.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Model Selection and Cost Management

What the exam tests

What to know about Model Selection and Cost Management

Be able to compute or reason about token-based cost: identify which factors (input size, output size, request count, model tier, cache hits) drive spend, and choose concrete Claude API features—model selection, context trimming, prompt caching—that lower it without losing required quality.

Selecting an appropriate Claude model tier (e.g., Haiku vs Sonnet vs Opus) for a workload's cost/quality tradeoff

Using prompt caching with cache_control to avoid re-charging full input tokens on repeated large context

Reducing token expenditure by trimming prompt context and limiting output length

Estimating monthly cost from input tokens, output tokens, and request volume

Watch out for

Common Model Selection and Cost Management exam traps

  • ▸Assuming prompt caching eliminates all cost—cache writes and reads still incur charges, and only repeated prefixes benefit.
  • ▸Choosing the largest model by default instead of matching model tier to task complexity and budget.
  • ▸Ignoring output tokens and max_tokens when estimating cost, focusing only on input prompt size.

Practice set

Model Selection and Cost Management questions

20 questions · select your answer, then reveal the explanation

A financial services firm is building a real-time sentiment analysis tool to process thousands of customer tweets per minute. The system requires low latency and the lowest possible operational cost, but the sentiment classification must still handle nuances like sarcasm accurately. Which model should the developer select for this production environment?

A developer is evaluating the Anthropic Batch API for a large-scale data labeling project. Which THREE statements accurately describe the characteristics or limitations of the Batch API?

When comparing Claude 3.5 Sonnet to Claude 3 Opus for an enterprise application, which TWO statements are true regarding model selection and cost?

An enterprise developer needs to process a massive backlog of 100,000 documents for sentiment analysis. The project is not time-sensitive but has a very strict budget limit. Which TWO strategies will result in the most significant cost savings for this specific workload?

A developer is using Claude 3.5 Sonnet to build a document analysis tool. They want to use Prompt Caching to save costs on a 5,000-token reference document that is included in every request. What is the minimum number of tokens required for a block to be eligible for caching on this specific model?

An enterprise application requires real-time document classification where latency must remain under 200 milliseconds, but the classification accuracy demands advanced reasoning capabilities. Which model configuration best aligns with these cost and performance constraints?

A developer is analyzing the cost of a Claude API integration that processes customer emails. The integration uses the Messages API with a system prompt, user messages, and assistant responses. Which TWO factors directly contribute to the total token cost of each API call? (Choose two.)

A developer is optimizing a Claude API integration that processes customer support tickets. Each ticket is analyzed by sending the full conversation history (up to 10,000 tokens) along with a fixed 2,000-token system prompt that contains detailed instructions and examples. The system prompt is identical for every request. The team wants to reduce costs without sacrificing response quality. Which two strategies should they implement? (Choose two.)

A developer needs to estimate the monthly budget for a new internal knowledge base application powered by Claude. Which TWO factors directly influence the total token consumption and subsequent cost of the API requests?

An enterprise is migrating a document processing pipeline that handles 50,000 PDFs daily. Each PDF is converted to text (approx. 2,000 tokens) and requires a summary. The project has a strict budget. Which approach provides the most significant cost reduction while utilizing Claude 3.5 Sonnet?

You are designing an AI agent that performs multi-step reasoning. The first step involves basic data extraction, while the second step requires complex logical deduction based on the extracted data. How should you select models to optimize for both performance and cost?

A developer is building a long-context application that processes 150,000 tokens per request. To manage costs and maintain performance, which TWO techniques should be prioritized?

Which of the following scenarios describes the most effective use of Prompt Caching for cost management?

A company needs to process 10 million short customer feedback snippets to identify 'bug reports' vs 'feature requests'. Speed and budget are the primary constraints, while the classification logic is straightforward. Which model provides the best throughput-to-cost ratio?

A developer is concerned about the 'token overhead' when using system prompts. How does Anthropic charge for tokens included in the system prompt compared to tokens in the user message?

A legal tech company is using Claude 3.5 Sonnet to generate 5,000-word contract drafts. They notice that the output costs are significantly higher than the input costs. What is the most effective way to reduce these specific output costs without switching to a lower-quality model?

In the context of Anthropic's pricing model, why is it generally recommended to provide only the necessary context rather than the entire available dataset in a single prompt?

A financial services firm is developing a customer support bot that requires low latency and high throughput for processing simple account inquiries. The firm expects over 500,000 requests per day and prioritizes cost-efficiency above the highest possible reasoning capabilities. Which model should the developer select to meet these specific business requirements?

Refer to the exhibit. A developer is implementing the provided JSON structure to optimize an application that repeatedly analyzes the same large report. What is the primary financial implication of using the 'cache_control' block in this specific API request?

Exhibit

{
  "model": "claude-3-5-sonnet-20240620",
  "max_tokens": 1024,
  "messages": [
    {
      "role": "user",
      "content": [
        {
          "type": "text",
          "text": "Analyze this report:",
          "cache_control": {"type": "ephemeral"}
        },
        {
          "type": "text",
          "text": "[Large Report Content Here...]"
        }
      ]
    }
  ]
}

A developer needs to select a model for a code generation tool that must handle complex multi-file refactoring and advanced algorithmic logic. Which Claude 3.5 model currently provides the highest level of intelligence and coding capability for this task?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Model Selection and Cost Management sessions

Start a Model Selection and Cost Management only practice session

Every question in these sessions is drawn from the Model Selection and Cost Management domain — nothing else.

Related practice questions

Related CCDV-F topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the CCDV-F exam test about Model Selection and Cost Management?
Be able to compute or reason about token-based cost: identify which factors (input size, output size, request count, model tier, cache hits) drive spend, and choose concrete Claude API features—model selection, context trimming, prompt caching—that lower it without losing required quality.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Model Selection and Cost Management questions in a focused session?
Yes — the session launcher on this page draws every question from the Model Selection and Cost Management domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other CCDV-F topics?
Use the topic links above to move to related areas, or go back to the CCDV-F question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the CCDV-F exam covers. They are not copied from any real exam or dump site.