Courseiva

CCNA Model Selection And Cost Management Questions

40 questions · Model Selection And Cost Management topic · All types, answers revealed

1
MCQmedium

Refer to the exhibit. This JSON object represents a single line in a JSONL file intended for the Anthropic Batch API. What is the primary advantage of submitting requests in this format rather than using the standard synchronous Messages API?

A.It allows the model to access a larger context window of 400k tokens.
B.It provides a 50% discount on the total token cost.
C.It enables real-time streaming of responses as they are generated.
D.It bypasses all Rate Limits (RPM and TPM) for the account.
AnswerB

The defining feature of the Anthropic Batch API is its pricing. By allowing Anthropic up to 24 hours to process the requests, users are charged only half of the standard rate for both input and output tokens. This makes it the most cost-effective choice for bulk data processing tasks.

Why this answer

The Batch API is a powerful tool for cost optimization. By grouping requests into a JSONL file and processing them asynchronously, Anthropic can optimize its compute usage, passing the savings to the user. This results in a 50% price reduction.

Developers use this for tasks that are not time-sensitive, effectively doubling their throughput per dollar spent on API credits.

Exam trap

Candidates often confuse asynchronous Batch API processing with real-time latency improvements, mistakenly thinking the format accelerates individual response times instead of providing a 50% cost discount.

2
Multi-Selecthard

An AI engineering team is auditing an expensive Claude API integration processing millions of customer support queries daily. Which TWO strategies should they implement to effectively reduce token expenditure without sacrificing core model intelligence? (Choose two)

Select 2 answers
A.Implement Anthropic prompt caching for large, recurring system prompts and reference knowledge bases.
B.Switch every production workload immediately from Claude 3.5 Sonnet to Claude 3 Opus to maximize parallel processing.
C.Refactor prompts to remove verbose formatting instructions, conversational filler, and redundant examples.
D.Disable response streaming to batch full payloads and reduce HTTP connection overhead.
E.Hardcode maximum output tokens to five thousand across all API calls regardless of task complexity.
AnswersA, C

Prompt caching allows developers to reuse previously processed prompt blocks, offering significant cost discounts and reduced time-to-first-token for recurring system instructions. This directly targets high-volume input token costs in applications with static reference materials or lengthy guidelines.

Why this answer

Optimizing context management and leveraging prompt caching are cornerstone strategies for cost reduction in high-volume Anthropic integrations. Caching static system prompts drastically lowers input token charges across repeated requests, while concise prompts eliminate redundant verbiage. Together, these practices protect infrastructure budgets while sustaining the semantic fidelity required by enterprise-grade support workflows.

Exam trap

Candidates often try to reduce costs by shortening model output lengths or lowering generation parameters, overlooking the greater impact of optimizing input context via caching and concise prompting.

3
Multi-Selecthard

A developer is building a high-traffic legal research tool. Which TWO conditions must be met for a content block to successfully return a 'cache hit' and reduce the cost of an API call?

Select 2 answers
A.The content must be identical, including all whitespace and formatting.
B.The cached block must be at the very beginning (prefix) of the prompt.
C.The request must be made within 5 minutes of the previous request.
D.The 'max_tokens' parameter must be identical across both requests.
E.The request must use the Claude 3 Opus model exclusively.
AnswersA, B

Anthropic's caching mechanism uses a cryptographic hash of the content. Even a single extra space or a different newline character will result in a different hash, causing a cache miss. Developers must ensure that the static part of their prompt is strictly consistent across all API requests to maintain savings.

Why this answer

Prompt Caching is sensitive to the exact structure of the request. For a cache hit to occur, the content must be bit-for-bit identical to a previously cached version, and it must be positioned at the exact same location in the prompt (the prefix). Understanding these strict requirements prevents developers from accidentally breaking their cache and incurring unexpected costs due to minor formatting changes.

Exam trap

Candidates often assume that semantic similarity alone is enough for a cache hit, forgetting that prompt caching requires an exact bit-for-bit string match positioned precisely at the prefix of the prompt.

4
MCQhard

A legal tech company is using Claude 3.5 Sonnet to generate 5,000-word contract drafts. They notice that the output costs are significantly higher than the input costs. What is the most effective way to reduce these specific output costs without switching to a lower-quality model?

A.Use prompt caching to store the generated contract for later use.
B.Refine the prompt to include strict length constraints and eliminate 'conversational filler'.
C.Switch to the Batch API to get a 50% discount on the output tokens.
D.Increase the 'max_tokens' parameter to allow the model more room to think.
AnswerB

By explicitly instructing the model to be concise and avoid unnecessary introductory or concluding remarks, the developer can reduce the total number of output tokens. Since output tokens are the most expensive part of this specific workload, even a 10% reduction in verbosity can lead to substantial financial savings across thousands of generated contracts.

Why this answer

Output tokens are almost always more expensive than input tokens. In applications where the generation is very long (like contract drafting), managing the volume of generated text is the primary cost-control lever. Since the developer wants to keep the high-quality Sonnet 3.5 model, they must focus on prompting techniques that encourage conciseness or use structured generation to avoid redundant or unnecessary verbiage.

Exam trap

Candidates often attempt to reduce costs by switching to a weaker model, which degrades output quality, rather than optimizing the prompt to generate only necessary, high-quality content.

5
Multi-Selectmedium

Which THREE criteria are most important when deciding whether to upgrade from Claude 3 Haiku to Claude 3.5 Sonnet for a production application?

Select 3 answers
A.The complexity of the reasoning or creative task required.
B.The acceptable latency for the end-user experience.
C.The availability of the 'stream' parameter.
D.The budget allocated for API token consumption.
E.The total number of API keys generated for the project.
AnswersA, B, D

Claude 3.5 Sonnet has significantly higher intelligence and reasoning capabilities than Haiku. If your application involves complex coding, nuanced translation, or multi-step logical deduction, the upgrade is often necessary because Haiku may produce lower-quality or inaccurate results that could negatively impact the utility of the tool.

Why this answer

Upgrading models involves a trade-off between performance and cost. Developers must evaluate if the task requires higher-level reasoning (where Haiku might fail), if the application can tolerate the slightly higher latency of Sonnet, and if the budget can accommodate the significant price increase. These three factors ensure that the model choice aligns with both technical and business constraints.

Exam trap

Test-takers often evaluate model upgrades solely on intelligence gains while ignoring critical operational constraints like latency and budget.

6
MCQeasy

A developer is using the Claude API to generate creative writing pieces. The application often sends the same 2,000-token instruction set as a prefix to every request. Which feature can reduce the cost of processing this repeated prefix?

A.Batch API
B.Prompt Caching
C.Streaming responses
D.Using a smaller model
AnswerB

Prompt Caching allows developers to cache a static prefix, such as a long instruction set, so that it is not re-processed and re-billed at full input token rates on subsequent calls. This directly reduces the cost of the repeated 2,000-token prefix, making it the ideal solution for this scenario.

Why this answer

Prompt Caching is specifically designed to reduce costs when the same prefix is reused across multiple API calls. By caching the 2,000-token instruction set, subsequent calls avoid reprocessing those tokens at full input rates, directly lowering the cost. Other features like streaming, Batch API, or model selection do not target the redundant processing of a repeated prefix.

Exam trap

The trap here is assuming that any cost-saving feature, such as using a smaller model or the Batch API, will address the specific issue of a repeated prefix, when Prompt Caching is the only feature that directly caches and reuses that prefix.

7
MCQhard

A developer is optimizing a Claude-powered chatbot that handles 50,000 daily conversations. Each conversation starts with a 1,200-token system prompt that includes detailed instructions and examples. The system prompt is identical across all conversations. The developer wants to reduce input token costs. Which technique should be applied?

A.Shorten the system prompt by removing examples.
B.Enable prompt caching for the system prompt.
C.Move the system prompt to the first user message.
D.Switch to a cheaper model for all conversations.
AnswerB

Prompt caching allows frequently used prefixes, such as a static system prompt, to be stored and reused across API calls. After the first request, subsequent calls that share the same cached prefix are charged at a reduced rate for those tokens. Since the 1,200-token system prompt is identical across all 50,000 conversations, caching it will significantly lower input token costs.

Why this answer

Prompt caching is designed for scenarios where a large, static prefix is reused across many requests. By caching the 1,200-token system prompt, the developer pays full price only once, and subsequent requests benefit from reduced input token costs. This directly targets the repetitive cost without sacrificing quality or altering the conversation flow.

Exam trap

The trap here is thinking that reducing the prompt size or switching models is the only way to cut costs, overlooking prompt caching as a targeted solution for repeated prefixes.

8
Multi-Selectmedium

Which THREE factors directly determine the total cost of a single API call to the Claude Messages endpoint?

Select 3 answers
A.The number of input tokens in the request.
B.The number of output tokens generated by the model.
C.The latency of the model response in milliseconds.
D.The presence of cached tokens (cache hits or writes).
E.The number of concurrent requests being processed.
AnswersA, B, D

Every token sent to the model as part of the system prompt and message history is billed at the model's specific input rate. This is usually the largest component of cost for long-context applications, making prompt engineering and truncation essential strategies for minimizing expenses in production environments.

Why this answer

Accurate cost management requires understanding the breakdown of billable components. The total cost is determined by the number of input tokens sent, the number of output tokens generated by the model, and whether any tokens were served from or written to the prompt cache. Knowing these factors allows developers to optimize their prompts and choose the right models for their budgets.

Exam trap

Test-takers often forget that prompt caching introduces separate billing components, overlooking cache write and cache hit fees alongside standard input and output token counts when calculating total API expenditure.

9
MCQeasy

A developer is using the Claude API to generate marketing copy for 50 different product lines. Each request includes a unique set of product details and a unique creative brief. The developer wants to minimize cost while ensuring high-quality output. Which model should they choose?

A.claude-3-opus-20240229
B.claude-3-sonnet-20240229
C.claude-2.1
D.claude-3-haiku-20240307
AnswerB

Sonnet offers a strong balance of quality and cost, making it well-suited for creative writing tasks like marketing copy. It is less expensive than Opus but still capable of producing high-quality output. For 50 unique requests, Sonnet provides the best cost-performance trade-off. This choice aligns with the objective to minimize cost while maintaining quality.

Why this answer

Marketing copy generation requires creativity and coherence but not the deep reasoning of Opus. Sonnet provides a balanced combination of quality and cost, making it the most suitable for this task. Haiku might sacrifice quality, while Opus would be overkill and more expensive.

Claude 2.1 is outdated. Therefore, Sonnet is the optimal choice to minimize cost while ensuring high-quality output.

Exam trap

The trap here is assuming that the most expensive model always yields the best results, or that the cheapest model is always sufficient.

10
Multi-Selecthard

A developer is building a long-context application that processes 150,000 tokens per request. To manage costs and maintain performance, which TWO techniques should be prioritized?

Select 2 answers
A.Implementing Prompt Caching for the static background context.
B.Using Retrieval-Augmented Generation (RAG) to only send relevant snippets.
C.Converting the text to a more compact binary format before sending to the API.
D.Hard-coding the model to Claude 3 Opus for better token compression.
E.Increasing the 'temperature' setting to reduce the length of the generated output.
AnswersA, B

Prompt caching allows the developer to store the 150,000-token context on Anthropic's servers after the first request. Subsequent requests that use the same context only pay a small 'cache hit' fee rather than the full input token price. This is the single most impactful feature for reducing costs in applications that repeatedly reference large documents or datasets.

Why this answer

Managing long-context applications requires a combination of architectural efficiency and cost-saving features. Developers must minimize redundant data processing and ensure the model remains focused on relevant information. Utilizing prompt caching for static data and implementing RAG to limit the context sent to the model are the two most effective strategies for scaling long-context applications while keeping costs under control.

Exam trap

Candidates often suggest summarizing the entire context before sending it, which ignores the efficiency of prompt caching for static data and the precision of RAG for dynamic data retrieval.

11
MCQeasy

A developer is concerned about the 'token overhead' when using system prompts. How does Anthropic charge for tokens included in the system prompt compared to tokens in the user message?

A.System prompt tokens are billed at a premium rate because they influence the model more.
B.System prompt tokens are free for the first 500 tokens in every request.
C.They are billed at the same standard input token rate as the user message.
D.System prompts are only billed if they exceed the 200,000-token context limit.
AnswerC

All input tokens are treated identically in the billing pipeline. Whether a token is part of a system instruction, a user query, or a previous turn in the conversation, it is billed at the uniform input rate for the selected model. This allows for a simple and transparent cost calculation based solely on total input volume.

Why this answer

Understanding the billing of different prompt components is essential for cost management. Anthropic treats all input tokens equally, regardless of their role in the message (system, user, or assistant). This means that a large system prompt containing complex instructions or few-shot examples will incur the same per-token cost as the actual user query or any historical context provided in the conversation.

Exam trap

Candidates often assume system prompts are 'free' or billed differently, leading to poor cost estimation when using massive system prompts or extensive few-shot learning examples.

12
MCQhard

A developer runs a nightly batch job that sends 8,000 independent product-description summarization requests to the Anthropic API. Each request shares an identical 12,000-token instruction block and a unique 300-token product description. The team wants to reduce input-token spend without changing output quality or model. Which approach best meets this goal?

A.Use prompt caching with cache_control on the shared 12,000-token instruction block and place it at the beginning of each request.
B.Send all 8,000 product descriptions in a single request so the instruction block is transmitted only once.
C.Switch the batch job to a smaller, cheaper model for all 8,000 requests.
D.Reduce the shared instruction block to 2,000 tokens so that every request costs less per call.
AnswerA

Prompt caching rewards a stable prefix: the 12,000-token instruction block is identical across all 8,000 calls, so marking it with cache_control lets subsequent requests read it from cache at the reduced cache-read rate instead of paying full input price every time. The unique 300-token description stays outside the cached prefix, preserving correctness while cutting the dominant cost.

Why this answer

The dominant cost is the repeated 12,000-token instruction block, not the small per-product text. Prompt caching is designed for exactly this pattern: a stable, large prefix reused across many calls. Marking that block with cache_control lets later requests read it at the reduced cache-read rate, while the variable 300-token description remains uncached.

This reduces spend without changing the model, the instructions, or the output quality.

Exam trap

The trap here is assuming cost reduction always means shortening the prompt or downgrading the model, rather than reusing an unchanged large prefix through prompt caching.

13
MCQmedium

A company needs to summarize 500 academic papers, each approximately 20,000 tokens long. The summaries are needed for a weekly report due in three days. Which approach provides the best balance of cost-efficiency and model capability for this specific task?

A.Real-time Claude 3.5 Sonnet with Prompt Caching.
B.Real-time Claude 3 Haiku.
C.Claude 3.5 Sonnet via the Message Batch API.
D.Claude 3 Opus via the Message Batch API.
AnswerC

The Batch API is perfect for this scenario because the deadline is three days away and the volume is high. It allows the company to use the highly capable 3.5 Sonnet model while benefiting from a 50% cost reduction, making it the most balanced choice for quality and budget.

Why this answer

This scenario involves high-volume, long-context processing with a flexible deadline. The Message Batch API is the ideal choice here, as it offers a 50% discount on standard rates. Since academic papers require strong reasoning, using a more capable model like Sonnet 3.5 via the Batch API ensures high-quality summaries at a significantly lower price point than real-time processing.

Exam trap

Candidates often choose real-time API calls for massive batch jobs, failing to recognize that the Message Batch API provides a 50% discount specifically designed for high-volume, non-time-sensitive tasks.

14
MCQhard

A developer notices that a Claude integration's monthly bill is far higher than expected. Logs show that each request includes a 15,000-token system prompt containing detailed style guidelines, plus a 200-token user question, and the model generates a 50-token answer. The system prompt is identical across all requests. Which change would MOST reduce cost without altering output quality?

A.Shorten the system prompt by removing style guidelines that the model already follows by default.
B.Move the user question before the system prompt so the model processes it first.
C.Enable Prompt Caching on the static system prompt so repeated requests pay a reduced cache-read rate for those tokens.
D.Switch to a model with a smaller context window to force the system prompt to be truncated automatically.
AnswerC

Because the 15,000-token system prompt is identical across requests, Prompt Caching lets the API serve those tokens from cache at a lower rate after the first write. This preserves the exact prompt content, so output quality is unchanged, while substantially cutting input token costs on the dominant portion of each request. It is the most direct cost reduction.

Why this answer

The dominant cost driver is the 15,000-token system prompt repeated on every request. Since it is static, Prompt Caching reduces the billed rate for those tokens on subsequent calls while keeping the prompt intact. This preserves output quality exactly and targets the largest cost component, making it the most effective change.

Exam trap

The trap here is focusing on trimming prompt content to save tokens, when caching the identical large prefix achieves larger savings without touching the instructions that shape output quality.

15
MCQmedium

A developer needs to monitor the costs of different departments using a single Anthropic API key. Which API feature should be utilized to categorize and track usage without creating multiple accounts or keys?

A.The 'system' prompt field.
B.The 'metadata' object with a 'user_id' or custom field.
C.The 'stop_sequences' parameter.
D.The 'top_p' sampling parameter.
AnswerB

The 'metadata' parameter in the Messages API request allows developers to include an object with custom fields. This data is passed through to Anthropic's billing systems, making it possible to filter and group costs by department, project, or individual user in the dashboard or via usage reports.

Why this answer

Tracking usage across different business units is a common requirement for cost management. Anthropic provides the 'metadata' field in the Messages API, allowing developers to attach custom identifiers like 'department_id'. These tags appear in usage logs and billing exports, enabling precise cost allocation and budget monitoring without the administrative overhead of managing multiple API keys.

Exam trap

Candidates often think they need to provision separate API keys or accounts for each department, missing the built-in metadata parameter that allows custom tagging directly inside requests.

16
MCQmedium

A financial services firm is developing a customer support bot that requires low latency and high throughput for processing simple account inquiries. The firm expects over 500,000 requests per day and prioritizes cost-efficiency above the highest possible reasoning capabilities. Which model should the developer select to meet these specific business requirements?

A.Claude 3.5 Sonnet
B.Claude 3 Haiku
C.Claude 3 Opus
D.Claude 2.1
AnswerB

This model is specifically designed for speed and cost-effectiveness in high-volume applications. It handles simple tasks with minimal latency and has the lowest price point in the Claude 3 family. This makes it ideal for real-time customer support scenarios where response time and budget management are the primary technical constraints.

Why this answer

Selecting the right model involves balancing intelligence, speed, and cost. For high-volume, low-complexity tasks like account inquiries, Claude 3 Haiku offers the best performance-to-price ratio. It provides near-instant responses and significantly lower per-token costs compared to Sonnet or Opus, making it the industry standard for high-scale, simple automation where latency is a critical factor for user satisfaction.

Exam trap

Candidates sometimes prioritize 'intelligence' benchmarks over business requirements, failing to realize that Haiku's low latency and cost are superior for simple, high-volume automated support interactions.

17
MCQmedium

A developer is building a customer-facing chatbot using the Claude API. The bot must respond within 1.5 seconds on average. The team initially selected claude-3-opus-20240229 for its high quality, but latency is consistently above 3 seconds. They need to reduce latency while maintaining acceptable response quality. Which action should the developer take?

A.Increase the max_tokens parameter to allow the model to generate more tokens per response.
B.Switch to claude-3-haiku-20240307 and evaluate response quality against the original model.
C.Add a system prompt instructing Claude to respond as quickly as possible.
D.Enable streaming in the API request so that tokens are delivered incrementally.
AnswerB

Haiku is Anthropic's fastest and most cost-effective model, designed for low-latency tasks. In a chatbot with a 1.5-second target, Opus is too slow. Switching to Haiku directly addresses the latency constraint while still providing competent conversational ability. The developer should then validate that the quality meets the product's bar, but this is the correct first step to meet the performance requirement.

Why this answer

The chatbot's strict latency requirement makes model selection critical. Opus prioritizes quality over speed, while Haiku is optimized for fast responses. Switching to Haiku directly addresses the performance bottleneck.

The other options either do not affect model inference speed or only change how output is delivered, not how quickly it is produced. The developer should then verify that Haiku's quality is sufficient for the use case.

Exam trap

The trap here is assuming that parameter tweaks or prompt instructions can overcome a model's inherent latency profile.

18
Multi-Selecthard

A developer is reaching the Rate Limits for their account tier while using Claude 3.5 Sonnet. Which TWO actions would help manage these limits while also potentially reducing costs?

Select 2 answers
A.Offload non-urgent, high-volume tasks to the Batch API.
B.Use Claude 3 Haiku for simple pre-processing or classification steps.
C.Implement a retry logic with exponential backoff.
D.Request a manual increase of the account's Tier level.
E.Set the 'temperature' parameter to 0 to make responses more predictable.
AnswersA, B

The Batch API has its own set of rate limits that are separate from the synchronous Messages API. By moving bulk processing to the Batch endpoint, you free up your synchronous TPM/RPM for real-time users while also taking advantage of the 50% cost discount offered for batch processing.

Why this answer

Rate limits are often tied to token throughput (TPM). By switching to the Batch API, developers can access separate, often higher, throughput limits for non-urgent tasks at a lower price. Similarly, using a smaller model like Haiku for less demanding sub-tasks reduces the TPM count against the Sonnet quota and costs less, effectively managing both constraints simultaneously.

Exam trap

Candidates frequently overlook the Batch API as a solution for rate limits, incorrectly assuming it is only for cost savings rather than a mechanism to bypass concurrent request throughput bottlenecks.

19
MCQeasy

A developer needs to select a model for a code generation tool that must handle complex multi-file refactoring and advanced algorithmic logic. Which Claude 3.5 model currently provides the highest level of intelligence and coding capability for this task?

A.Claude 3 Haiku
B.Claude 3.5 Sonnet
C.Claude 3 Opus
D.Claude 2.0
AnswerB

Claude 3.5 Sonnet is currently Anthropic's most advanced model for coding and reasoning. It outperforms Claude 3 Opus on industry benchmarks while maintaining higher speeds. This makes it the premier choice for developers building tools that require a deep understanding of complex codebases and logical structures.

Why this answer

In the Anthropic model hierarchy, Claude 3.5 Sonnet is currently positioned as the most capable model for coding and complex reasoning, even surpassing Claude 3 Opus in many benchmarks. For tasks involving multi-file refactoring and deep logic, it provides the best performance. Understanding model positioning is crucial for developers to ensure they are using the most capable tool for high-stakes technical work.

Exam trap

Candidates often assume that older flagship models like Claude 3 Opus are always the best for coding, missing that Claude 3.5 Sonnet is currently the superior model for technical tasks.

20
MCQhard

An application uses Claude 3.5 Sonnet and implements Prompt Caching for a 10,000-token system prompt. The application processes 1,000 requests per hour. If the cache is refreshed every request and never expires, how does the billing for the input tokens change after the very first request?

A.The first request is free, and subsequent requests are full price.
B.Every request is billed at a flat 50% discount.
C.The first request incurs a 'cache write' fee, and the next 999 incur 'cache hit' fees.
D.Billing remains the same as standard input for all requests.
AnswerC

The first time a prompt is cached, the developer pays the 'cache write' rate for those 10,000 tokens. Since the cache is refreshed and hit by every subsequent request in the hour, the remaining 999 requests only pay the 'cache hit' rate, which is significantly cheaper than standard input pricing.

Why this answer

Prompt Caching introduces a two-tier pricing model for input tokens. The first request 'writes' the tokens to the cache, which is billed at a 'cache write' rate (usually 25% more than standard input). All subsequent requests that hit this cache are billed at the 'cache hit' rate, which is about 10% of the standard cost.

This shift dramatically lowers the long-term cost of the application.

Exam trap

Candidates incorrectly assume all cached requests are billed at the discounted hit rate from the very beginning, forgetting that the initial request must write to the cache at a premium rate.

21
MCQhard

A developer is using the Claude Messages API to process a 50,000-token document for a summarization task. The summarization prompt and instructions add another 1,000 tokens. The developer wants to minimize output token costs while ensuring a comprehensive summary. Which strategy is most effective?

A.Split the document into smaller chunks and summarize each separately, then combine the summaries.
B.Set the max_tokens parameter to a very low value, such as 100.
C.Craft a prompt that instructs Claude to produce a concise summary within a specified token limit, such as 'Summarize in under 500 tokens.'
D.Use a smaller model like Claude 3 Haiku for summarization.
AnswerC

By explicitly instructing Claude to produce a summary within a token limit, the developer guides the model to generate a concise output, directly reducing output token costs. This approach maintains summarization quality while controlling length. The model will still capture essential information but avoid unnecessary verbosity, making it the most effective strategy for cost management.

Why this answer

To minimize output token costs while ensuring a comprehensive summary, the developer should control the output length through prompt instructions. Explicitly requesting a concise summary within a token limit guides the model to generate only the necessary content, reducing cost without sacrificing quality. Other strategies like chunking or using a smaller model may address different concerns but do not directly target output token reduction.

Exam trap

The trap here is focusing on reducing input costs or model size when the question specifically asks about minimizing output token costs, which is best addressed by controlling the generated response length.

22
MCQmedium

A team is estimating costs for a new feature that will send 1 million requests per month. Each request has a 2,000-token input and generates a 500-token output. The team wants to reduce the output token cost, which dominates the bill. Which strategy is MOST effective for reducing output token costs?

A.Enable Prompt Caching on the input to reduce the input token cost.
B.Increase the max_tokens parameter to allow longer responses.
C.Switch to a model with a larger context window to accommodate longer inputs.
D.Instruct the model to be concise and set a lower max_tokens limit to cap response length.
AnswerD

Output tokens are billed per generated token, so constraining response length directly reduces the dominant cost. Instructing conciseness and lowering max_tokens caps the worst-case output size. This approach targets the largest cost driver without sacrificing the feature's purpose, since the responses remain useful but shorter.

Why this answer

Because output tokens dominate the monthly bill, the most effective lever is limiting how many tokens the model generates. Lowering max_tokens and prompting for conciseness directly caps output length and thus cost. Optimizing input tokens or context capacity addresses smaller or irrelevant components of the bill.

Exam trap

The trap here is applying a familiar optimization like Prompt Caching reflexively, even when the scenario explicitly states that output tokens, not input tokens, are the dominant cost.

23
MCQmedium

A developer is building a real-time translation service that must respond within 1 second for short phrases. The service will handle thousands of requests per hour. Quality is important, but latency and cost are critical. Which Claude model is the most appropriate?

A.Claude 3 Haiku
B.Claude 3 Opus
C.Claude 3.5 Sonnet
D.Claude 3 Sonnet
AnswerA

Claude 3 Haiku is designed for speed and cost efficiency, making it ideal for real-time translation of short phrases. It can deliver responses within the 1-second latency requirement and handle thousands of requests per hour at a low cost. For straightforward translation, its quality is sufficient, and it meets the critical latency and cost constraints.

Why this answer

Claude 3 Haiku is optimized for low latency and low cost, making it the best fit for real-time translation of short phrases. It can meet the 1-second response requirement and handle high throughput without excessive expense. More powerful models like Sonnet or Opus would add latency and cost without necessary quality gains for this straightforward task.

Exam trap

The trap here is assuming that higher quality models are always needed, ignoring that latency and cost constraints may make a faster, cheaper model the correct choice.

24
MCQeasy

When calculating the estimated cost of a project using Claude, which metric is used by Anthropic to measure the volume of data processed and generated?

A.Characters (including spaces).
B.Total words in the prompt.
C.Tokens.
D.API call duration in seconds.
AnswerC

Tokens are the atomic unit of processing for Claude. Anthropic's pricing is strictly defined as a cost per million tokens. This includes both the input tokens sent by the user and the output tokens generated by the model. This is the standard metric for all cost and performance calculations.

Why this answer

Anthropic, like most LLM providers, uses 'tokens' as the fundamental unit of measurement for billing. Tokens represent chunks of text (roughly 3/4 of a word). Cost management involves estimating the total number of input tokens (prompt) and output tokens (response).

Understanding this unit is the first step in any cost-estimation exercise for a developer using Claude.

Exam trap

Candidates sometimes confuse billing metrics like word counts, character lengths, or API call frequencies with the actual fundamental unit used by Anthropic to measure data volume.

25
MCQeasy

Which of the following scenarios describes the most effective use of Prompt Caching for cost management?

A.A chatbot where every user query is unique and no history is maintained.
B.A translation service that processes single words one at a time.
C.An AI assistant that references a 50-page technical manual for every user query.
D.A daily report generator that uses a completely different dataset every morning.
AnswerC

This is the ideal use case for prompt caching. The 50-page manual acts as a large, static prefix that is sent with every request. By caching the manual, the developer only pays the full input price once, and all subsequent queries only pay for the much cheaper cache read tokens, resulting in massive long-term cost savings.

Why this answer

Prompt caching is most effective when a large amount of static information is reused across many different requests. It allows the model to 'remember' the prefix of a prompt, significantly reducing the cost of processing that prefix in subsequent calls. Identifying workloads with high prefix overlap is key to maximizing the financial benefits of this feature in production environments.

Exam trap

Candidates often apply prompt caching to dynamic or highly variable user inputs, failing to realize that caching is only financially beneficial when the same prefix is reused across many requests.

26
MCQmedium

An enterprise is migrating a document processing pipeline that handles 50,000 PDFs daily. Each PDF is converted to text (approx. 2,000 tokens) and requires a summary. The project has a strict budget. Which approach provides the most significant cost reduction while utilizing Claude 3.5 Sonnet?

A.Implementing client-side compression on the PDF text before sending.
B.Using the Anthropic Batch API for asynchronous processing.
C.Switching the entire pipeline to Claude 3 Haiku.
D.Reducing the 'max_tokens' parameter to 50 for every summary.
AnswerB

The Batch API allows developers to submit large groups of requests that are processed within a 24-hour window at a 50% discount compared to standard real-time API prices. For document processing pipelines where immediate results are not required, this is the most effective way to utilize the reasoning power of Claude 3.5 Sonnet within a limited budget.

Why this answer

Cost management in high-volume environments often involves leveraging specific API features designed for non-latency-sensitive workloads. The Anthropic Batch API is specifically designed for processing large volumes of data asynchronously at a significantly reduced price point. This allows developers to use high-intelligence models like Sonnet 3.5 for complex tasks while adhering to strict budgetary constraints that would be exceeded by standard API calls.

Exam trap

Candidates often recommend expensive real-time API calls for massive, non-urgent workloads instead of leveraging asynchronous processing features designed for cost savings.

27
MCQmedium

A developer wants to implement a 'Summary' feature for a long conversation history. As the conversation grows, the cost of sending the entire history with every new message increases. What is the most cost-effective architectural pattern to handle this?

A.Always send the full conversation history to maintain maximum context.
B.Use Prompt Caching for the entire dynamic conversation history.
C.Implement a sliding window that only sends the last 5 messages.
D.Periodically summarize the history and use the summary as context.
AnswerD

This pattern, known as context distillation or compression, involves replacing older messages with a concise summary. This drastically reduces the number of input tokens sent in subsequent requests while preserving the important information, making it the most cost-effective way to handle long-running, context-heavy sessions.

Why this answer

Managing long-running conversations requires balancing context and cost. Periodically summarizing the previous conversation and replacing the detailed history with that summary (Context Compression) keeps the input token count low. This limits the linear growth of costs as the session continues, ensuring the application remains affordable even during extended user interactions.

Exam trap

Candidates frequently select stateless strategies like completely truncating old messages or sending raw unlimited histories, overlooking how periodic summarization preserves crucial long-term context while strictly controlling linearly increasing token costs.

28
MCQmedium

A developer is comparing the cost of using Claude 3 Opus versus Claude 3.5 Sonnet for a task that requires complex reasoning. The task involves processing 1,000 requests, each with 500 input tokens and 200 output tokens. Which statement accurately reflects the cost consideration?

A.The cost is identical because both models are billed at the same rate for input and output tokens.
B.Claude 3 Opus is cheaper for output tokens but more expensive for input tokens compared to Claude 3.5 Sonnet.
C.Claude 3 Opus is always more cost-effective for complex reasoning because it requires fewer tokens to achieve the same result.
D.Claude 3.5 Sonnet is less expensive per token than Claude 3 Opus, but it may require more tokens to match Opus's reasoning quality.
AnswerD

Claude 3.5 Sonnet has a lower per-token cost than Claude 3 Opus, but for highly complex reasoning, it might need more tokens or additional prompting to achieve similar results. The total cost depends on both the per-token price and the number of tokens required. This statement accurately captures the trade-off developers must evaluate.

Why this answer

When choosing between Claude 3 Opus and Claude 3.5 Sonnet for complex reasoning, developers must weigh per-token cost against the number of tokens needed. Sonnet is cheaper per token, but Opus may deliver better results with fewer tokens. The correct statement acknowledges that Sonnet is less expensive per token but might require more tokens to match Opus's quality, making total cost dependent on the specific task and prompt engineering.

Exam trap

The trap here is assuming that a more powerful model like Opus is always more expensive overall, or that a cheaper model like Sonnet will always be cheaper in total, without considering how token usage might differ between models.

29
MCQmedium

A company needs to process 10 million short customer feedback snippets to identify 'bug reports' vs 'feature requests'. Speed and budget are the primary constraints, while the classification logic is straightforward. Which model provides the best throughput-to-cost ratio?

A.Claude 3.5 Sonnet
B.Claude 3 Opus
C.Claude 3 Haiku
D.Claude 2.0
AnswerC

Haiku is the fastest and least expensive model, making it the superior choice for high-volume classification. It can process millions of tokens for a very low cost while maintaining the accuracy needed for simple tasks like distinguishing between bug reports and feature requests. This maximizes the return on investment for the company's data processing pipeline.

Why this answer

When dealing with massive datasets and simple logic, the most important metric is the cost per million tokens. Claude 3 Haiku is specifically optimized for these 'utility' tasks, offering high throughput and the lowest pricing in the Claude 3 family. Selecting a larger model for such a simple, high-volume task would lead to unnecessary expenditures without providing a noticeable improvement in classification quality.

Exam trap

Candidates often default to the most capable model (like Sonnet or Opus) for simple classification tasks, ignoring that Haiku is specifically engineered for high-throughput, low-cost utility operations.

30
MCQmedium

A developer is building a real-time chat application where users expect responses in under two seconds. The prompts are short (under 200 tokens) and the responses are typically one or two sentences. Which Claude model should the developer choose to optimize for latency and cost?

A.Claude 3.5 Sonnet
B.Claude 3 Haiku
C.Claude 3 Opus
D.Claude 2.1
AnswerB

Claude 3 Haiku is Anthropic's fastest and most cost-effective model, optimized for near-instant responses on simple tasks. It handles short prompts and brief outputs efficiently, meeting the sub-two-second latency requirement while minimizing cost. For a real-time chat application with straightforward interactions, Haiku provides the best balance of speed and affordability.

Why this answer

For a real-time chat application with short prompts and simple responses, latency and cost are critical. Claude 3 Haiku is specifically designed to be the fastest and most affordable model in the Claude 3 family, making it ideal for this use case. More powerful models like Opus or Sonnet would add unnecessary cost and latency without providing meaningful benefits for such straightforward interactions.

Exam trap

The trap here is assuming that a newer or more powerful model like Claude 3.5 Sonnet is always the best choice, when in fact Haiku is optimized for speed and cost in simple, high-volume scenarios.

31
MCQmedium

You are building a summarization feature that processes 20,000 support tickets nightly. Each ticket is under 4,000 tokens and the output summary is around 300 tokens. The job must complete within a 6-hour window and you want to minimize cost. Which Claude model should you choose?

A.Claude 3 Sonnet
B.Claude 3 Opus
C.Claude 3.5 Sonnet
D.Claude 3 Haiku
AnswerD

Claude 3 Haiku is the fastest and most cost-effective model in the Claude 3 family. For summarizing short tickets under 4,000 tokens with a modest 300-token output, Haiku delivers adequate quality at the lowest price per token, making it ideal for high-volume, budget-sensitive batch jobs that must finish within a 6-hour window.

Why this answer

For high-volume, short-input summarization where cost is the primary constraint, Claude 3 Haiku provides the best balance of speed, quality, and price. Its lower per-token cost directly reduces the total spend for 20,000 requests, and its speed helps meet the 6-hour window. More expensive models like Sonnet or Opus are overkill for this task.

Exam trap

The trap here is assuming that a more capable model is always better, when the scenario explicitly prioritizes cost minimization for a simple task.

32
Multi-Selectmedium

A developer is building an application that uses the Claude Messages API to generate product descriptions. The application sends a system prompt of 1,500 tokens, a user prompt of 200 tokens, and receives a response of 300 tokens. The developer wants to reduce costs. Which two strategies would directly reduce the cost per API call? (Choose two.)

Select 2 answers
A.Use a smaller model like Claude 3 Haiku instead of Claude 3 Opus.
B.Enable streaming to receive tokens as they are generated.
C.Shorten the system prompt by removing redundant instructions.
D.Increase the max_tokens parameter to allow longer responses.
E.Cache the model's responses for identical prompts.
AnswersA, C

Claude 3 Haiku has a lower cost per token than Claude 3 Opus. Switching to Haiku for a task like generating product descriptions, which does not require the advanced reasoning of Opus, reduces the cost per API call. This is a direct cost-saving measure as long as the smaller model meets quality requirements.

Why this answer

The cost per API call is determined by the number of input and output tokens and the model's pricing. Shortening the system prompt reduces input tokens, and using a smaller model like Claude 3 Haiku lowers the per-token cost. Both directly decrease the cost of each call.

Increasing max_tokens, enabling streaming, or caching responses do not reduce the token-based cost of an individual call.

Exam trap

The trap here is confusing features that improve performance or reduce call volume, like streaming or caching, with those that directly lower the token-based cost of a single API call.

33
MCQeasy

A developer is using the Claude Messages API to generate a 500-token response. The input consists of a 200-token user message and a 100-token system prompt. Which factor directly determines the output token cost of this API call?

A.The number of tokens in the user message.
B.The number of tokens in the system prompt.
C.The total number of tokens in the request (input plus output).
D.The number of tokens in the model's response.
AnswerD

Output token cost is directly proportional to the number of tokens the model generates in its response. In this scenario, the response is 500 tokens, so the output cost is based on those 500 tokens. The input tokens (system prompt and user message) affect input cost separately, but the question asks specifically about output token cost.

Why this answer

Output token cost is based exclusively on the number of tokens the model generates in its response. In this case, the 500-token response defines the output cost. Input tokens, such as the system prompt and user message, are billed separately as input tokens.

Therefore, the response length is the direct determinant of output token cost.

Exam trap

The trap here is conflating total token count or input tokens with output token cost, when output cost depends solely on generated tokens.

34
MCQmedium

Refer to the exhibit. A developer is implementing the provided JSON structure to optimize an application that repeatedly analyzes the same large report. What is the primary financial implication of using the 'cache_control' block in this specific API request?

A.It eliminates the cost of the first 1024 tokens in every request.
B.It triggers a 50% discount on all output tokens for the session.
C.Subsequent requests with the same report will be billed at a reduced rate.
D.The request will be processed using the Batch API pricing model.
AnswerC

When a content block is marked with 'cache_control', the system stores the processed tokens. If a follow-up request contains the exact same text, the model reuses the cached state. These 'cache hits' are billed at a fraction of the cost of standard input tokens, leading to substantial savings for repetitive tasks.

Why this answer

The 'cache_control' block enables Prompt Caching, which is a key tool for cost management. By marking a block as ephemeral, Anthropic caches the preceding content. If the same content is sent again within the cache's lifetime, the developer is charged a significantly lower 'cache hit' rate instead of the full input token price.

This is essential for reducing costs in applications with static contexts.

Exam trap

Candidates often confuse the 'cache_control' block with general performance tuning, failing to identify that its primary purpose is enabling discounted pricing for repeated, static input context.

35
Multi-Selecthard

A developer is optimizing a Claude-powered document processing pipeline that sends large, mostly identical legal templates followed by short variable fields. They want to reduce input token costs while preserving output fidelity. Which TWO strategies are appropriate? (Choose two.)

Select 2 answers
A.Enable Prompt Caching on the static legal template so repeated requests reuse the cached prefix.
B.Remove the legal template entirely and rely on the model's pretrained knowledge of legal language.
C.Set max_tokens to a very low value to force the model to answer briefly.
D.Send only the variable fields and ask the model to reconstruct the template from memory.
E.Place the variable fields at the end of the prompt after the static template to maximize cache hits.
AnswersA, E

The legal template is large and mostly identical across requests, making it an ideal candidate for Prompt Caching. Cache reads are billed at a reduced rate, and the template content remains unchanged, so output fidelity is preserved. This directly targets the largest repeated input component and is a standard cost optimization for template-heavy pipelines.

Why this answer

The two effective strategies are caching the static template and ordering the prompt so the stable prefix comes first. Caching reduces the billed rate for the large repeated content, and correct ordering ensures the cache key remains valid across requests. Together they cut input token costs while leaving the template content and output quality intact.

Exam trap

The trap here is treating Prompt Caching as independent of prompt structure, when cache hits actually depend on keeping the stable prefix byte-identical and placing volatile content after it.

36
MCQeasy

A startup is prototyping a chatbot that handles simple FAQ responses for a small user base. The team wants the lowest possible cost per request and does not need advanced reasoning. Which Claude model selection strategy is MOST appropriate?

A.Use a different provider's cheapest model since all providers offer equivalent FAQ performance.
B.Use the largest, most capable model to ensure the highest quality answers regardless of cost.
C.Use a mid-tier model for all requests and rely on prompt engineering to reduce token usage.
D.Use a smaller, lower-cost model such as Claude Haiku, which is optimized for speed and cost on simpler tasks.
AnswerD

Claude Haiku is positioned as the fastest and most cost-effective model in the Claude family, making it ideal for high-volume, low-complexity tasks like FAQ responses. It provides sufficient quality for straightforward question answering while keeping per-token costs low. This aligns directly with the startup's goal of minimizing cost during prototyping.

Why this answer

Matching model capability to task complexity is the core principle of cost-effective model selection. Simple FAQ responses do not need frontier reasoning, so the smallest and cheapest Claude model provides adequate quality at the lowest per-token cost. Larger models would increase spend without meaningful quality improvement for this workload.

Exam trap

The trap here is defaulting to the most capable model out of caution, when the workload's low complexity makes a smaller model both sufficient and far cheaper.

37
MCQeasy

In the context of Anthropic's pricing model, why is it generally recommended to provide only the necessary context rather than the entire available dataset in a single prompt?

A.Because Anthropic charges a 'search fee' for every 1,000 tokens of context.
B.To minimize the input token count and keep the per-request cost low.
C.Because Claude models cannot process more than 10,000 tokens at a time.
D.To prevent the model from reaching its daily 'knowledge limit'.
AnswerB

Since billing is calculated per token, reducing the amount of context directly lowers the cost of the request. In a production environment with millions of calls, stripping away irrelevant data ensures that the budget is spent only on the information necessary for the model to produce a correct and helpful response for the user.

Why this answer

Token-based pricing means that every piece of information sent to the model has a direct financial cost. Large prompts not only increase the bill but can also lead to 'context stuffing,' which might degrade the model's focus. Efficient prompt engineering—sending only what is needed—is the fundamental practice for both cost management and ensuring high-quality, relevant responses from the AI.

Exam trap

Candidates believe dumping entire datasets into a prompt is safer than curating context, ignoring the financial and performance penalties.

38
Multi-Selecteasy

A developer needs to estimate the monthly budget for a new internal knowledge base application powered by Claude. Which TWO factors directly influence the total token consumption and subsequent cost of the API requests?

Select 2 answers
A.The total number of input tokens in the prompt
B.The number of concurrent users accessing the application
C.The total number of output tokens in the completion
D.The physical geographic location of the application server
E.The programming language used to make the API calls
AnswersA, C

Input tokens represent the data sent to the API, including system instructions, context, and user queries. Anthropic charges based on the quantity of these tokens, and large context windows or extensive document embeddings directly increase this portion of the bill. Monitoring input volume is essential for maintaining a predictable budget during scaling phases.

Why this answer

Estimating costs requires understanding the components of the billing model used by Anthropic. Total cost is derived from the volume of input tokens sent to the model and the volume of output tokens generated by the model. Developers must account for both segments because they are priced at different rates across the Claude 3 model family to reflect processing requirements.

Exam trap

Test-takers frequently forget to account for both input and output tokens separately, assuming flat-rate pricing or focusing only on prompt size.

39
MCQmedium

You are designing an AI agent that performs multi-step reasoning. The first step involves basic data extraction, while the second step requires complex logical deduction based on the extracted data. How should you select models to optimize for both performance and cost?

A.Use Claude 3 Opus for both steps to ensure maximum consistency.
B.Use Claude 3 Haiku for both steps to minimize total operational costs.
C.Use Claude 3 Haiku for extraction and Claude 3.5 Sonnet for logical deduction.
D.Use Claude 3.5 Sonnet for extraction and Claude 3 Haiku for logical deduction.
AnswerC

This tiered approach, known as model routing or chaining, optimizes for both cost and intelligence. Haiku handles the high-volume, low-complexity extraction task cheaply and quickly, while the more expensive Sonnet is reserved for the difficult reasoning phase. This balance ensures the agent remains reliable while keeping the overall cost per execution much lower than using Sonnet alone.

Why this answer

Model routing is a sophisticated cost management technique where different tasks within a single workflow are assigned to the most appropriate model. Using a 'one-size-fits-all' approach often leads to waste. By chaining models—using a cheaper model for simple tasks and a more powerful one for complex tasks—developers can achieve high-quality results while minimizing the total token spend for the entire process.

Exam trap

Test-takers often default to using a single expensive model for an entire multi-step workflow, wasting budget on simple preliminary tasks.

40
MCQmedium

A developer is building a Claude-powered coding assistant that must answer questions about a 60,000-token proprietary codebase on every user request. The codebase is static and updated only weekly. The developer wants to minimize per-request input token costs while keeping latency low. Which approach is MOST cost-effective?

A.Use the smallest available model and hope it can reason about the codebase without additional context.
B.Send only the file names and ask Claude to infer the implementation details from the names.
C.Enable Prompt Caching on the codebase context so subsequent requests reuse the cached prefix at a reduced input token rate.
D.Compress the codebase using a custom tokenizer before sending it to the Messages API.
AnswerC

Prompt Caching stores the static codebase prefix server-side for a short TTL, so repeated requests within that window are billed at a discounted cache-read rate instead of full input token price. Because the codebase rarely changes, the cache hit rate is high and latency improves. This directly reduces per-request input cost while preserving full context fidelity.

Why this answer

Prompt Caching is designed exactly for repeated large static prefixes such as a codebase that changes infrequently. Cache reads are billed at a lower rate than standard input tokens, and the cache persists for a short TTL, so a weekly update cadence yields very high hit rates. The other approaches either discard necessary context or rely on unsupported mechanisms.

Exam trap

The trap here is assuming that simply reducing token count by any means is a valid cost optimization, when the method must preserve both API compatibility and answer quality.

Ready to test yourself?

Try a timed practice session using only Model Selection And Cost Management questions.