Courseiva

CCDV-F Model Selection and Cost Management Practice Question

A developer is building a customer-facing chatbot using the Claude API. The bot must respond within 1.5 seconds on average. The team initially selected claude-3-opus-20240229 for its high quality, but latency is consistently above 3 seconds. They need to reduce latency while maintaining acceptable response quality. Which action should the developer take?

⚠ Common exam trap

The trap here is assuming that parameter tweaks or prompt instructions can overcome a model's inherent latency profile.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Switch to claude-3-haiku-20240307 and evaluate response quality against the original model.

The chatbot's strict latency requirement makes model selection critical. Opus prioritizes quality over speed, while Haiku is optimized for fast responses. Switching to Haiku directly addresses the performance bottleneck. The other options either do not affect model inference speed or only change how output is delivered, not how quickly it is produced. The developer should then verify that Haiku's quality is sufficient for the use case.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the max_tokens parameter to allow the model to generate more tokens per response.

    Why it's wrong here

    Increasing max_tokens does not reduce latency; it allows longer outputs, which generally increases total generation time. The chatbot's latency problem stems from the model's inherent speed, not from an output length cap. Raising max_tokens would likely worsen the average response time and is not a viable solution for meeting the 1.5-second target.

  • ✓

    Switch to claude-3-haiku-20240307 and evaluate response quality against the original model.

    Why this is correct

    Haiku is Anthropic's fastest and most cost-effective model, designed for low-latency tasks. In a chatbot with a 1.5-second target, Opus is too slow. Switching to Haiku directly addresses the latency constraint while still providing competent conversational ability. The developer should then validate that the quality meets the product's bar, but this is the correct first step to meet the performance requirement.

  • ✗

    Add a system prompt instructing Claude to respond as quickly as possible.

    Why it's wrong here

    System prompts guide content and style, not the underlying inference speed of the model. Claude cannot accelerate its own token generation based on an instruction. While a system prompt could encourage shorter answers, it does not change the model's fixed latency characteristics. The primary bottleneck remains the model choice, so this action is ineffective for the stated goal.

  • ✗

    Enable streaming in the API request so that tokens are delivered incrementally.

    Why it's wrong here

    Streaming improves perceived latency by delivering tokens as they are generated, but it does not reduce the total time to generate a full response. If the requirement is an average response time of 1.5 seconds for complete answers, streaming alone will not meet it. The underlying model speed is still the limiting factor, so this is not the correct solution.

About these practice questions

This CCDV-F question is part of Courseiva's 257-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCDV-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCDV-F exam.