Courseiva
Claude API Mechanics →hardMultiple Choice

CCDV-F Claude API Mechanics Practice Question

A high-traffic application is frequently sending the same large set of reference documents (100,000 tokens) in the system prompt for every request. Which Claude API feature would most effectively reduce both the latency and the cost of these requests?

⚠ Common exam trap

Candidates often try to optimize by manually truncating the prompt or using external vector databases, missing that built-in Prompt Caching is the specific, optimized mechanism for handling repetitive large context.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implementing Prompt Caching using the 'cache_control' metadata.

Prompt Caching is a specialized mechanic for handling repetitive, large-scale context. By marking parts of the prompt as cacheable, developers can avoid re-processing the same data across multiple requests. This lead to significant cost savings on input tokens and drastically reduces the time to first token for the model's generated response.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switching from Claude 3 Opus to Claude 3 Haiku.

    Why it's wrong here

    While Haiku is cheaper and faster, it does not solve the inefficiency of re-processing 100,000 tokens for every single request. The fundamental problem is the redundant processing of the input, which will still be charged and processed at full scale on any model without a specific caching mechanism.

  • ✓

    Implementing Prompt Caching using the 'cache_control' metadata.

    Why this is correct

    Prompt Caching allows the API to store the results of processing a prefix of the prompt. By tagging the large reference documents with 'cache_control: {"type": "ephemeral"}', subsequent requests that use the same prefix can reuse the cached state, leading to lower costs and much faster response times.

  • ✗

    Compressing the reference documents using a text summarization tool.

    Why it's wrong here

    Summarization can reduce the token count but at the cost of losing detail and nuance from the original documents. It also adds an extra processing step that increases complexity. Prompt caching is a superior solution because it allows you to keep the full context while still achieving performance and cost benefits.

  • ✗

    Increasing the 'max_tokens' limit to allow for larger responses.

    Why it's wrong here

    Increasing max_tokens only affects the length of the model's output; it has no impact on the efficiency or cost of processing the large input prompt. In fact, larger responses combined with large inputs might increase the risk of hitting timeouts or rate limits without addressing the underlying redundancy.

About these practice questions

Courseiva writes every CCDV-F question from scratch — 257 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCDV-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCDV-F exam.