CCDV-F Claude API Mechanics Practice Question
A developer is building a document Q&A service on the Anthropic Messages API. Their integration currently sends the entire 240,000-token knowledge base on every turn of a long conversation, and they are hitting the model's context window limit. They want to keep the full conversation history and the knowledge base available without exceeding the window. Which approach best addresses the problem?
⚠ Common exam trap
The trap here is assuming prompt caching or streaming expands the usable context window rather than only changing cost, latency, or delivery mechanics.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Summarize or truncate older conversation turns and retrieve only the most relevant knowledge-base chunks to include in each request.
The context window is a fixed ceiling on combined input and output tokens per request. To keep a long conversation and a large corpus usable, the developer must shrink what is actually sent: condense prior turns and inject only the retrieved passages relevant to the current question. Caching and streaming change cost or delivery mechanics, not the window itself, and max_tokens governs output only.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable prompt caching on the static knowledge base so repeated prefixes are served from cache and no longer count toward the context window.
Why it's wrong here
Prompt caching reduces latency and cost for repeated prefixes, but cached content still occupies the context window on each request. The model must be able to attend to the tokens, so caching does not raise the maximum token budget. A developer relying on caching to fit an oversized prompt will still receive a context-length error.
- ✗
Increase the max_tokens parameter so Claude can internally compress the knowledge base before answering.
Why it's wrong here
The max_tokens parameter caps only the generated output tokens, not the input. It has no effect on the size of the prompt and cannot cause Claude to compress the input. Raising it may even cause a request to fail if the combined input plus requested output exceeds the model's context limit.
- ✓
Summarize or truncate older conversation turns and retrieve only the most relevant knowledge-base chunks to include in each request.
Why this is correct
The context window is a hard cap on the total tokens sent and generated per request. Reducing the payload by condensing prior turns and injecting only relevant retrieved chunks keeps the request within the window while preserving useful information. This is the standard retrieval-plus-summarization pattern for long-running document assistants.
- ✗
Switch the request to stream: true so the server transmits the knowledge base in smaller chunks that bypass the context limit.
Why it's wrong here
Streaming controls how the response is delivered to the client, not how many tokens the model processes. The full prompt is still submitted and evaluated against the context window. Enabling streaming does not change token accounting and will not resolve a context-length error.
About these practice questions
This CCDV-F question is part of Courseiva's 257-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCDV-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCDV-F exam.