CCAR-F Context and Reliability Practice Question
You are architecting a customer-support assistant that uses Claude with a 200K-token context window. Each session accumulates roughly 40K tokens of chat history, and you also inject a 30K-token product manual on every turn. Latency has become unacceptable because the full payload is re-sent each call. Which change best reduces latency while preserving the assistant's knowledge of earlier turns?
⚠ Common exam trap
The trap here is assuming that context-window size or output limits drive latency, when the real cost in multi-turn applications is repeatedly reprocessing an unchanging input prefix.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable prompt caching on the stable prefix (manual plus prior turns) so repeated tokens are served from cache instead of being reprocessed on each call.
Prompt caching is the architectural lever for repeated large prefixes: the manual and accumulated history are stable across turns, so caching their processed state avoids re-encoding tens of thousands of tokens on every request. Caching preserves the full context the assistant needs, unlike truncation, and it targets input processing cost, unlike max_tokens or batching, which address output length or throughput rather than interactive latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Batch all user questions into a single nightly asynchronous request using the Message Batches API.
Why it's wrong here
The Message Batches API is designed for high-throughput offline workloads where latency is not a concern; results are typically returned within hours. A live customer-support assistant requires synchronous, interactive responses, so batching would make the experience far worse. It also does not reduce per-call input processing cost.
- ✓
Enable prompt caching on the stable prefix (manual plus prior turns) so repeated tokens are served from cache instead of being reprocessed on each call.
Why this is correct
Prompt caching stores the processed key/value state for a stable prefix, so subsequent calls that share that prefix skip re-encoding those tokens. Here the 30K-token manual and the growing history form a large, mostly stable prefix, so cache hits cut both time-to-first-token and cost while the assistant still sees all prior turns.
- ✗
Raise the max_tokens parameter so Claude can emit a longer response and finish the answer in fewer round trips.
Why it's wrong here
max_tokens caps the length of generated output; it does nothing to reduce the cost of re-reading the 70K-token input on every request. Increasing it may even lengthen responses and worsen perceived latency. The bottleneck in this scenario is input reprocessing, not output truncation, so this parameter change cannot address the problem.
- ✗
Switch to a smaller model with a shorter context window and truncate the manual to the first 8K tokens.
Why it's wrong here
Truncating the manual discards product information the assistant needs, degrading answer quality, and a smaller context window may force dropping earlier turns entirely. This trades correctness for speed rather than optimizing the existing workload. The scenario asks to preserve knowledge of earlier turns, which truncation directly violates.
About these practice questions
One of 271 original CCAR-F practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-F exam.