CCAR-P Practice Question: Developer Productivity and Operational Enablement
A developer wants to reduce the cost and latency of their LLM application, which sends repetitive, long context windows. Which optimization technique is most appropriate?
⚠ Common exam trap
Candidates often suggest reducing the context window or shortening the prompt, which degrades performance, instead of using prompt caching to optimize repetitive, large context segments.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement prompt caching for static segments of the context window.
Prompt Caching is designed specifically to handle large, static context segments by allowing the model to reuse the computed state. By caching parts of the prompt, the application avoids redundant processing, which directly reduces both latency and cost. This technique is a crucial operational enabler for developers building complex apps that rely on large knowledge bases or extensive instructions.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Aggressively compress all input text using standard text compression algorithms.
Why it's wrong here
LLMs do not natively understand compressed text formats like ZIP or GZIP. The model would require a full input stream to be decompressed, which doesn't address the cost or latency of the model inference itself. This adds unnecessary client-side processing without benefiting the LLM model performance.
- ✓
Implement prompt caching for static segments of the context window.
Why this is correct
Prompt Caching allows the developer to store and reuse the prefix of a prompt. This drastically reduces the number of tokens processed by the model on every call, leading to lower latency and significantly reduced costs for applications that involve large, unchanging instructions or documents.
- ✗
Switch to a smaller model version regardless of performance requirements.
Why it's wrong here
While switching to a smaller model reduces costs, it may cause a significant decline in output quality if the task requires high reasoning capabilities. Optimization techniques like Prompt Caching allow developers to maintain model performance while reducing costs, which is a superior engineering trade-off.
- ✗
Move all conversation processing to the client-side browser to offload the server.
Why it's wrong here
Moving LLM inference to the client-side exposes API keys, increases the risk of tampering, and is limited by the user's hardware performance. It does not solve the underlying cost of the inference service itself, as the requests must still be processed by the model provider's infrastructure.
About these practice questions
Courseiva writes every CCAR-P question from scratch — 262 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.