CCAR-P Practice Question: Developer Productivity and Operational Enablement
A developer productivity team is building an internal coding assistant that calls the Claude Messages API. During a spike in usage, the assistant starts failing with 429 responses and users see truncated answers. The team wants the assistant to degrade gracefully under load rather than fail outright, while keeping latency predictable for interactive use. Which change best meets these goals?
⚠ Common exam trap
The trap here is treating 429 errors as a signal to retry harder or to shrink the model, when the durable fix is to shape traffic so interactive requests keep a guaranteed share of rate-limit capacity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement exponential backoff with jitter on 429 responses, queue non-interactive requests, and reserve a dedicated rate-limit tier or header budget for interactive calls.
Rate-limit pressure is best handled by smoothing retries and separating traffic classes. Exponential backoff with jitter avoids retry storms, queueing non-interactive work frees capacity during spikes, and a reserved budget for interactive calls keeps latency predictable. This lets the assistant degrade gracefully by delaying background work instead of failing user-facing requests.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cache every prompt-response pair indefinitely in Redis and serve cached answers whenever the API returns a 429.
Why it's wrong here
Serving a cached answer for a different prompt produces incorrect or misleading responses and is unsafe for code suggestions. Caching also does not help the first request for a novel prompt, which is exactly when the user is waiting. This approach trades correctness for availability in a way that damages trust in the assistant.
- ✗
Increase max_tokens on every request so answers are never truncated, and retry failed requests immediately in a tight loop.
Why it's wrong here
Raising max_tokens increases per-request cost and latency and does nothing to reduce 429 errors caused by rate limits. Immediate retries in a tight loop amplify the load and can extend the throttling window. This combination worsens both the failure rate and the interactive latency the team is trying to protect.
- ✗
Switch every request to a smaller, faster model and remove retry logic to reduce the number of API calls.
Why it's wrong here
Removing retries means transient 429s become permanent user-visible failures, which is the opposite of graceful degradation. Downgrading every request to a smaller model sacrifices answer quality for interactive and non-interactive work alike. This does not address the underlying contention for rate-limit capacity during spikes.
- ✓
Implement exponential backoff with jitter on 429 responses, queue non-interactive requests, and reserve a dedicated rate-limit tier or header budget for interactive calls.
Why this is correct
Exponential backoff with jitter prevents synchronized retry storms, while queueing non-interactive work preserves capacity for interactive calls. Reserving a separate rate-limit budget for interactive traffic ensures users get predictable latency even when batch jobs are running. Together these mechanisms let the assistant degrade gracefully instead of failing outright during usage spikes.
About these practice questions
One of 262 original CCAR-P practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.