Courseiva

CCAR-P Practice Question: Developer Productivity and Operational Enablement

A team runs a Claude-based document summarization service. They notice that during peak hours, API requests occasionally fail with rate limit errors, causing user-visible failures. They want to improve reliability without over-provisioning. Which strategy is most effective?

⚠ Common exam trap

The trap here is thinking that changing model parameters or caching alone can solve rate limiting, when the core need is to manage request flow and retries.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implement exponential backoff with jitter and retry on rate limit errors, and use a queue to smooth bursts.

Exponential backoff with jitter and a queue are the most robust ways to handle rate limits. Backoff with jitter avoids synchronized retries, and a queue absorbs spikes so the API is called at a sustainable rate. This improves reliability without over-provisioning, and it is a best practice for integrating with Anthropic's API.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the max_tokens parameter for each request to get more done per call.

    Why it's wrong here

    Increasing max_tokens does not reduce the number of requests and may actually increase token consumption, making rate limits worse. Rate limits are typically based on requests per minute and tokens per minute; larger responses consume more tokens, potentially hitting token limits faster. This approach does not address the root cause of rate limit errors and can degrade performance.

  • ✗

    Switch to a smaller, faster model to reduce latency and avoid rate limits.

    Why it's wrong here

    A smaller model may have different rate limits, but it does not eliminate the need to handle rate limit errors. It could also reduce output quality, which may be unacceptable for summarization. The scenario asks for reliability without over-provisioning; changing models is a trade-off that doesn't inherently solve rate limiting and may introduce new issues.

  • ✗

    Cache all responses indefinitely so repeated requests never hit the API.

    Why it's wrong here

    Indefinite caching is impractical for dynamic documents and can serve stale content. It also does not help with unique requests during peak hours. While caching can reduce load, it is not a complete solution for rate limit errors and may violate data freshness requirements. The scenario requires handling bursts, which caching alone cannot guarantee.

  • ✓

    Implement exponential backoff with jitter and retry on rate limit errors, and use a queue to smooth bursts.

    Why this is correct

    Exponential backoff with jitter prevents thundering herd problems by spreading retries, while a queue decouples request spikes from API consumption. This combination handles transient rate limits gracefully and maintains throughput within limits. It is a standard reliability pattern that avoids over-provisioning and directly addresses peak-hour failures.

About these practice questions

One of 262 original CCAR-P practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.