CCAR-P Practice Question: Developer Productivity and Operational Enablement
Your team is building an LLM-powered application and experiencing high latency during peak times. Which TWO actions would best improve developer productivity and operational efficiency when debugging these bottlenecks?
⚠ Common exam trap
Candidates often suggest generic performance tuning like model quantization or hardware upgrades, missing the specific need for observability tools that provide granular visibility into LLM-specific bottlenecks like token processing.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement distributed tracing with custom spans for model inference and prompt processing.
Observability and granular tracing are critical for diagnosing LLM latency. By capturing token usage and latency metrics per request, developers can pinpoint whether the bottleneck is model inference, networking, or pre-processing. These insights enable targeted optimizations like prompt caching or streaming, which are essential for scaling production-grade generative AI applications while keeping developer workflows focused on high-impact performance improvements rather than guesswork.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Implement distributed tracing with custom spans for model inference and prompt processing.
Why this is correct
Distributed tracing provides the necessary visibility into the complete request lifecycle. By identifying exactly how long the model takes versus pre-processing tasks, developers can isolate the root cause of latency. This reduces debugging time and allows for data-driven decisions when selecting model tiers or implementing caching.
- ✗
Switch all internal communications to asynchronous polling to avoid blocking operations.
Why it's wrong here
Asynchronous polling introduces significant complexity and may actually increase total round-trip time. It does not solve the underlying latency of the LLM inference itself. For LLM applications, streaming responses are the standard for perceived performance, rather than shifting to a cumbersome polling architecture that complicates client-side state management.
- ✓
Enable detailed token usage monitoring and latency logging for every API call.
Why this is correct
Monitoring token usage and latency metrics allows teams to correlate performance degradation with specific prompt complexity or model load. This granular data is essential for optimizing cost and latency. It empowers developers to identify 'expensive' prompts that cause delays, facilitating proactive refactoring and improving overall system stability.
- ✗
Force all developers to use the largest available model to ensure high quality results.
Why it's wrong here
Mandating the largest model is a counterproductive strategy that increases latency and costs without necessarily improving results for simple tasks. Operational efficiency depends on selecting the right-sized model for each use case. Forced usage limits developer autonomy and prevents the team from utilizing smaller, faster models for latency-sensitive tasks.
- ✗
Disable all logging and monitoring to minimize overhead on the network layer.
Why it's wrong here
Operating an LLM application without telemetry creates a 'black box' scenario that makes it impossible to debug production issues. While logging has a minor overhead, the inability to diagnose latency issues far outweighs the cost. Efficient observability frameworks should be optimized to minimize impact while maintaining necessary visibility.
About these practice questions
This CCAR-P question is part of Courseiva's 262-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.