CCAR-F Agentic Architecture and Orchestration Practice Question
An architect is optimizing a multi-agent system for a legal firm. The system uses a 'Router' agent followed by several 'Specialist' agents. Which THREE techniques will most effectively minimize latency in this specific architecture?
⚠ Common exam trap
Candidates often focus on model intelligence for all steps, ignoring that routing should be fast and that prompt caching is essential for minimizing latency in multi-agent systems.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement prompt caching for the specialist agents' long system prompts
Latency is a primary concern in multi-agent systems due to the sequential nature of LLM calls. Prompt caching significantly reduces time-to-first-token for large system instructions, while selecting faster models like Claude 3.5 Haiku for routing tasks and executing independent specialist tasks in parallel ensures the overall workflow remains responsive and efficient for the end-user.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Implement prompt caching for the specialist agents' long system prompts
Why this is correct
Specialist agents often require extensive context and detailed instructions that remain static across many requests. By using prompt caching, the architect ensures that these large prompts are not re-processed for every turn, drastically reducing the latency and cost of the specialist phase in a multi-agent conversation.
- ✓
Use Claude 3.5 Haiku for the initial routing decision
Why this is correct
The routing task is often a classification problem that does not require the high reasoning power of Sonnet or Opus. Using Claude 3.5 Haiku, which is optimized for speed, allows the system to identify the correct specialist almost instantly, reducing the overhead of the orchestration layer before the main work begins.
- ✓
Execute multiple specialist agents in parallel when their tasks are independent
Why this is correct
If the router identifies that a query requires input from multiple domains (e.g., contract law and tax law), executing those specialists simultaneously rather than sequentially reduces the total wall-clock time. This parallel execution is a core optimization for sophisticated agentic architectures that handle multifaceted user requests.
- ✗
Set 'max_tokens' to 1 for the specialist agents' responses
Why it's wrong here
Setting 'max_tokens' to 1 would prevent the specialists from providing any meaningful legal analysis or output. While it would technically reduce latency, it renders the system useless. Optimization must balance performance with the functional requirements of the task, and specialist agents need sufficient token limits to generate high-quality results.
- ✗
Force the model to use the 'computer_20241022' beta tool for all lookups
Why it's wrong here
The 'computer_20241022' tool is specifically for computer use (GUI interaction) and is much slower than standard API tool calls or text-based lookups. Using it for all lookups would significantly increase latency and introduce unnecessary complexity, as it is designed for tasks requiring visual navigation rather than data retrieval.
About these practice questions
This CCAR-F question is part of Courseiva's 271-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-F exam.