CCAR-F Agentic Architecture and Orchestration Practice Question
An architect is designing a research agent whose context window is filling with raw web page HTML. The agent must keep working across dozens of sources without exceeding the context limit. Which architecture best preserves reasoning quality over a long session?
⚠ Common exam trap
The trap here is treating context limits as a tuning problem solved by temperature or max_tokens, rather than as an architecture problem solved by sub-agents and summarization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run a sub-agent that reads each page and returns a structured summary, and keep only those summaries in the main agent's context.
Context management in agentic systems is an architectural concern, not a parameter tweak. Sub-agents act as context firewalls: they consume noisy inputs in their own window and return only distilled results to the orchestrator. This keeps the primary reasoning thread focused on high-signal tokens and enables the session to scale across many sources without truncation or attention dilution.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Store full page text in the conversation and rely on the model's long-context ability to attend to everything.
Why it's wrong here
Dumping raw HTML into the conversation rapidly exhausts the context window and degrades attention on the most relevant passages. Long-context attention is not free; irrelevant markup competes with signal. Without compaction, summarization, or retrieval, the agent will hit the limit mid-task and lose earlier findings, harming reasoning quality across dozens of sources.
- ✗
Increase the `max_tokens` output parameter so Claude can compress the HTML itself in one pass.
Why it's wrong here
max_tokens caps generated output, not input size. Raising it does not shrink the accumulated HTML already in the transcript and can even increase cost by allowing longer generations. Compression must be architectural, through summarization, retrieval, or sub-agents, not through a generation-length parameter.
- ✓
Run a sub-agent that reads each page and returns a structured summary, and keep only those summaries in the main agent's context.
Why this is correct
Delegating page reading to a sub-agent with its own context window keeps the main thread lean. The sub-agent absorbs the raw HTML, extracts structured findings, and returns a compact result. The orchestrator retains only distilled knowledge, preserving attention on high-value tokens and enabling the session to span far more sources than a flat transcript would allow.
- ✗
Set the model's temperature to zero so responses become shorter and consume fewer tokens.
Why it's wrong here
Temperature controls sampling randomness, not output length or context consumption. A deterministic setting does not shrink HTML inputs or prior turns, so the context window still fills. Length is governed by max_tokens and by how much content the orchestrator chooses to retain, neither of which temperature touches.
About these practice questions
This CCAR-F question is part of Courseiva's 271-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-F exam.