Courseiva

CCNA Advanced Agentic Architecture Questions

62 questions · Advanced Agentic Architecture · All types, answers revealed

1
Multi-Selecthard

Which TWO security measures are most effective at preventing 'Prompt Injection' attacks that target the arguments of tools used by an agent?

Select 2 answers
A.Adding 'Ignore all previous instructions' to the system prompt
B.Strict JSON schema validation for all tool inputs
C.Filtering the assistant's output for the word 'password'
D.Executing tools in a sandboxed, ephemeral environment
E.Using a secondary model to re-write every user query
AnswersB, D

By enforcing a rigid schema, the orchestrator ensures that the arguments passed to the tool conform to expected types and formats. This prevents attackers from injecting malicious scripts or unexpected commands into fields that should only contain simple data, such as numeric IDs or pre-defined string constants.

Why this answer

Prompt injection in tool arguments occurs when an agent processes untrusted data and passes it into a tool call that executes logic. Validating inputs against a strict schema and running the tool in an isolated environment are the primary defenses, ensuring that even if the agent is misled, the impact is contained.

Exam trap

Candidates often rely on 'system prompt instructions' to tell the model not to be hacked. This is ineffective against sophisticated prompt injection that bypasses textual constraints to manipulate tool arguments.

2
MCQmedium

What is the primary benefit of using a 'System Prompt' to define an agent's persona and constraints compared to embedding these in the user message?

A.System prompts are cached globally, reducing latency to zero.
B.User messages can be easily ignored, while system prompts cannot.
C.It establishes a distinct boundary between the agent's identity and user intent.
D.It is the only way to enable tool-use capabilities.
AnswerC

Separating identity and constraints from user inputs creates a robust architecture. This ensures that the agent's core instructions remain constant regardless of the user's conversational turns. This separation is fundamental to maintaining agent integrity and prevents users from coercing the agent into behaviors that violate defined policies.

Why this answer

System prompts are treated with higher priority by the model and are less susceptible to 'jailbreaking' or prompt injection from user inputs. By defining the agent's core identity and behavioral constraints in the system message, you create a stable, authoritative boundary that defines the model's behavior, ensuring consistent adherence to safety and operational guidelines throughout a long-lived conversation.

Exam trap

Candidates often assume the benefit is purely about token savings or context length, missing the critical security and authoritative boundary benefits of system prompts over user messages.

3
MCQhard

When designing an agent that must perform highly sensitive operations (e.g., deleting records), what architectural pattern is mandatory?

A.Fine-tune the model to never delete records.
B.Implement a mandatory human-in-the-loop confirmation step.
C.Use a low-temperature setting for sensitive tool calls.
D.Log the action to a secure bucket after execution.
AnswerB

The HITL pattern forces an external, verifiable authorization layer into the workflow. This ensures that no destructive or sensitive action can proceed without an explicit, audit-trailed human approval, which is the gold standard for secure agentic systems dealing with sensitive backend operations and data integrity.

Why this answer

A 'Human-in-the-loop' (HITL) gate for sensitive actions is non-negotiable for enterprise safety. The agent generates the plan and the proposed tool call, but the system architecture must intercept this and require explicit human verification before execution. This prevents accidental data loss or malicious exploitation of the agent's capabilities, balancing autonomous productivity with critical risk management in production environments.

Exam trap

Candidates often suggest relying on model-level security or fine-tuning to prevent errors, failing to recognize that human-in-the-loop is the only mandatory safeguard for high-stakes, irreversible actions.

4
MCQeasy

A team is building an agent that must call a payment provider's API. The API requires an idempotency key on every charge request, and the agent may retry a charge if a network error occurs mid-call. Which design correctly preserves financial correctness when retries happen?

A.Generate a fresh idempotency key for each retry attempt so the provider can distinguish a retry from a genuinely new charge.
B.Disable automatic retries for charge calls and surface every network error to a human operator for manual reconciliation.
C.Omit the idempotency key and instead check the account balance before each retry to see whether the charge already landed.
D.Derive the idempotency key deterministically from the logical charge (for example, a hash of order ID and amount) and reuse that same key across all retries of that charge.
AnswerD

A deterministic key ties every retry to the same logical charge, so the provider deduplicates repeats and returns the original result. Hashing stable fields such as order ID and amount means the key survives process restarts and agent replanning. This is the standard way to make an at-least-once retry loop behave as exactly-once at the provider boundary.

Why this answer

The idempotency key is the provider's contract for recognizing repeated attempts as one logical charge. Deriving it deterministically from stable charge attributes means every retry, even after a process restart or agent replan, presents the same key, so the provider returns the original outcome instead of charging again. This preserves correctness under at-least-once retry semantics without sacrificing resilience.

Exam trap

The trap here is assuming each retry needs a unique key to be distinguishable, when uniqueness per attempt is precisely what causes duplicate charges and defeats the idempotency contract.

5
MCQmedium

Which architectural approach is best for handling an agent's failure to retrieve information from a database tool?

A.Simply return the raw error message to the end user.
B.Configure the agent to automatically retry the exact same query 5 times.
C.Provide the error back to the agent as a new observation for re-planning.
D.Default to a hard-coded fallback value to satisfy the user.
AnswerC

Treating an error as an observation is a key principle of agentic architecture. By letting the agent 'see' what went wrong, it can apply its reasoning to correct the error, such as by broadening a query scope or sanitizing an input, leading to much higher success rates.

Why this answer

An 'error-handling wrapper' that catches the tool failure and feeds the error back into the agent's reasoning loop is essential. This allows the model to perform a 'self-correction' phase—perhaps by reformulating the query, searching a different table, or asking the user for clarification. This turns an error into a new data point, making the system significantly more robust and helpful than just reporting a failure.

Exam trap

Candidates often suggest showing the error to the user or simply retrying the same query, failing to leverage the model's ability to self-correct based on feedback.

6
MCQhard

An agent uses a retrieval tool that returns the top 20 chunks for any query. In production, the agent frequently cites irrelevant chunks and sometimes misses the correct answer even when it is present in the corpus. The corpus contains documents with overlapping terminology. Which architectural change most improves answer grounding without increasing the number of retrieved chunks?

A.Embed the entire corpus into the system prompt so the model always has full context available.
B.Increase the retrieval count to 50 so the correct chunk is more likely to be included somewhere in the context.
C.Lower the model's temperature to zero so it sticks more closely to the retrieved text.
D.Add a re-ranking stage that scores the retrieved chunks against the query and passes only the highest-scoring subset to the model.
AnswerD

Re-ranking applies a more precise relevance model to the candidate set, so the chunks most likely to contain the answer are promoted and the noisy ones are dropped. This improves grounding without retrieving more chunks, directly addressing both irrelevant citations and missed answers. It preserves the retrieval tool's contract while adding a precision layer that overlapping terminology makes necessary.

Why this answer

The retrieval step is high-recall but low-precision, which is why irrelevant chunks are cited and correct ones are missed amid overlapping terminology. Adding a re-ranking stage that scores candidates against the query and passes only the top subset improves precision without retrieving more chunks, directly improving grounding and citation quality.

Exam trap

The trap here is assuming that more retrieved context or a lower temperature will fix grounding, when the actual defect is ranking precision within a fixed candidate set.

7
MCQhard

An agentic system is designed to run long-lived 'background' tasks that may take several days to complete. How should the architect manage the agent's state to ensure it can resume correctly after a system reboot?

A.Pass the entire history in every API call
B.Use a very long 'max_tokens' setting
C.Externalize state and history to a persistent store
D.Rely on the model's internal memory
AnswerC

By saving the messages array and any orchestrator-level metadata (like step progress or tool IDs) to a database, the system becomes resilient. After a reboot, the orchestrator can reload the state, identify the last completed action, and resume the loop without losing progress or context.

Why this answer

Long-lived agents cannot rely on in-memory state. To ensure durability, the entire conversation history and any relevant internal state (like current goal or pending sub-tasks) must be externalized to a persistent database. This allows any worker instance to reconstruct the agent's 'mind' and continue the task from the last successful turn.

Exam trap

Candidates often assume the agent's session memory is sufficient. They ignore the reality that long-lived background tasks require persistent storage to survive system restarts, crashes, or scaling events.

8
MCQmedium

A document-processing agent must extract structured fields from thousands of PDFs. Some PDFs are scanned images, some are native text, and some are encrypted. The architect wants one pipeline that routes each document to the appropriate extractor and reports per-document confidence. Which design best fits?

A.Pre-classify each PDF by inspecting its structure, route native-text documents to a text extractor, scanned images to OCR, and encrypted files to a decryption step, then validate extracted fields and attach confidence scores.
B.Reject any PDF that is not native text and return an error, since scanned and encrypted documents are out of scope for automated processing.
C.Convert every PDF to images first, then run OCR on all of them, and skip classification because OCR handles any document.
D.Send every PDF to a single vision-language model and ask it to return the structured fields directly, ignoring document type.
AnswerA

Inspecting structure first lets each document take the cheapest and most accurate path: text extraction for native PDFs, OCR for scans, and a decryption stage for encrypted files. Validation and confidence scoring after extraction give the per-document reporting the architect wants. This matches processing to document type instead of forcing one method on all three, improving both cost and accuracy.

Why this answer

The three document classes need different handling, so the pipeline should inspect each file and route it: direct text extraction for native PDFs, OCR for scans, and a decryption stage for encrypted files. After extraction, validating fields and attaching confidence scores gives the per-document reporting the architect requires. This approach uses the cheapest viable method for each class and keeps confidence meaningful, rather than forcing one uniform method that is wrong for at least one class.

Exam trap

The trap here is treating PDF processing as a single uniform step, when routing by document structure is what makes extraction both accurate and cost-effective across mixed inputs.

9
Multi-Selectmedium

You are designing an agentic workflow that must reliably complete a multi-step refund process across an internal billing service and an external payment gateway. The workflow can fail at any step, and partial completion is unacceptable. Which two architectural mechanisms are required to guarantee that the workflow either completes fully or leaves no partial effect? (Choose two.)

Select 2 answers
A.A compensating action for each forward step, invoked in reverse order when the workflow cannot proceed to completion.
B.A semantic cache that stores prior successful refund transcripts and replays them on similar requests.
C.A higher model temperature on the planner so it can explore alternative step orderings when a step fails.
D.A durable execution log that records each step's intent and outcome so the orchestrator can resume or compensate after a crash.
E.A longer per-step timeout so that slow external calls are given more time to finish before the orchestrator gives up.
AnswersA, D

Because the two services cannot share a single transaction, atomicity must be simulated with compensation. Each forward step needs a defined inverse, such as voiding a charge when a subsequent refund fails. Invoking them in reverse order unwinds partial effects, which is what makes the workflow effectively all-or-nothing despite spanning independent systems.

Why this answer

Atomicity across independent services is achieved with a saga-style pattern: a durable log that records each step's intent and outcome, plus compensating actions that undo committed steps when the workflow cannot finish. The log enables correct recovery after a crash, and the compensations unwind partial effects in reverse order. Together they make the workflow effectively all-or-nothing even though no single transaction spans both systems.

Exam trap

The trap here is reaching for latency or model-tuning knobs such as timeouts and temperature, which change timing and variability but provide no mechanism for undoing a partially committed refund.

10
MCQmedium

You are architecting a customer-support agent for a SaaS platform. The agent must answer billing questions using a live invoice API and general policy questions using a static knowledge base. The invoice API is fast but occasionally returns stale data; the knowledge base is large and slow to search. Your design gives the agent a dedicated 'billing' sub-agent and a 'policy' sub-agent, coordinated by a router agent. After deployment, you observe the router sending nearly every query to both sub-agents in parallel, causing high latency. Which architectural change best addresses this while preserving answer quality?

A.Merge both sub-agents into a single agent with access to both the invoice API tool and the knowledge base search tool, and let it choose tools freely.
B.Have the router emit a structured routing decision with a confidence score, and only fan out to both sub-agents when the confidence falls below a configured threshold.
C.Cache the router's routing decisions for 24 hours so repeated identical queries skip the router on subsequent requests.
D.Replace the router with a fixed decision tree that maps every query containing the word 'invoice' to the billing sub-agent and all other queries to the policy sub-agent.
AnswerB

This preserves the router's semantic judgment while adding a confidence gate. High-confidence queries go to a single sub-agent, cutting latency; genuinely ambiguous queries still fan out, preserving answer quality. The threshold is tunable per environment, giving you an explicit latency/accuracy dial rather than an all-or-nothing routing rule.

Why this answer

The router is being overly cautious and fanning out on every query, which multiplies latency. Adding an explicit confidence signal to the routing decision lets the orchestrator fan out only when the router is genuinely unsure. High-confidence routing keeps the common case single-hop, while the fallback preserves correctness on ambiguous queries.

This is the only option that changes the router's decision policy rather than replacing it with a brittle rule or an unrelated caching layer.

Exam trap

The trap here is treating 'router fans out too often' as a routing-logic bug to be replaced with deterministic rules, when the real fix is to expose the router's confidence so the orchestrator can gate fan-out.

11
Multi-Selecthard

An agentic system uses a supervisor agent that delegates to specialized worker agents. During a long incident, the supervisor's own context fills with worker transcripts, and it begins losing track of which workers have completed and which are still running. Which TWO architectural changes best preserve the supervisor's ability to coordinate correctly? (Choose two.)

Select 2 answers
A.Restart the supervisor with a fresh context whenever its window fills, and rely on workers to resend their current status on request.
B.Have the supervisor maintain an external task ledger that records each delegated task, its assigned worker, and its status, and read from that ledger instead of retaining full worker transcripts.
C.Allow workers to message each other directly to resolve dependencies, removing the supervisor from the coordination path entirely.
D.Increase the supervisor's context window to the largest available model so it can retain all worker transcripts for the duration of the incident.
E.Have each worker return a compact structured status envelope containing task id, state, and a short result summary, and have the supervisor track tasks by these envelopes.
AnswersB, E

An external ledger keeps authoritative state about what was delegated and what finished, so the supervisor can coordinate without carrying every worker transcript in context. It survives context growth and lets the supervisor query only the fields it needs. This directly addresses losing track of completed versus running workers while keeping the supervisor's context bounded.

Why this answer

The supervisor's difficulty comes from mixing coordination state with raw worker output. Externalizing task state into a ledger and having workers return compact structured envelopes keeps the authoritative record of what is delegated and what has finished separate from bulky transcripts. The supervisor can then coordinate from small, reliable signals and pull detail only when needed, rather than trying to hold the entire incident in its context window.

Exam trap

The trap here is believing that a bigger context window or a restart can substitute for externalizing task state, when both leave the supervisor's coordination record coupled to transcript volume.

12
Multi-Selectmedium

You are architecting a Claude agent that orchestrates a long-running approval workflow spanning hours or days, where human reviewers may intervene between steps. Which TWO mechanisms are necessary to keep the workflow correct across these interruptions? (Choose two.)

Select 2 answers
A.Set a very short max_tokens on every turn to reduce the chance of context drift across days.
B.Persist the full message history and tool state to durable storage after each turn so the workflow can resume from the exact checkpoint.
C.Disable tool use entirely during human review windows and rely on the reviewer to re-enter all prior context manually.
D.Implement idempotency keys on all side-effecting tool calls so retries after an interruption do not duplicate approvals or notifications.
E.Use a stateless design where each invocation rebuilds context by summarizing all prior decisions into the system prompt.
AnswersB, D

Durable checkpointing is essential because the process may be suspended or crash between human interventions. Saving message history and tool state lets the workflow resume mid-conversation without re-running side effects or losing prior reasoning. Without it, a restart would either duplicate actions or force the agent to begin again, breaking the approval chain.

Why this answer

Long-running workflows with human-in-the-loop interruptions require durable checkpointing to resume exact state and idempotency keys to make retries safe. Together they ensure that an interruption or crash neither loses the reasoning chain nor duplicates side effects such as approvals or notifications. Stateless summarization, token limits, and disabling tools all fail to provide either guarantee.

Exam trap

The trap here is believing that reconstructing context from summaries is equivalent to persisted state, when only durable checkpoints preserve pending tool calls and exact arguments.

13
MCQhard

A research agent uses Claude with an extended thinking budget to analyze a 200-page regulatory filing. Mid-analysis it must call a `fetch_footnote` tool whose result is essential to the conclusion. The architect wants the tool result to be incorporated without discarding the model's prior reasoning. Which approach best achieves this?

A.Store the footnote in an external vector store and instruct the agent to retrieve it again later if needed.
B.Append the tool result as a new user message and restart the analysis from the beginning of the filing.
C.Summarize the footnote with a separate cheap model and paste the summary into the system prompt before the next call.
D.Insert the tool result into the conversation as a tool_result block associated with the original tool_use, then continue the same turn so the model resumes from its existing reasoning.
AnswerD

Returning the result as a tool_result block tied to the preceding tool_use preserves the message sequence the model already reasoned over, letting it continue the same turn without re-deriving prior steps. This is the canonical way to feed external data into an in-progress reasoning chain. It keeps the thinking budget intact and respects the API's tool-use contract.

Why this answer

Tool results must be returned as tool_result blocks paired with the originating tool_use so the conversation remains valid and the model can continue from its existing reasoning rather than restart. This preserves the thinking budget and the message sequence. Restarting, summarizing through another model, or deferring to retrieval all either discard reasoning or break the API contract, so they fail this scenario.

Exam trap

The trap here is treating a tool result as ordinary text that can be injected anywhere, when the API requires it to be paired with its tool_use block to keep the reasoning chain valid.

14
MCQhard

An architect needs to build an agent that handles complex, multi-step data migrations where each step depends on the output of the previous one. Which approach is most robust for ensuring the agent doesn't lose track of the long-term goal during execution?

A.Reactive Tool Calling (ReAct)
B.Zero-shot prompting with all tools
C.Few-shot prompting with migration examples
D.Parallel Sub-tasking
E.Plan-and-Execute Pattern
AnswerE

This pattern involves an initial planning phase to decompose the migration into discrete steps, followed by an execution loop that references and updates the plan. This separation of concerns ensures the agent maintains a stable roadmap, allowing it to verify each dependency before proceeding to the next migration phase.

Why this answer

For multi-dependency tasks, a Plan-and-Execute pattern is superior because it separates the high-level strategy from the low-level tool interactions. By maintaining an explicit plan that is updated after each step, the agent can track its progress against the original goal and adjust its future actions based on the results of completed tasks.

Exam trap

Candidates often suggest a simple ReAct or sequential chain. They fail to realize that complex, multi-dependency tasks require an explicit planning phase to prevent the agent from drifting off-target mid-execution.

15
MCQhard

A production agent uses a ReAct loop and frequently reaches its maximum step budget while still mid-task, then returns a partial answer that looks complete. Telemetry shows the agent often re-reads the same file and re-queries the same database row across consecutive steps. Which change most directly reduces wasted steps while preserving the agent's ability to finish?

A.Raise the maximum step budget from 15 to 60 so the agent always has enough room to finish any task.
B.Replace the ReAct loop with a single-shot prompt that asks the model to plan every step up front and then answer without further tool calls.
C.Add a step-level deduplication cache that returns the prior result when the agent issues an identical tool call with identical arguments, and inject a reminder of remaining budget into the loop.
D.Lower the maximum step budget to 8 so the agent is forced to be more decisive and avoid redundant work.
AnswerC

Caching identical tool calls with identical arguments removes the repeated file reads and database queries that consume steps, directly attacking the observed waste. Surfacing remaining budget lets the model prioritize finishing over redundant exploration. Together they cut wasted steps while leaving the agent free to complete the task, rather than artificially capping its work or raising the ceiling without improving efficiency.

Why this answer

The wasted steps come from repeating identical tool calls, so the most direct fix is to detect and short-circuit duplicates by caching results keyed on the call and its arguments. Pairing that with visibility into remaining budget helps the model spend its steps on novel actions and wrap up cleanly. Adjusting the budget up or down changes when the agent stops but not how much of its work is redundant, and collapsing the loop removes the observation-driven adaptation the agent depends on.

Exam trap

The trap here is treating the step budget as the problem, when the budget is only the boundary that exposes redundant tool calls as the real inefficiency.

16
Multi-Selectmedium

When evaluating the performance of a new 'Orchestrator-Worker' agentic architecture, which THREE metrics provide the most insight into the system's efficiency and reliability?

Select 3 answers
A.Task Success Rate (TSR)
B.Total number of tools defined in the schema
C.Average Steps per Task
D.The model's pre-training data cutoff date
E.Tool Call Accuracy Rate
AnswersA, C, E

TSR is the primary indicator of whether the agent is actually fulfilling its intended purpose. It measures the percentage of sessions where the agent reaches a correct and verified conclusion, providing a high-level view of the architecture's effectiveness across a diverse set of test cases or user queries.

Why this answer

Evaluating agents requires looking beyond simple accuracy to understand the operational characteristics of the loop. Task success rate measures the ultimate goal, while steps-per-task and tool accuracy pinpoint where the logic might be breaking down or becoming inefficient. These metrics collectively guide the architect in refining the agent's reasoning path.

Exam trap

Candidates often select business metrics like 'User Satisfaction' or 'Revenue Impact'. These are lagging indicators that do not provide the granular technical feedback necessary to debug agentic reasoning loops or tool usage.

17
MCQmedium

You are building a customer-support agent with Claude that must call a lookup_order tool, then a refund_order tool that depends on the order's status. During testing, the agent sometimes calls refund_order before the lookup_order result returns. Which change to your orchestration loop best enforces the required ordering?

A.Increase the max_tokens value so the model has more room to reason about the correct order before calling tools.
B.Expose only lookup_order until its result is returned, then add refund_order to the tools array for the subsequent request.
C.Include both tools in the same request's tools array and let the model choose the call order.
D.Set tool_choice to force a tool call on every turn so the agent never stalls between lookup and refund.
AnswerB

By withholding refund_order from the tools array until the lookup_order result is present in the conversation, the orchestration layer makes the premature call impossible rather than merely discouraged. The model cannot invoke a tool it was not offered. This state-gated tool exposure enforces the dependency deterministically while keeping the loop simple and auditable.

Why this answer

Deterministic ordering is achieved by controlling what the model can do at each step, not by prompting harder. Withholding the dependent tool until the prerequisite result is in the conversation removes the failure mode entirely, because the model cannot select a tool absent from the current tools array. This makes the dependency a property of the orchestration state machine rather than of model compliance.

Exam trap

The trap here is assuming that listing tools together or forcing tool use will make the model respect dependencies, when only gating tool availability on prior results can guarantee ordering.

18
MCQmedium

An agentic system often experiences 'goal drift' when managing long-running, multi-step tasks. Which architectural pattern most effectively mitigates this risk during recursive reasoning chains?

A.Increasing the context window size to include all previous turns.
B.Implementing a hard-coded decision tree for every possible action.
C.Integrating a reflection loop that evaluates progress against the initial goal.
D.Reducing the temperature parameter to zero for all model calls.
AnswerC

Reflection loops force the model to pause and assess the current state against its target objective. This meta-cognitive step allows the agent to identify deviations, prune irrelevant reasoning chains, and re-orient its strategy. It is essential for long-horizon task completion where errors naturally accumulate without periodic corrective oversight.

Why this answer

State-space re-grounding via periodic reflection loops allows the agent to compare current progress against original user intent. By forcing a dedicated 'evaluator' step that analyzes the chain of thought against the task definition, the system can self-correct before executing irreversible actions. This prevents the agent from spiraling into irrelevant sub-tasks that deviate from the primary objective, ensuring high task fidelity in complex autonomous workflows.

Exam trap

Candidates often suggest increasing the model's context window or using a more powerful model, ignoring that architectural patterns like reflection loops are required to fix logical drifting in multi-step chains.

19
MCQmedium

An agent must summarize a 400-page contract that exceeds the context window. The team wants a summary that references specific clauses without losing cross-references between distant sections. Which approach best preserves cross-reference integrity?

A.Map each section to a structured node with its clause identifiers and cross-references, summarize nodes, then synthesize using the reference graph.
B.Increase max_tokens on each summarization call so the model can produce longer per-chunk summaries.
C.Use a smaller model with a longer context window to process the whole contract in a single pass.
D.Split the contract into fixed-size chunks, summarize each independently, and concatenate the summaries in order.
AnswerA

Building an explicit reference graph lets the synthesis step follow links between distant sections, so a summary of one clause can resolve references to another. Structured nodes preserve clause identifiers as anchors, and the graph makes cross-references navigable rather than lost. This directly targets the requirement by carrying relational structure through summarization instead of discarding it at chunk boundaries.

Why this answer

Cross-reference integrity survives summarization only if the relational structure is modeled explicitly. Converting sections into nodes with clause identifiers and reference edges lets the synthesis stage traverse those links, resolving 'see Section X' pointers instead of severing them at chunk boundaries. This preserves meaning that flat chunking or larger output budgets cannot recover.

Exam trap

The trap here is equating a bigger context window or longer summaries with preserving relationships, when the real issue is that naive chunking destroys the reference graph between distant clauses.

20
MCQeasy

An architect is designing an agent to perform administrative tasks in a corporate environment. One of the tools allows the agent to delete employee records. What is the most important architectural safeguard to implement for this specific tool?

A.Strict regex validation on the employee ID
B.Using a smaller, more focused model for deletion
C.Increasing the temperature to allow for more flexibility
D.Human-in-the-Loop (HITL) approval step
AnswerD

Requiring a human to review and approve the agent's intent before executing a destructive action is the gold standard for agentic security. This ensures that a person validates the context and correctness of the operation, effectively mitigating the risks associated with autonomous decision-making in sensitive production systems.

Why this answer

In high-stakes agentic environments, Human-in-the-Loop (HITL) is the primary safety mechanism for irreversible actions. While automated checks are helpful, human oversight ensures that the agent's reasoning aligns with organizational policy and prevents catastrophic data loss from potential hallucinations or misinterpretations of user intent in sensitive contexts.

Exam trap

Candidates often choose automated validation or logging as the primary safeguard. They underestimate the risk of model hallucination in high-stakes actions, assuming that code-level checks are sufficient to prevent irreversible business data loss.

21
MCQhard

Your agent orchestrates a research task by spawning several subagents, each with its own Claude conversation, and merging their outputs. You observe that the final synthesis contradicts the subagents' findings. Which architectural change most reliably preserves fidelity when merging?

A.Pass each subagent's structured findings, including source citations and confidence, into the synthesizer as explicit quoted evidence.
B.Raise the synthesizer subagent's temperature so it can creatively reconcile conflicting viewpoints.
C.Have the synthesizer re-run each subagent's task itself to verify the results before merging.
D.Give all subagents access to a shared mutable scratchpad so they can overwrite each other's conclusions during execution.
AnswerA

When the synthesizer receives the subagents' outputs as attributed evidence rather than paraphrase, it can trace each claim to its origin and detect genuine conflicts instead of inventing a blended narrative. Structured findings with citations preserve provenance through the merge, letting the synthesizer reconcile or flag disagreements explicitly rather than silently overwriting them, which is what caused the contradiction.

Why this answer

Fidelity through a merge depends on provenance. When the synthesizer receives each subagent's findings as attributable, cited evidence, it can align claims to sources, surface real disagreements, and avoid fabricating a consensus. Structured handoff preserves the trail from subagent output to final answer, which is precisely what a paraphrased or mutable handoff destroys.

Exam trap

The trap here is treating synthesis as a creative blending step, when the contradiction stems from lost provenance and is fixed by passing attributed evidence rather than by re-running or loosening generation.

22
MCQmedium

A Claude agent orchestrates three specialized subagents: one for data extraction, one for validation, and one for report generation. In production, the validation subagent sometimes receives malformed input from the extraction subagent and silently produces a passing result. You need the orchestrator to detect and contain these failures without halting the entire pipeline. Which design change is most appropriate?

A.Increase the orchestration layer's retry count so that any subagent failure is retried several times before the pipeline reports an error to the caller.
B.Merge the extraction and validation subagents into a single agent so that malformed intermediate data never crosses a component boundary.
C.Instruct the validation subagent to always return a confidence score between zero and one, and have the orchestrator discard results below a fixed threshold.
D.Have each subagent return a structured result envelope containing a status, a schema-validated payload, and an error field, and make the orchestrator validate the envelope before routing onward.
AnswerD

A structured envelope with an explicit status and validated payload gives the orchestrator a deterministic contract to check, so malformed extraction output is caught before validation runs. Because the status is machine-readable, the orchestrator can retry, reroute, or quarantine just the failing branch instead of stopping the pipeline. This contains the fault at the boundary where it occurs.

Why this answer

Reliable multi-agent pipelines depend on explicit contracts at every handoff. A structured envelope with status, validated payload, and error field lets the orchestrator verify upstream output before trusting it downstream. That check enables targeted containment, such as retrying or quarantining one branch, rather than trusting a failing agent's self-assessment or retrying blindly.

Exam trap

The trap here is trusting a failing subagent to accurately report its own health through confidence scores, instead of validating the data contract at the boundary.

23
MCQeasy

A support agent built on Claude must decide whether a customer request is a billing issue, a technical issue, or a general inquiry before routing it. The categories are fixed and mutually exclusive, and the routing decision must be fast and cheap. Which approach is most appropriate?

A.Use a single Claude call with a constrained output format such as a tool definition or structured JSON that returns exactly one category label.
B.Fine-tune a separate model on historical tickets and deploy it alongside the Claude agent for classification.
C.Give the agent access to all internal tools and let it infer the category from whichever tool it chooses to call.
D.Build a multi-agent pipeline where a classifier agent debates a router agent until they agree on a category.
AnswerA

A fixed, mutually exclusive classification with speed and cost constraints is a natural fit for a single constrained call. Defining a tool or JSON schema with an enum of the three categories forces the model to emit one valid label. This avoids multi-step orchestration overhead and keeps latency and token cost low, which matches the stated requirements exactly.

Why this answer

Fixed, mutually exclusive categories with strict latency and cost goals are best handled by a single Claude call that constrains output to one of the allowed labels. Using a tool definition or JSON schema with an enum removes ambiguity and avoids the overhead of multi-agent orchestration or a separate fine-tuned model.

Exam trap

The trap here is equating higher architectural complexity with higher accuracy, when a constrained single call already guarantees a valid label for a fixed taxonomy.

24
MCQhard

You are building a system that requires strict adherence to a specific output format. Which approach provides the highest reliability in a high-traffic agentic environment?

A.Include a long prompt instruction asking the model to be very careful with formatting.
B.Use a post-generation regex parser to fix up malformed JSON responses.
C.Implement structured schema enforcement at the model interface layer.
D.Request the model to output the answer in XML because it is easier for agents.
AnswerC

Enforcing schema at the interface layer guarantees the output structure before the agent even completes its generation. This eliminates the need for expensive or unreliable post-processing and ensures that every response is ready for immediate consumption by downstream APIs, making the overall system significantly more robust and scalable.

Why this answer

Forcing structured output through system-level schema enforcement at the model interface is the only reliable way to guarantee format consistency in production. By utilizing constrained decoding or rigorous post-processing validation, you ensure that the agent consistently produces machine-readable formats like JSON. This architectural pattern is vital because it removes the variability of natural language, allowing downstream systems to process agent outputs without frequent parsing errors or exceptions.

Exam trap

Candidates frequently rely on prompt engineering or few-shot examples to enforce output formats, forgetting that these methods are probabilistic and fail frequently under high traffic compared to interface-level schema enforcement.

25
MCQmedium

When designing agentic systems that require human-in-the-loop (HITL) verification, which mechanism prevents the agent from stalling indefinitely while waiting for user input?

A.Implement a web-socket connection for instant notification.
B.Design a timeout handler that triggers an escalation flow.
C.Require the user to acknowledge receipt before the agent proceeds.
D.Use a higher model temperature to guess the user's intent.
AnswerB

A timeout handler provides a deterministic end-state for a waiting process. It allows the agent to break out of the 'wait' state if no human feedback arrives, ensuring the agent remains autonomous enough to take an alternative action or notify an administrator instead of stalling the entire workflow.

Why this answer

A timeout-driven fallback mechanism ensures the system retains agency even when a user is unavailable. By setting a predefined duration for waiting, the agent can trigger a 'default' or 'safe' path, such as escalating to a manager or pausing the task, rather than hanging. This maintains system uptime and keeps the workflow moving forward, which is critical for complex, real-world agent deployments where human response times are highly variable.

Exam trap

Candidates often rely on infinite asynchronous waiting states or manual user resets, failing to implement automated timeout handlers that proactively manage stalled human-in-the-loop workflows.

26
Multi-Selecthard

A claims-processing agent runs for hours and must survive process restarts without losing in-flight work. You are designing durable execution around Claude's stateless Messages API. Which TWO practices are required to make the agent resumable? (Choose two.)

Select 2 answers
A.Persist the full message history and pending tool_use ids to durable storage after each turn so the loop can be reconstructed.
B.Rely on the model's server-side session state to remember prior turns across restarts.
C.Cache every model response in a CDN so identical prompts return instantly after a restart.
D.Store a checkpoint that records which tool calls were dispatched but whose results were not yet appended.
E.Increase the context window by enabling extended thinking so the agent retains more history internally.
AnswersA, D

The Messages API is stateless, so the entire conversation including assistant tool_use blocks and their matching tool_result blocks must be stored externally. On restart, replaying this history lets the loop continue exactly where it stopped. Without persisting the tool_use ids, you cannot correctly pair results and the API will reject or misinterpret the reconstructed conversation.

Why this answer

Because the Messages API is stateless, resumability is entirely the orchestrator's responsibility. You must persist the conversation, including tool_use and tool_result pairings, and checkpoint in-flight tool dispatches so a restart does not replay side effects. Together these produce a replayable log that reconstructs the loop and reconciles uncertain operations, turning crash recovery into a deterministic continuation.

Exam trap

The trap here is assuming the API keeps conversation state or that caching responses provides durability, when only externally persisted history plus in-flight checkpoints enables safe resumption.

27
MCQmedium

When designing agents that interact with external APIs, which pattern best addresses the challenge of 'unreliable API latency' impacting the agent's reasoning chain?

A.Set the agent's request timeout to 60 seconds to ensure it eventually finishes.
B.Use an asynchronous 'request-poll' pattern to decouple execution from reasoning.
C.Hardcode a retry count of ten to force the API to respond faster.
D.Instruct the model to wait for a response in its thinking process.
AnswerB

The request-poll pattern allows the agent to initiate an action and then focus on other tasks or enter a non-blocking wait. The agent can then poll for results, ensuring that the reasoning engine is not blocked by slow network responses, resulting in a more resilient and performant architecture.

Why this answer

To manage unpredictable API latency, you must decouple the agent's reasoning from the API execution through an asynchronous task queue. By allowing the agent to request an action and then continue other tasks or enter a wait state, you prevent the reasoning process from timing out or becoming blocked. This is critical for building responsive, reliable agentic systems that can handle real-world network instability and variable service performance.

Exam trap

Candidates often try to solve latency by increasing the model's timeout settings. This leads to poor user experience and hangs the agent's reasoning process while it waits for slow responses.

28
MCQmedium

Refer to the exhibit. Your agent is receiving this error during a high-concurrency operation. Which implementation correctly handles this scenario?

A.Discard the task and return a failure message to the user.
B.Immediately retry the request without any delay.
C.Wait for the duration specified in 'retry_after' before retrying.
D.Switch to a backup API key and continue immediately.
AnswerC

Respecting the 'retry_after' duration is the standard way to handle rate limiting. It aligns your agent's behavior with the server's requirements, ensuring that your retry attempts occur after the limit has reset. This is the most efficient and polite way to interact with rate-limited external APIs.

Why this answer

The correct implementation is to respect the 'retry_after' header and back off before attempting another call. This prevents the system from entering a crash loop where it hammers the API with failed requests. By implementing an exponential backoff strategy that honors the provided delay, you ensure the agent remains resilient to transient load issues while complying with the service provider's rate constraints, ultimately increasing the reliability of the overall system.

Exam trap

Candidates mistakenly implement aggressive immediate retries upon hitting rate limits, which exacerbates server congestion and leads to extended application downtime.

29
MCQmedium

A financial services firm is deploying an AI agent to handle diverse requests including balance inquiries, market analysis, and loan applications. To minimize latency and maximize precision, the architect decides to implement a pattern where a primary model classifies the intent and delegates to specialized sub-agents. Which architectural pattern is being described?

A.Monolithic Chain
B.Sequential Pipeline
C.Router Pattern
D.Parallel Execution
AnswerC

This pattern effectively directs the flow to the most relevant sub-module based on the initial input classification. By scoping the context to only what is needed for a specific intent, it improves accuracy in tool selection and reduces the noise that usually degrades performance in large-scale LLM applications.

Why this answer

Implementing a router architecture allows the system to isolate logic into specialized modules, which significantly reduces the cognitive load on the LLM. By dynamically selecting only the necessary tools for each specific query, the architect ensures that the model remains focused on relevant parameters, leading to higher precision, lower token consumption, and a more maintainable codebase over time.

Exam trap

Candidates often confuse the Router Pattern with a Chain-of-Thought or Agentic Loop, failing to recognize that delegating tasks based on intent classification is specifically the Router Pattern.

30
MCQhard

A Claude agent maintains a long conversation with a user over weeks. The architect notices that the agent gradually forgets early constraints the user stated, even though the conversation is well within the model's context window. The team wants to fix this without re-summarizing the entire history on every turn. Which approach is most appropriate?

A.Raise the temperature slightly so the model explores more of the conversation when generating.
B.Increase the model's context window by switching to a larger variant so all history fits with room to spare.
C.Move persistent user constraints into a structured memory store that is retrieved and injected as a system-level block on each turn.
D.Insert the early constraints again at the very end of every user message as a reminder.
AnswerC

Extracting durable constraints into a structured memory store and re-injecting them each turn keeps them salient regardless of how the conversation grows, without re-summarizing everything. It directly addresses gradual forgetting by giving constraints a stable, high-priority position. This is the scalable pattern for long-lived agents with persistent user preferences.

Why this answer

Persistent constraints belong in a durable memory store that is retrieved and injected as a system-level block each turn, giving them stable salience without repeatedly summarizing the whole history. Larger context windows, repeated reminders, and temperature changes do not fix attention dilution over long conversations, and two of them are actively counterproductive. Structured memory is the right architectural layer for durable user preferences.

Exam trap

The trap here is assuming that a larger context window solves forgetting, when the scenario already excludes that and the real issue is salience of persistent constraints.

31
MCQmedium

An orchestration agent runs a nightly workflow that fans out to 12 subagents, each calling a partner REST API. Partner calls intermittently return HTTP 429 with a Retry-After header. The orchestrator currently retries immediately in a tight loop, causing cascading 429s and duplicate side effects on the partner systems. Which architectural change best addresses both the throttling and the duplicate side effects?

A.Reduce the fan-out to a single subagent that processes all 12 partners sequentially with no retry logic at all.
B.Switch every subagent to a larger context window so each can hold the full history of prior 429 responses and learn to avoid them.
C.Introduce a shared token-bucket rate limiter in front of the partner calls plus idempotency keys on every mutating request, with subagents respecting Retry-After.
D.Increase the orchestrator's max_tokens so the model can reason longer about each 429 before deciding to retry.
AnswerC

A shared token bucket caps aggregate concurrency across all 12 subagents so bursts stop triggering 429s, while Retry-After compliance paces retries. Idempotency keys let the partner recognize a repeated mutating call and return the original result instead of creating a second record, which directly eliminates duplicate side effects. Together these address throttling and duplication at the architecture layer rather than relying on model behavior.

Why this answer

Bursty fan-out against a rate-limited partner needs two coordinated controls: a shared limiter so the aggregate request rate stays under the partner's ceiling, and idempotency keys so a retried mutating call cannot create a second side effect. Honoring Retry-After aligns retry timing with the partner's guidance. Token or context changes operate on model text, not on outbound traffic, and serializing the workload removes parallelism without addressing throttling or duplication.

Exam trap

The trap here is assuming that making the model 'smarter' about errors, via more tokens or more context, substitutes for enforcing rate limits and idempotency at the orchestration layer.

32
MCQmedium

A developer is concerned about the high token cost and latency of an agent that has access to 50 different tools. What is the most effective architectural change to optimize this system?

A.Dynamically inject tools based on intent classification
B.Switch from JSON to XML tool definitions
C.Compress the tool descriptions into single words
D.Force the model to use only one tool per turn
AnswerA

Using a lightweight classifier to identify the user's intent allows the system to provide only a small subset of the 50 tools. This reduces the input token count, lowers latency, and improves the model's accuracy by removing irrelevant tool schemas that could cause confusion or lead to hallucinations.

Why this answer

Model performance and cost are directly impacted by the size of the system prompt and tool definitions. By implementing a dynamic tool selection mechanism, the architect can ensure that only the tools relevant to the user's current intent are loaded into the context, significantly reducing the prompt overhead for every turn.

Exam trap

Candidates often suggest fine-tuning the model to handle more tools, overlooking that dynamic injection is the standard architectural approach to reduce context bloat and improve performance.

33
MCQmedium

When designing an agent capable of multi-step tool use, what is the most important property to maintain across steps?

A.The exact same temperature parameter for every step.
B.The model's internal memory of its own personality.
C.Synchronization between the agent's world model and the environment.
D.Using the same model provider for every single step.
AnswerC

The 'world model' is the agent's mental map of the current environment state. If this map deviates from reality, the agent will choose incorrect tools or pass wrong arguments. Maintaining strict synchronization ensures that the agent's next action is always based on the most accurate, real-world data available.

Why this answer

Ensuring state consistency is the most important property. Each tool call changes the environment, and the agent must understand the new state to plan the next step. If the state becomes inconsistent—e.g., the agent thinks a file was deleted but it wasn't—the agent's plan will fail.

Architecting for state observability ensures that the model always has an accurate 'world model' to work from.

Exam trap

Candidates often assume that simply tracking tool execution logs is sufficient, failing to realize that discrepancies between the agent's internal assumptions and the actual remote environment cause silent multi-step execution failures.

34
MCQmedium

When implementing a 'Human-in-the-loop' (HITL) checkpoint, what is the best way to handle the state persistence during the wait period?

A.Keep the connection open to avoid re-initializing the agent.
B.Persist the state in an external database and terminate the process.
C.Store the state only in the client-side browser memory.
D.Retry the tool call periodically until the human responds.
AnswerB

Externalizing the state allows for scalability and durability. By saving the session context in a database, the system can remain idle without consuming compute resources. Once the human provides input, the system can reload the state, reconstructing the agent's context to resume the task seamlessly and reliably.

Why this answer

Externalizing state to a persistent database is essential. Agentic systems are often stateless by design, but a wait period requires 'pausing' execution. Saving the entire conversation history, current tool state, and reasoning chain allows the process to resume exactly where it left off once human approval is received.

This prevents data loss and maintains context continuity, which is critical for long-running, complex workflows.

Exam trap

Test-takers often assume that keeping the process alive in memory or using local process threads is sufficient for waiting states, ignoring that agentic architectures require externalized database persistence.

35
MCQeasy

Which of the following describes the role of the 'System Prompt' in an agentic architecture?

A.A temporary buffer for user-provided instructions.
B.A mechanism for defining the agent's identity and constraints.
C.An automated script for executing Python code.
D.A method to store long-term user history.
AnswerB

The system prompt is the primary location for defining 'who' the agent is and the boundaries it must respect. This is an essential architectural component that dictates how the agent processes information, adheres to safety protocols, and interacts with tools, forming the foundation of its autonomous behavior.

Why this answer

The system prompt serves as the 'DNA' of the agent, dictating its core behavior, constraints, and operational goals. By setting the context outside of the user's conversation stream, it ensures that the model adheres to defined roles and ethical boundaries consistently, regardless of user input. It provides the necessary structure that turns a raw LLM into a purposeful, reliable assistant within a larger system architecture.

Exam trap

Candidates frequently confuse the system prompt with dynamic few-shot learning examples, mistakenly thinking it handles turn-by-turn conversation memory rather than persistent global boundaries and agent identity.

36
MCQhard

You are architecting a Claude-based agent for a regulated financial client that performs long-running portfolio rebalancing workflows. A compliance requirement mandates that no single trade instruction may be executed unless it is cryptographically traceable to the exact model-generated intent that produced it. The agent uses the Messages API with tool use, and several downstream services consume tool calls asynchronously. Which architectural mechanism best satisfies this requirement while preserving agent autonomy?

A.Require the model to emit a SHA-256 hash of its own reasoning text inside the tool input, and verify that hash in the downstream trade service before execution.
B.Route all trades through a single orchestrator agent that logs a natural-language summary of each decision to an append-only ledger after the trade is confirmed.
C.Enable prompt caching on the system prompt so that every trade instruction inherits a stable cache key that can be presented to auditors as proof of origin.
D.Persist every assistant turn containing a tool_use block, its tool_use id, and the matching tool_result, forming an immutable audit chain keyed by the tool_use id.
AnswerD

The tool_use id is generated by Claude and echoed back in the corresponding tool_result, so pairing them creates a verifiable link between model intent and downstream execution. Persisting the full assistant turn preserves the reasoning context around the intent, and the id becomes the cryptographic join key that compliance can replay. This satisfies traceability without constraining how many tools the agent may call.

Why this answer

Binding each tool call to the platform-generated tool_use id and persisting the surrounding assistant turn creates a deterministic chain from model intent to executed instruction. Because the id appears in both the tool_use block and the tool_result, downstream services and auditors can reconcile them without trusting model-authored text. Caching, summaries, or self-hashes do not establish that binding.

Exam trap

The trap here is assuming that any unique value the model or cache layer produces can serve as an audit key, when only the platform-issued tool_use id is reliably echoed through the tool_result.

37
MCQmedium

A logistics agent needs to fetch shipping rates from five different carriers simultaneously to find the best price. Which architecture is best suited for this requirement?

A.Chain of Thought (CoT) prompting
B.Orchestrator-Worker Pattern
C.Recursive Reflection
D.Single-Agent Sequential Loop
AnswerB

This pattern allows a lead agent to delegate independent tasks to multiple specialized workers. In this scenario, the orchestrator can trigger five carrier-specific tool calls in parallel, collect the results, and then synthesize the findings to determine the best price, optimizing both speed and logical separation.

Why this answer

The Orchestrator-Worker pattern is ideal for tasks that can be parallelized. A central orchestrator identifies the need for multiple independent lookups and dispatches them to parallel workers (or tool calls). This significantly reduces total execution time compared to a sequential process where each carrier is queried one after another.

Exam trap

Candidates often choose a sequential loop or a single-agent architecture. They fail to recognize that parallelizing independent tasks is the primary driver of performance gains in modern agentic systems.

38
MCQeasy

In the Anthropic tool-use workflow, what is the primary purpose of the 'system prompt' relative to tool usage?

A.To provide the JSON schema for each tool
B.To store the results of previous tool executions
C.To define the API keys for the external services
D.To establish the agent's persona and tool-use guidelines
AnswerD

The system prompt is the correct place to define the agent's role (e.g., 'You are a helpful logistics assistant') and provide high-level instructions on how to prioritize different tools, how to handle errors, and how to interact with the user after receiving results from the tools.

Why this answer

The system prompt serves as the 'instruction manual' for the agent. It defines the agent's identity, its overall goals, and the constraints within which it must operate. While tool schemas define the 'how,' the system prompt provides the 'why' and 'when,' guiding the model's decision-making process for tool selection.

Exam trap

Candidates often confuse the system prompt with the actual JSON tool schemas, failing to realize the system prompt dictates the operational guidelines while schemas define syntax.

39
MCQmedium

An agent is engaged in a multi-hour troubleshooting session involving dozens of tool calls and thousands of lines of log data. The architect notices that the agent is starting to 'forget' early symptoms of the problem. Which strategy best addresses this while managing token costs?

A.Truncating the oldest messages once the limit is reached
B.Implementing a recursive summarization buffer
C.Switching to a model with a 1-million token window
D.Storing all tool results in an external vector database
AnswerB

Summarizing previous turns into a concise narrative preserves the essential findings and progress of the agent. By passing this summary forward into new turns, the agent maintains a continuous understanding of the session history without needing to re-process every individual token from the original, verbose tool outputs.

Why this answer

Managing context in long-running agentic sessions requires a balance between detail and capacity. A summarization strategy combined with a sliding window allows the agent to retain the 'gist' of historical turns while keeping the most recent, high-fidelity data available. This prevents context overflow while maintaining the logical continuity of the troubleshooting process.

Exam trap

Candidates frequently suggest simply increasing the context window or dumping all logs into the prompt. This ignores the exponential cost of token usage and the degradation of model focus over time.

40
MCQmedium

Which pattern is most suitable for an agent that must balance the need for speed (latency) versus the need for correctness in a customer support scenario?

A.Always prioritize speed by using the smallest, fastest model available.
B.Use a dual-path architecture with immediate streaming and background verification.
C.Force the user to wait for verification before showing any response.
D.Use a RAG-only architecture to guarantee factual accuracy.
AnswerB

This architecture provides the best of both worlds: immediate feedback for the user and a secondary validation layer that ensures the content is accurate. This pattern manages the trade-off between perceived latency and result quality, delivering a superior user experience without sacrificing the precision required for high-quality support.

Why this answer

The 'Dual-Path' pattern, where the system initiates an immediate low-latency response while simultaneously running a more comprehensive, high-correctness check in the background, optimizes for both user experience and accuracy. This ensures the user feels acknowledged immediately, while the eventual output remains verified and accurate. This balance is critical in customer support, where perceived responsiveness is as vital as providing correct, verified information for technical or complex queries.

Exam trap

Candidates often choose between either speed or accuracy. They mistakenly assume they must sacrifice one, ignoring that a dual-path architecture allows for both immediate user feedback and eventual verification.

41
MCQmedium

A customer-support agent must sometimes escalate to a human and sometimes resolve autonomously. The compliance team requires that any action touching billing be reviewed by a human before execution, while password resets may proceed automatically. The architect wants the model to decide routing without hardcoding every rule in the prompt. Which design best satisfies the requirement?

A.Fine-tune the model on historical escalations so it learns which actions compliance typically requires humans to approve.
B.Classify each proposed tool call by risk tier in the orchestration layer, auto-approve low-risk actions, and require human approval for billing-tier actions regardless of model output.
C.Give the model a single escalate tool and instruct it in the system prompt to use judgment about when human review is required.
D.Log every tool call the agent makes and have the compliance team review the logs weekly to catch any unauthorized billing actions.
AnswerB

Enforcing the risk tier outside the model makes the control deterministic and auditable: billing actions cannot execute without human approval even if the model proposes them. Password resets proceed automatically because they fall in the low-risk tier. The model still decides routing in the sense of proposing actions, but the orchestration layer holds the authoritative gate, satisfying compliance without hardcoding every conversational rule in the prompt.

Why this answer

The compliance requirement is a hard gate, so the decision to require human approval must live in deterministic orchestration code rather than in model judgment. Tiering tool calls by risk lets low-risk actions like password resets flow automatically while billing actions are held for approval before execution. The model can still propose actions, but the authoritative approval decision is enforced outside it, which makes the control auditable and immune to prompt drift or probabilistic misrouting.

Exam trap

The trap here is assuming that a well-written system prompt or a fine-tuned model can serve as a compliance control, when only deterministic enforcement outside the model can guarantee a gated action.

42
Multi-Selecthard

An autonomous agent is designed to browse the web and perform research. Which TWO mechanisms are most critical for preventing infinite loops and excessive API consumption during autonomous tool-calling cycles?

Select 2 answers
A.Increasing the context window size
B.Implementing a maximum iteration counter
C.Using a lower temperature setting
D.Hashing and comparing previous state snapshots
E.Enabling prompt caching for tool definitions
AnswersB, D

A hard limit on the number of turns an agent can take provides a definitive fail-safe against recursive behavior. This ensures that even if the model's reasoning fails to reach a conclusion, the process terminates after a pre-defined threshold, protecting the system from infinite execution and associated costs.

Why this answer

In autonomous agentic loops, the risk of 'stuck' states or recursive logic is high. Implementing both hard iteration limits and state-based detection ensures the system remains within operational bounds. These safeguards are essential for production-grade agents to prevent runaway costs and to provide a predictable user experience in non-deterministic environments.

Exam trap

Candidates often rely only on simple prompt instructions telling the agent not to loop, failing to implement hard programmatic safety limits like iteration counters and state hashing.

43
Multi-Selecthard

You are scaling an agentic system that uses external APIs. Which THREE design patterns prevent the agent from being blocked by third-party rate limits or latency?

Select 3 answers
A.Circuit breaker pattern to detect and isolate failing API endpoints.
B.Synchronous, real-time polling for every single tool invocation.
C.Asynchronous task queuing for long-running tool operations.
D.Request-response caching to serve recurring tool result patterns.
E.Increasing model temperature to provide more creative workarounds.
AnswersA, C, D

A circuit breaker monitors for failures and trips when a threshold is reached. This stops the agent from sending requests to a service known to be down, preventing wasted resources and allowing the service time to recover. It is essential for protecting the agent from external system instability.

Why this answer

Managing external dependency risk is critical for production agents. Implementing a circuit breaker prevents cascading failures by halting calls to failing services. A task queue enables asynchronous execution, decoupling the agent's reasoning from external latency.

Finally, caching common responses reduces total API calls, improving throughput and reliability. Combined, these patterns create a resilient boundary between the autonomous agent and the unpredictable nature of external network services.

Exam trap

Candidates often focus on increasing API rate limits or scaling infrastructure, failing to recognize that architectural patterns like circuit breakers and caching are the standard for handling third-party instability.

44
MCQmedium

In a swarm of specialized agents, what is the primary benefit of using a 'Blackboard' pattern for inter-agent communication?

A.It forces all agents to run sequentially, which improves execution speed.
B.It eliminates the need for any form of agent orchestration or management.
C.It enables decoupled, non-linear collaboration between specialized agents.
D.It increases security by isolating agent memory into individual silos.
AnswerC

By utilizing a shared blackboard, agents can contribute findings independently, allowing for a non-linear problem-solving process. This decoupling is crucial because it allows individual agents to focus on their specific tasks while providing a unified view of the current progress, enabling more sophisticated emergent reasoning across the entire multi-agent swarm.

Why this answer

The Blackboard pattern allows multiple agents to contribute their expertise to a shared data structure, or blackboard, which serves as a common knowledge base. This promotes decoupling, as agents do not need to know about each other's existence, only how to read from or write to the blackboard. This architecture is essential for complex, multi-faceted problems where multiple specialized agents must collaborate without creating rigid, brittle dependency chains.

Exam trap

Candidates often confuse the Blackboard pattern with a centralized controller. They mistakenly believe the pattern is about command-and-control, rather than the decoupling of agents through a shared knowledge space.

45
Multi-Selectmedium

When designing an agentic system, which TWO of these 'observability' metrics are most crucial for monitoring the health of the agent's reasoning process?

Select 2 answers
A.Total Step Count per Task.
B.Average latency of the user's internet connection.
C.Tool Call Success Rate.
D.The color profile of the agent's UI dashboard.
E.Number of times the user clicks the refresh button.
AnswersA, C

Monitoring the number of steps an agent takes helps identify inefficient workflows or runaway reasoning. If a task that should take 3 steps is taking 50, the agent is likely stuck in a loop or struggling to formulate a plan, indicating a need for better prompt instructions or debugging.

Why this answer

Tracking 'Step count per task' and 'Tool call success rate' provides immediate visibility into whether the agent is diverging or struggling. An unusually high step count indicates potential infinite loops or circular reasoning. A low tool call success rate reveals integration failures or ambiguous tool definitions.

By monitoring these, you can detect system degradation before it impacts the end-user, allowing for proactive debugging and iterative improvements to the agent's core instruction set.

Exam trap

Candidates frequently select vanity metrics like total token usage or wall-clock elapsed time, failing to realize that reasoning health is specifically evaluated via step counts and tool success rates.

46
Multi-Selecthard

Which THREE strategies should be employed to securely manage sensitive data within an agentic workflow?

Select 3 answers
A.Apply PII masking/redaction before passing text to the LLM.
B.Enable the model to have full administrative access to the file system.
C.Implement fine-grained access control (RBAC) on all tool calls.
D.Maintain immutable audit logs of all model inputs and outputs.
E.Embed all user data directly into the system prompt for faster access.
AnswersA, C, D

PII masking ensures that sensitive data never leaves your secure environment in a readable format. By replacing names, emails, or IDs with tokens, the model can still perform logic based on the pattern without ever having access to the real, private data, which is essential for GDPR/HIPAA compliance.

Why this answer

Securing agents requires a layered approach: PII masking before sending to the model keeps data private; granular access control ensures the agent only touches data it is authorized to see; and audit logging provides traceability for every decision made. These components ensure that even if the agent is compromised or hallucinates, the impact is contained, data privacy is maintained, and there is a clear record of actions taken for compliance and forensics.

Exam trap

Candidates often focus solely on encrypting data at rest while forgetting that sensitive PII must be actively scrubbed before being passed to external LLM APIs.

47
MCQmedium

You are operating a Claude-based support agent that must follow a strict refund policy: refunds over $500 require a manager approval code that is only obtainable through a separate internal API. The agent has access to a `get_manager_code` tool, but in production it occasionally issues refunds above $500 without calling the tool. Which architectural change most reliably prevents this?

A.Expand the system prompt to describe the $500 threshold in bold and add two few-shot examples of correct refund handling.
B.Increase the max_tokens setting so the agent has more room to reason about whether approval is needed.
C.Add a pre-tool-use hook that inspects the refund amount and blocks any refund tool call above $500 unless a valid manager code is present in the session state.
D.Lower the model temperature to 0 so the agent produces deterministic decisions.
AnswerC

A pre-tool-use hook runs deterministically before the tool executes, so it can inspect the refund payload and reject any call over $500 that lacks a verified manager code. This enforces the policy outside the model's probabilistic reasoning, which is exactly where a hard business rule belongs. It makes violation architecturally impossible rather than merely unlikely.

Why this answer

Hard business constraints must be enforced deterministically outside the model. A pre-tool-use hook inspects the pending refund call and rejects any amount above $500 that lacks a validated manager code, making the violation impossible rather than improbable. Prompting, temperature, and token limits only shift probabilities and cannot guarantee compliance with a policy that carries financial or legal consequences.

Exam trap

The trap here is assuming that stronger prompting or lower temperature can enforce a hard business rule, when only deterministic interception outside the model can guarantee it.

48
Multi-Selecthard

You are designing a long-running agent that executes a sequence of irreversible operations, such as issuing refunds and sending customer notifications. The agent runs unattended overnight. Which TWO architectural patterns best ensure that a partial failure does not leave the system in an inconsistent state? (Choose two.)

Select 2 answers
A.Cache all tool responses in memory so the agent can replay the session without re-calling tools after a restart.
B.Raise the max_tokens limit for each planning step so the agent can reason through more contingencies before acting.
C.Implement compensating actions so that each irreversible operation has a defined reversal or mitigation step.
D.Increase the model's temperature so it explores alternative execution orders and avoids getting stuck on a failing step.
E.Record each intended operation in a durable journal before executing it, and mark it complete only after the tool confirms success.
AnswersC, E

Compensating actions provide a defined path to undo or mitigate an operation that succeeded but belongs to a workflow that later failed. For refunds and notifications, this might mean issuing a reversal or a correction notice. Together with journaling, compensation ensures that a partially completed sequence can be brought back to a consistent state rather than left with orphaned side effects.

Why this answer

Unattended agents performing irreversible actions need both a durable record of intent and a way to reverse or mitigate completed steps. Journaling provides the audit trail and restart logic, while compensating actions provide the semantic undo. Together they allow the orchestrator to detect partial completion and bring the workflow back to a consistent state after a crash or timeout.

Exam trap

The trap here is reaching for model-side knobs like temperature or token limits to solve what is fundamentally a distributed-systems durability and compensation problem.

49
MCQhard

You operate a long-running research agent that maintains a scratchpad of findings across many turns. You notice that as the scratchpad grows, the agent begins ignoring recent tool results and repeating earlier conclusions. You cannot increase the context window. Which intervention most directly addresses the root cause of the recency failure?

A.Move the scratchpad out of the prompt entirely and store it in an external vector database, retrieving relevant entries only when the agent explicitly asks.
B.Append every tool result verbatim to the scratchpad so no information is lost, and instruct the agent to re-read the entire scratchpad before each decision.
C.Restructure the scratchpad into a fixed-size rolling summary that is regenerated each turn, with the most recent tool results kept verbatim in a dedicated, clearly delimited section.
D.Lower the model's temperature to zero so the agent becomes more deterministic and less likely to drift from the current evidence.
AnswerC

The root cause is that accumulated history pushes recent results into a diluted middle position. A rolling summary compresses old findings into a bounded block, while a dedicated recent-results section guarantees the newest evidence sits in a high-attention position. This keeps total context roughly constant, so recency is preserved without needing a larger window.

Why this answer

The agent is not forgetting recent results so much as losing them in an ever-growing block of text where early material dominates attention. Bounding the scratchpad with a regenerated rolling summary keeps total size stable, and reserving a clearly delimited section for the newest tool results puts current evidence in a position the model reliably attends to. This fixes the structural cause rather than the sampling or storage symptoms.

Exam trap

The trap here is diagnosing the recency failure as a memory or determinism problem and reaching for external storage or temperature changes, when the actual cause is where recent evidence sits inside a growing context.

50
MCQmedium

Refer to the exhibit. An agentic loop receives this response from the Anthropic API during a critical multi-step operation. Which strategy should the architect implement to ensure the agent completes its task successfully?

A.Immediately terminate the agent and alert the user
B.Restart the entire agentic loop from the first message
C.Implement exponential backoff with jitter in the orchestrator
D.Reduce the temperature of the next request to 0
AnswerC

This is the standard resilience pattern for distributed systems. Retrying the request after an increasing delay, combined with a small amount of randomness (jitter), helps mitigate congestion on the API server while allowing the agent to eventually proceed with its task once the service stabilizes.

Why this answer

Transient API errors like 'overloaded_error' are common in high-traffic environments. A robust agentic architecture must handle these gracefully using exponential backoff. This prevents the agent from failing the entire task due to a temporary service interruption and ensures that the long-running state of the agent's work is preserved and resumed.

Exam trap

Candidates often suggest increasing the model temperature or changing the system prompt. These do not solve transient infrastructure issues and will result in repeated failures during high-traffic periods.

51
Multi-Selecthard

You are designing a Claude agent that must complete a multi-hour research task spanning dozens of tool calls, and the transcript will eventually exceed the model's context window. You want the agent to keep making correct decisions without losing critical earlier findings. Which TWO architectural strategies best preserve decision quality across the compaction boundary? (Choose two.)

Select 2 answers
A.Summarize and discard the oldest transcript turns once a threshold is reached, retaining only the most recent turns verbatim.
B.Enable extended thinking on every turn so the model can reason through the full history internally without needing external state.
C.Raise the max_tokens parameter on every request so the model can hold more of the transcript in a single response.
D.Maintain an external structured scratchpad of confirmed findings, open questions, and decisions, and inject a curated summary of it into each new context window.
E.Persist a durable task state object and rehydrate the agent from it at each context boundary, treating the transcript as a disposable execution log.
AnswersD, E

An external scratchpad decouples durable knowledge from the ephemeral transcript, so compaction can drop raw turns without losing conclusions. Injecting a curated summary re-establishes the essential state each cycle, and because the scratchpad is structured, the agent can update specific fields rather than rewriting prose. This keeps decisions consistent even when the underlying conversation is truncated.

Why this answer

Durable continuity comes from moving authoritative state outside the transcript. A structured scratchpad preserves confirmed findings and open questions, while a persisted task state object lets the agent rehydrate deterministically at each boundary. Together they let compaction discard raw turns safely.

Token limits and internal reasoning do not expand context or survive eviction.

Exam trap

The trap here is conflating output length controls and reasoning effort with input context capacity, when none of them persists knowledge across a context reset.

52
MCQmedium

A Claude agent performs a multi-step deployment task. Step 3 calls a `deploy_service` tool that returns success, but the subsequent verification step fails because the service is not yet healthy. The agent currently treats any tool success as completion and ends the workflow. Which change best addresses this?

A.Instruct the model in the system prompt to always wait 60 seconds and re-check health after every deploy.
B.Increase the agent's max_tokens so it has more room to notice the verification failure in its reasoning.
C.Change the `deploy_service` tool to return a boolean instead of a status string so the agent can parse it more easily.
D.Add a post-tool-use hook that polls the service health endpoint and, if unhealthy, returns a structured error to the model so it can retry or escalate.
AnswerD

A post-tool-use hook runs after the deploy call and can independently verify health, converting a false success into a structured signal the model can act on. This closes the gap between the tool's reported success and the actual desired state. It keeps the orchestration deterministic at the verification boundary while letting the model decide the next step.

Why this answer

The agent conflates tool acceptance with desired end state, so verification must be moved into a deterministic post-tool-use hook that polls health and returns a structured error when the service is unhealthy. That lets the model retry or escalate instead of ending the workflow prematurely. Prompting, return-type changes, and token limits all leave the false-success gap intact.

Exam trap

The trap here is equating a tool's successful response with the workflow's desired outcome, when asynchronous systems often report acceptance before the effect is observable.

53
MCQhard

Which THREE components are essential for building a robust 'human-in-the-loop' (HITL) approval gate within an agentic workflow?

A.A state persistence layer to save the agent context during the pause.
B.An asynchronous messaging system to notify users of pending actions.
C.A direct feedback loop that forces the model to re-train after every human input.
D.A mechanism to resume execution by injecting the human's decision into the next prompt.
E.A mandatory requirement that humans provide a complete rewrite of the agent's plan.
AnswerA, B, D

Persistence is required to resume the agent's reasoning exactly where it left off. Without saving the state, the agent would lose its progress, context, and the reasoning chain that led to the approval request, rendering the workflow broken upon resumption. This layer is fundamental for durable, multi-step agentic processes.

Why this answer

A robust HITL gate requires a state suspension mechanism, a clear communication interface for the human, and an automated resumption process. By pausing the agent's execution, capturing the state, and allowing for human intervention, the system ensures that high-stakes decisions are verified. This is critical for preventing unauthorized or dangerous actions while maintaining the agent's context and momentum throughout the broader automated process.

Exam trap

Candidates often focus only on the notification system, neglecting the critical requirement of state persistence. Without saving the agent's internal state, the context is lost when the human intervenes.

54
MCQmedium

An agentic system is struggling with 'context fragmentation' over long-running sessions. What is the most effective architectural solution?

A.Increase the context window size of the model indefinitely.
B.Implement a hierarchical memory system with summarization.
C.Force the user to clear their history periodically.
D.Only use the most recent 10 messages of the history.
AnswerB

Hierarchical memory allows the agent to access both recent, detailed context and summarized long-term history. By managing memory at different levels of abstraction, you keep the agent's prompts concise and relevant, significantly improving its performance and goal-tracking accuracy over sessions that span hours or days.

Why this answer

Context fragmentation happens when the history becomes too large or disorganized. A 'summarization agent' or a 'memory manager' that periodically compresses the conversation history into a concise summary is the standard solution. This preserves core context while discarding transient details, ensuring the main agent remains focused on the long-term goal rather than getting lost in thousands of lines of previous chat history.

Exam trap

Test-takers might suggest increasing the model's max context window or clearing history entirely, failing to recognize that summarization preserves necessary long-term context.

55
MCQhard

A production agent uses the Claude Messages API with extended thinking enabled to solve multi-constraint scheduling problems. The agent must preserve its reasoning across several tool calls within a single user turn. A developer notices that after the second tool call, the model appears to forget earlier constraints it had already reasoned about. Which change best preserves the reasoning chain across tool calls?

A.Pass the full assistant message, including thinking blocks, back into the conversation history on each subsequent request.
B.Disable extended thinking for subsequent requests once the first tool call has returned, since the constraints are already known.
C.Increase the thinking budget_tokens value so the model has more room to reason on each request.
D.Summarize the thinking blocks into a single text note and append it as a user message before the next tool call.
AnswerA

With extended thinking, the reasoning is carried in thinking blocks attached to assistant messages. To keep that reasoning available across tool calls in the same turn, the orchestrator must return the complete assistant message, thinking blocks included, in the next request's messages array. Stripping them discards the model's prior reasoning and forces it to re-derive constraints, which explains the apparent forgetting after the second tool call.

Why this answer

Extended thinking stores reasoning in thinking blocks that are part of the assistant message. When an agent makes multiple tool calls within one turn, each follow-up request must include the prior assistant message with its thinking blocks intact. Omitting them causes the model to lose the constraints it already reasoned through, producing the observed regression after the second tool call.

Exam trap

The trap here is treating thinking budget as a memory setting, when the real requirement is to preserve the original thinking blocks in the conversation history across tool calls.

56
MCQhard

Refer to the exhibit. An agentic workflow encounters an error immediately after this message is generated by Claude. No further messages are sent to the API. What is the most likely cause of the failure in the orchestration logic?

A.The orchestrator failed to send a tool_result message
B.The input parameters do not match the JSON schema
C.The tool_use ID is missing the required prefix
D.The assistant content block lacks a stop_reason
AnswerA

The Messages API is a turn-based protocol where every tool_use request from the model must be acknowledged with a corresponding tool_result from the client. Without this result, the conversation state is incomplete, and the model cannot proceed to synthesize the information or determine the next appropriate action in the workflow.

Why this answer

The Anthropic Messages API requires a specific sequence for tool use. After the assistant generates a 'tool_use' block, the orchestration layer must execute the tool and return a 'tool_result' message to the model before the assistant can continue its turn. Failing to provide this required response breaks the synchronous nature of the tool-calling conversation flow.

Exam trap

Candidates assume the LLM will automatically proceed after generating a tool call, forgetting that the orchestrator must explicitly feed back a tool result message to continue.

57
MCQmedium

Your team operates a Claude agent that triages inbound customer support tickets. After several weeks in production, the agent begins approving refunds above the policy ceiling and citing outdated policy text. Investigation shows the system prompt embeds a policy document that was updated three weeks ago, but the deployed prompt was never regenerated. Which architectural practice most directly prevents this class of failure?

A.Move policy content out of the system prompt and retrieve the current policy at runtime via a versioned tool or retrieval step, recording the version used in each decision.
B.Add a second Claude agent that reviews the first agent's refund decisions and overrides any that appear inconsistent with its own recollection of company policy.
C.Increase the model's temperature slightly so the agent explores alternative interpretations of the policy text it has memorized.
D.Shorten the system prompt by summarizing the policy into a compact bullet list, which reduces the chance that outdated sections remain embedded.
AnswerA

Retrieving policy at runtime decouples the agent's behavior from a frozen prompt snapshot, so updates take effect without redeploying the prompt. Recording the retrieved version makes each decision reproducible and lets auditors see which policy text governed a given refund. This directly addresses staleness because the agent always reads the current authoritative source rather than an embedded copy.

Why this answer

Stale behavior traced to a frozen prompt is best solved by externalizing the mutable knowledge and fetching it at decision time from a versioned source. That way policy updates propagate immediately, and logging the retrieved version makes each action auditable. Compression, randomness, or peer review all leave the embedded copy as the de facto source of truth.

Exam trap

The trap here is treating the symptom as a prompt-engineering problem, when the real defect is that authoritative, frequently changing content was hard-coded into the prompt.

58
Multi-Selectmedium

An architect wants to improve the coherence of an agent that frequently makes 'leaps of logic' or misses obvious errors in its tool outputs. Which TWO techniques directly address this behavior?

Select 2 answers
A.Chain-of-Thought (CoT) in the assistant turns
B.Increasing the frequency of tool calls
C.Reducing the temperature to exactly 0.0
D.Using the 'tool_choice' parameter to force tool use
E.Implementing a self-reflection/critique loop
AnswersA, E

Encouraging the model to 'think out loud' before generating a tool call helps it process complex instructions more accurately. This explicit reasoning step allows the model to verify dependencies and logic internally, which significantly reduces the likelihood of making irrational or incorrect tool selections during complex tasks.

Why this answer

Coherence in agents is improved by forcing the model to externalize its reasoning and critique its own work. Chain-of-thought encourages the model to plan before acting, while self-reflection allows it to catch and correct its own mistakes in a subsequent turn, leading to much more reliable agentic behavior.

Exam trap

Candidates often suggest prompt engineering to 'tell the model to be smarter'. This is rarely effective for complex logic errors; structural improvements to the reasoning process are required instead.

59
MCQmedium

You are building a Claude-based agent that must parse unstructured customer emails, extract line-item order data, and then call a fulfillment tool with the extracted values. During testing, the agent occasionally calls the fulfillment tool with empty or garbled line items when an email contains a forwarded message with a different formatting style. Which architectural change most directly reduces this failure?

A.Reduce the number of tools available to the agent so it focuses only on fulfillment during extraction.
B.Add a validation step that checks the extracted line items against the tool's JSON schema and returns a structured error to the model before any tool call is allowed.
C.Switch the fulfillment tool to accept free-form natural language instead of structured parameters so the model can pass whatever it extracted.
D.Increase the max_tokens parameter on the extraction call so the model has more room to emit complete line items.
AnswerB

A schema validation gate on the extracted payload catches malformed or empty line items before the fulfillment tool is invoked. Returning a structured error to the model lets it re-read the email and repair the extraction. This addresses the root cause of garbled output rather than masking it, and it keeps the tool boundary safe from invalid inputs. It is the most direct architectural fix for the described failure.

Why this answer

The failure occurs because extracted line items are not verified before the tool boundary is crossed. A validation gate that compares the extraction against the tool's JSON schema and returns a structured error gives the model a chance to repair its output. This keeps invalid data out of downstream systems and converts a silent corruption into a retryable, observable event.

Exam trap

The trap here is assuming that a larger output budget or a simpler tool interface will fix malformed extractions, when the actual defect is the absence of a validation boundary before tool invocation.

60
MCQhard

An agent orchestrator delegates work to three specialist sub-agents: a 'search' agent that returns ranked documents, an 'extract' agent that pulls structured fields from those documents, and a 'verify' agent that checks extracted fields against source text. During evaluation you find that verify frequently approves fields that extract hallucinated, because verify receives only the extracted JSON, not the source passages. Which change to the orchestration contract most directly fixes this?

A.Route all extracted fields back through the search agent so it can re-rank the documents before verification.
B.Increase the verify agent's temperature so it is more likely to question the extracted fields it receives.
C.Add a fourth 'audit' agent that independently re-runs the extract agent on the same documents and compares the two JSON outputs.
D.Have the extract agent include, for each field, the source span it was derived from, and pass those spans to the verify agent alongside the extracted JSON.
AnswerD

Verification is only possible against evidence. By propagating the exact source spans that back each extracted field, the verify agent can compare the claimed value to the text it supposedly came from, which is the only way to detect fabrication. This changes the orchestration contract to carry provenance, directly closing the gap that let hallucinated fields pass.

Why this answer

The verifier approves hallucinations because it is asked to judge extracted values without access to the text they were derived from. Propagating per-field source spans gives the verifier the evidence needed to confirm or refute each value against its origin. This is a contract change between extract and verify, ensuring provenance flows with the data rather than being reconstructed or guessed downstream.

Exam trap

The trap here is treating verification failures as a reasoning or sampling problem and adding temperature, re-ranking, or a redundant extractor, when the verifier simply lacks the source evidence needed to detect fabrication.

61
MCQmedium

An enterprise agentic system using Claude needs to maintain strict state isolation across multiple concurrent user sessions while executing autonomous tool loops. Which architecture best ensures security and state integrity?

A.Maintain a single global conversation history array in memory and append all incoming user turns sequentially.
B.Store conversation history and intermediate scratchpads in an encrypted, session-scoped external database retrieved on each turn.
C.Rely entirely on Claude's native system prompt caching to automatically differentiate between distinct concurrent user sessions.
D.Encode all intermediate tool states directly into the client-side browser local storage and send the full payload to Claude.
AnswerB

Session-scoped external storage enforces isolation by keying every read and write to a unique session identifier, so concurrent agent loops cannot observe or mutate each other's scratchpads. Encryption at rest satisfies the security constraint, while per-turn retrieval keeps state authoritative rather than relying on in-context memory that could leak across sessions.

Why this answer

Isolating agent state at the session layer prevents context bleed and cross-user data contamination in enterprise environments. Stateless server designs combined with externalized transactional stores allow horizontal scaling while preserving deterministic tool execution history, which is critical for compliance and reliable multi-turn agent interactions.

Exam trap

Many candidates mistakenly choose to store session state within the agent's memory or prompt context, which leads to security risks like prompt injection or data leakage between concurrent user sessions.

62
MCQeasy

A team wants an agent to answer questions about a large internal corpus. They notice the agent invents details when the retrieved chunks are only loosely related to the question. Which change most directly reduces fabricated answers grounded in weak evidence?

A.Instruct the agent in the system prompt to answer only from provided documents and to state when evidence is insufficient.
B.Raise the temperature setting so the model explores more diverse phrasings of the answer.
C.Increase the number of retrieved chunks returned by the vector search to fifty per query.
D.Switch the embedding model to one with a larger vector dimension for finer similarity scoring.
AnswerA

Grounding instructions tie generation to the supplied context and give the model an explicit escape hatch when retrieval is weak. By authorizing an 'insufficient evidence' response, you remove the pressure to fabricate a plausible answer. This is the most direct lever because it changes behavior at generation time regardless of retrieval quality, and it composes with retrieval improvements later.

Why this answer

The reported failure is generation overreach under weak retrieval, so the highest-leverage fix is a grounding instruction that scopes answers to supplied documents and explicitly permits declining. This directly constrains what the model may assert and gives it a safe alternative to invention, independent of any retrieval tuning that might follow.

Exam trap

The trap here is treating every hallucination as a retrieval problem and reaching for embeddings or more chunks, when the described behavior is generation overreach best fixed by grounding instructions.

Ready to test yourself?

Try a timed practice session using only Advanced Agentic Architecture questions.