Courseiva

Claude Certified Developer (CCDV-F) — Questions 151–225

257 questions total · 4pages · All types, answers revealed

Page 2

Page 3 of 4

Page 4
151
MCQmedium

When should a developer choose to 'reject' an edit suggested by Claude Code?

A.Whenever the suggested change is more than ten lines.
B.When the changes introduce logical or security errors.
C.Only when the code fails to compile.
D.Every time the model uses a library the user dislikes.
AnswerB

Ensuring the code is logical and secure is the developer's primary responsibility. If the agent proposes code that fails testing, introduces vulnerabilities, or misunderstands business logic, the developer must reject the change to prevent the introduction of technical debt or critical bugs into the codebase.

Why this answer

Rejecting an edit is an essential part of the human-in-the-loop workflow. A developer should reject a suggestion if it violates architectural patterns, introduces security risks, fails to satisfy business requirements, or incorrectly interprets the codebase. Exercising this judgment ensures the quality of the project and prevents the AI from propagating errors, reinforcing the developer's role as the final authority on code quality and system architecture in the development lifecycle.

Exam trap

Candidates often think rejection is only for 'broken' code. They fail to realize that architectural misalignment or security concerns are equally valid reasons to reject an agent's suggestion.

152
MCQmedium

A financial firm needs to identify potential jailbreak attempts against their Claude-powered chatbot. Which approach provides the most effective real-time detection?

A.Use a static keyword-based filter on the user input.
B.Deploy a dedicated classification model to evaluate the safety of prompts and responses.
C.Audit the logs daily to search for suspicious query patterns.
D.Disable the ability for the model to access external URLs.
AnswerB

A dedicated classifier trained on adversarial examples can detect the nuances of jailbreak attempts far better than regex or simple logic. By checking inputs for patterns typical of prompt injection, this layer acts as an automated security filter that enhances the overall safety profile of the LLM application.

Why this answer

Real-time detection of jailbreak attempts is best achieved through a secondary, lightweight LLM or a specialized classifier that analyzes both the user input and the model's generated output. By evaluating the interaction for adversarial patterns—such as attempts to bypass safety filters—the security layer can intercept malicious prompts before they are fully processed or block harmful responses, providing a critical safety net that static rules cannot match.

Exam trap

Candidates choose static input sanitization filters, forgetting that sophisticated jailbreaks evolve dynamically and require real-time evaluation via a dedicated classification model.

153
MCQmedium

A developer is integrating the Messages API into a backend service. They want to send a multi-turn conversation where the assistant previously produced a tool call that returned a result. Which content structure should they send back to continue the conversation correctly?

A.A system message that appends the tool output to the original system prompt so it persists for all future turns.
B.A user message whose content array contains a tool_result block referencing the tool_use id, followed by any additional user text as separate blocks.
C.An assistant message echoing the tool name and its output so the model can read its own prior decision before answering.
D.A single user message containing the raw tool output as plain text, labeled with the tool name in parentheses.
AnswerB

The Messages API expects tool outputs delivered as tool_result blocks inside a user message, each referencing the tool_use_id from the assistant's prior tool_use block. Additional user text can coexist as separate content blocks in the same message. This preserves the call-result linkage and lets Claude continue reasoning with the returned data, which is the correct continuation pattern.

Why this answer

Tool results belong in a user-turn message as tool_result content blocks, each carrying the tool_use_id from the assistant's prior tool_use block. Additional user text can be included as sibling content blocks. This structured pairing preserves the call-result relationship the API expects, letting Claude incorporate the returned data and continue the conversation coherently.

Exam trap

The trap here is treating tool output as ordinary text in a user or assistant message, when the API requires a tool_result block linked by tool_use_id.

154
MCQmedium

Refer to the exhibit. A developer is configuring an MCP server to provide an agent with database access. What is the primary purpose of the 'args' field in this configuration block?

A.It defines the JSON Schema for the tools provided by the SQLite server.
B.It specifies the command-line flags and parameters used to initialize the server.
C.It lists the specific SQL queries the agent is permitted to execute.
D.It acts as a whitelist for authorized users of the SQLite server.
AnswerB

The 'args' array contains the specific flags—like the database path—required by the executable to function correctly. This is how the developer bridges the gap between the Agent SDK's request for a server and the specific local configuration needed to make that server's data accessible to the agent.

Why this answer

The 'args' field provides the necessary runtime parameters for the MCP server process. In this specific exhibit, it tells the 'uvx' command which package to run and which specific database file to mount. Correct configuration of these arguments is vital for the agent to successfully connect to the intended external resource.

Exam trap

Candidates often confuse 'args' with the tool's input parameters. In the context of MCP server configuration, 'args' refers to the command-line flags used to launch the server process itself.

155
MCQmedium

A developer is implementing a weather-reporting agent using Claude 3.5 Sonnet. After Claude generates a tool_use block for the 'get_weather' tool, the developer fetches the data. What is the correct next step to ensure Claude can finalize its response to the user?

A.Append the tool results directly to the end of the previous assistant message.
B.Send a new user message containing a 'tool_result' block with the 'tool_use_id'.
C.Update the system prompt with the new weather data and restart the session.
D.Issue a second assistant message that includes the weather data as a JSON string.
AnswerB

Sending a user message with a 'tool_result' block is the mandatory way to provide data back to Claude. The 'tool_use_id' must match the ID provided by Claude in the previous turn, enabling the model to link the result to the specific request it made, ensuring logical consistency in complex workflows.

Why this answer

The tool use lifecycle requires a multi-turn interaction where the model first requests a tool execution. Once the developer executes the tool locally, they must return the result to Claude in a new 'user' message with a 'tool_result' block. This allows Claude to incorporate the external data into its final synthesis and provide an accurate answer to the original query.

Exam trap

Candidates often try to send the tool result as a new 'assistant' message or a standard user message without the required 'tool_result' block, breaking the expected tool-use protocol.

156
Multi-Selecthard

A developer is building an MCP server that exposes several tools for managing a cloud infrastructure. The server will be used by multiple Claude clients. The developer wants to ensure that tool calls are handled efficiently and that the server can scale. Which two practices should the developer implement? (Choose two.)

Select 2 answers
A.Cache the results of expensive tool calls on the server side and return cached responses for identical requests.
B.Design tools to be idempotent where possible, so that repeated calls with the same parameters produce the same result.
C.Use a single global lock to serialize all tool calls, preventing concurrent execution.
D.Implement the server to handle each tool call statelessly, so that any instance can process any request.
E.Store session state in memory on each server instance to track client interactions.
AnswersB, D

Idempotency ensures that if a tool call is retried due to network issues or client retries, it does not cause unintended side effects like creating duplicate resources. This is crucial for infrastructure management where operations like provisioning or deletion should be safe to repeat. It improves reliability and simplifies error handling, contributing to efficient scaling.

Why this answer

The two best practices are statelessness and idempotency. Statelessness allows any server instance to handle any request, enabling horizontal scaling. Idempotency ensures that retries are safe, which is important for reliable infrastructure operations.

Together, they improve efficiency and scalability. Other options like caching, global locking, or in-memory session state either introduce complexity or hinder scalability.

Exam trap

The trap here is focusing on performance optimizations like caching or session tracking, when the fundamental practices for scalable MCP servers are statelessness and idempotency.

157
MCQhard

A developer is using the Claude Messages API to process a 50,000-token document for a summarization task. The summarization prompt and instructions add another 1,000 tokens. The developer wants to minimize output token costs while ensuring a comprehensive summary. Which strategy is most effective?

A.Split the document into smaller chunks and summarize each separately, then combine the summaries.
B.Set the max_tokens parameter to a very low value, such as 100.
C.Craft a prompt that instructs Claude to produce a concise summary within a specified token limit, such as 'Summarize in under 500 tokens.'
D.Use a smaller model like Claude 3 Haiku for summarization.
AnswerC

By explicitly instructing Claude to produce a summary within a token limit, the developer guides the model to generate a concise output, directly reducing output token costs. This approach maintains summarization quality while controlling length. The model will still capture essential information but avoid unnecessary verbosity, making it the most effective strategy for cost management.

Why this answer

To minimize output token costs while ensuring a comprehensive summary, the developer should control the output length through prompt instructions. Explicitly requesting a concise summary within a token limit guides the model to generate only the necessary content, reducing cost without sacrificing quality. Other strategies like chunking or using a smaller model may address different concerns but do not directly target output token reduction.

Exam trap

The trap here is focusing on reducing input costs or model size when the question specifically asks about minimizing output token costs, which is best addressed by controlling the generated response length.

158
MCQhard

You are debugging an agent that often fails to format its final response as JSON, despite being instructed in the system prompt. Which technique is most effective for ensuring consistent output structure?

A.Adding a 'Please provide JSON' phrase at the end of the prompt.
B.Utilizing a tool to force the output into a structured schema.
C.Lowering the system prompt length to reduce ambiguity.
D.Retrying the call automatically until valid JSON is returned.
AnswerB

Defining a tool that expects the final answer as its input forces the model to generate the JSON parameters that match the schema. This creates a hard constraint that the model must satisfy to execute the tool, ensuring the output is perfectly formatted.

Why this answer

Forcing the output structure through constrained generation is the most robust way to ensure valid JSON. While system prompts set expectations, they are soft constraints. Using tool-use or structured output features allows the agent to adhere to a schema, which guarantees the output is machine-readable and prevents parsing errors in downstream applications.

Exam trap

Candidates rely heavily on system prompt 'formatting' instructions, failing to realize that LLMs are not guaranteed to follow text-based constraints and require structured schema enforcement for reliable output.

159
MCQmedium

A developer wants Claude Code to remember a project-specific convention: all new database migration files must be placed in the `db/migrations` directory and named with a `YYYYMMDDHHMMSS_` prefix. The developer wants this rule applied automatically in every future Claude Code session for this repository, without having to repeat the instruction each time. What is the most appropriate way to achieve this?

A.Add the convention to a CLAUDE.md file at the repository root.
B.Run `claude config set migration.dir db/migrations` before each session.
C.Create a `.claudeignore` file listing `db/migrations` to protect the directory.
D.Paste the convention into the first message of every new session.
AnswerA

CLAUDE.md at the repository root is automatically loaded into context at the start of every Claude Code session in that repository, making it the correct place for persistent project conventions. It requires no per-session action and is version-controlled alongside the code, so all developers and all future sessions inherit the rule consistently.

Why this answer

Persistent, repository-scoped instructions belong in CLAUDE.md, which Claude Code reads automatically at session start. That file can be committed so every developer and every future session inherits the naming and placement rule without manual repetition. Ad-hoc pasting or non-existent config keys do not satisfy the requirement of automatic, durable application across sessions.

Exam trap

The trap here is assuming there is a dedicated config command for every project convention, when the supported mechanism for durable project instructions is the CLAUDE.md memory file.

160
MCQhard

A developer wants to implement 'pre-filling' to guide Claude's output toward a specific format. How should the 'messages' array be structured to accomplish this?

A.End the 'messages' array with a 'user' message containing the desired starting text.
B.End the 'messages' array with an 'assistant' message containing the desired starting text.
C.Include the desired starting text in the 'system' parameter with a 'prefix' label.
D.Add a 'prefill' field to the top-level API request object.
AnswerB

By ending the array with an 'assistant' message, the developer provides the initial tokens of the response. Claude will then continue generating from where that message left off. This is the standard and recommended way to steer the model's behavior and response format effectively.

Why this answer

Pre-filling is a powerful technique where the developer provides the beginning of the assistant's response. This forces Claude to continue from that point, which is highly effective for ensuring specific output formats like JSON or XML. It requires a specific message order that deviates from the typical user-only or user-assistant-user pattern.

Exam trap

Candidates often attempt to pre-fill by appending to the 'user' message or adding a new field, violating the API's structural requirement that the final role must be an 'assistant'.

161
Multi-Selecthard

A developer wants Claude Code to run a series of shell commands as part of a build task but is concerned about the assistant executing arbitrary or destructive commands. Which TWO practices reduce this risk while still allowing the needed commands to run? (Choose two.)

Select 2 answers
A.Add a PreToolUse hook on the Bash tool that inspects the command and blocks anything outside the approved set.
B.Instruct Claude Code in CLAUDE.md to never run dangerous commands.
C.Run the build inside a container with no network access and a read-only filesystem.
D.Enable `--dangerously-skip-permissions` to avoid prompts during the build.
E.Define an allow list in `.claude/settings.json` that permits only the specific build commands the task requires.
AnswersA, E

A PreToolUse hook can examine the proposed Bash command and reject it before execution, providing a programmatic guardrail beyond static allow lists. This lets the developer enforce custom rules, such as rejecting commands containing destructive flags, while still permitting the build commands. It is a flexible way to reduce risk during shell execution without removing the ability to run what is needed.

Why this answer

Least-privilege controls are the right answer. An allow list in settings.json restricts execution to the specific build commands, and a PreToolUse hook on the Bash tool can inspect and block disallowed commands before they run. Together they permit the intended build while constraining arbitrary or destructive execution.

Skipping permissions removes safeguards, memory-file advice is not enforced, and container hardening changes the environment rather than governing command selection.

Exam trap

The trap here is treating a written instruction in CLAUDE.md as a security control, when only enforced mechanisms like allow lists and hooks actually restrict which commands can execute.

162
MCQmedium

When sharing state between multiple agent turns, why is it recommended to use a managed session object instead of appending everything to a simple list of messages?

A.It guarantees that all messages are stored in a database.
B.It enables intelligent context pruning and summarization.
C.It automatically encrypts all user inputs for privacy.
D.It makes the model faster by skipping validation.
AnswerB

Managed sessions enable the application to intelligently decide which parts of history to drop or summarize. This keeps the prompt focused and within token limits, which is essential for maintaining model performance over long, multi-turn interactions where history might otherwise become unmanageable.

Why this answer

Managed state objects allow for intelligent context truncation and summarization. In complex agent workflows, message histories can quickly exceed the model's context window. A robust session manager handles pruning older turns, maintaining critical environment variables, and ensuring the model receives only the most relevant information for the current task, which is critical for long-running autonomous processes.

Exam trap

Candidates assume that storing all message history is sufficient, failing to account for the model's context window limits and the performance degradation caused by sending excessively long, unpruned message lists.

163
Multi-Selecthard

Your team is building an agentic workflow that interacts with internal databases. Which TWO security practices should be implemented to prevent prompt injection attacks that could lead to unauthorized data exfiltration?

Select 2 answers
A.Use hardcoded system prompts that are strictly enforced via fine-tuning.
B.Implement a strict allow-list of tools and functions the model can execute.
C.Wrap user input in XML tags or specific delimiters and instruct the model to treat content within tags as untrusted data.
D.Perform all API calls using a public-facing read-only database user.
E.Require human-in-the-loop approval for all model responses.
AnswersB, C

Limiting tool usage to a curated allow-list prevents the agent from calling unauthorized APIs or database functions. Even if an injection attack successfully manipulates the model, the model is unable to trigger unintended actions because the execution environment rejects any non-whitelisted function calls, effectively containing the potential damage.

Why this answer

Preventing prompt injection requires a defense-in-depth approach that separates user input from system instructions and enforces strict operational boundaries. By limiting the agent's capability to only necessary functions and using structured output formats, developers minimize the attack surface. These practices are critical because agentic workflows are highly susceptible to malicious instructions that override original system prompts, potentially leading to unauthorized data queries or unintended execution of dangerous internal operations.

Exam trap

Candidates mistakenly rely on the model's safety training alone, forgetting that agentic workflows need strict function allow-lists and delimiter wrapping to prevent malicious exfiltration.

164
Multi-Selectmedium

A developer is building an MCP client that connects to a remote MCP server over HTTP with Server-Sent Events. The server advertises tools that require user-specific permissions. Which TWO practices should the developer implement to keep the integration secure and functional? (Choose two.)

Select 2 answers
A.Validate the server's identity and TLS certificate, and reject connections that fail certificate verification.
B.Authenticate the connection using the server's supported OAuth flow and pass the resulting access token with each request.
C.Hard-code a shared service account token in the client binary so all users inherit the same permissions.
D.Disable TLS termination to reduce latency, since MCP messages are already structured JSON.
E.Cache all tool results on the client indefinitely so users avoid repeated permission checks.
AnswersA, B

Because the client sends credentials and receives tool results over the network, verifying the server's TLS certificate prevents man-in-the-middle attacks and credential theft. Rejecting connections that fail verification ensures the client only talks to the legitimate MCP server. This complements authentication by protecting the confidentiality and integrity of the session, which is essential when tools are permission-scoped.

Why this answer

Remote MCP servers with permission-scoped tools require two things: proof of who the user is, and protection of the channel carrying that proof. Completing the server's OAuth flow and sending the access token satisfies authentication, while validating TLS certificates prevents interception or impersonation. Shared credentials, plaintext transport, and indefinite caching all undermine per-user authorization and would make the integration both insecure and unreliable.

Exam trap

The trap here is treating MCP's JSON message format as if it provided security, leading to choices that skip authentication or TLS.

165
MCQeasy

A developer wants Claude Code to add input validation to a specific function in `src/api/users.py`. Which instruction is most likely to produce a correct, targeted change?

A.Check all the inputs everywhere and fix anything that looks wrong.
B.In `src/api/users.py`, add validation to the `create_user` function that rejects requests where `email` is missing or lacks an `@` character, and return a 400 with a clear message.
C.Make the code better.
D.Add a TODO comment in the users module about validation.
AnswerB

This instruction names the file, the function, the exact fields and conditions to validate, and the desired response behavior. That specificity lets Claude Code make a minimal, verifiable change without wandering into other files. It also establishes testable acceptance criteria, so the developer can immediately confirm the edit matches intent.

Why this answer

Targeted instructions that specify file, function, validation rules, and expected error behavior produce minimal, verifiable edits. Claude Code performs best when the scope and acceptance criteria are explicit, because it can act confidently without guessing. Vague, unbounded, or purely documentary prompts either miss the goal or create unnecessary churn across the repository.

Exam trap

The trap here is assuming any prompt mentioning validation will do, when the decisive factor is naming the exact file, function, conditions, and expected response.

166
MCQmedium

An agent developer is building a workflow where an agent must query a private database. To ensure the agent only performs read operations, which component of the Anthropic Agent SDK pattern should be prioritized?

A.A complex system prompt with strict negative constraints.
B.The Agent SDK's built-in session state management.
C.A strictly defined tool schema limited to read-only functions.
D.The model's internal safety filtering layer.
AnswerC

Defining a tool schema that only supports specific read queries ensures the model lacks the functional capacity to perform write or delete operations. By limiting the toolset, the developer hardcodes the security boundary, which is the most reliable way to prevent unauthorized data modification.

Why this answer

Tool definitions are the primary mechanism for restricting agent capabilities. By explicitly defining the schema and purpose of the tool, the developer enforces the principle of least privilege. The agent only executes what is provided in the tools list, making it vital to sanitize inputs and restrict access permissions at the API level rather than relying solely on natural language instructions within the system prompt.

Exam trap

Candidates rely on system prompt instructions to 'restrict' the agent, forgetting that system prompts are suggestions, not security boundaries; the tool schema itself must be the primary restriction mechanism.

167
MCQhard

A developer is using the Anthropic Messages API and wants to ensure that Claude's response is deterministic and reproducible for a given prompt. They set the temperature parameter to 0. However, they observe that repeated calls with the same input sometimes yield slightly different outputs. Which factor is the most likely cause of this non-determinism?

A.The model's internal computations involve non-deterministic operations such as floating-point arithmetic or parallel processing.
B.The API automatically injects a random seed into each request unless one is explicitly provided.
C.The top_p parameter is set to a value less than 1, introducing randomness even with temperature 0.
D.The temperature parameter is ignored when set to 0, and the default temperature is used instead.
AnswerA

Large language models like Claude run on distributed hardware and may use non-deterministic algorithms for efficiency, such as asynchronous floating-point operations or parallel reductions. Even with temperature 0, which selects the most likely token, the probability distribution itself can vary slightly due to these hardware-level nondeterminisms, leading to occasional different tokens. This is a known limitation.

Why this answer

Even with temperature set to 0, which makes the model choose the most probable next token, the underlying probability distribution can vary slightly between calls due to non-deterministic operations in the model's execution on distributed hardware. This includes floating-point arithmetic order and parallel processing. The API does not provide a seed parameter to enforce determinism.

Therefore, the most likely cause is the model's internal non-determinism, not misconfiguration of temperature or top_p.

Exam trap

The trap here is assuming that temperature 0 guarantees identical outputs, overlooking that hardware-level non-determinism can still cause slight variations.

168
MCQmedium

What role do 'project-aware' features play in Claude Code during a development session?

A.They allow the agent to automatically commit changes to Git.
B.They enable the agent to understand existing project conventions.
C.They force the agent to only use the most popular libraries.
D.They prevent the agent from accessing external online resources.
AnswerB

Project-awareness is the foundation for maintaining stylistic and architectural integrity. By analyzing existing patterns, the agent ensures that new code follows the established conventions of the project, which significantly reduces technical debt and ensures long-term maintainability for the team working on the codebase.

Why this answer

Project-awareness allows Claude Code to understand the structure, language conventions, and dependency relationships within the codebase. This context is crucial for generating code that is idiomatic and structurally sound. Instead of treating files as isolated units, the agent can reason about how changes in one file impact another, leading to more accurate refactoring and fewer integration issues across the entire project repository.

Exam trap

Users often treat the agent as a standalone code generator that lacks knowledge of the wider project. They fail to leverage project-awareness to ensure code style and dependency consistency.

169
MCQhard

A developer is using the Anthropic Messages API and wants to implement a retry mechanism for handling rate limit errors. They receive an HTTP 429 response with a 'retry-after' header. Which approach is the most appropriate for handling this error?

A.Immediately retry the request in a tight loop until it succeeds.
B.Wait for the number of seconds specified in the 'retry-after' header before retrying the request.
C.Switch to a different model that has higher rate limits and retry immediately.
D.Log the error and fail the request without retrying, as rate limits are permanent.
AnswerB

The 'retry-after' header provides the number of seconds to wait before making another request. Respecting this value is the correct way to handle rate limiting, as it aligns with the server's expectations and avoids overwhelming the API. This approach ensures the retry is made after the rate limit window has passed, increasing the chance of success.

Why this answer

When the API returns a 429 status with a 'retry-after' header, the best practice is to pause and wait for the specified number of seconds before retrying. This respects the server's rate limiting policy and avoids making unnecessary requests that could lead to further throttling. Immediate retries or switching models do not address the root cause.

Failing without retry is also not ideal for transient rate limits.

Exam trap

The trap here is ignoring the 'retry-after' header and either retrying immediately or giving up, rather than waiting the prescribed time.

170
MCQmedium

A developer is building a document Q&A service on the Anthropic Messages API. Their integration currently sends the entire 240,000-token knowledge base on every turn of a long conversation, and they are hitting the model's context window limit. They want to keep the full conversation history and the knowledge base available without exceeding the window. Which approach best addresses the problem?

A.Enable prompt caching on the static knowledge base so repeated prefixes are served from cache and no longer count toward the context window.
B.Increase the max_tokens parameter so Claude can internally compress the knowledge base before answering.
C.Summarize or truncate older conversation turns and retrieve only the most relevant knowledge-base chunks to include in each request.
D.Switch the request to stream: true so the server transmits the knowledge base in smaller chunks that bypass the context limit.
AnswerC

The context window is a hard cap on the total tokens sent and generated per request. Reducing the payload by condensing prior turns and injecting only relevant retrieved chunks keeps the request within the window while preserving useful information. This is the standard retrieval-plus-summarization pattern for long-running document assistants.

Why this answer

The context window is a fixed ceiling on combined input and output tokens per request. To keep a long conversation and a large corpus usable, the developer must shrink what is actually sent: condense prior turns and inject only the retrieved passages relevant to the current question. Caching and streaming change cost or delivery mechanics, not the window itself, and max_tokens governs output only.

Exam trap

The trap here is assuming prompt caching or streaming expands the usable context window rather than only changing cost, latency, or delivery mechanics.

171
MCQmedium

A team is estimating costs for a new feature that will send 1 million requests per month. Each request has a 2,000-token input and generates a 500-token output. The team wants to reduce the output token cost, which dominates the bill. Which strategy is MOST effective for reducing output token costs?

A.Enable Prompt Caching on the input to reduce the input token cost.
B.Increase the max_tokens parameter to allow longer responses.
C.Switch to a model with a larger context window to accommodate longer inputs.
D.Instruct the model to be concise and set a lower max_tokens limit to cap response length.
AnswerD

Output tokens are billed per generated token, so constraining response length directly reduces the dominant cost. Instructing conciseness and lowering max_tokens caps the worst-case output size. This approach targets the largest cost driver without sacrificing the feature's purpose, since the responses remain useful but shorter.

Why this answer

Because output tokens dominate the monthly bill, the most effective lever is limiting how many tokens the model generates. Lowering max_tokens and prompting for conciseness directly caps output length and thus cost. Optimizing input tokens or context capacity addresses smaller or irrelevant components of the bill.

Exam trap

The trap here is applying a familiar optimization like Prompt Caching reflexively, even when the scenario explicitly states that output tokens, not input tokens, are the dominant cost.

172
MCQeasy

A developer is building a Claude agent with the Anthropic SDK that must look up live order status. The agent has one tool named get_order_status, and its input_schema declares a required string property order_id. During testing, Claude returns a tool_use block with name "get_order_status" and input {"order_id": "A-4471"}. What must the developer's application do next to keep the agent loop running correctly?

A.Call the tool, then send the raw JSON output as a plain assistant text message so Claude can read the order details directly.
B.Append Claude's tool_use block to the conversation and immediately call the Messages API again with the same parameters to obtain the final answer.
C.Convert the tool_use block into a system message so the model treats the tool request as persistent instructions for the rest of the session.
D.Execute the lookup, then send a user-role message containing a tool_result block with the matching tool_use_id back to Claude.
AnswerD

This is the required half of the agent loop: the client executes the tool and returns its output to Claude as a tool_result content block inside a user-role message, with tool_use_id matching the originating tool_use block. Only then can Claude continue reasoning with the real data. This scenario's single tool call maps directly onto that pattern.

Why this answer

The agent loop alternates model reasoning with client-side tool execution. When Claude emits a tool_use block, the developer's code runs the named tool and returns the output as a tool_result content block in a user message, keyed by the same tool_use_id. That correlation lets Claude pair the result with its request and produce the final answer using real order data.

Exam trap

The trap here is assuming the model executes tools itself, when in fact the client application must run the tool and return a correlated tool_result block.

173
MCQmedium

When should you use the 'system' field versus putting instructions in the 'user' field?

A.Always put everything in the user field to keep the API call simple.
B.Use the system field for instructions that should remain constant throughout the entire conversation.
C.Use the system field only for the first message, then switch to user.
D.Use the system field to store user-specific information.
AnswerB

The system prompt is the foundation of the model's behavior. By keeping constant instructions in the system field, you ensure that the model remains aligned with its core directives regardless of how the conversation progresses, providing a stable, reliable, and secure experience for the end user.

Why this answer

The 'system' field is architecturally designed for global, persistent instructions that define the model's behavior, constraints, and identity throughout the session. The 'user' field is for specific, task-based requests. Mixing these up leads to poor instruction adherence, as the model differentiates between 'foundational rules' and 'conversational requests' based on these roles, ensuring a more robust and predictable experience.

Exam trap

Candidates often treat the system and user fields as interchangeable, failing to recognize that the model processes them with different levels of priority and foundational authority.

174
MCQmedium

A developer is building an agent using the Anthropic Python SDK. The agent's system prompt instructs it to answer questions about a company's internal policies by searching a vector database. The developer wants the agent to autonomously decide when to search and when to answer directly. Which combination of API features should the developer implement to achieve this?

A.Define a tool that queries the vector database, include it in the `tools` parameter of the Messages API request, and handle `tool_use` blocks in the response by executing the search and returning a `tool_result` block.
B.Use the `system` parameter to embed the entire vector database content as context, allowing the model to answer directly without tool use.
C.Implement a pre-processing step that always queries the vector database and appends the top result to the user's message before sending it to the model.
D.Set the `tool_choice` parameter to `{"type": "any"}` so the model always calls the vector search tool, and then parse the result to generate the final answer.
AnswerA

This approach correctly uses the Messages API's tool use capability: the model autonomously emits a `tool_use` block when it needs to search, and the developer provides the result via a `tool_result` block. The agent decides when to invoke the tool based on the system prompt and conversation context, fulfilling the requirement for autonomous decision-making.

Why this answer

The correct approach is to define the vector search as a tool, include it in the request, and handle the tool_use/tool_result cycle. This enables the model to decide when to search based on the conversation, which is the core of agentic behavior with the Anthropic SDK. Other options either force tool use, embed excessive context, or bypass the model's decision-making entirely.

Exam trap

The trap here is assuming that tool use must be forced or pre-processed, rather than allowing the model to autonomously decide when to invoke a tool based on its instructions.

175
MCQmedium

A developer is using the Anthropic SDK to build an agent that can answer questions by calling external APIs. The agent's tool returns a large JSON payload (over 100,000 tokens) as a `tool_result`. The developer notices that subsequent requests to the model fail with a context window error. What is the most effective way to resolve this issue?

A.Increase the `max_tokens` parameter in the request to accommodate the large tool result.
B.Set the `tool_choice` parameter to `{"type": "none"}` to prevent the model from calling the tool again.
C.Switch to a model with a larger context window, such as Claude 3.5 Sonnet, without modifying the tool result.
D.Truncate or summarize the tool result before including it in the conversation history.
AnswerD

The context window includes all messages, including tool results. A 100,000-token tool result will exceed the model's context limit. The most effective solution is to truncate or summarize the result to a manageable size, preserving only the necessary information. This reduces token usage and allows the conversation to continue without errors.

Why this answer

The most effective solution is to truncate or summarize the tool result before adding it to the conversation history. The context window includes all input tokens, so a massive tool result will cause errors. By reducing the size, the developer ensures the conversation stays within limits.

Other options do not address the root cause of the oversized input.

Exam trap

The trap here is confusing `max_tokens` with context window management, or assuming that switching models alone can handle arbitrarily large inputs.

176
MCQmedium

A developer is building a Claude-powered assistant that uses a tool called 'search_knowledge_base'. The tool definition includes a required parameter 'query' and an optional 'max_results' parameter. During testing, Claude sometimes calls the tool without providing 'max_results'. What should the developer do to ensure the tool call is handled correctly?

A.Add a description to the tool definition instructing Claude to always include 'max_results' with a value of 10.
B.In the tool implementation, check if 'max_results' is present and apply a sensible default value if it is missing.
C.Change the tool to accept a single string parameter that combines the query and max_results, such as 'query|max_results'.
D.Make 'max_results' a required parameter in the tool definition so Claude must always provide it.
AnswerB

Since 'max_results' is optional, Claude may omit it. The tool implementation should detect its absence and use a default value, such as 5 or 10, to ensure consistent behavior. This keeps the tool definition accurate while handling the missing parameter gracefully, preventing errors and providing predictable results.

Why this answer

Because 'max_results' is optional, Claude may omit it. The tool implementation should check for its presence and apply a default value, ensuring the tool works correctly regardless. This approach maintains the intended flexibility of the tool definition and prevents runtime errors.

Other options either incorrectly change the tool contract or rely on unreliable prompting.

Exam trap

The trap here is assuming that Claude will always include optional parameters or that a description can enforce it, rather than handling missing optional parameters with defaults in the tool code.

177
MCQmedium

When an agent is asked to perform a complex task, which approach minimizes latency while ensuring accuracy?

A.Requesting the model to solve the entire problem in one turn.
B.Decomposing the task into smaller, chained agent actions.
C.Using the maximum possible context window for every request.
D.Adding 'think step-by-step' to the final response request.
AnswerB

Task decomposition allows the model to verify its progress at each step. By chaining actions, the agent manages complexity effectively, reducing the likelihood of catastrophic errors while making the entire process more transparent and easier to monitor for performance.

Why this answer

Breaking down a complex task into smaller, sequential tool-use steps allows the model to reason incrementally. This approach reduces the load per turn, makes errors easier to isolate, and improves overall accuracy by allowing the model to process feedback at each stage. While it increases the total number of turns, it is often faster than forcing a single, massive inference that may fail entirely.

Exam trap

Candidates frequently choose a single massive prompt to solve complex tasks, believing it saves time, rather than breaking the problem down into chained agent actions.

178
MCQmedium

A developer is building a real-time translation service that must respond within 1 second for short phrases. The service will handle thousands of requests per hour. Quality is important, but latency and cost are critical. Which Claude model is the most appropriate?

A.Claude 3 Haiku
B.Claude 3 Opus
C.Claude 3.5 Sonnet
D.Claude 3 Sonnet
AnswerA

Claude 3 Haiku is designed for speed and cost efficiency, making it ideal for real-time translation of short phrases. It can deliver responses within the 1-second latency requirement and handle thousands of requests per hour at a low cost. For straightforward translation, its quality is sufficient, and it meets the critical latency and cost constraints.

Why this answer

Claude 3 Haiku is optimized for low latency and low cost, making it the best fit for real-time translation of short phrases. It can meet the 1-second response requirement and handle high throughput without excessive expense. More powerful models like Sonnet or Opus would add latency and cost without necessary quality gains for this straightforward task.

Exam trap

The trap here is assuming that higher quality models are always needed, ignoring that latency and cost constraints may make a faster, cheaper model the correct choice.

179
MCQeasy

When calculating the estimated cost of a project using Claude, which metric is used by Anthropic to measure the volume of data processed and generated?

A.Characters (including spaces).
B.Total words in the prompt.
C.Tokens.
D.API call duration in seconds.
AnswerC

Tokens are the atomic unit of processing for Claude. Anthropic's pricing is strictly defined as a cost per million tokens. This includes both the input tokens sent by the user and the output tokens generated by the model. This is the standard metric for all cost and performance calculations.

Why this answer

Anthropic, like most LLM providers, uses 'tokens' as the fundamental unit of measurement for billing. Tokens represent chunks of text (roughly 3/4 of a word). Cost management involves estimating the total number of input tokens (prompt) and output tokens (response).

Understanding this unit is the first step in any cost-estimation exercise for a developer using Claude.

Exam trap

Candidates sometimes confuse billing metrics like word counts, character lengths, or API call frequencies with the actual fundamental unit used by Anthropic to measure data volume.

180
MCQhard

An agent is designed to manage a user's calendar. During a session, the user asks to 'Schedule a meeting for tomorrow at 2 PM.' The agent must first check for conflicts and then create the event. How should the developer handle the 'observation' phase in the SDK to ensure the agent doesn't double-book?

A.Use a single tool that both checks for conflicts and creates the event.
B.Feed the 'check_conflicts' tool_result back to the model before it issues 'create_event'.
C.Parallelize both calls to improve the speed of the calendar update.
D.Prompt the model to assume there are no conflicts to save tokens.
AnswerB

By returning the conflict data as a 'tool_result', you allow Claude to see the current state of the calendar. The model can then reason: 'I see a conflict at 2 PM, I should not call create_event; instead, I will suggest 3 PM to the user.' This creates a safer and more helpful agent.

Why this answer

The observation phase involves feeding the actual results of a tool (the calendar check) back into the agent's context. If the agent doesn't receive the output of the conflict check before it attempts to create the event, it is 'flying blind.' Ensuring the orchestrator waits for the first tool's result before proceeding is key to logical consistency.

Exam trap

Candidates often assume the model can issue multiple tool calls simultaneously without waiting for results, missing the crucial step of feeding the first observation back into the context.

181
MCQeasy

A developer needs Claude to transform a list of product descriptions into a fixed XML schema that a downstream parser expects. The model sometimes adds a friendly introductory sentence before the XML. Which change most directly eliminates the extra prose?

A.Add a clear instruction that the response must begin immediately with the opening XML tag and contain nothing else.
B.Shorten the input descriptions so the model has less to respond to.
C.Increase temperature so the model varies its phrasing and may omit the introduction.
D.Ask the model to explain each transformation step before producing the XML.
AnswerA

A direct, explicit constraint on where the output starts is the most reliable way to suppress preamble. The model follows concrete formatting rules well when they are stated unambiguously and placed with the other output requirements. This addresses the exact failure mode without changing the task itself.

Why this answer

Explicit output-format constraints are the direct fix for unwanted preamble. Telling the model exactly where the response must begin, and that nothing may precede it, removes the ambiguity that lets conversational framing leak in. Temperature, input length, and added reasoning all fail to target the specific behavior.

Exam trap

The trap here is reaching for a generation parameter like temperature to fix a formatting problem that is actually solved by an explicit output constraint.

182
Multi-Selectmedium

A developer is building a Claude-powered agent that uses the Anthropic API with a tool-use loop. The agent can invoke a `fetch_url` tool that retrieves the contents of any URL supplied by the model. During a red-team exercise, an attacker embeds hidden instructions in a page the agent fetches, causing the agent to call `fetch_url` again with an attacker-controlled URL containing sensitive query parameters. Which TWO controls best reduce this tool-use loop risk? (Choose two.)

Select 2 answers
A.Log every tool invocation with its arguments and retain the logs for forensic analysis after an incident.
B.Increase the model's temperature setting so it is less likely to follow deterministic hidden instructions embedded in fetched content.
C.Restrict the `fetch_url` tool to an allowlist of approved domains and schemes, and reject any URL not on that list before the tool executes.
D.Raise the `max_tokens` value on each API request so the agent has enough context to distinguish legitimate instructions from injected ones.
E.Treat all tool output as untrusted data, and require human confirmation before any tool call that sends data to an external destination.
AnswersC, E

An allowlist constrains the tool's reachable surface to known-good destinations, so attacker-injected URLs cannot be fetched. It directly blocks the second-stage exfiltration step because the malicious URL is rejected before any network call, regardless of how persuasive the injected text appears to the model.

Why this answer

The attack works because fetched content is untrusted and can carry instructions that the agent treats as goals. Effective defense combines limiting where the tool can go with treating tool output as data rather than commands and requiring human approval for outbound calls. These controls stop the second-stage request before it can carry sensitive parameters to an attacker-controlled endpoint.

Exam trap

The trap here is assuming that model-side settings such as temperature or token limits can defend against prompt injection, when the durable fix is constraining the tool's capabilities and the trust level of its output.

183
MCQeasy

What is the primary purpose of the 'Claude Code' tool in a professional development environment?

A.To act as a cloud-based storage service for source code.
B.To automate the deployment of applications to production.
C.To assist with coding tasks by interacting with the local project.
D.To replace the need for version control systems like Git.
AnswerC

The primary goal of Claude Code is to act as a powerful assistant that understands the project context, navigates the file system, and executes code-related tasks. It enables developers to accelerate their workflow by automating repetitive or complex tasks while maintaining project alignment and high quality.

Why this answer

Claude Code functions as an autonomous, project-aware coding assistant designed to handle complex development tasks directly within the terminal. By integrating deeply with the codebase and local tools, it enhances developer productivity, handles repetitive tasks, and assists with complex debugging. Its professional utility lies in its ability to understand existing context, apply changes, and iterate based on feedback, bridging the gap between high-level intent and low-level code implementation.

Exam trap

Candidates often confuse Claude Code with a general-purpose chatbot. They fail to recognize its specific role as a terminal-integrated tool that directly manipulates the local project file system.

184
MCQeasy

Which of the following scenarios describes the most effective use of Prompt Caching for cost management?

A.A chatbot where every user query is unique and no history is maintained.
B.A translation service that processes single words one at a time.
C.An AI assistant that references a 50-page technical manual for every user query.
D.A daily report generator that uses a completely different dataset every morning.
AnswerC

This is the ideal use case for prompt caching. The 50-page manual acts as a large, static prefix that is sent with every request. By caching the manual, the developer only pays the full input price once, and all subsequent queries only pay for the much cheaper cache read tokens, resulting in massive long-term cost savings.

Why this answer

Prompt caching is most effective when a large amount of static information is reused across many different requests. It allows the model to 'remember' the prefix of a prompt, significantly reducing the cost of processing that prefix in subsequent calls. Identifying workloads with high prefix overlap is key to maximizing the financial benefits of this feature in production environments.

Exam trap

Candidates often apply prompt caching to dynamic or highly variable user inputs, failing to realize that caching is only financially beneficial when the same prefix is reused across many requests.

185
MCQmedium

An enterprise is migrating a document processing pipeline that handles 50,000 PDFs daily. Each PDF is converted to text (approx. 2,000 tokens) and requires a summary. The project has a strict budget. Which approach provides the most significant cost reduction while utilizing Claude 3.5 Sonnet?

A.Implementing client-side compression on the PDF text before sending.
B.Using the Anthropic Batch API for asynchronous processing.
C.Switching the entire pipeline to Claude 3 Haiku.
D.Reducing the 'max_tokens' parameter to 50 for every summary.
AnswerB

The Batch API allows developers to submit large groups of requests that are processed within a 24-hour window at a 50% discount compared to standard real-time API prices. For document processing pipelines where immediate results are not required, this is the most effective way to utilize the reasoning power of Claude 3.5 Sonnet within a limited budget.

Why this answer

Cost management in high-volume environments often involves leveraging specific API features designed for non-latency-sensitive workloads. The Anthropic Batch API is specifically designed for processing large volumes of data asynchronously at a significantly reduced price point. This allows developers to use high-intelligence models like Sonnet 3.5 for complex tasks while adhering to strict budgetary constraints that would be exceeded by standard API calls.

Exam trap

Candidates often recommend expensive real-time API calls for massive, non-urgent workloads instead of leveraging asynchronous processing features designed for cost savings.

186
MCQmedium

A developer wants to implement a 'Summary' feature for a long conversation history. As the conversation grows, the cost of sending the entire history with every new message increases. What is the most cost-effective architectural pattern to handle this?

A.Always send the full conversation history to maintain maximum context.
B.Use Prompt Caching for the entire dynamic conversation history.
C.Implement a sliding window that only sends the last 5 messages.
D.Periodically summarize the history and use the summary as context.
AnswerD

This pattern, known as context distillation or compression, involves replacing older messages with a concise summary. This drastically reduces the number of input tokens sent in subsequent requests while preserving the important information, making it the most cost-effective way to handle long-running, context-heavy sessions.

Why this answer

Managing long-running conversations requires balancing context and cost. Periodically summarizing the previous conversation and replacing the detailed history with that summary (Context Compression) keeps the input token count low. This limits the linear growth of costs as the session continues, ensuring the application remains affordable even during extended user interactions.

Exam trap

Candidates frequently select stateless strategies like completely truncating old messages or sending raw unlimited histories, overlooking how periodic summarization preserves crucial long-term context while strictly controlling linearly increasing token costs.

187
MCQmedium

A developer is comparing the cost of using Claude 3 Opus versus Claude 3.5 Sonnet for a task that requires complex reasoning. The task involves processing 1,000 requests, each with 500 input tokens and 200 output tokens. Which statement accurately reflects the cost consideration?

A.The cost is identical because both models are billed at the same rate for input and output tokens.
B.Claude 3 Opus is cheaper for output tokens but more expensive for input tokens compared to Claude 3.5 Sonnet.
C.Claude 3 Opus is always more cost-effective for complex reasoning because it requires fewer tokens to achieve the same result.
D.Claude 3.5 Sonnet is less expensive per token than Claude 3 Opus, but it may require more tokens to match Opus's reasoning quality.
AnswerD

Claude 3.5 Sonnet has a lower per-token cost than Claude 3 Opus, but for highly complex reasoning, it might need more tokens or additional prompting to achieve similar results. The total cost depends on both the per-token price and the number of tokens required. This statement accurately captures the trade-off developers must evaluate.

Why this answer

When choosing between Claude 3 Opus and Claude 3.5 Sonnet for complex reasoning, developers must weigh per-token cost against the number of tokens needed. Sonnet is cheaper per token, but Opus may deliver better results with fewer tokens. The correct statement acknowledges that Sonnet is less expensive per token but might require more tokens to match Opus's quality, making total cost dependent on the specific task and prompt engineering.

Exam trap

The trap here is assuming that a more powerful model like Opus is always more expensive overall, or that a cheaper model like Sonnet will always be cheaper in total, without considering how token usage might differ between models.

188
MCQmedium

When Claude Code is asked to perform a complex, multi-file task, how does it typically approach the problem?

A.It modifies all files simultaneously in one batch.
B.It breaks the task into logical, iterative steps.
C.It asks the developer to perform the file edits.
D.It ignores file dependencies to save time.
AnswerB

By breaking tasks into smaller steps, the agent can verify each modification. This allows it to identify issues early, manage project complexity effectively, and provide the developer with clear progress updates, ensuring that the overall goal is achieved safely and with higher accuracy than a batch-edit approach.

Why this answer

Claude Code is designed to decompose complex tasks into smaller, logical steps. It first analyzes the codebase, plans the necessary file modifications, and then executes these changes incrementally, often verifying intermediate steps through test runs. This iterative approach is crucial for managing complexity.

By breaking large tasks into digestible pieces, the agent can maintain state, identify conflicts, and ensure that each part of the refactor or implementation is correct before proceeding, reducing the risk of errors.

Exam trap

Many candidates believe the model attempts to rewrite an entire multi-file project in a single monolithic generation step, leading to high failure rates.

189
Multi-Selectmedium

A developer is building an agent that must decide when to call tools versus when to respond directly to the user. The agent uses the Anthropic SDK with tool definitions supplied in the request. Which TWO behaviors correctly describe how the model handles tool use in this setup? (Choose two.)

Select 2 answers
A.The model can return a normal text response instead of a tool_use block when it determines no tool is required.
B.The model emits a tool_use block containing the tool name and input arguments when it decides a tool is needed.
C.The model requires the tool's input_schema to be omitted so it can choose parameters freely.
D.The model automatically retries a failed tool call until it succeeds without developer intervention.
E.The model executes the tool itself and returns the result directly in the same response.
AnswersA, B

Tool use is optional from the model's perspective. If the user's request can be answered from existing context, Claude may reply with text and no tool_use block. The developer's loop should handle both outcomes, which is why checking the stop reason and content blocks is essential.

Why this answer

Claude decides whether to call a tool and, if so, emits a tool_use block with the name and arguments; it may also answer directly with text when no tool is needed. The application, not the model, performs execution and returns results, which is why the agent loop must handle both tool_use and plain text outcomes.

Exam trap

The trap here is assuming the model runs the tool and returns its output, when it only requests the tool and waits for the application to supply the result.

190
Multi-Selecthard

A developer is working on a large refactoring task with Claude Code. They want to ensure the agent has sufficient context about the project's architecture and coding standards. Which TWO of the following are recommended ways to provide this context? (Choose two.)

Select 2 answers
A.During the session, explicitly reference key files like `src/core/engine.ts` and ask Claude to read them for context.
B.Run `claude --add-context architecture.md` to load the architecture document into the session.
C.Place a `CLAUDE.md` file at the repository root with an overview of the architecture and coding conventions.
D.Include detailed comments in each source file describing the overall system design.
E.Use the `/context` command to manually upload a PDF of the architecture diagram.
AnswersA, C

Claude Code can read files on demand when prompted. By explicitly referencing important files, you direct the agent to incorporate their contents into its working context. This is useful for targeted context, especially when not all files are relevant. It complements `CLAUDE.md` for dynamic needs.

Why this answer

The recommended ways to provide project context are using a `CLAUDE.md` file at the root and explicitly referencing key files during the session. `CLAUDE.md` offers persistent, project-wide guidance, while on-demand file references allow the agent to focus on specific areas. Other options like `/context` command, `--add-context` flag, or relying solely on comments are not supported or effective.

Exam trap

The trap here is thinking that Claude Code has special commands like `/context` or `--add-context` for loading documents, when in fact `CLAUDE.md` and direct file references are the intended mechanisms.

191
MCQmedium

Why is it important to use the Claude Code CLI within an initialized Git repository?

A.It is required for the agent to connect to the internet.
B.It provides a safety net for reverting changes.
C.It automatically deploys the code to GitHub.
D.The agent refuses to work without a remote origin defined.
AnswerB

Git allows the developer to manage the changes made by the agent. By tracking the history, the developer can quickly review, verify, and rollback changes as needed. This is essential for maintaining project quality and preventing accidental regressions caused by an agent's code updates.

Why this answer

Git integration is fundamental to using Claude Code effectively. It allows the agent to track changes, easily undo mistakes, and see the 'diff' of what it has modified. This safety net encourages experimentation and allows the developer to revert changes easily if the agent's output is not what they intended, creating a professional development workflow that minimizes risk during automated code modifications.

Exam trap

Candidates might think Git is only for version control history, ignoring its crucial role as a safety net that lets automated agents rollback edits.

192
MCQmedium

A company needs to process 10 million short customer feedback snippets to identify 'bug reports' vs 'feature requests'. Speed and budget are the primary constraints, while the classification logic is straightforward. Which model provides the best throughput-to-cost ratio?

A.Claude 3.5 Sonnet
B.Claude 3 Opus
C.Claude 3 Haiku
D.Claude 2.0
AnswerC

Haiku is the fastest and least expensive model, making it the superior choice for high-volume classification. It can process millions of tokens for a very low cost while maintaining the accuracy needed for simple tasks like distinguishing between bug reports and feature requests. This maximizes the return on investment for the company's data processing pipeline.

Why this answer

When dealing with massive datasets and simple logic, the most important metric is the cost per million tokens. Claude 3 Haiku is specifically optimized for these 'utility' tasks, offering high throughput and the lowest pricing in the Claude 3 family. Selecting a larger model for such a simple, high-volume task would lead to unnecessary expenditures without providing a noticeable improvement in classification quality.

Exam trap

Candidates often default to the most capable model (like Sonnet or Opus) for simple classification tasks, ignoring that Haiku is specifically engineered for high-throughput, low-cost utility operations.

193
MCQmedium

A developer is building a customer-support bot with the Claude Messages API. The system prompt instructs Claude to answer concisely, but the bot also needs to return a stable JSON object with the fields 'intent', 'sentiment', and 'reply' on every turn. Where should the instruction to produce this JSON object be placed to most reliably control the response format?

A.In the system prompt, placed after the conciseness instruction.
B.Only in the first assistant message as an example.
C.Appended to each user message as a trailing reminder.
D.In a separate follow-up call after Claude replies in prose.
AnswerA

The system prompt is the highest-leverage location for persistent behavioral rules, including output format. Putting the JSON requirement there, after the tone rule, ensures it applies to every turn of the conversation without the caller having to resend it in each user message.

Why this answer

The system prompt is the correct place because it is sent with every Messages API request and reliably conditions all assistant turns. Interface requirements such as JSON keys should live there so they are not diluted by user content. Putting format rules in the user turn or in a lone assistant example makes them fragile and easy to override.

Exam trap

The trap here is assuming that any instruction in the user message carries the same weight as a system-prompt instruction, when system-level guidance is what persistently controls format across turns.

194
MCQmedium

A developer is building a Claude-powered agent that calls an internal `search_customer_notes` tool. The agent runs with a system prompt that includes a user-supplied `account_id`. A security review finds that an attacker can craft a prompt injection that convinces Claude to call the tool with a different `account_id` than the one in the system prompt. Which control most directly prevents this privilege escalation while keeping the agent functional?

A.Increase the tool's rate limit so that mass enumeration of account IDs is impractical.
B.Use a more capable Claude model with stronger instruction-following to reduce injection success.
C.Add a sentence to the system prompt instructing Claude to never change the `account_id` value.
D.Pass the authenticated `account_id` from the server-side session into the tool implementation and ignore any `account_id` Claude supplies in tool arguments.
AnswerD

The tool should derive authorization context from the server session, not from model output. Claude can be manipulated through prompt injection, so any parameter that controls data access must be bound server-side. Ignoring the model-supplied account_id and using the authenticated session value ensures the tool can only read records the caller is entitled to, while still allowing Claude to decide when to call the tool.

Why this answer

Authorization decisions must be enforced outside the model. Because Claude can be manipulated by prompt injection, any parameter that determines which records are accessible must come from a trusted server-side session rather than from model output. Binding the account_id server-side keeps the agent useful while ensuring it cannot be tricked into reading another customer's notes.

Exam trap

The trap here is assuming that a stronger system prompt or a more capable model turns a model-supplied parameter into a trusted authorization input.

195
MCQeasy

A developer is using the Messages API to have Claude return a JSON object describing a product. The response sometimes includes a conversational preamble such as 'Here is the JSON you requested:' before the object, which breaks the downstream parser. What is the most reliable way to eliminate the preamble?

A.Increase max_tokens so the model has enough room to output the full JSON without truncation.
B.Use the assistant turn prefill with an opening curly brace so the response must continue from that character.
C.Set the temperature to zero so the model always produces the same output.
D.Add 'Do not include any preamble' to the system prompt and hope the model complies.
AnswerB

Prefilling the assistant message with an opening brace constrains generation to continue from that exact point, so no preamble can appear before the JSON. The API returns the prefilled text plus the continuation, giving a clean object. This is a structural guarantee rather than a probabilistic instruction, which is why it is the most reliable fix.

Why this answer

Prefilling the assistant turn with an opening brace removes the space where a preamble could be generated, forcing the response to begin mid-JSON. Instructions and temperature settings only bias behavior probabilistically, and token limits address length rather than structure. Prefill is the only option that structurally guarantees the response starts with the expected character.

Exam trap

The trap here is treating a prompt instruction like 'no preamble' as equivalent to a structural constraint, when only prefill actually prevents leading text.

196
MCQhard

A developer is writing a tool called `send_email` for a Claude-based assistant. The tool's description currently reads: 'Sends an email.' During testing, Claude calls the tool for tasks like drafting an email or checking inbox contents. Which revision to the tool definition most directly improves Claude's decision about when to invoke `send_email`?

A.Set `input_schema` to `type: "object"` with `additionalProperties: false` to lock down the parameter set.
B.Rename the tool from `send_email` to `email` and rely on the shorter name to signal its broader role.
C.Rewrite the description to state precisely what the tool does, when to use it, and when not to use it, such as only when the user explicitly asks to send a message.
D.Add a `required` array listing `to`, `subject`, and `body` in the input schema.
AnswerC

Tool descriptions are the primary signal Claude uses to decide whether a tool applies. Stating the exact action, the trigger conditions, and explicit exclusions such as drafting or inbox retrieval gives the model the boundaries it needs. This directly addresses the over-invocation observed in testing, because Claude now has criteria to distinguish sending from drafting or reading messages.

Why this answer

The model chooses tools primarily from their descriptions, so a vague one like 'Sends an email' invites calls for drafting or inbox checks. A description that states the exact action, the conditions that should trigger it, and explicit non-triggers such as drafting or retrieving messages gives Claude the criteria it needs. Schema changes and renaming affect argument validity or naming, not the decision of when the tool applies.

Exam trap

The trap here is assuming stricter schemas or shorter tool names fix misuse, when the real lever is the natural-language description that guides tool selection.

197
MCQhard

Refer to the exhibit. An application suddenly begins receiving this error in production. What is the most immediate security-focused action to take?

A.Retry the request with an exponential backoff strategy.
B.Immediately revoke the current API key and generate a new one.
C.Check if the API billing limit has been reached.
D.Hardcode the master account key to restore service quickly.
AnswerB

Revoking a potentially compromised key is the standard response to an 'Invalid API key' error in production. This stops any unauthorized use of the credentials, protecting the organization from further risk. Replacing it with a new, securely managed key restores service while neutralizing the threat of an active attacker.

Why this answer

An authentication error indicates that the currently used key is either revoked, expired, or invalid. In a production environment, this is a major red flag that could signal a credential compromise. The most responsible action is to treat the key as compromised, revoke it immediately in the console, and rotate to a new key to protect the integrity of the application's API interactions.

Exam trap

Candidates often suggest debugging the code or checking network connectivity first, failing to recognize that an auth error in production is a high-priority security incident requiring immediate key rotation.

198
MCQmedium

A developer is building a real-time chat application where users expect responses in under two seconds. The prompts are short (under 200 tokens) and the responses are typically one or two sentences. Which Claude model should the developer choose to optimize for latency and cost?

A.Claude 3.5 Sonnet
B.Claude 3 Haiku
C.Claude 3 Opus
D.Claude 2.1
AnswerB

Claude 3 Haiku is Anthropic's fastest and most cost-effective model, optimized for near-instant responses on simple tasks. It handles short prompts and brief outputs efficiently, meeting the sub-two-second latency requirement while minimizing cost. For a real-time chat application with straightforward interactions, Haiku provides the best balance of speed and affordability.

Why this answer

For a real-time chat application with short prompts and simple responses, latency and cost are critical. Claude 3 Haiku is specifically designed to be the fastest and most affordable model in the Claude 3 family, making it ideal for this use case. More powerful models like Opus or Sonnet would add unnecessary cost and latency without providing meaningful benefits for such straightforward interactions.

Exam trap

The trap here is assuming that a newer or more powerful model like Claude 3.5 Sonnet is always the best choice, when in fact Haiku is optimized for speed and cost in simple, high-volume scenarios.

199
MCQeasy

What is the primary security benefit of using the Anthropic API in a Virtual Private Cloud (VPC) environment with a Private Link?

A.It increases the throughput of API requests.
B.It eliminates the need for API keys entirely.
C.It ensures that traffic remains within a private network path, avoiding the public internet.
D.It automatically encrypts the model responses locally.
AnswerC

Keeping traffic off the public internet prevents exposure to common network-based attacks. By routing requests through a private link, the traffic is encapsulated and remains protected by the cloud provider’s private network infrastructure, satisfying the highest levels of security and compliance requirements for sensitive enterprise data transfers.

Why this answer

Using a Private Link connects your VPC directly to the API service over a private network connection, bypassing the public internet. This significantly reduces the attack surface, protects data from man-in-the-middle attacks, and ensures that traffic remains within the provider's backbone. It is a fundamental architectural requirement for enterprises that must adhere to strict regulatory compliance regarding data isolation and network security.

Exam trap

Candidates confuse Virtual Private Cloud endpoints with application-layer encryption, overlooking how Private Link specifically isolates traffic from the public internet.

200
MCQeasy

What is the primary purpose of the 'claude' command-line interface tool?

A.To host a web server for displaying AI documentation.
B.To act as an autonomous agent for local code interaction.
C.To manage remote server configurations exclusively.
D.To replace the git version control system.
AnswerB

Claude Code is specifically architected as an autonomous agent capable of reading, writing, and executing code within a local environment. It functions as a partner to the developer, performing file operations and shell commands based on natural language instructions provided through the terminal interface.

Why this answer

The primary objective of the Claude Code CLI is to serve as an interactive AI agent that integrates directly into the developer's terminal. It bridges the gap between natural language intent and local file manipulation. By understanding the tool's purpose, developers can leverage it to automate repetitive coding tasks, debug issues in real-time, and maintain context across a project, significantly reducing the cognitive load involved in navigating complex codebases.

Exam trap

Many candidates confuse the CLI tool with a cloud-based API testing utility or a remote deployment dashboard rather than its actual role as an interactive local code agent.

201
MCQeasy

Which field in the tool definition is primarily responsible for helping Claude determine when a specific tool is appropriate to use for a given user query?

A.name
B.description
C.input_schema
D.tool_choice
AnswerB

The description field provides a natural language explanation of what the tool does. Claude uses this text to perform semantic mapping between the user's request and the tool's capabilities. Providing a clear, detailed description is the most effective way to improve the reliability of tool selection by the model.

Why this answer

The 'description' field is the primary source of information for Claude's reasoning. It provides the semantic context that the model uses to match user intent to the available functions. A well-written description is crucial for accurate tool selection and helps the model understand the utility and limitations of each tool.

Exam trap

Candidates frequently assume the 'name' or 'parameters' field is the primary driver for tool selection, overlooking the fact that the 'description' is the critical semantic signal Claude uses to evaluate relevance.

202
MCQmedium

A developer is building a customer support agent using the Anthropic SDK. The agent must call a 'get_order_status' tool that requires an order ID. During testing, the agent sometimes invents plausible-looking order IDs instead of asking the user. Which change to the tool definition will most directly reduce this behavior?

A.Increase the 'max_tokens' setting so the model has more room to reason before emitting the tool_use block.
B.Set 'input_schema' to an empty object so the model is forced to ask the user for all parameters.
C.Add a detailed 'description' field to the tool that explains when to use it and explicitly instructs the model not to guess missing parameters.
D.Change the tool name to 'lookup_order' so it sounds more like a retrieval operation.
AnswerC

The tool description is the primary place to encode usage guidance, including when the tool is appropriate and how to handle missing required inputs. Stating that the model must not fabricate an order ID and should instead ask the user directly addresses the observed failure mode without altering the schema or adding external logic.

Why this answer

The tool description is the model's instruction manual for when and how to invoke a tool. By explicitly stating that the agent must request an order ID from the user rather than guessing, the developer directly targets the observed hallucination. Schema changes, token limits, and renaming do not convey that behavioral rule.

Exam trap

The trap here is assuming that changing a tool's name or schema shape will fix a behavioral problem that is actually caused by missing usage instructions in the description.

203
MCQmedium

When Claude Code is asked to perform a complex task that might take a long time, how does the agent manage the execution flow to ensure task completion?

A.It executes the entire task as a single, atomic operation.
B.It prompts the user to verify every single line of code generated.
C.It breaks the goal into smaller, verifiable steps.
D.It pauses and requests human assistance for every file change.
AnswerC

Decomposition allows the agent to handle complex problems effectively. By isolating each step, Claude Code can validate outcomes, handle errors locally, and maintain context, which is fundamental to its ability to perform high-quality, reliable code generation and maintenance in sophisticated software projects.

Why this answer

Claude Code utilizes a sub-tasking strategy for complex requests. It breaks down large goals into smaller, manageable steps, executing them sequentially or in parallel when appropriate. This modular approach allows the agent to maintain focus, track progress across multiple files, and verify results at each step, significantly increasing the probability of successful task completion compared to attempting large, atomic changes without validation.

Exam trap

Candidates often think the agent performs large tasks as a single atomic operation. This leads to poor prompt design, as they fail to request step-by-step breakdowns for complex, multi-file refactoring tasks.

204
MCQhard

Refer to the exhibit. What is the primary technical benefit of the developer pre-filling the assistant's message with an opening curly brace?

A.It reduces the API latency by pre-calculating the first token of the response.
B.It forces Claude to skip conversational filler and start directly with the data payload.
C.It triggers the model's 'JSON mode' which is a dedicated high-performance inference path.
D.It allows the developer to bypass the system prompt entirely for better token economy.
AnswerB

When the assistant message is pre-filled, Claude continues the generation from that exact point. This effectively bypasses the model's tendency to include polite introductions or explanations, ensuring the output is immediately parseable by downstream code. This is the standard method for forcing strict adherence to structured data formats.

Why this answer

Pre-filling the assistant's response is a powerful technique to steer Claude toward a specific output format, such as JSON. By providing the start of the response, the developer constrains the model's first few tokens, which significantly increases the likelihood that the entire following sequence adheres to the desired schema and avoids unwanted conversational preamble like 'Certainly, here is the JSON'.

Exam trap

Candidates often confuse pre-filling with system prompts or stop sequences, missing how constraining the assistant's initial opening token directly enforces strict formatting schemas like JSON.

205
MCQmedium

Refer to the exhibit. An internal tool is configured to send user input directly to the API. Which security improvement should be applied to the architecture?

A.Increase the 'max_tokens' to ensure the deletion script is generated fully.
B.Implement a middleware to sanitize and block dangerous system commands in the user input.
C.Use a more advanced model for the same request.
D.Add a disclaimer in the system prompt that deleting files is prohibited.
AnswerB

A middleware layer acts as a gatekeeper, scanning for malicious keywords or intent that could result in dangerous system-level operations. By blocking requests that demand file deletions or system modifications before they reach the model, you prevent the LLM from being used as a weapon against the infrastructure.

Why this answer

The architecture is currently wide open to dangerous instruction execution. The application must include an input-validation or intent-classification layer before the request reaches the LLM. By checking if the request involves sensitive or destructive actions—such as file system deletion—the tool can block the request entirely, ensuring the model is never used to generate commands that could cause catastrophic system damage.

Exam trap

Candidates often assume that the LLM itself will act as a sufficient security filter, failing to realize that an LLM is easily manipulated into executing harmful system commands.

206
Multi-Selectmedium

Which THREE of the following are essential components of a well-defined tool for an agent? (Select exactly 3)

Select 3 answers
A.A clear, descriptive name and description.
B.A JSON schema defining the required parameters.
C.The implementation logic (function code) to perform the task.
D.A pre-trained model checkpoint for the tool.
E.A hardcoded response for every possible input.
AnswersA, B, C

The model relies entirely on the name and description to decide when a tool is appropriate for a given task. If these are vague or misleading, the model will struggle to select the correct tool at the right time, leading to execution errors.

Why this answer

Robust tools require clear documentation for the model, a strict schema for input validation, and a well-defined execution function. These three components work in concert to ensure the agent understands what the tool does, how to provide data to it, and how the system handles the resulting logic. Missing any of these leads to unpredictable agent behavior and failure to resolve tasks.

Exam trap

Candidates often select only the schema or the name and description, forgetting that the actual function implementation code is equally vital for the system to execute the requested task.

207
Multi-Selectmedium

A developer is tuning a customer-feedback classifier on the Anthropic API. The model currently mislabels sarcastic complaints as praise. The developer wants to improve accuracy using few-shot examples in the prompt. Which TWO practices should be applied? (Choose two.)

Select 2 answers
A.Use the same example repeatedly so the model strongly memorizes the sarcasm pattern.
B.Place all examples after the user's feedback text so the model reads the target first.
C.Include examples that cover the difficult sarcastic cases alongside typical positive and negative examples.
D.Provide as many examples as the context window allows, prioritizing quantity over diversity.
E.Wrap each example in XML tags that separate the input text from its label.
AnswersC, E

Few-shot examples teach the model the decision boundary, so including sarcastic edge cases directly targets the observed failure mode. Examples that only show easy positives and negatives leave the model guessing on sarcasm. Covering the difficult cases, with correct labels, gives the model concrete evidence of how to classify them, which is the most direct way to reduce the mislabeling described.

Why this answer

Effective few-shot prompting targets the observed failure mode with representative, correctly labeled edge cases, and structures each example so the input-to-label mapping is unambiguous. Covering sarcastic cases fixes the specific boundary error, while delimiter tags prevent confusion between example text and labels. Volume, repetition, and post-target placement do not broaden the decision boundary and can introduce new biases.

Exam trap

The trap here is assuming that more examples or repeating one example improves classification, when the real gains come from covering the difficult cases with clearly delimited, correctly labeled examples.

208
MCQhard

Refer to the exhibit. Claude Code has encountered a test failure while implementing a feature. How should the developer effectively guide the agent to resolve this error using the agent's internal capabilities?

A.Manually apply the fix to the code and restart the Claude Code session.
B.Instruct the agent to read the logs and apply a fix based on the stack trace.
C.Disable the agent's ability to run commands to prevent further errors.
D.Run 'npm install' to ensure all dependencies are resolved before proceeding.
AnswerB

Providing the error log as context allows Claude Code to trace the logic error back to the source. The agent can then identify the missing function definition or improper import, suggest a code modification, and execute the test suite again to verify the resolution effectively.

Why this answer

Claude Code excels at iterative debugging by leveraging its ability to read execution logs and apply fixes directly to code. The most efficient approach involves instructing the agent to analyze the stack trace, locate the undefined function call, and propose a corrective patch. This leverages the agent's autonomous capability to reason about project architecture rather than requiring manual intervention, keeping the development loop self-contained and highly efficient.

Exam trap

Developers frequently try to fix errors manually after seeing a stack trace. They miss the opportunity to have the agent analyze the logs itself, which is faster and more context-aware.

209
MCQhard

An agent built with the Claude Agent SDK runs a long research task and repeatedly hits the model's context window limit. The developer wants the agent to keep working without losing critical earlier findings. Which approach best addresses this?

A.Enable automatic context compaction so older turns are summarized into a condensed form while key findings are retained in the running context.
B.Raise the temperature so the model generates shorter responses and consumes fewer tokens per turn.
C.Restart the agent from scratch each time the limit is reached, relying on the model to re-derive earlier findings from the task prompt.
D.Increase the model's max_tokens parameter so each response can be longer and the context window is used more efficiently.
AnswerA

Context compaction summarizes or truncates older conversation content when the window fills, preserving essential information while freeing space for new turns. This lets a long-running agent continue past the raw token ceiling without discarding the findings it still needs, which is exactly the failure mode described in the scenario.

Why this answer

Long-running agents need a strategy for the finite context window. Compaction condenses older turns into summaries or retained highlights, freeing tokens while keeping the findings the agent depends on. Sampling parameters and response-length caps do not reclaim history, and restarting loses state, so compaction is the mechanism that lets the run continue coherently.

Exam trap

The trap here is confusing response-length controls like max_tokens or temperature with the separate problem of total accumulated conversation size.

210
Multi-Selectmedium

Which THREE strategies are effective for reducing hallucinations when Claude is asked to answer questions based on a large provided context?

Select 3 answers
A.Instructing the model to say 'I don't know' if the answer is not in the text.
B.Setting the temperature to 0.7 to allow for more creative synthesis of the facts.
C.Asking the model to provide direct quotes from the text to support its answer.
D.Using a two-step process: first extract relevant snippets, then answer the question.
E.Increasing the frequency_penalty to prevent the model from repeating context words.
AnswersA, C, D

Giving Claude an 'out' is one of the most effective ways to prevent it from making up information. By explicitly permitting the model to admit a lack of information, you reduce the pressure for it to be helpful at the expense of being truthful, which is a common cause of hallucinations.

Why this answer

Reducing hallucinations requires a combination of structural constraints and behavioral guidance. By allowing the model to express uncertainty and encouraging it to ground its answers in direct quotes, developers create a verification loop. These techniques ensure the model prioritizes accuracy over the helpfulness of providing an answer even when the information is missing.

Exam trap

Candidates frequently rely on vague phrases like 'be accurate,' ignoring the necessity of explicit verification loops like extraction steps and 'I don't know' permissions to curb hallucinations.

211
MCQmedium

A developer is streaming a response from the Anthropic Messages API using server-sent events and wants to assemble the final text on the client. They observe that each event delivers a small fragment of the answer. Which event type carries the incremental text delta that must be concatenated to reconstruct the full completion?

A.content_block_stop
B.message_start
C.content_block_delta
D.ping
AnswerC

With streaming enabled, Claude emits content_block_delta events whose delta object holds the incremental text (for example a text_delta with a partial string). Concatenating these fragments in arrival order reproduces the complete message, and the final message_delta and message_stop events signal the end of the stream.

Why this answer

Streaming responses arrive as a sequence of typed server-sent events, and the actual generated text is delivered piecewise inside content_block_delta events. A client must concatenate those deltas in order and treat message_start, content_block_stop, message_delta, and message_stop as structural markers rather than content.

Exam trap

The trap here is assuming the first event of a stream (message_start) or the block boundary events already contain the answer text, when only the delta events carry incremental output.

212
MCQhard

A large-scale deployment of Claude is causing intermittent spikes in latency. Which security-related monitoring practice helps differentiate between a DDoS attack and legitimate heavy usage?

A.Check the total cost of API usage in the Anthropic dashboard.
B.Analyze the request frequency and prompt diversity from specific origin IPs.
C.Limit all users to one request per minute.
D.Disable all external access to the API immediately.
AnswerB

DDoS attacks typically exhibit high-frequency requests from limited sources or bots, often with low diversity in prompts. Legitimate usage is generally more varied. By analyzing the diversity of queries and the origin of traffic, you can differentiate between normal growth and a targeted attempt to exhaust resources.

Why this answer

Differentiating between DDoS attacks and high legitimate load requires deep visibility into API request patterns. By tracking the distribution of source IPs, request frequency per token, and the complexity of prompts, security teams can identify anomalous patterns that signify an attack. This is crucial for maintaining availability without blocking legitimate users, ensuring that security measures are proportionate to the threat, rather than causing self-inflicted denial of service.

Exam trap

Candidates often suggest simple rate limiting, which fails to distinguish between a heavy legitimate user and a malicious actor, potentially blocking valid high-value business traffic.

213
MCQmedium

When using the Messages API, a developer wants to ensure that Claude does not use any tools and only provides a standard text response, even if tools are defined in the request. Which configuration should they use?

A.Omit the 'tools' parameter from the request entirely.
B.Set the 'tool_choice' parameter to {"type": "none"}.
C.Set the 'tool_choice' parameter to {"type": "auto"}.
D.Include a system prompt that says 'Do not use any tools in this conversation'.
AnswerB

Setting tool_choice to 'none' explicitly instructs Claude to ignore the tools provided in the 'tools' array. This ensures the model only generates a standard text response, which is essential when you want to use the same code base for both tool-enabled and text-only interactions with the model.

Why this answer

The tool_choice parameter provides granular control over when and how Claude uses the tools provided in the tools array. Setting it to 'none' is the definitive way to disable tool use for a specific request. This is useful for debugging or for scenarios where you want to provide tool definitions but decide dynamically to ignore them.

Exam trap

Test-takers often try to prevent tool usage by simply omitting the tools array entirely, forgetting that explicit control parameters are needed when tools are defined.

214
MCQeasy

Which object in the Claude API response body provides the exact count of tokens consumed by the prompt and the generated completion for billing and usage monitoring?

A.metadata
B.usage
C.token_details
D.consumption
AnswerB

The usage object is a top-level field in the JSON response that specifically contains input_tokens and output_tokens. This provides the definitive count of how many tokens were processed in the request and how many were generated in the response, which is the direct basis for Anthropic's usage-based pricing model.

Why this answer

Tracking token usage is essential for managing API costs and understanding the scale of data being processed. The Messages API returns a usage object that provides transparent metrics for every request. Developers use this data to implement internal billing, monitor for spikes in usage, and optimize their prompts to fit within budget and rate limit constraints.

Exam trap

Candidates often look for token counts in the top-level response object or metadata headers, missing the specific 'usage' object nested within the API response body that contains the actual metrics.

215
MCQmedium

You are building a summarization feature that processes 20,000 support tickets nightly. Each ticket is under 4,000 tokens and the output summary is around 300 tokens. The job must complete within a 6-hour window and you want to minimize cost. Which Claude model should you choose?

A.Claude 3 Sonnet
B.Claude 3 Opus
C.Claude 3.5 Sonnet
D.Claude 3 Haiku
AnswerD

Claude 3 Haiku is the fastest and most cost-effective model in the Claude 3 family. For summarizing short tickets under 4,000 tokens with a modest 300-token output, Haiku delivers adequate quality at the lowest price per token, making it ideal for high-volume, budget-sensitive batch jobs that must finish within a 6-hour window.

Why this answer

For high-volume, short-input summarization where cost is the primary constraint, Claude 3 Haiku provides the best balance of speed, quality, and price. Its lower per-token cost directly reduces the total spend for 20,000 requests, and its speed helps meet the 6-hour window. More expensive models like Sonnet or Opus are overkill for this task.

Exam trap

The trap here is assuming that a more capable model is always better, when the scenario explicitly prioritizes cost minimization for a simple task.

216
MCQmedium

An MCP Server is connected to a Host using the Stdio transport. If the server process crashes, how does the MCP architecture generally handle the reconnection logic?

A.The MCP Client sends a 'heartbeat' signal to automatically restart the server.
B.The MCP Server uses a sidecar process to monitor its own health and restart.
C.The Host application detects the closed stream and must manage the restart.
D.The Model will identify the crash and suggest a fix to the developer.
AnswerC

In a Stdio transport configuration, the Host is the parent process that spawned the Server. When the Server crashes, the Host's read/write pipes are broken. It is the Host's responsibility to handle this exception, inform the user, and decide whether to attempt to re-initialize the server process.

Why this answer

The Stdio transport relies on standard input and output streams for communication between the host and the server. Because this is a process-level connection, a crash usually results in the termination of the stream. The responsibility for detecting this failure and initiating a recovery or restart lies with the MCP Host application.

Exam trap

Candidates often assume the MCP server automatically restarts itself or that the protocol layer handles crash recovery independently of the host application.

217
MCQmedium

Which method is best for improving Claude's accuracy in a complex multi-step reasoning task?

A.Asking the model to provide only the final answer to save tokens.
B.Providing the model with a massive list of facts without any reasoning steps.
C.Instructing the model to 'think through this step-by-step' before providing the final answer.
D.Using a very high temperature to ensure the model finds a unique reasoning path.
AnswerC

This classic instruction triggers chain-of-thought reasoning. It prompts the model to break down complex problems into manageable logical steps. This drastically improves performance on tasks involving math, logical deduction, and multi-stage analysis, as it forces the model to maintain logical coherence throughout the entire reasoning sequence.

Why this answer

Chain-of-thought (CoT) prompting is the industry standard for improving accuracy in reasoning-heavy tasks. By forcing the model to articulate its logic step-by-step before arriving at a final answer, you minimize common reasoning errors. This process allows the model to 'show its work,' which helps in verifying its conclusions and significantly reduces the probability of reaching an incorrect final result through faulty assumptions.

Exam trap

Candidates frequently try to prompt for the final answer directly, failing to realize that complex reasoning tasks require the model to externalize its thought process to minimize logical errors.

218
MCQmedium

When evaluating LLM performance, why is it critical to use a 'hold-out' test set of prompts that the model was not trained on?

A.To increase the token limit for the evaluation process.
B.To prevent overfitting the prompt engineering to a specific set of inputs.
C.To ensure the model receives different system instructions every time.
D.To reduce the cost of API calls during the testing phase.
AnswerB

Overfitting in prompt engineering occurs when a prompt is tuned specifically to excel on a narrow set of inputs but fails on others. A hold-out set acts as a 'blind' test, confirming that the prompt structure is universally effective and not overly optimized for a specific set of examples.

Why this answer

Using a hold-out test set is essential for measuring generalization. If you evaluate a model using the same prompts used during development, you are testing for 'memorization' rather than true reasoning ability. A hold-out set ensures that the prompt engineering patterns you've developed are robust enough to work on unseen, novel inputs, providing a realistic prediction of performance in the production environment.

Exam trap

Candidates mistakenly believe that testing on the same set used for prompt development proves the prompt is 'ready,' ignoring that they have only tested for memorization rather than actual generalization.

219
MCQhard

When using tool use with streaming enabled in the Anthropic API, how does the client receive the tool arguments?

A.As a single, complete JSON object in the first stream event.
B.Through incremental 'input_json' deltas within content block events.
C.Via a separate WebSocket channel dedicated to tool execution.
D.As a Base64 encoded string at the end of the message stream.
AnswerB

This is the correct behavior for streaming. The client must listen for 'content_block_delta' events where the delta type is 'input_json_delta'. Each delta contains a piece of the JSON string. The client concatenates these strings until the block is finished, at which point the full JSON can be safely parsed.

Why this answer

Streaming tool use involves receiving the tool call in parts. The client receives a 'content_block_start' for the tool_use, followed by multiple 'content_block_delta' events containing fragments of the JSON in the 'input_json' field. The client must accumulate these fragments and parse the final JSON string only after the 'content_block_stop' event.

Exam trap

Candidates often expect the full tool arguments to appear in a single event, failing to account for the streaming nature of the API where JSON fragments arrive incrementally via delta events.

220
MCQhard

A developer is using Claude Code to refactor a function across multiple files. After a few edits, Claude Code proposes a change that would alter the behavior of a critical utility function used elsewhere. The developer wants to ensure the change is safe before applying it. What is the most appropriate action?

A.Use the `/undo` command to revert all previous edits and start the refactor from scratch.
B.Accept the change and run the full test suite afterward to catch any regressions.
C.Reject the change and ask Claude Code to propose an alternative that preserves the existing behavior.
D.Manually edit the utility function to match the proposed change and then continue.
AnswerC

Rejecting the change and requesting an alternative that maintains the original behavior ensures that the critical utility function remains stable. This approach leverages Claude Code's ability to iterate on suggestions while keeping the developer in control. It is the safest way to avoid unintended side effects in other parts of the codebase that depend on that function.

Why this answer

Rejecting the proposed change and asking for an alternative that preserves existing behavior is the safest approach when a critical utility function is involved. It keeps the developer in control and avoids unintended side effects. Other options either apply the risky change, discard too much work, or bypass the review process entirely.

Exam trap

The trap here is thinking that running tests after accepting a risky change is sufficient, when the real risk is that tests may not cover all usages of a critical function.

221
MCQmedium

When using Claude to process multiple independent documents in a single prompt, what is the best way to ensure Claude can refer to each document accurately in its response?

A.Separate the documents with simple newline characters and horizontal rules.
B.Provide each document in a separate user message within a single API call.
C.Wrap each document in tags like <document id="1"> and <document id="2">.
D.Bold the title of each document at the start of its respective section.
AnswerC

Using unique IDs within XML tags provides a clear, machine-readable structure that Claude excels at navigating. This allows you to instruct the model to 'Refer to document 1' or 'Compare document 1 and 2', and the model will have a clear anchor for those references, improving citation accuracy.

Why this answer

Giving each document a unique identifier within XML tags is the most effective way to help Claude disambiguate between multiple sources. By using attributes or specific tag names like <document id="1">, you provide a clear reference system that the model can use in its reasoning and output, preventing it from mixing up details between different texts.

Exam trap

Candidates often provide multiple documents without identifiers, forcing the model to infer context, which leads to frequent cross-contamination of information and incorrect attribution of facts.

222
MCQeasy

When constructing a request for the Messages API, where should instructions that guide Claude's personality, tone, and global constraints be placed for optimal performance and architectural clarity?

A.As the first message in the messages array with the role set to 'user'.
B.In the system top-level parameter outside of the messages array.
C.As a hidden field within each individual content block of the user role.
D.Appended to the end of the last message with the role set to 'assistant'.
AnswerB

The system parameter is specifically designed to hold high-level instructions that define the model's behavior and constraints. Using this top-level field ensures that Claude treats the content as a foundational framework for the entire interaction, which improves adherence to complex rules and maintains a consistent persona throughout the dialogue.

Why this answer

The Messages API distinguishes between the conversational flow and the operational constraints of the model. Placing instructions in the system parameter allows the model to separate the 'how' of the response from the 'what' of the user query. This structural separation is a core mechanic for building robust, steerable applications that behave consistently across different sessions.

Exam trap

Candidates frequently embed persona and tone guidelines directly into the user message or conversation history instead of using the designated structural parameter.

223
MCQhard

A team wants Claude Code to run a custom script that inspects the repository for license headers before any commit. The script must run automatically each time Claude Code is about to create a commit, and the team wants the same hook to apply to every developer without manual setup. Which mechanism should they use?

A.Ask each developer to add the script to their personal shell profile as an alias.
B.Configure a cron job on each developer machine to run the script every five minutes.
C.Add a note in CLAUDE.md instructing Claude Code to run the script before committing.
D.Define a hook in the project's Claude Code settings that triggers on the commit event and runs the script.
AnswerD

Claude Code supports hooks configured in project settings that fire on specific lifecycle events such as pre-commit, letting the team run the license-header script automatically. Because the configuration lives in the repository, every developer inherits the same behavior without individual setup, which satisfies both the automatic execution and shared-application requirements.

Why this answer

Hooks defined in project-level Claude Code settings run deterministically on lifecycle events such as pre-commit and are shared through the repository, so every developer gets identical behavior automatically. Advisory notes in CLAUDE.md and per-machine aliases or cron jobs are either non-enforcing or require individual setup, and none reliably fires at the exact moment Claude Code is about to commit.

Exam trap

The trap here is relying on CLAUDE.md instructions or shell aliases for enforcement, when only a configured hook guarantees the script runs on the commit event for every developer.

224
MCQmedium

In the context of MCP, what is the primary purpose of the 'List Resources' capability of a server?

A.To provide a directory of available functions the model can execute.
B.To allow the client to discover data sources that can be read by the model.
C.To enumerate the hardware specifications of the host machine.
D.To list the active connections currently handled by the MCP Client.
AnswerB

Resources represent the 'data' layer of MCP. Listing resources allows an AI assistant to see what information is available in its environment. For example, a server could list all the CSV files in a folder as resources, which the model can then selectively read to answer user questions.

Why this answer

Resources in MCP are a way for servers to expose data that is not necessarily an executable tool. The 'List Resources' capability allows the client to see what data sources are available—such as log files, database tables, or documentation—so the model can then 'read' those resources to gain context for its tasks.

Exam trap

Candidates often confuse resources with executable tools, incorrectly thinking that listing resources allows Claude to run functions rather than read data.

225
Multi-Selecthard

A developer is hardening a remote MCP server that exposes a 'run_query' tool against a production analytics database. The server will be reachable over HTTP by multiple Claude clients. Which TWO practices are MOST important to prevent the tool from being abused? (Choose two.)

Select 2 answers
A.Embed the production database connection string in the tool description so Claude can decide when the query is safe to run.
B.Require authentication and authorization on the MCP server so only approved clients can invoke 'run_query', and scope the database credentials to read-only.
C.Validate and parameterize all SQL inside the tool handler instead of concatenating model-supplied strings into the query.
D.Increase the 'max_tokens' value on every Claude request so the model has enough room to reason about query safety before calling the tool.
E.Return the full database error message, including schema names and stack traces, to Claude so it can self-correct failed queries.
AnswersB, C

A remote HTTP MCP server is an API surface; without authentication, anyone who can reach it can call the tool. Pairing client authentication with least-privilege database credentials ensures that even a compromised client cannot drop tables or read unrelated schemas. This combination addresses both who may call the tool and what the tool can do once called, which is essential for a production analytics endpoint exposed to multiple clients.

Why this answer

Hardening a remote MCP tool requires controls at the trust boundary and at the data layer. Authenticating clients and scoping database credentials to read-only limits who can invoke the tool and what it can do. Parameterizing and validating SQL ensures model-supplied arguments cannot alter query structure.

Together these reduce the impact of prompt injection, hallucinated arguments, and unauthorized access to a production analytics database.

Exam trap

The trap here is treating model reasoning or token budget as a security control, when MCP tool safety must be enforced by authentication and input handling on the server.

Page 2

Page 3 of 4

Page 4

All pages