Courseiva

Claude Certified Associate (CCAO-F) — Questions 226–259

259 questions total · 4pages · All types, answers revealed

Page 3

Page 4 of 4

226
MCQmedium

Which of the following best describes the 'Shared Responsibility Model' for safety when using Anthropic's API?

A.Anthropic is 100% responsible for any output the model generates in any context.
B.The user is 100% responsible for all model behavior, and Anthropic provides no safety features.
C.Anthropic secures the model and base safety, while the customer secures their application and use case.
D.Both parties are responsible for sharing the financial costs of any legal disputes.
AnswerC

This correctly defines the shared responsibility. Anthropic ensures the foundation model is aligned with safety principles, while the customer is responsible for the 'last mile' of safety—including how prompts are handled, how sensitive data is managed, and ensuring the application complies with relevant industry regulations and laws.

Why this answer

The Shared Responsibility Model clarifies that both Anthropic and the customer have roles in ensuring safety. Anthropic provides a secure and aligned model, while the customer is responsible for how they implement and use that model in their specific application. This partnership is essential for the safe and ethical deployment of AI across different industries.

Exam trap

Candidates often assume Anthropic is responsible for all safety aspects, failing to realize that the customer is responsible for how they configure and deploy the model in their specific application.

227
MCQhard

A developer is building a Claude-powered chatbot for a mental health support app. During testing, a user prompt expresses suicidal ideation. The developer wants to ensure the chatbot responds safely and appropriately. Which of the following is the best course of action according to Anthropic's safety guidelines?

A.Have the chatbot express empathy, encourage the user to seek professional help, and provide crisis resources without attempting to diagnose or treat.
B.Program the chatbot to offer a diagnosis of depression and suggest coping strategies.
C.Configure the chatbot to immediately provide a list of emergency hotlines and end the conversation.
D.Instruct the chatbot to change the subject to a positive topic to de-escalate the situation.
AnswerA

This approach aligns with Anthropic's safety guidelines for sensitive topics: show empathy, avoid giving medical advice, and direct users to professional resources. It maintains a supportive tone while recognizing the chatbot's limitations. Providing crisis resources is critical, and encouraging professional help ensures the user gets appropriate care.

Why this answer

The best action is for the chatbot to express empathy, encourage professional help, and provide crisis resources without diagnosing or treating. This aligns with Anthropic's guidelines for handling sensitive mental health topics, ensuring the user feels heard while directing them to appropriate care. It balances safety with helpfulness and avoids overstepping the chatbot's role.

Exam trap

The trap here is thinking that immediately providing hotlines and ending the conversation or changing the subject is safe, when empathetic engagement and professional referral are required.

228
Multi-Selecthard

A developer wants to utilize 'Tool Use' (function calling) with Claude. Which TWO steps are necessary to properly define a tool in the API request?

Select 2 answers
A.Provide a 'description' field explaining what the tool does.
B.Include the tool's source code in the 'implementation' field.
C.Define an 'input_schema' using JSON Schema format.
D.Set the 'role' of the message to 'function'.
E.Upload a CSV file containing example tool outputs.
AnswersA, C

The description is vital because it is the primary way Claude understands the tool's purpose. Without a clear description, the model may not know when it is appropriate to call the tool or what real-world action the tool represents, leading to poor tool selection and incorrect workflow execution.

Why this answer

To use tools, developers must provide a 'tools' array in the request. Each tool needs a 'name', a 'description' (which helps Claude understand when to use it), and an 'input_schema'. The 'input_schema' must be a valid JSON Schema object, defining the parameters the tool expects.

This structured approach allows Claude to generate valid arguments that the developer's code can then execute.

Exam trap

Many candidates forget to include an explicit description for each tool, wrongly assuming that the input schema alone is enough for Claude to know when to call it.

229
MCQeasy

When designing a prompt for Claude, which technique is most effective at reducing the risk of 'hallucination' or factually incorrect information?

A.Setting the temperature parameter to 0.
B.Instructing the model to admit when it lacks information.
C.Using the most expensive model variant.
D.Increasing the max_tokens limit.
AnswerB

Instructing the model to refrain from guessing and to indicate when information is missing is a highly effective grounding strategy. This creates a safety boundary, preventing the model from inventing facts when the source material is inadequate, which is essential for maintaining trust and accuracy in enterprise applications.

Why this answer

Providing clear, constrained instructions and grounding the model in provided context is the most effective way to minimize hallucinations. By explicitly instructing the model to reply 'I don't know' if the information is not present in the provided source text, developers can significantly improve the factual reliability of the system. This practice forces the model to prioritize provided data over its internal training parameters, improving accuracy in domain-specific tasks.

Exam trap

Candidates often select 'giving the model more creative freedom' or 'increasing the temperature', which directly contradicts the goal of factual accuracy and hallucination reduction.

230
MCQeasy

A user wants Claude to provide a response in Markdown format with specific headers. Which approach is best for achieving this consistently?

A.Request the format in the system prompt and provide a Markdown template.
B.Only mention 'Markdown' once at the beginning of the user prompt.
C.Use the word 'IMPORTANT' in all caps before the Markdown request.
D.Hope that the model defaults to Markdown since it is common.
AnswerA

Combining a high-level instruction in the system prompt with a concrete template provides both the 'what' and the 'how'. This dual approach is highly effective for structural tasks, as it sets the expectation and then provides a visual reference for the model to emulate.

Why this answer

Clear formatting instructions combined with structural examples are the most effective way to guide Claude's output format. By explicitly naming the desired format (Markdown) and providing a template or list of required headers, the developer gives the model a clear blueprint to follow during the generation process.

Exam trap

Candidates often try to instruct the model to use Markdown in the user message. This is less effective than defining the format in the system prompt for consistent application.

231
Multi-Selectmedium

An enterprise is scaling their Claude integration and needs to manage their rate limits effectively. Which TWO strategies are recommended by Anthropic to handle rate limiting gracefully?

Select 2 answers
A.Implement exponential backoff for retries
B.Request a limit increase immediately upon the first 429 error
C.Monitor 'anthropic-ratelimit' headers in responses
D.Use multiple API keys to multiply the available rate limits
E.Switch to a smaller model only when a rate limit is hit
AnswersA, C

Exponential backoff is a standard error-handling technique where the client waits longer between each successive retry of a failed request. This prevents 'thundering herd' problems where many clients overwhelm the server at once. It is the most effective way to recover from 429 errors while respecting the server's capacity limits.

Why this answer

Rate limits are a reality of high-scale API usage. To handle them, developers should implement exponential backoff, which increases the wait time between retries after each failure. Additionally, they should monitor the rate limit headers in the API response to adjust their request frequency dynamically.

These practices prevent the application from being blocked and ensure smoother overall performance.

Exam trap

Candidates often rely purely on static hardcoded sleep timers for rate limiting, ignoring response headers and exponential backoff best practices.

232
MCQmedium

Refer to the exhibit. Why might the model struggle to answer the final question if you are not using a stateful chat implementation?

A.The model's internal memory resets every time you send a new API request.
B.The model has a maximum token limit that prevents it from remembering history.
C.The system prompt is preventing the model from accessing past information for safety.
D.The model is not optimized for question-answering tasks and lacks long-term recall.
AnswerA

The Anthropic API is stateless. Each request is a standalone interaction. To maintain the appearance of a conversation, you must send the entire message history (user and assistant turns) in the 'messages' parameter of the API request so the model can see the full thread and provide consistent answers.

Why this answer

In a stateless API architecture, the model does not automatically remember previous turns. To ensure the model has access to the full conversation history, you must include the full chat thread in every API request. If you only send the most recent user message, the model loses the context of previous messages, making it impossible to answer questions that depend on historical interactions.

Exam trap

Candidates incorrectly believe the model has a persistent 'session' memory. They fail to realize that each API call is isolated and requires the full history to maintain context.

233
MCQmedium

A developer is writing an assistant that must always answer in strict JSON with a fixed set of keys. They want to guarantee the model's output is parseable without writing custom repair logic. Which approach best fits the Messages API?

A.Set 'temperature' to 0 and rely on the model to output valid JSON.
B.Append the word 'JSON' to the system prompt and trust the model's compliance.
C.Define a tool with an input schema and require the model to call it.
D.Use the 'stop_sequences' parameter to halt generation at the closing brace.
AnswerC

Tool definitions carry a JSON Schema for their input, and when the model is required to use a tool, its arguments arrive as structured data conforming to that schema. This gives a reliable, machine-parseable shape for the fixed-key JSON the assistant must return, removing the need for hand-written repair logic.

Why this answer

The Messages API supports tool use, where each tool declares an input_schema in JSON Schema. When the model is required to call that tool, the returned arguments conform to the declared schema, yielding reliably structured output. Sampling settings, stop sequences, and prompt wording all shape behavior probabilistically but cannot guarantee a parseable, fixed-key JSON object.

Exam trap

The trap here is believing that temperature 0 or a strong prompt instruction guarantees valid JSON, when only a schema-constrained mechanism does.

234
MCQeasy

A developer is writing code that calls the Anthropic Messages API and needs to authenticate each request. The developer has retrieved the API key from a secure secret manager at runtime. Where should the API key be placed in the HTTP request?

A.In the URL as a query parameter such as ?api_key=sk-ant-...
B.In the request body as a top-level 'api_key' field alongside 'model' and 'messages'.
C.In the 'Authorization' header using the scheme 'Basic' with the key base64-encoded.
D.In the 'x-api-key' HTTP header on each request.
AnswerD

The Anthropic Messages API authenticates requests using the x-api-key header, and the value is the API key retrieved from the secret manager. This keeps the credential out of the URL and the JSON body, reducing leakage risk. Sending it this way on every request is the documented and expected authentication method for direct HTTP integrations with the API.

Why this answer

Direct HTTP calls to the Anthropic Messages API authenticate by including the API key in the x-api-key request header. This keeps the credential out of URLs and payloads, which is important for security and log hygiene. Body fields, query parameters, and Basic authentication schemes are not recognized by the API for this purpose, so requests using them will fail authentication.

Exam trap

The trap here is assuming the API uses the generic Authorization header or a body field, when it specifically requires the x-api-key header.

235
Multi-Selecthard

A developer is designing a Claude-powered assistant that must maintain a coherent conversation across many turns while controlling cost and context limits. Which TWO practices are appropriate according to Claude model fundamentals? (Choose two.)

Select 2 answers
A.Summarize or truncate older turns to keep the context within limits while preserving recent, relevant exchange.
B.Increase max_tokens on every request to allow Claude to store more conversation history internally.
C.Place durable persona and formatting rules in the system prompt so they persist across turns without repetition.
D.Rely on Claude to remember previous turns automatically between separate API calls.
E.Send the full conversation history with each request so Claude has complete context.
AnswersA, C

Because the API is stateless and context is finite, condensing or dropping older turns controls token growth and keeps requests within the context window. Summarization preserves important facts while reducing size, and retaining recent turns maintains conversational continuity. This is a recommended way to manage long conversations without losing coherence or driving up cost unnecessarily.

Why this answer

Because the Messages API is stateless, long conversations require deliberate context management. Summarizing or truncating older turns keeps requests within the context window and controls cost while preserving recent continuity. Durable persona, tone, and formatting rules belong in the system prompt, where they persist across turns without repetition.

Relying on automatic memory or misusing max_tokens misunderstands how Claude processes each independent request.

Exam trap

The trap here is believing Claude remembers prior turns automatically, when each Messages API call is stateless and history must be explicitly managed or it will be lost.

236
MCQhard

Refer to the exhibit. An engineer wants to use Prompt Caching to optimize this request. What is the correct way to modify the request body to enable this?

A.Add a 'cache_control' field to the root level of the JSON body.
B.Nest a 'cache_control' object within the content block of the message.
C.Rename the 'system' field to 'system_cached'.
D.Enable 'caching=true' in the request headers.
AnswerB

Prompt caching is activated by adding a cache_control block to the content of a message. This instructs the Anthropic API to store the preceding tokens in the cache. This is the correct structural way to enable the caching feature as defined in the Anthropic API technical documentation.

Why this answer

Prompt caching requires the inclusion of a 'cache_control' block within the message content. This tells the API to store the processed state of that specific chunk. This is critical for high-latency or high-cost prompts because it allows for faster processing of subsequent requests that share the same cached prefix, significantly reducing latency and cost for repetitive tasks like data analysis.

Exam trap

Candidates often try to enable caching by adding a top-level parameter or a system prompt flag, forgetting that 'cache_control' must be placed inside the specific content block.

237
MCQhard

Refer to the exhibit. When using 'tool_choice': {'type': 'auto'}, how does Claude determine whether to use the 'get_stock_price' tool?

A.It executes the tool automatically if the word 'stock' appears in the prompt.
B.It only uses the tool if the user explicitly mentions the tool name.
C.The model decides based on whether the tool's description matches the user's request intent.
D.It randomly selects a tool whenever the user asks a question about finance.
AnswerC

When tool_choice is set to 'auto', Claude uses the description and schema provided for each tool to judge if calling that tool would help it respond to the user. This involves a high-level reasoning step where the model balances its internal knowledge with the available external tools.

Why this answer

In 'auto' mode, Claude evaluates the user's intent against the provided tool descriptions and schemas. If the model determines that a tool is necessary to fulfill the request, it will generate a tool-use block. This process is driven by the model's internal reasoning and its understanding of the tool's utility in context.

Exam trap

Candidates often assume the model uses external search or database queries automatically. They fail to recognize that the model relies solely on the provided tool descriptions to determine relevance.

238
MCQmedium

A marketing team wants Claude to generate blog posts that strictly adhere to a specific brand voice. They find that the model occasionally deviates into a generic tone. Which strategy is most effective for ensuring consistent adherence to the brand voice?

A.Add 'Do not be generic' to the system prompt.
B.Provide five examples of existing high-quality blog posts within XML tags.
C.Set the temperature to 1.0 to encourage creative writing.
D.Use a single-sentence instruction at the very end of the prompt.
AnswerB

Providing concrete examples allows the model to map the desired tone, vocabulary, and structure directly from the context. This few-shot technique is a fundamental pillar of prompt engineering for Claude, as it provides a clear reference point that significantly outperforms zero-shot instructions when trying to capture complex styles.

Why this answer

Few-shot prompting provides the model with specific patterns and stylistic nuances that are difficult to describe through instructions alone. By including several high-quality examples, the developer leverages Claude’s ability to perform in-context learning. This approach stabilizes the output quality and ensures the brand voice is maintained throughout the generated content, reducing stylistic drift during long generations.

Exam trap

Candidates often rely solely on generic system instructions like 'be professional' instead of using concrete few-shot examples within XML tags to eliminate stylistic drift.

239
MCQmedium

A developer is building a sensitive financial advice application using Claude. To ensure compliance with safety standards and prevent the model from providing harmful advice, which architectural approach provides the most robust defense?

A.Relying exclusively on a robust system prompt to define all financial boundaries.
B.Using few-shot prompting to demonstrate correct financial advice patterns to Claude.
C.Integrating a secondary validation layer that inspects model output for prohibited content.
D.Setting the temperature parameter to zero to ensure deterministic, safe outputs.
AnswerC

External validation acts as a circuit breaker, inspecting outputs for harmful or non-compliant content before delivery. This method ensures that even if the LLM produces unexpected or prohibited advice, the system prevents it from reaching the end user, maintaining a critical layer of safety and regulatory compliance.

Why this answer

Effective AI safety in financial contexts requires a multi-layered approach. While system prompts set boundaries, they are vulnerable to jailbreaks. Implementing external guardrails, such as PII filtering and semantic checks, ensures that model outputs are validated against regulatory requirements before reaching the user.

This strategy is critical because it decouples model generation from output validation, creating a verifiable safety perimeter that remains intact even if the model's internal heuristics are bypassed or manipulated.

Exam trap

Candidates often believe that a robust system prompt is sufficient for safety, ignoring the reality that prompt injection or edge cases can bypass internal safety heuristics.

240
MCQmedium

A product manager is drafting a system prompt for a customer-facing Claude assistant. They want the assistant to maintain a professional tone, avoid discussing competitors, and always end responses with a link to the help center. They also want to allow users to override the tone for casual conversations. Which statement about system prompts best guides this design?

A.The system prompt should be left empty, and all rules including tone and competitor avoidance should be embedded in the first user message of every session.
B.The system prompt should instruct Claude to ignore all user attempts to change its behavior, ensuring the professional tone and competitor rules are never overridden.
C.The system prompt should set the global role and rules, and the assistant can be instructed to allow user requests to adjust tone within defined bounds.
D.The system prompt should contain only the tone and competitor rules, while the help center link instruction belongs in each user turn to ensure it is followed.
AnswerC

System prompts establish persistent behavior and constraints for the assistant, making them the right place for global rules like professionalism and competitor avoidance. They can also define how user instructions may modify tone within limits, giving the desired flexibility. This design keeps consistent guardrails while honoring the product manager's intent to allow casual tone when users request it.

Why this answer

System prompts are the appropriate place for persistent, application-controlled behavior such as role, tone, and global restrictions, and they can also specify how user requests may modify certain aspects within defined limits. This provides consistent guardrails while allowing the desired flexibility. Relying on user messages, leaving the system prompt empty, or forbidding all overrides either weakens the guardrails or removes the intended flexibility.

Exam trap

The trap here is treating the system prompt as all-or-nothing, when it can both enforce fixed rules and define bounded flexibility for user-driven adjustments.

241
MCQhard

An HR department is using Claude to screen resumes for a high-volume engineering role. They are concerned that the model might inadvertently favor candidates from certain universities over others, reflecting historical biases in the training data. Which concept in AI safety does this concern directly address?

A.Model Hallucination.
B.Prompt Injection.
C.Algorithmic Bias and Fairness.
D.Data Sovereignty.
AnswerC

This concern relates directly to algorithmic bias, where the model's outputs reflect and potentially amplify societal prejudices found in its training data. Ensuring fairness in high-stakes areas like employment is a key part of Anthropic's responsible use mission, requiring developers to audit their AI systems for discriminatory patterns.

Why this answer

Algorithmic bias and fairness are critical components of responsible AI use. In hiring, bias can manifest as a preference for certain demographics or backgrounds that were overrepresented in the training data. Addressing this involves careful testing, monitoring, and potentially using debiasing techniques to ensure the model makes equitable decisions based on merit rather than historical prejudice.

Exam trap

Candidates often incorrectly label this as 'Model Hallucination' or 'Prompt Injection,' failing to recognize that historical data bias is a distinct ethical concern related to fairness and representation.

242
MCQhard

Refer to the exhibit. What is the specific effect of the 'tool_choice' parameter as configured in this request?

A.It allows Claude to choose between using the tool or answering normally.
B.It forces Claude to use the 'get_weather' tool specifically.
C.It prevents Claude from using any tools during this request.
D.It enables Claude to use multiple tools in a single response.
AnswerB

By explicitly naming the tool in the tool_choice object, the developer overrides the model's default behavior. Claude will be compelled to generate a tool_use block for 'get_weather', ensuring that the application receives the structured data it needs to proceed with the weather-related logic, regardless of the prompt's phrasing.

Why this answer

The 'tool_choice' parameter allows developers to control Claude's tool-calling behavior. By setting it to type 'tool' and specifying a 'name', the developer is forcing Claude to use that specific tool, even if it thinks it could answer without it. This is useful for specialized workflows where a specific function must always be executed as the next step in the application logic.

Exam trap

Candidates often mistake 'tool_choice' for a general suggestion, not realizing that 'tool_choice: {type: "tool", name: "..."}' forces the model to use that specific tool.

243
MCQhard

When using Claude as a tool-calling engine, what is the best strategy to handle scenarios where the model needs to call multiple tools in sequence?

A.Force the model to output all tool calls in a single JSON array at the very beginning of the response.
B.Provide clear descriptions for each tool that explain when and why they should be used.
C.Implement a 'Human-in-the-loop' check for every single tool call to ensure safety.
D.Use a system prompt to tell the model to use all tools in a specific hard-coded order regardless of the input.
AnswerB

Clear descriptions are vital for the model to correctly identify which tool to use in which situation. If the descriptions are vague or redundant, the model will struggle to select the appropriate tool or sequence. Well-defined tool interfaces are the prerequisite for reliable automated tool-use workflows in production environments.

Why this answer

The model's ability to call tools is dependent on clear documentation and well-structured tool definitions. By providing concise descriptions and expected input schemas, the model can infer the sequence required to solve a problem. Ensuring that tools are independent where possible, or clearly linked in intent, allows the model to chain them effectively to achieve complex tasks without manual intervention from the user.

Exam trap

Candidates often over-engineer prompts by trying to manually dictate the exact tool sequence, rather than trusting the model's ability to infer the correct order from clear tool descriptions and intent.

244
MCQeasy

A product team is integrating Claude into a children's educational app. Before launch, they must ensure the model does not produce age-inappropriate content. Which action best aligns with Anthropic's safety guidance for this deployment?

A.Implement additional input and output filtering, plus human review, specifically tuned for the children's audience.
B.Rely on the model's default safety filters without additional testing, since Anthropic already ensures all outputs are child-safe.
C.Only test with adult users first and assume the same behavior will hold for children.
D.Disable all safety features to avoid false refusals, because educational content is inherently safe.
AnswerA

Deployers are expected to add context-appropriate safeguards beyond the model's baseline. For a children's app, that means extra filtering and human oversight to catch age-inappropriate content, reflecting Anthropic's guidance that safety is a shared responsibility between model provider and deployer.

Why this answer

Deploying Claude in a child-facing context requires going beyond baseline protections. The correct approach is to add audience-specific input/output filtering and human review, as Anthropic emphasizes that safety is a shared responsibility. Relying solely on defaults, disabling safeguards, or testing only with adults all fail to address the unique risks of a children's educational app.

Exam trap

The trap here is assuming that Anthropic's built-in safety features are sufficient for any audience, when deployers must add context-specific safeguards.

245
MCQmedium

You are building an application using Claude to extract structured data from messy user emails. The model occasionally ignores your schema constraints when the input text is ambiguous. Which strategy most effectively ensures strict adherence to the requested format?

A.Increase the temperature parameter to 1.0 to allow the model more creative flexibility in formatting.
B.Add a preamble asking the model to act as a helpful assistant that tries its best to output JSON.
C.Enclose the user input in <user_input> tags and include a few-shot example of the expected JSON structure within the system prompt.
D.Ask the model to output the result in a markdown code block without providing a schema definition.
AnswerC

XML tags provide clear delimiters that help Claude distinguish between instructions and data, reducing instruction leakage. Incorporating few-shot examples provides a concrete pattern for the model to follow, which is a highly effective way to enforce strict schema adherence, even when the input content is messy or ambiguous.

Why this answer

Using XML tags to delineate user input from instructions, coupled with a few-shot example of the desired JSON output, significantly improves schema adherence. This technique leverages Claude’s architectural strength in processing structured markup, reducing the likelihood of the model conflating input data with system instructions. Consistent formatting is vital in production pipelines where downstream systems expect rigid data schemas, as failure to comply can cause parsing errors and data loss.

Exam trap

Candidates frequently rely solely on system prompt instructions to enforce schema, failing to realize that few-shot examples and XML tagging are far more effective at anchoring the model.

246
MCQeasy

A support team wants Claude to classify incoming customer emails into one of five categories: Billing, Technical, Account, Feature Request, or Other. The team needs consistent, machine-readable output that a downstream script can parse reliably. Which prompt design best meets this requirement?

A.Ask Claude to choose the best category and, if multiple categories seem to apply, list them all separated by commas.
B.Ask Claude to explain its reasoning in a paragraph and end with the category name in bold.
C.Provide three examples of emails and their categories, then ask Claude to respond with the category and a confidence score between 0 and 100.
D.Instruct Claude to output only the category label, chosen from the predefined list, with no additional text.
AnswerD

Constraining the output to exactly one label from a predefined list produces consistent, easily parsed results. It eliminates ambiguity and reduces the chance of extraneous text breaking the downstream script. This directly satisfies the need for machine-readable output and consistent classification, making it the most reliable design for an automated pipeline.

Why this answer

When output feeds an automated script, the prompt should constrain the response to a single value from a predefined set and forbid extra text. This minimizes parsing complexity and ambiguity. Few-shot examples can improve consistency, but adding confidence scores or allowing multiple labels reintroduces variability.

Narrative explanations with formatting are human-friendly but not reliably machine-readable, so they fail the stated integration requirement.

Exam trap

The trap here is treating a human-readable answer with bold formatting or reasoning as machine-readable, when automation actually requires a strictly constrained, single-token-style label.

247
MCQmedium

A customer support team is building a Claude-powered assistant that must answer questions using a 300-page product manual. They want to avoid sending the entire manual with every request because of latency and cost. Which approach best leverages Claude's capabilities while keeping responses grounded in the manual?

A.Use retrieval-augmented generation: embed the manual, retrieve the most relevant passages, and include them in the prompt.
B.Lower the temperature to 0 so Claude reproduces the manual verbatim from memory.
C.Increase the max_tokens parameter so Claude can generate longer, more detailed answers directly from its training data.
D.Fine-tune Claude on the manual so the knowledge is embedded in the model weights.
AnswerA

RAG keeps the authoritative manual outside the model and injects only the top matching passages into the context window. This reduces token usage and latency, allows the manual to be updated without retraining, and gives Claude the exact source text to cite. It is the recommended pattern for grounding answers in a large, evolving document set such as a 300-page product manual.

Why this answer

Retrieval-augmented generation is the standard architecture for grounding Claude in a large, private, frequently updated document such as a product manual. Embedding the manual and retrieving only the most relevant passages keeps prompts small, reduces latency and cost, and gives Claude authoritative source text to answer from. Fine-tuning, max_tokens, and temperature do not supply the missing manual content and therefore cannot ensure grounded support answers.

Exam trap

The trap here is assuming that fine-tuning or a lower temperature can teach Claude a private 300-page manual, when grounding actually requires retrieving and injecting the relevant passages at request time.

248
Multi-Selectmedium

A developer is preparing to deploy Claude in an application that generates creative fiction. They want to ensure the application is used responsibly and complies with Anthropic's Usage Policies. Which TWO practices should they implement? (Choose two.)

Select 2 answers
A.Require users to sign a legal waiver acknowledging that AI-generated content may be offensive.
B.Include a clear disclosure to end users that content is AI-generated.
C.Fine-tune Claude to never generate content that could be considered controversial.
D.Add a mechanism for users to report harmful or policy-violating outputs.
E.Implement a content filter to block any violent or sexual themes in the generated fiction.
AnswersB, D

Disclosing that content is AI-generated promotes transparency and helps users understand the nature of the output. This aligns with responsible use principles and reduces the risk of deception. It is a straightforward practice that supports policy compliance without limiting the creative feature.

Why this answer

Disclosing AI-generated content and providing a reporting mechanism are two concrete practices that support transparency and accountability. They align with responsible use expectations for creative applications without unnecessarily restricting legitimate expression. The other options are either overly broad, ineffective, or not required by policy.

Exam trap

The trap here is assuming that responsible use requires blocking all sensitive themes, when in fact transparency and user reporting are the more appropriate controls for creative fiction.

249
Multi-Selectmedium

A developer is building a Claude-based assistant that must answer questions about a 90,000-token product manual. The manual is provided in the prompt on every request. The developer wants to improve answer accuracy and reduce the chance of the model overlooking relevant sections. Which TWO techniques are most appropriate? (Choose two.)

Select 2 answers
A.Ask Claude to summarize the entire manual first, then answer the question based only on that summary.
B.Instruct Claude to answer only from the provided manual and to say 'I don't know' if the answer is not present.
C.Place the long manual content before the user's question and the specific instructions, so the model reads the source material first.
D.Increase the temperature to 1.0 so Claude considers a wider range of interpretations from the manual.
E.Split the manual into random fragments and include only a few fragments per request to keep the prompt short.
AnswersB, C

Explicitly grounding the answer in the provided manual and allowing a 'I don't know' response reduces hallucination and encourages the model to rely on the source. This is especially important with long documents where the model might otherwise fill gaps with general knowledge. It also gives users a clear signal when the manual lacks the needed information.

Why this answer

Long-context accuracy improves when the source material is placed before the question and instructions, and when the model is explicitly told to answer only from that source with a not-found fallback. These two techniques together reduce overlooking and hallucination. Raising temperature, summarizing first, or randomly sampling fragments either adds noise, loses detail, or risks omitting the relevant content entirely.

Exam trap

The trap here is assuming that shortening a long prompt by random sampling or summarizing is always beneficial, when it can remove the exact evidence needed to answer correctly.

250
MCQeasy

What is the primary purpose of a 'System Prompt' in the Anthropic Claude API?

A.To provide a place for users to input their variable data.
B.To define the model's persona and overarching behavioral rules.
C.To reduce the latency of the API response for small tasks.
D.To encrypt sensitive information before it reaches the model.
AnswerB

System prompts are specifically designed to set the context for how the model should behave. This includes defining its role, such as a 'helpful coding assistant' or a 'formal legal researcher,' and setting global rules that the model must follow throughout the entire interaction.

Why this answer

The system prompt serves as a foundation for the model's behavior, establishing its persona, operational constraints, and high-level goals. By separating these instructions from the user's specific request, the system prompt provides a consistent framework that persists across a multi-turn conversation, guiding the model's overall 'personality' and boundaries.

Exam trap

Candidates often place role definitions and overarching guardrails inside the user message, rather than utilizing the dedicated system prompt parameter.

251
MCQhard

In the context of the Anthropic API, what is the primary benefit of using Streaming?

A.It reduces the total number of tokens consumed.
B.It drastically increases the accuracy of the response.
C.It improves the perceived response time for users.
D.It allows the model to handle more input tokens.
AnswerC

Streaming provides immediate feedback by showing text as it is generated. This dramatically improves the perceived latency, as users can start reading or interacting with the answer as it appears, rather than staring at a loading spinner for the entire duration of the model's generation process.

Why this answer

Streaming allows the model's output to be delivered to the client piece-by-piece as it is generated, rather than waiting for the entire response to be finished. This significantly reduces the 'Time to First Token' (TTFT) perceived by the user, making applications feel much more responsive. This is a critical technique for real-time chat interfaces where waiting for a multi-paragraph response to finish would lead to an unacceptable user experience.

Exam trap

Candidates often incorrectly state that streaming 'reduces total token cost' or 'increases overall model intelligence', confusing the user-facing latency benefit with backend operational efficiency.

252
MCQmedium

A developer is building a backend service that calls the Claude Messages API to summarize user-submitted articles. The service must enforce a hard limit: no summary should exceed 500 tokens. The developer sets max_tokens to 500. During testing, a response returns stop_reason: "max_tokens". What is the most accurate interpretation of this result?

A.The input article exceeded the context window, so the model could not process the entire document.
B.The model finished its summary naturally and the value indicates the summary is exactly 500 tokens long.
C.The response was truncated because it reached the max_tokens limit before the model naturally finished its output.
D.The model encountered an internal error and stopped generating tokens prematurely.
AnswerC

stop_reason "max_tokens" means the model hit the configured output token ceiling and the response was cut off, so the summary may be incomplete. The developer should treat this as a truncation signal, possibly increase max_tokens, or shorten the requested output, and then verify the content is complete before returning it to the user.

Why this answer

The stop_reason field reports why the model stopped generating. A value of "max_tokens" indicates the output was cut off because it reached the configured ceiling, meaning the content may be incomplete. Developers should treat this as a truncation signal and decide whether to raise max_tokens, shorten the prompt, or handle partial output gracefully.

Exam trap

The trap here is confusing an output-side truncation signal (stop_reason "max_tokens") with an input-side context window overflow error.

253
MCQeasy

Which capability is a primary benefit of using Claude 3.5 Sonnet compared to smaller, legacy models when processing complex, multi-step instructions?

A.It can be trained locally on customer-provided hardware.
B.It maintains higher instruction-following performance on complex, multi-step tasks.
C.It allows for unlimited parallel API requests without rate limiting.
D.It completely removes the need for systematic prompt engineering.
AnswerB

Claude 3.5 Sonnet is specifically optimized for advanced reasoning and instruction-following, allowing it to navigate complex, multi-step logic without losing track of constraints. This is a significant improvement over earlier models, making it ideal for tasks like code generation, complex document analysis, and multi-stage workflow execution in production environments.

Why this answer

Claude 3.5 Sonnet exhibits superior steerability and logical reasoning, allowing it to maintain consistency across long, multi-step prompts. This capability is vital for complex workflows where instructions are layered. Understanding model evolution helps developers choose the right tool for tasks that require high-level reasoning and instruction following, ensuring that complex business logic remains intact throughout the interaction, reducing the need for iterative prompting and debugging.

Exam trap

Candidates tend to think legacy models can handle layered logic just as well, underestimating how architectural evolution specifically improves complex multi-step reasoning.

254
MCQhard

An engineer is choosing between Claude 3.5 Sonnet and Claude 3 Opus for a pipeline that extracts structured fields from scanned invoices. The pipeline processes thousands of documents per hour and must balance accuracy against cost. Which statement best reflects the appropriate model-selection reasoning?

A.Claude 3 Opus should be used only for the first document and Sonnet for the rest, because Opus can teach Sonnet the schema at runtime.
B.Claude 3 Opus should always be chosen because it is the most capable model and therefore the most cost-effective at scale.
C.Model choice is irrelevant because all Claude 3 models have identical pricing and performance characteristics.
D.Claude 3.5 Sonnet is often the better balance because it provides strong accuracy on structured extraction at lower cost and higher throughput than Opus.
AnswerD

Claude 3.5 Sonnet delivers strong reasoning and extraction quality while costing less and offering better throughput than Opus. For high-volume invoice field extraction, that combination usually satisfies accuracy requirements without the premium price of Opus. This makes Sonnet the pragmatic default for a pipeline that must process thousands of documents per hour, with Opus reserved for the hardest edge cases.

Why this answer

Model selection should match capability to task and volume. Claude 3.5 Sonnet offers strong structured-extraction accuracy at a lower price and higher throughput than Claude 3 Opus, making it the sensible default for a high-volume invoice pipeline. Reserving Opus for the hardest cases balances quality and cost.

Treating all Claude 3 models as identical, or assuming runtime knowledge transfer between models, misstates how the API and model tiers actually work.

Exam trap

The trap here is defaulting to the most capable model for every task, when high-volume structured extraction usually favors a mid-tier model that balances accuracy with cost and throughput.

255
MCQmedium

A developer is integrating Claude into a customer-facing chatbot for a bank. The chatbot must not provide specific investment advice. During testing, a user asks, 'Should I buy Tesla stock now?' Claude responds with a detailed recommendation. Which action should the developer take to best align with responsible use?

A.Add a system prompt that instructs Claude to avoid giving financial advice and to respond with a disclaimer when asked for recommendations.
B.Allow Claude to give advice but add a footer to every response stating that the information is not financial advice.
C.Fine-tune Claude on a dataset of financial disclaimers so it learns to refuse all financial questions automatically.
D.Block all user messages containing the word 'stock' or 'invest' to prevent any financial discussion.
AnswerA

A system prompt is the primary way to steer Claude's behavior for a specific application. Instructing it to avoid financial advice and to provide a disclaimer directly addresses the requirement. While not foolproof, it significantly reduces the chance of inappropriate responses. This is a standard, responsible practice for domain-specific constraints, especially in regulated industries like banking.

Why this answer

A system prompt that instructs Claude to avoid financial advice and to disclaim when asked is the most direct and flexible way to enforce the bank's requirement. It guides the model's behavior at the point of generation. Disclaimers alone, fine-tuning, or keyword blocking are either insufficient or overly broad, and they do not reliably prevent specific investment recommendations.

Exam trap

The trap here is thinking that a disclaimer footer or keyword blocking is sufficient, when the more effective approach is to instruct Claude via system prompt to avoid giving financial advice altogether.

256
MCQhard

A developer wants Claude to call an internal inventory function when a user asks about stock levels, but the function must only be invoked when the model decides it is needed. The application will execute the function and return the result. Which sequence correctly implements this with the Messages API?

A.Send the tools definition; if the response contains a tool_use block, run the function and send a new request including a tool_result block referencing the tool_use id.
B.Pre-execute the inventory function on every request and inject its output into the system prompt.
C.Define the function in the system prompt as a description and parse the model's prose reply for a function name.
D.Send the tools definition and, upon receiving a tool_use block, immediately send a tool_result with an empty payload to acknowledge it.
AnswerA

This matches the tool use contract: tools are declared in the request, the model may return a tool_use content block with an id, and the application executes the function and returns a tool_result block that references that same id in a follow-up request, allowing the model to produce a final answer grounded in real data.

Why this answer

Tool use is a round trip: the application declares tools, the model optionally returns a tool_use block, the application executes the named function, and it sends back a tool_result block that references the tool_use id so the model can incorporate the real output. Pre-running functions, prose-based parsing, or empty acknowledgements all break this contract.

Exam trap

The trap here is believing the model itself executes the function, when in fact the application runs it and must return the result keyed to the tool_use id.

257
MCQmedium

A product team drafts a system prompt for a customer-facing assistant. The draft is 4,000 words and mixes persona description, tone guidance, formatting rules, escalation policy, and several anecdotes about past incidents. Reviewers find that Claude follows the formatting rules but frequently ignores the escalation policy. Which change best improves adherence to the escalation policy?

A.Shorten the system prompt by deleting the formatting rules so the escalation policy receives more attention.
B.Restructure the system prompt into clearly labeled sections and place the escalation policy in its own section with explicit trigger conditions.
C.Convert the entire system prompt into a numbered list of every rule in the order it was originally written.
D.Move the escalation policy to the last line of the system prompt without changing its wording.
AnswerB

Isolating the escalation policy in a labeled section with concrete triggers makes it a distinct, findable instruction rather than one clause buried among thousands of words of narrative. Structured system prompts reduce interference between unrelated instructions, so the policy is more likely to be applied when a triggering condition appears in the conversation.

Why this answer

The escalation policy is being lost among unrelated narrative content. Giving it a dedicated, labeled section with explicit trigger conditions turns it into an unambiguous, retrievable instruction. Structural organization reduces interference between competing directives, which is why the fix is about clarity and grouping rather than length alone or ordering alone.

Exam trap

The trap here is assuming that a long system prompt is the root cause and that shortening it will fix compliance, when the actual defect is that the escalation policy lacks structure and explicit trigger conditions.

258
Multi-Selectmedium

A developer is building a robust error-handling wrapper for the Messages API. Which THREE HTTP status codes should specifically trigger a retry logic with exponential backoff in a production environment?

Select 3 answers
A.400 - Bad Request
B.429 - Too Many Requests
C.500 - Internal Server Error
D.529 - Overloaded
E.401 - Unauthorized
AnswersB, C, D

This code is returned when the client has exceeded their rate limit. It is a classic 'transient' error that should be handled with exponential backoff. Retrying after a short delay allows the rate limit bucket to refill, ensuring the application can eventually complete its task without failing permanently for the user.

Why this answer

Handling API errors correctly is essential for application stability. 429 errors mean you are being rate-limited and should wait. 500 errors indicate a general server-side issue, while 529 errors mean the server is currently overloaded. All three represent temporary conditions where a retry might succeed. Conversely, errors like 400 or 401 represent client-side issues that retrying will not fix.

Exam trap

Candidates often include 400 (Bad Request) or 401 (Unauthorized) in their retry logic, not realizing that these client-side errors indicate invalid requests that will never succeed upon retry.

259
MCQhard

A researcher is using Claude to analyze sensitive personal data from a study. They want to ensure the data is handled responsibly and in line with Anthropic's guidelines. Which action is most appropriate before sending the data to Claude?

A.Encrypt the data and send it as a base64-encoded string in the prompt.
B.Anonymize or de-identify the personal data before including it in the prompt.
C.Obtain verbal consent from the study participants to use their data with an AI system.
D.Use a custom fine-tuned model that has been trained on similar sensitive data.
AnswerB

Anonymizing or de-identifying data reduces privacy risks and aligns with responsible data handling. It allows analysis to proceed while protecting individuals. This is a standard practice for sensitive data and directly addresses the concern about handling personal information in line with guidelines.

Why this answer

Anonymizing or de-identifying personal data before sending it to Claude is the most direct way to protect privacy and comply with responsible data handling guidelines. Encryption without decryption, consent alone, or fine-tuning on sensitive data do not adequately address the risk. De-identification allows analysis while minimizing exposure.

Exam trap

The trap here is thinking that encryption or consent is enough, when the core responsible-use step is to remove or mask personal identifiers before processing.

Page 3

Page 4 of 4

All pages