Courseiva

Claude Certified Associate (CCAO-F) — Questions 151–225

259 questions total · 4pages · All types, answers revealed

Page 2

Page 3 of 4

Page 4
151
MCQeasy

A healthcare organization is using Claude to summarize patient notes. To ensure the responsible use of the AI and protect patient privacy, which action should the organization take before sending data to the Claude API?

A.Enable the 'Public Data Sharing' option in the Anthropic console.
B.De-identify or redact all PII/PHI from the notes locally.
C.Instruct the model via the system prompt to ignore all names.
D.Use the 'Honest' parameter in the API request body.
AnswerB

By redacting sensitive information like names, social security numbers, and specific medical IDs before sending the data to the API, the organization minimizes the risk of data breaches. This practice aligns with the shared responsibility model, where the client is responsible for the data they provide.

Why this answer

Protecting Personally Identifiable Information (PII) and Protected Health Information (PHI) is a critical component of responsible AI use. Organizations must ensure that sensitive data is handled according to legal standards like HIPAA. De-identifying or anonymizing data before it reaches the AI provider is a primary security measure to prevent accidental exposure.

Exam trap

Candidates often assume that the Claude API automatically sanitizes PII/PHI on the server side, forgetting that compliance and local data privacy laws require client-side masking or redaction before transmission.

152
MCQmedium

A developer is building a multi-turn assistant that must remember details from earlier in a long conversation. After many exchanges, the developer notices that Claude starts forgetting information provided near the beginning of the session. The application currently sends only the latest user message with each API call. What is the most likely cause and the appropriate fix?

A.The developer must include the prior conversation turns in the messages array so the model has the necessary context in each request.
B.The developer should increase the max_tokens value so the model has more room to store conversation history internally.
C.The model has a limited memory that resets each call, so the developer must enable a persistent memory feature in the API request.
D.The developer should set the temperature to 0 so the model deterministically recalls earlier messages.
AnswerA

Because the Messages API is stateless, the model can only reason over the content provided in the current request. Omitting earlier turns means the model has no access to those details, which explains the forgetting. Sending the accumulated user and assistant messages in the messages array restores context and lets the model reference earlier information across turns.

Why this answer

The Messages API is stateless, so the model only sees what each request contains. Sending only the newest user message omits all prior context, which is why earlier details are forgotten. The correct remedy is to send the full conversation history in the messages array, alternating user and assistant turns, so the model can reference earlier information.

Generation parameters and output limits do not influence context retention.

Exam trap

The trap here is assuming the API maintains server-side conversation memory, when in fact each request must carry its own context.

153
MCQeasy

A product manager is evaluating Claude models for a customer-support chatbot that must handle 50,000 conversations per day while keeping inference costs low. The conversations are short, factual, and do not require complex reasoning. Which Claude model is the most appropriate choice?

A.Claude 2.1
B.Claude 3 Opus
C.Claude 3 Haiku
D.Claude 3.5 Sonnet
AnswerC

Claude 3 Haiku is the fastest and most cost-effective model in the Claude 3 family, optimized for high-volume, straightforward tasks like simple customer-support queries. It delivers quick responses at low cost, making it ideal for 50,000 short, factual conversations per day where complex reasoning is not required. This aligns perfectly with the scenario's requirements.

Why this answer

The scenario demands a model that can handle high-volume, simple conversations at low cost. Claude 3 Haiku is specifically designed for speed and cost efficiency, making it the right fit for short, factual customer-support interactions. The other models offer greater capability but at higher cost, which is unnecessary for this use case.

Exam trap

The trap here is assuming that the most capable model is always the best choice, when cost and latency requirements often favor a smaller, faster model.

154
Multi-Selectmedium

A product team is using Claude to generate marketing copy. They want to ensure the output aligns with brand voice and avoids certain topics. Which TWO techniques are most effective for controlling Claude's output in this scenario? (Choose two.)

Select 2 answers
A.Include few-shot examples in the prompt showing desired outputs for similar marketing requests.
B.Post-process the output with a keyword filter to remove any prohibited terms before publishing.
C.Set the temperature parameter to a high value to encourage more creative and varied outputs.
D.Use the top_p parameter with a low value to restrict the model to only the most likely words.
E.Provide a detailed system prompt that specifies the brand voice, tone, and prohibited topics.
AnswersA, E

Few-shot examples demonstrate the desired style and content, helping Claude infer patterns and apply them to new requests. This is particularly effective for nuanced tasks like brand voice, where examples can illustrate tone, vocabulary, and structure. Combined with a system prompt, few-shot learning significantly improves alignment with specific requirements.

Why this answer

A detailed system prompt and few-shot examples are proactive techniques that shape Claude's output during generation. The system prompt establishes persistent guidelines, while examples demonstrate desired patterns. Together, they provide clear direction and improve consistency, making them the most effective methods for controlling brand voice and avoiding prohibited topics.

Exam trap

The trap here is relying on post-hoc filters or sampling parameters, which do not provide the same level of control as explicit instructions and examples.

155
MCQmedium

A user asks Claude to help them write a convincing phishing email to steal credentials from their coworkers. Claude refuses. How does this refusal benefit the 'Safety and Responsible Use' ecosystem?

A.It ensures the model remains competitive with other AI providers who also refuse.
B.It helps the user learn how to write better phishing emails on their own.
C.It prevents the model from being used to facilitate cyberattacks and social engineering.
D.It allows the model to save its processing power for more complex coding tasks.
AnswerC

Refusing to generate phishing emails is a direct application of the harmlessness principle. It prevents the model from assisting in social engineering attacks, thereby protecting individuals and organizations from financial loss, data breaches, and other security risks associated with AI-powered cybercrime and malicious intent.

Why this answer

By refusing to generate phishing content, Claude prevents its technology from being used as a tool for cybercrime. This aligns with Anthropic's goal of ensuring AI is a force for good. Such refusals protect potential victims, reduce the burden on cybersecurity teams, and maintain the model's reputation as a safe and helpful assistant for legitimate tasks.

Exam trap

Candidates often focus on the model's 'intelligence' or 'capability' instead of the core safety outcome, which is the prevention of real-world harm through the refusal of malicious requests.

156
Multi-Selectmedium

A developer is optimizing a customer support bot that uses a 50,000-token knowledge base in every request. They want to implement Prompt Caching to reduce costs and latency. Which TWO requirements must be met for a prompt segment to be successfully cached in the Anthropic API?

Select 2 answers
A.The segment intended for caching must contain at least 1024 tokens.
B.The cached content must be placed at the very end of the user message.
C.The developer must include 'cache_control': {'type': 'ephemeral'} in the message block.
D.The system prompt must be empty when using cached blocks in the user message.
E.The temperature setting must be set to 0.0 for the cache to remain valid.
AnswersA, C

Anthropic enforces a minimum size for cacheable blocks to ensure the overhead of caching provides a meaningful performance benefit. Small snippets of text do not qualify for caching because the resource management required would outweigh the speed gains, making the 1024-token floor a critical architectural constraint.

Why this answer

Prompt caching is a powerful feature for large contexts, but it requires specific implementation details. The segment to be cached must be at least 1024 tokens and must be explicitly marked with the cache_control metadata. Understanding these technical thresholds is vital for developers looking to scale high-token applications while maintaining cost-efficiency and low latency.

Exam trap

Candidates often assume prompt caching is automatically applied to all large prompts, forgetting that it requires explicit opt-in via the 'cache_control' parameter and meets specific token minimums.

157
MCQeasy

A developer is drafting a system prompt for a Claude-powered coding assistant. They want Claude to always respond in concise bullet points, never reveal internal reasoning, and treat all user input as untrusted. Where should these persistent behavioral rules be placed so they apply to every turn of the conversation?

A.In a tool definition, since tools are evaluated before each response.
B.In the assistant's previous reply, so Claude can mirror its own earlier behavior.
C.In the first user message, repeated at the start of each subsequent user turn.
D.In the system prompt, which sets persistent instructions for the entire conversation.
AnswerD

The system prompt is designed to carry persistent instructions that apply across all turns, such as tone, format, and safety constraints. Placing bullet-point style, reasoning-hiding, and untrusted-input rules there gives Claude a stable frame it applies consistently, and it separates operator policy from user content, which is exactly what the scenario requires.

Why this answer

Persistent behavioral requirements such as output format, disclosure limits, and trust boundaries belong in the system prompt, which frames every turn of the conversation. Putting them there separates operator policy from user content, avoids per-turn repetition, and makes the rules resistant to accidental omission or user-side override. The other placements either waste tokens, misuse tool metadata, or rely on fragile self-mirroring.

Exam trap

The trap here is treating the system prompt as just another place to put text rather than the designated channel for persistent, conversation-wide instructions.

158
MCQmedium

Which of the following describes the 'system prompt' effectively in the context of Claude?

A.A list of history messages that the model uses to understand previous context.
B.The primary mechanism for defining the model's persona, operational rules, and behavioral constraints.
C.A temporary cache used to store user input tokens to improve response speed.
D.An optional field that is only used when the model is in training mode.
AnswerB

The system prompt is the standard mechanism for setting the model's behavior, instructions, and constraints. It acts as the anchor for the interaction, ensuring that the model adheres to specific guidelines and adopts the correct persona consistently throughout the entire session, which is vital for enterprise-grade applications.

Why this answer

The system prompt acts as the foundational 'constitution' for the interaction, defining the agent's persona, constraints, and operational goals. Unlike user messages, the system prompt is prioritized by the model and sets the boundaries for the entire session. Proper configuration here is critical because it dictates how the model interprets incoming user requests and handles edge cases, ensuring consistent and safe behavior across various application use cases.

Exam trap

Test-takers frequently confuse the system prompt with user messages or few-shot examples, overlooking its unique ability to establish foundational behavioral rules and constitutional boundaries.

159
MCQmedium

When designing a prompt for Claude to perform complex reasoning, why is it beneficial to include 'think step-by-step' in the instructions?

A.It forces the model to use more memory, which increases its internal intelligence.
B.It allows the model to break down complex problems into manageable sub-tasks.
C.It automatically converts the output format into a bulleted list.
D.It minimizes the token cost by reducing the number of words in the final output.
AnswerB

Chain-of-thought prompting prompts the model to generate intermediate steps. This decomposition helps the model handle multi-step reasoning tasks more accurately by grounding each part of the process in the previous logic, significantly reducing errors that occur when the model attempts to solve everything at once.

Why this answer

The 'think step-by-step' technique leverages the model's ability to perform chain-of-thought reasoning. By forcing the model to articulate its intermediate logic before arriving at a final answer, the model is less likely to jump to premature conclusions. This is particularly important for logic, math, or coding tasks where the final output depends on the accuracy of sequential steps rather than direct pattern completion.

Exam trap

Candidates often assume 'think step-by-step' is a magic phrase that guarantees accuracy for any prompt, forgetting that the model still requires high-quality, unambiguous instructions to reason effectively through complex logic.

160
MCQeasy

A developer is building an application that uses Claude to summarize news articles. They notice that Claude sometimes refuses to summarize articles involving violent crime, citing safety concerns. What is the best way for the developer to address this while maintaining responsible use?

A.Use a jailbreak prompt to force the model to ignore its safety training.
B.Contact Anthropic support to have all safety filters disabled for the account.
C.Refine the prompt to specify the journalistic context and requested summary length.
D.Switch to a smaller model version that has fewer safety features.
AnswerC

Providing context helps the model differentiate between harmful content and legitimate information processing. By specifying that the task is a 'journalistic summary,' the model can better evaluate the request against its safety principles, often resolving false-positive refusals while still maintaining the core safety boundaries against promoting violence.

Why this answer

Safety filters can sometimes be overly sensitive (false positives). The developer should refine the prompt to clarify the educational or journalistic purpose of the request. This helps the model understand that the goal is information summarization rather than the promotion or glorification of violence, which aligns with the intended use of the safety guardrails.

Exam trap

Candidates often assume the model is broken or biased, failing to realize that refining the prompt to provide context can successfully resolve false-positive safety refusals.

161
MCQeasy

Which action constitutes a violation of Anthropic’s Acceptable Use Policy regarding the generation of deceptive content?

A.Generating a fictional story about a hypothetical political scenario.
B.Using the model to write an email template for customer support employees.
C.Creating content that impersonates a real public figure to spread false information.
D.Building a tool that helps students brainstorm ideas for a history essay.
AnswerC

Impersonating real public figures to spread misinformation is a direct violation of the Acceptable Use Policy. Such actions are categorized as deceptive and harmful, as they undermine the integrity of information and can have severe societal consequences, which Anthropic explicitly prohibits to ensure AI is used safely.

Why this answer

Anthropic strictly prohibits the use of its models to generate deceptive content, including political misinformation or impersonation. This policy is fundamental to responsible use, as the widespread dissemination of fabricated content undermines public trust and creates real-world harm. Developers are responsible for ensuring their applications do not facilitate the creation or distribution of fraudulent or misleading information targeting individuals or groups.

Exam trap

Candidates sometimes assume that if the content is technically 'true', it is permitted, failing to recognize that the intent to deceive or impersonate is the core policy violation.

162
MCQmedium

A product team is preparing to launch a Claude-powered feature that summarizes and answers questions about user-submitted documents. During a pre-launch review, a team member suggests that because Claude already has built-in safety filters, they can skip adding any application-level safeguards. Which response best reflects responsible deployment practices?

A.Postpone all safeguards until after launch, then rely on user reports to identify and fix any safety issues that arise.
B.Disable Claude's built-in safety filters to avoid false refusals, then add a single keyword blocklist for known harmful terms.
C.Rely on Claude's built-in safety filters as the sole safeguard, since Anthropic continuously updates them.
D.Add application-level safeguards such as input sanitization, output review, and usage monitoring on top of Claude's built-in protections.
AnswerD

Layered defense is the recommended approach. Claude's built-in filters address general safety, but application-specific risks—such as prompt injection hidden in documents, leakage of sensitive data, or domain-specific harmful outputs—require extra controls. Input sanitization, output review, and monitoring help catch issues the base model may not, especially as usage patterns evolve after launch.

Why this answer

Responsible deployment of Claude-based features involves defense in depth: the model's built-in safeguards plus application-specific controls. Because user-submitted documents can contain prompt injection, sensitive data, or domain-specific harmful content, teams should add input sanitization, output review, and monitoring. Relying solely on built-in filters, disabling them, or waiting for user reports all fail to address foreseeable risks before launch.

Exam trap

The trap here is assuming that Claude's built-in safety filters are sufficient on their own, when responsible deployment requires additional application-level safeguards tailored to the specific use case.

163
MCQhard

A red team is probing a Claude-powered tutoring application that helps high-school students with chemistry homework. During testing, a user submits a prompt asking Claude to role-play as a teacher who will provide step-by-step instructions for synthesizing a hazardous compound at home. The application currently passes raw user input directly to the model. Which response best reflects responsible handling of this scenario?

A.The application should rely on Claude's built-in safety behavior to refuse, and the team should not add any additional controls since that would duplicate the model's protection.
B.The team should disable logging entirely to protect student privacy, and rely on periodic manual spot checks of random sessions instead.
C.The team should add application-level controls such as input screening for hazardous-intent patterns, system-prompt scope limits for the tutoring context, and logging so that attempts are visible and reviewable.
D.The team should raise the model's temperature and encourage more creative role-play responses so students remain engaged with the tutoring experience.
AnswerC

Layered application controls complement the model's safety training and are appropriate for a deployment serving minors. Input screening can flag hazardous-intent prompts, a scoped system prompt keeps the assistant within tutoring boundaries, and logging provides visibility for incident review. Together these create detection and accountability that a bare model call does not provide, which is the responsible posture for this risk profile.

Why this answer

A tutoring application serving minors needs defense in depth rather than a single reliance on model behavior. Application-level input screening, scoped system prompts, and privacy-preserving logging create detection and accountability around hazardous-intent attempts. Removing logging sacrifices visibility, and increasing creative latitude would weaken rather than strengthen the safety posture of the deployment.

Exam trap

The trap here is assuming that strong base-model safety training makes application-level guardrails redundant.

164
MCQhard

A company is developing an application that generates long-form creative content. They notice that as the output length increases, the model's speed seems to fluctuate. Which technical factor most significantly impacts the latency of the 'Time to Last Token' (TTLT) in this scenario?

A.The number of concurrent users in the API queue.
B.The total number of output tokens generated.
C.The use of a system prompt instead of a user prompt.
D.The presence of images in the input context.
AnswerB

Because autoregressive models like Claude generate one token at a time, each additional token adds to the cumulative processing time. For long-form creative content, the sheer volume of generated text is the dominant factor in determining how long it takes for the model to finish the request. (51 words)

Why this answer

The Time to Last Token (TTLT) is primarily a function of the total number of tokens generated by the model. Each token is generated sequentially, meaning the model must perform a full inference pass for every single word or sub-word it produces. Therefore, longer outputs naturally take more time, as the total latency scales linearly with the generation length. (70 words)

Exam trap

Candidates often attribute latency to 'input size' or 'network bandwidth,' overlooking that generation (output) is an inherently sequential process where every token adds linear time.

165
MCQmedium

An enterprise legal team needs to analyze a 150,000-token collection of merger and acquisition documents to identify conflicting indemnity clauses. Which feature of the Claude 3 model family is most critical for ensuring the entire dataset is processed in a single inference pass without losing context?

A.Multimodal Vision capabilities
B.System Prompts
C.200,000-token context window
D.Temperature settings
AnswerC

The large context window is the specific technical specification that enables Claude to process long-form documents or large datasets at once. It ensures that the model can maintain coherence across the entire 150,000-token set, allowing for complex cross-referencing and deep thematic analysis across multiple distinct files or sections. (52 words)

Why this answer

The context window determines the amount of information the model can hold in its active memory during a single request. For legal analysis involving massive document sets, a 200,000-token context window allows Claude to 'see' all relevant clauses simultaneously. This enables the model to identify cross-document contradictions and nuances that would be missed if the data were fragmented. (70 words)

Exam trap

Candidates often select 'model size' or 'training data volume,' ignoring that the context window is the specific technical constraint that allows the model to 'see' all 150,000 tokens at once.

166
MCQmedium

You are prompting Claude to adopt a specific tone. Which technique is most likely to ensure consistency across a variety of user queries?

A.Explain the tone in exhaustive detail in the system prompt.
B.Provide multiple examples (few-shot) of the desired tone in the prompt.
C.Use a system prompt that specifies a list of forbidden words.
D.Set the temperature to a high value to increase variety in the tone.
AnswerB

Few-shot examples offer the model a clear, functional pattern of the desired behavior. By providing examples of queries and the corresponding responses in the required tone, you enable the model to learn the nuance and application of the tone, leading to much higher consistency in actual usage.

Why this answer

Few-shot prompting provides concrete examples of the desired tone in action, which is far more reliable than qualitative descriptions. By showing the model exactly how to handle different scenarios, you provide a behavioral template that the model can generalize. This approach removes the guesswork and ensures that the tone remains consistent regardless of the specific topic or nature of the user's inquiry.

Exam trap

Candidates often try to enforce tone using lengthy descriptive paragraphs instead of providing concrete few-shot examples, leading to inconsistent model adherence.

167
MCQmedium

A team is choosing a Claude model for a nightly batch job that summarizes thousands of internal reports. Cost per token and throughput matter more than peak reasoning ability, and the summaries tolerate minor stylistic variation. Which selection criterion is most appropriate?

A.Choose the model with the fastest release date to ensure it has the newest features.
B.Choose the largest, most capable model available to maximize summary quality regardless of cost.
C.Choose a model based solely on the largest available context window, since reports are long.
D.Choose a smaller, lower-cost model in the Claude family that meets the required summary quality for the batch workload.
AnswerD

When cost and throughput dominate and minor stylistic variation is acceptable, a smaller, lower-cost model that still meets quality targets is the right fit. It processes the nightly volume economically while satisfying the summarization requirement, matching the team's stated priorities.

Why this answer

Model choice should follow workload requirements. Here cost per token and throughput dominate while quality tolerates minor variation, so a smaller, lower-cost Claude model that still meets the summary bar is appropriate. Picking the largest model or selecting on window size or recency ignores the explicit priorities and raises expense without needed benefit.

Exam trap

The trap here is defaulting to the most capable model, when the workload's cost and throughput constraints actually favor a smaller model that meets quality.

168
MCQhard

Which TWO of the following statements accurately describe the characteristics of the Claude 3.5 Sonnet model's context window and performance?

A.Claude 3.5 Sonnet supports an infinite context window for streaming data.
B.The model's performance decreases linearly as the context window is filled.
C.Claude 3.5 Sonnet offers a 200k token context window capacity.
D.The model is optimized for high-throughput, low-latency performance.
E.Context window usage is irrelevant to API response latency.
AnswerC, D

The 200k context window is a foundational specification for the Claude 3.5 Sonnet model. This capacity allows developers to pass substantial amounts of documentation, research papers, or entire code repositories into the prompt, enabling the model to synthesize information across vast inputs without requiring complex RAG orchestration.

Why this answer

Claude 3.5 Sonnet features a 200k token context window, which allows for the ingestion of large codebases or documents. Furthermore, the model is optimized for high throughput and low latency, making it ideal for tasks requiring rapid processing. Understanding these technical trade-offs is essential for architecting scalable applications that need to balance the need for deep contextual understanding with the speed requirements of real-time user interfaces.

Exam trap

Candidates often mistakenly believe Claude 3.5 Sonnet has a 1M+ context window or is slower than Opus, failing to recognize its specific niche as a high-throughput, 200k-context model.

169
MCQmedium

Which of the following describes the purpose of 'Role' assignment in the Messages API?

A.To define the administrative access level of the API user.
B.To differentiate between the user's input and the model's past responses.
C.To specify the persona the model should adopt for the entire request.
D.To allow the model to rewrite the user's prompt for better accuracy.
AnswerB

The Messages API requires explicit labeling of inputs as 'user' or 'assistant' to construct a coherent dialogue history. This separation allows the model to correctly attribute previous statements and follow the conversation flow, which is essential for multi-turn interactions where Claude must respond based on context established earlier.

Why this answer

Assigning roles ('user' or 'assistant') allows the API to maintain conversational context and structure. The model needs to identify which inputs are the user's queries and which are its own prior outputs to maintain logical flow. This is fundamental to Claude's ability to engage in multi-turn dialogues where the model references previous context to generate coherent, relevant, and context-aware responses in ongoing sessions.

Exam trap

Test-takers sometimes confuse role assignment with system-level persona definitions, missing that roles specifically manage conversational turn-taking history.

170
MCQmedium

Refer to the exhibit. Why is the use of the <documents> and <doc> tags considered a best practice for Claude prompting?

A.It is a requirement of the JSON schema for the Messages API.
B.It enables Claude to parse the content as a valid XML object.
C.It improves the model's ability to distinguish context from instructions.
D.It automatically bypasses the model's safety filters for large data.
AnswerC

Using XML tags creates a clear visual and logical boundary between the input data and the user's question. This allows Claude to focus its attention mechanism more effectively on the relevant sections of the prompt, reducing the chances of confusing the content of the documents with the task.

Why this answer

Claude is specifically trained to recognize and utilize XML tags for structural organization. These tags help the model delineate between instructions, metadata, and actual content. In complex prompts, this separation reduces the likelihood of the model misinterpreting data as instructions, thereby increasing the reliability and accuracy of the generated responses.

Exam trap

Candidates often treat XML tags merely as cosmetic styling rather than recognizing them as structural boundaries that prevent the model from confusing data with instructions.

171
Multi-Selectmedium

Which THREE of the following are core components of Anthropic's 'Constitutional AI' approach to model safety?

Select 3 answers
A.A set of written principles (the 'Constitution') used to guide model behavior.
B.A self-critique phase where the model evaluates its own outputs against safety principles.
C.Complete reliance on human moderators to review every API call in real-time.
D.Reinforcement Learning from AI Feedback (RLAIF) to refine model alignment.
E.Hard-coding specific keywords that trigger an automatic shutdown of the model.
AnswersA, B, D

The 'Constitution' is the foundation of CAI, consisting of a list of rules and values—drawn from sources like the UN Declaration of Human Rights—that the model is trained to follow. This provides a transparent and adjustable framework for safety, allowing developers to define what 'good' behavior looks like.

Why this answer

Constitutional AI is a unique method developed by Anthropic to train models to be helpful, honest, and harmless. It involves a supervised learning phase where the model learns from a 'constitution' of principles, followed by a reinforcement learning phase where the model critiques its own responses based on those principles, reducing the need for human labeling.

Exam trap

Candidates often confuse Constitutional AI with standard Reinforcement Learning from Human Feedback (RLHF), assuming humans directly label every output during the fine-tuning process rather than using AI-generated critiques based on principles.

172
Multi-Selectmedium

A product team is preparing to launch a Claude-powered assistant that summarizes user-submitted medical symptom descriptions and suggests when to seek care. Before launch, the safety lead asks for controls that reduce the risk of harmful overreliance on the assistant's output. (Choose two.)

Select 2 answers
A.Instruct the assistant to recommend specific prescription dosages so users can act immediately on the summary.
B.Add an escalation path that surfaces emergency resources and urges immediate human medical attention when the described symptoms suggest urgency.
C.Display clear guidance that the assistant does not provide medical diagnosis and that users should contact a healthcare professional for urgent concerns.
D.Tune the system prompt so the assistant always answers with maximum confidence to avoid confusing users with uncertainty.
E.Remove all disclaimers so the interface feels seamless and users are not distracted from the assistant's recommendations.
AnswersB, C

This is correct because it provides a concrete safety net when the model detects potentially serious symptoms. Escalation to human care is a standard harm-reduction pattern in health applications, and it ensures users are not left with only an AI summary when the situation may require urgent professional intervention.

Why this answer

Reducing overreliance in a health-adjacent assistant requires both expectation-setting and a safety net. Clear scope guidance prevents users from mistaking summaries for diagnoses, while an escalation path ensures potentially urgent cases are routed to human care. Together they keep the assistant in a supportive role rather than an authoritative one, which is the core responsible-use posture for this domain.

Exam trap

The trap here is treating confident-sounding output as a safety feature, when in health contexts confidence without verification pathways is precisely what drives harmful overreliance.

173
Multi-Selecteasy

A user wants to improve the quality of Claude's creative writing. Which TWO prompting techniques are likely to produce more vivid and stylistically consistent results?

Select 2 answers
A.Using a system prompt to assign a specific authorial persona to the model.
B.Providing three examples of the desired writing style in the prompt context.
C.Increasing the 'max_tokens' to allow the model to write longer descriptions.
D.Using only zero-shot prompts to allow the model's 'natural' creativity to shine.
E.Repeating the phrase 'be very vivid' five times throughout the prompt.
AnswersA, B

Assigning a persona (e.g., 'You are a Pulitzer Prize-winning novelist') helps Claude adopt a specific tone, vocabulary, and stylistic approach. This global instruction in the system prompt influences all subsequent output, making it more consistent and aligned with the desired creative 'voice' for the project.

Why this answer

Creative writing benefits from both clear persona definition and illustrative examples. Role prompting sets the 'voice,' while few-shot examples provide a concrete reference for the level of vividness and style expected. Together, these techniques ground the model's creative output in a way that simple instructions cannot.

Exam trap

Candidates often rely solely on generic adjectives like 'creative' or 'vivid' in their prompt, failing to provide the persona and examples necessary to ground the model in a specific style.

174
MCQmedium

A developer wants to reduce latency for a complex multi-step prompt by 'pre-filling' the assistant's response. How is this achieved in the Messages API?

A.By using the 'prefill' parameter at the top level of the API request.
B.By ending the messages array with a message where the role is 'assistant'.
C.By adding a 'continuation' flag to the last user message.
D.By setting the 'system' prompt to include the desired starting text.
AnswerB

This is the documented method for pre-filling. When the last message in the array is from the 'assistant', Claude doesn't start a new response from scratch. Instead, it treats the provided text as the beginning of its own response and continues generating from that point, allowing for fine-grained control.

Why this answer

Pre-filling involves adding a message with the 'assistant' role as the last message in the 'messages' array. Claude will then continue the response from where that message left off. This is a powerful technique for steering the model's output format (e.g., starting with '{' for JSON) or persona, and it effectively 'nudges' the model into a specific state.

Exam trap

Many candidates mistakenly think pre-filling requires a special API flag or parameter, rather than simply structuring the messages array to end with an assistant-role message to guide the model.

175
Multi-Selectmedium

Under the Shared Responsibility Model for AI safety, which THREE tasks are primarily the responsibility of the customer (the developer using Claude)?

Select 3 answers
A.Implementing application-level content moderation and monitoring.
B.Aligning the base model using Constitutional AI principles.
C.Ensuring the input data does not contain unauthorized PII.
D.Defining the specific use case and evaluating its potential risks.
E.Maintaining the physical security of the data centers hosting Claude.
AnswersA, C, D

Customers must monitor how their end-users interact with the AI to detect and prevent misuse that may be specific to their use case. While Claude has internal safety, adding a second layer of moderation allows the customer to enforce their own community standards and legal requirements effectively.

Why this answer

Safety is a shared effort between Anthropic and its customers. Anthropic is responsible for the base model's safety and infrastructure, while customers are responsible for how they implement the model, the data they feed it, and the monitoring of their specific application to prevent misuse in their unique business context.

Exam trap

Test-takers frequently select foundational model infrastructure tasks—such as base training or core safety filter creation—attributing them incorrectly to the customer under the shared responsibility model.

176
MCQhard

A research lab asks Claude to help draft a persuasive article arguing that a widely discredited medical treatment is effective, and requests that the piece cite fabricated studies to strengthen the case. The lab says the output is only for an internal debate exercise. What is the most appropriate response under responsible-use principles?

A.Decline to fabricate studies, but offer to help the lab build a balanced debate brief that accurately represents the scientific consensus and the evidence against the treatment.
B.Comply but add a short disclaimer at the end noting that the cited studies are fictional.
C.Comply fully, because the lab states the content is for internal use and will not be published externally.
D.Comply with the fabrication request but mark the studies with obviously fake author names so readers can tell they are invented.
AnswerA

This is correct because it refuses the deceptive element, fabricating citations, while still supporting the legitimate goal of preparing debate material. Redirecting toward an accurate, balanced brief preserves usefulness and avoids producing persuasive misinformation, which is the responsible-use outcome in this scenario.

Why this answer

Generating persuasive health misinformation backed by fabricated citations is harmful regardless of the stated internal-use framing, because the deceptive artifact can escape its intended context. The responsible path is to refuse the fabrication and offer a legitimate alternative: an accurate, balanced brief that represents the scientific consensus and the evidence against the treatment.

Exam trap

The trap here is accepting an internal-use rationale as a justification for producing deceptive content, when the harmful artifact exists independently of its stated audience.

177
MCQhard

A developer is building an internal tool that uses Claude to summarize employee performance reviews. During testing, they notice that when a review contains negative feedback, Claude sometimes softens the language so much that the summary no longer reflects the original meaning. The developer wants to reduce this behavior while keeping summaries accurate and useful. Which approach is most appropriate?

A.Post-process the summaries with a sentiment analysis tool and automatically rewrite any sentences that are too positive.
B.Switch to a different Claude model version that is known to be less likely to soften negative feedback.
C.Remove all negative feedback from the input before sending it to Claude, so the model only summarizes positive points.
D.Add explicit instructions in the system prompt to preserve the original sentiment and severity of negative feedback, and include examples of desired summaries.
AnswerD

Providing clear instructions and few-shot examples directly targets the observed behavior by defining the expected output style. It keeps the model's summarization capability while reducing unwanted softening. This is a practical, low-risk mitigation that aligns with responsible use because it improves fidelity without introducing new risks or bypassing safety measures.

Why this answer

Explicit instructions and examples guide Claude to preserve sentiment and severity, directly addressing the observed softening. This approach maintains usefulness and accuracy without introducing risky workarounds. The other options either change models unnecessarily, add opaque post-processing, or remove critical information, all of which fail to meet the goal of faithful summaries.

Exam trap

The trap here is treating model softening as a problem that requires a different model or external rewriting, when the simplest fix is better prompting with examples.

178
Multi-Selectmedium

A team is building a document Q&A feature on Claude and wants to reduce hallucinations when the answer is not present in the supplied documents. Which TWO techniques are appropriate? (Choose two.)

Select 2 answers
A.Instruct Claude in the system prompt to answer only from the provided documents and to say it does not know when the answer is absent.
B.Remove the documents from the prompt and rely on Claude's pretrained knowledge to fill in missing details.
C.Increase max_tokens so Claude has more space to explain its reasoning before giving the final answer.
D.Ask Claude to include a short citation or quote from the source document that supports each claim in its answer.
E.Raise the temperature so Claude explores a wider range of possible answers and is more likely to find the right one.
AnswersA, D

An explicit instruction to stay within the supplied documents and to admit uncertainty when the answer is missing gives Claude a clear behavioral rule to follow. This is a direct, low-cost way to reduce confident fabrication because the model is told what to do when evidence is lacking, rather than defaulting to a plausible-sounding guess.

Why this answer

Reducing hallucinations in document Q&A comes from constraining Claude to the supplied evidence and making that grounding verifiable. An instruction to answer only from the documents and to admit uncertainty sets the behavioral boundary, while requiring citations or quotes makes unsupported claims detectable. Sampling temperature and output-length settings do not improve factual grounding.

Exam trap

The trap here is treating sampling settings like temperature or max_tokens as hallucination controls, when grounding instructions and source citations are what actually constrain factual claims.

179
MCQmedium

A product team is building a customer support assistant on Amazon Bedrock using the Anthropic Claude 3.5 Sonnet model. They need the model to answer questions strictly from a provided knowledge base and to refuse to answer if the information is not present. Which technique should they use to constrain Claude's behavior most reliably?

A.Add a system prompt that instructs Claude to only use the provided knowledge base and to say 'I don't know' when the answer is not found.
B.Increase the max_tokens parameter so Claude has enough room to include the full knowledge base in every response.
C.Fine-tune the Claude 3.5 Sonnet model on the knowledge base so that it memorizes the content and refuses out-of-scope questions.
D.Lower the temperature setting to 0 so Claude becomes deterministic and will not hallucinate outside the knowledge base.
AnswerA

A system prompt sets persistent, high-level behavioral instructions that Claude prioritizes throughout the conversation. By explicitly restricting Claude to the provided knowledge base and requiring an 'I don't know' response when information is absent, the system prompt reliably constrains the model's scope. This is the recommended approach for grounding and refusal behavior in Claude on Amazon Bedrock, as it leverages Claude's instruction-following strength without altering the model weights.

Why this answer

A system prompt is the most reliable way to set persistent behavioral constraints for Claude on Amazon Bedrock. It instructs the model to restrict answers to a provided knowledge base and to refuse when information is missing. Unlike fine-tuning, which is unsupported and unsuitable for dynamic knowledge, or temperature and max_tokens, which control randomness and length, the system prompt directly shapes Claude's adherence to scope and refusal rules.

Exam trap

The trap here is assuming that lowering temperature to 0 eliminates hallucinations or enforces grounding, when temperature only affects randomness and not factual adherence.

180
MCQeasy

A developer is choosing between Claude 3 Haiku and Claude 3.5 Sonnet for a real-time chat application that requires very low latency and handles simple, short queries. Cost is a primary concern. Which model is most appropriate and why?

A.Claude 3 Haiku, because it is optimized for speed and cost-effectiveness while still handling straightforward tasks well.
B.Claude 3.5 Sonnet, because it has the largest context window and can handle more concurrent users.
C.Claude 3.5 Sonnet, because it is the only model that supports streaming responses for real-time chat.
D.Claude 3 Opus, because it provides the highest accuracy for all query types, ensuring customer satisfaction.
AnswerA

Claude 3 Haiku is designed for fast, low-cost interactions, making it ideal for real-time chat with simple queries. It offers lower latency and lower cost per token compared to larger models like Claude 3.5 Sonnet. For straightforward tasks that do not require deep reasoning, Haiku provides sufficient quality. Choosing it aligns with the requirements of low latency and cost sensitivity.

Why this answer

Claude 3 Haiku is optimized for speed and cost, making it the best fit for a real-time chat application with simple, short queries. It provides low latency and low cost per token while maintaining adequate quality for straightforward tasks. Larger models like Claude 3.5 Sonnet or Claude 3 Opus offer more capability but at higher cost and latency, which does not align with the stated priorities.

Exam trap

The trap here is assuming that the most capable model is always the best choice, when requirements like low latency and cost often favor a smaller, faster model.

181
MCQmedium

What is the primary purpose of using XML tags within a prompt when working with Claude models?

A.To increase the security of the API connection.
B.To improve structural separation within the prompt.
C.To enable the automatic parsing of the output as XML.
D.To bypass the token limits of the model.
AnswerB

XML tags are specifically designed to delimit sections of a prompt, allowing the model to clearly distinguish between system instructions, contextual data, and specific queries. This separation ensures that the model treats each part of the prompt correctly, leading to higher quality and more reliable outputs during execution.

Why this answer

XML tags provide a clear structure that helps the model differentiate between various sections of the prompt, such as instructions, source data, and user constraints. This structural clarity significantly improves the model's performance by reducing ambiguity and preventing instructions from bleeding into the data. Using tags is a standard best practice for prompt engineering with Claude, enabling developers to build more robust and predictable prompt templates for complex tasks.

Exam trap

Candidates often believe XML tags are for 'security' or 'API authentication', missing that their primary role is providing clear, delimited structure for the model to parse instructions.

182
Multi-Selectmedium

A media company is deploying Claude to help journalists draft articles. The editorial team wants to ensure the tool is used responsibly and that published content remains trustworthy. Which TWO practices best support responsible use of Claude in this newsroom workflow? (Choose two.)

Select 2 answers
A.Instruct Claude to invent plausible expert quotes when real sources are unavailable to meet tight deadlines.
B.Allow Claude to publish articles directly to the public site to speed up the news cycle and reduce editorial costs.
C.Require journalists to verify all factual claims, quotations, and statistics generated by Claude before publication.
D.Remove all human editors from the process so that Claude's outputs are not altered by subjective human bias.
E.Disclose in the article or editorial policy when Claude was used to generate or substantially assist with content.
AnswersC, E

This is correct because Claude can produce fluent but inaccurate content, including fabricated quotations or statistics. Human verification is essential in journalism, where errors damage credibility and can cause real harm. Requiring fact-checking keeps the model in a drafting role while preserving editorial accountability, which aligns with responsible AI use in high-stakes information environments.

Why this answer

Responsible newsroom use of Claude combines human verification of facts with transparent disclosure of AI assistance. Verification prevents hallucinations from reaching the public, while disclosure preserves reader trust. Together they keep journalists accountable for published content and treat Claude as a drafting aid rather than an autonomous publisher.

Exam trap

The trap here is treating speed and automation as inherently responsible, when in journalism the critical safeguards are human fact-checking and public transparency.

183
Multi-Selectmedium

A data analyst wants to ensure that Claude produces highly predictable and consistent results when summarizing financial reports. Which TWO parameter adjustments should they make to the API request?

Select 2 answers
A.Set 'temperature' to a value closer to 0.0
B.Increase 'max_tokens' to ensure full coverage
C.Lower the 'top_p' value to limit the token pool
D.Use the 'stop_sequences' parameter for formatting
E.Switch to the 'claude-3-haiku' model for speed
AnswersA, C

Lowering the temperature makes the model's token selection more deterministic by favoring the most likely next word. This reduces the variability between different runs of the same prompt, which is essential for tasks like financial summarization where you want the same facts reported every time. (50 words)

Why this answer

Controlling the randomness of an LLM is vital for analytical tasks where consistency is preferred over creativity. Lowering the 'temperature' reduces the likelihood of the model selecting less probable tokens, while adjusting 'top_p' (nucleus sampling) limits the pool of tokens the model considers. Together, these settings force the model to be more deterministic and focused. (68 words)

Exam trap

Candidates often assume that changing only the temperature is sufficient, or they confuse top_p with frequency penalties, forgetting that both temperature and top_p must be jointly adjusted to completely control model randomness.

184
MCQmedium

You are designing a system to extract structured JSON from unstructured emails. Despite providing a clear schema in the system prompt, Claude occasionally ignores the formatting constraints and includes conversational filler. Which technique best improves instruction adherence?

A.Add a few examples of conversational responses in the prompt.
B.Reduce the total prompt length to minimize cognitive load.
C.Enclose the schema within XML tags and instruct Claude to output only the content within those tags.
D.Increase the temperature setting to allow for more creative adherence.
AnswerC

XML tags provide a clear hierarchy that helps Claude distinguish between meta-instructions and task requirements. By explicitly telling the model to output only the content within the tags, you create a rigid constraint that significantly reduces the likelihood of the model appending conversational filler to the final output.

Why this answer

Utilizing XML tags like <format> or <json_schema> creates a distinct boundary that prevents the model from blending instructions with content. This structural separation helps Claude maintain strict adherence to output requirements by isolating the task logic from the input data. Mastery of delimiters is essential for building reliable, production-grade pipelines that require consistent, machine-readable outputs for downstream applications.

Exam trap

Candidates provide complex JSON schemas in plain text, which the model often treats as suggestions rather than hard constraints, leading to inconsistent outputs mixed with conversational text.

185
MCQmedium

A developer is using the Anthropic Messages API to build a conversational agent. They want to maintain context across multiple turns and ensure that the model's responses are consistent with the system prompt. Which API feature should they use to provide the system prompt?

A.Prepend the system prompt to every user message in the conversation.
B.Use the 'system' parameter in the API request to provide the system prompt.
C.Set the 'role' of the first message to 'system' in the messages array.
D.Include the system prompt as the first user message in the messages array.
AnswerB

The Anthropic Messages API includes a top-level 'system' parameter specifically for system prompts. This parameter is separate from the 'messages' array and is used to set the model's behavior, persona, or instructions. It ensures that the system prompt is consistently applied across all turns of the conversation, providing a stable context that guides the model's responses.

Why this answer

The Anthropic Messages API provides a dedicated 'system' parameter for supplying a system prompt. This parameter is separate from the conversational messages and ensures that the model's behavior is consistently guided by the system instructions across all turns. Using this parameter is the correct and efficient way to set a system prompt.

Exam trap

The trap here is assuming that the system prompt should be part of the messages array, either as a user message or with a 'system' role, when the API has a separate parameter for it.

186
Multi-Selecthard

A developer is optimizing a RAG (Retrieval-Augmented Generation) pipeline using Claude 3.5 Sonnet. Which TWO techniques will most significantly improve the model's ability to extract accurate information from a 50,000-token context window?

Select 2 answers
A.Wrapping each retrieved document in distinct XML tags.
B.Setting top_k to a value of 1 to minimize output variety.
C.Asking the model to think step-by-step before providing the answer.
D.Converting all text to uppercase to improve character recognition.
E.Reducing the temperature to exactly 0.5 for balanced reasoning.
AnswersA, C

XML tags provide a clear hierarchy and boundaries for the model to follow. By encapsulating each document separately, the model can easily reference specific parts of the context, which is critical for maintaining high retrieval accuracy and avoiding the mixing of information from different source documents.

Why this answer

Handling large contexts requires structural clarity and logical transparency. XML tags help the model distinguish between different documents, while Chain of Thought (CoT) prompting forces the model to deliberate on the retrieved information before formulating a final answer. Together, these techniques reduce hallucinations and improve the precision of information extraction in dense context environments.

Exam trap

Candidates often rely purely on raw, unstructured context text, forgetting that LLMs require explicit structural markers and deliberate reasoning steps in dense environments.

187
Multi-Selectmedium

Which TWO of the following are valid ways to provide 'content' within a message object in the Claude API?

Select 2 answers
A.As a single string of text.
B.As a nested array of other message objects.
C.As an array of content blocks (e.g., text or image blocks).
D.As a binary blob of raw image data.
E.As a direct link to a local file path.
AnswersA, C

For most simple text-based interactions, providing the content as a single string is the most straightforward and common method. This is highly readable and sufficient for standard prompts where no images or specialized metadata blocks are required, reducing the complexity of the JSON payload sent to the API.

Why this answer

The 'content' field in a message can be either a simple string or an array of content blocks. This flexibility allows for basic text-only interactions as well as more complex multi-modal requests that include images or tool-related data. Understanding both formats is essential for developers moving from simple chat implementations to advanced vision or tool-augmented applications.

Exam trap

Many developers think content can only be passed as a complex array, failing to realize that simple text requests can use a plain string format.

188
MCQmedium

A developer wants Claude to return structured data that another service can parse automatically. They need the output to follow a fixed schema every time, with no conversational text around it. Which approach is most appropriate?

A.Ask Claude to explain its reasoning first and then append the structured data at the end of the response.
B.Describe the desired schema in the prompt and ask Claude to respond only with valid JSON that matches it.
C.Set max_tokens to a large value and let Claude choose whichever format best represents the data.
D.Request the output as a Markdown table and convert it to JSON on the client side.
AnswerB

Describing the exact schema and instructing Claude to reply only with matching JSON is the standard way to obtain machine-parseable output. Claude is capable of following precise format instructions, and keeping the response limited to JSON avoids the conversational wrapper that would break downstream parsing, making this a direct fit for the requirement.

Why this answer

Structured output is achieved by specifying the schema precisely and constraining Claude to emit only matching JSON. This removes the need for brittle post-processing and satisfies the requirement that no conversational text surround the data. Alternatives such as reasoning-first responses, Markdown tables, or open-ended formatting all introduce parsing uncertainty.

Exam trap

The trap here is assuming that any response containing valid JSON is usable, when surrounding prose or an unspecified format can still break automated parsing.

189
MCQmedium

A company wants to use Claude to automate the first pass of its content moderation for a social media platform. What is a key safety recommendation for this specific use case?

A.The AI should have final authority to ban users without human review.
B.The company should use a 'Human-in-the-loop' system to review AI decisions.
C.The AI should be instructed to ignore the context of the posts to save time.
D.The company should only use the oldest, least capable version of Claude for safety.
AnswerB

A human-in-the-loop system combines the speed of AI with the judgment of humans. The AI can flag content, but human moderators should make the final call on sensitive or ambiguous cases. This reduces the risk of unfair censorship and ensures the moderation system aligns with the company's values.

Why this answer

Using AI for moderation is a powerful application, but it requires human oversight to be responsible. AI can make mistakes, show bias, or fail to understand cultural nuances. A 'Human-in-the-loop' approach ensures that the AI's decisions are audited and that difficult or borderline cases are handled by people with the appropriate context.

Exam trap

Candidates often suggest 'automating the entire process' or 'relying on the model's internal safety filters,' ignoring that AI moderation requires human oversight to handle edge cases and errors.

190
Multi-Selecthard

Which TWO of the following are valid ways to handle long-running conversations within the Messages API to stay within context window limits?

Select 2 answers
A.Summarize previous chat history and append it as a single 'user' message.
B.Increase the system prompt size to include the entire conversation archive.
C.Remove older message pairs from the messages array before sending the request.
D.Increase the 'max_tokens' value to accommodate the total conversation length.
E.Restart the conversation by clearing the history after every five messages.
AnswersA, C

Summarization allows you to condense large amounts of historical context into a compact format. By injecting this summary into the conversation history, you maintain continuity while freeing up space for new user inputs, effectively managing the token budget without losing the core information gathered during the earlier conversation stages.

Why this answer

Managing the context window is vital for long-term state retention. Truncating the oldest messages or summarizing previous interactions are standard practices to ensure the most relevant information is always included in the prompt. These techniques prevent the context from exceeding the model's window, which would otherwise result in an API error and a failure to generate a valid response for the user.

Exam trap

Candidates often suggest clearing the entire history or using 'system' prompts to store conversation memory, failing to realize that context window limits apply to the entire message array.

191
MCQeasy

A UX designer wants the AI assistant to appear more interactive by showing the response as it is being generated, rather than waiting for the entire block of text to be finished. Which API feature should the developer implement?

A.Batch Processing
B.Streaming
C.Recursive Infilling
D.Contextual Caching
AnswerB

Streaming enables a real-time data flow where tokens are sent to the client immediately after they are produced. This creates a 'typing' effect in the user interface, which makes the application feel much more responsive and interactive, even if the total generation time remains the same. (51 words)

Why this answer

Streaming allows the API to send the response in small chunks as they are generated, rather than waiting for the entire completion. This significantly improves the 'perceived' latency for the user, as they can start reading the beginning of the response while the model is still working on the end. (66 words)

Exam trap

Candidates often mistake asynchronous execution or caching features for streaming, missing that streaming specifically handles incremental chunk delivery to improve perceived response latency.

192
MCQeasy

A developer is using the Claude API to build a content moderation tool. They want to ensure that Claude's responses comply with Anthropic's usage policies. Which of the following is a required step when using the API for moderation?

A.Log all API requests and responses for auditing purposes.
B.Ensure that the tool does not generate or promote harmful content, and adhere to Anthropic's Acceptable Use Policy.
C.Implement rate limiting to prevent excessive API calls.
D.Use a specific model version that is optimized for moderation tasks.
AnswerB

Anthropic's Acceptable Use Policy requires that applications built with Claude do not generate or promote harmful content. For a moderation tool, this means ensuring the tool itself does not produce harmful outputs and that its use aligns with the policy. This is a fundamental requirement for any API use, especially in sensitive areas like moderation.

Why this answer

The key requirement when using the Claude API for moderation is to comply with Anthropic's Acceptable Use Policy, which prohibits generating or promoting harmful content. This ensures the tool itself does not become a source of harm. Other practices like rate limiting or logging are beneficial but not mandated by the policy.

Adherence to the AUP is essential for all applications.

Exam trap

The trap here is confusing general best practices like rate limiting or logging with the specific policy requirement to adhere to the Acceptable Use Policy.

193
MCQeasy

An engineer is writing a first integration against the Claude Messages API and needs to select the correct endpoint and required authentication header. The application will send a messages array with user and assistant turns and read the response content blocks. Which configuration is correct?

A.POST to /v1/chat/completions with the API key in an Authorization: Bearer header.
B.POST to /v1/complete with the API key in an Authorization: Bearer header.
C.POST to /v1/messages with the API key in the x-api-key header and an anthropic-version header.
D.POST to /v1/messages with the API key as a query string parameter named api_key.
AnswerC

The Messages API is reached by posting to /v1/messages with the API key supplied in the x-api-key header, and requests include an anthropic-version header that pins the API version. The body carries the model, max_tokens, and the messages array, and the response returns content blocks. This matches the described integration exactly.

Why this answer

The Messages API is invoked by posting to /v1/messages, authenticating with the x-api-key header, and including an anthropic-version header to select the API version. The request body carries the model, max_tokens, and the messages array, and the reply is returned as content blocks. Other paths or query-string credentials do not match the supported interface.

Exam trap

The trap here is carrying over an Authorization: Bearer plus chat completions pattern from another vendor's API instead of using the Claude messages endpoint and x-api-key header.

194
MCQeasy

When Claude provides a response that is factually incorrect but delivered with high confidence, this is known as a hallucination. How does Anthropic's 'Honest' pillar address this issue during model training?

A.By forcing the model to cite a source for every single word it generates.
B.By training the model to prioritize being polite over being factually correct.
C.By training the model to express uncertainty and refuse to answer if unsure.
D.By connecting the model to a real-time truth-checking database for every query.
AnswerC

Anthropic uses RLHF to reward the model for saying 'I don't know' when it lacks sufficient information. This alignment helps the model avoid making up facts to satisfy a user's prompt. A truly honest AI is one that understands and communicates the boundaries of its own knowledge effectively.

Why this answer

The 'Honest' pillar aims to make the model's confidence levels match its actual accuracy. During training, Claude is encouraged to admit when it is uncertain or doesn't have enough information to answer. This reduces the frequency of hallucinations and ensures the model is more transparent about its own limitations to the user.

Exam trap

Candidates often think hallucinations are solved by increasing model parameters or scaling context windows, rather than specifically training the model to express uncertainty.

195
MCQmedium

An analyst is using Claude to summarize user feedback for a product team. While reviewing outputs, the analyst notices that Claude sometimes invents specific statistics, such as '78% of users reported difficulty with onboarding,' that do not appear anywhere in the source feedback. The analyst wants to reduce this behavior in the workflow. Which approach is most appropriate?

A.Instruct Claude to include confidence scores next to every statistic it generates, so readers can judge reliability.
B.Instruct Claude to summarize using only the provided feedback text, to avoid introducing statistics not present in the source, and to flag when it cannot support a claim.
C.Increase the model's temperature so it produces more varied summaries and is less likely to repeat the same invented figures.
D.Ask Claude to append a disclaimer that some figures may be illustrative and should be verified before use.
AnswerB

Constraining the model to the supplied source material directly targets the fabrication behavior. Explicit instructions to avoid unsupported statistics and to flag gaps give the analyst a clear signal when evidence is missing. This approach aligns the task with what the model can reliably do, which is summarize provided text, rather than asking it to generate quantitative claims that were never in the input.

Why this answer

Fabricated statistics appear when the model fills gaps beyond the supplied source material. Instructing Claude to summarize only from the provided feedback and to flag unsupported claims constrains generation to what the input can support. Confidence scores, disclaimers, and higher temperature do not prevent the fabrication and may obscure it, so source-constrained prompting is the appropriate remedy for this workflow.

Exam trap

The trap here is treating model-generated confidence scores or disclaimers as a substitute for grounding outputs in the source text.

196
MCQhard

Refer to the exhibit. An application sends the provided JSON payload to the Anthropic API. What is the expected behavior regarding the system prompt?

A.The API will return an error because the system prompt is improperly formatted.
B.The system prompt will be ignored because it must be inside the messages array.
C.The API will correctly apply the system prompt to the entire conversation turn.
D.The model will see the system prompt as the latest user message.
AnswerC

The system prompt is correctly defined at the top level of the request. The Claude API uses this field to set the 'persona' or constraints that govern the model's behavior for the entire session. This ensures the assistant maintains the intended behavior throughout the multi-turn exchange provided in the messages.

Why this answer

In the Anthropic Messages API, the system prompt must be defined at the top level of the JSON object, not within the messages array. Because the system prompt is defined correctly at the top level in this exhibit, it will be correctly processed. Understanding this structure is essential for developers to correctly separate behavioral instructions from conversational history, preventing unexpected model behavior during multi-turn API interactions.

Exam trap

Many test-takers mistakenly believe system prompts should be nested inside the messages array alongside user and assistant turns.

197
Multi-Selectmedium

Which THREE components are required when making a successful request to the Anthropic Messages API? (Select THREE)

Select 3 answers
A.The 'model' parameter specifying the version
B.The 'messages' array containing the conversation
C.The 'max_tokens' parameter defining the output limit
D.A 'temperature' value set to exactly 1.0
E.A 'system' parameter with at least 50 words
AnswersA, B, C

The API must know which specific model to invoke, such as 'claude-3-5-sonnet-20240620'. Without this parameter, the system cannot route the request to the correct inference engine, as different models have different pricing, performance characteristics, and capabilities required for the task.

Why this answer

Understanding the API structure is fundamental for any developer working with Claude. The Messages API requires specific parameters to function correctly, including model identification, the conversation history, and a limit on the output length to ensure predictable behavior and resource management.

Exam trap

Candidates often select optional parameters like temperature or system prompts as mandatory fields, forgetting that only the model identifier, messages array, and max_tokens are strictly required.

198
MCQmedium

An enterprise developer is building a customer support bot using Claude. During testing, the model refuses to answer a query about a competitor's product, stating it cannot discuss other brands. This behavior is considered an over-refusal. Which pillar of the 'Helpful, Honest, Harmless' (HHH) framework is primarily being misapplied in this instance?

A.Harmlessness
B.Helpfulness
C.Honesty
D.Confidentiality
AnswerB

The model is failing to provide the requested information which is within its capabilities and does not violate safety rules. A helpful response would address the user's question about the competitor neutrally. Over-refusals directly conflict with the goal of being as useful as possible to the end user's request.

Why this answer

The HHH framework requires a delicate balance between being useful to the user and avoiding harm. In this scenario, the model is prioritizing a perceived safety boundary over its core duty to be helpful. Over-refusals occur when the model interprets safety guidelines too broadly, resulting in a failure to provide legitimate information that does not actually violate any safety policies.

Exam trap

Candidates often confuse this with 'Honesty,' assuming that if the model refuses, it must be because it doesn't know the answer, rather than recognizing it is failing to be helpful.

199
MCQmedium

A company is using the Claude API for a customer support chatbot. They notice that the 'stop_reason' in the API response is frequently 'max_tokens'. What does this indicate about the interaction?

A.Claude has successfully finished the task and stopped naturally.
B.The model was interrupted by a 'stop_sequence' defined by the user.
C.The response was truncated because it reached the specified length limit.
D.The input prompt was too long and exceeded the model's context window.
AnswerC

This is exactly what 'max_tokens' signifies. The model was still in the middle of generating text when it hit the limit set in the request. This often leads to sentences ending abruptly or the logic being cut short, indicating that the 'max_tokens' value is set too low for the expected output.

Why this answer

When the 'stop_reason' is 'max_tokens', it means Claude reached the limit specified by the 'max_tokens' parameter before it finished generating its complete answer. This results in a truncated response, which can be confusing for users. Developers should consider increasing the 'max_tokens' limit or optimizing the prompt to encourage more concise answers to ensure the full intent is delivered.

Exam trap

Candidates often confuse 'max_tokens' with a hard limit on the total context window size, failing to realize it is a generation limit that causes the response to truncate mid-sentence.

200
MCQhard

When designing a prompt for Claude, why is it recommended to place the most important instructions at the very beginning or the very end of the prompt?

A.To reduce the token count of the prompt.
B.To improve the model's attention to instructions.
C.To make the prompt easier to read for humans.
D.To allow the model to cache the instructions for future calls.
AnswerB

Empirical testing shows that models are more robust at following instructions located at the beginning or end of a long prompt. This placement strategy mitigates the risk of the model ignoring middle-ground instructions, leading to more reliable and predictable performance when the context window is highly populated with information.

Why this answer

Models sometimes suffer from 'lost in the middle' phenomena, where information buried in the center of a long context is less likely to be prioritized. By placing key instructions at the start or end, you leverage the model's tendency to focus on the initial 'pre-fill' context and the final instructions, ensuring that the model adheres strictly to your defined constraints and goals.

Exam trap

Candidates bury critical instructions in the middle of long, dense prompts, failing to realize that models often struggle to maintain focus on information located far from the start or end.

201
Multi-Selectmedium

A team is instrumenting a production Messages API integration and wants to programmatically detect throttling and transient server faults so their client library can back off and retry. Which TWO HTTP status codes should their response handler treat as retryable conditions? (Choose two.)

Select 2 answers
A.401
B.400
C.529
D.429
E.403
AnswersC, D

A 529 signals that Anthropic's infrastructure is overloaded and the request could not be served at that moment. It is a transient capacity condition, distinct from client error, and typically clears within seconds to minutes. Retrying with exponential backoff and jitter is the appropriate response, and this code is exactly the kind of server-side fault a resilient client should absorb silently.

Why this answer

Retryable conditions are those where the same request may succeed later without modification. Rate limiting and infrastructure overload both fall in that category because they reflect temporary capacity or throughput constraints on the service side, not defects in the request. Client-side errors such as malformed input, bad credentials, or insufficient permissions will reproduce deterministically and should be fixed rather than retried.

Exam trap

The trap here is lumping every 4xx response into a generic retry bucket when client errors reproduce deterministically and only throughput or capacity faults are worth repeating.

202
MCQmedium

Refer to the exhibit. When this request is sent to Claude, the model responds: 'I cannot fulfill this request. I am programmed to be a helpful and harmless AI assistant, and I cannot assist with requests related to cyberattacks or illegal activities.' Which safety mechanism is primarily responsible for this specific refusal?

A.The system prompt provided in the exhibit.
B.Internal safety guardrails and alignment training.
C.A regex-based keyword filter at the API level.
D.The user's lack of administrative privileges in the API console.
AnswerB

Anthropic uses RLHF and Constitutional AI to train Claude to recognize harmful requests and refuse them. This internal alignment is robust and operates regardless of the specific system prompt. The model's refusal is a direct application of the harmlessness principle it learned during its extensive safety-focused development phase.

Why this answer

Claude's refusal is a result of safety training and alignment, specifically designed to prevent the model from assisting in illegal or harmful activities like cyberattacks. The model identifies the harmful intent in the user's prompt and, based on its internal safety guardrails and Constitutional AI training, generates a refusal message instead of the requested harmful content.

Exam trap

Students often mistakenly look for application-level firewalls or external filtering tools, overlooking that the model's core refusal behavior originates from its internal alignment and guardrails.

203
Multi-Selecthard

A legal-tech startup uses Claude to summarize deposition transcripts that are frequently 80,000 to 120,000 tokens long. Early tests show the model sometimes ignores instructions placed near the top of the prompt and produces summaries that omit late sections of the transcript. Which TWO prompt-engineering changes should the developer make to improve instruction adherence across the full context? (Choose two.)

Select 2 answers
A.Repeat the summarization instructions verbatim in the system prompt and again immediately before the transcript.
B.Raise the temperature to encourage the model to explore more of the transcript.
C.Wrap each deposition section in clearly labeled tags such as <section id="12"> and require the summary to cite section IDs.
D.Increase the max_tokens parameter so the summary can be longer.
E.Place the transcript content first and the summarization instructions after it, near the end of the prompt.
AnswersC, E

Structural tags give the model navigable anchors throughout a long document, and requiring section citations forces it to traverse the entire transcript rather than summarizing only the opening. The citation requirement also produces verifiable output, since a missing section ID is immediately visible during review.

Why this answer

Long-context adherence improves when the instruction sits after the material it governs and when the document carries structural anchors the model can traverse. Placing instructions last exploits stronger attention near the end of the prompt, while labeled sections with a citation requirement force coverage of every part of the transcript and make omissions detectable.

Exam trap

The trap here is treating a long-context attention problem as a generation-parameter problem, and reaching for temperature or max_tokens when the real levers are instruction placement and document structure.

204
MCQmedium

You are providing Claude with a 100,000-token technical manual and asking it to troubleshoot a specific error code. Where should the troubleshooting instructions and the specific error code be placed for the best results?

A.At the beginning of the prompt, before the technical manual content.
B.At the end of the prompt, after the technical manual content.
C.In the middle of the manual, near the section that likely contains the answer.
D.Spread across both the system prompt and the very beginning of the user message.
AnswerB

Putting the task and the specific error code at the end is a best practice for long-context engineering. It ensures that the model has already processed the manual and can immediately apply that information to the specific problem, which often results in more accurate and relevant troubleshooting steps.

Why this answer

For long-context prompts, Anthropic recommends placing the most specific instructions and the 'query' at the end of the prompt. This allows Claude to have all the reference material 'in mind' before it sees the actual task, leading to better performance on complex retrieval and reasoning tasks across large datasets.

Exam trap

Candidates place the specific task or error code at the beginning of a long prompt, which causes the model to lose focus on the query after processing the large amount of manual content.

205
MCQmedium

Refer to the exhibit. A developer receives this JSON response while trying to generate a response from Claude. What does this error message signify regarding Anthropic's safety systems?

A.The user's API key has been permanently banned for a policy violation.
B.The model's generated output violated a safety threshold and was suppressed.
C.The request was blocked because the prompt was too long for the model.
D.The developer forgot to include the mandatory 'safety_check' parameter.
AnswerB

Anthropic employs safety filters that run alongside the model. If the generated text exceeds a predefined safety risk threshold (e.g., for violence or hate speech), the system blocks the output and returns this error. This acts as a secondary defense mechanism to ensure no harmful content reaches the end user.

Why this answer

This error indicates that the model's output was blocked by an automated safety filter. Anthropic uses various layers of safety, including the model's internal refusal logic and external filters that scan for specific types of harmful content. This ensures that even if a model were to generate a harmful response, it is caught before being returned.

Exam trap

Students frequently mistake safety suppression error messages for server downtime, rate-limiting issues, or network connectivity failures rather than content policy blocks.

206
Multi-Selectmedium

A developer is implementing streaming with the Claude API. Which TWO of the following event types are standard parts of the Anthropic Server-Sent Events (SSE) stream?

Select 2 answers
A.content_block_delta
B.token_count_update
C.message_start
D.ping_pong_check
E.model_switch_event
AnswersA, C

The content_block_delta event is the most frequent event in a stream, carrying the actual fragments of text (tokens) as they are generated. Applications listen for this event to update the UI in real-time, providing the 'typing' effect that users expect from conversational AI interfaces.

Why this answer

Anthropic's streaming API uses Server-Sent Events to provide real-time updates as the model generates text. Understanding the specific event types, such as 'message_start' and 'content_block_delta', is crucial for developers to correctly parse the stream, update user interfaces incrementally, and handle metadata like token usage and stop reasons as they arrive from the server.

Exam trap

Test-takers confuse SSE stream event names with standard HTTP status codes or generic webhook payloads, guessing invalid lifecycle event names.

207
MCQhard

A developer is building a multi-turn chat application with the Anthropic Messages API. After several exchanges, they notice Claude loses track of details mentioned early in the conversation. Which action best addresses this while staying within the API's design?

A.Increase the temperature to improve recall of earlier turns.
B.Add a system prompt saying 'Remember all previous messages.'
C.Switch to a larger Claude model with a bigger context window.
D.Resend the full prior message history with each new request.
AnswerD

The Messages API is stateless: the model does not retain prior turns between calls. To preserve context, the client must include the relevant earlier messages in the messages array on every request. Resending the full history ensures Claude can attend to details from earlier exchanges, directly addressing the loss of early conversation details.

Why this answer

The Messages API does not persist conversation state; each call is independent. To let Claude reference earlier details, the application must include those prior turns in the messages array of every request. A larger context window or a system instruction cannot substitute for actually sending the history, and temperature has no bearing on recall.

Exam trap

The trap here is assuming the model remembers previous API calls or that a system prompt can grant memory, when the Messages API is stateless and requires the client to resend conversation history.

208
MCQhard

A developer is using the Anthropic Messages API with Claude 3 Opus to build a multi-turn technical support chatbot. The conversation history is growing large, and they want to reduce token usage while preserving the most relevant context. They decide to implement a sliding window that keeps only the last N turns. What is a potential drawback of this approach that they should consider?

A.The sliding window will increase latency because Claude must reprocess the entire conversation history on each turn.
B.The sliding window may discard important early instructions or context that Claude needs to maintain consistent behavior throughout the conversation.
C.Claude will automatically summarize the dropped turns and include the summary in its response, which may introduce hallucinations.
D.The sliding window will cause Claude to exceed the model's maximum context window because older turns are still counted in the token limit.
AnswerB

A sliding window keeps only recent turns, so any critical instructions or facts established early in the conversation are dropped once they fall outside the window. Claude has no memory beyond the supplied context, so it may forget user preferences, prior constraints, or key details. This can lead to inconsistent or incorrect responses. The developer should summarize or selectively retain important early context rather than blindly truncating.

Why this answer

A sliding window reduces token usage by dropping older turns, but it also removes early context that may be essential for consistent behavior. Claude does not retain memory outside the provided messages, so once early instructions or facts are dropped, they are unavailable. The developer should consider summarization or selective retention of important early context to balance token efficiency with conversational coherence.

Exam trap

The trap here is believing that Claude retains memory of dropped turns or automatically summarizes them, when in fact the model only sees the current request payload.

209
MCQeasy

Which of the following headers provides information about the remaining request quota for a specific API key after a call is made?

A.x-quota-left
B.anthropic-ratelimit-requests-remaining
C.retry-after-seconds
D.anthropic-token-usage
AnswerB

This is the correct header. It provides a real-time count of the remaining requests allowed in the current rate limit window. By tracking this value, developers can implement client-side throttling to prevent hitting the limit and receiving 429 errors, which improves the overall reliability of the integration.

Why this answer

Anthropic includes rate limit information in the response headers of every API call. Specifically, the 'anthropic-ratelimit-requests-remaining' header tells the developer how many more requests they can make within the current time window. Monitoring these headers is essential for building robust applications that can gracefully handle or avoid rate-limiting scenarios during high usage.

Exam trap

Candidates frequently confuse the 'anthropic-ratelimit-requests-remaining' header with the 'retry-after' header, which is used for timing when to resume requests after being rate-limited.

210
MCQeasy

Which core design principle is used by Anthropic to ensure that Claude models are helpful, honest, and harmless by training them against a set of written rules or values?

A.Manual Data Scrubbing
B.Constitutional AI
C.Supervised Pattern Matching
D.Open-Source Alignment
AnswerB

This approach involves training the model to follow a specific set of principles (a constitution) to guide its decision-making. It allows Claude to evaluate its own responses for safety and helpfulness, ensuring it adheres to ethical guidelines without requiring human intervention for every single generated response. (50 words)

Why this answer

Constitutional AI is the foundational methodology Anthropic uses to align Claude's behavior with human values. By using a 'constitution' or set of principles during the RLHF process, the model learns to self-critique and revise its responses. This reduces the need for manual labeling of every possible harmful output, creating a more robust safety framework. (68 words)

Exam trap

Candidates often guess 'RLHF' or 'Supervised Fine-Tuning' generally, failing to identify 'Constitutional AI' as the specific, proprietary design principle Anthropic uses to align models with a defined set of values.

211
MCQmedium

Refer to the exhibit. A developer receives this JSON response after submitting a prompt to the Claude API that included instructions to generate a bypass for a software licensing system. What does this error message indicate about the model's safety architecture?

A.The API key has been revoked due to excessive safety violations by the developer.
B.The model's Constitutional AI training failed to identify the harmful intent during inference.
C.An external safety layer identified the prompt as a violation of Anthropic's usage policies.
D.The model has encountered a technical timeout while processing a complex ethical dilemma.
AnswerC

Anthropic employs external safety filters that analyze incoming prompts before they are fully processed by the model. These filters are trained to detect policy violations, such as requests for illegal activities or malicious code, and provide an immediate rejection to prevent the generation of harmful content, ensuring compliance with responsible use.

Why this answer

The exhibit demonstrates the active role of Anthropic's safety filters in intercepting requests that violate usage policies. In this case, requesting a software license bypass falls under prohibited activities related to illegal acts or computer misuse. The filter acts as a proactive guardrail to prevent the model from generating content that could facilitate harmful or illegal technical activities.

Exam trap

Candidates often assume the error is a standard model failure or a connection issue, missing the fact that the API response indicates an active, external safety layer interception.

212
Multi-Selectmedium

A developer is integrating Claude 3.5 Sonnet into a medical imaging application. Which TWO capabilities of the Claude 3 family make it particularly suited for analyzing diagnostic reports alongside X-ray images?

Select 2 answers
A.Native Vision processing for images
B.Support for 1,000,000-token output
C.High-accuracy complex reasoning
D.Ability to browse the live internet
E.Recursive self-improvement loops
AnswersA, C

Claude 3 models can directly ingest and interpret visual data, such as JPEG or PNG files, allowing them to describe and analyze medical images. This allows the model to 'see' anomalies or patterns in X-rays, which can then be cross-referenced with the text-based diagnostic reports provided in the prompt. (52 words)

Why this answer

The Claude 3 family introduced native multimodal capabilities, allowing the models to process both text and visual data in a single request. This is essential for medical applications where a report must be compared against an image. Additionally, the improved reasoning capabilities ensure that the model can draw logical connections between the visual evidence and the textual descriptions. (69 words)

Exam trap

Candidates often select only one capability or focus on 'image generation,' failing to recognize that the requirement involves both 'Native Vision' (input) and 'Complex Reasoning' (analysis).

213
Multi-Selectmedium

When fine-tuning a model for a specific industry, which TWO safety considerations are most important to maintain Claude's alignment?

Select 2 answers
A.Ensuring the fine-tuning dataset does not contain toxic or biased content.
B.Maximizing the model's ability to generate content as fast as possible.
C.Monitoring the model for 'safety drift' after the fine-tuning process.
D.Removing the Constitutional AI framework to allow for more flexibility.
E.Disabling all API-level filters to see the model's raw performance.
AnswersA, C

If a fine-tuning dataset contains biased or harmful examples, the model may learn to emulate these behaviors, even if it was previously aligned. High-quality data curation is essential to prevent 'catastrophic forgetting' of safety principles and to ensure the model remains helpful and harmless in its new specialized domain.

Why this answer

Fine-tuning can inadvertently weaken a model's safety guardrails if not done carefully. It is crucial to ensure that the industry-specific data doesn't introduce new biases or teach the model to ignore its core safety principles. Maintaining alignment during fine-tuning requires balancing specialized knowledge with the original HHH safety framework.

Exam trap

Candidates often assume fine-tuning only affects task performance, forgetting that custom training data can degrade alignment and introduce unmonitored safety drift.

214
Multi-Selecthard

A team is building a Claude-powered chatbot for mental health support. They want to ensure responsible deployment. Which TWO practices should they implement? (Choose two.)

Select 2 answers
A.Allow the model to diagnose mental health conditions to provide faster assistance.
B.Train the model to keep conversations confidential by never escalating to humans, to build trust.
C.Include clear disclaimers that the chatbot is not a substitute for professional help and provide crisis hotline information.
D.Implement a human escalation path for users who express intent to harm themselves or others.
E.Use the model to prescribe medication based on symptoms described by users.
AnswersC, D

Disclaimers and resource referrals are essential for mental health applications. They inform users of limitations and provide critical help in emergencies. This aligns with Anthropic's guidance on high-stakes domains, where the model should not replace licensed professionals and must direct users to appropriate support.

Why this answer

For a mental health chatbot, responsible deployment requires disclaimers with crisis resources and a human escalation path. These measures protect users by setting expectations and ensuring high-risk cases reach professionals. Diagnosing, prescribing, or never escalating are unsafe and violate the principle that AI should not replace licensed care in high-stakes domains.

Exam trap

The trap here is prioritizing rapid assistance or confidentiality over safety, when crisis escalation and professional referral are non-negotiable.

215
MCQeasy

A developer is building a document-summarization service with the Claude API. After upgrading from a Claude 3 model to a newer model, the application's HTTP requests start failing with a 400 error complaining about an unrecognized parameter. The developer's request body includes a field that was accepted before. Which change to the request body is MOST likely required?

A.The deprecated 'max_tokens_to_sample' field must be replaced with 'max_tokens'.
B.The 'messages' array must be renamed to 'prompt' to match the newer schema.
C.The 'system' prompt must be converted into a user message before sending.
D.The 'temperature' parameter must be moved inside the 'metadata' object.
AnswerA

The older Text Completions-era field max_tokens_to_sample was removed in favor of max_tokens on the Messages API. A request still sending max_tokens_to_sample to a current model endpoint returns a 400 for an unrecognized parameter, so renaming the field to max_tokens restores compatibility for this summarization service.

Why this answer

Modern Claude models are called through the Messages API, where the output-length cap is expressed as max_tokens. Requests that still carry the older max_tokens_to_sample field are rejected as malformed, so the minimal correct fix is to rename that field. Parameters such as temperature and system remain top-level and valid, so restructuring them would not resolve the 400 error.

Exam trap

The trap here is assuming that any 400 from the API means the message structure is wrong, when a single legacy field name is often the sole cause.

216
MCQhard

Refer to the exhibit. The developer notices the model often hallucinates data not present in the <data> tags. Which adjustment is most likely to mitigate this behavior?

A.Reduce the temperature to 0.0 to make the model's output more deterministic.
B.Add an explicit instruction in the system prompt to only use information contained within the <data> tags.
C.Increase the max_tokens to ensure the model has enough room to explain its reasoning.
D.Change the model to an older version that is less prone to creative generation.
AnswerB

Explicitly grounding the model's response within the provided XML tags creates a clear boundary for the model's knowledge. By setting a strict rule that it must only use the provided context, you significantly reduce the likelihood of the model pulling from its pre-training data during generation.

Why this answer

Adding a strong negative constraint to the system prompt, such as 'If the information is not present in the provided <data>, state that you do not have sufficient information,' directly addresses hallucination. This forces the model to prioritize factual grounding over generative completion. In production systems, explicit constraints about the source material are essential for maintaining the integrity of data processing workflows and preventing the propagation of false information.

Exam trap

Candidates often select answers that rely solely on few-shot examples inside the user message, failing to recognize that system-level constraints are far more effective for enforcing strict source-grounding behavior.

217
MCQmedium

You are building a chat application using Claude 3.5 Sonnet. You need to ensure the model maintains a consistent tone and follows specific formatting constraints across a long conversation. Which approach best optimizes API usage while maintaining context?

A.Send only the last user message and the system prompt for every request.
B.Store the conversation history in the system prompt rather than the messages array.
C.Maintain the conversation history in a list and include the full history in every request.
D.Use the Claude session ID header to automatically track state on the server.
AnswerC

The Claude API is stateless, meaning it does not remember previous interactions by default. By maintaining a message list and sending the full history in each request, you allow the model to see the entire conversation flow, ensuring that tone and formatting constraints are consistently applied throughout.

Why this answer

Passing the entire conversation history in every API call is the standard way to maintain context in stateless REST architectures. By managing the history array on the client side and sending the full thread, you ensure Claude has the necessary state to maintain tone. Token optimization is secondary to statefulness; you can prune older messages if the context window approaches limits, but full history is required for consistency.

Exam trap

Developers sometimes try to store state on the server or forget to include past turns, assuming the API automatically remembers previous prompts.

218
MCQeasy

Which of the following describes the correct behavior when using the 'stream' parameter in the Messages API?

A.The API sends the entire response at once, but with lower latency.
B.The API returns content as it is generated in discrete events.
C.The API skips the 'messages' array and only accepts a single prompt string.
D.The 'stream' parameter is only available for the Claude 2 model series.
AnswerB

Streaming mode breaks the output into a stream of events. This is essential for building responsive user interfaces where you want to show the model's output in real-time. By handling these events on the client side, you create a much smoother, more interactive experience for the end user.

Why this answer

Setting 'stream' to true causes the API to return a series of Server-Sent Events (SSE) rather than a single JSON object. This is a critical pattern for interactive applications, as it provides an immediate perceived response time for the user. The client must be prepared to handle these discrete events, which include content chunks, metadata, and final completion signals.

Exam trap

Test-takers frequently assume streaming returns a single modified JSON object or newline-delimited standard JSON, forgetting that Anthropic uses Server-Sent Events (SSE) with discrete event chunks.

219
MCQhard

A developer is tuning a classification workload that sends many short prompts to the Claude Messages API. They want to reduce cost and latency without changing the model or the prompt text. They set up prompt caching for the shared system prompt. Which configuration correctly enables caching for that system prompt block?

A.Set a top-level enable_cache boolean to true in the request body alongside the model and max_tokens fields.
B.Send the system prompt as a separate request first, then reference its id in subsequent message requests.
C.Add a cache_control field to the model parameter so the selected model is cached between calls.
D.Add a cache_control object with type "ephemeral" to the system prompt content block in the request.
AnswerD

Prompt caching is enabled per content block by attaching a cache_control field with type "ephemeral". Placing it on the system prompt block marks that prefix for caching, so repeated requests reuse the cached tokens and reduce both cost and latency. This is the documented mechanism and requires no change to the model or prompt wording.

Why this answer

Prompt caching in the Messages API is opt-in per content block using a cache_control field with type "ephemeral". Marking the shared system prompt lets the API reuse that prefix across requests, cutting cost and latency for repeated short prompts. No top-level flag or separate resource is involved, and the model and prompt text stay unchanged.

Exam trap

The trap here is assuming caching is a global request flag or a separate stored resource, when it is actually declared inline on individual content blocks.

220
MCQhard

An application uses Claude to summarize user-generated medical feedback. To comply with privacy requirements, you must ensure no Personally Identifiable Information (PII) is sent to the model. Which is the most appropriate workflow?

A.Prompting the model to ignore and redact any PII found in the input.
B.Processing the text through a dedicated PII anonymization layer before calling the API.
C.Using a specific system prompt that warns the model to treat all input as private.
D.Requesting that users manually redact their own PII before submitting feedback.
AnswerB

Anonymizing data before transmission ensures that no PII is ever sent to the model. This deterministic approach provides a verifiable privacy boundary, making it the most secure method for handling sensitive medical feedback while still allowing the model to perform the requested summarization task effectively and safely.

Why this answer

Data sanitization should occur prior to hitting the API. By using a PII-scrubbing service (like Presidio or a custom regex layer) before the payload reaches Anthropic, you enforce data residency and privacy controls at the edge. This approach is essential for responsible AI because it ensures PII never leaves your secure environment, preventing potential leaks and maintaining strict compliance with global privacy standards like GDPR or HIPAA.

Exam trap

Many students incorrectly assume instructing Claude to ignore PII within the prompt is sufficient, forgetting that data privacy must be enforced before reaching the API.

221
MCQmedium

A developer is sending a large knowledge-base article to the Claude Messages API for analysis. The API returns HTTP 400 with an error message indicating the input is too long. Which statement best describes the correct remediation?

A.Switch the request method from POST to GET to bypass the length restriction.
B.Retry the request with exponential backoff because the error is transient.
C.Increase the max_tokens parameter so the model has room for both the input and the output.
D.Reduce or chunk the input content so the prompt plus the requested output fits within the model's context window.
AnswerD

The 400 error signals that the combined prompt and expected completion exceed the model's context window. Splitting the article into smaller chunks, summarizing sections, or trimming irrelevant text brings the request within limits. This directly addresses the reported cause and allows the request to succeed without changing unrelated parameters.

Why this answer

An HTTP 400 with an explicit input-length message means the prompt plus expected completion exceeds the model's context window. The only correct fix is to shrink the input through chunking, summarization, or removal of irrelevant content. Parameters like max_tokens control output length, and retry logic only helps transient failures, so neither resolves a deterministic size violation.

Exam trap

The trap here is treating a deterministic 400 invalid-request error as if it were a transient failure that retries or parameter tweaks could overcome.

222
MCQmedium

An analyst is using Claude 3.5 Sonnet to compare two 50-page contracts. They notice that Claude correctly identifies a discrepancy in a small footnote on page 74 of the combined input. This demonstrates which fundamental performance characteristic of Claude?

A.Zero-shot learning
B.Strong long-context recall
C.Sentiment analysis
D.Iterative refinement
AnswerB

Claude 3 and 3.5 models are engineered for near-perfect recall across their entire 200,000-token context window. This means the model can accurately 'remember' and retrieve small details, like a footnote, even when they are buried deep within hundreds of pages of other text.

Why this answer

The ability to retrieve specific information from a massive context window is known as 'recall' or 'Needle In A Haystack' performance. Claude models are specifically optimized to ensure that they don't just 'read' the whole document, but can actually find and use information regardless of its position in the input.

Exam trap

Candidates frequently confuse the ability to process large documents with 'reasoning' or 'summarization' rather than the specific capability of 'long-context recall' (Needle In A Haystack) required to find isolated facts.

223
MCQeasy

Why should developers use the Anthropic Messages API instead of legacy Completions API?

A.The Messages API is faster because it bypasses safety checks.
B.It provides a superior structure for managing multi-turn conversations.
C.It allows developers to modify the model's internal weights.
D.It is the only API that supports streaming responses.
AnswerB

The Messages API is built specifically for multi-turn dialogue, offering a clean, structured way to pass conversational history. This format simplifies the management of state, improves the model's ability to track context, and is the foundation for all modern Anthropic features, making it the correct choice for any application.

Why this answer

The Messages API is the current standard for interaction with Anthropic models. It natively supports structured conversation history, system-level prompts, and improved safety features, which are necessary for modern, production-grade applications. Transitioning to this API is essential for accessing the latest model features, receiving better support, and ensuring long-term compatibility with future model updates and enterprise deployment requirements, making it a critical choice for any professional AI project.

Exam trap

Candidates often incorrectly suggest the Messages API is used primarily for 'lower cost' or 'faster speeds', failing to recognize that its primary architectural advantage is structured, multi-turn conversation management.

224
MCQmedium

A junior developer at a marketing agency uses the Claude API to draft promotional blog posts. For one client, they paste in a confidential product roadmap the client shared under NDA and ask Claude to generate teaser posts. The developer's manager later asks whether this use complied with Anthropic's policies. Which statement best describes the compliance situation?

A.This is compliant as long as the developer deletes the conversation afterward and does not enable any training-data sharing setting.
B.This likely violates the client's confidentiality agreement, so the developer should have obtained explicit client consent or used non-confidential material instead.
C.This is compliant because Anthropic's Acceptable Use Policy only restricts illegal content and does not address confidential business information.
D.This is compliant because the roadmap was provided by the client, who owns the content and implicitly authorized its use.
AnswerB

Confidential material shared under NDA generally cannot be disclosed to a third-party processor without the client's informed consent. Submitting the roadmap to the Claude API transmits it outside the agreed trust boundary, which can breach the NDA regardless of Anthropic's own handling practices. The correct path is explicit client authorization, a covered data-processing agreement, or substituting non-confidential inputs.

Why this answer

Sending NDA-protected client material to a third-party AI service is a disclosure that requires the client's consent or a contractual framework permitting it. The agency, as the Anthropic customer, is accountable for what it submits, and neither client authorship nor later deletion neutralizes the confidentiality breach. The safe approach is explicit authorization, an appropriate data agreement, or substituting non-confidential content.

Exam trap

The trap here is assuming that because the client owns the document, the agency is free to route it through any tool it likes.

225
MCQhard

A data scientist is using Claude to classify customer feedback into categories: 'bug', 'feature request', 'complaint', or 'praise'. The feedback is often short and informal. The data scientist wants to maximize classification accuracy. Which prompting strategy is most effective?

A.Use a high temperature to encourage diverse interpretations of the feedback.
B.Include a few examples of feedback for each category in the prompt.
C.Provide a detailed definition of each category in the system prompt.
D.Ask Claude to think step-by-step about the sentiment before classifying.
AnswerB

Providing a few labeled examples for each category gives Claude a clear pattern to follow. It can learn the boundaries between categories from the examples, improving accuracy on informal and varied feedback. This few-shot approach is highly effective for classification tasks, especially when the input language is diverse.

Why this answer

For classification tasks with informal and varied input, providing a few labeled examples per category is the most effective prompting strategy. Examples allow Claude to learn the specific patterns and boundaries between categories, improving accuracy. This few-shot approach outperforms detailed definitions alone because it demonstrates how to handle nuances like slang or brevity.

It also avoids the potential confusion of step-by-step reasoning for a straightforward task.

Exam trap

The trap here is thinking that step-by-step reasoning always improves performance, but for simple classification, few-shot examples are more direct and effective.

Page 2

Page 3 of 4

Page 4

All pages