Courseiva

CCNA Using The Claude Api Questions

65 questions · Using The Claude Api topic · All types, answers revealed

1
Multi-Selecthard

A team is instrumenting its Claude Messages API integration to monitor usage and cost. Which TWO response fields should the developer log to track token consumption per request? (Choose two.)

Select 2 answers
A.stop_reason
B.model
C.usage.output_tokens
D.id
E.usage.input_tokens
AnswersC, E

The usage object also reports output_tokens, the tokens generated in the completion. Because output tokens are typically priced differently from input tokens, logging this separately is necessary for precise cost accounting. Combined with input tokens, it gives a complete per-request consumption picture from the API itself.

Why this answer

The usage object in a Messages API response exposes input_tokens and output_tokens, which together quantify consumption for each call. Because input and output tokens are priced separately, both must be logged for accurate cost attribution. Other fields like stop_reason, id, and model provide context or tracing value but do not measure token usage.

Exam trap

The trap here is logging descriptive response fields such as the model name or message id and assuming they convey token usage when only the usage object reports counts.

2
MCQmedium

A developer maintains a long-running support session and wants to keep the conversation coherent without exceeding the model's context window. They currently resend the entire transcript on every turn. Which approach best addresses the context limit while preserving conversational continuity?

A.Summarize or truncate older turns and send only the relevant recent messages plus a summary.
B.Increase 'max_tokens' so the model can hold more of the transcript in memory.
C.Set a higher 'temperature' so the model compresses prior context automatically.
D.Switch the conversation to a single user message that concatenates all prior turns into one string.
AnswerA

The context window bounds total input plus output tokens. Condensing older turns into a summary or dropping stale messages keeps the payload within that bound while retaining the gist needed for continuity, which directly solves the growing-transcript problem for this support session.

Why this answer

The context window limits the combined tokens of input and output. Because the developer resends the full transcript each turn, the payload grows until it overflows. Condensing older turns into a summary or trimming stale messages, while keeping recent exchanges intact, preserves continuity and keeps each request within the token budget.

Exam trap

The trap here is confusing max_tokens, which caps generated output, with the context window, which bounds the entire input plus output.

3
MCQeasy

A company needs to summarize thousands of short customer feedback snippets every hour. Speed and low cost are their primary requirements, while the complexity of each task is very low. Which Claude 3 model should they choose for this specific API integration?

A.Claude 3 Opus
B.Claude 3 Sonnet
C.Claude 3 Haiku
D.Claude 2.1
AnswerC

Haiku is the fastest and most cost-effective model in the Claude 3 family. It is designed for near-instantaneous responses and can handle lightweight tasks like classification or simple data extraction. Using Haiku for high-frequency small tasks ensures the application remains responsive while significantly reducing overall API expenditure compared to larger models.

Why this answer

Selecting the appropriate model involves balancing performance requirements against cost and latency constraints. Claude 3 Haiku is specifically optimized for speed and affordability, making it ideal for high-volume, simple tasks. Understanding these trade-offs is critical for architecting scalable AI solutions that remain cost-effective while meeting strict response time service level agreements.

Exam trap

Candidates often choose Opus for all tasks assuming 'more capable' is always better, ignoring explicit prompt requirements for speed, high volume, and low cost where Haiku is designed.

4
MCQmedium

When migrating from legacy Claude models to the Claude 3.5 Sonnet Messages API, what is the most important structural change you must implement?

A.Switching the API endpoint from '/v1/messages' to '/v1/completions'.
B.Adopting the structured 'messages' array format with alternating roles.
C.Hardcoding the temperature to 0.0 for all migrated requests.
D.Removing the need for a 'model' identifier in the request body.
AnswerB

The Messages API requires a structured list of message objects, each with 'role' and 'content' keys. This structure is fundamentally different from the legacy completion API, which took a single string. This shift allows for much better context management and is the foundation for all modern Anthropic API interactions.

Why this answer

The Messages API introduced a formal separation between system instructions and conversational content. Unlike older APIs that might have relied on a flat string format, the Messages API requires a structured JSON body with defined 'user' and 'assistant' roles. This shift is essential to allow for more complex multi-turn interactions, tool use, and improved safety alignment.

Exam trap

Candidates often attempt to pass history as a single concatenated string, failing to realize that the Messages API requires a structured array of alternating role-based objects.

5
MCQhard

When using Claude 3.5 Sonnet to process a high-resolution image via the API, which requirement must be met for the image input to be accepted?

A.The image must be provided as a direct URL to a public Amazon S3 bucket.
B.The image must be base64-encoded and wrapped in a content block of type 'image'.
C.The image must be converted to a grayscale 256x256 thumbnail before transmission.
D.The image data should be sent as a separate multipart/form-data part of the HTTP request.
AnswerB

This is the standard requirement for vision tasks in the Messages API. By encoding the image in base64, the binary data is safely transmitted as part of the JSON request. The 'image' type content block then explicitly tells Claude to process that specific part of the message as visual data.

Why this answer

Claude's vision capabilities require specific input formatting. Images must be provided as base64-encoded strings within a content block of type 'image'. Additionally, the developer must specify the media type (e.g., 'image/jpeg', 'image/png').

Failing to provide the correct base64 encoding or the supported media type will result in a 400 error, as the API cannot interpret the raw binary data.

Exam trap

Candidates often attempt to pass raw image file paths or binary data directly into the API, causing immediate bad request errors.

6
MCQeasy

A developer writes a Python script that calls the Claude Messages API. The script currently reads the API key from a hardcoded string in the source file and the file is committed to a public repository. Which change best addresses the security concern while keeping the script functional?

A.Move the key into an environment variable and read it at runtime with os.environ, keeping the key out of source control.
B.Base64-encode the hardcoded key so it is not readable as plain text in the repository.
C.Obfuscate the key by splitting it across several string concatenations in the code before passing it to the client.
D.Add the source file to a .gitignore entry so future commits exclude it, leaving the existing key in the file.
AnswerA

Storing secrets in environment variables keeps them out of the codebase and repository history, which directly removes the exposure caused by committing the key. The script still works because it reads the value at runtime. This is the standard practice for API credentials and allows different keys per environment without code changes.

Why this answer

Credentials must never live in source control. Reading the API key from an environment variable at runtime removes it from the codebase and repository history, so the exposure is eliminated while the script continues to authenticate normally. Any key that has already been committed should also be rotated.

Exam trap

The trap here is treating encoding or obfuscation as if it protects a secret, when only removing the credential from source and rotating it resolves the exposure.

7
MCQeasy

A developer wants to ensure that Claude always responds in a professional, concise tone and never uses emojis. What is the most effective way to implement this across all API calls in their application?

A.Append the instructions to the end of every user message
B.Use the 'system' parameter to define these behavioral constraints
C.Set the 'temperature' to 0 to prevent creative emoji use
D.Configure the 'stop_sequences' to include all common emoji characters
AnswerB

The system parameter is specifically designed for high-level instructions that govern the model's entire response. It is the ideal place for tone, style, and constraint definitions. This ensures that Claude maintains the professional, emoji-free persona consistently across different user interactions, providing a reliable experience for the application's end-users.

Why this answer

Consistent behavior and tone are best handled through the system prompt in the Messages API. By defining these constraints once in the system parameter, the developer ensures the model follows these rules regardless of the specific user input. This centralizes the 'persona' of the application and reduces the need to repeat instructions in every individual user message.

Exam trap

Candidates often repeat instructions in every user message, which is inefficient and inconsistent, failing to use the 'system' parameter designed specifically for global behavioral overrides.

8
MCQeasy

What is the primary function of the 'temperature' parameter in the Claude API?

A.To set the maximum number of tokens for the response.
B.To control the randomness of token generation.
C.To increase the speed of token generation.
D.To enforce safety and compliance filters.
AnswerB

Temperature acts as a scalar on the logit values before they are converted into a probability distribution for the next token. By adjusting this value, you directly influence how likely the model is to choose less probable tokens, thereby controlling the diversity of the generated text across multiple potential outputs.

Why this answer

Temperature controls the 'randomness' or 'creativity' of the model's output by adjusting the probability distribution of the next token. A temperature of 0 results in the most deterministic, consistent output, which is ideal for analytical tasks. Higher temperatures increase the diversity of the output, which is better for brainstorming or creative writing tasks where variety is desired.

Exam trap

Candidates often confuse temperature with execution speed or maximum token length, assuming higher temperatures make the model run faster or process more data, rather than recognizing its role in controlling randomness and creativity.

9
MCQhard

A developer is building a document-summarization service that sends very large PDFs to Claude. The first request returns a 413 error with a message indicating the request entity is too large. The developer wants to keep using the Messages API directly. Which change best resolves the error while preserving the ability to summarize the full document?

A.Set the stream parameter to true so the API streams the request body and bypasses the size limit.
B.Lower the temperature to 0 so the model processes the input more efficiently and accepts the larger payload.
C.Split the document into smaller chunks and send them in multiple Messages API calls, then combine the partial summaries.
D.Increase the max_tokens value in the request to allow the API to accept more input.
AnswerC

A 413 indicates the request body exceeds the API's size limit. Chunking the document into smaller segments keeps each request within the limit, and aggregating the per-chunk summaries preserves coverage of the full document. This approach works within the Messages API's constraints and is a standard pattern for handling large inputs without switching to a different product or violating limits.

Why this answer

An HTTP 413 signals that the request body exceeds the API's permitted size. The Messages API does not accept arbitrarily large inputs, so the developer must reduce each request's payload. Chunking the document into smaller pieces and summarizing each, then combining results, keeps every request within limits and still covers the whole document.

Output-length parameters, streaming flags, and sampling settings do not influence request size limits.

Exam trap

The trap here is confusing generation controls such as max_tokens or stream with the request body size limit that actually triggers the 413.

10
Multi-Selectmedium

When constructing a request for the Messages API, which TWO roles are valid within the 'messages' array?

Select 2 answers
A.system
B.user
C.assistant
D.admin
E.context
AnswersB, C

The user role represents the human or the external system interacting with Claude. This role is required for the first message in the array and typically alternates with the assistant role. It provides the input that Claude processes to generate a response, forming the core of the conversational interaction.

Why this answer

The Messages API is strictly structured to support conversational turns between two primary entities. The roles defined in the API schema ensure that the model understands who is speaking at any given time. Correct usage of these roles is essential for maintaining the conversation flow and allowing the model to generate appropriate, context-aware responses based on previous interactions.

Exam trap

Candidates often mistakenly include a 'system' role inside the 'messages' array, confusing the top-level system parameter with conversational turn roles.

11
Multi-Selectmedium

A developer needs to configure a Claude 3 model to generate a wide variety of creative and unpredictable marketing slogans for a new product line. Which TWO parameter adjustments would best support this goal?

Select 2 answers
A.Set temperature to 1.0
B.Set temperature to 0.0
C.Set top_p to 1.0
D.Set top_p to 0.1
E.Set max_tokens to a very low value
AnswersA, C

Higher temperature values increase the randomness of the output by allowing the model to choose tokens that are not the most statistically likely. For creative tasks like slogan generation, a temperature near 1.0 is effective. This helps the model avoid repetitive or generic phrasing and explore more unique linguistic combinations.

Why this answer

To increase creativity and variety in model outputs, developers must adjust the sampling parameters. Increasing temperature makes the probability distribution of the next token flatter, allowing for less likely choices. Similarly, a high top_p value ensures a larger pool of potential tokens is considered.

These settings are ideal for creative tasks where consistency and predictability are less important than novelty.

Exam trap

Candidates often lower the temperature to reduce errors, forgetting that creative tasks require maximized sampling parameters like temperature and top_p set to 1.0.

12
MCQhard

A developer is implementing a retrieval-augmented generation pipeline with the Claude Messages API. Retrieved documents are inserted into the user turn, and the developer wants the model to cite which document supports each claim. Which technique most directly improves the model's ability to attribute statements to specific source documents?

A.Increase max_tokens so the model has room to quote each source document verbatim in the answer.
B.Number or tag each retrieved document in the prompt and instruct the model to reference those identifiers in its answer.
C.Concatenate all retrieved documents into one continuous block with no separators to keep the prompt compact.
D.Raise the temperature so the model explores more of the retrieved context before answering.
AnswerB

Giving each document a stable identifier and asking the model to cite it creates an explicit mapping between claims and sources. The model can then emit something like a document number or tag alongside each statement, which the application can verify against the retrieved set. This structured labeling is the most direct way to improve attribution accuracy in a RAG prompt.

Why this answer

Attribution improves when the prompt makes source boundaries explicit. Labeling each retrieved document with a stable identifier and instructing the model to cite that identifier gives it a concrete token to attach to each claim, and gives the application something to validate. Sampling settings and output length do not create source-to-claim mappings, and removing document separators actively destroys the structure needed for citation.

Exam trap

The trap here is treating attribution as a generation-quality problem solvable with temperature or token limits, when it is really a prompt-structure problem that requires explicit document identifiers.

13
MCQhard

When designing a high-throughput application using Claude 3, a developer is concerned about the costs of repeatedly sending a 20,000-token context in every request. Which API feature should they implement to optimize both cost and performance?

A.Context Compression
B.Prompt Caching
C.Batch Processing
D.Token Distillation
AnswerB

Prompt Caching allows the API to store frequently used context on the server side. When a new request starts with the same cached prefix, it is processed much faster and at a lower cost. This is ideal for scenarios where a large knowledge base is queried multiple times, directly addressing the developer's concerns about throughput and expense.

Why this answer

Prompt Caching is a powerful feature for applications that reuse large amounts of context, such as documentation or legal files. By caching the prefix of a prompt, the developer only pays a reduced rate for the cached tokens in subsequent requests. This also reduces processing time, as the model does not need to re-encode the same information repeatedly, significantly improving efficiency.

Exam trap

Candidates often try to manually truncate or summarize context to save costs, ignoring the built-in 'Prompt Caching' feature designed to handle large, static context efficiently.

14
MCQmedium

A developer is building a document Q&A service on the Claude API. Each request sends the full 180-page PDF as base64 in a single text content block, and most calls now fail with a 400 error stating the request exceeds the maximum allowed size. The developer wants to keep using the same model and continue asking questions about the whole document. Which approach should the developer take?

A.Split the document into chunks, send each chunk in its own request, and merge the returned answers inside the application.
B.Move the PDF into the system prompt so that the document bytes are not counted against the user message size.
C.Upload the PDF with the Files API, then reference the returned file identifier in the message content instead of embedding base64.
D.Compress the base64 string with gzip and place the compressed value in the text content block.
AnswerC

The Files API accepts a document once, returns an identifier, and lets later Messages requests reference that identifier rather than re-transmitting the encoded bytes. This removes the oversized payload from each call while keeping the entire document available as context, so whole-document questions still work. It is the intended path for large documents that exceed inline request size limits.

Why this answer

Large documents should be ingested once through the Files API and then referenced by identifier in subsequent Messages calls, which keeps each request's inline payload small while preserving full-document context. Re-sending base64, relocating it to the system prompt, or compressing it all leave the oversized or unreadable payload inside the request, so the size rejection or loss of usable content persists.

Exam trap

The trap here is assuming that moving large content into the system prompt exempts it from the request size limits, when the system prompt is part of the same payload.

15
MCQmedium

A developer wants Claude to always reply in strict JSON matching a provided schema for an internal data-extraction pipeline. They consider the tool use feature as a way to constrain output. Which statement best describes how to use tool use to reliably obtain schema-conformant output?

A.Include the JSON schema in the system prompt and rely on the model to follow it without any tool definitions.
B.Set the response_format parameter to "json_schema" and supply the schema so the API validates the output before returning it.
C.Define a tool whose input_schema describes the desired fields, then read the structured input from the tool_use block the model returns.
D.Enable strict mode by setting the temperature parameter to 0, which forces valid JSON output.
AnswerC

Defining a tool with an input_schema that mirrors the target structure, and optionally forcing selection with tool_choice, makes the model emit a tool_use block whose input conforms to that schema. Parsing that input yields structured data directly, which is a common and reliable pattern for schema-constrained extraction with the Messages API.

Why this answer

Tool use provides a structural contract: define a tool whose input_schema matches the target fields, and the model returns a tool_use block containing structured input. Reading that input gives schema-shaped data without parsing free text. This is the intended pattern for schema-constrained extraction in the Messages API.

Exam trap

The trap here is expecting a dedicated JSON response format parameter, when structured output is obtained through tool definitions and the resulting tool_use block.

16
MCQeasy

When calling the Claude API, what is the primary benefit of setting a 'max_tokens' value that is close to the expected output length?

A.It improves the reasoning capabilities of the model for complex tasks.
B.It forces the model to be more creative and less repetitive.
C.It prevents the model from generating unnecessary text and saves costs.
D.It ensures the model always uses the full context window provided.
AnswerC

The 'max_tokens' parameter serves as a safety buffer and cost control mechanism. Since API billing depends on the number of output tokens, limiting unnecessary verbosity ensures that you only pay for the content you actually need, while also improving response latency by reducing the time spent on generation.

Why this answer

Setting 'max_tokens' accurately helps manage latency and control costs by preventing the model from generating excessively long or rambling responses. While the model may stop before this limit if it reaches a natural conclusion, defining a reasonable ceiling prevents runaway generation in edge cases. This is essential for maintaining a predictable user experience and ensuring that API usage remains within your predefined budget and throughput expectations for production applications.

Exam trap

Candidates often believe that setting max_tokens strictly controls the exact output length or improves model intelligence, confusing token limits with generation quality parameters.

17
MCQeasy

A developer needs to stop Claude from generating text as soon as it produces a double newline sequence ('\n\n'). Which API parameter should be used to implement this?

A.max_tokens
B.stop_sequences
C.system_prompt
D.top_p
AnswerB

The stop_sequences parameter is an array of strings that act as 'cut-off' triggers. When Claude generates any of the specified sequences, it stops exactly at that point. This is the correct tool for ensuring the model stops after a double newline, providing precise control over the completion boundaries.

Why this answer

The 'stop_sequences' parameter allows developers to define a list of strings that, if generated by the model, will cause it to immediately cease production of further tokens. This is particularly useful for controlling output format, preventing the model from rambling, or ensuring that the assistant does not simulate a user response in a few-shot prompting scenario.

Exam trap

Candidates often confuse the 'stop_sequences' parameter with model instructions or system prompts, incorrectly assuming the model will stop on its own if told to do so in the prompt text.

18
MCQeasy

When using the Claude API, which parameter should be adjusted if you want to influence the model's creativity and randomness?

A.max_tokens
B.top_p
C.temperature
D.system
AnswerC

Temperature is the standard parameter used to control the randomness of the model's output. By increasing the temperature, the model becomes more exploratory, while decreasing it makes the output more predictable and focused. It is the direct control for balancing creativity versus consistency in generated responses.

Why this answer

The 'temperature' parameter is the primary lever for controlling the stochastic nature of the model's output. By adjusting this value, you directly influence the probability distribution of the next token selection. A lower temperature makes the model more deterministic and focused, whereas a higher temperature introduces more variation, which is essential for creative writing or brainstorming tasks in AI applications.

Exam trap

Candidates confuse 'temperature' with token limits or max_tokens, incorrectly thinking length parameters control output creativity and randomness.

19
Multi-Selectmedium

Which TWO of the following are true about Anthropic's 'usage' metadata returned in the API response?

Select 2 answers
A.It includes 'input_tokens' and 'output_tokens' fields.
B.It can be used to override the model's internal temperature.
C.It is only available when streaming is disabled.
D.It helps monitor the cost of the request accurately.
E.It provides the latency in milliseconds for each model layer.
AnswersA, D

The usage object explicitly breaks down the token count into input and output sections. This separation is necessary for billing, as input and output tokens are often priced differently. Developers rely on these fields to keep track of their spending per request and to optimize the length of their prompts.

Why this answer

The usage metadata is vital for tracking your API costs and understanding prompt token consumption, especially when dealing with long history or caches. It provides clear counts for both input and output tokens. This data allows for accurate budget forecasting and helps developers optimize prompts by identifying which parts of the input are consuming the most tokens during the model's processing phase.

Exam trap

Test-takers frequently look for pricing metrics directly in prompt text or response content, forgetting that actual token counts and cost tracking are found in the usage metadata.

20
Multi-Selecthard

A team is building a support triage assistant on the Claude API. They want the model to classify each ticket into a fixed set of categories and also return a short justification, while guaranteeing that the category value is always one of five allowed strings. They are choosing between tool use with a JSON schema and free-form text output parsed with a regex. Which TWO statements correctly describe the advantages of the tool use approach in this scenario? (Choose two.)

Select 2 answers
A.Tool use guarantees the model will never produce an invalid category, so the application can skip validating the returned value before storing it.
B.The tool input schema constrains the model's output to the declared properties and types, so the category field can be defined as an enum of the five allowed values.
C.Enabling tool use automatically reduces the token cost of each request because structured outputs are billed at a lower rate than free-form text.
D.The tool call response arrives as a structured content block with the parsed input object, removing the need to write brittle regex parsing for the category and justification.
E.Tool use forces the model to return only the tool call and suppresses any accompanying natural-language explanation, which is why the justification must be placed inside the tool input.
AnswersB, D

When a tool is defined with an input schema, the model's tool call is generated to conform to that schema, so a property declared as an enum restricts the category to the five permitted strings. The justification can be a sibling string property. This gives the application a structured, validated payload instead of prose that must be pattern-matched, which directly satisfies the guaranteed-value requirement.

Why this answer

Defining a tool whose input schema declares the category as an enum and a justification as a string gives the application a structured, schema-shaped payload with the allowed values baked into the contract, and it removes brittle regex scraping. The remaining claims are false: schema guidance still warrants validation, structured output is not billed differently, and natural-language text is not forcibly suppressed around a tool call.

Exam trap

The trap here is treating schema-guided tool output as a hard guarantee that eliminates the need for any application-side validation of the returned values.

21
MCQhard

An engineer is designing a multi-tenant SaaS feature that lets each customer supply their own Anthropic API key, which the backend stores encrypted and uses when calling the Messages API on that tenant's behalf. A security review asks how requests should be attributed so usage can be billed back and abuse isolated per tenant. Which practice best meets this requirement?

A.Have the backend call the Messages API with each tenant's own stored key so rate limits, usage, and audit trails are scoped to that tenant.
B.Proxy all tenant traffic through a single key and reconcile billing by counting tokens in the responses your backend receives.
C.Issue each tenant the same platform key but append a unique tenant ID to the user field of every message.
D.Send every tenant's request using a single shared platform API key and rely on the prompt contents to identify the tenant.
AnswerA

Using the tenant's own key means Anthropic's rate limiting, usage reporting, and logs are naturally partitioned per credential, which gives clean billing attribution and contains abuse to a single tenant. The backend still controls the call path, so it can enforce its own quotas and redact sensitive fields before forwarding. This aligns the credential boundary with the tenant boundary, which is exactly what the security review needs.

Why this answer

Credential boundaries are the cleanest way to achieve per-tenant attribution and isolation. When each tenant's traffic flows through that tenant's own API key, Anthropic's rate limiting, usage records, and audit logs partition automatically, and a compromised or abusive tenant cannot exhaust capacity belonging to others. The backend retains control over request construction and can layer additional quotas.

Exam trap

The trap here is treating a message-level identifier such as the user field as a substitute for credential-level separation of rate limits and billing.

22
Multi-Selecthard

A developer is implementing retry logic for calls to the Anthropic Messages API in a production service. The service must handle transient failures gracefully without overwhelming the API. Which TWO practices should the developer implement? (Choose two.)

Select 2 answers
A.Retry requests that fail with HTTP 429 and HTTP 500 status codes, using exponential backoff with jitter between attempts.
B.Disable all retries and rely on the API's internal queueing to eventually deliver failed requests.
C.Retry all failed requests immediately in a tight loop until they succeed, to minimize latency for the end user.
D.Set a maximum retry count or total elapsed time budget so the service stops retrying after a defined threshold.
E.Retry requests that fail with HTTP 400 and HTTP 401 status codes, since these often resolve on a second attempt.
AnswersA, D

Status 429 indicates rate limiting and 500 indicates a server-side error, both of which are typically transient. Retrying with exponential backoff and jitter spaces out attempts and avoids synchronized retry storms. Jitter prevents many clients from retrying simultaneously. This combination is a standard resilient pattern that improves success rates without hammering the API during periods of contention.

Why this answer

Resilient retry logic targets transient failures such as 429 and 500 responses, spacing attempts with exponential backoff plus jitter to avoid synchronized retry storms. A retry cap, whether by count or total time budget, prevents runaway loops and bounds latency. Client errors like 400 and 401 are not transient and should not be retried, and the API does not redeliver failed requests, so some client-side retry strategy is necessary.

Exam trap

The trap here is treating all HTTP errors as retryable, when client errors such as 400 and 401 will never succeed on a repeated identical request.

23
MCQmedium

When integrating Claude into an application that processes PII (Personally Identifiable Information), what is the most recommended approach to maintaining data privacy?

A.Send the raw data and rely on the model's system prompt to ignore PII.
B.Mask PII locally before making the API request.
C.Request a private VPC deployment for all Claude API interactions.
D.Enable the 'hide-pii' flag in the request body.
AnswerB

Local data masking is the most reliable way to maintain privacy. By transforming sensitive identifiers into generic tokens or hashes on your server, you ensure that the raw PII never leaves your control, effectively mitigating risks even if the data was somehow exposed. This is the standard procedure for secure LLM workflows.

Why this answer

Anonymizing or masking PII before sending data to the API is a critical security best practice. By stripping or obfuscating sensitive data locally, you minimize the risks associated with data processing in third-party environments. This approach aligns with industry standards for data protection and ensures your application complies with common privacy regulations like GDPR or HIPAA by design.

Exam trap

Candidates often assume that using the API in a private VPC or secure connection is sufficient, forgetting that data must be sanitized before it enters the model's processing context.

24
Multi-Selecthard

A developer is designing a tool use workflow with the Claude Messages API. After the model returns a tool_use content block, the application executes the tool and must send the result back. Which TWO of the following are required for the follow-up request to be processed correctly? (Choose two.)

Select 2 answers
A.Preserve the prior assistant turn, including the tool_use block, in the messages array of the follow-up request.
B.Include a user message containing a tool_result content block whose tool_use_id matches the id of the model's tool_use block.
C.Re-declare the full tools array in the follow-up request even though tools were already provided in the first request.
D.Set the stop_sequences parameter to include the tool name so the model knows when to stop calling tools.
E.Convert the tool_result content into plain assistant text so the model can read it as part of its own previous reply.
AnswersA, B

The conversation history must include the assistant message that contained the tool_use block, because the tool_result must reference a tool call that exists in context. Sending only the tool_result without the originating assistant turn breaks the required pairing and the API rejects the request. The full exchange forms the basis for the model's next reasoning step.

Why this answer

Completing a tool use loop requires sending the tool's output back as a tool_result content block inside a user message, with a tool_use_id that matches the model's tool_use block id, while keeping the original assistant turn in the messages array. This pairing lets the model associate the result with its earlier call and continue reasoning.

Exam trap

The trap here is assuming the tool result can be sent as free-form text or without the originating assistant turn, when the API requires the structured tool_use/tool_result pairing with matching ids.

25
MCQmedium

Refer to the exhibit. Based on the headers returned in the API response, what is the most immediate constraint the developer should be concerned about for their next few requests?

A.The application is about to run out of token capacity.
B.The API key has expired and needs to be refreshed by the reset time.
C.The request count limit is nearly reached.
D.The server is overloaded and will reset at the specified time.
AnswerC

The 'requests-remaining' header shows only 5 units left. This is a very low number compared to the total limit of 1,000. If the developer continues sending requests at the same rate, they will soon receive 429 Too Many Requests errors. They should monitor this header to implement proactive throttling within their application logic.

Why this answer

The Anthropic API returns rate limit information in the response headers. In this exhibit, the 'anthropic-ratelimit-requests-remaining' value is 5, while the 'tokens-remaining' is 350,000. This indicates that the client is very close to exhausting their allowed number of requests, even though they have plenty of token capacity left.

The developer needs to slow down their request frequency.

Exam trap

Candidates frequently confuse token limits with request rate limits, focusing on the high remaining token count while ignoring the dangerously low number of remaining requests indicated in the headers.

26
Multi-Selecthard

A developer is integrating the Claude API into a production pipeline and needs to handle a response where the model stopped because it hit the output token ceiling before finishing its answer. Which TWO actions are appropriate? (Choose two.)

Select 2 answers
A.Assume the response is complete and parse it as-is, since a 200 status was returned.
B.Inspect the 'stop_reason' field and treat a value indicating the token limit as a truncation signal.
C.Retry the identical request unchanged, expecting the model to finish within the same limit.
D.Lower 'temperature' to 0 so the model produces a shorter answer next time.
E.Raise 'max_tokens' or continue the response by sending the partial output back as an assistant turn.
AnswersB, E

The response's stop_reason tells the application why generation ended. A value indicating the output token limit means the answer was cut off, so checking this field lets the pipeline detect truncation programmatically and decide how to respond, rather than silently consuming an incomplete result.

Why this answer

When generation stops because it reached the output token ceiling, the response still returns successfully but the content is incomplete. The application should detect this via stop_reason and then either allow more output tokens or continue from the partial text. Retrying unchanged, adjusting temperature, or trusting the HTTP status all fail to detect or resolve the truncation.

Exam trap

The trap here is treating a 200 response as proof of completeness, when stop_reason is what actually signals whether output was truncated.

27
MCQeasy

A developer is writing the first integration test against the Claude Messages API using the official Anthropic SDK. The test must send a single user turn and read the model's text reply. Which request structure correctly represents the required Messages API input?

A.A messages array with alternating user and assistant turns, and the model name omitted because it is inferred from the endpoint.
B.A messages array with a single object whose role is "system" and content set to the prompt, plus max_tokens.
C.A prompt parameter containing the text, plus the model name, with no messages array.
D.A messages array containing one object with role "user" and content set to the prompt string, plus the model and max_tokens parameters.
AnswerD

The Messages API requires a model, a max_tokens value, and a messages array. Each message has a role of "user" or "assistant" and content that is a string or a list of content blocks. A single user message with a string prompt is the minimal valid request, so this structure satisfies the API contract for a first integration test.

Why this answer

The Messages API expects a required model, a required max_tokens, and a messages array whose entries use role "user" or "assistant" with string or block content. A single user message with a string prompt is the smallest valid request. System instructions are passed through a separate top-level system field, and the model must always be named explicitly, so the other structures would be rejected before generation.

Exam trap

The trap here is carrying over the legacy completions habit of sending a top-level prompt string, or treating the system prompt as a message role, instead of using the messages array with user and assistant roles.

28
Multi-Selecthard

A developer is building a customer-facing assistant with the Claude Messages API and must implement multi-turn conversations that stay within context limits while remaining coherent. Which TWO practices are appropriate for managing the conversation history? (Choose two.)

Select 2 answers
A.Store durable facts and user preferences in a separate summary that is injected into the system prompt on each call.
B.Resend the full messages array on every request, trimming or summarizing the oldest turns when the total approaches the context window.
C.Rely on the API to remember previous turns server-side using a conversation identifier returned in each response.
D.Delete all prior assistant turns and keep only user messages to halve the token usage.
E.Increase max_tokens to the model's maximum on every request so no earlier turn is ever dropped.
AnswersA, B

Extracting stable facts into a summary and injecting them via the system prompt preserves important context even after raw turns are dropped. It keeps the token cost low while maintaining coherence, because the model still sees the durable information. This is an appropriate technique for managing long conversations within context limits.

Why this answer

Because the Messages API keeps no server-side session, the client owns conversation state and must resend it. The practical approaches are to manage the growing messages array by trimming or summarizing old turns, and to preserve durable facts in a compact summary injected through the system prompt. Neither raising max_tokens nor deleting assistant turns addresses the input-size limit, and the API does not remember prior turns on its own.

Exam trap

The trap here is assuming the Messages API retains conversation state between calls, when each request is stateless and only the content the client resends is visible to the model.

29
MCQmedium

A developer is adding a retrieval step so Claude can answer questions over a 200-page internal policy manual. The full manual exceeds the context window, so the application must select relevant sections to include in each Messages API request. Which strategy best keeps answers accurate while staying within the context limit?

A.Embed the manual into chunks, retrieve the top-matching chunks for each question, and include only those excerpts in the request.
B.Truncate the manual to the first N tokens that fit and always send that same prefix with every question.
C.Increase the max_tokens parameter on each request so the model can internally hold more of the manual.
D.Summarize the entire manual once with Claude, cache the summary, and send the summary with every question instead of the source text.
AnswerA

Chunking plus similarity retrieval places the passages most likely to contain the answer into the context window while keeping total tokens bounded, which is the standard retrieval-augmented pattern for documents larger than the context limit. It scales to manuals of any size because only the retrieved excerpts are sent. Accuracy depends on chunk sizing and retrieval quality, but this directly addresses both the size constraint and the grounding requirement.

Why this answer

When source material exceeds the context window, retrieval-augmented generation is the appropriate pattern: split the document into retrievable chunks, select the passages most similar to the incoming question, and place only those excerpts in the request. This bounds token usage while keeping the answer grounded in the passages most likely to contain the relevant policy.

Exam trap

The trap here is confusing max_tokens with input capacity, when max_tokens only limits generated output and never enlarges the model's context window.

30
MCQmedium

An application is processing very long documents through the Claude API. The developer notices that some responses are being cut off before they are naturally finished. Which property in the API response should they inspect to determine if the truncation was caused by reaching a length limit?

A.finish_status
B.truncation_flag
C.stop_reason
D.usage.input_tokens
AnswerC

This field indicates why the model stopped generating. A value of 'max_tokens' confirms that the response was truncated due to the limit set in the request. If the value is 'end_turn', the model finished its thought naturally. Checking this value allows the application to respond appropriately, such as by prompting the model to continue.

Why this answer

The API response includes metadata that describes why the model stopped generating text. The stop_reason field is the primary indicator of this behavior. If this field contains 'max_tokens', it signifies that the model had more to say but was interrupted because it reached the limit specified in the request.

Understanding this allows developers to programmatically decide whether to request more tokens.

Exam trap

Candidates often look for a 'status' or 'error' field in the body, failing to realize that the 'stop_reason' field is the specific metadata indicator for model generation limits.

31
MCQmedium

A developer is optimizing a high-volume classification workload on the Claude Messages API. Every request shares a long, static set of instructions and few-shot examples, followed by a short variable user input. The developer wants to cut cost and latency without changing output quality. Which feature should the developer apply?

A.Move the static instructions into the system parameter and the few-shot examples into the final user turn.
B.Enable prompt caching on the shared instruction and example prefix, placing the variable input after the cached portion.
C.Switch to a smaller model and remove the few-shot examples to compensate for the reduced capability.
D.Lower max_tokens to the smallest value that still fits a classification label.
AnswerB

Prompt caching stores the processed prefix so subsequent requests reuse it instead of reprocessing those tokens. Cache reads are billed at a reduced rate and reduce time to first token. Because the instructions and examples are identical across requests, caching that prefix while keeping the variable input after it directly lowers cost and latency without altering the model's output.

Why this answer

When many requests share an identical leading block of tokens, prompt caching lets the API reuse the processed prefix. Cache reads cost less and reduce latency, and because the cached content is unchanged, output quality is preserved. The variable input must come after the cached prefix so the cacheable portion stays contiguous.

Repositioning content, shrinking output limits, or swapping models does not achieve the same effect without side effects.

Exam trap

The trap here is assuming any prompt reorganization yields caching benefits, when caching only applies to an identical contiguous prefix and is invalidated by even small changes within it.

32
MCQmedium

A developer is building a document summarization service that calls the Claude Messages API. The service must always respond in valid JSON containing exactly two fields: "summary" and "confidence". The developer wants to maximize the chance of receiving a valid JSON object without writing a custom parser. Which approach should the developer take?

A.Send the request with temperature set to 0 and count the number of braces in the response to validate the JSON.
B.Use the Messages API with a tool definition whose input_schema describes the two required fields, and require the tool to be called.
C.Set the system prompt to instruct Claude to reply only with JSON, and include a single example of the desired JSON structure.
D.Append the word "JSON" to the end of the user message and set max_tokens to a value large enough for the expected output.
AnswerB

Defining a tool whose input_schema is a JSON Schema with "summary" and "confidence", then requiring tool use, constrains Claude's output to arguments that conform to that schema. The API returns a structured tool_use block, so the developer receives a validated JSON object without building a custom parser, which directly satisfies the reliability requirement.

Why this answer

Tool use with a JSON Schema input_schema is the Messages API's structured-output mechanism. By declaring the two fields as required properties and forcing the tool call, the developer makes Claude return arguments that conform to the schema, eliminating custom parsing. Instruction-only approaches and temperature tuning reduce but do not remove the risk of malformed or extra output, so they are less reliable for a strict two-field contract.

Exam trap

The trap here is assuming that telling Claude to "respond in JSON" in a prompt is equivalent to schema-enforced structured output, when only a tool definition with an input_schema actually constrains the response shape.

33
MCQmedium

Which HTTP header is required in every request to the Claude API to specify the version of the API being used?

A.API-Version
B.anthropic-version
C.X-Claude-Version
D.Accept-Version
AnswerB

This is the correct, mandatory header. As of the current documentation, the value '2023-06-01' is frequently used. This header allows the developer to pin their application to a specific version of the API, ensuring stability even as Anthropic evolves the platform and adds new features or parameters.

Why this answer

The 'anthropic-version' header is a mandatory requirement for all requests to the Claude API. It ensures that the client is compatible with the specific API versioning schema and allows Anthropic to introduce updates or breaking changes without affecting older implementations. Without this header, the API will return an error because it cannot determine which schema to validate the request against.

Exam trap

Candidates often forget the 'anthropic-version' header entirely or use an incorrect date format, causing the API to reject the request due to missing version context.

34
MCQhard

Refer to the exhibit. An application receives this JSON response from the Claude API. Which action should the developer take to handle this specific error effectively?

A.Modify the prompt to reduce the total number of input tokens.
B.Immediately resend the request until a 200 OK status is received.
C.Implement a retry mechanism with exponential backoff.
D.Check the API key and ensure it has not expired or been revoked.
AnswerC

Exponential backoff involves waiting for progressively longer periods between retries. This is the industry-standard approach for handling transient 5xx errors like 'overloaded_error'. It balances the need for the application to complete its task with the need to be a 'good citizen' by reducing traffic during peak congestion.

Why this answer

The 'overloaded_error' (HTTP 529) indicates that Anthropic's servers are currently experiencing high traffic and cannot process the request. Unlike client-side errors, this is a transient server issue. The recommended approach is to implement a retry strategy with exponential backoff, which prevents the client from further stressing the system while allowing the request to eventually succeed when capacity becomes available.

Exam trap

Developers often treat server overload errors (HTTP 529) like permanent client-side 400 errors, failing to implement proper retry logic and crashing the app.

35
Multi-Selectmedium

You are building a high-throughput application. Which TWO of the following strategies are best for optimizing your API costs and efficiency?

Select 2 answers
A.Use the largest available model for all tasks to ensure accuracy.
B.Cache frequently used static system prompts or common context.
C.Set the temperature to 0 for all production API calls.
D.Select the most cost-efficient model that meets the latency and task quality requirements.
E.Increase the 'max_tokens' to the maximum allowed for every request.
AnswersB, D

Prompt caching allows you to store long, static portions of your prompt, reducing the number of tokens processed in subsequent requests. This drastically decreases latency and lowers cost for applications that frequently reuse large amounts of reference documentation or complex instruction sets across many API calls.

Why this answer

Optimizing API usage involves balancing model choice with intelligent prompt management. Choosing the right model for the task (Sonnet vs. Haiku) and ensuring input tokens are minimized through efficient prompting are the most effective ways to reduce operational overhead.

These practices are fundamental to scaling Claude-based applications while maintaining a sustainable cost structure and ensuring that the API responds with low latency to handle high-volume user traffic.

Exam trap

Candidates often prioritize model fine-tuning or complex prompt engineering over basic architectural efficiencies like model selection and prompt caching, which provide more immediate cost benefits.

36
MCQmedium

A developer is building an application that needs to maintain a specific tone and set of behavioral constraints across multiple turns of a conversation. Where should these instructions be placed in the Messages API call to ensure the most consistent adherence by the model?

A.Inside the first object of the messages array with the role set to 'user'.
B.As a standalone 'system' parameter at the top level of the request body.
C.Inside every assistant message to remind the model of its persona.
D.In a metadata field within the request body to be parsed by the API.
AnswerB

The top-level system parameter provides a dedicated space for instructions that guide the model's behavior throughout the session. This separation of concerns allows Claude to prioritize these instructions differently than message content, ensuring the persona and constraints remain active and influential even as the dialogue history becomes complex.

Why this answer

In the Messages API, the system parameter is specifically designed for high-level instructions, personas, and behavioral constraints. Placing these instructions in the system prompt rather than the first user message helps Claude distinguish between the developer's rules and the user's input, leading to better instruction following and reduced likelihood of the model ignoring constraints during long interactions.

Exam trap

Many candidates mistakenly believe instructions should go inside the user message or conversation history to maintain context, forgetting that the top-level system parameter is explicitly designed for persistent behavioral guidelines.

37
MCQmedium

A software engineer is building a real-time customer support chatbot using the Claude Messages API. To improve the user experience, they want the assistant's response to appear gradually on the screen as it is being generated. Which parameter must be set to true in the API request to enable this functionality?

A.interactive
B.stream
C.incremental
D.real_time
AnswerB

Setting this boolean parameter to true triggers the API to use Server-Sent Events for the response. This allows the client to receive and process text fragments as they are generated by the model. It is the standard method for minimizing time-to-first-token in web applications, providing a more fluid and responsive chat interface.

Why this answer

Streaming is a critical feature for interactive applications because it reduces the perceived latency for the end-user. By enabling the stream parameter, the Anthropic API sends partial message increments via Server-Sent Events. This allows developers to display content immediately as it becomes available, rather than waiting several seconds for the entire completion to be finished and returned in a single block.

Exam trap

Candidates often confuse the 'stream' parameter with 'streaming_mode' or other non-existent fields, forgetting that it is a simple boolean flag in the request body.

38
MCQmedium

A developer needs Claude to always respond with a valid JSON object containing specific fields for a downstream parser. The team wants the strongest guarantee that the model's reply will conform to a defined structure. Which capability should the developer use?

A.Set temperature to 0 so the output becomes deterministic and therefore valid JSON.
B.Add the phrase 'respond only in JSON' to the system prompt and hope the model complies.
C.Use the tool use feature with a defined input schema to force structured arguments.
D.Increase max_tokens so the model has enough room to finish the JSON object.
AnswerC

Tool use lets the developer declare a tool with a JSON schema for its input, and the model returns tool-call arguments that conform to that schema. This provides a strong structural guarantee and integrates cleanly with a downstream parser expecting specific fields, making it the most reliable way to obtain structured output from Claude.

Why this answer

Tool use with a JSON schema gives the strongest structural guarantee because the model emits arguments validated against the declared input schema. Prompt instructions and temperature settings influence behavior but cannot enforce format, and max_tokens only bounds length. When a downstream parser demands exact fields, schema-backed tool calls are the dependable choice.

Exam trap

The trap here is believing that adding a formatting instruction to the prompt, or lowering temperature, is equivalent to enforcing a schema.

39
MCQmedium

A developer is sending a single request to the Anthropic Messages API with a system prompt and a user message. The application needs Claude to return a JSON object that strictly conforms to a predefined schema without any explanatory prose. Which request configuration should the developer use to maximize the likelihood of receiving only valid JSON?

A.Set the temperature parameter to 1.0 and rely on the model's natural tendency to produce structured output.
B.Add the instruction 'Return only JSON, nothing else' to the system prompt and set max_tokens to a small value.
C.Use the tool_choice parameter set to {"type": "tool", "name": "your_tool"} with an input_schema defining the desired JSON structure.
D.Set the stop_sequences parameter to ['}'] so the response terminates immediately after the closing brace.
AnswerC

Forcing a specific tool with tool_choice directs the model to emit a tool_use block whose input conforms to the supplied input_schema, effectively producing structured JSON. This is the documented mechanism for schema-constrained outputs in the Messages API. It reliably yields a parseable JSON object matching the schema rather than free-form prose, which satisfies the strict conformance requirement.

Why this answer

The Messages API provides a structural mechanism for schema-constrained output: defining a tool with an input_schema and forcing its use via tool_choice. This makes the model emit a tool_use block whose input matches the schema, yielding a clean JSON object without surrounding prose. Sampling parameters, prompt instructions, and stop sequences influence generation but do not enforce structure, so they cannot reliably satisfy the strict-JSON requirement.

Exam trap

The trap here is assuming that a strongly worded prompt instruction such as 'return only JSON' is equivalent to a structural constraint enforced by the API.

40
MCQhard

A developer wants Claude to answer questions strictly from a provided internal policy document and to refuse when the answer is not present. The team must ensure the model treats the document as authoritative reference material rather than as instructions to follow. Which approach best achieves this?

A.Raise max_tokens so the model can read the entire document before answering.
B.Place the policy document inside the system prompt so it is treated as a higher-priority instruction.
C.Include the document in a user message wrapped in clear delimiters and instruct the model to answer only from that content and refuse otherwise.
D.Set temperature to a high value so the model explores the document more thoroughly before responding.
AnswerC

Placing the document in a user turn with explicit delimiters and pairing it with a refusal instruction keeps the content in the data channel. The model is told to treat the enclosed text as reference material and to decline when an answer is absent, satisfying both the grounding and refusal requirements.

Why this answer

Grounding requires that reference content stay in the data channel and be explicitly framed as material to consult, not instructions to obey. Placing the document in a delimited user message with a refusal directive achieves both goals. Promoting the document to the system prompt grants it instruction authority, and output-length or temperature settings do not affect how content is interpreted.

Exam trap

The trap here is assuming that moving reference content into the system prompt makes the model more faithful, when it actually risks the document being followed as instructions.

41
MCQmedium

You are processing large documents with Claude. If the document exceeds the context window, which strategy is most effective for maintaining quality results?

A.Increase the max_tokens to accommodate the entire document.
B.Use a RAG approach to retrieve and send only relevant chunks to the model.
C.Split the document into chunks and send them in parallel as separate API calls.
D.Request the model to summarize the document in sections using a loop.
AnswerB

RAG is the best practice for handling documents that are too large for the context window. By retrieving only the most relevant parts of the document, you stay well within token limits and provide the model with high-signal content, which improves accuracy and performance for large-scale analysis.

Why this answer

When dealing with documents exceeding the context window, a RAG (Retrieval-Augmented Generation) approach is the standard solution. By segmenting the document into chunks and retrieving only the most relevant sections for a specific query, you ensure the model focuses on pertinent information. This avoids truncation issues while keeping the input within the model's limits, thereby maintaining the quality and relevance of the generated responses.

Exam trap

Candidates often suggest increasing the context window or simply summarizing the whole document, ignoring that RAG is the standard architectural pattern for handling data larger than the context limit.

42
MCQmedium

A developer is building a multi-turn assistant that must remember details from earlier in a long conversation. After many exchanges, the developer notices that Claude starts forgetting information provided near the beginning of the session. The application currently sends only the latest user message with each API call. What is the most likely cause and the appropriate fix?

A.The developer must include the prior conversation turns in the messages array so the model has the necessary context in each request.
B.The developer should increase the max_tokens value so the model has more room to store conversation history internally.
C.The model has a limited memory that resets each call, so the developer must enable a persistent memory feature in the API request.
D.The developer should set the temperature to 0 so the model deterministically recalls earlier messages.
AnswerA

Because the Messages API is stateless, the model can only reason over the content provided in the current request. Omitting earlier turns means the model has no access to those details, which explains the forgetting. Sending the accumulated user and assistant messages in the messages array restores context and lets the model reference earlier information across turns.

Why this answer

The Messages API is stateless, so the model only sees what each request contains. Sending only the newest user message omits all prior context, which is why earlier details are forgotten. The correct remedy is to send the full conversation history in the messages array, alternating user and assistant turns, so the model can reference earlier information.

Generation parameters and output limits do not influence context retention.

Exam trap

The trap here is assuming the API maintains server-side conversation memory, when in fact each request must carry its own context.

43
MCQmedium

A developer wants to reduce latency for a complex multi-step prompt by 'pre-filling' the assistant's response. How is this achieved in the Messages API?

A.By using the 'prefill' parameter at the top level of the API request.
B.By ending the messages array with a message where the role is 'assistant'.
C.By adding a 'continuation' flag to the last user message.
D.By setting the 'system' prompt to include the desired starting text.
AnswerB

This is the documented method for pre-filling. When the last message in the array is from the 'assistant', Claude doesn't start a new response from scratch. Instead, it treats the provided text as the beginning of its own response and continues generating from that point, allowing for fine-grained control.

Why this answer

Pre-filling involves adding a message with the 'assistant' role as the last message in the 'messages' array. Claude will then continue the response from where that message left off. This is a powerful technique for steering the model's output format (e.g., starting with '{' for JSON) or persona, and it effectively 'nudges' the model into a specific state.

Exam trap

Many candidates mistakenly think pre-filling requires a special API flag or parameter, rather than simply structuring the messages array to end with an assistant-role message to guide the model.

44
Multi-Selectmedium

Which TWO of the following are valid ways to provide 'content' within a message object in the Claude API?

Select 2 answers
A.As a single string of text.
B.As a nested array of other message objects.
C.As an array of content blocks (e.g., text or image blocks).
D.As a binary blob of raw image data.
E.As a direct link to a local file path.
AnswersA, C

For most simple text-based interactions, providing the content as a single string is the most straightforward and common method. This is highly readable and sufficient for standard prompts where no images or specialized metadata blocks are required, reducing the complexity of the JSON payload sent to the API.

Why this answer

The 'content' field in a message can be either a simple string or an array of content blocks. This flexibility allows for basic text-only interactions as well as more complex multi-modal requests that include images or tool-related data. Understanding both formats is essential for developers moving from simple chat implementations to advanced vision or tool-augmented applications.

Exam trap

Many developers think content can only be passed as a complex array, failing to realize that simple text requests can use a plain string format.

45
Multi-Selecthard

Which TWO of the following are valid ways to handle long-running conversations within the Messages API to stay within context window limits?

Select 2 answers
A.Summarize previous chat history and append it as a single 'user' message.
B.Increase the system prompt size to include the entire conversation archive.
C.Remove older message pairs from the messages array before sending the request.
D.Increase the 'max_tokens' value to accommodate the total conversation length.
E.Restart the conversation by clearing the history after every five messages.
AnswersA, C

Summarization allows you to condense large amounts of historical context into a compact format. By injecting this summary into the conversation history, you maintain continuity while freeing up space for new user inputs, effectively managing the token budget without losing the core information gathered during the earlier conversation stages.

Why this answer

Managing the context window is vital for long-term state retention. Truncating the oldest messages or summarizing previous interactions are standard practices to ensure the most relevant information is always included in the prompt. These techniques prevent the context from exceeding the model's window, which would otherwise result in an API error and a failure to generate a valid response for the user.

Exam trap

Candidates often suggest clearing the entire history or using 'system' prompts to store conversation memory, failing to realize that context window limits apply to the entire message array.

46
MCQeasy

An engineer is writing a first integration against the Claude Messages API and needs to select the correct endpoint and required authentication header. The application will send a messages array with user and assistant turns and read the response content blocks. Which configuration is correct?

A.POST to /v1/chat/completions with the API key in an Authorization: Bearer header.
B.POST to /v1/complete with the API key in an Authorization: Bearer header.
C.POST to /v1/messages with the API key in the x-api-key header and an anthropic-version header.
D.POST to /v1/messages with the API key as a query string parameter named api_key.
AnswerC

The Messages API is reached by posting to /v1/messages with the API key supplied in the x-api-key header, and requests include an anthropic-version header that pins the API version. The body carries the model, max_tokens, and the messages array, and the response returns content blocks. This matches the described integration exactly.

Why this answer

The Messages API is invoked by posting to /v1/messages, authenticating with the x-api-key header, and including an anthropic-version header to select the API version. The request body carries the model, max_tokens, and the messages array, and the reply is returned as content blocks. Other paths or query-string credentials do not match the supported interface.

Exam trap

The trap here is carrying over an Authorization: Bearer plus chat completions pattern from another vendor's API instead of using the Claude messages endpoint and x-api-key header.

47
MCQmedium

A company is using the Claude API for a customer support chatbot. They notice that the 'stop_reason' in the API response is frequently 'max_tokens'. What does this indicate about the interaction?

A.Claude has successfully finished the task and stopped naturally.
B.The model was interrupted by a 'stop_sequence' defined by the user.
C.The response was truncated because it reached the specified length limit.
D.The input prompt was too long and exceeded the model's context window.
AnswerC

This is exactly what 'max_tokens' signifies. The model was still in the middle of generating text when it hit the limit set in the request. This often leads to sentences ending abruptly or the logic being cut short, indicating that the 'max_tokens' value is set too low for the expected output.

Why this answer

When the 'stop_reason' is 'max_tokens', it means Claude reached the limit specified by the 'max_tokens' parameter before it finished generating its complete answer. This results in a truncated response, which can be confusing for users. Developers should consider increasing the 'max_tokens' limit or optimizing the prompt to encourage more concise answers to ensure the full intent is delivered.

Exam trap

Candidates often confuse 'max_tokens' with a hard limit on the total context window size, failing to realize it is a generation limit that causes the response to truncate mid-sentence.

48
MCQhard

When designing a prompt for Claude, why is it recommended to place the most important instructions at the very beginning or the very end of the prompt?

A.To reduce the token count of the prompt.
B.To improve the model's attention to instructions.
C.To make the prompt easier to read for humans.
D.To allow the model to cache the instructions for future calls.
AnswerB

Empirical testing shows that models are more robust at following instructions located at the beginning or end of a long prompt. This placement strategy mitigates the risk of the model ignoring middle-ground instructions, leading to more reliable and predictable performance when the context window is highly populated with information.

Why this answer

Models sometimes suffer from 'lost in the middle' phenomena, where information buried in the center of a long context is less likely to be prioritized. By placing key instructions at the start or end, you leverage the model's tendency to focus on the initial 'pre-fill' context and the final instructions, ensuring that the model adheres strictly to your defined constraints and goals.

Exam trap

Candidates bury critical instructions in the middle of long, dense prompts, failing to realize that models often struggle to maintain focus on information located far from the start or end.

49
Multi-Selectmedium

A team is instrumenting a production Messages API integration and wants to programmatically detect throttling and transient server faults so their client library can back off and retry. Which TWO HTTP status codes should their response handler treat as retryable conditions? (Choose two.)

Select 2 answers
A.401
B.400
C.529
D.429
E.403
AnswersC, D

A 529 signals that Anthropic's infrastructure is overloaded and the request could not be served at that moment. It is a transient capacity condition, distinct from client error, and typically clears within seconds to minutes. Retrying with exponential backoff and jitter is the appropriate response, and this code is exactly the kind of server-side fault a resilient client should absorb silently.

Why this answer

Retryable conditions are those where the same request may succeed later without modification. Rate limiting and infrastructure overload both fall in that category because they reflect temporary capacity or throughput constraints on the service side, not defects in the request. Client-side errors such as malformed input, bad credentials, or insufficient permissions will reproduce deterministically and should be fixed rather than retried.

Exam trap

The trap here is lumping every 4xx response into a generic retry bucket when client errors reproduce deterministically and only throughput or capacity faults are worth repeating.

50
Multi-Selectmedium

A developer is implementing streaming with the Claude API. Which TWO of the following event types are standard parts of the Anthropic Server-Sent Events (SSE) stream?

Select 2 answers
A.content_block_delta
B.token_count_update
C.message_start
D.ping_pong_check
E.model_switch_event
AnswersA, C

The content_block_delta event is the most frequent event in a stream, carrying the actual fragments of text (tokens) as they are generated. Applications listen for this event to update the UI in real-time, providing the 'typing' effect that users expect from conversational AI interfaces.

Why this answer

Anthropic's streaming API uses Server-Sent Events to provide real-time updates as the model generates text. Understanding the specific event types, such as 'message_start' and 'content_block_delta', is crucial for developers to correctly parse the stream, update user interfaces incrementally, and handle metadata like token usage and stop reasons as they arrive from the server.

Exam trap

Test-takers confuse SSE stream event names with standard HTTP status codes or generic webhook payloads, guessing invalid lifecycle event names.

51
MCQeasy

Which of the following headers provides information about the remaining request quota for a specific API key after a call is made?

A.x-quota-left
B.anthropic-ratelimit-requests-remaining
C.retry-after-seconds
D.anthropic-token-usage
AnswerB

This is the correct header. It provides a real-time count of the remaining requests allowed in the current rate limit window. By tracking this value, developers can implement client-side throttling to prevent hitting the limit and receiving 429 errors, which improves the overall reliability of the integration.

Why this answer

Anthropic includes rate limit information in the response headers of every API call. Specifically, the 'anthropic-ratelimit-requests-remaining' header tells the developer how many more requests they can make within the current time window. Monitoring these headers is essential for building robust applications that can gracefully handle or avoid rate-limiting scenarios during high usage.

Exam trap

Candidates frequently confuse the 'anthropic-ratelimit-requests-remaining' header with the 'retry-after' header, which is used for timing when to resume requests after being rate-limited.

52
MCQeasy

A developer is building a document-summarization service with the Claude API. After upgrading from a Claude 3 model to a newer model, the application's HTTP requests start failing with a 400 error complaining about an unrecognized parameter. The developer's request body includes a field that was accepted before. Which change to the request body is MOST likely required?

A.The deprecated 'max_tokens_to_sample' field must be replaced with 'max_tokens'.
B.The 'messages' array must be renamed to 'prompt' to match the newer schema.
C.The 'system' prompt must be converted into a user message before sending.
D.The 'temperature' parameter must be moved inside the 'metadata' object.
AnswerA

The older Text Completions-era field max_tokens_to_sample was removed in favor of max_tokens on the Messages API. A request still sending max_tokens_to_sample to a current model endpoint returns a 400 for an unrecognized parameter, so renaming the field to max_tokens restores compatibility for this summarization service.

Why this answer

Modern Claude models are called through the Messages API, where the output-length cap is expressed as max_tokens. Requests that still carry the older max_tokens_to_sample field are rejected as malformed, so the minimal correct fix is to rename that field. Parameters such as temperature and system remain top-level and valid, so restructuring them would not resolve the 400 error.

Exam trap

The trap here is assuming that any 400 from the API means the message structure is wrong, when a single legacy field name is often the sole cause.

53
MCQmedium

You are building a chat application using Claude 3.5 Sonnet. You need to ensure the model maintains a consistent tone and follows specific formatting constraints across a long conversation. Which approach best optimizes API usage while maintaining context?

A.Send only the last user message and the system prompt for every request.
B.Store the conversation history in the system prompt rather than the messages array.
C.Maintain the conversation history in a list and include the full history in every request.
D.Use the Claude session ID header to automatically track state on the server.
AnswerC

The Claude API is stateless, meaning it does not remember previous interactions by default. By maintaining a message list and sending the full history in each request, you allow the model to see the entire conversation flow, ensuring that tone and formatting constraints are consistently applied throughout.

Why this answer

Passing the entire conversation history in every API call is the standard way to maintain context in stateless REST architectures. By managing the history array on the client side and sending the full thread, you ensure Claude has the necessary state to maintain tone. Token optimization is secondary to statefulness; you can prune older messages if the context window approaches limits, but full history is required for consistency.

Exam trap

Developers sometimes try to store state on the server or forget to include past turns, assuming the API automatically remembers previous prompts.

54
MCQeasy

Which of the following describes the correct behavior when using the 'stream' parameter in the Messages API?

A.The API sends the entire response at once, but with lower latency.
B.The API returns content as it is generated in discrete events.
C.The API skips the 'messages' array and only accepts a single prompt string.
D.The 'stream' parameter is only available for the Claude 2 model series.
AnswerB

Streaming mode breaks the output into a stream of events. This is essential for building responsive user interfaces where you want to show the model's output in real-time. By handling these events on the client side, you create a much smoother, more interactive experience for the end user.

Why this answer

Setting 'stream' to true causes the API to return a series of Server-Sent Events (SSE) rather than a single JSON object. This is a critical pattern for interactive applications, as it provides an immediate perceived response time for the user. The client must be prepared to handle these discrete events, which include content chunks, metadata, and final completion signals.

Exam trap

Test-takers frequently assume streaming returns a single modified JSON object or newline-delimited standard JSON, forgetting that Anthropic uses Server-Sent Events (SSE) with discrete event chunks.

55
MCQhard

A developer is tuning a classification workload that sends many short prompts to the Claude Messages API. They want to reduce cost and latency without changing the model or the prompt text. They set up prompt caching for the shared system prompt. Which configuration correctly enables caching for that system prompt block?

A.Set a top-level enable_cache boolean to true in the request body alongside the model and max_tokens fields.
B.Send the system prompt as a separate request first, then reference its id in subsequent message requests.
C.Add a cache_control field to the model parameter so the selected model is cached between calls.
D.Add a cache_control object with type "ephemeral" to the system prompt content block in the request.
AnswerD

Prompt caching is enabled per content block by attaching a cache_control field with type "ephemeral". Placing it on the system prompt block marks that prefix for caching, so repeated requests reuse the cached tokens and reduce both cost and latency. This is the documented mechanism and requires no change to the model or prompt wording.

Why this answer

Prompt caching in the Messages API is opt-in per content block using a cache_control field with type "ephemeral". Marking the shared system prompt lets the API reuse that prefix across requests, cutting cost and latency for repeated short prompts. No top-level flag or separate resource is involved, and the model and prompt text stay unchanged.

Exam trap

The trap here is assuming caching is a global request flag or a separate stored resource, when it is actually declared inline on individual content blocks.

56
MCQmedium

A developer is sending a large knowledge-base article to the Claude Messages API for analysis. The API returns HTTP 400 with an error message indicating the input is too long. Which statement best describes the correct remediation?

A.Switch the request method from POST to GET to bypass the length restriction.
B.Retry the request with exponential backoff because the error is transient.
C.Increase the max_tokens parameter so the model has room for both the input and the output.
D.Reduce or chunk the input content so the prompt plus the requested output fits within the model's context window.
AnswerD

The 400 error signals that the combined prompt and expected completion exceed the model's context window. Splitting the article into smaller chunks, summarizing sections, or trimming irrelevant text brings the request within limits. This directly addresses the reported cause and allows the request to succeed without changing unrelated parameters.

Why this answer

An HTTP 400 with an explicit input-length message means the prompt plus expected completion exceeds the model's context window. The only correct fix is to shrink the input through chunking, summarization, or removal of irrelevant content. Parameters like max_tokens control output length, and retry logic only helps transient failures, so neither resolves a deterministic size violation.

Exam trap

The trap here is treating a deterministic 400 invalid-request error as if it were a transient failure that retries or parameter tweaks could overcome.

57
Multi-Selecthard

A developer wants to utilize 'Tool Use' (function calling) with Claude. Which TWO steps are necessary to properly define a tool in the API request?

Select 2 answers
A.Provide a 'description' field explaining what the tool does.
B.Include the tool's source code in the 'implementation' field.
C.Define an 'input_schema' using JSON Schema format.
D.Set the 'role' of the message to 'function'.
E.Upload a CSV file containing example tool outputs.
AnswersA, C

The description is vital because it is the primary way Claude understands the tool's purpose. Without a clear description, the model may not know when it is appropriate to call the tool or what real-world action the tool represents, leading to poor tool selection and incorrect workflow execution.

Why this answer

To use tools, developers must provide a 'tools' array in the request. Each tool needs a 'name', a 'description' (which helps Claude understand when to use it), and an 'input_schema'. The 'input_schema' must be a valid JSON Schema object, defining the parameters the tool expects.

This structured approach allows Claude to generate valid arguments that the developer's code can then execute.

Exam trap

Many candidates forget to include an explicit description for each tool, wrongly assuming that the input schema alone is enough for Claude to know when to call it.

58
Multi-Selectmedium

An enterprise is scaling their Claude integration and needs to manage their rate limits effectively. Which TWO strategies are recommended by Anthropic to handle rate limiting gracefully?

Select 2 answers
A.Implement exponential backoff for retries
B.Request a limit increase immediately upon the first 429 error
C.Monitor 'anthropic-ratelimit' headers in responses
D.Use multiple API keys to multiply the available rate limits
E.Switch to a smaller model only when a rate limit is hit
AnswersA, C

Exponential backoff is a standard error-handling technique where the client waits longer between each successive retry of a failed request. This prevents 'thundering herd' problems where many clients overwhelm the server at once. It is the most effective way to recover from 429 errors while respecting the server's capacity limits.

Why this answer

Rate limits are a reality of high-scale API usage. To handle them, developers should implement exponential backoff, which increases the wait time between retries after each failure. Additionally, they should monitor the rate limit headers in the API response to adjust their request frequency dynamically.

These practices prevent the application from being blocked and ensure smoother overall performance.

Exam trap

Candidates often rely purely on static hardcoded sleep timers for rate limiting, ignoring response headers and exponential backoff best practices.

59
MCQmedium

A developer is writing an assistant that must always answer in strict JSON with a fixed set of keys. They want to guarantee the model's output is parseable without writing custom repair logic. Which approach best fits the Messages API?

A.Set 'temperature' to 0 and rely on the model to output valid JSON.
B.Append the word 'JSON' to the system prompt and trust the model's compliance.
C.Define a tool with an input schema and require the model to call it.
D.Use the 'stop_sequences' parameter to halt generation at the closing brace.
AnswerC

Tool definitions carry a JSON Schema for their input, and when the model is required to use a tool, its arguments arrive as structured data conforming to that schema. This gives a reliable, machine-parseable shape for the fixed-key JSON the assistant must return, removing the need for hand-written repair logic.

Why this answer

The Messages API supports tool use, where each tool declares an input_schema in JSON Schema. When the model is required to call that tool, the returned arguments conform to the declared schema, yielding reliably structured output. Sampling settings, stop sequences, and prompt wording all shape behavior probabilistically but cannot guarantee a parseable, fixed-key JSON object.

Exam trap

The trap here is believing that temperature 0 or a strong prompt instruction guarantees valid JSON, when only a schema-constrained mechanism does.

60
MCQeasy

A developer is writing code that calls the Anthropic Messages API and needs to authenticate each request. The developer has retrieved the API key from a secure secret manager at runtime. Where should the API key be placed in the HTTP request?

A.In the URL as a query parameter such as ?api_key=sk-ant-...
B.In the request body as a top-level 'api_key' field alongside 'model' and 'messages'.
C.In the 'Authorization' header using the scheme 'Basic' with the key base64-encoded.
D.In the 'x-api-key' HTTP header on each request.
AnswerD

The Anthropic Messages API authenticates requests using the x-api-key header, and the value is the API key retrieved from the secret manager. This keeps the credential out of the URL and the JSON body, reducing leakage risk. Sending it this way on every request is the documented and expected authentication method for direct HTTP integrations with the API.

Why this answer

Direct HTTP calls to the Anthropic Messages API authenticate by including the API key in the x-api-key request header. This keeps the credential out of URLs and payloads, which is important for security and log hygiene. Body fields, query parameters, and Basic authentication schemes are not recognized by the API for this purpose, so requests using them will fail authentication.

Exam trap

The trap here is assuming the API uses the generic Authorization header or a body field, when it specifically requires the x-api-key header.

61
MCQhard

Refer to the exhibit. An engineer wants to use Prompt Caching to optimize this request. What is the correct way to modify the request body to enable this?

A.Add a 'cache_control' field to the root level of the JSON body.
B.Nest a 'cache_control' object within the content block of the message.
C.Rename the 'system' field to 'system_cached'.
D.Enable 'caching=true' in the request headers.
AnswerB

Prompt caching is activated by adding a cache_control block to the content of a message. This instructs the Anthropic API to store the preceding tokens in the cache. This is the correct structural way to enable the caching feature as defined in the Anthropic API technical documentation.

Why this answer

Prompt caching requires the inclusion of a 'cache_control' block within the message content. This tells the API to store the processed state of that specific chunk. This is critical for high-latency or high-cost prompts because it allows for faster processing of subsequent requests that share the same cached prefix, significantly reducing latency and cost for repetitive tasks like data analysis.

Exam trap

Candidates often try to enable caching by adding a top-level parameter or a system prompt flag, forgetting that 'cache_control' must be placed inside the specific content block.

62
MCQhard

Refer to the exhibit. What is the specific effect of the 'tool_choice' parameter as configured in this request?

A.It allows Claude to choose between using the tool or answering normally.
B.It forces Claude to use the 'get_weather' tool specifically.
C.It prevents Claude from using any tools during this request.
D.It enables Claude to use multiple tools in a single response.
AnswerB

By explicitly naming the tool in the tool_choice object, the developer overrides the model's default behavior. Claude will be compelled to generate a tool_use block for 'get_weather', ensuring that the application receives the structured data it needs to proceed with the weather-related logic, regardless of the prompt's phrasing.

Why this answer

The 'tool_choice' parameter allows developers to control Claude's tool-calling behavior. By setting it to type 'tool' and specifying a 'name', the developer is forcing Claude to use that specific tool, even if it thinks it could answer without it. This is useful for specialized workflows where a specific function must always be executed as the next step in the application logic.

Exam trap

Candidates often mistake 'tool_choice' for a general suggestion, not realizing that 'tool_choice: {type: "tool", name: "..."}' forces the model to use that specific tool.

63
MCQmedium

A developer is building a backend service that calls the Claude Messages API to summarize user-submitted articles. The service must enforce a hard limit: no summary should exceed 500 tokens. The developer sets max_tokens to 500. During testing, a response returns stop_reason: "max_tokens". What is the most accurate interpretation of this result?

A.The input article exceeded the context window, so the model could not process the entire document.
B.The model finished its summary naturally and the value indicates the summary is exactly 500 tokens long.
C.The response was truncated because it reached the max_tokens limit before the model naturally finished its output.
D.The model encountered an internal error and stopped generating tokens prematurely.
AnswerC

stop_reason "max_tokens" means the model hit the configured output token ceiling and the response was cut off, so the summary may be incomplete. The developer should treat this as a truncation signal, possibly increase max_tokens, or shorten the requested output, and then verify the content is complete before returning it to the user.

Why this answer

The stop_reason field reports why the model stopped generating. A value of "max_tokens" indicates the output was cut off because it reached the configured ceiling, meaning the content may be incomplete. Developers should treat this as a truncation signal and decide whether to raise max_tokens, shorten the prompt, or handle partial output gracefully.

Exam trap

The trap here is confusing an output-side truncation signal (stop_reason "max_tokens") with an input-side context window overflow error.

64
MCQhard

A developer wants Claude to call an internal inventory function when a user asks about stock levels, but the function must only be invoked when the model decides it is needed. The application will execute the function and return the result. Which sequence correctly implements this with the Messages API?

A.Send the tools definition; if the response contains a tool_use block, run the function and send a new request including a tool_result block referencing the tool_use id.
B.Pre-execute the inventory function on every request and inject its output into the system prompt.
C.Define the function in the system prompt as a description and parse the model's prose reply for a function name.
D.Send the tools definition and, upon receiving a tool_use block, immediately send a tool_result with an empty payload to acknowledge it.
AnswerA

This matches the tool use contract: tools are declared in the request, the model may return a tool_use content block with an id, and the application executes the function and returns a tool_result block that references that same id in a follow-up request, allowing the model to produce a final answer grounded in real data.

Why this answer

Tool use is a round trip: the application declares tools, the model optionally returns a tool_use block, the application executes the named function, and it sends back a tool_result block that references the tool_use id so the model can incorporate the real output. Pre-running functions, prose-based parsing, or empty acknowledgements all break this contract.

Exam trap

The trap here is believing the model itself executes the function, when in fact the application runs it and must return the result keyed to the tool_use id.

65
Multi-Selectmedium

A developer is building a robust error-handling wrapper for the Messages API. Which THREE HTTP status codes should specifically trigger a retry logic with exponential backoff in a production environment?

Select 3 answers
A.400 - Bad Request
B.429 - Too Many Requests
C.500 - Internal Server Error
D.529 - Overloaded
E.401 - Unauthorized
AnswersB, C, D

This code is returned when the client has exceeded their rate limit. It is a classic 'transient' error that should be handled with exponential backoff. Retrying after a short delay allows the rate limit bucket to refill, ensuring the application can eventually complete its task without failing permanently for the user.

Why this answer

Handling API errors correctly is essential for application stability. 429 errors mean you are being rate-limited and should wait. 500 errors indicate a general server-side issue, while 529 errors mean the server is currently overloaded. All three represent temporary conditions where a retry might succeed. Conversely, errors like 400 or 401 represent client-side issues that retrying will not fix.

Exam trap

Candidates often include 400 (Bad Request) or 401 (Unauthorized) in their retry logic, not realizing that these client-side errors indicate invalid requests that will never succeed upon retry.

Ready to test yourself?

Try a timed practice session using only Using The Claude Api questions.