Courseiva

CCNA Claude API Mechanics Questions

34 questions · Claude API Mechanics · All types, answers revealed

1
MCQhard

A team is building a retrieval-augmented assistant that sends a large document set inside the system prompt on every turn. To reduce cost, they enable prompt caching. They notice cache_read_input_tokens is high on most turns but intermittently drops to zero even though the system prompt text has not changed. Which explanation best fits this behavior?

A.The API randomly invalidates cache entries per account to distribute load, so intermittent cache misses are expected and cannot be controlled.
B.Prompt caching only applies to user messages, so any content placed in the system prompt bypasses the cache and is billed at full input rate.
C.The cache has a short time-to-live, so if enough time passes between requests the cached prefix expires and the next request pays full input cost before re-populating the cache.
D.Cache reads only register when max_tokens is large enough to cover the entire cached prefix, so smaller responses skip the cache.
AnswerC

Prompt caching uses a time-to-live window; if no request references the cached prefix before it expires, the entry is evicted. The next request then misses the cache, shows zero cache reads, and re-creates the entry, which matches the intermittent pattern described. Steady traffic keeps the cache warm, while idle gaps cause the drop to zero.

Why this answer

Prompt caching entries expire after a time-to-live if they are not referenced. Bursty or idle traffic patterns cause the prefix to be evicted between requests, so the next call misses the cache and reports zero cache reads before re-creating the entry. Keeping steady traffic, or reducing gaps between calls, stabilizes cache_read_input_tokens and preserves the intended cost savings.

Exam trap

The trap here is assuming cache misses must be caused by changing content or a broken marker, when the real cause is often the cache entry expiring during an idle gap.

2
MCQmedium

A developer wants Claude to return a fixed set of user-profile fields extracted from free-text bios. They want the response to be valid JSON they can parse without regex. The model occasionally adds a friendly sentence before the JSON. Which Messages API feature most reliably constrains the output?

A.Increase max_tokens so Claude has more room to format the JSON correctly.
B.Define a tool whose input_schema describes the profile fields, then set tool_choice to force Claude to call it.
C.Set stop_sequences to "}" so generation halts exactly at the end of the JSON object.
D.Add the word "JSON" to the prompt and set temperature to 0.
AnswerB

Tool use with a JSON input schema gives Claude a structured contract, and forcing tool_choice makes it emit a tool_use block whose input conforms to that schema. This yields parseable structured data without preamble. It is the API-native way to obtain schema-constrained extraction, avoiding regex cleanup of free text.

Why this answer

Structured extraction is best enforced by defining a tool whose input_schema specifies the exact fields and types, then using tool_choice to require that the model call it. Claude then returns a tool_use block whose input is JSON matching the schema, eliminating stray prose. Prompt wording, temperature, stop sequences, and token limits influence style or length but cannot enforce a schema.

Exam trap

The trap here is relying on prompt phrasing or temperature to guarantee valid JSON when only a schema-backed tool contract enforces structure.

3
MCQmedium

Refer to the exhibit. A developer encounters this error when trying to send a conversation history. What must the developer do to fix the 'messages' array structure?

A.Insert a blank 'assistant' message between the two 'user' messages.
B.Merge the content of the consecutive 'user' messages into a single message object.
C.Change the 'role' of the second 'user' message to 'system'.
D.Enable the 'allow_consecutive_roles' flag in the request headers.
AnswerB

Merging the content is the most robust way to resolve this error. By combining the text from both user inputs into one 'user' role message, the developer maintains the logical flow while adhering to the API's structural requirement for alternating roles between participants in the conversation.

Why this answer

The Messages API requires a strict alternation of roles to maintain a clear dialogue structure. A 'user' message must be followed by an 'assistant' message, and vice versa. If a developer has multiple consecutive pieces of information from the same role, they must be combined into a single message or separated by a response from the other role.

Exam trap

Candidates often try to insert an empty assistant message to separate two user messages, which is an invalid workaround that causes API errors due to improper turn alternation.

4
MCQhard

An application needs to ensure that Claude never uses its pre-trained knowledge to answer questions, but only uses the provided context. After placing instructions in the system prompt, the developer still sees Claude occasionally using outside info. What is the most effective API-level mechanical adjustment to further constrain Claude?

A.Set 'temperature' to 1.0 to encourage more creative following of the system instructions.
B.Use pre-filling by ending the 'messages' array with an 'assistant' message like 'Based ONLY on the context provided, I will answer...'
C.Increase the 'top_p' value to 1.0 to ensure all possible valid answers are considered.
D.Add a 'stop_sequence' for the word 'I' to prevent the model from speaking in the first person.
AnswerB

Pre-filling the assistant response is the strongest way to guide Claude. By starting its response with a commitment to use only the provided context, the model's internal attention mechanism is heavily weighted towards that constraint for the remainder of the generation, significantly reducing the likelihood of outside knowledge leakage.

Why this answer

While system prompts are the primary way to give instructions, pre-filling the assistant's response is a more forceful mechanical constraint. By starting the assistant's response with a specific phrase, you lock the model into a particular persona or logical path, making it much harder for the model to deviate into its default behaviors or general knowledge.

Exam trap

Candidates over-rely on system prompts for strict grounding, failing to realize that pre-filling the assistant response is a more effective mechanical constraint to prevent the model from hallucinating.

5
MCQhard

A high-traffic application is frequently sending the same large set of reference documents (100,000 tokens) in the system prompt for every request. Which Claude API feature would most effectively reduce both the latency and the cost of these requests?

A.Switching from Claude 3 Opus to Claude 3 Haiku.
B.Implementing Prompt Caching using the 'cache_control' metadata.
C.Compressing the reference documents using a text summarization tool.
D.Increasing the 'max_tokens' limit to allow for larger responses.
AnswerB

Prompt Caching allows the API to store the results of processing a prefix of the prompt. By tagging the large reference documents with 'cache_control: {"type": "ephemeral"}', subsequent requests that use the same prefix can reuse the cached state, leading to lower costs and much faster response times.

Why this answer

Prompt Caching is a specialized mechanic for handling repetitive, large-scale context. By marking parts of the prompt as cacheable, developers can avoid re-processing the same data across multiple requests. This lead to significant cost savings on input tokens and drastically reduces the time to first token for the model's generated response.

Exam trap

Candidates often try to optimize by manually truncating the prompt or using external vector databases, missing that built-in Prompt Caching is the specific, optimized mechanism for handling repetitive large context.

6
MCQmedium

A developer is building a document-summarization service on the Anthropic Messages API. The service must always return a summary followed by exactly one structured JSON block containing metadata (e.g., word count, sentiment). To guarantee the model never emits any text after the JSON block, the developer adds the closing brace of the JSON structure to the 'stop_sequences' array. What is the effect of this configuration on the API response?

A.The API returns an error because stop_sequences cannot contain characters that might appear inside the model's generated JSON.
B.The response's 'stop_reason' is 'max_tokens' because stop sequences only apply to the final assistant turn, not to content inside a single message.
C.The API silently strips the entire JSON block from the response, returning only the summary text to the caller.
D.The response's 'stop_reason' is 'stop_sequence' and the text field contains the JSON block with the stop string trimmed off.
AnswerD

When a stop sequence is encountered, generation halts and the matched stop string is excluded from the returned text. The 'stop_reason' field is set to 'stop_sequence' so callers can distinguish this from a natural end_turn. The developer gets the summary plus JSON with the trailing brace removed, which matters for downstream parsing.

Why this answer

A stop sequence halts generation as soon as the matching string is produced, and the matched string itself is not included in the returned text. The response's stop_reason becomes 'stop_sequence', letting the application know a configured terminator fired rather than the model finishing naturally. This is why developers who place a closing JSON brace in stop_sequences see the brace omitted from the output.

Exam trap

The trap here is assuming stop sequences act like output filters that remove content, when they actually truncate generation at the match point and drop the matched string itself.

7
MCQeasy

When using the Messages API, what happens if the 'messages' array provided in the request is empty?

A.Claude will generate a random greeting message to start the conversation.
B.The API will return a 400 Bad Request error indicating that at least one message is required.
C.The API will process the 'system' prompt and return a response based solely on those instructions.
D.Claude will return the last response it generated for that specific API key.
AnswerB

The API validation layer requires the 'messages' array to be non-empty. If a developer sends an empty array, the server will return a 400 error. This ensures that every request contains at least one piece of user input (or assistant pre-fill) for the model to process.

Why this answer

The Messages API has strict validation rules for the 'messages' array. It is the core of the request, representing the conversation that Claude is meant to continue. An empty array is logically equivalent to asking a model to respond to nothing, which is not supported by the API's current design and validation logic.

Exam trap

Test-takers sometimes assume an empty messages array will be treated as an empty initial prompt or simply ignored by the API, rather than triggering an immediate validation failure.

8
Multi-Selectmedium

In the Messages API 'messages' array, which TWO roles are currently supported for maintaining conversation history?

Select 2 answers
A.user
B.system
C.assistant
D.function
E.admin
AnswersA, C

The 'user' role represents instructions or queries provided by the human interacting with the model. It is a required role for the first message in the array (unless the assistant response is being pre-filled) and is used to provide the context that Claude must respond to.

Why this answer

The Messages API enforces a strict structure for conversation history to ensure the model correctly understands the dialogue flow. Unlike some other APIs that allow custom roles, Claude specifically recognizes two primary roles within the messages array. Correctly utilizing these roles is fundamental to building a coherent chat history that the model can process.

Exam trap

Candidates frequently try to insert 'system' or 'developer' roles directly into the messages array, forgetting that the Messages API strictly requires alternating 'user' and 'assistant' roles for the conversation history.

9
MCQmedium

When configuring an API call to generate a specific JSON object, a developer adds the string '}' to the 'stop_sequences' array. What is the most likely outcome of this configuration?

A.Claude will successfully generate the full JSON object and then stop.
B.The API will return an error because stop sequences cannot be single characters.
C.Claude will generate the JSON object, but the final '}' will be missing from the response.
D.The model will ignore the stop sequence if it occurs within a code block.
AnswerC

Stop sequences work by terminating generation the moment the sequence is matched. The matching sequence itself is not included in the response text. Therefore, the model will stop right after it intends to close the JSON, but the actual '}' will be absent from the payload.

Why this answer

The 'stop_sequences' parameter tells Claude to stop generating text as soon as a specific string is produced. If a developer uses a character that is required for the structural integrity of the output—like the closing brace of a JSON object—the model will terminate immediately upon producing that character, often leaving the output incomplete or invalid for parsers.

Exam trap

Candidates often assume stop sequences are ignored if they are critical to the output format, failing to realize the API stops immediately upon matching the character, causing structural JSON truncation.

10
MCQmedium

A developer is building a Python application that streams responses from the Anthropic Messages API. They notice that when they set stream=True, the response object is an iterator of server-sent events. They want to extract only the incremental text chunks as they arrive, without waiting for the full message. Which event type should they filter for to get the text deltas?

A.message_delta
B.message_start
C.content_block_delta
D.ping
AnswerC

The content_block_delta event is emitted each time a new piece of text (or tool input) is generated within a content block. It carries a delta object that includes the incremental text. By filtering for this event, the developer can accumulate the partial text chunks as they stream, which is exactly what is needed for real-time display or processing without waiting for the full message.

Why this answer

When streaming with the Messages API, the server sends a series of events. The content_block_delta event is specifically designed to carry incremental text deltas as Claude generates the response. By listening for this event type, the developer can process text as it arrives, enabling real-time user experiences.

Other events like message_start, message_delta, and ping serve different purposes such as initialization, final metadata, or keep-alive.

Exam trap

The trap here is confusing message_delta with content_block_delta, assuming that any event with 'delta' in the name carries the streamed text, when in fact message_delta only carries final metadata.

11
MCQeasy

An application needs to ensure that Claude stops generating text as soon as it produces a specific character sequence, such as 'END_OF_REPORT'. Which API feature should be used to implement this behavior?

A.The 'stop_sequences' top-level parameter.
B.Setting 'max_tokens' to the exact length of the expected report.
C.Using a 'system' prompt to tell Claude to stop at 'END_OF_REPORT'.
D.Setting 'temperature' to 0 to make the output more predictable.
AnswerA

The stop_sequences parameter allows you to define up to 20 custom strings that will signal Claude to stop generating. When the model generates any of these sequences, it terminates the response immediately. The sequence itself is not included in the final output, making it perfect for clean text termination.

Why this answer

Stop sequences are a fundamental tool for controlling the termination of Claude's output. By providing a list of strings, developers can force the model to cease generation immediately upon producing those strings. This is vital for maintaining the structure of generated documents and ensuring that the model does not continue into unwanted or redundant text.

Exam trap

Candidates often confuse stop sequences with max_tokens or attempt to instruct Claude via system prompts to stop generating, which is less reliable.

12
MCQmedium

Refer to the exhibit. When submitting this request via a standard HTTP client, which header is mandatory to specify the API version and ensure compatibility with the Messages API?

A.x-api-version: 2023-06-01
B.anthropic-version: 2023-06-01
C.version: claude-v3
D.anthropic-model-version: 2024-06-20
AnswerB

This is the correct, mandatory header required for all calls to the Anthropic Messages API. It informs the server which version of the API logic to execute. The value '2023-06-01' is the current standard version string used for the Claude 3 family and Messages API interactions.

Why this answer

Anthropic requires a specific versioning header for all requests to the Messages API to ensure that developers are using the intended API contract. This prevents breaking changes from affecting existing integrations when the API evolves. Without this header, the API will return a 400 error, as it cannot determine which schema and logic to apply to the request.

Exam trap

Developers often omit the versioning header or use an incorrect date format, assuming the API defaults to the latest version, which results in immediate 400 Bad Request errors.

13
Multi-Selectmedium

When a developer uses Prompt Caching, which TWO metrics are specifically returned in the 'usage' object of the API response to help track cache performance?

Select 2 answers
A.cache_creation_input_tokens
B.cache_read_input_tokens
C.cache_expiry_timestamp
D.cache_hit_ratio
E.total_cached_tokens_stored
AnswersA, B

This metric counts the number of tokens that were written to the cache for the first time during the current request. These tokens are billed at the standard input rate but will contribute to future savings if the same prefix is used in subsequent requests to the API.

Why this answer

Prompt Caching introduces new usage categories to the API response. Monitoring these metrics is vital for developers to calculate their actual costs and verify that their caching strategy is working as expected. These fields allow for a granular breakdown of how many tokens were read from the cache versus how many were processed normally.

Exam trap

Candidates often look for generic token counters like input_tokens or output_tokens, forgetting the specific prompt caching usage keys returned in the API response.

14
MCQhard

Refer to the exhibit. What is the expected behavior of Claude when receiving this specific API request configuration?

A.Claude will provide a conversational text response about London's typical weather without using any tools.
B.Claude will return an error because the user did not provide a specific API key for the weather service.
C.Claude will bypass the tool and ask the user for more clarification about which part of London they mean.
D.Claude will immediately generate a tool_use block for the 'get_weather' tool with 'London' as the location.
AnswerD

The tool_choice parameter with type 'tool' and a specific name forces Claude to use that exact tool. Since the user mentioned London, Claude will populate the required 'location' parameter in the JSON output, fulfilling the instruction to use the tool as the first and only action.

Why this answer

The tool_choice parameter allows developers to override Claude's natural decision-making process. By setting it to a specific tool, you force the model to generate a tool-use block for that tool, regardless of whether it thinks it's necessary. This is a powerful mechanic for building rigid workflows where a specific step must always result in a structured function call.

Exam trap

Candidates assume the model will use its 'reasoning' to decide whether to call the tool, ignoring that the 'tool_choice' parameter forces execution regardless of the model's actual internal assessment.

15
MCQmedium

Refer to the exhibit. An application monitoring system captures this response from the Anthropic API. Which strategy is the most mechanically sound approach for the application to take to resolve this specific error and continue processing?

A.Immediately resending the request with a new API key to bypass the limits.
B.Switching the model parameter to a larger model like Claude 3 Opus.
C.Implementing an exponential backoff algorithm before retrying the request.
D.Increasing the max_tokens parameter to ensure the request is prioritized.
AnswerC

Exponential backoff involves waiting for an increasing amount of time before each retry attempt. This gives the API's rate-limiting window time to reset without flooding the service with repetitive requests. It is the industry-standard method for managing transient errors and ensuring the stability of distributed systems during high load.

Why this answer

Rate limit errors occur when a developer exceeds the allocated requests per minute or tokens per minute for their tier. Handling these gracefully via exponential backoff is a fundamental skill for API integration. This approach ensures that the application doesn't overwhelm the server further, allowing the rate limit bucket to refill naturally while maintaining the best possible user experience.

Exam trap

Candidates often confuse rate limit errors (429) with server overload errors (529) or mistakenly assume that immediate retries without delays are acceptable, which further overwhelms the API infrastructure during traffic spikes.

16
MCQmedium

A developer is using 'top_k' to control the diversity of Claude's responses. If they set 'top_k' to 1, what is the expected behavior of the model during token generation?

A.Claude will only use words from the top 1% of its vocabulary.
B.Claude will pick the most likely next token every time, similar to 'greedy' decoding.
C.Claude will prioritize tokens that have only appeared once in the preceding text.
D.Claude will stop generating after exactly one token is produced.
AnswerB

When 'top_k' is 1, the model is restricted to a single choice: the token with the highest probability. This effectively disables the randomness of sampling, making the output predictable and deterministic. This is identical to the behavior of 'greedy' decoding used in many linguistic models.

Why this answer

The 'top_k' parameter limits the model's choice of the next token to the 'k' most likely options. Setting this value to 1 is a restrictive configuration that forces the model to always pick the single most probable token. This results in highly deterministic and 'greedy' sampling, which can be useful for tasks requiring high precision.

Exam trap

Candidates often believe 'top_k' at 1 increases randomness or creativity, failing to recognize that it forces the model to choose only the single most probable token every time.

17
MCQmedium

A developer is streaming a long Claude response using the Messages API with stream: true. Their client code reads Server-Sent Events and appends text to the UI. Mid-stream, the connection drops and the client reconnects by re-issuing the same request from scratch. Users complain that the answer restarts from the beginning. What is the most accurate explanation of what is happening?

A.Streaming responses are stateless; a dropped connection cannot be resumed, so re-issuing the request starts a brand new generation and the partial text already shown is discarded.
B.The client must call the count_tokens endpoint before reconnecting so the server can splice the new stream onto the previous tokens.
C.The Messages API supports resuming a stream from the last received event, so the client should have sent the last event ID back on reconnect.
D.Setting a higher max_tokens on the reconnect lets the API detect the earlier partial output and continue from where it stopped.
AnswerA

Streaming in the Messages API is a one-shot delivery of an in-progress generation. If the HTTP connection drops, the server-side generation is not resumable by the client, so re-issuing the request begins a new generation from the same prompt. This is why the answer restarts. Clients must persist partial output and design idempotent retries rather than expecting continuation.

Why this answer

Streaming responses from the Messages API are not resumable: each HTTP request produces an independent generation delivered as Server-Sent Events. When the transport fails, the client cannot ask the server to continue from the last token it saw. The correct mitigation is to persist partial output locally, avoid blindly restarting the UI, and design retries that tolerate duplicate or truncated content.

Exam trap

The trap here is assuming that streaming supports resumable delivery with an event ID like some message-queue protocols, when in fact each request is a fresh, non-resumable generation.

18
MCQeasy

A developer is using the Messages API and wants Claude to reply in strict JSON matching a schema their downstream service expects. They have already written a clear instruction in the system prompt describing the schema. Which additional API feature most directly enforces that the output conforms to the schema?

A.Defining a tool whose input_schema describes the desired fields, then instructing Claude to call that tool for its response.
B.Setting temperature to 0 to eliminate all variation in the response format.
C.Increasing max_tokens so the model has room to complete the full JSON object without truncation.
D.Adding a stop_sequence equal to the closing brace so the API halts generation exactly at the end of the JSON.
AnswerA

Tools accept a JSON Schema via input_schema, and when Claude calls a tool the API returns structured input conforming to that schema. Using a tool as the response channel gives the strongest structural guarantee available in the Messages API. The system prompt describes intent, but the tool's schema is what constrains the shape of the emitted arguments, making downstream parsing reliable.

Why this answer

Tool definitions carry a JSON Schema in input_schema, and when Claude invokes a tool, the API returns arguments that conform to that schema. Routing the model's answer through a tool call therefore gives the strongest structural guarantee in the Messages API. Prompt instructions help guide behavior but do not enforce shape, so pairing them with a tool schema is the reliable pattern for strict JSON.

Exam trap

The trap here is treating system-prompt instructions or low temperature as schema enforcement, when only the tool's input_schema actually constrains the response structure.

19
Multi-Selecthard

A developer is implementing a real-time streaming interface for Claude. Which THREE event types are standard components of the Server-Sent Events (SSE) stream provided by the Messages API?

Select 3 answers
A.message_start
B.content_block_delta
C.message_stop
D.token_heartbeat
E.session_keep_alive
AnswersA, B, C

This event is the first one sent in a successful stream and contains the 'message' object with initial metadata. It provides the 'id', 'role', and 'model' information before any actual content blocks are generated, allowing the client to initialize the UI state for the incoming assistant response.

Why this answer

Understanding the lifecycle of a Claude API stream is essential for creating responsive user interfaces. The Messages API uses specific SSE event types to signal the start of the message, the incremental delivery of content, and the finalization of the response. Developers must handle these events correctly to parse the JSON data and reconstruct the full assistant message for the end user.

Exam trap

Candidates often assume 'message_complete' is a standard SSE event type, failing to realize that the API uses 'message_stop' to signal the end of the streaming sequence.

20
MCQhard

A developer sends a Messages API request that includes a system prompt and several prior turns, and Claude replies with stop_reason set to "max_tokens" even though the conversation is short. They want Claude to finish its thought rather than truncate. Which change most directly addresses the truncation?

A.Move the instructions from the system prompt into the first user turn to free output tokens.
B.Lower the temperature to 0 so the response becomes deterministic and naturally shorter.
C.Add a stop_sequence that matches the end of the expected answer so Claude knows when to stop.
D.Increase the max_tokens value in the request so Claude has more room to complete its response.
AnswerD

A stop_reason of max_tokens means generation halted because the output token ceiling was reached, not because Claude finished. Raising max_tokens gives the model additional output budget so it can complete the response, and the developer should also confirm the value does not exceed the model's remaining context window.

Why this answer

The stop_reason value max_tokens is a direct signal that the output token limit was reached before the model finished. The remedy is to raise max_tokens within the model's context constraints; stop sequences, temperature, and prompt placement do not affect how many output tokens are permitted.

Exam trap

The trap here is treating truncation as a prompt-quality or determinism problem when the stop_reason explicitly identifies the output token ceiling as the cause.

21
MCQeasy

A developer is preparing a Messages API request and wants to give Claude a persistent persona: "You are a concise legal assistant. Never give advice outside contract review." Where should this instruction be placed so it applies to the whole conversation?

A.In the system parameter, as the top-level system prompt for the request.
B.As the first user message in the messages array, prefixed with "System:".
C.Inside a metadata object attached to the request body.
D.In the stop_sequences parameter so Claude stops when it deviates from the persona.
AnswerA

The system parameter is the designated top-level field for global instructions that shape Claude's role, tone, and constraints across the entire exchange. Placing the persona there ensures it frames every user and assistant turn without being treated as a user message. This is the intended location for standing behavioral guidance.

Why this answer

The system parameter is purpose-built for top-level instructions that define Claude's role, tone, and boundaries for the whole request. Placing the legal-assistant persona there gives it consistent influence over all turns and separates standing guidance from conversational content. Other fields either carry no model-visible text or serve unrelated control functions.

Exam trap

The trap here is putting standing instructions into a user message or metadata field instead of the dedicated system parameter.

22
MCQmedium

A developer is integrating the Messages API into a backend service. They want to send a multi-turn conversation where the assistant previously produced a tool call that returned a result. Which content structure should they send back to continue the conversation correctly?

A.A system message that appends the tool output to the original system prompt so it persists for all future turns.
B.A user message whose content array contains a tool_result block referencing the tool_use id, followed by any additional user text as separate blocks.
C.An assistant message echoing the tool name and its output so the model can read its own prior decision before answering.
D.A single user message containing the raw tool output as plain text, labeled with the tool name in parentheses.
AnswerB

The Messages API expects tool outputs delivered as tool_result blocks inside a user message, each referencing the tool_use_id from the assistant's prior tool_use block. Additional user text can coexist as separate content blocks in the same message. This preserves the call-result linkage and lets Claude continue reasoning with the returned data, which is the correct continuation pattern.

Why this answer

Tool results belong in a user-turn message as tool_result content blocks, each carrying the tool_use_id from the assistant's prior tool_use block. Additional user text can be included as sibling content blocks. This structured pairing preserves the call-result relationship the API expects, letting Claude incorporate the returned data and continue the conversation coherently.

Exam trap

The trap here is treating tool output as ordinary text in a user or assistant message, when the API requires a tool_result block linked by tool_use_id.

23
MCQhard

A developer wants to implement 'pre-filling' to guide Claude's output toward a specific format. How should the 'messages' array be structured to accomplish this?

A.End the 'messages' array with a 'user' message containing the desired starting text.
B.End the 'messages' array with an 'assistant' message containing the desired starting text.
C.Include the desired starting text in the 'system' parameter with a 'prefix' label.
D.Add a 'prefill' field to the top-level API request object.
AnswerB

By ending the array with an 'assistant' message, the developer provides the initial tokens of the response. Claude will then continue generating from where that message left off. This is the standard and recommended way to steer the model's behavior and response format effectively.

Why this answer

Pre-filling is a powerful technique where the developer provides the beginning of the assistant's response. This forces Claude to continue from that point, which is highly effective for ensuring specific output formats like JSON or XML. It requires a specific message order that deviates from the typical user-only or user-assistant-user pattern.

Exam trap

Candidates often attempt to pre-fill by appending to the 'user' message or adding a new field, violating the API's structural requirement that the final role must be an 'assistant'.

24
MCQhard

A developer is using the Anthropic Messages API and wants to ensure that Claude's response is deterministic and reproducible for a given prompt. They set the temperature parameter to 0. However, they observe that repeated calls with the same input sometimes yield slightly different outputs. Which factor is the most likely cause of this non-determinism?

A.The model's internal computations involve non-deterministic operations such as floating-point arithmetic or parallel processing.
B.The API automatically injects a random seed into each request unless one is explicitly provided.
C.The top_p parameter is set to a value less than 1, introducing randomness even with temperature 0.
D.The temperature parameter is ignored when set to 0, and the default temperature is used instead.
AnswerA

Large language models like Claude run on distributed hardware and may use non-deterministic algorithms for efficiency, such as asynchronous floating-point operations or parallel reductions. Even with temperature 0, which selects the most likely token, the probability distribution itself can vary slightly due to these hardware-level nondeterminisms, leading to occasional different tokens. This is a known limitation.

Why this answer

Even with temperature set to 0, which makes the model choose the most probable next token, the underlying probability distribution can vary slightly between calls due to non-deterministic operations in the model's execution on distributed hardware. This includes floating-point arithmetic order and parallel processing. The API does not provide a seed parameter to enforce determinism.

Therefore, the most likely cause is the model's internal non-determinism, not misconfiguration of temperature or top_p.

Exam trap

The trap here is assuming that temperature 0 guarantees identical outputs, overlooking that hardware-level non-determinism can still cause slight variations.

25
MCQhard

A developer is using the Anthropic Messages API and wants to implement a retry mechanism for handling rate limit errors. They receive an HTTP 429 response with a 'retry-after' header. Which approach is the most appropriate for handling this error?

A.Immediately retry the request in a tight loop until it succeeds.
B.Wait for the number of seconds specified in the 'retry-after' header before retrying the request.
C.Switch to a different model that has higher rate limits and retry immediately.
D.Log the error and fail the request without retrying, as rate limits are permanent.
AnswerB

The 'retry-after' header provides the number of seconds to wait before making another request. Respecting this value is the correct way to handle rate limiting, as it aligns with the server's expectations and avoids overwhelming the API. This approach ensures the retry is made after the rate limit window has passed, increasing the chance of success.

Why this answer

When the API returns a 429 status with a 'retry-after' header, the best practice is to pause and wait for the specified number of seconds before retrying. This respects the server's rate limiting policy and avoids making unnecessary requests that could lead to further throttling. Immediate retries or switching models do not address the root cause.

Failing without retry is also not ideal for transient rate limits.

Exam trap

The trap here is ignoring the 'retry-after' header and either retrying immediately or giving up, rather than waiting the prescribed time.

26
MCQmedium

A developer is building a document Q&A service on the Anthropic Messages API. Their integration currently sends the entire 240,000-token knowledge base on every turn of a long conversation, and they are hitting the model's context window limit. They want to keep the full conversation history and the knowledge base available without exceeding the window. Which approach best addresses the problem?

A.Enable prompt caching on the static knowledge base so repeated prefixes are served from cache and no longer count toward the context window.
B.Increase the max_tokens parameter so Claude can internally compress the knowledge base before answering.
C.Summarize or truncate older conversation turns and retrieve only the most relevant knowledge-base chunks to include in each request.
D.Switch the request to stream: true so the server transmits the knowledge base in smaller chunks that bypass the context limit.
AnswerC

The context window is a hard cap on the total tokens sent and generated per request. Reducing the payload by condensing prior turns and injecting only relevant retrieved chunks keeps the request within the window while preserving useful information. This is the standard retrieval-plus-summarization pattern for long-running document assistants.

Why this answer

The context window is a fixed ceiling on combined input and output tokens per request. To keep a long conversation and a large corpus usable, the developer must shrink what is actually sent: condense prior turns and inject only the retrieved passages relevant to the current question. Caching and streaming change cost or delivery mechanics, not the window itself, and max_tokens governs output only.

Exam trap

The trap here is assuming prompt caching or streaming expands the usable context window rather than only changing cost, latency, or delivery mechanics.

27
MCQmedium

A developer is streaming a response from the Anthropic Messages API using server-sent events and wants to assemble the final text on the client. They observe that each event delivers a small fragment of the answer. Which event type carries the incremental text delta that must be concatenated to reconstruct the full completion?

A.content_block_stop
B.message_start
C.content_block_delta
D.ping
AnswerC

With streaming enabled, Claude emits content_block_delta events whose delta object holds the incremental text (for example a text_delta with a partial string). Concatenating these fragments in arrival order reproduces the complete message, and the final message_delta and message_stop events signal the end of the stream.

Why this answer

Streaming responses arrive as a sequence of typed server-sent events, and the actual generated text is delivered piecewise inside content_block_delta events. A client must concatenate those deltas in order and treat message_start, content_block_stop, message_delta, and message_stop as structural markers rather than content.

Exam trap

The trap here is assuming the first event of a stream (message_start) or the block boundary events already contain the answer text, when only the delta events carry incremental output.

28
MCQmedium

When using the Messages API, a developer wants to ensure that Claude does not use any tools and only provides a standard text response, even if tools are defined in the request. Which configuration should they use?

A.Omit the 'tools' parameter from the request entirely.
B.Set the 'tool_choice' parameter to {"type": "none"}.
C.Set the 'tool_choice' parameter to {"type": "auto"}.
D.Include a system prompt that says 'Do not use any tools in this conversation'.
AnswerB

Setting tool_choice to 'none' explicitly instructs Claude to ignore the tools provided in the 'tools' array. This ensures the model only generates a standard text response, which is essential when you want to use the same code base for both tool-enabled and text-only interactions with the model.

Why this answer

The tool_choice parameter provides granular control over when and how Claude uses the tools provided in the tools array. Setting it to 'none' is the definitive way to disable tool use for a specific request. This is useful for debugging or for scenarios where you want to provide tool definitions but decide dynamically to ignore them.

Exam trap

Test-takers often try to prevent tool usage by simply omitting the tools array entirely, forgetting that explicit control parameters are needed when tools are defined.

29
MCQeasy

Which object in the Claude API response body provides the exact count of tokens consumed by the prompt and the generated completion for billing and usage monitoring?

A.metadata
B.usage
C.token_details
D.consumption
AnswerB

The usage object is a top-level field in the JSON response that specifically contains input_tokens and output_tokens. This provides the definitive count of how many tokens were processed in the request and how many were generated in the response, which is the direct basis for Anthropic's usage-based pricing model.

Why this answer

Tracking token usage is essential for managing API costs and understanding the scale of data being processed. The Messages API returns a usage object that provides transparent metrics for every request. Developers use this data to implement internal billing, monitor for spikes in usage, and optimize their prompts to fit within budget and rate limit constraints.

Exam trap

Candidates often look for token counts in the top-level response object or metadata headers, missing the specific 'usage' object nested within the API response body that contains the actual metrics.

30
MCQeasy

When constructing a request for the Messages API, where should instructions that guide Claude's personality, tone, and global constraints be placed for optimal performance and architectural clarity?

A.As the first message in the messages array with the role set to 'user'.
B.In the system top-level parameter outside of the messages array.
C.As a hidden field within each individual content block of the user role.
D.Appended to the end of the last message with the role set to 'assistant'.
AnswerB

The system parameter is specifically designed to hold high-level instructions that define the model's behavior and constraints. Using this top-level field ensures that Claude treats the content as a foundational framework for the entire interaction, which improves adherence to complex rules and maintains a consistent persona throughout the dialogue.

Why this answer

The Messages API distinguishes between the conversational flow and the operational constraints of the model. Placing instructions in the system parameter allows the model to separate the 'how' of the response from the 'what' of the user query. This structural separation is a core mechanic for building robust, steerable applications that behave consistently across different sessions.

Exam trap

Candidates frequently embed persona and tone guidelines directly into the user message or conversation history instead of using the designated structural parameter.

31
MCQeasy

A developer is using the Anthropic Messages API and wants to limit the maximum number of tokens that Claude can generate in its response. Which parameter should they set in the request body?

A.token_limit
B.max_length
C.stop_sequences
D.max_tokens
AnswerD

The max_tokens parameter is a required top-level parameter in the Messages API that specifies the maximum number of tokens to generate in the response. It acts as a hard limit; the model will stop once it reaches this number or when it finishes naturally. Setting it appropriately prevents excessively long responses and controls cost. This is the correct parameter for the scenario.

Why this answer

The max_tokens parameter is specifically designed to cap the number of tokens generated in the response. It is a required field in the Messages API request. Other parameters like stop_sequences control stopping based on content, not token count.

Parameters such as max_length or token_limit are not part of the API. Therefore, setting max_tokens is the correct approach to limit response length.

Exam trap

The trap here is confusing max_tokens with stop_sequences, or assuming that other APIs' parameter names like max_length apply to the Anthropic API.

32
Multi-Selecthard

A developer is diagnosing a production Messages API workload where some requests fail with an overloaded_error and others return stop_reason "max_tokens". They want to handle both conditions correctly. Which TWO actions are appropriate? (Choose two.)

Select 2 answers
A.Implement exponential backoff with jitter and retry requests that return overloaded_error.
B.Convert overloaded_error responses into successful responses by caching the last valid completion for the same prompt.
C.Retry every request that returns stop_reason "max_tokens" with the identical parameters until it completes.
D.Lower max_tokens to 1 whenever overloaded_error occurs so the server has less work to do.
E.Treat stop_reason "max_tokens" as a truncation signal and either raise max_tokens or shorten the prompt or conversation.
AnswersA, E

overloaded_error signals transient capacity pressure on the server and is designed to be retried. Exponential backoff with jitter spreads retries so a client does not hammer the endpoint, improving the chance of success. This is the recommended handling for that error class.

Why this answer

The two conditions require different handling. Transient overloaded_error should be retried with exponential backoff and jitter because capacity pressure is temporary. A max_tokens stop reason indicates the output was truncated by the configured ceiling, so the developer must raise max_tokens or reduce input to leave room for a complete answer.

Retrying truncation unchanged, caching around an unprocessed request, or shrinking output for server load all fail to address the actual cause.

Exam trap

The trap here is conflating a transient server error with a deterministic output-limit signal and applying the same retry logic to both.

33
MCQeasy

A developer wants to reduce latency and costs for a high-traffic application that sends a large, static set of instructions in every request. Which API feature should they implement to achieve this?

A.Batch API processing.
B.Prompt Caching using cache_control.
C.Top-k sampling reduction.
D.System prompt compression.
AnswerB

Prompt Caching allows the developer to mark static content with a 'cache_control' block. When subsequent requests share the same cached prefix, Claude can skip the computation for those tokens. This reduces the time-to-first-token and provides a substantial discount on the input token costs for the cached portion.

Why this answer

Prompt Caching is a specialized feature designed to optimize performance for repetitive content. By marking specific parts of the prompt as cacheable, developers can significantly reduce the processing time for the 'prefill' phase and take advantage of lower pricing for cached tokens. This is particularly effective for large system prompts or complex context documents.

Exam trap

Candidates often try to implement custom caching layers at the application level instead of utilizing the built-in 'cache_control' feature, leading to higher latency and unnecessary token costs.

34
MCQhard

A developer's agent calls the Messages API with several custom tools defined. During testing, the response repeatedly returns stop_reason "tool_use" with a tool_use block, but the application crashes because it expects a text block. The developer wants to handle this correctly so the agent can continue. What should the application do when it receives stop_reason "tool_use"?

A.Set tool_choice to "none" and resend the request so Claude answers directly without invoking tools.
B.Resend the original request with the same parameters, because a tool_use stop reason indicates a transient server condition.
C.Ignore the tool_use block and read the accompanying text field, since Claude always includes a natural-language summary alongside tool calls.
D.Parse the tool_use block, execute the tool, then send a new request containing the assistant's tool_use turn and a user turn with the corresponding tool_result block.
AnswerD

When stop_reason is tool_use, the model has paused and is waiting for tool output. The correct loop is to execute the requested tool and return the result as a tool_result content block in a user turn, preserving the assistant's tool_use turn first. This lets the model continue reasoning with the result.

Why this answer

A tool_use stop reason means the model has deliberately halted to request external execution. The application must run the named tool with the provided input, then continue the conversation by appending the assistant's tool_use turn followed by a user turn carrying the matching tool_result block. Only after receiving that result will the model resume and produce its final answer.

Exam trap

The trap here is treating stop_reason "tool_use" as an error to retry rather than a request for the application to execute a tool and return its result.

Ready to test yourself?

Try a timed practice session using only Claude API Mechanics questions.