CCDV-F · domain
Claude API Mechanics
This domain covers the mechanics of calling Claude through the Messages API: how requests are structured, how system instructions and tools are configured, how responses report token usage, and how caching reduces cost and latency. Questions present concrete developer scenarios and ask you to choose the correct parameter, object, or feature rather than recall marketing claims.
Focused practice
Practice Claude API Mechanics questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Claude API Mechanics
Be able to assemble a correct Messages API request: system prompt for global behavior, tools plus tool_choice for tool control, and streaming or non-streaming as needed. Then read the usage object for billing. The single most important thing is knowing where each control lives in the request and response.
Placing persona, tone, and global constraints in the top-level system parameter rather than a user turn
Using tool_choice to control whether Claude calls tools or returns plain text
Reading the usage object in the response for input and output token counts
Applying prompt caching to large, static instruction prefixes to cut latency and cost
Watch out for
Common Claude API Mechanics exam traps
- ▸Putting personality and global rules inside the first user message instead of the system parameter, weakening instruction priority
- ▸Assuming defining tools forces Claude to call one, when default tool_choice still permits a text-only reply
- ▸Looking for token counts at the top level of the response instead of inside the usage object
Question index
All Claude API Mechanics questions (34)
Click any question to see the full explanation, or start a practice session above.
A team is building a retrieval-augmented assistant that sends a large document set inside the system prompt on every turn. To reduce cost, they enable prompt caching. They notice cache_read_input_tokens is high on most turns but intermittently drops to zero even though the system prompt text has not changed. Which explanation best fits this behavior?
Hard2A developer wants Claude to return a fixed set of user-profile fields extracted from free-text bios. They want the response to be valid JSON they can parse without regex. The model occasionally adds a friendly sentence before the JSON. Which Messages API feature most reliably constrains the output?
Medium3Refer to the exhibit. A developer encounters this error when trying to send a conversation history. What must the developer do to fix the 'messages' array structure?
Medium4An application needs to ensure that Claude never uses its pre-trained knowledge to answer questions, but only uses the provided context. After placing instructions in the system prompt, the developer still sees Claude occasionally using outside info. What is the most effective API-level mechanical adjustment to further constrain Claude?
Hard5A high-traffic application is frequently sending the same large set of reference documents (100,000 tokens) in the system prompt for every request. Which Claude API feature would most effectively reduce both the latency and the cost of these requests?
Hard6A developer is building a document-summarization service on the Anthropic Messages API. The service must always return a summary followed by exactly one structured JSON block containing metadata (e.g., word count, sentiment). To guarantee the model never emits any text after the JSON block, the developer adds the closing brace of the JSON structure to the 'stop_sequences' array. What is the effect of this configuration on the API response?
Medium7When using the Messages API, what happens if the 'messages' array provided in the request is empty?
Easy8In the Messages API 'messages' array, which TWO roles are currently supported for maintaining conversation history?
Medium9When configuring an API call to generate a specific JSON object, a developer adds the string '}' to the 'stop_sequences' array. What is the most likely outcome of this configuration?
Medium10A developer is building a Python application that streams responses from the Anthropic Messages API. They notice that when they set stream=True, the response object is an iterator of server-sent events. They want to extract only the incremental text chunks as they arrive, without waiting for the full message. Which event type should they filter for to get the text deltas?
Medium11An application needs to ensure that Claude stops generating text as soon as it produces a specific character sequence, such as 'END_OF_REPORT'. Which API feature should be used to implement this behavior?
Easy12Refer to the exhibit. When submitting this request via a standard HTTP client, which header is mandatory to specify the API version and ensure compatibility with the Messages API?
Medium13When a developer uses Prompt Caching, which TWO metrics are specifically returned in the 'usage' object of the API response to help track cache performance?
Medium14Refer to the exhibit. What is the expected behavior of Claude when receiving this specific API request configuration?
Hard15Refer to the exhibit. An application monitoring system captures this response from the Anthropic API. Which strategy is the most mechanically sound approach for the application to take to resolve this specific error and continue processing?
Medium16A developer is using 'top_k' to control the diversity of Claude's responses. If they set 'top_k' to 1, what is the expected behavior of the model during token generation?
Medium17A developer is streaming a long Claude response using the Messages API with stream: true. Their client code reads Server-Sent Events and appends text to the UI. Mid-stream, the connection drops and the client reconnects by re-issuing the same request from scratch. Users complain that the answer restarts from the beginning. What is the most accurate explanation of what is happening?
Medium18A developer is using the Messages API and wants Claude to reply in strict JSON matching a schema their downstream service expects. They have already written a clear instruction in the system prompt describing the schema. Which additional API feature most directly enforces that the output conforms to the schema?
Easy19A developer is implementing a real-time streaming interface for Claude. Which THREE event types are standard components of the Server-Sent Events (SSE) stream provided by the Messages API?
Hard20A developer sends a Messages API request that includes a system prompt and several prior turns, and Claude replies with stop_reason set to "max_tokens" even though the conversation is short. They want Claude to finish its thought rather than truncate. Which change most directly addresses the truncation?
Hard21A developer is preparing a Messages API request and wants to give Claude a persistent persona: "You are a concise legal assistant. Never give advice outside contract review." Where should this instruction be placed so it applies to the whole conversation?
Easy22A developer is integrating the Messages API into a backend service. They want to send a multi-turn conversation where the assistant previously produced a tool call that returned a result. Which content structure should they send back to continue the conversation correctly?
Medium23A developer wants to implement 'pre-filling' to guide Claude's output toward a specific format. How should the 'messages' array be structured to accomplish this?
Hard24A developer is using the Anthropic Messages API and wants to ensure that Claude's response is deterministic and reproducible for a given prompt. They set the temperature parameter to 0. However, they observe that repeated calls with the same input sometimes yield slightly different outputs. Which factor is the most likely cause of this non-determinism?
Hard25A developer is using the Anthropic Messages API and wants to implement a retry mechanism for handling rate limit errors. They receive an HTTP 429 response with a 'retry-after' header. Which approach is the most appropriate for handling this error?
Hard26A developer is building a document Q&A service on the Anthropic Messages API. Their integration currently sends the entire 240,000-token knowledge base on every turn of a long conversation, and they are hitting the model's context window limit. They want to keep the full conversation history and the knowledge base available without exceeding the window. Which approach best addresses the problem?
Medium27A developer is streaming a response from the Anthropic Messages API using server-sent events and wants to assemble the final text on the client. They observe that each event delivers a small fragment of the answer. Which event type carries the incremental text delta that must be concatenated to reconstruct the full completion?
Medium28When using the Messages API, a developer wants to ensure that Claude does not use any tools and only provides a standard text response, even if tools are defined in the request. Which configuration should they use?
Medium29Which object in the Claude API response body provides the exact count of tokens consumed by the prompt and the generated completion for billing and usage monitoring?
Easy30When constructing a request for the Messages API, where should instructions that guide Claude's personality, tone, and global constraints be placed for optimal performance and architectural clarity?
Easy31A developer is using the Anthropic Messages API and wants to limit the maximum number of tokens that Claude can generate in its response. Which parameter should they set in the request body?
Easy32A developer is diagnosing a production Messages API workload where some requests fail with an overloaded_error and others return stop_reason "max_tokens". They want to handle both conditions correctly. Which TWO actions are appropriate? (Choose two.)
Hard33A developer wants to reduce latency and costs for a high-traffic application that sends a large, static set of instructions in every request. Which API feature should they implement to achieve this?
Easy34A developer's agent calls the Messages API with several custom tools defined. During testing, the response repeatedly returns stop_reason "tool_use" with a tool_use block, but the application crashes because it expects a text block. The developer wants to handle this correctly so the agent can continue. What should the application do when it receives stop_reason "tool_use"?
HardOther domains
All CCDV-F exam domains
Frequently asked questions
- What does the Claude API Mechanics domain cover on the CCDV-F exam?
- Be able to assemble a correct Messages API request: system prompt for global behavior, tools plus tool_choice for tool control, and streaming or non-streaming as needed. Then read the usage object for billing. The single most important thing is knowing where each control lives in the request and response.
- How many questions are in this domain?
- This page lists all 34 Claude API Mechanics questions in the CCDV-F question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Claude API Mechanics questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.