Anthropic · Free Practice Questions · Last reviewed May 2026
30real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
An application uses Claude 3.5 Sonnet to summarize legal documents. Occasionally, the model hallucinates clauses not present in the source text. What is the most effective architectural approach to ground the model's output?
Increase the temperature setting to 1.0 to ensure the model explores more creative possibilities.
Fine-tune the model on the full legal corpus to embed the documents directly into its weights.
Use RAG to inject the relevant document snippets into the system prompt and instruct the model to only use that data.
RAG provides the specific, authoritative source material directly within the context window for every request. By instructing the model to rely exclusively on this provided context and return 'I don't know' if the information is missing, you effectively minimize the model's propensity to generate ungrounded, hallucinated legal clauses.
Reduce the maximum tokens allowed for the response to prevent the model from generating extra text.
When designing a multi-turn conversation, which practice best maintains context reliability over long interactions?
Always send the entire conversation history, regardless of length.
Periodically summarize the interaction history and include the summary in the next prompt.
Summarizing interaction history keeps the context window focused and relevant. It provides a condensed, accurate representation of previous turns, which prevents the model from losing the thread of the conversation or getting distracted by outdated information, thereby significantly improving the reliability of the model in long, multi-turn interactions.
Use the system prompt to store all variables to save space.
Randomly drop older turns to manage context size.
Which of the following is the most important factor in maintaining reliability when using Claude?
Having the largest possible context window.
A clear, concise, and specific system prompt.
The system prompt is the fundamental blueprint for model behavior. Clear, specific instructions regarding role, task, and constraints provide the necessary guardrails for consistent performance. Without a high-quality system prompt, the model lacks the guidance needed to consistently deliver reliable results, regardless of its underlying capabilities or context window.
Always enabling verbose logging for debugging.
Setting the temperature to exactly 0.5.
You are building an RAG system. Which THREE factors most significantly impact the reliability of the retrieved information?
Chunking strategy (size and overlap).
Effective chunking ensures that relevant context is captured within the window limits without losing critical cross-reference information. Proper overlap is essential for maintaining continuity between chunks. Poor chunking leads to fragmented information, causing the model to miss key details, which directly undermines the reliability of the RAG system.
The model's temperature setting.
Embedding model quality.
The embedding model determines how accurately the system understands the semantic meaning of both the user query and the stored documents. If the embedding is inaccurate, the retrieval process will pull irrelevant documents, resulting in a model response that is hallucinated or logically disconnected from the actual user intent.
Relevance and accuracy of the vector search results.
The retrieval step is the gatekeeper of RAG performance. If the search results are irrelevant, the model is provided with 'garbage' context, which inevitably leads to inaccurate or hallucinated answers. Ensuring high-precision search results is the most critical step in maintaining the factual reliability of any RAG-based LLM application.
The language of the user query.
Your model is consistently ignoring specific safety formatting rules during long multi-turn conversations. What is the most robust architectural solution?
Increase the frequency of full-conversation restarts.
Include a system instruction that explicitly defines the formatting rules and use few-shot examples.
Few-shot examples provide concrete, unambiguous demonstrations of the expected behavior, which are much harder for the model to ignore than abstract instructions alone. By combining clear formatting rules with concrete examples in the system prompt, you create a robust anchor that keeps the model compliant across long-running turns.
Switch to a larger, more 'intelligent' model family.
Send the formatting rules as a user message at every turn.
Refer to the exhibit. The model output is a list, but it also includes introductory text like 'Here is the summary you requested:'. How can you ensure the output contains ONLY the list?
Set the temperature to 0.
Add a negative constraint: 'Do not include conversational filler or introductory text.'
This is the most direct and effective way to control the model's behavior. By explicitly banning conversational filler, you provide a clear boundary that the model can follow. This architectural pattern is essential for data-driven applications that require clean, predictable outputs from the model without any extraneous conversational noise.
Truncate the output in the application code.
Use a regex filter on the input document.
Want more Context and Reliability practice?
Practice this domainAn architect is building a legal research assistant using Claude 3.5 Sonnet. To ensure the model adopts a formal, authoritative tone and strictly adheres to statutory interpretations, where is the most effective location to define this persona?
At the end of the user message using a specific 'tone' tag.
Within the system prompt field.
System prompts are specifically designed to provide high-level instructions and context that govern the model's behavior throughout the interaction. By establishing the legal researcher persona here, Claude internalizes the constraints more effectively, resulting in higher fidelity to the desired authoritative tone and statutory focus.
Inside a JSON schema definition for the output.
By setting a specific stop sequence for informal words.
When prompting Claude to process a large document and perform multiple tasks like summarization and sentiment analysis, what is the recommended way to separate the source text from the instructions?
Using Markdown headers like #Source and #Task.
Enclosing the source text in XML tags like <document>.
Anthropic models are highly optimized to parse and respect XML tags as delimiters. Using tags like <document> and <instructions> creates a clear hierarchy, allowing Claude to distinguish between the data it must process and the commands it must follow, which significantly improves overall task accuracy.
Separating sections with several empty newline characters.
Indenting the source text by four spaces.
Which prompting technique involves providing Claude with a few examples of the desired input-output pairs to improve performance on a specific task?
Zero-shot prompting.
Chain of Thought prompting.
Few-shot prompting.
Few-shot prompting specifically refers to the practice of including a small number of examples in the prompt to demonstrate the desired behavior. This technique is highly effective for teaching Claude new formats or specific stylistic requirements that might be difficult to describe with instructions alone.
Role prompting.
Refer to the exhibit. The developer observes that Claude occasionally includes conversational filler before the JSON block. What is the most reliable way to ensure the output starts immediately with the JSON object?
Include a strict JSON schema in the system prompt.
Prefill the assistant response with the '{' character.
Starting the assistant's message with an opening curly brace forces the model to complete the JSON object immediately. This technical maneuver bypasses the model's tendency to explain its actions, ensuring that the very first character of the generated response is part of the structured data payload.
Set the temperature to 0.0 for deterministic output.
Use a higher frequency penalty to avoid repetitions.
A company wants to prevent Claude from answering questions outside the scope of their internal knowledge base. Which prompt engineering strategy is most effective for enforcing this constraint?
Increase the top-p value to allow more diverse answers.
Use few-shot examples of Claude answering general knowledge.
Include explicit 'stay in character' instructions in the user message.
Set strict boundary instructions in the system prompt.
The system prompt is the ideal location for establishing 'guardrails' and boundary conditions. By instructing the model to only use the provided context and to state 'I do not have information on that' for other topics, you create a robust mechanism for ensuring the model stays on-task.
A data scientist is adjusting generation parameters for a creative writing application. They want to ensure a wide variety of vocabulary while preventing the model from selecting highly improbable words. Which TWO parameters should be adjusted?
Temperature.
Temperature controls the 'randomness' of the output by scaling the probability distribution of the next token. A higher temperature increases variety by making less likely tokens more probable, which is essential for creative writing, although it must be balanced to maintain the logical flow of the narrative.
Max Tokens.
Top-P (Nucleus Sampling).
Top-P restricts the model to the smallest set of tokens whose cumulative probability exceeds a certain threshold. By adjusting this, the data scientist can ensure the model only considers reasonably likely words, preventing the 'hallucination' of nonsensical vocabulary that sometimes occurs with high temperature alone.
Stop Sequences.
Presence Penalty.
Want more Prompt Engineering and Structured Output practice?
Practice this domainWhich component is responsible for executing the logic associated with an MCP tool?
The Anthropic model itself.
The MCP Client application.
The MCP Server.
The MCP server is specifically designed to host tools, manage their lifecycle, and execute the backend logic when invoked. This architectural separation ensures that tools remain modular, secure, and easily maintainable. By keeping execution on the server, you protect sensitive backend resources from direct access by the client.
The JSON Schema definition.
An architect is concerned about prompt injection attacks targeting an MCP tool. Which strategy provides the strongest defense?
Implement a natural language filter to detect malicious intent before tool invocation.
Use strict JSON Schema validation and server-side parameter sanitization.
Strict schema validation acts as a structural firewall, while parameter sanitization ensures that any dynamic data is safely handled. This combined approach is the industry-standard defense against injection. It prevents the model from injecting unintended characters or commands into the tool's backend execution logic, maintaining overall system integrity.
Ask the model to verify if the user's prompt is malicious before executing.
Disable all tool execution if the prompt contains special characters.
When designing an MCP server, which TWO of these factors primarily influence the model's performance in selecting the right tool?
The total amount of system memory allocated to the MCP server.
The semantic quality and uniqueness of the 'name' field.
The 'name' field is the primary identifier the model uses to categorize and evaluate the tool's utility. Unique, semantically meaningful names allow the model to distinguish between similar operations, significantly improving the precision of the selection process. This is a critical design element for maintaining high tool-use accuracy.
The clarity and verbosity of the 'description' field.
The 'description' field provides the context the model needs to understand when and how to use the tool. Highly descriptive text that includes usage examples or boundary conditions allows the model to reason effectively about tool applicability. Clear documentation is the most impactful factor in ensuring correct tool selection.
The character count of the internal code implementation.
The version number of the MCP server protocol.
An architect needs to implement a tool that allows the LLM to search a file system. What is the most secure way to present this capability?
Allow the model to provide any absolute path as an argument.
Implement a 'root_directory' parameter that the model can change dynamically.
Design the tool to accept relative paths within a predefined server-side root.
This approach uses a server-side enforced root, which prevents the LLM from accessing files outside the designated workspace. By resolving paths against a fixed base, you eliminate directory traversal risks. This design pattern ensures the tool is useful while maintaining a strict, secure boundary against unauthorized file system access.
Provide a tool that lets the model browse all directories starting from '/'
How should an MCP server communicate a non-recoverable error to the client?
Return a 500 Internal Server Error status code via the transport layer.
Send a JSON error object that adheres to the MCP Error specification.
The MCP protocol defines specific error objects for reporting issues during tool execution. Using this structure ensures that the client interprets the error correctly, allowing for programmatic handling of the failure. This is the correct, standard way to communicate issues while maintaining protocol compliance and system-wide interoperability.
Return a string message indicating the failure as the result.
Close the connection abruptly to signal a critical failure.
When designing tools for a highly regulated environment, what is the most important architectural goal?
Maximizing the model's autonomy to perform complex, multi-step tasks.
Ensuring the model can access as much data as possible to be effective.
Providing complete observability and strict access controls for every tool call.
Observability and strict access controls are the cornerstones of compliant architecture. By logging every call and enforcing precise authorization, the system provides the proof of compliance required by auditors. This setup allows the organization to leverage LLM capabilities while maintaining the security posture demanded by strictly regulated operational environments.
Optimizing for the lowest possible latency in every tool interaction.
Want more Tool Design and MCP Integration practice?
Practice this domainDuring an interactive session, Claude Code identifies a bug and proposes a fix that requires executing a shell command to install a new dependency. How does the default permission model handle this request?
Claude Code automatically executes the command if it is deemed safe by the internal classifier.
The tool prompts the user for manual approval before executing any shell command.
Manual approval is the standard security gate for shell execution in Claude Code. This workflow ensures that developers can review the exact command, understand its implications, and verify its correctness before it runs. It strikes a balance between agentic productivity and the necessity of human-in-the-loop security verification.
The command is blocked unless the user has pre-authorized the specific package manager in config.
Claude Code executes the command silently and only reports the output if an error occurs.
Which command is used to start a new Claude Code session in the current directory and begin indexing the local files?
claude init
claude start
claude
Running the 'claude' command in the terminal launches the interactive agent. It immediately checks the current directory for a repository, reads any configuration files, and prepares the indexing service so that the model can answer questions about the code right away. This simplicity is a core feature of the tool's design.
claude session --new
Refer to the exhibit. A developer sees this error log in the terminal output. What is the most likely cause of this failure in the Claude Code workflow?
The developer does not have write permissions for the file 'src/auth.ts'.
The model is attempting to edit a file that has been excluded by .claudeignore.
The model's context for 'src/auth.ts' is outdated or incorrect.
The error 'old_string not found' occurs when the model tries to replace a block of code that doesn't exactly match what is on the disk. This happens if the file changed since Claude last read it. The correct workflow is for the model to re-read the file to synchronize its state.
The 'edit_file' tool has been disabled in the project's .claude.json configuration.
How does Claude Code handle large repositories that exceed the standard context window size during initial indexing?
It uploads the entire repository to a vector database and uses RAG for every query.
It truncates the repository and only indexes the first 500 files detected.
It creates a local index of file names and symbols to enable targeted file reading.
Claude Code builds a local index that helps it navigate the codebase efficiently. When a user asks a question, the tool can search this index to identify which files are most likely to contain the answer. It then uses its 'read_file' tool to bring only the necessary code into the context.
It requires the user to manually specify which directories to index using the /index command.
What is the primary purpose of the 'compact' workflow in Claude Code configuration?
To compress the source code on disk to save storage space.
To reduce the token count of the conversation history while retaining key context.
Compaction is a context-management technique. By summarizing the dialogue and results achieved so far, the tool can 'forget' the verbose details while 'remembering' the important outcomes. This allows for much longer sessions without hitting the hard limits of the model's context window.
To minify JavaScript and CSS files automatically before deployment.
To encrypt the session logs before they are sent to Anthropic's servers.
A developer wants to use Claude Code to perform a large refactor across multiple files. What is the most efficient workflow to ensure the model has the necessary context without manual file-by-file reading?
Manually run 'cat' on every file in the project and paste the output into the prompt.
Use the /read-all command to force Claude to load every file into memory at once.
Provide a high-level description of the refactor and let the model use its search and read tools.
The most effective way to work with Claude Code is to leverage its agentic nature. By describing the goal (e.g., 'Refactor the Auth module to use the new DB schema'), the model can use its search tools to find relevant files and then selectively read and edit them as needed.
Copy the entire project into a single .txt file and upload it using the /upload command.
Want more Claude Code Configuration and Workflows practice?
Practice this domainRefer to the exhibit. The provided JSON represents a partial response from Claude. According to the Anthropic API specification for tool use, what is the mandatory next step for the orchestrator to continue the agentic loop?
Send a new user message asking the model to ignore the previous tool request.
Submit a message with the role 'user' containing a 'tool_result' block matching the tool_use id.
The orchestrator must execute the 'query_inventory' function locally and then return the data to Claude. This is done by adding a new message to the history where the role is 'user' and the content contains a 'tool_result' type with the matching 'tool_use_id'.
Call the 'messages' endpoint again with the same prompt to see if Claude changes its mind.
Update the system prompt to include the inventory data discovered by the tool.
A company is deploying an agent that interacts with a sensitive financial database. To maintain security and accuracy, the architect wants to ensure that any 'delete' or 'transfer' actions are reviewed by a human before execution. Which agentic pattern should be used?
Parallel Tool Execution to allow the model to run multiple checks simultaneously.
The ReAct pattern to allow the model to think before it acts on the database.
An 'Interrupt-and-Approve' orchestration loop for specific tool definitions.
This pattern involves the application logic intercepting specific tool calls identified as 'sensitive.' The orchestrator pauses the agent's progress, presents the proposed tool arguments to a human user, and only proceeds to send the 'tool_result' back to Claude once the human has provided explicit approval.
Increasing the frequency of system prompt injections to remind the model of the rules.
When designing a multi-agent system, what is the primary benefit of the 'Orchestrator-Worker' pattern compared to a single-agent architecture with 20 tools?
It eliminates the need for a system prompt for the worker agents.
It reduces the complexity of the orchestrator's context window by offloading details to workers.
By delegating tasks, the orchestrator only needs to know which worker to call and the final result they provide. This keeps the orchestrator's context window clean of the minute details and intermediate tool outputs used by the workers, reducing noise and the risk of the model getting distracted.
It allows the model to bypass the per-token pricing of the Anthropic API.
It guarantees that the model will never produce a hallucinated tool name.
An agent tasked with data analysis repeatedly fails because the 'tool_result' it receives is too large for the context window, causing subsequent calls to truncate. What is the most effective architectural adjustment to resolve this?
Switch to a model with a smaller context window to force more efficient coding.
Use the 'tool_result' to return a summary or a file path instead of the full raw data.
Returning a summary allows Claude to understand the key characteristics of the data without overwhelming its context window. If the model needs specific details, it can then use a different tool to query a specific subset of the data, maintaining a high signal-to-noise ratio in the conversation.
Format the tool output as a series of multiple user messages to spread the load.
Request that Claude only uses the last 100 lines of the tool output in its reasoning.
An orchestrator is designed to handle 'Parallel Tool Use' for a travel booking agent. Which TWO conditions must be met for the architect to successfully implement this feature?
The tools being called must not have sequential dependencies on each other's outputs.
Parallel execution is only possible if the model can generate all necessary arguments at once. If Tool B requires an ID generated by Tool A, they must be called sequentially across two different turns of the agentic loop, as the model cannot know the result of Tool A beforehand.
The model must be instructed to use a single 'tool_result' block for all outputs.
The orchestrator must be capable of processing the model's response which contains multiple 'tool_use' blocks.
When parallel tool use is enabled, Claude may return an array of content blocks where several items have the 'tool_use' type. The orchestrator's logic must be able to iterate through this array, execute all requested tools (potentially in parallel), and collect all results for the next turn.
The 'temperature' must be set to 0 to ensure the tools are called in alphabetical order.
A human must manually approve the parallel execution via a dedicated CLI tool.
In the context of agentic orchestration, what is the primary purpose of a 'ReAct' (Reason-Act) loop?
To allow the model to rewrite its own system prompt dynamically during execution.
To encourage the model to verbalize its internal reasoning before executing a tool.
By forcing Claude to 'think' out loud, the ReAct pattern leverages the model's autoregressive nature. The generated 'Thought' tokens act as a form of working memory, helping the model stay on track and reducing errors by ensuring the reasoning logic is established before the tool call is generated.
To reduce the latency of the API response by skipping the output of text blocks.
To provide a way for the model to access historical data that was not included in the context window.
Want more Agentic Architecture and Orchestration practice?
Practice this domainThe CCAR-F exam has 60–90 questions and must be completed in 120 minutes. The passing score is 700/1000.
Scenario-based questions covering exam objectives with detailed answer explanations.
The exam covers 5 domains: Context and Reliability, Prompt Engineering and Structured Output, Tool Design and MCP Integration, Claude Code Configuration and Workflows, Agentic Architecture and Orchestration. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official Anthropic CCAR-F exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.