Be able to build a prompt template that injects variables inside XML tags, add few-shot examples for a target task, and design RAG prompts that force answers to stay grounded in retrieved context. The single most important thing: structure prompts so instructions and data are unambiguous to Claude.
Start practicing
Prompt Engineering and Structured Output — choose a session length
Free · No account required
Domain overview
This domain covers how you write prompts and constrain Claude's responses on the CCAR-F exam. It tests prompt templates with variable injection, XML tag structuring, few-shot examples, chain-of-thought, and hallucination reduction in RAG pipelines. Questions are scenario-based: pick the technique or tag structure that reliably produces parseable, grounded output from Claude.
Exam objectives
Using XML tags to delimit injected variables and instructions in prompt templates
Few-shot prompting with input-output example pairs to shape Claude's task behavior
Hallucination reduction in RAG via grounding, citations, and explicit uncertainty instructions
Sandwich prompting: placing instructions before and after long context for recall
Assuming plain text labels like 'Input:' separate variables as reliably as XML tags do
Adding more few-shot examples than needed instead of improving example quality and relevance
Letting RAG answers rely on model memory instead of retrieved context, causing unsupported claims
Click any question to see the full explanation and answer options, or start a focused practice session above.
An architect is building a legal research assistant using Claude 3.5 Sonnet. To ensure the model adopts a formal, authoritative tone and strictly adheres to statutory interpretations, where is the most effective location to define this persona?
2When prompting Claude to process a large document and perform multiple tasks like summarization and sentiment analysis, what is the recommended way to separate the source text from the instructions?
3Which prompting technique involves providing Claude with a few examples of the desired input-output pairs to improve performance on a specific task?
4Refer to the exhibit. The developer observes that Claude occasionally includes conversational filler before the JSON block. What is the most reliable way to ensure the output starts immediately with the JSON object?
5A company wants to prevent Claude from answering questions outside the scope of their internal knowledge base. Which prompt engineering strategy is most effective for enforcing this constraint?
6A data scientist is adjusting generation parameters for a creative writing application. They want to ensure a wide variety of vocabulary while preventing the model from selecting highly improbable words. Which TWO parameters should be adjusted?
7Refer to the exhibit. When using variable injection in a prompt template for Claude, why is the use of XML tags preferred over simple text labels?
8When designing a prompt for high-reliability JSON extraction from unstructured medical reports, which THREE strategies are considered best practices for Anthropic models?
9A developer wants Claude to generate a list of items but stop immediately after the fifth item is completed to save on token costs. Which configuration change is most appropriate?
10In the context of structured output, what is the primary benefit of 'prefilling' the assistant's response in the API call?
11To reduce the risk of Claude hallucinating information when it doesn't know the answer, what instruction should be added to the prompt?
12Refer to the exhibit. The developer wants to ensure the summary is concise and follows a specific structure. What is the most effective way to improve this prompt?
13You are building a system that must extract entities from messy transcripts. The model often misses entities at the end of the text. What is the most likely cause and architectural fix?
14Which of the following is the recommended approach for handling PII (Personally Identifiable Information) in prompts?
15You notice that your prompt, which works perfectly with Claude 3.5 Sonnet, fails when you switch to Claude 3 Haiku. What is the most likely reason?
16Which strategy most effectively handles long documents that exceed the model's immediate focus in a single prompt?
17When designing a prompt, why is it important to define a clear persona for the model?
18Which of the following describes the 'Sandwich' prompting technique?
19Which THREE strategies are effective for reducing hallucinations in RAG (Retrieval-Augmented Generation) systems?
20Refer to the exhibit. The model output is inconsistently professional. What is the best architectural improvement?
21Which of the following is an advantage of using XML tags in Anthropic prompts?
22You are building a data extraction pipeline using Claude 3.5 Sonnet. You need the model to return a valid JSON object representing user feedback metrics. Which technique best ensures the model adheres to your specific schema?
23You are designing a system to process customer support tickets. You want to ensure Claude provides accurate summaries and extracts the priority level. Which TWO strategies should you implement to improve output reliability?
24You are using Claude to generate complex JSON configurations. The model frequently truncates the output because the response exceeds the token limit. What is the most robust architectural solution to handle this?
25Which of the following describes the 'Chain of Thought' prompting strategy, and why should it be used for complex logic tasks?
26You are building an application that extracts data from messy, unstructured text. Which THREE practices help maximize the accuracy and consistency of the extraction?
27Refer to the exhibit. You are using few-shot prompting to guide the model's extraction behavior. What is the purpose of including this specific interaction in your prompt?
28You are building an agent that uses tool calls to interact with an external API. The agent is failing to select the correct tool. What is the most effective way to improve the agent's tool selection performance?
29When designing prompts, what is the 'Persona' strategy and how does it improve performance?
30Refer to the exhibit. Your API call is returning incomplete JSON because it hits a token limit. What is the most robust architectural fix for this error?
31An application requires Claude to consistently return JSON data for a configuration parser. Which technique most effectively ensures the model adheres to a specific schema?
32You maintain a production pipeline where Claude extracts line items from scanned invoices and must return them as strict JSON conforming to a supplied schema. Intermittent outputs include a short apology sentence before the JSON object, which breaks your downstream parser. Logs show the preamble appears only when the OCR text contains contradictory totals. Which change most reliably eliminates the preamble while preserving extraction quality?
33You are building an application that uses the Anthropic API to classify customer support tickets into exactly one of five categories: Billing, Technical, Account, Feature Request, or Other. You require the output to be a JSON object with a single field "category" and a value from that list. During testing, you notice that for ambiguous tickets the model sometimes returns a different key name, such as "ticket_category", or wraps the JSON in markdown code fences, causing parsing failures. You need to enforce the schema reliably while keeping latency and cost low. Which approach is most effective?
34A team uses Claude to generate release notes from git commit messages. The output must be a JSON array of objects with 'title' and 'summary' fields. The model sometimes wraps the JSON in markdown fences or adds a conversational preamble. Which approach most reliably yields parseable JSON?
35A developer is extracting line items from invoices. The prompt includes a schema and examples, yet Claude occasionally invents field names not in the schema. Which change most directly prevents hallucinated field names?
36You are using the Anthropic Claude API in a customer-support application. Your prompt asks Claude to classify each incoming ticket into exactly one of four categories and return the result. Occasionally, Claude adds a friendly sentence before the category, which breaks your parser. Which change most directly prevents this extra text?
37An architect is designing a prompt that asks Claude to analyze a contract and return a risk rating plus supporting clauses. The response must be machine-readable and include a confidence score. Which design most reliably produces both a rating and a confidence score in a single response?
38You are designing a prompt for Claude to convert a customer email into a JSON object with fields "customer_name", "order_id", and "issue_summary". During testing, about 15% of outputs include the JSON inside markdown code fences (```json ... ```) or add a friendly sentence before the JSON. You need the raw JSON object only, with no extra text, every time. Which approach is most effective?
39A developer is building a contract-analysis pipeline. The prompt instructs Claude to return JSON with fields 'party_a', 'party_b', and 'effective_date'. In testing, Claude sometimes wraps the JSON in a Markdown code fence or adds a trailing comment, causing json.loads to fail. Which approach best ensures the response can be parsed directly as JSON?
40A developer is building a structured output pipeline where Claude must return a list of support tickets, each with 'id', 'priority', and 'category'. The pipeline must be robust against malformed responses. Which TWO practices most directly improve the reliability of this structured output? (Choose two.)
41A team is using the Anthropic Messages API to extract line items from vendor invoices. They define a tool named `record_line_items` with an input_schema that requires `description`, `quantity`, and `unit_price`. In production, Claude sometimes returns a text message describing the items instead of calling the tool, especially on invoices with unusual layouts. They want to guarantee a tool call every time. Which change is most effective?
42You are iterating on a prompt that summarizes technical articles. The summaries are sometimes too long and sometimes miss the required 'key_takeaways' section. You want to diagnose whether the problem is the instruction wording or the input content. Which practice best supports systematic prompt iteration?
43A financial analyst wants Claude to extract the total amount due from each invoice in a batch of 300 PDF-extracted text blobs, returning results as JSON with fields "invoice_id", "amount_due", and "currency". Some invoices list multiple line items and a subtotal, while others show only a single grand total. Which prompt design most reliably produces consistent, machine-parseable extraction across all invoices?
44A financial analyst uses Claude via the Messages API to extract line items from scanned invoices. The response must be valid JSON conforming to a fixed schema with fields invoice_number, date, and total. During testing, the model occasionally adds a trailing sentence after the JSON object or wraps the JSON in a markdown code fence. Which approach best guarantees the response parses cleanly while preserving extraction accuracy?
45A developer is building a classification pipeline that routes support tickets into one of eight categories. They want Claude to return a category label and a confidence score between 0 and 1. The output must be parseable by a JSON parser without post-processing. They plan to use the Anthropic Messages API. Which approach best ensures the output is directly parseable JSON with the required fields?
46A team is extracting medication names and dosages from clinical notes. The notes contain many similar-looking terms, and the team needs high precision to avoid false positives. They plan to use few-shot examples. Which design of the few-shot examples best supports high precision in this scenario?
47A platform team is building a Claude-powered service that returns structured records to a downstream database. During testing, roughly 4% of responses include a friendly sentence before the JSON, and another 3% omit a required field entirely. Which TWO changes most directly reduce these failure modes? (Choose two.)
48A support-engineering team is prototyping a Claude-powered assistant that must always reply in a consistent tone and never invent refund policies. They want a single, reusable instruction block that applies to every conversation regardless of the user's message. Where should this persistent behavioral guidance be placed in a Messages API request?
49You are building a document-processing pipeline with the Anthropic Messages API. Claude must extract a fixed set of fields from each document and return them as JSON. In production, you observe two problems: (1) some responses include explanatory prose before the JSON, and (2) some responses omit required fields when the document is ambiguous. You need to make the pipeline more reliable. (Choose two.)
50You are designing a prompt that asks Claude to review a pull request and return findings in a strict XML structure with a severity attribute for each finding. Your parser depends on the structure being consistent. Which TWO techniques most directly improve the consistency of the XML output? (Choose two.)
51A platform team ships a Claude-based service that classifies incoming tickets into one of eight categories and returns a JSON object with category and confidence. In production they observe two failure modes: the model sometimes returns a category outside the eight allowed values, and occasionally returns confidence as the string "high" instead of a number. Which TWO changes most directly reduce these failures while keeping the pipeline automated? (Choose two.)
52A developer is prompting Claude to classify 5,000 customer support tickets into one of six fixed categories. The first 200 classifications look correct, but the developer notices that category labels come back with inconsistent capitalization and occasional trailing punctuation. What is the simplest change that makes the labels uniform and directly usable as database keys?
53A developer wants Claude to summarize a long support thread and return the result as a JSON object with a single field "summary". They want to minimize the chance that Claude adds commentary outside the JSON. Which is the most direct Anthropic-recommended technique to apply?
54A data-engineering team uses Claude to normalize product descriptions into a fixed set of attributes. Their few-shot examples all come from the electronics catalog, and accuracy is high there, but when the same prompt is applied to apparel listings the model mislabels attributes. They want to improve apparel accuracy without degrading electronics results. Which approach best addresses the problem?
55A team is using Claude to generate a structured incident report from raw log excerpts. The report must contain a 'summary' paragraph, a 'root_cause' field, and a 'remediation' list. During review, engineers find that remediation steps sometimes appear inside the summary paragraph, and summaries sometimes contain root-cause speculation. Which technique most reliably keeps each piece of content in its designated section?
56A developer builds a Claude-powered assistant that answers questions over a 300-page policy manual. When the whole manual is placed in a single prompt, answers to questions about clauses in the middle of the document are frequently wrong or vague, while answers about the opening and closing sections are accurate. The manual must stay available in full, and the team wants to keep latency reasonable. Which technique most effectively improves accuracy for the mid-document clauses?
57An engineering team is building a Claude-powered contract analyzer that must return a JSON object with a 'parties' array, an 'effective_date' string, and a 'governing_law' string. Legal reviewers report that for about one in twenty contracts, the model invents a governing law when the contract is silent. The team wants to stop fabricated values without losing valid extractions. Which approach best addresses this?
Be able to build a prompt template that injects variables inside XML tags, add few-shot examples for a target task, and design RAG prompts that force answers to stay grounded in retrieved context. The single most important thing: structure prompts so instructions and data are unambiguous to Claude.
The Courseiva CCAR-F question bank contains 57 questions in the Prompt Engineering and Structured Output domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Prompt Engineering and Structured Output domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included