Courseiva

CCNA Claude Model Fundamentals Questions

67 questions · Claude Model Fundamentals · All types, answers revealed

1
MCQeasy

Which term describes the fundamental unit of text that Claude processes, which can be a single character, a part of a word, or a whole word?

A.Bit
B.Token
C.Sentence
D.Prompt
AnswerB

Tokens are the basic units of text for Claude. A token can represent a common word, a part of a word, or even just punctuation. Anthropic's models use a tokenizer to convert raw text into these numerical tokens, which the model then uses for inference and generation.

Why this answer

Understanding tokens is essential for grasping how LLMs like Claude process information and how costs are calculated. Tokens are the 'atomic' units of the model's vocabulary, and the efficiency of the tokenizer directly impacts the model's context window usage and pricing.

Exam trap

Candidates often confuse tokens with words or characters, forgetting that a token is an arbitrary sub-word, character, or whole word chunk defined by the model's specific tokenizer.

2
MCQmedium

Refer to the exhibit. In the context of Claude Model Fundamentals, what is the primary purpose of the 'system' field shown in this API configuration?

A.To provide the actual question the user wants to ask.
B.To specify the hardware resources allocated to the request.
C.To set the behavior, persona, and constraints for the model.
D.To bypass the safety filters for sensitive queries.
AnswerC

The 'system' prompt allows developers to define a specific persona or set of rules that Claude must follow. In the exhibit, it is used to constrain the model's output format to JSON and remove conversational filler, providing a consistent framework for the model's responses.

Why this answer

The system prompt is a powerful tool for defining the model's personality, constraints, and output format. It acts as a set of 'guardrails' or 'operating instructions' that the model follows throughout the entire session, ensuring consistent behavior across multiple user turns in a conversation.

Exam trap

Candidates often mistake the system field for a conversational history container or a memory cache, missing that its primary purpose is establishing permanent behavioral constraints and personas.

3
MCQmedium

A developer is building a RAG application and notices Claude 3.5 Sonnet occasionally hallucinates when provided with a large context window. Which architectural adjustment is most effective for improving factual fidelity?

A.Increase the temperature parameter to 1.0 to allow for more creative exploration.
B.Switch to Claude 3 Haiku to benefit from its faster token processing speeds.
C.Implement a semantic reranking step to filter out low-relevance chunks before prompt insertion.
D.Remove the system prompt to allow the model to operate without behavioral constraints.
AnswerC

Semantic reranking significantly improves RAG performance by ensuring that only the most contextually relevant chunks are fed into the prompt. This reduces the cognitive load on Claude, allowing it to focus on synthesized information rather than filtering through potentially irrelevant or misleading document snippets.

Why this answer

Improving factual fidelity in RAG applications relies on reducing the noise-to-signal ratio within the context window. By implementing a reranking step, the developer ensures that only the most relevant document chunks reach the model. Claude performs significantly better when the prompt is constrained to highly pertinent information, as this minimizes the risk of the model prioritizing distractor content or irrelevant retrieval artifacts during the generation phase.

Exam trap

Candidates often suggest increasing the model's temperature or adding more context, failing to realize that excessive, irrelevant information actually increases the likelihood of hallucinations in RAG systems.

4
MCQhard

Which THREE factors are primary considerations when calculating the cost of using the Anthropic API in a production environment?

A.The total number of input tokens sent per request.
B.The number of concurrent API users active.
C.The total number of output tokens generated.
D.The specific model tier (e.g., Haiku vs. Sonnet vs. Opus).
E.The latency of the model response time.
AnswerA, C, D

Input tokens constitute a major portion of the cost per request. Every character and system instruction counts toward this limit, so optimizing prompt length, system instructions, and history management is vital for maintaining a predictable budget as the application scales to handle a larger number of user requests.

Why this answer

Cost management in production relies on understanding the input token count, output token count, and the specific model selected. Since pricing is transparently based on these variables, optimizing the prompt length and response size is the most direct way to control expenditure. Failing to account for these three variables can lead to significant cost spikes as traffic scales or if prompts become excessively verbose over time.

Exam trap

Candidates often overlook the 'output tokens' cost, focusing only on 'input tokens', or they forget that different model tiers (Haiku vs Opus) have vastly different price points.

5
MCQhard

An application requires Claude to analyze a 180,000-token legal document and find a specific clause. Why might Claude 3 Opus be a better choice for this task than a smaller model like Haiku, even though both have a 200,000-token window?

A.Haiku's context window is only 200,000 tokens for output, not input.
B.Opus has higher 'Needle In A Haystack' recall at large context sizes.
C.Opus can process 200,000 tokens in less than one second.
D.Smaller models automatically truncate inputs over 100,000 tokens.
AnswerB

As the context window fills up, smaller models can sometimes lose 'focus' on information buried in the middle of the text. Opus is engineered to maintain near-perfect recall across the full 200,000 tokens, making it more reliable for finding specific details in very large legal documents.

Why this answer

Context window size is only one factor in model performance. For very large inputs, the model's ability to maintain focus and accurately retrieve information from the middle of the text (recall) is critical. Higher-tier models like Opus are specifically optimized for better performance at these extreme context limits.

Exam trap

Candidates assume that sharing the same maximum token window means smaller models perform retrieval tasks identically to flagship models, ignoring retrieval accuracy disparities.

6
MCQmedium

Which of the following is an effective technique for reducing hallucination in Claude when answering fact-based questions?

A.Instruct the model to always provide an answer, even if unsure.
B.Include a system instruction that explicitly allows the model to use external knowledge.
C.Provide the context and require the model to answer 'I don't know' if the answer is missing.
D.Run the same prompt three times and take the most frequent answer.
AnswerC

This technique forces the model to treat the context as the sole source of truth. By explicitly providing a fallback strategy for missing information, you prevent the model from attempting to invent an answer, which is the most effective way to ensure factual reliability and reduce potential hallucinations.

Why this answer

Grounding is the primary method to combat hallucination. By providing the model with a 'grounding document' and explicitly instructing it to state 'I don't know' if the answer isn't present in the provided text, you shift the model's behavior from generative to extractive. This pattern is essential for high-stakes enterprise applications, where accuracy and honesty about limitations are more valuable than a guess, protecting the application's integrity and user trust.

Exam trap

Candidates often mistakenly choose 'increasing the temperature' or 'adding more examples', which can actually increase hallucination risk instead of using grounding techniques to force honesty.

7
MCQmedium

A developer wants Claude to act as a specialized technical support assistant for a specific product. Which component of the API call is most appropriate for defining the assistant's persona and product boundaries?

A.The 'user' message role.
B.The 'system' prompt field.
C.The 'assistant' message role in the first turn.
D.The 'max_tokens' configuration parameter.
AnswerB

The 'system' prompt is the correct and intended place for defining the model's persona, constraints, and instructions. It separates the behavioral instructions from the dynamic user input, ensuring the model adheres to its assigned role and product scope throughout the entire session, regardless of the conversation's depth.

Why this answer

The system prompt is the designated field for defining the assistant's persona, expertise, and operational boundaries. By clearly articulating these in the system prompt, you ensure consistency throughout the conversation. This is vital for maintaining a professional, reliable brand presence in AI-powered customer support, as it keeps the model grounded in its designated role and prevents it from straying into off-topic or unauthorized assistance.

Exam trap

Test-takers sometimes try to define persistent personas within the user message, failing to use the dedicated top-level parameter.

8
MCQmedium

A developer needs to ensure that Claude 3.5 Sonnet consistently follows a specific JSON schema for structured data extraction. Which implementation strategy provides the highest level of deterministic output format control?

A.Instructing the model to output valid JSON within a system prompt.
B.Utilizing Few-Shot prompting with XML tags.
C.Defining a structured schema using the tool_use parameter.
D.Appending a post-processing script to parse the output as JSON.
AnswerC

The tool use feature allows developers to specify an exact JSON schema that the model must satisfy. By leveraging this, the Anthropic API forces the model to generate content that conforms to the defined input schema, significantly reducing parsing failures and ensuring high-quality, structured data extraction for integration.

Why this answer

Using tool use (function calling) with a defined schema is the industry-standard approach for structured data extraction with Anthropic models. By defining specific JSON schemas in the tools parameter, the model is constrained to generate valid, parsable objects that match the expected structure. This method minimizes hallucinations compared to prompt-based formatting, ensuring downstream systems can reliably consume the output without manual parsing errors or type mismatches.

Exam trap

Candidates often suggest prompt engineering techniques like 'asking for JSON output', which is non-deterministic, instead of utilizing the 'tool_use' parameter for robust, schema-enforced output structure.

9
MCQeasy

What is the primary function of the 'stop_sequences' parameter in the Claude API?

A.To specify the maximum number of tokens the model is permitted to generate.
B.To define the delimiters that signal the end of a model response.
C.To filter out harmful or restricted content from the model output.
D.To reset the model state to its original pre-trained condition.
AnswerB

Stop sequences explicitly define strings that trigger an immediate cessation of the generation process. By setting these, developers ensure that the model stops at logical points, which is essential for structured data extraction or ensuring the model does not continue generating after finishing a specific task.

Why this answer

The stop_sequences parameter allows developers to define specific strings that, when encountered by the model, force it to cease generation immediately. This is crucial for controlling output length and structure, particularly when integrating Claude into automated workflows where the model needs to stop at specific boundaries, such as a closing delimiter or a specific marker, preventing unnecessary token usage and ensuring predictable program execution.

Exam trap

Candidates frequently mistake stop_sequences for a tool to manage model behavior or tone, rather than recognizing it as a strictly functional mechanism for terminating generation at specific character boundaries.

10
MCQeasy

A developer is choosing a Claude model for a high-volume classification task that must meet a strict monthly budget. The task is straightforward and does not require deep reasoning. Which selection principle best fits this scenario?

A.Choose a model based on its maximum context window size, since a larger window guarantees lower cost per request.
B.Choose a model randomly and rely on prompt engineering to compensate for any capability gaps.
C.Choose the largest, most capable Claude model available to maximize classification accuracy regardless of cost.
D.Choose the smallest, fastest Claude model that meets the accuracy bar, then validate with a sample of real inputs.
AnswerD

Matching model capability to task complexity keeps costs low while still satisfying accuracy needs. A small, fast model is typically cheapest per token and sufficient for straightforward classification. Validating against real inputs confirms the choice before committing, which is the disciplined way to balance budget against quality.

Why this answer

Cost-effective model selection starts by matching capability to task difficulty. A simple classification task rarely needs the most capable model, so the smallest model that clears the accuracy bar is the right default, validated against representative inputs. Choosing on context window or at random ignores the actual drivers of cost and quality.

Exam trap

The trap here is equating more capable models with better outcomes on every task, when simple workloads are usually served more economically by smaller models.

11
MCQhard

Refer to the exhibit. A company is building a translation service for live subtitles where latency must be under 200ms per segment. Which Claude 3 model family member is the only viable candidate based on the provided performance and cost data?

A.Claude 3 Opus
B.Claude 3 Sonnet
C.Claude 3 Haiku
D.All Claude 3 models are equally viable.
AnswerC

Haiku is specifically designed for speed. The exhibit lists it as 'Ultra-Low' latency, making it the most capable of meeting the 200ms requirement for live subtitles. Additionally, its low cost makes it economically viable for the high token volume generated by continuous live translation.

Why this answer

In real-time applications like live subtitling, latency is the primary constraint. Even if a model is more intelligent, if it cannot meet the timing requirements, it is unusable. Claude 3 Haiku's ultra-low latency makes it the only choice for near-instantaneous feedback loops required for live media.

Exam trap

Candidates often choose the 'most intelligent' model (Opus) instead of the 'fastest' model, ignoring the specific constraint that latency must be under 200ms, which is the defining requirement for live services.

12
MCQmedium

When evaluating model output, what does the term 'latency' specifically refer to in the context of the Anthropic API?

A.The number of tokens the model can process per second.
B.The total delay from sending a request to receiving the response.
C.The frequency of API rate limit errors.
D.The cost per token based on request time.
AnswerB

Latency is the end-to-end measure of how long the user waits for the API response. This includes network transit time, server-side processing time, and the time taken for the model to generate the tokens. It is the primary metric for ensuring a responsive, high-quality user experience in applications.

Why this answer

Latency refers to the total time elapsed between sending the request to the API and receiving the response. In production systems, measuring this is critical for user experience, as high latency can make applications feel sluggish or unresponsive. Understanding that latency is influenced by factors like input/output token counts and system load helps developers design more performant features, such as streaming or pre-computing content for the user.

Exam trap

Candidates often confuse latency with throughput or time-to-first-token, failing to recognize that latency specifically measures the total end-to-end delay from request transmission to final response reception.

13
MCQmedium

A data engineer is using Amazon Bedrock to invoke Claude 3 Haiku for classifying customer feedback into categories. They notice that the model sometimes returns categories that are not in the predefined list. Which change to the prompt is most likely to improve adherence to the allowed categories?

A.Increase the temperature to 1.0 so Claude explores more creative category assignments.
B.Add a stop sequence that matches the first invalid category so Claude stops generating when it deviates.
C.Reduce the max_tokens parameter so Claude has less room to generate invalid categories.
D.Provide a few-shot examples in the prompt showing correct classifications for similar feedback, including the exact category labels.
AnswerD

Few-shot examples demonstrate the desired input-output mapping and reinforce the exact category labels. Claude learns from the pattern and is more likely to output only the allowed categories. This is a prompt engineering technique that improves adherence without changing model parameters. Including examples that cover edge cases and explicitly show the format helps the model generalize correctly to new feedback.

Why this answer

Few-shot examples in the prompt show Claude the exact category labels and the expected format, which significantly improves adherence to a predefined set. By demonstrating correct classifications, the model learns the pattern and is less likely to invent new categories. Other parameters like temperature, stop sequences, and max_tokens affect randomness, truncation, and length, but do not teach the model which labels are valid.

Exam trap

The trap here is thinking that lowering max_tokens or adding stop sequences can enforce a category list, when only prompt-level guidance like few-shot examples reliably shapes label selection.

14
MCQhard

An engineering team wants Claude to classify thousands of support tickets into categories. They need the model to always return one of five exact category labels. Which approach most reliably constrains the output to those labels?

A.Use a high temperature so the model explores all categories evenly.
B.List the five allowed labels in the system prompt and instruct the model to output only one of them.
C.Post-process the model's free-text output with a keyword search for category names.
D.Ask the model to 'pick the best category' in the user message.
AnswerB

Enumerating the exact allowed labels and instructing the model to return only one of them gives a precise constraint. The system prompt applies to every request, so the model consistently sees the closed set. This directly satisfies the requirement that outputs match one of five fixed strings, making downstream parsing reliable.

Why this answer

To force output into a closed set of labels, the prompt must enumerate the allowed values and instruct the model to return exactly one. Placing this in the system prompt ensures the constraint applies to every classification request. Vague instructions, high temperature, or downstream keyword matching all allow invalid or ambiguous outputs that break automated pipelines.

Exam trap

The trap here is relying on post-hoc parsing or vague wording to fix classification output, when the robust solution is to define the exact allowed labels and constrain generation up front.

15
MCQeasy

An insurance firm wants to use Claude to extract data from scanned PDF claim forms that contain both handwritten text and printed tables. Which fundamental capability of the Claude 3 model family makes this workflow possible without using external OCR software?

A.Large Context Window
B.Vision Capabilities
C.Constitutional AI
D.JSON Mode
AnswerB

Claude 3 models feature sophisticated vision capabilities that allow them to process images, charts, and technical drawings. This native multimodal support enables the model to 'see' the scanned insurance forms, interpret handwriting, and extract data from tables directly from the visual input provided in the prompt.

Why this answer

The Claude 3 family was built with native multimodal capabilities, allowing the models to process and understand visual information directly. This eliminates the need for separate Optical Character Recognition (OCR) tools, as the model can interpret pixels and text simultaneously to extract structured data from complex document layouts.

Exam trap

Candidates often assume that processing complex documents like scanned PDFs always requires an external OCR pre-processing pipeline, overlooking the native multimodal vision capabilities built into models like Claude 3.

16
MCQmedium

A financial analyst uses Claude 3.5 Sonnet to extract line items from quarterly PDFs. They want the model to output only a JSON array without any conversational filler. Which feature should they configure to enforce this output format most reliably?

A.Use the system prompt to instruct 'Output only a valid JSON array with no other text.'
B.Set the temperature to 0.
C.Enable streaming responses.
D.Increase the max_tokens parameter to 4096.
AnswerA

System prompts are the correct mechanism to set persistent behavioral constraints, including output format. Instructing the model to emit only a JSON array without conversational filler directly addresses the requirement and is honored across turns. While no prompt guarantees perfect syntax, this is the most reliable native control for the described scenario.

Why this answer

The system prompt is the primary control for persistent behavioral instructions, including strict output formatting. Telling Claude to return only a valid JSON array with no additional text directly shapes the response structure. Other parameters like max_tokens, temperature, and streaming influence length, randomness, and delivery, but not the format contract the analyst needs.

Exam trap

The trap here is assuming that lowering temperature to 0 or raising max_tokens will make Claude output clean JSON, when only explicit formatting instructions in the system prompt reliably constrain structure.

17
MCQmedium

Refer to the exhibit. What is the specific purpose of the 'stop_sequences' parameter in this JSON payload?

A.It forces the model to ignore user inputs after that sequence.
B.It instructs the model to stop generating text when it encounters that sequence.
C.It increases the probability of the model using that sequence.
D.It reduces the model's temperature to 0.
AnswerB

The 'stop_sequences' parameter tells the model to halt generation as soon as it produces any of the provided strings. This is a common technique to prevent the model from entering a conversation loop or generating content beyond the intended response, ensuring the output is clean and ready for integration.

Why this answer

Stop sequences allow developers to force the model to cease generation when it hits a specific string. This is essential for controlling output length and preventing the model from hallucinating a continuing conversation (e.g., trying to generate the next 'Human:' turn). Mastering this parameter is key to integrating Claude into existing chat UI architectures where the developer needs precise control over when the model hands control back to the user.

Exam trap

Candidates often confuse 'stop_sequences' with 'max_tokens', assuming the former is for limiting response length rather than defining specific exit points to prevent the model from continuing into unwanted conversational turns.

18
MCQeasy

A support team wants Claude to answer customer questions using only the company's internal help articles. They plan to paste the relevant article text into the prompt before the user's question. Which prompting technique does this describe?

A.Zero-shot prompting
B.Retrieval-augmented generation (RAG)
C.Chain-of-thought prompting
D.Few-shot prompting
AnswerB

RAG combines a retrieval step that fetches relevant documents with generation, where the model answers using that fetched context. Pasting internal help articles before the user question is exactly this pattern: external knowledge is injected into the prompt to ground the response. It reduces hallucination by anchoring answers in provided source material.

Why this answer

Inserting relevant source documents into the prompt so the model answers from them is retrieval-augmented generation. The retrieval step supplies factual context, and the generation step produces an answer grounded in that context. This differs from zero-shot, few-shot, and chain-of-thought, which address examples, demonstrations, and reasoning traces respectively.

Exam trap

The trap here is conflating any prompt that includes extra text with few-shot prompting, when the key distinction is whether the added text is reference material (RAG) or solved examples (few-shot).

19
Multi-Selectmedium

A product team is evaluating Claude for a customer-facing assistant that must refuse to give medical diagnoses. Which TWO techniques should they use to make the refusal behavior consistent across many different user phrasings? (Choose two.)

Select 2 answers
A.Set temperature to 1.0 to encourage diverse safe responses.
B.Increase max_tokens so the model has room to explain refusals.
C.Include a few-shot example showing the assistant declining a diagnosis request.
D.Place a clear refusal policy in the system prompt.
E.Disable the system prompt and rely only on user instructions.
AnswersC, D

Few-shot examples demonstrate the exact desired behavior, including tone and wording for refusals. By showing a sample user request and the assistant's decline, the model learns the pattern and can generalize to new phrasings. This complements a system-prompt policy by giving a concrete template, improving consistency when users vary their wording.

Why this answer

Consistent refusal behavior comes from persistent instructions plus concrete demonstrations. A system-prompt policy establishes the rule for every turn, while a few-shot example shows exactly how a refusal should look, helping the model generalize across varied user phrasings. Temperature, max_tokens, and removing the system prompt do not reinforce policy adherence and can weaken it.

Exam trap

The trap here is thinking that raising temperature or output length makes the assistant more flexible or thorough, when refusal consistency actually depends on persistent instructions and demonstrated examples.

20
MCQmedium

Refer to the exhibit. What will happen if the user adds a new message to the 'messages' array?

A.The model will see only the latest message and forget the previous ones.
B.The model will correctly interpret the new input based on the full conversation history.
C.The system prompt will be overwritten by the new user message.
D.The API will return an error because messages cannot be appended.
AnswerB

The Claude API is designed to consume the entire sequence of messages as a single context. As long as the user correctly appends the new message to the existing history array, the model will maintain continuity, allowing for natural, multi-turn dialogue where previous context informs the current response.

Why this answer

The API requires that the 'messages' array represents the full turn-by-turn history. To continue the conversation, the application should append the new user message to the existing list while maintaining the original messages in the sequence. This ensures the model has access to the full context, allowing it to maintain conversational coherence and remember previous turns, which is crucial for building natural, fluid user experiences in AI applications.

Exam trap

Candidates often think appending a message replaces the history or requires resetting the API state from scratch each turn.

21
Multi-Selectmedium

A developer is integrating Claude into a legal document review system. The system must process lengthy contracts and answer questions about specific clauses. Which TWO of the following techniques help mitigate the risk of Claude hallucinating details not present in the document? (Choose two.)

Select 2 answers
A.Use a higher temperature setting to encourage more creative responses.
B.Instruct Claude to answer only based on the provided document and to say 'I don't know' if the information is not present.
C.Set max_tokens to a very low value to limit the response length.
D.Provide the full contract text in the prompt and ask Claude to cite the specific section for each answer.
E.Fine-tune the model on a dataset of legal contracts.
AnswersB, D

Explicitly instructing Claude to ground its answers in the provided document and to admit uncertainty when information is missing reduces hallucinations. This prompt engineering technique encourages the model to rely on the context rather than its internal knowledge, which is crucial for legal documents where accuracy is paramount. It sets clear boundaries for the model's responses.

Why this answer

To reduce hallucinations when reviewing legal documents, the most effective methods are to instruct Claude to rely solely on the provided text and to require citations for its answers. These techniques ground the model's responses in the source material, making it less likely to invent details and easier to verify accuracy.

Exam trap

The trap here is thinking that adjusting generation parameters like temperature or max_tokens can prevent hallucinations, when grounding instructions and citations are the key.

22
MCQmedium

An engineering team is designing a RAG system using Claude 3.5 Sonnet. They need to ensure the model focuses exclusively on the provided context without hallucinating external knowledge. Which architectural approach best ensures high adherence to provided context?

A.Increase the temperature setting to 1.0 to ensure maximum creativity.
B.Rely solely on the model's internal training weights for domain-specific queries.
C.Use XML tags to structure the input and define strict instructions in the system prompt.
D.Reduce the maximum token limit to prevent the model from generating long, incorrect answers.
AnswerC

XML tags provide clear delimiters that help Claude distinguish between instructions and context documents. Combining this structure with a system prompt that explicitly restricts the model to only use provided information forces the model to ignore its internal knowledge, effectively reducing hallucinations and increasing factual grounding in the retrieved data.

Why this answer

Grounding models in provided context requires clear instructions via system prompts and specific document delimiting. By using XML tags to isolate the context and providing explicit constraints in the system prompt, you define the boundaries of the model's knowledge. This architectural pattern is crucial for enterprise applications where accuracy is prioritized over creative generation, minimizing the risk of the model relying on its internal pre-training data.

Exam trap

Test-takers frequently choose general prompting techniques instead of leveraging specific structural delimiters like XML tags combined with strict system constraints for precise context grounding.

23
MCQmedium

When dealing with extremely large documents, what is the best strategy to maximize Claude's accuracy in information extraction?

A.Send the document in multiple separate API calls.
B.Use XML tags to delimit document sections and ask the model to reference those tags.
C.Force the model to summarize the document before extraction.
D.Increase the temperature to 2.0 to force creativity.
AnswerB

XML tags provide structure that helps the model navigate the context window. By labeling sections and instructing the model to look within specific tags, you significantly improve its ability to locate relevant information and produce accurate, grounded answers, which is especially effective for very long or dense input files.

Why this answer

The 'needle in a haystack' problem refers to finding a specific fact within a large volume of text. By utilizing strategic XML tagging to chunk the document and instructing the model to search within those specific tags, you guide its attention. This is a vital skill for enterprise document processing, where models must parse through hundreds of pages of documentation to identify critical, specific data points without getting overwhelmed.

Exam trap

Test-takers frequently rely on the model to scan huge documents without guidance, forgetting to use structural delimiters to direct attention.

24
MCQeasy

What does the 'Temperature' parameter control when configuring a request to a Claude model?

A.The speed at which the model processes the input.
B.The probability distribution of the next token selection.
C.The maximum number of tokens the model can generate.
D.The number of concurrent requests allowed.
AnswerB

Temperature modulates the probability distribution of the next token. By scaling the logits before the softmax operation, it flattens or sharpens the distribution. This directly impacts the randomness of the model's output, allowing users to move between highly predictable, logical responses and more diverse, creative generation styles.

Why this answer

Temperature is a hyperparameter that controls the randomness or 'creativity' of the model's output. A lower temperature leads to more deterministic and focused responses, while a higher temperature increases the probability of selecting less likely tokens, resulting in more varied and creative text. This setting is crucial for tuning the model behavior for specific use cases, such as coding (lower) versus creative writing (higher).

Exam trap

Candidates often confuse temperature with token length limits or repetition penalties, failing to recognize its role in probability distributions.

25
MCQhard

A developer is building a customer support chatbot using Claude. The chatbot must remember details from earlier in the conversation, such as the customer's order number and issue, to provide coherent responses. The conversation can last for many turns. Which implementation strategy best ensures Claude maintains context without exceeding token limits?

A.Summarize the conversation periodically and include the summary plus recent messages in subsequent requests.
B.Send the entire conversation history with each request, including all previous messages, to preserve full context.
C.Store the conversation in a vector database and retrieve relevant past messages based on the current query.
D.Rely on Claude's built-in memory feature to automatically remember previous interactions across sessions.
AnswerA

Summarizing older parts of the conversation condenses essential information into a compact form, reducing token usage while retaining context. Recent messages are kept verbatim for immediate coherence. This approach scales to long conversations and is a recommended pattern for managing context in Claude applications, balancing detail and efficiency.

Why this answer

Periodically summarizing the conversation and including the summary with recent messages keeps essential context within token limits. This method preserves coherence over many turns without unbounded growth. It is a standard pattern for long-running conversations in the Anthropic Messages API, ensuring Claude has the necessary background to respond appropriately.

Exam trap

The trap here is assuming Claude has persistent memory across API calls, when in fact each request must include all necessary context explicitly.

26
MCQmedium

When fine-tuning or optimizing prompts for Claude, what is the impact of excessive 'System Prompt' length?

A.It improves the model's ability to ignore user inputs.
B.It can lead to 'prompt drift' where the model loses focus on core instructions.
C.It causes the model to generate responses significantly faster.
D.It forces the model to use more creative, less deterministic tokens.
AnswerB

Extremely long system prompts can lead to a decrease in the model's ability to adhere to core constraints. As the length increases, the model may weigh instructions unevenly or lose track of critical directives, resulting in less consistent behavior and potentially lower quality responses compared to a concise, optimized prompt.

Why this answer

While Claude supports large context windows, excessively long system prompts can lead to dilution of focus, where the model may prioritize specific instructions over others or become less sensitive to the user's immediate input. Maintaining concise, high-impact system prompts is a best practice in AI engineering. It ensures the model remains responsive and accurate, reducing the noise-to-signal ratio and preventing degradation in instruction-following performance over time.

Exam trap

Candidates often assume that because models have large context windows, adding more instructions to the system prompt is always better, ignoring the risk of focus dilution and prompt drift.

27
Multi-Selecthard

A developer is building an application that uses the Anthropic Messages API with Claude 3.5 Sonnet to generate structured JSON output for a data pipeline. They need to ensure the output is valid JSON and conforms to a specific schema. Which TWO strategies should they use to maximize reliability? (Choose two.)

Select 2 answers
A.Set the temperature to 1.0 to encourage Claude to explore different JSON structures and pick the most valid one.
B.Include a clear instruction in the system prompt that the response must be valid JSON matching the provided schema, and provide the schema in the prompt.
C.Increase the max_tokens to the maximum allowed so Claude has enough space to include all schema fields.
D.Use a prefill technique by starting the assistant's response with an opening brace '{' to guide Claude into generating JSON.
E.Use a stop sequence of '}' to ensure Claude stops immediately after closing the JSON object.
AnswersB, D

Explicitly instructing Claude in the system prompt to output valid JSON and providing the schema gives the model a precise target. Claude is trained to follow detailed formatting instructions, so this significantly increases the likelihood of schema-conformant output. It is a fundamental step in structured generation with the Messages API, and it works alongside other techniques like validation and retries.

Why this answer

To maximize reliability for JSON output with Claude 3.5 Sonnet, combine explicit schema instructions in the system prompt with the prefill technique of starting the assistant response with an opening brace. These two strategies directly guide the model toward valid, schema-conformant JSON. Other parameters like temperature, max_tokens, and stop sequences do not enforce structure and may even harm output validity.

Exam trap

The trap here is assuming that a stop sequence on a closing brace guarantees complete JSON, when braces can appear in nested structures and cause premature truncation.

28
MCQhard

A user is experiencing 'Model Refusal' when processing a document that contains sensitive (but safe) medical information. What is the most likely cause?

A.The document is too long for the context window.
B.The model's safety guardrails are misinterpreting the intent due to sensitive keywords.
C.The model has reached its internal limit for medical-related queries.
D.The API key has expired, triggering a default security lockdown.
AnswerB

Safety filters often trigger on high-risk topics like medical data. If the prompt does not clearly state the benign intent, the model may default to a refusal to avoid providing potentially harmful advice. Adding context that emphasizes the professional or research-based nature of the request often resolves this issue.

Why this answer

Claude has built-in safety guardrails designed to prevent the generation of harmful content. Sometimes these filters can trigger on sensitive topics even when the user's intent is benign. This is known as a false positive.

Recognizing this behavior is critical for developers to adjust their prompts to provide more clear, benign context, which helps the model's safety systems distinguish between dangerous content and legitimate, safe professional use cases.

Exam trap

Candidates often assume safe medical or legal texts will never trigger refusals, misinterpreting false positives as true safety violations.

29
MCQhard

A machine learning engineer is comparing Claude 3 Opus and Claude 3.5 Sonnet for a complex mathematical reasoning task. The task involves multi-step proofs and requires the highest possible accuracy. Cost is not a primary concern. Which statement accurately describes the trade-off between these models for this use case?

A.Claude 3.5 Sonnet is strictly more capable than Claude 3 Opus in all reasoning tasks.
B.Claude 3.5 Sonnet is always faster and more accurate than Claude 3 Opus, making it the best choice regardless of task complexity.
C.Both models have identical reasoning capabilities, so the choice should be based solely on cost.
D.Claude 3 Opus is designed for the most complex reasoning tasks and generally offers higher accuracy than Claude 3.5 Sonnet on such tasks, though at higher cost and latency.
AnswerD

Claude 3 Opus is the most powerful model in the Claude 3 family, intended for highly complex reasoning where accuracy is paramount. While Claude 3.5 Sonnet is faster and more cost-effective, Opus typically provides superior performance on the most challenging reasoning tasks. Given that cost is not a concern and accuracy is critical, Opus is the appropriate choice.

Why this answer

For a complex mathematical reasoning task where accuracy is paramount and cost is not a concern, Claude 3 Opus is the best choice. It is the most capable model in the Claude 3 family, designed for the most demanding reasoning tasks, and generally provides higher accuracy than Claude 3.5 Sonnet, albeit with higher cost and latency.

Exam trap

The trap here is assuming that the newest model (Claude 3.5 Sonnet) is always superior in every aspect, overlooking that Opus remains the top-tier model for the most complex reasoning.

30
MCQmedium

A healthcare provider wants to use Claude to summarize patient-doctor conversations. They are concerned about the model 'hallucinating' or making up medical facts. Which Claude 3 feature or design principle directly addresses this concern by ensuring the model is honest and admits when it doesn't know an answer?

A.Increased Context Window
B.Vision Support
C.Constitutional AI Training
D.High Tokens-Per-Second
AnswerC

Claude is trained using Constitutional AI, which includes principles that reward the model for being honest and harmless. This training process specifically targets the reduction of hallucinations by teaching the model to prioritize accuracy and to admit uncertainty rather than providing a false but confident-sounding answer.

Why this answer

Anthropic uses Constitutional AI and specific training techniques to ensure that Claude models are not just helpful but also honest. This reduces the frequency of hallucinations and encourages the model to be 'calibrated,' meaning it expresses uncertainty when it is not confident in its answer.

Exam trap

Candidates often confuse 'Constitutional AI' with 'Model Fine-tuning' or 'Data Masking,' missing that the honesty and uncertainty expression is a direct outcome of the Constitutional AI training process.

31
MCQmedium

Which of the following describes the correct behavior of the Anthropic API regarding the 'system' prompt?

A.It is ignored if a user message is also provided.
B.It can be dynamically changed during a single request.
C.It provides foundational behavioral instructions to the model.
D.It is automatically appended to the end of the user message.
AnswerC

The system prompt serves as the anchor for the model's behavior, establishing the identity, rules, and context that the model adheres to. By separating these instructions from user input, it ensures that the model maintains its intended focus, reducing the risk of 'jailbreaking' or deviation from instructions.

Why this answer

The system prompt is designed to set the behavior, persona, and constraints of the model before it processes user-provided inputs. It is treated with higher priority than the user message, making it the ideal place for defining core safety guidelines or operational instructions. Correct utilization of the system field is a fundamental security practice, as it helps enforce behavioral boundaries consistently throughout the conversation, regardless of user attempts to influence the model.

Exam trap

Candidates often confuse the system prompt with user messages or memory storage, mistakenly thinking it dynamically changes based on user input during the conversation rather than remaining a fixed, high-priority foundational instruction.

32
MCQmedium

Refer to the exhibit. A developer is testing the vision capabilities of Claude 3.5 Sonnet. Based on the provided JSON request, which statement accurately describes how the model will process this input?

A.The model will fail because images must be sent in a separate system prompt.
B.The model will analyze the image and text together to provide the receipt total.
C.The request will error because Claude 3.5 Sonnet does not support base64 images.
D.The model will ignore the text and only provide a description of the image.
AnswerB

Claude 3.5 Sonnet is a multimodal model that can process text and image blocks simultaneously. By providing the image data and the text question in the same message, the developer allows the model to use its vision capabilities to extract the requested information from the receipt.

Why this answer

The Claude API supports multimodal inputs by allowing a mix of text and image content blocks within the messages array. In this specific configuration, the model receives both a base64-encoded image and a text query, enabling it to apply its vision capabilities to answer a specific question about the visual data.

Exam trap

Candidates often assume the model requires a separate 'Vision API' or 'OCR tool', failing to realize the native multimodal capability of the standard Messages API.

33
MCQeasy

A developer wants Claude to always respond in a strict, terse style and never use emojis, regardless of how users phrase their requests. Where should this persistent behavioral instruction be placed in a request to the Anthropic Messages API?

A.In the assistant's first turn as a pre-filled response.
B.In the system parameter, as a top-level instruction that applies to the whole conversation.
C.Appended to every user message as a trailing reminder.
D.As a metadata field attached to the request payload.
AnswerB

The system parameter is designed for persistent, cross-turn instructions such as tone, persona, and formatting rules. Placing the terse-style and no-emoji directive there applies it to every assistant turn without repeating it in each user message. This is the recommended way to enforce consistent behavior in the Messages API and keeps user turns focused on their actual questions.

Why this answer

Persistent behavioral rules such as tone and formatting belong in the system parameter of the Messages API. The system prompt is applied across all turns, so the terse, no-emoji requirement does not need to be repeated in each user message. Appending reminders, pre-filling an assistant turn, or using metadata do not provide the same durable, conversation-wide control over Claude's behavior.

Exam trap

The trap here is treating the system parameter as optional and instead repeating style rules in every user message, which is token-wasteful and less reliable than a single persistent system instruction.

34
MCQmedium

A product team is building a customer-support assistant on Claude. They want Claude to answer only from a fixed set of help-center articles and to refuse any question outside that scope. They also need to update the article set frequently without retraining a model. Which approach best meets these requirements?

A.Lower the temperature to 0 and rely on Claude's built-in knowledge of common support topics to produce consistent answers.
B.Fine-tune a Claude model on the current help-center articles so that it memorizes the answers and cannot go off-topic.
C.Use a very large max_tokens value so Claude has room to reproduce entire help-center articles inside each response.
D.Place the help-center articles in the system prompt and instruct Claude to answer only from those articles, then update the system prompt when the article set changes.
AnswerD

Putting the approved articles in the system prompt gives Claude authoritative grounding for every turn, and the instruction to stay within scope is enforced by the model's instruction-following behavior. Because the system prompt is supplied at request time, the team can swap in updated articles instantly without any model training, which directly satisfies the frequent-update requirement.

Why this answer

Grounding Claude in the approved articles through the system prompt satisfies both the scoping requirement and the need for frequent updates, because the content is provided at request time rather than baked into the model. Fine-tuning is poorly suited to a changing knowledge base, temperature does not enforce scope, and max_tokens only affects output length.

Exam trap

The trap here is assuming that fine-tuning is the right way to give Claude a specific, changeable body of knowledge, when prompt-time grounding is what actually enables fast updates and strict scoping.

35
MCQhard

An engineering firm is using Claude to help design a complex micro-architecture for a new processor. The task requires deep logical reasoning, knowledge of hardware description languages (Verilog), and the ability to handle highly abstract concepts. Which model should be used for the highest possible accuracy?

A.Claude 3 Haiku
B.Claude 3 Sonnet
C.Claude 3 Opus
D.Claude 2.0
AnswerC

Opus is the top-tier model in the Claude 3 family, offering the highest level of performance on complex reasoning benchmarks. Its superior ability to synthesize information and handle specialized technical language makes it the correct choice for advanced research and development in fields like micro-processor architecture. (51 words)

Why this answer

Claude 3 Opus is Anthropic's most advanced model, specifically designed for tasks that require deep reasoning and expert-level knowledge. For specialized engineering tasks like micro-architecture design, the model's ability to navigate complex, multi-step logical problems and its extensive training on technical datasets provide a level of accuracy and nuance that smaller models cannot match. (68 words)

Exam trap

Candidates frequently choose faster, cheaper models like Haiku or Sonnet, assuming standard tasks apply, but fail to recognize that deeply specialized architectural tasks demand Opus's superior reasoning.

36
MCQmedium

A team is using Claude to summarize legal contracts. They need the summaries to reflect only the contract text, not any outside assumptions. Which prompting technique best reduces the chance that Claude introduces information not present in the source document?

A.Ask Claude to summarize from its general knowledge of contract law first, then compare with the document.
B.Remove the document entirely and ask Claude to generate a typical contract summary template.
C.Instruct Claude to answer only from the provided document and to say when the document does not contain the answer.
D.Raise the temperature so Claude explores more possible interpretations of the contract.
AnswerC

Explicitly constraining Claude to the supplied document and requiring it to acknowledge missing information is a direct, effective guard against unsupported claims. It sets a clear behavioral boundary in the prompt and gives Claude a safe fallback instead of inventing details. For legal summaries where fidelity to the source is critical, this grounding instruction materially reduces hallucination risk.

Why this answer

Grounding Claude in the provided document, and instructing it to say when the answer is absent, is the most direct way to keep summaries faithful to the source. It constrains the model to the supplied text and provides a safe response when information is missing. Raising temperature, relying on general legal knowledge, or removing the document all increase the risk of unsupported content rather than reducing it.

Exam trap

The trap here is assuming that a more capable model automatically avoids hallucination, when faithful summarization actually depends on explicit grounding instructions and a safe fallback for missing information.

37
MCQmedium

An AI researcher is concerned about 'hallucinations' when Claude summarizes internal technical specifications. Which approach leverages Claude's fundamental design to minimize the risk of the model inventing non-existent features?

A.Increasing the temperature to 1.0
B.Using Claude 3 Haiku for higher precision
C.Providing the documents in the context and requesting citations
D.Setting 'max_tokens' to a very low value
AnswerC

By placing the technical specifications in the prompt and asking the model to cite specific passages, you force Claude to ground its response in the provided text. This 'RAG-style' approach ensures the model focuses on the evidence at hand, reducing the likelihood of generating outside or invented information. (52 words)

Why this answer

Hallucinations occur when a model generates plausible but incorrect information. To mitigate this, grounding the model in the provided context is the most effective strategy. By explicitly instructing Claude to use only the provided text and to cite its sources, the model's reasoning is constrained to the verified data, significantly improving the factual accuracy of the summary. (69 words)

Exam trap

Candidates often select 'increasing the model temperature' or 'using a larger model,' which actually increases the risk of hallucination rather than grounding the model in the provided technical specifications.

38
MCQeasy

A product manager is evaluating Claude models for a customer-support chatbot that must handle 50,000 conversations per day while keeping inference costs low. The conversations are short, factual, and do not require complex reasoning. Which Claude model is the most appropriate choice?

A.Claude 2.1
B.Claude 3 Opus
C.Claude 3 Haiku
D.Claude 3.5 Sonnet
AnswerC

Claude 3 Haiku is the fastest and most cost-effective model in the Claude 3 family, optimized for high-volume, straightforward tasks like simple customer-support queries. It delivers quick responses at low cost, making it ideal for 50,000 short, factual conversations per day where complex reasoning is not required. This aligns perfectly with the scenario's requirements.

Why this answer

The scenario demands a model that can handle high-volume, simple conversations at low cost. Claude 3 Haiku is specifically designed for speed and cost efficiency, making it the right fit for short, factual customer-support interactions. The other models offer greater capability but at higher cost, which is unnecessary for this use case.

Exam trap

The trap here is assuming that the most capable model is always the best choice, when cost and latency requirements often favor a smaller, faster model.

39
Multi-Selectmedium

A product team is using Claude to generate marketing copy. They want to ensure the output aligns with brand voice and avoids certain topics. Which TWO techniques are most effective for controlling Claude's output in this scenario? (Choose two.)

Select 2 answers
A.Include few-shot examples in the prompt showing desired outputs for similar marketing requests.
B.Post-process the output with a keyword filter to remove any prohibited terms before publishing.
C.Set the temperature parameter to a high value to encourage more creative and varied outputs.
D.Use the top_p parameter with a low value to restrict the model to only the most likely words.
E.Provide a detailed system prompt that specifies the brand voice, tone, and prohibited topics.
AnswersA, E

Few-shot examples demonstrate the desired style and content, helping Claude infer patterns and apply them to new requests. This is particularly effective for nuanced tasks like brand voice, where examples can illustrate tone, vocabulary, and structure. Combined with a system prompt, few-shot learning significantly improves alignment with specific requirements.

Why this answer

A detailed system prompt and few-shot examples are proactive techniques that shape Claude's output during generation. The system prompt establishes persistent guidelines, while examples demonstrate desired patterns. Together, they provide clear direction and improve consistency, making them the most effective methods for controlling brand voice and avoiding prohibited topics.

Exam trap

The trap here is relying on post-hoc filters or sampling parameters, which do not provide the same level of control as explicit instructions and examples.

40
MCQmedium

When designing a prompt for Claude to perform complex reasoning, why is it beneficial to include 'think step-by-step' in the instructions?

A.It forces the model to use more memory, which increases its internal intelligence.
B.It allows the model to break down complex problems into manageable sub-tasks.
C.It automatically converts the output format into a bulleted list.
D.It minimizes the token cost by reducing the number of words in the final output.
AnswerB

Chain-of-thought prompting prompts the model to generate intermediate steps. This decomposition helps the model handle multi-step reasoning tasks more accurately by grounding each part of the process in the previous logic, significantly reducing errors that occur when the model attempts to solve everything at once.

Why this answer

The 'think step-by-step' technique leverages the model's ability to perform chain-of-thought reasoning. By forcing the model to articulate its intermediate logic before arriving at a final answer, the model is less likely to jump to premature conclusions. This is particularly important for logic, math, or coding tasks where the final output depends on the accuracy of sequential steps rather than direct pattern completion.

Exam trap

Candidates often assume 'think step-by-step' is a magic phrase that guarantees accuracy for any prompt, forgetting that the model still requires high-quality, unambiguous instructions to reason effectively through complex logic.

41
MCQhard

A company is developing an application that generates long-form creative content. They notice that as the output length increases, the model's speed seems to fluctuate. Which technical factor most significantly impacts the latency of the 'Time to Last Token' (TTLT) in this scenario?

A.The number of concurrent users in the API queue.
B.The total number of output tokens generated.
C.The use of a system prompt instead of a user prompt.
D.The presence of images in the input context.
AnswerB

Because autoregressive models like Claude generate one token at a time, each additional token adds to the cumulative processing time. For long-form creative content, the sheer volume of generated text is the dominant factor in determining how long it takes for the model to finish the request. (51 words)

Why this answer

The Time to Last Token (TTLT) is primarily a function of the total number of tokens generated by the model. Each token is generated sequentially, meaning the model must perform a full inference pass for every single word or sub-word it produces. Therefore, longer outputs naturally take more time, as the total latency scales linearly with the generation length. (70 words)

Exam trap

Candidates often attribute latency to 'input size' or 'network bandwidth,' overlooking that generation (output) is an inherently sequential process where every token adds linear time.

42
MCQmedium

An enterprise legal team needs to analyze a 150,000-token collection of merger and acquisition documents to identify conflicting indemnity clauses. Which feature of the Claude 3 model family is most critical for ensuring the entire dataset is processed in a single inference pass without losing context?

A.Multimodal Vision capabilities
B.System Prompts
C.200,000-token context window
D.Temperature settings
AnswerC

The large context window is the specific technical specification that enables Claude to process long-form documents or large datasets at once. It ensures that the model can maintain coherence across the entire 150,000-token set, allowing for complex cross-referencing and deep thematic analysis across multiple distinct files or sections. (52 words)

Why this answer

The context window determines the amount of information the model can hold in its active memory during a single request. For legal analysis involving massive document sets, a 200,000-token context window allows Claude to 'see' all relevant clauses simultaneously. This enables the model to identify cross-document contradictions and nuances that would be missed if the data were fragmented. (70 words)

Exam trap

Candidates often select 'model size' or 'training data volume,' ignoring that the context window is the specific technical constraint that allows the model to 'see' all 150,000 tokens at once.

43
MCQmedium

A team is choosing a Claude model for a nightly batch job that summarizes thousands of internal reports. Cost per token and throughput matter more than peak reasoning ability, and the summaries tolerate minor stylistic variation. Which selection criterion is most appropriate?

A.Choose the model with the fastest release date to ensure it has the newest features.
B.Choose the largest, most capable model available to maximize summary quality regardless of cost.
C.Choose a model based solely on the largest available context window, since reports are long.
D.Choose a smaller, lower-cost model in the Claude family that meets the required summary quality for the batch workload.
AnswerD

When cost and throughput dominate and minor stylistic variation is acceptable, a smaller, lower-cost model that still meets quality targets is the right fit. It processes the nightly volume economically while satisfying the summarization requirement, matching the team's stated priorities.

Why this answer

Model choice should follow workload requirements. Here cost per token and throughput dominate while quality tolerates minor variation, so a smaller, lower-cost Claude model that still meets the summary bar is appropriate. Picking the largest model or selecting on window size or recency ignores the explicit priorities and raises expense without needed benefit.

Exam trap

The trap here is defaulting to the most capable model, when the workload's cost and throughput constraints actually favor a smaller model that meets quality.

44
MCQhard

Which TWO of the following statements accurately describe the characteristics of the Claude 3.5 Sonnet model's context window and performance?

A.Claude 3.5 Sonnet supports an infinite context window for streaming data.
B.The model's performance decreases linearly as the context window is filled.
C.Claude 3.5 Sonnet offers a 200k token context window capacity.
D.The model is optimized for high-throughput, low-latency performance.
E.Context window usage is irrelevant to API response latency.
AnswerC, D

The 200k context window is a foundational specification for the Claude 3.5 Sonnet model. This capacity allows developers to pass substantial amounts of documentation, research papers, or entire code repositories into the prompt, enabling the model to synthesize information across vast inputs without requiring complex RAG orchestration.

Why this answer

Claude 3.5 Sonnet features a 200k token context window, which allows for the ingestion of large codebases or documents. Furthermore, the model is optimized for high throughput and low latency, making it ideal for tasks requiring rapid processing. Understanding these technical trade-offs is essential for architecting scalable applications that need to balance the need for deep contextual understanding with the speed requirements of real-time user interfaces.

Exam trap

Candidates often mistakenly believe Claude 3.5 Sonnet has a 1M+ context window or is slower than Opus, failing to recognize its specific niche as a high-throughput, 200k-context model.

45
MCQmedium

Which of the following describes the purpose of 'Role' assignment in the Messages API?

A.To define the administrative access level of the API user.
B.To differentiate between the user's input and the model's past responses.
C.To specify the persona the model should adopt for the entire request.
D.To allow the model to rewrite the user's prompt for better accuracy.
AnswerB

The Messages API requires explicit labeling of inputs as 'user' or 'assistant' to construct a coherent dialogue history. This separation allows the model to correctly attribute previous statements and follow the conversation flow, which is essential for multi-turn interactions where Claude must respond based on context established earlier.

Why this answer

Assigning roles ('user' or 'assistant') allows the API to maintain conversational context and structure. The model needs to identify which inputs are the user's queries and which are its own prior outputs to maintain logical flow. This is fundamental to Claude's ability to engage in multi-turn dialogues where the model references previous context to generate coherent, relevant, and context-aware responses in ongoing sessions.

Exam trap

Test-takers sometimes confuse role assignment with system-level persona definitions, missing that roles specifically manage conversational turn-taking history.

46
Multi-Selectmedium

A team is building a document Q&A feature on Claude and wants to reduce hallucinations when the answer is not present in the supplied documents. Which TWO techniques are appropriate? (Choose two.)

Select 2 answers
A.Instruct Claude in the system prompt to answer only from the provided documents and to say it does not know when the answer is absent.
B.Remove the documents from the prompt and rely on Claude's pretrained knowledge to fill in missing details.
C.Increase max_tokens so Claude has more space to explain its reasoning before giving the final answer.
D.Ask Claude to include a short citation or quote from the source document that supports each claim in its answer.
E.Raise the temperature so Claude explores a wider range of possible answers and is more likely to find the right one.
AnswersA, D

An explicit instruction to stay within the supplied documents and to admit uncertainty when the answer is missing gives Claude a clear behavioral rule to follow. This is a direct, low-cost way to reduce confident fabrication because the model is told what to do when evidence is lacking, rather than defaulting to a plausible-sounding guess.

Why this answer

Reducing hallucinations in document Q&A comes from constraining Claude to the supplied evidence and making that grounding verifiable. An instruction to answer only from the documents and to admit uncertainty sets the behavioral boundary, while requiring citations or quotes makes unsupported claims detectable. Sampling temperature and output-length settings do not improve factual grounding.

Exam trap

The trap here is treating sampling settings like temperature or max_tokens as hallucination controls, when grounding instructions and source citations are what actually constrain factual claims.

47
MCQmedium

A product team is building a customer support assistant on Amazon Bedrock using the Anthropic Claude 3.5 Sonnet model. They need the model to answer questions strictly from a provided knowledge base and to refuse to answer if the information is not present. Which technique should they use to constrain Claude's behavior most reliably?

A.Add a system prompt that instructs Claude to only use the provided knowledge base and to say 'I don't know' when the answer is not found.
B.Increase the max_tokens parameter so Claude has enough room to include the full knowledge base in every response.
C.Fine-tune the Claude 3.5 Sonnet model on the knowledge base so that it memorizes the content and refuses out-of-scope questions.
D.Lower the temperature setting to 0 so Claude becomes deterministic and will not hallucinate outside the knowledge base.
AnswerA

A system prompt sets persistent, high-level behavioral instructions that Claude prioritizes throughout the conversation. By explicitly restricting Claude to the provided knowledge base and requiring an 'I don't know' response when information is absent, the system prompt reliably constrains the model's scope. This is the recommended approach for grounding and refusal behavior in Claude on Amazon Bedrock, as it leverages Claude's instruction-following strength without altering the model weights.

Why this answer

A system prompt is the most reliable way to set persistent behavioral constraints for Claude on Amazon Bedrock. It instructs the model to restrict answers to a provided knowledge base and to refuse when information is missing. Unlike fine-tuning, which is unsupported and unsuitable for dynamic knowledge, or temperature and max_tokens, which control randomness and length, the system prompt directly shapes Claude's adherence to scope and refusal rules.

Exam trap

The trap here is assuming that lowering temperature to 0 eliminates hallucinations or enforces grounding, when temperature only affects randomness and not factual adherence.

48
MCQeasy

A developer is choosing between Claude 3 Haiku and Claude 3.5 Sonnet for a real-time chat application that requires very low latency and handles simple, short queries. Cost is a primary concern. Which model is most appropriate and why?

A.Claude 3 Haiku, because it is optimized for speed and cost-effectiveness while still handling straightforward tasks well.
B.Claude 3.5 Sonnet, because it has the largest context window and can handle more concurrent users.
C.Claude 3.5 Sonnet, because it is the only model that supports streaming responses for real-time chat.
D.Claude 3 Opus, because it provides the highest accuracy for all query types, ensuring customer satisfaction.
AnswerA

Claude 3 Haiku is designed for fast, low-cost interactions, making it ideal for real-time chat with simple queries. It offers lower latency and lower cost per token compared to larger models like Claude 3.5 Sonnet. For straightforward tasks that do not require deep reasoning, Haiku provides sufficient quality. Choosing it aligns with the requirements of low latency and cost sensitivity.

Why this answer

Claude 3 Haiku is optimized for speed and cost, making it the best fit for a real-time chat application with simple, short queries. It provides low latency and low cost per token while maintaining adequate quality for straightforward tasks. Larger models like Claude 3.5 Sonnet or Claude 3 Opus offer more capability but at higher cost and latency, which does not align with the stated priorities.

Exam trap

The trap here is assuming that the most capable model is always the best choice, when requirements like low latency and cost often favor a smaller, faster model.

49
MCQmedium

What is the primary purpose of using XML tags within a prompt when working with Claude models?

A.To increase the security of the API connection.
B.To improve structural separation within the prompt.
C.To enable the automatic parsing of the output as XML.
D.To bypass the token limits of the model.
AnswerB

XML tags are specifically designed to delimit sections of a prompt, allowing the model to clearly distinguish between system instructions, contextual data, and specific queries. This separation ensures that the model treats each part of the prompt correctly, leading to higher quality and more reliable outputs during execution.

Why this answer

XML tags provide a clear structure that helps the model differentiate between various sections of the prompt, such as instructions, source data, and user constraints. This structural clarity significantly improves the model's performance by reducing ambiguity and preventing instructions from bleeding into the data. Using tags is a standard best practice for prompt engineering with Claude, enabling developers to build more robust and predictable prompt templates for complex tasks.

Exam trap

Candidates often believe XML tags are for 'security' or 'API authentication', missing that their primary role is providing clear, delimited structure for the model to parse instructions.

50
Multi-Selectmedium

A data analyst wants to ensure that Claude produces highly predictable and consistent results when summarizing financial reports. Which TWO parameter adjustments should they make to the API request?

Select 2 answers
A.Set 'temperature' to a value closer to 0.0
B.Increase 'max_tokens' to ensure full coverage
C.Lower the 'top_p' value to limit the token pool
D.Use the 'stop_sequences' parameter for formatting
E.Switch to the 'claude-3-haiku' model for speed
AnswersA, C

Lowering the temperature makes the model's token selection more deterministic by favoring the most likely next word. This reduces the variability between different runs of the same prompt, which is essential for tasks like financial summarization where you want the same facts reported every time. (50 words)

Why this answer

Controlling the randomness of an LLM is vital for analytical tasks where consistency is preferred over creativity. Lowering the 'temperature' reduces the likelihood of the model selecting less probable tokens, while adjusting 'top_p' (nucleus sampling) limits the pool of tokens the model considers. Together, these settings force the model to be more deterministic and focused. (68 words)

Exam trap

Candidates often assume that changing only the temperature is sufficient, or they confuse top_p with frequency penalties, forgetting that both temperature and top_p must be jointly adjusted to completely control model randomness.

51
MCQmedium

A developer is using the Anthropic Messages API to build a conversational agent. They want to maintain context across multiple turns and ensure that the model's responses are consistent with the system prompt. Which API feature should they use to provide the system prompt?

A.Prepend the system prompt to every user message in the conversation.
B.Use the 'system' parameter in the API request to provide the system prompt.
C.Set the 'role' of the first message to 'system' in the messages array.
D.Include the system prompt as the first user message in the messages array.
AnswerB

The Anthropic Messages API includes a top-level 'system' parameter specifically for system prompts. This parameter is separate from the 'messages' array and is used to set the model's behavior, persona, or instructions. It ensures that the system prompt is consistently applied across all turns of the conversation, providing a stable context that guides the model's responses.

Why this answer

The Anthropic Messages API provides a dedicated 'system' parameter for supplying a system prompt. This parameter is separate from the conversational messages and ensures that the model's behavior is consistently guided by the system instructions across all turns. Using this parameter is the correct and efficient way to set a system prompt.

Exam trap

The trap here is assuming that the system prompt should be part of the messages array, either as a user message or with a 'system' role, when the API has a separate parameter for it.

52
MCQmedium

A developer wants Claude to return structured data that another service can parse automatically. They need the output to follow a fixed schema every time, with no conversational text around it. Which approach is most appropriate?

A.Ask Claude to explain its reasoning first and then append the structured data at the end of the response.
B.Describe the desired schema in the prompt and ask Claude to respond only with valid JSON that matches it.
C.Set max_tokens to a large value and let Claude choose whichever format best represents the data.
D.Request the output as a Markdown table and convert it to JSON on the client side.
AnswerB

Describing the exact schema and instructing Claude to reply only with matching JSON is the standard way to obtain machine-parseable output. Claude is capable of following precise format instructions, and keeping the response limited to JSON avoids the conversational wrapper that would break downstream parsing, making this a direct fit for the requirement.

Why this answer

Structured output is achieved by specifying the schema precisely and constraining Claude to emit only matching JSON. This removes the need for brittle post-processing and satisfies the requirement that no conversational text surround the data. Alternatives such as reasoning-first responses, Markdown tables, or open-ended formatting all introduce parsing uncertainty.

Exam trap

The trap here is assuming that any response containing valid JSON is usable, when surrounding prose or an unspecified format can still break automated parsing.

53
MCQeasy

A UX designer wants the AI assistant to appear more interactive by showing the response as it is being generated, rather than waiting for the entire block of text to be finished. Which API feature should the developer implement?

A.Batch Processing
B.Streaming
C.Recursive Infilling
D.Contextual Caching
AnswerB

Streaming enables a real-time data flow where tokens are sent to the client immediately after they are produced. This creates a 'typing' effect in the user interface, which makes the application feel much more responsive and interactive, even if the total generation time remains the same. (51 words)

Why this answer

Streaming allows the API to send the response in small chunks as they are generated, rather than waiting for the entire completion. This significantly improves the 'perceived' latency for the user, as they can start reading the beginning of the response while the model is still working on the end. (66 words)

Exam trap

Candidates often mistake asynchronous execution or caching features for streaming, missing that streaming specifically handles incremental chunk delivery to improve perceived response latency.

54
MCQhard

Refer to the exhibit. An application sends the provided JSON payload to the Anthropic API. What is the expected behavior regarding the system prompt?

A.The API will return an error because the system prompt is improperly formatted.
B.The system prompt will be ignored because it must be inside the messages array.
C.The API will correctly apply the system prompt to the entire conversation turn.
D.The model will see the system prompt as the latest user message.
AnswerC

The system prompt is correctly defined at the top level of the request. The Claude API uses this field to set the 'persona' or constraints that govern the model's behavior for the entire session. This ensures the assistant maintains the intended behavior throughout the multi-turn exchange provided in the messages.

Why this answer

In the Anthropic Messages API, the system prompt must be defined at the top level of the JSON object, not within the messages array. Because the system prompt is defined correctly at the top level in this exhibit, it will be correctly processed. Understanding this structure is essential for developers to correctly separate behavioral instructions from conversational history, preventing unexpected model behavior during multi-turn API interactions.

Exam trap

Many test-takers mistakenly believe system prompts should be nested inside the messages array alongside user and assistant turns.

55
Multi-Selectmedium

Which THREE components are required when making a successful request to the Anthropic Messages API? (Select THREE)

Select 3 answers
A.The 'model' parameter specifying the version
B.The 'messages' array containing the conversation
C.The 'max_tokens' parameter defining the output limit
D.A 'temperature' value set to exactly 1.0
E.A 'system' parameter with at least 50 words
AnswersA, B, C

The API must know which specific model to invoke, such as 'claude-3-5-sonnet-20240620'. Without this parameter, the system cannot route the request to the correct inference engine, as different models have different pricing, performance characteristics, and capabilities required for the task.

Why this answer

Understanding the API structure is fundamental for any developer working with Claude. The Messages API requires specific parameters to function correctly, including model identification, the conversation history, and a limit on the output length to ensure predictable behavior and resource management.

Exam trap

Candidates often select optional parameters like temperature or system prompts as mandatory fields, forgetting that only the model identifier, messages array, and max_tokens are strictly required.

56
MCQhard

A developer is building a multi-turn chat application with the Anthropic Messages API. After several exchanges, they notice Claude loses track of details mentioned early in the conversation. Which action best addresses this while staying within the API's design?

A.Increase the temperature to improve recall of earlier turns.
B.Add a system prompt saying 'Remember all previous messages.'
C.Switch to a larger Claude model with a bigger context window.
D.Resend the full prior message history with each new request.
AnswerD

The Messages API is stateless: the model does not retain prior turns between calls. To preserve context, the client must include the relevant earlier messages in the messages array on every request. Resending the full history ensures Claude can attend to details from earlier exchanges, directly addressing the loss of early conversation details.

Why this answer

The Messages API does not persist conversation state; each call is independent. To let Claude reference earlier details, the application must include those prior turns in the messages array of every request. A larger context window or a system instruction cannot substitute for actually sending the history, and temperature has no bearing on recall.

Exam trap

The trap here is assuming the model remembers previous API calls or that a system prompt can grant memory, when the Messages API is stateless and requires the client to resend conversation history.

57
MCQhard

A developer is using the Anthropic Messages API with Claude 3 Opus to build a multi-turn technical support chatbot. The conversation history is growing large, and they want to reduce token usage while preserving the most relevant context. They decide to implement a sliding window that keeps only the last N turns. What is a potential drawback of this approach that they should consider?

A.The sliding window will increase latency because Claude must reprocess the entire conversation history on each turn.
B.The sliding window may discard important early instructions or context that Claude needs to maintain consistent behavior throughout the conversation.
C.Claude will automatically summarize the dropped turns and include the summary in its response, which may introduce hallucinations.
D.The sliding window will cause Claude to exceed the model's maximum context window because older turns are still counted in the token limit.
AnswerB

A sliding window keeps only recent turns, so any critical instructions or facts established early in the conversation are dropped once they fall outside the window. Claude has no memory beyond the supplied context, so it may forget user preferences, prior constraints, or key details. This can lead to inconsistent or incorrect responses. The developer should summarize or selectively retain important early context rather than blindly truncating.

Why this answer

A sliding window reduces token usage by dropping older turns, but it also removes early context that may be essential for consistent behavior. Claude does not retain memory outside the provided messages, so once early instructions or facts are dropped, they are unavailable. The developer should consider summarization or selective retention of important early context to balance token efficiency with conversational coherence.

Exam trap

The trap here is believing that Claude retains memory of dropped turns or automatically summarizes them, when in fact the model only sees the current request payload.

58
MCQeasy

Which core design principle is used by Anthropic to ensure that Claude models are helpful, honest, and harmless by training them against a set of written rules or values?

A.Manual Data Scrubbing
B.Constitutional AI
C.Supervised Pattern Matching
D.Open-Source Alignment
AnswerB

This approach involves training the model to follow a specific set of principles (a constitution) to guide its decision-making. It allows Claude to evaluate its own responses for safety and helpfulness, ensuring it adheres to ethical guidelines without requiring human intervention for every single generated response. (50 words)

Why this answer

Constitutional AI is the foundational methodology Anthropic uses to align Claude's behavior with human values. By using a 'constitution' or set of principles during the RLHF process, the model learns to self-critique and revise its responses. This reduces the need for manual labeling of every possible harmful output, creating a more robust safety framework. (68 words)

Exam trap

Candidates often guess 'RLHF' or 'Supervised Fine-Tuning' generally, failing to identify 'Constitutional AI' as the specific, proprietary design principle Anthropic uses to align models with a defined set of values.

59
Multi-Selectmedium

A developer is integrating Claude 3.5 Sonnet into a medical imaging application. Which TWO capabilities of the Claude 3 family make it particularly suited for analyzing diagnostic reports alongside X-ray images?

Select 2 answers
A.Native Vision processing for images
B.Support for 1,000,000-token output
C.High-accuracy complex reasoning
D.Ability to browse the live internet
E.Recursive self-improvement loops
AnswersA, C

Claude 3 models can directly ingest and interpret visual data, such as JPEG or PNG files, allowing them to describe and analyze medical images. This allows the model to 'see' anomalies or patterns in X-rays, which can then be cross-referenced with the text-based diagnostic reports provided in the prompt. (52 words)

Why this answer

The Claude 3 family introduced native multimodal capabilities, allowing the models to process both text and visual data in a single request. This is essential for medical applications where a report must be compared against an image. Additionally, the improved reasoning capabilities ensure that the model can draw logical connections between the visual evidence and the textual descriptions. (69 words)

Exam trap

Candidates often select only one capability or focus on 'image generation,' failing to recognize that the requirement involves both 'Native Vision' (input) and 'Complex Reasoning' (analysis).

60
MCQmedium

An analyst is using Claude 3.5 Sonnet to compare two 50-page contracts. They notice that Claude correctly identifies a discrepancy in a small footnote on page 74 of the combined input. This demonstrates which fundamental performance characteristic of Claude?

A.Zero-shot learning
B.Strong long-context recall
C.Sentiment analysis
D.Iterative refinement
AnswerB

Claude 3 and 3.5 models are engineered for near-perfect recall across their entire 200,000-token context window. This means the model can accurately 'remember' and retrieve small details, like a footnote, even when they are buried deep within hundreds of pages of other text.

Why this answer

The ability to retrieve specific information from a massive context window is known as 'recall' or 'Needle In A Haystack' performance. Claude models are specifically optimized to ensure that they don't just 'read' the whole document, but can actually find and use information regardless of its position in the input.

Exam trap

Candidates frequently confuse the ability to process large documents with 'reasoning' or 'summarization' rather than the specific capability of 'long-context recall' (Needle In A Haystack) required to find isolated facts.

61
MCQeasy

Why should developers use the Anthropic Messages API instead of legacy Completions API?

A.The Messages API is faster because it bypasses safety checks.
B.It provides a superior structure for managing multi-turn conversations.
C.It allows developers to modify the model's internal weights.
D.It is the only API that supports streaming responses.
AnswerB

The Messages API is built specifically for multi-turn dialogue, offering a clean, structured way to pass conversational history. This format simplifies the management of state, improves the model's ability to track context, and is the foundation for all modern Anthropic features, making it the correct choice for any application.

Why this answer

The Messages API is the current standard for interaction with Anthropic models. It natively supports structured conversation history, system-level prompts, and improved safety features, which are necessary for modern, production-grade applications. Transitioning to this API is essential for accessing the latest model features, receiving better support, and ensuring long-term compatibility with future model updates and enterprise deployment requirements, making it a critical choice for any professional AI project.

Exam trap

Candidates often incorrectly suggest the Messages API is used primarily for 'lower cost' or 'faster speeds', failing to recognize that its primary architectural advantage is structured, multi-turn conversation management.

62
MCQeasy

When designing a prompt for Claude, which technique is most effective at reducing the risk of 'hallucination' or factually incorrect information?

A.Setting the temperature parameter to 0.
B.Instructing the model to admit when it lacks information.
C.Using the most expensive model variant.
D.Increasing the max_tokens limit.
AnswerB

Instructing the model to refrain from guessing and to indicate when information is missing is a highly effective grounding strategy. This creates a safety boundary, preventing the model from inventing facts when the source material is inadequate, which is essential for maintaining trust and accuracy in enterprise applications.

Why this answer

Providing clear, constrained instructions and grounding the model in provided context is the most effective way to minimize hallucinations. By explicitly instructing the model to reply 'I don't know' if the information is not present in the provided source text, developers can significantly improve the factual reliability of the system. This practice forces the model to prioritize provided data over its internal training parameters, improving accuracy in domain-specific tasks.

Exam trap

Candidates often select 'giving the model more creative freedom' or 'increasing the temperature', which directly contradicts the goal of factual accuracy and hallucination reduction.

63
Multi-Selecthard

A developer is designing a Claude-powered assistant that must maintain a coherent conversation across many turns while controlling cost and context limits. Which TWO practices are appropriate according to Claude model fundamentals? (Choose two.)

Select 2 answers
A.Summarize or truncate older turns to keep the context within limits while preserving recent, relevant exchange.
B.Increase max_tokens on every request to allow Claude to store more conversation history internally.
C.Place durable persona and formatting rules in the system prompt so they persist across turns without repetition.
D.Rely on Claude to remember previous turns automatically between separate API calls.
E.Send the full conversation history with each request so Claude has complete context.
AnswersA, C

Because the API is stateless and context is finite, condensing or dropping older turns controls token growth and keeps requests within the context window. Summarization preserves important facts while reducing size, and retaining recent turns maintains conversational continuity. This is a recommended way to manage long conversations without losing coherence or driving up cost unnecessarily.

Why this answer

Because the Messages API is stateless, long conversations require deliberate context management. Summarizing or truncating older turns keeps requests within the context window and controls cost while preserving recent continuity. Durable persona, tone, and formatting rules belong in the system prompt, where they persist across turns without repetition.

Relying on automatic memory or misusing max_tokens misunderstands how Claude processes each independent request.

Exam trap

The trap here is believing Claude remembers prior turns automatically, when each Messages API call is stateless and history must be explicitly managed or it will be lost.

64
MCQmedium

A customer support team is building a Claude-powered assistant that must answer questions using a 300-page product manual. They want to avoid sending the entire manual with every request because of latency and cost. Which approach best leverages Claude's capabilities while keeping responses grounded in the manual?

A.Use retrieval-augmented generation: embed the manual, retrieve the most relevant passages, and include them in the prompt.
B.Lower the temperature to 0 so Claude reproduces the manual verbatim from memory.
C.Increase the max_tokens parameter so Claude can generate longer, more detailed answers directly from its training data.
D.Fine-tune Claude on the manual so the knowledge is embedded in the model weights.
AnswerA

RAG keeps the authoritative manual outside the model and injects only the top matching passages into the context window. This reduces token usage and latency, allows the manual to be updated without retraining, and gives Claude the exact source text to cite. It is the recommended pattern for grounding answers in a large, evolving document set such as a 300-page product manual.

Why this answer

Retrieval-augmented generation is the standard architecture for grounding Claude in a large, private, frequently updated document such as a product manual. Embedding the manual and retrieving only the most relevant passages keeps prompts small, reduces latency and cost, and gives Claude authoritative source text to answer from. Fine-tuning, max_tokens, and temperature do not supply the missing manual content and therefore cannot ensure grounded support answers.

Exam trap

The trap here is assuming that fine-tuning or a lower temperature can teach Claude a private 300-page manual, when grounding actually requires retrieving and injecting the relevant passages at request time.

65
MCQhard

In the context of the Anthropic API, what is the primary benefit of using Streaming?

A.It reduces the total number of tokens consumed.
B.It drastically increases the accuracy of the response.
C.It improves the perceived response time for users.
D.It allows the model to handle more input tokens.
AnswerC

Streaming provides immediate feedback by showing text as it is generated. This dramatically improves the perceived latency, as users can start reading or interacting with the answer as it appears, rather than staring at a loading spinner for the entire duration of the model's generation process.

Why this answer

Streaming allows the model's output to be delivered to the client piece-by-piece as it is generated, rather than waiting for the entire response to be finished. This significantly reduces the 'Time to First Token' (TTFT) perceived by the user, making applications feel much more responsive. This is a critical technique for real-time chat interfaces where waiting for a multi-paragraph response to finish would lead to an unacceptable user experience.

Exam trap

Candidates often incorrectly state that streaming 'reduces total token cost' or 'increases overall model intelligence', confusing the user-facing latency benefit with backend operational efficiency.

66
MCQeasy

Which capability is a primary benefit of using Claude 3.5 Sonnet compared to smaller, legacy models when processing complex, multi-step instructions?

A.It can be trained locally on customer-provided hardware.
B.It maintains higher instruction-following performance on complex, multi-step tasks.
C.It allows for unlimited parallel API requests without rate limiting.
D.It completely removes the need for systematic prompt engineering.
AnswerB

Claude 3.5 Sonnet is specifically optimized for advanced reasoning and instruction-following, allowing it to navigate complex, multi-step logic without losing track of constraints. This is a significant improvement over earlier models, making it ideal for tasks like code generation, complex document analysis, and multi-stage workflow execution in production environments.

Why this answer

Claude 3.5 Sonnet exhibits superior steerability and logical reasoning, allowing it to maintain consistency across long, multi-step prompts. This capability is vital for complex workflows where instructions are layered. Understanding model evolution helps developers choose the right tool for tasks that require high-level reasoning and instruction following, ensuring that complex business logic remains intact throughout the interaction, reducing the need for iterative prompting and debugging.

Exam trap

Candidates tend to think legacy models can handle layered logic just as well, underestimating how architectural evolution specifically improves complex multi-step reasoning.

67
MCQhard

An engineer is choosing between Claude 3.5 Sonnet and Claude 3 Opus for a pipeline that extracts structured fields from scanned invoices. The pipeline processes thousands of documents per hour and must balance accuracy against cost. Which statement best reflects the appropriate model-selection reasoning?

A.Claude 3 Opus should be used only for the first document and Sonnet for the rest, because Opus can teach Sonnet the schema at runtime.
B.Claude 3 Opus should always be chosen because it is the most capable model and therefore the most cost-effective at scale.
C.Model choice is irrelevant because all Claude 3 models have identical pricing and performance characteristics.
D.Claude 3.5 Sonnet is often the better balance because it provides strong accuracy on structured extraction at lower cost and higher throughput than Opus.
AnswerD

Claude 3.5 Sonnet delivers strong reasoning and extraction quality while costing less and offering better throughput than Opus. For high-volume invoice field extraction, that combination usually satisfies accuracy requirements without the premium price of Opus. This makes Sonnet the pragmatic default for a pipeline that must process thousands of documents per hour, with Opus reserved for the hardest edge cases.

Why this answer

Model selection should match capability to task and volume. Claude 3.5 Sonnet offers strong structured-extraction accuracy at a lower price and higher throughput than Claude 3 Opus, making it the sensible default for a high-volume invoice pipeline. Reserving Opus for the hardest cases balances quality and cost.

Treating all Claude 3 models as identical, or assuming runtime knowledge transfer between models, misstates how the API and model tiers actually work.

Exam trap

The trap here is defaulting to the most capable model for every task, when high-volume structured extraction usually favors a mid-tier model that balances accuracy with cost and throughput.

Ready to test yourself?

Try a timed practice session using only Claude Model Fundamentals questions.