CCAR-F Prompt Engineering and Structured Output Practice Question
A platform team ships a Claude-based service that classifies incoming tickets into one of eight categories and returns a JSON object with category and confidence. In production they observe two failure modes: the model sometimes returns a category outside the eight allowed values, and occasionally returns confidence as the string "high" instead of a number. Which TWO changes most directly reduce these failures while keeping the pipeline automated? (Choose two.)
⚠ Common exam trap
The trap here is believing that a clear prose instruction alone fixes type and vocabulary drift, when only schema constraints plus validation reliably enforce them.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Add a validation step that parses the JSON, checks category against the allowed list and confidence against a numeric range, and re-prompts with the specific error when either check fails.
Schema-constrained tool input directly prevents out-of-set categories and non-numeric confidence by enumerating and typing the fields, while a validating re-prompt loop catches any residual drift and feeds the exact error back for correction. Together they address both failure modes structurally. Sampling and length parameters do not constrain vocabulary or type, and unverified prose instructions are too weak for production.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Add a validation step that parses the JSON, checks category against the allowed list and confidence against a numeric range, and re-prompts with the specific error when either check fails.
Why this is correct
A post-generation validator that inspects category membership and confidence type catches exactly the two failure modes and, by feeding the concrete error back on a retry, gives the model a targeted correction. This closes the loop so invalid outputs never reach downstream systems, and it works even when the model ignores schema hints. It complements schema constraints by handling residual drift.
- ✓
Define a tool whose input_schema enumerates the eight categories for the category property and types confidence as a number, then read the structured tool input.
Why this is correct
An input_schema with an enum for category and a numeric type for confidence constrains the structured tool input to the allowed vocabulary and data types, directly preventing both observed failures. The model must fill the schema fields, so an out-of-set category or a string confidence cannot appear in the parsed tool input. This is the most direct structural fix for both symptoms.
- ✗
Ask the model in the prompt to 'be careful and always use one of the eight categories', and trust the instruction without further checks.
Why it's wrong here
Polite prose instructions are soft guidance the model can still violate, especially under ambiguous tickets, so they do not reliably eliminate out-of-set categories or string confidences. Relying on them without validation leaves the pipeline exposed to the same failures. This option lacks the structural enforcement that the two observed defects demand, making it insufficient on its own.
- ✗
Increase max_tokens so the model has more room to produce the JSON object without truncation.
Why it's wrong here
Truncation would cut off the JSON entirely, but the reported failures are a wrong category value and a string instead of a number, which are semantic and type errors rather than length problems. Raising max_tokens does nothing to restrict the category vocabulary or enforce numeric confidence. This change addresses an unrelated failure mode and leaves both observed issues untouched.
- ✗
Raise the temperature to 0.7 so the model explores more phrasing and is less likely to repeat a bad category.
Why it's wrong here
Higher temperature increases variability, which makes out-of-set categories and type drift more likely, not less. Classification into a fixed label set benefits from low-temperature, constrained decoding rather than exploration. This option moves in the opposite direction from what the failure modes require and would likely worsen both the category and confidence problems.
About these practice questions
One of 271 original CCAR-F practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-F exam.