Courseiva

CCAO-F Prompting and Context Engineering Practice Question

When engineering a prompt for a multimodal model like Claude 3.5 Sonnet that includes both text and images, what is the recommended way to handle the relationship between the two types of content?

⚠ Common exam trap

Candidates frequently upload images without explicit textual references, assuming the model will automatically link the visual data to specific parts of the prompt, leading to vague or disconnected analysis.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use text to explicitly refer to image content, such as 'In the first image...'.

Multimodal prompting requires clear associations between text and visual data. Placing the text instructions near the images they refer to, and using descriptive language to link them, helps the model understand the spatial and semantic relationships between the visual elements and the task it is being asked to perform.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Place all images at the very end of the message after all text instructions.

    Why it's wrong here

    While this can work, it isn't always the 'best' practice. For prompts with multiple images and specific instructions for each, interleaving text and images—or using XML tags to label images—is often more effective because it provides the model with immediate context for what each image represents.

  • ✓

    Use text to explicitly refer to image content, such as 'In the first image...'.

    Why this is correct

    Explicitly referencing the images in your text (e.g., 'Look at the chart in Image 1 and compare it to the table in Image 2') is a best practice. This creates a strong link between the visual and textual data, ensuring the model knows exactly which image to analyze for a given instruction.

  • ✗

    Always convert images to text descriptions before sending them to the model.

    Why it's wrong here

    This is unnecessary and defeats the purpose of using a multimodal model. Claude is designed to 'see' and analyze images directly. Converting them to text is not only time-consuming but often loses critical visual information—like spatial layout or subtle color cues—that the model could have used directly.

  • ✗

    Provide the images in the system prompt to establish them as permanent context.

    Why it's wrong here

    Anthropic's current API does not support images in the system prompt; they must be provided in the 'messages' array. Even if it were supported, images are typically part of a specific user query and belong in the user message where they can be properly associated with specific instructions.

About these practice questions

One of 259 original CCAO-F practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCAO-F practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAO-F exam.