Courseiva

AIF-C01 Applications of Foundation Models Practice Question

A company is building a multi-modal application that processes images and text to answer questions about product defects. Which foundation model approach is BEST?

⚠ Common exam trap

AWS often tests the misconception that combining two separate single-modal models (Option D) is equivalent to a true multi-modal model, but the trap is that late fusion lacks the joint embedding and cross-attention mechanisms needed for coherent multi-modal reasoning.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a multi-modal foundation model that processes both images and text

Multi-modal foundation models (e.g., CLIP, Flamingo, GPT-4V) are specifically designed to jointly process and reason over images and text in a unified architecture. This allows the model to directly correlate visual defects with textual descriptions without intermediate lossy transformations, making it the most effective approach for a multi-modal QA task.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use an image captioning model and then analyze the caption text

    Why it's wrong here

    Captioning collapses the image into a single text string, discarding spatial and visual detail needed to judge defects, and errors propagate into the text stage. It is tempting because it reuses a text model, but captioning suits search or accessibility indexing, not fine-grained visual question answering.

  • ✗

    Use a text-to-image generation model and analyze the generated image

    Why it's wrong here

    A text-to-image generator produces synthetic pictures from prompts; it cannot inspect a real photograph for defects. It is tempting because it handles both modalities, but generation runs in the opposite direction from analysis. The correct approach uses a multimodal model that accepts image and text input and emits an answer.

  • ✓

    Use a multi-modal foundation model that processes both images and text

    Why this is correct

    A multi-modal foundation model encodes images and text into a shared embedding space, letting one model jointly reason over both modalities. This directly satisfies the stem's requirement to answer defect questions from combined image and text input, unlike text-only or vision-only models that cannot correlate the two.

  • ✗

    Use a separate image analysis model and a text model, then combine outputs

    Why it's wrong here

    Chaining a separate image model and text model requires bespoke glue code and loses cross-modal grounding, so the model cannot jointly reason over pixels and words. It is tempting because combining specialist models works when the two inputs are genuinely independent, but defect questions need one model trained on aligned image-text pairs.

About these practice questions

One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.