AIF-C01 Applications of Foundation Models Practice Question
A company is building a multi-modal application that processes images and text to answer questions about product defects. Which foundation model approach is BEST?
⚠ Common exam trap
AWS often tests the misconception that combining two separate single-modal models (Option D) is equivalent to a true multi-modal model, but the trap is that late fusion lacks the joint embedding and cross-attention mechanisms needed for coherent multi-modal reasoning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a multi-modal foundation model that processes both images and text
Multi-modal foundation models (e.g., CLIP, Flamingo, GPT-4V) are specifically designed to jointly process and reason over images and text in a unified architecture. This allows the model to directly correlate visual defects with textual descriptions without intermediate lossy transformations, making it the most effective approach for a multi-modal QA task.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use an image captioning model and then analyze the caption text
Why it's wrong here
Captioning collapses the image into a single text string, discarding spatial and visual detail needed to judge defects, and errors propagate into the text stage. It is tempting because it reuses a text model, but captioning suits search or accessibility indexing, not fine-grained visual question answering.
- ✗
Use a text-to-image generation model and analyze the generated image
Why it's wrong here
A text-to-image generator produces synthetic pictures from prompts; it cannot inspect a real photograph for defects. It is tempting because it handles both modalities, but generation runs in the opposite direction from analysis. The correct approach uses a multimodal model that accepts image and text input and emits an answer.
- ✓
Use a multi-modal foundation model that processes both images and text
Why this is correct
A multi-modal foundation model encodes images and text into a shared embedding space, letting one model jointly reason over both modalities. This directly satisfies the stem's requirement to answer defect questions from combined image and text input, unlike text-only or vision-only models that cannot correlate the two.
- ✗
Use a separate image analysis model and a text model, then combine outputs
Why it's wrong here
Chaining a separate image model and text model requires bespoke glue code and loses cross-modal grounding, so the model cannot jointly reason over pixels and words. It is tempting because combining specialist models works when the two inputs are genuinely independent, but defect questions need one model trained on aligned image-text pairs.
Go deeper
Related to this question
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.