AIF-C01 Fundamentals of AI and ML Practice Question
A media company stores thousands of hours of unlabeled video footage and wants to build a searchable index that lets editors retrieve clips by describing their content in natural language. The team has no annotated dataset and no budget to label one. Which machine learning approach is the MOST appropriate starting point?
⚠ Common exam trap
The trap here is assuming that a large unlabeled video collection must be turned into a supervised classification dataset, when pretrained multimodal embeddings can support retrieval directly.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a pretrained multimodal embedding model to encode video frames and text into a shared vector space, then retrieve clips by similarity.
With no labels and a need to search footage using free-form language, a pretrained multimodal embedding model is the most practical foundation. It places video and text in a shared vector space so that descriptive queries match relevant clips by similarity, enabling semantic retrieval without any custom annotation effort.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a pretrained multimodal embedding model to encode video frames and text into a shared vector space, then retrieve clips by similarity.
Why this is correct
A pretrained multimodal model maps both video content and text descriptions into a common embedding space, so a natural-language query can be compared directly against indexed clips without any custom labels. This leverages existing pretrained knowledge, requires no annotation budget, and supports flexible search. It is the standard foundation for semantic video retrieval systems.
- ✗
Train a supervised video classifier from scratch using the footage, treating each video file as its own class.
Why it's wrong here
Treating each video file as a separate class would create thousands of classes with a single example each, which provides no meaningful signal for a supervised classifier to generalize. The footage is unlabeled and the goal is open-ended retrieval, not assigning files to fixed categories. This approach would also require substantial compute while failing to support natural-language queries.
- ✗
Apply k-means clustering to raw pixel values of every frame to produce searchable cluster identifiers.
Why it's wrong here
Clustering raw pixels groups frames by low-level visual statistics such as color and brightness, not by semantic content, so an editor searching for a described scene would not get useful matches. Cluster identifiers also carry no language meaning, making natural-language querying impossible. This approach ignores the rich pretrained representations that make semantic retrieval feasible.
- ✗
Build a reinforcement learning agent that learns to select the clip most likely to satisfy an editor through trial and error.
Why it's wrong here
Reinforcement learning needs a well-defined environment, actions, and a reward signal that can be sampled repeatedly. Here there is no simulator for editor satisfaction and no reward to optimize, so the agent has nothing to learn from. The retrieval goal is better served by similarity search over embeddings, which requires no interactive training loop.
Go deeper
Related to this question
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.