A media company stores thousands of hours of unlabeled video footage and wants to build a searchable index that lets editors retrieve clips by describing their content in natural language. The team has no annotated dataset and no budget to label one. Which machine learning approach is the MOST appropriate starting point?
A pretrained multimodal model maps both video content and text descriptions into a common embedding space, so a natural-language query can be compared directly against indexed clips without any custom labels. This leverages existing pretrained knowledge, requires no annotation budget, and supports flexible search. It is the standard foundation for semantic video retrieval systems.
Why this answer
With no labels and a need to search footage using free-form language, a pretrained multimodal embedding model is the most practical foundation. It places video and text in a shared vector space so that descriptive queries match relevant clips by similarity, enabling semantic retrieval without any custom annotation effort.
Exam trap
The trap here is assuming that a large unlabeled video collection must be turned into a supervised classification dataset, when pretrained multimodal embeddings can support retrieval directly.