AIF-C01 Applications of Foundation Models Practice Question
A media company wants to automatically generate short video captions from uploaded audio files using a foundation model on AWS. The solution should transcribe speech and then produce concise captions. Which combination of AWS services should they use?
⚠ Common exam trap
Candidates often confuse text-to-speech with speech-to-text, or assuming a foundation model can accept raw audio input directly.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Amazon Transcribe to convert audio to text, then Amazon Bedrock to generate captions from the transcript.
The task requires two capabilities: speech-to-text and generative text creation. Amazon Transcribe provides accurate automatic speech recognition, and Amazon Bedrock supplies foundation models that turn the transcript into concise captions. Together they form a managed, scalable pipeline that meets the scenario without custom model development.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Amazon Kinesis Data Streams to ingest audio, then Amazon Bedrock to generate captions directly from the audio stream.
Why it's wrong here
Kinesis Data Streams is for real-time data streaming and does not perform speech-to-text conversion. Amazon Bedrock foundation models accept text and image inputs, not raw audio, so they cannot generate captions directly from an audio stream without a transcription step. This design omits the required speech recognition component and would not produce usable captions.
- ✗
Amazon Rekognition to analyze the audio, then Amazon Bedrock to generate captions.
Why it's wrong here
Amazon Rekognition analyzes images and video for objects, faces, and activities; it is not designed for speech transcription from audio files. Using it as the first step would not produce an accurate transcript for the foundation model to work from. The scenario requires converting speech to text, which calls for a dedicated speech recognition service such as Amazon Transcribe.
- ✗
Amazon Polly to convert audio to text, then Amazon Comprehend to generate captions.
Why it's wrong here
Amazon Polly is a text-to-speech service that turns written text into lifelike audio, which is the reverse of what is needed. Amazon Comprehend performs natural language processing tasks such as entity and sentiment detection, not generative caption writing. This combination cannot transcribe uploaded audio or produce creative caption text, so it fails the scenario requirements.
- ✓
Amazon Transcribe to convert audio to text, then Amazon Bedrock to generate captions from the transcript.
Why this is correct
Amazon Transcribe performs automatic speech recognition to convert audio into text, and Amazon Bedrock provides access to foundation models that can summarize or rewrite that transcript into concise captions. This two-step pipeline matches the requirement: transcription followed by generative caption creation. Both are managed AWS services, so the company avoids building and operating its own speech and language models.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.