AI0-001 AI Infrastructure and Technologies Practice Question
A media company stores thousands of hours of raw broadcast footage in a cloud object storage bucket. A data engineering team needs a training dataset that contains only the short clips where a goal is scored, so they must locate and extract those specific time ranges from the video files before training. Which technology should the team use to extract the required segments from the video objects?
⚠ Common exam trap
The trap here is assuming that a data storage or table format can perform media manipulation, when extracting video segments requires a codec-aware media tool.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use FFmpeg to decode the video and cut the required time ranges into new clip files.
Trimming video to precise time ranges is a media-processing task, and FFmpeg is the purpose-built tool for demuxing, decoding, and re-encoding streams with start and end timestamps. Object storage and table formats organize files but cannot parse video containers, so the extraction must be performed by a codec-aware utility before the clips enter the training pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use FFmpeg to decode the video and cut the required time ranges into new clip files.
Why this is correct
FFmpeg is the standard open-source toolkit for demuxing, decoding, and re-encoding media streams, and it supports precise time-based trimming with parameters such as -ss and -to. Running it over the objects lets the team extract exactly the goal segments and write new clip files that become clean training data for the model pipeline.
- ✗
Use a data lakehouse table format with schema evolution to store the video files as Delta tables.
Why it's wrong here
Lakehouse table formats such as Delta Lake manage tabular data, transactions, and schema changes; they do not decode video containers or seek to specific timestamps inside a media file. Storing the footage as a Delta table does not produce trimmed clips, so the team still lacks the goal-only training dataset this scenario requires.
- ✗
Use OpenCV's VideoCapture with a GPU-accelerated codec to retrain the object detection model directly on the raw footage.
Why it's wrong here
VideoCapture reads frames sequentially for analysis, and retraining a detector does not create trimmed clips. The scenario asks for extraction of specific time ranges into a dataset, not for model training. Even GPU-accelerated decoding leaves the team with frames in memory rather than the goal-only clip files they need.
- ✗
Use Apache Parquet to store each frame as a column and filter on the goal timestamp column.
Why it's wrong here
Parquet is a columnar storage format for structured records, not a media processing library. It cannot open an MP4 container, decode H.264 frames, or split a stream at a timestamp. Writing frames as Parquet columns would require the frames to already be decoded and extracted, which is precisely the work that must happen first.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.