AI0-001 AI Models and Data Engineering Practice Question
A media company is building a recommendation model from user clickstream logs. The raw data arrives as millions of small JSON files in an object store, and nightly training jobs currently take over ten hours because the training cluster reads thousands of tiny files per second. The team wants to reduce training time without changing the model or the underlying data values. Which data engineering approach is most appropriate?
⚠ Common exam trap
The trap here is assuming that adding more compute or switching file formats without consolidating files will fix an I/O-bound small-file problem.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Compact the small JSON files into larger columnar files such as Parquet and have the training pipeline read those instead.
The bottleneck is the small-file pattern in object storage, not compute capacity or parsing speed. Repacking the same records into fewer large columnar files cuts listing, open, and read overhead and enables column pruning and compression, so the training job moves the same data values with far less I/O. This addresses the root cause while preserving the model and data semantics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert the JSON files to CSV because CSV parsing is faster than JSON parsing.
Why it's wrong here
CSV parsing can be marginally faster per record, but converting millions of small JSON objects into millions of small CSV objects leaves the dominant problem unchanged: too many individual files and object store requests. The team would also lose nested structure in the JSON clickstream events, which can alter the feature set the model expects.
- ✗
Increase the number of worker nodes in the training cluster so more files can be read in parallel.
Why it's wrong here
Adding workers increases aggregate read throughput, but each worker still pays the fixed cost of opening, listing, and parsing tiny JSON objects, and the object store may throttle request rates. This scales cost linearly while leaving the small-file bottleneck intact, and the nightly job may still exceed its window even with a much larger cluster.
- ✗
Cache the JSON files on local SSD storage on each worker node before training begins.
Why it's wrong here
Local caching can help on repeated epochs, but the initial population of the cache still requires reading every small file from the object store, so the first epoch remains slow. Cache capacity on worker nodes is also limited, and the approach adds operational complexity without fixing the underlying small-file layout problem.
- ✓
Compact the small JSON files into larger columnar files such as Parquet and have the training pipeline read those instead.
Why this is correct
Consolidating many small JSON files into fewer larger Parquet files preserves all the underlying records while dramatically reducing per-file overhead, metadata operations, and list calls against the object store. Columnar layout also lets the training job read only the columns it needs and enables efficient compression, so the same data values feed the model with far less I/O time.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.