AI0-001 AI Models and Data Engineering Practice Question
A data engineer needs to store training data in a format that supports columnar pruning during model training. Which storage format should they use?
⚠ Common exam trap
CompTIA often tests the misconception that JSON or CSV are acceptable for columnar pruning because they are common and human-readable, but the trap here is that only columnar formats like Parquet or ORC support efficient column-level access, while row-oriented formats require full record scans.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Parquet
Parquet is the correct choice because it is a columnar storage format that enables column pruning, allowing the training process to read only the columns needed for model training rather than entire rows. This reduces I/O and speeds up data loading, which is critical for large-scale AI/ML workloads. Unlike row-oriented formats, Parquet stores data by columns, making it efficient for analytical queries and feature selection.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Parquet
Why this is correct
Parquet stores data column-wise with per-column statistics and row-group metadata, letting readers skip irrelevant columns and row groups entirely. This columnar pruning reduces I/O during training, directly meeting the stem's requirement for efficient columnar access.
- ✗
XML
Why it's wrong here
XML nests tagged elements per record, so column pruning requires parsing the full document tree; there are no per-column chunks or min/max metadata to skip. It suits hierarchical document interchange with schema validation, not wide analytical training tables.
- ✗
JSON
Why it's wrong here
JSON stores semi-structured key-value records; readers must parse each object to reach a field, and no column statistics permit skipping. It fits API payloads and nested document exchange, where schema flexibility outweighs columnar scan performance.
- ✗
CSV
Why it's wrong here
CSV stores rows as delimited plain text, so reading any subset of columns still requires parsing every field in each record; no column-level statistics exist to skip data. It is tempting for quick exports and human-readable interchange, where universal tool support matters more than scan efficiency.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.