DA0-002 Data Acquisition and Preparation Practice Question
A data engineer is ingesting a 40 GB JSON event log into a columnar analytics platform. The file contains deeply nested arrays of user actions, and queries only ever filter on three top-level fields: event_id, event_type, and event_timestamp. The ingestion is currently slow and queries scan excessive data. Which preparation approach is MOST appropriate?
⚠ Common exam trap
The trap here is treating JSON flattening as all-or-nothing and either parsing at query time or fully normalizing, instead of selectively promoting the columns the workload actually filters on.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Flatten the top-level fields into typed columns, extract the nested arrays into a separate child table, and partition or cluster on event_timestamp.
Matching the storage layout to the access pattern is the decisive factor. Promoting the three filtered scalars to typed columns enables column pruning and partition pruning, while relocating the unused nested arrays keeps the hot table narrow. Parsing strings or fully normalizing both impose costs the workload does not justify.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Load the JSON as a single string column and rely on the query engine to parse it at read time.
Why it's wrong here
Keeping the payload as an opaque string forces every query to parse JSON at execution time, which wastes CPU and prevents the columnar engine from pruning data by the three filter fields. Predicate pushdown becomes impossible because the engine cannot see inside the string. This approach also retains the full nested structure that the workload never uses, so scan volume stays high.
- ✓
Flatten the top-level fields into typed columns, extract the nested arrays into a separate child table, and partition or cluster on event_timestamp.
Why this is correct
Extracting the frequently filtered scalars into native typed columns lets the columnar engine read only those columns and apply partition pruning on event_timestamp. Moving the rarely queried nested arrays into a child table keeps the main fact table narrow and fast. This matches the access pattern precisely and reduces both ingestion cost and scan volume.
- ✗
Store the file in a row-oriented relational table with indexes on all nested array paths.
Why it's wrong here
Row-oriented storage reads whole rows even when only three columns are needed, defeating the columnar platform's main advantage. Indexing nested array paths is expensive to maintain and irrelevant because the workload never filters on those paths. This choice increases storage and write overhead while leaving the scan problem unresolved.
- ✗
Normalize every nested array element into its own row in a fully relational schema with foreign keys.
Why it's wrong here
Full normalization explodes the event grain into many rows per event and adds join overhead to every query, which is counterproductive for an analytics workload that only needs three scalar fields. It solves a transactional-integrity problem that is not present here. The extra joins and row multiplication would likely make the slow queries even slower.
Go deeper
Related to this question
About these practice questions
One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.