DA0-002 Data Acquisition and Preparation Practice Question
An analyst needs to combine two datasets from different sources that share a common key but have different levels of granularity. Dataset A has daily sales per store, Dataset B has hourly foot traffic per store. The analyst wants to analyze correlation. Which approach is appropriate?
⚠ Common exam trap
CompTIA often tests the misconception that disaggregating (splitting) the coarser dataset is acceptable, but this introduces artificial data and violates the assumption of uniform distribution, whereas aggregation preserves the actual measured values.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Aggregate Dataset B to daily level before merging
Aggregating Dataset B (hourly foot traffic) to the daily level ensures both datasets share the same granularity before merging on the common key (store and date). This allows a valid correlation analysis between daily sales and daily foot traffic without introducing artificial patterns or data duplication. Merging at mismatched granularities would violate the assumption that each row represents a comparable unit of observation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Aggregate Dataset B to daily level before merging
Why this is correct
Correlation requires matching granularity; hourly foot traffic cannot align one-to-one with daily sales rows. Aggregating Dataset B to daily totals creates a shared daily key per store, enabling a valid join and meaningful correlation analysis across the two sources.
- ✗
Use an outer join and keep all rows
Why it's wrong here
An outer join on store and date matches each daily row against every hourly row for that day, producing a many-to-many cartesian expansion that duplicates sales figures. Aggregating foot traffic to daily totals first gives one row per store-day. Outer joins suit preserving unmatched records, not reconciling granularity.
- ✗
Disaggregate Dataset A to hourly level by dividing daily sales by hours
Why it's wrong here
Dividing daily sales evenly across hours fabricates uniform intra-day distribution that the data does not contain, distorting any correlation with hourly foot traffic. Aggregating Dataset B to daily totals preserves real values. Disaggregation suits cases where a known hourly weighting exists.
- ✗
Join on store and date without aggregation
Why it's wrong here
Joining on store and date alone leaves daily sales repeated across each of the 24 hourly rows, so correlation is computed on duplicated values and is meaningless. Summing hourly foot traffic to daily level yields one row per store-day. Unaggregated joins suit same-granularity datasets.
Go deeper
Related to this question
About these practice questions
This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.