DA0-002 Data Acquisition and Preparation Practice Question
A data analyst is tasked with collecting data from multiple spreadsheets provided by different departments. Each spreadsheet has different column names and formats. What is the best first step?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Develop a data dictionary and standardize column names
Developing a data dictionary and standardizing column names ensures consistency across all data sources before loading, reducing errors and facilitating integration. Immediately loading data can cause inconsistencies. Discarding mismatched data loses potentially valuable information. Using a machine learning model is an unnecessary and complex first step.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Develop a data dictionary and standardize column names
Why this is correct
Differing column names and formats across departmental spreadsheets create schema conflicts that must be resolved before any merging or analysis. A data dictionary documents each source field and its meaning, enabling analysts to map and standardise names consistently, satisfying the stem's requirement for a first step addressing heterogeneous source structures.
- ✗
Discard any mismatched data
Why it's wrong here
Discarding mismatched data destroys records that may hold required values, creating gaps and bias before any reconciliation. It is tempting because removing inconsistent rows appears to simplify cleaning, but the correct first step is profiling and mapping the differing columns and formats so all department data is retained.
- ✗
Use a machine learning model to clean data
Why it's wrong here
A machine learning model requires clean, labelled training data and cannot resolve unknown schema differences on its own. It is tempting because automated cleaning promises speed, but the first step is manual profiling and mapping of column names and formats, which the model cannot infer from inconsistent sources.
- ✗
Immediately load all data into a database
Why it's wrong here
Loading raw, inconsistent spreadsheets into a database propagates mismatched column names and formats into the target schema. It is tempting because centralising data is the eventual goal, but the necessary first step is profiling and standardising the sources so their schemas align before any load occurs.
Go deeper
Related to this question
About these practice questions
This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.