DA0-002 Data Acquisition and Preparation Practice Question
A data analyst is performing data acquisition from multiple source files. Which TWO data profiling tasks should the analyst complete before loading the data into the target system?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Verify data types and formats
Option B (Verify data types and formats) is correct because data profiling must confirm that each source column's actual data type (e.g., integer, string, date) and format (e.g., YYYY-MM-DD, decimal separators) match what the target system expects, preventing load failures or silent corruption. Option E (Identify missing values and nulls) is correct because profiling must detect NULLs, empty strings, and placeholder values so the analyst can decide on imputation, defaulting, or rejection rules before loading. Option A (Create a dashboard for stakeholders) is a downstream reporting activity, not a pre-load profiling task. Option C (Build a linear regression model) is predictive analytics performed after data is cleansed and loaded, not profiling. Option D (Perform cluster analysis) is an unsupervised modeling technique also done post-load, so it does not belong in pre-load data profiling.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create a dashboard for stakeholders
Why it's wrong here
A dashboard presents aggregated results to stakeholders, producing no metadata about null counts, key uniqueness or data types across the source files. Dashboards are the right choice once data is loaded and modelled, not during acquisition, when the analyst must first understand each file's structure and quality.
- ✓
Verify data types and formats
Why this is correct
Verifying data types and formats confirms each source column matches the target schema, catching mismatches such as strings in numeric fields or inconsistent date layouts before load. This satisfies the pre-load profiling requirement, preventing type-conversion failures or silent corruption during ingestion into the target system.
- ✗
Build a linear regression model
Why it's wrong here
Regression modelling predicts a continuous target from historical data, so it cannot assess source-file completeness, duplicates or value distributions before loading. It is tempting because regression is a core analytics technique, but it belongs after profiling, when relationships between cleaned variables are being quantified.
- ✗
Perform cluster analysis
Why it's wrong here
Cluster analysis groups records by similarity, which requires already-loaded, cleansed and scaled data; it reveals nothing about column-level completeness, duplicates or format inconsistencies. Clustering is correct later for segmentation or anomaly discovery, but profiling must precede any such modelling.
- ✓
Identify missing values and nulls
Why this is correct
Identifying missing values and nulls reveals gaps that would violate target constraints or skew aggregates once loaded. Profiling these before ingestion satisfies the pre-load requirement, letting the analyst decide whether to impute, default or reject affected records rather than discovering the problem post-load.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.