DA0-002 Data Analysis Practice Question
A data analyst at a healthcare organization is analyzing patient readmission rates. The analyst has a dataset with patient demographics, diagnosis codes, and length of stay. Before performing any statistical modeling, the analyst must address data quality issues. Which TWO of the following actions are most appropriate for ensuring the dataset is ready for analysis? (Choose two.)
⚠ Common exam trap
The trap here is assuming that any missing data must be removed or imputed with a simple mean, when in fact proper handling depends on the missingness mechanism and the analysis goal.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Identify and remove duplicate patient records based on a unique patient identifier and admission date.
Ensuring data quality involves identifying and correcting issues that could bias results. Removing duplicate patient records prevents overcounting admissions, and standardizing diagnosis codes ensures consistency for analysis. Mean imputation can distort data, binning numerical variables is not a quality fix, and listwise deletion can introduce bias and reduce sample size. These two actions directly address common data quality problems in healthcare datasets.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Remove all records with any missing values to ensure a complete dataset.
Why it's wrong here
Listwise deletion can introduce bias if missingness is not completely at random and can drastically reduce sample size. In healthcare, missing data is often informative; removing records may exclude important patient subgroups. A better approach is to assess the pattern of missingness and apply appropriate imputation or sensitivity analysis. Thus, this action is not recommended as a general data quality step.
- ✓
Identify and remove duplicate patient records based on a unique patient identifier and admission date.
Why this is correct
Duplicate records can skew readmission rates and other analyses. Using a unique patient identifier combined with admission date helps distinguish between multiple admissions for the same patient versus true duplicates. Removing duplicates ensures each admission is counted once, which is critical for accurate readmission metrics. This is a standard data cleaning step and directly addresses a common data quality issue in healthcare datasets.
- ✗
Convert all numerical columns to categorical bins to simplify the analysis.
Why it's wrong here
Converting numerical columns to categorical bins loses information and can obscure relationships. For readmission analysis, continuous variables like length of stay or age may have important nonlinear relationships with the outcome. Binning should be done thoughtfully, not as a blanket data quality step. This action does not address data quality issues and could harm the analysis.
- ✗
Impute missing values in the length of stay column using the mean length of stay for all patients.
Why it's wrong here
Mean imputation is a common technique, but it can distort the distribution and reduce variability, especially if missingness is not random. In healthcare data, length of stay often varies by diagnosis, so a global mean may introduce bias. A more appropriate approach would consider grouping by diagnosis or using a model-based imputation. Therefore, this action is not the best choice for ensuring data quality in this scenario.
- ✓
Standardize diagnosis codes to a consistent format (e.g., ICD-10) and validate against a reference list.
Why this is correct
Diagnosis codes may be entered in different formats or with typos, leading to inconsistent grouping. Standardizing to a single coding system like ICD-10 and validating against a reference list ensures that codes are accurate and comparable. This improves the reliability of any analysis involving diagnoses, such as risk adjustment or comorbidity indices. It is a fundamental data quality step for healthcare data.
Go deeper
Related to this question
About these practice questions
This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.