DA0-002 Data Analysis Practice Question
Which THREE of the following are appropriate methods to handle outliers in a dataset?
⚠ Common exam trap
Candidates often confuse data preprocessing techniques like binning or imputation with outlier handling methods, but binning is for discretization and mean imputation is not robust for outliers, while the correct methods (transformation, removal, capping) directly address outlier impact.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Transforming the data using log transformation
Log transformation (A) is correct because applying a log (or similar power) transform compresses the scale of extreme values, reducing the leverage of right-skewed outliers and making the distribution more symmetric for downstream models. Removing the outlier records (B) is correct when outliers are confirmed to be data-entry errors or otherwise invalid observations, since deleting them prevents distortion of statistics such as the mean and variance. Capping the outlier values at a certain percentile (C), also called winsorizing, is correct because it clips extreme values to a chosen percentile boundary (e.g., the 1st and 99th percentiles), preserving the record while limiting the outlier's influence. Binning continuous variables (D) is not a dedicated outlier-handling method; it discretizes values into bins and merely masks extremes rather than addressing them. Imputing outliers with the mean (E) is inappropriate because the mean itself is distorted by the outliers, so replacing them with it introduces bias and does not resolve the underlying problem.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Transforming the data using log transformation
Why this is correct
A log transformation compresses the upper tail of a right-skewed distribution, pulling extreme values closer to the bulk of the data. This reduces the leverage outliers exert on model estimates without deleting records, making it an appropriate handling method.
- ✓
Removing the outlier records
Why this is correct
Deleting outlier records removes their distorting influence on statistics such as the mean and standard deviation, satisfying the requirement to handle extreme values. This suits cases where outliers stem from data-entry errors, though it sacrifices sample size and risks discarding legitimate rare observations.
- ✓
Capping the outlier values at a certain percentile
Why this is correct
Capping replaces values beyond a chosen percentile with that boundary, limiting an outlier's leverage on means and models while retaining every record. This satisfies the handling requirement without discarding data, which matters when sample size is small or outliers are genuine.
- ✗
Binning continuous variables
Why it's wrong here
Binning converts continuous values into discrete intervals, which conceals extreme values rather than resolving them, so outliers remain inside the bins and skew nothing visibly. It is tempting because binning is genuinely used to smooth noisy data or reduce granularity for modelling, and it would suit a scenario needing categorical bands rather than outlier treatment.
- ✗
Imputing outliers with the mean
Why it's wrong here
Imputing outliers with the mean retains their distorted values and pulls the mean itself, biasing results; outliers should be removed, capped or transformed instead. Mean imputation is correct for missing values under a missing-at-random assumption, not for extreme observations.
Visual reference
Go deeper
Related to this question
About these practice questions
This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.