Courseiva
Data Analysis →hardMultiple Select

DA0-002 Data Analysis Practice Question

Which THREE of the following are appropriate methods to handle outliers in a dataset?

⚠ Common exam trap

Candidates often confuse data preprocessing techniques like binning or imputation with outlier handling methods, but binning is for discretization and mean imputation is not robust for outliers, while the correct methods (transformation, removal, capping) directly address outlier impact.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Transforming the data using log transformation

Log transformation (A) is correct because applying a log (or similar power) transform compresses the scale of extreme values, reducing the leverage of right-skewed outliers and making the distribution more symmetric for downstream models. Removing the outlier records (B) is correct when outliers are confirmed to be data-entry errors or otherwise invalid observations, since deleting them prevents distortion of statistics such as the mean and variance. Capping the outlier values at a certain percentile (C), also called winsorizing, is correct because it clips extreme values to a chosen percentile boundary (e.g., the 1st and 99th percentiles), preserving the record while limiting the outlier's influence. Binning continuous variables (D) is not a dedicated outlier-handling method; it discretizes values into bins and merely masks extremes rather than addressing them. Imputing outliers with the mean (E) is inappropriate because the mean itself is distorted by the outliers, so replacing them with it introduces bias and does not resolve the underlying problem.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Transforming the data using log transformation

    Why this is correct

    A log transformation compresses the upper tail of a right-skewed distribution, pulling extreme values closer to the bulk of the data. This reduces the leverage outliers exert on model estimates without deleting records, making it an appropriate handling method.

  • ✓

    Removing the outlier records

    Why this is correct

    Deleting outlier records removes their distorting influence on statistics such as the mean and standard deviation, satisfying the requirement to handle extreme values. This suits cases where outliers stem from data-entry errors, though it sacrifices sample size and risks discarding legitimate rare observations.

  • ✓

    Capping the outlier values at a certain percentile

    Why this is correct

    Capping replaces values beyond a chosen percentile with that boundary, limiting an outlier's leverage on means and models while retaining every record. This satisfies the handling requirement without discarding data, which matters when sample size is small or outliers are genuine.

  • ✗

    Binning continuous variables

    Why it's wrong here

    Binning converts continuous values into discrete intervals, which conceals extreme values rather than resolving them, so outliers remain inside the bins and skew nothing visibly. It is tempting because binning is genuinely used to smooth noisy data or reduce granularity for modelling, and it would suit a scenario needing categorical bands rather than outlier treatment.

  • ✗

    Imputing outliers with the mean

    Why it's wrong here

    Imputing outliers with the mean retains their distorted values and pulls the mean itself, biasing results; outliers should be removed, capped or transformed instead. Mean imputation is correct for missing values under a missing-at-random assumption, not for extreme observations.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.