Courseiva
Data Analysis →mediumMultiple Select

DA0-002 Data Analysis Practice Question

A dataset contains outliers in a feature that will be used for linear regression. Which two outlier treatment methods are appropriate? (Choose TWO)

⚠ Common exam trap

DA0-002 often tests the confusion that normalization or mean-imputation 'handles' outliers, when in fact only capping, trimming, or robust methods (e.g., Huber regression) actually mitigate their influence.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cap the outliers at a percentile (e.g., 99th percentile)

Option A (Cap the outliers at a percentile, e.g., 99th percentile) is correct because winsorizing/capping replaces extreme values with a boundary value at a chosen percentile, which preserves the data points while reducing the leverage that extreme values exert on the least-squares fit, thereby limiting their disproportionate influence on the regression coefficients. Option D (Remove the outlier rows) is correct because deleting the offending observations eliminates their influence entirely, which is a standard and valid treatment when outliers are identified as erroneous or unrepresentative and the remaining sample is still adequate for fitting the model. Option B (Use min-max normalization) is not appropriate because rescaling features to a [0,1] range is a transformation that preserves the relative positions and leverage of outliers, so it does nothing to mitigate their effect on the regression. Option C (Increase the sample size) does not treat the outliers themselves; adding more data does not reduce the influence of existing extreme points and may even introduce more outliers. Option E (Replace outliers with the mean) is generally inappropriate because substituting the mean distorts the distribution, artificially reduces variance, and biases the regression estimates rather than properly handling the extreme values.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Cap the outliers at a percentile (e.g., 99th percentile)

    Why this is correct

    Capping at the 99th percentile winsorises extreme values to a boundary, retaining the observation while limiting its leverage on the regression coefficients. This preserves sample size and satisfies the need to reduce outlier influence without discarding data.

  • ✗

    Use min-max normalization

    Why it's wrong here

    Min-max normalization rescales values using the minimum and maximum, so a single extreme outlier compresses all other points into a narrow band; it does not remove or cap outliers. It is tempting because it is a standard preprocessing step for distance-based algorithms, but linear regression needs outliers treated, not rescaled.

  • ✗

    Increase the sample size

    Why it's wrong here

    Adding rows does not alter the existing extreme values, so the outliers still distort the least-squares fit and inflate residual variance. Increasing sample size is tempting when seeking statistical robustness, and it would help if the outliers were genuine rare events needing better representation, but it does not treat them.

  • ✓

    Remove the outlier rows

    Why this is correct

    Deleting outlier rows removes extreme values entirely, preventing them from skewing the ordinary least squares fit. This satisfies the linear regression assumption of no influential points, though it reduces sample size and must be justified by data quality.

  • ✗

    Replace outliers with the mean

    Why it's wrong here

    Mean substitution drags the mean toward the outlier and shrinks variance, distorting the regression coefficients the model estimates. It tempts as a quick fix, but it suits roughly symmetric distributions with mild anomalies; for genuine outliers, trimming, winsorising or median imputation preserves the linear relationship.

About these practice questions

This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.