Courseiva
Data Analysis →mediumMultiple Choice

DA0-002 Data Analysis Practice Question

A data analyst is examining a dataset of customer orders and notices that the 'order_total' column contains values ranging from $5 to $50,000, with most orders under $200. The analyst wants to reduce the influence of extreme values on a linear regression model predicting customer lifetime value. Which technique should the analyst apply?

⚠ Common exam trap

The trap here is assuming that standardization alone will handle outliers, when it actually preserves their relative extremity.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply a logarithmic transformation to the order_total variable.

The logarithmic transformation is effective because it compresses the range of large values, making the distribution more symmetric and reducing the influence of outliers on the regression model. Standardization, removal, or binning either do not address the skewness or introduce other issues such as information loss or bias.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert order_total to a categorical variable by binning into deciles.

    Why it's wrong here

    Binning a continuous variable into categories loses information and assumes a non-linear step relationship, which may not reflect the true effect. While it can reduce outlier impact, it is less precise than a transformation like log for mitigating skewness in regression.

  • ✗

    Standardize the order_total variable to have a mean of 0 and standard deviation of 1.

    Why it's wrong here

    Standardization rescales the variable but does not reduce the influence of outliers; extreme values remain extreme in terms of z-scores. The shape of the distribution is unchanged, so a single $50,000 order still exerts disproportionate leverage on the regression coefficients.

  • ✗

    Remove all orders with a total greater than $1,000 from the dataset.

    Why it's wrong here

    Removing high-value orders arbitrarily discards legitimate data and can introduce bias, especially if high spenders are a meaningful customer segment. This approach is not a principled transformation and could distort the model's predictions for valuable customers.

  • ✓

    Apply a logarithmic transformation to the order_total variable.

    Why this is correct

    A logarithmic transformation compresses the scale of large values, reducing the leverage of extreme high order totals on the regression model. This makes the relationship between order_total and the target more linear and stabilizes variance, which is appropriate when data is positively skewed with outliers.

About these practice questions

One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.