Courseiva

DA0-002 Data Acquisition and Preparation Practice Question

An analyst is profiling a newly acquired customer table and wants a quick summary of the central tendency and spread of the 'annual_income' column, which contains a few extreme outliers from data entry errors. Which combination of descriptive statistics is most appropriate to report the typical income while limiting the influence of those outliers?

⚠ Common exam trap

The trap here is defaulting to mean and standard deviation for any numeric column, even when the scenario states that extreme outliers are present.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Median and interquartile range

For skewed data with known outliers, robust statistics are preferred. The median resists extreme values and represents the middle of the distribution, while the interquartile range measures spread using the middle 50 percent of records. Mean and standard deviation, mode and range, and sum and count either amplify outlier influence or fail to describe central tendency and dispersion together.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Sum and count

    Why it's wrong here

    Sum and count describe volume and total, not central tendency or spread. They cannot tell the analyst what a typical customer earns or how dispersed incomes are. While count is useful for completeness checks during profiling, using sum and count as the summary of the income distribution would omit the shape information needed to assess data quality and choose later transformations.

  • ✓

    Median and interquartile range

    Why this is correct

    The median is the middle value and is resistant to extreme outliers, so it reflects the typical income even when a few records are corrupted. The interquartile range describes the spread of the middle half of the data and likewise ignores extreme tails. Together they give a robust central tendency and dispersion summary that suits skewed income data with known entry errors.

  • ✗

    Mean and standard deviation

    Why it's wrong here

    Mean and standard deviation are both sensitive to extreme values, so a handful of erroneous high incomes will pull the mean upward and inflate the standard deviation. That makes the reported 'typical' income misleading for the majority of customers. While these measures are useful for symmetric, outlier-free distributions, they are the wrong choice when the scenario explicitly warns of extreme data entry errors.

  • ✗

    Mode and range

    Why it's wrong here

    The mode identifies the most frequent income value, which for continuous income data may be arbitrary or not meaningful, and the range is defined entirely by the minimum and maximum, making it the least robust spread measure possible. A single erroneous maximum would define the range. This pairing does not summarize typical income or its variation for the bulk of customers.

About these practice questions

This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.