Courseiva

DA0-002 Data Acquisition and Preparation Practice Question

A data analyst is using Python pandas to perform exploratory data analysis. Which THREE methods are commonly used to assess data quality and distributions?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

df.describe()

describe() gives summary statistics, info() shows data types and non-null counts, and value_counts() shows frequency distributions.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    df.transpose()

    Why it's wrong here

    transpose() swaps rows and columns, reshaping orientation only; it reveals nothing about missing values, duplicates or distribution shape. It tempts when analysts want a wider layout for readability, but assessing data quality and distributions requires describe(), info() or value_counts() instead.

  • ✓

    df.describe()

    Why this is correct

    df.describe() returns count, mean, standard deviation, minimum, quartiles and maximum for numeric columns, exposing outliers, skew and missing values. This single call summarises distribution shape and completeness, making it a standard first step in data-quality assessment.

  • ✓

    df.info()

    Why this is correct

    `df.info()` prints a concise summary of the DataFrame: column names, non-null counts, dtypes and memory usage. Non-null counts directly expose missing values, satisfying the data-quality assessment the stem requires, while dtypes confirm whether numeric columns are usable for distribution analysis.

  • ✗

    df.sort_values()

    Why it's wrong here

    sort_values() orders rows by column values, which aids inspection but does not itself quantify missingness, duplicates or distribution spread. It tempts because sorting surfaces outliers visually, yet the stem asks for methods assessing quality and distributions, such as describe() or isnull().sum().

  • ✓

    df.value_counts()

    Why this is correct

    `df.value_counts()` tallies occurrences of each distinct value within a Series, exposing frequency distributions and surfacing unexpected or dominant categories during exploratory analysis. This directly satisfies the stem's requirement to assess distributions, revealing skewed or rare values that signal data-quality problems such as miscoding or inconsistent entries.

About these practice questions

This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.