Courseiva

DA0-002 · topic practice

Data Analysis practice questions

Data Analysis is 24% of CompTIA Data+ (DA0-002). It covers data exploration, descriptive and inferential statistics, hypothesis testing, correlation versus causation, and clustering methods such as k-means. Questions present realistic business scenarios and ask you to select the correct technique, interpret output, or identify what a result does and does not prove.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Data Analysis

What the exam tests

What to know about Data Analysis

Be able to pick the right analysis for a scenario, run and interpret descriptive and inferential statistics, and read clustering output. The single most important thing: correctly interpret results, including significance, effect size, and the limits of what the data can prove.

Selecting descriptive statistics and appropriate visualizations during exploratory data analysis

Choosing inferential tests such as two-sample t-tests for comparing group means

Interpreting p-values, confidence intervals, and significance when evaluating a hypothesis

Applying k-means clustering and using WCSS elbow plots to choose cluster count

Watch out for

Common Data Analysis exam traps

  • ▸Treating a statistically significant result as proof of practical importance or causation without considering effect size and study design.
  • ▸Running k-means on raw variables with very different scales instead of standardizing or normalizing features first, which distorts distance calculations.
  • ▸Confusing correlation with causation, or ignoring confounding variables and sampling bias when interpreting an observed relationship.

Practice set

Data Analysis questions

20 questions · select your answer, then reveal the explanation

Question 1easymultiple choice
Read the full Data Analysis explanation →

A data analyst needs to combine two datasets that have the same columns but different rows. Which operation should they use?

Question 2mediummultiple choice
Read the full Data Analysis explanation →

A data analyst needs to visualize the distribution of a continuous variable across different categories. Which chart type is most suitable?

Which TWO of the following are common assumptions of linear regression?

A data analyst is building a linear regression model to predict sales based on advertising spend across TV, radio, and newspaper channels. Which TWO diagnostics should the analyst perform to validate the model assumptions?

Question 5hardmultiple choice
Read the full Data Analysis explanation →

A data analyst is preparing a logistic regression model to predict customer churn. After examining the exhibit, which data quality issue should the analyst address first?

Network Topology
|Refer to the exhibit.Table: customer_churn

Match each data governance role to its responsibility.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Ensures data quality and adherence to policies

Manages technical environment and data access

Has accountability for specific data assets

Sets strategic direction for data management

Designs data structures and integration processes

Question 7mediummultiple choice
Read the full Data Analysis explanation →

A data analyst is preparing a dataset for a predictive model. The dataset contains a feature 'age' with values ranging from 18 to 80, and a feature 'income' ranging from 20,000 to 200,000. To ensure both features contribute equally to distance-based algorithms, which transformation should the analyst apply?

Question 8hardmultiple choice
Read the full Data Analysis explanation →

A data scientist is tuning a decision tree model to prevent overfitting. The model currently has a high variance. Which hyperparameter adjustment is most effective?

Question 9hardmultiple choice
Read the full Data Analysis explanation →

Refer to the exhibit. Which data quality dimension is being violated?

Exhibit

2024-01-15 10:23:45 ERROR: DataTypeMismatchException - Column 'age' contains mixed data types: INT and VARCHAR. Pipeline 'user_profile_etl' failed.
Question 10hardmultiple choice
Read the full Data Analysis explanation →

A data scientist is building a classification model to predict customer churn. The dataset has 10,000 records with 500 churners. The scientist uses logistic regression and achieves 98% accuracy, but the precision for churn class is only 15%. Which of the following is the most likely cause?

Question 11hardmultiple choice
Read the full Data Analysis explanation →

An analyst is fitting a polynomial regression model and wants to choose the degree that minimizes overfitting. Which technique should the analyst use?

Question 12easymultiple choice
Read the full Data Analysis explanation →

A data analyst needs to create a visual that shows the distribution of customer ages across different regions. Which chart type is most appropriate?

A data scientist is cleaning a dataset and notices missing values in several columns. Which THREE techniques are appropriate for handling missing data? (Select THREE.)

Which THREE of the following are examples of descriptive statistics? (Select THREE.)

Question 15mediummultiple choice
Read the full Data Analysis explanation →

Refer to the exhibit. An analyst runs the following query: SELECT product_id, AVG(quantity) FROM sales GROUP BY product_id HAVING AVG(quantity) > 8; Which product_id(s) will be returned?

Network Topology
+Refer to the exhibit.SQL query output:```
Question 16hardmultiple choice
Read the full Data Analysis explanation →

Refer to the exhibit. Before running the code, the original salary column had 50 missing values. The median was calculated as 52000. After imputation, which of the following statements is true?

Exhibit

Refer to the exhibit.

Python pandas code and output:
```
import pandas as pd
df = pd.read_csv('employees.csv')
df['salary'].fillna(df['salary'].median(), inplace=True)
print(df['salary'].describe())
```
Output:
```
count    1000.000000
mean     55000.000000
std      15000.000000
min      25000.000000
25%      45000.000000
50%      52000.000000
75%      65000.000000
max     120000.000000
Name: salary, dtype: float64
```
Question 17easymultiple choice
Read the full Data Analysis explanation →

Refer to the exhibit. A data analyst wants to grant read access to an entire cloud storage bucket named 'data-lake'. Which of the following best describes what this policy does?

Exhibit

Refer to the exhibit.

JSON policy:
```
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "s3:GetObject",
      "Resource": "arn:aws:s3:::data-lake/*"
    }
  ]
}
```
Question 18mediummultiple choice
Read the full Data Analysis explanation →

A retail company wants to forecast monthly sales for the next 12 months. Sales data shows a clear upward trend and seasonal patterns that repeat yearly. Which time series model is most appropriate?

Question 19hardmultiple choice
Read the full Data Analysis explanation →

A data scientist is analyzing a dataset with 100 variables and 5,000 records. The dataset has several missing values and a few extreme outliers. The goal is to build a regression model to predict a continuous target. Which combination of preprocessing steps is most likely to improve model performance?

Question 20easymultiple choice
Read the full Data Analysis explanation →

A dataset contains a column 'Income' with values in different scales (some in thousands, some in hundreds). What is the best way to standardize this column for use in a machine learning model?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Data Analysis sessions

Start a Data Analysis only practice session

Every question in these sessions is drawn from the Data Analysis domain — nothing else.

Related practice questions

Related DA0-002 topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the DA0-002 exam test about Data Analysis?
Be able to pick the right analysis for a scenario, run and interpret descriptive and inferential statistics, and read clustering output. The single most important thing: correctly interpret results, including significance, effect size, and the limits of what the data can prove.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Data Analysis questions in a focused session?
Yes — the session launcher on this page draws every question from the Data Analysis domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other DA0-002 topics?
Use the topic links above to move to related areas, or go back to the DA0-002 question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the DA0-002 exam covers. They are not copied from any real exam or dump site.