Courseiva

DA0-002 · domain

Data Analysis

Data Analysis is 24% of CompTIA Data+ (DA0-002). It covers data exploration, descriptive and inferential statistics, hypothesis testing, correlation versus causation, and clustering methods such as k-means. Questions present realistic business scenarios and ask you to select the correct technique, interpret output, or identify what a result does and does not prove.

238 questions58 easy108 medium72 hard

Focused practice

Practice Data Analysis questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Data Analysis

Be able to pick the right analysis for a scenario, run and interpret descriptive and inferential statistics, and read clustering output. The single most important thing: correctly interpret results, including significance, effect size, and the limits of what the data can prove.

Selecting descriptive statistics and appropriate visualizations during exploratory data analysis

Choosing inferential tests such as two-sample t-tests for comparing group means

Interpreting p-values, confidence intervals, and significance when evaluating a hypothesis

Applying k-means clustering and using WCSS elbow plots to choose cluster count

Watch out for

Common Data Analysis exam traps

  • ▸Treating a statistically significant result as proof of practical importance or causation without considering effect size and study design.
  • ▸Running k-means on raw variables with very different scales instead of standardizing or normalizing features first, which distorts distance calculations.
  • ▸Confusing correlation with causation, or ignoring confounding variables and sampling bias when interpreting an observed relationship.

Question index

All Data Analysis questions (238)

Click any question to see the full explanation, or start a practice session above.

1

An analyst runs an A/B test with 1000 users per group and observes a conversion rate of 5% in the control and 6% in the treatment. The p-value is 0.12. What should the analyst conclude?

Hard
2

In A/B testing, the null hypothesis typically states that:

Easy
3

A data analyst receives a sales dataset where the 'order_date' column is stored as text in the format '2024-03-15'. The analyst needs to calculate the number of days between each order date and the current date. Which action should the analyst take first?

Easy
4

An analyst computed the mean, median, and mode of a dataset and found they are all equal. Which of the following best describes the distribution?

Easy
5

A time series of monthly sales data exhibits a clear upward trend over several years, with consistent peaks each December. Which components are present in this series?

Hard
6

A dataset contains a feature with values ranging from 10 to 1000. The analyst applies min-max normalization to scale the feature between 0 and 1. What is the normalized value of 520?

Hard
7

An analyst is comparing the average sales of two different store locations using a t-test. The p-value obtained is 0.03, and the significance level is 0.05. What should the analyst conclude?

Medium
8

A data analyst is preparing to run an A/B test comparing two email subject lines. Which TWO of the following should the analyst define before the test begins?

Medium
9

A retailer wants to test if a new website layout increases the average time spent on the site. They split traffic: control group (old layout) and treatment group (new layout). Which statistical test is most appropriate to compare the average time spent between the two groups?

Medium
10

A data analyst is working with a dataset that contains a column for 'order_date' stored as text in the format 'YYYY-MM-DD'. The analyst needs to calculate the average time between orders for each customer. Which of the following steps should the analyst take first?

Easy
11

A data analyst is performing a hypothesis test with a significance level of 0.05. The p-value obtained is 0.03. What should the analyst conclude?

Medium
12

A data analyst runs an A/B test on a new website layout. The test yields a p-value of 0.04 with the null hypothesis being no difference in conversion rates. The significance threshold is α=0.05. Which of the following is the correct conclusion?

Hard
13

A data analyst is performing a chi-square test of independence on a 2x2 contingency table. The p-value is 0.04. At α=0.05, which THREE of the following statements are correct?

Medium
14

Which TWO of the following data quality dimensions are most directly affected by duplicate records?

Medium
15

In a dataset with variables on different scales (e.g., age in years and income in dollars), which preprocessing step is necessary before applying k-means clustering?

Medium
16

A data analyst at a retail company is building a multiple linear regression model to forecast weekly sales. The dataset contains 50 predictor variables, including store size, promotional spend, holiday indicators, and many others. After training the model, the analyst observes an R-squared of 0.99 on the training set but only 0.55 on the holdout test set. Which action should the analyst take first to address this discrepancy?

Hard
17

A data analyst is evaluating a classification model that predicts whether a customer will churn. The model's confusion matrix shows 80 true positives, 20 false negatives, 30 false positives, and 120 true negatives. Which TWO of the following metrics can be directly calculated from this confusion matrix? (Choose two.)

Medium
18

Which THREE of the following are appropriate methods to handle outliers in a dataset?

Hard
19

An analyst is preparing data for a clustering algorithm that uses Euclidean distance. Which TWO data preprocessing techniques should be applied to ensure all features contribute equally?

Medium
20

An analyst calculates a Pearson correlation coefficient of -0.8 between advertising spend and customer churn rate. Which interpretation is correct?

Medium
21

A data analyst wants to segment customers into groups based on their purchasing behavior. The dataset includes numerical features such as annual income and purchase frequency. Which algorithm is most appropriate for this task?

Medium
22

A data analyst needs to determine whether the mean sales of two different regions are significantly different. The samples are independent and the data is normally distributed. Which statistical test should be used?

Medium
23

A data scientist is conducting an A/B test with a significance level of 0.05. Which three factors should be considered when calculating the required sample size? (Choose THREE)

Hard
24

A logistic regression model is used to predict the probability of customer churn. The model's coefficient for the feature 'customer support calls' is 0.8 with a p-value of 0.001. Which interpretation is correct?

Hard
25

In time series decomposition, a data analyst separates a retail sales series into trend, seasonal, and residual components. After decomposition, the residual component shows no pattern and is random. Which of the following best describes the seasonal component?

Hard
26

A data analyst calculates the mean, median, and mode of a dataset. Which measure of central tendency is most affected by extreme outliers?

Easy
27

A retail company wants to identify customer segments based on purchase history and demographics. Which technique is most appropriate for this task?

Medium
28

A data scientist is using K-means clustering with k=3. After the first iteration, the centroids are recalculated. Which step occurs next in the algorithm?

Medium
29

In a multiple regression model, one predictor has a high p-value (0.45). What should the analyst consider doing?

Medium
30

A data analyst needs to join two tables in a SQL database: Orders and Customers. The analyst wants to include all orders, even if there is no matching customer record. Which type of join should be used?

Easy
31

In a multiple regression model with three predictors, the coefficient for one predictor is 5.2 with a p-value of 0.001. Which of the following is the best interpretation?

Hard
32

A company wants to determine if there is a significant difference in the average sales revenue between two different store layouts. They collect sales data from 30 stores with Layout A and 30 stores with Layout B. Which statistical test is most appropriate for comparing the means of these two independent groups?

Medium
33

A data analyst is examining a dataset of customer orders and notices that the 'order_total' column contains values ranging from $5 to $50,000, with most orders under $200. The analyst wants to reduce the influence of extreme values on a linear regression model predicting customer lifetime value. Which technique should the analyst apply?

Medium
34

A retail company wants to test whether a new website layout increases the conversion rate compared to the current layout. They randomly assign visitors to either the control or treatment group. Which statistical test is most appropriate to compare the conversion rates?

Medium
35

A data scientist builds a logistic regression model to predict customer churn (yes/no). The model outputs a probability of 0.75 for a particular customer. Which of the following best describes this output?

Hard
36

A data analyst at a regional bank is examining monthly account balances for 5,000 customers. The distribution is strongly right-skewed because a small number of high-net-worth clients hold very large balances. The analyst must report a measure of central tendency that best represents the typical customer without being pulled upward by those few extreme accounts. Which measure should the analyst report?

Medium
37

A data analyst at a marketing firm is tasked with segmenting customers based on their purchasing behavior. The dataset contains 10,000 customers with features such as annual spend, frequency of purchases, recency of last purchase, and average order value. The analyst decides to use k-means clustering. After standardizing the features, the analyst runs k-means with k=3, k=4, and k=5, and computes the silhouette score for each: k=3: 0.45, k=4: 0.52, k=5: 0.48. The analyst also plots the elbow curve and observes that the within-cluster sum of squares (WCSS) decreases sharply from k=2 to k=4, then levels off. Based on these results, what is the most appropriate number of clusters?

Medium
38

A data analyst is building a time series forecasting model for monthly sales data. The analyst observes that the data has a clear upward trend and a seasonal pattern that repeats every 12 months. After differencing the data once to remove the trend, the analyst notices that the seasonal pattern remains. Which of the following should the analyst do next to properly model the data?

Hard
39

A data analyst is building a predictive model to forecast customer churn. The dataset includes a categorical feature 'payment_method' with categories: credit card, debit card, PayPal, and bank transfer. The analyst decides to use one-hot encoding. After encoding, the analyst notices that the model's performance on the training set is excellent but poor on the test set. Which issue is most likely contributing to this problem?

Hard
40

A data analyst is reviewing a dataset of customer transactions and notices that the 'transaction_amount' column contains several negative values. The analyst suspects these are refunds rather than errors. Which data validation technique should the analyst apply to confirm this?

Easy
41

A data analyst is working with a dataset containing customer ages. The ages range from 18 to 90, but the analyst notices that the distribution is heavily skewed to the right. To better understand the central tendency, the analyst decides to calculate a measure that is resistant to outliers. Which measure of central tendency should the analyst use?

Easy
42

A data analyst at a healthcare organization is analyzing patient readmission rates. The analyst has a dataset with patient demographics, diagnosis codes, and length of stay. Before performing any statistical modeling, the analyst must address data quality issues. Which TWO of the following actions are most appropriate for ensuring the dataset is ready for analysis? (Choose two.)

Hard
43

A dataset contains the ages of 100 customers. The analyst wants to transform the ages to a 0-1 range for use in a distance-based algorithm. Which technique should be used?

Easy
44

A data analyst is cleaning a dataset and finds missing values in a categorical variable representing customer region. Which imputation method is most appropriate?

Easy
45

An analyst is conducting an A/B test on a new website layout. Which TWO of the following must be defined before the test begins?

Hard
46

A marketing team runs an A/B test comparing two webpage designs. The null hypothesis states there is no difference in conversion rates. The p-value is 0.08 at α=0.05. Which is the correct interpretation?

Medium
47

Which measure best describes the spread of the middle 50% of a dataset?

Easy
48

An analyst uses K-means clustering on customer purchase data. After plotting the within-cluster sum of squares for different values of k, they observe an elbow at k=4. What is the most appropriate number of clusters?

Medium
49

A data analyst is analyzing a time series of monthly sales for a retail store. The analyst observes a repeating pattern every 12 months. Which component of the time series is the analyst observing?

Medium
50

A dataset has missing values in the 'age' column. The distribution of age is approximately normal with few outliers. Which imputation method is most appropriate?

Medium
51

A data analyst is comparing the average test scores of students who attended a tutoring program versus those who did not. Which statistical test is most appropriate for determining if there is a significant difference between the means of these two independent groups?

Easy
52

Match each data sampling method to its description.

Medium
53

In a linear regression model predicting house prices, the coefficient for the number of bedrooms is $30,000 and the intercept is $50,000. If a house has 3 bedrooms, what is the predicted price?

Hard
54

In time series analysis, which component represents regular patterns that repeat over fixed periods, such as daily or yearly?

Medium
55

After a marketing campaign, sales increased by 15%. The analyst wants to understand which customer segment contributed most to the increase. Which type of analysis is this?

Easy
56

A data analyst is working with a dataset that contains a column 'income' with a highly skewed distribution. The analyst wants to apply a transformation to make the distribution more symmetric for use in a linear regression model. Which transformation is most appropriate?

Medium
57

A data scientist is building a K-means clustering model for customer segmentation. After plotting the within-cluster sum of squares (WCSS) against the number of clusters (k), she observes that the WCSS decreases sharply until k=5 and then levels off. Which value of k should she choose based on the elbow method?

Hard
58

A data analyst is preparing a dataset for a machine learning algorithm that assumes normally distributed features. Which TWO data transformation methods should the analyst consider to achieve this?

Medium
59

A dataset contains employee salaries ranging from $30,000 to $200,000. An analyst wants to scale the salaries to a range of 0 to 1 for use in a distance-based clustering algorithm. Which method should they use?

Medium
60

In an A/B test, the null hypothesis states that there is no difference between the conversion rates of the control and treatment groups. After collecting data, the p-value is 0.03. Using a significance level α = 0.05, what should the analyst conclude?

Easy
61

A data analyst is designing a data model for a sales data warehouse. The model should optimize query performance for aggregations by minimizing joins and duplicating data where necessary. Which schema design should the analyst use?

Easy
62

A data analyst uses linear regression to model the relationship between advertising spend and sales. The residual plot shows a clear U-shaped pattern. What assumption is violated?

Hard
63

A healthcare analytics team is building a predictive model to identify patients at high risk of readmission within 30 days of discharge. The dataset includes 50,000 patient records with 200 features, including demographics, vital signs, lab results, and historical admissions. The target variable is binary (readmitted or not). The team uses a logistic regression model and achieves an AUC of 0.72 on the test set. However, the model's calibration is poor: for patients predicted to have a 70% risk, the actual readmission rate is only 40%. The team wants to improve calibration without significantly reducing discrimination (AUC). The data scientist suggests applying Platt scaling. However, the team lead is concerned that Platt scaling may reduce the model's ability to rank patients correctly. Which of the following is the best course of action?

Hard
64

A data analyst is examining a dataset of customer transactions and notices that the 'transaction_amount' column contains negative values. The analyst determines that these negative values represent refunds. Which data quality dimension is most directly relevant to this finding?

Easy
65

A data scientist is performing a hypothesis test with a significance level α=0.05. The p-value obtained is 0.03. What should the scientist conclude?

Medium
66

A healthcare analytics team is building a classification model to predict patient readmission within 30 days. The dataset contains 10,000 records with 30 features, including demographics, vital signs, lab results, and medication history. The target variable is imbalanced: 85% no readmission, 15% readmission. The team used logistic regression with default settings and achieved an accuracy of 85%, but the model predicted 'no readmission' for all patients. The lead analyst suspects the model is not learning due to class imbalance. The team has time to implement one corrective action before the next model review. Which action should the team take?

Medium
67

A data analyst is preparing a dataset for a machine learning model and needs to handle missing values in several columns. The analyst wants to choose appropriate imputation methods. Which TWO of the following are valid considerations when selecting an imputation technique? (Choose two.)

Medium
68

Which TWO of the following are components of time series data?

Medium
69

A data analyst wants to understand the relationship between advertising spend and sales revenue. The analyst calculates a Pearson correlation coefficient of 0.85. Which of the following is the best interpretation?

Medium
70

A company is analyzing customer feedback sentiment. The dataset is highly imbalanced with 95% positive and 5% negative comments. Which technique should the analyst use to address class imbalance before modeling?

Hard
71

Which TWO of the following are measures of central tendency?

Easy
72

A data analyst wants to use a Z-score to standardize a dataset. The variable has a mean of 50 and a standard deviation of 10. What is the Z-score for a raw value of 70?

Easy
73

A data analyst notices that a dataset of customer ages has several missing values. Which method for handling missing data is most appropriate if the data is missing completely at random and the analyst wants to preserve sample size?

Medium
74

A data analyst is building a logistic regression model to predict whether a customer will churn (yes/no). Which TWO statements about logistic regression are correct?

Hard
75

A data analyst is evaluating a classification model that predicts customer churn. The model has an accuracy of 95%, but the analyst suspects it is not performing well on the minority class (churners). The dataset is highly imbalanced, with only 5% churners. Which metric should the analyst prioritize to assess the model's ability to correctly identify churners?

Hard
76

Which TWO of the following are valid techniques for validating the performance of a predictive model?

Hard
77

An analyst runs a simple linear regression with an R² value of 0.85. Which interpretation is correct?

Medium
78

An analyst is performing K-means clustering on customer data. The elbow method shows a clear bend at k=4. Which THREE of the following are true about K-means clustering with k=4?

Hard
79

A marketing analyst wants to predict whether a customer will churn (yes/no) based on account age and monthly charges. Which regression technique is most appropriate?

Medium
80

A data scientist is analyzing a dataset with multiple features and wants to apply k-means clustering to segment customers. She chooses k = 4 based on the elbow method. During the iteration process, which of the following correctly describes a step in the k-means algorithm?

Hard
81

Refer to the exhibit. Which type of ensemble method is being used?

Medium
82

A data analyst is reviewing a SQL query that joins three large tables. The query takes over an hour to run. The analyst notices that the WHERE clause filters on indexed columns in only two tables. Which of the following should the analyst do first to improve performance?

Medium
83

An analyst compares average sales across three different store locations using a statistical test. Which test is most appropriate?

Medium
84

A data analyst is evaluating the quality of a customer database. Which THREE of the following are dimensions of data quality?

Hard
85

Which data quality dimension ensures that data represents the real-world object or event correctly?

Easy
86

A data analyst is performing time series analysis on monthly sales data and notices a consistent pattern of higher sales every December. Which component of time series does this represent?

Medium
87

A data analyst is cleaning a dataset and identifies several outliers. Which TWO methods are appropriate for handling outliers?

Hard
88

A data analyst is examining sales data for a retail chain and notices that the mean monthly sales is $50,000 while the median is $35,000. Which of the following best describes the distribution of the sales data?

Medium
89

A data analyst is cleaning a dataset and finds that 5% of values in the 'income' column are missing. The analyst decides to impute missing values using the mean of the non-missing values. Which potential issue should the analyst be most concerned about?

Hard
90

A dataset contains a variable 'Income' with many missing values. The analyst decides to impute missing values with the median income of the non-missing values. Which type of imputation is this?

Medium
91

A data analyst is reviewing a dashboard that tracks monthly sales. The analyst notices that the 'Total Sales' metric for the current month is significantly lower than previous months, but the underlying daily sales data appears normal. The dashboard uses a live connection to the database. Which of the following is the most likely cause of the discrepancy?

Easy
92

A data analyst needs to identify outliers in a dataset. Which of the following is a common method based on the interquartile range (IQR)?

Easy
93

Which THREE of the following are assumptions of linear regression? (Select THREE).

Hard
94

Which TWO of the following are examples of supervised learning algorithms?

Easy
95

A data analyst is examining the relationship between advertising spend (in dollars) and revenue (in dollars). The Pearson correlation coefficient r is calculated as +0.92. Which of the following interpretations is correct?

Medium
96

A data analyst is examining a dataset of daily website visits collected over the past year. The analyst wants to determine whether the distribution of visits on weekdays differs significantly from the distribution on weekends, without assuming normality. The analyst has already separated the data into two independent groups. Which statistical test should the analyst use?

Medium
97

A data analyst needs to identify the most frequently occurring value in a dataset. Which measure of central tendency should they use?

Easy
98

A data analyst is examining a dataset of employee salaries and notices that the distribution is heavily right-skewed due to a few executives with very high salaries. The analyst wants to apply a transformation to make the distribution more symmetric for further analysis. Which transformation is most appropriate?

Medium
99

A data analyst is performing a chi-square test of independence on a contingency table of customer satisfaction (satisfied vs. dissatisfied) and product type (A, B, C). The test yields a p-value of 0.04 with α = 0.05. What is the correct conclusion?

Hard
100

A company wants to segment its customers into distinct groups based on purchasing behavior. Which algorithm is best suited for this task?

Medium
101

A data analyst is preparing a dataset for analysis and needs to ensure data quality. Which TWO of the following are dimensions of data quality?

Medium
102

A data analyst wants to test if the proportion of customers who prefer Product A over Product B is different from 50%. She surveys 200 customers and finds that 120 prefer Product A. Which statistical test should she use?

Medium
103

A data analyst is exploring a dataset of 8,000 customer transactions and notices that the 'transaction_amount' column has a mean of $120 but a median of $45. A small number of transactions exceed $10,000. Which measure of central tendency should the analyst report to describe the typical transaction?

Medium
104

A retail company wants to predict sales based on advertising spend and season. Which data modeling technique should the analyst use?

Medium
105

A data team is preparing data for a clustering analysis. Which THREE of the following steps are commonly part of data cleaning?

Medium
106

A data scientist is performing K-means clustering on customer data. She plots the within-cluster sum of squares (WCSS) for different values of k and observes an 'elbow' at k=4. What does this indicate?

Medium
107

A data analyst is evaluating a multiple regression model with three predictors. The R² value is 0.85. Which of the following is the best interpretation of R²?

Medium
108

Which TWO of the following are dimensional modeling techniques commonly used in data warehouses?

Easy
109

A data scientist trains a regression model and observes high variance with low bias. Which technique is most appropriate to reduce variance?

Hard
110

A logistic regression model predicts customer churn (0=no churn, 1=churn). The model outputs probabilities. Which THREE of the following statements about logistic regression are correct?

Hard
111

A data analyst is creating a visualization to show the proportion of total sales contributed by each of five product categories. The analyst wants to emphasize the relative size of each category as part of a whole. Which type of chart is most appropriate?

Easy
112

A data analyst is performing a chi-square test for independence between two categorical variables. Which THREE of the following are necessary conditions for the test to be valid?

Hard
113

A retail company wants to analyze monthly sales data over the past three years to identify long-term trends. Which component of time series analysis is most relevant for this goal?

Easy
114

A company is planning an A/B test to compare two website designs. Which THREE of the following must be determined before the test begins to ensure valid results? (Select three.)

Hard
115

A marketing analyst wants to segment customers based on their purchase history, including total spent, number of transactions, and average order value. The analyst runs k-means clustering with k=5 on the raw data but notices that the cluster assignments change significantly every time the algorithm is executed. What should the analyst do first to obtain consistent and meaningful clusters?

Easy
116

Refer to the exhibit. Which clause is used to aggregate the data by department?

Easy
117

You are a data analyst at a logistics company. The operations manager wants to reduce delivery delays. You have historical data including order date, delivery date, distance, weather conditions, and driver ID. Initial analysis shows that the average delivery time has increased over the past six months. You suspect that weather is a contributing factor, but you need to confirm. The company also wants to build a model to predict delivery times to better manage customer expectations. The data contains missing values for weather conditions in about 10% of records, and some driver IDs are incorrect. You have limited time and resources. What should you do first?

Easy
118

A data analyst is comparing the means of two independent groups using a t-test. The sample sizes are small and the data is not normally distributed. Which condition is violated for a valid t-test?

Hard
119

A data analyst needs to summarize customer satisfaction scores. The data contains a few extremely low scores that skew the distribution. Which measure of central tendency is most appropriate?

Easy
120

A data analyst is building a model to predict customer churn. The dataset has 10,000 records with 500 churned customers. The model predicts churn with 95% accuracy, but only identifies 10% of actual churners. Which metric best highlights this issue?

Hard
121

In an A/B test, the null hypothesis states that there is no difference between the control and treatment groups. After running the test, the p-value is 0.04. Assuming α = 0.05, what is the correct conclusion?

Easy
122

After building a binary classification model, the data analyst obtains the following confusion matrix: True Positives=80, True Negatives=100, False Positives=20, False Negatives=30. What is the F1 score?

Hard
123

A data analyst is preparing to build a predictive model. Which TWO steps are essential to ensure model validity? (Choose two.)

Easy
124

A data analyst is preparing a dataset for analysis and needs to address data quality issues. Which TWO of the following are common data cleaning tasks?

Medium
125

A marketing team wants to segment customers into distinct groups based on purchasing behavior. The data includes numeric features such as frequency, monetary value, and recency. Which unsupervised learning algorithm should be used?

Easy
126

A data analyst is cleaning a customer dataset. Which two actions are appropriate for handling duplicate records? (Choose TWO)

Medium
127

A data analyst at a hospital network is exploring the relationship between patient age and length of stay for 400 discharged patients. A scatterplot shows a roughly linear upward trend, but the analyst wants a single number summarizing both the strength and direction of the association before reporting to clinicians. Which measure should the analyst calculate?

Medium
128

A data analyst is testing whether the average sales amount differs between two regions. Which statistical test is most appropriate?

Medium
129

A data scientist builds a simple linear regression model to predict house prices based on square footage. The model yields an R-squared value of 0.85. Which statement accurately interprets this result?

Medium
130

A data analyst is performing a chi-square test of independence on a contingency table of customer satisfaction (satisfied, neutral, dissatisfied) by region (North, South, East, West). Which THREE of the following are necessary assumptions for the test?

Hard
131

A dataset contains customer records with a column for 'Phone Number' that should be unique. However, the analyst finds several duplicate phone numbers. Which data quality dimension is primarily affected?

Easy
132

A marketing team runs an A/B test on email subject lines. The p-value is 0.03 with α = 0.05. Which of the following is the correct interpretation?

Medium
133

In logistic regression, the output is a probability between 0 and 1. If the predicted probability for a customer churning is 0.7 and the decision threshold is 0.5, what is the predicted class?

Hard
134

Drag and drop the steps to implement a data classification policy in the correct order.

Medium
135

A data analyst is working with a dataset that includes a categorical variable 'product_category' with 50 unique values. The analyst wants to reduce dimensionality before clustering. Which technique should the analyst use?

Hard
136

Which TWO of the following are true about Pearson correlation coefficient (r)?

Medium
137

A data analyst is analyzing customer purchase amounts. The dataset contains several extreme high values due to luxury purchases. Which measure of central tendency is most robust to these outliers?

Medium
138

A data analyst at a logistics company is analyzing delivery times for three different shipping carriers. The analyst wants to determine whether the mean delivery time differs across carriers. The data are normally distributed, and the variances across carriers are assumed equal. Which statistical test should the analyst use?

Medium
139

A financial analyst wants to compare the mean annual returns of three different investment strategies. Which statistical test is most appropriate?

Medium
140

During ETL, a data analyst discovers that a date column contains values like '01/02/2023' and '2023-01-02'. Which of the following is the best practice to ensure consistent date format before analysis?

Easy
141

Which data cleaning method involves replacing a missing value with the average of the available values in that column?

Easy
142

A company has a dataset with 100 features. The data analyst wants to reduce dimensionality while preserving as much variance as possible. Which technique should be used?

Medium
143

A data analyst calculates a correlation coefficient of -0.85 between temperature and heating costs. What does this indicate?

Easy
144

A data analyst is performing a multiple linear regression with three predictors. The model output shows an R-squared of 0.85 and an adjusted R-squared of 0.80. Which of the following is the best interpretation of the difference between these two values?

Hard
145

A marketing team uses K-means clustering to segment customers based on purchase history. To determine the optimal number of clusters, they plot the within-cluster sum of squares (WCSS) against k and look for an elbow. What is the purpose of this method?

Medium
146

After training a decision tree, the tree has depth 20 and 100% accuracy on training data but only 60% on test data. Which hyperparameter adjustment is most likely to improve generalization?

Hard
147

A data analyst at a subscription media company is asked to determine whether a new recommendation engine increased average watch time. The analyst plans a two-sample t-test comparing a randomly assigned control group against a randomly assigned treatment group, with watch time recorded in minutes per user per week. Which TWO conditions must hold for the pooled-variance two-sample t-test to be statistically valid here? (Choose two.)

Medium
148

An analyst wants to compare the mean sales revenue across three different store regions. The data is normally distributed and variances are equal. Which statistical test is most appropriate?

Medium
149

A researcher is designing an A/B test to compare two website layouts. Which TWO elements are essential for determining the required sample size?

Medium
150

A data analyst is cleaning a dataset and finds that some records have duplicate entries based on customer ID. Which data quality dimension is most directly affected by these duplicates?

Hard
151

A data analyst is building a supervised learning model to predict customer churn. The target variable is binary (churn = yes/no). Which TWO modeling techniques are appropriate for this task? (Select two.)

Medium
152

A data analyst at an e-commerce company is profiling a transactions table with ten million rows. The 'order_total' column is stored as text and contains values like '$1,299.00', a few entries of 'N/A', and several negative amounts that represent refunds. The analyst must produce a numeric column suitable for summation in a revenue report. Which sequence of operations best prepares the column?

Hard
153

In time series decomposition, a pattern that repeats at regular intervals (e.g., weekly, yearly) is called:

Hard
154

A simple linear regression model predicts sales (y) from advertising spend (x). The equation is y = 2.5x + 10, and R² = 0.81. Which interpretation is correct?

Medium
155

A data analyst is reviewing a dataset and notices that the 'ProductCategory' column contains inconsistent entries such as 'Electronics', 'electronics', and 'ELECTRONICS'. Which data cleaning technique should the analyst apply to standardize these values?

Easy
156

In a time series analysis, a retail analyst observes consistent peaks in sales every December and troughs every February. This pattern repeats annually. Which component of time series does this represent?

Medium
157

A dataset contains outliers in a feature that will be used for linear regression. Which two outlier treatment methods are appropriate? (Choose TWO)

Medium
158

A marketing analyst wants to segment customers based on purchasing behavior and demographics. The dataset includes continuous variables (spending amount, frequency) and categorical variables (region, gender). The analyst decides to use k-means clustering. What should the analyst do to prepare the data?

Hard
159

A data analyst wants to predict customer churn based on categorical features like region and plan type, and continuous features like usage and tenure. Which regression type should be used?

Easy
160

During data exploration, an analyst notices that the target variable has a heavily right-skewed distribution. Which data transformation would be most appropriate to make the distribution more symmetric?

Easy
161

A company’s marketing team wants to segment customers based on purchase history, demographics, and website behavior. The data includes both numeric and categorical variables. Which clustering algorithm is best suited for handling mixed data types?

Medium
162

A data analyst is working with a dataset that contains a column for 'customer satisfaction rating' on a scale from 1 to 5, where 1 is very dissatisfied and 5 is very satisfied. The analyst wants to summarize the central tendency of this data. Which measure of central tendency is most appropriate for this ordinal data?

Easy
163

A data analyst is working with a dataset containing house prices. After building a multiple linear regression model, the analyst observes that the model performs well on training data but poorly on validation data. Which technique is most appropriate to address this issue?

Hard
164

A data analyst is working with a dataset containing a categorical variable 'Education Level' with categories: High School, Bachelor's, Master's, PhD. The analyst wants to include this variable in a regression model. Which encoding technique should the analyst use?

Medium
165

A data analyst is evaluating a binary classification model for loan default prediction. The model achieves 98% accuracy on the test set, but the analyst notices that only 2% of loans in the dataset actually defaulted. The analyst is concerned that accuracy is misleading. Which metric should the analyst examine to better assess the model's ability to identify actual defaulters?

Hard
166

A data analyst has a time series of monthly sales data. They observe that sales are consistently higher every December and lower every January. Which component of time series does this pattern represent?

Hard
167

A data analyst is compiling a monthly sales dashboard. The dataset contains a 'Revenue' column with a highly right-skewed distribution due to a few large enterprise transactions. The analyst wants a measure of central tendency that is robust to these outliers. Which measure should the analyst use?

Medium
168

Given the linear regression output, which independent variable has the strongest effect on price, based on standardized coefficients?

Hard
169

An analyst is conducting an A/B test to compare two website designs. The null hypothesis is that there is no difference in conversion rates. The p-value obtained is 0.03, and the significance threshold is 0.05. What should the analyst conclude?

Medium
170

A data analyst calculates the mean, median, and mode of a dataset. Which of the following best describes how these measures are used in descriptive statistics?

Easy
171

A data analyst at an online retailer is examining a dataset of customer orders. The 'order_total' column has a mean of $85 and a median of $62. The analyst wants to describe the typical order amount for a presentation to the marketing team. Which measure of central tendency is most appropriate to report as the typical value, and why?

Medium
172

A data analyst is evaluating a classification model that predicts whether customers will churn. The dataset is highly imbalanced, with only 5% of customers churning. The analyst wants to choose a metric that focuses on the model's ability to correctly identify actual churners. Which metric should the analyst prioritize?

Hard
173

A stock analyst is analyzing monthly sales data for a retail company and observes a consistent pattern of high sales every December. This pattern is most likely an example of which time series component?

Medium
174

In simple linear regression, the coefficient of determination R² measures:

Easy
175

In a regression analysis, the coefficient of determination (R²) is 0.85. How should this value be interpreted?

Easy
176

A data analyst is working with a dataset that contains a column 'purchase_date' stored as a string in the format 'YYYY-MM-DD'. The analyst needs to calculate the number of days between each purchase and the current date. Which data transformation is required?

Medium
177

A data analyst is building a binary classification model to predict customer churn. The dataset is imbalanced, with only 10% churners. The analyst wants to evaluate model performance with a focus on correctly identifying churners. Which metric is most appropriate?

Hard
178

A retail company wants to predict future sales based on historical data. Which modeling approach is most appropriate if the data shows a clear seasonal pattern?

Medium
179

A retail company wants to segment its customers based on purchase history. Which THREE methods are appropriate for customer segmentation?

Medium
180

The exhibit shows an SQL query executed on an 'orders' table that contains 'order_id', 'customer_id', and 'order_date'. What is the purpose of this query?

Medium
181

An analyst is preparing data for an A/B test and wants to ensure valid results. Which TWO of the following should be considered when calculating the required sample size?

Medium
182

A data analyst wants to compare the means of three different training methods on employee productivity. Which statistical test is most appropriate?

Easy
183

A data analyst is examining the relationship between two continuous variables: temperature and ice cream sales. The analyst wants to quantify the strength and direction of their linear association. Which statistical measure should the analyst use?

Easy
184

An analyst is planning an A/B test to compare two website designs. Which TWO factors should be considered when calculating the required sample size?

Medium
185

A data analyst is cleaning a dataset with missing values in a time series of daily temperatures. The missing values occur sporadically. Which imputation method is most appropriate to maintain the temporal trend?

Hard
186

A data analyst is reviewing a dataset containing house prices. The mean price is $350,000 and the median is $280,000. Which of the following best describes the distribution of house prices?

Medium
187

Which TWO of the following are appropriate uses of min-max normalisation?

Medium
188

A data analyst is examining the relationship between advertising spend and sales revenue across 50 regions. The analyst calculates a Pearson correlation coefficient of 0.85. Which of the following conclusions is most appropriate?

Hard
189

Which data quality dimension ensures that data represents the real-world scenario correctly and without errors?

Easy
190

A data analyst is analyzing survey responses where respondents rated satisfaction on a scale of 1-5. The analyst wants to visualize the distribution of responses. Which chart type is most appropriate?

Medium
191

In A/B testing, which factor is increased by having a larger sample size?

Hard
192

A data analyst is using a decision tree to classify loan applications as approved or denied. To avoid overfitting, the analyst plans to prune the tree. Which TWO of the following are valid pruning techniques? (Choose two.)

Medium
193

A data analyst is creating a visualization to show the relationship between two continuous variables: advertising spend and sales revenue. The analyst wants to display the strength and direction of the linear relationship. Which type of chart should the analyst use?

Easy
194

Which THREE of the following are common steps in data cleaning?

Medium
195

Which data quality dimension is violated if a customer record has a missing phone number?

Easy
196

A data analyst is working with a dataset that includes a categorical variable 'education_level' with four categories: High School, Bachelor's, Master's, and PhD. The analyst wants to include this variable in a linear regression model. Which encoding method should the analyst use to avoid the dummy variable trap?

Medium
197

A marketing team wants to segment customers into groups based on purchasing behavior without prior labels. Which algorithm should the data analyst use?

Medium
198

A data analyst is working with a dataset that contains a column 'region' with values such as 'North', 'South', 'East', 'West', and 'N/A'. The analyst needs to prepare this column for a machine learning model. Which of the following is the most appropriate approach to handle the 'N/A' values?

Medium
199

A data scientist is working with a dataset containing 1000 features and 500 samples. The goal is to build a predictive model. Which technique should be used to reduce the number of features while retaining most of the variance?

Hard
200

A data analyst is building a dashboard that tracks daily active users. The source system records one event row per user session, and a user can have multiple sessions in a day. The analyst needs to display the count of unique users per day. Which approach produces the correct daily active user count?

Hard
201

A data analyst is compiling a quarterly report for a retail chain and needs to show how sales revenue is distributed across five product categories. The analyst wants a single visualization that makes it easy to compare each category's share of total revenue and also includes a separate small panel displaying total revenue for the quarter. Which visualization approach best meets these requirements?

Medium
202

A data analyst is examining a dataset of customer transactions and notices that the distribution of transaction amounts is highly skewed to the right. The analyst wants to apply a transformation to make the distribution more symmetric for use in a linear regression model. Which of the following transformations is most appropriate?

Medium
203

In a simple linear regression model y = 2.5 + 1.2x, what is the predicted value of y when x = 10?

Easy
204

A data analyst is validating a regression model that predicts monthly churn probability for 50,000 subscribers. The analyst wants to detect whether the model is overfitting before deploying it. (Choose two.)

Hard
205

A data analyst wants to segment customers based on purchasing behavior such as frequency, monetary value, and recency. Which TWO clustering evaluation methods can help determine the optimal number of clusters? (Select two.)

Medium
206

A data analyst is asked to compare the average sales across three different store locations. The data is normally distributed and variances are approximately equal. Which statistical test is most appropriate?

Hard
207

An analyst is performing a logistic regression to predict customer churn (yes/no). The model outputs a probability of 0.75 for a particular customer. Which of the following best describes the interpretation?

Hard
208

An analyst is conducting an A/B test on a new checkout process. To calculate sample size, which THREE factors must be considered?

Medium
209

A data analyst is performing data cleaning. Which THREE steps are part of this process? (Choose three.)

Hard
210

Which statistical test should be used to determine if there is a significant association between two categorical variables, such as gender and product preference?

Easy
211

A data scientist runs a linear regression model to predict customer spending based on income. The R-squared value is 0.45 and the p-value for the slope coefficient is 0.03. At a significance level of α=0.05, which of the following conclusions is correct?

Hard
212

A data analyst is building a linear regression model to predict sales based on advertising spend. The analyst notices that the residuals are not normally distributed and have a non‑constant variance. Which of the following transformations is most appropriate to apply to the dependent variable?

Easy
213

Which data quality dimension is most concerned with whether data values fall within a defined domain or acceptable range?

Easy
214

A data scientist applies K-means clustering to a customer dataset. The elbow method suggests using 4 clusters. After running K-means with k=4, the within-cluster sum of squares (WCSS) is plotted against k, and the elbow is at k=4. What does this indicate?

Hard
215

A data analyst is examining the relationship between advertising spend (in thousands) and sales (in thousands). The Pearson correlation coefficient is computed as r = -0.85. Which of the following interpretations is correct?

Medium
216

A data analyst is evaluating a classification model for predicting customer churn. The model's confusion matrix shows 100 true positives, 20 false positives, 30 false negatives, and 850 true negatives. The analyst wants to assess the model's ability to correctly identify actual churners. Which metric should the analyst use?

Hard
217

A data analyst trains a complex model that achieves 99% accuracy on training data but only 65% on new data. What is the most likely issue?

Hard
218

A data analyst is cleaning a dataset and finds that the 'age' column has several missing values. Which of the following is a valid method for handling missing numerical data?

Easy
219

A data analyst is using a decision tree to predict customer churn. The tree is very deep, with many branches, and achieves 99% accuracy on the training data but only 70% on the test data. Which technique should the analyst use to address this issue?

Hard
220

A data analyst is summarizing the central tendency of a dataset with extreme outliers. Which measure is most robust to outliers?

Easy
221

A data analyst is cleaning a dataset and finds that a numeric field has several missing values. The variable is normally distributed. Which imputation method is most appropriate?

Medium
222

A data analyst calculates the mean, median, and mode of a dataset. Which of the following measures of central tendency is least affected by extreme outliers?

Easy
223

A data analyst uses the elbow method to determine the number of clusters for k-means. The plot shows a sharp bend at k=3 and a small bend at k=5. What is the recommended number of clusters?

Hard
224

A data analyst is building a decision tree to classify whether customers will churn. The analyst wants to prevent the tree from overfitting the training data. Which technique should the analyst use?

Hard
225

A data analyst is performing hypothesis testing to compare the mean sales of two store locations. Which TWO conditions must be satisfied to use a two‑sample t‑test? (Select TWO.)

Medium
226

A data analyst is preparing a dataset for a machine learning model and notices that several numerical features have different scales (e.g., age ranges from 18-90, income from 20,000-200,000). The analyst decides to apply feature scaling. Which TWO of the following statements about scaling techniques are correct? (Choose two.)

Hard
227

An analyst is performing a linear regression and obtains an R-squared value of 0.85. Which of the following is the best interpretation?

Medium
228

A data analyst is preparing a dataset for a machine learning model to predict customer churn. The dataset contains missing values in several columns. The analyst wants to handle missing data in a way that preserves as much information as possible and avoids introducing bias. Which TWO of the following techniques are appropriate? (Choose two.)

Hard
229

A data scientist is building a model to predict customer churn (yes/no). After training a logistic regression model, the coefficient for 'monthly charges' is 0.05 with a p-value of 0.03. Which interpretation is correct at α=0.05?

Hard
230

Drag and drop the steps to normalize a database table from 1NF to 3NF in the correct order.

Medium
231

A data analyst is performing K-means clustering on customer data. Which THREE of the following are steps in the K-means algorithm?

Hard
232

A dataset contains a column 'Age' with values: [22, 25, 25, 30, 35, 40, 45]. What is the interquartile range (IQR)?

Easy
233

A data scientist is preparing data for a K-means clustering algorithm. The dataset contains features measured in different units (e.g., income in dollars and age in years). Which preprocessing step is most critical before running K-means?

Medium
234

A data scientist is building a predictive model to forecast monthly sales. The data shows a linear trend with no seasonality. Which regression technique is most appropriate?

Medium
235

In a logistic regression model predicting customer churn (1 = churn, 0 = not churn), the coefficient for 'contract length' is -0.5. Which of the following is the correct interpretation?

Medium
236

A data analyst is preparing a dataset for analysis and needs to handle outliers. Which TWO of the following are common methods for treating outliers?

Medium
237

A healthcare analytics team is analyzing patient readmission rates. They have a dataset with thousands of records including patient age, diagnosis, length of stay, number of prior admissions, and discharge date. The goal is to identify key factors influencing readmission and create a model to predict high-risk patients. The data is imbalanced: only 5% of patients are readmitted within 30 days. The team plans to use logistic regression. What is the most appropriate approach?

Medium
238

A data analyst at a logistics company receives a shipment-tracking table in which the carrier identifier column contains values such as 'FEDX', 'FedEx', 'fed ex', and 'FX'. The analyst needs to consolidate these into a single carrier value before building a delivery-performance dashboard. Which data preparation step should the analyst perform?

Easy

Frequently asked questions

What does the Data Analysis domain cover on the DA0-002 exam?
Be able to pick the right analysis for a scenario, run and interpret descriptive and inferential statistics, and read clustering output. The single most important thing: correctly interpret results, including significance, effect size, and the limits of what the data can prove.
How many questions are in this domain?
This page lists all 238 Data Analysis questions in the DA0-002 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Data Analysis questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
data-plus DATA-PLUS dap analysing data Practice Questions