Courseiva

CCNA Data Analysis Questions

75 of 238 questions · Page 1/4 · Data Analysis · Answers revealed

1
MCQhard

An analyst runs an A/B test with 1000 users per group and observes a conversion rate of 5% in the control and 6% in the treatment. The p-value is 0.12. What should the analyst conclude?

A.The difference is not statistically significant at the 0.05 level.
B.The sample size is too small to detect an effect.
C.The treatment significantly outperforms control.
D.There is a 12% chance the treatment is better.
AnswerA

With a p-value of 0.12 exceeding the 0.05 significance threshold, the observed 1% conversion lift is within the range expected from random variation. The analyst cannot reject the null hypothesis, so the difference lacks statistical significance.

Why this answer

A p-value of 0.12 is greater than the conventional significance level of 0.05, so the null hypothesis of no difference cannot be rejected. The observed difference between 5% and 6% conversion rates is not statistically significant at the 0.05 level. The analyst should conclude that there is insufficient evidence to claim the treatment outperforms control.

Exam trap

DA0-002 often tests the misinterpretation of p-values, especially the false belief that a non-significant result proves no effect or that p-value equals the probability the treatment is better.

How to eliminate wrong answers

Option B is wrong because the sample size may or may not be too small; the p-value alone does not determine that, and with 1000 per group the test has reasonable power to detect large effects. Option C is wrong because a p-value of 0.12 does not indicate statistical significance; claiming the treatment significantly outperforms control is incorrect. Option D is wrong because a p-value is not the probability that the treatment is better; it is the probability of observing such data if the null hypothesis were true.

2
MCQeasy

In A/B testing, the null hypothesis typically states that:

A.There is no difference between the control and treatment groups
B.The treatment group will perform better than the control group
C.The sample size is sufficient for the test
D.There is a significant difference between the control and treatment groups
AnswerA

The null hypothesis in A/B testing posits zero effect: any observed difference between control and treatment arises from random variation alone. This satisfies the stem's requirement by defining the baseline against which statistical significance is measured, so rejecting it indicates the treatment genuinely influenced the metric rather than chance.

Why this answer

The null hypothesis (H0) is a statement of no effect or no difference between groups.

3
MCQeasy

A data analyst receives a sales dataset where the 'order_date' column is stored as text in the format '2024-03-15'. The analyst needs to calculate the number of days between each order date and the current date. Which action should the analyst take first?

A.Encode the 'order_date' column using one-hot encoding
B.Convert the 'order_date' column from text to a date data type
C.Apply a logarithm transformation to the 'order_date' column values
D.Impute missing values in the 'order_date' column with the column mean
AnswerB

Date arithmetic such as subtracting one date from another requires a true date or datetime data type, because text values are compared lexically rather than chronologically. Converting the column first ensures that functions like DATEDIFF or date subtraction produce correct day counts, and it also enables proper sorting and filtering by date.

Why this answer

Calculating elapsed days requires the order date to be a genuine date type so the platform can perform chronological subtraction. While the values look like dates, storing them as text means comparisons and arithmetic operate on character strings, producing errors or incorrect results. Converting the column to a date type is the prerequisite step that unlocks accurate day-difference calculations and correct date-based sorting.

Exam trap

The trap here is assuming that text formatted like a date behaves like a date, when arithmetic on text values fails or returns incorrect results.

4
MCQeasy

An analyst computed the mean, median, and mode of a dataset and found they are all equal. Which of the following best describes the distribution?

A.Bimodal
B.Negatively skewed
C.Positively skewed
D.Symmetric
AnswerD

Equal mean, median and mode indicates a symmetric distribution, where values balance evenly around the centre and the mode coincides with that central point. Skewed distributions pull the mean away from the median and mode, so this equality satisfies the stem's condition.

Why this answer

When mean, median, and mode are equal, the distribution is symmetric and unimodal, often resembling a normal distribution.

5
MCQhard

A time series of monthly sales data exhibits a clear upward trend over several years, with consistent peaks each December. Which components are present in this series?

A.Trend and seasonality
B.Cyclical and irregular components only
C.Seasonality and cyclical components only
D.Trend and irregular components only
AnswerA

The sustained upward movement over years is the trend component, while the recurring December peaks represent seasonality with a fixed period of twelve months. Both components coexist in the series, so trend and seasonality are present.

Why this answer

The time series exhibits an upward trend over several years and consistent peaks each December, which is seasonality. Therefore, the components present are trend and seasonality. Cyclical components are not mentioned, and irregular components are always present but not the focus.

Exam trap

DA0-002 often tests the identification of time series components, and candidates may confuse seasonality with cyclical patterns or overlook the trend when seasonality is prominent.

How to eliminate wrong answers

Option B is wrong because cyclical and irregular components only would not account for the clear upward trend and seasonal peaks. Option C is wrong because seasonality and cyclical components only would miss the trend. Option D is wrong because trend and irregular components only would miss the seasonality.

6
MCQhard

A dataset contains a feature with values ranging from 10 to 1000. The analyst applies min-max normalization to scale the feature between 0 and 1. What is the normalized value of 520?

A.0.515
B.0.510
C.0.480
D.0.520
AnswerA

Min-max normalisation applies (x − min) / (max − min), giving (520 − 10) / (1000 − 10) = 510 / 990 = 0.515. This satisfies the stem's constraint of scaling the 10–1000 range onto 0–1, correctly mapping 520 to 0.515.

Why this answer

Min-max normalization scales values using the formula (x - min) / (max - min). With min=10 and max=1000, the normalized value of 520 is (520 - 10) / (1000 - 10) = 510 / 990 = 0.51515..., which rounds to 0.515. Option A correctly applies this formula.

Exam trap

The trap here is using the maximum value (1000) as the denominator instead of the range (max - min = 990), which leads to option B or D.

How to eliminate wrong answers

Option B (0.510) is wrong because it results from dividing 510 by 1000 instead of by the range (990), a common mistake of using the raw max rather than max minus min. Option C (0.480) is wrong because it appears to use an incorrect range or subtract the wrong minimum, producing a value that does not match the formula. Option D (0.520) is wrong because it simply divides 520 by 1000, ignoring the min offset entirely.

7
MCQmedium

An analyst is comparing the average sales of two different store locations using a t-test. The p-value obtained is 0.03, and the significance level is 0.05. What should the analyst conclude?

A.Fail to reject the null hypothesis; no significant difference
B.The test is inconclusive because the p-value is too low
C.Reject the null hypothesis; there is a significant difference
D.Accept the null hypothesis; the means are equal
AnswerC

A p-value of 0.03 falls below the 0.05 significance level, so the null hypothesis of no difference is rejected. The result is statistically significant, indicating the two store locations' average sales genuinely differ rather than reflecting random sampling variation.

Why this answer

When the p-value (0.03) is less than the significance level (0.05), the analyst rejects the null hypothesis. This means there is statistically significant evidence of a difference between the average sales of the two store locations.

Exam trap

DA0-002 often tests the p-value vs. α comparison — candidates confuse 'low p-value' with 'inconclusive' or mistakenly say you 'accept' the null hypothesis instead of 'fail to reject' it.

How to eliminate wrong answers

Option A is wrong because failing to reject requires p > α; here 0.03 < 0.05, so the null is rejected. Option B is wrong because a low p-value is not 'inconclusive' — it is the evidence needed to reject the null. Option D is wrong because you never 'accept' the null hypothesis in frequentist testing; you either reject or fail to reject it, and here the null is rejected anyway.

8
Multi-Selectmedium

A data analyst is preparing to run an A/B test comparing two email subject lines. Which TWO of the following should the analyst define before the test begins?

Select 2 answers
A.The exact lift in conversion rate
B.The time series decomposition
C.The p-value after the test
D.The null and alternative hypotheses
E.The sample size required for the desired power
AnswersD, E

Stating the null and alternative hypotheses before the A/B test begins fixes exactly what is being tested and which outcome counts as significant. This pre-registration prevents post-hoc interpretation of results and satisfies the stem's requirement to define parameters before the test starts.

Why this answer

Before A/B testing, define null and alternative hypotheses, and determine sample size needed for desired statistical power and effect size.

9
MCQmedium

A retailer wants to test if a new website layout increases the average time spent on the site. They split traffic: control group (old layout) and treatment group (new layout). Which statistical test is most appropriate to compare the average time spent between the two groups?

A.ANOVA
B.Pearson correlation
C.Chi-square test
D.Two-sample t-test
AnswerD

A two-sample t-test compares the means of two independent groups, matching the control and treatment split. It tests whether the difference in average time spent is statistically significant, which is exactly the retailer's question about whether the new layout changes the mean.

Why this answer

The question compares the means of a continuous outcome (average time spent) between two independent groups (control vs. treatment), which is exactly the two-sample t-test's purpose. The two-sample t-test evaluates whether the difference in group means is statistically significant, assuming approximately normal distributions or sufficiently large samples. It is the standard test for a two-group A/B comparison on a continuous metric.

Exam trap

The trap is confusing the two-sample t-test with the paired t-test — candidates must check whether the two groups are independent (different subjects) or paired (same subjects measured twice), since that determines which t-test applies.

How to eliminate wrong answers

Option A is wrong because ANOVA is used to compare means across three or more groups; with only two groups it is equivalent to a t-test but is not the most appropriate or standard choice. Option B is wrong because Pearson correlation measures the linear relationship between two continuous variables, not the difference in means between two groups. Option C is wrong because the chi-square test is used for categorical data (comparing observed vs. expected frequencies), not for comparing continuous means like time spent.

10
MCQeasy

A data analyst is working with a dataset that contains a column for 'order_date' stored as text in the format 'YYYY-MM-DD'. The analyst needs to calculate the average time between orders for each customer. Which of the following steps should the analyst take first?

A.Calculate the difference between consecutive 'order_date' values using string manipulation.
B.Group the data by customer and compute the average of the 'order_date' column.
C.Convert the 'order_date' column to a date data type using a function like TO_DATE or PARSE_DATE.
D.Sort the dataset by the 'order_date' column in ascending order.
AnswerC

Converting the text column to a proper date data type is essential before performing any date arithmetic. Functions like TO_DATE (in SQL) or PARSE_DATE (in BigQuery) transform the string into a date value, enabling calculations such as differences between dates. Without this step, the analyst cannot accurately compute time intervals, as text comparison would yield incorrect results.

Why this answer

To calculate the average time between orders, the analyst must first convert the text-based date column into a proper date data type. This enables the use of date functions to compute intervals between consecutive orders. Without this conversion, any arithmetic on dates would be invalid or highly error-prone.

Once converted, the analyst can sort by customer and date, then compute differences.

Exam trap

The trap here is assuming that sorting or string manipulation can substitute for proper date conversion, but accurate date arithmetic requires a date data type.

11
MCQmedium

A data analyst is performing a hypothesis test with a significance level of 0.05. The p-value obtained is 0.03. What should the analyst conclude?

A.Reject the null hypothesis
B.Fail to reject the null hypothesis
C.Accept the null hypothesis
D.The result is practically significant
AnswerA

The p-value of 0.03 falls below the 0.05 significance level, meaning the observed result is unlikely under the null hypothesis. The analyst therefore rejects the null hypothesis in favour of the alternative, accepting the result as statistically significant.

Why this answer

Since the p-value (0.03) is less than the significance level (0.05), the result is statistically significant. This means the observed data provides sufficient evidence to reject the null hypothesis in favor of the alternative hypothesis. The analyst should conclude that there is a statistically significant effect or difference.

Exam trap

The trap here is that candidates often confuse 'fail to reject' with 'accept' the null hypothesis, or they mistakenly think a p-value less than α means the null hypothesis is proven false with certainty, rather than just providing sufficient evidence to reject it.

How to eliminate wrong answers

Option B is wrong because failing to reject the null hypothesis occurs only when the p-value is greater than or equal to the significance level (p ≥ 0.05), not when it is smaller. Option C is wrong because hypothesis testing never 'accepts' the null hypothesis; we either reject it or fail to reject it, as acceptance implies proof of truth, which is not a valid statistical conclusion. Option D is wrong because practical significance is a separate consideration from statistical significance; a statistically significant result (p < 0.05) does not automatically imply practical importance, and the question only asks about the hypothesis test conclusion.

12
MCQhard

A data analyst runs an A/B test on a new website layout. The test yields a p-value of 0.04 with the null hypothesis being no difference in conversion rates. The significance threshold is α=0.05. Which of the following is the correct conclusion?

A.The result is not significant; accept the alternative hypothesis.
B.Reject the null hypothesis; the new layout is proven to increase conversions.
C.Reject the null hypothesis; there is a statistically significant difference in conversion rates.
D.Fail to reject the null hypothesis; there is no evidence of a difference.
AnswerC

With p = 0.04 falling below the α = 0.05 threshold, the result lies inside the rejection region, so the null hypothesis of no difference is rejected. This satisfies the stem's significance criterion, supporting a statistically significant difference in conversion rates between the layouts.

Why this answer

Since p-value (0.04) < α (0.05), we reject the null hypothesis and conclude there is a statistically significant difference. However, statistical significance does not guarantee practical significance.

13
Multi-Selectmedium

A data analyst is performing a chi-square test of independence on a 2x2 contingency table. The p-value is 0.04. At α=0.05, which THREE of the following statements are correct?

Select 3 answers
A.There is a statistically significant association between the two variables.
B.The test indicates a strong association between variables.
C.The variables are not independent.
D.The null hypothesis is rejected.
E.The result is not statistically significant.
AnswersA, C, D

With p=0.04 below the α=0.05 threshold, the null hypothesis of independence is rejected, so a statistically significant association exists between the two variables. The test detects a relationship, though it does not quantify its strength or direction.

Why this answer

Option A is correct because with p = 0.04 < α = 0.05, the chi-square test of independence yields a statistically significant result, indicating evidence of an association between the two variables. Option C is correct because rejecting the null hypothesis of independence means the data support that the two variables are not independent (i.e., they are associated). Option D is correct because the decision rule for a chi-square test of independence is to reject the null hypothesis when p < α, and 0.04 < 0.05.

Option B is not correct because statistical significance does not imply a strong association; strength would require an effect size such as Cramér's V or the phi coefficient. Option E is not correct because p = 0.04 is below α = 0.05, so the result is statistically significant, not non-significant.

Exam trap

DA0-002 often tests the misconception that a statistically significant p-value implies a strong or large effect, conflating significance with effect size.

14
Multi-Selectmedium

Which TWO of the following data quality dimensions are most directly affected by duplicate records?

Select 2 answers
A.Timeliness
B.Consistency
C.Uniqueness
D.Accuracy
E.Completeness
AnswersC, D

Duplicate records mean the same real-world entity appears more than once, directly violating the uniqueness dimension, which measures whether each record occurs only once. This is the dimension most immediately degraded by duplication, independent of whether the underlying values are otherwise correct.

Why this answer

Duplicate records most directly undermine Uniqueness (C), because uniqueness is the data quality dimension that measures whether each real-world entity appears exactly once in a dataset; duplicates violate this one-record-per-entity rule by definition. They also directly affect Accuracy (D), since duplicated rows distort counts, aggregations, and the true representation of the entity, making the stored data an incorrect reflection of reality. Timeliness (A) concerns whether data is current and available when needed, which duplicates do not inherently affect.

Consistency (B) refers to agreement of the same data across systems or formats, which is a separate issue from repeated rows. Completeness (E) measures whether required data is present, and duplicates add data rather than omit it, so it is not the dimension most directly impacted.

Exam trap

The trap is assuming that duplicates only affect uniqueness, but they also impact accuracy because the data no longer correctly represents the real-world entity.

15
MCQmedium

In a dataset with variables on different scales (e.g., age in years and income in dollars), which preprocessing step is necessary before applying k-means clustering?

A.Feature selection
B.Dimensionality reduction
C.Normalization (scaling)
D.One-hot encoding
AnswerC

Normalization ensures each feature contributes equally to distance calculations.

Why this answer

K-means clustering relies on Euclidean distance to measure similarity between data points. When variables like age (in years) and income (in dollars) are on different scales, the variable with larger numeric values (income) will dominate the distance calculation, skewing the clustering results. Normalization (scaling), such as min-max scaling or z-score standardization, rescales all features to a comparable range (e.g., [0,1] or mean=0, variance=1), ensuring each feature contributes equally to the distance computation.

Exam trap

The trap here is that candidates may confuse normalization with other preprocessing steps like feature selection or dimensionality reduction, thinking that removing irrelevant features or reducing dimensions will automatically fix scale differences, but k-means specifically requires scaling to ensure equal feature influence in distance calculations.

How to eliminate wrong answers

Option A is wrong because feature selection is about choosing a subset of relevant features to reduce noise or improve model performance, but it does not address the issue of differing scales among features, which is required before k-means. Option B is wrong because dimensionality reduction (e.g., PCA) reduces the number of features, but it does not inherently scale the data; scaling is typically performed before dimensionality reduction, not as a substitute for it. Option D is wrong because one-hot encoding is used to convert categorical variables into numerical format, not to handle numerical variables on different scales; applying one-hot encoding to already numerical features would be incorrect and does not solve the scaling problem.

16
MCQhard

A data analyst at a retail company is building a multiple linear regression model to forecast weekly sales. The dataset contains 50 predictor variables, including store size, promotional spend, holiday indicators, and many others. After training the model, the analyst observes an R-squared of 0.99 on the training set but only 0.55 on the holdout test set. Which action should the analyst take first to address this discrepancy?

A.Remove highly correlated predictor variables and apply regularization (e.g., Ridge or Lasso).
B.Add more predictor variables to increase the training R-squared further.
C.Use k-fold cross-validation with a different random seed to get a more reliable test set estimate.
D.Increase the number of hidden layers in the model to capture more complexity.
AnswerA

The 0.99 versus 0.55 gap indicates overfitting from 50 predictors, many correlated. Removing correlated variables reduces multicollinearity and dimensionality, while Ridge or Lasso regularisation penalises large coefficients, shrinking variance and improving holdout generalisation before other remedies are attempted.

Why this answer

The high R-squared of 0.99 on training data versus 0.55 on test data is a classic sign of overfitting, where the model has learned noise and specific patterns in the training set that do not generalize. Removing highly correlated predictors reduces multicollinearity and model complexity, while regularization (Ridge or Lasso) penalizes large coefficients, shrinking them to prevent overfitting. This is the most direct first step to improve generalization.

Exam trap

The trap here is that candidates may think a high R-squared is always good, or they may confuse overfitting with underfitting and choose to add more complexity (Option D) or more data (Option B), rather than recognizing the need to reduce model complexity and apply regularization.

How to eliminate wrong answers

Option B is wrong because adding more predictor variables would increase the training R-squared but worsen overfitting, making the test set performance even lower. Option C is wrong because k-fold cross-validation with a different random seed does not address the fundamental overfitting issue; it only provides a different estimate of test error but does not change the model's tendency to overfit. Option D is wrong because increasing the number of hidden layers (a neural network technique) is irrelevant for a multiple linear regression model and would introduce unnecessary complexity, likely exacerbating overfitting.

17
Multi-Selectmedium

A data analyst is evaluating a classification model that predicts whether a customer will churn. The model's confusion matrix shows 80 true positives, 20 false negatives, 30 false positives, and 120 true negatives. Which TWO of the following metrics can be directly calculated from this confusion matrix? (Choose two.)

Select 2 answers
A.Accuracy
B.R-squared
C.Root Mean Square Error (RMSE)
D.Precision
E.Adjusted R-squared
AnswersA, D

Accuracy is the proportion of correct predictions: (TP + TN) / total. Here, (80 + 120) / (80+20+30+120) = 200/250 = 0.8. It can be directly calculated from the confusion matrix, making it a valid metric for this scenario. It provides an overall measure of correct classifications.

Why this answer

Accuracy and precision are both classification metrics that can be computed directly from the counts in a confusion matrix. Accuracy gives the overall correct prediction rate, while precision focuses on the reliability of positive predictions. R-squared, RMSE, and adjusted R-squared are regression metrics and cannot be calculated from a confusion matrix, so they are not applicable.

Exam trap

The trap here is selecting regression metrics like R-squared or RMSE for a classification problem, confusing the evaluation metrics of different model types.

18
Multi-Selecthard

Which THREE of the following are appropriate methods to handle outliers in a dataset?

Select 3 answers
A.Transforming the data using log transformation
B.Removing the outlier records
C.Capping the outlier values at a certain percentile
D.Binning continuous variables
E.Imputing outliers with the mean
AnswersA, B, C

A log transformation compresses the upper tail of a right-skewed distribution, pulling extreme values closer to the bulk of the data. This reduces the leverage outliers exert on model estimates without deleting records, making it an appropriate handling method.

Why this answer

Log transformation (A) is correct because applying a log (or similar power) transform compresses the scale of extreme values, reducing the leverage of right-skewed outliers and making the distribution more symmetric for downstream models. Removing the outlier records (B) is correct when outliers are confirmed to be data-entry errors or otherwise invalid observations, since deleting them prevents distortion of statistics such as the mean and variance. Capping the outlier values at a certain percentile (C), also called winsorizing, is correct because it clips extreme values to a chosen percentile boundary (e.g., the 1st and 99th percentiles), preserving the record while limiting the outlier's influence.

Binning continuous variables (D) is not a dedicated outlier-handling method; it discretizes values into bins and merely masks extremes rather than addressing them. Imputing outliers with the mean (E) is inappropriate because the mean itself is distorted by the outliers, so replacing them with it introduces bias and does not resolve the underlying problem.

Exam trap

The trap here is that candidates may confuse data preprocessing techniques like binning or imputation with outlier handling methods, but binning is for discretization and mean imputation is not robust for outliers, while the correct methods (transformation, removal, capping) directly address outlier impact.

19
Multi-Selectmedium

An analyst is preparing data for a clustering algorithm that uses Euclidean distance. Which TWO data preprocessing techniques should be applied to ensure all features contribute equally?

Select 2 answers
A.Min-max normalization
B.Z-score standardization
C.Log transformation
D.Principal component analysis
E.One-hot encoding
AnswersA, B

Min-max normalization rescales each feature to a fixed 0–1 range, removing unit and magnitude disparities before Euclidean distance is computed. This equalises each feature's contribution, satisfying the requirement that no single large-range variable dominates the clustering metric.

Why this answer

Min-max normalization (A) is correct because it rescales each feature to a common range, typically [0, 1], via (x - min) / (max - min), so that no feature dominates the Euclidean distance calculation due to a larger scale. Z-score standardization (B) is also correct because it transforms each feature to mean 0 and standard deviation 1 using (x - μ) / σ, which equalizes the variance and ensures all features contribute equally to Euclidean distance. Log transformation (C) is not a scale-equalizing technique; it only compresses skewed distributions and does not guarantee equal feature contribution.

Principal component analysis (D) is a dimensionality-reduction method, not a preprocessing step for equalizing feature scales. One-hot encoding (E) is for converting categorical variables into binary vectors and does not address differing numeric scales.

Exam trap

DA0-002 often tests the confusion between scaling techniques and other preprocessing methods like log transformation or PCA. Candidates might think log transformation scales features, but it only reduces skewness; or that PCA is a scaling method, but it's for dimensionality reduction.

20
MCQmedium

An analyst calculates a Pearson correlation coefficient of -0.8 between advertising spend and customer churn rate. Which interpretation is correct?

A.There is a weak positive relationship.
B.Advertising spend causes churn to decrease.
C.64% of the variance in churn is explained by spend.
D.Increasing advertising spend is associated with decreasing churn rate.
AnswerD

A coefficient of -0.8 indicates a strong negative linear association: as advertising spend rises, churn rate tends to fall. This satisfies the stem's requirement to interpret the sign and magnitude correctly, since the negative sign denotes an inverse relationship rather than causation or a weak correlation.

Why this answer

A Pearson correlation of -0.8 indicates a strong negative linear relationship, meaning that as advertising spend increases, customer churn rate tends to decrease. The correct interpretation is directional association, not causation — the coefficient describes the strength and direction of the linear relationship between the two variables.

Exam trap

DA0-002 often tests the correlation-versus-causation distinction and the r-versus-r² confusion, so candidates who see a strong coefficient jump to a causal claim (B) or misapply the variance-explained percentage (C).

How to eliminate wrong answers

Option A is wrong because -0.8 is a strong correlation (|r| close to 1), not weak, and the negative sign indicates an inverse relationship, not positive. Option B is wrong because correlation does not establish causation — the analyst cannot conclude that advertising spend causes churn to decrease without a controlled experiment or causal inference method. Option C is wrong because 64% (r² = 0.64) is the coefficient of determination for variance explained, but the question asks for the correct interpretation of the correlation coefficient itself, and stating '64% of variance' misapplies the statistic to the correlation value rather than describing the relationship.

21
MCQmedium

A data analyst wants to segment customers into groups based on their purchasing behavior. The dataset includes numerical features such as annual income and purchase frequency. Which algorithm is most appropriate for this task?

A.Linear regression
B.K-means clustering
C.Logistic regression
D.Chi-square test
AnswerB

K-means clustering partitions numerical observations into k groups by minimising within-cluster variance, which suits segmenting customers on annual income and purchase frequency. It satisfies the stem's unsupervised grouping requirement, unlike classification algorithms that need labelled outcomes.

Why this answer

K-means clustering is an unsupervised learning algorithm that partitions data points into k groups based on similarity of numerical features, which is exactly what customer segmentation requires. Since the task is to discover natural groupings in unlabeled data using features like annual income and purchase frequency, K-means is the standard and most appropriate choice.

Exam trap

DA0-002 often tests the confusion between supervised and unsupervised algorithms — candidates see 'segment customers' and mistakenly reach for classification (logistic regression) instead of clustering.

How to eliminate wrong answers

Option A is wrong because linear regression is a supervised algorithm that predicts a continuous target variable, not one that discovers groups in unlabeled data. Option C is wrong because logistic regression is a supervised classification algorithm that predicts a categorical outcome from labeled training data, which is not the case here. Option D is wrong because the chi-square test is a statistical hypothesis test for independence between categorical variables, not a clustering or segmentation algorithm.

22
MCQmedium

A data analyst needs to determine whether the mean sales of two different regions are significantly different. The samples are independent and the data is normally distributed. Which statistical test should be used?

A.Chi-square test for independence
B.ANOVA
C.Independent samples t-test
D.Paired t-test
AnswerC

Two independent, normally distributed samples compared on a continuous mean call for the independent samples t-test, which assesses whether the difference between group means exceeds sampling variability. A paired test would require matched observations, which the stem excludes.

Why this answer

The independent samples t-test is the correct choice because the scenario involves comparing the means of two independent groups (two different regions) with normally distributed data. This test specifically assesses whether the difference between the two sample means is statistically significant, assuming equal or unequal variances as determined by Levene's test.

Exam trap

CompTIA often tests the distinction between independent and paired t-tests, trapping candidates who overlook the 'independent samples' condition and mistakenly choose the paired t-test for any two-group comparison.

How to eliminate wrong answers

Option A is wrong because the Chi-square test for independence is used for categorical data to assess associations between two variables, not for comparing means of continuous data. Option B is wrong because ANOVA is used to compare means among three or more groups, not exactly two independent groups. Option D is wrong because the paired t-test is used for dependent samples (e.g., before-and-after measurements on the same subjects), not for independent samples from different regions.

23
Multi-Selecthard

A data scientist is conducting an A/B test with a significance level of 0.05. Which three factors should be considered when calculating the required sample size? (Choose THREE)

Select 3 answers
A.Seasonality of the data
B.Statistical power (e.g., 0.80)
C.Minimum detectable effect size
D.Number of clusters in k-means
E.Significance level (α)
AnswersB, C, E

Statistical power, typically 0.80, fixes the tolerated Type II error rate and determines how large a sample is needed to detect a real effect. It is a required input alongside significance level and effect size.

Why this answer

Option B is correct because statistical power (typically 0.80 or 80%) directly determines sample size: higher power requires a larger sample to reliably detect a true effect if one exists. Option C is correct because the minimum detectable effect size is a core input to sample-size formulas (e.g., n ≈ 16σ²/Δ² for a two-sample t-test at α=0.05, power=0.80); smaller effects demand much larger samples. Option E is correct because the significance level α (here 0.05) sets the Type I error threshold and appears in every sample-size calculation, with smaller α requiring larger samples.

Option A is not a direct input to the standard sample-size formula, though seasonality may inform variance estimates or test design. Option D is unrelated: the number of clusters in k-means is an unsupervised clustering parameter and has no bearing on A/B test sample-size determination.

Exam trap

The trap is including operational or data-context factors (seasonality, clustering hyperparameters) that feel relevant to experiment design but are not inputs to the statistical power calculation — only power, effect size, and α are.

24
MCQhard

A logistic regression model is used to predict the probability of customer churn. The model's coefficient for the feature 'customer support calls' is 0.8 with a p-value of 0.001. Which interpretation is correct?

A.For each additional support call, the log-odds of churn increase by 0.8, and this effect is statistically significant.
B.The odds of churn are multiplied by 0.8 for each additional call.
C.Support calls have no significant effect on churn.
D.For each additional support call, the probability of churn increases by 80%.
AnswerA

In logistic regression, coefficients are expressed in log-odds, so a coefficient of 0.8 means each additional support call raises the log-odds of churn by 0.8. The p-value of 0.001, below the 0.05 threshold, confirms this effect is statistically significant rather than due to chance.

Why this answer

In logistic regression, coefficients are expressed in log-odds units. A coefficient of 0.8 for 'customer support calls' means each additional call increases the log-odds of churn by 0.8, holding other predictors constant. The p-value of 0.001 is below the conventional 0.05 threshold, so the effect is statistically significant.

Exam trap

DA0-002 often tests the confusion between log-odds coefficients and odds ratios — candidates must remember that the raw coefficient is in log-odds, and only e^β gives the multiplicative odds change.

How to eliminate wrong answers

Option B is wrong because it confuses the log-odds coefficient with an odds ratio — the odds ratio would be e^0.8 ≈ 2.23, meaning odds multiply by about 2.23, not 0.8. Option C is wrong because a p-value of 0.001 is highly significant, directly contradicting the claim of no effect. Option D is wrong because a coefficient of 0.8 does not translate to an 80% probability increase; probability changes depend on the baseline probability and the logistic curve, not a linear 80% shift.

25
MCQhard

In time series decomposition, a data analyst separates a retail sales series into trend, seasonal, and residual components. After decomposition, the residual component shows no pattern and is random. Which of the following best describes the seasonal component?

A.Cyclical variations lasting more than a year.
B.Irregular fluctuations that cannot be predicted.
C.Regular patterns that repeat at fixed intervals.
D.A long-term increase or decrease in sales.
AnswerC

The seasonal component captures periodic fluctuations that recur at fixed intervals, such as weekly or yearly cycles, distinct from the trend's long-term direction and the residual's random noise. The stem's random residual confirms seasonality is the regular, repeating pattern.

Why this answer

The seasonal component captures regular, repeating patterns at fixed intervals — for example, higher retail sales every December or every weekend. Because the residual shows no pattern (random noise), the decomposition has successfully isolated the systematic periodic structure into the seasonal component. This is the defining characteristic of seasonality in time series decomposition.

Exam trap

The trap is conflating the four decomposition components — candidates must distinguish seasonal (fixed repeating period) from cyclical (longer, irregular), trend (long-term direction), and residual (random noise).

How to eliminate wrong answers

Option A is wrong because cyclical variations lasting more than a year are a distinct component (the 'cyclical' component), not the seasonal component — seasonality has a fixed, known period (e.g., weekly, monthly, quarterly). Option B is wrong because irregular, unpredictable fluctuations describe the residual (noise) component, which the question states is random. Option D is wrong because a long-term increase or decrease describes the trend component, not the seasonal component.

26
MCQeasy

A data analyst calculates the mean, median, and mode of a dataset. Which measure of central tendency is most affected by extreme outliers?

A.Mean
B.Range
C.Median
D.Mode
AnswerA

Every value contributes arithmetically to the mean, so a single extreme outlier shifts it substantially. The median depends only on positional rank, and the mode only on frequency, leaving both largely resistant to extreme values.

Why this answer

The mean is sensitive to extreme values because it includes all data points in its calculation, whereas median and mode are more robust.

27
MCQmedium

A retail company wants to identify customer segments based on purchase history and demographics. Which technique is most appropriate for this task?

A.Linear regression
B.K-means clustering
C.Chi-square test
D.Logistic regression
AnswerB

K-means clustering partitions unlabelled data into k groups by minimising within-cluster variance, directly satisfying the requirement to identify customer segments from purchase history and demographics without predefined labels. Unlike classification, it discovers natural groupings, making it appropriate for exploratory segmentation of retail customers.

Why this answer

K-means clustering is an unsupervised learning algorithm that partitions observations into k groups based on similarity across multiple features, making it ideal for segmenting customers by purchase history and demographics. It doesn't require labeled outcomes, which matches the exploratory nature of customer segmentation. The algorithm iteratively assigns points to the nearest centroid and recalculates centroids until convergence.

Exam trap

The trap is confusing supervised classification (logistic regression) with unsupervised clustering (K-means) — candidates who see 'segments' and think 'categories' pick logistic regression, but segmentation discovers groups rather than predicting predefined labels.

How to eliminate wrong answers

Option A is wrong because linear regression predicts a continuous outcome from input features — it's supervised and requires a target variable, which segmentation doesn't have. Option C is wrong because a chi-square test assesses association between categorical variables; it tests hypotheses rather than grouping observations. Option D is wrong because logistic regression predicts a binary outcome (e.g., churn yes/no) and is supervised — it classifies into predefined categories rather than discovering natural groupings.

28
MCQmedium

A data scientist is using K-means clustering with k=3. After the first iteration, the centroids are recalculated. Which step occurs next in the algorithm?

A.Calculate the sum of squared errors
B.Stop the algorithm because k is fixed
C.Compute the elbow curve
D.Assign each point to the nearest centroid
AnswerD

Reassigning every point to its nearest centroid follows each recalculation, since K-means alternates between these two phases until convergence. After the first iteration's centroid update, the algorithm must reassign points before recalculating again, satisfying the stem's "which step occurs next" constraint.

Why this answer

After centroids are recalculated, K-means returns to the assignment step, reassigning each data point to the nearest (typically Euclidean distance) centroid. This assignment-then-update loop repeats until convergence — either centroids stop moving, assignments stabilize, or a maximum iteration count is reached. So the next step after recalculation is reassignment.

Exam trap

The trap is thinking the algorithm stops after centroid recalculation — candidates who confuse 'k is fixed' with 'algorithm is done' pick option B, but K-means iterates assignment and update until convergence.

How to eliminate wrong answers

Option A is wrong because sum of squared errors (inertia) is computed to monitor convergence or select k via the elbow method, but it's not the immediate next algorithmic step after centroid recalculation. Option B is wrong because fixing k does not terminate the algorithm — k is a hyperparameter set before training; the algorithm stops only on convergence criteria. Option C is wrong because the elbow curve is a diagnostic plot generated across multiple k values before or after training, not a step inside a single K-means run.

29
MCQmedium

In a multiple regression model, one predictor has a high p-value (0.45). What should the analyst consider doing?

A.Transform the predictor
B.Keep the predictor regardless
C.Remove the predictor from the model
D.Increase the sample size
AnswerC

A p-value of 0.45 exceeds typical significance thresholds, indicating the predictor's coefficient is not statistically distinguishable from zero given the other variables. Removing it simplifies the model and reduces multicollinearity, though the analyst should first check theoretical relevance.

Why this answer

A p-value of 0.45 is far above the conventional significance threshold (typically 0.05), meaning there is no statistically significant evidence that this predictor is associated with the outcome variable after accounting for the other predictors. Removing non-significant predictors simplifies the model, reduces multicollinearity risk, and improves interpretability without meaningful loss of predictive power. This is the standard first step in backward elimination-style model refinement.

Exam trap

The trap here is confusing statistical insignificance with a need for data transformation or more data — candidates often assume any problematic predictor requires a technical fix rather than simple removal from the model.

How to eliminate wrong answers

Option A is wrong because transforming a predictor is used to fix non-linearity or heteroscedasticity issues, not to address statistical insignificance — a transformation won't make an irrelevant variable relevant. Option B is wrong because keeping a high-p-value predictor inflates model complexity, can introduce noise, and may destabilize coefficient estimates through multicollinearity. Option D is wrong because increasing sample size addresses statistical power for detecting small effects, but a p-value of 0.45 indicates the effect is essentially absent, not merely underpowered.

30
MCQeasy

A data analyst needs to join two tables in a SQL database: Orders and Customers. The analyst wants to include all orders, even if there is no matching customer record. Which type of join should be used?

A.RIGHT JOIN
B.FULL OUTER JOIN
C.LEFT JOIN
D.INNER JOIN
AnswerC

A LEFT JOIN returns every row from the left table (Orders) and matches rows from Customers where they exist, filling NULLs otherwise. Because the stem requires all orders to appear even without a matching customer, this outer join preserves the unmatched left-side rows.

Why this answer

A LEFT JOIN returns all rows from the left table (Orders) and the matching rows from the right table (Customers). If there is no match, NULL values are returned for the right table's columns. This satisfies the requirement to include all orders, even those without a matching customer record.

Exam trap

The trap here is that candidates often confuse LEFT JOIN with RIGHT JOIN, mistakenly thinking they need to 'keep all customers' instead of 'keep all orders,' or they overcomplicate the requirement by choosing FULL OUTER JOIN when only one side needs to be preserved.

How to eliminate wrong answers

Option A (RIGHT JOIN) is wrong because it returns all rows from the right table (Customers) and matching rows from the left table (Orders), which would include all customers, not all orders. Option B (FULL OUTER JOIN) is wrong because it returns all rows from both tables, including unmatched rows from both sides, which is unnecessary when the requirement is specifically to keep all orders. Option D (INNER JOIN) is wrong because it returns only rows where there is a match in both tables, which would exclude orders without a matching customer record.

31
MCQhard

In a multiple regression model with three predictors, the coefficient for one predictor is 5.2 with a p-value of 0.001. Which of the following is the best interpretation?

A.The predictor explains 5.2% of the variance in the dependent variable.
B.A one-unit increase in the predictor decreases the dependent variable by 5.2 units, on average.
C.The model is not a good fit because one predictor is significant.
D.The predictor has a statistically significant effect on the dependent variable, controlling for other variables.
AnswerD

A p-value of 0.001 falls below the 0.05 threshold, so the null hypothesis is rejected. The coefficient of 5.2 indicates the predictor's effect on the dependent variable while the other two predictors are held constant.

Why this answer

A coefficient of 5.2 with a p-value of 0.001 means that, holding the other predictors constant, a one-unit increase in this predictor is associated with a 5.2-unit increase in the dependent variable, and the effect is statistically significant at conventional thresholds (p < 0.05). The phrase 'controlling for other variables' is key in multiple regression — the coefficient represents the partial effect of this predictor given the others in the model.

Exam trap

The trap is misreading the coefficient as a proportion of variance explained (R²) or ignoring the sign — candidates must remember that a regression coefficient is a per-unit effect, not a percentage, and that its sign indicates direction.

How to eliminate wrong answers

Option A is wrong because the coefficient value (5.2) is not a percentage of variance explained — variance explained is measured by R², and the coefficient's magnitude depends on the units of the variables. Option B is wrong because the sign is positive (5.2), indicating an increase, not a decrease, in the dependent variable. Option C is wrong because a significant predictor is generally a sign of a useful model term, not evidence that the model is a poor fit — model fit is assessed via R², adjusted R², or F-statistics.

32
MCQmedium

A company wants to determine if there is a significant difference in the average sales revenue between two different store layouts. They collect sales data from 30 stores with Layout A and 30 stores with Layout B. Which statistical test is most appropriate for comparing the means of these two independent groups?

A.ANOVA
B.Chi-square test
C.Paired t-test
D.Two-sample t-test
AnswerD

A two-sample t-test compares the means of two independent groups when the population standard deviation is unknown, which fits the 30 Layout A and 30 Layout B stores. It determines whether the observed revenue difference is statistically significant.

Why this answer

The scenario compares the means of two independent groups (Layout A vs. Layout B) on a continuous outcome (sales revenue), which is precisely what the two-sample t-test is designed for. With 30 stores per group, the sample size is adequate for the t-test's assumptions, and the groups are independent (different stores).

This is the standard test for a two-group mean comparison.

Exam trap

The trap is confusing the two-sample t-test with the paired t-test or ANOVA — candidates must check whether the groups are independent (two-sample), matched (paired), or more than two (ANOVA).

How to eliminate wrong answers

Option A is wrong because ANOVA is used to compare means across three or more groups; with only two groups, the two-sample t-test is the appropriate and more specific choice. Option B is wrong because the chi-square test is for categorical data (comparing observed vs. expected frequencies), not for comparing continuous means like sales revenue. Option C is wrong because the paired t-test requires matched or repeated measurements on the same subjects — here the stores in Layout A and Layout B are different, independent groups.

33
MCQmedium

A data analyst is examining a dataset of customer orders and notices that the 'order_total' column contains values ranging from $5 to $50,000, with most orders under $200. The analyst wants to reduce the influence of extreme values on a linear regression model predicting customer lifetime value. Which technique should the analyst apply?

A.Convert order_total to a categorical variable by binning into deciles.
B.Standardize the order_total variable to have a mean of 0 and standard deviation of 1.
C.Remove all orders with a total greater than $1,000 from the dataset.
D.Apply a logarithmic transformation to the order_total variable.
AnswerD

A logarithmic transformation compresses the scale of large values, reducing the leverage of extreme high order totals on the regression model. This makes the relationship between order_total and the target more linear and stabilizes variance, which is appropriate when data is positively skewed with outliers.

Why this answer

The logarithmic transformation is effective because it compresses the range of large values, making the distribution more symmetric and reducing the influence of outliers on the regression model. Standardization, removal, or binning either do not address the skewness or introduce other issues such as information loss or bias.

Exam trap

The trap here is assuming that standardization alone will handle outliers, when it actually preserves their relative extremity.

34
MCQmedium

A retail company wants to test whether a new website layout increases the conversion rate compared to the current layout. They randomly assign visitors to either the control or treatment group. Which statistical test is most appropriate to compare the conversion rates?

A.Two-sample t-test
B.Chi-square test
C.ANOVA
D.Logistic regression
AnswerB

Conversion rate is a binary outcome (converted or not) across two independent groups, so the chi-square test compares observed versus expected frequencies in a contingency table. It satisfies the stem's requirement to test whether the new layout's conversion rate differs from the control.

Why this answer

The outcome variable is conversion (yes/no), a categorical binary variable, and the predictor is group (control vs treatment), also categorical. The chi-square test of independence is designed to compare observed versus expected frequencies across categories, making it the correct choice for comparing two conversion rates.

Exam trap

DA0-002 often tests the distinction between tests for means (t-test, ANOVA) and tests for proportions/frequencies (chi-square) — candidates pick t-test because they see 'two groups' and forget the outcome is binary, not continuous.

How to eliminate wrong answers

Option A is wrong because a two-sample t-test compares means of a continuous variable, not proportions of a binary outcome. Option C is wrong because ANOVA compares means across three or more groups and still assumes a continuous dependent variable. Option D is wrong because logistic regression models the probability of a binary outcome as a function of predictors — it can be used, but it is not the most appropriate simple test for comparing two proportions; chi-square is the standard hypothesis test for this design.

35
MCQhard

A data scientist builds a logistic regression model to predict customer churn (yes/no). The model outputs a probability of 0.75 for a particular customer. Which of the following best describes this output?

A.The customer will definitely churn.
B.There is a 75% chance the customer will churn.
C.The odds of churning are 0.75 to 1.
D.The model is 75% accurate.
AnswerB

Logistic regression outputs a probability between 0 and 1 for the positive class. A value of 0.75 means the model estimates a 75% probability that this customer will churn, expressed as a likelihood rather than a certainty.

Why this answer

The output of a logistic regression model is a probability between 0 and 1. A value of 0.75 means the model estimates a 75% probability that the customer will churn (the positive class). This is a probabilistic prediction, not a certainty.

Exam trap

DA0-002 often tests the interpretation of logistic regression output, confusing probability with odds or accuracy. Candidates might think 0.75 means 75% accuracy or odds of 0.75 to 1.

How to eliminate wrong answers

Option A is wrong because the model outputs a probability, not a definitive prediction; 0.75 does not mean the customer will definitely churn. Option C is wrong because the output is a probability, not odds; odds would be p/(1-p), which here is 0.75/0.25 = 3 to 1. Option D is wrong because the output is not a measure of model accuracy; accuracy is a performance metric computed over many predictions, not a single output.

36
MCQmedium

A data analyst at a regional bank is examining monthly account balances for 5,000 customers. The distribution is strongly right-skewed because a small number of high-net-worth clients hold very large balances. The analyst must report a measure of central tendency that best represents the typical customer without being pulled upward by those few extreme accounts. Which measure should the analyst report?

A.The arithmetic mean of the monthly balances.
B.The standard deviation of the monthly balances.
C.The median of the monthly balances.
D.The range between the minimum and maximum monthly balances.
AnswerC

The median is the middle value when balances are ordered, so it depends only on position and is unaffected by how large the extreme accounts become. For a right-skewed distribution it sits closer to the bulk of customers and therefore better represents a typical balance. This is exactly the robust measure of central tendency the analyst needs when outliers would distort the mean.

Why this answer

With a strong right skew driven by a few very large balances, the mean is pulled upward while the median stays anchored near the middle of the ordered data. Because the request is for a representative central value that resists outliers, the median is the appropriate choice. Standard deviation and the range describe dispersion, not central tendency, so they cannot answer the question.

Exam trap

The trap here is defaulting to the mean because it is familiar, without recognizing that extreme values in a skewed distribution drag it away from the typical observation.

37
MCQmedium

A data analyst at a marketing firm is tasked with segmenting customers based on their purchasing behavior. The dataset contains 10,000 customers with features such as annual spend, frequency of purchases, recency of last purchase, and average order value. The analyst decides to use k-means clustering. After standardizing the features, the analyst runs k-means with k=3, k=4, and k=5, and computes the silhouette score for each: k=3: 0.45, k=4: 0.52, k=5: 0.48. The analyst also plots the elbow curve and observes that the within-cluster sum of squares (WCSS) decreases sharply from k=2 to k=4, then levels off. Based on these results, what is the most appropriate number of clusters?

A.k=4
B.k=2
C.k=3
D.k=5
AnswerA

k=4 yields the highest silhouette score (0.52), indicating the tightest, best-separated clusters, and the elbow curve's WCSS levelling off after k=4 corroborates this. Both diagnostics converge, satisfying the stem's requirement to select the most appropriate cluster count.

Why this answer

The silhouette score is highest at k=4 (0.52), indicating that clusters are well-separated and cohesive. The elbow curve shows WCSS decreasing sharply up to k=4 and then leveling off, suggesting that k=4 captures the optimal trade-off between model complexity and variance explained. Together, these metrics point to k=4 as the most appropriate number of clusters.

Exam trap

The trap here is that candidates might rely solely on the elbow curve and pick k=3 or k=5, ignoring the silhouette score which directly measures cluster quality and clearly favors k=4.

How to eliminate wrong answers

Option B (k=2) is wrong because the elbow curve shows a sharp decrease in WCSS from k=2 to k=4, meaning k=2 would underfit the data and miss meaningful segmentation. Option C (k=3) is wrong because its silhouette score (0.45) is lower than k=4 (0.52), indicating poorer cluster separation and cohesion. Option D (k=5) is wrong because its silhouette score (0.48) is lower than k=4, and the elbow curve shows WCSS leveling off after k=4, so adding a fifth cluster introduces unnecessary complexity without significant improvement.

38
MCQhard

A data analyst is building a time series forecasting model for monthly sales data. The analyst observes that the data has a clear upward trend and a seasonal pattern that repeats every 12 months. After differencing the data once to remove the trend, the analyst notices that the seasonal pattern remains. Which of the following should the analyst do next to properly model the data?

A.Include a moving average term of order 12.
B.Fit a non-seasonal ARIMA model with only autoregressive terms.
C.Apply a seasonal difference with lag 12.
D.Apply a second round of non-seasonal differencing.
AnswerC

A seasonal difference with lag 12 subtracts the value from the same month in the previous year, effectively removing the seasonal pattern. This is the standard approach when seasonality persists after non-seasonal differencing. It aligns the data for modeling with seasonal ARIMA (SARIMA) components, allowing the analyst to capture both trend and seasonality.

Why this answer

After removing the trend with non-seasonal differencing, the persistent seasonal pattern indicates the need for seasonal differencing. Taking a difference at lag 12 removes the yearly seasonal effect, making the series stationary in both trend and seasonality. This step is crucial before fitting a SARIMA model, which can then include seasonal autoregressive and moving average terms to capture the remaining structure.

Exam trap

The trap here is applying another non-seasonal difference or adding a moving average term instead of using seasonal differencing to address the recurring pattern.

39
MCQhard

A data analyst is building a predictive model to forecast customer churn. The dataset includes a categorical feature 'payment_method' with categories: credit card, debit card, PayPal, and bank transfer. The analyst decides to use one-hot encoding. After encoding, the analyst notices that the model's performance on the training set is excellent but poor on the test set. Which issue is most likely contributing to this problem?

A.The one-hot encoding caused data leakage from the test set into the training set.
B.The one-hot encoding failed to capture the ordinal relationship among payment methods.
C.The one-hot encoding introduced multicollinearity among the payment method dummy variables.
D.The model is overfitting due to high dimensionality from one-hot encoding combined with a small dataset.
AnswerD

One-hot encoding increases dimensionality, especially if the categorical feature has many categories. With a small dataset, this can lead to overfitting because the model may learn noise specific to the training set. The symptom of excellent training performance but poor test performance is classic overfitting. High dimensionality from encoding can exacerbate this, particularly if the model is complex. Thus, this is the most likely issue.

Why this answer

Overfitting is indicated by strong training performance and weak test performance. One-hot encoding expands the feature space, and with a small dataset, this can lead to a model that memorizes training noise. The other options do not directly explain overfitting: multicollinearity affects interpretation, data leakage would inflate both training and test scores, and payment methods are nominal so ordinality is irrelevant.

Thus, high dimensionality from encoding is the most plausible cause.

Exam trap

The trap here is assuming that any encoding issue causes overfitting, when in fact the primary driver is often the increased dimensionality relative to sample size.

40
MCQeasy

A data analyst is reviewing a dataset of customer transactions and notices that the 'transaction_amount' column contains several negative values. The analyst suspects these are refunds rather than errors. Which data validation technique should the analyst apply to confirm this?

A.Calculate the mean and standard deviation of transaction_amount
B.Remove rows with negative transaction amounts
C.Replace all negative values with zero
D.Cross-reference negative amounts with a separate refunds table
AnswerD

Cross-referencing with a refunds table directly verifies whether negative transaction amounts correspond to legitimate refunds. This validation technique compares data across sources to confirm accuracy and meaning. It is the most reliable way to distinguish refunds from data entry errors. By matching transaction IDs or timestamps, the analyst can confirm the negative values are intentional and correctly recorded.

Why this answer

Cross-referencing negative transaction amounts with a separate refunds table is a validation technique that confirms whether the negative values are legitimate refunds. It compares the dataset against an authoritative source to verify accuracy. This approach preserves data integrity and ensures that any subsequent analysis correctly accounts for refunds.

Other options either summarize, alter, or delete data without confirming the values' meaning.

Exam trap

The trap here is confusing data cleaning with data validation; removing or replacing negative values might seem like a quick fix, but validation requires confirming the values' legitimacy first.

41
MCQeasy

A data analyst is working with a dataset containing customer ages. The ages range from 18 to 90, but the analyst notices that the distribution is heavily skewed to the right. To better understand the central tendency, the analyst decides to calculate a measure that is resistant to outliers. Which measure of central tendency should the analyst use?

A.Median
B.Mean
C.Mode
D.Range
AnswerA

The median is the middle value when data is ordered, making it resistant to outliers and skewness. In a right-skewed distribution, the median remains a robust measure of central tendency, better representing the typical customer age. It is not influenced by extreme values, so it provides a more accurate reflection of the center for skewed data.

Why this answer

The median is the appropriate measure because it is resistant to outliers and skewness, providing a better representation of the central value in a right-skewed distribution. The mean would be inflated by high ages, while the mode and range do not measure central tendency effectively for continuous skewed data.

Exam trap

The trap here is assuming the mean is always the best measure of center; in skewed distributions, the median is more representative.

42
Multi-Selecthard

A data analyst at a healthcare organization is analyzing patient readmission rates. The analyst has a dataset with patient demographics, diagnosis codes, and length of stay. Before performing any statistical modeling, the analyst must address data quality issues. Which TWO of the following actions are most appropriate for ensuring the dataset is ready for analysis? (Choose two.)

Select 2 answers
A.Remove all records with any missing values to ensure a complete dataset.
B.Identify and remove duplicate patient records based on a unique patient identifier and admission date.
C.Convert all numerical columns to categorical bins to simplify the analysis.
D.Impute missing values in the length of stay column using the mean length of stay for all patients.
E.Standardize diagnosis codes to a consistent format (e.g., ICD-10) and validate against a reference list.
AnswersB, E

Duplicate records can skew readmission rates and other analyses. Using a unique patient identifier combined with admission date helps distinguish between multiple admissions for the same patient versus true duplicates. Removing duplicates ensures each admission is counted once, which is critical for accurate readmission metrics. This is a standard data cleaning step and directly addresses a common data quality issue in healthcare datasets.

Why this answer

Ensuring data quality involves identifying and correcting issues that could bias results. Removing duplicate patient records prevents overcounting admissions, and standardizing diagnosis codes ensures consistency for analysis. Mean imputation can distort data, binning numerical variables is not a quality fix, and listwise deletion can introduce bias and reduce sample size.

These two actions directly address common data quality problems in healthcare datasets.

Exam trap

The trap here is assuming that any missing data must be removed or imputed with a simple mean, when in fact proper handling depends on the missingness mechanism and the analysis goal.

43
MCQeasy

A dataset contains the ages of 100 customers. The analyst wants to transform the ages to a 0-1 range for use in a distance-based algorithm. Which technique should be used?

A.Square root transformation
B.Log transformation
C.Z-score normalization
D.Min-max normalization
AnswerD

Min-max normalization rescales each value using (x − min) / (max − min), mapping the age range linearly onto 0–1. This satisfies the stem's requirement for a bounded 0–1 range, which distance-based algorithms such as k-means or k-NN need so no single feature dominates the distance calculation.

Why this answer

Min-max normalization transforms features to a fixed range, typically 0 to 1, by subtracting the minimum and dividing by the range. This is exactly what is needed to scale ages to a 0-1 range for distance-based algorithms.

Exam trap

The trap is confusing normalization with standardization. Candidates might choose Z-score normalization because it's commonly used, but it does not produce a 0-1 range. The question explicitly asks for 0-1 range, so min-max is correct.

How to eliminate wrong answers

Option A is wrong because square root transformation is used to reduce skewness, not to scale to a specific range. Option B is wrong because log transformation is used to handle skewed data or exponential relationships, not for range scaling. Option C is wrong because Z-score normalization standardizes data to have mean 0 and standard deviation 1, which does not guarantee a 0-1 range.

44
MCQeasy

A data analyst is cleaning a dataset and finds missing values in a categorical variable representing customer region. Which imputation method is most appropriate?

A.Drop rows with missing values
B.Mode imputation
C.Mean imputation
D.Median imputation
AnswerB

Mode imputation fills missing categorical entries with the variable's most frequent region, preserving the existing distribution and requiring no numeric assumptions. It suits categorical data, where mean or median imputation is meaningless, satisfying the requirement to handle missing region values sensibly.

Why this answer

Mode imputation is the most appropriate method for a categorical variable because it replaces missing values with the most frequently occurring category, preserving the distribution of the data. Unlike mean or median imputation, which are designed for numerical data, mode imputation maintains the categorical nature of the variable and avoids introducing invalid values. This approach is simple and effective when missing data is random and the category is well-represented.

Exam trap

The trap here is that candidates often confuse imputation methods across data types, incorrectly applying mean or median imputation to categorical variables because they focus on central tendency without considering data type appropriateness.

How to eliminate wrong answers

Option A is wrong because dropping rows with missing values can lead to significant data loss and potential bias, especially if the missingness is not completely random, reducing the dataset's representativeness. Option C is wrong because mean imputation is only appropriate for numerical data, not categorical variables, as calculating the mean of categories is meaningless and would produce non-categorical values. Option D is wrong because median imputation is also designed for numerical data and cannot be applied to categorical variables, as the median requires ordered numerical values to compute.

45
Multi-Selecthard

An analyst is conducting an A/B test on a new website layout. Which TWO of the following must be defined before the test begins?

Select 2 answers
A.The final conversion rates for each group
B.The actual p-value from the test
C.The confidence interval for the lift
D.The significance threshold (alpha)
E.The sample size required for adequate statistical power
AnswersD, E

Fixing alpha before the test begins controls the false-positive rate and prevents p-hacking after results arrive. The stem demands pre-defined parameters; without a declared significance threshold, the analyst cannot objectively judge whether the observed difference between layouts is statistically meaningful.

Why this answer

Option D is correct because the significance threshold (alpha), typically set at 0.05, must be fixed in advance to determine the probability of a Type I error and to provide the criterion against which the resulting p-value will be judged. Option E is correct because the required sample size must be calculated beforehand, based on the desired statistical power (commonly 0.80), the baseline conversion rate, and the minimum detectable effect, to ensure the test can reliably detect a real difference. Options A, B, and C are not pre-test definitions: the final conversion rates and the actual p-value are outcomes observed only after the test concludes, and the confidence interval for the lift is likewise a post-hoc computed estimate rather than an input specified before the test begins.

Exam trap

DA0-002 often tests the confusion between pre-test design parameters (alpha, sample size, MDE) and post-test outputs (p-value, confidence interval, observed conversion rates), so candidates must recognize which values can only exist after data collection.

46
MCQmedium

A marketing team runs an A/B test comparing two webpage designs. The null hypothesis states there is no difference in conversion rates. The p-value is 0.08 at α=0.05. Which is the correct interpretation?

A.The null hypothesis is rejected, indicating the designs are different.
B.The alternative hypothesis is accepted, showing the new design is better.
C.There is insufficient evidence to conclude a difference between the designs.
D.There is a statistically significant difference between the designs.
AnswerC

At α=0.05, a p-value of 0.08 exceeds the significance threshold, so the null hypothesis is not rejected. The test simply lacks statistical evidence of a conversion-rate difference; it does not prove the designs perform identically.

Why this answer

A p-value of 0.08 exceeds the significance level α=0.05, so we fail to reject the null hypothesis. Failing to reject the null means there is not enough statistical evidence to conclude that a real difference exists between the two webpage designs. The result is not statistically significant at the 5% threshold.

Exam trap

The trap here is confusing 'fail to reject the null' with 'the null is true' or 'the alternative is accepted' — candidates often pick A or B by treating a non-significant p-value as proof of no effect or as evidence for the alternative.

How to eliminate wrong answers

Option A is wrong because rejecting the null requires p ≤ α (0.08 > 0.05), so rejection is not warranted. Option B is wrong because we never 'accept' the alternative hypothesis — we only fail to reject the null; also, the test does not establish that the new design is better. Option D is wrong because a p-value above α indicates the difference is not statistically significant, so claiming significance misreads the threshold.

47
MCQeasy

Which measure best describes the spread of the middle 50% of a dataset?

A.IQR
B.Range
C.Standard deviation
D.Variance
AnswerA

The interquartile range is the difference between the third and first quartiles (Q3 minus Q1), so it spans exactly the middle 50% of ordered observations. Unlike range or standard deviation, it ignores extreme values, describing central spread precisely.

Why this answer

Interquartile range (IQR) is the range between Q1 and Q3, covering the middle 50%.

48
MCQmedium

An analyst uses K-means clustering on customer purchase data. After plotting the within-cluster sum of squares for different values of k, they observe an elbow at k=4. What is the most appropriate number of clusters?

A.4
B.6
C.5
D.3
AnswerA

An elbow at k=4 indicates that adding further clusters yields diminishing reductions in within-cluster sum of squares, so four captures the data's structure without overfitting. Selecting four directly satisfies the stem's elbow-method criterion, giving the most appropriate cluster count for the purchase data.

Why this answer

The elbow method plots within-cluster sum of squares (WCSS) against k and looks for the 'elbow' — the point where adding more clusters yields diminishing reductions in WCSS. An elbow at k=4 means the marginal gain from k=4 to k=5 is small, so k=4 is the recommended number of clusters. Choosing k=4 balances model simplicity with fit quality.

Exam trap

DA0-002 often tests the misconception that the elbow value itself is the answer versus confusing it with the 'optimal' k from silhouette analysis — candidates sometimes pick a value adjacent to the elbow or assume more clusters is always better.

How to eliminate wrong answers

Option B (6) is wrong because it is past the elbow — adding clusters beyond the elbow yields minimal WCSS reduction and risks overfitting. Option C (5) is wrong because the elbow is at 4, not 5; the analyst would be adding an unnecessary cluster. Option D (3) is wrong because it is before the elbow — WCSS would still drop substantially by moving to 4, indicating underfitting.

49
MCQmedium

A data analyst is analyzing a time series of monthly sales for a retail store. The analyst observes a repeating pattern every 12 months. Which component of the time series is the analyst observing?

A.Noise
B.Trend
C.Cyclicality
D.Seasonality
AnswerD

Seasonality is a repeating pattern that occurs at fixed intervals, such as monthly, quarterly, or yearly. The 12-month repeating pattern in monthly sales data is a classic example of seasonality, often driven by holidays, weather, or school schedules. This component is predictable and consistent in timing. The analyst is observing seasonality because the pattern repeats every 12 months.

Why this answer

Seasonality is the time series component that captures repeating patterns at fixed intervals, such as every 12 months in monthly data. The analyst's observation of a yearly repeating pattern in sales is a direct example of seasonality. Trend, cyclicality, and noise do not describe fixed-period repetitions.

Therefore, seasonality is the correct component.

Exam trap

The trap here is confusing seasonality with cyclicality; both are repeating patterns, but seasonality has a fixed period, while cyclicality is irregular and often longer term.

50
MCQmedium

A dataset has missing values in the 'age' column. The distribution of age is approximately normal with few outliers. Which imputation method is most appropriate?

A.Mean imputation
B.Forward-fill
C.Delete all rows with missing data
D.Mode imputation
AnswerA

Age is approximately normal with few outliers, so the mean is a stable, representative estimate and preserves the sample mean. Mean imputation satisfies this distributional constraint better than median or mode, which suit skewed data or categorical fields respectively.

Why this answer

Mean imputation is appropriate when the data is approximately normally distributed and has few outliers, because the mean is a representative measure of central tendency for symmetric distributions. It preserves the overall mean of the variable and is simple to implement. Forward-fill and mode imputation are better for time-series or categorical data, respectively, and deleting rows can introduce bias and reduce sample size.

Exam trap

DA0-002 often tests the assumption that mean imputation is always appropriate, but the trap is failing to recognize that it is only suitable for continuous, normally distributed data without outliers; mode is for categorical, forward-fill for time-series.

How to eliminate wrong answers

Option B is wrong because forward-fill is used for time-series data where the next value is likely similar to the previous one, not for independent observations like age. Option C is wrong because deleting all rows with missing data can lead to significant data loss and bias if the missingness is not completely random. Option D is wrong because mode imputation is used for categorical variables, not continuous variables like age, and it can distort the distribution.

51
MCQeasy

A data analyst is comparing the average test scores of students who attended a tutoring program versus those who did not. Which statistical test is most appropriate for determining if there is a significant difference between the means of these two independent groups?

A.Paired t-test
B.Chi-square test
C.Two-sample t-test
D.ANOVA
AnswerC

The two-sample t-test compares means drawn from two independent groups, which matches the tutoring versus non-tutoring cohorts. It partitions variance between and within groups to produce a p-value for the null hypothesis of equal population means, satisfying the requirement to test significance between two unrelated samples.

Why this answer

The independent samples t-test is used to compare the means of two independent groups.

52
Matchingmedium

Match each data sampling method to its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Each member has equal chance of selection

Population divided into subgroups; random sample from each

Randomly select entire groups (clusters)

Select every k-th element from a list

Sample based on ease of access

Why these pairings

In matching, ensure you distinguish between methods: simple random gives equal chance; stratified divides by characteristic; cluster divides into groups and picks whole groups; systematic uses a fixed interval.

53
MCQhard

In a linear regression model predicting house prices, the coefficient for the number of bedrooms is $30,000 and the intercept is $50,000. If a house has 3 bedrooms, what is the predicted price?

A.$80,000
B.$150,000
C.$90,000
D.$140,000
AnswerD

Applying the regression equation gives $50,000 + ($30,000 × 3) = $140,000. The intercept supplies the baseline price, and the bedroom coefficient scales linearly with the count, so three bedrooms add $90,000 to the base, satisfying the stem's prediction request.

Why this answer

The linear regression equation is Price = Intercept + (Coefficient × Bedrooms) = $50,000 + ($30,000 × 3) = $50,000 + $90,000 = $140,000. This is a direct application of the slope-intercept form of a simple linear regression model with one predictor.

Exam trap

DA0-002 often tests whether candidates correctly apply the regression equation, and the trap is forgetting to add the intercept or misreading the coefficient as the total price rather than the per-unit change.

How to eliminate wrong answers

Option A is wrong because $80,000 results from adding only $30,000 to the intercept, which would correspond to one bedroom, not three. Option B is wrong because $150,000 results from multiplying $50,000 by 3, incorrectly treating the intercept as a per-bedroom value. Option C is wrong because $90,000 is just the coefficient times bedrooms ($30,000 × 3) without adding the intercept, omitting the base price.

54
MCQmedium

In time series analysis, which component represents regular patterns that repeat over fixed periods, such as daily or yearly?

A.Seasonality
B.Trend
C.Cyclical
D.Irregular
AnswerA

Seasonality describes fluctuations that recur at fixed intervals, such as daily, weekly or yearly cycles, and is modelled separately from trend and residual noise. It is the component capturing the regular repeating pattern the stem describes, not the long-term direction or random variation.

Why this answer

Seasonality refers to patterns that repeat at regular intervals.

55
MCQeasy

After a marketing campaign, sales increased by 15%. The analyst wants to understand which customer segment contributed most to the increase. Which type of analysis is this?

A.Predictive analysis
B.Diagnostic analysis
C.Prescriptive analysis
D.Descriptive analysis
AnswerB

Diagnostic analysis examines why an outcome occurred by drilling into contributing factors. Here it isolates which customer segment drove the 15% sales lift, satisfying the requirement to explain the increase's source rather than merely describing it or forecasting future values.

Why this answer

Diagnostic analysis is used to understand the root cause of an event or change. In this scenario, the analyst already knows sales increased by 15% and wants to determine which customer segment drove that increase, which is a classic diagnostic question. This type of analysis goes beyond describing what happened to explain why it happened.

Exam trap

The trap here is confusing diagnostic analysis with descriptive analysis, as both deal with past data, but descriptive only summarizes what happened while diagnostic explains why it happened.

How to eliminate wrong answers

Option A is wrong because predictive analysis uses historical data to forecast future outcomes, not to explain past changes. Option C is wrong because prescriptive analysis recommends actions or decisions to achieve a desired outcome, not to diagnose the cause of a past event. Option D is wrong because descriptive analysis summarizes what happened (e.g., 'sales increased by 15%') but does not investigate which segment contributed most to the increase.

56
MCQmedium

A data analyst is working with a dataset that contains a column 'income' with a highly skewed distribution. The analyst wants to apply a transformation to make the distribution more symmetric for use in a linear regression model. Which transformation is most appropriate?

A.Min-max normalization
B.Binning into equal-width intervals
C.Standardization (z-score normalization)
D.Logarithmic transformation
AnswerD

A logarithmic transformation is commonly used to reduce right skewness in positive data. Income data often has a long right tail, and taking the log compresses larger values more than smaller ones, making the distribution more symmetric. This can improve the performance of linear regression by making the relationship more linear and stabilizing variance. Thus, it is the most appropriate choice for this scenario.

Why this answer

A logarithmic transformation is effective for reducing right skewness in positive data like income. It compresses the upper tail, making the distribution more symmetric and improving linearity for regression. Standardization and min-max normalization only rescale without changing shape.

Binning discards information and does not achieve symmetry. Therefore, the logarithmic transformation is the correct choice.

Exam trap

The trap here is confusing scaling transformations (like standardization) with shape-changing transformations (like log). Scaling does not fix skewness.

57
MCQhard

A data scientist is building a K-means clustering model for customer segmentation. After plotting the within-cluster sum of squares (WCSS) against the number of clusters (k), she observes that the WCSS decreases sharply until k=5 and then levels off. Which value of k should she choose based on the elbow method?

A.k=5
B.k=6
C.k=4
D.k=3
AnswerA

The elbow method selects the k where WCSS reduction transitions from steep to shallow, forming the bend. At k=5 the curve flattens, meaning additional clusters yield diminishing compactness gains. Choosing k=5 satisfies the stem's observed inflection point, balancing model simplicity against within-cluster variance.

Why this answer

The elbow method plots WCSS against k and looks for the 'elbow' — the point where the rate of decrease sharply changes from steep to shallow. Since WCSS drops sharply until k=5 and then levels off, k=5 is the optimal choice because additional clusters beyond that yield diminishing returns.

Exam trap

DA0-002 often tests the misconception that the lowest WCSS (highest k) is best; candidates must recognize that the elbow method selects the inflection point, not the minimum WCSS.

How to eliminate wrong answers

Option B (k=6) is wrong because it is past the elbow, adding a cluster with minimal WCSS reduction and risking overfitting. Option C (k=4) is wrong because it is before the elbow, meaning the model is under-clustered and WCSS is still decreasing steeply. Option D (k=3) is wrong for the same reason — it is even further from the elbow and would leave significant within-cluster variance unexplained.

58
Multi-Selectmedium

A data analyst is preparing a dataset for a machine learning algorithm that assumes normally distributed features. Which TWO data transformation methods should the analyst consider to achieve this?

Select 2 answers
A.Square root transformation
B.Log transformation
C.One-hot encoding
D.Z-score standardization
E.Min-max normalization
AnswersA, B

The square root transformation compresses right-skewed data, reducing the influence of large values and pulling the distribution toward normality. It satisfies the algorithm's normality assumption for moderately skewed features, particularly count data. Unlike logarithms, it handles zero values and is milder, making it suitable when skew is moderate rather than severe.

Why this answer

The square root transformation (A) is correct because it is a variance-stabilizing transformation that compresses right-skewed data and can make moderately skewed distributions more symmetric and closer to normal, which suits algorithms assuming normally distributed features. The log transformation (B) is also correct because it strongly reduces right skew and pulls in large outliers, making positively skewed data (e.g., exponential or multiplicative data) approximate a normal distribution. One-hot encoding (C) is not appropriate here because it converts categorical variables into binary indicator columns and does not change the distribution shape of numeric features toward normality.

Z-score standardization (D) only rescales features to mean 0 and standard deviation 1 without altering skewness or the underlying distribution shape. Min-max normalization (E) merely rescales values to a fixed range such as [0,1] and likewise does not make a non-normal distribution normal.

Exam trap

The trap here is confusing transformations that change distribution shape (e.g., square root, log) with those that only rescale (e.g., standardization, normalization), leading candidates to select scaling methods when the goal is to achieve normality.

59
MCQmedium

A dataset contains employee salaries ranging from $30,000 to $200,000. An analyst wants to scale the salaries to a range of 0 to 1 for use in a distance-based clustering algorithm. Which method should they use?

A.Log transformation
B.Robust scaling
C.Min-max normalization
D.Z-score standardization
AnswerC

Min-max normalization rescales each value using (x − min)/(max − min), mapping the $30,000–$200,000 salary range linearly onto 0–1. This preserves relative distances, which distance-based clustering requires, unlike z-score standardisation, which centres on the mean with unbounded output.

Why this answer

Min-max normalization rescales values linearly to a fixed range, typically [0, 1], using the formula (x − min) / (max − min). This is exactly what the analyst needs for a distance-based clustering algorithm, where features on different scales would otherwise dominate the distance metric. It preserves the relative ordering and shape of the distribution while bounding all values between 0 and 1.

Exam trap

DA0-002 often tests normalization vs. standardization — candidates pick z-score because it is commonly used, missing that the question explicitly requires a 0-to-1 bounded range that only min-max normalization provides.

How to eliminate wrong answers

Option A is wrong because a log transformation compresses skewed data but does not bound values to [0, 1], so it fails the stated range requirement. Option B is wrong because robust scaling uses the median and IQR, centering data around zero with no fixed upper bound — it handles outliers but does not produce a 0–1 range. Option D is wrong because z-score standardization produces a mean of 0 and standard deviation of 1, yielding values that can be negative or exceed 1, not a bounded [0, 1] range.

60
MCQeasy

In an A/B test, the null hypothesis states that there is no difference between the conversion rates of the control and treatment groups. After collecting data, the p-value is 0.03. Using a significance level α = 0.05, what should the analyst conclude?

A.Reject the null hypothesis; there is a significant difference
B.Accept the alternative hypothesis; the treatment is better
C.The test is inconclusive
D.Fail to reject the null hypothesis; no significant difference
AnswerA

The p-value of 0.03 falls below the 0.05 significance level, so the observed difference is unlikely under the null hypothesis. Rejecting the null is therefore the statistically correct conclusion, indicating a significant difference between control and treatment conversion rates.

Why this answer

Since the p-value (0.03) is less than α (0.05), the null hypothesis is rejected, indicating a statistically significant difference between the groups.

61
MCQeasy

A data analyst is designing a data model for a sales data warehouse. The model should optimize query performance for aggregations by minimizing joins and duplicating data where necessary. Which schema design should the analyst use?

A.Entity-relationship model
B.Snowflake schema
C.3NF normalized model
D.Star schema
AnswerD

A star schema places a central fact table surrounded by denormalised dimension tables, so aggregation queries join fewer tables and scan pre-joined data. This deliberately duplicates attributes to minimise joins, matching the stated performance requirement.

Why this answer

A star schema is the correct choice because it organizes data into a central fact table surrounded by denormalized dimension tables, which minimizes the number of joins required for aggregation queries. By duplicating dimension attributes rather than normalizing them, the star schema trades storage for query speed, making it ideal for data warehouse workloads that emphasize analytical performance. This design directly supports the requirement to optimize aggregations while reducing join complexity.

Exam trap

DA0-002 often tests the misconception that normalization always improves performance, but in data warehousing, denormalization via star schema is preferred for analytical query speed.

How to eliminate wrong answers

Option A is wrong because an entity-relationship model is a conceptual modeling technique used for OLTP systems, not a physical schema optimized for analytical query performance. Option B is wrong because a snowflake schema normalizes dimension tables into multiple related tables, which increases the number of joins and slows aggregation queries. Option C is wrong because a 3NF normalized model eliminates redundancy and is designed for transactional integrity, not for minimizing joins in analytical queries.

62
MCQhard

A data analyst uses linear regression to model the relationship between advertising spend and sales. The residual plot shows a clear U-shaped pattern. What assumption is violated?

A.Independence of residuals
B.Homoscedasticity
C.Normality of residuals
D.Linearity
AnswerD

A U-shaped residual pattern means the model systematically under- and over-predicts across the predictor range, indicating the true relationship is curved rather than straight. The linearity assumption, that predictors relate to the outcome additively in a straight line, is therefore violated.

Why this answer

The U-shaped pattern in the residual plot indicates that the relationship between advertising spend and sales is not linear; the model fails to capture the curvature in the data. Linear regression assumes a straight-line relationship between predictors and the response, so a systematic pattern like a U-shape directly violates the linearity assumption. This means the model is misspecified and requires a transformation or a nonlinear modeling approach.

Exam trap

CompTIA often tests the distinction between residual pattern shapes and their corresponding assumptions, so the trap here is that candidates confuse a curved pattern (nonlinearity) with heteroscedasticity or non-normality, leading them to pick B or C instead of D.

How to eliminate wrong answers

Option A is wrong because independence of residuals refers to errors being uncorrelated with each other, often violated in time-series data, but a U-shaped pattern does not imply autocorrelation. Option B is wrong because homoscedasticity means constant variance of residuals across fitted values, which would appear as a funnel or cone shape, not a U-shaped curve. Option C is wrong because normality of residuals concerns the distribution of errors (checked via Q-Q plot or histogram), not the pattern of residuals versus fitted values; a U-shaped pattern does not directly indicate non-normality.

63
MCQhard

A healthcare analytics team is building a predictive model to identify patients at high risk of readmission within 30 days of discharge. The dataset includes 50,000 patient records with 200 features, including demographics, vital signs, lab results, and historical admissions. The target variable is binary (readmitted or not). The team uses a logistic regression model and achieves an AUC of 0.72 on the test set. However, the model's calibration is poor: for patients predicted to have a 70% risk, the actual readmission rate is only 40%. The team wants to improve calibration without significantly reducing discrimination (AUC). The data scientist suggests applying Platt scaling. However, the team lead is concerned that Platt scaling may reduce the model's ability to rank patients correctly. Which of the following is the best course of action?

A.Remove poorly calibrated predictions by discarding all patients with predicted risk between 0.3 and 0.7.
B.Ignore calibration because AUC is the only metric that matters for readmission risk models.
C.Apply Platt scaling on a held-out validation set to recalibrate the predicted probabilities without refitting the original model.
D.Switch to a random forest model, which inherently produces better-calibrated probabilities.
AnswerC

Platt scaling fits a logistic regression on the model's raw scores using a held-out validation set, correcting probability estimates while leaving the underlying model and its ranking untouched. Because it is monotonic, discrimination and AUC are preserved, satisfying the calibration goal without refitting.

Why this answer

Platt scaling is a post-processing technique that fits a logistic regression model on the predicted probabilities from the original model using a held-out validation set. This recalibrates the probabilities without altering the ranking of patients (the AUC remains unchanged), directly addressing the poor calibration while preserving discrimination. Option C correctly describes this procedure.

Exam trap

The trap here is that candidates may think Platt scaling changes the model's ranking (AUC), but in reality it applies a monotonic transformation that preserves rank order, so discrimination is unaffected.

How to eliminate wrong answers

Option A is wrong because discarding patients with predicted risk between 0.3 and 0.7 removes a large portion of the data and does not fix the underlying miscalibration; it merely hides the problem and reduces the model's utility. Option B is wrong because AUC measures only rank ordering, not probability accuracy; for clinical risk models, well-calibrated probabilities are critical for decision-making (e.g., resource allocation). Option D is wrong because random forest models are known to produce poorly calibrated probabilities due to their averaging of decision tree outputs, often requiring their own calibration (e.g., isotonic regression) and do not inherently guarantee better calibration than logistic regression.

64
MCQeasy

A data analyst is examining a dataset of customer transactions and notices that the 'transaction_amount' column contains negative values. The analyst determines that these negative values represent refunds. Which data quality dimension is most directly relevant to this finding?

A.Consistency
B.Accuracy
C.Validity
D.Completeness
AnswerC

Validity ensures data conforms to defined business rules or constraints. If the business rule expects transaction amounts to be positive (e.g., sales), then negative values violate that rule. However, if refunds are allowed, the rule may need to accommodate negatives. The analyst's realization that negatives represent refunds highlights a validity check: are these values valid according to the expected domain? This makes validity the most relevant dimension.

Why this answer

Validity is about whether data adheres to defined rules or constraints. Negative transaction amounts may be valid if refunds are permitted, but the analyst must verify that such values are allowed by business rules. Accuracy, completeness, and consistency do not directly address the conformity to rules.

Thus, validity is the most relevant dimension when encountering unexpected negative values that might indicate either a data error or a legitimate business process.

Exam trap

The trap here is confusing validity with accuracy. Negative values can be accurate (they correctly represent refunds) but may still violate a validity rule if the system expects only positive amounts.

65
MCQmedium

A data scientist is performing a hypothesis test with a significance level α=0.05. The p-value obtained is 0.03. What should the scientist conclude?

A.Reject the null hypothesis because the p-value is less than the significance level.
B.Fail to reject the null hypothesis because the p-value is greater than 0.01.
C.The test is inconclusive, need a larger sample size.
D.Accept the null hypothesis because the p-value is small.
AnswerA

With α=0.05, a p-value of 0.03 falls inside the rejection region, so the null hypothesis is rejected. This satisfies the stem's stated significance level and p-value, indicating the observed result is statistically significant at that threshold.

Why this answer

The decision rule for hypothesis testing is: if p-value < α, reject the null hypothesis. Here p = 0.03 and α = 0.05, so 0.03 < 0.05, meaning the result is statistically significant and the null hypothesis should be rejected in favor of the alternative. This indicates the observed effect is unlikely to have occurred by chance alone at the 5% significance level.

Exam trap

DA0-002 often tests the misconception that a small p-value means 'accept the null' or that p-values should be compared to a value other than the stated α — candidates confuse rejection logic or misread the threshold.

How to eliminate wrong answers

Option B is wrong because it compares the p-value to 0.01, which is not the stated significance level — the analyst set α=0.05, and the correct comparison is 0.03 < 0.05, leading to rejection, not failure to reject. Option C is wrong because the test is not inconclusive — the p-value clearly falls below α, so a definitive decision can be made without a larger sample. Option D is wrong because the logic is inverted: a small p-value leads to rejecting the null hypothesis, not accepting it; also, hypothesis tests never 'accept' the null, they only fail to reject it.

66
MCQmedium

A healthcare analytics team is building a classification model to predict patient readmission within 30 days. The dataset contains 10,000 records with 30 features, including demographics, vital signs, lab results, and medication history. The target variable is imbalanced: 85% no readmission, 15% readmission. The team used logistic regression with default settings and achieved an accuracy of 85%, but the model predicted 'no readmission' for all patients. The lead analyst suspects the model is not learning due to class imbalance. The team has time to implement one corrective action before the next model review. Which action should the team take?

A.Remove features with low variance to reduce noise
B.Apply SMOTE to oversample the readmission class
C.Use accuracy as the evaluation metric to monitor improvement
D.Switch to a random forest model with default settings
AnswerB

SMOTE synthesises minority-class readmission examples, rebalancing the 85/15 split so logistic regression no longer converges on the majority class. This directly counters the imbalance causing the all-negative predictions, addressing the stated constraint within one corrective action.

Why this answer

SMOTE (Synthetic Minority Oversampling Technique) directly addresses the class imbalance by generating synthetic samples for the minority class (readmission). This forces the logistic regression model to learn decision boundaries that separate the two classes, rather than defaulting to the majority class prediction. With 85% majority and 15% minority, accuracy alone is misleading, and SMOTE is a proven technique to improve recall for the minority class.

Exam trap

The trap here is that candidates often choose accuracy as a metric (Option C) because it seems intuitive, but in imbalanced datasets, accuracy is misleading and does not reflect model performance for the minority class.

How to eliminate wrong answers

Option A is wrong because removing low-variance features does not address class imbalance; it only reduces noise or redundant features, but the model will still predict the majority class if the imbalance is not handled. Option C is wrong because using accuracy as the evaluation metric is exactly the problem—it will remain high (85%) even if the model predicts all 'no readmission', so it does not monitor improvement for the minority class. Option D is wrong because switching to a random forest model with default settings does not inherently solve class imbalance; random forest can also be biased toward the majority class without techniques like class weighting or resampling.

67
Multi-Selectmedium

A data analyst is preparing a dataset for a machine learning model and needs to handle missing values in several columns. The analyst wants to choose appropriate imputation methods. Which TWO of the following are valid considerations when selecting an imputation technique? (Choose two.)

Select 2 answers
A.Imputation should always use the mean for numerical variables to preserve the distribution.
B.The proportion of missing data in a column affects the reliability of imputation.
C.The missing data mechanism (MCAR, MAR, MNAR) influences the choice of imputation method.
D.Imputation should be performed before splitting data into training and test sets to avoid data leakage.
E.The choice of imputation method should be based solely on computational efficiency.
AnswersB, C

A high proportion of missing values (e.g., >50%) can make imputation unreliable and may warrant dropping the column or using advanced techniques. The amount of missingness impacts the confidence in imputed values and the potential for bias. Thus, it is a valid consideration when selecting an imputation method.

Why this answer

The missing data mechanism determines whether imputation can be unbiased, and the proportion of missing data affects reliability. These are key statistical considerations. Mean imputation is not always appropriate, imputation should occur after train-test split to avoid leakage, and efficiency alone is insufficient.

Therefore, the mechanism and proportion are the valid considerations.

Exam trap

The trap here is assuming that mean imputation is always safe and that imputation can be done before splitting; both can lead to biased models.

68
Multi-Selectmedium

Which TWO of the following are components of time series data?

Select 2 answers
A.Mean
B.Variance
C.Trend
D.Seasonality
E.Median
AnswersC, D

Trend is the long-term directional movement of a series over time, rising, falling or flat, ignoring short-term fluctuations. It is a core time series component alongside seasonality, cyclical variation and irregular residuals, so it satisfies the question's requirement.

Why this answer

Trend (C) is a core component of time series data because it represents the long-term, systematic increase or decrease in the series level over time, which is exactly what decomposition methods (e.g., additive or multiplicative decomposition) isolate. Seasonality (D) is also a core component because it captures regular, calendar-linked repeating patterns (e.g., daily, weekly, monthly, or quarterly cycles) that recur with a fixed period. By contrast, Mean (A), Variance (B), and Median (E) are descriptive statistics or summary measures of a distribution, not structural components of a time series; they describe the data's central tendency or spread rather than the temporal dynamics that decomposition separates.

Exam trap

DA0-002 often tests whether candidates confuse statistical summary measures (mean, median, variance) with the structural components of time series (trend, seasonality, cyclicality, irregular), so candidates must recognize that the question asks for components, not statistics.

69
MCQmedium

A data analyst wants to understand the relationship between advertising spend and sales revenue. The analyst calculates a Pearson correlation coefficient of 0.85. Which of the following is the best interpretation?

A.There is a strong positive linear relationship between advertising spend and sales.
B.85% of the variation in sales is explained by advertising spend.
C.Increasing advertising spend by $1 will increase sales by $0.85.
D.There is a strong negative linear relationship between advertising spend and sales.
AnswerA

A coefficient of 0.85 sits near the top of the -1 to +1 range, and its positive sign confirms that as advertising spend rises, sales revenue tends to rise too. The magnitude indicates a strong, near-linear association, satisfying the stem's request to interpret the calculated value.

Why this answer

A Pearson correlation coefficient of 0.85 indicates a strong positive linear relationship between the two variables. The value is close to +1, meaning as advertising spend increases, sales revenue tends to increase in a linear fashion. Correlation measures the strength and direction of a linear association, not causation or predictive proportion.

Exam trap

The trap here is confusing the correlation coefficient r with the coefficient of determination r², causing candidates to select the '85% of variation' answer.

How to eliminate wrong answers

Option B is wrong because 85% of variation explained would require squaring the correlation (r² = 0.7225, or ~72%), and even that describes coefficient of determination, not the raw r value. Option C is wrong because correlation does not imply a slope of 0.85; that would require regression coefficients, and correlation is unitless. Option D is wrong because a positive value of 0.85 indicates a positive relationship, not a negative one.

70
MCQhard

A company is analyzing customer feedback sentiment. The dataset is highly imbalanced with 95% positive and 5% negative comments. Which technique should the analyst use to address class imbalance before modeling?

A.Use accuracy as the evaluation metric
B.Undersample the majority class
C.Oversample the majority class
D.Use SMOTE
AnswerD

SMOTE generates synthetic minority-class samples by interpolating between existing nearest neighbours, rebalancing the 95:5 split before training. This satisfies the stem's requirement to address class imbalance, letting the model learn negative-class patterns instead of defaulting to the majority class.

Why this answer

SMOTE (Synthetic Minority Oversampling Technique) is the correct choice because it generates synthetic samples for the minority class (negative comments) by interpolating between existing minority instances, rather than simply duplicating them. This addresses the 95:5 imbalance without the information loss of undersampling or the overfitting risk of naive oversampling.

Exam trap

The trap here is that candidates often confuse oversampling the minority class with oversampling the majority class, or they incorrectly assume that simply using a different evaluation metric (like accuracy) can fix the imbalance problem without modifying the dataset.

How to eliminate wrong answers

Option A is wrong because accuracy is a misleading metric for imbalanced datasets; a model predicting all comments as positive would achieve 95% accuracy but fail to identify any negative comments. Option B is wrong because undersampling the majority class discards a large amount of potentially useful data, which can lead to loss of important patterns and reduced model performance. Option C is wrong because oversampling the majority class would exacerbate the imbalance, making the model even more biased toward the majority class.

71
Multi-Selecteasy

Which TWO of the following are measures of central tendency?

Select 2 answers
A.Median
B.Range
C.Variance
D.Standard deviation
E.Mean
AnswersA, E

The median is a measure of central tendency because it identifies the middle value of an ordered dataset, dividing it into two equal halves. Unlike the mean, it resists distortion by extreme outliers, satisfying the stem's requirement for a positional average rather than a measure of dispersion or spread.

Why this answer

The median (A) is a measure of central tendency because it identifies the middle value of a dataset when ordered, representing the central point that divides the data into two equal halves. The mean (E) is also a measure of central tendency, calculated as the arithmetic average (sum of all values divided by the number of values), and it indicates the typical or central value of the data. The range (B) is a measure of dispersion, showing the difference between the maximum and minimum values, not a central value.

Variance (C) and standard deviation (D) are both measures of spread that quantify how far data points deviate from the mean, so they do not describe central tendency.

Exam trap

DA0-002 often tests the confusion between measures of central tendency (mean, median, mode) and measures of dispersion (range, variance, standard deviation), so candidates who see 'statistical measure' and pick variance or standard deviation fall into the trap.

72
MCQeasy

A data analyst wants to use a Z-score to standardize a dataset. The variable has a mean of 50 and a standard deviation of 10. What is the Z-score for a raw value of 70?

A.0.5
B.20
C.-2
D.2
AnswerD

Applying the Z-score formula (x minus mean, divided by standard deviation) gives (70−50)/10 = 2. This standardised value states the raw score sits two standard deviations above the mean, satisfying the stem's requirement to standardise using the given mean of 50 and standard deviation of 10.

Why this answer

The Z-score formula is Z = (X - μ) / σ, where X is the raw value, μ is the mean, and σ is the standard deviation. Plugging in the given values: Z = (70 - 50) / 10 = 20 / 10 = 2. Thus, the raw value of 70 is 2 standard deviations above the mean, corresponding to a Z-score of 2.

Exam trap

The trap here is confusing the difference between the raw value and the mean (20) with the Z-score, or incorrectly reversing the numerator to get a negative Z-score, which would misrepresent the direction from the mean.

How to eliminate wrong answers

Option A (0.5) is wrong because it results from dividing the standard deviation by the difference (10/20) instead of the correct order, or from misapplying the formula as (μ - X)/σ? Actually (50-70)/10 = -2, not 0.5. Option B (20) is wrong because it is simply the difference between the raw value and the mean (70 - 50 = 20) without dividing by the standard deviation. Option C (-2) is wrong because it reverses the sign, computing (50 - 70)/10 = -2, which would indicate the value is below the mean, but 70 is above the mean of 50.

73
MCQmedium

A data analyst notices that a dataset of customer ages has several missing values. Which method for handling missing data is most appropriate if the data is missing completely at random and the analyst wants to preserve sample size?

A.Forward-fill using the previous value
B.Impute with the mean age
C.Replace missing values with zero
D.Delete all rows with missing data
AnswerB

Mean imputation replaces each missing age with the variable's average, retaining every record and therefore preserving sample size. Because the data is missing completely at random, the missingness is unrelated to any variable, so mean substitution introduces minimal bias compared with deletion methods.

Why this answer

When data is missing completely at random (MCAR) and the analyst wants to preserve sample size, mean imputation is the standard approach — it replaces missing values with the average of the observed values, retaining all rows and avoiding the bias that deletion would introduce. For MCAR data, mean imputation produces unbiased estimates of the mean (though it reduces variance).

Exam trap

DA0-002 often tests the trade-off between preserving sample size and introducing bias — candidates pick deletion for 'cleanliness' or zero-fill for simplicity without considering the distortion each introduces.

How to eliminate wrong answers

Option A is wrong because forward-fill is appropriate for time-series or ordered data where the previous value is a reasonable proxy — customer ages have no inherent order, so forward-fill would introduce arbitrary values. Option C is wrong because replacing missing ages with zero is nonsensical (age zero is a newborn) and would severely distort the distribution and any downstream statistics. Option D is wrong because deleting rows with missing data reduces sample size and, if the missingness is not truly random, introduces selection bias — the question explicitly states the analyst wants to preserve sample size.

74
Multi-Selecthard

A data analyst is building a logistic regression model to predict whether a customer will churn (yes/no). Which TWO statements about logistic regression are correct?

Select 2 answers
A.It is used only for time series forecasting.
B.The dependent variable is continuous.
C.The output is a probability between 0 and 1.
D.It requires normally distributed errors.
E.It assumes a linear relationship between predictors and the log-odds of the outcome.
AnswersC, E

Logistic regression applies the sigmoid function to a linear combination of predictors, squashing its output into the 0 to 1 interval. That value is interpreted as the probability of the positive class, directly supporting the churn yes/no prediction the analyst requires.

Why this answer

Option C is correct because logistic regression applies the sigmoid (logistic) function to a linear combination of predictors, producing an output that is a probability bounded between 0 and 1, which is exactly what is needed to model the churn probability (yes/no). Option E is correct because logistic regression assumes linearity on the logit scale: the log-odds of the outcome, ln(p/(1-p)), is modeled as a linear function of the predictor variables. Option A is incorrect because logistic regression is a classification method for binary outcomes, not a time series forecasting technique.

Option B is incorrect because the dependent variable is binary/categorical (churn yes/no), not continuous. Option D is incorrect because logistic regression does not require normally distributed errors; that assumption belongs to linear regression, whereas logistic regression uses maximum likelihood estimation with a binomial error distribution.

Exam trap

DA0-002 often tests the confusion between linear and logistic regression assumptions — candidates incorrectly apply OLS assumptions (normal errors, continuous outcome) to logistic regression or miss that the linearity assumption applies to log-odds, not raw probabilities.

75
MCQhard

A data analyst is evaluating a classification model that predicts customer churn. The model has an accuracy of 95%, but the analyst suspects it is not performing well on the minority class (churners). The dataset is highly imbalanced, with only 5% churners. Which metric should the analyst prioritize to assess the model's ability to correctly identify churners?

A.Recall
B.Specificity
C.Accuracy
D.Precision
AnswerA

Recall measures the proportion of actual churners that were correctly identified. In an imbalanced dataset where churners are rare, high accuracy can be misleading because the model may simply predict the majority class. Recall directly assesses the model's sensitivity to the minority class, which is critical when missing a churner is costly.

Why this answer

Recall is the appropriate metric because it quantifies how many actual churners the model successfully identifies. In imbalanced settings, accuracy and specificity can be high even when the model fails to detect the minority class. Precision is also important but secondary when the goal is to capture as many churners as possible.

Exam trap

The trap here is assuming high accuracy indicates good performance, ignoring the class imbalance and the need to evaluate minority class detection.

Page 1 of 4 · 238 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Data Analysis questions.