Courseiva

CCNA Data Analysis Questions

75 of 238 questions · Page 2/4 · Data Analysis · Answers revealed

76
Multi-Selecthard

Which TWO of the following are valid techniques for validating the performance of a predictive model?

Select 2 answers
A.Bootstrapping
B.Feature scaling
C.Train-test split
D.K-fold cross-validation
E.Increasing training data
AnswersC, D

Splitting data into training and testing sets is a basic validation approach.

Why this answer

The train-test split (Option C) is a fundamental technique for validating predictive model performance by partitioning the dataset into separate training and testing subsets, ensuring the model is evaluated on unseen data to gauge generalization. This method directly addresses overfitting and provides an unbiased estimate of model accuracy, making it a standard practice in supervised learning workflows.

Exam trap

CompTIA often tests the distinction between data preprocessing techniques (like feature scaling) and actual model validation methods, leading candidates to mistakenly select feature scaling as a validation technique because it is a common step in the modeling pipeline.

77
MCQmedium

An analyst runs a simple linear regression with an R² value of 0.85. Which interpretation is correct?

A.85% of the variance in the dependent variable is explained by the independent variable.
B.The slope of the regression line is 0.85.
C.The independent variable is 85% correlated with the dependent variable.
D.85% of the data points lie on the regression line.
AnswerA

R² measures the proportion of variance in the dependent variable accounted for by the independent variable in the fitted model. A value of 0.85 therefore means 85% of that variance is explained, with the remaining 15% attributable to other factors.

Why this answer

R², the coefficient of determination, measures the proportion of variance in the dependent variable that is predictable from the independent variable(s). An R² of 0.85 means that 85% of the variance in the dependent variable is explained by the model, leaving 15% unexplained. This is the standard interpretation in regression analysis.

Exam trap

DA0-002 often tests the confusion between R² and the correlation coefficient r, or misinterpreting R² as the slope or the proportion of points on the line.

How to eliminate wrong answers

Option B is wrong because the slope of the regression line is given by the coefficient estimate (β₁), not R²; R² is a unitless measure of fit. Option C is wrong because the correlation coefficient r is the square root of R² (for simple linear regression), so r = √0.85 ≈ 0.92, not 0.85. Option D is wrong because R² does not indicate the proportion of data points on the line; it measures variance explained, and even a perfect fit (R²=1) does not require all points to lie exactly on the line unless the data are perfectly linear.

78
Multi-Selecthard

An analyst is performing K-means clustering on customer data. The elbow method shows a clear bend at k=4. Which THREE of the following are true about K-means clustering with k=4?

Select 3 answers
A.The number of clusters is determined to be 4.
B.The algorithm will always produce the same clusters regardless of initial centroids.
C.The centroids are recomputed iteratively until convergence.
D.Categorical variables should be standardised before clustering.
E.The algorithm minimises the sum of squared distances between points and their assigned centroid.
AnswersA, C, E

The elbow method plots within-cluster sum of squares against k; the clear bend at k=4 indicates diminishing returns beyond four clusters, so the analyst selects four as the cluster count. This directly sets the k parameter for the algorithm.

Why this answer

Option A is correct because the elbow method identifies the optimal number of clusters at the point where the within-cluster sum of squares (inertia) begins to level off, and a clear bend at k=4 indicates that 4 is the chosen number of clusters. Option C is correct because K-means is an iterative algorithm: after assigning points to the nearest centroid, it recomputes each centroid as the mean of the points in its cluster, repeating until assignments and centroids stabilise (convergence). Option E is correct because the objective function K-means minimises is the sum of squared Euclidean distances between each point and its assigned centroid (inertia).

Option B is not correct because K-means results depend on initial centroid placement and can converge to different local optima, which is why k-means++ or multiple restarts are used. Option D is not correct because categorical variables are not suitable for K-means (which relies on Euclidean distance and means); they require encoding or a different algorithm such as K-modes, and standardisation applies to numerical features, not categorical ones.

Exam trap

The trap is thinking K-means always gives the same result (it doesn't due to random initialization) and that categorical variables can be standardised (they need encoding instead).

79
MCQmedium

A marketing analyst wants to predict whether a customer will churn (yes/no) based on account age and monthly charges. Which regression technique is most appropriate?

A.Logistic regression
B.Simple linear regression
C.Multiple linear regression
D.K-means clustering
AnswerA

Churn is a binary yes/no outcome, and logistic regression models the probability of a categorical dependent variable using a sigmoid function. Linear regression would predict continuous values outside 0–1, making it unsuitable for this classification scenario.

Why this answer

Logistic regression is specifically designed for binary classification problems, such as predicting churn (yes/no). It models the probability of the outcome using a logistic function, making it appropriate for this scenario. Simple and multiple linear regression are for continuous outcomes, and K-means is for clustering, not prediction.

Exam trap

The trap is confusing regression techniques: candidates might think any regression can predict binary outcomes, but only logistic regression is suited for classification. Linear regression outputs continuous values, which are not probabilities.

How to eliminate wrong answers

Option B is wrong because simple linear regression predicts a continuous dependent variable, not a binary outcome. Option C is wrong because multiple linear regression also predicts a continuous outcome and cannot directly model a binary response. Option D is wrong because K-means clustering is an unsupervised learning algorithm used for grouping data, not for predicting a binary target.

80
MCQhard

A data scientist is analyzing a dataset with multiple features and wants to apply k-means clustering to segment customers. She chooses k = 4 based on the elbow method. During the iteration process, which of the following correctly describes a step in the k-means algorithm?

A.Compute the covariance matrix and use principal components to initialize centroids.
B.Use hierarchical clustering to determine initial centroids.
C.Randomly assign centroids and then compute distances to the cluster medians.
D.Assign each point to the nearest centroid based on Euclidean distance, then update centroids as the mean of points in each cluster.
AnswerD

K-means alternates two steps: each point is assigned to the cluster whose centroid is nearest by Euclidean distance, then every centroid is recomputed as the mean of its assigned points. Iteration repeats until assignments stabilise, satisfying the k=4 segmentation.

Why this answer

In the k-means algorithm, after initializing centroids, each data point is assigned to the nearest centroid based on Euclidean distance, and then centroids are recomputed as the mean of all points in the cluster. This iterative process continues until convergence. Option D accurately describes this step.

Exam trap

The trap is mixing up k-means with other algorithms (e.g., PCA for initialization) or using median instead of mean. Candidates might also confuse the update step with using medians (as in k-medians).

How to eliminate wrong answers

Option A is wrong because computing the covariance matrix and using principal components is part of PCA, not k-means initialization. Option B is wrong because hierarchical clustering is a different clustering method and is not a standard step in k-means. Option C is wrong because k-means updates centroids as the mean, not the median, and it does not use cluster medians.

81
MCQmedium

Refer to the exhibit. Which type of ensemble method is being used?

A.Boosting
B.Stacking
C.Voting
D.Bagging
AnswerD

Bagging trains multiple instances of the same algorithm on bootstrap samples drawn with replacement, then aggregates their predictions by voting or averaging. This parallel, variance-reducing structure distinguishes it from boosting, which trains sequentially on reweighted data.

Why this answer

The exhibit shows multiple base models (Model 1, Model 2, Model 3) trained in parallel on bootstrap samples of the data, and their predictions are combined via averaging (regression) or majority voting (classification). This parallel training with resampled data and equal-weight aggregation is the defining characteristic of bagging (Bootstrap Aggregating).

Exam trap

CompTIA often tests the distinction between bagging and boosting by showing parallel vs. sequential training diagrams, and the trap here is confusing the parallel bootstrap resampling with the sequential error-correction approach of boosting.

How to eliminate wrong answers

Option A is wrong because boosting trains models sequentially, where each subsequent model focuses on correcting the errors of the previous one, not in parallel on bootstrap samples. Option B is wrong because stacking uses a meta-learner to combine predictions from diverse base models, not simple averaging or majority voting. Option C is wrong because voting typically combines predictions from different model types (e.g., logistic regression, SVM) trained on the same dataset, not from the same model type trained on bootstrap samples.

82
MCQmedium

A data analyst is reviewing a SQL query that joins three large tables. The query takes over an hour to run. The analyst notices that the WHERE clause filters on indexed columns in only two tables. Which of the following should the analyst do first to improve performance?

A.Use subqueries instead of joins
B.Check the query execution plan and optimize join order
C.Add indexes to all columns used in joins
D.Increase server memory
AnswerB

The execution plan reveals how the optimiser orders and joins the three tables, including scan and join methods. Since only two tables have indexed filter columns, examining the plan identifies whether join order or a missing index causes the hour-long runtime.

Why this answer

The query execution plan reveals how the database engine processes joins and filters. By checking the plan, the analyst can identify the most selective filter and rearrange the join order to reduce the number of rows processed early, which is the most impactful first step. Optimizing join order leverages existing indexes without requiring schema changes or hardware upgrades.

Exam trap

CompTIA often tests the misconception that adding indexes or hardware is the immediate fix, when in fact analyzing the execution plan and adjusting join order is the cheapest and most effective first step.

How to eliminate wrong answers

Option A is wrong because subqueries often perform worse than joins in large-table scenarios, as they can lead to correlated subquery execution and repeated scans. Option C is wrong because adding indexes to all join columns is unnecessary and may degrade write performance; the analyst should first verify if existing indexes are being used efficiently via the execution plan. Option D is wrong because increasing server memory is a reactive, costly measure that does not address the root cause of inefficient query processing, such as poor join order or missing index usage.

83
MCQmedium

An analyst compares average sales across three different store locations using a statistical test. Which test is most appropriate?

A.ANOVA
B.t-test
C.Correlation analysis
D.Chi-square test
AnswerA

ANOVA compares means across three or more independent groups on one continuous dependent variable, here average sales per store location. It tests whether at least one group mean differs, controlling family-wise error. A t-test suits only two groups, so ANOVA is the appropriate choice for three locations.

Why this answer

ANOVA compares means across three or more groups.

84
Multi-Selecthard

A data analyst is evaluating the quality of a customer database. Which THREE of the following are dimensions of data quality?

Select 3 answers
A.Completeness
B.Correlation
C.Timeliness
D.Accuracy
E.Variance
AnswersA, C, D

Completeness measures whether required customer attributes are populated, directly assessing missing values in the database. It is a recognised data quality dimension because gaps in mandatory fields undermine segmentation, contactability and downstream analytics, satisfying the evaluation's need to quantify absent data.

Why this answer

Accuracy, completeness, and timeliness are standard data quality dimensions.

85
MCQeasy

Which data quality dimension ensures that data represents the real-world object or event correctly?

A.Accuracy
B.Completeness
C.Consistency
D.Timeliness
AnswerA

Accuracy verifies that values faithfully reflect the real-world object or event they describe, directly satisfying the stem's requirement for correct representation. Unlike validity, which only confirms conformance to defined formats or rules, accuracy measures correspondence with reality itself, making it the dimension that ensures data genuinely depicts what it claims to represent.

Why this answer

Accuracy is the data quality dimension that measures how closely data values reflect the real-world object, event, or condition they are intended to describe. If a customer's address, a transaction amount, or a sensor reading is recorded incorrectly, the data is inaccurate even if it is complete, consistent, and timely. Accuracy is therefore the dimension specifically concerned with correctness of representation.

Exam trap

The trap here is confusing accuracy with completeness or consistency — candidates often pick completeness because the data 'looks full,' but the question asks specifically about correctly representing the real-world object or event.

How to eliminate wrong answers

Option B is wrong because completeness only measures whether all required data values are present, not whether those values are correct — a record can be 100% populated yet contain wrong values. Option C is wrong because consistency concerns whether the same data is represented uniformly across systems, formats, or time periods, not whether it matches reality. Option D is wrong because timeliness addresses whether data is available and up to date when needed, which says nothing about whether the values themselves are correct.

86
MCQmedium

A data analyst is performing time series analysis on monthly sales data and notices a consistent pattern of higher sales every December. Which component of time series does this represent?

A.Trend
B.Irregular component
C.Seasonality
D.Cyclical
AnswerC

Seasonality describes a repeating pattern that recurs at fixed intervals shorter than a year, such as December sales peaks. The consistent annual December uplift satisfies this definition, distinguishing it from trend, which is a longer-term directional movement, and from irregular noise.

Why this answer

Seasonality refers to a repeating pattern that occurs at fixed, predictable intervals within a year, such as higher sales every December or increased traffic every Monday. Because the December spike recurs annually at the same calendar point, it is a seasonal component. Seasonality is distinguished from trend (long-term direction) and cyclical (irregular multi-year fluctuations) by its fixed, calendar-based period.

Exam trap

DA0-002 often tests the seasonality vs. cyclical distinction — candidates pick cyclical because both repeat, but only seasonality has a fixed, calendar-based period like December.

How to eliminate wrong answers

Option A is wrong because trend describes the long-term upward or downward movement of the series over many periods, not a recurring annual spike. Option B is wrong because the irregular component is the random, unpredictable residual left after trend, seasonality, and cyclical effects are removed — December sales are predictable, not random. Option D is wrong because cyclical components fluctuate over longer, non-fixed periods (often tied to economic cycles) and are not tied to a specific calendar month like December.

87
Multi-Selecthard

A data analyst is cleaning a dataset and identifies several outliers. Which TWO methods are appropriate for handling outliers?

Select 2 answers
A.Capping
B.Mean imputation
C.Removal
D.Min-max normalization
E.Forward-fill
AnswersA, C

Capping replaces extreme values with a defined threshold, such as the 1st or 99th percentile, retaining the record while limiting its influence. This satisfies the scenario's need to handle outliers without discarding data, unlike deletion, which removes rows entirely and risks losing valid observations.

Why this answer

Capping (A) is correct because it winsorizes extreme values by replacing outliers with a boundary value such as the 1st/99th percentile or a value derived from the IQR (e.g., Q1 − 1.5×IQR, Q3 + 1.5×IQR), preserving the record while limiting the outlier's influence. Removal (C) is correct because dropping outlier rows (or excluding them from analysis) is a standard, defensible approach when the values are confirmed to be erroneous or when their influence must be eliminated. Mean imputation (B) is not an outlier-handling method; it replaces missing values with the column mean and would actually be distorted by the very outliers present.

Min-max normalization (D) merely rescales all values to the [0,1] range and does not reduce or eliminate outlier effects. Forward-fill (E) is a time-series missing-value technique that propagates the last valid observation and does nothing to address outliers.

Exam trap

The trap here is confusing data transformation techniques (like normalization or imputation) with outlier-specific handling methods; candidates might select mean imputation or min-max normalization because they are common preprocessing steps, but they do not directly address outliers.

88
MCQmedium

A data analyst is examining sales data for a retail chain and notices that the mean monthly sales is $50,000 while the median is $35,000. Which of the following best describes the distribution of the sales data?

A.The distribution is right-skewed.
B.The distribution is bimodal.
C.The distribution is left-skewed.
D.The distribution is symmetrical.
AnswerA

A mean of $50,000 exceeding the median of $35,000 indicates positive skew: extreme high-value months pull the mean upward while the median stays resistant. This satisfies the stem's constraint, so the distribution is right-skewed, with the longer tail extending toward higher sales values.

Why this answer

When the mean ($50,000) is greater than the median ($35,000), the distribution has a long tail on the right side pulling the mean upward, which is the definition of a right-skewed (positively skewed) distribution. In right-skewed data, a few high values inflate the mean above the median. Therefore the sales data is right-skewed.

Exam trap

DA0-002 often tests the mean-vs-median relationship, and candidates frequently reverse the skew direction; the rule is mean > median means right skew, mean < median means left skew.

How to eliminate wrong answers

Option B is wrong because bimodality refers to two distinct peaks in the distribution and is not indicated by a mean-median gap; the question gives no evidence of two modes. Option C is wrong because left-skewed (negatively skewed) distributions have the mean less than the median, which is the opposite of the given values. Option D is wrong because in a symmetrical distribution the mean and median are approximately equal, but here they differ by $15,000.

89
MCQhard

A data analyst is cleaning a dataset and finds that 5% of values in the 'income' column are missing. The analyst decides to impute missing values using the mean of the non-missing values. Which potential issue should the analyst be most concerned about?

A.The imputation may reduce the variance and distort the distribution.
B.The imputation is not valid because the missing rate is too low.
C.The imputation will increase the standard deviation of the variable.
D.The imputation will create outliers.
AnswerA

Mean imputation replaces missing values with the column average, pulling extreme values toward the centre. This artificially shrinks variance and narrows the distribution, biasing subsequent statistical analysis and understating the true spread of income data.

Why this answer

Mean imputation replaces missing values with the average of observed values, which pulls imputed values toward the center of the distribution. This artificially reduces the variance and can distort the shape of the distribution (e.g., making it more peaked), biasing downstream statistics and models. The analyst should be most concerned about this distortion rather than the missing rate itself.

Exam trap

DA0-002 often tests whether candidates recognize that mean imputation reduces variance and distorts distribution, rather than mistakenly thinking it increases variability or creates outliers.

How to eliminate wrong answers

Option B is wrong because a 5% missing rate is not too low to impute; low missingness is generally favorable, and there is no rule that invalidates imputation at low rates. Option C is wrong because mean imputation decreases, not increases, the standard deviation since imputed values cluster at the mean. Option D is wrong because mean imputation does not create outliers; it actually reduces extreme values by replacing missing entries with the central mean.

90
MCQmedium

A dataset contains a variable 'Income' with many missing values. The analyst decides to impute missing values with the median income of the non-missing values. Which type of imputation is this?

A.Interpolation
B.Deletion
C.Median imputation
D.Forward-fill imputation
AnswerC

Median imputation replaces each missing entry with the median calculated from the observed, non-missing incomes, satisfying the stem's requirement to use the median of non-missing values. Unlike mean imputation, the median resists distortion from outliers and skew, making it robust for income data, which is typically right-skewed.

Why this answer

Imputing missing values with the median of the non-missing values is exactly median imputation. It is a common method for handling missing data, especially when the distribution is skewed or contains outliers.

Exam trap

The trap is confusing median imputation with other imputation methods like interpolation or forward-fill. Candidates might also think deletion is a form of imputation, but it is not.

How to eliminate wrong answers

Option A is wrong because interpolation estimates missing values using surrounding data points, often in time series, not a global median. Option B is wrong because deletion removes missing data, not imputes it. Option D is wrong because forward-fill imputation carries forward the last observed value, typically used in time series, not a median.

91
MCQeasy

A data analyst is reviewing a dashboard that tracks monthly sales. The analyst notices that the 'Total Sales' metric for the current month is significantly lower than previous months, but the underlying daily sales data appears normal. The dashboard uses a live connection to the database. Which of the following is the most likely cause of the discrepancy?

A.The database connection is using a stale cache that has not refreshed since last month.
B.The dashboard's date filter is set to include only the first week of the current month.
C.The 'Total Sales' metric is calculated using a moving average that lags by one month.
D.The sales data for the current month has not been fully loaded into the database yet.
AnswerB

A date filter that restricts data to only the first week would cause the total to appear lower than a full month, even if daily sales are normal. This is a common dashboard configuration issue where filters are not updated to reflect the full current month-to-date period.

Why this answer

The most likely cause is a date filter limiting the data to a subset of the month, such as the first week, which would reduce the total while daily sales remain normal. Other options would affect the daily data or show different patterns, making them less consistent with the scenario.

Exam trap

The trap here is overlooking dashboard filter settings and assuming a data pipeline issue when the problem is often in the visualization configuration.

92
MCQeasy

A data analyst needs to identify outliers in a dataset. Which of the following is a common method based on the interquartile range (IQR)?

A.Values more than 2 standard deviations from the mean
B.Values that are negative
C.Values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR
D.Values below the 5th percentile or above the 95th percentile
AnswerC

The IQR method flags observations falling outside Tukey's fences: Q1 minus 1.5 times IQR and Q3 plus 1.5 times IQR. Values beyond these bounds lie far from the central 50% of the distribution, making them statistical outliers.

Why this answer

A common rule is to consider any data point below Q1 - 1.5*IQR or above Q3 + 1.5*IQR as an outlier.

93
Multi-Selecthard

Which THREE of the following are assumptions of linear regression? (Select THREE).

Select 3 answers
A.Normal distribution of independent variables
B.Multicollinearity among independent variables
C.Independence of errors
D.Homoscedasticity (constant variance of errors)
E.Linearity between independent and dependent variables
AnswersC, D, E

Independence of errors means each residual is uncorrelated with the others, so no observation's error predicts another's. Violations, common in time-series or clustered data, bias standard errors and invalidate inference, making this a core linear regression assumption.

Why this answer

Option C is correct because linear regression assumes the residuals (errors) are independent of one another, meaning there is no autocorrelation; violations such as serial correlation in time-series data bias standard errors. Option D is correct because homoscedasticity—constant variance of the error terms across all levels of the predictors—is a core Gauss-Markov assumption; heteroscedasticity makes ordinary least squares estimates inefficient and invalidates inference. Option E is correct because the model assumes a linear relationship between each independent variable and the dependent variable, so the expected value of the outcome is a linear function of the coefficients.

Option A is not an assumption: normality applies to the error terms (and is only strictly needed for small-sample inference), not to the independent variables themselves. Option B is not an assumption but a problem to avoid—multicollinearity among predictors inflates coefficient variances and should be absent, not assumed.

Exam trap

The trap here is that candidates confuse the normality assumption for errors with a normality assumption for the independent variables, leading them to incorrectly select Option A.

94
Multi-Selecteasy

Which TWO of the following are examples of supervised learning algorithms?

Select 2 answers
A.Linear regression
B.K-means clustering
C.Principal component analysis (PCA)
D.Decision trees
E.Apriori algorithm
AnswersA, D

Linear regression is supervised because it trains on labelled data, learning a mapping from input features to a continuous target value by minimising error against known outcomes. This labelled-target dependency is precisely what the question's supervised learning criterion requires, unlike unsupervised methods that find structure without labels.

Why this answer

Linear regression (A) is a supervised learning algorithm because it learns a mapping from labeled input-output pairs by fitting a line (or hyperplane) that minimizes a loss such as mean squared error, making it a classic regression technique. Decision trees (D) are also supervised: they are trained on labeled data, recursively splitting features (e.g., via Gini impurity or information gain) to predict a target class or value. In contrast, K-means clustering (B) is unsupervised, as it groups unlabeled data by minimizing within-cluster variance without target labels.

Principal component analysis (C) is an unsupervised dimensionality-reduction technique that finds orthogonal components maximizing variance, requiring no labels. The Apriori algorithm (E) is unsupervised association-rule mining that discovers frequent itemsets from unlabeled transactional data, so it is not supervised learning.

Exam trap

CompTIA often tests the distinction between supervised and unsupervised learning by including clustering (K-means) and association (Apriori) as distractors, which candidates mistakenly think are supervised because they involve pattern discovery.

95
MCQmedium

A data analyst is examining the relationship between advertising spend (in dollars) and revenue (in dollars). The Pearson correlation coefficient r is calculated as +0.92. Which of the following interpretations is correct?

A.There is a strong negative linear relationship.
B.There is no linear relationship.
C.There is a strong positive linear relationship.
D.92% of the variation in revenue is explained by advertising spend.
AnswerC

A coefficient of +0.92 sits close to +1, so the two variables move together almost perfectly along a straight line, with advertising spend rising as revenue rises. The positive sign confirms the direction is increasing rather than inverse, and the magnitude confirms the relationship is strong and linear.

Why this answer

A Pearson correlation coefficient of +0.92 indicates a strong positive linear relationship between advertising spend and revenue. The positive sign means as one variable increases, the other tends to increase, and the magnitude (close to 1) indicates a strong linear association.

Exam trap

The trap is confusing the correlation coefficient r with the coefficient of determination r²; candidates often mistakenly think r = 0.92 means 92% of the variation is explained, but it's actually r² that gives that percentage.

How to eliminate wrong answers

Option A is wrong because a negative relationship would have a negative r value (e.g., -0.92). Option B is wrong because an r near 0 would indicate no linear relationship, but 0.92 is far from 0. Option D is wrong because r^2 (coefficient of determination) represents the proportion of variance explained; r^2 = 0.92^2 = 0.8464, so about 84.6% of the variation is explained, not 92%.

96
MCQmedium

A data analyst is examining a dataset of daily website visits collected over the past year. The analyst wants to determine whether the distribution of visits on weekdays differs significantly from the distribution on weekends, without assuming normality. The analyst has already separated the data into two independent groups. Which statistical test should the analyst use?

A.One-sample t-test
B.Mann-Whitney U test
C.Paired samples t-test
D.Chi-square goodness-of-fit test
AnswerB

The Mann-Whitney U test is a nonparametric alternative to the independent samples t-test, comparing whether two independent groups come from the same distribution. It does not assume normality, making it appropriate for skewed visit counts. It ranks all observations and assesses whether one group tends to have higher values. This directly addresses the analyst's question of whether weekday and weekend visit distributions differ without distributional assumptions.

Why this answer

The Mann-Whitney U test is a nonparametric method that compares two independent groups without requiring normality. It assesses whether one group tends to have larger values than the other by ranking all observations. Since the analyst wants to compare weekday and weekend website visits and cannot assume normality, this test is the correct choice.

It is specifically designed for independent groups and continuous or ordinal data.

Exam trap

The trap here is assuming that any nonparametric test for two groups works, but the paired t-test and one-sample t-test require different data structures, while the chi-square test is for categorical frequencies, not continuous distributions.

97
MCQeasy

A data analyst needs to identify the most frequently occurring value in a dataset. Which measure of central tendency should they use?

A.Mode
B.Standard deviation
C.Median
D.Mean
AnswerA

The mode identifies the value appearing most frequently, directly satisfying the requirement to find the most frequently occurring value. Unlike the mean, which averages all values, or the median, which finds the middle position, the mode alone measures frequency of occurrence, making it the appropriate measure of central tendency for this scenario.

Why this answer

The mode is the measure of central tendency that identifies the most frequently occurring value in a dataset. Unlike the mean or median, the mode directly counts the frequency of each distinct value and returns the value with the highest count, making it the correct choice for this specific requirement.

Exam trap

The trap here is that candidates often confuse 'most frequently occurring' with 'average' or 'middle value' and incorrectly choose mean or median, especially when the dataset is numeric and they assume central tendency always refers to mean.

How to eliminate wrong answers

Option B (Standard deviation) is wrong because it measures the dispersion or spread of data points around the mean, not the frequency of occurrence of any single value. Option C (Median) is wrong because it identifies the middle value when the dataset is sorted, which does not indicate which value appears most often. Option D (Mean) is wrong because it calculates the arithmetic average of all values, which can be skewed by outliers and does not reflect frequency of occurrence.

98
MCQmedium

A data analyst is examining a dataset of employee salaries and notices that the distribution is heavily right-skewed due to a few executives with very high salaries. The analyst wants to apply a transformation to make the distribution more symmetric for further analysis. Which transformation is most appropriate?

A.Square root transformation
B.Logarithmic transformation
C.Square transformation
D.Exponential transformation
AnswerB

The logarithmic transformation is highly effective for right-skewed data with positive values, as it compresses the upper tail and reduces the influence of extreme values. For salary data with a few very high earners, taking the log of salaries will make the distribution more symmetric and closer to normal. This is a common practice before applying statistical techniques that assume normality.

Why this answer

For right-skewed data with positive values, the logarithmic transformation is the most appropriate to reduce skewness and make the distribution more symmetric. It compresses the upper tail, mitigating the influence of extreme high values. Square root is milder, while square and exponential transformations would increase skewness.

Therefore, the log transformation is the best choice.

Exam trap

The trap here is confusing transformations that increase skewness (square, exponential) with those that reduce it (log, square root).

99
MCQhard

A data analyst is performing a chi-square test of independence on a contingency table of customer satisfaction (satisfied vs. dissatisfied) and product type (A, B, C). The test yields a p-value of 0.04 with α = 0.05. What is the correct conclusion?

A.There is no evidence of an association between satisfaction and product type.
B.There is a significant association between satisfaction and product type.
C.The test is invalid because the expected counts are too low.
D.Satisfaction and product type are independent.
AnswerB

With p = 0.04 below the α = 0.05 threshold, the null hypothesis of independence is rejected, so satisfaction and product type are statistically associated. The chi-square test of independence detects whether the two categorical variables' observed cell frequencies deviate from those expected under independence.

Why this answer

With a p-value of 0.04 and α = 0.05, the p-value is less than the significance level, so we reject the null hypothesis of independence. This means there is statistically significant evidence of an association between customer satisfaction and product type. The correct conclusion is that a significant association exists.

Exam trap

DA0-002 often tests the interpretation of p-values versus α, and candidates frequently confuse 'fail to reject' with 'accept the null' or misinterpret a significant result as proving causation rather than association.

How to eliminate wrong answers

Option A is wrong because it states there is no evidence of association, which would be the conclusion if we failed to reject the null hypothesis (p ≥ α). Option C is wrong because the question does not provide information about expected counts; while chi-square requires expected counts ≥ 5 in each cell, the p-value alone does not indicate invalidity. Option D is wrong because it states the variables are independent, which is the null hypothesis we rejected.

100
MCQmedium

A company wants to segment its customers into distinct groups based on purchasing behavior. Which algorithm is best suited for this task?

A.Decision tree
B.Logistic regression
C.K-means clustering
D.Linear regression
AnswerC

K-means clustering partitions unlabelled records into k groups by minimising within-cluster variance, using distance between feature vectors. It suits segmentation on purchasing behaviour, where no predefined labels exist, satisfying the requirement to form distinct customer groups.

Why this answer

K-means clustering is an unsupervised learning algorithm that partitions data into K distinct clusters based on feature similarity, making it ideal for segmenting customers by purchasing behavior without predefined labels. It groups customers who exhibit similar purchasing patterns, enabling the company to identify natural segments for targeted marketing.

Exam trap

The trap here is that candidates often confuse supervised learning algorithms (like decision trees or logistic regression) with unsupervised clustering, mistakenly thinking that any algorithm that 'groups' data can be used for segmentation without recognizing the need for unlabeled data.

How to eliminate wrong answers

Option A is wrong because a decision tree is a supervised learning algorithm used for classification or regression, requiring labeled training data to predict outcomes, not for discovering unknown groupings in unlabeled data. Option B is wrong because logistic regression is a supervised classification algorithm for binary or multinomial outcomes, relying on labeled target variables, and cannot perform unsupervised clustering. Option D is wrong because linear regression is a supervised regression algorithm that models the relationship between a dependent variable and one or more independent variables, and it is not designed to segment data into distinct groups without predefined categories.

101
Multi-Selectmedium

A data analyst is preparing a dataset for analysis and needs to ensure data quality. Which TWO of the following are dimensions of data quality?

Select 2 answers
A.Volume
B.Velocity
C.Consistency
D.Variety
E.Accuracy
AnswersC, E

Consistency is a recognised data quality dimension, confirming that values remain uniform across systems and records without contradictory entries. It satisfies the stem's requirement for genuine quality dimensions, alongside accuracy, completeness, timeliness and validity, so it qualifies as one of the two correct selections.

Why this answer

Consistency (C) is a core dimension of data quality because it ensures that the same data values are represented uniformly across different datasets, systems, or records, preventing contradictions that would corrupt analysis results. Accuracy (E) is also a fundamental data quality dimension, as it verifies that data correctly reflects the real-world entities or events it is meant to describe, which is essential for trustworthy analysis. In contrast, Volume (A), Velocity (B), and Variety (D) are the three defining characteristics of big data (the '3 Vs'), describing the scale, speed, and diversity of data rather than its quality.

Therefore, only Consistency and Accuracy belong to the set of data quality dimensions.

Exam trap

DA0-002 often mixes big-data Vs (volume, velocity, variety) with data quality dimensions (accuracy, consistency, completeness), so candidates must distinguish characteristics of data from measures of its quality.

102
MCQmedium

A data analyst wants to test if the proportion of customers who prefer Product A over Product B is different from 50%. She surveys 200 customers and finds that 120 prefer Product A. Which statistical test should she use?

A.Chi-square test of independence
B.One-sample z-test for proportions
C.ANOVA
D.Two-sample t-test
AnswerB

The scenario compares one observed sample proportion against a hypothesised population proportion of 0.50, with a large sample of 200. A one-sample z-test for proportions is the appropriate parametric test for this single-proportion hypothesis, unlike chi-square or two-sample alternatives.

Why this answer

The analyst wants to test whether the proportion of customers preferring Product A differs from 50%, using a single sample of 200 customers with 120 preferring A. This is a one-sample test of a proportion against a hypothesized value (0.5), so the one-sample z-test for proportions is appropriate. It compares the observed proportion (0.60) to the null hypothesis proportion (0.50).

Exam trap

DA0-002 often tests test selection by scenario; candidates confuse proportion tests with chi-square or t-tests, especially when the word 'prefer' suggests a comparison between two products rather than a single proportion against a benchmark.

How to eliminate wrong answers

Option A is wrong because the chi-square test of independence compares two categorical variables across a contingency table, not a single proportion against a hypothesized value. Option C is wrong because ANOVA compares means across three or more groups, not proportions. Option D is wrong because a two-sample t-test compares means of two independent groups, not a single proportion to a target.

103
MCQmedium

A data analyst is exploring a dataset of 8,000 customer transactions and notices that the 'transaction_amount' column has a mean of $120 but a median of $45. A small number of transactions exceed $10,000. Which measure of central tendency should the analyst report to describe the typical transaction?

A.The range, because it shows the spread between the smallest and largest transactions
B.The mean, because it uses every value in the dataset
C.The mode, because it identifies the most frequently occurring transaction amount
D.The median, because it is resistant to the influence of extreme values
AnswerD

The median is the middle value when data is ordered, so a handful of transactions above $10,000 do not shift it. With a mean far above the median, the distribution is right-skewed, and the median better represents the typical transaction amount experienced by most customers.

Why this answer

A large gap between the mean and median signals skew, and here the high-value transactions drag the mean well above the middle of the distribution. The median marks the point where half the transactions are smaller and half larger, so it reflects the typical customer experience far better. For skewed monetary data, the median is the more representative measure of central tendency.

Exam trap

The trap here is defaulting to the mean as the summary of a numeric column without checking whether extreme values have distorted it.

104
MCQmedium

A retail company wants to predict sales based on advertising spend and season. Which data modeling technique should the analyst use?

A.Simple linear regression
B.Multiple linear regression
C.Logistic regression
D.K-means clustering
AnswerB

Multiple linear regression models a continuous outcome, sales, as a linear function of two or more predictors, advertising spend and season. Season enters as a categorical dummy variable, satisfying the requirement to predict sales from both numeric and categorical inputs.

Why this answer

Multiple linear regression is the correct technique because the analyst needs to model a continuous outcome (sales) based on two or more predictor variables: advertising spend (continuous) and season (categorical, typically encoded as dummy variables). This allows the model to capture the independent effect of each predictor on sales, which simple linear regression cannot do because it only handles one predictor.

Exam trap

The trap here is that candidates often confuse simple linear regression with multiple linear regression, thinking that 'linear regression' alone suffices, but the exam specifically tests whether you recognize that multiple predictors require multiple regression.

How to eliminate wrong answers

Option A is wrong because simple linear regression can only model the relationship between one independent variable and the dependent variable, but here we have two predictors (advertising spend and season). Option C is wrong because logistic regression is used for binary or categorical outcome variables (e.g., yes/no), not for continuous outcomes like sales. Option D is wrong because K-means clustering is an unsupervised learning technique used to group similar data points, not to predict a continuous target variable.

105
Multi-Selectmedium

A data team is preparing data for a clustering analysis. Which THREE of the following steps are commonly part of data cleaning?

Select 3 answers
A.Removing duplicate records
B.Imputing missing values
C.Calculating the mean
D.Training a regression model
E.Capping outliers at the 5th and 95th percentiles
AnswersA, B, E

Removing duplicate records eliminates redundant observations that would otherwise distort distance calculations, causing clustering algorithms to over-weight repeated points. This directly satisfies the stem's data-cleaning requirement by ensuring each entity contributes once, preventing artificial density concentrations that skew centroid placement and cluster assignment.

Why this answer

Option A (Removing duplicate records) is correct because duplicate rows distort distance calculations in clustering, causing the same observation to be counted multiple times and biasing cluster centroids, so deduplication is a standard data-cleaning step. Option B (Imputing missing values) is correct because clustering algorithms such as k-means cannot handle nulls, so missing entries must be filled via mean/median/mode imputation, k-NN, or similar methods before analysis. Option E (Capping outliers at the 5th and 95th percentiles) is correct because winsorizing extreme values limits their disproportionate influence on distance metrics and centroid placement, which is a recognized cleaning technique.

Option C (Calculating the mean) is not a cleaning step but a descriptive statistic or profiling operation, and Option D (Training a regression model) is a modeling task, not data preparation, so neither belongs to data cleaning.

Exam trap

The trap here is confusing data cleaning with data exploration or modeling — candidates see 'calculating the mean' and think it's part of preparation because it's a common early step, but it doesn't clean anything.

106
MCQmedium

A data scientist is performing K-means clustering on customer data. She plots the within-cluster sum of squares (WCSS) for different values of k and observes an 'elbow' at k=4. What does this indicate?

A.The optimal number of clusters is 4
B.The algorithm should be run with k=3 to avoid overfitting
C.The data contains exactly 4 outliers
D.The WCSS is minimized at k=4, indicating perfect clustering
AnswerA

The elbow marks where adding clusters stops yielding meaningful WCSS reduction, so k=4 balances model complexity against fit. Beyond four, each extra cluster captures only marginal variance, indicating diminishing returns. This inflection point therefore identifies four as the optimal cluster count for the customer dataset.

Why this answer

The elbow method suggests that adding more clusters beyond k=4 yields diminishing returns, so k=4 is a suitable number of clusters.

107
MCQmedium

A data analyst is evaluating a multiple regression model with three predictors. The R² value is 0.85. Which of the following is the best interpretation of R²?

A.85% of the variance in the outcome is explained by the predictors.
B.85% of the predicted values are correct.
C.The model has a high bias.
D.The model has a strong correlation of 0.85.
AnswerA

R² measures the proportion of total variance in the dependent variable accounted for by the three predictors collectively. A value of 0.85 therefore means 85% of outcome variance is explained, with the remaining 15% attributable to other factors or error.

Why this answer

R² (coefficient of determination) measures the proportion of variance in the dependent variable explained by the independent variables in the model. An R² of 0.85 means 85% of the variability in the outcome is accounted for by the three predictors, with the remaining 15% attributable to factors not in the model or random noise. This is a goodness-of-fit measure, not an accuracy percentage or correlation coefficient.

Exam trap

DA0-002 often tests the misconception that R² is a percentage of correct predictions or a correlation value — candidates confuse goodness-of-fit with accuracy or with Pearson's r.

How to eliminate wrong answers

Option B is wrong because R² does not measure the percentage of correct predictions — that would be classification accuracy or a metric like MAPE/RMSE for regression, and R² can even be negative for poorly fit models. Option C is wrong because high bias refers to underfitting (a model too simple to capture the pattern), which is unrelated to the R² value itself; a high R² typically indicates low bias. Option D is wrong because R² is not a correlation coefficient — the correlation r would be the square root of R² (≈0.92) only in simple linear regression, and even then R² measures explained variance, not correlation strength.

108
Multi-Selecteasy

Which TWO of the following are dimensional modeling techniques commonly used in data warehouses?

Select 2 answers
A.Entity-relationship diagram
B.Snowflake schema
C.Star schema
D.Scatter plot
E.Histogram
AnswersB, C

Snowflake schema normalises dimension tables into multiple related tables, reducing redundancy and storage at the cost of more joins. It is a recognised dimensional modelling technique, alongside star schema, used to structure data warehouse dimensions.

Why this answer

Snowflake schema (B) is a dimensional modeling technique in which dimension tables are normalized into multiple related tables, reducing redundancy while preserving the fact-dimension structure of a data warehouse. Star schema (C) is the classic dimensional modeling technique where a central fact table joins directly to denormalized dimension tables, optimizing query performance for analytical workloads. Both are standard approaches described by Kimball for organizing data marts and warehouses around facts and dimensions.

Entity-relationship diagram (A) is a conceptual/logical modeling notation for OLTP-style normalized databases, not a dimensional technique. Scatter plot (D) and histogram (E) are data visualization or statistical analysis tools, not data warehouse modeling techniques.

Exam trap

The trap here is that candidates may confuse general data modeling concepts (like ERDs) or data visualization tools (like scatter plots and histograms) with specific dimensional modeling techniques used in data warehouses.

109
MCQhard

A data scientist trains a regression model and observes high variance with low bias. Which technique is most appropriate to reduce variance?

A.Apply Ridge regularization
B.Increase polynomial features
C.Use a smaller training set
D.Remove correlated features
AnswerA

High variance with low bias indicates overfitting, where coefficients fit training noise. Ridge adds an L2 penalty on squared coefficient magnitudes, shrinking them and constraining model complexity, which reduces variance at the cost of slightly increased bias.

Why this answer

Ridge regularization (L2) reduces variance by adding a penalty term proportional to the square of the coefficients, which shrinks them toward zero without eliminating them. This directly addresses high variance (overfitting) by constraining the model's complexity, while low bias indicates the model fits the training data well. The regularization parameter λ controls the trade-off between bias and variance.

Exam trap

CompTIA often tests the misconception that reducing variance requires removing features or simplifying the model, but Ridge regularization is the correct technique because it penalizes coefficient magnitude without discarding predictors.

How to eliminate wrong answers

Option B is wrong because increasing polynomial features adds higher-order terms, which increases model complexity and typically increases variance, not reduces it. Option C is wrong because using a smaller training set reduces the amount of data available for learning, which generally increases variance due to less stable coefficient estimates. Option D is wrong because removing correlated features can reduce multicollinearity but does not directly penalize coefficient magnitudes; it may even increase variance if important predictors are dropped.

110
Multi-Selecthard

A logistic regression model predicts customer churn (0=no churn, 1=churn). The model outputs probabilities. Which THREE of the following statements about logistic regression are correct?

Select 3 answers
A.The model output is a probability between 0 and 1.
B.The coefficient of determination R² is used to assess model fit.
C.The coefficients represent the change in log-odds for a one-unit change in the predictor.
D.Logistic regression is used for binary classification.
E.The model uses the linear regression equation y = mx + b directly.
AnswersA, C, D

The sigmoid function maps any linear combination of predictors onto the interval between 0 and 1, so each output is interpretable as the estimated probability of churn. A threshold, commonly 0.5, then converts that probability into a class label.

Why this answer

Option A is correct because logistic regression applies the sigmoid (logistic) function to a linear combination of predictors, producing an output strictly between 0 and 1 that can be interpreted as the probability of the positive class (churn = 1). Option C is correct because the model is linear in the log-odds: each coefficient βj gives the change in log-odds of the outcome for a one-unit increase in predictor Xj, holding other predictors constant (equivalently, e^βj is the odds ratio). Option D is correct because logistic regression is specifically designed for binary classification, here distinguishing churn (1) from no churn (0) by thresholding the predicted probability.

Option B is not correct because R² (coefficient of determination) is a goodness-of-fit measure for ordinary least squares linear regression, not for logistic regression, which instead uses measures like log-likelihood, deviance, AIC/BIC, or pseudo-R². Option E is not correct because logistic regression does not use the linear equation y = mx + b directly; it models the log-odds as a linear function and then applies the logistic function to obtain probabilities, rather than predicting y linearly.

Exam trap

DA0-002 often tests logistic regression by mixing in linear regression concepts — candidates who assume R² applies or that the linear equation is used directly pick the wrong statements, missing that logistic regression uses log-odds and the sigmoid function.

111
MCQeasy

A data analyst is creating a visualization to show the proportion of total sales contributed by each of five product categories. The analyst wants to emphasize the relative size of each category as part of a whole. Which type of chart is most appropriate?

A.Histogram
B.Line chart
C.Scatter plot
D.Pie chart
AnswerD

A pie chart is designed to show parts of a whole, with each slice representing a proportion of the total. It is ideal for displaying the relative contribution of a small number of categories, such as five product categories, to a total. It visually emphasizes the size of each category in relation to the whole.

Why this answer

A pie chart is the most suitable because it visually represents each category as a slice of the whole, making it easy to compare proportions. Line charts, scatter plots, and histograms serve different purposes and do not effectively communicate part-to-whole relationships.

Exam trap

The trap here is confusing part-to-whole visualization with other chart types; pie charts are specifically designed for proportions, but they are often misused for too many categories.

112
Multi-Selecthard

A data analyst is performing a chi-square test for independence between two categorical variables. Which THREE of the following are necessary conditions for the test to be valid?

Select 3 answers
A.Variances are equal across groups
B.Data is normally distributed
C.Sample is randomly selected
D.Observations are independent
E.Expected frequency in each cell is at least 5
AnswersC, D, E

Random selection ensures the sample represents the population, so observed cell counts estimate population proportions without selection bias. This satisfies the chi-square validity condition that expected frequencies reflect genuine population distributions rather than a biased subset.

Why this answer

Option C is correct because a chi-square test for independence requires that the sample be randomly selected from the population, ensuring the results can be generalized and that the expected counts reflect the underlying distribution. Option D is correct because the test assumes observations are independent; each subject or case must contribute to only one cell of the contingency table, and correlated or repeated observations violate the chi-square model. Option E is correct because the chi-square approximation is valid only when the expected frequency in each cell is at least 5 (or, in larger tables, when no more than 20% of cells have expected counts below 5 and none below 1).

Option A is not required because chi-square is a nonparametric test of frequencies and does not assume homogeneity of variances. Option B is not required because chi-square does not assume normally distributed data; it operates on counts of categorical outcomes.

Exam trap

DA0-002 often tests the confusion between parametric assumptions (normality, equal variance) and chi-square's non-parametric requirements — candidates incorrectly apply t-test or ANOVA assumptions to chi-square.

113
MCQeasy

A retail company wants to analyze monthly sales data over the past three years to identify long-term trends. Which component of time series analysis is most relevant for this goal?

A.Irregular component
B.Cyclical component
C.Seasonality
D.Trend
AnswerD

Trend captures the long-term direction of a series after removing seasonal and irregular fluctuations, which directly matches the three-year monthly sales goal of identifying sustained movement rather than repeating yearly patterns. Decomposition isolates this component, so analysts can quantify whether sales are genuinely rising or falling across the full period.

Why this answer

The trend component of a time series represents the long-term direction or underlying movement of the data over an extended period, which is exactly what the retail company wants to identify across three years of monthly sales. Trend captures gradual increases, decreases, or stagnation that persist beyond short-term fluctuations. Seasonality and cyclical components are periodic but shorter-term or irregular in period, while the irregular component is random noise.

Exam trap

DA0-002 often tests the confusion between trend and cyclical components — candidates pick 'cyclical' because it sounds long-term, but trend specifically refers to the persistent direction, while cycles are irregular economic waves.

How to eliminate wrong answers

Option A is wrong because the irregular component represents random, unpredictable fluctuations (e.g., a one-off supply chain disruption) that obscure rather than reveal long-term patterns. Option B is wrong because the cyclical component refers to wave-like movements tied to economic cycles (typically 2-10 years) that are not fixed in period — while relevant to long-term analysis, the question emphasizes identifying the underlying direction, which is trend. Option C is wrong because seasonality captures repeating patterns within a fixed period (e.g., monthly or quarterly), which is short-term and would mask the long-term trend if not decomposed.

114
Multi-Selecthard

A company is planning an A/B test to compare two website designs. Which THREE of the following must be determined before the test begins to ensure valid results? (Select three.)

Select 3 answers
A.The desired effect size
B.The p-value of the test
C.Which hypothesis is true
D.The minimum sample size required
E.The significance level (α)
AnswersA, D, E

The desired effect size defines the smallest lift worth detecting, which drives sample-size and power calculations. Specifying it before testing satisfies the stem's validity requirement, preventing an underpowered test or a post-hoc redefinition of success.

Why this answer

Option A (the desired effect size) must be determined beforehand because it defines the minimum difference between the two designs that the test should be able to detect, and it directly drives the sample-size calculation; without it, the test may be underpowered or impractically large. Option D (the minimum sample size required) is essential because it ensures the test has adequate statistical power to detect the specified effect, preventing premature conclusions from too little data. Option E (the significance level, α) must be set in advance because it defines the threshold for rejecting the null hypothesis (commonly 0.05) and controls the false-positive rate, keeping the analysis objective.

Option B (the p-value) is not determined before the test; it is an output computed from the collected data. Option C (which hypothesis is true) cannot be known in advance—it is precisely what the A/B test is designed to infer, so it is not a pre-test determination.

Exam trap

DA0-002 often tests the confusion between pre-test design parameters (effect size, sample size, α) and post-test outputs (p-value, hypothesis truth).

115
MCQeasy

A marketing analyst wants to segment customers based on their purchase history, including total spent, number of transactions, and average order value. The analyst runs k-means clustering with k=5 on the raw data but notices that the cluster assignments change significantly every time the algorithm is executed. What should the analyst do first to obtain consistent and meaningful clusters?

A.Normalize the features and set a fixed random seed for the initial centroids.
B.Switch to hierarchical clustering, which does not require specifying k.
C.Increase the number of clusters to k=10 to capture more detail.
D.Use principal component analysis (PCA) to reduce the number of features to two.
AnswerA

Features on different scales distort Euclidean distance, so k-means centroids shift with each random initialisation. Normalising the features and fixing the random seed stabilises initial centroid selection, satisfying the requirement for consistent, meaningful cluster assignments across runs.

Why this answer

The instability in cluster assignments is caused by the algorithm's sensitivity to the scale of features and the random initialization of centroids. Normalizing the features ensures that each variable contributes equally to the distance calculations, while setting a fixed random seed makes the initial centroid selection deterministic, leading to reproducible results.

Exam trap

The trap here is that candidates may think the instability is due to the choice of k or the algorithm itself, rather than recognizing that k-means is sensitive to feature scaling and random initialization, which are the first things to address for consistency.

How to eliminate wrong answers

Option B is wrong because hierarchical clustering does not require specifying k, but it still suffers from sensitivity to data scaling and does not address the core issue of random initialization causing variability. Option C is wrong because increasing k to 10 would likely increase instability and overfit noise, not resolve the fundamental problem of non-deterministic centroids. Option D is wrong because PCA reduces dimensionality but does not stabilize the k-means algorithm; the cluster assignments would still vary with different random seeds unless combined with normalization and a fixed seed.

116
MCQeasy

Refer to the exhibit. Which clause is used to aggregate the data by department?

A.HAVING
B.WHERE
C.ORDER BY
D.GROUP BY
AnswerD

GROUP BY partitions rows sharing the same department value into groups, enabling aggregate functions such as COUNT, SUM or AVG to be computed per department. It is the clause that performs the grouping the question requires.

Why this answer

The GROUP BY clause is used to aggregate data by department because it groups rows that have the same values in the specified column(s), allowing aggregate functions like SUM, AVG, or COUNT to be applied per group. In SQL, without GROUP BY, aggregate functions would operate on the entire result set, not per department.

Exam trap

CompTIA often tests the distinction between WHERE (row-level filter) and HAVING (group-level filter), leading candidates to confuse HAVING with GROUP BY when the question asks for the clause that performs aggregation.

How to eliminate wrong answers

Option A is wrong because HAVING is used to filter groups after aggregation, not to define the grouping itself. Option B is wrong because WHERE filters individual rows before aggregation and cannot group data by department. Option C is wrong because ORDER BY sorts the result set but does not perform any aggregation or grouping.

117
MCQeasy

You are a data analyst at a logistics company. The operations manager wants to reduce delivery delays. You have historical data including order date, delivery date, distance, weather conditions, and driver ID. Initial analysis shows that the average delivery time has increased over the past six months. You suspect that weather is a contributing factor, but you need to confirm. The company also wants to build a model to predict delivery times to better manage customer expectations. The data contains missing values for weather conditions in about 10% of records, and some driver IDs are incorrect. You have limited time and resources. What should you do first?

A.Immediately focus on time series analysis to look for patterns
B.Start by cleaning the data: correct driver IDs and decide how to handle missing weather data, then perform exploratory data analysis
C.Collect more data to fill missing values
D.Build a predictive model using all available data after imputing missing weather data
AnswerB

Cleaning the incorrect driver IDs and deciding a treatment for the 10% missing weather values must precede any analysis, because exploratory data analysis and the predictive model both depend on accurate, complete inputs; with limited time, this directly satisfies the stem's constraint of confirming weather's role reliably.

Why this answer

Data cleaning and exploratory data analysis (EDA) are foundational steps before any modeling or time series work. With missing weather data (10%) and incorrect driver IDs, proceeding without cleaning would introduce bias and errors. EDA will reveal patterns, correlations, and data quality issues, enabling informed decisions on imputation and feature engineering for the predictive model.

Exam trap

CompTIA often tests the misconception that you can jump directly to modeling or advanced analysis without first ensuring data quality, ignoring the 'garbage in, garbage out' principle.

How to eliminate wrong answers

Option A is wrong because time series analysis assumes clean, consistent data; applying it directly with missing values and incorrect IDs would yield unreliable patterns and waste resources. Option C is wrong because collecting more data is time-consuming and does not address the existing incorrect driver IDs or the need to understand current data quality; it also assumes missing values are random, which may not hold. Option D is wrong because building a predictive model on uncleaned data with imputed weather values without prior EDA risks overfitting, misinterpretation of feature importance, and propagation of errors from incorrect IDs.

118
MCQhard

A data analyst is comparing the means of two independent groups using a t-test. The sample sizes are small and the data is not normally distributed. Which condition is violated for a valid t-test?

A.Normality
B.Equal variances
C.Independence of observations
D.Sample size larger than 30
AnswerA

The t-test assumes the sampling distribution of the mean is normal. With small samples, the central limit theorem does not apply, so non-normal data violates the normality assumption. Independence and equal variances are separate assumptions; normality is the condition breached here.

Why this answer

A t-test assumes that the sampling distribution of the mean is approximately normal. With small sample sizes, the Central Limit Theorem does not guarantee normality, so the data itself should be approximately normal. Since the data is not normally distributed and sample sizes are small, the normality assumption is violated, making the t-test invalid.

Exam trap

DA0-002 often tests the confusion between the assumptions of a t-test, leading candidates to select 'equal variances' or 'sample size larger than 30' when the actual violated condition is normality due to small, non-normal samples.

How to eliminate wrong answers

Option B is wrong because equal variances is a separate assumption (homogeneity of variance) that can be relaxed with Welch's t-test; it is not the condition violated here. Option C is wrong because independence of observations is an assumption, but the question states the groups are independent, so it is not violated. Option D is wrong because a sample size larger than 30 is not a requirement for a t-test; it is a rule of thumb for invoking the Central Limit Theorem, and the question says sample sizes are small, so this is not a violated condition but rather a reason normality matters.

119
MCQeasy

A data analyst needs to summarize customer satisfaction scores. The data contains a few extremely low scores that skew the distribution. Which measure of central tendency is most appropriate?

A.Range
B.Mode
C.Median
D.Mean
AnswerC

Extremely low scores pull the mean downward, so it no longer represents typical satisfaction. The median resists this skew because it depends only on positional rank, not magnitude, satisfying the need to summarise a distribution distorted by outliers.

Why this answer

The median is the most appropriate measure of central tendency when data contains extreme outliers, such as the very low customer satisfaction scores described. Unlike the mean, the median is resistant to skew because it depends only on the middle value(s) of the sorted dataset, not on the magnitude of extreme values. This makes it the standard choice for summarizing ordinal or skewed interval/ratio data in data analysis.

Exam trap

The trap here is that candidates often default to the mean as the 'average' without considering outlier impact, but CompTIA Data+ tests the understanding that the mean is non-robust and the median is the correct choice for skewed data in the Analyzing and Modeling domain.

How to eliminate wrong answers

Option A (Range) is wrong because it is a measure of dispersion (the difference between the maximum and minimum values), not a measure of central tendency, and it is heavily influenced by outliers. Option B (Mode) is wrong because it identifies the most frequently occurring score, which may not represent the center of the distribution and can be misleading when outliers are present but not frequent. Option D (Mean) is wrong because it is sensitive to extreme values; the few extremely low scores will pull the arithmetic mean downward, misrepresenting the typical customer satisfaction experience.

120
MCQhard

A data analyst is building a model to predict customer churn. The dataset has 10,000 records with 500 churned customers. The model predicts churn with 95% accuracy, but only identifies 10% of actual churners. Which metric best highlights this issue?

A.Accuracy
B.F1 score
C.Recall
D.Precision
AnswerC

Recall measures the proportion of actual churners correctly identified, so 10% recall exposes the model's failure to catch churn despite 95% accuracy. Accuracy is misleading here because the 500 churners are a small minority of 10,000 records.

Why this answer

Recall (also known as sensitivity or true positive rate) measures the proportion of actual positives correctly identified. With only 10% of actual churners detected, the model has a recall of 0.1, which directly highlights the failure to capture churners despite high overall accuracy.

Exam trap

The trap here is that candidates may choose accuracy because it is a familiar and seemingly high value (95%), failing to recognize that in imbalanced datasets, accuracy can be deceptive and does not reflect poor performance on the minority class.

How to eliminate wrong answers

Option A is wrong because accuracy (95%) is misleading in imbalanced datasets; it can be high even if the model fails to detect churners, as the majority class (non-churners) dominates. Option B is wrong because the F1 score is the harmonic mean of precision and recall; while it would be low here, it does not directly isolate the issue of missing churners—recall is the metric that specifically measures detection of the positive class. Option D is wrong because precision measures the proportion of predicted churners that are actual churners; it does not reflect how many actual churners were missed, which is the core problem.

121
MCQeasy

In an A/B test, the null hypothesis states that there is no difference between the control and treatment groups. After running the test, the p-value is 0.04. Assuming α = 0.05, what is the correct conclusion?

A.Fail to reject the null hypothesis
B.Reject the null hypothesis
C.Accept the null hypothesis
D.The test is invalid because the p-value is too low
AnswerB

The p-value of 0.04 falls below the significance level of 0.05, so the observed difference is statistically significant. The null hypothesis of no difference between control and treatment is rejected, supporting the conclusion that the treatment had an effect.

Why this answer

The p-value of 0.04 is less than the significance level α = 0.05, so we reject the null hypothesis. This means there is statistically significant evidence to suggest a difference between the control and treatment groups. The correct conclusion is to reject the null hypothesis.

Exam trap

The trap is misinterpreting the p-value as the probability that the null hypothesis is true, or thinking that a low p-value means the test is invalid; candidates might also incorrectly choose 'accept the null' instead of 'fail to reject'.

How to eliminate wrong answers

Option A is wrong because failing to reject would occur if the p-value were greater than α. Option C is wrong because we never 'accept' the null hypothesis; we only fail to reject it. Option D is wrong because a low p-value does not invalidate the test; it indicates strong evidence against the null hypothesis.

122
MCQhard

After building a binary classification model, the data analyst obtains the following confusion matrix: True Positives=80, True Negatives=100, False Positives=20, False Negatives=30. What is the F1 score?

A.0.76
B.0.73
C.0.80
D.0.69
AnswerA

F1 balances precision and recall via their harmonic mean. Precision is 80/100 = 0.80; recall is 80/110 ≈ 0.727. F1 = 2 × (0.80 × 0.727) / (0.80 + 0.727) ≈ 0.762, which rounds to 0.76.

Why this answer

The F1 score is the harmonic mean of precision and recall. Precision = TP/(TP+FP) = 80/(80+20) = 0.80. Recall = TP/(TP+FN) = 80/(80+30) ≈ 0.7273.

F1 = 2 * (0.80 * 0.7273) / (0.80 + 0.7273) ≈ 0.7619, which rounds to 0.76. Option A is correct.

Exam trap

CompTIA often tests the distinction between precision, recall, and F1, and the trap here is that candidates mistakenly use accuracy or a simple average instead of the harmonic mean, or they confuse recall with F1.

How to eliminate wrong answers

Option B (0.73) is wrong because it approximates recall (0.727) instead of computing the harmonic mean. Option C (0.80) is wrong because it uses precision alone, ignoring recall. Option D (0.69) is wrong because it likely results from a miscalculation, such as averaging precision and recall arithmetically (0.80+0.727)/2 ≈ 0.76, not 0.69, or from an incorrect formula like (TP+TN)/(TP+TN+FP+FN) = 180/230 ≈ 0.78, which is accuracy, not F1.

123
Multi-Selecteasy

A data analyst is preparing to build a predictive model. Which TWO steps are essential to ensure model validity? (Choose two.)

Select 2 answers
A.Increase model complexity
B.Perform cross-validation
C.Avoid feature selection
D.Use the entire dataset for training
E.Split data into training and testing sets
AnswersB, E

Cross-validation partitions the dataset into complementary training and validation folds, so performance is estimated on data the model has not seen. This directly satisfies the stem's validity requirement by detecting overfitting and yielding a generalisable accuracy estimate rather than an optimistic fit to the training set alone.

Why this answer

Option B (Perform cross-validation) is correct because cross-validation, such as k-fold or stratified k-fold, partitions the data into multiple train/validation folds to estimate how well the model generalizes and to detect overfitting, which is essential for establishing model validity. Option E (Split data into training and testing sets) is correct because holding out an independent test set ensures the model is evaluated on data it has never seen, giving an unbiased estimate of predictive performance and guarding against data leakage. Option A is not correct because increasing model complexity can cause overfitting and does not by itself ensure validity.

Option C is not correct because skipping feature selection can introduce irrelevant or noisy variables that degrade model performance. Option D is not correct because training on the entire dataset leaves no independent data for evaluation, making it impossible to assess generalization.

Exam trap

The trap here is that candidates may think using the entire dataset for training (Option D) is acceptable because it maximizes data for learning, but they overlook the necessity of a separate testing set to validate model performance and avoid overfitting.

124
Multi-Selectmedium

A data analyst is preparing a dataset for analysis and needs to address data quality issues. Which TWO of the following are common data cleaning tasks?

Select 2 answers
A.Performing hypothesis testing
B.Imputing missing values
C.Building a regression model
D.Calculating correlation coefficients
E.Deduplicating records
AnswersB, E

Imputing missing values replaces nulls with substituted estimates such as mean, median or model-predicted figures, satisfying the stem's data quality remediation goal. It preserves row counts for analysis rather than discarding incomplete records, directly addressing missingness as a cleaning task.

Why this answer

Option B (Imputing missing values) is correct because missing data is a classic data quality problem, and imputation—filling gaps using methods like mean, median, mode, or model-based estimates—is a standard data cleaning step that makes the dataset complete and usable for analysis. Option E (Deduplicating records) is correct because duplicate rows or records distort counts, aggregates, and model results, so identifying and removing or merging duplicates is a core data cleaning task. The unmarked options do not belong because hypothesis testing (A), building a regression model (C), and calculating correlation coefficients (D) are all downstream analytical or statistical modeling activities performed on already-cleaned data, not cleaning operations themselves.

Exam trap

The trap is mixing analysis activities (hypothesis testing, regression, correlation) with cleaning activities — candidates must distinguish preprocessing/data-quality remediation from downstream statistical modeling.

125
MCQeasy

A marketing team wants to segment customers into distinct groups based on purchasing behavior. The data includes numeric features such as frequency, monetary value, and recency. Which unsupervised learning algorithm should be used?

A.Decision tree
B.K-means clustering
C.Linear regression
D.Association rules
AnswerB

K-means clustering partitions numeric feature space into k groups by minimising within-cluster variance, directly satisfying the requirement to segment customers on frequency, monetary value and recency. It handles continuous, unlabelled data without predefined categories, making it appropriate for this unsupervised behavioural segmentation task.

Why this answer

K-means clustering is the correct choice because it is an unsupervised learning algorithm that partitions data into K distinct clusters based on feature similarity. For segmenting customers by purchasing behavior (frequency, monetary value, recency), K-means groups customers with similar numeric patterns without requiring labeled outcomes, making it ideal for exploratory segmentation.

Exam trap

The trap here is that candidates may confuse unsupervised clustering (K-means) with supervised classification (decision tree) or regression (linear regression), mistakenly thinking any algorithm that 'groups' data must be supervised, or that association rules are for segmentation rather than transaction pattern mining.

How to eliminate wrong answers

Option A is wrong because a decision tree is a supervised learning algorithm used for classification or regression, requiring labeled target variables, not for unsupervised segmentation of unlabeled customer data. Option C is wrong because linear regression is a supervised learning algorithm that models the relationship between independent and dependent variables, predicting a continuous output, not for discovering hidden groups in unlabeled data. Option D is wrong because association rules are used for market basket analysis to find frequent itemsets and co-occurrence patterns (e.g., products bought together), not for clustering customers into distinct groups based on numeric features.

126
Multi-Selectmedium

A data analyst is cleaning a customer dataset. Which two actions are appropriate for handling duplicate records? (Choose TWO)

Select 2 answers
A.Impute missing values with mean
B.Delete any row with a duplicate email address
C.Remove all rows with identical values in every field
D.Apply Z-score standardization
E.Use a fuzzy matching algorithm to identify near-duplicates
AnswersC, E

Exact-row deduplication removes records where every field matches, eliminating true duplicates while preserving legitimate distinct customers who happen to share some values. This directly satisfies the cleaning goal, since identical rows across all fields carry no additional information and inflate counts and aggregates.

Why this answer

Option C is correct because removing rows whose values are identical across every field eliminates exact duplicate records, which is a standard and safe deduplication step in data cleaning. Option E is correct because fuzzy matching algorithms (e.g., Levenshtein distance or Jaro-Winkler similarity) identify near-duplicates that differ slightly due to typos, formatting, or abbreviations, allowing the analyst to review and consolidate them. Option A is incorrect because mean imputation addresses missing values, not duplicate records.

Option B is incorrect because deleting every row with a duplicate email address is overly aggressive and may remove legitimate distinct customers who share an email. Option D is incorrect because Z-score standardization is a scaling technique for numeric features and does not handle duplicates.

Exam trap

The trap here is confusing data-cleaning categories — candidates may pick mean imputation or Z-score standardization because they sound like 'cleaning' steps, but those address missing values and scaling, not duplicates.

127
MCQmedium

A data analyst at a hospital network is exploring the relationship between patient age and length of stay for 400 discharged patients. A scatterplot shows a roughly linear upward trend, but the analyst wants a single number summarizing both the strength and direction of the association before reporting to clinicians. Which measure should the analyst calculate?

A.The Pearson correlation coefficient between age and length of stay.
B.The coefficient of determination from a model predicting age from length of stay.
C.The chi-square statistic from a contingency table of age and length of stay.
D.The covariance between age and length of stay.
AnswerA

The Pearson correlation coefficient is a unitless value between -1 and 1 that captures both the direction and the strength of a linear association. With a roughly linear scatterplot, it directly answers the analyst's question and is easily communicated to clinicians. Its sign reveals whether older patients tend to stay longer, and its magnitude indicates how tightly the points follow a line.

Why this answer

For two continuous variables with an approximately linear relationship, the Pearson correlation coefficient is the standard single-number summary: it is bounded between -1 and 1, carries direction through its sign, and reflects strength through its magnitude. Covariance is scale-dependent, while R-squared and chi-square discard direction or require arbitrary categorization, so none of them answers the clinician-facing question as directly.

Exam trap

The trap here is confusing covariance with correlation; covariance signals direction but its magnitude changes with measurement units, so it cannot be read as a strength score.

128
MCQmedium

A data analyst is testing whether the average sales amount differs between two regions. Which statistical test is most appropriate?

A.Chi-square test
B.ANOVA
C.Two-sample t-test
D.Paired t-test
AnswerC

A two-sample t-test compares the means of two independent groups, matching the two regions and continuous sales measure. It tests whether the difference in average sales is statistically significant, unlike a paired test, which requires matched observations.

Why this answer

A two-sample t-test compares the means of two independent groups.

129
MCQmedium

A data scientist builds a simple linear regression model to predict house prices based on square footage. The model yields an R-squared value of 0.85. Which statement accurately interprets this result?

A.The slope of the regression line is 0.85
B.85% of the data points lie exactly on the regression line
C.The model explains 85% of the variability in house prices
D.There is a 85% chance that square footage causes higher prices
AnswerC

R-squared measures the proportion of variance in the dependent variable explained by the model. A value of 0.85 means square footage accounts for 85% of house price variability, directly satisfying the stem's constraint of interpreting the reported R-squared value.

Why this answer

R-squared (the coefficient of determination) measures the proportion of variance in the dependent variable explained by the independent variable(s). An R² of 0.85 means 85% of the variability in house prices is accounted for by the square footage in this linear model. It is a goodness-of-fit measure, not a probability, slope, or count of points on the line.

Exam trap

The trap is treating R² as a probability or a count of points on the line — candidates often misread it as '85% chance' or '85% of points fit exactly,' when it strictly measures explained variance.

How to eliminate wrong answers

Option A is wrong because 0.85 is R², not the slope (β₁); the slope is a separate regression coefficient with its own units (price per square foot). Option B is wrong because R² does not measure how many points lie exactly on the line — in real data almost none do; it measures explained variance, not point-on-line counts. Option D is wrong because R² is not a probability and regression does not establish causation — correlation between square footage and price does not prove square footage causes higher prices.

130
Multi-Selecthard

A data analyst is performing a chi-square test of independence on a contingency table of customer satisfaction (satisfied, neutral, dissatisfied) by region (North, South, East, West). Which THREE of the following are necessary assumptions for the test?

Select 3 answers
A.The two variables are categorical
B.The sample size is greater than 30
C.Expected frequencies in each cell are at least 5 (or most cells)
D.The observations are independent
E.The data must be normally distributed
AnswersA, C, D

Chi-square of independence compares observed against expected counts across categories, so both variables must be categorical. Satisfaction and region are nominal groupings, not continuous measurements; treating them as such would violate the test's foundation and invalidate the computed statistic.

Why this answer

Option A is correct because the chi-square test of independence requires both variables to be categorical (nominal or ordinal), which holds here since satisfaction level and region are categorical variables. Option C is correct because the test relies on the chi-square approximation, which is valid when expected cell frequencies are sufficiently large—typically at least 5 in each cell, or in most cells (with no cell below 1) for larger tables. Option D is correct because each observation must be independent, meaning each respondent contributes to only one cell of the contingency table, with no repeated or paired measurements.

Option B is incorrect because there is no fixed sample-size threshold of 30; adequacy is judged by expected frequencies, not raw n. Option E is incorrect because chi-square tests make no normality assumption—they operate on counts of categorical data, not continuous normally distributed variables.

Exam trap

The trap is importing parametric assumptions (normality, n > 30) from t-tests and ANOVA into chi-square, which is a non-parametric test concerned only with categorical counts and expected frequencies.

131
MCQeasy

A dataset contains customer records with a column for 'Phone Number' that should be unique. However, the analyst finds several duplicate phone numbers. Which data quality dimension is primarily affected?

A.Completeness
B.Accuracy
C.Uniqueness
D.Consistency
AnswerC

Duplicate phone numbers violate the requirement that each value in the column be distinct, which is precisely what the uniqueness dimension measures. Completeness, accuracy and consistency concern missing, wrong or conflicting values, not repeated ones.

Why this answer

Uniqueness measures whether each real-world entity appears exactly once in the dataset. Since 'Phone Number' is intended to be a unique identifier per customer, finding duplicate values directly violates that expectation, so the affected dimension is uniqueness. Completeness, accuracy, and consistency describe different properties (missing values, correctness, and uniformity across sources) and are not what duplicate keys violate.

Exam trap

The trap here is conflating uniqueness with accuracy or consistency — candidates see 'duplicate' and think 'wrong data,' but duplicates are a cardinality problem, not a correctness problem.

How to eliminate wrong answers

Option A is wrong because completeness concerns missing or null values, not repeated values — a duplicated phone number is present, not absent. Option B is wrong because accuracy concerns whether a value correctly reflects reality (e.g., a wrong digit), whereas duplicates can each be individually accurate. Option D is wrong because consistency concerns agreement of the same data across systems or formats, not the cardinality of values within one column.

132
MCQmedium

A marketing team runs an A/B test on email subject lines. The p-value is 0.03 with α = 0.05. Which of the following is the correct interpretation?

A.The result is not statistically significant at the 95% confidence level.
B.The probability that the null hypothesis is true is 3%.
C.Fail to reject the null hypothesis; no significant difference.
D.Reject the null hypothesis; there is a statistically significant difference.
AnswerD

With α = 0.05, a p-value of 0.03 falls below the significance threshold, so the null hypothesis is rejected. The result is statistically significant, meaning the observed difference in subject-line performance is unlikely to have arisen from chance alone.

Why this answer

With a p-value of 0.03 and α = 0.05, the p-value is less than the significance level, so we reject the null hypothesis. This indicates a statistically significant difference at the 95% confidence level. The p-value is not the probability that the null hypothesis is true.

Exam trap

DA0-002 often tests the interpretation of p-values, and candidates frequently misinterpret the p-value as the probability that the null hypothesis is true or confuse significance with practical importance.

How to eliminate wrong answers

Option A is wrong because a p-value of 0.03 is less than 0.05, so the result is statistically significant. Option B is wrong because the p-value is not the probability that the null hypothesis is true; it is the probability of observing the data (or more extreme) assuming the null hypothesis is true. Option C is wrong because we reject, not fail to reject, the null hypothesis when p < α.

133
MCQhard

In logistic regression, the output is a probability between 0 and 1. If the predicted probability for a customer churning is 0.7 and the decision threshold is 0.5, what is the predicted class?

A.Not churn (class 0)
B.Churn (class 1)
C.Both classes equally likely
D.Uncertain, need more data
AnswerB

Logistic regression applies a decision threshold to the predicted probability to assign a class. Since 0.7 exceeds the 0.5 threshold, the observation is classified as the positive outcome, churn (class 1). The 0.5 cutoff is the axis separating class 0 from class 1.

Why this answer

In logistic regression, the predicted class is determined by comparing the predicted probability to the decision threshold. Here the probability is 0.7 and the threshold is 0.5; since 0.7 ≥ 0.5, the model predicts the positive class, which is churn (class 1).

Exam trap

The trap is overcomplicating a simple threshold comparison — candidates sometimes think 0.7 is 'uncertain' or requires more data, forgetting that any probability above the 0.5 threshold maps deterministically to the positive class.

How to eliminate wrong answers

Option A is wrong because predicting 'not churn' would require the probability to be below the 0.5 threshold, but 0.7 exceeds it. Option C is wrong because 'both classes equally likely' corresponds to a probability of exactly 0.5, not 0.7. Option D is wrong because the decision rule is deterministic once the probability and threshold are known — no additional data is needed to assign the class.

134
Drag & Dropmedium

Drag and drop the steps to implement a data classification policy in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Classification involves defining levels, assigning ownership, labeling, access control, and training.

135
MCQhard

A data analyst is working with a dataset that includes a categorical variable 'product_category' with 50 unique values. The analyst wants to reduce dimensionality before clustering. Which technique should the analyst use?

A.Multiple correspondence analysis (MCA)
B.One-hot encoding
C.Factor analysis
D.Principal component analysis (PCA)
AnswerA

Multiple correspondence analysis is specifically designed to reduce dimensionality of categorical data by transforming categories into a lower-dimensional numerical space. It captures associations between categories and can handle variables with many levels. This makes it ideal for the analyst's goal of reducing the 50 product categories before clustering. MCA preserves the categorical structure while enabling the use of distance-based clustering algorithms.

Why this answer

Multiple correspondence analysis is a dimensionality reduction technique tailored for categorical data. It converts categories into numerical dimensions that capture the underlying structure, making it suitable for clustering. Unlike one-hot encoding, it reduces rather than expands the feature space.

PCA and factor analysis are designed for continuous data, so they are not directly applicable to a categorical variable with many levels.

Exam trap

The trap here is assuming PCA can be applied to any data after encoding, but PCA on one-hot encoded data may not effectively reduce dimensionality and can lose interpretability.

136
Multi-Selectmedium

Which TWO of the following are true about Pearson correlation coefficient (r)?

Select 2 answers
A.An r of 0 means no relationship exists
B.It ranges from 0 to 1
C.It measures the strength and direction of a linear relationship
D.A value of +1 indicates a perfect positive linear relationship
E.It can be used for categorical variables
AnswersC, D

Pearson's r quantifies both the strength and direction of a linear association between two continuous variables, ranging from -1 to +1. It captures linearity specifically, so non-linear relationships can yield r near zero despite a genuine association.

Why this answer

Option C is correct because the Pearson correlation coefficient r is specifically designed to quantify both the strength and the direction (positive or negative) of a linear association between two continuous variables. Option D is correct because r = +1 represents the maximum possible value, meaning the two variables lie exactly on a straight line with a positive slope, i.e., a perfect positive linear relationship. Option A is incorrect because r = 0 indicates no linear relationship, but a strong non-linear (e.g., curvilinear) relationship may still exist.

Option B is incorrect because r ranges from -1 to +1, not 0 to 1; negative values indicate an inverse linear relationship. Option E is incorrect because Pearson r requires continuous, interval- or ratio-scaled variables, whereas categorical variables call for other measures such as chi-square, Cramér's V, or point-biserial correlation.

Exam trap

DA0-002 often tests the misconception that r = 0 means 'no relationship' and that r only ranges from 0 to 1 — candidates forget the negative range and the linear-only scope of Pearson's r.

137
MCQmedium

A data analyst is analyzing customer purchase amounts. The dataset contains several extreme high values due to luxury purchases. Which measure of central tendency is most robust to these outliers?

A.Range
B.Mean
C.Mode
D.Median
AnswerD

The median depends solely on positional order, so extreme luxury purchase values shift it minimally, unlike the mean, which sums all values and is dragged upward by outliers. It therefore satisfies the requirement for a central tendency measure robust to extreme high values.

Why this answer

The median is the middle value when data is ordered, so extreme high values (luxury purchases) shift it only slightly — it is a robust measure of central tendency. Unlike the mean, which is pulled upward by outliers because it sums all values, the median depends only on rank position. This makes it the preferred measure for skewed distributions like income or purchase amounts.

Exam trap

The trap is defaulting to the mean as 'the average' out of habit — candidates must recognize that 'robust to outliers' specifically signals the median, not the arithmetic mean.

How to eliminate wrong answers

Option A is wrong because the range is a measure of dispersion (max − min), not central tendency, and it is extremely sensitive to outliers — a single luxury purchase inflates it dramatically. Option B is wrong because the mean is the least robust measure here; it incorporates every value, so extreme highs drag it upward and misrepresent the typical purchase. Option C is wrong because the mode identifies the most frequent value, which may not reflect the center at all in continuous data and is unaffected by but also unrepresentative of outliers.

138
MCQmedium

A data analyst at a logistics company is analyzing delivery times for three different shipping carriers. The analyst wants to determine whether the mean delivery time differs across carriers. The data are normally distributed, and the variances across carriers are assumed equal. Which statistical test should the analyst use?

A.Independent two-sample t-test
B.One-way ANOVA
C.Paired t-test
D.Chi-square test of independence
AnswerB

One-way ANOVA compares the means of three or more independent groups to determine if at least one group mean is significantly different. Here, the analyst has three carriers, normal data, and equal variances, which satisfies the assumptions for ANOVA. It is the appropriate test to assess whether mean delivery times differ across the carriers.

Why this answer

One-way ANOVA is designed to compare means across three or more independent groups under assumptions of normality and equal variances. The scenario involves three carriers, continuous delivery times, and the goal of detecting mean differences, making ANOVA the correct choice. Other tests either handle only two groups or require categorical data, which does not fit the analysis objective.

Exam trap

The trap here is assuming that an independent two-sample t-test can be extended to three groups without adjusting for multiple comparisons, which increases the risk of a false positive.

139
MCQmedium

A financial analyst wants to compare the mean annual returns of three different investment strategies. Which statistical test is most appropriate?

A.Chi-square test
B.Paired t-test
C.One-way ANOVA
D.Two-sample t-test
AnswerC

One-way ANOVA compares the means of three or more independent groups on a single continuous dependent variable, testing whether at least one strategy's mean annual return differs. It suits this scenario because the analyst has one factor with three levels.

Why this answer

ANOVA is used to compare means of three or more groups.

140
MCQeasy

During ETL, a data analyst discovers that a date column contains values like '01/02/2023' and '2023-01-02'. Which of the following is the best practice to ensure consistent date format before analysis?

A.Keep both formats and handle during analysis
B.Use regular expressions to parse and convert each format
C.Remove records with inconsistent date formats
D.Apply a standardized date parsing function to convert all dates
AnswerD

A standardised parsing function explicitly interprets each source pattern, then emits one canonical representation, satisfying the consistency constraint. Unlike locale-dependent casting, it resolves ambiguity between day-first and ISO 8601 input deterministically, so '01/02/2023' and '2023-01-02' become comparable values before analysis.

Why this answer

Applying a standardized date parsing function (e.g., `TO_DATE` in SQL or `pd.to_datetime` in Python) ensures all date values are converted to a single, consistent format regardless of the original representation. This is a fundamental ETL best practice to avoid ambiguity and enable accurate date-based filtering, aggregation, and joins during analysis.

Exam trap

The trap here is that candidates may choose Option B (regular expressions) thinking it offers fine-grained control, but they overlook that dedicated date parsing functions are more reliable, simpler, and handle edge cases like leap years or time zones that regex cannot easily manage.

How to eliminate wrong answers

Option A is wrong because keeping both formats forces the analyst to handle multiple date patterns during every query, increasing complexity and risk of errors in comparisons or calculations. Option B is wrong because using regular expressions to parse dates is fragile, error-prone, and unnecessary when dedicated date parsing functions exist that handle locale and format variations robustly. Option C is wrong because removing records with inconsistent date formats discards potentially valid data, leading to incomplete analysis and biased results.

141
MCQeasy

Which data cleaning method involves replacing a missing value with the average of the available values in that column?

A.Mean imputation
B.Interpolation
C.Listwise deletion
D.Forward-fill
AnswerA

Mean imputation calculates the arithmetic average of a column's non-missing entries and substitutes that value for each gap, satisfying the stem's requirement to replace missing values with the column average. It preserves the column's central tendency, though it shrinks variance and can distort relationships between variables.

Why this answer

Mean imputation is the data cleaning technique that replaces missing values in a column with the arithmetic mean (average) of the non-missing values in that same column. This preserves the column's central tendency and keeps the overall sample size intact, which is why it is the standard answer when the question specifies 'replacing a missing value with the average of the available values in that column.'

Exam trap

The trap here is confusing mean imputation with interpolation or forward-fill, since all three 'fill in' missing values — candidates must recognize that only mean imputation uses the column-wide average rather than neighboring values.

How to eliminate wrong answers

Option B is wrong because interpolation estimates missing values from surrounding data points (typically in time-series or ordered data) using linear or spline methods, not from the column-wide average. Option C is wrong because listwise deletion removes entire rows containing missing values rather than substituting them, reducing sample size. Option D is wrong because forward-fill propagates the last observed value forward to fill gaps, which is a carry-forward technique, not an averaging technique.

142
MCQmedium

A company has a dataset with 100 features. The data analyst wants to reduce dimensionality while preserving as much variance as possible. Which technique should be used?

A.PCA (Principal Component Analysis)
B.LDA (Linear Discriminant Analysis)
C.Autoencoders
D.t-SNE
AnswerA

Principal Component Analysis projects the 100 features onto orthogonal components ordered by explained variance, so retaining the leading components captures maximum variance in fewer dimensions. This directly satisfies the goal of reducing dimensionality while preserving as much variance as possible.

Why this answer

PCA is the correct choice because it is an unsupervised linear dimensionality reduction technique that projects the data onto orthogonal components ordered by the variance they capture. By selecting the top principal components, the analyst can retain the maximum possible variance in the dataset while reducing the number of features from 100 to a smaller set, directly addressing the goal of preserving variance.

Exam trap

The trap here is that candidates often confuse PCA with LDA because both are linear transformations, but LDA requires labeled data and maximizes class separation, not variance, making it unsuitable for this unsupervised variance-preservation goal.

How to eliminate wrong answers

Option B (LDA) is wrong because LDA is a supervised technique that maximizes class separability, not variance preservation, and requires labeled target classes, which are not mentioned in the scenario. Option C (Autoencoders) is wrong because while autoencoders can reduce dimensionality, they are neural-network-based, require significant tuning and data, and are not the standard first-choice technique for simple variance-preserving reduction; PCA is more straightforward and computationally efficient for this task. Option D (t-SNE) is wrong because t-SNE is a nonlinear visualization technique primarily used for exploring high-dimensional data in 2D or 3D plots; it does not preserve global variance structure and cannot be used to transform new data or reduce dimensionality for modeling.

143
MCQeasy

A data analyst calculates a correlation coefficient of -0.85 between temperature and heating costs. What does this indicate?

A.No correlation
B.Strong positive correlation
C.Strong negative correlation
D.Weak negative correlation
AnswerC

A coefficient of -0.85 indicates a strong inverse relationship: as temperature rises, heating costs tend to fall substantially. The magnitude (0.85) approaches -1, signalling a tight linear association, while the negative sign gives the direction. This satisfies the stem's requirement to interpret both strength and direction of the correlation.

Why this answer

A correlation coefficient of -0.85 indicates a strong negative linear relationship between temperature and heating costs. As temperature increases, heating costs decrease significantly, and the magnitude of 0.85 (close to -1) confirms the strength of this inverse association.

Exam trap

CompTIA often tests the misinterpretation of the sign of the correlation coefficient, where candidates confuse a strong negative correlation with a weak one or mistakenly think a negative value implies no relationship.

How to eliminate wrong answers

Option A is wrong because a correlation coefficient of -0.85 is far from 0, indicating a clear relationship, not no correlation. Option B is wrong because a positive correlation would have a coefficient greater than 0, but -0.85 is negative, showing an inverse relationship. Option D is wrong because a weak negative correlation would have a coefficient closer to 0 (e.g., -0.2 to -0.4), whereas -0.85 is near -1, indicating a strong negative correlation.

144
MCQhard

A data analyst is performing a multiple linear regression with three predictors. The model output shows an R-squared of 0.85 and an adjusted R-squared of 0.80. Which of the following is the best interpretation of the difference between these two values?

A.The model is overfitted, so all predictors should be removed
B.The model has high multicollinearity
C.The residuals are not normally distributed
D.One or more predictors may not be contributing meaningfully
AnswerD

Adjusted R-squared penalises each added predictor, so a drop from 0.85 to 0.80 signals that some of the three predictors add little explanatory power relative to the degrees of freedom they consume. This satisfies the stem's request to interpret the gap between the two metrics.

Why this answer

R-squared (0.85) measures how much variance the model explains, while adjusted R-squared (0.80) penalizes for the number of predictors and only increases if a new predictor improves the model more than chance would predict. A notable drop between the two (0.05) suggests that at least one predictor is not contributing meaningfully and may be adding noise rather than explanatory power.

Exam trap

DA0-002 often tests whether candidates understand that adjusted R-squared penalizes complexity, so the trap is picking 'overfitting' (A) when the correct interpretation is simply that some predictors lack meaningful contribution.

How to eliminate wrong answers

Option A is wrong because a 0.05 gap does not indicate severe overfitting requiring removal of all predictors — the model still explains 80% of adjusted variance, which is strong. Option B is wrong because multicollinearity is detected through variance inflation factors (VIF) or correlation matrices among predictors, not by comparing R-squared to adjusted R-squared. Option C is wrong because residual normality is assessed through Q-Q plots, histograms, or Shapiro-Wilk tests, not through the R-squared/adjusted R-squared gap.

145
MCQmedium

A marketing team uses K-means clustering to segment customers based on purchase history. To determine the optimal number of clusters, they plot the within-cluster sum of squares (WCSS) against k and look for an elbow. What is the purpose of this method?

A.To find the point where the rate of decrease in WCSS slows down
B.To identify the value of k that minimizes WCSS
C.To determine the initial centroids for the algorithm
D.To ensure all clusters have equal size
AnswerA

The elbow method identifies the k where WCSS reduction sharply decelerates, marking diminishing returns from adding clusters. Plotting WCSS against k, the inflection point balances model complexity against fit, so the marketing team selects the smallest k beyond which further segmentation yields negligible within-cluster variance improvement.

Why this answer

The elbow method helps choose k where adding more clusters yields diminishing returns in reducing variance.

146
MCQhard

After training a decision tree, the tree has depth 20 and 100% accuracy on training data but only 60% on test data. Which hyperparameter adjustment is most likely to improve generalization?

A.Increase number of estimators
B.Decrease minimum samples per split
C.Increase minimum samples per leaf
D.Increase maximum depth
AnswerC

Raising minimum samples per leaf prunes splits that isolate tiny, noisy subsets, directly countering the depth-20 overfitting that produces 100% training but 60% test accuracy. This pre-pruning constraint caps tree complexity, trading some training fit for better generalisation.

Why this answer

The model is overfitting: 100% training accuracy vs. 60% test accuracy with a depth-20 tree. Increasing minimum samples per leaf forces the tree to be simpler by requiring more samples in each leaf, reducing variance and improving generalization. This directly combats the overfitting caused by the overly deep tree.

Exam trap

The trap here is that candidates often confuse hyperparameters that reduce overfitting with those that increase model complexity, mistakenly choosing options like 'increase maximum depth' or 'decrease minimum samples per split' thinking they will improve accuracy.

How to eliminate wrong answers

Option A is wrong because increasing the number of estimators applies to ensemble methods like Random Forest or Gradient Boosting, not to a single decision tree; it would not affect this tree's overfitting. Option B is wrong because decreasing minimum samples per split allows the tree to split on smaller subsets, making it even more complex and worsening overfitting. Option D is wrong because increasing maximum depth would allow the tree to grow even deeper, exacerbating the overfitting problem rather than reducing it.

147
Multi-Selectmedium

A data analyst at a subscription media company is asked to determine whether a new recommendation engine increased average watch time. The analyst plans a two-sample t-test comparing a randomly assigned control group against a randomly assigned treatment group, with watch time recorded in minutes per user per week. Which TWO conditions must hold for the pooled-variance two-sample t-test to be statistically valid here? (Choose two.)

Select 2 answers
A.Both groups must contain at least 30,000 users to satisfy the central limit theorem.
B.Watch time must be measured on a ratio scale with a true zero and no upper bound.
C.The population variances must be equal, or the sample sizes must be large enough for the test's robustness.
D.The control and treatment groups must have identical means before the experiment begins.
E.The two samples must be independent of each other.
AnswersC, E

The pooled-variance t-test assumes equal population variances. When group sizes are similar and reasonably large, the test is fairly robust to modest variance differences. However, if variances are markedly unequal and group sizes differ greatly, the pooled estimate is biased and a Welch correction is preferable. Verifying this condition protects the validity of the reported significance level.

Why this answer

The pooled two-sample t-test rests on independent observations and an assumption about variance equality that is relaxed by large, balanced samples. Random assignment to control and treatment satisfies independence, and checking variance similarity or group balance safeguards the test's nominal error rate. Sample-size thresholds and measurement-scale restrictions are not assumptions of this procedure, so they do not belong in the analyst's validation checklist.

Exam trap

The trap here is assuming a fixed sample-size cutoff like 30 or 30,000 is itself a t-test assumption, rather than recognizing that independence and the variance condition are what actually govern validity.

148
MCQmedium

An analyst wants to compare the mean sales revenue across three different store regions. The data is normally distributed and variances are equal. Which statistical test is most appropriate?

A.Two-sample t-test
B.ANOVA
C.Paired t-test
D.Chi-square test
AnswerB

ANOVA compares means across three or more independent groups in a single test, keeping the Type I error rate controlled. With normal distributions and equal variances, the parametric F-test assumptions hold, making it appropriate for the three store regions.

Why this answer

ANOVA (Analysis of Variance) is the correct test for comparing means across three or more groups when data is normally distributed and variances are equal (homogeneity of variance). It tests the null hypothesis that all group means are equal using an F-statistic comparing between-group to within-group variance.

Exam trap

The trap is reaching for a t-test when comparing more than two groups — candidates must recognize that three or more independent group means require ANOVA, not repeated t-tests.

How to eliminate wrong answers

Option A is wrong because a two-sample t-test compares only two groups, not three regions — using it repeatedly inflates Type I error. Option C is wrong because a paired t-test requires matched/dependent samples (e.g., before-and-after on the same subjects), but the three regions are independent groups. Option D is wrong because chi-square tests association between categorical variables, not differences in means of a continuous variable like revenue.

149
Multi-Selectmedium

A researcher is designing an A/B test to compare two website layouts. Which TWO elements are essential for determining the required sample size?

Select 2 answers
A.Sample mean
B.Statistical power
C.Confidence interval width
D.Desired effect size
E.P-value
AnswersB, D

Statistical power, typically 80%, is the probability of detecting a real difference if one exists. It directly determines sample size: higher power requires more observations, satisfying the stem's requirement for an essential input alongside effect size and significance level.

Why this answer

Statistical power (B) is essential because it defines the probability of detecting a true effect when one exists, and standard sample-size formulas require a target power (typically 0.80) as an input. Desired effect size (D) is also essential because the minimum detectable or expected difference between the two layouts directly determines how many observations are needed—smaller effects require larger samples. Together with the significance level (alpha), power and effect size are the core parameters in sample-size calculations for A/B tests.

The sample mean (A) is an outcome estimated from data, not an input for planning sample size. Confidence interval width (C) is a result that depends on sample size rather than a prerequisite for computing it. The p-value (E) is also a post-hoc result of the test, not a design parameter used to determine required sample size.

Exam trap

DA0-002 often tests whether candidates confuse post-hoc statistics like p-value and sample mean with pre-experiment design inputs, so the trap is selecting observed outcomes instead of the design parameters power and effect size.

150
MCQhard

A data analyst is cleaning a dataset and finds that some records have duplicate entries based on customer ID. Which data quality dimension is most directly affected by these duplicates?

A.Timeliness
B.Consistency
C.Accuracy
D.Uniqueness
AnswerD

Duplicate records sharing the same customer ID violate the requirement that each entity appears once, which is precisely the uniqueness dimension. Completeness, accuracy, consistency and timeliness concern missing values, correctness, agreement across sources and currency respectively, so duplicates map directly to uniqueness.

Why this answer

Uniqueness is the data quality dimension that measures whether each real-world entity appears exactly once in a dataset. Duplicate records based on customer ID directly violate uniqueness because the same customer is represented multiple times. Timeliness, consistency, and accuracy describe other dimensions and are not the primary issue when duplicates exist.

Exam trap

DA0-002 often tests the overlap between accuracy and uniqueness, so candidates must recognize that duplicates are a uniqueness violation even when the duplicated values themselves are accurate.

How to eliminate wrong answers

Option A is wrong because timeliness concerns whether data is current and available when needed, not whether records are duplicated. Option B is wrong because consistency concerns whether values agree across systems or formats (e.g., 'NY' vs 'New York'), not whether rows are repeated. Option C is wrong because accuracy concerns whether values correctly reflect reality; a duplicate row can be perfectly accurate in content yet still violate uniqueness.

← PreviousPage 2 of 4 · 238 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Data Analysis questions.