Courseiva

CCNA Data Analysis Questions

75 of 238 questions · Page 3/4 · Data Analysis · Answers revealed

151
Multi-Selectmedium

A data analyst is building a supervised learning model to predict customer churn. The target variable is binary (churn = yes/no). Which TWO modeling techniques are appropriate for this task? (Select two.)

Select 2 answers
A.K-means clustering
B.Linear regression
C.Logistic regression
D.Decision trees
E.Apriori algorithm
AnswersC, D

Logistic regression models the probability of a binary outcome by applying the logistic (sigmoid) function to a linear combination of predictors, bounding output between 0 and 1. This directly suits the churn yes/no target, unlike ordinary linear regression.

Why this answer

Logistic regression (C) is correct because it is a supervised classification technique that models the probability of a binary outcome (churn yes/no) using a sigmoid function, making it a standard choice for binary targets. Decision trees (D) are also correct because they are supervised classifiers that recursively split features to predict categorical class labels, and they handle binary targets naturally while offering interpretability. K-means clustering (A) is wrong because it is an unsupervised algorithm that groups unlabeled data and cannot predict a labeled binary target.

Linear regression (B) is wrong because it predicts continuous numeric values rather than class probabilities or discrete categories, so it is unsuitable for binary classification. The Apriori algorithm (E) is wrong because it is an unsupervised association-rule mining method for finding frequent itemsets, not a predictive classification model.

Exam trap

The trap here is that candidates may confuse unsupervised clustering (K-means) or association rule mining (Apriori) with supervised classification, or mistakenly think linear regression can be adapted for binary outcomes without transformation.

152
MCQhard

A data analyst at an e-commerce company is profiling a transactions table with ten million rows. The 'order_total' column is stored as text and contains values like '$1,299.00', a few entries of 'N/A', and several negative amounts that represent refunds. The analyst must produce a numeric column suitable for summation in a revenue report. Which sequence of operations best prepares the column?

A.Strip the currency symbol and thousands separators, convert non-numeric entries to null or a flagged value, preserve negative signs for refunds, then cast to decimal.
B.Cast the column directly to a decimal type, letting the database coerce invalid strings to nulls.
C.Replace all non-numeric characters with zeros, then sum the column as text.
D.Take the absolute value of every entry first, then remove the dollar signs and cast to integer.
AnswerA

This sequence removes formatting characters that block conversion, explicitly handles the 'N/A' entries rather than letting them fail silently, keeps negative amounts so refunds reduce revenue correctly, and performs a deliberate cast to a numeric type. It also leaves a clear audit trail of which rows were invalid, which matters when the summed total feeds a financial report.

Why this answer

The column mixes formatting characters, non-numeric placeholders, and meaningful negative values. A robust cleanup removes currency symbols and separators, explicitly maps invalid entries to nulls or a flagged value, retains the negative sign for refunds, and then casts to a decimal type. Direct casting and character-stripping shortcuts either lose rows silently or corrupt the sign and precision of the amounts.

Exam trap

The trap here is trusting an implicit cast to clean currency-formatted text, when the dollar signs, commas, and 'N/A' entries actually cause errors or silent nulls that understate revenue.

153
MCQhard

In time series decomposition, a pattern that repeats at regular intervals (e.g., weekly, yearly) is called:

A.Cyclical
B.Irregular
C.Trend
D.Seasonality
AnswerD

Seasonality describes a repeating pattern occurring at fixed intervals, such as weekly or yearly cycles, distinct from trend and irregular components. This matches the stem's definition of a pattern repeating at regular intervals within time series decomposition.

Why this answer

Seasonality refers to a pattern that repeats at fixed, calendar-based intervals such as weekly, monthly, quarterly, or yearly. Because the question specifies 'regular intervals' like weekly or yearly, this is the defining characteristic of the seasonal component in time series decomposition. Seasonality is distinct from trend (long-term direction) and cyclical (irregular multi-year fluctuations).

Exam trap

The trap here is confusing cyclical with seasonal — both are repeating patterns, but cyclical has an irregular, multi-year period while seasonality has a fixed calendar period, and the exam relies on candidates glossing over that distinction.

How to eliminate wrong answers

Option A is wrong because cyclical patterns occur over longer, irregular periods (typically 2+ years) tied to economic or business cycles, not fixed calendar intervals. Option B is wrong because the irregular component represents random, unpredictable residual noise after trend, seasonal, and cyclical components are removed. Option C is wrong because trend describes the long-term upward or downward movement of the series over time, not a repeating pattern at fixed intervals.

154
MCQmedium

A simple linear regression model predicts sales (y) from advertising spend (x). The equation is y = 2.5x + 10, and R² = 0.81. Which interpretation is correct?

A.The correlation between sales and advertising is 0.81.
B.When advertising is $0, sales are $2.5.
C.81% of the variation in sales is explained by advertising spend.
D.For every $1 increase in advertising, sales increase by $10 on average.
AnswerC

R-squared is the coefficient of determination: 0.81 means 81% of the variance in the dependent variable (sales) is accounted for by the independent variable (advertising spend). The remaining 19% is unexplained by the model, so this interpretation matches the statistic precisely.

Why this answer

Slope indicates that each unit increase in x increases y by 2.5 units. R² of 0.81 means 81% of variance in y is explained by x.

155
MCQeasy

A data analyst is reviewing a dataset and notices that the 'ProductCategory' column contains inconsistent entries such as 'Electronics', 'electronics', and 'ELECTRONICS'. Which data cleaning technique should the analyst apply to standardize these values?

A.Standardization
B.Normalization
C.Case normalization
D.One-hot encoding
AnswerC

Case normalization involves converting all text to a consistent case, such as lowercase or uppercase. This resolves inconsistencies like 'Electronics', 'electronics', and 'ELECTRONICS' by making them identical. It is a simple and effective data cleaning step for standardizing categorical text values. After case normalization, the entries would all become 'electronics', for example.

Why this answer

Case normalization is the appropriate technique to standardize text values that differ only by case. By converting all entries to a consistent case, the analyst ensures that 'Electronics', 'electronics', and 'ELECTRONICS' are treated as the same category. Other techniques like normalization, standardization, and one-hot encoding are not designed to fix case inconsistencies in text data.

Exam trap

The trap here is confusing data normalization or standardization with text case normalization, applying numerical scaling methods to a categorical text issue.

156
MCQmedium

In a time series analysis, a retail analyst observes consistent peaks in sales every December and troughs every February. This pattern repeats annually. Which component of time series does this represent?

A.Irregular
B.Seasonality
C.Trend
D.Cyclical
AnswerB

Seasonality captures repeating fluctuations within a fixed period shorter than a year, so the annual December peaks and February troughs satisfy that constraint precisely. Unlike trend, which shows long-term direction, or cyclical variation, which lacks a fixed period, seasonality is defined by calendar-linked regularity.

Why this answer

Seasonality refers to regular patterns that repeat over fixed periods, such as months or quarters.

157
Multi-Selectmedium

A dataset contains outliers in a feature that will be used for linear regression. Which two outlier treatment methods are appropriate? (Choose TWO)

Select 2 answers
A.Cap the outliers at a percentile (e.g., 99th percentile)
B.Use min-max normalization
C.Increase the sample size
D.Remove the outlier rows
E.Replace outliers with the mean
AnswersA, D

Capping at the 99th percentile winsorises extreme values to a boundary, retaining the observation while limiting its leverage on the regression coefficients. This preserves sample size and satisfies the need to reduce outlier influence without discarding data.

Why this answer

Option A (Cap the outliers at a percentile, e.g., 99th percentile) is correct because winsorizing/capping replaces extreme values with a boundary value at a chosen percentile, which preserves the data points while reducing the leverage that extreme values exert on the least-squares fit, thereby limiting their disproportionate influence on the regression coefficients. Option D (Remove the outlier rows) is correct because deleting the offending observations eliminates their influence entirely, which is a standard and valid treatment when outliers are identified as erroneous or unrepresentative and the remaining sample is still adequate for fitting the model. Option B (Use min-max normalization) is not appropriate because rescaling features to a [0,1] range is a transformation that preserves the relative positions and leverage of outliers, so it does nothing to mitigate their effect on the regression.

Option C (Increase the sample size) does not treat the outliers themselves; adding more data does not reduce the influence of existing extreme points and may even introduce more outliers. Option E (Replace outliers with the mean) is generally inappropriate because substituting the mean distorts the distribution, artificially reduces variance, and biases the regression estimates rather than properly handling the extreme values.

Exam trap

DA0-002 often tests the confusion that normalization or mean-imputation 'handles' outliers, when in fact only capping, trimming, or robust methods (e.g., Huber regression) actually mitigate their influence.

158
MCQhard

A marketing analyst wants to segment customers based on purchasing behavior and demographics. The dataset includes continuous variables (spending amount, frequency) and categorical variables (region, gender). The analyst decides to use k-means clustering. What should the analyst do to prepare the data?

A.Use raw data because k-means works with mixed types
B.Standardize continuous variables and one-hot encode categorical variables
C.Apply PCA first to reduce dimensionality
D.Remove categorical variables entirely
AnswerB

K-means computes Euclidean distance, so unstandardised continuous variables with larger ranges dominate. Standardising spending and frequency equalises their influence, while one-hot encoding converts region and gender into numeric binary columns, since k-means cannot process categorical labels directly.

Why this answer

K-means clustering relies on Euclidean distance, which is sensitive to the scale of features. Standardizing continuous variables (e.g., spending amount, frequency) ensures they contribute equally to distance calculations, while one-hot encoding categorical variables (e.g., region, gender) converts them into numerical form without implying ordinal relationships, allowing k-means to process mixed data types correctly.

Exam trap

The trap here is that candidates assume k-means can natively handle mixed data types because it is a common clustering algorithm, but it strictly requires numerical input and scale normalization to avoid skewed distance calculations.

How to eliminate wrong answers

Option A is wrong because k-means cannot directly handle categorical variables; it requires numerical input and assumes continuous features, so using raw mixed-type data would produce meaningless distance calculations. Option C is wrong because PCA is a dimensionality reduction technique applied after preprocessing, not a substitute for standardizing and encoding; it may be used optionally but is not the required preparation step. Option D is wrong because removing categorical variables discards valuable demographic information that could improve segmentation, and k-means can incorporate them after proper encoding.

159
MCQeasy

A data analyst wants to predict customer churn based on categorical features like region and plan type, and continuous features like usage and tenure. Which regression type should be used?

A.Logistic regression
B.Ridge regression
C.Linear regression
D.Lasso regression
AnswerA

Logistic regression predicts a binary outcome, churn versus no churn, satisfying the categorical target requirement. It handles mixed predictors: region and plan type via dummy encoding, usage and tenure as continuous terms. Linear regression would model a continuous response, which churn is not, so logistic regression fits this classification scenario.

Why this answer

Logistic regression is the correct choice because the target variable, customer churn, is binary (churn vs. no churn). Logistic regression models the probability of a binary outcome using a sigmoid function, making it suitable for classification tasks with both categorical and continuous predictors.

Exam trap

CompTIA often tests the misconception that 'regression' in the option name implies it is only for continuous outcomes, leading candidates to overlook logistic regression as a valid classification technique.

How to eliminate wrong answers

Option B (Ridge regression) is wrong because it is a regularized form of linear regression used for continuous outcomes, not binary classification. Option C (Linear regression) is wrong because it predicts a continuous value and is inappropriate for a binary dependent variable; it can produce probabilities outside [0,1] and violates the assumption of normally distributed errors. Option D (Lasso regression) is wrong because, like Ridge, it is a regularized linear regression for continuous targets and performs feature selection via L1 penalty, but it does not handle binary classification.

160
MCQeasy

During data exploration, an analyst notices that the target variable has a heavily right-skewed distribution. Which data transformation would be most appropriate to make the distribution more symmetric?

A.Log transformation
B.Reciprocal transformation
C.No transformation needed
D.Square root transformation
AnswerA

Log transformation compresses the long right tail by rescaling values on a multiplicative scale, pulling extreme high observations closer to the bulk. This reduces positive skewness, making the target distribution more symmetric and better suited to techniques assuming normality.

Why this answer

Log transformation is appropriate for heavily right-skewed distributions because it compresses the high values and spreads out the low values, making the distribution more symmetric. Square root transformation is better for moderate skew, and reciprocal transformation is for severe skew. Therefore, option A (Log transformation) is correct.

161
MCQmedium

A company’s marketing team wants to segment customers based on purchase history, demographics, and website behavior. The data includes both numeric and categorical variables. Which clustering algorithm is best suited for handling mixed data types?

A.Hierarchical clustering with Gower distance
B.K-modes clustering
C.DBSCAN with Euclidean distance
D.K-means clustering
AnswerA

Hierarchical clustering with Gower distance computes pairwise dissimilarity across numeric and categorical attributes simultaneously, so mixed-type records can be segmented without arbitrary encoding. It satisfies the stem's mixed data constraint, unlike k-means, which relies on Euclidean distance and requires numeric, scaled input.

Why this answer

Hierarchical clustering with Gower distance is best suited for mixed data types because Gower distance computes a dissimilarity measure that handles both numeric and categorical variables by normalizing numeric differences and using a simple matching coefficient for categorical ones. This allows the algorithm to create a distance matrix that equally weights all variable types, making it ideal for segmenting customers with purchase history, demographics, and website behavior data.

Exam trap

The trap here is that candidates often assume K-means or DBSCAN can handle mixed data by simply encoding categorical variables, but they overlook that Euclidean distance on encoded data distorts the geometry and fails to preserve the natural dissimilarity structure of categorical variables.

How to eliminate wrong answers

Option B (K-modes clustering) is wrong because it is designed exclusively for categorical data and cannot handle numeric variables like purchase history or website behavior metrics. Option C (DBSCAN with Euclidean distance) is wrong because Euclidean distance is only meaningful for numeric data and cannot properly measure dissimilarity between categorical variables, leading to distorted clusters. Option D (K-means clustering) is wrong because it relies on Euclidean distance and assumes numeric, continuous data; it cannot directly incorporate categorical variables without encoding, and even with encoding, it is sensitive to scaling and does not naturally handle mixed types.

162
MCQeasy

A data analyst is working with a dataset that contains a column for 'customer satisfaction rating' on a scale from 1 to 5, where 1 is very dissatisfied and 5 is very satisfied. The analyst wants to summarize the central tendency of this data. Which measure of central tendency is most appropriate for this ordinal data?

A.Range
B.Mode
C.Mean
D.Median
AnswerD

The median is the middle value when data is ordered, making it appropriate for ordinal data because it does not assume equal intervals. For satisfaction ratings, the median represents the central rating without implying that the difference between ratings is consistent. It is robust to outliers and skewed distributions, providing a more accurate summary of the typical response.

Why this answer

For ordinal data such as satisfaction ratings, the median is the most appropriate measure of central tendency because it does not assume equal intervals between categories. It identifies the middle value when the data is ordered, providing a robust summary that is not influenced by extreme ratings. The mean assumes equal intervals, and the mode may not represent the center, while the range measures spread.

Exam trap

The trap here is assuming that the mean is always the best measure of central tendency, but for ordinal data, the median is more appropriate due to unequal intervals.

163
MCQhard

A data analyst is working with a dataset containing house prices. After building a multiple linear regression model, the analyst observes that the model performs well on training data but poorly on validation data. Which technique is most appropriate to address this issue?

A.Decrease the training data size
B.Use a polynomial transformation
C.Increase the number of features
D.Apply L2 regularization (Ridge)
AnswerD

L2 regularization (Ridge) adds a penalty proportional to the squared magnitude of coefficients, shrinking them towards zero and reducing variance without eliminating features. This directly counters the overfitting causing the train-validation performance gap, satisfying the scenario's need to improve generalisation on unseen house price data.

Why this answer

The model is overfitting the training data, as evidenced by high performance on training data but poor performance on validation data. L2 regularization (Ridge) adds a penalty term proportional to the square of the coefficients, which shrinks them and reduces model complexity, thereby improving generalization to unseen data.

Exam trap

CompTIA often tests the distinction between overfitting and underfitting, and candidates mistakenly choose polynomial transformation or adding features thinking they will improve fit, when in fact they increase model complexity and worsen overfitting.

How to eliminate wrong answers

Option A is wrong because decreasing the training data size would exacerbate overfitting by providing the model with even less information to learn generalizable patterns. Option B is wrong because polynomial transformation increases model complexity and feature interactions, which typically worsens overfitting rather than addressing it. Option C is wrong because increasing the number of features adds more predictors, which increases the risk of overfitting and does not directly penalize large coefficients.

164
MCQmedium

A data analyst is working with a dataset containing a categorical variable 'Education Level' with categories: High School, Bachelor's, Master's, PhD. The analyst wants to include this variable in a regression model. Which encoding technique should the analyst use?

A.Binary encoding
B.Label encoding
C.Ordinal encoding
D.One-hot encoding
AnswerD

One-hot encoding creates binary columns for each category, avoiding implying any ordinal relationship. For regression, it is appropriate for nominal categorical variables like Education Level, which has no inherent order. It allows the model to estimate a separate coefficient for each category, capturing differences without assuming a linear progression.

Why this answer

One-hot encoding is the preferred method for nominal categorical variables in regression because it avoids imposing an artificial order or equal spacing. It creates separate binary indicators, allowing the model to estimate distinct effects for each education level. Label, binary, and ordinal encoding all introduce numerical relationships that may misrepresent the categorical nature of the data.

Exam trap

The trap here is using label encoding for a nominal variable, which forces the model to treat categories as ordered numbers and can lead to incorrect coefficient estimates.

165
MCQhard

A data analyst is evaluating a binary classification model for loan default prediction. The model achieves 98% accuracy on the test set, but the analyst notices that only 2% of loans in the dataset actually defaulted. The analyst is concerned that accuracy is misleading. Which metric should the analyst examine to better assess the model's ability to identify actual defaulters?

A.Mean Absolute Error (MAE)
B.Adjusted R-squared
C.Root Mean Squared Error (RMSE)
D.Recall
AnswerD

Recall, also called sensitivity or true positive rate, measures the proportion of actual defaulters that the model correctly identifies. With only 2% defaults, a model predicting 'no default' for all cases would achieve 98% accuracy but zero recall. Therefore, recall is the appropriate metric to assess the model's ability to catch actual defaulters, directly addressing the analyst's concern.

Why this answer

In imbalanced classification, accuracy can be misleading because a model that always predicts the majority class achieves high accuracy but fails to identify the minority class. Recall focuses on the minority class by measuring how many actual positives were correctly identified. For loan default prediction, missing a defaulter is costly, so recall is a key metric.

The other metrics are for regression or model fit, not classification performance.

Exam trap

The trap here is assuming that high accuracy always indicates a good model, even when the data is imbalanced and the cost of missing the minority class is high.

166
MCQhard

A data analyst has a time series of monthly sales data. They observe that sales are consistently higher every December and lower every January. Which component of time series does this pattern represent?

A.Irregular
B.Cyclical
C.Seasonality
D.Trend
AnswerC

Seasonality describes repeating, calendar-linked fluctuations within a fixed period, such as December peaks and January troughs each year. The consistent annual recurrence in the monthly sales series identifies this component rather than trend or cyclical variation.

Why this answer

A pattern that repeats every December (high) and every January (low) is tied to the calendar year, which is the defining characteristic of seasonality. Seasonal components repeat at fixed, known periods — here, an annual cycle with monthly data.

Exam trap

The trap is picking cyclical because both involve repetition — but cyclical spans irregular multi-year periods while seasonality is locked to a fixed calendar period like December, and the exam expects you to spot the calendar anchor.

How to eliminate wrong answers

Option A is wrong because the irregular component is random, unpredictable noise that cannot be attributed to a repeating calendar pattern. Option B is wrong because cyclical patterns span multiple years with irregular timing tied to economic cycles, not a fixed annual December/January rhythm. Option D is wrong because trend represents the long-term direction of the series (growth or decline over years), not a recurring within-year pattern.

167
MCQmedium

A data analyst is compiling a monthly sales dashboard. The dataset contains a 'Revenue' column with a highly right-skewed distribution due to a few large enterprise transactions. The analyst wants a measure of central tendency that is robust to these outliers. Which measure should the analyst use?

A.Mode
B.Mean
C.Median
D.Range
AnswerC

The median is the middle value when data is ordered, so it is not affected by extreme values. With right-skewed revenue data, the median remains representative of the typical transaction, making it the robust measure of central tendency. It accurately reflects the center without being distorted by the large enterprise transactions.

Why this answer

The median is the appropriate measure of central tendency when data is skewed or contains outliers because it depends only on the middle position. The mean would be inflated by the large enterprise transactions, while the mode and range do not serve as robust central measures. Thus, the median best represents the typical revenue.

Exam trap

The trap here is assuming the mean is always the best measure of central tendency, ignoring its sensitivity to outliers in skewed distributions.

168
MCQhard

Given the linear regression output, which independent variable has the strongest effect on price, based on standardized coefficients?

A.bathrooms
B.sqft_living
C.Intercept
D.bedrooms
AnswerB

Standardised coefficients express each predictor's effect in standard-deviation units, so they are directly comparable across variables measured on different scales. The variable with the largest absolute standardised coefficient, sqft_living, therefore exerts the strongest effect on price.

Why this answer

Standardized coefficients (beta weights) allow comparison of the relative strength of independent variables by measuring the number of standard deviations the dependent variable changes per one standard deviation change in the predictor. In the regression output, sqft_living has the highest absolute standardized coefficient, indicating it has the strongest effect on price. The intercept is not an independent variable and its coefficient is not standardized for comparison.

Exam trap

The trap here is that candidates mistakenly compare unstandardized coefficients or p-values instead of standardized coefficients, leading them to choose a variable like bathrooms or bedrooms that appears significant but has a weaker standardized effect.

How to eliminate wrong answers

Option A is wrong because bathrooms may have a statistically significant coefficient, but its standardized coefficient is smaller than that of sqft_living, meaning it has a weaker relative effect on price. Option C is wrong because the intercept is a constant term representing the predicted price when all independent variables are zero; it is not an independent variable and its coefficient is not standardized for effect comparison. Option D is wrong because bedrooms, while possibly significant, has a lower absolute standardized coefficient than sqft_living, indicating a weaker influence on price per standard deviation change.

169
MCQmedium

An analyst is conducting an A/B test to compare two website designs. The null hypothesis is that there is no difference in conversion rates. The p-value obtained is 0.03, and the significance threshold is 0.05. What should the analyst conclude?

A.Reject the null hypothesis; there is a significant difference.
B.Accept the alternative hypothesis that the new design is better.
C.The test is inconclusive; need a larger sample size.
D.Fail to reject the null hypothesis; there is no significant difference.
AnswerA

A p-value of 0.03 falls below the 0.05 significance threshold, so the null hypothesis of equal conversion rates is rejected. The result is statistically significant, indicating the observed difference between the two website designs is unlikely to arise from chance alone.

Why this answer

In hypothesis testing, if the p-value (0.03) is less than the significance level (0.05), you reject the null hypothesis. This indicates that there is statistically significant evidence of a difference in conversion rates between the two designs. The correct conclusion is to reject the null hypothesis.

Exam trap

The trap is interpreting a significant p-value as proof that the alternative hypothesis is true in a specific direction (e.g., new design is better), when it only indicates a difference.

How to eliminate wrong answers

Option B is wrong because rejecting the null hypothesis does not automatically mean the new design is better; it only indicates a significant difference, not the direction. Option C is wrong because the test is not inconclusive; the p-value is below the threshold, so a conclusion can be drawn. Option D is wrong because failing to reject the null hypothesis would require a p-value greater than 0.05.

170
MCQeasy

A data analyst calculates the mean, median, and mode of a dataset. Which of the following best describes how these measures are used in descriptive statistics?

A.To identify outliers using standard deviation
B.To test hypotheses about population parameters
C.To describe the central tendency of the data
D.To determine the probability of an event
AnswerC

Mean, median and mode each summarise a dataset with a single representative value, describing where data clusters. They are the standard measures of central tendency in descriptive statistics, distinct from dispersion measures such as range or standard deviation.

Why this answer

Mean, median, and mode are measures of central tendency, which are used in descriptive statistics to summarize the center point of a dataset. They describe where the data tends to cluster. The correct answer is C because these measures describe the central tendency.

Exam trap

The trap is mixing up descriptive statistics (summarizing data) with inferential statistics (drawing conclusions about populations), leading candidates to choose hypothesis testing or probability options.

How to eliminate wrong answers

Option A is wrong because identifying outliers using standard deviation is a measure of dispersion, not central tendency. Option B is wrong because testing hypotheses about population parameters is inferential statistics, not descriptive. Option D is wrong because determining the probability of an event is probability theory, not descriptive statistics.

171
MCQmedium

A data analyst at an online retailer is examining a dataset of customer orders. The 'order_total' column has a mean of $85 and a median of $62. The analyst wants to describe the typical order amount for a presentation to the marketing team. Which measure of central tendency is most appropriate to report as the typical value, and why?

A.The mean, because it uses all data points and is the standard measure of central tendency.
B.The median, because it is resistant to extreme values and better represents the typical order in a skewed distribution.
C.The mode, because it identifies the most frequent order amount and is easy to understand.
D.The range, because it shows the spread of order amounts and highlights variability.
AnswerB

The median is the middle value when data is ordered, so it is not influenced by a few extremely large orders. In a right-skewed distribution where mean > median, the median better reflects the typical order amount for the majority of customers, making it the appropriate measure to report.

Why this answer

When a distribution is skewed, the mean is pulled toward the tail, while the median remains a robust measure of the center. Here, the mean ($85) is substantially higher than the median ($62), signaling right skew. The median therefore better represents the typical order amount for most customers.

Exam trap

The trap here is assuming the mean is always the best measure of central tendency, overlooking the impact of outliers in skewed data.

172
MCQhard

A data analyst is evaluating a classification model that predicts whether customers will churn. The dataset is highly imbalanced, with only 5% of customers churning. The analyst wants to choose a metric that focuses on the model's ability to correctly identify actual churners. Which metric should the analyst prioritize?

A.Specificity
B.Recall
C.Precision
D.Accuracy
AnswerB

Recall, also known as sensitivity or true positive rate, measures the proportion of actual positives that are correctly identified. In this scenario, the analyst wants to correctly identify actual churners, so recall directly addresses that goal. High recall means the model captures most churners, which is crucial when the cost of missing a churner is high.

Why this answer

Recall is the metric that quantifies the model's ability to find all relevant instances of the positive class. Since the analyst wants to correctly identify actual churners, recall is the most appropriate metric. Accuracy and specificity focus on the negative class or overall correctness, and precision does not capture missed positives.

Exam trap

The trap here is defaulting to accuracy as the primary metric; in imbalanced datasets, it can be deceptively high even when the model fails to predict the minority class.

173
MCQmedium

A stock analyst is analyzing monthly sales data for a retail company and observes a consistent pattern of high sales every December. This pattern is most likely an example of which time series component?

A.Irregular
B.Cyclical
C.Seasonality
D.Trend
AnswerC

Seasonality describes a repeating pattern that recurs at fixed intervals shorter than a year, such as peak sales every December. This matches the stem's consistent annual December spike, distinguishing it from trend, which is a long-term directional movement.

Why this answer

A consistent spike every December is a calendar-driven repeating pattern, which is the definition of seasonality in time series analysis. The fixed annual period (12 months) and predictable timing make this a classic seasonal component.

Exam trap

The trap is confusing seasonality with trend — a December spike repeats every year (seasonal), whereas trend would be a steady multi-year rise or fall in sales, and the exam tests whether you anchor on the calendar periodicity.

How to eliminate wrong answers

Option A is wrong because irregular components are random shocks (e.g., a one-off supply disruption) that do not repeat predictably each December. Option B is wrong because cyclical patterns are multi-year, irregular fluctuations tied to business cycles, not a reliable annual December spike. Option D is wrong because trend captures the long-term direction of sales over many years, not the recurring within-year December peak.

174
MCQeasy

In simple linear regression, the coefficient of determination R² measures:

A.The probability that the slope is zero
B.The slope of the regression line
C.The proportion of variance in the dependent variable explained by the independent variable
D.The strength and direction of the linear relationship
AnswerC

R² quantifies the share of total variance in the dependent variable accounted for by the fitted regression on the independent variable, expressed as a proportion between 0 and 1. It therefore directly satisfies the stem's requirement to measure explained variance rather than correlation strength or slope.

Why this answer

R², the coefficient of determination, quantifies the proportion of variance in the dependent variable that is explained by the independent variable(s) in the regression model. It ranges from 0 to 1, where 1 means the model explains all variance and 0 means it explains none.

Exam trap

The trap is confusing R² with Pearson's r — R² measures explained variance (0 to 1) while r measures strength and direction (−1 to +1), and the exam offers both as plausible-sounding options.

How to eliminate wrong answers

Option A is wrong because the probability that the slope is zero is assessed via the p-value of the slope coefficient's t-test, not R². Option B is wrong because the slope of the regression line is the coefficient β₁ (or b), which represents the change in Y per unit change in X — R² is a goodness-of-fit measure, not a slope. Option D is wrong because the strength and direction of a linear relationship is measured by the correlation coefficient r (Pearson's r), which ranges from −1 to +1; R² is r squared and loses direction information.

175
MCQeasy

In a regression analysis, the coefficient of determination (R²) is 0.85. How should this value be interpreted?

A.85% of the data points lie on the regression line
B.The slope of the regression line is 0.85
C.85% of the variance in the dependent variable is explained by the model
D.85% of the independent variables are significant
AnswerC

R² measures the proportion of variance in the dependent variable accounted for by the regression model. A value of 0.85 means the model explains 85% of that variance, leaving 15% attributable to other factors or random error.

Why this answer

R², the coefficient of determination, measures the proportion of variance in the dependent variable that is explained by the independent variables in the regression model. An R² of 0.85 means 85% of the variance in the dependent variable is accounted for by the model, indicating a strong fit.

Exam trap

DA0-002 often tests the misconception that R² represents the percentage of points on the line or the slope — candidates must remember it is the proportion of variance explained.

How to eliminate wrong answers

Option A is wrong because R² does not measure the percentage of data points lying on the regression line — that would be a different concept, and in most real regressions very few points lie exactly on the line. Option B is wrong because R² is not the slope; the slope is a separate coefficient (β₁) in the regression equation. Option D is wrong because R² says nothing about the significance of individual independent variables — that is assessed via t-tests and p-values for each coefficient.

176
MCQmedium

A data analyst is working with a dataset that contains a column 'purchase_date' stored as a string in the format 'YYYY-MM-DD'. The analyst needs to calculate the number of days between each purchase and the current date. Which data transformation is required?

A.Apply one-hot encoding to the string.
B.Normalize the string values to a 0-1 range.
C.Convert the string to an integer data type.
D.Convert the string to a date data type.
AnswerD

To calculate the difference between dates, the string must be converted to a date data type so that date arithmetic can be performed. Many tools and languages provide functions to parse strings into dates, enabling subtraction to yield the number of days.

Why this answer

Calculating the number of days between dates requires date arithmetic, which is only possible if the values are in a date data type. Converting the string to a date allows the use of built-in functions to compute differences. Other transformations like integer conversion or encoding do not support date calculations.

Exam trap

The trap here is thinking that string manipulation alone can compute date differences, but without a proper date type, calculations are error-prone and complex.

177
MCQhard

A data analyst is building a binary classification model to predict customer churn. The dataset is imbalanced, with only 10% churners. The analyst wants to evaluate model performance with a focus on correctly identifying churners. Which metric is most appropriate?

A.Recall (sensitivity)
B.F1-score
C.Precision
D.Accuracy
AnswerA

Recall measures the proportion of actual churners correctly identified, directly matching the stated focus on catching churners. With only 10% churners, accuracy is misleading because predicting 'no churn' everywhere scores 90%, whereas recall exposes missed churners.

Why this answer

Recall (sensitivity) is the most appropriate metric because it measures the proportion of actual churners correctly identified by the model. Since the dataset is imbalanced (only 10% churners) and the analyst's focus is on correctly identifying churners, recall directly addresses the cost of missing positive cases (false negatives). Accuracy would be misleading due to class imbalance, while precision and F1-score prioritize different trade-offs.

Exam trap

The trap here is that candidates often default to accuracy as the default metric, failing to recognize that class imbalance renders accuracy misleading, and that the question's explicit focus on 'correctly identifying churners' points directly to recall, not precision or F1-score.

How to eliminate wrong answers

Option B (F1-score) is wrong because it balances precision and recall, but the analyst's primary goal is to maximize identification of churners, not to balance false positives and false negatives; F1-score would penalize a model that achieves high recall at the expense of precision, which may be acceptable in this scenario. Option C (Precision) is wrong because it measures the proportion of predicted churners that are actual churners, focusing on false positives rather than false negatives; the analyst wants to minimize missed churners, not necessarily avoid false alarms. Option D (Accuracy) is wrong because with only 10% churners, a naive model predicting all non-churners would achieve 90% accuracy, masking poor performance on the minority class; accuracy is inappropriate for imbalanced classification problems.

178
MCQmedium

A retail company wants to predict future sales based on historical data. Which modeling approach is most appropriate if the data shows a clear seasonal pattern?

A.Linear regression
B.Time series analysis
C.K-means clustering
D.Logistic regression
AnswerB

Time series analysis explicitly models sequential dependence and seasonality through components such as trend, seasonal, and residual terms, satisfying the stem's clear seasonal pattern requirement. Unlike regression, it uses autocorrelation and prior-period values, so forecasts of future sales account for recurring cycles rather than treating observations as independent.

Why this answer

Time series analysis is specifically designed to model data points indexed in time order, making it ideal for capturing and forecasting seasonal patterns. Unlike regression models, it accounts for autocorrelation, trends, and seasonality components, which are critical for accurate sales prediction from historical data.

Exam trap

The trap here is that candidates see 'predict future sales' and mistakenly choose linear regression, overlooking that time series methods are required when data has temporal dependencies and seasonality.

How to eliminate wrong answers

Option A is wrong because linear regression assumes independence of observations and cannot model time-dependent structures like seasonality or autocorrelation. Option C is wrong because K-means clustering is an unsupervised learning method used for grouping similar data points, not for forecasting future values. Option D is wrong because logistic regression is used for binary classification problems, not for predicting continuous numeric sales figures.

179
Multi-Selectmedium

A retail company wants to segment its customers based on purchase history. Which THREE methods are appropriate for customer segmentation?

Select 3 answers
A.RFM analysis
B.Linear regression
C.K-means clustering
D.t-test
E.Hierarchical clustering
AnswersA, C, E

RFM analysis scores customers on recency, frequency and monetary value derived from purchase history, directly satisfying the segmentation requirement. It groups customers into actionable tiers such as best, loyal or at-risk based on transactional behaviour.

Why this answer

RFM analysis (A) is correct because it segments customers by Recency, Frequency, and Monetary value of purchases, directly using purchase history to group customers into meaningful tiers. K-means clustering (C) is correct because it partitions customers into k groups based on feature similarity such as purchase behavior, making it a standard unsupervised segmentation technique. Hierarchical clustering (E) is correct because it builds a dendrogram of nested customer clusters, allowing segmentation at different granularity levels without pre-specifying the number of clusters.

Linear regression (B) is not appropriate because it predicts a continuous outcome rather than assigning customers to segments. The t-test (D) is not appropriate because it is a hypothesis test comparing means between two groups, not a segmentation method.

Exam trap

The trap is that regression and hypothesis-testing methods sound analytical and data-driven, so candidates may select them for segmentation even though they predict values or test differences rather than grouping customers.

180
MCQmedium

The exhibit shows an SQL query executed on an 'orders' table that contains 'order_id', 'customer_id', and 'order_date'. What is the purpose of this query?

A.Count total orders per customer regardless of date
B.Calculate average order count per customer for 2023
C.Find products with more than 5 orders in 2023
D.Identify customers who placed more than 5 orders in 2023
AnswerD

The query groups orders by customer_id, filters order_date to the 2023 range, and applies a HAVING count greater than five. This returns customers whose 2023 order count exceeds five, satisfying the stated purpose of identifying high-frequency customers.

Why this answer

The query groups orders by customer_id and filters using a HAVING clause with COUNT(*) > 5, which counts the number of orders per customer. The WHERE clause restricts orders to those placed in 2023, so the result identifies customers who placed more than 5 orders in that year. This matches option D exactly.

Exam trap

CompTIA often tests the distinction between WHERE and HAVING, and the trap here is confusing a count of orders per customer with a count of products or an average, leading candidates to pick option B or C.

How to eliminate wrong answers

Option A is wrong because the WHERE clause filters for order_date in 2023, so the count is not regardless of date. Option B is wrong because the query counts orders per customer, not the average order count per customer. Option C is wrong because the query operates on an 'orders' table with no product-related column; it counts orders per customer, not products.

181
Multi-Selectmedium

An analyst is preparing data for an A/B test and wants to ensure valid results. Which TWO of the following should be considered when calculating the required sample size?

Select 2 answers
A.Data dimensionality
B.Desired effect size
C.Skewness of data
D.Number of features
E.Statistical power
AnswersB, E

Effect size quantifies the minimum difference the test must detect between variants. Smaller effects demand larger samples, so specifying it directly drives the sample size calculation and satisfies the validity constraint, preventing an underpowered test that misses genuine differences.

Why this answer

Sample size calculation depends on desired effect size and statistical power, among other factors like significance level.

182
MCQeasy

A data analyst wants to compare the means of three different training methods on employee productivity. Which statistical test is most appropriate?

A.Correlation analysis
B.ANOVA
C.Chi-square test
D.t-test
AnswerB

ANOVA compares means across three or more independent groups in one test, controlling the family-wise error rate that repeated t-tests would inflate. With three training methods as the grouping factor, it fits the scenario directly.

Why this answer

ANOVA (Analysis of Variance) is used to compare means of three or more groups.

183
MCQeasy

A data analyst is examining the relationship between two continuous variables: temperature and ice cream sales. The analyst wants to quantify the strength and direction of their linear association. Which statistical measure should the analyst use?

A.Covariance
B.Spearman rank correlation
C.Pearson correlation coefficient
D.Chi-square test of independence
AnswerC

The Pearson correlation coefficient measures the strength and direction of a linear relationship between two continuous variables. Temperature and ice cream sales are both continuous, and the analyst seeks a linear association, making Pearson correlation the appropriate measure. It ranges from -1 to 1, indicating perfect negative to perfect positive linear relationships.

Why this answer

Pearson correlation coefficient is the standard measure for quantifying the strength and direction of a linear relationship between two continuous variables. It is scale-independent, ranging from -1 to 1, and directly addresses the analyst's goal. Other options either measure different types of association or are not suitable for continuous data.

Exam trap

The trap here is selecting covariance because it also measures linear relationship, but covariance lacks standardization and does not convey strength on a fixed scale.

184
Multi-Selectmedium

An analyst is planning an A/B test to compare two website designs. Which TWO factors should be considered when calculating the required sample size?

Select 2 answers
A.Data type of the outcome variable
B.Desired effect size
C.Statistical power
D.Color scheme of the designs
E.Number of missing values
AnswersB, C

The desired effect size is the minimum detectable difference between designs that the test must reliably detect. Smaller effects require substantially larger samples, making it a core input to any sample size calculation for the A/B test.

Why this answer

The required sample size for an A/B test depends directly on the desired effect size (B), because smaller effects are harder to detect and demand more observations to distinguish a real difference from noise. It also depends on statistical power (C), conventionally set at 0.80, since higher power (lower Type II error risk) requires a larger sample to reliably detect the effect when it truly exists. The data type of the outcome variable (A) affects the choice of statistical test, not the sample-size formula's core inputs, and the color scheme (D) is merely the design variation being tested, not a computational factor.

The number of missing values (E) is a data-quality issue handled during cleaning or imputation and is not a standard parameter in sample-size calculation.

Exam trap

DA0-002 often tests the four inputs to sample size (α, power, effect size, variance) — candidates pick data type or missing values because they sound statistical, but only effect size and power are among the core parameters.

185
MCQhard

A data analyst is cleaning a dataset with missing values in a time series of daily temperatures. The missing values occur sporadically. Which imputation method is most appropriate to maintain the temporal trend?

A.Forward-fill
B.Mean imputation
C.Median imputation
D.Interpolation
AnswerD

Interpolation estimates missing points from neighbouring known values along the time axis, preserving the daily temperature trend and seasonality. Mean or median imputation flattens local variation, whereas interpolation maintains temporal continuity, satisfying the requirement to keep the series' trend intact.

Why this answer

Interpolation estimates missing values by using the values immediately before and after the gap, producing a smooth curve that preserves the temporal trend of a time series. For sporadically missing daily temperatures, linear interpolation between neighboring days is the most faithful reconstruction. It respects the ordered, continuous nature of the data better than simple fill methods.

Exam trap

The trap is that mean and median imputation are the most commonly taught missing-value fixes, so candidates default to them without recognizing that time-ordered data requires methods that respect temporal continuity.

How to eliminate wrong answers

Option A is wrong because forward-fill copies the last observed value into the gap, creating flat plateaus that distort the trend and are especially poor when values change steadily over time. Option B is wrong because mean imputation replaces missing values with the overall average, ignoring time order and flattening seasonal or trend variation. Option C is wrong because median imputation has the same flaw as mean imputation — it ignores temporal position and injects a constant that does not reflect local conditions.

186
MCQmedium

A data analyst is reviewing a dataset containing house prices. The mean price is $350,000 and the median is $280,000. Which of the following best describes the distribution of house prices?

A.The distribution is right-skewed.
B.The distribution is symmetric.
C.The distribution is left-skewed.
D.The distribution is bimodal.
AnswerA

A mean above the median indicates right-skew, where high-value outliers pull the mean upward while the median stays resistant. This satisfies the stem's constraint of mean ($350,000) exceeding median ($280,000), confirming a positively skewed distribution of house prices.

Why this answer

When the mean ($350,000) is greater than the median ($280,000), the distribution is right-skewed (positively skewed). In a right-skewed distribution, a long tail of high values pulls the mean upward above the median. This is common with house prices, where a few very expensive homes inflate the mean.

Exam trap

DA0-002 often tests the mean-vs-median relationship: candidates who forget that mean > median indicates right skew (not left) pick the wrong direction.

How to eliminate wrong answers

Option B is wrong because in a symmetric distribution the mean and median are approximately equal, but here they differ by $70,000. Option C is wrong because left-skewness (negative skew) produces a mean lower than the median, the opposite of what is observed. Option D is wrong because bimodality refers to two distinct peaks in the distribution and cannot be inferred from mean and median alone.

187
Multi-Selectmedium

Which TWO of the following are appropriate uses of min-max normalisation?

Select 2 answers
A.Transforming data to have mean 0 and standard deviation 1
B.Scaling features to a range of 0 to 1
C.Preparing data for linear regression with normally distributed residuals
D.Preparing data for k-nearest neighbours algorithm
E.Handling missing values
AnswersB, D

Min-max normalisation linearly transforms each feature to a fixed 0–1 range using the minimum and maximum values, satisfying the requirement to bound features within a common scale. This suits algorithms sensitive to magnitude, such as k-nearest neighbours or neural networks, where unbounded inputs distort distance calculations.

Why this answer

Option B is correct because min-max normalisation rescales each feature to a fixed range, typically [0, 1], using the formula (x − min)/(max − min), which is exactly its defining purpose. Option D is correct because k-nearest neighbours relies on distance calculations (e.g., Euclidean distance), so features on different scales would dominate the distance metric; min-max normalisation puts all features on a comparable 0–1 scale, improving the algorithm's behaviour. Option A is not a use of min-max normalisation but of standardisation (z-score scaling), which produces mean 0 and standard deviation 1.

Option C is not specific to min-max normalisation; linear regression with normally distributed residuals concerns the error distribution, not feature scaling, and standardisation is more commonly associated with such assumptions. Option E is incorrect because missing values are handled by imputation or deletion techniques, not by min-max normalisation, which requires complete numeric data.

Exam trap

DA0-002 often tests the confusion between min-max normalisation and z-score standardisation; candidates may incorrectly select the z-score description as a use of min-max.

188
MCQhard

A data analyst is examining the relationship between advertising spend and sales revenue across 50 regions. The analyst calculates a Pearson correlation coefficient of 0.85. Which of the following conclusions is most appropriate?

A.Increasing advertising spend causes an increase in sales revenue.
B.The relationship is not statistically significant because the correlation is less than 0.90.
C.Advertising spend explains 85% of the variation in sales revenue.
D.There is a strong positive linear relationship between advertising spend and sales revenue.
AnswerD

A Pearson correlation of 0.85 indicates a strong positive linear association. This means that as advertising spend increases, sales revenue tends to increase linearly. It does not imply causation, but it does summarize the strength and direction of the linear relationship. This is the most appropriate conclusion based solely on the correlation coefficient.

Why this answer

A Pearson correlation of 0.85 indicates a strong positive linear relationship between the two variables. It does not establish causation, nor does it directly state the proportion of variance explained. The correlation coefficient must be squared to get the coefficient of determination, and significance requires a hypothesis test.

Thus, the appropriate conclusion is the strong positive linear association.

Exam trap

The trap here is interpreting a high correlation as evidence of causation or confusing the correlation coefficient with the coefficient of determination.

189
MCQeasy

Which data quality dimension ensures that data represents the real-world scenario correctly and without errors?

A.Completeness
B.Consistency
C.Accuracy
D.Timeliness
AnswerC

Accuracy verifies that values correctly reflect the real-world entities or events they describe, directly satisfying the stem's requirement for data representing reality without errors. Unlike completeness, which checks for missing values, or consistency, which checks formatting agreement across sources, accuracy measures correctness against an authoritative reference.

Why this answer

Accuracy is the data quality dimension that measures whether data correctly reflects the real-world object or event it describes, free from errors. If a customer's address is stored incorrectly or a transaction amount is wrong, the data is inaccurate even if it is complete and consistent. Completeness, consistency, and timeliness address different aspects and do not capture correctness of representation.

Exam trap

DA0-002 often tests the overlap between accuracy and completeness, tricking candidates into choosing completeness when the scenario describes incorrect values rather than missing values.

How to eliminate wrong answers

Option A is wrong because completeness measures whether all required data is present — a record can be complete but still contain incorrect values. Option B is wrong because consistency measures whether data agrees across systems or formats — data can be consistently wrong in multiple places. Option D is wrong because timeliness measures whether data is available when needed and up to date — stale data can still be accurate for the period it represents.

190
MCQmedium

A data analyst is analyzing survey responses where respondents rated satisfaction on a scale of 1-5. The analyst wants to visualize the distribution of responses. Which chart type is most appropriate?

A.Box plot
B.Scatter plot
C.Line chart
D.Histogram
AnswerD

A histogram bins the numeric 1–5 ratings into adjacent intervals and plots frequency per bin, directly revealing the distribution's shape, spread and skew. Bar charts suit categorical counts, not ordered numeric scales, so they obscure distributional patterns the analyst needs.

Why this answer

A histogram is the most appropriate chart for visualizing the distribution of a single discrete variable, such as satisfaction ratings on a 1-5 scale. It groups the responses into bins (each rating value) and displays the frequency of each bin using bars, clearly showing the shape, central tendency, and spread of the data.

Exam trap

The trap here is that candidates often confuse a histogram with a bar chart, but the key distinction is that a histogram is used for quantitative (ordinal or continuous) data where bin order matters, while a bar chart is for categorical (nominal) data with no inherent order.

How to eliminate wrong answers

Option A is wrong because a box plot summarizes data using five-number statistics (min, Q1, median, Q3, max) and is better for comparing distributions across groups, not for showing the detailed frequency distribution of a single ordinal variable. Option B is wrong because a scatter plot is used to visualize the relationship between two continuous variables, not the distribution of a single categorical or ordinal variable. Option C is wrong because a line chart is typically used to display trends over time or sequential data, not the frequency distribution of discrete survey responses.

191
MCQhard

In A/B testing, which factor is increased by having a larger sample size?

A.P-value
B.Effect size
C.Type I error rate
D.Statistical power
AnswerD

Larger samples shrink the standard error of the estimated effect, which directly raises statistical power — the probability of detecting a true difference when one exists. This satisfies the scenario's need to distinguish genuine treatment effects from random variation.

Why this answer

Statistical power is the probability of correctly rejecting a false null hypothesis (i.e., detecting a true effect). Increasing sample size reduces the standard error, making it easier to detect a true effect and thus increasing power. This is a fundamental principle in hypothesis testing and A/B testing.

Exam trap

The trap is conflating p-value with power or thinking that larger samples increase effect size; candidates often pick p-value because they associate it with significance, but the exam expects understanding that power is the probability of detecting a true effect and is directly boosted by sample size.

How to eliminate wrong answers

Option A is wrong because the p-value is the probability of observing the data given that the null hypothesis is true; it is not directly increased by sample size—in fact, with larger samples, smaller effects can become statistically significant, but the p-value itself is a calculated value from the test, not a factor that increases. Option B is wrong because effect size is a measure of the magnitude of the phenomenon (e.g., Cohen's d) and is independent of sample size; it is a property of the population, not the study design. Option C is wrong because the Type I error rate (alpha) is the probability of rejecting a true null hypothesis and is set by the researcher (typically 0.05); it is not affected by sample size.

192
Multi-Selectmedium

A data analyst is using a decision tree to classify loan applications as approved or denied. To avoid overfitting, the analyst plans to prune the tree. Which TWO of the following are valid pruning techniques? (Choose two.)

Select 2 answers
A.Post-pruning by removing branches that do not improve validation performance.
B.Adding more training data to the dataset.
C.Pre-pruning by limiting the maximum depth of the tree.
D.Increasing the number of features used at each split.
E.Using a random forest ensemble instead of a single decision tree.
AnswersA, C

Post-pruning involves growing the full tree and then removing branches that provide little improvement in validation performance. This technique, such as cost-complexity pruning, simplifies the tree and reduces overfitting. It is effective because it uses validation data to guide the pruning process, ensuring that only branches that contribute to predictive accuracy are retained.

Why this answer

Pruning techniques aim to simplify a decision tree to reduce overfitting. Pre-pruning stops tree growth early, for example by limiting maximum depth. Post-pruning removes branches after the tree is fully grown, using validation performance to guide removal.

Both are standard methods. Increasing features, adding data, or using ensembles are not pruning techniques for a single tree.

Exam trap

The trap here is confusing ensemble methods or data augmentation with pruning, which specifically modifies the tree structure.

193
MCQeasy

A data analyst is creating a visualization to show the relationship between two continuous variables: advertising spend and sales revenue. The analyst wants to display the strength and direction of the linear relationship. Which type of chart should the analyst use?

A.Pie chart
B.Bar chart
C.Scatter plot
D.Histogram
AnswerC

A scatter plot displays individual data points on a two-dimensional plane, with one variable on the x-axis and the other on the y-axis. It allows the analyst to visually assess the strength, direction, and form of the relationship between two continuous variables. It is the standard tool for examining correlation and identifying patterns or outliers.

Why this answer

A scatter plot is the appropriate visualization for examining the relationship between two continuous variables. It plots each observation as a point, allowing the analyst to see if there is a positive, negative, or no correlation, and to spot outliers. Bar charts, pie charts, and histograms are designed for categorical or univariate data and cannot effectively display bivariate relationships.

Exam trap

The trap here is confusing a histogram, which shows one variable's distribution, with a scatter plot, which shows the relationship between two variables.

194
Multi-Selectmedium

Which THREE of the following are common steps in data cleaning?

Select 3 answers
A.Removing outliers without justification
B.Imputing missing values
C.Standardizing data formats
D.Removing duplicate records
E.Increasing sample size
AnswersB, C, D

Missing values are often imputed to maintain dataset completeness.

Why this answer

Imputing missing values is a common data cleaning step because real-world datasets often have gaps due to data collection errors or system failures. Techniques like mean/median imputation, regression imputation, or using algorithms like k-NN help preserve sample size and avoid bias that would result from simply dropping rows. This ensures the dataset remains usable for analysis without introducing significant distortion.

Exam trap

CompTIA often tests the distinction between data cleaning steps and data collection or preprocessing steps, so the trap here is confusing 'increasing sample size' (a data augmentation or collection activity) with actual cleaning tasks like imputation, standardization, and deduplication.

195
MCQeasy

Which data quality dimension is violated if a customer record has a missing phone number?

A.Consistency
B.Accuracy
C.Completeness
D.Validity
AnswerC

Completeness measures whether all required values are present in a record. A missing phone number means the attribute holds no value, directly violating this dimension. Other dimensions, such as accuracy or consistency, assume a value exists but is wrong or conflicting; here nothing is recorded at all.

Why this answer

Completeness refers to the extent to which data is not missing.

196
MCQmedium

A data analyst is working with a dataset that includes a categorical variable 'education_level' with four categories: High School, Bachelor's, Master's, and PhD. The analyst wants to include this variable in a linear regression model. Which encoding method should the analyst use to avoid the dummy variable trap?

A.Label encoding
B.Binary encoding
C.One-hot encoding with three categories (dropping one)
D.One-hot encoding with all four categories
AnswerC

One-hot encoding with k-1 categories (here, three) avoids the dummy variable trap by preventing perfect multicollinearity. The dropped category becomes the reference level, and the coefficients for the other categories represent the difference from that reference. This is the standard approach for including nominal categorical variables in linear regression.

Why this answer

The dummy variable trap occurs when all dummy variables are included in a model with an intercept, causing perfect multicollinearity. To avoid it, one category is omitted as the reference. One-hot encoding with three categories (k-1) achieves this.

Label encoding and binary encoding impose ordinality, which is inappropriate for nominal data. Including all four categories would cause multicollinearity.

Exam trap

The trap here is thinking that one-hot encoding all categories is fine; it actually creates perfect multicollinearity and unstable estimates.

197
MCQmedium

A marketing team wants to segment customers into groups based on purchasing behavior without prior labels. Which algorithm should the data analyst use?

A.K-means clustering
B.K-nearest neighbors
C.Linear regression
D.Decision tree
AnswerA

K-means clustering partitions unlabelled data into k groups by minimising within-cluster variance, directly satisfying the stem's requirement for segmentation without prior labels. Unlike supervised methods such as logistic regression or decision trees, it needs no target variable, making it the appropriate choice for discovering behavioural customer segments.

Why this answer

K-means clustering is the correct choice because it is an unsupervised learning algorithm that groups unlabeled data into clusters based on feature similarity. Since the marketing team has no prior labels for customer segments, K-means can partition customers by purchasing behavior patterns, such as frequency and monetary value, without needing predefined categories.

Exam trap

The trap here is that candidates often confuse unsupervised clustering (K-means) with supervised classification (K-nearest neighbors) because both involve 'K' and grouping, but KNN requires labeled data and predicts labels, while K-means discovers inherent structures without labels.

How to eliminate wrong answers

Option B is wrong because K-nearest neighbors is a supervised learning algorithm that requires labeled training data to classify or predict outcomes, making it unsuitable for unlabeled segmentation. Option C is wrong because linear regression is a supervised regression algorithm used to predict a continuous target variable, not to discover hidden groupings in unlabeled data. Option D is wrong because decision trees are typically used for supervised classification or regression tasks, relying on labeled data to split on features, and cannot perform unsupervised clustering without prior labels.

198
MCQmedium

A data analyst is working with a dataset that contains a column 'region' with values such as 'North', 'South', 'East', 'West', and 'N/A'. The analyst needs to prepare this column for a machine learning model. Which of the following is the most appropriate approach to handle the 'N/A' values?

A.Replace 'N/A' with the mode of the 'region' column.
B.Treat 'N/A' as a separate category and encode it as its own level.
C.Encode 'N/A' as a numeric value of 0 and other regions as 1-4.
D.Remove all rows where 'region' is 'N/A'.
AnswerB

Treating 'N/A' as a distinct category preserves the information that the region is missing or not applicable, which can be predictive. This approach avoids introducing false assumptions and allows the model to learn any signal associated with missingness. It is a common and valid strategy for categorical variables.

Why this answer

Treating 'N/A' as a separate category is the best approach because it retains the information that the region is missing or not applicable, which could be informative. Imputing with the mode or removing rows can introduce bias or lose data, and numeric encoding creates a false ordinal relationship.

Exam trap

The trap here is assuming that missing values must be imputed or removed, when sometimes they represent a valid category that should be preserved.

199
MCQhard

A data scientist is working with a dataset containing 1000 features and 500 samples. The goal is to build a predictive model. Which technique should be used to reduce the number of features while retaining most of the variance?

A.Ridge regression
B.Forward selection
C.Principal Component Analysis (PCA)
D.Lasso regression
AnswerC

PCA projects the 1000 features onto orthogonal principal components ordered by explained variance, letting you keep the top components that retain most variance while discarding the rest, reducing dimensionality despite having fewer samples than features.

Why this answer

Principal Component Analysis (PCA) is an unsupervised dimensionality reduction technique that transforms the original features into a set of orthogonal components, ordered by the variance they capture. Given 1000 features and only 500 samples, PCA is ideal because it reduces the feature space while retaining the maximum variance, helping to avoid overfitting and the curse of dimensionality.

Exam trap

CompTIA often tests the distinction between supervised feature selection (Lasso, Forward selection) and unsupervised dimensionality reduction (PCA), trapping candidates who confuse regularization with variance-based reduction.

How to eliminate wrong answers

Option A is wrong because Ridge regression is a regularization technique that shrinks coefficients but does not reduce the number of features; it retains all features with penalized weights. Option B is wrong because Forward selection is a supervised feature selection method that selects features based on their predictive power, not on variance retention, and it can be computationally expensive with 1000 features. Option D is wrong because Lasso regression performs feature selection by shrinking some coefficients to zero, but it is a supervised method that selects features based on target correlation, not on maximizing variance retention, and may not be optimal for unsupervised dimensionality reduction.

200
MCQhard

A data analyst is building a dashboard that tracks daily active users. The source system records one event row per user session, and a user can have multiple sessions in a day. The analyst needs to display the count of unique users per day. Which approach produces the correct daily active user count?

A.Sum the maximum session count per user across all dates
B.Average the number of session rows per user for each date
C.Count all session rows grouped by date
D.Count distinct user identifiers grouped by date
AnswerD

Applying a distinct count to the user identifier within each date collapses multiple sessions for the same user into one, producing the true number of unique users active that day. This directly matches the definition of daily active users and prevents session volume from distorting the metric.

Why this answer

Daily active users is defined as the number of distinct users who take an action on a given day. Because the source grain is one row per session, the analyst must deduplicate users within each date before counting. A distinct count of the user identifier grouped by date performs exactly that deduplication and yields the correct metric, whereas counting rows or averaging sessions measures activity volume instead of reach.

Exam trap

The trap here is treating the row grain of the source table as the unit of analysis, when the business question asks about distinct people rather than events.

201
MCQmedium

A data analyst is compiling a quarterly report for a retail chain and needs to show how sales revenue is distributed across five product categories. The analyst wants a single visualization that makes it easy to compare each category's share of total revenue and also includes a separate small panel displaying total revenue for the quarter. Which visualization approach best meets these requirements?

A.A dashboard with a pie chart showing category shares and a separate KPI card displaying total revenue.
B.A stacked bar chart showing revenue by category, with the total revenue represented by the height of the entire stack.
C.A donut chart with each category's share, plus a small table in a separate panel listing the total revenue.
D.A pie chart showing each category's percentage of total revenue, with the total revenue displayed as the chart title.
AnswerA

A dashboard allows multiple visual elements to be combined. A pie chart clearly shows each category's proportion of total revenue, while a separate KPI card prominently displays the overall total. This directly satisfies both requirements: comparing shares and having a distinct panel for the total. It is the most effective and standard approach for this type of reporting need.

Why this answer

The requirement is to compare each category's share of total revenue and to have a separate small panel for total revenue. A dashboard with a pie chart and a KPI card meets both: the pie chart shows proportions, and the KPI card is a distinct panel for the total. Other options either embed the total in the chart or use less effective visuals for comparing shares.

Exam trap

The trap here is assuming that displaying the total in a chart title or as part of a stacked bar satisfies the need for a separate panel, when the scenario explicitly calls for a distinct panel.

202
MCQmedium

A data analyst is examining a dataset of customer transactions and notices that the distribution of transaction amounts is highly skewed to the right. The analyst wants to apply a transformation to make the distribution more symmetric for use in a linear regression model. Which of the following transformations is most appropriate?

A.Box-Cox transformation with lambda = 2
B.Square root transformation
C.Reciprocal transformation
D.Logarithmic transformation
AnswerD

The logarithmic transformation is highly effective for right-skewed data with positive values, as it compresses large values more than small ones, reducing skewness and stabilizing variance. It is commonly used for monetary amounts and can make the distribution more symmetric, improving the performance of linear models. It also helps with heteroscedasticity.

Why this answer

For highly right-skewed positive data like transaction amounts, the logarithmic transformation is the standard choice to reduce skewness and make the distribution more symmetric. It compresses the upper tail and can linearize relationships, which benefits linear regression. Other transformations like square root are weaker, while reciprocal and squaring are either too strong or counterproductive.

Exam trap

The trap here is selecting a transformation that is either too weak or that increases skew, such as squaring, instead of the logarithmic transformation suited for right-skewed monetary data.

203
MCQeasy

In a simple linear regression model y = 2.5 + 1.2x, what is the predicted value of y when x = 10?

A.12.0
B.13.7
C.14.5
D.10.0
AnswerC

Substituting x = 10 into the fitted equation gives y = 2.5 + 1.2(10) = 2.5 + 12 = 14.5. The slope of 1.2 contributes twelve units across the ten-unit increase, and the intercept of 2.5 adds the baseline, yielding the predicted response.

Why this answer

Plug x=10: y = 2.5 + 1.2*10 = 2.5 + 12 = 14.5.

204
Multi-Selecthard

A data analyst is validating a regression model that predicts monthly churn probability for 50,000 subscribers. The analyst wants to detect whether the model is overfitting before deploying it. (Choose two.)

Select 2 answers
A.Remove all records with missing values from the validation set only
B.Increase the number of predictor variables until training accuracy reaches 100 percent
C.Apply a log transformation to the target churn probability before evaluation
D.Compare the model's error on the training set with its error on a held-out validation set
E.Use k-fold cross-validation to estimate performance across multiple data splits
AnswersD, E

A large gap between training error and validation error is the classic symptom of overfitting, because the model has memorized training noise that does not generalize. Comparing the two error values directly quantifies that gap and tells the analyst whether the model's complexity is justified, making this a core diagnostic step before deployment.

Why this answer

Overfitting is detected by measuring how well a model performs on data it did not train on. Holding out a validation set and comparing its error with training error reveals the generalization gap, while k-fold cross-validation repeats that comparison across multiple partitions to produce a stable estimate. Together they show whether the churn model has learned signal versus memorized noise, which is exactly what must be established before deployment.

Exam trap

The trap here is confusing actions that increase model complexity with techniques that actually diagnose overfitting.

205
Multi-Selectmedium

A data analyst wants to segment customers based on purchasing behavior such as frequency, monetary value, and recency. Which TWO clustering evaluation methods can help determine the optimal number of clusters? (Select two.)

Select 2 answers
A.Correlation coefficient
B.ANOVA
C.Silhouette score
D.t-test
E.Elbow method
AnswersC, E

The silhouette score measures how similar each point is to its own cluster versus the nearest other cluster, producing a coefficient between -1 and 1. The cluster count with the highest average silhouette indicates the best-separated segmentation.

Why this answer

The elbow method uses within-cluster sum of squares, and the silhouette score measures cohesion and separation. Both help choose k. Correlation coefficient is for association, not clustering.

ANOVA and t-test are for hypothesis testing.

206
MCQhard

A data analyst is asked to compare the average sales across three different store locations. The data is normally distributed and variances are approximately equal. Which statistical test is most appropriate?

A.ANOVA
B.Pearson correlation
C.Chi-square test
D.Two-sample t-test
AnswerA

ANOVA compares means across three or more groups simultaneously while controlling family-wise error. With normal distributions and roughly equal variances, its assumptions hold, making it correct for comparing average sales across the three store locations.

Why this answer

ANOVA (Analysis of Variance) is the correct test because it compares the means of three or more independent groups (store locations) while controlling the Type I error rate. The assumptions of normality and equal variances are satisfied, which are key prerequisites for ANOVA. A t-test would be inappropriate for more than two groups as it would require multiple pairwise comparisons, inflating the risk of false positives.

Exam trap

The trap here is confusing the number of groups: candidates might select a t-test for comparing means, but t-tests are limited to two groups; ANOVA is required for three or more.

How to eliminate wrong answers

Option B is wrong because Pearson correlation measures the linear relationship between two continuous variables, not differences in means across groups. Option C is wrong because the chi-square test is used for categorical data to test independence or goodness-of-fit, not for comparing means of continuous variables. Option D is wrong because a two-sample t-test only compares two groups; using it for three groups would require multiple tests, increasing the chance of a Type I error.

207
MCQhard

An analyst is performing a logistic regression to predict customer churn (yes/no). The model outputs a probability of 0.75 for a particular customer. Which of the following best describes the interpretation?

A.The model predicts that the customer will not churn
B.There is a 75% chance that the customer will churn
C.The customer will definitely churn because the probability is above 0.5
D.The odds of churning are 0.75 to 1
AnswerB

Logistic regression outputs a calibrated probability of the positive class, so 0.75 means a 75% estimated likelihood of churn for that customer. This satisfies the stem's interpretation requirement, distinguishing probabilistic output from a deterministic classification decision.

Why this answer

In logistic regression, the model outputs a probability between 0 and 1 that the positive class (churn = yes) occurs. A value of 0.75 means the model estimates a 75% probability that the customer will churn, given the input features. This is a probabilistic prediction, not a deterministic one, and it does not mean the customer will definitely churn.

Exam trap

DA0-002 often tests the confusion between probability and odds, and between probabilistic and deterministic predictions, causing candidates to misinterpret the output of logistic regression.

How to eliminate wrong answers

Option A is wrong because 0.75 is greater than 0.5, so the model predicts the positive class (churn), not the negative class. Option C is wrong because logistic regression outputs probabilities, not certainties; a 0.75 probability still implies a 25% chance of not churning, so 'definitely' is incorrect. Option D is wrong because 0.75 is a probability, not odds; odds would be calculated as p/(1-p) = 0.75/0.25 = 3, meaning 3 to 1 odds of churning, not 0.75 to 1.

208
Multi-Selectmedium

An analyst is conducting an A/B test on a new checkout process. To calculate sample size, which THREE factors must be considered?

Select 3 answers
A.Number of control groups
B.Desired effect size
C.Significance level (alpha)
D.Statistical power
E.Population standard deviation
AnswersB, C, D

Desired effect size is the minimum lift worth detecting between control and variant checkout processes. Smaller effects need larger samples, so effect size is a required input when calculating sample size for the A/B test.

Why this answer

Statistical power, significance level (alpha), and desired effect size (minimum detectable effect) are essential for sample size calculation.

209
Multi-Selecthard

A data analyst is performing data cleaning. Which THREE steps are part of this process? (Choose three.)

Select 3 answers
A.Correcting inconsistent data
B.Normalization
C.Handling missing values
D.Feature engineering
E.Removing duplicate records
AnswersA, C, E

Correcting inconsistent data resolves conflicting formats, units and values so records agree across sources. This is a core data cleaning activity, satisfying the stem's requirement to identify steps that standardise and reconcile raw data before analysis.

Why this answer

Data cleaning is the process of detecting and correcting (or removing) corrupt or inaccurate records from a dataset, so option A (Correcting inconsistent data) is correct because fixing mismatched formats, units, or values (e.g., 'NY' vs 'New York') is a core cleaning task. Option C (Handling missing values) is correct because cleaning must address nulls or blanks through imputation, deletion, or flagging so downstream analysis is not skewed. Option E (Removing duplicate records) is correct because duplicate rows inflate counts and distort aggregates, and deduplication is a standard cleaning step.

Option B (Normalization) is not part of cleaning; it is a data transformation/scaling technique (e.g., min-max or z-score) typically applied during preprocessing/modeling. Option D (Feature engineering) is also not cleaning; it is the creation of new derived variables for modeling, which occurs after cleaning.

Exam trap

The trap here is that candidates confuse data cleaning with data transformation or feature engineering, leading them to select normalization or feature engineering as cleaning steps, when in fact cleaning strictly addresses data quality issues like consistency, completeness, and uniqueness.

210
MCQeasy

Which statistical test should be used to determine if there is a significant association between two categorical variables, such as gender and product preference?

A.ANOVA
B.Chi-square test
C.Pearson correlation
D.t-test
AnswerB

The chi-square test of independence compares observed versus expected frequencies across a contingency table, determining whether two categorical variables are associated. Gender and product preference are both nominal categories, so this test satisfies the requirement for assessing significant association between them.

Why this answer

The chi-square test is specifically designed to test for independence between two categorical variables, such as gender (male/female) and product preference (e.g., product A/B/C). It compares observed frequencies in each category combination to the frequencies expected if the variables were independent, using the chi-square statistic. A significant result indicates that the variables are associated, not independent.

Exam trap

The trap here is confusing tests for means (t-test, ANOVA) with tests for association between categorical variables, or mixing up correlation (for continuous variables) with chi-square (for categorical variables).

How to eliminate wrong answers

Option A is wrong because ANOVA compares means of a continuous dependent variable across three or more groups defined by a categorical independent variable, not associations between two categorical variables. Option C is wrong because Pearson correlation measures the linear relationship between two continuous variables, not categorical ones. Option D is wrong because a t-test compares means between two groups (or one group against a known value) for a continuous outcome, not categorical associations.

211
MCQhard

A data scientist runs a linear regression model to predict customer spending based on income. The R-squared value is 0.45 and the p-value for the slope coefficient is 0.03. At a significance level of α=0.05, which of the following conclusions is correct?

A.The slope is not statistically significant, and the model explains 55% of the variance.
B.The slope is statistically significant, and the model explains 45% of the variance.
C.The slope is statistically significant, and the model explains 55% of the variance.
D.The slope is not statistically significant, and the model explains 45% of the variance.
AnswerB

With α=0.05, the slope p-value of 0.03 falls below the threshold, so the slope is statistically significant. R-squared of 0.45 means income explains 45% of the variance in spending. Both conditions in the stem are satisfied simultaneously.

Why this answer

The p-value (0.03) is less than α (0.05), so the slope is statistically significant. R²=0.45 means the model explains 45% of the variance.

212
MCQeasy

A data analyst is building a linear regression model to predict sales based on advertising spend. The analyst notices that the residuals are not normally distributed and have a non‑constant variance. Which of the following transformations is most appropriate to apply to the dependent variable?

A.Standardization (z-score)
B.Normalization (min-max scaling)
C.Logarithmic transformation
D.Square root transformation
AnswerC

Non-constant variance and non-normal residuals violate linear regression assumptions. A logarithmic transformation of the dependent variable compresses the scale of large values, stabilising variance and pulling the residual distribution towards normality, which is the standard remedy for this pattern of heteroscedasticity.

Why this answer

The logarithmic transformation is the most appropriate choice because it stabilizes non‑constant variance (heteroscedasticity) and helps make the residuals more normally distributed, which are key assumptions for linear regression. By compressing the scale of the dependent variable (sales), it reduces the impact of large values and often linearizes multiplicative relationships, such as diminishing returns from advertising spend.

Exam trap

CompTIA often tests the misconception that any scaling technique (standardization or normalization) can fix heteroscedasticity or non‑normality, but these methods only change the range or center of the data, not the shape of the residual distribution or the variance structure.

How to eliminate wrong answers

Option A is wrong because standardization (z-score) centers and scales the data to mean 0 and standard deviation 1, but it does not address heteroscedasticity or non‑normal residuals; it merely changes the units of the dependent variable without altering the shape of the distribution. Option B is wrong because normalization (min-max scaling) rescales the data to a fixed range (e.g., 0 to 1), which also fails to correct non‑constant variance or non‑normality; it is primarily used for feature scaling in algorithms like neural networks, not for satisfying regression assumptions. Option D is wrong because the square root transformation is typically used for count data (e.g., Poisson-distributed outcomes) to stabilize variance, but it is less effective than the log transformation when the variance increases proportionally with the mean, which is common in sales data; the log transformation is the standard choice for multiplicative relationships and heteroscedasticity.

213
MCQeasy

Which data quality dimension is most concerned with whether data values fall within a defined domain or acceptable range?

A.Completeness
B.Consistency
C.Validity
D.Accuracy
AnswerC

Validity checks that values conform to a defined domain, format or acceptable range, directly matching the stem's requirement. Unlike accuracy, which concerns correspondence to real-world truth, or consistency, which concerns agreement across sources, validity enforces permitted value sets and boundaries, so out-of-range or malformed entries fail the check.

Why this answer

Validity refers to whether data values conform to defined rules or constraints.

214
MCQhard

A data scientist applies K-means clustering to a customer dataset. The elbow method suggests using 4 clusters. After running K-means with k=4, the within-cluster sum of squares (WCSS) is plotted against k, and the elbow is at k=4. What does this indicate?

A.Increasing k beyond 4 would not significantly reduce WCSS.
B.The data naturally forms 4 clusters with no noise.
C.The algorithm converged to a local minimum.
D.The model has overfit the data.
AnswerA

The elbow marks where additional clusters stop yielding meaningful WCSS reduction. At k=4 the curve flattens, so moving to k=5 or beyond adds complexity without materially lowering within-cluster sum of squares, confirming four clusters as the sensible choice.

Why this answer

The elbow method plots the within-cluster sum of squares (WCSS) against the number of clusters k. The 'elbow' point indicates where the rate of decrease in WCSS sharply changes, meaning that adding more clusters beyond that point yields diminishing returns. Therefore, at k=4, increasing k further would not significantly reduce WCSS, suggesting 4 is a reasonable choice for the number of clusters.

Exam trap

The trap is interpreting the elbow as a definitive indication of the true number of clusters; candidates must remember that the elbow method is a heuristic and does not guarantee that the data naturally forms that many clusters.

How to eliminate wrong answers

Option B is wrong because the elbow method does not guarantee that the data naturally forms exactly 4 clusters with no noise; it only suggests a good trade-off between model complexity and fit. Option C is wrong because the elbow method is about choosing k, not about convergence to a local minimum; K-means can converge to a local minimum regardless of k. Option D is wrong because overfitting would be indicated by a very low WCSS with many clusters, but the elbow method helps avoid overfitting by selecting a simpler model.

215
MCQmedium

A data analyst is examining the relationship between advertising spend (in thousands) and sales (in thousands). The Pearson correlation coefficient is computed as r = -0.85. Which of the following interpretations is correct?

A.There is no linear relationship.
B.There is a strong positive linear relationship between advertising spend and sales.
C.There is a weak negative linear relationship.
D.There is a strong negative linear relationship.
AnswerD

The magnitude of r, 0.85, exceeds the conventional 0.7 threshold for strong association, and the negative sign indicates that as advertising spend rises, sales tend to fall along a linear trend. Direction and strength are therefore both captured by this interpretation.

Why this answer

Pearson r measures linear correlation: -0.85 indicates a strong negative linear relationship (as one increases, the other decreases). The magnitude |0.85| is close to 1, so strong.

216
MCQhard

A data analyst is evaluating a classification model for predicting customer churn. The model's confusion matrix shows 100 true positives, 20 false positives, 30 false negatives, and 850 true negatives. The analyst wants to assess the model's ability to correctly identify actual churners. Which metric should the analyst use?

A.Precision
B.Accuracy
C.Specificity
D.Recall
AnswerD

Recall, also known as sensitivity, is the proportion of actual positives that are correctly identified, calculated as TP/(TP+FN) = 100/130 ≈ 0.769. It measures the model's ability to find all actual churners, which aligns with the analyst's goal. A high recall means few churners are missed, which is often critical in churn prediction to avoid losing customers.

Why this answer

The analyst wants to assess the model's ability to correctly identify actual churners, which is the definition of recall (sensitivity). Recall is calculated as true positives divided by the sum of true positives and false negatives. In this scenario, recall is 100/(100+30) ≈ 0.769.

Precision and specificity focus on different aspects, and accuracy can be misleading in imbalanced datasets.

Exam trap

The trap here is confusing recall with precision; recall measures how many actual positives were found, while precision measures how many predicted positives were correct.

217
MCQhard

A data analyst trains a complex model that achieves 99% accuracy on training data but only 65% on new data. What is the most likely issue?

A.Underfitting
B.Overfitting
C.Multicollinearity
D.High bias
AnswerB

Overfitting occurs when a model learns noise and idiosyncrasies in the training set rather than generalisable patterns, producing the 99% versus 65% gap. The stem's constraint — high training accuracy with poor unseen-data performance — is the defining signature of variance-dominated overfitting, so regularisation, pruning or more data would be required.

Why this answer

The model performs exceptionally well on training data (99% accuracy) but poorly on new data (65% accuracy), which is the classic symptom of overfitting. Overfitting occurs when the model learns noise and specific patterns in the training data rather than generalizing to unseen data, often due to excessive complexity (e.g., too many parameters or deep layers). This results in high variance and poor performance on validation or test sets.

Exam trap

CompTIA often tests the distinction between overfitting and underfitting by presenting a large gap between training and test accuracy, tempting candidates to choose high bias or multicollinearity due to confusion about bias-variance tradeoff or correlation issues.

How to eliminate wrong answers

Option A is wrong because underfitting would show poor performance on both training and new data (e.g., low accuracy on both), not high training accuracy with low test accuracy. Option C is wrong because multicollinearity refers to high correlation among predictor variables in regression models, which inflates coefficient standard errors but does not directly cause a large gap between training and test accuracy. Option D is wrong because high bias typically leads to underfitting, where the model is too simple and performs poorly on both training and test data, not the specific pattern of high training accuracy and low test accuracy seen here.

218
MCQeasy

A data analyst is cleaning a dataset and finds that the 'age' column has several missing values. Which of the following is a valid method for handling missing numerical data?

A.Delete the entire column
B.Ignore the missing values
C.Impute with the mean
D.Replace with zeros
AnswerC

Imputing with the mean replaces each missing 'age' entry with the column's arithmetic average, preserving the existing sample size and keeping the dataset's overall mean unchanged. This satisfies the stem's requirement for a valid numerical handling method, since the mean is calculable only on quantitative data such as age.

Why this answer

Imputing missing numerical values with the mean is a standard and valid statistical approach for handling missing data in a numerical column. It preserves the overall distribution's central tendency and allows the analyst to retain all other rows for analysis. This method is especially appropriate when the missingness is random and the column is roughly normally distributed, minimizing bias introduced by dropping records.

Exam trap

The trap here is confusing 'ignoring' missing values with a valid handling method, or assuming that any replacement (like zeros) is acceptable without considering the data's meaning.

How to eliminate wrong answers

Option A is wrong because deleting the entire column discards all data in that feature, including the non-missing values, which is rarely justified unless the column is entirely unusable. Option B is wrong because ignoring missing values is not a handling method; many algorithms and calculations will fail or produce incorrect results with NaN values, and it does not resolve the missingness. Option D is wrong because replacing missing ages with zeros introduces a false value (age 0) that is not a valid age for the dataset, severely distorting the mean, variance, and any downstream analysis.

219
MCQhard

A data analyst is using a decision tree to predict customer churn. The tree is very deep, with many branches, and achieves 99% accuracy on the training data but only 70% on the test data. Which technique should the analyst use to address this issue?

A.Use a larger training dataset
B.Decrease the size of the test set
C.Increase the maximum depth of the tree
D.Apply pruning to the decision tree
AnswerD

Pruning reduces the size of the tree by removing branches that provide little predictive power, which helps combat overfitting. By simplifying the model, pruning improves generalization to unseen data. In this scenario, the large gap between training and test accuracy indicates overfitting, so pruning is the appropriate technique to enhance test performance.

Why this answer

Pruning is a technique specifically designed to reduce overfitting in decision trees by removing unnecessary branches. The large discrepancy between training and test accuracy indicates the model is too complex and has memorized the training data. Pruning simplifies the tree, improving its ability to generalize.

Other options either worsen overfitting or do not directly address the model complexity.

Exam trap

The trap here is thinking that more data or a deeper tree will solve overfitting, when the real issue is model complexity that requires pruning.

220
MCQeasy

A data analyst is summarizing the central tendency of a dataset with extreme outliers. Which measure is most robust to outliers?

A.Standard deviation
B.Median
C.Mean
D.Range
AnswerB

The median is positional, not arithmetic: it depends only on the middle value's rank, so extreme outliers cannot drag it. The mean incorporates every value's magnitude, making it highly sensitive to skew. For skewed data, the median better represents typical central tendency.

Why this answer

The median is the middle value when data is ordered, so extreme outliers (whether very high or very low) do not affect its position. Unlike the mean, which incorporates every value and can be pulled strongly by outliers, the median remains stable. Therefore, for a dataset with extreme outliers, the median is the most robust measure of central tendency.

Exam trap

The trap here is confusing measures of central tendency with measures of dispersion, and assuming that the mean is always the best measure of center. Candidates often pick the mean out of habit, ignoring its sensitivity to outliers.

How to eliminate wrong answers

Option A is wrong because standard deviation is a measure of dispersion, not central tendency, and it is highly sensitive to outliers since it squares deviations from the mean. Option C is wrong because the mean is the arithmetic average and is directly influenced by every value, so extreme outliers can skew it dramatically. Option D is wrong because the range is the difference between the maximum and minimum values, making it entirely dependent on outliers and not a measure of central tendency at all.

221
MCQmedium

A data analyst is cleaning a dataset and finds that a numeric field has several missing values. The variable is normally distributed. Which imputation method is most appropriate?

A.Median imputation
B.Mean imputation
C.Mode imputation
D.Forward-fill
AnswerB

For a normally distributed variable, the mean preserves the central tendency and keeps the overall distribution shape intact, unlike median or mode imputation. Since the data is symmetric, the mean is the best unbiased estimate for replacing missing numeric entries, satisfying the normality constraint in the stem.

Why this answer

For a normally distributed numeric variable, mean imputation is the most appropriate method because the mean is the central tendency that best represents the typical value in a symmetric distribution. Replacing missing values with the mean preserves the overall mean of the variable and is statistically sound when the data is normally distributed. Other methods like median or mode are better for skewed or categorical data, respectively.

Exam trap

The trap is that candidates might choose median imputation thinking it's always more robust, but for a normal distribution, mean imputation is the standard; the question specifies normal distribution to guide you to the mean.

How to eliminate wrong answers

Option A is wrong because median imputation is more appropriate for skewed distributions or when outliers are present, as the median is robust to extreme values; for a normal distribution, the mean is preferred. Option C is wrong because mode imputation is used for categorical data or discrete variables, not continuous numeric variables; it would be inappropriate here. Option D is wrong because forward-fill is used for time-series data to carry forward the last observed value, which assumes a temporal order and is not suitable for a general numeric variable without such structure.

222
MCQeasy

A data analyst calculates the mean, median, and mode of a dataset. Which of the following measures of central tendency is least affected by extreme outliers?

A.Median
B.Range
C.Mode
D.Mean
AnswerA

The median depends only on positional order, so extreme values shift it minimally. The mean sums every value and is dragged strongly by outliers, while the mode reflects only the most frequent value and can be unstable in continuous data.

Why this answer

The median is the middle value when data is ordered, so it depends only on the rank position of values, not their magnitude. Extreme outliers, even if arbitrarily large or small, do not change which value sits in the middle unless they cross the 50th percentile. This makes the median a robust measure of central tendency, unlike the mean, which incorporates every value and is pulled toward outliers.

Exam trap

The trap here is confusing measures of central tendency with measures of dispersion, leading candidates to select 'Range' because it sounds like a statistical measure, or to assume the mode is always the most robust since it ignores magnitudes—when in fact the median is the standard robust measure of center tested on DA0-002.

How to eliminate wrong answers

Option B is wrong because the range is not a measure of central tendency at all—it is a measure of dispersion calculated as max minus min, and it is extremely sensitive to outliers since a single extreme value directly sets the maximum or minimum. Option C is wrong because the mode is the most frequent value and, while not affected by the magnitude of outliers, it can be unstable or uninformative in continuous data and is not considered a robust measure of center in the same way; moreover, an outlier could theoretically become the mode if it repeats, but more importantly, the question asks for the least affected measure among central tendency options, and the median is the standard robust choice. Option D is wrong because the mean sums all values and divides by n, so a single extreme outlier shifts the mean substantially, making it the most outlier-sensitive measure of central tendency.

223
MCQhard

A data analyst uses the elbow method to determine the number of clusters for k-means. The plot shows a sharp bend at k=3 and a small bend at k=5. What is the recommended number of clusters?

A.5
B.The method is inconclusive.
C.2
D.3
AnswerD

The elbow method selects the k where the within-cluster sum of squares drops sharply before levelling off. The sharp bend at k=3 marks that inflection point, so three clusters is recommended; the minor bend at k=5 is a secondary, weaker signal.

Why this answer

The elbow method suggests choosing k where the decrease in inertia becomes marginal; the sharp bend at 3 indicates the optimal k.

224
MCQhard

A data analyst is building a decision tree to classify whether customers will churn. The analyst wants to prevent the tree from overfitting the training data. Which technique should the analyst use?

A.Increasing the maximum depth of the tree
B.Setting the minimum samples per leaf to 1
C.Pruning the tree after full growth
D.Using all available features without selection
AnswerC

Pruning involves growing a full tree and then removing branches that provide little predictive power, reducing complexity and overfitting. This technique, such as cost-complexity pruning, balances bias and variance, improving generalization. For churn classification, pruning helps avoid capturing noise in the training data, leading to more robust predictions on new customers.

Why this answer

Pruning is a standard technique to reduce decision tree overfitting by removing branches that do not improve predictive accuracy on validation data. It simplifies the model and enhances generalization. Increasing depth, using all features, or allowing single-sample leaves all increase complexity and overfitting risk, making them incorrect choices for this scenario.

Exam trap

The trap here is thinking that a fully grown tree is always best, but without pruning it will overfit and perform poorly on new data.

225
Multi-Selectmedium

A data analyst is performing hypothesis testing to compare the mean sales of two store locations. Which TWO conditions must be satisfied to use a two‑sample t‑test? (Select TWO.)

Select 2 answers
A.The data is paired between the two locations
B.The sample sizes are equal
C.The data is approximately normally distributed
D.The variances of the two populations are equal
E.The two samples are independent of each other
AnswersC, E

Normality is assumed for the t-test, though it is robust for large samples.

Why this answer

The two-sample t-test assumes that the data in each group are approximately normally distributed. This is a key parametric assumption; if the sample sizes are large (typically n > 30), the Central Limit Theorem can relax this requirement, but for smaller samples, normality must hold to ensure valid test statistics and p-values.

Exam trap

CompTIA often tests the misconception that equal sample sizes or equal variances are required for a two-sample t-test, but the actual core assumptions are independence and normality (or large sample sizes via CLT).

← PreviousPage 3 of 4 · 238 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Data Analysis questions.