Courseiva

CCNA Data Analysis Questions

13 of 238 questions · Page 4/4 · Data Analysis · Answers revealed

226
Multi-Selecthard

A data analyst is preparing a dataset for a machine learning model and notices that several numerical features have different scales (e.g., age ranges from 18-90, income from 20,000-200,000). The analyst decides to apply feature scaling. Which TWO of the following statements about scaling techniques are correct? (Choose two.)

Select 2 answers
A.Standardization (Z-score normalization) transforms features to have a mean of 0 and a standard deviation of 1, and is less affected by outliers than min-max scaling.
B.Min-max normalization transforms features to a fixed range, typically [0, 1], and is sensitive to outliers.
C.Feature scaling is unnecessary for distance-based algorithms like k-nearest neighbors (KNN) because they are scale-invariant.
D.Standardization scales features to a range of [0, 1] and is preferred when features have different units.
E.Min-max normalization preserves the shape of the original distribution and is robust to outliers.
AnswersA, B

Standardization subtracts the mean and divides by the standard deviation, resulting in a distribution with mean 0 and standard deviation 1. It does not bound values to a specific range, so outliers have less impact on the scaling of the majority of data compared to min-max normalization.

Why this answer

Min-max normalization scales to a fixed range but is sensitive to outliers, while standardization centers data with unit variance and is less affected by outliers. Both are common scaling techniques. The other statements misrepresent properties of scaling or incorrectly claim that distance-based algorithms are scale-invariant.

Exam trap

The trap here is confusing the properties of min-max normalization and standardization, especially regarding outlier sensitivity and output range.

227
MCQmedium

An analyst is performing a linear regression and obtains an R-squared value of 0.85. Which of the following is the best interpretation?

A.85% of the residuals are zero.
B.85% of the data points lie on the regression line.
C.There is an 85% chance that the relationship is causal.
D.The model explains 85% of the variability in the dependent variable.
AnswerD

R-squared measures the proportion of variance in the dependent variable accounted for by the regression model. A value of 0.85 therefore means the predictors explain 85% of that variability, leaving 15% unexplained by the model.

Why this answer

R-squared (R²) is the coefficient of determination, representing the proportion of variance in the dependent variable that is predictable from the independent variable(s). An R² of 0.85 means that 85% of the variability in the dependent variable is explained by the regression model, indicating a strong fit. This is the standard statistical interpretation.

Exam trap

DA0-002 often tests the interpretation of R-squared, and candidates frequently confuse it with correlation, causation, or the proportion of data points on the line.

How to eliminate wrong answers

Option A is wrong because R-squared does not measure the proportion of residuals that are zero — residuals are the differences between observed and predicted values, and R² does not imply any specific number of zero residuals. Option B is wrong because R-squared does not indicate the percentage of data points that lie exactly on the regression line; it measures explained variance, not exact fit. Option C is wrong because R-squared does not imply causation — correlation does not equal causation, and R² is a measure of association, not causal inference.

228
Multi-Selecthard

A data analyst is preparing a dataset for a machine learning model to predict customer churn. The dataset contains missing values in several columns. The analyst wants to handle missing data in a way that preserves as much information as possible and avoids introducing bias. Which TWO of the following techniques are appropriate? (Choose two.)

Select 2 answers
A.Listwise deletion
B.K-nearest neighbors (KNN) imputation
C.Mean imputation for numerical features
D.Replacing missing values with zero
E.Multiple imputation by chained equations (MICE)
AnswersB, E

KNN imputation estimates missing values by finding the most similar complete observations and averaging their values. It leverages correlations between features and can preserve local data structure. This method is non-parametric and can handle both numerical and categorical data with appropriate distance metrics. It is suitable for churn prediction as it maintains information and reduces bias compared to simple imputation.

Why this answer

Multiple imputation by chained equations and KNN imputation are both advanced techniques that preserve data structure and reduce bias. MICE models each variable with missing data using other variables and creates multiple imputations to account for uncertainty. KNN imputation uses similarity between observations to estimate missing values.

Both are superior to mean imputation, listwise deletion, or zero replacement, which either discard information or introduce bias.

Exam trap

The trap here is assuming that any imputation method is fine, but simple methods like mean imputation or zero replacement can bias the model, while listwise deletion reduces sample size and may introduce bias.

229
MCQhard

A data scientist is building a model to predict customer churn (yes/no). After training a logistic regression model, the coefficient for 'monthly charges' is 0.05 with a p-value of 0.03. Which interpretation is correct at α=0.05?

A.The model's R-squared is 0.05.
B.For every unit increase in monthly charges, the odds of churn increase by about 5%.
C.Monthly charges decrease the probability of churn.
D.Monthly charges have no significant effect on churn.
AnswerB

A one-unit rise in monthly charges multiplies the odds of churn by e^0.05 ≈ 1.051, an increase of roughly 5%. The p-value of 0.03 falls below α=0.05, so the coefficient is statistically significant and the predictor's effect on churn odds is supported.

Why this answer

In logistic regression, coefficients are in log-odds units. A coefficient of 0.05 means that for a one-unit increase in monthly charges, the log-odds of churn increase by 0.05. Exponentiating gives e^0.05 ≈ 1.051, so the odds increase by about 5.1%.

The p-value of 0.03 is less than α=0.05, so the effect is statistically significant.

Exam trap

The trap is misinterpreting the coefficient as a direct probability change or confusing it with R-squared; candidates may also ignore the p-value and incorrectly claim no significance.

How to eliminate wrong answers

Option A is wrong because R-squared is not derived from a coefficient value; logistic regression typically uses pseudo-R-squared measures, and 0.05 here is the coefficient, not R-squared. Option C is wrong because the coefficient is positive, meaning higher monthly charges are associated with higher odds of churn, not lower. Option D is wrong because the p-value of 0.03 is below 0.05, so the effect is statistically significant, not insignificant.

230
Drag & Dropmedium

Drag and drop the steps to normalize a database table from 1NF to 3NF in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Normalization proceeds from 1NF to 2NF to 3NF, then table creation and foreign keys.

231
Multi-Selecthard

A data analyst is performing K-means clustering on customer data. Which THREE of the following are steps in the K-means algorithm?

Select 3 answers
A.Perform eigenvalue decomposition.
B.Calculate the correlation matrix.
C.Initialize k centroids randomly.
D.Update centroids by computing the mean of all points assigned to each centroid.
E.Assign each data point to the nearest centroid.
AnswersC, D, E

K-means begins by choosing k initial centroids, typically at random positions in the feature space, before any assignment occurs. This initialisation step satisfies the stem's requirement to identify actual algorithm steps, establishing the starting points that subsequent assignment and update iterations refine.

Why this answer

Option C is correct because the standard K-means algorithm begins by initializing k centroids, typically at random positions in the feature space, before any assignment occurs. Option E is correct because after initialization, each data point is assigned to the nearest centroid, usually measured by Euclidean distance, forming the initial clusters. Option D is correct because the algorithm then recomputes each centroid as the mean (average) of all points assigned to it, and this assignment-and-update cycle repeats until convergence.

Options A and B are not steps in K-means: eigenvalue decomposition is used in dimensionality-reduction techniques such as PCA, and calculating a correlation matrix is a preprocessing or exploratory analysis step, not part of the iterative K-means procedure.

Exam trap

The trap here is confusing K-means with other algorithms like PCA or correlation-based clustering, leading candidates to select eigenvalue decomposition or correlation matrix calculation as steps.

232
MCQeasy

A dataset contains a column 'Age' with values: [22, 25, 25, 30, 35, 40, 45]. What is the interquartile range (IQR)?

A.15
B.10
C.20
D.25
AnswerA

Ordered values give Q1 at 25 and Q3 at 40, so IQR equals 40 minus 25, which is 15. This satisfies the stem by measuring the spread of the middle half of the dataset, excluding the lowest and highest values.

Why this answer

The interquartile range (IQR) is Q3 minus Q1. For the dataset [22, 25, 25, 30, 35, 40, 45], the median is 30. Q1 is the median of the lower half [22, 25, 25] = 25.

Q3 is the median of the upper half [35, 40, 45] = 40. IQR = 40 - 25 = 15.

Exam trap

DA0-002 often tests quartile calculation with odd-numbered datasets — the trap is including the median in both halves, which shifts Q1 and Q3 and produces an incorrect IQR.

How to eliminate wrong answers

Option B (10) is wrong because it would result from miscalculating Q1 or Q3 — for example, using 30 as Q1 and 40 as Q3, or misidentifying quartile positions. Option C (20) is wrong because it likely comes from subtracting the minimum from the median or miscomputing Q3 as 45 and Q1 as 25. Option D (25) is wrong because it equals Q1 itself or results from subtracting the minimum (22) from Q3 (40) rounded, which is not the IQR definition.

233
MCQmedium

A data scientist is preparing data for a K-means clustering algorithm. The dataset contains features measured in different units (e.g., income in dollars and age in years). Which preprocessing step is most critical before running K-means?

A.Remove outliers
B.Encode categorical variables
C.Standardize or normalize the features
D.Perform feature selection
AnswerC

K-means relies on Euclidean distance, so features in different units let larger-scale variables such as income dominate the distance calculation. Standardising or normalising puts every feature on a comparable scale, satisfying the requirement for meaningful cluster assignment.

Why this answer

K-means clustering uses Euclidean distance to measure similarity between data points. If features are on different scales (e.g., income in dollars vs. age in years), the feature with the larger range will dominate the distance calculation, leading to biased clusters. Standardizing (z-score normalization) or normalizing (min-max scaling) the features ensures that all features contribute equally to the distance metric, which is critical for K-means to produce meaningful clusters.

Exam trap

The trap is thinking that removing outliers or encoding categorical variables is the most critical step, but the question specifically highlights features in different units, which directly points to scaling. Candidates might also confuse feature selection with preprocessing necessity.

How to eliminate wrong answers

Option A is wrong because while removing outliers can be beneficial, it is not as critical as scaling; outliers can affect clustering but the primary issue is feature scale. Option B is wrong because encoding categorical variables is necessary if the dataset contains categorical features, but the question specifies features measured in different units (numerical), so encoding is not the most critical step. Option D is wrong because feature selection can improve performance but does not address the scale disparity that directly impacts K-means distance calculations.

234
MCQmedium

A data scientist is building a predictive model to forecast monthly sales. The data shows a linear trend with no seasonality. Which regression technique is most appropriate?

A.Polynomial regression
B.Logistic regression
C.Linear regression
D.Ridge regression
AnswerC

Linear regression models a straight-line relationship between predictors and a continuous target, matching data with a linear trend and no seasonality. Seasonal techniques such as SARIMA or decomposition would add unnecessary parameters, while tree ensembles forgo the interpretable linear fit the pattern warrants.

Why this answer

Linear regression is the most appropriate technique because the data shows a linear trend with no seasonality, making a straight-line model the simplest and most effective fit. It directly models the relationship between the independent variable (e.g., time) and the dependent variable (monthly sales) using a linear equation, minimizing the sum of squared residuals.

Exam trap

The trap here is that candidates often confuse 'linear trend' with 'linear in parameters' and incorrectly choose polynomial regression, thinking it adds flexibility, when the question explicitly states no seasonality and a linear trend, making simple linear regression the optimal choice.

How to eliminate wrong answers

Option A is wrong because polynomial regression introduces higher-degree terms (e.g., x², x³) to model curvature, which is unnecessary and risks overfitting when the trend is explicitly linear. Option B is wrong because logistic regression is used for binary classification problems (e.g., predicting yes/no outcomes), not for forecasting continuous numeric values like monthly sales. Option D is wrong because ridge regression is a regularization technique designed to handle multicollinearity or overfitting by adding an L2 penalty, but it is not a distinct regression type for linear trends and would be overkill when a simple linear model suffices.

235
MCQmedium

In a logistic regression model predicting customer churn (1 = churn, 0 = not churn), the coefficient for 'contract length' is -0.5. Which of the following is the correct interpretation?

A.For each unit increase in contract length, the log-odds of churn decrease by 0.5.
B.Longer contract length increases the odds of churn.
C.The probability of churn decreases by 50% for each unit increase in contract length.
D.Contract length is not a significant predictor.
AnswerA

In logistic regression the coefficient operates on the log-odds scale, so a -0.5 coefficient means each one-unit rise in contract length lowers the log-odds of churn by 0.5, holding other predictors constant. This directly satisfies the stem's request for the correct interpretation of that coefficient.

Why this answer

In logistic regression, the coefficient represents the change in log-odds of the outcome for a one-unit increase in the predictor, holding other variables constant. A coefficient of -0.5 means that for each unit increase in contract length, the log-odds of churn decrease by 0.5. This is the correct interpretation of the model's output.

Exam trap

DA0-002 often tests the confusion between log-odds and probability, or between odds and probability, leading candidates to incorrectly interpret coefficients as direct probability changes.

How to eliminate wrong answers

Option B is wrong because a negative coefficient indicates a decrease in log-odds, not an increase. Option C is wrong because it misinterprets the coefficient as a direct probability change; the change in probability is not constant and depends on the current probability. Option D is wrong because the coefficient value does not determine significance; significance is assessed via p-values or confidence intervals, which are not provided.

236
Multi-Selectmedium

A data analyst is preparing a dataset for analysis and needs to handle outliers. Which TWO of the following are common methods for treating outliers?

Select 2 answers
A.Removal
B.Capping
C.Normalization
D.Imputation
E.Standardization
AnswersA, B

Removal deletes or filters out records whose values fall beyond a defined threshold, such as three standard deviations or the interquartile range fence. This satisfies the need to treat outliers by eliminating their distortion, though it reduces sample size and risks discarding legitimate extreme observations.

Why this answer

Removal (A) is a common outlier treatment because it deletes the extreme data points from the dataset, which is appropriate when outliers are errors or when their influence must be eliminated before analysis. Capping (B), also called winsorizing, is a common treatment that replaces extreme values with a defined threshold, such as the 1st/99th percentile or a value at a set number of standard deviations, preserving the record while limiting the outlier's effect. Normalization (C) is a scaling technique that rescales values to a fixed range like 0–1; it does not identify or treat outliers, so it is not a treatment method.

Imputation (D) fills in missing values rather than addressing extreme values, so it is not an outlier treatment. Standardization (E) transforms data to have mean 0 and standard deviation 1, which is also a scaling method and does not by itself treat outliers.

Exam trap

The trap is confusing data scaling techniques (normalization, standardization) with outlier treatment methods; candidates must distinguish between transforming the scale and actually handling extreme values.

237
MCQmedium

A healthcare analytics team is analyzing patient readmission rates. They have a dataset with thousands of records including patient age, diagnosis, length of stay, number of prior admissions, and discharge date. The goal is to identify key factors influencing readmission and create a model to predict high-risk patients. The data is imbalanced: only 5% of patients are readmitted within 30 days. The team plans to use logistic regression. What is the most appropriate approach?

A.Use the dataset as is because logistic regression handles imbalance
B.Remove most of the non-readmitted patients to balance the dataset
C.Use accuracy as the evaluation metric
D.Apply oversampling techniques like SMOTE to the training set
AnswerD

With only 5% readmissions, logistic regression would bias toward the majority class. SMOTE synthesises minority-class examples in the training set, balancing class distribution so the model learns readmission patterns rather than predicting 'no readmission' constantly.

Why this answer

With imbalanced data, logistic regression can be biased toward the majority class. Oversampling the minority class (e.g., SMOTE) helps the model learn patterns for readmission. Using accuracy as a metric would be misleading.

Removing majority samples discards valuable data. Using data as-is often fails to predict the minority class.

238
MCQeasy

A data analyst at a logistics company receives a shipment-tracking table in which the carrier identifier column contains values such as 'FEDX', 'FedEx', 'fed ex', and 'FX'. The analyst needs to consolidate these into a single carrier value before building a delivery-performance dashboard. Which data preparation step should the analyst perform?

A.One-hot encode the carrier column into binary indicator fields.
B.Standardize and normalize the categorical values through data cleansing.
C.Impute the missing carrier values using the column mean.
D.Apply min-max normalization to rescale the carrier column.
AnswerB

These entries are the same carrier recorded inconsistently, so the analyst should cleanse the column by trimming whitespace, unifying case, and mapping known aliases to one canonical value. Consolidating variants into a single label prevents the dashboard from splitting one carrier's shipments across four bogus categories, which would distort on-time rates and volume comparisons.

Why this answer

The carrier column suffers from inconsistent representation of the same entity, a classic data cleansing issue. Trimming whitespace, normalizing case, and mapping aliases to a canonical label collapses the variants into one category. Numeric rescaling, missing-value imputation, and one-hot encoding all address different problems and would leave the duplicated carrier categories intact in the delivery dashboard.

Exam trap

The trap here is reaching for a transformation like one-hot encoding or normalization because the column looks messy, when the actual defect is inconsistent labels that only cleansing and standardization can merge.

← PreviousPage 4 of 4 · 238 questions total

Ready to test yourself?

Try a timed practice session using only Data Analysis questions.