Courseiva

CompTIA Data+ (DA0-002) (DA0-002) — Questions 301–375

1004 questions total · 14pages · All types, answers revealed

Page 4

Page 5 of 14

Page 6
301
MCQhard

The exhibit shows a SQL query result intended for a bar chart of revenue by region. However, the chart shows only the top 10 regions, but the query returns all regions. What is the most likely cause?

A.The GROUP BY clause is incorrect
B.The visualization tool has a default limit on the number of categories displayed
C.The query is missing a WHERE clause
D.The ORDER BY clause is ignored in the chart
AnswerB

A default category cap in the visualisation tool truncates the axis to the first ten regions, leaving the remaining rows unrendered despite the query returning them. This satisfies the stem's constraint that the chart shows only the top 10 regions while the underlying SQL result contains all regions, so the query itself is not at fault.

Why this answer

The SQL query itself returns all regions because there is no LIMIT clause. However, the visualization tool has a built-in default limit on the number of categories displayed, such as top 10, which truncates the data in the chart. This is the most likely cause, making Option B correct.

Options A, C, and D are incorrect: the GROUP BY clause is correctly specified, a WHERE clause is not required to get all regions, and the ORDER BY clause may be applied for sorting but the tool's limit overrides the full result set.

302
MCQmedium

A manager in operations needs a real-time dashboard showing production line status, including machine uptime and error counts. Which type of report is most appropriate?

A.Analytical report
B.Ad hoc report
C.Scheduled report
D.Operational report
AnswerD

Operational reports track day-to-day activity in real time, covering metrics such as machine uptime and error counts. This matches the manager's need for a live production line dashboard, unlike tactical or strategic reports, which address medium- and long-term planning.

Why this answer

An operational report is designed to support day-to-day monitoring of ongoing business processes, providing near real-time visibility into metrics like machine uptime and error counts. It is the correct choice for a production line dashboard that operations managers check continuously. Operational reports focus on current status rather than historical trends or deep analysis.

Exam trap

DA0-002 often tests the confusion between 'operational' and 'analytical' reports — candidates pick analytical because it sounds more sophisticated, missing that real-time monitoring is the defining trait of operational reporting.

How to eliminate wrong answers

Option A is wrong because analytical reports focus on historical trends, patterns, and root-cause analysis rather than real-time operational status. Option B is wrong because ad hoc reports are one-off, user-initiated queries for a specific question, not continuous dashboards. Option C is wrong because scheduled reports run at fixed intervals (daily, weekly) and are not designed for real-time monitoring.

303
MCQmedium

A data analyst must build a one-page dashboard for a hospital's bed-management team. The team needs to know, at a glance, how many beds are free right now, how that compares with the same hour yesterday, and whether the trend is worsening. Which combination of visual elements is MOST appropriate for the top band of this dashboard?

A.A single large number (KPI card) for current free beds, a smaller comparison value for yesterday, and a sparkline showing the last 24 hours.
B.A pie chart of free beds by ward, a stacked bar of occupied versus free beds, and a scatter plot of admissions against discharges.
C.A detailed table listing every bed, its ward, and its last-cleaned timestamp, sorted by ward name.
D.A geographic map of the hospital campus with color-coded buildings and a legend explaining each ward's capacity.
AnswerA

A KPI card gives the instantaneous value the team acts on, the comparison value supplies the like-for-like benchmark against the same hour yesterday, and a sparkline encodes the recent direction of travel without consuming much space. Together they answer the three questions the team asked in one compact band.

Why this answer

The request is for a glanceable status band answering three linked questions: current value, comparison to a prior equivalent period, and direction. A KPI card paired with a comparison figure and a sparkline delivers exactly that combination with minimal reading effort. Composition, correlation, row-level detail, and spatial views all answer different questions and would slow the team's decision.

Exam trap

The trap here is assuming a visually rich chart such as a pie or map is automatically better for a dashboard, when the requirement is a fast, precise status read that only a KPI-plus-comparison-plus-sparkline band delivers.

304
MCQeasy

A data analyst needs to retrieve all unique job titles from the employees table. Which SQL clause should be used with the SELECT statement?

A.TOP
B.UNIQUE
C.DISTINCT
D.ORDER BY
AnswerC

DISTINCT removes duplicate rows from the result set, so each job title appears once. The stem explicitly requires all unique job titles, and DISTINCT operates on the selected column values, satisfying that uniqueness constraint directly. GROUP BY would also deduplicate but returns grouped aggregates rather than a simple unique list.

Why this answer

The DISTINCT keyword is used to return only distinct (different) values.

305
Multi-Selecthard

A data analyst is building a logistic regression model to predict whether a customer will churn (yes/no). Which TWO statements about logistic regression are correct?

Select 2 answers
A.It is used only for time series forecasting.
B.The dependent variable is continuous.
C.The output is a probability between 0 and 1.
D.It requires normally distributed errors.
E.It assumes a linear relationship between predictors and the log-odds of the outcome.
AnswersC, E

Logistic regression applies the sigmoid function to a linear combination of predictors, squashing its output into the 0 to 1 interval. That value is interpreted as the probability of the positive class, directly supporting the churn yes/no prediction the analyst requires.

Why this answer

Option C is correct because logistic regression applies the sigmoid (logistic) function to a linear combination of predictors, producing an output that is a probability bounded between 0 and 1, which is exactly what is needed to model the churn probability (yes/no). Option E is correct because logistic regression assumes linearity on the logit scale: the log-odds of the outcome, ln(p/(1-p)), is modeled as a linear function of the predictor variables. Option A is incorrect because logistic regression is a classification method for binary outcomes, not a time series forecasting technique.

Option B is incorrect because the dependent variable is binary/categorical (churn yes/no), not continuous. Option D is incorrect because logistic regression does not require normally distributed errors; that assumption belongs to linear regression, whereas logistic regression uses maximum likelihood estimation with a binomial error distribution.

Exam trap

DA0-002 often tests the confusion between linear and logistic regression assumptions — candidates incorrectly apply OLS assumptions (normal errors, continuous outcome) to logistic regression or miss that the linearity assumption applies to log-odds, not raw probabilities.

306
MCQhard

A data analyst is evaluating a classification model that predicts customer churn. The model has an accuracy of 95%, but the analyst suspects it is not performing well on the minority class (churners). The dataset is highly imbalanced, with only 5% churners. Which metric should the analyst prioritize to assess the model's ability to correctly identify churners?

A.Recall
B.Specificity
C.Accuracy
D.Precision
AnswerA

Recall measures the proportion of actual churners that were correctly identified. In an imbalanced dataset where churners are rare, high accuracy can be misleading because the model may simply predict the majority class. Recall directly assesses the model's sensitivity to the minority class, which is critical when missing a churner is costly.

Why this answer

Recall is the appropriate metric because it quantifies how many actual churners the model successfully identifies. In imbalanced settings, accuracy and specificity can be high even when the model fails to detect the minority class. Precision is also important but secondary when the goal is to capture as many churners as possible.

Exam trap

The trap here is assuming high accuracy indicates good performance, ignoring the class imbalance and the need to evaluate minority class detection.

307
MCQmedium

In Tableau, an analyst wants to create a calculated field that returns the average sales per customer only for customers who have made more than five purchases. Which Tableau function or approach would be most efficient?

A.Use a table calculation for running sum
B.Use a context filter on the number of records
C.Use a Level of Detail expression to count purchases per customer, then filter
D.Create a parameter to filter customers
AnswerC

A Level of Detail expression computes the purchase count at customer granularity, independent of the view's dimensions, so filtering on that fixed aggregate correctly restricts results to customers exceeding five purchases. This satisfies the per-customer threshold constraint without altering the underlying data source.

Why this answer

A Level of Detail (LOD) expression such as {FIXED [Customer] : COUNT([Order ID])} computes the purchase count per customer independently of the view's granularity, which can then be used in a filter or conditional calculation. This is the canonical Tableau approach for row-level filtering based on aggregated per-entity metrics. It is efficient because it pushes the aggregation to the data source level.

Exam trap

The trap is reaching for table calculations or context filters when the requirement is a per-entity aggregate filter, which only LOD expressions handle correctly in Tableau.

How to eliminate wrong answers

Option A is wrong because a running sum table calculation operates on the visible view and cannot filter customers by their total purchase count before aggregation. Option B is wrong because a context filter on number of records filters individual rows, not aggregated counts per customer, so it would not isolate customers with more than five purchases. Option D is wrong because parameters are user-driven inputs, not mechanisms for computing per-customer aggregates to drive filtering.

308
Multi-Selecthard

Which TWO of the following are valid techniques for validating the performance of a predictive model?

Select 2 answers
A.Bootstrapping
B.Feature scaling
C.Train-test split
D.K-fold cross-validation
E.Increasing training data
AnswersC, D

Splitting data into training and testing sets is a basic validation approach.

Why this answer

The train-test split (Option C) is a fundamental technique for validating predictive model performance by partitioning the dataset into separate training and testing subsets, ensuring the model is evaluated on unseen data to gauge generalization. This method directly addresses overfitting and provides an unbiased estimate of model accuracy, making it a standard practice in supervised learning workflows.

Exam trap

CompTIA often tests the distinction between data preprocessing techniques (like feature scaling) and actual model validation methods, leading candidates to mistakenly select feature scaling as a validation technique because it is a common step in the modeling pipeline.

309
MCQmedium

An analyst runs a simple linear regression with an R² value of 0.85. Which interpretation is correct?

A.85% of the variance in the dependent variable is explained by the independent variable.
B.The slope of the regression line is 0.85.
C.The independent variable is 85% correlated with the dependent variable.
D.85% of the data points lie on the regression line.
AnswerA

R² measures the proportion of variance in the dependent variable accounted for by the independent variable in the fitted model. A value of 0.85 therefore means 85% of that variance is explained, with the remaining 15% attributable to other factors.

Why this answer

R², the coefficient of determination, measures the proportion of variance in the dependent variable that is predictable from the independent variable(s). An R² of 0.85 means that 85% of the variance in the dependent variable is explained by the model, leaving 15% unexplained. This is the standard interpretation in regression analysis.

Exam trap

DA0-002 often tests the confusion between R² and the correlation coefficient r, or misinterpreting R² as the slope or the proportion of points on the line.

How to eliminate wrong answers

Option B is wrong because the slope of the regression line is given by the coefficient estimate (β₁), not R²; R² is a unitless measure of fit. Option C is wrong because the correlation coefficient r is the square root of R² (for simple linear regression), so r = √0.85 ≈ 0.92, not 0.85. Option D is wrong because R² does not indicate the proportion of data points on the line; it measures variance explained, and even a perfect fit (R²=1) does not require all points to lie exactly on the line unless the data are perfectly linear.

310
Multi-Selecthard

A data analyst at a healthcare insurer is building a new analytics dataset that combines claims data with member demographic data. The compliance team requires the analyst to apply data governance controls that support accountability and traceability. Which TWO actions best satisfy this requirement? (Choose two.)

Select 2 answers
A.Document the source systems, transformation logic, and refresh schedule for each field in the dataset.
B.Convert all numeric claim amounts to a different currency to standardize reporting across regions.
C.Increase the storage capacity of the data warehouse to accommodate future growth of the combined dataset.
D.Schedule a weekly full backup of the analytics database to a secondary data center.
E.Assign a named data steward who is responsible for approving access requests and resolving data quality issues.
AnswersA, E

Documenting source systems, transformations, and refresh schedules creates data lineage, which is a core governance control for traceability. It allows auditors and data stewards to trace any value back to its origin and understand how it was changed, directly supporting accountability for the combined claims and demographic dataset.

Why this answer

Accountability and traceability in data governance are achieved through documented lineage and clearly assigned ownership. Recording source systems, transformation logic, and refresh schedules provides traceability, while naming a data steward who approves access and resolves quality issues provides accountability. Together these two actions directly satisfy the compliance team's requirement.

Exam trap

The trap here is selecting operational or infrastructure tasks such as backups or storage expansion, which feel responsible but do not create governance accountability or traceability.

311
MCQhard

A data governance team is drafting a policy for handling personally identifiable information (PII). According to data governance best practices, which document should define the classification levels and handling procedures?

A.Data dictionary
B.Data classification policy
C.Data quality report
D.Data flow diagram
AnswerB

A data classification policy defines the tiers of sensitivity and the handling rules for each, directly satisfying the stem's requirement for a document specifying PII classification levels and procedures. It governs how data is labelled and protected, unlike retention schedules or access-control standards, which address different governance concerns.

Why this answer

The data classification policy is the authoritative document that defines classification levels (e.g., public, internal, confidential, restricted) and specifies handling procedures for each category, including PII. This aligns with data governance best practices, as it establishes the rules for labeling, storing, transmitting, and disposing of sensitive data. A data dictionary describes metadata and schema, not classification rules.

Exam trap

The trap here is that candidates confuse the data dictionary (which describes data structure) with the data classification policy (which governs data sensitivity and handling), leading them to select the dictionary as the document that defines classification levels.

How to eliminate wrong answers

Option A is wrong because a data dictionary documents metadata such as field names, data types, and definitions, but it does not define classification levels or handling procedures for PII. Option C is wrong because a data quality report measures data accuracy, completeness, and consistency, not security or classification policies. Option D is wrong because a data flow diagram visually maps how data moves between systems, but it does not prescribe classification levels or handling rules.

312
MCQmedium

A data analyst creates a dashboard for operational metrics. The operations team reports that the dashboard is confusing because it shows too many metrics on one screen. Which design principle should the analyst apply?

A.Apply progressive disclosure
B.Increase white space
C.Use a single chart type
D.Add more filters
AnswerA

Progressive disclosure initially presents only high-level metrics, revealing detail on demand, which directly resolves the stem's complaint of too many metrics crowding one screen. It reduces cognitive load while preserving access to deeper data, unlike simply removing metrics.

Why this answer

Progressive disclosure is a design principle that presents only the most critical information initially, with the option to reveal additional details as needed. This directly addresses the operations team's complaint of too many metrics on one screen by reducing cognitive load and allowing users to drill down into specific metrics when required. In dashboard design, this is often implemented through expandable sections, hover-over tooltips, or click-through layers.

Exam trap

The trap here is that candidates often confuse 'reducing clutter' (white space) with 'reducing information overload' (progressive disclosure), or they mistakenly believe that adding more filters will simplify the initial view, when in fact filters only change what is shown without addressing the core issue of too many metrics displayed at once.

How to eliminate wrong answers

Option B is wrong because increasing white space improves visual clarity and reduces clutter, but it does not solve the problem of too many metrics being displayed simultaneously; it merely spaces them out. Option C is wrong because using a single chart type does not reduce the number of metrics shown; it may even force inappropriate visualization of diverse data types, leading to misinterpretation. Option D is wrong because adding more filters gives users control over what data is displayed, but it does not address the initial overload of visible metrics; filters are a complementary feature, not a primary solution for reducing on-screen complexity.

313
MCQeasy

A data analyst wants to identify customers whose last name starts with 'Mc' from the 'customers' table. Which WHERE clause condition should be used?

A.last_name LIKE 'Mc_'
B.last_name LIKE 'Mc%'
C.last_name IN ('Mc%')
D.last_name = 'Mc%'
AnswerB

The LIKE operator with the wildcard pattern 'Mc%' matches any last name beginning with the literal characters 'Mc', satisfying the prefix constraint. The percent sign represents any sequence of zero or more characters, so surnames such as 'McDonald' or 'McIntyre' are returned while others are excluded.

Why this answer

The LIKE operator with the '%' wildcard matches any sequence of zero or more characters, so 'Mc%' correctly finds all last names beginning with 'Mc' (e.g., 'McDonald', 'McIntyre', 'Mc'). The underscore '_' matches exactly one character, and the equals operator requires an exact literal match, so neither can express a prefix pattern.

Exam trap

The trap here is confusing the two LIKE wildcards — candidates often pick '_' thinking it means 'any characters' when it actually matches exactly one character, while '%' is the multi-character wildcard.

How to eliminate wrong answers

Option A is wrong because the underscore wildcard '_' matches exactly one character, so 'Mc_' would only match three-character names like 'McX' and miss 'McDonald'. Option C is wrong because IN is a set-membership operator that compares against literal values, not patterns — it would look for the literal string 'Mc%' and never perform wildcard matching. Option D is wrong because '=' performs an exact string comparison, so it would only match a last name literally equal to 'Mc%' and not any name starting with 'Mc'.

314
MCQeasy

A data analyst needs to communicate the findings of a marketing campaign analysis to the Vice President of Marketing. The VP typically only reads the first paragraph and wants the key takeaway immediately. Which format should the analyst use?

A.A raw data extract
B.A detailed technical report
C.An executive summary
D.A data dictionary
AnswerC

An executive summary condenses the entire analysis into a brief opening that states the key takeaway and recommendations first. This satisfies the VP's constraint of reading only the first paragraph while still capturing the campaign's essential findings, unlike a detailed report or dashboard requiring further exploration.

Why this answer

An executive summary is a concise, standalone document that presents the most important findings and recommendations at the very beginning, tailored for senior leaders who need the key takeaway immediately. It allows the VP to grasp the campaign's outcome and recommended actions without reading the full report. This format directly matches the VP's stated preference for reading only the first paragraph.

Exam trap

DA0-002 often tests the distinction between communication formats for different audiences, and candidates mistakenly choose a detailed technical report or data dictionary thinking more detail is always better, ignoring the executive's need for brevity.

How to eliminate wrong answers

Option A is wrong because a raw data extract contains unprocessed data with no narrative or interpretation, forcing the VP to analyze it personally, which is the opposite of a quick takeaway. Option B is wrong because a detailed technical report is lengthy and method-heavy, designed for analysts or engineers, not for an executive who reads only the first paragraph. Option D is wrong because a data dictionary documents metadata (field names, types, definitions), not analytical findings or recommendations, so it does not communicate campaign results.

315
Multi-Selecthard

An analyst is performing K-means clustering on customer data. The elbow method shows a clear bend at k=4. Which THREE of the following are true about K-means clustering with k=4?

Select 3 answers
A.The number of clusters is determined to be 4.
B.The algorithm will always produce the same clusters regardless of initial centroids.
C.The centroids are recomputed iteratively until convergence.
D.Categorical variables should be standardised before clustering.
E.The algorithm minimises the sum of squared distances between points and their assigned centroid.
AnswersA, C, E

The elbow method plots within-cluster sum of squares against k; the clear bend at k=4 indicates diminishing returns beyond four clusters, so the analyst selects four as the cluster count. This directly sets the k parameter for the algorithm.

Why this answer

Option A is correct because the elbow method identifies the optimal number of clusters at the point where the within-cluster sum of squares (inertia) begins to level off, and a clear bend at k=4 indicates that 4 is the chosen number of clusters. Option C is correct because K-means is an iterative algorithm: after assigning points to the nearest centroid, it recomputes each centroid as the mean of the points in its cluster, repeating until assignments and centroids stabilise (convergence). Option E is correct because the objective function K-means minimises is the sum of squared Euclidean distances between each point and its assigned centroid (inertia).

Option B is not correct because K-means results depend on initial centroid placement and can converge to different local optima, which is why k-means++ or multiple restarts are used. Option D is not correct because categorical variables are not suitable for K-means (which relies on Euclidean distance and means); they require encoding or a different algorithm such as K-modes, and standardisation applies to numerical features, not categorical ones.

Exam trap

The trap is thinking K-means always gives the same result (it doesn't due to random initialization) and that categorical variables can be standardised (they need encoding instead).

316
MCQmedium

A data analyst needs to present findings to a non-technical executive audience. Which visualization type is most appropriate to communicate a clear comparison of sales performance across multiple regions for the current quarter?

A.Scatter plot
B.Line chart
C.Bar chart
D.Heatmap
AnswerC

A bar chart encodes each region's sales as a separate bar on a common axis, so executives can compare magnitudes at a glance without interpreting scales or trends. Position and length are the most accurately judged visual encodings, making categorical comparison across regions immediate and unambiguous.

Why this answer

A bar chart is the most appropriate choice because it excels at comparing discrete categories (regions) using a common baseline, making it easy for a non-technical audience to quickly see which regions performed best or worst in the current quarter. The vertical or horizontal bars provide a clear, direct visual comparison of sales performance without requiring interpretation of trends or correlations.

Exam trap

The trap here is that candidates often choose a line chart (Option B) because they associate sales data with time series, but the question specifies a single quarter comparison across regions, not a trend over time.

How to eliminate wrong answers

Option A is wrong because a scatter plot is designed to show the relationship or correlation between two continuous variables, not to compare discrete categories like regions; it would confuse a non-technical audience with unnecessary data point dispersion. Option B is wrong because a line chart is best for showing trends over time, but the question asks for a comparison across regions for a single time period (current quarter), making the line chart misleading as it implies a temporal sequence. Option D is wrong because a heatmap uses color intensity to represent values in a matrix, which is effective for spotting patterns in large datasets but is less intuitive for direct, side-by-side comparisons of a single metric across a small number of categories.

317
MCQmedium

A marketing analyst wants to predict whether a customer will churn (yes/no) based on account age and monthly charges. Which regression technique is most appropriate?

A.Logistic regression
B.Simple linear regression
C.Multiple linear regression
D.K-means clustering
AnswerA

Churn is a binary yes/no outcome, and logistic regression models the probability of a categorical dependent variable using a sigmoid function. Linear regression would predict continuous values outside 0–1, making it unsuitable for this classification scenario.

Why this answer

Logistic regression is specifically designed for binary classification problems, such as predicting churn (yes/no). It models the probability of the outcome using a logistic function, making it appropriate for this scenario. Simple and multiple linear regression are for continuous outcomes, and K-means is for clustering, not prediction.

Exam trap

The trap is confusing regression techniques: candidates might think any regression can predict binary outcomes, but only logistic regression is suited for classification. Linear regression outputs continuous values, which are not probabilities.

How to eliminate wrong answers

Option B is wrong because simple linear regression predicts a continuous dependent variable, not a binary outcome. Option C is wrong because multiple linear regression also predicts a continuous outcome and cannot directly model a binary response. Option D is wrong because K-means clustering is an unsupervised learning algorithm used for grouping data, not for predicting a binary target.

318
Multi-Selecthard

A data analyst is investigating a correlation between two continuous variables. Which THREE of the following are appropriate steps in this exploratory data analysis? (Select THREE.)

Select 3 answers
A.Calculate the Pearson correlation coefficient
B.Create a scatter plot
C.Perform a t-test
D.Check for outliers using box plots
E.Create a contingency table
AnswersA, B, D

Calculating the Pearson correlation coefficient quantifies the strength and direction of a linear relationship between two continuous variables, directly satisfying the stem's correlation investigation. It assumes interval or ratio data, linearity, and approximately normal distributions, making it the standard parametric measure for this exploratory step.

Why this answer

Option A is correct because the Pearson correlation coefficient (r) is the standard statistic for quantifying the strength and direction of a linear relationship between two continuous variables, which is exactly the analyst's goal. Option B is correct because a scatter plot visually reveals the form, direction, and strength of the relationship between the two continuous variables and can expose non-linearity that a single correlation value would hide. Option D is correct because box plots (or their underlying IQR-based rules) identify outliers that can disproportionately distort the Pearson correlation coefficient, so checking for them is a necessary data-quality step before trusting r.

Option C is not appropriate here because a t-test compares means between groups (or against a hypothesized mean), not the association between two continuous variables. Option E is not appropriate because a contingency table summarizes counts of categorical variables, whereas both variables in this scenario are continuous.

Exam trap

DA0-002 often tests the distinction between correlation analysis (continuous variables, Pearson/scatter/outliers) and group comparison or categorical analysis (t-test, contingency table) — candidates who pick the t-test confuse 'comparing' with 'correlating'.

319
MCQhard

An organization wants to ensure that all reports use the same definitions for metrics like 'Active Customer' to avoid confusion. Which data governance element should be implemented?

A.Row-level security
B.Data dictionary
C.Single version of truth
D.Data lineage
AnswerB

A data dictionary documents standard metric definitions such as 'Active Customer', giving every report author one agreed meaning. This directly satisfies the requirement for consistent definitions, unlike access controls or retention policies, which govern permissions and lifecycle rather than semantics.

Why this answer

A data dictionary is a centralized repository that documents the names, definitions, data types, and business rules for data elements such as 'Active Customer.' By defining the metric once and publishing it, all reports and analysts reference the same meaning, eliminating ambiguity. This directly addresses the requirement for consistent metric definitions across the organization.

Exam trap

DA0-002 often tests the confusion between data governance artifacts, and candidates pick 'single version of truth' because it sounds like the goal, missing that the question asks for the specific element to implement—the data dictionary.

How to eliminate wrong answers

Option A is wrong because row-level security controls which users can access which rows based on roles or attributes; it does not standardize metric definitions. Option C is wrong because 'single version of truth' is a desired outcome or principle, not an implementable governance element; a data dictionary is the concrete artifact that helps achieve it. Option D is wrong because data lineage tracks the origin and transformation of data through systems, which supports auditability but does not define business metrics.

320
MCQhard

A data analyst is merging two datasets: one containing employee details (employee_id, name, department) and another containing salary information (employee_id, salary). The employee_id in the first dataset is stored as an integer, while in the second dataset it is stored as a string with leading zeros (e.g., '00123'). The analyst attempts to join the tables on employee_id but gets no matches. What is the most likely cause of the join failure?

A.The employee_id data types are different, causing implicit conversion issues.
B.The employee_id column has NULL values in both tables.
C.The join condition is missing a necessary filter on department.
D.The employee_id column contains duplicate values in one of the tables.
AnswerA

When joining on columns with different data types, the database may perform implicit conversion, but it can lead to unexpected results or errors. Here, integer vs. string with leading zeros means '123' does not equal '00123'. This mismatch causes the join to fail. Explicitly converting both to the same type and format is necessary for a successful join.

Why this answer

The join fails because the employee_id values are stored differently: one as integer, the other as string with leading zeros. Even if the numeric values are the same, the string representation differs, so equality comparison fails. Converting both to a consistent type and format resolves the issue.

Exam trap

The trap here is overlooking data type and format differences, assuming that '123' and '00123' are equivalent when they are not in a join condition.

321
MCQhard

A data scientist is analyzing a dataset with multiple features and wants to apply k-means clustering to segment customers. She chooses k = 4 based on the elbow method. During the iteration process, which of the following correctly describes a step in the k-means algorithm?

A.Compute the covariance matrix and use principal components to initialize centroids.
B.Use hierarchical clustering to determine initial centroids.
C.Randomly assign centroids and then compute distances to the cluster medians.
D.Assign each point to the nearest centroid based on Euclidean distance, then update centroids as the mean of points in each cluster.
AnswerD

K-means alternates two steps: each point is assigned to the cluster whose centroid is nearest by Euclidean distance, then every centroid is recomputed as the mean of its assigned points. Iteration repeats until assignments stabilise, satisfying the k=4 segmentation.

Why this answer

In the k-means algorithm, after initializing centroids, each data point is assigned to the nearest centroid based on Euclidean distance, and then centroids are recomputed as the mean of all points in the cluster. This iterative process continues until convergence. Option D accurately describes this step.

Exam trap

The trap is mixing up k-means with other algorithms (e.g., PCA for initialization) or using median instead of mean. Candidates might also confuse the update step with using medians (as in k-medians).

How to eliminate wrong answers

Option A is wrong because computing the covariance matrix and using principal components is part of PCA, not k-means initialization. Option B is wrong because hierarchical clustering is a different clustering method and is not a standard step in k-means. Option C is wrong because k-means updates centroids as the mean, not the median, and it does not use cluster medians.

322
MCQmedium

Refer to the exhibit. Which type of ensemble method is being used?

A.Boosting
B.Stacking
C.Voting
D.Bagging
AnswerD

Bagging trains multiple instances of the same algorithm on bootstrap samples drawn with replacement, then aggregates their predictions by voting or averaging. This parallel, variance-reducing structure distinguishes it from boosting, which trains sequentially on reweighted data.

Why this answer

The exhibit shows multiple base models (Model 1, Model 2, Model 3) trained in parallel on bootstrap samples of the data, and their predictions are combined via averaging (regression) or majority voting (classification). This parallel training with resampled data and equal-weight aggregation is the defining characteristic of bagging (Bootstrap Aggregating).

Exam trap

CompTIA often tests the distinction between bagging and boosting by showing parallel vs. sequential training diagrams, and the trap here is confusing the parallel bootstrap resampling with the sequential error-correction approach of boosting.

How to eliminate wrong answers

Option A is wrong because boosting trains models sequentially, where each subsequent model focuses on correcting the errors of the previous one, not in parallel on bootstrap samples. Option B is wrong because stacking uses a meta-learner to combine predictions from diverse base models, not simple averaging or majority voting. Option C is wrong because voting typically combines predictions from different model types (e.g., logistic regression, SVM) trained on the same dataset, not from the same model type trained on bootstrap samples.

323
MCQmedium

A healthcare analyst is creating a report on patient outcomes. To comply with privacy regulations, which action should be taken before publishing?

A.Anonymize personally identifiable information
B.Include patient names for context
C.Provide raw data to all stakeholders
D.Aggregate data without anonymization
AnswerA

Anonymising removes identifiers such as names, addresses and dates, so published patient outcome data cannot be traced back to individuals. This directly satisfies the privacy regulation constraint in the stem, which requires de-identification before release. Unlike pseudonymisation, anonymisation is irreversible, ensuring the report itself carries no personally identifiable information.

Why this answer

Anonymizing PII protects patient identities and ensures compliance with privacy laws like HIPAA.

324
MCQmedium

A data analyst is reviewing a SQL query that joins three large tables. The query takes over an hour to run. The analyst notices that the WHERE clause filters on indexed columns in only two tables. Which of the following should the analyst do first to improve performance?

A.Use subqueries instead of joins
B.Check the query execution plan and optimize join order
C.Add indexes to all columns used in joins
D.Increase server memory
AnswerB

The execution plan reveals how the optimiser orders and joins the three tables, including scan and join methods. Since only two tables have indexed filter columns, examining the plan identifies whether join order or a missing index causes the hour-long runtime.

Why this answer

The query execution plan reveals how the database engine processes joins and filters. By checking the plan, the analyst can identify the most selective filter and rearrange the join order to reduce the number of rows processed early, which is the most impactful first step. Optimizing join order leverages existing indexes without requiring schema changes or hardware upgrades.

Exam trap

CompTIA often tests the misconception that adding indexes or hardware is the immediate fix, when in fact analyzing the execution plan and adjusting join order is the cheapest and most effective first step.

How to eliminate wrong answers

Option A is wrong because subqueries often perform worse than joins in large-table scenarios, as they can lead to correlated subquery execution and repeated scans. Option C is wrong because adding indexes to all join columns is unnecessary and may degrade write performance; the analyst should first verify if existing indexes are being used efficiently via the execution plan. Option D is wrong because increasing server memory is a reactive, costly measure that does not address the root cause of inefficient query processing, such as poor join order or missing index usage.

325
MCQmedium

A data team is creating a dashboard to monitor real-time sales. What design principle is critical?

A.Provide downloadable raw data
B.Use auto-refresh and clear alert thresholds
C.Include all historical data
D.Minimize use of color
AnswerB

Auto-refresh keeps the dashboard current without manual intervention, satisfying the real-time monitoring requirement, while clear alert thresholds convert raw sales figures into actionable signals. Together they ensure the team detects anomalies promptly rather than reviewing stale data, which is the critical design principle for a real-time sales dashboard.

Why this answer

For a real-time sales dashboard, the critical design principle is to ensure data freshness and immediate actionability. Option B is correct because auto-refresh keeps the dashboard current without manual intervention, and clear alert thresholds enable the team to instantly identify when sales metrics deviate from expected ranges, which is essential for real-time monitoring.

Exam trap

The trap here is that candidates often confuse general dashboard design principles (like minimizing color or providing raw data) with the specific, non-negotiable requirements of a real-time monitoring system, where data freshness and alerting are paramount.

How to eliminate wrong answers

Option A is wrong because providing downloadable raw data is a feature for offline analysis or auditing, not a critical principle for real-time monitoring; it can even introduce latency and security risks. Option C is wrong because including all historical data would overwhelm the dashboard's performance and cognitive load, contradicting the need for real-time, focused insights. Option D is wrong because minimizing color use is a general design best practice for accessibility, but it is not the critical principle for a real-time dashboard; color can be effectively used to highlight alerts and thresholds.

326
MCQmedium

An analyst compares average sales across three different store locations using a statistical test. Which test is most appropriate?

A.ANOVA
B.t-test
C.Correlation analysis
D.Chi-square test
AnswerA

ANOVA compares means across three or more independent groups on one continuous dependent variable, here average sales per store location. It tests whether at least one group mean differs, controlling family-wise error. A t-test suits only two groups, so ANOVA is the appropriate choice for three locations.

Why this answer

ANOVA compares means across three or more groups.

327
MCQeasy

A data analyst wants to compare the total sales of four different product categories for a single year. Which chart type is most appropriate for this comparison?

A.Pie chart
B.Scatter plot
C.Line chart
D.Bar chart
AnswerD

A bar chart encodes each category's total as a separate bar on a common axis, so four product categories can be compared side by side for one year. Length differences are read accurately, unlike pie charts, which obscure close values.

Why this answer

A bar chart is the most appropriate choice for comparing discrete categorical totals such as total sales across four product categories. Bar length directly encodes the magnitude, making comparisons easy and accurate. It handles a small number of categories cleanly without the perceptual issues of pie charts.

Exam trap

The trap is choosing a pie chart because the question mentions 'categories,' but pie charts are for parts-of-whole, not direct magnitude comparison across categories.

How to eliminate wrong answers

Option A is wrong because pie charts are poor for precise comparison and become hard to read beyond a few slices; they emphasize parts-of-whole rather than direct magnitude comparison. Option B is wrong because scatter plots show relationships between two continuous variables, not totals across categories. Option C is wrong because line charts are designed for trends over time, not for comparing static category totals.

328
MCQeasy

A retail company stores customer transaction data in a relational database. They want to analyze purchasing patterns over time. Which type of data structure best supports this analysis?

A.Relational table
B.Graph database
C.Document store
D.Key-value store
AnswerA

A relational table stores transactions as rows with typed columns, so SQL aggregation and time-series grouping can reveal purchasing trends. Its fixed schema and join support satisfy the requirement to analyse historical transaction records over time, which non-relational or unstructured stores handle less directly.

Why this answer

A relational table is the correct choice because it organizes transaction data into structured rows and columns with defined schemas, enabling efficient SQL-based queries for time-series analysis (e.g., aggregating purchases by date, customer, or product). The relational model supports ACID transactions and joins across related tables (e.g., customers, products, transactions), which is essential for analyzing purchasing patterns over time while maintaining data integrity.

Exam trap

The trap here is that candidates may confuse 'analyzing purchasing patterns over time' with needing a graph database for relationships, but the key requirement is structured time-series aggregation, which is a core strength of relational tables, not graph or NoSQL stores.

How to eliminate wrong answers

Option B (Graph database) is wrong because graph databases excel at modeling relationships between entities (e.g., social networks or recommendation engines) but are not optimized for time-series aggregation or range queries on structured transaction data; they lack native support for SQL-style GROUP BY and window functions. Option C (Document store) is wrong because document stores (e.g., MongoDB) store semi-structured JSON-like documents, which can lead to data duplication and complex aggregation pipelines for time-based analysis, and they typically do not enforce strict schemas or support ACID transactions across multiple collections. Option D (Key-value store) is wrong because key-value stores (e.g., Redis) provide fast lookups by a single key but cannot efficiently query on multiple attributes (e.g., date range, product category) or perform relational joins, making them unsuitable for analytical queries on purchasing patterns.

329
Multi-Selecthard

A data analyst is evaluating the quality of a customer database. Which THREE of the following are dimensions of data quality?

Select 3 answers
A.Completeness
B.Correlation
C.Timeliness
D.Accuracy
E.Variance
AnswersA, C, D

Completeness measures whether required customer attributes are populated, directly assessing missing values in the database. It is a recognised data quality dimension because gaps in mandatory fields undermine segmentation, contactability and downstream analytics, satisfying the evaluation's need to quantify absent data.

Why this answer

Accuracy, completeness, and timeliness are standard data quality dimensions.

330
MCQeasy

A data analyst at a marketing agency is working with a dataset containing customer demographics, purchase history, and social media engagement metrics. The agency wants to perform sentiment analysis on unstructured social media comments to identify brand perception. The dataset also includes structured fields like age, income, and purchase amounts. The analyst needs to choose a storage and processing platform that can handle both structured and unstructured data efficiently without requiring extensive schema definition upfront. Which platform should the analyst recommend?

A.Relational database (RDBMS)
B.Data lake
C.Data warehouse
D.NoSQL document database
AnswerB

A data lake stores raw structured and unstructured data without upfront schema definition, satisfying the stem's schema-on-read constraint. Unlike a data warehouse, which demands schema-on-write modelling, it ingests social media comments and demographic fields together, letting the analyst apply sentiment analysis later.

Why this answer

A data lake is the correct choice because it can store both structured data (e.g., age, income, purchase amounts) and unstructured data (e.g., social media comments) in its native format without requiring a predefined schema. This flexibility allows the analyst to ingest raw social media text for sentiment analysis and later apply schema-on-read for structured queries, avoiding the upfront schema definition needed by other platforms.

Exam trap

The trap here is that candidates often confuse a data warehouse with a data lake, assuming both can handle unstructured data, but a data warehouse requires structured, transformed data and cannot natively store raw social media comments without prior schema definition.

How to eliminate wrong answers

Option A is wrong because a relational database (RDBMS) requires a rigid, predefined schema and is optimized for structured data, making it inefficient for storing and processing unstructured social media comments without extensive ETL. Option C is wrong because a data warehouse is designed for structured, processed data and typically uses a schema-on-write approach, which cannot natively handle unstructured text like social media comments without significant transformation. Option D is wrong because a NoSQL document database can store semi-structured data (e.g., JSON) but is not optimized for large-scale, raw unstructured text and lacks the integrated processing capabilities (e.g., Apache Spark or Hadoop) that a data lake provides for sentiment analysis.

331
MCQeasy

Which data quality dimension ensures that data represents the real-world object or event correctly?

A.Accuracy
B.Completeness
C.Consistency
D.Timeliness
AnswerA

Accuracy verifies that values faithfully reflect the real-world object or event they describe, directly satisfying the stem's requirement for correct representation. Unlike validity, which only confirms conformance to defined formats or rules, accuracy measures correspondence with reality itself, making it the dimension that ensures data genuinely depicts what it claims to represent.

Why this answer

Accuracy is the data quality dimension that measures how closely data values reflect the real-world object, event, or condition they are intended to describe. If a customer's address, a transaction amount, or a sensor reading is recorded incorrectly, the data is inaccurate even if it is complete, consistent, and timely. Accuracy is therefore the dimension specifically concerned with correctness of representation.

Exam trap

The trap here is confusing accuracy with completeness or consistency — candidates often pick completeness because the data 'looks full,' but the question asks specifically about correctly representing the real-world object or event.

How to eliminate wrong answers

Option B is wrong because completeness only measures whether all required data values are present, not whether those values are correct — a record can be 100% populated yet contain wrong values. Option C is wrong because consistency concerns whether the same data is represented uniformly across systems, formats, or time periods, not whether it matches reality. Option D is wrong because timeliness addresses whether data is available and up to date when needed, which says nothing about whether the values themselves are correct.

332
MCQeasy

Which of the following is a characteristic of structured data?

A.It conforms to a fixed schema with rows and columns.
B.It has a flexible schema that can vary per record.
C.It cannot be analyzed using SQL.
D.It is stored as blobs in a data lake.
AnswerA

Structured data is organised into a predefined, fixed schema of rows and columns, typically stored in relational databases and queried with SQL. This tabular rigidity satisfies the stem's characteristic, distinguishing it from semi-structured formats like JSON and unstructured content such as text.

Why this answer

Structured data is defined by its adherence to a fixed schema, typically organized into rows and columns within relational databases. This rigid structure enables efficient querying and manipulation using SQL, as each field has a predefined data type and constraints. The correct answer highlights this fundamental characteristic, which distinguishes structured data from semi-structured or unstructured formats.

Exam trap

The trap here is that candidates often confuse semi-structured data (which has some organizational tags but no fixed schema) with structured data, leading them to select Option B, or they mistakenly think SQL cannot analyze structured data, falling for Option C.

How to eliminate wrong answers

Option B is wrong because a flexible schema that can vary per record describes semi-structured data (e.g., JSON, XML), not structured data. Option C is wrong because structured data is specifically designed to be analyzed using SQL, which is the primary query language for relational databases. Option D is wrong because storing data as blobs in a data lake is characteristic of unstructured data (e.g., images, videos), not structured data, which is stored in tables with defined schemas.

333
MCQmedium

A data analyst is performing time series analysis on monthly sales data and notices a consistent pattern of higher sales every December. Which component of time series does this represent?

A.Trend
B.Irregular component
C.Seasonality
D.Cyclical
AnswerC

Seasonality describes a repeating pattern that recurs at fixed intervals shorter than a year, such as December sales peaks. The consistent annual December uplift satisfies this definition, distinguishing it from trend, which is a longer-term directional movement, and from irregular noise.

Why this answer

Seasonality refers to a repeating pattern that occurs at fixed, predictable intervals within a year, such as higher sales every December or increased traffic every Monday. Because the December spike recurs annually at the same calendar point, it is a seasonal component. Seasonality is distinguished from trend (long-term direction) and cyclical (irregular multi-year fluctuations) by its fixed, calendar-based period.

Exam trap

DA0-002 often tests the seasonality vs. cyclical distinction — candidates pick cyclical because both repeat, but only seasonality has a fixed, calendar-based period like December.

How to eliminate wrong answers

Option A is wrong because trend describes the long-term upward or downward movement of the series over many periods, not a recurring annual spike. Option B is wrong because the irregular component is the random, unpredictable residual left after trend, seasonality, and cyclical effects are removed — December sales are predictable, not random. Option D is wrong because cyclical components fluctuate over longer, non-fixed periods (often tied to economic cycles) and are not tied to a specific calendar month like December.

334
MCQhard

In a table 'sales_team' with columns 'salesperson', 'quarter', and 'revenue', an analyst wants to assign a rank to each salesperson within their quarter based on revenue, with the highest revenue getting rank 1. However, if two salespeople have the same revenue, they should receive the same rank, and the next rank should be the next consecutive integer (no gaps). Which window function should be used?

A.RANK()
B.NTILE(4)
C.DENSE_RANK()
D.ROW_NUMBER()
AnswerC

DENSE_RANK() assigns identical ranks to tied revenues and continues with the next consecutive integer, producing no gaps. RANK() would skip numbers after a tie, and ROW_NUMBER() would break ties arbitrarily, so neither meets the stated requirement.

Why this answer

DENSE_RANK() assigns ranks without gaps, so if two salespeople tie for rank 1, the next rank is 2. This matches the requirement that ties receive the same rank and the next rank is consecutive. It is the correct window function for this scenario.

Exam trap

DA0-002 often tests the distinction between RANK, DENSE_RANK, and ROW_NUMBER; candidates may choose RANK() when no gaps are required, forgetting that RANK() leaves gaps.

How to eliminate wrong answers

Option A is wrong because RANK() leaves gaps after ties (e.g., 1,1,3). Option B is wrong because NTILE(4) divides rows into four buckets, not ranking within quarters. Option D is wrong because ROW_NUMBER() assigns unique sequential numbers, ignoring ties.

335
MCQmedium

A company ingests customer clickstream data from its website. The data arrives continuously in JSON format and must be stored for real-time analytics. Which type of data source is being described?

A.Transactional database
B.Flat file
C.Data warehouse
D.Streaming data
AnswerD

Streaming data matches because the clickstream arrives continuously as discrete JSON events requiring real-time analytics, rather than being collected in scheduled batches. This satisfies the stem's ingestion constraint: an unbounded, ongoing flow of records processed as they arrive, distinguishing it from batch or static file sources.

Why this answer

The description matches a streaming data source because clickstream data arrives continuously in JSON format and must be stored for real-time analytics. Streaming data sources, such as Apache Kafka or Amazon Kinesis, ingest unbounded data in real time, enabling immediate processing and analytics without batch delays.

Exam trap

CompTIA Data+ often tests the distinction between 'streaming data' and 'data warehouse' by describing continuous ingestion, leading candidates to mistakenly choose 'data warehouse' because they associate analytics with warehousing, ignoring the real-time requirement.

How to eliminate wrong answers

Option A is wrong because a transactional database (e.g., OLTP system) is designed for ACID-compliant transaction processing, not for ingesting continuous, high-velocity streaming data. Option B is wrong because a flat file (e.g., CSV or text file) is a static, batch-oriented storage format that cannot handle real-time, continuous ingestion without manual intervention or scheduled loads. Option C is wrong because a data warehouse is optimized for structured, historical analytics and typically relies on batch ETL processes, not real-time streaming ingestion from clickstream sources.

336
Multi-Selecthard

A financial services firm is establishing a data governance program for its customer analytics platform. The chief data officer wants to ensure that data quality issues are detected and resolved systematically. Which two of the following practices are most appropriate for maintaining data quality on an ongoing basis? (Choose two.)

Select 2 answers
A.Establishing a data quality issue log with defined severity levels, owners, and service-level agreements for resolution.
B.Encrypting all customer data at rest and in transit to prevent unauthorized access.
C.Implementing automated data quality rules that validate completeness, uniqueness, and referential integrity during ETL loads.
D.Assigning a data steward to manually review a random sample of records each quarter and document findings in a spreadsheet.
E.Relying on business users to report data errors through ad hoc emails to the IT help desk.
AnswersA, C

A formal issue log with severity levels, assigned owners, and SLAs ensures that detected quality problems are tracked, prioritized, and resolved within agreed timeframes. This creates accountability and a repeatable process, which is a core component of an ongoing data governance program for customer analytics.

Why this answer

Automated data quality rules during ETL and a formal issue log with severity levels and SLAs together provide a systematic, repeatable approach to detecting and resolving data quality problems. The automated rules catch issues early, while the issue log ensures accountability and timely resolution, which are essential for an ongoing governance program in financial services.

Exam trap

The trap here is selecting manual sampling or ad hoc reporting as quality practices, when they are reactive and not systematic enough for ongoing governance.

337
Multi-Selecthard

A data analyst is cleaning a dataset and identifies several outliers. Which TWO methods are appropriate for handling outliers?

Select 2 answers
A.Capping
B.Mean imputation
C.Removal
D.Min-max normalization
E.Forward-fill
AnswersA, C

Capping replaces extreme values with a defined threshold, such as the 1st or 99th percentile, retaining the record while limiting its influence. This satisfies the scenario's need to handle outliers without discarding data, unlike deletion, which removes rows entirely and risks losing valid observations.

Why this answer

Capping (A) is correct because it winsorizes extreme values by replacing outliers with a boundary value such as the 1st/99th percentile or a value derived from the IQR (e.g., Q1 − 1.5×IQR, Q3 + 1.5×IQR), preserving the record while limiting the outlier's influence. Removal (C) is correct because dropping outlier rows (or excluding them from analysis) is a standard, defensible approach when the values are confirmed to be erroneous or when their influence must be eliminated. Mean imputation (B) is not an outlier-handling method; it replaces missing values with the column mean and would actually be distorted by the very outliers present.

Min-max normalization (D) merely rescales all values to the [0,1] range and does not reduce or eliminate outlier effects. Forward-fill (E) is a time-series missing-value technique that propagates the last valid observation and does nothing to address outliers.

Exam trap

The trap here is confusing data transformation techniques (like normalization or imputation) with outlier-specific handling methods; candidates might select mean imputation or min-max normalization because they are common preprocessing steps, but they do not directly address outliers.

338
MCQmedium

A data team created a dashboard for executives. The dashboard updates daily and includes several KPIs. Executives complain that they cannot quickly identify the most critical issues. Which design change would best address this?

A.Use a single aggregated metric to simplify.
B.Increase the refresh rate to every hour.
C.Incorporate conditional formatting with color alerts.
D.Add more detailed charts to each KPI.
AnswerC

Conditional formatting applies rule-based colour coding directly to KPI values, so executives spot threshold breaches without reading every figure. This satisfies the stem's constraint of quickly identifying critical issues, since colour draws attention pre-attentively. Static daily updates remain unchanged; only the visual encoding of exceptions is added.

Why this answer

Conditional formatting with color alerts (e.g., red for critical thresholds, yellow for warnings) directly addresses the executives' need to quickly identify critical issues at a glance. This design change leverages pre-attentive visual processing, allowing users to spot anomalies without manually scanning each KPI. It is a standard best practice in dashboard design for executive reporting, as it reduces cognitive load and speeds up decision-making.

Exam trap

The trap here is that candidates may confuse 'increasing data freshness' (Option B) with 'improving data interpretability,' when in fact the core issue is about visual salience and rapid issue detection, not data latency.

How to eliminate wrong answers

Option A is wrong because using a single aggregated metric oversimplifies the data and hides the specific KPIs that executives need to monitor, potentially masking critical issues in individual metrics. Option B is wrong because increasing the refresh rate to every hour does not help executives quickly identify critical issues; it only updates data more frequently, which could even cause confusion if alerts are not visually highlighted. Option D is wrong because adding more detailed charts to each KPI increases visual clutter and cognitive load, making it harder for executives to quickly spot the most critical issues, contrary to the goal of rapid identification.

339
MCQmedium

A data analyst is examining sales data for a retail chain and notices that the mean monthly sales is $50,000 while the median is $35,000. Which of the following best describes the distribution of the sales data?

A.The distribution is right-skewed.
B.The distribution is bimodal.
C.The distribution is left-skewed.
D.The distribution is symmetrical.
AnswerA

A mean of $50,000 exceeding the median of $35,000 indicates positive skew: extreme high-value months pull the mean upward while the median stays resistant. This satisfies the stem's constraint, so the distribution is right-skewed, with the longer tail extending toward higher sales values.

Why this answer

When the mean ($50,000) is greater than the median ($35,000), the distribution has a long tail on the right side pulling the mean upward, which is the definition of a right-skewed (positively skewed) distribution. In right-skewed data, a few high values inflate the mean above the median. Therefore the sales data is right-skewed.

Exam trap

DA0-002 often tests the mean-vs-median relationship, and candidates frequently reverse the skew direction; the rule is mean > median means right skew, mean < median means left skew.

How to eliminate wrong answers

Option B is wrong because bimodality refers to two distinct peaks in the distribution and is not indicated by a mean-median gap; the question gives no evidence of two modes. Option C is wrong because left-skewed (negatively skewed) distributions have the mean less than the median, which is the opposite of the given values. Option D is wrong because in a symmetrical distribution the mean and median are approximately equal, but here they differ by $15,000.

340
MCQhard

A financial institution needs to acquire credit transaction data from multiple sources while ensuring compliance with data privacy regulations. What is the most critical step?

A.Data replication for redundancy
B.Data enrichment with external sources
C.Data compression for storage
D.Data anonymization during extraction
AnswerD

Anonymising data during extraction prevents personal identifiers from entering downstream storage or processing, satisfying privacy regulation requirements at the earliest point. Applying it at extraction rather than later limits exposure and supports purpose limitation and data minimisation obligations.

Why this answer

Data anonymization during extraction is the most critical step because it ensures that personally identifiable information (PII) is irreversibly masked or removed before the data enters the processing pipeline, directly addressing compliance with regulations such as GDPR and PCI DSS. Without this step, even if other measures are applied later, the initial exposure of sensitive data violates privacy mandates and increases breach risk.

Exam trap

The trap here is that candidates confuse operational efficiency measures (replication, compression) or data enhancement (enrichment) with privacy compliance, overlooking that anonymization must be applied at the earliest point of data acquisition to satisfy regulatory requirements.

How to eliminate wrong answers

Option A is wrong because data replication for redundancy focuses on high availability and disaster recovery, not on privacy compliance; it does not prevent exposure of sensitive credit transaction data. Option B is wrong because data enrichment with external sources typically adds more data attributes, which can increase privacy risk and regulatory exposure rather than ensuring compliance. Option C is wrong because data compression for storage reduces storage footprint and may improve I/O performance but has no effect on data privacy or regulatory compliance.

341
MCQeasy

A company receives daily sales data in CSV format. The data includes a 'Date' column in MM/DD/YYYY format. To load this into a database that expects YYYY-MM-DD, the analyst should:

A.Manually edit the CSV files before loading
B.Change the database schema to accept MM/DD/YYYY
C.Ignore the date column and use a default date
D.Use a data transformation tool to convert the date format during ETL
AnswerD

Transforming the date format during ETL converts MM/DD/YYYY strings into the YYYY-MM-DD structure the target database requires, satisfying the format-mismatch constraint. Handling this in the transformation layer preserves source data integrity and loads correctly typed values.

Why this answer

The correct approach is to use a data transformation tool to convert the date format during ETL. This ensures the data is standardized to the database's expected YYYY-MM-DD format without manual intervention, preserving data integrity and enabling automated, repeatable loads. Transformation tools can parse MM/DD/YYYY and reformat it consistently, which is a core ETL function.

Exam trap

DA0-002 often tests the misconception that manual editing or schema changes are acceptable solutions for data format mismatches, when in fact automated transformation during ETL is the standard best practice.

How to eliminate wrong answers

Option A is wrong because manually editing CSV files is error-prone, not scalable, and defeats the purpose of automated ETL. Option B is wrong because changing the database schema to accept MM/DD/YYYY would require altering the database design and could break other applications or queries that expect the standard format. Option C is wrong because ignoring the date column and using a default date would result in loss of critical sales data and inaccurate analysis.

342
MCQmedium

An executive dashboard needs to display key performance indicators (KPIs) such as sales growth and customer satisfaction. Which design principle is most important?

A.Consistent color encoding for similar metrics
B.Use gauges for each KPI
C.Include all raw data in the dashboard
D.Animate charts to draw attention
AnswerA

Consistent colour encoding lets executives map each metric to a fixed hue, so sales growth and customer satisfaction are compared across tiles and refreshes without relearning the legend, directly supporting rapid KPI interpretation on the dashboard.

Why this answer

Consistent color encoding for similar metrics lets viewers instantly associate a color with a metric category (for example, green for growth, red for decline) across the dashboard, reducing cognitive load and misinterpretation. This is a core data visualization principle for executive dashboards where quick, accurate comprehension is critical. It supports preattentive processing, allowing users to spot trends without reading every label.

Exam trap

DA0-002 often tests whether candidates confuse flashy visualization features (gauges, animation) with fundamental design principles like consistent encoding, leading them to pick visually appealing but ineffective options.

How to eliminate wrong answers

Option B is wrong because gauges are often criticized for wasting space and being hard to compare; they are not the most important design principle and can mislead when scales differ. Option C is wrong because including all raw data overwhelms executives and defeats the purpose of a KPI dashboard, which should show summarized, actionable metrics. Option D is wrong because animation can distract and slow comprehension; it is a stylistic choice, not a fundamental design principle for KPI clarity.

343
MCQmedium

Refer to the exhibit. A data analyst notices that direct S3 access to files outside the "incoming/" prefix is blocked. Which data governance principle does this policy enforce?

A.Data colocation
B.Data retention
C.Data access control
D.Data encryption
AnswerC

Restricting direct S3 access to the incoming/ prefix enforces data access control, the governance principle governing who and what may reach specific datasets. Other principles such as retention or quality do not determine prefix-level read permissions.

Why this answer

The policy blocks direct S3 access to files outside the 'incoming/' prefix, which restricts which users or roles can read or write objects in specific S3 prefixes. This is a classic implementation of data access control, as it enforces permissions based on the resource path, ensuring only authorized operations are allowed on designated data. In AWS S3, such restrictions are typically applied via bucket policies or IAM policies that use conditions like `s3:prefix` to limit access.

Exam trap

CompTIA often tests the distinction between access control and encryption by presenting a policy that restricts access based on a path or condition, leading candidates to confuse it with data encryption, which is about scrambling data rather than authorizing access.

How to eliminate wrong answers

Option A is wrong because data colocation refers to physically or logically placing related data together for performance or compliance, not to restricting access based on a prefix. Option B is wrong because data retention governs how long data is kept (e.g., lifecycle policies or retention periods), not who can access it. Option D is wrong because data encryption protects data at rest or in transit (e.g., using SSE-S3 or TLS), but the policy described does not mention encryption keys, algorithms, or any cryptographic controls.

344
Multi-Selectmedium

A data analyst is evaluating data quality issues during acquisition. Which TWO issues are most likely to arise from merging data from different sources? (Select exactly 2)

Select 2 answers
A.User access permissions
B.Duplicate records
C.Slow network speed
D.High storage cost
E.Formatting inconsistencies
AnswersB, E

Merging sources that share overlapping entities, such as the same customer appearing in two systems, produces duplicate records unless deduplication keys are defined. This directly satisfies the stem's constraint of combining data from different sources, where repeated identifiers create redundancy.

Why this answer

Option B (Duplicate records) is correct because merging data from multiple sources commonly produces the same entity appearing in more than one source, and without deduplication via matching keys or fuzzy matching, redundant rows inflate counts and skew analysis. Option E (Formatting inconsistencies) is correct because different sources often represent the same field differently — for example dates as MM/DD/YYYY versus ISO 8601 YYYY-MM-DD, or units in metric versus imperial — requiring normalization before the data can be combined reliably. Option A (User access permissions) is an authorization/security concern rather than a data quality issue arising from the merge itself.

Option C (Slow network speed) is a performance/infrastructure factor that affects transfer time, not the quality of the merged data. Option D (High storage cost) is a cost/resource consideration, not a data quality defect introduced by combining sources.

Exam trap

DA0-002 often tests the difference between data-quality issues (duplicates, formatting) and non-quality concerns (permissions, cost, network), causing candidates to select operational or security issues as data-quality problems.

345
MCQmedium

A financial analyst is working with a dataset of stock prices. The dataset contains a column 'Price' that records the closing price of each stock in USD. The analyst wants to calculate the percentage change in price from the previous day for each stock. Which data type is most appropriate for the 'Price' column?

A.Integer
B.Float
C.String
D.Boolean
AnswerB

Float data types represent real numbers with fractional components, which is necessary for storing stock prices like $123.45. The analyst needs to compute percentage changes, which involve division and multiplication, so preserving decimal precision is critical. Float types accommodate the continuous nature of price data and support the required arithmetic operations.

Why this answer

Stock prices require decimal precision to accurately represent fractional dollar amounts and to perform arithmetic for percentage changes. Float data types store real numbers with decimals, making them the correct choice. Integer, string, and Boolean types cannot properly capture or compute with the continuous numeric values needed for this analysis.

Exam trap

The trap here is selecting integer because prices are often thought of as whole numbers, but stock prices include cents that are essential for accurate calculations.

346
MCQhard

A data audit reveals that some numbers in the "Revenue" column were manually entered from PDF invoices. This introduces potential errors. Which data concept is being addressed?

A.Data lineage
B.Data quality
C.Data security
D.Data governance
AnswerB

Manual transcription from PDF invoices introduces typographical and transcription errors, directly degrading accuracy, a core data quality dimension. Data quality addresses fitness for purpose, so the audit's concern with erroneous revenue figures maps to this concept.

Why this answer

The scenario describes a data audit that identifies potential errors from manual data entry from PDF invoices. This directly relates to data quality, which assesses aspects like accuracy, completeness, and consistency. The audit is highlighting a quality concern (potential errors), not lineage tracking.

Data lineage focuses on the origin and transformation of data, but the primary issue here is the risk of inaccuracies, making data quality the correct concept.

Exam trap

Candidates may see the mention of 'audit' and 'origin' and incorrectly choose data lineage. However, the key issue is the potential errors from manual entry, which is a data quality concern. The audit is identifying a quality problem, not tracing the data's path.

How to eliminate wrong answers

Option A is wrong because data lineage tracks the origin, movement, and transformation of data through its lifecycle, not the potential errors from manual entry. Option C is wrong because data security focuses on protecting data from unauthorized access, breaches, or corruption, not on the accuracy of manually entered values. Option D is wrong because data governance defines policies, roles, and procedures for managing data assets, but the specific issue of manual entry errors falls under data quality assessment, not governance frameworks.

347
MCQhard

A data analyst is using pandas to clean a DataFrame. They need to replace missing values in the 'age' column with the median age. Which method should they use?

A.df['age'].replace(np.nan, df['age'].mean())
B.df['age'].dropna()
C.df['age'].fillna(df['age'].median())
D.df['age'].interpolate()
AnswerC

fillna substitutes missing entries, and passing the column's median supplies a single imputation value computed across non-null ages. This directly satisfies the requirement to replace NaN values in 'age' with the median, preserving row count without dropping records.

Why this answer

fillna() with median() fills NaN values with the median of the column.

348
MCQhard

A data analyst is cleaning a dataset and finds that 5% of values in the 'income' column are missing. The analyst decides to impute missing values using the mean of the non-missing values. Which potential issue should the analyst be most concerned about?

A.The imputation may reduce the variance and distort the distribution.
B.The imputation is not valid because the missing rate is too low.
C.The imputation will increase the standard deviation of the variable.
D.The imputation will create outliers.
AnswerA

Mean imputation replaces missing values with the column average, pulling extreme values toward the centre. This artificially shrinks variance and narrows the distribution, biasing subsequent statistical analysis and understating the true spread of income data.

Why this answer

Mean imputation replaces missing values with the average of observed values, which pulls imputed values toward the center of the distribution. This artificially reduces the variance and can distort the shape of the distribution (e.g., making it more peaked), biasing downstream statistics and models. The analyst should be most concerned about this distortion rather than the missing rate itself.

Exam trap

DA0-002 often tests whether candidates recognize that mean imputation reduces variance and distorts distribution, rather than mistakenly thinking it increases variability or creates outliers.

How to eliminate wrong answers

Option B is wrong because a 5% missing rate is not too low to impute; low missingness is generally favorable, and there is no rule that invalidates imputation at low rates. Option C is wrong because mean imputation decreases, not increases, the standard deviation since imputed values cluster at the mean. Option D is wrong because mean imputation does not create outliers; it actually reduces extreme values by replacing missing entries with the central mean.

349
MCQmedium

A healthcare analytics team is building a patient readmission risk model. They have a dataset containing admission date, discharge date, primary diagnosis code, and patient age. To predict readmission within 30 days, the team needs to derive a new field that represents the number of days a patient stayed in the hospital. Which data transformation technique should they apply to create this derived field?

A.Date arithmetic
B.Data imputation
C.Data aggregation
D.Data normalization
AnswerA

Date arithmetic subtracts the admission date from the discharge date to produce the length of stay in days. This directly creates the derived field the team needs. It is the standard transformation for calculating durations between two temporal columns and is appropriate for a per-row calculation in a predictive model.

Why this answer

The team needs a derived field representing the number of days between admission and discharge. Date arithmetic is the transformation that subtracts one date from another to yield a duration. Aggregation summarizes groups of rows, normalization rescales values, and imputation fills missing data, none of which produce a per-patient length-of-stay field.

Exam trap

The trap here is confusing data transformation techniques and assuming that any manipulation of date fields qualifies as normalization or aggregation.

350
MCQeasy

Which chart type is best for showing the sales pipeline from leads to closed deals, illustrating how many prospects drop off at each stage?

A.Treemap
B.Funnel chart
C.Waterfall chart
D.Stacked bar chart
AnswerB

A funnel chart directly encodes sequential stage attrition, plotting descending bars for each pipeline phase from leads to closed deals. Its tapering shape makes drop-off volume visually explicit at every stage, satisfying the stem's requirement to show how many prospects are lost between stages. Other chart types cannot represent ordered stage reduction as clearly.

Why this answer

A funnel chart is designed to show a sequential process where values decrease through stages, making it ideal for visualizing a sales pipeline from leads to closed deals and highlighting drop-off at each stage. Its tapering shape directly communicates conversion and attrition, which is exactly what the question asks for.

Exam trap

DA0-002 often tests the confusion between funnel charts and waterfall or stacked bar charts — candidates must recognize that funnel charts specifically depict sequential stage drop-off, while waterfall charts show cumulative positive/negative contributions.

How to eliminate wrong answers

Option A is wrong because a treemap displays hierarchical data as nested rectangles sized by value, which is suited to part-to-whole comparisons, not sequential stage drop-off. Option C is wrong because a waterfall chart shows how an initial value is increased or decreased by a series of positive and negative contributions, typically for financial variance analysis, not pipeline conversion stages. Option D is wrong because a stacked bar chart shows composition within categories across a dimension, but it does not inherently convey sequential drop-off or conversion rates through stages the way a funnel does.

351
MCQmedium

An organization is integrating data from multiple sources into a data warehouse. They need to handle differences in data granularity (e.g., daily vs. hourly sales data). Which technique is most appropriate?

A.Data aggregation
B.Data normalization
C.Data deduplication
D.Data profiling
AnswerA

Aggregation rolls hourly sales records up to daily totals, aligning finer-grained source data with the coarser warehouse grain. This resolves the daily-versus-hourly mismatch by summarising detail, which is precisely the granularity conflict the scenario describes.

Why this answer

Data aggregation is the correct technique because it allows the organization to roll up hourly sales data to a daily granularity, ensuring consistency when integrating sources with different levels of detail. By applying aggregation functions (e.g., SUM, AVG) during the ETL process, the data warehouse can store all data at a common grain, which is essential for accurate reporting and analysis.

Exam trap

The trap here is that candidates may confuse data normalization (a schema design concept) with the need to standardize data granularity, leading them to incorrectly select normalization instead of aggregation.

How to eliminate wrong answers

Option B is wrong because data normalization is a database design technique used to reduce redundancy and dependency by organizing columns and tables, not to reconcile differences in data granularity. Option C is wrong because data deduplication focuses on identifying and removing duplicate records, which does not address the mismatch in time-based granularity between daily and hourly data. Option D is wrong because data profiling is an exploratory process to assess data quality and structure, but it does not transform or harmonize data to a common granularity level.

352
Multi-Selecteasy

An organization is implementing a data lake to store raw data from various sources. Which THREE characteristics are typically associated with a data lake compared to a data warehouse?

Select 3 answers
A.Supports batch and real-time processing
B.Stores data in its native format
C.Schema-on-read approach
D.Supports only structured data
E.Requires data transformation before loading
AnswersA, B, C

A data lake ingests streams and files through the same storage layer, so it handles both batch loads and real-time processing. A data warehouse typically relies on scheduled ETL batches, making this a defining capability of the lake architecture.

Why this answer

A is correct because a data lake is designed to handle both batch and real-time/streaming ingestion and processing, unlike a traditional data warehouse that is primarily optimized for batch ETL workloads. B is correct because a data lake stores data in its native/raw format (e.g., JSON, Parquet, CSV, images, logs) without forcing an upfront conversion. C is correct because a data lake applies schema-on-read, meaning the schema is defined when the data is queried rather than when it is written.

D is incorrect because data lakes support structured, semi-structured, and unstructured data, not only structured data. E is incorrect because a data lake typically loads raw data as-is and defers transformation until read/query time, whereas a data warehouse usually requires transformation before loading.

Exam trap

CompTIA often tests the misconception that data lakes require data transformation before loading (schema-on-write), when in fact they use schema-on-read, allowing raw data storage without upfront transformation.

353
MCQhard

In Tableau, you want to create a view showing the total sales per region, but also want to allow users to filter by year without losing the ability to see all regions. What feature should you use to compute the total sales that ignores the year filter?

A.Table calculation running total
B.Parameter action
C.LOD expression: { FIXED [Region] : SUM([Sales]) }
D.Calculated field using SUM([Sales])
AnswerC

A FIXED level-of-detail expression computes SUM([Sales]) grouped solely by [Region], independent of any dimension filters applied in the view. This satisfies the stem's constraint: totals per region persist even when users filter by year, because FIXED LODs evaluate before dimension filters unless those dimensions are included in the expression's scope.

Why this answer

A Level of Detail (LOD) expression with FIXED can compute a value at the region level, ignoring other dimensions like year.

354
MCQmedium

A dataset contains a variable 'Income' with many missing values. The analyst decides to impute missing values with the median income of the non-missing values. Which type of imputation is this?

A.Interpolation
B.Deletion
C.Median imputation
D.Forward-fill imputation
AnswerC

Median imputation replaces each missing entry with the median calculated from the observed, non-missing incomes, satisfying the stem's requirement to use the median of non-missing values. Unlike mean imputation, the median resists distortion from outliers and skew, making it robust for income data, which is typically right-skewed.

Why this answer

Imputing missing values with the median of the non-missing values is exactly median imputation. It is a common method for handling missing data, especially when the distribution is skewed or contains outliers.

Exam trap

The trap is confusing median imputation with other imputation methods like interpolation or forward-fill. Candidates might also think deletion is a form of imputation, but it is not.

How to eliminate wrong answers

Option A is wrong because interpolation estimates missing values using surrounding data points, often in time series, not a global median. Option B is wrong because deletion removes missing data, not imputes it. Option D is wrong because forward-fill imputation carries forward the last observed value, typically used in time series, not a median.

355
MCQeasy

A data analyst is reviewing a dashboard that tracks monthly sales. The analyst notices that the 'Total Sales' metric for the current month is significantly lower than previous months, but the underlying daily sales data appears normal. The dashboard uses a live connection to the database. Which of the following is the most likely cause of the discrepancy?

A.The database connection is using a stale cache that has not refreshed since last month.
B.The dashboard's date filter is set to include only the first week of the current month.
C.The 'Total Sales' metric is calculated using a moving average that lags by one month.
D.The sales data for the current month has not been fully loaded into the database yet.
AnswerB

A date filter that restricts data to only the first week would cause the total to appear lower than a full month, even if daily sales are normal. This is a common dashboard configuration issue where filters are not updated to reflect the full current month-to-date period.

Why this answer

The most likely cause is a date filter limiting the data to a subset of the month, such as the first week, which would reduce the total while daily sales remain normal. Other options would affect the daily data or show different patterns, making them less consistent with the scenario.

Exam trap

The trap here is overlooking dashboard filter settings and assuming a data pipeline issue when the problem is often in the visualization configuration.

356
MCQmedium

A business analyst needs to explain to a sales director that the reported revenue has a 95% confidence interval of ±2%. Which concept is being communicated?

A.Data governance
B.Uncertainty communication
C.Single version of truth
D.Data lineage
AnswerB

A 95% confidence interval of ±2% communicates uncertainty: the true revenue plausibly lies within that range around the estimate. Framing it this way tells the sales director the precision of the reported figure, rather than presenting a single value as exact.

Why this answer

Confidence intervals quantify the uncertainty around a metric, conveying that the true value may vary.

357
MCQmedium

A data analyst wants to generate a report showing employee names and their department names, but some employees are not assigned to any department. The analyst wants to include all employees. Which JOIN type should be used?

A.INNER JOIN
B.LEFT JOIN
C.CROSS JOIN
D.RIGHT JOIN
AnswerB

A LEFT JOIN returns every row from the employees table and matches department rows where they exist, producing NULLs for unassigned staff. An INNER JOIN would drop those employees, so LEFT JOIN satisfies the requirement to include all employees.

Why this answer

LEFT JOIN includes all rows from the left table (employees) even if no match in departments.

358
MCQeasy

A data analyst needs to identify outliers in a dataset. Which of the following is a common method based on the interquartile range (IQR)?

A.Values more than 2 standard deviations from the mean
B.Values that are negative
C.Values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR
D.Values below the 5th percentile or above the 95th percentile
AnswerC

The IQR method flags observations falling outside Tukey's fences: Q1 minus 1.5 times IQR and Q3 plus 1.5 times IQR. Values beyond these bounds lie far from the central 50% of the distribution, making them statistical outliers.

Why this answer

A common rule is to consider any data point below Q1 - 1.5*IQR or above Q3 + 1.5*IQR as an outlier.

359
MCQeasy

A data analyst needs to count the number of distinct product categories in a table named 'products'. Which SQL function should be used in the SELECT clause?

A.COUNT(category)
B.DISTINCT COUNT(category)
C.COUNT(DISTINCT category)
D.COUNT(*) WHERE category IS NOT NULL
AnswerC

COUNT(DISTINCT category) returns the number of unique category values, satisfying the requirement to count distinct categories rather than all rows. Plain COUNT(category) would include duplicates and ignore nulls, so DISTINCT is essential to the stated goal.

Why this answer

The correct syntax to count unique non-null values in a column is COUNT(DISTINCT column_name). This function first eliminates duplicate values in the specified column and then counts the remaining distinct entries. In this scenario, COUNT(DISTINCT category) will return the number of unique product categories present in the 'products' table, ignoring any NULLs.

Exam trap

The trap here is confusing the correct placement of DISTINCT within the COUNT function; many candidates mistakenly write DISTINCT COUNT() or COUNT() DISTINCT, or they forget that COUNT(*) with a WHERE clause does not deduplicate.

How to eliminate wrong answers

Option A is wrong because COUNT(category) counts all non-null rows in the category column, including duplicates, so it would return the total number of products with a non-null category, not the number of distinct categories. Option B is wrong because DISTINCT COUNT(category) is not valid SQL syntax; DISTINCT is not a function and cannot be used as a prefix to COUNT in this manner. Option D is wrong because COUNT(*) WHERE category IS NOT NULL counts all rows where category is not null, but it does not eliminate duplicates, so it would give the total count of non-null categories, not the distinct count.

360
Multi-Selecthard

Which THREE of the following are assumptions of linear regression? (Select THREE).

Select 3 answers
A.Normal distribution of independent variables
B.Multicollinearity among independent variables
C.Independence of errors
D.Homoscedasticity (constant variance of errors)
E.Linearity between independent and dependent variables
AnswersC, D, E

Independence of errors means each residual is uncorrelated with the others, so no observation's error predicts another's. Violations, common in time-series or clustered data, bias standard errors and invalidate inference, making this a core linear regression assumption.

Why this answer

Option C is correct because linear regression assumes the residuals (errors) are independent of one another, meaning there is no autocorrelation; violations such as serial correlation in time-series data bias standard errors. Option D is correct because homoscedasticity—constant variance of the error terms across all levels of the predictors—is a core Gauss-Markov assumption; heteroscedasticity makes ordinary least squares estimates inefficient and invalidates inference. Option E is correct because the model assumes a linear relationship between each independent variable and the dependent variable, so the expected value of the outcome is a linear function of the coefficients.

Option A is not an assumption: normality applies to the error terms (and is only strictly needed for small-sample inference), not to the independent variables themselves. Option B is not an assumption but a problem to avoid—multicollinearity among predictors inflates coefficient variances and should be absent, not assumed.

Exam trap

The trap here is that candidates confuse the normality assumption for errors with a normality assumption for the independent variables, leading them to incorrectly select Option A.

361
Multi-Selecteasy

Which TWO of the following are examples of supervised learning algorithms?

Select 2 answers
A.Linear regression
B.K-means clustering
C.Principal component analysis (PCA)
D.Decision trees
E.Apriori algorithm
AnswersA, D

Linear regression is supervised because it trains on labelled data, learning a mapping from input features to a continuous target value by minimising error against known outcomes. This labelled-target dependency is precisely what the question's supervised learning criterion requires, unlike unsupervised methods that find structure without labels.

Why this answer

Linear regression (A) is a supervised learning algorithm because it learns a mapping from labeled input-output pairs by fitting a line (or hyperplane) that minimizes a loss such as mean squared error, making it a classic regression technique. Decision trees (D) are also supervised: they are trained on labeled data, recursively splitting features (e.g., via Gini impurity or information gain) to predict a target class or value. In contrast, K-means clustering (B) is unsupervised, as it groups unlabeled data by minimizing within-cluster variance without target labels.

Principal component analysis (C) is an unsupervised dimensionality-reduction technique that finds orthogonal components maximizing variance, requiring no labels. The Apriori algorithm (E) is unsupervised association-rule mining that discovers frequent itemsets from unlabeled transactional data, so it is not supervised learning.

Exam trap

CompTIA often tests the distinction between supervised and unsupervised learning by including clustering (K-means) and association (Apriori) as distractors, which candidates mistakenly think are supervised because they involve pattern discovery.

362
MCQeasy

A dashboard should be designed so that the most important metric is prominently displayed. This is an example of which design principle?

A.Data-ink ratio
B.Consistent color coding
C.Appropriate precision
D.Visual hierarchy
AnswerD

Visual hierarchy arranges size, colour, position and contrast so the eye lands first on the most important metric. This satisfies the stem's requirement that the key metric be prominently displayed rather than buried among secondary visuals.

Why this answer

Visual hierarchy is the design principle that dictates arranging elements to guide the viewer's eye to the most important information first. By making the key metric the most prominent element (e.g., larger, bolder, or top-left), the dashboard ensures immediate comprehension of the primary data point, which is critical for operational decision-making.

Exam trap

In the data-plus exam, visual hierarchy is often contrasted with data-ink ratio. Candidates may confuse 'making a metric prominent' with 'reducing non-data ink,' but the former is about emphasis while the latter is about eliminating visual clutter.

How to eliminate wrong answers

Option A is wrong because the data-ink ratio focuses on minimizing non-data ink (e.g., gridlines, borders) to maximize the proportion of ink used for actual data, not on emphasizing a specific metric. Option B is wrong because consistent color coding ensures that the same color represents the same category across visualizations, aiding comparison, but it does not inherently prioritize one metric over others. Option C is wrong because appropriate precision refers to displaying data with the correct level of detail (e.g., rounding to whole numbers when decimals are unnecessary), not to the visual prominence of a metric.

363
MCQeasy

A company needs to visualize the trend of monthly sales revenue over the past two years. Which chart type is most appropriate?

A.Bar chart
B.Scatter plot
C.Line chart
D.Pie chart
AnswerC

A line chart plots revenue against a continuous time axis, so the two-year monthly trend and its direction are read directly from the connecting segments. Categorical charts such as bar or pie cannot show temporal continuity or rate of change as clearly.

Why this answer

A line chart is the most appropriate for visualizing the trend of monthly sales revenue over time because it plots data points connected by lines, clearly showing the direction and pattern of change across a continuous time series. The x-axis represents time (months), and the y-axis represents revenue, allowing for easy identification of upward or downward trends, seasonality, and fluctuations. Unlike other chart types, line charts are specifically designed to highlight trends and changes over intervals.

Exam trap

The trap here is confusing the purpose of different chart types: candidates might select a bar chart because it also shows time on the x-axis, but the question emphasizes 'trend,' which is best represented by a line chart.

How to eliminate wrong answers

Option A is wrong because a bar chart is better suited for comparing discrete categories or showing rankings, not for emphasizing trends over time; while it can display time series, it is less effective for highlighting continuous trends. Option B is wrong because a scatter plot is used to show the relationship between two continuous variables (correlation), not to visualize a time-based trend. Option D is wrong because a pie chart is designed to show parts of a whole (proportions) at a single point in time, not to display trends over time.

364
MCQmedium

A data analyst is examining the relationship between advertising spend (in dollars) and revenue (in dollars). The Pearson correlation coefficient r is calculated as +0.92. Which of the following interpretations is correct?

A.There is a strong negative linear relationship.
B.There is no linear relationship.
C.There is a strong positive linear relationship.
D.92% of the variation in revenue is explained by advertising spend.
AnswerC

A coefficient of +0.92 sits close to +1, so the two variables move together almost perfectly along a straight line, with advertising spend rising as revenue rises. The positive sign confirms the direction is increasing rather than inverse, and the magnitude confirms the relationship is strong and linear.

Why this answer

A Pearson correlation coefficient of +0.92 indicates a strong positive linear relationship between advertising spend and revenue. The positive sign means as one variable increases, the other tends to increase, and the magnitude (close to 1) indicates a strong linear association.

Exam trap

The trap is confusing the correlation coefficient r with the coefficient of determination r²; candidates often mistakenly think r = 0.92 means 92% of the variation is explained, but it's actually r² that gives that percentage.

How to eliminate wrong answers

Option A is wrong because a negative relationship would have a negative r value (e.g., -0.92). Option B is wrong because an r near 0 would indicate no linear relationship, but 0.92 is far from 0. Option D is wrong because r^2 (coefficient of determination) represents the proportion of variance explained; r^2 = 0.92^2 = 0.8464, so about 84.6% of the variation is explained, not 92%.

365
MCQmedium

A data analyst wants to extract the year from a date column 'order_date' in a SQL database. Which function should be used?

A.YEAR(order_date)
B.DATEADD(year, order_date, 0)
C.DATEDIFF(year, order_date, GETDATE())
D.GETDATE()
AnswerA

The YEAR() function extracts the year component from a date or timestamp value, returning an integer. Applied to 'order_date', it isolates the year portion, satisfying the requirement to extract the year from that column in a single SQL expression.

Why this answer

The YEAR() function is a standard SQL date function that extracts the year from a date or datetime expression. It takes a single argument, the date column, and returns an integer representing the year. This is the most direct and appropriate function for the task.

Exam trap

DA0-002 often tests the confusion between functions that extract date parts and those that manipulate or compare dates, leading candidates to pick DATEADD or DATEDIFF.

How to eliminate wrong answers

Option B is wrong because DATEADD is used to add or subtract a specified time interval from a date, not to extract a component. Option C is wrong because DATEDIFF calculates the difference between two dates in a specified unit, not extract a part of a date. Option D is wrong because GETDATE() returns the current system date and time, not a component of a given column.

366
Multi-Selectmedium

A data steward at a retail company is classifying data assets. The company wants to identify which data elements are considered structured data. Which two of the following are examples of structured data? (Choose two.)

Select 2 answers
A.A table of customer orders with columns OrderID, CustomerName, and OrderDate.
B.A collection of emails stored as plain text files.
C.A spreadsheet with rows for employees and columns for EmployeeID, Name, and Department.
D.A JSON document containing nested product details and customer reviews.
E.A video recording of a customer service call.
AnswersA, C

This is structured data because it is organized into a tabular format with predefined columns and data types. Each row represents a single order, and each column holds a specific attribute. The rigid schema makes it easy to query, sort, and analyze using SQL or other relational tools, which is the hallmark of structured data.

Why this answer

Structured data is organized into a predefined format, typically tables with rows and columns, making it easily searchable and analyzable. The customer orders table and the employee spreadsheet both fit this description. Emails, JSON documents, and video recordings lack the rigid schema and tabular organization required for structured data.

Exam trap

The trap here is assuming that any data with labels or tags is structured, but semi-structured formats like JSON still lack the fixed tabular schema that defines structured data.

367
Multi-Selectmedium

A data governance council at an insurance company is establishing a data stewardship program. The council wants to define the core responsibilities that data stewards will own within their assigned domains. Which two responsibilities are most appropriate for data stewards? (Choose two.)

Select 2 answers
A.Negotiating vendor contracts for third-party data purchases
B.Approving the enterprise-wide data governance budget
C.Defining and maintaining data quality rules and standards for their domain
D.Resolving data quality issues and documenting root causes within their domain
E.Building the physical ETL pipelines that load data into the warehouse
AnswersC, D

Data stewards are accountable for the fitness of data in their domains, which includes authoring quality rules such as validity ranges, completeness thresholds, and referential integrity checks. They translate business expectations into measurable criteria and monitor adherence. This custodial responsibility directly supports the council's goal of embedding accountability for data health across domains.

Why this answer

Data stewards own the health of data in their domains, which centers on defining quality rules and standards and resolving quality issues with documented root causes. Budget approval, pipeline engineering, and vendor contracting are responsibilities of executives, engineers, and procurement respectively. Keeping stewards focused on quality definition and remediation preserves clear accountability within the governance model.

Exam trap

The trap here is expanding the steward role into technical or commercial territory, when stewardship centers on defining and defending data quality within a domain.

368
MCQmedium

A data analyst is examining a dataset of daily website visits collected over the past year. The analyst wants to determine whether the distribution of visits on weekdays differs significantly from the distribution on weekends, without assuming normality. The analyst has already separated the data into two independent groups. Which statistical test should the analyst use?

A.One-sample t-test
B.Mann-Whitney U test
C.Paired samples t-test
D.Chi-square goodness-of-fit test
AnswerB

The Mann-Whitney U test is a nonparametric alternative to the independent samples t-test, comparing whether two independent groups come from the same distribution. It does not assume normality, making it appropriate for skewed visit counts. It ranks all observations and assesses whether one group tends to have higher values. This directly addresses the analyst's question of whether weekday and weekend visit distributions differ without distributional assumptions.

Why this answer

The Mann-Whitney U test is a nonparametric method that compares two independent groups without requiring normality. It assesses whether one group tends to have larger values than the other by ranking all observations. Since the analyst wants to compare weekday and weekend website visits and cannot assume normality, this test is the correct choice.

It is specifically designed for independent groups and continuous or ordinal data.

Exam trap

The trap here is assuming that any nonparametric test for two groups works, but the paired t-test and one-sample t-test require different data structures, while the chi-square test is for categorical frequencies, not continuous distributions.

369
MCQeasy

Refer to the exhibit. What data quality issue is indicated?

A.Data inconsistency
B.Non-standardized data entry
C.Outlier
D.Data duplication
AnswerB

The exhibit shows the same values recorded in inconsistent formats and spellings, indicating non-standardised data entry. This inconsistency stems from missing input validation or controlled vocabularies at the point of capture rather than duplication or completeness problems.

Why this answer

Option B is correct because the exhibit shows the same logical value entered in multiple inconsistent formats (e.g., 'NY', 'New York', 'new york', 'N.Y.'), which is the hallmark of non-standardized data entry. The data is not duplicated (different spellings), not an outlier (no extreme numeric value), and not a referential inconsistency (no conflicting values across tables) — it is a formatting/standardization problem at the point of capture.

Exam trap

The trap here is confusing non-standardized data entry (formatting variants of the same value) with data duplication (identical repeated records) — candidates often pick duplication because both involve 'the same thing appearing multiple times'.

How to eliminate wrong answers

Option A is wrong because data inconsistency refers to conflicting values for the same attribute across systems or records (e.g., a customer's birthdate differing between CRM and billing), not multiple spellings of the same value. Option C is wrong because an outlier is a data point that lies far outside the expected distribution (e.g., an age of 250), which is a statistical anomaly, not a formatting issue. Option D is wrong because data duplication means the same record appears more than once with identical or near-identical values, whereas here the values differ in format, indicating entry standardization failure rather than duplication.

370
MCQeasy

A company needs to store raw, unprocessed data from IoT sensors for future machine learning experiments. The data is in various formats and schemas are not yet defined. Which storage solution is most appropriate?

A.Data lake
B.Data mart
C.Data warehouse
D.Operational database
AnswerA

A data lake stores raw data in native formats without requiring a predefined schema, so varied IoT sensor payloads can be landed as-is. Schema-on-read defers structure until analysis, satisfying the undefined-schema, future-machine-learning requirement that a warehouse or relational store cannot meet.

Why this answer

A data lake is the correct choice because it stores raw, unprocessed data in its native format (structured, semi-structured, or unstructured) without requiring a predefined schema. This aligns perfectly with the need to ingest IoT sensor data in various formats for future machine learning experiments, where schemas are not yet defined. Unlike data warehouses or data marts, a data lake supports schema-on-read, allowing the data to be transformed and queried later as needed.

Exam trap

CompTIA often tests the misconception that 'raw data' belongs in a data warehouse because it is 'data,' but the trap is that data warehouses require structured, processed data with a fixed schema, while a data lake is specifically designed for raw, schema-less data storage.

How to eliminate wrong answers

Option B is wrong because a data mart is a subset of a data warehouse designed for a specific business line or department, requiring pre-defined schemas and processed data, not raw unprocessed data. Option C is wrong because a data warehouse stores structured, cleaned, and transformed data optimized for business intelligence and reporting, not raw data in various formats. Option D is wrong because an operational database (e.g., OLTP system) is designed for real-time transaction processing with strict schemas and ACID compliance, not for storing large volumes of raw, schema-less IoT data for future analytics.

371
Multi-Selectmedium

A data governance team is defining roles and responsibilities for data management. Which TWO of the following are common data governance roles? (Select TWO).

Select 2 answers
A.Data steward
B.Data owner
C.Data scientist
D.Database administrator
E.Data analyst
AnswersA, B

The data steward executes day-to-day governance: maintaining metadata, enforcing quality rules, resolving data issues and applying the owner's policies. This operational role satisfies the need for hands-on responsibility, bridging policy set by owners and the practical management of data assets.

Why this answer

Option A, Data steward, is correct because a data steward is a core data governance role responsible for day-to-day management of data assets, enforcing data policies, maintaining data quality, and resolving data issues within a defined domain. Option B, Data owner, is correct because the data owner is an accountable governance role, typically a senior business leader, who is responsible for the data's classification, access decisions, quality, and compliance with policies for a specific data domain. These two roles are explicitly defined in governance frameworks such as DAMA-DMBOK, which distinguishes accountable owners from custodial stewards.

Option C, Data scientist, is not a governance role but an analytics role focused on building models and deriving insights from data. Option D, Database administrator, is an operational/technical role managing database systems, not a governance role. Option E, Data analyst, is also an analytics role focused on querying and interpreting data rather than defining governance responsibilities.

Exam trap

DA0-002 often tests the distinction between governance roles and operational/analytical roles; candidates may incorrectly select data scientist or database administrator because they are familiar technical roles, but they are not governance-specific.

372
MCQeasy

In a dataset of customer orders, you need to count the number of distinct customers who have placed orders. Which SQL aggregate function should you use?

A.DISTINCT COUNT(customer_id)
B.COUNT(customer_id)
C.COUNT(DISTINCT customer_id)
D.COUNT(*)
AnswerC

COUNT(DISTINCT customer_id) eliminates duplicate customer identifiers before tallying, returning the number of unique customers rather than total order rows. This directly satisfies the requirement to count distinct customers, whereas plain COUNT(customer_id) would include repeat purchasers multiple times and overstate the customer base.

Why this answer

The correct syntax to count unique non-null values in a column is COUNT(DISTINCT column_name). In this case, COUNT(DISTINCT customer_id) returns the number of different customers who have placed at least one order. This is the standard SQL aggregate function designed for exactly this purpose, and it ignores NULLs in the customer_id column.

Exam trap

The trap here is confusing the syntax of DISTINCT with COUNT, leading candidates to pick 'DISTINCT COUNT(customer_id)' instead of the correct 'COUNT(DISTINCT customer_id)'. Many also mistakenly think COUNT(customer_id) automatically deduplicates, but it does not.

How to eliminate wrong answers

Option A is wrong because DISTINCT COUNT(customer_id) is not valid SQL syntax; DISTINCT is a keyword used inside the COUNT function, not a standalone function. Option B is wrong because COUNT(customer_id) counts all non-null customer_id values, including duplicates, so it returns the total number of orders (with a customer_id) rather than the number of distinct customers. Option D is wrong because COUNT(*) counts all rows in the result set, including rows with NULL customer_id, and does not deduplicate, so it gives the total number of orders, not distinct customers.

373
MCQeasy

A data analyst needs to identify the most frequently occurring value in a dataset. Which measure of central tendency should they use?

A.Mode
B.Standard deviation
C.Median
D.Mean
AnswerA

The mode identifies the value appearing most frequently, directly satisfying the requirement to find the most frequently occurring value. Unlike the mean, which averages all values, or the median, which finds the middle position, the mode alone measures frequency of occurrence, making it the appropriate measure of central tendency for this scenario.

Why this answer

The mode is the measure of central tendency that identifies the most frequently occurring value in a dataset. Unlike the mean or median, the mode directly counts the frequency of each distinct value and returns the value with the highest count, making it the correct choice for this specific requirement.

Exam trap

The trap here is that candidates often confuse 'most frequently occurring' with 'average' or 'middle value' and incorrectly choose mean or median, especially when the dataset is numeric and they assume central tendency always refers to mean.

How to eliminate wrong answers

Option B (Standard deviation) is wrong because it measures the dispersion or spread of data points around the mean, not the frequency of occurrence of any single value. Option C (Median) is wrong because it identifies the middle value when the dataset is sorted, which does not indicate which value appears most often. Option D (Mean) is wrong because it calculates the arithmetic average of all values, which can be skewed by outliers and does not reflect frequency of occurrence.

374
Multi-Selectmedium

Which TWO chart types are best suited for visualizing the distribution of a single continuous variable? (Select two.)

Select 2 answers
A.Scatter plot
B.Box plot
C.Line chart
D.Histogram
E.Pie chart
AnswersB, D

A box plot summarises a continuous variable's distribution through its median, quartiles and whiskers, directly satisfying the stem's requirement to show distribution of one continuous variable. It also exposes skewness and outliers, which histograms alone can obscure, making it a valid selection alongside the histogram.

Why this answer

A box plot (B) is correct because it summarizes the distribution of a single continuous variable using the median, quartiles, and potential outliers, making spread and skewness directly visible. A histogram (D) is correct because it bins a single continuous variable into intervals and displays frequencies, revealing the shape, center, and spread of the distribution. A scatter plot (A) is not appropriate here because it requires two continuous variables to show their relationship.

A line chart (C) is used to show trends over an ordered sequence, typically time, not the distribution of one variable. A pie chart (E) displays proportions of a whole for categorical data, not the distribution of a continuous variable.

Exam trap

DA0-002 often tests the confusion between charts for distribution (histogram, box plot) and charts for relationships (scatter plot) or trends (line chart), so candidates must match the chart type to the analytical purpose.

375
MCQmedium

A data analyst is examining a dataset of employee salaries and notices that the distribution is heavily right-skewed due to a few executives with very high salaries. The analyst wants to apply a transformation to make the distribution more symmetric for further analysis. Which transformation is most appropriate?

A.Square root transformation
B.Logarithmic transformation
C.Square transformation
D.Exponential transformation
AnswerB

The logarithmic transformation is highly effective for right-skewed data with positive values, as it compresses the upper tail and reduces the influence of extreme values. For salary data with a few very high earners, taking the log of salaries will make the distribution more symmetric and closer to normal. This is a common practice before applying statistical techniques that assume normality.

Why this answer

For right-skewed data with positive values, the logarithmic transformation is the most appropriate to reduce skewness and make the distribution more symmetric. It compresses the upper tail, mitigating the influence of extreme high values. Square root is milder, while square and exponential transformations would increase skewness.

Therefore, the log transformation is the best choice.

Exam trap

The trap here is confusing transformations that increase skewness (square, exponential) with those that reduce it (log, square root).

Page 4

Page 5 of 14

Page 6