Courseiva

CompTIA Data+ (DA0-002) (DA0-002) — Questions 901–975

1004 questions total · 14pages · All types, answers revealed

Page 12

Page 13 of 14

Page 14
901
MCQeasy

Which data quality dimension is most concerned with whether data values fall within a defined domain or acceptable range?

A.Completeness
B.Consistency
C.Validity
D.Accuracy
AnswerC

Validity checks that values conform to a defined domain, format or acceptable range, directly matching the stem's requirement. Unlike accuracy, which concerns correspondence to real-world truth, or consistency, which concerns agreement across sources, validity enforces permitted value sets and boundaries, so out-of-range or malformed entries fail the check.

Why this answer

Validity refers to whether data values conform to defined rules or constraints.

902
Multi-Selecthard

A retail analytics team is building a data catalog and must document the metadata for a new sales fact table. The team needs to record structural metadata that describes how the data is organized. Which two items qualify as structural metadata for this table? (Choose two.)

Select 2 answers
A.The name of the business owner accountable for the sales data
B.The primary key constraint that enforces uniqueness on the sales transaction identifier
C.The list of columns in the table with each column's declared data type
D.The timestamp recording when the table was last refreshed from the source system
E.A free-text business definition explaining what a completed sale means to the merchandising team
AnswersB, C

A primary key constraint defines how rows are uniquely identified and how the table's structure enforces integrity, which is structural metadata. It tells consumers which column or columns distinguish records and informs join planning and validation. Ownership and load timestamps describe governance and operations, not the table's internal organization, so the constraint is the structural element.

Why this answer

Structural metadata describes how data is organized: the columns and their data types, plus constraints such as primary keys that define row identity. Ownership, refresh timestamps, and business definitions describe accountability, operational state, and semantics respectively, which fall under administrative, operational, or descriptive categories. Documenting the column list with types and the primary key constraint therefore satisfies the structural requirement for the catalog.

Exam trap

The trap here is treating any catalog entry as structural metadata, when ownership, freshness, and business definitions describe governance, operations, and semantics rather than the table's organization.

903
MCQhard

In Tableau, a data analyst creates a calculated field to compute the average sales per customer. The analyst wants this calculation to remain constant regardless of the level of detail in the view. Which Tableau feature should be used?

A.Table calculation
B.Level of Detail expression
C.Parameter
D.Filter
AnswerB

A Level of Detail expression fixes the aggregation at a declared granularity, such as customer, independently of the dimensions in the view. This satisfies the requirement that the average sales per customer stays constant regardless of the view's level of detail.

Why this answer

A Level of Detail (LOD) expression fixes the granularity of a calculation independently of the dimensions in the view, so the average sales per customer stays constant no matter how the view is sliced. LOD expressions like {FIXED [Customer] : AVG([Sales])} compute at the specified level and then can be reused at any view level of detail.

Exam trap

The trap is confusing table calculations with LOD expressions — both can compute averages, but only LOD expressions are independent of the view's level of detail, which is exactly what the question requires.

How to eliminate wrong answers

Option A is wrong because table calculations operate on the data already aggregated in the view and depend on the view's level of detail, so they change as dimensions change. Option C is wrong because parameters are user-input values that substitute into calculations or filters; they do not control aggregation granularity. Option D is wrong because filters restrict rows, not the level at which a calculation is computed, and can even alter LOD results depending on filter order.

904
MCQhard

A data scientist applies K-means clustering to a customer dataset. The elbow method suggests using 4 clusters. After running K-means with k=4, the within-cluster sum of squares (WCSS) is plotted against k, and the elbow is at k=4. What does this indicate?

A.Increasing k beyond 4 would not significantly reduce WCSS.
B.The data naturally forms 4 clusters with no noise.
C.The algorithm converged to a local minimum.
D.The model has overfit the data.
AnswerA

The elbow marks where additional clusters stop yielding meaningful WCSS reduction. At k=4 the curve flattens, so moving to k=5 or beyond adds complexity without materially lowering within-cluster sum of squares, confirming four clusters as the sensible choice.

Why this answer

The elbow method plots the within-cluster sum of squares (WCSS) against the number of clusters k. The 'elbow' point indicates where the rate of decrease in WCSS sharply changes, meaning that adding more clusters beyond that point yields diminishing returns. Therefore, at k=4, increasing k further would not significantly reduce WCSS, suggesting 4 is a reasonable choice for the number of clusters.

Exam trap

The trap is interpreting the elbow as a definitive indication of the true number of clusters; candidates must remember that the elbow method is a heuristic and does not guarantee that the data naturally forms that many clusters.

How to eliminate wrong answers

Option B is wrong because the elbow method does not guarantee that the data naturally forms exactly 4 clusters with no noise; it only suggests a good trade-off between model complexity and fit. Option C is wrong because the elbow method is about choosing k, not about convergence to a local minimum; K-means can converge to a local minimum regardless of k. Option D is wrong because overfitting would be indicated by a very low WCSS with many clusters, but the elbow method helps avoid overfitting by selecting a simpler model.

905
MCQhard

Refer to the exhibit. What is the best corrective action to resolve this error?

A.Convert the 'revenue' column to numeric data type during ETL
B.Change the chart type to a bar chart
C.Remove the 'revenue' column from the visualization
D.Use a string-compatible chart type
AnswerA

The error stems from the 'revenue' column being stored as text, so aggregation and arithmetic operations fail or sort lexically. Casting it to a numeric type during ETL ensures values load correctly and downstream calculations behave as expected.

Why this answer

The error shown in the exhibit is a type mismatch: the 'revenue' column contains string values, but the visualization expects a numeric measure for aggregation and charting. The correct fix is to convert the column to a numeric data type during the ETL process so that calculations and visual encodings work as intended. Changing the chart type or removing the column only masks the symptom rather than resolving the underlying data quality issue.

Exam trap

DA0-002 often tests the confusion between fixing a data quality issue at the source (ETL type conversion) versus applying cosmetic workarounds in the visualization layer, which do not resolve the root cause.

How to eliminate wrong answers

Option B is wrong because changing the chart type does not fix the data type mismatch; the numeric aggregation will still fail or produce incorrect results. Option C is wrong because removing the 'revenue' column eliminates a critical business metric from the visualization instead of correcting its data type. Option D is wrong because using a string-compatible chart type treats revenue as a categorical label, which prevents any mathematical aggregation and defeats the purpose of the analysis.

906
Multi-Selecthard

Which THREE of the following are properties of ratio data? (Choose THREE.)

Select 3 answers
A.Data can be categorized into groups
B.Allows negative values
C.Supports multiplication and division
D.Intervals between values are equal
E.Has a meaningful zero point
AnswersC, D, E

Ratio data allows meaningful ratios (e.g., twice as heavy).

Why this answer

Ratio data supports multiplication and division because it has a true, meaningful zero point that indicates the absence of the measured attribute. This allows ratios to be computed (e.g., one value is twice another), which is a defining property of ratio scales in measurement theory.

Exam trap

The trap here is that candidates confuse the 'meaningful zero' property with the ability to have negative values, or they think categorization is a defining feature of ratio data, when it is actually a property shared by all measurement scales.

907
MCQmedium

A data analyst is examining the relationship between advertising spend (in thousands) and sales (in thousands). The Pearson correlation coefficient is computed as r = -0.85. Which of the following interpretations is correct?

A.There is no linear relationship.
B.There is a strong positive linear relationship between advertising spend and sales.
C.There is a weak negative linear relationship.
D.There is a strong negative linear relationship.
AnswerD

The magnitude of r, 0.85, exceeds the conventional 0.7 threshold for strong association, and the negative sign indicates that as advertising spend rises, sales tend to fall along a linear trend. Direction and strength are therefore both captured by this interpretation.

Why this answer

Pearson r measures linear correlation: -0.85 indicates a strong negative linear relationship (as one increases, the other decreases). The magnitude |0.85| is close to 1, so strong.

908
MCQhard

A data analyst is evaluating a classification model for predicting customer churn. The model's confusion matrix shows 100 true positives, 20 false positives, 30 false negatives, and 850 true negatives. The analyst wants to assess the model's ability to correctly identify actual churners. Which metric should the analyst use?

A.Precision
B.Accuracy
C.Specificity
D.Recall
AnswerD

Recall, also known as sensitivity, is the proportion of actual positives that are correctly identified, calculated as TP/(TP+FN) = 100/130 ≈ 0.769. It measures the model's ability to find all actual churners, which aligns with the analyst's goal. A high recall means few churners are missed, which is often critical in churn prediction to avoid losing customers.

Why this answer

The analyst wants to assess the model's ability to correctly identify actual churners, which is the definition of recall (sensitivity). Recall is calculated as true positives divided by the sum of true positives and false negatives. In this scenario, recall is 100/(100+30) ≈ 0.769.

Precision and specificity focus on different aspects, and accuracy can be misleading in imbalanced datasets.

Exam trap

The trap here is confusing recall with precision; recall measures how many actual positives were found, while precision measures how many predicted positives were correct.

909
MCQhard

A data analyst trains a complex model that achieves 99% accuracy on training data but only 65% on new data. What is the most likely issue?

A.Underfitting
B.Overfitting
C.Multicollinearity
D.High bias
AnswerB

Overfitting occurs when a model learns noise and idiosyncrasies in the training set rather than generalisable patterns, producing the 99% versus 65% gap. The stem's constraint — high training accuracy with poor unseen-data performance — is the defining signature of variance-dominated overfitting, so regularisation, pruning or more data would be required.

Why this answer

The model performs exceptionally well on training data (99% accuracy) but poorly on new data (65% accuracy), which is the classic symptom of overfitting. Overfitting occurs when the model learns noise and specific patterns in the training data rather than generalizing to unseen data, often due to excessive complexity (e.g., too many parameters or deep layers). This results in high variance and poor performance on validation or test sets.

Exam trap

CompTIA often tests the distinction between overfitting and underfitting by presenting a large gap between training and test accuracy, tempting candidates to choose high bias or multicollinearity due to confusion about bias-variance tradeoff or correlation issues.

How to eliminate wrong answers

Option A is wrong because underfitting would show poor performance on both training and new data (e.g., low accuracy on both), not high training accuracy with low test accuracy. Option C is wrong because multicollinearity refers to high correlation among predictor variables in regression models, which inflates coefficient standard errors but does not directly cause a large gap between training and test accuracy. Option D is wrong because high bias typically leads to underfitting, where the model is too simple and performs poorly on both training and test data, not the specific pattern of high training accuracy and low test accuracy seen here.

910
MCQmedium

An e-commerce company is acquiring product data from multiple supplier APIs. The APIs return JSON with inconsistent field naming conventions. Which data acquisition technique should be applied?

A.Data compression
B.Data mapping and transformation
C.Data deduplication
D.Data aggregation
AnswerB

Mapping reconciles differing field names across supplier schemas into one canonical structure, while transformation standardises values. This satisfies the stem's inconsistent naming constraint, since raw ingestion would yield mismatched columns that cannot be joined or queried consistently.

Why this answer

Data mapping and transformation is the correct technique because the JSON responses from different supplier APIs use inconsistent field naming conventions (e.g., 'product_id' vs. 'ProductID'). This technique defines a schema to map source fields to a standardized target format, ensuring data consistency before loading into the company's system. Without transformation, downstream processes like analytics or inventory management would fail due to mismatched field names.

Exam trap

The trap here is that candidates confuse data transformation with data aggregation or deduplication, assuming any processing step can fix schema inconsistencies, but only mapping and transformation directly address field naming and structure mismatches.

How to eliminate wrong answers

Option A is wrong because data compression reduces storage size or transfer bandwidth, but does not address structural inconsistencies in field naming. Option C is wrong because data deduplication removes duplicate records based on content, but does not reconcile different field names or schemas. Option D is wrong because data aggregation summarizes or combines data (e.g., sums, averages), but does not resolve naming conflicts or schema mismatches.

911
MCQeasy

A data analyst is cleaning a dataset and finds that the 'age' column has several missing values. Which of the following is a valid method for handling missing numerical data?

A.Delete the entire column
B.Ignore the missing values
C.Impute with the mean
D.Replace with zeros
AnswerC

Imputing with the mean replaces each missing 'age' entry with the column's arithmetic average, preserving the existing sample size and keeping the dataset's overall mean unchanged. This satisfies the stem's requirement for a valid numerical handling method, since the mean is calculable only on quantitative data such as age.

Why this answer

Imputing missing numerical values with the mean is a standard and valid statistical approach for handling missing data in a numerical column. It preserves the overall distribution's central tendency and allows the analyst to retain all other rows for analysis. This method is especially appropriate when the missingness is random and the column is roughly normally distributed, minimizing bias introduced by dropping records.

Exam trap

The trap here is confusing 'ignoring' missing values with a valid handling method, or assuming that any replacement (like zeros) is acceptable without considering the data's meaning.

How to eliminate wrong answers

Option A is wrong because deleting the entire column discards all data in that feature, including the non-missing values, which is rarely justified unless the column is entirely unusable. Option B is wrong because ignoring missing values is not a handling method; many algorithms and calculations will fail or produce incorrect results with NaN values, and it does not resolve the missingness. Option D is wrong because replacing missing ages with zeros introduces a false value (age 0) that is not a valid age for the dataset, severely distorting the mean, variance, and any downstream analysis.

912
MCQhard

A dataset contains transaction amounts with a few extremely high values. The analyst wants to reduce the impact of these outliers on the average. Which measure of central tendency is most robust?

A.Mean
B.Median
C.Mode
D.Standard deviation
AnswerB

The median is a positional measure, so extreme high transaction values shift it far less than the mean, which sums every value. This robustness to outliers satisfies the stem's requirement to reduce their impact on the average.

Why this answer

Median is not affected by extreme values, while mean is sensitive.

913
Multi-Selectmedium

A data analyst is using Tableau to build a dashboard. Which THREE features are available in Tableau for creating interactive dashboards?

Select 3 answers
A.Dashboard actions
B.Calculated fields
C.DAX measures
D.Parameters
E.Power Query
AnswersA, B, D

Dashboard actions add interactivity by triggering behaviour from user clicks, such as filter, highlight, URL, or parameter actions. This satisfies the stem's requirement for interactive dashboard features, letting viewers drive exploration rather than viewing a static image.

Why this answer

Dashboard actions (A) are a Tableau feature that lets a dashboard respond to user interaction—such as filter, highlight, URL, or go-to-sheet actions—making the dashboard interactive. Calculated fields (B) are Tableau expressions (using functions like IF, SUM, and LOD expressions) that create new derived data, and they can be exposed as interactive controls through parameters or actions. Parameters (D) are Tableau's user-controlled input values (for example, a slider or dropdown) that can drive calculated fields, filters, and reference lines, directly enabling interactivity.

DAX measures (C) belong to Microsoft Power BI, not Tableau, and Power Query (E) is a Power BI/Excel data-transformation tool, so neither is available in Tableau for building dashboards.

Exam trap

DA0-002 often tests whether candidates can distinguish Tableau-native features (actions, parameters, calculated fields) from Power BI/Excel features (DAX, Power Query) — remember DAX and Power Query belong to Microsoft, not Tableau.

914
MCQhard

A data analyst is using a decision tree to predict customer churn. The tree is very deep, with many branches, and achieves 99% accuracy on the training data but only 70% on the test data. Which technique should the analyst use to address this issue?

A.Use a larger training dataset
B.Decrease the size of the test set
C.Increase the maximum depth of the tree
D.Apply pruning to the decision tree
AnswerD

Pruning reduces the size of the tree by removing branches that provide little predictive power, which helps combat overfitting. By simplifying the model, pruning improves generalization to unseen data. In this scenario, the large gap between training and test accuracy indicates overfitting, so pruning is the appropriate technique to enhance test performance.

Why this answer

Pruning is a technique specifically designed to reduce overfitting in decision trees by removing unnecessary branches. The large discrepancy between training and test accuracy indicates the model is too complex and has memorized the training data. Pruning simplifies the tree, improving its ability to generalize.

Other options either worsen overfitting or do not directly address the model complexity.

Exam trap

The trap here is thinking that more data or a deeper tree will solve overfitting, when the real issue is model complexity that requires pruning.

915
MCQhard

A dataset contains a column 'Education Level' with values: 'High School', 'Bachelor', 'Master', 'PhD'. An analyst computes the average by assigning numbers 1-4. Which data concept is being violated?

A.Misclassifying data as structured
B.Treating ordinal data as interval
C.Treating nominal data as ordinal
D.Treating ratio data as interval
AnswerB

Education levels are ordinal: order matters but gaps between categories are not equal, so distances between 1 and 2 versus 3 and 4 are meaningless. Averaging the numeric codes treats those ranks as interval data with equal spacing, producing a statistically invalid mean.

Why this answer

The analyst assigned numeric values (1-4) to 'Education Level' categories and computed an average. This treats the ordinal data as if it were interval data, assuming equal spacing between categories (e.g., the difference between 'High School' and 'Bachelor' is the same as between 'Master' and 'PhD'), which is not valid. Ordinal data only preserves order, not magnitude or equal intervals, so calculating a mean is inappropriate.

Exam trap

CompTIA often tests the distinction between ordinal and interval scales by presenting a scenario where a mean is computed on ranked categories, tempting candidates to think the error is about nominal vs. ordinal (Option C) rather than the misuse of arithmetic operations on ordinal data.

How to eliminate wrong answers

Option A is wrong because misclassifying data as structured refers to incorrectly labeling unstructured data (e.g., text) as structured, but the dataset already has a structured column; the violation is about measurement scale, not structure. Option C is wrong because treating nominal data as ordinal would involve imposing an order on unordered categories (e.g., colors), but 'Education Level' already has a natural order, so the error is not about misordering but about assuming equal intervals. Option D is wrong because treating ratio data as interval would ignore a true zero point (e.g., income), but 'Education Level' has no meaningful zero, so the violation is not about ratio vs. interval but about ordinal vs. interval.

916
Multi-Selectmedium

A data architect is designing a system to handle a workload that requires strong consistency and complex multi-table joins for financial reporting. Which two characteristics should the architect prioritize when selecting the database technology? (Choose two.)

Select 2 answers
A.Denormalized wide-column storage to optimize write throughput
B.A schema-less document model to accommodate evolving report formats
C.A relational schema with well-defined foreign key relationships
D.ACID compliance to guarantee transaction integrity
E.Eventual consistency to maximize availability during network partitions
AnswersC, D

A relational schema with foreign keys enforces referential integrity and supports the complex multi-table joins required for financial reporting. Well-defined relationships allow the query optimizer to execute joins efficiently and ensure that related records remain consistent. This characteristic aligns directly with both the strong consistency and complex join requirements in the scenario.

Why this answer

Financial reporting demands strong consistency and the ability to join data across multiple related tables. ACID compliance guarantees that transactions complete reliably without partial updates, and a relational schema with foreign keys enforces the relationships that make complex joins both possible and efficient. Together these characteristics satisfy the core requirements, while eventual consistency, wide-column storage, and document models each sacrifice one or both of these needs.

Exam trap

The trap here is assuming that any modern database can handle financial reporting, when the combination of strong consistency and complex joins specifically points to ACID-compliant relational systems.

917
MCQmedium

During exploratory data analysis, you calculate the IQR for a numeric column and find that several data points fall below Q1 - 1.5*IQR. These points are likely:

A.Normal variations within the distribution
B.The mode of the dataset
C.The median of the dataset
D.Outliers
AnswerD

Values below Q1 - 1.5*IQR sit outside Tukey's fence, the standard IQR rule for flagging extreme observations. Since the stem specifies this exact threshold, those points are classified as outliers, though they may still be legitimate values requiring investigation rather than automatic removal.

Why this answer

The 1.5×IQR rule is the standard Tukey fence for identifying outliers in exploratory data analysis. Any value below Q1 − 1.5×IQR (or above Q3 + 1.5×IQR) is flagged as a potential outlier. Since the question states points fall below Q1 − 1.5×IQR, they are by definition outliers.

Exam trap

The trap here is confusing the IQR outlier rule with measures of central tendency — candidates may pick 'mode' or 'median' because those are familiar summary statistics, but the 1.5×IQR fence is specifically the outlier detection rule.

How to eliminate wrong answers

Option A is wrong because normal variations within the distribution fall inside the fences (between Q1 − 1.5×IQR and Q3 + 1.5×IQR); points outside the fence are flagged as unusual, not normal. Option B is wrong because the mode is the most frequently occurring value and has no relationship to the IQR fence calculation. Option C is wrong because the median is the middle value (Q2) and is always inside the IQR; it cannot be identified by falling below Q1 − 1.5×IQR.

918
MCQhard

You are a data analyst at a retail company. Your team has deployed a new dashboard that displays real-time inventory levels across all warehouses. The dashboard uses a traffic light color scheme (green = sufficient, yellow = low, red = out of stock). After one week, the warehouse managers complain that the dashboard is overwhelming because it shows all 500+ products at once, and they cannot quickly identify which items need immediate attention. They request a way to prioritize items that are red or yellow. Additionally, the dashboard updates every second, causing the screen to flicker constantly, which they find distracting. You need to redesign the dashboard to address these issues. Which course of action should you take?

A.Add a filter to show only items with red or yellow status, and reduce the refresh rate to every 5 minutes.
B.Replace the dashboard with a static table that shows all products sorted alphabetically, and increase the refresh rate to every 2 seconds.
C.Remove the color scheme and display only numerical values, and keep the refresh rate at 1 second.
D.Add additional colors to distinguish more inventory levels and increase the refresh rate to every 500 milliseconds.
AnswerA

Filtering to red and yellow status surfaces only items needing action, while a five-minute refresh removes the flicker caused by per-second updates. Together these satisfy both stated constraints: prioritising urgent stock and eliminating the distracting constant refresh.

Why this answer

It directly addresses both user complaints: adding a filter to show only red/yellow items reduces cognitive overload by hiding the 500+ products that are not critical, and reducing the refresh rate from every second to every 5 minutes eliminates distracting screen flicker while still providing near-real-time inventory updates. This approach balances usability with the need for timely alerts without overwhelming the user.

Exam trap

The trap here is that candidates may think increasing refresh rate improves data accuracy, but the real issue is user distraction and information overload, not data latency.

How to eliminate wrong answers

Option B is wrong because sorting alphabetically does not help prioritize critical items (red/yellow), and increasing the refresh rate to every 2 seconds would worsen the flicker issue. Option C is wrong because removing the color scheme removes the quick visual prioritization that the traffic light system provides, and keeping the 1-second refresh rate does not solve the flicker problem. Option D is wrong because adding more colors would increase visual complexity, not reduce it, and increasing the refresh rate to every 500 milliseconds would make the flicker even more distracting.

919
MCQeasy

A data analyst is summarizing the central tendency of a dataset with extreme outliers. Which measure is most robust to outliers?

A.Standard deviation
B.Median
C.Mean
D.Range
AnswerB

The median is positional, not arithmetic: it depends only on the middle value's rank, so extreme outliers cannot drag it. The mean incorporates every value's magnitude, making it highly sensitive to skew. For skewed data, the median better represents typical central tendency.

Why this answer

The median is the middle value when data is ordered, so extreme outliers (whether very high or very low) do not affect its position. Unlike the mean, which incorporates every value and can be pulled strongly by outliers, the median remains stable. Therefore, for a dataset with extreme outliers, the median is the most robust measure of central tendency.

Exam trap

The trap here is confusing measures of central tendency with measures of dispersion, and assuming that the mean is always the best measure of center. Candidates often pick the mean out of habit, ignoring its sensitivity to outliers.

How to eliminate wrong answers

Option A is wrong because standard deviation is a measure of dispersion, not central tendency, and it is highly sensitive to outliers since it squares deviations from the mean. Option C is wrong because the mean is the arithmetic average and is directly influenced by every value, so extreme outliers can skew it dramatically. Option D is wrong because the range is the difference between the maximum and minimum values, making it entirely dependent on outliers and not a measure of central tendency at all.

920
MCQmedium

An analyst is creating a report that includes multiple charts. To ensure the audience quickly grasps the key insight, which principle of data storytelling should be applied?

A.Avoid using titles to reduce clutter
B.Include a legend for every chart
C.Use a title that states the main insight
D.Place the chart before any explanation
AnswerC

A declarative title delivers the insight before the audience interprets the chart, satisfying the requirement that they grasp the key message quickly. Descriptive titles merely label axes; an insight title states the finding itself, directing attention to the conclusion the data supports.

Why this answer

In data storytelling, chart titles should communicate the insight, not just describe the data. A title like 'Q3 revenue dropped 12% in EMEA' tells the audience the takeaway immediately, whereas 'Q3 Revenue by Region' forces them to interpret the chart. This aligns with the principle of reducing cognitive load and leading with the message.

Exam trap

The trap is equating minimalism with effectiveness — candidates may think fewer labels means cleaner charts, but the exam rewards clarity of insight, which titles deliver.

How to eliminate wrong answers

Option A is wrong because removing titles increases ambiguity — titles are essential for orienting the audience and stating the point, not clutter. Option B is wrong because a legend is only needed when multiple series require identification; adding a legend to every chart can add noise when a single series or direct labeling suffices. Option D is wrong because placing a chart before any explanation forces the audience to guess the point; effective storytelling frames the insight first (or in the title) so the visual supports a clear message.

921
MCQhard

A data team is ingesting JSON event data from a mobile application into a columnar warehouse. Each event has a nested array of product objects, and analysts frequently need to report on individual products within those events. The team wants to avoid repeated manual parsing in every query. Which approach best prepares the data for efficient product-level analysis?

A.Convert the entire event to a single delimited text field and split it into columns by position.
B.Keep the nested array as-is and rely on the warehouse's native semi-structured query functions for every report.
C.Flatten the nested product array into a separate table with one row per product and a foreign key back to the event, during the load process.
D.Store the raw JSON as a single string column and parse it at query time using string functions.
AnswerC

Shredding the nested array into a child table normalizes the data so each product becomes a queryable row, which columnar warehouses handle efficiently. Analysts can then join or aggregate without repeating parsing logic. This approach also supports indexing and consistent semantics, directly meeting the requirement for efficient product-level reporting.

Why this answer

The requirement is reusable, efficient product-level analysis from nested JSON. Shredding the array into a child table during load normalizes the structure, lets the columnar warehouse store each product as a row, and removes repeated parsing from reports. Storing raw JSON, querying nested arrays each time, or converting to positional text all defer or damage the preparation work.

Exam trap

The trap here is assuming modern warehouses make nested JSON queries so convenient that no load-time transformation is needed, when repeated parsing still harms performance and consistency.

922
Multi-Selectmedium

Which TWO are valid data acquisition methods? (Select two.)

Select 2 answers
A.Web scraping
B.Data normalization
C.API calls
D.Data encryption
E.Data profiling
AnswersA, C

Web scraping programmatically extracts data from web pages, making it a recognised acquisition method for gathering external, publicly available datasets. It satisfies the stem's requirement for valid acquisition methods, alongside APIs, by pulling content directly from source sites.

Why this answer

Web scraping (A) is a valid data acquisition method because it programmatically extracts data from web pages, typically via HTTP requests and HTML parsing (e.g., using libraries like BeautifulSoup or Scrapy), making it a recognized way to collect data from sources that lack a formal data feed. API calls (C) are also a valid acquisition method, since they retrieve structured data directly from a service's application programming interface using protocols such as REST/HTTP or SOAP, often returning JSON or XML payloads. In contrast, data normalization (B) is a data preparation/transformation step that scales or restructures values, not a way to obtain data.

Data encryption (D) is a security control that protects data confidentiality, and data profiling (E) is an analysis technique that examines data quality and structure — neither acquires new data.

Exam trap

DA0-002 often tests the confusion between data acquisition (collecting data) and data preparation activities (normalizing, profiling, encrypting), which are downstream steps in the data pipeline.

923
MCQeasy

You are using pandas in Python to clean a dataset. You notice several rows with missing values in the 'age' column. Which method would you use to remove those rows?

A.df.drop_duplicates()
B.df.dropna()
C.df.fillna(0)
D.df.isna()
AnswerB

`df.dropna()` removes every row containing any null value, directly satisfying the requirement to eliminate rows with missing 'age' entries. Its default `axis=0` and `how='any'` parameters target rows rather than columns, so the cleaned DataFrame retains only complete records without additional filtering logic.

Why this answer

The df.dropna() method is specifically designed to remove rows (or columns) that contain missing values (NaN). By default, it drops any row where at least one NaN is present, which directly addresses the requirement to remove rows with missing 'age' values. This is the standard pandas approach for handling incomplete records when deletion is preferred over imputation.

Exam trap

The trap here is confusing methods for detecting missing values (isna) with those for removing them (dropna), or mistakenly thinking fillna removes rows when it actually replaces values.

How to eliminate wrong answers

Option A is wrong because df.drop_duplicates() removes duplicate rows based on all columns (or a subset), not rows with missing values; it does not consider NaN as a criterion for removal. Option C is wrong because df.fillna(0) replaces missing values with 0 (or a specified value) rather than removing the rows, which would keep the rows but alter the data. Option D is wrong because df.isna() returns a boolean DataFrame indicating which cells are missing, but it does not modify the DataFrame or remove any rows; it is typically used for detection, not cleaning.

924
MCQmedium

A data analyst is creating a report on customer satisfaction scores. The analyst wants to ensure that regional managers can only see data for their own region. Which security measure should be applied?

A.Data encryption
B.Row-level security
C.Single version of truth
D.Data masking
AnswerB

Row-level security filters table rows by the user's identity, so each regional manager's query returns only their own region's satisfaction scores. This directly satisfies the stem's constraint that managers see solely their region's data, enforced at query time within the semantic model rather than through separate reports or workspace permissions.

Why this answer

Row-level security (RLS) filters query results at the row level based on the identity of the user running the report, so each regional manager automatically sees only the rows belonging to their region. This is the standard mechanism in BI platforms (e.g., Power BI, Tableau, Salesforce) for enforcing per-user data segmentation without duplicating reports. Encryption, masking, and single-version-of-truth do not restrict which rows a user can retrieve.

Exam trap

The trap here is confusing data protection techniques (encryption, masking) with access-control techniques (row-level security) — candidates often pick encryption because it sounds like the strongest security measure, but it does not restrict which rows a user can see.

How to eliminate wrong answers

Option A is wrong because data encryption protects data confidentiality at rest or in transit but does not filter which rows a given user can query — an authorized manager could still decrypt and view other regions' data. Option C is wrong because 'single version of truth' is a data governance principle about having one authoritative dataset, not an access-control mechanism. Option D is wrong because data masking obscures sensitive field values (e.g., showing only last four digits of an SSN) but does not restrict row visibility by user identity.

925
MCQmedium

A database administrator is designing a normalized database to reduce data redundancy. They have a table with columns: OrderID, ProductID, ProductName, and Quantity. The table is currently in 1NF. To move to 2NF, which issue must be resolved?

A.The table has repeating groups
B.ProductName depends only on ProductID, causing a partial dependency
C.Quantity depends on both OrderID and ProductID
D.The table has a transitive dependency
AnswerB

ProductName depends only on ProductID, not on the full composite key (OrderID, ProductID), which is a partial dependency violating 2NF. Resolving it means moving ProductName into a separate Products table keyed by ProductID, leaving Quantity and OrderID in the order line table.

Why this answer

To move from 1NF to 2NF, the table must have no partial dependencies. A partial dependency occurs when a non-key attribute depends on only part of a composite primary key. Here, the composite key is (OrderID, ProductID).

ProductName depends only on ProductID, not on the full key, so it is a partial dependency. Option A (repeating groups) is a violation of 1NF, not 2NF, and the table is already in 1NF. Option C is incorrect because Quantity depends on both OrderID and ProductID (it is fully functionally dependent on the composite key).

Option D is incorrect because a transitive dependency (where a non-key attribute depends on another non-key attribute) is a 3NF issue, not 2NF. Therefore, the correct answer is B.

Exam trap

CompTIA Data+ often tests the distinction between partial dependencies (2NF) and transitive dependencies (3NF), so candidates mistakenly choose a transitive dependency when the real issue is a partial dependency on a composite key.

How to eliminate wrong answers

Option A is wrong because repeating groups are a 1NF violation, and the table is already stated to be in 1NF, so this issue is already resolved. Option C is wrong because Quantity depending on both OrderID and ProductID is a full functional dependency on the composite key, which is acceptable and does not violate 2NF. Option D is wrong because a transitive dependency (where a non-key column depends on another non-key column) is a 3NF violation, not a 2NF issue.

926
MCQmedium

A data analyst is cleaning a dataset and finds that a numeric field has several missing values. The variable is normally distributed. Which imputation method is most appropriate?

A.Median imputation
B.Mean imputation
C.Mode imputation
D.Forward-fill
AnswerB

For a normally distributed variable, the mean preserves the central tendency and keeps the overall distribution shape intact, unlike median or mode imputation. Since the data is symmetric, the mean is the best unbiased estimate for replacing missing numeric entries, satisfying the normality constraint in the stem.

Why this answer

For a normally distributed numeric variable, mean imputation is the most appropriate method because the mean is the central tendency that best represents the typical value in a symmetric distribution. Replacing missing values with the mean preserves the overall mean of the variable and is statistically sound when the data is normally distributed. Other methods like median or mode are better for skewed or categorical data, respectively.

Exam trap

The trap is that candidates might choose median imputation thinking it's always more robust, but for a normal distribution, mean imputation is the standard; the question specifies normal distribution to guide you to the mean.

How to eliminate wrong answers

Option A is wrong because median imputation is more appropriate for skewed distributions or when outliers are present, as the median is robust to extreme values; for a normal distribution, the mean is preferred. Option C is wrong because mode imputation is used for categorical data or discrete variables, not continuous numeric variables; it would be inappropriate here. Option D is wrong because forward-fill is used for time-series data to carry forward the last observed value, which assumes a temporal order and is not suitable for a general numeric variable without such structure.

927
Multi-Selecthard

Which TWO of the following are best practices when creating a data visualization for a presentation to stakeholders?

Select 2 answers
A.Including interactive filters when possible to allow stakeholders to explore data
B.Using a consistent color scheme that aligns with company branding
C.Placing the legend at the bottom of the chart
D.Using 3D effects to make the chart visually appealing
E.Using only one type of chart throughout the presentation for consistency
AnswersA, B

Interactivity lets stakeholders focus on relevant data.

Why this answer

Interactive filters (e.g., slicers, parameter controls) empower stakeholders to drill down into specific data subsets, ask ad-hoc questions, and uncover insights without overwhelming the initial view. This aligns with the principle of user-driven exploration, which is a best practice for stakeholder presentations where different audience members may have varying interests. Including interactivity also reduces the need for multiple static charts, keeping the presentation concise while still offering depth.

Exam trap

CompTIA often tests the misconception that visual embellishments like 3D effects or consistent chart types are always beneficial, when in fact they can obscure data accuracy and limit communicative effectiveness.

928
Multi-Selecthard

A data analyst is designing a database for a retail application. Which TWO of the following are valid reasons to use a NoSQL document database like MongoDB instead of a relational database? (Select 2)

Select 2 answers
A.The application requires high-speed transactional consistency
B.The data structure evolves frequently
C.The data is hierarchical, such as orders with line items
D.The data has a fixed schema with many relationships
E.The application needs complex joins across multiple tables
AnswersB, C

Document stores allow schema flexibility.

Why this answer

NoSQL document databases like MongoDB are schema-flexible, allowing the data structure to evolve over time without requiring migrations or downtime. This is ideal for agile development where application requirements change frequently, as documents can have varying fields without breaking existing records.

Exam trap

The trap here is that candidates often assume NoSQL databases are always faster or more consistent, but the exam tests the specific trade-offs: document databases excel at flexible schemas and hierarchical data, not at transactional consistency or complex joins.

929
MCQmedium

A data analyst is profiling a dataset of customer records and discovers that 12% of the rows have missing values in the "postal_code" column, while the "customer_id" column is fully populated. The analyst needs to document this finding for a data quality report. Which data quality dimension does the missing postal code values primarily violate?

A.Uniqueness
B.Completeness
C.Accuracy
D.Consistency
AnswerB

Completeness measures the extent to which required data is present without missing values. With 12% of postal codes absent, the column fails to fully represent all customer records, directly violating completeness. Documenting this gap helps downstream processes assess whether imputation or source correction is needed before analysis.

Why this answer

Completeness assesses whether all required values are present. The postal_code column has 12% missing entries, meaning the dataset does not fully capture that attribute for all customers, which is the defining characteristic of a completeness violation. Accuracy, consistency, and uniqueness address correctness, agreement, and duplication respectively, none of which describe absent values.

Exam trap

The trap here is conflating missing values with incorrect values, which leads to selecting accuracy instead of recognizing that absence of data is a completeness issue.

930
MCQeasy

A data analyst calculates the mean, median, and mode of a dataset. Which of the following measures of central tendency is least affected by extreme outliers?

A.Median
B.Range
C.Mode
D.Mean
AnswerA

The median depends only on positional order, so extreme values shift it minimally. The mean sums every value and is dragged strongly by outliers, while the mode reflects only the most frequent value and can be unstable in continuous data.

Why this answer

The median is the middle value when data is ordered, so it depends only on the rank position of values, not their magnitude. Extreme outliers, even if arbitrarily large or small, do not change which value sits in the middle unless they cross the 50th percentile. This makes the median a robust measure of central tendency, unlike the mean, which incorporates every value and is pulled toward outliers.

Exam trap

The trap here is confusing measures of central tendency with measures of dispersion, leading candidates to select 'Range' because it sounds like a statistical measure, or to assume the mode is always the most robust since it ignores magnitudes—when in fact the median is the standard robust measure of center tested on DA0-002.

How to eliminate wrong answers

Option B is wrong because the range is not a measure of central tendency at all—it is a measure of dispersion calculated as max minus min, and it is extremely sensitive to outliers since a single extreme value directly sets the maximum or minimum. Option C is wrong because the mode is the most frequent value and, while not affected by the magnitude of outliers, it can be unstable or uninformative in continuous data and is not considered a robust measure of center in the same way; moreover, an outlier could theoretically become the mode if it repeats, but more importantly, the question asks for the least affected measure among central tendency options, and the median is the standard robust choice. Option D is wrong because the mean sums all values and divides by n, so a single extreme outlier shifts the mean substantially, making it the most outlier-sensitive measure of central tendency.

931
MCQhard

A data analyst uses the elbow method to determine the number of clusters for k-means. The plot shows a sharp bend at k=3 and a small bend at k=5. What is the recommended number of clusters?

A.5
B.The method is inconclusive.
C.2
D.3
AnswerD

The elbow method selects the k where the within-cluster sum of squares drops sharply before levelling off. The sharp bend at k=3 marks that inflection point, so three clusters is recommended; the minor bend at k=5 is a secondary, weaker signal.

Why this answer

The elbow method suggests choosing k where the decrease in inertia becomes marginal; the sharp bend at 3 indicates the optimal k.

932
MCQhard

An analyst is creating a data story about the decline in customer satisfaction scores. The data shows that satisfaction dropped after a software update. Which narrative structure best helps the audience understand the cause and effect?

A.Start with the current scores, then show historical trend
B.Begin with the software update event, then show satisfaction before and after
C.Present all data points without a story
D.Use a pie chart of satisfaction categories
AnswerB

Opening with the software update establishes the causal event, then before-and-after satisfaction figures demonstrate its effect. This chronological cause-then-effect ordering directly satisfies the requirement to make the relationship between the update and the decline understandable.

Why this answer

It follows a chronological cause-and-effect narrative, starting with the software update event and then showing satisfaction scores before and after. This structure directly maps to the audience's need to understand the causal relationship, as it highlights the intervention point and the resulting change in the metric. In data storytelling, this is known as the 'before-and-after' or 'change-over-time' narrative, which is most effective for demonstrating impact.

Exam trap

The trap here is that candidates often choose Option A because they think starting with the current state is more engaging, but CompTIA Data+ tests the understanding that a chronological cause-and-effect narrative is required to clearly demonstrate the impact of a specific event, not just a general trend.

How to eliminate wrong answers

Option A is wrong because starting with current scores and then showing historical trend presents the data in reverse chronological order, which obscures the cause-and-effect relationship and forces the audience to infer the impact of the software update from the end of the story. Option C is wrong because presenting all data points without a story fails to provide context or a narrative arc, leaving the audience to interpret the cause-and-effect on their own, which is inefficient and can lead to misinterpretation. Option D is wrong because a pie chart of satisfaction categories shows a static distribution of current satisfaction levels, not the temporal change before and after the software update, and thus cannot illustrate cause and effect.

933
MCQeasy

A hospital's patient records system must process thousands of small transactions per second. Which type of database system is best suited for this workload?

A.Data mart
B.OLTP
C.Data warehouse
D.OLAP
AnswerB

OLTP systems are optimised for high volumes of small, concurrent read/write transactions with strong consistency and row-level locking, matching the hospital's thousands-per-second small transaction requirement. OLAP instead suits large analytical queries, so it cannot meet this throughput pattern.

Why this answer

OLTP (Online Transaction Processing) systems are designed to handle a high volume of small, concurrent transactions with low latency and high concurrency. This makes them ideal for a hospital patient records system that must process thousands of small transactions per second, such as patient check-ins, prescription updates, and billing entries.

Exam trap

The trap here is that candidates often confuse OLTP with OLAP, mistakenly thinking that 'processing many transactions' implies analytical processing, when in fact OLTP is the correct choice for high-frequency, small, write-heavy workloads.

How to eliminate wrong answers

Option A is wrong because a data mart is a subset of a data warehouse focused on a specific business line (e.g., cardiology), not designed for high-throughput transactional processing. Option C is wrong because a data warehouse is optimized for complex analytical queries on large historical datasets, not for handling thousands of small, real-time transactions per second. Option D is wrong because OLAP (Online Analytical Processing) is used for multidimensional analysis and reporting, not for high-frequency transactional workloads.

934
MCQmedium

A hospital's analytics team is designing a new repository for electronic health records. The records include patient demographics, lab results, and physician notes, and the schema must remain flexible because new lab test types are added frequently. The team needs to enforce relationships between patients, encounters, and lab orders while keeping the ability to evolve the schema. Which data model best fits these requirements?

A.Relational model
B.Document model
C.Key-value model
D.Graph model
AnswerA

The relational model stores data in tables with primary and foreign keys, which directly enforces the patient-to-encounter-to-lab-order relationships the team requires. Schema evolution is supported through migrations such as adding new tables or columns for new lab test types. This combination of strong referential integrity and structured change management matches the hospital's need to keep records consistent while the schema grows.

Why this answer

The requirement to enforce relationships between patients, encounters, and lab orders points to a model with primary and foreign keys, which is the relational model. It also supports schema evolution through controlled migrations when new lab test types appear. Key-value, graph, and document models offer flexibility or traversal but do not natively enforce the structured relationships clinical records demand.

Exam trap

The trap here is focusing only on schema flexibility and choosing a document store, while overlooking the explicit requirement to enforce relationships between entities.

935
MCQhard

A data analyst is building a decision tree to classify whether customers will churn. The analyst wants to prevent the tree from overfitting the training data. Which technique should the analyst use?

A.Increasing the maximum depth of the tree
B.Setting the minimum samples per leaf to 1
C.Pruning the tree after full growth
D.Using all available features without selection
AnswerC

Pruning involves growing a full tree and then removing branches that provide little predictive power, reducing complexity and overfitting. This technique, such as cost-complexity pruning, balances bias and variance, improving generalization. For churn classification, pruning helps avoid capturing noise in the training data, leading to more robust predictions on new customers.

Why this answer

Pruning is a standard technique to reduce decision tree overfitting by removing branches that do not improve predictive accuracy on validation data. It simplifies the model and enhances generalization. Increasing depth, using all features, or allowing single-sample leaves all increase complexity and overfitting risk, making them incorrect choices for this scenario.

Exam trap

The trap here is thinking that a fully grown tree is always best, but without pruning it will overfit and perform poorly on new data.

936
MCQhard

A data scientist has a dataset with 50 variables and wants to identify clusters of similar observations. Which visualization technique is most suitable for reducing dimensionality to 2D while preserving cluster structure?

A.Heatmap of correlations
B.Scatter matrix (pairplot)
C.Parallel coordinates plot
D.Scatter plot of first two principal components
AnswerD

A scatter plot of the first two principal components projects the 50-variable dataset onto two orthogonal axes of greatest variance, preserving cluster separation. This satisfies the dimensionality-reduction constraint: PCA compresses correlated variables into uncorrelated components, letting similar observations group visibly in 2D.

Why this answer

A scatter plot of the first two principal components applies PCA to project 50-dimensional data onto two axes that capture the most variance, preserving the global structure and cluster separation. It is the standard dimensionality-reduction visualization for cluster exploration because distances between points in PC space approximate distances in the original feature space. Other listed techniques either show pairwise relationships or all dimensions without reducing to 2D.

Exam trap

The trap is choosing a visualization that shows all variables (pairplot, parallel coordinates) instead of one that actually reduces dimensionality to 2D while preserving cluster structure — only PCA scatter does both.

How to eliminate wrong answers

Option A is wrong because a correlation heatmap shows pairwise variable relationships, not observation-level clusters, and does not reduce dimensionality to 2D. Option B is wrong because a scatter matrix plots every pair of variables, producing a 50x50 grid that is unreadable and does not perform dimensionality reduction. Option C is wrong because a parallel coordinates plot displays all 50 dimensions as axes, which becomes visually cluttered and does not project to 2D.

937
MCQhard

An analyst needs to compute a running total of sales for each department, ordered by date. Which window function is most appropriate?

A.ROW_NUMBER() OVER (PARTITION BY department ORDER BY date)
B.SUM(sales) OVER (ORDER BY date)
C.SUM(sales) OVER (PARTITION BY department ORDER BY date)
D.LAG(sales, 1) OVER (PARTITION BY department ORDER BY date)
AnswerC

SUM with PARTITION BY department and ORDER BY date produces a cumulative running total within each department, because the default frame spans from the first row to the current row. Partitioning resets the accumulation per department, while the date ordering guarantees the running sequence the analyst requires.

Why this answer

The running total of sales for each department, ordered by date, requires a window function that sums sales over a partition by department and orders by date. SUM(sales) OVER (PARTITION BY department ORDER BY date) computes a cumulative sum within each department as the date progresses, which is exactly a running total.

Exam trap

DA0-002 often tests the omission of PARTITION BY in window functions, causing candidates to compute a running total across all departments instead of per department.

How to eliminate wrong answers

Option A is wrong because ROW_NUMBER() assigns a sequential integer to each row within the partition, not a running total. Option B is wrong because SUM(sales) OVER (ORDER BY date) computes a running total across the entire result set, not per department; it lacks the PARTITION BY clause. Option D is wrong because LAG(sales, 1) returns the sales value from the previous row, which is not a running total.

938
MCQhard

A logistics company maintains a master customer table used by billing, operations, and analytics. Two source systems submit conflicting addresses for the same customer, and reports now disagree. Which governance practice should be applied to resolve the conflict and prevent recurrence?

A.Archive both source records and report only the customer identifier without an address attribute
B.Allow each consuming system to keep its own address and add a note explaining the discrepancy
C.Increase the frequency of the nightly extract so conflicting addresses are overwritten sooner
D.Define a survivorship rule with a documented system of record and apply it during master data consolidation
AnswerD

Survivorship rules specify which source value prevails when records conflict, and designating a system of record anchors that decision in an authoritative source. Applying the rule during consolidation produces one consistent golden record, so billing, operations, and analytics stop diverging. This directly resolves the conflict and establishes a repeatable standard that prevents the same dispute from recurring.

Why this answer

Conflicting source values are resolved through survivorship rules that encode which system or attribute wins, coupled with an explicitly designated system of record. Applying those rules during consolidation yields a single golden record that all consumers share, eliminating disagreement and giving stewards a repeatable decision framework. Faster refreshes, per-system copies, or removing the attribute each leave the underlying authority question unanswered.

Exam trap

The trap here is assuming that a technical fix, such as more frequent loading, resolves conflicting data when the real issue is undefined authority and survivorship.

939
MCQeasy

An e-commerce company wants to provide real-time personalized product recommendations based on customer browsing behavior. Currently, they have a traditional data warehouse that processes batch updates every night. The marketing team complains that recommendations are outdated within hours because customers see yesterday's data. The data engineer needs to modify the architecture to support near-real-time analytics. The budget is limited, and the existing warehouse infrastructure must be reused as much as possible. Which architectural change would best meet the requirement?

A.Replace the warehouse with an in-memory database for real-time processing.
B.Add more nodes to the warehouse cluster to speed up batch processing.
C.Implement a streaming data pipeline (e.g., Apache Kafka) that feeds a real-time recommendation engine.
D.Increase the frequency of batch load from nightly to every hour.
AnswerC

Kafka ingests events continuously and pushes them to the recommendation engine within seconds, eliminating the nightly batch latency that made recommendations stale. It satisfies the near-real-time requirement while the existing warehouse remains in place, keeping costs within the limited budget.

Why this answer

Implementing a streaming data pipeline like Apache Kafka enables the ingestion and processing of customer browsing events in near real-time, feeding a dedicated recommendation engine that can update recommendations within seconds or minutes. This approach reuses the existing data warehouse for historical analytics and batch reporting while adding a lightweight streaming layer for low-latency recommendations, aligning with the limited budget and reuse requirement.

Exam trap

The trap here is that candidates may assume increasing batch frequency (Option D) is sufficient for near-real-time needs, but the Data+ exam tests the understanding that 'near-real-time' typically requires sub-minute latency, which batch processing cannot achieve due to scheduling overhead and resource contention.

How to eliminate wrong answers

Option A is wrong because replacing the warehouse with an in-memory database would discard the existing infrastructure entirely, incurring high migration costs and losing the warehouse's batch processing capabilities for other workloads, which violates the constraint to reuse the existing warehouse. Option B is wrong because adding more nodes to the warehouse cluster only improves the throughput of batch processing, but does not reduce the latency of data freshness—recommendations would still be based on data that is at least hours old, failing the near-real-time requirement. Option D is wrong because increasing batch frequency to every hour still introduces a delay of up to 60 minutes, which is insufficient for real-time personalization; moreover, frequent batch loads can cause resource contention and degrade warehouse performance for other queries.

940
Multi-Selectmedium

A data analyst is performing hypothesis testing to compare the mean sales of two store locations. Which TWO conditions must be satisfied to use a two‑sample t‑test? (Select TWO.)

Select 2 answers
A.The data is paired between the two locations
B.The sample sizes are equal
C.The data is approximately normally distributed
D.The variances of the two populations are equal
E.The two samples are independent of each other
AnswersC, E

Normality is assumed for the t-test, though it is robust for large samples.

Why this answer

The two-sample t-test assumes that the data in each group are approximately normally distributed. This is a key parametric assumption; if the sample sizes are large (typically n > 30), the Central Limit Theorem can relax this requirement, but for smaller samples, normality must hold to ensure valid test statistics and p-values.

Exam trap

CompTIA often tests the misconception that equal sample sizes or equal variances are required for a two-sample t-test, but the actual core assumptions are independence and normality (or large sample sizes via CLT).

941
MCQmedium

A data analyst is performing data profiling on a customer table. Which metric provides the number of unique values in a column?

A.Row count
B.Cardinality
C.Standard deviation
D.Null count
AnswerB

Cardinality counts the distinct values present in a column, directly satisfying the requirement for the number of unique values during data profiling. Unlike row count, which totals all records including duplicates, cardinality reveals value distribution and repetition, exposing low-variance or high-uniqueness columns that affect indexing and query planning decisions.

Why this answer

Cardinality refers to the number of distinct (unique) values in a column, which is exactly the metric described. It is a fundamental data profiling statistic used to assess column uniqueness, identify candidate keys, and inform indexing or partitioning decisions. High cardinality columns (e.g., primary keys) have many unique values; low cardinality columns (e.g., status flags) have few.

Exam trap

The trap is confusing cardinality with row count or null count — candidates often assume 'unique values' means total rows, but cardinality specifically counts distinct values, which can be far fewer than the row count.

How to eliminate wrong answers

Option A is wrong because row count measures the total number of records in the table, not the number of distinct values in a specific column. Option C is wrong because standard deviation measures the dispersion or spread of numeric values around the mean, which is unrelated to uniqueness. Option D is wrong because null count measures how many rows have missing values in a column, which is a completeness metric, not a uniqueness metric.

942
Multi-Selecthard

A data analyst is preparing a dataset for a machine learning model and notices that several numerical features have different scales (e.g., age ranges from 18-90, income from 20,000-200,000). The analyst decides to apply feature scaling. Which TWO of the following statements about scaling techniques are correct? (Choose two.)

Select 2 answers
A.Standardization (Z-score normalization) transforms features to have a mean of 0 and a standard deviation of 1, and is less affected by outliers than min-max scaling.
B.Min-max normalization transforms features to a fixed range, typically [0, 1], and is sensitive to outliers.
C.Feature scaling is unnecessary for distance-based algorithms like k-nearest neighbors (KNN) because they are scale-invariant.
D.Standardization scales features to a range of [0, 1] and is preferred when features have different units.
E.Min-max normalization preserves the shape of the original distribution and is robust to outliers.
AnswersA, B

Standardization subtracts the mean and divides by the standard deviation, resulting in a distribution with mean 0 and standard deviation 1. It does not bound values to a specific range, so outliers have less impact on the scaling of the majority of data compared to min-max normalization.

Why this answer

Min-max normalization scales to a fixed range but is sensitive to outliers, while standardization centers data with unit variance and is less affected by outliers. Both are common scaling techniques. The other statements misrepresent properties of scaling or incorrectly claim that distance-based algorithms are scale-invariant.

Exam trap

The trap here is confusing the properties of min-max normalization and standardization, especially regarding outlier sensitivity and output range.

943
MCQmedium

An analyst is performing a linear regression and obtains an R-squared value of 0.85. Which of the following is the best interpretation?

A.85% of the residuals are zero.
B.85% of the data points lie on the regression line.
C.There is an 85% chance that the relationship is causal.
D.The model explains 85% of the variability in the dependent variable.
AnswerD

R-squared measures the proportion of variance in the dependent variable accounted for by the regression model. A value of 0.85 therefore means the predictors explain 85% of that variability, leaving 15% unexplained by the model.

Why this answer

R-squared (R²) is the coefficient of determination, representing the proportion of variance in the dependent variable that is predictable from the independent variable(s). An R² of 0.85 means that 85% of the variability in the dependent variable is explained by the regression model, indicating a strong fit. This is the standard statistical interpretation.

Exam trap

DA0-002 often tests the interpretation of R-squared, and candidates frequently confuse it with correlation, causation, or the proportion of data points on the line.

How to eliminate wrong answers

Option A is wrong because R-squared does not measure the proportion of residuals that are zero — residuals are the differences between observed and predicted values, and R² does not imply any specific number of zero residuals. Option B is wrong because R-squared does not indicate the percentage of data points that lie exactly on the regression line; it measures explained variance, not exact fit. Option C is wrong because R-squared does not imply causation — correlation does not equal causation, and R² is a measure of association, not causal inference.

944
MCQhard

A data analyst is working with a relational database that contains a table of customer orders. To optimize query performance for a report that filters by order date and customer ID, the analyst wants to create an index. Which type of index would be most effective for queries that filter on both columns?

A.B-tree index on order_date
B.Hash index on customer_id
C.Composite index on (order_date, customer_id)
D.Clustered index on order_id
AnswerC

A composite index stores the two key columns together in a defined order, so the database can satisfy the combined order_date and customer_id filter from a single index structure rather than intersecting separate single-column indexes or scanning the table.

Why this answer

A composite B-tree index on (order_date, customer_id) allows the database to efficiently satisfy equality and range predicates on both columns in a single index scan. B-tree indexes support ordered traversal and range lookups, making them ideal for date-based filtering combined with an equality filter on customer_id. This index structure minimizes the number of rows scanned by leveraging the index's leading column for the date range and the second column for the customer ID match.

Exam trap

The trap here is that candidates often choose a single-column index (A or B) thinking it will be sufficient, not realizing that a composite index is required to avoid a 'filter' step that scans many rows after the index lookup.

How to eliminate wrong answers

Option A is wrong because a single-column B-tree index on order_date can only efficiently filter by date; any additional filter on customer_id would require a separate lookup or a full scan of the date-matched rows, leading to poor performance. Option B is wrong because a hash index on customer_id only supports equality lookups and cannot handle range queries on order_date, making it unsuitable for date-range filtering. Option D is wrong because a clustered index on order_id physically reorders the table by order_id, which does not help with filtering on order_date or customer_id and may even degrade performance for these queries due to unnecessary key lookups.

945
MCQmedium

After presenting a data analysis to stakeholders, they ask for additional details on a specific trend. The analyst needs to provide a follow-up report. What is the best approach?

A.Create a new dashboard with all the data.
B.Send the raw data file to the stakeholders.
C.Provide a detailed written report with charts and explanations.
D.Hold another meeting to verbally explain the trend.
AnswerC

A detailed written report with charts and explanations captures the requested trend analysis asynchronously, letting stakeholders review supporting evidence at their own pace. It documents methodology and findings thoroughly, which a verbal follow-up cannot, satisfying the request for additional detail.

Why this answer

The stakeholders requested additional details on a specific trend, so the follow-up should be focused, clear, and self-contained. A detailed written report with charts and explanations directly addresses their request by providing context, visual evidence, and interpretation they can review asynchronously. This approach also creates a reusable artifact that documents the analysis for future reference, which is a best practice in data analytics communication.

Exam trap

DA0-002 often tests the distinction between data presentation and data communication, trapping candidates who equate 'more data' with 'better answers' or who overlook the need for tailored, explanatory follow-up materials.

How to eliminate wrong answers

Option A is wrong because creating a new dashboard with all the data is overkill and does not focus on the specific trend; it may overwhelm stakeholders with unrelated information and lacks the narrative explanation needed for a follow-up. Option B is wrong because sending the raw data file shifts the burden of analysis to the stakeholders, who likely lack the technical skills or context to interpret it correctly, and it provides no explanations or visualizations. Option D is wrong because holding another meeting to verbally explain the trend is inefficient, does not create a lasting record, and may not suit stakeholders' schedules; it also risks miscommunication without written documentation.

946
MCQhard

An analyst creates a histogram of customer transaction amounts but observes that the distribution looks bimodal. Upon review, the analyst realizes that two different customer segments (retail and wholesale) were combined. Which action best addresses this?

A.Create two separate histograms, one for each segment
B.Use a single histogram with two colors for segments
C.Use a box plot instead of a histogram
D.Increase the number of bins to see more detail
AnswerA

The bimodality is an artefact of mixing two populations with different spending patterns. Splitting retail and wholesale into separate histograms removes the confounding grouping variable, revealing each segment's true unimodal distribution instead of one misleading combined shape.

Why this answer

Creating two separate histograms, one for each customer segment (retail and wholesale), best addresses the bimodal distribution because it allows the analyst to see the underlying distributions of each segment clearly. Combining them into one histogram obscures the distinct patterns, while separating them reveals the true characteristics of each group. This is a fundamental principle of data visualization: when data contains subgroups, disaggregating can provide more meaningful insights.

Exam trap

The trap is thinking that a single visualization with colors or more bins can solve the issue, but the root cause is mixing two populations; the best practice is to disaggregate.

How to eliminate wrong answers

Option B is wrong because using a single histogram with two colors still overlays the distributions, which can make it difficult to interpret the shape of each segment, especially if they overlap. Option C is wrong because a box plot summarizes distribution but does not show modality or detailed shape like a histogram. Option D is wrong because increasing the number of bins might show more detail but will not resolve the bimodality caused by combining two distinct segments; it may even make the bimodal nature more apparent but still not separate the segments.

947
MCQmedium

A data analyst is building a self-service reporting environment. Which of the following is the primary benefit of this approach?

A.It ensures all reports use the same data source.
B.It reduces the number of ad hoc report requests to the analytics team.
C.It automatically generates executive summaries.
D.It improves data security by limiting access.
AnswerB

Self-service reporting shifts routine question answering to business users working from governed datasets, so the analytics team fields fewer one-off requests. This directly delivers the reduced ad hoc workload the stem identifies as the primary benefit.

Why this answer

Self-service reporting empowers business users to build and run their own reports against governed data, which directly reduces the volume of ad hoc report requests that would otherwise be funneled to the analytics team. The core value proposition is democratizing data access while freeing skilled analysts for higher-value work. This is the primary benefit cited in BI maturity models and analytics enablement frameworks.

Exam trap

DA0-002 often tests the distinction between the primary benefit of self-service BI (reducing analyst workload/ad hoc requests) and secondary governance outcomes like consistent data sources, which candidates mistakenly select as the 'main' benefit.

How to eliminate wrong answers

Option A is wrong because a single shared data source is a data governance/consistency goal, not the defining benefit of self-service — self-service can actually increase source sprawl if ungoverned. Option C is wrong because automated executive summaries are a feature of specific BI tools (e.g., narrative generation), not an inherent property of self-service reporting. Option D is wrong because self-service typically broadens access rather than restricting it, and security is a governance concern rather than the primary benefit of the approach.

948
Multi-Selecthard

A data analyst is preparing a dataset for a machine learning model to predict customer churn. The dataset contains missing values in several columns. The analyst wants to handle missing data in a way that preserves as much information as possible and avoids introducing bias. Which TWO of the following techniques are appropriate? (Choose two.)

Select 2 answers
A.Listwise deletion
B.K-nearest neighbors (KNN) imputation
C.Mean imputation for numerical features
D.Replacing missing values with zero
E.Multiple imputation by chained equations (MICE)
AnswersB, E

KNN imputation estimates missing values by finding the most similar complete observations and averaging their values. It leverages correlations between features and can preserve local data structure. This method is non-parametric and can handle both numerical and categorical data with appropriate distance metrics. It is suitable for churn prediction as it maintains information and reduces bias compared to simple imputation.

Why this answer

Multiple imputation by chained equations and KNN imputation are both advanced techniques that preserve data structure and reduce bias. MICE models each variable with missing data using other variables and creates multiple imputations to account for uncertainty. KNN imputation uses similarity between observations to estimate missing values.

Both are superior to mean imputation, listwise deletion, or zero replacement, which either discard information or introduce bias.

Exam trap

The trap here is assuming that any imputation method is fine, but simple methods like mean imputation or zero replacement can bias the model, while listwise deletion reduces sample size and may introduce bias.

949
MCQhard

A data scientist is building a model to predict customer churn (yes/no). After training a logistic regression model, the coefficient for 'monthly charges' is 0.05 with a p-value of 0.03. Which interpretation is correct at α=0.05?

A.The model's R-squared is 0.05.
B.For every unit increase in monthly charges, the odds of churn increase by about 5%.
C.Monthly charges decrease the probability of churn.
D.Monthly charges have no significant effect on churn.
AnswerB

A one-unit rise in monthly charges multiplies the odds of churn by e^0.05 ≈ 1.051, an increase of roughly 5%. The p-value of 0.03 falls below α=0.05, so the coefficient is statistically significant and the predictor's effect on churn odds is supported.

Why this answer

In logistic regression, coefficients are in log-odds units. A coefficient of 0.05 means that for a one-unit increase in monthly charges, the log-odds of churn increase by 0.05. Exponentiating gives e^0.05 ≈ 1.051, so the odds increase by about 5.1%.

The p-value of 0.03 is less than α=0.05, so the effect is statistically significant.

Exam trap

The trap is misinterpreting the coefficient as a direct probability change or confusing it with R-squared; candidates may also ignore the p-value and incorrectly claim no significance.

How to eliminate wrong answers

Option A is wrong because R-squared is not derived from a coefficient value; logistic regression typically uses pseudo-R-squared measures, and 0.05 here is the coefficient, not R-squared. Option C is wrong because the coefficient is positive, meaning higher monthly charges are associated with higher odds of churn, not lower. Option D is wrong because the p-value of 0.03 is below 0.05, so the effect is statistically significant, not insignificant.

950
MCQhard

A data governance council at an insurance company is reviewing an incident in which an analyst exported a customer dataset to a personal cloud drive. The council wants a control that detects and blocks sensitive data leaving the managed environment regardless of which application initiates the transfer. Which control type should the council implement?

A.Role-based access control applied to the data warehouse
B.Column-level encryption of the customer dataset at rest
C.A data retention schedule that deletes customer records after seven years
D.Data loss prevention policies enforced on endpoints and network egress
AnswerD

Data loss prevention inspects content and context on endpoints, email, and network egress to detect and block sensitive data leaving the managed environment, regardless of the initiating application. This matches the council's requirement to catch exports to personal cloud drives. DLP can identify regulated data patterns and stop the transfer or alert security teams in real time.

Why this answer

Detecting and blocking sensitive data leaving the managed environment requires content-aware inspection at the points where data exits, which is the role of data loss prevention on endpoints, email, and network egress. Access control, encryption at rest, and retention schedules address different risks and cannot stop a user from exporting data during an active session.

Exam trap

The trap here is equating access control or encryption with exfiltration prevention, when neither monitors or blocks outbound transfers initiated by an authorized user.

951
Drag & Dropmedium

Drag and drop the steps to normalize a database table from 1NF to 3NF in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Normalization proceeds from 1NF to 2NF to 3NF, then table creation and foreign keys.

952
MCQeasy

A marketing analytics team maintains a customer contact table that includes email addresses and mobile numbers. The governance policy states that consent must be recorded and honored for each channel before outreach. A campaign manager wants to email everyone in the table, but the analyst notices that many records lack a consent timestamp. Which governance principle should guide the analyst's decision?

A.Data retention, because contact records should be deleted after a fixed period
B.Data accuracy, because email addresses must be verified before sending
C.Data minimization, because only the fields needed for the campaign should be queried
D.Purpose limitation and consent, because outreach must match the permission each contact granted
AnswerD

Purpose limitation and consent require that data be used only for the purposes the individual authorized. Records without a consent timestamp have no documented permission for email outreach, so they must be excluded. This principle directly governs the campaign manager's request and protects the organization from contacting people who never agreed to receive marketing messages.

Why this answer

Purpose limitation and consent require that outreach match the permission each individual granted. Records lacking a consent timestamp have no documented authorization for email, so they must be excluded regardless of other governance principles. This protects the organization and respects the contact's choices while still allowing the campaign to proceed for consented records.

Exam trap

The trap here is reaching for data minimization or accuracy when the actual blocker is the absence of documented consent for a specific outreach channel.

953
Multi-Selecthard

A data analyst is performing K-means clustering on customer data. Which THREE of the following are steps in the K-means algorithm?

Select 3 answers
A.Perform eigenvalue decomposition.
B.Calculate the correlation matrix.
C.Initialize k centroids randomly.
D.Update centroids by computing the mean of all points assigned to each centroid.
E.Assign each data point to the nearest centroid.
AnswersC, D, E

K-means begins by choosing k initial centroids, typically at random positions in the feature space, before any assignment occurs. This initialisation step satisfies the stem's requirement to identify actual algorithm steps, establishing the starting points that subsequent assignment and update iterations refine.

Why this answer

Option C is correct because the standard K-means algorithm begins by initializing k centroids, typically at random positions in the feature space, before any assignment occurs. Option E is correct because after initialization, each data point is assigned to the nearest centroid, usually measured by Euclidean distance, forming the initial clusters. Option D is correct because the algorithm then recomputes each centroid as the mean (average) of all points assigned to it, and this assignment-and-update cycle repeats until convergence.

Options A and B are not steps in K-means: eigenvalue decomposition is used in dimensionality-reduction techniques such as PCA, and calculating a correlation matrix is a preprocessing or exploratory analysis step, not part of the iterative K-means procedure.

Exam trap

The trap here is confusing K-means with other algorithms like PCA or correlation-based clustering, leading candidates to select eigenvalue decomposition or correlation matrix calculation as steps.

954
MCQhard

A data analyst is tasked with presenting the results of a customer churn analysis to a group of marketing managers. The model identified that customers with low engagement scores are 3 times more likely to churn. The analyst wants to effectively communicate this finding to persuade the managers to focus on engagement. Which of the following is the most compelling way to present this insight?

A.Show a table with raw churn rates by engagement score
B.Use a bar chart comparing churn rates of low vs. high engagement
C.Show a scatter plot of engagement score vs. churn
D.Calculate the potential revenue increase if engagement improves by 10%
AnswerD

Quantifying the revenue uplift from a 10% engagement improvement translates the statistical finding into financial impact, which is the metric marketing managers are accountable for. This satisfies the stem's persuasion constraint by framing churn risk as potential gain, making the engagement investment decision concrete rather than abstract.

Why this answer

To persuade marketing managers, the most compelling presentation is to quantify the business impact—calculating the potential revenue increase if engagement improves by 10% translates the churn insight into financial terms that resonate with managers. This directly ties the analysis to a business outcome, making it more persuasive than raw statistics or charts.

Exam trap

The trap is choosing a visually appealing chart or detailed table over a business-impact metric, confusing technical presentation with persuasive communication.

How to eliminate wrong answers

Option A is wrong because a table of raw churn rates by engagement score is too detailed and does not directly convey the business impact or urgency. Option B is wrong because a bar chart comparing churn rates is informative but still focuses on the metric rather than the financial consequence. Option C is wrong because a scatter plot of engagement vs. churn shows correlation but does not persuade managers to act without a clear business implication.

955
MCQeasy

A dataset contains a column 'Age' with values: [22, 25, 25, 30, 35, 40, 45]. What is the interquartile range (IQR)?

A.15
B.10
C.20
D.25
AnswerA

Ordered values give Q1 at 25 and Q3 at 40, so IQR equals 40 minus 25, which is 15. This satisfies the stem by measuring the spread of the middle half of the dataset, excluding the lowest and highest values.

Why this answer

The interquartile range (IQR) is Q3 minus Q1. For the dataset [22, 25, 25, 30, 35, 40, 45], the median is 30. Q1 is the median of the lower half [22, 25, 25] = 25.

Q3 is the median of the upper half [35, 40, 45] = 40. IQR = 40 - 25 = 15.

Exam trap

DA0-002 often tests quartile calculation with odd-numbered datasets — the trap is including the median in both halves, which shifts Q1 and Q3 and produces an incorrect IQR.

How to eliminate wrong answers

Option B (10) is wrong because it would result from miscalculating Q1 or Q3 — for example, using 30 as Q1 and 40 as Q3, or misidentifying quartile positions. Option C (20) is wrong because it likely comes from subtracting the minimum from the median or miscomputing Q3 as 45 and Q1 as 25. Option D (25) is wrong because it equals Q1 itself or results from subtracting the minimum (22) from Q3 (40) rounded, which is not the IQR definition.

956
MCQmedium

A retail analytics team publishes a dashboard showing average order value by store. A regional manager notices that the store with the highest average order value is also the store with the fewest orders, and questions whether that store is genuinely the best performer. Which dashboard design change would BEST let the manager judge that store fairly?

A.Switch the metric from average order value to total revenue by store and rebuild the ranking on that basis.
B.Remove the lowest-volume stores from the dashboard so the ranking reflects only established locations.
C.Add order-count context to the measure, for example by displaying the number of orders alongside the average and sizing or annotating the mark accordingly.
D.Apply a logarithmic scale to the axis showing average order value so differences between stores appear smaller.
AnswerC

Average order value alone hides sample size, so a store with a handful of unusually large orders can top the ranking by chance. Showing the order count alongside the average, or encoding it visually, lets the manager discount small-sample results and compare stores on equal footing. It addresses the exact concern raised without discarding the requested metric.

Why this answer

The manager's objection is fundamentally about sample size: an average built on few orders is unstable and easy to top by chance. Pairing the average with the order count, or encoding volume in the mark itself, gives the reader the information needed to weigh the result appropriately while preserving the requested metric. Substituting a different metric, transforming the scale, or filtering out small stores all fail to expose reliability.

Exam trap

The trap here is responding to a data-reliability concern by changing the metric or filtering data, when the fix is to add the missing context that exposes sample size.

957
MCQmedium

A data analyst creates a report showing sales by product category. The analyst notices that one category has a very high sales figure due to a one-time bulk order. Which of the following is the best way to communicate this insight to stakeholders?

A.Delete the bulk order from the dataset.
B.Add a note to the chart explaining the bulk order.
C.Remove the category with the bulk order from the report.
D.Use a pie chart to show the proportion of each category.
AnswerB

Annotating the chart identifies the one-time bulk order as the cause of the spike, preventing stakeholders from misreading it as sustained demand. This preserves the accurate sales figure while supplying the context needed for sound decisions.

Why this answer

It maintains data integrity while providing necessary context. Adding a note to the chart allows stakeholders to understand the anomaly without distorting the underlying data. This approach aligns with best practices in data communication, where transparency about outliers is critical for accurate interpretation.

Exam trap

The trap here is that candidates may think removing or hiding data (options A and C) simplifies the message, but the exam tests the principle that data communication must be transparent and preserve the full context for decision-making.

How to eliminate wrong answers

Option A is wrong because deleting the bulk order from the dataset introduces bias and misrepresents actual sales performance, violating data integrity principles. Option C is wrong because removing the entire category hides the outlier entirely, preventing stakeholders from understanding its impact on overall trends. Option D is wrong because a pie chart would visually exaggerate the proportion of the bulk order category, potentially misleading stakeholders about normal sales distribution.

958
MCQmedium

A dashboard designer is creating a sales performance dashboard and wants to minimize non-data ink to improve clarity. Which action best follows the principle of maximizing the data-ink ratio?

A.Removing gridlines that do not add value to the chart
B.Including a detailed company logo on every chart
C.Adding a background image to make the dashboard more visually appealing
D.Using 3-D effects on bars to make them stand out
AnswerA

Gridlines are non-data ink: they consume pixels without encoding values. Removing those that add no interpretive value raises the data-ink ratio, letting the plotted data dominate the chart and improving clarity, exactly as the principle requires.

Why this answer

It directly follows Tufte's data-ink ratio principle by removing non-data ink (gridlines that don't aid interpretation) while preserving the core data. This maximizes the proportion of ink devoted to actual sales metrics, improving clarity without sacrificing information.

Exam trap

The trap here is that candidates may confuse 'visually appealing' (options B, C, D) with effective data communication, not realizing that decorative elements reduce the data-ink ratio and can obscure insights in a professional dashboard.

How to eliminate wrong answers

Option B is wrong because a detailed company logo on every chart adds non-data ink that distracts from the sales performance data, violating the data-ink ratio principle. Option C is wrong because a background image introduces decorative non-data ink that reduces the clarity of the data visualization, contrary to the principle. Option D is wrong because 3-D effects on bars add chartjunk (non-data ink) that can distort perception of bar heights and make comparisons harder, directly opposing the goal of maximizing data-ink.

959
MCQmedium

You are a data analyst at a logistics company. You have created a dashboard to monitor delivery performance. The dashboard includes a scatter plot showing delivery time (hours) vs. distance (miles) for each delivery, with points colored by delivery region (A, B, C, D, E). Users have reported that the scatter plot is cluttered because there are over 10,000 points, making it hard to see patterns. Additionally, the color legend for the five regions uses similar shades of blue, making it difficult to distinguish which region a point belongs to. You need to improve the scatter plot to reduce overplotting and improve region differentiation. Which approach is most effective?

A.Use a hexagonal binning plot (hexbin) with color representing region
B.Create five separate scatter plots (small multiples) for each region
C.Set point opacity to 30% and use a qualitative color palette with distinct hues for each region
D.Convert to a bubble chart by adding package weight as bubble size
AnswerC

Reducing point opacity to 30% lets overlapping marks accumulate into darker density, revealing clusters across 10,000 points. A qualitative palette assigns each of the five regions a perceptually distinct hue, replacing the near-identical blues that currently prevent region differentiation.

Why this answer

Reducing opacity (alpha blending) mitigates overplotting by making overlapping points more transparent, while switching to a qualitative color palette (e.g., distinct hues like red, green, blue) ensures each of the five regions is easily distinguishable. This directly addresses both user complaints without losing the overall distribution context.

Exam trap

The trap here is that candidates often choose small multiples (Option B) thinking they reduce clutter, but the question specifically asks to improve differentiation and reduce overplotting in a single view, and small multiples fragment the data, making cross-region comparison harder.

How to eliminate wrong answers

Option A is wrong because a hexagonal binning plot aggregates points into hexagons based on density, but it cannot encode a categorical variable like region with color inside a single plot—hexbin typically uses a single color gradient for density, not multiple regions. Option B is wrong because creating five separate scatter plots (small multiples) forces users to mentally compare across separate charts, losing the ability to see relative spatial patterns and interactions between regions in a single view. Option D is wrong because converting to a bubble chart by adding package weight as bubble size does not reduce overplotting (it can worsen it with larger circles) and does not improve region differentiation; it introduces a third continuous variable that is not related to the reported issues.

960
MCQmedium

An organization uses a data warehouse for analytics. The data team wants to load data from source systems into the warehouse. They choose to load raw data first and then perform transformations within the warehouse. Which approach are they using?

A.ELT
B.Data lake
C.Data mart
D.ETL
AnswerA

ELT extracts raw source data, loads it into the warehouse unchanged, then transforms it using the warehouse's own compute. This matches the described sequence exactly, distinguishing it from ETL, where transformation occurs in a staging area before loading.

Why this answer

ELT (Extract, Load, Transform) is correct because the raw data is extracted from source systems and loaded directly into the data warehouse before any transformation occurs. Transformations are then performed inside the warehouse using its own compute engine (e.g., SQL, dbt, or Snowflake virtual warehouses). This is the defining characteristic of the ELT pattern, which leverages modern cloud warehouse scalability.

Exam trap

DA0-002 often tests the order of operations — candidates confuse ETL and ELT by assuming transformation always happens before loading, but the phrase 'load raw data first, then transform' is the explicit ELT signal.

How to eliminate wrong answers

Option B is wrong because a data lake is a storage repository for raw structured, semi-structured, and unstructured data — it describes where data lives, not the load-then-transform sequence described. Option C is wrong because a data mart is a subset of a warehouse focused on a specific business line or department, not a data integration methodology. Option D is wrong because ETL transforms data in a separate staging or transformation engine before loading it into the warehouse, which is the reverse order of the scenario described.

961
MCQmedium

A data analyst receives a dataset with inconsistent date formats (e.g., "01/02/2023", "2023-01-02", "Jan 2, 2023"). Which data quality dimension is most directly affected?

A.Accuracy
B.Consistency
C.Completeness
D.Timeliness
AnswerB

Consistency measures whether values conform to the same format and units across a dataset. Mixed representations of the same date attribute violate that uniformity, whereas accuracy concerns correctness of values and completeness concerns missing data.

Why this answer

Consistency refers to the uniformity of data representation. Inconsistent date formats violate consistency, not accuracy, completeness, or timeliness.

962
MCQmedium

An e-commerce company stores product descriptions that contain accented characters and emoji in customer reviews. The database currently fails to store these characters correctly, replacing them with question marks. Which action should the data engineer take to resolve this issue?

A.Increase the maximum length of the VARCHAR column
B.Convert the column to a BLOB data type to store raw bytes
C.Set the column's character set to UTF-8 and ensure the connection uses UTF-8 encoding
D.Change the column data type from VARCHAR to CHAR
AnswerC

Accented characters and emoji require Unicode support. Changing the column character set to UTF-8 and aligning the client connection encoding ensures that multibyte characters are transmitted and stored without loss. When the connection or column uses a limited encoding such as Latin-1, characters outside that set are replaced with question marks, which exactly matches the symptom described.

Why this answer

The symptom of accented characters and emoji becoming question marks is a classic encoding problem. Unicode character sets such as UTF-8 can represent virtually all characters, while legacy encodings cannot. Aligning both the column character set and the client connection encoding to UTF-8 ensures end-to-end representation and storage of multibyte characters without corruption.

Exam trap

The trap here is focusing on the length or type of the column when the actual problem is that the character encoding cannot represent the characters being stored.

963
MCQhard

An analyst creates a dashboard with a large amount of decorative graphics and background images. According to CompTIA Data+, which design principle is being violated?

A.Visual hierarchy
B.Data-ink ratio
C.Appropriate precision
D.Consistent color coding
AnswerB

Decorative graphics and background images consume ink without conveying data, lowering the data-ink ratio that Tufte defines as data-ink divided by total ink. CompTIA Data+ expects dashboards to maximise this ratio by removing non-data pixels, so heavy ornamentation directly violates the principle the stem describes.

Why this answer

The data-ink ratio, a concept from Edward Tufte, states that a visualization should maximize the proportion of ink used to represent actual data relative to total ink used. By adding excessive decorative graphics and background images, the analyst reduces this ratio, making the dashboard less efficient and potentially obscuring the data. CompTIA Data+ emphasizes minimizing non-data ink to improve clarity and insight extraction.

Exam trap

The trap here is that candidates confuse 'decorative graphics' with 'visual hierarchy' or 'color coding,' but CompTIA Data+ specifically tests the data-ink ratio as a principle for minimizing non-data elements to enhance readability.

How to eliminate wrong answers

Option A is wrong because visual hierarchy refers to the arrangement of elements to guide the viewer's eye by importance, which is not directly violated by decorative graphics; the issue is about ink efficiency, not order. Option C is wrong because appropriate precision concerns the level of detail in data values (e.g., rounding or decimal places), not the use of decorative elements. Option D is wrong because consistent color coding ensures that colors have the same meaning across the dashboard, which is unrelated to the overuse of non-data ink.

964
Multi-Selecthard

A data governance team is establishing policies. Which three activities are part of data governance? (Select THREE.)

Select 3 answers
A.Data quality management
B.Data ownership assignment
C.Data indexing
D.Data steward designation
E.Data normalization
AnswersA, B, D

Ensuring data quality is a core governance function.

Why this answer

Data quality management is a core activity of data governance because it ensures that data meets defined standards for accuracy, completeness, consistency, and timeliness. Governance policies mandate monitoring and remediation processes to maintain data quality across the organization.

Exam trap

CompTIA Data+ often tests the distinction between data governance (policies, roles, quality) and data management (technical implementation like indexing and normalization), leading candidates to confuse operational tasks with governance activities.

965
Multi-Selecteasy

A data analyst is performing data acquisition from multiple source files. Which TWO data profiling tasks should the analyst complete before loading the data into the target system?

Select 2 answers
A.Create a dashboard for stakeholders
B.Verify data types and formats
C.Build a linear regression model
D.Perform cluster analysis
E.Identify missing values and nulls
AnswersB, E

Verifying data types and formats confirms each source column matches the target schema, catching mismatches such as strings in numeric fields or inconsistent date layouts before load. This satisfies the pre-load profiling requirement, preventing type-conversion failures or silent corruption during ingestion into the target system.

Why this answer

Option B (Verify data types and formats) is correct because data profiling must confirm that each source column's actual data type (e.g., integer, string, date) and format (e.g., YYYY-MM-DD, decimal separators) match what the target system expects, preventing load failures or silent corruption. Option E (Identify missing values and nulls) is correct because profiling must detect NULLs, empty strings, and placeholder values so the analyst can decide on imputation, defaulting, or rejection rules before loading. Option A (Create a dashboard for stakeholders) is a downstream reporting activity, not a pre-load profiling task.

Option C (Build a linear regression model) is predictive analytics performed after data is cleansed and loaded, not profiling. Option D (Perform cluster analysis) is an unsupervised modeling technique also done post-load, so it does not belong in pre-load data profiling.

966
MCQmedium

A data analyst is building a dashboard for a regional sales director who needs to monitor daily performance across 12 territories. The director wants to instantly see which territories are below target, with the ability to click a territory to view its individual store details. The analyst decides to use a bullet chart for each territory. Which design choice best supports the director's need for immediate exception identification and drill-down?

A.Use a stacked bar chart with actual and target values stacked on top of each other, and add a tooltip showing the variance.
B.Replace the bullet charts with a single pie chart showing each territory's percentage of total sales.
C.Use a uniform color for all bullet charts and rely on the target marker position to indicate performance.
D.Apply a conditional color rule that highlights bullet charts in red when actual sales fall below the target, and enable drill-through to a store-level detail page.
AnswerD

Conditional coloring makes below-target territories immediately visible without reading each chart, and drill-through provides the requested store-level detail. This combination directly addresses both the need for instant exception identification and the ability to investigate a territory. The bullet chart itself shows actual versus target, and the color adds a preattentive alert.

Why this answer

The director needs to instantly spot territories below target and then drill into store details. Conditional formatting on bullet charts uses color to preattentively flag exceptions, while drill-through provides the required detail. Together, these features meet both requirements without forcing the director to interpret each chart individually.

Exam trap

The trap here is assuming that showing actual and target values on the same chart is sufficient for exception identification, when preattentive color encoding is needed for rapid scanning.

967
MCQeasy

A small business wants to acquire customer feedback through a short questionnaire emailed after purchase. Which data acquisition method does this represent?

A.Transaction log
B.Interview
C.Survey
D.Observation
AnswerC

A short questionnaire emailed after purchase is a survey: structured questions distributed to respondents for self-completion. It differs from observation, which records behaviour directly, and from interviews, which are conducted interactively. This satisfies the stem's requirement for acquiring customer feedback at scale.

Why this answer

A survey is a structured data collection method where respondents answer predefined questions, typically via a form or questionnaire. In this scenario, the business is using a short questionnaire emailed after purchase to gather customer feedback, which directly aligns with the definition of a survey as a data acquisition method.

Exam trap

The trap here is that candidates may confuse a survey with a transaction log because both can be automated and delivered electronically, but a transaction log captures system events, not user-provided feedback.

How to eliminate wrong answers

Option A is wrong because a transaction log records system-level events such as database changes, user logins, or API calls, not subjective customer feedback via a questionnaire. Option B is wrong because an interview involves a direct, synchronous conversation between an interviewer and a respondent, often with open-ended questions, whereas the scenario describes an asynchronous, self-administered questionnaire. Option D is wrong because observation involves watching and recording behavior or events without direct interaction, whereas the scenario explicitly involves asking customers for their opinions through a questionnaire.

968
MCQeasy

A junior data analyst is asked to create a chart that shows total monthly revenue for the past 24 months so the finance team can see whether revenue is generally rising or falling over time. Which chart type is most appropriate?

A.A stacked bar chart with one stack per month and a segment for each sales region
B.A line chart with month on the horizontal axis and total revenue on the vertical axis
C.A pie chart with one slice per month showing each month's share of total revenue
D.A scatter plot with month on the horizontal axis and revenue on the vertical axis
AnswerB

A line chart connects ordered time points, making the direction and shape of the trend easy to perceive. With 24 monthly values, the line reveals whether revenue is rising, falling, or fluctuating. This directly answers the finance team's question about the overall trend across two years.

Why this answer

A line chart is the standard choice for showing change over a continuous, ordered time dimension because the connected line makes direction and rate of change immediately visible. The finance team's question is about trend, not composition or correlation, so a pie chart, stacked bar chart, or scatter plot would obscure the answer. Month on the horizontal axis and revenue on the vertical axis aligns with conventional time-series design.

Exam trap

The trap here is treating any chart that includes all the months as valid, even when the chart type cannot express direction of change.

969
MCQhard

A data analyst is using a public API to collect historical weather data. The API has a rate limit of 100 requests per minute, but the analyst needs to retrieve 10,000 records as quickly as possible. What strategy should be used?

A.Increase the request rate
B.Use multiple API keys
C.Paginate with appropriate delays
D.Download a precompiled dataset
AnswerC

Pagination splits the 10,000 records into smaller result sets, while delays keep request frequency at or below the API's 100-per-minute rate limit. This satisfies the throughput constraint without triggering throttling errors, retrieving everything as quickly as the limit permits.

Why this answer

Paginating with appropriate delays respects the API's rate limit of 100 requests per minute while maximizing throughput. By splitting the 10,000 records into pages (e.g., 100 records per page) and sending requests at a rate just under the limit (e.g., one request every 0.6 seconds), the analyst can retrieve all data in approximately 100 minutes without triggering HTTP 429 rate-limit errors.

Exam trap

The trap here is that candidates may assume 'as quickly as possible' means sending requests as fast as possible (Option A) or using multiple keys (Option B), overlooking that rate limits are enforced per key or IP and that proper pagination with delays is the only compliant way to maximize throughput.

How to eliminate wrong answers

Option A is wrong because increasing the request rate beyond 100 requests per minute would violate the API's rate limit, resulting in HTTP 429 (Too Many Requests) responses or temporary IP bans. Option B is wrong because using multiple API keys to circumvent rate limits violates the API's terms of service and could lead to account suspension or revocation of access. Option D is wrong because downloading a precompiled dataset may not be available, may not contain the specific historical weather data needed, or may be outdated, and the question explicitly states the analyst is using a public API to collect data.

970
MCQhard

A data analyst is designing a dashboard that will be displayed on a large monitor in a call center. The dashboard shows real-time metrics such as average handle time, call volume, and customer satisfaction. The analyst wants to ensure that the most critical metric—average handle time—is immediately noticeable when it exceeds a threshold. Which design technique best achieves this?

A.Use a red background color for the entire dashboard when average handle time exceeds the threshold.
B.Display the average handle time as a gauge with a red zone beyond the threshold, and add a blinking animation when the threshold is exceeded.
C.Apply a conditional formatting rule that changes the metric's background to a bright color (e.g., red) and displays an alert icon when the threshold is exceeded.
D.Place the average handle time metric in the top-left corner and use a larger font size than other metrics.
AnswerC

Conditional formatting with a bright background and an alert icon creates a clear, localized visual cue that immediately draws attention to the metric in exception. It does not disrupt the entire dashboard and avoids the pitfalls of blinking or full-screen color changes. This technique is a standard best practice for real-time monitoring dashboards.

Why this answer

For real-time monitoring, the most critical metric should have a dynamic visual alert when it crosses a threshold. Conditional formatting with a bright background and an alert icon localizes the cue to the metric itself, making it immediately noticeable without distracting from other information. This approach balances salience with usability.

Exam trap

The trap here is equating prominence (size, position) with alerting, when the requirement is to signal a threshold breach dynamically.

971
MCQmedium

A business analyst needs to show the monthly sales trend for the past three years. Which chart type is most appropriate?

A.Histogram
B.Pie chart
C.Line chart
D.Scatter plot
AnswerC

A line chart plots continuous time-series data along a chronological axis, making month-to-month movement across three years immediately visible. Its connected points reveal trend direction and seasonality, which categorical charts such as bar or pie cannot express as clearly for this continuous temporal requirement.

Why this answer

A line chart is the standard visualization for showing a continuous trend over time, with time on the x-axis and the measured value on the y-axis. Monthly sales across three years is a time series, so a line chart clearly reveals trends, seasonality, and growth patterns. It is the most appropriate choice for this analytical goal.

Exam trap

DA0-002 often tests chart selection by intent, so the trap is choosing a histogram because it sounds statistical, when the question explicitly asks for a trend over time, which requires a line chart.

How to eliminate wrong answers

Option A is wrong because a histogram shows the frequency distribution of a single continuous variable, not a trend over time. Option B is wrong because a pie chart shows parts of a whole at a single point in time and cannot effectively display 36 monthly data points or trends. Option D is wrong because a scatter plot shows the relationship between two continuous variables, not a time-based trend.

972
MCQeasy

A data analyst is preparing a dataset for analysis and notices that the 'age' column has a significant number of missing values. The analyst decides to impute the missing values using the mean age. Which data preparation technique is being applied?

A.Data discretization
B.Data deduplication
C.Data imputation
D.Data normalization
AnswerC

Data imputation is the process of replacing missing values with substituted values. Using the mean age to fill missing entries is a common imputation method. This technique helps maintain dataset size and can reduce bias if the missingness is random, though it may underestimate variability.

Why this answer

Imputing missing values with the mean is a standard data preparation technique known as data imputation. It allows the analyst to retain records that would otherwise be excluded, enabling more complete analysis. This method is simple but should be used with caution as it can affect statistical properties.

Exam trap

The trap here is confusing imputation with normalization or other data preparation steps, but the key is that missing values are being filled with a calculated value.

973
MCQhard

A data analyst is using a window function to assign a unique rank to each employee within their department based on salary, with ties receiving the same rank and leaving gaps. Which function should be used?

A.DENSE_RANK()
B.RANK()
C.ROW_NUMBER()
D.NTILE()
AnswerB

RANK() assigns identical ranks to tied salary values but skips subsequent numbers, producing gaps such as 1, 2, 2, 4. This matches the stem's requirement for ties sharing a rank while leaving gaps, unlike DENSE_RANK or ROW_NUMBER.

Why this answer

RANK() assigns a rank to each row within a partition based on the ORDER BY clause, giving tied values the same rank and then skipping subsequent ranks (e.g., 1, 2, 2, 4). This matches the requirement of 'ties receiving the same rank and leaving gaps.' DENSE_RANK() would not leave gaps, ROW_NUMBER() would not assign ties, and NTILE() divides rows into buckets rather than ranking them.

Exam trap

The trap here is confusing RANK() with DENSE_RANK(); candidates often forget that RANK() leaves gaps after ties, while DENSE_RANK() does not.

How to eliminate wrong answers

Option A is wrong because DENSE_RANK() assigns consecutive ranks without gaps after ties (e.g., 1, 2, 2, 3), which contradicts the 'leaving gaps' requirement. Option C is wrong because ROW_NUMBER() assigns a unique sequential number to every row regardless of ties, so tied salaries would receive different ranks. Option D is wrong because NTILE() distributes rows into a specified number of roughly equal groups (buckets) and does not produce a ranking with ties and gaps.

974
MCQeasy

A retail company operates 50 stores across the country. The data analyst has been asked to create a dashboard to visualize monthly sales trends over the past two years and compare the performance of the top 5 stores. The dataset includes store name, date, and daily sales amount. Initial exploration reveals that some stores have missing sales data for certain months due to system outages, and there are occasional extreme values caused by promotional events (e.g., Black Friday sales are 10x normal). The analyst needs to choose an appropriate visualization approach that accurately represents the trends and comparisons while handling these data quality issues. What should the analyst do to best meet the requirements?

A.Use a line chart for each store's monthly sales and exclude outliers from the dataset.
B.Use a bar chart showing monthly sales for each month and filter out stores with missing data.
C.Use a line chart for the overall monthly sales trend (aggregated across all stores) and a separate bar chart for the top 5 stores by total sales, with tooltips explaining outlier values.
D.Use a stacked bar chart with all 50 stores, each bar representing a month.
AnswerC

Aggregating daily sales into monthly totals smooths outage gaps, while a line chart shows trend and a bar chart ranks the top 5 stores. Tooltips flag promotional outliers without distorting the plotted values, satisfying both trend and comparison requirements.

Why this answer

The analyst should use a line chart for the overall monthly sales trend to show the aggregated pattern over two years, and a separate bar chart for the top 5 stores to compare their performance. Tooltips can explain outlier values (e.g., Black Friday spikes) without distorting the overall trend. This approach handles missing data by aggregating (which can smooth over gaps) and addresses outliers by providing context rather than removing them, which could bias the analysis.

Exam trap

DA0-002 often tests the trade-off between data cleaning and information loss, tricking candidates into choosing options that exclude outliers or missing data rather than using visualization techniques to handle them.

How to eliminate wrong answers

Option A is wrong because excluding outliers removes valuable information about promotional events and can misrepresent actual sales performance. Option B is wrong because filtering out stores with missing data reduces the dataset and may exclude stores that are otherwise important, and a bar chart for each month across all stores would be cluttered. Option D is wrong because a stacked bar chart with all 50 stores would be visually overwhelming and make it difficult to compare trends or identify top performers.

975
MCQmedium

A data scientist is preparing data for a K-means clustering algorithm. The dataset contains features measured in different units (e.g., income in dollars and age in years). Which preprocessing step is most critical before running K-means?

A.Remove outliers
B.Encode categorical variables
C.Standardize or normalize the features
D.Perform feature selection
AnswerC

K-means relies on Euclidean distance, so features in different units let larger-scale variables such as income dominate the distance calculation. Standardising or normalising puts every feature on a comparable scale, satisfying the requirement for meaningful cluster assignment.

Why this answer

K-means clustering uses Euclidean distance to measure similarity between data points. If features are on different scales (e.g., income in dollars vs. age in years), the feature with the larger range will dominate the distance calculation, leading to biased clusters. Standardizing (z-score normalization) or normalizing (min-max scaling) the features ensures that all features contribute equally to the distance metric, which is critical for K-means to produce meaningful clusters.

Exam trap

The trap is thinking that removing outliers or encoding categorical variables is the most critical step, but the question specifically highlights features in different units, which directly points to scaling. Candidates might also confuse feature selection with preprocessing necessity.

How to eliminate wrong answers

Option A is wrong because while removing outliers can be beneficial, it is not as critical as scaling; outliers can affect clustering but the primary issue is feature scale. Option B is wrong because encoding categorical variables is necessary if the dataset contains categorical features, but the question specifies features measured in different units (numerical), so encoding is not the most critical step. Option D is wrong because feature selection can improve performance but does not address the scale disparity that directly impacts K-means distance calculations.

Page 12

Page 13 of 14

Page 14