Courseiva

CompTIA Data+ (DA0-002) (DA0-002) — Questions 1–75

1004 questions total · 14pages · All types, answers revealed

Page 1 of 14

Page 2
1
MCQhard

A data analyst has built a dashboard for a retail chain that shows sales by product category. During user testing, store managers report that they cannot tell which categories are the top performers because the bars are arranged in the order the categories appear in the database, and the y-axis starts at zero with a very high maximum that compresses all bars. Which change should the analyst make to improve the chart's communicative value?

A.Sort the bars alphabetically by category name and add data labels to each bar
B.Replace the bar chart with a pie chart sorted by slice size and add percentage labels
C.Keep the database order but change the y-axis to start at the minimum bar value to emphasize differences
D.Sort the bars in descending order by sales and keep the y-axis starting at zero
AnswerD

Sorting bars by descending sales places the top performers at the top, so ranking becomes immediately visible without scanning. Keeping the y-axis at zero preserves the true proportional relationship between bars, avoiding exaggeration. Together these changes address both complaints while maintaining honest visual encoding.

Why this answer

Descending sort by sales surfaces the ranking directly, and maintaining a zero baseline keeps bar lengths proportional and honest. The original problems were ordering by an irrelevant database sequence and a scale that compressed differences; sorting fixes the first, and a zero-based axis with an appropriate maximum fixes the second without distortion. Alphabetical sorting, axis truncation, or switching to a pie chart each fails to resolve the ranking or scale issue.

Exam trap

The trap here is believing that truncating the y-axis is a quick fix for compressed bars when it actually distorts the data.

2
MCQeasy

A business analyst creates a one-time report to analyze customer churn for a special project. Which report type is being used?

A.Ad hoc report
B.Scheduled report
C.Self-service report
D.Analytical report
AnswerA

Ad hoc reporting is purpose-built for one-off analysis outside routine reporting cycles, matching this special churn project. It lets the analyst define fields and filters freely without committing to a scheduled or standard report structure.

Why this answer

An ad hoc report is a one-time, on-demand report created to answer a specific question or support a particular project, exactly as described. It is not scheduled, not part of a recurring cycle, and is typically built quickly by a user or analyst. The phrase 'one-time report for a special project' is the textbook definition of ad hoc reporting.

Exam trap

The trap is confusing 'ad hoc' with 'self-service' or 'analytical': candidates may pick self-service because a business analyst created it, but the defining clue is 'one-time' — that is ad hoc.

How to eliminate wrong answers

Option B is wrong because a scheduled report runs automatically at predefined intervals (daily, weekly, monthly) and is recurring, not one-time. Option C is wrong because self-service reporting refers to business users creating their own reports via tools, which describes the method of creation, not the one-time nature of the report. Option D is wrong because 'analytical report' is a broad category describing the report's purpose (analysis), not its frequency; an ad hoc report can be analytical, but the question asks for the type defined by its one-time nature.

3
MCQeasy

A healthcare provider needs to integrate patient data from multiple clinics into a single data warehouse. Which process is used to extract, transform, and load the data?

A.ELT
B.ETL
C.OLAP
D.OLTP
AnswerB

ETL extracts data from the separate clinic sources, transforms it into a consistent format, and loads it into the central warehouse. This three-stage pipeline directly satisfies the integration requirement, unlike ELT, which loads raw data before transforming.

Why this answer

ETL (Extract, Transform, Load) is the correct process because the healthcare provider must first extract data from multiple source clinics, then transform it (e.g., standardize formats, clean duplicates, apply business rules) before loading it into the target data warehouse. This ensures data quality and consistency, which is critical for clinical analytics and reporting.

Exam trap

The trap here is confusing ETL with ELT, where candidates assume ELT is always better due to modern big data tools, but the question explicitly describes a traditional data warehouse integration requiring pre-load transformations.

How to eliminate wrong answers

Option A is wrong because ELT (Extract, Load, Transform) loads raw data into the target system first and transforms it later, which is less suitable for a data warehouse requiring pre-integrated, clean data from multiple sources; it is more common in big data environments like Hadoop. Option C is wrong because OLAP (Online Analytical Processing) is a category of database systems optimized for complex queries and multidimensional analysis, not a data integration process. Option D is wrong because OLTP (Online Transaction Processing) is designed for high-volume transactional operations (e.g., recording patient visits), not for extracting, transforming, and loading data into a warehouse.

4
MCQhard

An analyst runs an A/B test with 1000 users per group and observes a conversion rate of 5% in the control and 6% in the treatment. The p-value is 0.12. What should the analyst conclude?

A.The difference is not statistically significant at the 0.05 level.
B.The sample size is too small to detect an effect.
C.The treatment significantly outperforms control.
D.There is a 12% chance the treatment is better.
AnswerA

With a p-value of 0.12 exceeding the 0.05 significance threshold, the observed 1% conversion lift is within the range expected from random variation. The analyst cannot reject the null hypothesis, so the difference lacks statistical significance.

Why this answer

A p-value of 0.12 is greater than the conventional significance level of 0.05, so the null hypothesis of no difference cannot be rejected. The observed difference between 5% and 6% conversion rates is not statistically significant at the 0.05 level. The analyst should conclude that there is insufficient evidence to claim the treatment outperforms control.

Exam trap

DA0-002 often tests the misinterpretation of p-values, especially the false belief that a non-significant result proves no effect or that p-value equals the probability the treatment is better.

How to eliminate wrong answers

Option B is wrong because the sample size may or may not be too small; the p-value alone does not determine that, and with 1000 per group the test has reasonable power to detect large effects. Option C is wrong because a p-value of 0.12 does not indicate statistical significance; claiming the treatment significantly outperforms control is incorrect. Option D is wrong because a p-value is not the probability that the treatment is better; it is the probability of observing such data if the null hypothesis were true.

5
MCQhard

A financial services company is implementing a data governance program. The chief data officer wants to ensure that data is classified according to its sensitivity and that access is restricted accordingly. The company must comply with GDPR and internal policies. Which data governance component is primarily responsible for defining and enforcing data classification and access controls?

A.Data security and access management
B.Data quality management
C.Data stewardship
D.Master data management
AnswerA

Data security and access management encompasses the policies, processes, and technologies that classify data based on sensitivity and enforce access controls. It directly addresses the requirement to restrict access according to classification and comply with regulations like GDPR. This component includes authentication, authorization, and encryption.

Why this answer

Data security and access management is the governance component that defines data classification and enforces access controls. It ensures that sensitive data is protected according to regulations like GDPR. Data stewardship, data quality management, and master data management address different aspects of governance and do not primarily handle classification and access enforcement.

Exam trap

The trap here is confusing data stewardship, which assigns accountability, with data security, which enforces access controls.

6
MCQeasy

In A/B testing, the null hypothesis typically states that:

A.There is no difference between the control and treatment groups
B.The treatment group will perform better than the control group
C.The sample size is sufficient for the test
D.There is a significant difference between the control and treatment groups
AnswerA

The null hypothesis in A/B testing posits zero effect: any observed difference between control and treatment arises from random variation alone. This satisfies the stem's requirement by defining the baseline against which statistical significance is measured, so rejecting it indicates the treatment genuinely influenced the metric rather than chance.

Why this answer

The null hypothesis (H0) is a statement of no effect or no difference between groups.

7
MCQeasy

A data analyst wants to show the relative proportions of defects by type in a manufacturing process. There are 6 defect types. Which chart is most appropriate?

A.Line chart
B.Pie chart
C.Stacked bar chart
D.Scatter plot
AnswerB

A pie chart encodes each defect type as a proportional slice of a single whole, making relative share immediately visible. With only six categories, the slices remain distinguishable, satisfying the requirement to show proportions by type.

Why this answer

A pie chart is standard for showing parts of a whole, especially with few categories. Bar chart is also possible but not the best for proportions. Scatter and line charts are not for proportions.

8
MCQmedium

A retail company wants to analyze customer purchase patterns over time. The data is stored in a relational database with tables for Customers, Orders, and Products. Which database concept should be used to ensure that each order references a valid customer?

A.View
B.Index
C.Primary key
D.Foreign key
AnswerD

A foreign key constrains the Orders table's customer reference to values existing in the Customers primary key, enforcing referential integrity so no order can point at an invalid customer. This directly prevents orphaned order records.

Why this answer

A foreign key constraint enforces referential integrity by ensuring that every value in the 'customer_id' column of the Orders table matches a valid primary key value in the Customers table. This prevents orphaned records and guarantees that each order references an existing customer.

Exam trap

CompTIA often tests the distinction between a primary key (which enforces uniqueness within a table) and a foreign key (which enforces relationships between tables), leading candidates to mistakenly choose primary key when the question asks about cross-table validation.

How to eliminate wrong answers

Option A is wrong because a view is a virtual table based on a query and does not enforce any constraints between tables. Option B is wrong because an index speeds up data retrieval but does not enforce referential integrity or validate relationships. Option C is wrong because a primary key uniquely identifies rows within its own table and cannot enforce relationships between different tables.

9
MCQeasy

An analyst wants to show the trend of monthly sales over the past two years. Which chart type is most appropriate?

A.Pie chart
B.Bar chart
C.Scatter plot
D.Line chart
AnswerD

A line chart plots values against a continuous time axis, so consecutive monthly points connect into a readable trend line. This makes direction, seasonality and turning points across the two-year span immediately visible, which categorical charts cannot convey.

Why this answer

A line chart is best for showing trends over time.

10
Multi-Selectmedium

Which TWO of the following are benefits of database normalization to 3NF? (Select 2)

Select 2 answers
A.Improves query performance for all queries
B.Reduces data redundancy
C.Simplifies complex joins
D.Eliminates all data anomalies
E.Increases data integrity
AnswersB, E

Third normal form eliminates repeating groups and transitive dependencies, so each fact is stored once in its owning table. This satisfies the redundancy benefit by removing duplicate copies that would otherwise need synchronising across the schema.

Why this answer

Normalization to 3NF eliminates transitive dependencies, which directly reduces data redundancy by ensuring each non-key attribute depends only on the primary key. This reduction in redundancy also increases data integrity because updates, inserts, and deletes are less likely to create inconsistencies or anomalies. In a relational database, 3NF achieves this without sacrificing the ability to reconstruct the original data via joins.

Exam trap

The trap here is that candidates confuse normalization with denormalization, assuming that reducing redundancy always improves query performance, when in fact normalization often increases join complexity and can slow down read queries.

11
Multi-Selecteasy

Which TWO of the following are effective techniques for presenting data to a non-technical audience?

Select 2 answers
A.Explain the statistical methods used in the analysis.
B.Include detailed data tables for reference.
C.Highlight the most important insights using callouts.
D.Use many different colors to distinguish data points.
E.Use simple language and avoid jargon.
AnswersC, E

Callouts place short annotations directly beside the relevant data points, directing a non-technical audience's attention to the key insight without requiring them to interpret axes or legends. This satisfies the constraint of presenting effectively to viewers lacking analytical background, reducing cognitive effort.

Why this answer

Option C is correct because callouts draw a non-technical audience's attention directly to the key insights, letting them grasp the main message without needing to interpret raw data or charts themselves. Option E is correct because using simple language and avoiding jargon makes the presentation accessible to listeners who lack a technical or statistical background, ensuring the findings are understood rather than obscured. Options A and B do not belong because explaining statistical methods and providing detailed data tables cater to a technical or analytical audience and tend to overwhelm non-technical listeners with unnecessary detail.

Option D is also incorrect because using many different colors can create visual clutter and confusion rather than clarifying the data for a general audience.

Exam trap

CompTIA often tests the misconception that non-technical audiences need more data (tables, statistics) to understand insights, when in fact they need less—focusing on simplicity, visual emphasis, and clear language—so candidates mistakenly choose options A, B, or D thinking they are thorough.

12
MCQeasy

A data analyst receives a sales dataset where the 'order_date' column is stored as text in the format '2024-03-15'. The analyst needs to calculate the number of days between each order date and the current date. Which action should the analyst take first?

A.Encode the 'order_date' column using one-hot encoding
B.Convert the 'order_date' column from text to a date data type
C.Apply a logarithm transformation to the 'order_date' column values
D.Impute missing values in the 'order_date' column with the column mean
AnswerB

Date arithmetic such as subtracting one date from another requires a true date or datetime data type, because text values are compared lexically rather than chronologically. Converting the column first ensures that functions like DATEDIFF or date subtraction produce correct day counts, and it also enables proper sorting and filtering by date.

Why this answer

Calculating elapsed days requires the order date to be a genuine date type so the platform can perform chronological subtraction. While the values look like dates, storing them as text means comparisons and arithmetic operate on character strings, producing errors or incorrect results. Converting the column to a date type is the prerequisite step that unlocks accurate day-difference calculations and correct date-based sorting.

Exam trap

The trap here is assuming that text formatted like a date behaves like a date, when arithmetic on text values fails or returns incorrect results.

13
MCQeasy

An analyst computed the mean, median, and mode of a dataset and found they are all equal. Which of the following best describes the distribution?

A.Bimodal
B.Negatively skewed
C.Positively skewed
D.Symmetric
AnswerD

Equal mean, median and mode indicates a symmetric distribution, where values balance evenly around the centre and the mode coincides with that central point. Skewed distributions pull the mean away from the median and mode, so this equality satisfies the stem's condition.

Why this answer

When mean, median, and mode are equal, the distribution is symmetric and unimodal, often resembling a normal distribution.

14
MCQhard

A data analyst creates a dashboard for executives that shows a key metric trending downward. The analyst notices that the metric is highly volatile day-to-day. Which visualization type is most appropriate to show the underlying trend?

A.7-day moving average line chart.
B.Pie chart showing proportion of days.
C.Scatter plot with regression line.
D.Daily bar chart.
AnswerA

A 7-day moving average smooths daily volatility by averaging each point with the preceding six days, revealing the underlying directional trend. Executives see the genuine trajectory rather than noise, directly addressing the highly volatile day-to-day readings in the stem.

Why this answer

A 7-day moving average line chart smooths out day-to-day volatility by averaging each point with the preceding six days, revealing the underlying directional trend that executives need to see. Because the metric is highly volatile, plotting raw daily values would obscure the trend in noise; the moving average makes the downward trajectory visually obvious while preserving the time-series nature of the data.

Exam trap

The trap here is confusing 'showing the data' with 'showing the trend' — candidates pick the daily bar chart because it is the most detailed, but the question explicitly asks for the underlying trend, which requires smoothing.

How to eliminate wrong answers

Option B is wrong because a pie chart shows parts of a whole at a single point in time and cannot represent a trend over time at all — it is a categorical composition chart, not a time-series chart. Option C is wrong because a scatter plot with regression line is used to show correlation between two continuous variables, not to display a single metric's trend over time; it also does not smooth volatility. Option D is wrong because a daily bar chart plots every raw data point, so the day-to-day volatility the analyst already noticed would dominate the visual and hide the trend.

15
MCQmedium

An analyst needs to present quarterly sales data to the board. The CEO wants to see both overall trend and breakdown by region. Which dashboard layout is most effective?

A.A single line chart with all regions
B.A KPI card with total sales
C.A combination of a line chart for total and a stacked area chart for regional breakdown
D.A table with all quarterly figures
AnswerC

A line chart shows the overall quarterly trend while a stacked area chart decomposes the same timeline by region, satisfying the CEO's dual requirement. Both share one time axis, letting the board read total and regional contribution together without switching views.

Why this answer

It simultaneously satisfies the CEO's dual requirement: a line chart clearly shows the overall quarterly sales trend, while a stacked area chart breaks down total sales by region, allowing the board to see both the aggregate performance and the contribution of each region over time. This combination leverages the strengths of each chart type—line for trend clarity and stacked area for part-to-whole relationships—without overloading the viewer with data.

Exam trap

The trap here is that candidates often choose a single line chart (Option A) thinking it shows both trend and breakdown, but they overlook that multiple overlapping lines make it hard to see the aggregate trend, which is the CEO's primary need.

How to eliminate wrong answers

Option A is wrong because a single line chart with all regions would create visual clutter and make it difficult to discern the overall trend from the regional lines, especially if regions have overlapping values; it fails to provide a clear aggregate view. Option B is wrong because a KPI card with total sales only shows a single number, which cannot convey the quarterly trend or regional breakdown required by the CEO. Option D is wrong because a table with all quarterly figures forces the board to manually parse numbers to identify trends and regional contributions, which is inefficient for a high-level presentation and violates the principle of data visualization for quick insight.

16
MCQhard

A time series of monthly sales data exhibits a clear upward trend over several years, with consistent peaks each December. Which components are present in this series?

A.Trend and seasonality
B.Cyclical and irregular components only
C.Seasonality and cyclical components only
D.Trend and irregular components only
AnswerA

The sustained upward movement over years is the trend component, while the recurring December peaks represent seasonality with a fixed period of twelve months. Both components coexist in the series, so trend and seasonality are present.

Why this answer

The time series exhibits an upward trend over several years and consistent peaks each December, which is seasonality. Therefore, the components present are trend and seasonality. Cyclical components are not mentioned, and irregular components are always present but not the focus.

Exam trap

DA0-002 often tests the identification of time series components, and candidates may confuse seasonality with cyclical patterns or overlook the trend when seasonality is prominent.

How to eliminate wrong answers

Option B is wrong because cyclical and irregular components only would not account for the clear upward trend and seasonal peaks. Option C is wrong because seasonality and cyclical components only would miss the trend. Option D is wrong because trend and irregular components only would miss the seasonality.

17
MCQhard

A multinational retailer must comply with GDPR for its European customers. The governance team is mapping where personal data originates, how it flows between the CRM, marketing automation platform, and analytics warehouse, and which transformations occur at each hop. Which governance capability does this mapping effort primarily establish?

A.Data lineage
B.Data quality profiling
C.Data stewardship assignment
D.Master data management
AnswerA

Data lineage documents the origin, movement, and transformation of data across systems. By tracing personal data from the CRM through marketing automation into the analytics warehouse, the team builds the end-to-end visibility GDPR requires for subject access requests, deletion propagation, and breach impact analysis. This capability directly answers where data comes from and what happens to it.

Why this answer

Mapping origins, flows, and transformations of personal data across systems is the definition of data lineage. GDPR requires organizations to know where personal data resides and how it moves to fulfill access requests, propagate erasure, and assess breach scope. Lineage provides that traceability, whereas quality profiling, master data management, and stewardship address different governance concerns.

Exam trap

The trap here is equating any governance mapping exercise with master data management or stewardship, when the specific trace of origin, flow, and transformation is lineage.

18
MCQhard

A dataset contains a feature with values ranging from 10 to 1000. The analyst applies min-max normalization to scale the feature between 0 and 1. What is the normalized value of 520?

A.0.515
B.0.510
C.0.480
D.0.520
AnswerA

Min-max normalisation applies (x − min) / (max − min), giving (520 − 10) / (1000 − 10) = 510 / 990 = 0.515. This satisfies the stem's constraint of scaling the 10–1000 range onto 0–1, correctly mapping 520 to 0.515.

Why this answer

Min-max normalization scales values using the formula (x - min) / (max - min). With min=10 and max=1000, the normalized value of 520 is (520 - 10) / (1000 - 10) = 510 / 990 = 0.51515..., which rounds to 0.515. Option A correctly applies this formula.

Exam trap

The trap here is using the maximum value (1000) as the denominator instead of the range (max - min = 990), which leads to option B or D.

How to eliminate wrong answers

Option B (0.510) is wrong because it results from dividing 510 by 1000 instead of by the range (990), a common mistake of using the raw max rather than max minus min. Option C (0.480) is wrong because it appears to use an incorrect range or subtract the wrong minimum, producing a value that does not match the formula. Option D (0.520) is wrong because it simply divides 520 by 1000, ignoring the min offset entirely.

19
MCQmedium

A data analyst needs to count the number of orders placed by each customer, but only for customers who have placed more than 5 orders. Which SQL clause should be used to filter the aggregated results?

A.FILTER
B.HAVING
C.WHERE
D.LIMIT
AnswerB

HAVING filters groups after aggregation, so it can test COUNT(*) > 5 per customer. WHERE cannot reference aggregate functions, making HAVING the clause that satisfies the requirement to restrict results to customers exceeding five orders.

Why this answer

HAVING is used to filter groups after aggregation. The query would use GROUP BY customer_id, then HAVING COUNT(*) > 5.

20
MCQhard

A government agency's data analyst is commissioned to produce a report on public transportation usage trends. The report will be read by policymakers, transit planners, and the general public. The data includes ridership numbers, delay rates, and demographic breakdowns. The analyst needs to ensure the report is accessible and persuasive, especially to non-technical readers. The goal is to advocate for increased funding in underserved areas. The report must be data-driven but also tell a compelling story. What strategy should the analyst prioritize?

A.Provide raw data in appendices only.
B.Create a narrative that highlights the impact of delayed trains on low-income commuters.
C.Use complex statistical analysis to show significance of trends.
D.Focus solely on ridership numbers without context.
AnswerB

Framing delay rates around low-income commuters converts ridership and demographic data into a human impact story, satisfying the stem's demand for a persuasive, non-technical narrative that advocates funding for underserved areas. Policymakers and the public grasp consequences faster than raw statistics, making the funding case compelling.

Why this answer

It directly addresses the need to make data accessible and persuasive to non-technical readers by weaving a narrative around a specific, relatable impact (delayed trains on low-income commuters). This approach aligns with the goal of advocating for increased funding in underserved areas, as it humanizes the data and creates a compelling story that policymakers and the public can understand and act upon, without requiring technical expertise.

Exam trap

The trap here is that candidates often choose Option C (complex statistical analysis) because they equate 'data-driven' with technical rigor, failing to recognize that the exam's focus on 'communicating data insights' prioritizes accessibility and persuasion over statistical complexity for non-technical stakeholders.

How to eliminate wrong answers

Option A is wrong because providing raw data only in appendices fails to make the report accessible or persuasive; it buries the key insights and requires readers to perform their own analysis, which is ineffective for non-technical audiences. Option C is wrong because using complex statistical analysis (e.g., p-values, regression coefficients) would alienate non-technical readers like the general public and many policymakers, making the report inaccessible and undermining its persuasive power. Option D is wrong because focusing solely on ridership numbers without context (e.g., demographic breakdowns, delay rates) provides no narrative or actionable insight, failing to tell a compelling story or advocate for specific funding needs.

21
MCQhard

A company uses a dashboard to monitor server uptime. The data is collected every minute, but the dashboard only refreshes every hour. Users see gaps in the line chart. What is the most likely cause, and how should it be fixed?

A.The line chart should interpolate missing data points
B.Switch to a bar chart to avoid gaps
C.Increase the dashboard refresh rate to match data collection frequency
D.Use a different data series with the same refresh rate
AnswerC

The chart plots minute-level data but redraws hourly, so intermediate points are never rendered, producing visible gaps. Raising the dashboard refresh rate to match the one-minute collection interval ensures each recorded value appears, eliminating the apparent discontinuity.

Why this answer

The dashboard refreshes hourly while data arrives every minute, so between refreshes the chart shows only the last-fetched snapshot and appears to have gaps or stale segments. Aligning the dashboard refresh interval with the data collection frequency (or at least much closer to it) ensures the chart reflects the continuous stream of one-minute data points.

Exam trap

DA0-002 often tests whether candidates blame the chart type or missing data when the real issue is a refresh-cadence mismatch between the data source and the visualization.

How to eliminate wrong answers

Option A is wrong because interpolating missing points fabricates data that was never collected — it hides the real problem (stale refresh) and misleads viewers. Option B is wrong because switching to a bar chart does not address the underlying refresh mismatch; the gaps would still appear as missing bars. Option D is wrong because changing the data series does not fix the refresh cadence — if the new series also updates every minute but the dashboard still refreshes hourly, the same gaps persist.

22
MCQmedium

A retail analyst must build a single dashboard that executives will view on a large conference-room display and regional managers will open on tablets in the field. The executive layout uses a dense 4x3 grid of small KPI tiles. Which design approach best preserves readability across both viewing contexts?

A.Keep one dense layout but increase the font size of every label until it is legible on the tablet, accepting overlap on the display.
B.Publish two entirely separate dashboard files and ask users to remember which URL corresponds to their device type.
C.Design a responsive layout with breakpoints so the tile grid reflows to fewer columns and larger touch targets on smaller screens.
D.Fix the dashboard to a single wide canvas and instruct tablet users to rotate to landscape and scroll horizontally.
AnswerC

Responsive breakpoints let one dashboard serve both contexts: the wide conference display keeps the dense 4x3 grid, while the narrower tablet viewport reflows to fewer columns with larger visual elements and touch-friendly targets. This preserves the executives' information density without forcing field managers to pinch-zoom or misread compressed KPI tiles. It is the standard approach for multi-device reporting.

Why this answer

A responsive layout with defined breakpoints lets a single dashboard artifact adapt its column count, element sizing, and touch targets to the viewport actually in use. The conference display retains dense KPI tiles for executive scanning, while the tablet view reflows into fewer, larger tiles that remain legible and tappable. This keeps one governed source of truth instead of divergent copies.

Exam trap

The trap here is assuming that enlarging fonts globally on a fixed grid solves multi-device readability, when the real fix is letting the layout reflow at breakpoints.

23
Multi-Selectmedium

A data analyst is creating a dashboard for a retail company. The dashboard should provide insights into sales performance across multiple dimensions. Which TWO chart types are best suited for showing the contribution of each product category to total sales?

Select 2 answers
A.Scatter plot
B.Histogram
C.Pie chart
D.Line chart
E.Stacked bar chart
AnswersC, E

A pie chart encodes each product category as a proportional slice of the whole, directly showing its percentage contribution to total sales. This satisfies the stem's requirement to display contribution across categories, since the slices sum to 100% of the sales total.

Why this answer

A pie chart is ideal for showing the contribution of each product category to total sales because it visually represents parts of a whole, making it easy to compare proportions at a glance. The stacked bar chart also effectively shows category contributions within a total, allowing for both absolute and relative comparisons across different time periods or segments. Both chart types directly address the need to visualize proportional breakdowns of a single aggregate metric.

Exam trap

The trap here is that candidates often choose a line chart (Option D) for any sales data because they associate sales with trends, overlooking that the question specifically asks for contribution to total sales, not change over time.

24
MCQeasy

Which chart type is best for comparing the distribution of a continuous variable across different categories?

A.Pie chart
B.Box plot
C.Histogram
D.Treemap
AnswerB

A box plot encodes median, quartiles and outliers on one continuous axis, so several categories can be compared side by side. This directly satisfies the stem's requirement to compare distribution across categories, which a bar or pie chart cannot show.

Why this answer

A box plot is ideal for comparing the distribution of a continuous variable across different categories because it displays the median, quartiles, and outliers for each category side by side. This allows for quick visual comparison of central tendency, spread, and skewness across groups. Histograms show distribution for a single variable, not across categories.

Exam trap

DA0-002 often tests the confusion between histograms and box plots, where candidates might choose histogram for comparing distributions across categories, but histograms are for a single variable's distribution.

How to eliminate wrong answers

Option A is wrong because a pie chart shows proportions of a whole for categorical data, not distribution of a continuous variable. Option C is wrong because a histogram displays the distribution of a single continuous variable, not across multiple categories. Option D is wrong because a treemap visualizes hierarchical data using nested rectangles, not distributions.

25
MCQeasy

A data analyst wants to show the distribution of customer ages for a retail store. The ages are continuous and the analyst needs to visualize the frequency of different age ranges. Which chart type is most appropriate?

A.Bar chart
B.Line chart
C.Pie chart
D.Histogram
AnswerD

A histogram bins continuous age values into ranges and plots frequency per bin, directly showing the distribution's shape. It satisfies the stem's constraint that ages are continuous and the analyst needs frequency across age ranges.

Why this answer

A histogram is used to display the distribution of a continuous variable by grouping data into bins.

26
Multi-Selecteasy

A data analyst discovers an anomaly in a dataset. Which two actions should be taken before reporting? (Choose TWO.)

Select 2 answers
A.Assume the anomaly is real and report it
B.Immediately alert all stakeholders
C.Verify the data source and extraction process
D.Check for data entry errors or technical glitches
E.Remove the anomaly without documentation
AnswersC, D

Verifying the source and extraction process confirms the anomaly reflects real data rather than a pipeline fault, satisfying the need to validate findings before reporting. Checking lineage, transformation logic and load errors distinguishes genuine outliers from artefacts introduced during ETL, preventing false alarms reaching stakeholders.

Why this answer

Option C is correct because verifying the data source and extraction process ensures the anomaly is not an artifact of how the data was pulled, transformed, or joined before it is escalated. Option D is correct because checking for data entry errors or technical glitches (e.g., sensor faults, ETL failures, or duplicate records) confirms whether the anomaly reflects a genuine signal or a data quality issue. Both steps establish validity and reproducibility of the finding, which is essential prior to reporting.

Option A is wrong because assuming the anomaly is real without validation risks reporting a false positive. Option B is wrong because alerting all stakeholders immediately is premature before the anomaly is verified. Option E is wrong because removing the anomaly without documentation destroys evidence and violates data integrity and auditability practices.

Exam trap

The trap here is that candidates may confuse 'immediate reporting' with proactive communication, but CompTIA Data+ expects the understanding that data validation must precede any stakeholder notification to maintain data credibility.

27
MCQhard

A data analyst needs to combine customer data from two tables: Customers (CustomerID, Name) and Orders (OrderID, CustomerID, Amount). Only customers who have placed at least one order should be included. Which JOIN type should be used?

A.INNER JOIN
B.LEFT JOIN
C.FULL OUTER JOIN
D.RIGHT JOIN
AnswerA

An INNER JOIN returns only rows where CustomerID matches in both tables, excluding customers with no orders. This directly satisfies the requirement that only customers who have placed at least one order appear in the result, while unmatched customer rows are discarded.

Why this answer

An INNER JOIN returns only rows where there is a match in both tables. Since the requirement is to include only customers who have placed at least one order, the INNER JOIN on CustomerID will filter out any customer without a matching order record, exactly meeting the condition.

Exam trap

The trap here is that candidates often choose LEFT JOIN thinking it 'includes all customers' without realizing it also includes customers with no orders, which fails the explicit condition of 'only customers who have placed at least one order'.

How to eliminate wrong answers

Option B (LEFT JOIN) is wrong because it would include all customers, even those with no orders, with NULL values for order columns, which violates the 'only customers who have placed at least one order' requirement. Option C (FULL OUTER JOIN) is wrong because it would include customers without orders and orders without customers, both of which are not needed. Option D (RIGHT JOIN) is wrong because it would include all orders, potentially including orders with no matching customer, and still would not restrict customers to only those with orders.

28
MCQhard

A large retail company is integrating customer data from two separate CRM systems into a new data warehouse. System A stores customer IDs as integers (e.g., 12345), while System B stores them as alphanumeric strings (e.g., 'CUST-12345-X'). Additionally, some customers exist in both systems but with slight name variations (e.g., 'John Smith' vs 'Jon Smith'). The data warehouse requires a unified customer table with a single unique identifier for each customer. The analyst needs to design the data acquisition process. Which of the following is the most appropriate first step?

A.Use a simple crosswalk table based on exact name matches to link records
B.Load all data from both systems into a staging table, then run a fuzzy matching algorithm to identify duplicates
C.Perform data profiling to analyze data distributions, data types, and quality issues in each source
D.Standardize all customer IDs to a common format (e.g., UUIDs) and then merge the tables
AnswerC

Data profiling first exposes the exact type mismatch between integer and string customer IDs, plus the scale of near-duplicate names such as 'John Smith' versus 'Jon Smith'. This evidence, covering distributions, formats and quality, is required before choosing a matching or standardisation strategy for the unified customer table.

Why this answer

Data profiling is the foundational first step in any data integration project. It systematically assesses source data types, formats, completeness, and quality issues (e.g., integer vs. alphanumeric IDs, name variations) before designing transformation logic. Without profiling, subsequent steps like fuzzy matching or ID standardization risk being built on incorrect assumptions about the data.

Exam trap

The trap here is that candidates often jump to a technical solution (fuzzy matching or ID standardization) without recognizing that data profiling is the prerequisite step that validates source assumptions and prevents costly rework.

How to eliminate wrong answers

Option A is wrong because exact name matches cannot resolve the known name variations (e.g., 'John Smith' vs 'Jon Smith'), leading to missed linkages and duplicate customers. Option B is wrong because loading all data into a staging table before profiling risks propagating unknown data quality issues (e.g., inconsistent ID formats, nulls) into the staging area, making fuzzy matching less reliable and harder to tune. Option D is wrong because standardizing IDs to a common format (e.g., UUIDs) without first profiling the source data ignores the need to understand existing relationships and quality issues, and may break referential integrity if applied prematurely.

29
MCQhard

A data analyst is creating a report that compares the performance of five different marketing campaigns over the past year. The analyst wants to show both the total revenue generated by each campaign and the proportion of total revenue that each campaign represents. Which combination of visualizations is most effective for this purpose?

A.A treemap showing revenue by campaign, with rectangle sizes proportional to revenue.
B.A line chart showing revenue trends for each campaign over the year and a scatter plot showing revenue versus campaign cost.
C.A bar chart showing total revenue by campaign and a pie chart showing the percentage of total revenue for each campaign.
D.A stacked bar chart showing revenue by campaign, with each bar segmented by month.
AnswerC

A bar chart is excellent for comparing total revenue across campaigns, as it allows precise comparison of lengths. A pie chart, while limited for many categories, effectively shows part-to-whole relationships for a small number of categories. With five campaigns, a pie chart can clearly display each campaign's proportion of total revenue. Together, they meet both requirements.

Why this answer

The most effective combination is a bar chart for total revenue and a pie chart for proportion. A bar chart enables accurate comparison of revenue totals across campaigns, while a pie chart clearly illustrates each campaign's share of the total when there are only five categories. This dual approach directly addresses both analytical goals without requiring viewers to perform mental calculations.

Exam trap

The trap here is selecting a single chart that attempts to show both total and proportion, when often two complementary charts are clearer, especially for a small number of categories.

30
MCQmedium

An analyst is comparing the average sales of two different store locations using a t-test. The p-value obtained is 0.03, and the significance level is 0.05. What should the analyst conclude?

A.Fail to reject the null hypothesis; no significant difference
B.The test is inconclusive because the p-value is too low
C.Reject the null hypothesis; there is a significant difference
D.Accept the null hypothesis; the means are equal
AnswerC

A p-value of 0.03 falls below the 0.05 significance level, so the null hypothesis of no difference is rejected. The result is statistically significant, indicating the two store locations' average sales genuinely differ rather than reflecting random sampling variation.

Why this answer

When the p-value (0.03) is less than the significance level (0.05), the analyst rejects the null hypothesis. This means there is statistically significant evidence of a difference between the average sales of the two store locations.

Exam trap

DA0-002 often tests the p-value vs. α comparison — candidates confuse 'low p-value' with 'inconclusive' or mistakenly say you 'accept' the null hypothesis instead of 'fail to reject' it.

How to eliminate wrong answers

Option A is wrong because failing to reject requires p > α; here 0.03 < 0.05, so the null is rejected. Option B is wrong because a low p-value is not 'inconclusive' — it is the evidence needed to reject the null. Option D is wrong because you never 'accept' the null hypothesis in frequentist testing; you either reject or fail to reject it, and here the null is rejected anyway.

31
Multi-Selecteasy

A data analyst is working with a dataset that includes customer names, email addresses, and purchase history. The analyst wants to ensure that each customer is uniquely identified. Which TWO database concepts should be used to enforce uniqueness and link related data?

Select 2 answers
A.Foreign key
B.Normalization
C.View
D.Primary key
E.Index
AnswersA, D

A foreign key links related data by referencing the primary key of another table, enforcing referential integrity between customer records and their purchase history. Combined with a primary key, which uniquely identifies each customer, it satisfies the stem's requirement to enforce uniqueness and link related data.

Why this answer

A primary key uniquely identifies each row in a table, ensuring no duplicate customer records. A foreign key links related data across tables by referencing the primary key of another table, enforcing referential integrity. Together, they guarantee uniqueness and enable relational joins between customer and purchase history tables.

Exam trap

The trap here is that candidates often confuse normalization with a constraint or think an index enforces uniqueness, when only primary and foreign keys provide the required referential integrity and unique identification.

32
MCQeasy

A data analyst creates a line chart showing monthly sales over the past year. The chart uses a y-axis starting at $100,000 instead of zero. What is the most likely misinterpretation a viewer might have?

A.The differences between months are exaggerated, making small changes look large.
B.The sales appear to be decreasing when they are actually increasing.
C.The chart is correctly scaled, so no misinterpretation occurs.
D.The sales appear to be increasing when they are actually decreasing.
AnswerA

A truncated y-axis compresses the visual distance that represents zero, so equal dollar gaps occupy unequal chart heights. Monthly variations therefore appear steeper than they are, exaggerating differences and making minor fluctuations look like significant shifts.

Why this answer

Starting the y-axis at $100,000 instead of zero truncates the baseline, which visually exaggerates the relative differences between monthly sales values. This is a common data visualization pitfall that can mislead viewers into perceiving small fluctuations as significant trends, violating the principle of using a zero baseline for bar and line charts to accurately represent proportional change.

Exam trap

The trap here is that candidates may think a truncated y-axis only affects bar charts or that it reverses trends, but CompTIA often tests the specific misinterpretation that small changes appear exaggerated due to the loss of a zero baseline, not that the direction of the trend is flipped.

How to eliminate wrong answers

Option B is wrong because a truncated y-axis does not inherently reverse the direction of a trend; it only amplifies the visual magnitude of changes, so sales that are actually increasing would still appear to increase, just more dramatically. Option C is wrong because the chart is not correctly scaled for accurate proportional interpretation; starting the y-axis at a non-zero value is a deliberate distortion that can mislead viewers, and best practices for data visualization recommend a zero baseline for line charts showing magnitude. Option D is wrong because a truncated y-axis does not reverse the direction of a trend; if sales are actually decreasing, they would still appear to decrease, but the visual drop would be exaggerated, not inverted.

33
Multi-Selectmedium

A data analyst is preparing to run an A/B test comparing two email subject lines. Which TWO of the following should the analyst define before the test begins?

Select 2 answers
A.The exact lift in conversion rate
B.The time series decomposition
C.The p-value after the test
D.The null and alternative hypotheses
E.The sample size required for the desired power
AnswersD, E

Stating the null and alternative hypotheses before the A/B test begins fixes exactly what is being tested and which outcome counts as significant. This pre-registration prevents post-hoc interpretation of results and satisfies the stem's requirement to define parameters before the test starts.

Why this answer

Before A/B testing, define null and alternative hypotheses, and determine sample size needed for desired statistical power and effect size.

34
MCQmedium

In a dataset of employee salaries, the analyst notices one value that is significantly higher than the rest. Using the IQR method, which values are typically considered outliers?

A.Values beyond Q1 - 3*IQR or Q3 + 3*IQR
B.Values beyond Q1 - 1.5*IQR or Q3 + 1.5*IQR
C.Values beyond mean ± 2 standard deviations
D.Values beyond min and max
AnswerB

Values falling below Q1 − 1.5×IQR or above Q3 + 1.5×IQR are flagged as outliers, satisfying the stem's requirement to identify the salary that sits significantly higher than the rest. Tukey's fences define this boundary, so an extreme high salary exceeding Q3 + 1.5×IQR is correctly detected.

Why this answer

Outliers are values less than Q1 - 1.5*IQR or greater than Q3 + 1.5*IQR.

35
MCQmedium

A retailer wants to test if a new website layout increases the average time spent on the site. They split traffic: control group (old layout) and treatment group (new layout). Which statistical test is most appropriate to compare the average time spent between the two groups?

A.ANOVA
B.Pearson correlation
C.Chi-square test
D.Two-sample t-test
AnswerD

A two-sample t-test compares the means of two independent groups, matching the control and treatment split. It tests whether the difference in average time spent is statistically significant, which is exactly the retailer's question about whether the new layout changes the mean.

Why this answer

The question compares the means of a continuous outcome (average time spent) between two independent groups (control vs. treatment), which is exactly the two-sample t-test's purpose. The two-sample t-test evaluates whether the difference in group means is statistically significant, assuming approximately normal distributions or sufficiently large samples. It is the standard test for a two-group A/B comparison on a continuous metric.

Exam trap

The trap is confusing the two-sample t-test with the paired t-test — candidates must check whether the two groups are independent (different subjects) or paired (same subjects measured twice), since that determines which t-test applies.

How to eliminate wrong answers

Option A is wrong because ANOVA is used to compare means across three or more groups; with only two groups it is equivalent to a t-test but is not the most appropriate or standard choice. Option B is wrong because Pearson correlation measures the linear relationship between two continuous variables, not the difference in means between two groups. Option C is wrong because the chi-square test is used for categorical data (comparing observed vs. expected frequencies), not for comparing continuous means like time spent.

36
MCQeasy

A data analyst is working with a dataset that contains a column for 'order_date' stored as text in the format 'YYYY-MM-DD'. The analyst needs to calculate the average time between orders for each customer. Which of the following steps should the analyst take first?

A.Calculate the difference between consecutive 'order_date' values using string manipulation.
B.Group the data by customer and compute the average of the 'order_date' column.
C.Convert the 'order_date' column to a date data type using a function like TO_DATE or PARSE_DATE.
D.Sort the dataset by the 'order_date' column in ascending order.
AnswerC

Converting the text column to a proper date data type is essential before performing any date arithmetic. Functions like TO_DATE (in SQL) or PARSE_DATE (in BigQuery) transform the string into a date value, enabling calculations such as differences between dates. Without this step, the analyst cannot accurately compute time intervals, as text comparison would yield incorrect results.

Why this answer

To calculate the average time between orders, the analyst must first convert the text-based date column into a proper date data type. This enables the use of date functions to compute intervals between consecutive orders. Without this conversion, any arithmetic on dates would be invalid or highly error-prone.

Once converted, the analyst can sort by customer and date, then compute differences.

Exam trap

The trap here is assuming that sorting or string manipulation can substitute for proper date conversion, but accurate date arithmetic requires a date data type.

37
MCQmedium

A data analyst is performing a hypothesis test with a significance level of 0.05. The p-value obtained is 0.03. What should the analyst conclude?

A.Reject the null hypothesis
B.Fail to reject the null hypothesis
C.Accept the null hypothesis
D.The result is practically significant
AnswerA

The p-value of 0.03 falls below the 0.05 significance level, meaning the observed result is unlikely under the null hypothesis. The analyst therefore rejects the null hypothesis in favour of the alternative, accepting the result as statistically significant.

Why this answer

Since the p-value (0.03) is less than the significance level (0.05), the result is statistically significant. This means the observed data provides sufficient evidence to reject the null hypothesis in favor of the alternative hypothesis. The analyst should conclude that there is a statistically significant effect or difference.

Exam trap

The trap here is that candidates often confuse 'fail to reject' with 'accept' the null hypothesis, or they mistakenly think a p-value less than α means the null hypothesis is proven false with certainty, rather than just providing sufficient evidence to reject it.

How to eliminate wrong answers

Option B is wrong because failing to reject the null hypothesis occurs only when the p-value is greater than or equal to the significance level (p ≥ 0.05), not when it is smaller. Option C is wrong because hypothesis testing never 'accepts' the null hypothesis; we either reject it or fail to reject it, as acceptance implies proof of truth, which is not a valid statistical conclusion. Option D is wrong because practical significance is a separate consideration from statistical significance; a statistically significant result (p < 0.05) does not automatically imply practical importance, and the question only asks about the hypothesis test conclusion.

38
MCQeasy

A data analyst needs to visualize the relationship between advertising spend and sales revenue for 50 products. Which chart type is most appropriate?

A.Histogram
B.Bar chart
C.Pie chart
D.Scatter plot
AnswerD

A scatter plot places advertising spend on one axis and sales revenue on the other, plotting each of the 50 products as a point. This reveals correlation, clusters and outliers between two continuous variables, which bar or line charts cannot show.

Why this answer

A scatter plot is the most appropriate chart type for visualizing the relationship between two continuous variables—advertising spend and sales revenue—across 50 products. It allows the data analyst to assess correlation, trends, and potential outliers by plotting each product as a point on a Cartesian plane, with advertising spend on the x-axis and sales revenue on the y-axis.

Exam trap

The trap here is that candidates often confuse a bar chart or histogram with a scatter plot when the question involves two numeric variables, but a bar chart is designed for categorical comparisons and a histogram for single-variable distributions, not for bivariate relationships.

How to eliminate wrong answers

Option A is wrong because a histogram is used to display the distribution of a single continuous variable by grouping data into bins, not to show the relationship between two variables. Option B is wrong because a bar chart compares categorical data or discrete values, not the continuous relationship between two numeric variables like advertising spend and sales revenue. Option C is wrong because a pie chart shows proportions of a whole for categorical data, making it unsuitable for visualizing the correlation between two continuous variables.

39
MCQhard

A data analyst runs an A/B test on a new website layout. The test yields a p-value of 0.04 with the null hypothesis being no difference in conversion rates. The significance threshold is α=0.05. Which of the following is the correct conclusion?

A.The result is not significant; accept the alternative hypothesis.
B.Reject the null hypothesis; the new layout is proven to increase conversions.
C.Reject the null hypothesis; there is a statistically significant difference in conversion rates.
D.Fail to reject the null hypothesis; there is no evidence of a difference.
AnswerC

With p = 0.04 falling below the α = 0.05 threshold, the result lies inside the rejection region, so the null hypothesis of no difference is rejected. This satisfies the stem's significance criterion, supporting a statistically significant difference in conversion rates between the layouts.

Why this answer

Since p-value (0.04) < α (0.05), we reject the null hypothesis and conclude there is a statistically significant difference. However, statistical significance does not guarantee practical significance.

40
MCQmedium

A healthcare analytics team is ingesting a nightly CSV extract of patient encounters. The extract occasionally contains rows where the 'discharge_date' is earlier than the 'admission_date'. The team wants the pipeline to flag these rows for review rather than silently load them. Which data preparation action best addresses this requirement?

A.Swap the values in admission_date and discharge_date whenever discharge_date is earlier.
B.Load the rows as-is and rely on downstream report filters to exclude invalid dates.
C.Apply a validation rule that compares admission_date and discharge_date and routes failing rows to a quarantine table with a reason code.
D.Drop all rows where discharge_date is earlier than admission_date before loading.
AnswerC

A validation rule that checks the chronological relationship and diverts violations to quarantine preserves the suspect records while preventing them from entering the trusted layer. Attaching a reason code supports triage and root-cause analysis. This matches the requirement to flag rather than silently load or delete, and it is a standard data quality control in healthcare ETL pipelines.

Why this answer

The requirement is to identify and isolate suspect records for human review. A validation rule comparing the two dates and routing failures to a quarantine table with a reason code achieves this without deleting data or contaminating the trusted layer. Dropping, passing through, or auto-correcting records all fail because they either lose information, hide the problem, or make unverified changes to clinical data.

Exam trap

The trap here is treating a data quality exception as something to fix or delete automatically, when the stated requirement is to flag it for review.

41
MCQhard

A business intelligence analyst is designing a dashboard for executives. The dashboard includes revenue, profit margin, and customer satisfaction score. To maximize the data-ink ratio, the analyst should:

A.Add a background image of the company logo
B.Remove gridlines, use minimal colors, and avoid decorative elements
C.Use different font styles for each metric label
D.Use 3D pie charts with shadow effects
AnswerB

Removing gridlines, minimising colour and stripping decoration directly raises the data-ink ratio, since every drop of ink then encodes revenue, profit margin or satisfaction data rather than ornament. This satisfies the stem's constraint of maximising data-ink ratio, a principle Tufte defined as data-ink divided by total ink used.

Why this answer

Maximizing the data-ink ratio means every drop of ink should convey data, so removing gridlines, minimizing color, and stripping decorative elements directly increases the proportion of ink devoted to actual data. This aligns with Edward Tufte's principle that non-data ink and redundant ink should be eliminated or minimized. Executives need fast, clear signal, so a clean chart with minimal chrome communicates revenue, margin, and CSAT more effectively.

Exam trap

DA0-002 often tests the misconception that visual appeal (logos, 3D effects, varied fonts) improves dashboards, when it actually lowers the data-ink ratio.

How to eliminate wrong answers

Option A is wrong because a background logo is pure decoration (non-data ink) that reduces the data-ink ratio and can obscure data points. Option C is wrong because varying font styles per metric adds visual noise and inconsistency without conveying additional data, hurting readability. Option D is wrong because 3D pie charts with shadows distort proportions and add redundant ink, making comparisons less accurate and violating data-ink ratio principles.

42
MCQeasy

A data analyst needs to combine two datasets: one contains customer information (customer_id, name, address) and the other contains order information (order_id, customer_id, order_date). The analyst wants to include all customers, even those who have not placed orders. Which type of join should be used?

A.FULL OUTER JOIN
B.INNER JOIN
C.LEFT JOIN
D.RIGHT JOIN
AnswerC

A LEFT JOIN returns every row from the customer table plus matching order rows, emitting NULLs where no orders exist. This satisfies the stem's requirement to retain customers without orders, unlike an INNER JOIN which would exclude them.

Why this answer

A LEFT JOIN returns all rows from the left table (customers) and the matching rows from the right table (orders). If a customer has no orders, the order columns will contain NULLs. This satisfies the requirement to include all customers, even those without orders.

Exam trap

The trap here is that candidates often confuse LEFT JOIN with INNER JOIN, assuming all customers must have orders, or they pick FULL OUTER JOIN thinking it includes all customers, but it also includes unmatched orders, which is not required.

How to eliminate wrong answers

Option A is wrong because a FULL OUTER JOIN returns all rows from both tables, which would include unmatched orders (if any) — unnecessary for this requirement. Option B is wrong because an INNER JOIN returns only rows with matching keys in both tables, excluding customers who have never placed an order. Option D is wrong because a RIGHT JOIN returns all rows from the right table (orders) and only matching rows from the left table (customers), which would exclude customers without orders.

43
Multi-Selecthard

A data analyst is creating a data story about sales performance. Which THREE elements are essential for effective data storytelling? (Choose THREE.)

Select 3 answers
A.Raw data tables for reference.
B.A clear narrative with a beginning, middle, and end.
C.Context and background information.
D.Use of multiple chart types to show variety.
E.A call to action.
AnswersB, C, E

A clear narrative with a beginning, middle and end gives the sales data a logical structure, guiding the audience from the situation through the analysis to the conclusion, which is essential for coherent, persuasive data storytelling.

Why this answer

Option B is correct because effective data storytelling requires a clear narrative arc — a beginning that sets up the situation, a middle that presents the analysis and insights, and an end that resolves the story — which guides the audience logically from data to meaning. Option C is correct because context and background information are essential to make the data interpretable, helping the audience understand the business situation, benchmarks, and why the findings matter. Option E is correct because a call to action directs the audience toward a specific decision or next step, which is the ultimate purpose of a data story and distinguishes storytelling from mere data reporting.

Option A does not belong because raw data tables are supporting evidence, not an essential storytelling element; dumping raw data can overwhelm the audience and obscure the narrative. Option D does not belong because using multiple chart types is a stylistic choice for variety, not a requirement — a single well-chosen visualization can be more effective than several charts used merely for the sake of variety.

Exam trap

CompTIA often tests the distinction between supporting elements (like raw data tables or chart variety) and the core structural components (narrative, context, call to action) that define effective data storytelling.

44
MCQmedium

Refer to the exhibit. What does the query return?

A.All orders grouped by customer ID.
B.Customers who have placed at least 5 orders.
C.Customers who have placed more than 5 orders.
D.All customers who have placed orders.
AnswerC

The HAVING clause filters aggregated groups rather than individual rows, so COUNT(order_id) grouped by customer is compared against 5. Only customers whose order count exceeds that threshold are returned, excluding those with five or fewer orders.

Why this answer

The query uses a HAVING clause with COUNT(*) > 5, which filters groups (by customer ID) to only those with more than 5 orders. The GROUP BY customer ID ensures the count is per customer, so the result is customers who have placed more than 5 orders. Option C is correct because the condition is strictly greater than 5, not at least 5.

Exam trap

CompTIA often tests the distinction between 'at least' (>=) and 'more than' (>) in HAVING clauses, and candidates may misread the condition as including exactly 5 orders.

How to eliminate wrong answers

Option A is wrong because the query does not return all orders; it returns aggregated results (counts) per customer, not individual order rows. Option B is wrong because the condition is COUNT(*) > 5, not COUNT(*) >= 5; 'at least 5' would include exactly 5, which is excluded by the strict greater-than operator. Option D is wrong because the HAVING clause filters out customers with 5 or fewer orders; the query does not return all customers who have placed orders, only those exceeding the threshold.

45
Matchingmedium

Match each data quality dimension to its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Degree to which data correctly reflects real-world values

Extent to which all required data is present

Absence of contradictions across data sources

Data is up-to-date and available when needed

No duplicate records exist within the dataset

Why these pairings

The correct matches are: Accuracy - data correctly reflects real-world values; Completeness - all required data is present; Consistency - data values are the same across systems; Timeliness - data is available when needed. Common confusions include mixing timeliness with accuracy and consistency with completeness.

46
Multi-Selectmedium

A data analyst is performing a chi-square test of independence on a 2x2 contingency table. The p-value is 0.04. At α=0.05, which THREE of the following statements are correct?

Select 3 answers
A.There is a statistically significant association between the two variables.
B.The test indicates a strong association between variables.
C.The variables are not independent.
D.The null hypothesis is rejected.
E.The result is not statistically significant.
AnswersA, C, D

With p=0.04 below the α=0.05 threshold, the null hypothesis of independence is rejected, so a statistically significant association exists between the two variables. The test detects a relationship, though it does not quantify its strength or direction.

Why this answer

Option A is correct because with p = 0.04 < α = 0.05, the chi-square test of independence yields a statistically significant result, indicating evidence of an association between the two variables. Option C is correct because rejecting the null hypothesis of independence means the data support that the two variables are not independent (i.e., they are associated). Option D is correct because the decision rule for a chi-square test of independence is to reject the null hypothesis when p < α, and 0.04 < 0.05.

Option B is not correct because statistical significance does not imply a strong association; strength would require an effect size such as Cramér's V or the phi coefficient. Option E is not correct because p = 0.04 is below α = 0.05, so the result is statistically significant, not non-significant.

Exam trap

DA0-002 often tests the misconception that a statistically significant p-value implies a strong or large effect, conflating significance with effect size.

47
MCQhard

A data analyst notices that a column labeled 'Income' contains values like '$50,000' and '$75,000', but also 'High' and 'Low'. What data concept issue is occurring?

A.Mixing quantitative and qualitative data
B.Mixing discrete and continuous data
C.Mixing nominal and ordinal data
D.Mixing structured and unstructured data
AnswerA

The Income column holds numeric currency amounts alongside categorical labels such as High and Low, so a single field contains both quantitative and qualitative values. This mixed-type inconsistency prevents valid aggregation or statistical analysis of that column.

Why this answer

The 'Income' column contains both numeric values (e.g., '$50,000', '$75,000') which are quantitative data, and categorical labels ('High', 'Low') which are qualitative data. Mixing these two distinct data types in a single column violates data consistency principles and prevents proper statistical analysis or machine learning processing. This is a classic example of mixing quantitative and qualitative data.

Exam trap

CompTIA often tests the distinction between data type categories (quantitative vs. qualitative) versus subtypes (discrete/continuous or nominal/ordinal), so candidates mistakenly pick a subtype option when the core issue is the fundamental type mismatch.

How to eliminate wrong answers

Option B is wrong because discrete and continuous data are both subtypes of quantitative data (e.g., number of children vs. height), but the issue here is mixing numbers with text labels, not distinguishing between countable and measurable values. Option C is wrong because nominal and ordinal data are both categorical (qualitative) subtypes (e.g., colors vs. rankings), but the column includes actual numeric income values, not just ordered categories. Option D is wrong because structured data refers to organized formats like tables (which this column is part of), while unstructured data refers to free-form text or media; the problem is not about format but about inconsistent data types within a structured field.

48
Multi-Selectmedium

A data analyst at a healthcare organization is preparing a dataset for analysis that includes patient identifiers. The organization must comply with HIPAA and internal data governance policies. Which two actions should the analyst take to protect patient privacy while enabling analysis? (Choose two.)

Select 2 answers
A.Aggregate all patient data into summary statistics before any analysis is performed.
B.Replace patient names and Social Security numbers with pseudonyms before sharing the dataset with analysts.
C.Store the dataset on a shared network drive with read permissions for the entire department.
D.Remove or mask direct identifiers such as names, addresses, and phone numbers from the dataset.
E.Encrypt the dataset using AES-256 and share the decryption key with all analysts.
AnswersB, D

Pseudonymization replaces direct identifiers with artificial identifiers, reducing privacy risk while preserving the ability to link records. It is a recognized HIPAA-safe harbor technique when applied correctly. This allows analysts to work with the data without exposing protected health information directly, supporting both privacy and utility.

Why this answer

To protect patient privacy while enabling analysis, the analyst should apply de-identification techniques. Masking or removing direct identifiers prevents immediate identification, and pseudonymization further reduces risk by replacing identifiers with codes. Together, these actions support HIPAA compliance and internal governance by minimizing exposed protected health information while preserving data utility for analysis.

Exam trap

The trap here is thinking that encryption alone or broad access controls are sufficient, when the core need is to de-identify the data itself.

49
MCQhard

A company needs to store user session data for a web application. Each session has a unique session ID, and the data must be retrieved very quickly by session ID. The data does not require complex relationships or transactions. Which type of NoSQL database is most appropriate?

A.Key-value store (e.g., Redis)
B.Wide-column store (e.g., Cassandra)
C.Document store (e.g., MongoDB)
D.Graph database (e.g., Neo4j)
AnswerA

A key-value store maps each session ID directly to its session data, giving the fastest possible lookup by key. It satisfies the stem's requirement for quick retrieval by session ID without complex relationships or transactions.

Why this answer

A key-value store like Redis is the most appropriate choice because it is optimized for extremely fast lookups by a unique key (session ID) and does not require complex relationships or transactions. Redis stores data in memory, providing sub-millisecond retrieval times ideal for session management, and supports built-in expiration (TTL) to automatically clean up stale sessions.

Exam trap

The trap here is that candidates often choose a document store like MongoDB because they associate 'session data' with JSON objects, overlooking that key-value stores are purpose-built for the exact use case of fast, simple key-based retrieval without the overhead of document querying.

How to eliminate wrong answers

Option B (Wide-column store, e.g., Cassandra) is wrong because it is designed for high-volume, distributed writes and complex query patterns over column families, not for simple, low-latency key-based lookups; its eventual consistency model and overhead for single-key reads make it overkill for session storage. Option C (Document store, e.g., MongoDB) is wrong because it stores semi-structured JSON-like documents with rich querying capabilities, which adds unnecessary complexity and latency for simple session data that only needs key-based retrieval. Option D (Graph database, e.g., Neo4j) is wrong because it is purpose-built for traversing relationships between entities (nodes and edges), which is irrelevant for session data that has no relational structure.

50
MCQeasy

A data analyst needs to present findings about customer churn to business stakeholders. The analysis identified that churn is highest among customers who have called customer support more than three times in the last month. Which of the following is the best way to communicate this insight?

A.A scatter plot to show the relationship between support calls and churn.
B.A pie chart showing the proportion of churned vs. retained customers.
C.A bar chart comparing churn rates for different support call counts.
D.A table of raw churn data by customer ID.
AnswerC

A bar chart plotting churn rate against support call count directly exposes the threshold relationship, showing churn rising beyond three calls. This visual encoding satisfies the need to communicate a quantitative pattern to business stakeholders clearly and immediately.

Why this answer

A bar chart directly compares churn rates across discrete categories of support call counts (e.g., 0, 1, 2, 3, 4+ calls), making it easy for stakeholders to see the spike at 'more than three calls'. This aligns with the insight that churn is highest among customers with >3 support calls, and a bar chart is the standard visualization for comparing a continuous metric (churn rate) across categorical bins.

Exam trap

The trap here is that candidates may choose a scatter plot (Option A) because they think it shows 'relationship', but they fail to recognize that a scatter plot is inappropriate for a binary dependent variable and discrete independent variable, whereas a bar chart is the correct choice for comparing rates across categories.

How to eliminate wrong answers

Option A is wrong because a scatter plot is used to show the relationship between two continuous variables, but here the independent variable (number of support calls) is discrete and the dependent variable (churn) is binary, so a scatter plot would produce overlapping points and fail to clearly communicate the categorical threshold of 'more than three calls'. Option B is wrong because a pie chart only shows the overall proportion of churned vs. retained customers, which does not convey the relationship between support call frequency and churn, missing the key insight entirely. Option D is wrong because a table of raw churn data by customer ID presents unaggregated, granular data that obscures the pattern and is not suitable for a high-level stakeholder presentation; it would require the audience to manually compute churn rates per call count.

51
MCQeasy

A marketing team needs to store customer feedback from social media posts, including text, images, and emojis. Which data concept is most appropriate for this storage?

A.Unstructured data in a NoSQL document database
B.Structured data in a relational database
C.Unstructured data in a relational database
D.Semi-structured data in an XML database
AnswerA

NoSQL document databases store each post as a self-describing document, holding text, image references and emoji characters together without a fixed relational schema. This satisfies the requirement to capture mixed-format social media feedback, which varies in structure between posts.

Why this answer

Customer feedback from social media includes text, images, and emojis, which lack a predefined schema and are best stored as unstructured data. NoSQL document databases (e.g., MongoDB) store such data in flexible JSON-like documents, allowing each record to have varying fields and data types without requiring a fixed schema.

Exam trap

CompTIA often tests the misconception that 'unstructured data' cannot be stored in any database, when in fact NoSQL document databases are purpose-built for it, while relational databases require rigid schemas that fail with variable content.

How to eliminate wrong answers

Option B is wrong because structured data in a relational database requires a fixed schema with predefined columns and data types, which cannot efficiently handle variable-length text, images, and emojis without complex workarounds like BLOBs. Option C is wrong because relational databases are designed for structured data; storing unstructured data in them forces schema rigidity and poor performance for heterogeneous content. Option D is wrong because XML databases are semi-structured and impose hierarchical markup, which is unnecessary overhead for social media posts that are naturally schema-less and better served by document stores.

52
Multi-Selectmedium

Which TWO of the following data quality dimensions are most directly affected by duplicate records?

Select 2 answers
A.Timeliness
B.Consistency
C.Uniqueness
D.Accuracy
E.Completeness
AnswersC, D

Duplicate records mean the same real-world entity appears more than once, directly violating the uniqueness dimension, which measures whether each record occurs only once. This is the dimension most immediately degraded by duplication, independent of whether the underlying values are otherwise correct.

Why this answer

Duplicate records most directly undermine Uniqueness (C), because uniqueness is the data quality dimension that measures whether each real-world entity appears exactly once in a dataset; duplicates violate this one-record-per-entity rule by definition. They also directly affect Accuracy (D), since duplicated rows distort counts, aggregations, and the true representation of the entity, making the stored data an incorrect reflection of reality. Timeliness (A) concerns whether data is current and available when needed, which duplicates do not inherently affect.

Consistency (B) refers to agreement of the same data across systems or formats, which is a separate issue from repeated rows. Completeness (E) measures whether required data is present, and duplicates add data rather than omit it, so it is not the dimension most directly impacted.

Exam trap

The trap is assuming that duplicates only affect uniqueness, but they also impact accuracy because the data no longer correctly represents the real-world entity.

53
Multi-Selecthard

A data architect is evaluating storage engines for a new analytics platform. The platform must support storing large volumes of structured historical sales data and must efficiently handle queries that aggregate a few columns across billions of rows. The architect is considering column-oriented storage. Which two characteristics accurately describe column-oriented storage in this context? (Choose two.)

Select 2 answers
A.It requires reading all columns of a row even when only one column is needed for a query
B.It typically achieves better data compression than row-oriented storage for repetitive column values
C.It is optimized for high-frequency single-row inserts and updates typical of transactional workloads
D.It eliminates the need for indexing because every column is stored in sorted order automatically
E.It stores each column's values contiguously, improving compression and scan efficiency for analytical aggregates
AnswersB, E

Because values within a column share the same data type and often repeat or fall within narrow ranges, column-oriented storage can apply dictionary, run-length, and delta encoding effectively. This yields higher compression than row-oriented formats, where heterogeneous row values limit encoding options. Reduced storage and I/O directly benefit large-scale analytical scans.

Why this answer

Column-oriented storage groups values by column, which enables efficient compression and allows queries to read only the columns they need. These two properties directly support aggregating a few columns across billions of rows. The remaining statements describe row-oriented behavior, overstate write performance, or misrepresent indexing requirements, so they do not accurately characterize column-oriented storage for this analytics platform.

Exam trap

The trap here is assuming that column-oriented storage improves all workloads equally, when its advantages apply to analytical scans and compression while row-oriented engines remain superior for frequent single-row writes.

54
Multi-Selecthard

A data engineer is designing a data pipeline for a retail company. The source system is an OLTP database that records sales transactions. The target is a data warehouse used for reporting. The engineer is evaluating whether to use ETL or ELT. Which three factors would favor using ELT over ETL? (Select THREE)

Select 3 answers
A.The transformation logic requires proprietary functions not available in the warehouse
B.The business analysts need access to raw data for ad-hoc exploration
C.The target data warehouse has massive compute power (e.g., Snowflake) that can handle transformations efficiently
D.Data must be cleansed and validated before loading into the warehouse
E.The source data volume is very large and the warehouse can scale resources on demand
AnswersB, C, E

ELT loads raw data, allowing analysts to explore it before transformation.

Why this answer

ELT loads raw data into the warehouse first, allowing business analysts to perform ad-hoc exploration directly on the source data without pre-transformation. This flexibility is a key advantage of ELT over ETL, where transformations are applied before loading.

Exam trap

The trap here is that candidates often confuse the direction of data flow, mistakenly thinking that ELT requires transformations before loading, when in fact ELT defers transformations until after data is in the warehouse.

55
MCQhard

A company uses a NoSQL document database to store product catalogs. Each product document includes fields like product_id, name, category, and price. The operations team frequently queries by product_id and by category. Which type of NoSQL database is being used, and what should be created to optimize queries by category?

A.Key-value store; create a secondary index on category
B.Graph database; create a relationship between products
C.Document database; create an index on category
D.Wide-column store; create a column family for category
AnswerC

A document database stores each product as a self-describing JSON-like document, matching the schema described. Indexing the category field builds a secondary lookup structure, so queries filtering by category avoid full collection scans and return results efficiently.

Why this answer

The question explicitly states a document database is used, and the operations team frequently queries by category. In a document database like MongoDB, creating an index on the category field optimizes these queries by allowing the database to quickly locate documents without scanning every document in the collection. This is the standard approach for improving query performance on non-primary-key fields in document stores.

Exam trap

CompTIA Data+ may test the misconception that all NoSQL databases support secondary indexes similarly. The trap here is that document databases natively support secondary indexes, while key-value and wide-column stores require different optimization strategies.

How to eliminate wrong answers

Option A is wrong because a key-value store does not support secondary indexes on fields like category; it only allows lookups by the primary key (product_id), making it unsuitable for the described query pattern. Option B is wrong because a graph database is designed for relationship-heavy data (e.g., social networks), not for product catalogs with simple field queries, and creating relationships between products does not optimize category-based lookups. Option D is wrong because a wide-column store organizes data by column families, not by documents, and creating a column family for category would not provide the index-based optimization needed for document-style queries.

56
MCQmedium

In a dataset with variables on different scales (e.g., age in years and income in dollars), which preprocessing step is necessary before applying k-means clustering?

A.Feature selection
B.Dimensionality reduction
C.Normalization (scaling)
D.One-hot encoding
AnswerC

Normalization ensures each feature contributes equally to distance calculations.

Why this answer

K-means clustering relies on Euclidean distance to measure similarity between data points. When variables like age (in years) and income (in dollars) are on different scales, the variable with larger numeric values (income) will dominate the distance calculation, skewing the clustering results. Normalization (scaling), such as min-max scaling or z-score standardization, rescales all features to a comparable range (e.g., [0,1] or mean=0, variance=1), ensuring each feature contributes equally to the distance computation.

Exam trap

The trap here is that candidates may confuse normalization with other preprocessing steps like feature selection or dimensionality reduction, thinking that removing irrelevant features or reducing dimensions will automatically fix scale differences, but k-means specifically requires scaling to ensure equal feature influence in distance calculations.

How to eliminate wrong answers

Option A is wrong because feature selection is about choosing a subset of relevant features to reduce noise or improve model performance, but it does not address the issue of differing scales among features, which is required before k-means. Option B is wrong because dimensionality reduction (e.g., PCA) reduces the number of features, but it does not inherently scale the data; scaling is typically performed before dimensionality reduction, not as a substitute for it. Option D is wrong because one-hot encoding is used to convert categorical variables into numerical format, not to handle numerical variables on different scales; applying one-hot encoding to already numerical features would be incorrect and does not solve the scaling problem.

57
MCQhard

A data analyst at a retail company is building a multiple linear regression model to forecast weekly sales. The dataset contains 50 predictor variables, including store size, promotional spend, holiday indicators, and many others. After training the model, the analyst observes an R-squared of 0.99 on the training set but only 0.55 on the holdout test set. Which action should the analyst take first to address this discrepancy?

A.Remove highly correlated predictor variables and apply regularization (e.g., Ridge or Lasso).
B.Add more predictor variables to increase the training R-squared further.
C.Use k-fold cross-validation with a different random seed to get a more reliable test set estimate.
D.Increase the number of hidden layers in the model to capture more complexity.
AnswerA

The 0.99 versus 0.55 gap indicates overfitting from 50 predictors, many correlated. Removing correlated variables reduces multicollinearity and dimensionality, while Ridge or Lasso regularisation penalises large coefficients, shrinking variance and improving holdout generalisation before other remedies are attempted.

Why this answer

The high R-squared of 0.99 on training data versus 0.55 on test data is a classic sign of overfitting, where the model has learned noise and specific patterns in the training set that do not generalize. Removing highly correlated predictors reduces multicollinearity and model complexity, while regularization (Ridge or Lasso) penalizes large coefficients, shrinking them to prevent overfitting. This is the most direct first step to improve generalization.

Exam trap

The trap here is that candidates may think a high R-squared is always good, or they may confuse overfitting with underfitting and choose to add more complexity (Option D) or more data (Option B), rather than recognizing the need to reduce model complexity and apply regularization.

How to eliminate wrong answers

Option B is wrong because adding more predictor variables would increase the training R-squared but worsen overfitting, making the test set performance even lower. Option C is wrong because k-fold cross-validation with a different random seed does not address the fundamental overfitting issue; it only provides a different estimate of test error but does not change the model's tendency to overfit. Option D is wrong because increasing the number of hidden layers (a neural network technique) is irrelevant for a multiple linear regression model and would introduce unnecessary complexity, likely exacerbating overfitting.

58
MCQeasy

A marketing analyst needs to combine customer data from a CRM database with social media engagement data from a third-party API. Which data acquisition method is most appropriate?

A.Web scraping
B.Manual data entry
C.API integration
D.Batch file upload
AnswerC

API integration lets the analyst pull social media engagement data programmatically from the third-party service and join it with CRM records, satisfying the need to combine two distinct sources. It handles authentication, scheduled retrieval and structured responses that flat-file exports cannot provide reliably.

Why this answer

API integration is the most appropriate method because it allows the analyst to programmatically retrieve structured social media engagement data directly from the third-party service's RESTful or GraphQL API endpoints. This approach ensures real-time or near-real-time data synchronization, supports authentication (e.g., OAuth 2.0), and returns data in standardized formats like JSON or XML, which can be directly ingested into the CRM system without manual intervention.

Exam trap

The trap here is that candidates may confuse web scraping with API integration, assuming both can retrieve web data, but the question specifically requires combining structured data from a third-party API, where web scraping would be unreliable, unauthorized, and technically inappropriate for programmatic data acquisition.

How to eliminate wrong answers

Option A is wrong because web scraping is used to extract unstructured data from HTML pages, which is inefficient, brittle, and often violates the third-party API's terms of service; it is not designed for reliable, authenticated access to structured social media metrics. Option B is wrong because manual data entry is error-prone, time-consuming, and impractical for large volumes of social media engagement data, and it lacks any automated validation or consistency checks. Option D is wrong because batch file upload assumes the data is already exported into a file (e.g., CSV) and delivered manually, which introduces latency and requires the third-party to support file exports, whereas the API provides direct, on-demand access to live data.

59
MCQeasy

A data analyst is working with a dataset that contains a column for customer satisfaction ratings on a scale of 1 to 5, where 1 means 'very dissatisfied' and 5 means 'very satisfied'. The analyst wants to calculate the median satisfaction score. Which level of measurement does this data represent?

A.Ratio
B.Ordinal
C.Nominal
D.Interval
AnswerB

Ordinal data has a meaningful order but the intervals between values are not necessarily equal. Satisfaction ratings from 1 to 5 are ordered, but the difference between 1 and 2 may not equal the difference between 4 and 5. The median is a valid measure of central tendency for ordinal data, making this the correct classification.

Why this answer

Customer satisfaction ratings on a 1-5 scale are ordered but the intervals between points are not equal, so they are ordinal. The median is the appropriate measure of central tendency for ordinal data. Nominal data lacks order, interval data requires equal intervals, and ratio data requires a true zero, none of which apply here.

Exam trap

The trap here is assuming that numeric labels always indicate interval or ratio data, when in fact ordered categories like satisfaction ratings are ordinal.

60
Multi-Selecthard

Which THREE are best practices for acquiring data via web scraping? (Select exactly 3)

Select 3 answers
A.Use multiple IP addresses
B.Respect robots.txt
C.Identify yourself with a user-agent
D.Scrape all data without regard to terms
E.Limit request rate
AnswersB, C, E

Honouring robots.txt satisfies the legal and ethical constraint of scraping only permitted paths. The file declares which directories crawlers may access, so parsing it before requesting pages prevents the crawler from breaching site owner directives and reduces the risk of being blocked or facing legal action.

Why this answer

Option B is correct because robots.txt is the standard mechanism (the Robots Exclusion Protocol) by which site owners declare which paths crawlers may access, and honoring it keeps scraping ethical and within the site's stated permissions. Option C is correct because setting a descriptive User-Agent header identifies your crawler and provides contact information, letting administrators recognize and reach you rather than blocking you as an anonymous bot. Option E is correct because limiting request rate (e.g., throttling to a modest number of requests per second and honoring Retry-After on 429/503 responses) avoids overloading servers and reduces the chance of being rate-limited or banned.

Option A does not belong because rotating multiple IP addresses to evade blocks is an evasion technique, not a best practice, and can violate terms of service. Option D does not belong because ignoring terms of service and scraping indiscriminately is unethical and often unlawful, contrary to responsible scraping.

61
Multi-Selectmedium

A data analyst is evaluating a classification model that predicts whether a customer will churn. The model's confusion matrix shows 80 true positives, 20 false negatives, 30 false positives, and 120 true negatives. Which TWO of the following metrics can be directly calculated from this confusion matrix? (Choose two.)

Select 2 answers
A.Accuracy
B.R-squared
C.Root Mean Square Error (RMSE)
D.Precision
E.Adjusted R-squared
AnswersA, D

Accuracy is the proportion of correct predictions: (TP + TN) / total. Here, (80 + 120) / (80+20+30+120) = 200/250 = 0.8. It can be directly calculated from the confusion matrix, making it a valid metric for this scenario. It provides an overall measure of correct classifications.

Why this answer

Accuracy and precision are both classification metrics that can be computed directly from the counts in a confusion matrix. Accuracy gives the overall correct prediction rate, while precision focuses on the reliability of positive predictions. R-squared, RMSE, and adjusted R-squared are regression metrics and cannot be calculated from a confusion matrix, so they are not applicable.

Exam trap

The trap here is selecting regression metrics like R-squared or RMSE for a classification problem, confusing the evaluation metrics of different model types.

62
MCQhard

A data analyst is creating a report on customer satisfaction scores across different regions. The analyst wants to highlight regions that are significantly below average. Which of the following statistical methods is most appropriate for identifying these outliers?

A.Bar chart with average line.
B.Pie chart of satisfaction categories.
C.Box plot with interquartile range (IQR) to identify outliers.
D.Scatter plot of satisfaction vs. region.
AnswerC

The IQR defines the middle 50% of scores; values falling below Q1 minus 1.5×IQR sit statistically apart from the distribution, flagging regions genuinely below average. A box plot displays this spread and its outliers directly, satisfying the requirement to highlight underperforming regions.

Why this answer

A box plot with interquartile range (IQR) is the most appropriate method because it explicitly identifies outliers as data points falling below Q1 - 1.5*IQR or above Q3 + 1.5*IQR. This directly addresses the analyst's goal of highlighting regions significantly below average, as the IQR method is a standard statistical technique for detecting extreme values in a distribution.

Exam trap

The trap here is that candidates may choose a bar chart with an average line (Option A) because it visually shows deviations, but it lacks a formal statistical criterion to define 'significantly below average,' which the IQR-based box plot provides.

How to eliminate wrong answers

Option A is wrong because a bar chart with an average line only shows the mean and individual region values, but does not provide a statistical threshold to determine which regions are significantly below average; it merely visualizes deviations without identifying outliers. Option B is wrong because a pie chart of satisfaction categories shows proportions of categorical data, not numerical scores across regions, and cannot identify outliers or deviations from the mean. Option D is wrong because a scatter plot of satisfaction vs. region treats region as a categorical variable on one axis, which does not produce a meaningful distribution for outlier detection; it would simply plot points per region without any statistical measure of dispersion or outlier boundaries.

63
MCQmedium

During data profiling, an analyst wants to identify the number of distinct values in a column. Which SQL function should be used?

A.DISTINCT(column)
B.COUNT(DISTINCT column)
C.COUNT(*)
D.COUNT(column)
AnswerB

COUNT(DISTINCT column) returns the cardinality of a column by eliminating duplicate values before counting, directly satisfying the requirement to identify the number of distinct values during data profiling. Unlike COUNT(*), which counts every row including repeats, this aggregate yields the unique-value total the analyst needs.

Why this answer

COUNT(DISTINCT column) returns the number of unique non-null values in a column, which is exactly what data profiling requires to measure cardinality. It is the standard SQL aggregate for distinct value counts.

Exam trap

DA0-002 often tests the difference between COUNT(*), COUNT(column), and COUNT(DISTINCT column) — candidates must remember that only COUNT(DISTINCT) deduplicates values.

How to eliminate wrong answers

Option A is wrong because DISTINCT(column) is not a valid aggregate function — DISTINCT is a keyword used with SELECT or inside aggregate functions like COUNT, not a standalone function. Option C is wrong because COUNT(*) returns the total number of rows, including duplicates and nulls, not the number of distinct values. Option D is wrong because COUNT(column) returns the number of non-null values in the column, which still counts duplicates.

64
MCQhard

A data analyst is writing a query to rank products by total sales amount within each category. They want ties to have the same rank and no gaps in the ranking sequence. Which window function should they use?

A.ROW_NUMBER()
B.DENSE_RANK()
C.RANK()
D.NTILE()
AnswerB

DENSE_RANK() assigns identical ranks to tied sales totals yet never skips subsequent numbers, satisfying the no-gaps constraint. RANK() would leave gaps after ties, while ROW_NUMBER() breaks ties arbitrarily. Partitioning by category and ordering by total sales amount yields the required per-category sequence.

Why this answer

DENSE_RANK() assigns the same rank to tied rows and does not skip subsequent ranks, so if two products tie for rank 1, the next product receives rank 2 (no gap). This matches the requirement for ties to share a rank with no gaps in the sequence. It is the correct window function for this ranking behavior.

Exam trap

DA0-002 often tests the distinction between RANK() (gaps after ties) and DENSE_RANK() (no gaps), and candidates frequently pick RANK() when the question specifies 'no gaps'.

How to eliminate wrong answers

Option A (ROW_NUMBER()) is wrong because it assigns a unique sequential number to every row, even ties, so tied products would receive different ranks. Option C (RANK()) is wrong because it assigns the same rank to ties but skips subsequent ranks (e.g., 1,1,3), creating gaps. Option D (NTILE()) is wrong because it divides rows into a specified number of roughly equal buckets, not a ranking based on sales amount.

65
Multi-Selecthard

Which THREE of the following are appropriate methods to handle outliers in a dataset?

Select 3 answers
A.Transforming the data using log transformation
B.Removing the outlier records
C.Capping the outlier values at a certain percentile
D.Binning continuous variables
E.Imputing outliers with the mean
AnswersA, B, C

A log transformation compresses the upper tail of a right-skewed distribution, pulling extreme values closer to the bulk of the data. This reduces the leverage outliers exert on model estimates without deleting records, making it an appropriate handling method.

Why this answer

Log transformation (A) is correct because applying a log (or similar power) transform compresses the scale of extreme values, reducing the leverage of right-skewed outliers and making the distribution more symmetric for downstream models. Removing the outlier records (B) is correct when outliers are confirmed to be data-entry errors or otherwise invalid observations, since deleting them prevents distortion of statistics such as the mean and variance. Capping the outlier values at a certain percentile (C), also called winsorizing, is correct because it clips extreme values to a chosen percentile boundary (e.g., the 1st and 99th percentiles), preserving the record while limiting the outlier's influence.

Binning continuous variables (D) is not a dedicated outlier-handling method; it discretizes values into bins and merely masks extremes rather than addressing them. Imputing outliers with the mean (E) is inappropriate because the mean itself is distorted by the outliers, so replacing them with it introduces bias and does not resolve the underlying problem.

Exam trap

The trap here is that candidates may confuse data preprocessing techniques like binning or imputation with outlier handling methods, but binning is for discretization and mean imputation is not robust for outliers, while the correct methods (transformation, removal, capping) directly address outlier impact.

66
MCQeasy

A marketing team wants to collect data on competitor pricing for similar products. Which data source is most appropriate?

A.Customer surveys
B.Internal ERP system
C.External public web scraping
D.Internal sales data
AnswerC

External public web scraping directly satisfies the competitor-pricing constraint, since competitor prices sit outside the organisation's own systems and cannot be obtained from internal transactional databases or CRM records. Scraping public pages supplies the required market data without purchasing third-party reports, making it the most appropriate source for this competitive-intelligence scenario.

Why this answer

External public web scraping is the most appropriate data source because competitor pricing is publicly available on websites, and web scraping allows automated extraction of this structured or unstructured data. This approach directly addresses the need for external competitive intelligence without relying on internal or customer-reported data.

Exam trap

The trap here is that candidates may confuse internal data sources (ERP, sales) with external data needs, or mistakenly think customer surveys can provide accurate, unbiased competitor pricing data.

How to eliminate wrong answers

Option A is wrong because customer surveys collect subjective opinions and self-reported data, not objective, real-time competitor pricing from external sources. Option B is wrong because an internal ERP system contains only the company's own operational and financial data, not competitor pricing information. Option D is wrong because internal sales data reflects the company's own transactions and pricing, not competitor pricing.

67
Multi-Selectmedium

An analyst is preparing data for a clustering algorithm that uses Euclidean distance. Which TWO data preprocessing techniques should be applied to ensure all features contribute equally?

Select 2 answers
A.Min-max normalization
B.Z-score standardization
C.Log transformation
D.Principal component analysis
E.One-hot encoding
AnswersA, B

Min-max normalization rescales each feature to a fixed 0–1 range, removing unit and magnitude disparities before Euclidean distance is computed. This equalises each feature's contribution, satisfying the requirement that no single large-range variable dominates the clustering metric.

Why this answer

Min-max normalization (A) is correct because it rescales each feature to a common range, typically [0, 1], via (x - min) / (max - min), so that no feature dominates the Euclidean distance calculation due to a larger scale. Z-score standardization (B) is also correct because it transforms each feature to mean 0 and standard deviation 1 using (x - μ) / σ, which equalizes the variance and ensures all features contribute equally to Euclidean distance. Log transformation (C) is not a scale-equalizing technique; it only compresses skewed distributions and does not guarantee equal feature contribution.

Principal component analysis (D) is a dimensionality-reduction method, not a preprocessing step for equalizing feature scales. One-hot encoding (E) is for converting categorical variables into binary vectors and does not address differing numeric scales.

Exam trap

DA0-002 often tests the confusion between scaling techniques and other preprocessing methods like log transformation or PCA. Candidates might think log transformation scales features, but it only reduces skewness; or that PCA is a scaling method, but it's for dimensionality reduction.

68
MCQmedium

An analyst calculates a Pearson correlation coefficient of -0.8 between advertising spend and customer churn rate. Which interpretation is correct?

A.There is a weak positive relationship.
B.Advertising spend causes churn to decrease.
C.64% of the variance in churn is explained by spend.
D.Increasing advertising spend is associated with decreasing churn rate.
AnswerD

A coefficient of -0.8 indicates a strong negative linear association: as advertising spend rises, churn rate tends to fall. This satisfies the stem's requirement to interpret the sign and magnitude correctly, since the negative sign denotes an inverse relationship rather than causation or a weak correlation.

Why this answer

A Pearson correlation of -0.8 indicates a strong negative linear relationship, meaning that as advertising spend increases, customer churn rate tends to decrease. The correct interpretation is directional association, not causation — the coefficient describes the strength and direction of the linear relationship between the two variables.

Exam trap

DA0-002 often tests the correlation-versus-causation distinction and the r-versus-r² confusion, so candidates who see a strong coefficient jump to a causal claim (B) or misapply the variance-explained percentage (C).

How to eliminate wrong answers

Option A is wrong because -0.8 is a strong correlation (|r| close to 1), not weak, and the negative sign indicates an inverse relationship, not positive. Option B is wrong because correlation does not establish causation — the analyst cannot conclude that advertising spend causes churn to decrease without a controlled experiment or causal inference method. Option C is wrong because 64% (r² = 0.64) is the coefficient of determination for variance explained, but the question asks for the correct interpretation of the correlation coefficient itself, and stating '64% of variance' misapplies the statistic to the correlation value rather than describing the relationship.

69
MCQeasy

In Tableau, which of the following is used to create a calculated field that aggregates data at a different level of detail than the view?

A.LOD expressions
B.Marks
C.Parameters
D.Dashboard actions
AnswerA

LOD expressions compute aggregations at a declared dimension level, independent of the view's granularity. This satisfies the stem's constraint of aggregating at a different level of detail than the view, using FIXED, INCLUDE or EXCLUDE scopes.

Why this answer

LOD (Level of Detail) expressions in Tableau allow you to compute aggregations at a specified granularity that can be finer or coarser than the view's level of detail. They are written using curly braces with keywords FIXED, INCLUDE, or EXCLUDE, and they enable calculations like average sales per customer regardless of the dimensions in the view. This directly addresses the need to aggregate data at a different level of detail than the view.

Exam trap

The trap here is confusing LOD expressions with other Tableau features that also allow some level of control over aggregation, such as table calculations or parameters, but only LOD expressions are specifically designed to aggregate at a different level of detail than the view.

How to eliminate wrong answers

Option B is wrong because Marks are visual properties (such as color, size, shape) used to represent data points in a view, not to perform aggregations at a different level of detail. Option C is wrong because Parameters are dynamic values that can be used in calculations or filters but do not themselves change the level of detail of an aggregation. Option D is wrong because Dashboard actions are interactive elements that respond to user actions (like clicks) to filter or highlight data, not to compute aggregations at a different granularity.

70
MCQmedium

A data analyst has built a dashboard tracking monthly active users (MAU) for a mobile app over the past two years. The executive sponsor complains that the line chart is hard to read because the y-axis starts at zero, making month-to-month variation look small. The analyst wants to emphasize the trend and small fluctuations while still being honest about the data. Which change to the y-axis is most appropriate?

A.Set the y-axis minimum to zero and add a secondary axis showing percentage change.
B.Apply a logarithmic scale to the y-axis starting at one.
C.Set the y-axis minimum to a value slightly below the lowest MAU value in the period.
D.Convert the line chart to a stacked area chart with the y-axis starting at zero.
AnswerC

Truncating the y-axis to start slightly below the minimum MAU value magnifies the month-to-month variation, making the trend and fluctuations visible. This preserves the relative shape of the data and is acceptable when the audience understands the axis does not start at zero, as long as it is clearly labeled and not used to mislead.

Why this answer

Truncating the y-axis to start just below the lowest value highlights the changes the executive wants to see without falsifying the data. The other options either keep the chart flat, introduce unnecessary complexity, or use scales that obscure the trend. Clear labeling and context are essential when using a non-zero baseline.

Exam trap

The trap here is assuming that any non-zero y-axis is always misleading, when in fact a clearly labeled truncated axis can be appropriate for emphasizing variation.

71
MCQmedium

A data analyst wants to segment customers into groups based on their purchasing behavior. The dataset includes numerical features such as annual income and purchase frequency. Which algorithm is most appropriate for this task?

A.Linear regression
B.K-means clustering
C.Logistic regression
D.Chi-square test
AnswerB

K-means clustering partitions numerical observations into k groups by minimising within-cluster variance, which suits segmenting customers on annual income and purchase frequency. It satisfies the stem's unsupervised grouping requirement, unlike classification algorithms that need labelled outcomes.

Why this answer

K-means clustering is an unsupervised learning algorithm that partitions data points into k groups based on similarity of numerical features, which is exactly what customer segmentation requires. Since the task is to discover natural groupings in unlabeled data using features like annual income and purchase frequency, K-means is the standard and most appropriate choice.

Exam trap

DA0-002 often tests the confusion between supervised and unsupervised algorithms — candidates see 'segment customers' and mistakenly reach for classification (logistic regression) instead of clustering.

How to eliminate wrong answers

Option A is wrong because linear regression is a supervised algorithm that predicts a continuous target variable, not one that discovers groups in unlabeled data. Option C is wrong because logistic regression is a supervised classification algorithm that predicts a categorical outcome from labeled training data, which is not the case here. Option D is wrong because the chi-square test is a statistical hypothesis test for independence between categorical variables, not a clustering or segmentation algorithm.

72
MCQmedium

To consolidate data from multiple operational databases into a central repository for reporting, a company decides to transform data before loading it into the target system. Which data integration approach is being used?

A.ETL (Extract, Transform, Load)
B.Data virtualization
C.Change data capture
D.ELT (Extract, Load, Transform)
AnswerA

ETL transforms data in a staging area before writing to the target, so the central repository receives cleansed, conformed records. This matches the stem's constraint that data is transformed before loading, unlike ELT, which loads raw data first.

Why this answer

The scenario describes transforming data before loading it into the target system, which is the defining characteristic of ETL (Extract, Transform, Load). In ETL, data is extracted from source systems, transformed in a staging area (e.g., cleaning, aggregating, joining), and then loaded into the central repository. This approach is commonly used when the target system (e.g., a data warehouse) requires pre-processed, high-quality data for reporting.

Exam trap

The trap here is that candidates often confuse ETL with ELT, assuming that any transformation before loading is ELT, but the key distinction is that ELT loads raw data first and transforms it later inside the target system, whereas ETL transforms data before it reaches the target.

How to eliminate wrong answers

Option B (Data virtualization) is wrong because it does not physically move or transform data before loading; instead, it creates a virtual layer that queries source systems in real-time, leaving data in place. Option C (Change data capture) is wrong because it is a technique for identifying and capturing only changed data from source systems, not a complete integration approach that includes transformation before loading. Option D (ELT) is wrong because it loads raw data into the target system first and then transforms it within the target, which contradicts the 'transform before loading' requirement in the question.

73
MCQmedium

When designing a dashboard for executives, which principle is most important to follow to ensure key information is immediately visible?

A.Apply consistent color coding across all charts.
B.Create a visual hierarchy with the most important metric prominently displayed.
C.Use as many colors as possible to highlight different data points.
D.Include all available data to provide complete context.
AnswerB

A visual hierarchy directs the executive's eye to the primary KPI first, using size, position and contrast to rank information by importance. This satisfies the requirement that key information be immediately visible, rather than forcing viewers to scan equal-weight tiles to find the metric that matters most.

Why this answer

Visual hierarchy places the most important metric in the prime viewing position (typically top-left or center-top) with the largest size and strongest contrast, so an executive scanning for five seconds immediately sees the headline number. This is the defining principle of executive dashboard design because executives consume dashboards in brief, infrequent glances rather than deep analytical sessions.

Exam trap

The trap is conflating 'consistent color coding' with 'visual hierarchy' — both are dashboard best practices, but only hierarchy addresses the specific requirement of making key information immediately visible to executives.

How to eliminate wrong answers

Option A is wrong because consistent color coding is a supporting best practice for interpretability, not the primary mechanism for making key information immediately visible — it aids comprehension but doesn't establish prominence. Option C is wrong because using as many colors as possible creates visual noise and violates preattentive processing principles; color should encode meaning, not decorate. Option D is wrong because including all available data produces information overload and buries the key metric — executive dashboards should surface 3-5 KPIs, not exhaustive detail.

74
Multi-Selecthard

Which THREE of the following are NoSQL database types?

Select 3 answers
A.Document
B.Hierarchical
C.Relational
D.Key-Value
E.Graph
AnswersA, D, E

Document stores (e.g., MongoDB) are NoSQL.

Why this answer

Document databases, such as MongoDB, store data in flexible, JSON-like documents (BSON in MongoDB's case). This allows for nested structures and schema-less designs, making them a core NoSQL category distinct from relational models.

Exam trap

CompTIA often tests the distinction between legacy database models (hierarchical) and modern NoSQL categories, leading candidates to mistakenly include hierarchical as a NoSQL type due to its non-relational nature.

75
Multi-Selectmedium

Which TWO of the following are common methods for acquiring data from external sources?

Select 2 answers
A.Data warehousing
B.Manual data entry
C.Public APIs
D.Web scraping
E.Direct database connection to an internal server
AnswersC, D

Public APIs expose structured endpoints that return JSON or XML over HTTP, letting scripts pull external data directly without manual exports. This satisfies the stem's requirement for acquiring data from external sources, since the API provider hosts the data and your tooling retrieves it programmatically on demand.

Why this answer

Public APIs (C) are a standard method for acquiring external data, since they expose structured endpoints (typically REST/HTTP or GraphQL) that applications can query programmatically to pull data from third-party services. Web scraping (D) is also a common external data acquisition method, using tools such as HTTP clients and HTML parsers to extract data from public web pages when no API is available. Both are specifically designed to bring in data from outside sources, which matches the scenario.

Data warehousing (A) is a storage and analytics architecture, not an acquisition method. Manual data entry (B) is a way to input data, usually from internal or human sources, not a typical external acquisition technique. Direct database connection to an internal server (E) targets internal systems, so it is not an external data source method.

Exam trap

The trap here is that candidates may confuse data warehousing (a storage/management process) with data acquisition methods, or think manual data entry is a valid external acquisition method, when the exam specifically tests automated, programmatic techniques for pulling data from outside the organization.

Page 1 of 14

Page 2