Courseiva

CompTIA Data+ (DA0-002) (DA0-002) — Questions 376–450

1004 questions total · 14pages · All types, answers revealed

Page 5

Page 6 of 14

Page 7
376
MCQhard

A data analyst is performing a chi-square test of independence on a contingency table of customer satisfaction (satisfied vs. dissatisfied) and product type (A, B, C). The test yields a p-value of 0.04 with α = 0.05. What is the correct conclusion?

A.There is no evidence of an association between satisfaction and product type.
B.There is a significant association between satisfaction and product type.
C.The test is invalid because the expected counts are too low.
D.Satisfaction and product type are independent.
AnswerB

With p = 0.04 below the α = 0.05 threshold, the null hypothesis of independence is rejected, so satisfaction and product type are statistically associated. The chi-square test of independence detects whether the two categorical variables' observed cell frequencies deviate from those expected under independence.

Why this answer

With a p-value of 0.04 and α = 0.05, the p-value is less than the significance level, so we reject the null hypothesis of independence. This means there is statistically significant evidence of an association between customer satisfaction and product type. The correct conclusion is that a significant association exists.

Exam trap

DA0-002 often tests the interpretation of p-values versus α, and candidates frequently confuse 'fail to reject' with 'accept the null' or misinterpret a significant result as proving causation rather than association.

How to eliminate wrong answers

Option A is wrong because it states there is no evidence of association, which would be the conclusion if we failed to reject the null hypothesis (p ≥ α). Option C is wrong because the question does not provide information about expected counts; while chi-square requires expected counts ≥ 5 in each cell, the p-value alone does not indicate invalidity. Option D is wrong because it states the variables are independent, which is the null hypothesis we rejected.

377
MCQmedium

A company wants to share a dataset with external partners via an API. Which API type is typically used for web services and uses XML or JSON for messaging?

A.REST API
B.GraphQL API
C.SOAP API
D.WebSocket API
AnswerA

REST APIs use HTTP methods and URIs to expose resources, returning representations in XML or JSON as the stem requires. Their stateless, web-native design makes them the standard choice for partner-facing web service integration, satisfying both the external sharing constraint and the XML/JSON messaging requirement.

Why this answer

REST (Representational State Transfer) is an architectural style for web services that typically uses HTTP methods and supports XML or JSON for message payloads. It is the most common API type for sharing data over the web due to its simplicity, statelessness, and scalability. SOAP also uses XML but is heavier and less common for modern web APIs; GraphQL and WebSocket serve different purposes.

Exam trap

The trap is that SOAP also uses XML, so candidates may pick SOAP thinking it's the standard for web services, but the question emphasizes 'typically used for web services' and 'XML or JSON'—REST is the modern default.

How to eliminate wrong answers

Option B is wrong because GraphQL is a query language for APIs that allows clients to request specific data, but it is not typically described as using XML or JSON for messaging—it uses a single endpoint and a query language, often over HTTP with JSON responses. Option C is wrong because SOAP is a protocol that uses XML for messaging but is not the typical choice for modern web services due to its complexity and overhead. Option D is wrong because WebSocket provides full-duplex communication channels over a single TCP connection, not a request-response API style for sharing datasets.

378
MCQeasy

A dashboard designer needs to ensure that color choices are accessible to users with color vision deficiencies. Which practice should be followed?

A.Use a rainbow color palette to maximize differentiation
B.Convert all charts to grayscale
C.Use red and green to indicate positive and negative values
D.Combine color with patterns or labels to convey information
AnswerD

Colour alone fails for users with colour vision deficiencies, so redundant encoding is required. Pairing colour with patterns, labels or shapes ensures information remains readable when hues are indistinguishable, meeting accessibility requirements without discarding colour entirely.

Why this answer

Combining color with patterns or labels ensures that information is conveyed through multiple channels, not just color. This makes the visualization accessible to users with color vision deficiencies, who may not distinguish certain colors. It also benefits users in grayscale printing or low-quality displays.

This practice aligns with WCAG guidelines for using color as a supplementary, not sole, means of conveying information.

Exam trap

The trap is assuming that using a color-blind-safe palette alone is sufficient. The exam often tests that redundancy (patterns/labels) is required, not just palette choice. Candidates may pick 'use a rainbow palette' thinking it maximizes differentiation, but it actually worsens accessibility.

How to eliminate wrong answers

Option A is wrong because a rainbow palette can be problematic for color vision deficiencies and often creates false boundaries in data; it does not ensure accessibility. Option B is wrong because converting all charts to grayscale removes color entirely, which may reduce the ability to differentiate categories for users with normal vision and is not a best practice; it's an overcorrection. Option C is wrong because red and green are the most common colors confused by people with deuteranopia or protanopia, so using them to indicate positive/negative is a classic accessibility failure.

379
MCQmedium

A company wants to segment its customers into distinct groups based on purchasing behavior. Which algorithm is best suited for this task?

A.Decision tree
B.Logistic regression
C.K-means clustering
D.Linear regression
AnswerC

K-means clustering partitions unlabelled records into k groups by minimising within-cluster variance, using distance between feature vectors. It suits segmentation on purchasing behaviour, where no predefined labels exist, satisfying the requirement to form distinct customer groups.

Why this answer

K-means clustering is an unsupervised learning algorithm that partitions data into K distinct clusters based on feature similarity, making it ideal for segmenting customers by purchasing behavior without predefined labels. It groups customers who exhibit similar purchasing patterns, enabling the company to identify natural segments for targeted marketing.

Exam trap

The trap here is that candidates often confuse supervised learning algorithms (like decision trees or logistic regression) with unsupervised clustering, mistakenly thinking that any algorithm that 'groups' data can be used for segmentation without recognizing the need for unlabeled data.

How to eliminate wrong answers

Option A is wrong because a decision tree is a supervised learning algorithm used for classification or regression, requiring labeled training data to predict outcomes, not for discovering unknown groupings in unlabeled data. Option B is wrong because logistic regression is a supervised classification algorithm for binary or multinomial outcomes, relying on labeled target variables, and cannot perform unsupervised clustering. Option D is wrong because linear regression is a supervised regression algorithm that models the relationship between a dependent variable and one or more independent variables, and it is not designed to segment data into distinct groups without predefined categories.

380
Multi-Selecthard

A data analyst is designing a dashboard for executives to monitor company performance. Which THREE practices should the analyst follow to ensure effective storytelling with data? (Select three.)

Select 3 answers
A.Annotate key events on time-series charts
B.Include as much data as possible to avoid missing details
C.Always use pie charts for part-to-whole comparisons
D.Choose the right chart type for the message
E.Use a narrative arc: situation, complication, resolution
AnswersA, D, E

Annotating key events on time-series charts links metric movements to the causes behind them, such as a campaign launch or outage. Executives immediately see why performance shifted, satisfying the storytelling requirement to explain context rather than present numbers alone.

Why this answer

Option A is correct because annotating key events (such as product launches, outages, or policy changes) directly on time-series charts gives executives immediate context for spikes and dips, turning raw trends into an explanatory story rather than an unexplained line. Option D is correct because selecting the chart type that matches the intended message — for example, line charts for trends over time, bar charts for comparisons across categories, and scatter plots for correlation — ensures the visual encoding reinforces the insight instead of obscuring it. Option E is correct because structuring the dashboard around a narrative arc (situation, complication, resolution) mirrors how executives reason about performance: it establishes the baseline, highlights the deviation or problem, and points to the recommended action.

Option B is not appropriate because cramming in as much data as possible creates clutter and cognitive overload, which undermines clarity and focus on the key message. Option C is not appropriate because pie charts are generally poor for part-to-whole comparisons with many categories and are widely discouraged in favor of bar charts, so 'always' using them is not an effective storytelling practice.

Exam trap

The trap here is that candidates may think more data always leads to better insights, or that pie charts are the default for part-to-whole comparisons, but effective storytelling prioritizes clarity and context over volume or conventional but flawed visuals.

381
MCQhard

A data analyst is working with a sales table that contains columns: sale_id, product_id, sale_date, and amount. They need to calculate a 7-day moving average of sales amount for each product, ordered by sale_date. Which window function syntax should they use?

A.AVG(amount) OVER (PARTITION BY product_id ORDER BY sale_date ROWS BETWEEN 6 PRECEDING AND CURRENT ROW)
B.AVG(amount) OVER (PARTITION BY product_id ORDER BY sale_date)
C.AVG(amount) OVER (ORDER BY sale_date ROWS BETWEEN 6 PRECEDING AND CURRENT ROW)
D.SUM(amount) OVER (PARTITION BY product_id ORDER BY sale_date ROWS BETWEEN 6 PRECEDING AND CURRENT ROW)
AnswerA

ROWS BETWEEN 6 PRECEDING AND CURRENT ROW defines a seven-row frame ending at the current row, giving a 7-day moving average. PARTITION BY product_id restarts the window per product, and ORDER BY sale_date ensures chronological ordering.

Why this answer

The correct syntax uses AVG with OVER, partitioning by product_id to calculate per product, ordering by sale_date, and specifying ROWS BETWEEN 6 PRECEDING AND CURRENT ROW to include the current row and the previous six rows, yielding a 7-day moving average.

Exam trap

The trap is forgetting to partition by product_id or using the default frame, which results in a cumulative average instead of a moving average, or confusing SUM with AVG.

How to eliminate wrong answers

Option B is wrong because without a frame clause, the default frame is RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, which computes a cumulative average, not a moving average. Option C is wrong because it lacks PARTITION BY product_id, so it would compute a moving average across all products combined, not per product. Option D is wrong because it uses SUM instead of AVG, which would give a rolling sum, not an average.

382
Multi-Selectmedium

Which TWO are best practices for data storytelling?

Select 2 answers
A.Start with the conclusion
B.Use complex jargon to show expertise
C.Tailor the story to the audience
D.Use a single visualization to avoid confusion
E.Include all data points for completeness
AnswersA, C

Beginning with the conclusion respects the audience's limited attention, delivering the key insight before supporting detail. This satisfies the stem's storytelling best-practise requirement by structuring the narrative around the decision the data must inform, rather than chronological analysis order, which buries the actionable message.

Why this answer

Option A (Start with the conclusion) is correct because effective data storytelling follows the inverted-pyramid principle: leading with the key insight or recommendation ensures the audience immediately grasps the 'so what' before diving into supporting evidence, which is especially critical for time-constrained executive stakeholders. Option C (Tailor the story to the audience) is correct because the same data must be framed differently depending on the audience's technical depth, role, and decision-making needs — for example, executives need business impact and KPIs, while analysts need methodology and granular metrics — so audience adaptation is a core best practice. Option B is incorrect because complex jargon obscures meaning and alienates non-specialist stakeholders rather than demonstrating genuine expertise.

Option D is incorrect because relying on a single visualization can oversimplify or hide important patterns; best practice is to use multiple complementary charts that each serve a specific narrative purpose. Option E is incorrect because including every data point creates cognitive overload and dilutes the message — storytellers should curate only the data that supports the narrative.

Exam trap

CompTIA often tests the misconception that data storytelling should prioritize completeness or technical complexity over audience comprehension, leading candidates to select options like 'include all data points' or 'use complex jargon' instead of focusing on clarity and narrative flow.

383
Multi-Selectmedium

An analyst is presenting findings to stakeholders. Which TWO techniques effectively communicate uncertainty in data? (Choose TWO.)

Select 2 answers
A.Include confidence intervals.
B.Use only point estimates.
C.Use error bars on charts.
D.Remove all outliers from the data.
E.State exact numbers without ranges.
AnswersA, C

Confidence intervals quantify the range within which a population parameter likely falls, directly satisfying the need to communicate statistical uncertainty to stakeholders. Unlike point estimates, they expose sampling variability and precision, letting decision-makers judge reliability rather than assuming false certainty from a single figure.

Why this answer

Confidence intervals (option A) are the standard statistical technique for expressing uncertainty around an estimate, giving stakeholders a range of plausible values at a stated confidence level (e.g., 95%), which directly communicates the precision of the finding. Error bars on charts (option C) visually convey uncertainty—such as standard deviation, standard error, or confidence intervals—around plotted data points, making variability and reliability immediately apparent to a non-technical audience. Together these techniques quantify and visualize uncertainty rather than hiding it.

In contrast, using only point estimates (B) and stating exact numbers without ranges (E) falsely imply perfect precision and conceal variability, while removing all outliers (D) is a data-cleaning decision that can bias results and does not communicate uncertainty at all.

Exam trap

CompTIA often tests the distinction between measures of central tendency (point estimates) and measures of variability (confidence intervals, error bars), trapping candidates who think stating exact numbers or removing outliers is a valid way to handle uncertainty.

384
MCQeasy

Which of the following data types best describes a JSON file containing customer orders with varying fields per record?

A.Unstructured data
B.Structured data
C.Relational data
D.Semi-structured data
AnswerD

JSON stores data as nested key-value pairs without a fixed schema, so each order record can carry different fields. That schema-on-read flexibility is the defining trait of semi-structured data, distinguishing it from rigidly tabular structured data and from unstructured raw text.

Why this answer

JSON (JavaScript Object Notation) is a text-based format that uses key-value pairs and nested objects/arrays, providing a flexible schema where each record can have different fields. This self-describing structure—where data carries its own metadata via keys—is the hallmark of semi-structured data. Unlike structured data, it doesn't require a rigid, predefined schema, but unlike unstructured data, it still has an organized, machine-readable format.

Therefore, a JSON file with varying fields per record is best classified as semi-structured data.

Exam trap

The trap here is confusing semi-structured data with unstructured data because both lack a fixed schema; candidates often overlook that JSON's self-describing key-value pairs provide enough structure to separate it from truly unstructured formats like images or free text.

How to eliminate wrong answers

Option A is wrong because unstructured data (e.g., images, audio, free text) lacks any predefined data model or consistent organization, whereas JSON has a clear hierarchical structure with keys and values. Option B is wrong because structured data (e.g., relational tables) adheres to a strict, fixed schema where every record has the same fields and data types, which contradicts the varying fields described. Option C is wrong because relational data is a subset of structured data stored in tables with rows and columns and relationships defined by foreign keys; JSON is not inherently relational and does not enforce such relationships.

385
MCQeasy

A data analyst needs to present the correlation between advertising spend and website traffic. Which chart type is most appropriate?

A.Bar chart
B.Line chart
C.Pie chart
D.Scatter plot
AnswerD

A scatter plot encodes each observation as a point on two continuous axes, revealing the strength, direction, and shape of the relationship between advertising spend and website traffic. Correlation is a pairwise association, so this two-variable encoding satisfies the requirement precisely.

Why this answer

A scatter plot is the most appropriate chart type for visualizing the correlation between two continuous variables, such as advertising spend and website traffic. It displays individual data points on a Cartesian plane, allowing the analyst to assess the strength, direction, and form of the relationship (e.g., linear, non-linear, or no correlation). This aligns with the DA0-001 objective of selecting the correct visualization for bivariate analysis.

Exam trap

The trap here is that candidates often choose a line chart because they mistakenly think 'correlation' implies a trend over time, but the DA0-001 exam specifically tests that scatter plots are the standard for bivariate correlation analysis without a temporal component.

How to eliminate wrong answers

Option A is wrong because a bar chart is used to compare categorical data or discrete values, not to show the relationship between two continuous variables; it would obscure the correlation pattern. Option B is wrong because a line chart is best for displaying trends over time or sequential data, not for revealing the correlation between two independent continuous variables; it implies a temporal order that may not exist. Option C is wrong because a pie chart is designed to show proportions of a whole for categorical data, making it completely unsuitable for visualizing the correlation between two numeric variables.

386
MCQeasy

A company wants to show the number of products sold across different categories: Electronics, Clothing, Home Goods, and Books. Which chart type is most appropriate?

A.Bar chart
B.Scatter plot
C.Line chart
D.Pie chart
AnswerA

A bar chart compares discrete categorical values using rectangular bars whose lengths encode quantity, making it ideal for showing product counts across Electronics, Clothing, Home Goods, and Books. Categories sit on one axis, so relative sales volumes are compared at a glance.

Why this answer

A bar chart is ideal for comparing a quantitative measure (number of products sold) across a small number of discrete categories (Electronics, Clothing, Home Goods, Books). The length of each bar directly encodes the value, making comparisons easy and accurate.

Exam trap

The trap is confusing 'number sold' (a count) with 'proportion of total' (a part-to-whole); candidates who see categories and reach for a pie chart fall into it.

How to eliminate wrong answers

Option B is wrong because a scatter plot shows the relationship between two continuous variables, not counts across categories. Option C is wrong because a line chart implies a continuous progression (typically over time) and would falsely suggest an ordered relationship among unrelated categories. Option D is wrong because a pie chart shows parts of a whole (proportions) and becomes hard to read with more than a few slices; the question asks for counts, not proportions.

387
MCQeasy

A data analyst at a marketing firm is creating a weekly performance report for the marketing team. The report includes metrics like click-through rates, conversion rates, and cost per acquisition. The team prefers a quick overview of the week's performance to identify trends and make decisions in their Monday morning meeting. The meeting is only 30 minutes, and the team has limited time to review data. The analyst wants to provide a report that is concise and actionable. What format should the analyst use?

A.A one-page executive summary with key metrics highlighted.
B.A live dashboard with interactive filters.
C.A detailed spreadsheet with all data.
D.A 20-slide presentation.
AnswerA

A one-page executive summary surfaces click-through, conversion and cost-per-acquisition figures with highlights, letting the team absorb performance in minutes. This satisfies the stem's 30-minute Monday meeting constraint by being concise and actionable rather than a detailed multi-page report.

Why this answer

A one-page executive summary with key metrics highlighted provides the marketing team with a concise, actionable overview that can be quickly reviewed in a 30-minute meeting. This format aligns with the requirement for a quick overview to identify trends and make decisions without overwhelming the team with excessive detail or requiring interactive exploration.

Exam trap

The trap here is that candidates often confuse 'interactive' with 'efficient,' choosing a live dashboard (Option B) because it seems modern and flexible, but they overlook the specific constraint of a 30-minute meeting where pre-digested, static summaries are more actionable than tools requiring active exploration.

How to eliminate wrong answers

Option B is wrong because a live dashboard with interactive filters, while powerful for ad-hoc analysis, requires time to explore and manipulate, which is not suitable for a quick 30-minute meeting where the team needs a pre-digested overview. Option C is wrong because a detailed spreadsheet with all data presents raw, unsummarized information that would take too long to parse and interpret, defeating the goal of conciseness and quick trend identification. Option D is wrong because a 20-slide presentation is too lengthy and detailed for a 30-minute meeting, likely leading to information overload and insufficient time for discussion and decision-making.

388
Multi-Selectmedium

A data analyst is creating a dashboard for a sales team. Which TWO actions adhere to dashboard design best practices?

Select 2 answers
A.Including 3D effects to make charts look more professional
B.Using consistent color coding across charts
C.Adding detailed axis labels and titles to every chart
D.Placing the most important metric at the top-left or most prominent position
E.Using as many bright colors as possible to make the dashboard attractive
AnswersB, D

Consistent colour coding lets viewers map a single hue to one meaning across every chart, so the sales team reads categories without rechecking legends. This satisfies the stem's design best-practice constraint by reducing cognitive load and preventing misinterpretation, which matters when non-technical users scan the dashboard quickly.

Why this answer

Option B is correct because consistent color coding across charts lets viewers instantly associate a color with the same category or metric everywhere on the dashboard, reducing cognitive load and preventing misinterpretation. Option D is correct because placing the most important metric at the top-left or most prominent position follows the natural reading pattern (F-pattern/Z-pattern) and ensures the key insight is seen first, which is a core dashboard design best practice. Option A is wrong because 3D effects distort data perception and add visual clutter rather than improving professionalism.

Option C is wrong because detailed axis labels and titles on every chart create redundancy and clutter; concise, clear labels are preferred, with detail available via tooltips or drill-downs. Option E is wrong because using many bright colors overwhelms viewers and destroys the meaningful, consistent color encoding that dashboards require.

389
MCQhard

A data team is using web scraping to collect competitor pricing data. The target website has anti-scraping measures like CAPTCHAs and rate limiting. Which approach is most effective?

A.Use a single IP address
B.Disregard robots.txt
C.Use rotating proxies and respectful delays
D.Increase request frequency
AnswerC

Rotating proxies distribute requests across many IP addresses, defeating IP-based rate limiting, while respectful delays reduce request frequency to avoid triggering CAPTCHAs and detection. Together they sustain collection without overwhelming the target site or breaching its defensive thresholds.

Why this answer

Using rotating proxies and respectful delays is the most effective approach because it distributes requests across multiple IP addresses, avoiding rate limiting and IP bans, while delays reduce the load on the target server and mimic human behavior, helping to bypass anti-scraping measures like CAPTCHAs and rate limiting.

Exam trap

DA0-002 often tests the misconception that increasing request frequency or using a single IP can overcome anti-scraping measures, when in fact these approaches worsen blocking; the correct strategy involves rotation and throttling.

How to eliminate wrong answers

Option A is wrong because using a single IP address makes the scraper easily detectable and rate-limited or blocked, as all requests originate from one source. Option B is wrong because disregarding robots.txt is unethical and may lead to legal issues, and it does not help bypass technical anti-scraping measures; it can also result in IP bans. Option D is wrong because increasing request frequency exacerbates rate limiting and triggers more aggressive anti-scraping responses, making the scraper less effective.

390
MCQhard

A data analyst creates a bubble chart showing country GDP (x-axis), life expectancy (y-axis), and population (bubble size). However, large bubbles overlap and obscure many data points. Which corrective action should the analyst take?

A.Increase the chart canvas size
B.Set bubble opacity to 70%
C.Reduce all bubble sizes uniformly
D.Remove outlier countries with large populations
AnswerB

Setting bubble opacity to 70% lets overlapping marks remain partially visible, so obscured data points stay readable while bubble size still encodes population. This directly addresses the stem's constraint that large bubbles overlap and hide points, without altering the GDP or life expectancy axes.

Why this answer

Setting bubble opacity to 70% allows overlapping bubbles to become semi-transparent, so data points underneath remain visible. This technique preserves the original data representation (GDP, life expectancy, and population) without altering the chart's scale or removing data. It is a standard visualization practice for handling overplotting in dense scatter plots and bubble charts.

Exam trap

The trap here is that candidates often choose to reduce bubble sizes uniformly (Option C) thinking it solves overlap, but this distorts the proportional encoding of population, whereas opacity preserves the original data relationships while improving visibility.

How to eliminate wrong answers

Option A is wrong because increasing the canvas size does not resolve the fundamental issue of overlapping bubbles; it only spreads them out slightly, and large bubbles will still obscure others if their sizes are disproportionate. Option C is wrong because uniformly reducing all bubble sizes changes the visual encoding of population, potentially making small populations invisible and distorting the data's relative comparison. Option D is wrong because removing outlier countries with large populations eliminates valid data points, which introduces bias and violates the principle of representing the full dataset; the goal is to visualize all data, not discard it.

391
MCQeasy

Which of the following best describes a data mart?

A.A repository for raw, unprocessed data
B.An OLTP system for transaction processing
C.A subject-specific subset of a data warehouse
D.A tool for extract, transform, and load processes
AnswerC

A data mart is a subject-specific subset of a data warehouse, scoped to one department or business function. It inherits warehouse data but serves a narrower analytical audience, distinguishing it from the enterprise-wide warehouse itself.

Why this answer

A data mart is a subject-specific subset of a data warehouse, designed to serve the analytical needs of a particular department or business function (e.g., sales, finance, marketing). It contains a focused set of data extracted from the enterprise data warehouse or other sources, optimized for query and reporting by a specific user group. This aligns with the definition of a data mart as a smaller, more specialized version of a data warehouse.

Exam trap

The trap here is confusing data marts with data lakes or ETL tools, as all are related to data warehousing but serve different purposes; candidates might incorrectly choose 'raw, unprocessed data' (data lake) or 'ETL tool' due to familiarity with those terms.

How to eliminate wrong answers

Option A is wrong because a repository for raw, unprocessed data describes a data lake or staging area, not a data mart; data marts contain curated, transformed data for analysis. Option B is wrong because an OLTP system is designed for transactional processing (inserts, updates, deletes) and is not optimized for analytical queries, whereas a data mart is an analytical construct. Option D is wrong because ETL is a process or tool used to move and transform data, not a data mart itself; a data mart is the target repository, not the tool.

392
Multi-Selectmedium

Which THREE elements should be included in a data insight report to ensure it is actionable? (Choose three.)

Select 3 answers
A.The level of confidence or statistical significance.
B.The source code of the analysis scripts.
C.All raw data used in the analysis.
D.A clear recommendation based on the insight.
E.An estimate of the potential business impact.
AnswersA, D, E

Confidence helps stakeholders assess reliability.

Why this answer

An actionable data insight report must include the level of confidence or statistical significance to allow decision-makers to assess the reliability of the findings. Without this, stakeholders cannot determine whether the observed patterns are likely to be real or due to random chance, which is critical for making data-driven decisions.

Exam trap

CompTIA often tests the distinction between technical artifacts (like source code or raw data) and actionable business insights, so candidates mistakenly include all supporting materials instead of focusing on elements that directly drive decision-making.

393
Multi-Selectmedium

A data analyst is preparing a dataset for analysis and needs to ensure data quality. Which TWO of the following are dimensions of data quality?

Select 2 answers
A.Volume
B.Velocity
C.Consistency
D.Variety
E.Accuracy
AnswersC, E

Consistency is a recognised data quality dimension, confirming that values remain uniform across systems and records without contradictory entries. It satisfies the stem's requirement for genuine quality dimensions, alongside accuracy, completeness, timeliness and validity, so it qualifies as one of the two correct selections.

Why this answer

Consistency (C) is a core dimension of data quality because it ensures that the same data values are represented uniformly across different datasets, systems, or records, preventing contradictions that would corrupt analysis results. Accuracy (E) is also a fundamental data quality dimension, as it verifies that data correctly reflects the real-world entities or events it is meant to describe, which is essential for trustworthy analysis. In contrast, Volume (A), Velocity (B), and Variety (D) are the three defining characteristics of big data (the '3 Vs'), describing the scale, speed, and diversity of data rather than its quality.

Therefore, only Consistency and Accuracy belong to the set of data quality dimensions.

Exam trap

DA0-002 often mixes big-data Vs (volume, velocity, variety) with data quality dimensions (accuracy, consistency, completeness), so candidates must distinguish characteristics of data from measures of its quality.

394
MCQmedium

A data analyst wants to test if the proportion of customers who prefer Product A over Product B is different from 50%. She surveys 200 customers and finds that 120 prefer Product A. Which statistical test should she use?

A.Chi-square test of independence
B.One-sample z-test for proportions
C.ANOVA
D.Two-sample t-test
AnswerB

The scenario compares one observed sample proportion against a hypothesised population proportion of 0.50, with a large sample of 200. A one-sample z-test for proportions is the appropriate parametric test for this single-proportion hypothesis, unlike chi-square or two-sample alternatives.

Why this answer

The analyst wants to test whether the proportion of customers preferring Product A differs from 50%, using a single sample of 200 customers with 120 preferring A. This is a one-sample test of a proportion against a hypothesized value (0.5), so the one-sample z-test for proportions is appropriate. It compares the observed proportion (0.60) to the null hypothesis proportion (0.50).

Exam trap

DA0-002 often tests test selection by scenario; candidates confuse proportion tests with chi-square or t-tests, especially when the word 'prefer' suggests a comparison between two products rather than a single proportion against a benchmark.

How to eliminate wrong answers

Option A is wrong because the chi-square test of independence compares two categorical variables across a contingency table, not a single proportion against a hypothesized value. Option C is wrong because ANOVA compares means across three or more groups, not proportions. Option D is wrong because a two-sample t-test compares means of two independent groups, not a single proportion to a target.

395
MCQeasy

Which of the following is an example of semi-structured data?

A.A CSV file without header
B.An image file
C.A table in a relational database
D.A JSON file
AnswerD

Semi-structured data carries organisational tags or keys but lacks a rigid relational schema. A JSON file uses nested key-value pairs and arrays, fitting this definition, whereas tables and CSV files are structured and free text or images are unstructured.

Why this answer

Semi-structured data has tags or markers to separate data elements, like JSON or XML.

396
MCQmedium

A data analyst is building a dashboard that includes a KPI card showing the current customer satisfaction (CSAT) score. The target is 90%. The analyst wants the card to immediately convey whether the current score is meeting, approaching, or below target. Which design approach best achieves this?

A.Display the CSAT score with conditional formatting: green if >= 90%, yellow if 85-89.9%, and red if < 85%.
B.Display the CSAT score with a comparison to the previous period, such as an up or down arrow with percentage change.
C.Display the CSAT score as a gauge chart with a needle pointing to the current value on a 0-100% scale.
D.Display the CSAT score in a large font with a small line chart showing the trend over the past 30 days.
AnswerA

Conditional formatting with color thresholds directly maps the score to performance status. Green, yellow, and red instantly communicate whether the score is meeting, approaching, or below target. This is a best practice for KPI cards because it reduces cognitive load and supports rapid decision-making. The thresholds should be clearly defined and accessible.

Why this answer

Conditional formatting with color thresholds immediately signals performance status relative to the target. It leverages pre-attentive processing, allowing viewers to grasp the situation without mental calculation. The other options provide context or trends but fail to deliver an instant, unambiguous status indicator, which is the core need for a KPI card.

Exam trap

The trap here is equating trend indicators or gauges with status indicators; only explicit threshold-based formatting provides immediate performance context.

397
MCQeasy

A data analyst needs to count the number of customers who have placed at least one order. Which SQL query should be used?

A.SELECT DISTINCT COUNT(customer_id) FROM orders
B.SELECT SUM(customer_id) FROM orders
C.SELECT COUNT(customer_id) FROM orders
D.SELECT COUNT(DISTINCT customer_id) FROM orders
AnswerD

COUNT(DISTINCT customer_id) eliminates duplicate customer identifiers before counting, so each customer contributes exactly one to the total regardless of how many orders they placed. This directly satisfies the stem's requirement to count customers with at least one order, rather than counting order rows themselves.

Why this answer

COUNT(DISTINCT customer_id) counts each unique customer exactly once, which is precisely what is needed to answer 'how many customers placed at least one order' — a customer with multiple orders is counted only once. This is the standard SQL pattern for counting distinct entities in a fact table.

Exam trap

DA0-002 often tests the placement of DISTINCT inside the COUNT function — candidates who write SELECT DISTINCT COUNT(...) believe they are deduplicating, but DISTINCT applies to the aggregate result, not the input rows, so the count remains inflated.

How to eliminate wrong answers

Option A is wrong because SELECT DISTINCT COUNT(customer_id) applies DISTINCT to the result of COUNT(), which is already a single scalar value — it does not deduplicate customer_id before counting, so it returns the same result as COUNT(customer_id), i.e., the total number of order rows, not unique customers. Option B is wrong because SUM(customer_id) adds up the numeric values of customer_id, which is meaningless for counting customers and produces a nonsensical large number. Option C is wrong because COUNT(customer_id) counts every non-NULL order row, so a customer with five orders is counted five times, overstating the number of distinct customers.

398
Multi-Selectmedium

A data analyst is assessing the quality of a newly acquired dataset from an external source. The analyst needs to evaluate two dimensions of data quality that directly affect the dataset's fitness for use in analysis. Which two dimensions should the analyst prioritize? (Choose two.)

Select 2 answers
A.Completeness
B.Accessibility
C.Portability
D.Scalability
E.Accuracy
AnswersA, E

Completeness measures the extent to which all required data is present. Missing values can lead to biased or invalid analysis results. In an externally acquired dataset, completeness is critical because missing data may indicate collection issues or gaps that affect the reliability of conclusions. Prioritizing completeness ensures that analyses are based on a full picture.

Why this answer

Completeness and accuracy are fundamental data quality dimensions that directly affect the validity of analysis. Completeness ensures no critical data is missing, while accuracy ensures the data correctly represents reality. These two dimensions are essential when evaluating an external dataset because they determine whether the data can be trusted for decision-making.

Other dimensions like accessibility, portability, and scalability are more about operational aspects.

Exam trap

The trap here is selecting operational or technical dimensions like accessibility or scalability instead of core data quality dimensions that directly impact analytical results.

399
MCQmedium

A company uses a data warehouse for reporting. They need to extract data from multiple sources, load it into a staging area, and then transform it before moving to the warehouse. This process is known as:

A.ELT
B.ETL
C.Data replication
D.Data ingestion
AnswerB

ETL extracts from multiple sources, loads into a staging area, then transforms before loading into the warehouse, matching the stem's stated sequence exactly. ELT would transform after loading, so it does not fit the described staging-then-transform order.

Why this answer

The process described—extracting data from multiple sources, loading it into a staging area, and then transforming it before moving to the warehouse—is the classic definition of ETL (Extract, Transform, Load). In ETL, transformation occurs after extraction but before loading into the target system, which is exactly what the staging area is used for. This contrasts with ELT, where transformation happens after loading into the warehouse.

Exam trap

The trap here is that candidates confuse the order of operations in ETL versus ELT, assuming that because modern cloud warehouses support ELT, the described staging-area process must be ELT, when in fact the staging area is a hallmark of traditional ETL.

How to eliminate wrong answers

Option A is wrong because ELT (Extract, Load, Transform) loads raw data into the target system first and transforms it later, which is the opposite of the described sequence where transformation occurs before moving to the warehouse. Option C is wrong because data replication refers to copying data from one system to another for redundancy or availability, not a multi-stage pipeline with transformation. Option D is wrong because data ingestion is a broad term covering the initial import of data into a system, but it does not specifically include the staging and transformation steps described in the question.

400
MCQmedium

A data analyst is exploring a dataset of 8,000 customer transactions and notices that the 'transaction_amount' column has a mean of $120 but a median of $45. A small number of transactions exceed $10,000. Which measure of central tendency should the analyst report to describe the typical transaction?

A.The range, because it shows the spread between the smallest and largest transactions
B.The mean, because it uses every value in the dataset
C.The mode, because it identifies the most frequently occurring transaction amount
D.The median, because it is resistant to the influence of extreme values
AnswerD

The median is the middle value when data is ordered, so a handful of transactions above $10,000 do not shift it. With a mean far above the median, the distribution is right-skewed, and the median better represents the typical transaction amount experienced by most customers.

Why this answer

A large gap between the mean and median signals skew, and here the high-value transactions drag the mean well above the middle of the distribution. The median marks the point where half the transactions are smaller and half larger, so it reflects the typical customer experience far better. For skewed monetary data, the median is the more representative measure of central tendency.

Exam trap

The trap here is defaulting to the mean as the summary of a numeric column without checking whether extreme values have distorted it.

401
MCQhard

Refer to the exhibit. A data analyst is creating a report that includes customer transaction data from 6 years ago. According to the policy, what should the analyst do?

A.Anonymize the data before inclusion
B.Flag the data for review
C.Exclude the data because it exceeds the retention period
D.Include the data since it is valuable for analysis
AnswerC

The retention policy caps how long customer transaction data may be kept, and six-year-old records exceed that limit. Excluding them complies with the stated retention period, avoiding the regulatory and privacy breach that retaining or reporting expired personal data would create.

Why this answer

The data retention policy specifies that customer transaction data must be retained for only 5 years. Since the data is from 6 years ago, it exceeds the retention period and must be excluded from the report to comply with data governance and regulatory requirements. Including or modifying such data would violate policy and potentially expose the organization to legal or compliance risks.

Exam trap

The trap here is that candidates may assume data can be retained or modified (e.g., anonymized) if it is valuable for analysis, but the policy strictly prohibits using data beyond its retention period, regardless of its potential value or transformation.

How to eliminate wrong answers

Option A is wrong because anonymizing the data does not address the policy violation; the data has already exceeded the retention period and should not be used at all, regardless of anonymization. Option B is wrong because flagging the data for review implies it might still be used after evaluation, but the policy is clear that data beyond the retention period must be excluded, not reviewed for potential inclusion. Option D is wrong because including the data for its analytical value directly violates the retention policy, which prioritizes compliance over data utility.

402
MCQmedium

A retail company wants to predict sales based on advertising spend and season. Which data modeling technique should the analyst use?

A.Simple linear regression
B.Multiple linear regression
C.Logistic regression
D.K-means clustering
AnswerB

Multiple linear regression models a continuous outcome, sales, as a linear function of two or more predictors, advertising spend and season. Season enters as a categorical dummy variable, satisfying the requirement to predict sales from both numeric and categorical inputs.

Why this answer

Multiple linear regression is the correct technique because the analyst needs to model a continuous outcome (sales) based on two or more predictor variables: advertising spend (continuous) and season (categorical, typically encoded as dummy variables). This allows the model to capture the independent effect of each predictor on sales, which simple linear regression cannot do because it only handles one predictor.

Exam trap

The trap here is that candidates often confuse simple linear regression with multiple linear regression, thinking that 'linear regression' alone suffices, but the exam specifically tests whether you recognize that multiple predictors require multiple regression.

How to eliminate wrong answers

Option A is wrong because simple linear regression can only model the relationship between one independent variable and the dependent variable, but here we have two predictors (advertising spend and season). Option C is wrong because logistic regression is used for binary or categorical outcome variables (e.g., yes/no), not for continuous outcomes like sales. Option D is wrong because K-means clustering is an unsupervised learning technique used to group similar data points, not to predict a continuous target variable.

403
MCQmedium

A healthcare organization maintains a database of patient records. The database has a table 'patients' with columns: patient_id (primary key), first_name, last_name, date_of_birth, gender, and last_visit_date. A data analyst is tasked with creating a report that lists all patients who have not visited in the last two years. The analyst writes a query: SELECT * FROM patients WHERE last_visit_date < DATEADD(year, -2, GETDATE()); However, the query returns zero rows, even though the analyst knows there are patients who have not visited for over two years. Upon inspection, the analyst discovers that the last_visit_date column contains NULL values for patients who have never visited. Which modification to the query should the analyst make to include patients with NULL last_visit_date?

A.Remove the WHERE clause entirely.
B.Add OR last_visit_date IS NULL to the WHERE clause.
C.Use COALESCE(last_visit_date, '1900-01-01') in the WHERE clause.
D.Add AND last_visit_date IS NOT NULL to the WHERE clause.
AnswerB

SQL's three-valued logic evaluates NULL comparisons as UNKNOWN, so last_visit_date < DATEADD(...) never matches NULL rows. Adding OR last_visit_date IS NULL explicitly includes patients who have never visited, satisfying the report's requirement to list all inactive patients.

Why this answer

The original query uses a WHERE clause that compares last_visit_date to a computed date, but NULL comparisons in SQL always yield UNKNOWN, so rows with NULL last_visit_date are excluded. Adding OR last_visit_date IS NULL explicitly includes those rows, ensuring patients who have never visited are listed in the report.

Exam trap

The trap here is that candidates often forget that NULL comparisons in SQL do not return TRUE, leading them to incorrectly think the original query already handles NULLs, and they may choose Option C (COALESCE) as a workaround instead of the simpler and correct IS NULL check.

How to eliminate wrong answers

Option A is wrong because removing the WHERE clause entirely would return all rows, including those with recent visits, which fails to filter for patients who have not visited in two years. Option C is wrong because COALESCE(last_visit_date, '1900-01-01') would replace NULL with a very old date, making the comparison work, but it is not the standard or most efficient approach; the correct method is to use IS NULL to handle NULLs directly. Option D is wrong because AND last_visit_date IS NOT NULL would explicitly exclude rows with NULL last_visit_date, which is the opposite of what is needed.

404
Multi-Selectmedium

A data team is preparing data for a clustering analysis. Which THREE of the following steps are commonly part of data cleaning?

Select 3 answers
A.Removing duplicate records
B.Imputing missing values
C.Calculating the mean
D.Training a regression model
E.Capping outliers at the 5th and 95th percentiles
AnswersA, B, E

Removing duplicate records eliminates redundant observations that would otherwise distort distance calculations, causing clustering algorithms to over-weight repeated points. This directly satisfies the stem's data-cleaning requirement by ensuring each entity contributes once, preventing artificial density concentrations that skew centroid placement and cluster assignment.

Why this answer

Option A (Removing duplicate records) is correct because duplicate rows distort distance calculations in clustering, causing the same observation to be counted multiple times and biasing cluster centroids, so deduplication is a standard data-cleaning step. Option B (Imputing missing values) is correct because clustering algorithms such as k-means cannot handle nulls, so missing entries must be filled via mean/median/mode imputation, k-NN, or similar methods before analysis. Option E (Capping outliers at the 5th and 95th percentiles) is correct because winsorizing extreme values limits their disproportionate influence on distance metrics and centroid placement, which is a recognized cleaning technique.

Option C (Calculating the mean) is not a cleaning step but a descriptive statistic or profiling operation, and Option D (Training a regression model) is a modeling task, not data preparation, so neither belongs to data cleaning.

Exam trap

The trap here is confusing data cleaning with data exploration or modeling — candidates see 'calculating the mean' and think it's part of preparation because it's a common early step, but it doesn't clean anything.

405
Matchingmedium

Match each data analysis tool to its primary function.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Query and manipulate structured data in databases

General-purpose language for data analysis and modeling

Statistical computing and graphics

Interactive data visualization and dashboards

Spreadsheet for data manipulation and basic analysis

Why these pairings

Excel is for spreadsheets, SQL for databases, Python for programming/ML, and Tableau for visualization. Common confusions include assigning predictive modeling to Excel or dashboard creation to SQL.

406
MCQeasy

A data analyst is creating a report that will be printed in black and white for a quarterly business review. The report includes a bar chart comparing revenue across five product categories. To ensure the chart is interpretable without color, which design element should the analyst prioritize?

A.Use a different shade of gray for each bar and include a legend.
B.Add data labels directly on each bar showing the revenue value.
C.Sort the bars from highest to lowest revenue and omit the legend.
D.Use a distinct pattern (e.g., solid, striped, dotted) for each bar and include a legend or direct labels.
AnswerD

Patterns provide a non-color visual cue that remains clear in black and white, even with poor print quality. Combining patterns with a legend or direct labels ensures the viewer can map each bar to its category. This approach directly addresses the need for interpretability without relying on color, making it the most robust choice for a printed report.

Why this answer

In a black-and-white print environment, color cannot be used to distinguish categories. Patterns are a reliable non-color encoding that works even with low-quality printing. Pairing patterns with a legend or direct labels ensures the viewer can correctly identify each product category, satisfying the requirement for interpretability.

Exam trap

The trap here is assuming that grayscale shades are sufficient for differentiation in print, when patterns are more robust against poor contrast and copying.

407
MCQmedium

A financial analyst wants to create a dashboard that shows the monthly profit and loss, highlighting how each component (revenue, cost of goods sold, operating expenses) contributes to the final net profit. Which chart type is most appropriate?

A.Pie chart
B.Stacked bar chart
C.Area chart
D.Waterfall chart
AnswerD

A waterfall chart shows how sequentially added or subtracted components bridge an opening balance to a closing total, so revenue, cost of goods sold and operating expenses visibly build to net profit. This matches the requirement to highlight each component's contribution rather than just the final figure.

Why this answer

A waterfall chart is purpose-built to show how an initial value is incrementally increased or decreased by a series of positive and negative contributors to arrive at a final total. For a P&L, it starts at revenue, subtracts COGS and operating expenses as downward steps, and lands on net profit — making each component's contribution visually explicit. This is exactly the 'bridge' or 'walk' visualization finance teams use for variance and P&L analysis.

Exam trap

DA0-002 often tests the confusion between charts that show composition (pie, stacked bar) and charts that show sequential contribution to a final value (waterfall), so candidates who see 'components contributing to net profit' and reach for a pie or stacked bar fall into the trap.

How to eliminate wrong answers

Option A is wrong because a pie chart shows parts of a single whole at one point in time and cannot represent sequential additions and subtractions leading to a net result, nor can it show negative values. Option B is wrong because a stacked bar chart shows composition of totals across categories but does not visually bridge from a starting value to an ending value through intermediate gains and losses. Option C is wrong because an area chart emphasizes cumulative magnitude and trends over time, not the discrete step-by-step contributions that reconcile revenue down to net profit.

408
MCQhard

In Power BI, an analyst wants to create a measure that calculates the total sales for the current year up to the latest date in the data. Which DAX function should be used?

A.SAMEPERIODLASTYEAR
B.CALCULATE
C.SUMX
D.TOTALYTD
AnswerD

TOTALYTD aggregates a measure from the start of the year to the latest date present in the data, matching the stem's requirement for year-to-date sales. Unlike DATESYTD, which returns a date table, TOTALYTD evaluates the expression directly, producing the cumulative sales figure the analyst needs.

Why this answer

TOTALYTD is a time intelligence function that sums values for the year up to the last date in the filter context (or specified end date).

409
MCQeasy

A data analyst needs to retrieve all unique job titles from an employees table. Which SQL keyword should be used in the SELECT clause?

A.UNIQUE
B.REMOVE DUPLICATES
C.DISTINCT
D.FILTER
AnswerC

DISTINCT eliminates duplicate rows from the result set, returning only unique job titles as the stem requires. Applied directly in the SELECT clause, it collapses repeated values across the specified column, satisfying the uniqueness constraint without needing GROUP BY or aggregate functions.

Why this answer

The DISTINCT keyword in a SELECT clause eliminates duplicate rows from the result set, returning only unique combinations of the selected columns. When applied to a single column like job_title, it returns each job title exactly once, which is precisely what the analyst needs. DISTINCT operates on the entire row of selected columns, so SELECT DISTINCT job_title FROM employees yields the unique list of titles.

Exam trap

The trap here is confusing the DDL constraint UNIQUE with the DML query keyword DISTINCT — candidates who have seen UNIQUE in CREATE TABLE statements may reflexively choose it for a SELECT query.

How to eliminate wrong answers

Option A is wrong because UNIQUE is not a SQL SELECT keyword — it is a constraint used in DDL (CREATE TABLE / ALTER TABLE) to enforce uniqueness on a column, not a query modifier. Option B is wrong because 'REMOVE DUPLICATES' is not valid SQL syntax in any major RDBMS (Oracle, SQL Server, PostgreSQL, MySQL); it is a conceptual description, not a keyword. Option D is wrong because FILTER is used in PostgreSQL as a clause on aggregate functions (e.g., COUNT(*) FILTER (WHERE ...)) and in some engines as a WHERE-like construct, but it does not deduplicate rows.

410
MCQeasy

A healthcare organization collects patient questionnaire data via paper forms at clinics. The forms are scanned and sent to a central office, where staff manually enter data into an electronic system. This process is slow and error-prone. The organization wants to reduce manual entry errors and speed up data availability. Which method should they adopt?

A.Continue manual entry but double-check all entries
B.Use optical character recognition (OCR) to digitize the forms and automatically populate the database
C.Send forms to an external data processing company
D.Require patients to fill out forms online at home
AnswerB

OCR converts scanned glyph images into machine-encoded text, then template or field mapping populates the database directly, eliminating the transcription step where staff introduce keystroke errors. This satisfies both stated constraints: fewer manual entry errors and faster data availability, since digitisation occurs at scan time rather than awaiting central-office keying.

Why this answer

Using optical character recognition (OCR) to digitize the paper forms and automatically populate the database directly addresses the goal of reducing manual entry errors and speeding up data availability. OCR converts scanned images of text into machine-readable data, eliminating the need for manual transcription and enabling faster processing. This method is well-suited for structured forms like patient questionnaires.

Exam trap

The trap here is assuming that any automation (like outsourcing or online forms) solves the problem, but the question specifically asks for reducing manual entry errors and speeding data availability, which OCR directly addresses.

How to eliminate wrong answers

Option A is wrong because continuing manual entry with double-checking still relies on human transcription, which remains slow and error-prone. Option C is wrong because sending forms to an external data processing company may reduce internal effort but still involves manual entry (unless they use OCR), introduces data privacy risks, and does not inherently speed up availability. Option D is wrong because requiring patients to fill out forms online at home may not be feasible for all patients and does not address the existing paper-based workflow; it also shifts the burden to patients and may not integrate with current systems.

411
MCQmedium

A data scientist is performing K-means clustering on customer data. She plots the within-cluster sum of squares (WCSS) for different values of k and observes an 'elbow' at k=4. What does this indicate?

A.The optimal number of clusters is 4
B.The algorithm should be run with k=3 to avoid overfitting
C.The data contains exactly 4 outliers
D.The WCSS is minimized at k=4, indicating perfect clustering
AnswerA

The elbow marks where adding clusters stops yielding meaningful WCSS reduction, so k=4 balances model complexity against fit. Beyond four, each extra cluster captures only marginal variance, indicating diminishing returns. This inflection point therefore identifies four as the optimal cluster count for the customer dataset.

Why this answer

The elbow method suggests that adding more clusters beyond k=4 yields diminishing returns, so k=4 is a suitable number of clusters.

412
MCQeasy

Which of the following is a key difference between a Key Performance Indicator (KPI) and a metric?

A.Metrics are always quantitative, while KPIs can be qualitative
B.There is no difference; the terms are interchangeable
C.KPIs are tied to strategic objectives, while metrics are broader operational measurements
D.KPIs are always lagging indicators, while metrics are leading indicators
AnswerC

KPIs measure progress toward specific strategic goals, whereas metrics are any quantifiable operational measurements. This axis of difference distinguishes the two, since a metric only becomes a KPI when it is explicitly linked to a business objective.

Why this answer

A KPI is a metric that is explicitly tied to a strategic business objective — it measures progress toward a goal the organization cares about (e.g., customer retention rate). A metric is any quantifiable measurement (e.g., page views, CPU utilization) that may or may not be strategically important. The key distinction is the linkage to strategy, not the data type or leading/lagging nature.

Exam trap

DA0-002 often tests the misconception that KPIs and metrics differ by data type (quantitative vs. qualitative) or by leading/lagging status — the real differentiator is strategic alignment, and candidates who overlook that pick the wrong option.

How to eliminate wrong answers

Option A is wrong because both metrics and KPIs are quantitative by definition — a KPI is a specific type of metric, not a qualitative statement. Option B is wrong because the terms are not interchangeable; conflating them leads to 'vanity metrics' that look good but do not reflect strategic progress. Option D is wrong because KPIs can be either leading or lagging indicators — for example, 'number of qualified leads' is a leading KPI, while 'quarterly revenue' is a lagging KPI; the leading/lagging distinction is orthogonal to the metric/KPI distinction.

413
Multi-Selecteasy

Which TWO of the following are characteristics of a data lake?

Select 2 answers
A.Retains raw data in native format
B.Optimized for OLTP
C.Stores only structured data
D.Enforces ACID transactions
E.Uses schema-on-read
AnswersA, E

Data lakes store data as-is without transformation.

Why this answer

A data lake retains raw data in its native format, meaning data is ingested without transformation or schema enforcement. This allows storage of structured, semi-structured, and unstructured data as-is, preserving fidelity for future analytics. Unlike a data warehouse, a data lake does not require upfront schema definition, enabling flexible exploration and machine learning workloads.

Exam trap

The trap here is that candidates confuse data lakes with data warehouses, assuming all enterprise data stores enforce ACID and schema-on-write, when in fact data lakes prioritize raw storage and schema flexibility.

414
Multi-Selectmedium

Which TWO of the following are characteristics of structured data? (Choose TWO.)

Select 2 answers
A.Has a defined schema
B.Requires NoSQL databases for storage
C.Often contains natural language text
D.Cannot be queried using SQL
E.Organized in rows and columns
AnswersA, E

Schema defines structure.

Why this answer

Structured data is defined by having a predefined schema, which specifies the data types, constraints, and relationships for each field. This schema ensures consistency and allows for efficient querying and validation. Option A is correct because a defined schema is a fundamental characteristic of structured data, as seen in relational database tables where each column has a specific data type and constraints.

Exam trap

The trap here is that candidates often confuse structured data with semi-structured data (e.g., JSON or XML) and incorrectly assume that structured data cannot be queried with SQL or that it requires NoSQL databases.

415
MCQmedium

A data analyst is evaluating a multiple regression model with three predictors. The R² value is 0.85. Which of the following is the best interpretation of R²?

A.85% of the variance in the outcome is explained by the predictors.
B.85% of the predicted values are correct.
C.The model has a high bias.
D.The model has a strong correlation of 0.85.
AnswerA

R² measures the proportion of total variance in the dependent variable accounted for by the three predictors collectively. A value of 0.85 therefore means 85% of outcome variance is explained, with the remaining 15% attributable to other factors or error.

Why this answer

R² (coefficient of determination) measures the proportion of variance in the dependent variable explained by the independent variables in the model. An R² of 0.85 means 85% of the variability in the outcome is accounted for by the three predictors, with the remaining 15% attributable to factors not in the model or random noise. This is a goodness-of-fit measure, not an accuracy percentage or correlation coefficient.

Exam trap

DA0-002 often tests the misconception that R² is a percentage of correct predictions or a correlation value — candidates confuse goodness-of-fit with accuracy or with Pearson's r.

How to eliminate wrong answers

Option B is wrong because R² does not measure the percentage of correct predictions — that would be classification accuracy or a metric like MAPE/RMSE for regression, and R² can even be negative for poorly fit models. Option C is wrong because high bias refers to underfitting (a model too simple to capture the pattern), which is unrelated to the R² value itself; a high R² typically indicates low bias. Option D is wrong because R² is not a correlation coefficient — the correlation r would be the square root of R² (≈0.92) only in simple linear regression, and even then R² measures explained variance, not correlation strength.

416
MCQhard

A financial institution wants to analyze transaction networks to detect fraud rings. Which database type is best suited for this analysis?

A.Wide-column store
B.Graph database
C.Key-value store
D.Document store
AnswerB

Graph databases store entities as nodes and relationships as edges, so traversing transaction links between accounts is a native pointer hop rather than an expensive join. This directly satisfies the fraud-ring detection requirement, where identifying circular money flows depends on multi-hop relationship traversal.

Why this answer

A graph database is designed to store and traverse relationships between entities, making it ideal for analyzing transaction networks where connections between accounts, merchants, and transactions reveal fraud rings. Its native graph model (nodes and edges) allows efficient pattern matching and pathfinding queries, such as detecting circular transactions or shared attributes, which are common in fraud detection.

Exam trap

CompTIA often tests the misconception that any NoSQL database can handle relationship-heavy workloads, but the trap here is that only graph databases are purpose-built for deep relationship traversal and pattern matching, while other NoSQL types sacrifice relationship performance for scalability or flexibility.

How to eliminate wrong answers

Option A is wrong because wide-column stores (e.g., Cassandra, HBase) are optimized for high-volume, low-latency reads/writes on sparse data with flexible schemas, but they lack native relationship traversal capabilities, making multi-hop queries across transaction networks slow and complex. Option C is wrong because key-value stores (e.g., Redis, DynamoDB) provide fast lookups by primary key but cannot efficiently model or query the interconnected relationships between transactions and entities, requiring application-level joins that degrade performance. Option D is wrong because document stores (e.g., MongoDB, Couchbase) store semi-structured data as JSON-like documents and support indexing, but they do not have built-in graph traversal algorithms, so analyzing fraud rings would require expensive recursive queries or external graph processing.

417
MCQhard

An analyst is profiling a table of laboratory test results. The ResultValue column stores numeric readings, but a subset of rows contains the literal text "N/A" in that column, and a separate ResultUnit column records units such as mg/dL or mmol/L. The analyst must compare average ResultValue across two hospital sites. Which data issue most directly prevents a valid comparison, and what is the appropriate first step?

A.The average must be computed as a median because laboratory results are always skewed, so the analyst should replace ResultValue with a rank-based statistic
B.The presence of "N/A" text in a numeric column makes the column string-typed, so the analyst must drop the entire ResultValue column before comparing sites
C.The ResultUnit column contains multiple unit systems, so the analyst must convert all readings to a single unit before averaging
D.The "N/A" text forces the column to be interpreted as text, so numeric averages silently ignore or misorder values, and the analyst must convert or exclude those values before averaging
AnswerD

When a numeric column contains text sentinels, many tools coerce the whole column to a string type, so AVG either fails or sorts lexicographically, producing meaningless results. The analyst must identify the non-numeric rows, decide whether to convert them to nulls or exclude them, and then compute the average on a clean numeric type. Only then is the two-site comparison trustworthy.

Why this answer

A numeric column polluted with text values like "N/A" typically gets coerced to a string type, which breaks arithmetic aggregation and can cause silent misordering or failed averages. The analyst must first isolate the non-numeric rows and convert or exclude them so the column is genuinely numeric, then compute the average for each site. Only after that remediation does the two-site comparison become valid.

Exam trap

The trap here is assuming an average still works on a column that contains text, when the presence of a non-numeric value can change the column's type and silently corrupt the aggregation.

418
MCQhard

A data analyst is troubleshooting a report that shows unusually high sales for a specific product. Upon investigation, the analyst finds that the product was returned by several customers, but the returns were recorded in a separate system and not reflected in the sales data. Which data integration concept was likely missing?

A.ETL (Extract, Transform, Load)
B.Data reconciliation
C.Data profiling
D.Data governance
AnswerB

Sales figures were never adjusted for returns held in a separate system, so the two datasets disagreed. Data reconciliation compares and aligns records across sources to detect and correct such mismatches, which would have surfaced the unreflected returns before the report ran.

Why this answer

The core issue is that the sales data and returns data are inconsistent because they were not cross-verified. Data reconciliation is the process of comparing datasets to ensure they are in agreement and identifying discrepancies, such as returns not being reflected in sales figures. Without reconciliation, the analyst would not detect that the high sales number is inflated by unrecorded returns.

Exam trap

The trap here is that candidates confuse the data movement process (ETL) with the data validation process (reconciliation), assuming that simply extracting and loading data will automatically ensure consistency between separate systems.

How to eliminate wrong answers

Option A is wrong because ETL (Extract, Transform, Load) is a process for moving and transforming data from source to target systems, but it does not inherently include a step to compare or verify data consistency between separate systems; the missing concept here is not about data movement but about data agreement. Option C is wrong because data profiling focuses on examining data quality, structure, and content (e.g., nulls, duplicates, data types), not on cross-system consistency checks; the problem is not about the quality of the sales data itself but about its mismatch with returns data. Option D is wrong because data governance refers to the overall management of data availability, usability, integrity, and security through policies and standards, not a specific technical process for reconciling discrepancies between two systems.

419
MCQmedium

A business needs to store large volumes of raw data in its native format for future analytics. Which storage architecture is most appropriate?

A.Relational database
B.Data lake
C.Operational data store
D.Data warehouse
AnswerB

A data lake stores raw data in its native format, such as JSON, Parquet or CSV, without requiring a predefined schema. This satisfies the requirement to retain large volumes of unprocessed data for future, undetermined analytics workloads.

Why this answer

A data lake is designed to store large volumes of raw data in its native format (structured, semi-structured, or unstructured) without requiring a predefined schema. This makes it ideal for future analytics where the data schema may not yet be known, as it supports schema-on-read rather than schema-on-write.

Exam trap

The trap here is that candidates confuse a data warehouse with a data lake, assuming both are for analytics, but the key differentiator is that a data warehouse requires schema-on-write and processed data, while a data lake stores raw data in native format.

How to eliminate wrong answers

Option A is wrong because a relational database enforces a strict schema-on-write and is optimized for transactional processing (OLTP), not for storing raw, unprocessed data at scale. Option C is wrong because an operational data store (ODS) is used for integrating data from multiple operational systems for near-real-time reporting, not for storing raw data in native format for future analytics. Option D is wrong because a data warehouse stores cleansed, transformed, and structured data optimized for query performance and business intelligence, not raw data in its native format.

420
MCQmedium

A data analyst is using pandas in Python to merge two DataFrames: sales (columns: sale_id, product_id, amount) and products (columns: product_id, product_name). Which pandas function should they use to combine these DataFrames on the 'product_id' column?

A.combine()
B.merge()
C.join()
D.concat()
AnswerB

merge() performs a database-style join on a shared key, so passing on='product_id' combines sales and products into one DataFrame with product_name attached to each sale. concat() only stacks frames, and join() defaults to index alignment, neither matching this key-based requirement.

Why this answer

The pandas merge function is used to combine DataFrames on common columns. The syntax is pd.merge(sales, products, on='product_id').

421
MCQhard

You are a data analyst for an e-commerce company. Your team has built a dashboard to monitor daily sales performance across five regions: North, South, East, West, and Central. The dashboard includes a bar chart showing total sales per region, a line chart showing daily sales trend over the past 30 days, and a pie chart showing sales distribution by product category (Clothing, Electronics, Home, Books, Sports). Recently, stakeholders have complained that the pie chart is hard to interpret because the Sports category has very small sales and is barely visible. Also, the bar chart uses a rainbow color scheme that makes it difficult to compare bar heights because the colors are not ordered by magnitude. The line chart is fine. You need to redesign the dashboard to address these issues. Which combination of changes is most appropriate?

A.Replace the pie chart with a stacked bar chart and use a categorical color scheme for the bar chart
B.Explode the Sports slice in the pie chart and use a monochromatic color scheme for the bar chart
C.Change the pie chart to a 3D pie chart and use a diverging color scheme for the bar chart
D.Group small categories into an 'Other' slice in the pie chart and use a sequential color scheme ordered by sales for the bar chart
AnswerD

Grouping tiny categories into 'Other' makes the Sports slice legible, while a sequential scheme ordered by sales encodes magnitude in colour, letting viewers rank bar heights accurately. Both changes directly fix the reported readability and comparison problems.

Why this answer

Grouping small categories into an 'Other' slice directly addresses the pie chart's readability issue by consolidating negligible values, and using a sequential color scheme ordered by sales for the bar chart improves the ability to compare bar heights by encoding magnitude through color intensity. This combination follows best practices for data visualization: avoid cluttering with tiny slices and use ordered, perceptually uniform colors to facilitate accurate comparisons.

Exam trap

CompTIA often tests the misconception that simply highlighting or separating a small slice (exploding or 3D) fixes pie chart readability, when in fact it does not address the fundamental issue of angle comparison for tiny values.

How to eliminate wrong answers

Option A is wrong because replacing the pie chart with a stacked bar chart does not solve the problem of a barely visible category; it may still compress small values into thin segments, and a categorical color scheme for the bar chart does not order colors by magnitude, leaving the comparison of bar heights difficult. Option B is wrong because exploding the Sports slice in the pie chart only draws attention to it without improving the overall readability of the pie chart for small slices, and a monochromatic color scheme for the bar chart lacks the ordered intensity needed to compare bar heights effectively. Option C is wrong because a 3D pie chart distorts proportions and makes interpretation even harder, and a diverging color scheme is designed for data with a meaningful midpoint (e.g., positive/negative values), not for ordering bars by magnitude.

422
MCQeasy

A market research firm collects survey responses where customers rate satisfaction on a scale of 'Very Unsatisfied', 'Unsatisfied', 'Neutral', 'Satisfied', 'Very Satisfied'. What type of data is being collected?

A.Interval
B.Ordinal
C.Ratio
D.Nominal
AnswerB

Ordinal data has a meaningful order but unequal intervals between categories. The satisfaction scale runs from Very Unsatisfied to Very Satisfied in ranked sequence, yet the gap between adjacent labels is not numerically defined, ruling out interval or nominal classification.

Why this answer

The data is ordinal because the satisfaction levels have a clear, ordered ranking from 'Very Unsatisfied' to 'Very Satisfied', but the intervals between categories are not necessarily equal. This type of categorical data preserves the order without assuming a consistent numerical difference between each level.

Exam trap

The trap here is that candidates mistakenly treat ordered categorical data as interval data because they assume the numeric labels (e.g., 1 to 5) imply equal spacing, but the exam expects you to recognize that the underlying measurement scale lacks guaranteed equal intervals.

How to eliminate wrong answers

Option A is wrong because interval data requires equal, measurable intervals between values (e.g., temperature in Celsius), but the satisfaction scale does not guarantee equal psychological distance between categories. Option C is wrong because ratio data requires a true, meaningful zero point (e.g., income, height), and 'Very Unsatisfied' does not represent an absolute absence of satisfaction. Option D is wrong because nominal data is unordered categorical data (e.g., colors, gender), but the satisfaction scale has a natural order that must be preserved.

423
Multi-Selecteasy

Which TWO of the following are dimensional modeling techniques commonly used in data warehouses?

Select 2 answers
A.Entity-relationship diagram
B.Snowflake schema
C.Star schema
D.Scatter plot
E.Histogram
AnswersB, C

Snowflake schema normalises dimension tables into multiple related tables, reducing redundancy and storage at the cost of more joins. It is a recognised dimensional modelling technique, alongside star schema, used to structure data warehouse dimensions.

Why this answer

Snowflake schema (B) is a dimensional modeling technique in which dimension tables are normalized into multiple related tables, reducing redundancy while preserving the fact-dimension structure of a data warehouse. Star schema (C) is the classic dimensional modeling technique where a central fact table joins directly to denormalized dimension tables, optimizing query performance for analytical workloads. Both are standard approaches described by Kimball for organizing data marts and warehouses around facts and dimensions.

Entity-relationship diagram (A) is a conceptual/logical modeling notation for OLTP-style normalized databases, not a dimensional technique. Scatter plot (D) and histogram (E) are data visualization or statistical analysis tools, not data warehouse modeling techniques.

Exam trap

The trap here is that candidates may confuse general data modeling concepts (like ERDs) or data visualization tools (like scatter plots and histograms) with specific dimensional modeling techniques used in data warehouses.

424
MCQhard

A data scientist trains a regression model and observes high variance with low bias. Which technique is most appropriate to reduce variance?

A.Apply Ridge regularization
B.Increase polynomial features
C.Use a smaller training set
D.Remove correlated features
AnswerA

High variance with low bias indicates overfitting, where coefficients fit training noise. Ridge adds an L2 penalty on squared coefficient magnitudes, shrinking them and constraining model complexity, which reduces variance at the cost of slightly increased bias.

Why this answer

Ridge regularization (L2) reduces variance by adding a penalty term proportional to the square of the coefficients, which shrinks them toward zero without eliminating them. This directly addresses high variance (overfitting) by constraining the model's complexity, while low bias indicates the model fits the training data well. The regularization parameter λ controls the trade-off between bias and variance.

Exam trap

CompTIA often tests the misconception that reducing variance requires removing features or simplifying the model, but Ridge regularization is the correct technique because it penalizes coefficient magnitude without discarding predictors.

How to eliminate wrong answers

Option B is wrong because increasing polynomial features adds higher-order terms, which increases model complexity and typically increases variance, not reduces it. Option C is wrong because using a smaller training set reduces the amount of data available for learning, which generally increases variance due to less stable coefficient estimates. Option D is wrong because removing correlated features can reduce multicollinearity but does not directly penalize coefficient magnitudes; it may even increase variance if important predictors are dropped.

425
MCQhard

A data architect is designing a system to store customer support tickets. The tickets are written in free-form text and include attachments such as screenshots and PDFs. The system must allow support agents to search for tickets by keywords within the text and attachments. The architect expects the volume of tickets to grow to millions and requires fast, full-text search capabilities. Which storage solution is most appropriate?

A.Relational database with BLOB columns
B.Document store with inverted index
C.Graph database with full-text search plugin
D.Key-value store with secondary indexes
AnswerB

A document store with an inverted index, such as Elasticsearch, is optimized for full-text search across large volumes of text and metadata. It can index the content of tickets and extract text from attachments (via ingest pipelines) to enable keyword search. This architecture scales horizontally and provides fast, relevant search results, directly addressing the requirements.

Why this answer

A document store with an inverted index is purpose-built for full-text search, enabling fast keyword queries across millions of text documents and attachments. It supports text extraction from various file types and scales horizontally. Relational, key-value, and graph databases lack the specialized indexing and search features needed for efficient full-text retrieval at this scale.

Exam trap

The trap here is assuming that any database with indexing can handle full-text search, but only systems with an inverted index provide the tokenization and ranking required for efficient keyword queries.

426
Multi-Selecthard

A data analyst is communicating insights about a sales forecast to stakeholders. Which three of the following should the analyst include to build trust and clarity? (Select THREE.)

Select 3 answers
A.Only the most optimistic scenario.
B.The raw data used for the forecast.
C.The confidence intervals around the forecast.
D.The assumptions made in the forecast model.
E.A discussion of potential risks and uncertainties.
AnswersC, D, E

Confidence intervals quantify the uncertainty around the forecast, directly satisfying the need for clarity about prediction reliability. Stakeholders can judge the plausible range of sales outcomes rather than treating a single point estimate as certain, which builds trust through honest disclosure of model precision and variability.

Why this answer

Option C is correct because confidence intervals quantify the range within which the true forecast value is likely to fall, giving stakeholders an honest measure of the forecast's precision rather than a single misleading point estimate. Option D is correct because stating the assumptions made in the forecast model (e.g., growth rates, seasonality, market conditions) lets stakeholders judge whether those premises match reality and understand how the results were derived. Option E is correct because discussing potential risks and uncertainties prepares stakeholders for scenarios where the forecast may not hold, which builds credibility and supports better contingency planning.

Option A does not belong because presenting only the most optimistic scenario hides downside risk and biases decisions, undermining trust. Option B does not belong because dumping raw data on stakeholders adds noise without interpretation; the analyst should present summarized, relevant evidence instead of the full raw dataset.

Exam trap

CompTIA often tests the distinction between transparency and information overload, so the trap here is that candidates think sharing raw data (Option B) is always good practice, but in stakeholder communication, raw data without context or summary statistics can confuse rather than clarify.

427
Multi-Selectmedium

A data analyst frequently receives ad hoc requests for the same type of analysis. Which TWO approaches could reduce the number of ad hoc requests?

Select 2 answers
A.Increase data freshness to real-time
B.Create a scheduled report that covers the common analysis
C.Add more security to the data
D.Ignore the requests until they become urgent
E.Encourage users to create their own reports using a self-service BI tool
AnswersB, E

A scheduled report delivers the recurring analysis automatically at set intervals, so requesters self-serve instead of raising tickets. This satisfies the goal of reducing repeated ad hoc requests by converting a predictable, recurring need into a standing deliverable.

Why this answer

Option B is correct because a scheduled report that covers the common analysis proactively delivers the recurring results the analyst keeps being asked for, eliminating the need for users to submit one-off requests each time. Option E is correct because a self-service BI tool lets business users build and run their own reports against governed data, shifting routine ad hoc demand away from the analyst. Option A is not correct because increasing data freshness to real-time addresses latency, not the volume of repetitive requests.

Option C is not correct because adding security controls does not reduce how often users ask for the same analysis. Option D is not correct because ignoring requests is unprofessional and does nothing to eliminate the underlying recurring demand.

Exam trap

DA0-002 often tests the difference between reactive and proactive approaches, causing candidates to choose options that do not actually reduce request volume (like increasing data freshness or security) instead of automation and self-service.

428
MCQmedium

An executive dashboard must display high-level KPIs such as current revenue, profit margin, and customer count. Which visualization type is most appropriate for each KPI?

A.Pie chart
B.Sparkline for each KPI
C.KPI card showing value and variance
D.Gauge chart
AnswerC

A KPI card displays a single aggregated value with its variance against target, giving executives immediate status without requiring interpretation of axes or trends. This satisfies the requirement for high-level revenue, margin and customer-count indicators on a dashboard.

Why this answer

A KPI card showing value and variance is the most appropriate visualization for executive dashboards displaying high-level KPIs such as current revenue, profit margin, and customer count. KPI cards present a single, clear metric with its current value and comparison to a target or previous period (variance), enabling executives to quickly assess performance. This format avoids the clutter of charts and focuses attention on the key number.

Exam trap

DA0-002 often tests the distinction between KPI cards (value + variance) and other visualizations like gauges or sparklines, tempting candidates to choose a more 'visual' chart when a simple card is the best practice for executive KPIs.

How to eliminate wrong answers

Option A is wrong because a pie chart is used to show parts of a whole (proportions) and is not suitable for displaying a single KPI value with variance. Option B is wrong because a sparkline shows a trend over time but lacks the current value and variance context needed for executive KPIs; it is better for supplementary trend indication. Option D is wrong because a gauge chart displays a value within a range but does not typically show variance against a target and can be harder to read for multiple KPIs on a dashboard.

429
MCQhard

A data analyst is creating a self-service reporting environment. Which data governance practice ensures users see only data relevant to their department?

A.Data lineage
B.Row-level security
C.Data dictionary
D.Single version of truth
AnswerB

Row-level security filters individual rows by predicate, so each user's query returns only records matching their department. Unlike object-level permissions, which grant whole tables, RLS enforces the departmental constraint inside the query itself, satisfying the self-service requirement without duplicating datasets.

Why this answer

Row-level security (RLS) is a data governance practice that restricts data access at the row level based on user attributes such as department, role, or region. It ensures that users see only the data relevant to their department by dynamically filtering rows according to predefined policies. This is implemented in many BI and database platforms (e.g., Power BI, Tableau, Snowflake) and directly addresses the requirement of self-service reporting with department-specific data visibility.

Exam trap

The trap here is confusing data governance practices that manage metadata or consistency (like data lineage or single version of truth) with those that enforce access control (like row-level security). Candidates might pick 'data dictionary' thinking it restricts access, but it only documents data.

How to eliminate wrong answers

Option A is wrong because data lineage tracks the origin and transformation of data but does not control access. Option C is wrong because a data dictionary documents metadata (definitions, relationships) but does not enforce row-level access. Option D is wrong because a single version of truth ensures consistency across the organization but does not restrict data visibility by department.

430
Multi-Selectmedium

An analyst is creating a report in Power BI and needs to calculate year-to-date total sales compared to the same period last year. Which TWO DAX functions should be used? (Choose two.)

Select 2 answers
A.FILTER
B.RELATED
C.SAMEPERIODLASTYEAR
D.SUMX
E.TOTALYTD
AnswersC, E

SAMEPERIODLASTYEAR returns the equivalent date range one year prior, shifting the current filter context backwards. Combined with TOTALYTD, it supplies the prior-year comparison for the year-to-date sales measure, satisfying the requirement to compare against the same period last year.

Why this answer

TOTALYTD is correct because it is a time-intelligence function that evaluates an expression over the year-to-date interval, e.g., TOTALYTD(SUM(Sales[Amount]), 'Date'[Date]), which directly produces the year-to-date total sales the analyst needs. SAMEPERIODLASTYEAR is correct because it is a time-intelligence function that shifts the current date context back one year, returning the equivalent period from the prior year so the YTD figure can be compared to the same period last year. Together they satisfy the two comparison requirements in the scenario.

FILTER is not a time-intelligence function; it returns a filtered table and would not by itself compute YTD or prior-year periods. RELATED is used to fetch a value from the many-side of a relationship, not for time comparisons. SUMX is an iterator that sums an expression row by row and does not handle year-to-date or prior-year logic.

Exam trap

The trap is confusing time intelligence functions with general aggregation functions like SUMX or FILTER, which do not handle date shifts.

431
MCQhard

A data scientist is building a machine learning model to predict customer churn. The dataset includes both numerical features (age, income) and categorical features (gender, marital status). Which data concept describes the process of converting categorical features into numerical values that can be used by the algorithm?

A.Data sampling
B.Encoding
C.Feature scaling
D.Dimensionality reduction
AnswerB

Encoding maps categorical values such as gender and marital status into numeric representations, for example one-hot or ordinal vectors, which the algorithm can process. Numerical features like age and income need no conversion, so encoding is the concept that addresses the categorical constraint.

Why this answer

Encoding is the correct data concept because it transforms categorical features (like gender and marital status) into numerical representations (e.g., one-hot encoding, label encoding) that machine learning algorithms can process. Unlike feature scaling or dimensionality reduction, encoding directly addresses the incompatibility of non-numeric data with mathematical model operations.

Exam trap

CompTIA often tests the distinction between encoding and feature scaling, where candidates mistakenly think scaling applies to categorical data, but scaling only adjusts numeric ranges and cannot convert text labels to numbers.

How to eliminate wrong answers

Option A is wrong because data sampling refers to selecting a subset of data for training/testing, not converting categorical data to numeric. Option C is wrong because feature scaling normalizes numerical ranges (e.g., via min-max scaling or z-score standardization) and does not handle categorical-to-numeric conversion. Option D is wrong because dimensionality reduction (e.g., PCA, t-SNE) reduces the number of features, but it assumes all input features are already numeric and does not address the encoding of categorical variables.

432
Multi-Selecthard

A logistic regression model predicts customer churn (0=no churn, 1=churn). The model outputs probabilities. Which THREE of the following statements about logistic regression are correct?

Select 3 answers
A.The model output is a probability between 0 and 1.
B.The coefficient of determination R² is used to assess model fit.
C.The coefficients represent the change in log-odds for a one-unit change in the predictor.
D.Logistic regression is used for binary classification.
E.The model uses the linear regression equation y = mx + b directly.
AnswersA, C, D

The sigmoid function maps any linear combination of predictors onto the interval between 0 and 1, so each output is interpretable as the estimated probability of churn. A threshold, commonly 0.5, then converts that probability into a class label.

Why this answer

Option A is correct because logistic regression applies the sigmoid (logistic) function to a linear combination of predictors, producing an output strictly between 0 and 1 that can be interpreted as the probability of the positive class (churn = 1). Option C is correct because the model is linear in the log-odds: each coefficient βj gives the change in log-odds of the outcome for a one-unit increase in predictor Xj, holding other predictors constant (equivalently, e^βj is the odds ratio). Option D is correct because logistic regression is specifically designed for binary classification, here distinguishing churn (1) from no churn (0) by thresholding the predicted probability.

Option B is not correct because R² (coefficient of determination) is a goodness-of-fit measure for ordinary least squares linear regression, not for logistic regression, which instead uses measures like log-likelihood, deviance, AIC/BIC, or pseudo-R². Option E is not correct because logistic regression does not use the linear equation y = mx + b directly; it models the log-odds as a linear function and then applies the logistic function to obtain probabilities, rather than predicting y linearly.

Exam trap

DA0-002 often tests logistic regression by mixing in linear regression concepts — candidates who assume R² applies or that the linear equation is used directly pick the wrong statements, missing that logistic regression uses log-odds and the sigmoid function.

433
MCQeasy

A data analyst is creating a visualization to show the proportion of total sales contributed by each of five product categories. The analyst wants to emphasize the relative size of each category as part of a whole. Which type of chart is most appropriate?

A.Histogram
B.Line chart
C.Scatter plot
D.Pie chart
AnswerD

A pie chart is designed to show parts of a whole, with each slice representing a proportion of the total. It is ideal for displaying the relative contribution of a small number of categories, such as five product categories, to a total. It visually emphasizes the size of each category in relation to the whole.

Why this answer

A pie chart is the most suitable because it visually represents each category as a slice of the whole, making it easy to compare proportions. Line charts, scatter plots, and histograms serve different purposes and do not effectively communicate part-to-whole relationships.

Exam trap

The trap here is confusing part-to-whole visualization with other chart types; pie charts are specifically designed for proportions, but they are often misused for too many categories.

434
MCQeasy

A data analyst is performing data profiling on a customer table. Which metric would best help identify missing values in the 'phone' column?

A.Cardinality
B.Null count
C.Mean
D.Row count
AnswerB

Null count directly quantifies absent entries in the phone column, satisfying the profiling goal of identifying missing values. Unlike distinct count or data type checks, it measures completeness per attribute, exposing the exact volume of nulls requiring remediation before analysis.

Why this answer

The null count metric directly measures the number of missing (NULL) values in a column, which is exactly what the analyst needs to identify missing phone numbers. Data profiling tools report null count per column as a standard completeness metric. Other metrics like cardinality or mean do not reveal missingness.

Exam trap

The trap is confusing cardinality with completeness — candidates see 'cardinality' and think it measures how many values are present, but it actually measures distinct values, not missing ones.

How to eliminate wrong answers

Option A is wrong because cardinality measures the number of distinct values in a column, which tells you about uniqueness, not missingness. Option C is wrong because mean is an arithmetic average of numeric values and is undefined or meaningless for a phone column, and it does not indicate missing values. Option D is wrong because row count gives the total number of rows in the table, not the number of missing values in a specific column.

435
MCQhard

A DBA wants to improve query performance on a large table that is frequently filtered on two columns: department_id and hire_date. The table has millions of rows. Which index strategy would be most effective?

A.Create a composite B-tree index on (department_id, hire_date)
B.Create a bitmap index on hire_date
C.Create a hash index on department_id only
D.Create two separate B-tree indexes, one on each column
AnswerA

A composite B-tree index on (department_id, hire_date) lets the optimiser seek directly to a department and then range-scan hire_date within it, satisfying both filter predicates in one index. Separate single-column indexes would require bitmap merges or scans, which scale poorly across millions of rows.

Why this answer

A composite B-tree index on (department_id, hire_date) is most effective because it allows the database to satisfy equality and range predicates on both columns in a single index scan. B-tree indexes are optimized for high-cardinality columns and support efficient multi-column filtering when the leading column matches the query's equality condition, followed by the range condition on hire_date.

Exam trap

The trap here is that candidates often assume two separate single-column indexes are equivalent to a composite index, but they fail to realize that the database cannot efficiently combine them for range predicates without a costly index merge operation.

How to eliminate wrong answers

Option B is wrong because bitmap indexes are designed for low-cardinality columns (e.g., gender or status) and perform poorly with high-cardinality columns like hire_date, leading to excessive bitmap merge overhead and poor query performance. Option C is wrong because a hash index on department_id only supports equality lookups, not range queries on hire_date, and cannot be used for filtering on both columns simultaneously. Option D is wrong because two separate B-tree indexes would force the optimizer to choose one index and then filter the other column via a table access (or perform an expensive index merge), which is less efficient than a single composite index that can directly satisfy both predicates.

436
Multi-Selecthard

A data analyst is performing a chi-square test for independence between two categorical variables. Which THREE of the following are necessary conditions for the test to be valid?

Select 3 answers
A.Variances are equal across groups
B.Data is normally distributed
C.Sample is randomly selected
D.Observations are independent
E.Expected frequency in each cell is at least 5
AnswersC, D, E

Random selection ensures the sample represents the population, so observed cell counts estimate population proportions without selection bias. This satisfies the chi-square validity condition that expected frequencies reflect genuine population distributions rather than a biased subset.

Why this answer

Option C is correct because a chi-square test for independence requires that the sample be randomly selected from the population, ensuring the results can be generalized and that the expected counts reflect the underlying distribution. Option D is correct because the test assumes observations are independent; each subject or case must contribute to only one cell of the contingency table, and correlated or repeated observations violate the chi-square model. Option E is correct because the chi-square approximation is valid only when the expected frequency in each cell is at least 5 (or, in larger tables, when no more than 20% of cells have expected counts below 5 and none below 1).

Option A is not required because chi-square is a nonparametric test of frequencies and does not assume homogeneity of variances. Option B is not required because chi-square does not assume normally distributed data; it operates on counts of categorical outcomes.

Exam trap

DA0-002 often tests the confusion between parametric assumptions (normality, equal variance) and chi-square's non-parametric requirements — candidates incorrectly apply t-test or ANOVA assumptions to chi-square.

437
MCQeasy

A retail company wants to analyze monthly sales data over the past three years to identify long-term trends. Which component of time series analysis is most relevant for this goal?

A.Irregular component
B.Cyclical component
C.Seasonality
D.Trend
AnswerD

Trend captures the long-term direction of a series after removing seasonal and irregular fluctuations, which directly matches the three-year monthly sales goal of identifying sustained movement rather than repeating yearly patterns. Decomposition isolates this component, so analysts can quantify whether sales are genuinely rising or falling across the full period.

Why this answer

The trend component of a time series represents the long-term direction or underlying movement of the data over an extended period, which is exactly what the retail company wants to identify across three years of monthly sales. Trend captures gradual increases, decreases, or stagnation that persist beyond short-term fluctuations. Seasonality and cyclical components are periodic but shorter-term or irregular in period, while the irregular component is random noise.

Exam trap

DA0-002 often tests the confusion between trend and cyclical components — candidates pick 'cyclical' because it sounds long-term, but trend specifically refers to the persistent direction, while cycles are irregular economic waves.

How to eliminate wrong answers

Option A is wrong because the irregular component represents random, unpredictable fluctuations (e.g., a one-off supply chain disruption) that obscure rather than reveal long-term patterns. Option B is wrong because the cyclical component refers to wave-like movements tied to economic cycles (typically 2-10 years) that are not fixed in period — while relevant to long-term analysis, the question emphasizes identifying the underlying direction, which is trend. Option C is wrong because seasonality captures repeating patterns within a fixed period (e.g., monthly or quarterly), which is short-term and would mask the long-term trend if not decomposed.

438
Multi-Selecthard

A company is planning an A/B test to compare two website designs. Which THREE of the following must be determined before the test begins to ensure valid results? (Select three.)

Select 3 answers
A.The desired effect size
B.The p-value of the test
C.Which hypothesis is true
D.The minimum sample size required
E.The significance level (α)
AnswersA, D, E

The desired effect size defines the smallest lift worth detecting, which drives sample-size and power calculations. Specifying it before testing satisfies the stem's validity requirement, preventing an underpowered test or a post-hoc redefinition of success.

Why this answer

Option A (the desired effect size) must be determined beforehand because it defines the minimum difference between the two designs that the test should be able to detect, and it directly drives the sample-size calculation; without it, the test may be underpowered or impractically large. Option D (the minimum sample size required) is essential because it ensures the test has adequate statistical power to detect the specified effect, preventing premature conclusions from too little data. Option E (the significance level, α) must be set in advance because it defines the threshold for rejecting the null hypothesis (commonly 0.05) and controls the false-positive rate, keeping the analysis objective.

Option B (the p-value) is not determined before the test; it is an output computed from the collected data. Option C (which hypothesis is true) cannot be known in advance—it is precisely what the A/B test is designed to infer, so it is not a pre-test determination.

Exam trap

DA0-002 often tests the confusion between pre-test design parameters (effect size, sample size, α) and post-test outputs (p-value, hypothesis truth).

439
MCQeasy

A marketing analyst wants to segment customers based on their purchase history, including total spent, number of transactions, and average order value. The analyst runs k-means clustering with k=5 on the raw data but notices that the cluster assignments change significantly every time the algorithm is executed. What should the analyst do first to obtain consistent and meaningful clusters?

A.Normalize the features and set a fixed random seed for the initial centroids.
B.Switch to hierarchical clustering, which does not require specifying k.
C.Increase the number of clusters to k=10 to capture more detail.
D.Use principal component analysis (PCA) to reduce the number of features to two.
AnswerA

Features on different scales distort Euclidean distance, so k-means centroids shift with each random initialisation. Normalising the features and fixing the random seed stabilises initial centroid selection, satisfying the requirement for consistent, meaningful cluster assignments across runs.

Why this answer

The instability in cluster assignments is caused by the algorithm's sensitivity to the scale of features and the random initialization of centroids. Normalizing the features ensures that each variable contributes equally to the distance calculations, while setting a fixed random seed makes the initial centroid selection deterministic, leading to reproducible results.

Exam trap

The trap here is that candidates may think the instability is due to the choice of k or the algorithm itself, rather than recognizing that k-means is sensitive to feature scaling and random initialization, which are the first things to address for consistency.

How to eliminate wrong answers

Option B is wrong because hierarchical clustering does not require specifying k, but it still suffers from sensitivity to data scaling and does not address the core issue of random initialization causing variability. Option C is wrong because increasing k to 10 would likely increase instability and overfit noise, not resolve the fundamental problem of non-deterministic centroids. Option D is wrong because PCA reduces dimensionality but does not stabilize the k-means algorithm; the cluster assignments would still vary with different random seeds unless combined with normalization and a fixed seed.

440
MCQeasy

Refer to the exhibit. Which clause is used to aggregate the data by department?

A.HAVING
B.WHERE
C.ORDER BY
D.GROUP BY
AnswerD

GROUP BY partitions rows sharing the same department value into groups, enabling aggregate functions such as COUNT, SUM or AVG to be computed per department. It is the clause that performs the grouping the question requires.

Why this answer

The GROUP BY clause is used to aggregate data by department because it groups rows that have the same values in the specified column(s), allowing aggregate functions like SUM, AVG, or COUNT to be applied per group. In SQL, without GROUP BY, aggregate functions would operate on the entire result set, not per department.

Exam trap

CompTIA often tests the distinction between WHERE (row-level filter) and HAVING (group-level filter), leading candidates to confuse HAVING with GROUP BY when the question asks for the clause that performs aggregation.

How to eliminate wrong answers

Option A is wrong because HAVING is used to filter groups after aggregation, not to define the grouping itself. Option B is wrong because WHERE filters individual rows before aggregation and cannot group data by department. Option C is wrong because ORDER BY sorts the result set but does not perform any aggregation or grouping.

441
MCQmedium

An analyst wants to visualize the relationship between advertising spend (x-axis) and revenue (y-axis) for 100 different products. Each product is in one of three categories. Which chart type best displays this data?

A.Scatter plot with points colored by category
B.Bubble chart
C.Stacked bar chart
D.Line chart with three lines
AnswerA

A scatter plot places advertising spend and revenue on two numeric axes, revealing correlation and outliers across 100 products. Colouring points by category adds a third categorical dimension, exposing whether relationships differ between the three groups — something a line or bar chart cannot convey.

Why this answer

A scatter plot with color-coded categories effectively shows relationships between two continuous variables and a third categorical dimension.

442
MCQmedium

A data analyst creates a weekly KPI dashboard for executives. The analyst notes that the data is updated as of the previous day. Which report quality element should be included?

A.Data dictionary
B.Data lineage
C.Row-level security
D.Data freshness timestamp
AnswerD

A data freshness timestamp states exactly when the underlying data was last refreshed, so executives know the dashboard reflects the previous day. This satisfies the stem's stated lag and prevents stale figures being mistaken for current ones.

Why this answer

When a dashboard shows data as of the previous day, the report must communicate that staleness to executives. A data freshness timestamp explicitly states when the underlying data was last refreshed, so consumers know the currency of the metrics. This is a core report quality element for trust and decision-making.

Exam trap

DA0-002 often tests the confusion between data lineage (where data came from) and data freshness (how current it is), since both are metadata about the data.

How to eliminate wrong answers

Option A is wrong because a data dictionary documents field definitions and metadata, not the recency of the data shown. Option B is wrong because data lineage traces the origin and transformation path of data, which supports auditability but does not indicate how current the dashboard is. Option C is wrong because row-level security restricts which rows a user can see based on permissions, which is an access control concern, not a freshness indicator.

443
MCQmedium

A data engineer is ingesting JSON data from an IoT sensor network. The JSON records contain nested arrays and objects. The engineer needs to flatten the structure to load it into a relational table. Which approach is most appropriate?

A.Use a JSON parsing function to extract each field into separate columns, manually specifying the path for each nested element.
B.Store the entire JSON as a single string column and parse it during analysis using string functions.
C.Convert the JSON to XML first, then use XML parsing functions to extract data.
D.Use a SQL function like OPENJSON or JSON_TABLE to shred the JSON and return a relational rowset.
AnswerD

Functions like OPENJSON (SQL Server) or JSON_TABLE (MySQL, Oracle) are designed to parse JSON and output relational rows and columns. They can handle nested arrays by specifying paths and using CROSS APPLY or LATERAL joins. This approach is scalable and adapts to schema changes with minimal code changes.

Why this answer

The most appropriate method is to use a dedicated JSON shredding function like OPENJSON or JSON_TABLE. These functions are built to parse nested JSON and output a relational rowset, handling arrays and objects efficiently. They integrate with SQL queries and allow joining and filtering, making them ideal for loading into relational tables.

Exam trap

The trap here is underestimating the complexity of nested JSON and attempting to parse it with basic string functions, which fails with nested structures.

444
MCQeasy

You are a data analyst at a logistics company. The operations manager wants to reduce delivery delays. You have historical data including order date, delivery date, distance, weather conditions, and driver ID. Initial analysis shows that the average delivery time has increased over the past six months. You suspect that weather is a contributing factor, but you need to confirm. The company also wants to build a model to predict delivery times to better manage customer expectations. The data contains missing values for weather conditions in about 10% of records, and some driver IDs are incorrect. You have limited time and resources. What should you do first?

A.Immediately focus on time series analysis to look for patterns
B.Start by cleaning the data: correct driver IDs and decide how to handle missing weather data, then perform exploratory data analysis
C.Collect more data to fill missing values
D.Build a predictive model using all available data after imputing missing weather data
AnswerB

Cleaning the incorrect driver IDs and deciding a treatment for the 10% missing weather values must precede any analysis, because exploratory data analysis and the predictive model both depend on accurate, complete inputs; with limited time, this directly satisfies the stem's constraint of confirming weather's role reliably.

Why this answer

Data cleaning and exploratory data analysis (EDA) are foundational steps before any modeling or time series work. With missing weather data (10%) and incorrect driver IDs, proceeding without cleaning would introduce bias and errors. EDA will reveal patterns, correlations, and data quality issues, enabling informed decisions on imputation and feature engineering for the predictive model.

Exam trap

CompTIA often tests the misconception that you can jump directly to modeling or advanced analysis without first ensuring data quality, ignoring the 'garbage in, garbage out' principle.

How to eliminate wrong answers

Option A is wrong because time series analysis assumes clean, consistent data; applying it directly with missing values and incorrect IDs would yield unreliable patterns and waste resources. Option C is wrong because collecting more data is time-consuming and does not address the existing incorrect driver IDs or the need to understand current data quality; it also assumes missing values are random, which may not hold. Option D is wrong because building a predictive model on uncleaned data with imputed weather values without prior EDA risks overfitting, misinterpretation of feature importance, and propagation of errors from incorrect IDs.

445
Multi-Selecthard

Which THREE actions improve the accessibility of data visualizations for users with visual impairments? (Select exactly three.)

Select 3 answers
A.Provide text alternatives for charts (e.g., data tables).
B.Use only color to convey information.
C.Use clear and descriptive labels.
D.Ensure sufficient color contrast.
E.Add animated transitions between views.
AnswersA, C, D

Text alternatives such as data tables expose the underlying values in a machine-readable form, letting screen readers convey the chart's content to blind users. This satisfies the accessibility requirement by removing dependence on the graphical rendering.

Why this answer

Option A is correct because providing text alternatives such as data tables gives screen reader users and others who cannot perceive the chart a non-visual equivalent of the data, satisfying WCAG 1.1.1 Non-text Content. Option C is correct because clear, descriptive labels (titles, axis names, legends, and data point labels) let users understand what the visualization represents without relying on visual inference or color alone. Option D is correct because sufficient color contrast between text/graphical elements and their background meets WCAG 1.4.3 and 1.4.11, making charts legible for users with low vision or color vision deficiencies.

Option B does not belong because using only color to convey information fails WCAG 1.4.1 Use of Color and excludes users who cannot distinguish those colors. Option E does not belong because animated transitions can trigger vestibular issues and distract users, and they do not improve accessibility of the underlying data.

Exam trap

DA0-002 often tests the misconception that color alone can convey information or that animations enhance accessibility, when in fact both are accessibility anti-patterns.

446
MCQhard

An analyst creates a stacked bar chart showing quarterly sales by product category. The chart becomes hard to read because some categories have very small contributions. Which redesign is most effective?

A.Combine small categories into an 'Other' group
B.Change to a pie chart for each quarter
C.Increase the width of each bar
D.Switch to a 3D stacked column chart
AnswerA

Combining small categories into an 'Other' group reduces the number of segments competing for limited bar height, so each remaining category occupies a larger, readable portion. This directly addresses the stem's constraint: tiny contributions that become illegible in a stacked bar. Aggregation preserves the total while restoring visual clarity.

Why this answer

Combining small categories into an 'Other' group reduces visual clutter and improves readability by aggregating negligible contributions into a single bar segment. This technique preserves the overall trend while eliminating the noise from many tiny slices that make the stacked bar chart hard to interpret.

Exam trap

The trap here is that candidates often think adding more visual elements (3D, wider bars) or changing chart types (pie) will fix readability, when the real solution is data aggregation to reduce cognitive load.

How to eliminate wrong answers

Option B is wrong because using a pie chart for each quarter does not solve the problem of small categories; it merely shifts the same issue to a different chart type, where tiny slices are even harder to compare across quarters. Option C is wrong because increasing bar width does not address the core problem of too many small segments; it only stretches the visual horizontally without reducing the number of categories. Option D is wrong because switching to a 3D stacked column chart introduces perspective distortion and occlusion, making small contributions even more difficult to discern and violating best practices for accurate data visualization.

447
MCQhard

A data analyst is comparing the means of two independent groups using a t-test. The sample sizes are small and the data is not normally distributed. Which condition is violated for a valid t-test?

A.Normality
B.Equal variances
C.Independence of observations
D.Sample size larger than 30
AnswerA

The t-test assumes the sampling distribution of the mean is normal. With small samples, the central limit theorem does not apply, so non-normal data violates the normality assumption. Independence and equal variances are separate assumptions; normality is the condition breached here.

Why this answer

A t-test assumes that the sampling distribution of the mean is approximately normal. With small sample sizes, the Central Limit Theorem does not guarantee normality, so the data itself should be approximately normal. Since the data is not normally distributed and sample sizes are small, the normality assumption is violated, making the t-test invalid.

Exam trap

DA0-002 often tests the confusion between the assumptions of a t-test, leading candidates to select 'equal variances' or 'sample size larger than 30' when the actual violated condition is normality due to small, non-normal samples.

How to eliminate wrong answers

Option B is wrong because equal variances is a separate assumption (homogeneity of variance) that can be relaxed with Welch's t-test; it is not the condition violated here. Option C is wrong because independence of observations is an assumption, but the question states the groups are independent, so it is not violated. Option D is wrong because a sample size larger than 30 is not a requirement for a t-test; it is a rule of thumb for invoking the Central Limit Theorem, and the question says sample sizes are small, so this is not a violated condition but rather a reason normality matters.

448
MCQmedium

A data analyst finds that the "Age" column contains values like "N/A", "unknown", and negative numbers. Which data quality dimension is primarily affected?

A.Accuracy
B.Consistency
C.Validity
D.Completeness
AnswerC

Validity concerns whether values conform to the defined format, type, and permitted range for a field. Text placeholders and negative ages violate Age's expected numeric, non-negative domain, so the dimension breached is validity rather than completeness or accuracy.

Why this answer

Validity refers to whether data values conform to the defined format, type, and allowable range for a field. 'N/A', 'unknown', and negative numbers in an Age column violate the expected numeric, non-negative domain, so validity is the primary dimension affected.

Exam trap

DA0-002 often tests the overlap between validity and accuracy, causing candidates to choose accuracy when the issue is really format/domain conformance rather than truthfulness.

How to eliminate wrong answers

Option A is wrong because accuracy concerns whether a value correctly reflects reality (e.g., a real age of 30 recorded as 31), not whether it fits the allowed format. Option B is wrong because consistency concerns whether the same data is represented uniformly across systems or records, not whether individual values are permissible. Option D is wrong because completeness concerns missing values, whereas here values are present but invalid.

449
MCQhard

A data engineer is loading a large CSV file into a relational staging table. Several columns contain numeric values with thousands separators, such as '1,234.56', and a few rows contain the text 'N/A' in those columns. The target columns are defined as DECIMAL. Which approach best prepares the data for a successful load while preserving the ability to audit rejected values?

A.Strip thousands separators, convert 'N/A' to NULL, validate the remaining values as numeric, and load rejected rows into an error table.
B.Load the values as strings into VARCHAR columns, then cast them during every downstream query.
C.Replace 'N/A' with 0 and load all values directly into the DECIMAL columns.
D.Set the database session to a locale that interprets commas as decimal points and load the file unchanged.
AnswerA

Removing thousands separators produces parseable numeric strings, mapping 'N/A' to NULL respects the DECIMAL nullability, and validating before load prevents type errors. Routing rejected rows to an error table preserves auditability. This combination satisfies the goal of a clean load while retaining visibility into records that could not be converted, which is essential for data quality monitoring and reprocessing.

Why this answer

Preparing numeric text with thousands separators and non-numeric tokens requires cleaning the format, mapping non-values to NULL, validating conversion, and isolating failures. Stripping separators and converting 'N/A' to NULL allows valid rows to load as DECIMAL, while an error table preserves rejected rows for audit and reprocessing. String storage, zero substitution, and locale changes either defer errors, fabricate data, or corrupt magnitudes.

Exam trap

The trap here is treating 'N/A' as equivalent to zero or assuming locale settings can safely reinterpret thousands separators without corrupting numeric magnitude.

450
MCQhard

An IT operations team monitors 200 servers. Each server reports CPU utilization (0-100%) every five minutes for the past year. The team wants to visualize the data to identify servers that are consistently over 80% utilization and detect any unusual spikes. They have a large dataset with 100,000+ records per server. The current visualization is a single scatter plot with CPU utilization on the y-axis, time on the x-axis, and each server as a different colored point. The chart is extremely cluttered, with points overlapping and colors indistinguishable. What should the team do to improve the visualization?

A.Use a heatmap showing CPU utilization over time per server, or create small multiple charts (one per server)
B.Switch to a line chart with each server as a separate line
C.Add a trend line to each server's data and remove the individual points
D.Increase the size of the data points to make them more visible
AnswerA

Aggregating each server into its own small-multiple panel, or binning utilisation into colour cells on a heatmap, removes the overplotting caused by 200 overlapping series on shared axes. Per-server panels expose sustained over-80% periods and spikes that a single scatter plot hides.

Why this answer

With 200 servers and 100,000+ records per server, a single scatter plot is too cluttered. A heatmap can show CPU utilization over time per server using color intensity, or small multiples (one chart per server) can separate the data. Both approaches reduce overlap and make patterns like consistent over-80% utilization and spikes visible.

Exam trap

DA0-002 often tests visualization best practices, and candidates may choose a line chart or trend line thinking it simplifies, but the correct answer addresses overplotting with aggregation or separation.

How to eliminate wrong answers

Option B is wrong because a line chart with 200 lines would still be cluttered and colors indistinguishable, similar to the scatter plot. Option C is wrong because adding a trend line and removing points would hide the spikes and detailed fluctuations, which are important for detecting unusual spikes. Option D is wrong because increasing point size would worsen the clutter and overlap.

Page 5

Page 6 of 14

Page 7