Courseiva

CompTIA Data+ (DA0-002) (DA0-002) — Questions 751–825

1004 questions total · 14pages · All types, answers revealed

Page 10

Page 11 of 14

Page 12
751
MCQeasy

Refer to the exhibit. A stakeholder complains that the line chart exaggerates the changes in sales. What is the most likely cause?

A.The y-axis does not start at zero
B.There are too few data points
C.The data labels are incorrect
D.The chart type should be a bar chart
AnswerA

A truncated y-axis compresses the visible range, so small absolute changes occupy a large proportion of the plotted height, steepening the apparent slope. The stem's complaint of exaggerated variation is satisfied precisely because the baseline sits above zero, inflating the visual gradient of the sales line.

Why this answer

A line chart can exaggerate changes if the y-axis does not start at zero. By truncating the y-axis (e.g., starting at 80 instead of 0), small fluctuations appear as large peaks and valleys, misleading viewers. This is a common data visualization pitfall that distorts the perception of magnitude.

Exam trap

DA0-002 often tests the principle that truncated y-axes exaggerate changes, so candidates may blame the chart type or data points instead of recognizing the axis scaling issue as the root cause.

How to eliminate wrong answers

Option B is wrong because having too few data points would make the chart sparse but would not inherently exaggerate changes; it might make trends less reliable but not visually exaggerated. Option C is wrong because incorrect data labels would show wrong numbers, but the complaint is about exaggeration of changes, which is a scaling issue. Option D is wrong because changing to a bar chart would not fix the exaggeration if the y-axis is still truncated; the issue is the axis scale, not the chart type.

752
MCQhard

While reviewing a dashboard, an analyst notices that the data in a trend line chart does not match the underlying data due to a filter setting. The dashboard is used for weekly executive meetings. What should the analyst do?

A.Ignore the discrepancy if it is small.
B.Wait for someone to complain before acting.
C.Immediately remove the dashboard and send raw data.
D.Document the issue and fix the filter before the next meeting.
AnswerD

The filter misconfiguration corrupts the trend line, so executives would review inaccurate figures. Documenting the discrepancy preserves an audit trail, and correcting the filter restores data integrity before the weekly meeting, satisfying the need for trustworthy reporting.

Why this answer

The analyst has identified a data integrity issue caused by a filter setting that directly impacts the accuracy of the trend line chart. The dashboard is used for weekly executive meetings, so the analyst must document the discrepancy and correct the filter before the next meeting to ensure data-driven decisions are based on accurate visualizations. This aligns with best practices in data governance and the principle of maintaining trust in reporting tools.

Exam trap

The trap here is that candidates may assume small discrepancies are acceptable or that waiting for complaints is a valid approach, but the exam emphasizes proactive data integrity and the importance of maintaining accurate visualizations for scheduled stakeholder meetings.

How to eliminate wrong answers

Option A is wrong because ignoring even a small discrepancy in a dashboard used for executive decision-making can lead to compounded errors in trend analysis and erode trust in the data; any deviation from the underlying data must be investigated and corrected. Option B is wrong because waiting for someone to complain is reactive and unprofessional; the analyst should proactively ensure data accuracy, especially for a recurring weekly meeting where stakeholders rely on consistent, correct visualizations. Option C is wrong because immediately removing the dashboard and sending raw data disrupts the established reporting workflow and forces executives to interpret unaggregated data, which is inefficient and likely to introduce new errors; the proper action is to fix the filter and restore the correct trend line chart.

753
Multi-Selectmedium

A data analyst needs to identify duplicate customer records. Which TWO methods are commonly used? (Select two.)

Select 2 answers
A.Fuzzy matching using Levenshtein distance
B.Sorting and comparing adjacent rows
C.Visual inspection of random sample
D.Using a hash function on primary key
E.Exact match on all fields
AnswersA, B

Levenshtein distance measures the minimum single-character edits between strings, so records differing by typos or transpositions still match. This satisfies the need to catch near-duplicates that exact key comparison would miss across customer names and addresses.

Why this answer

Option A (Fuzzy matching using Levenshtein distance) is correct because it measures the minimum number of single-character edits (insertions, deletions, substitutions) needed to transform one string into another, making it ideal for catching near-duplicate customer records that differ slightly due to typos or formatting variations. Option B (Sorting and comparing adjacent rows) is correct because once records are sorted by a key field such as name or email, duplicate or near-duplicate entries naturally cluster together, allowing efficient pairwise comparison of neighboring rows to flag matches. Option C is not a reliable, scalable method since random sampling cannot guarantee detection of all duplicates and is subjective.

Option D does not help because hashing a primary key, which is unique by definition, will never reveal duplicates. Option E is too strict, as exact matching on all fields will miss duplicates that differ in even one attribute, such as a middle initial or apartment number.

Exam trap

The trap here is that candidates often choose 'Exact match on all fields' (Option E) thinking it is a reliable deduplication method, but in practice it fails to catch real-world duplicates that have any minor variation, and the exam expects you to recognize that fuzzy matching and sorted adjacency comparisons are the standard techniques for duplicate detection.

754
MCQmedium

A marketing team wants a single view comparing this quarter's campaign performance against the same quarter last year across five channels, with the ability to drill from channel totals down to individual campaign rows. Which visualization structure best supports both the comparison and the drill path?

A.A treemap sized by campaign spend with channels as top-level rectangles.
B.A matrix with channels as rows and the two quarters as columns, configured with a drill-down hierarchy from channel to campaign.
C.A bar chart of this quarter's channel totals with a separate bar chart of last year's totals placed beside it.
D.A line chart plotting campaign rows over time with one line per channel.
AnswerB

A matrix aligns each channel on a row and places this quarter and last year side by side as columns, so the comparison is read horizontally in one glance. Because matrices support hierarchical row groups, expanding a channel row reveals its individual campaigns, delivering the drill path without leaving the visual. One artifact satisfies both the comparison and the drill requirement.

Why this answer

A matrix handles both demands at once: channels as rows give a stable axis for comparing the two quarter columns side by side, and the built-in row hierarchy lets a user expand a channel to reveal its campaigns. The comparison stays in one visual while the drill path descends the same hierarchy, avoiding a jump to a separate page.

Exam trap

The trap here is choosing a chart that shows magnitude or composition when the requirement is a precise period-over-period comparison combined with hierarchical drill-down.

755
MCQeasy

A data analyst receives the above JSON snippet from a web API. The analyst needs to extract the email addresses for all customers. Which JSONPath expression should be used?

A.$.customers[0].email
B.$..email
C.$.customers[*].email
D.$.customers.email
AnswerC

The wildcard `[*]` iterates every element of the `customers` array, while `.email` selects that key from each object, returning all addresses in one expression. This satisfies the stem's requirement to extract email addresses for all customers, not just a single indexed entry.

Why this answer

The JSONPath expression `$.customers[*].email` uses the wildcard `[*]` to select all elements in the `customers` array and then accesses the `email` property of each element. This matches the requirement to extract email addresses for all customers from the JSON snippet.

Exam trap

The trap here is that candidates often confuse the deep scan operator `..` with the array wildcard `[*]`, thinking `$..email` will neatly extract all customer emails, but it actually retrieves every `email` property at any depth, including from non-customer objects, leading to incorrect data extraction.

How to eliminate wrong answers

Option A is wrong because `$.customers[0].email` only retrieves the email address of the first customer in the array, not all customers. Option B is wrong because `$..email` uses the deep scan operator `..` which recursively searches the entire JSON tree for any property named `email`, potentially returning emails from nested objects or arrays that are not customers (e.g., from an `orders` or `address` object), leading to incorrect or extra results. Option D is wrong because `$.customers.email` attempts to access `email` directly on the `customers` array object, but arrays in JSONPath do not have a property named `email`; this expression would return `null` or an empty result unless the array itself has an `email` property, which it does not.

756
MCQmedium

In Looker Studio, what is the difference between dimensions and metrics?

A.Dimensions are used for aggregation; metrics are for grouping
B.Dimensions are numerical; metrics are categorical
C.Both can be categorical or numerical
D.Dimensions are categorical; metrics are numerical
AnswerD

Dimensions group and label data as categories such as product name or country, while metrics are the numeric measures aggregated by those dimensions, like revenue or session count. This categorical-versus-numerical split is the defining axis that determines how each field is used in charts.

Why this answer

In Looker Studio, dimensions are fields that contain categorical data (e.g., text, dates, or geographic names) used to group and segment data, while metrics are numerical fields (e.g., counts, sums, averages) that can be aggregated. Option D is correct because this distinction is fundamental to how Looker Studio processes and visualizes data: dimensions define the rows or categories in a chart, and metrics provide the quantitative values to be measured.

Exam trap

The trap here is that candidates often confuse the general data types (numeric vs. string) with the semantic roles in Looker Studio, leading them to choose option C, but the exam expects you to know that dimensions are always used for grouping (categorical) and metrics for aggregation (numerical) in the context of this tool.

How to eliminate wrong answers

Option A is wrong because it reverses the roles: dimensions are used for grouping and segmenting data, not aggregation, while metrics are the fields that are aggregated (e.g., SUM, COUNT, AVG). Option B is wrong because it incorrectly states that dimensions are numerical and metrics are categorical; in reality, dimensions are typically categorical (text, date, boolean) and metrics are numerical. Option C is wrong because while both can technically be categorical or numerical in raw data, Looker Studio enforces a strict semantic distinction: dimensions are treated as grouping keys (categorical) and metrics as aggregatable values (numerical), and mixing them leads to incorrect chart behavior.

757
MCQmedium

A sales manager wants to see the conversion rates at each stage of the sales pipeline, from initial contact to closed deal. Which chart type is most appropriate?

A.Waterfall chart
B.Funnel chart
C.Treemap
D.Bar chart
AnswerB

A funnel chart plots sequential, progressively narrowing stages, so each pipeline phase from initial contact to closed deal displays as a descending segment. This directly shows conversion rates and drop-off between stages, which no other chart type represents as clearly.

Why this answer

A funnel chart is purpose-built to show sequential stages of a process and the progressive reduction in count or value at each stage, making it ideal for visualizing sales pipeline conversion from initial contact to closed deal. The tapering shape visually communicates drop-off at each stage, which is exactly what the sales manager needs.

Exam trap

The trap is selecting a bar chart because it can display the same stage counts, but bar charts lack the sequential, tapering visual metaphor that communicates conversion drop-off inherent to a funnel chart.

How to eliminate wrong answers

Option A is wrong because a waterfall chart shows how an initial value is increased or decreased by a series of intermediate positive/negative contributions (e.g., revenue bridge), not sequential stage conversion. Option C is wrong because a treemap displays hierarchical part-to-whole relationships using nested rectangles, not stage-by-stage progression. Option D is wrong because a bar chart compares discrete categories side by side but does not inherently convey sequential flow or conversion drop-off between stages.

758
MCQeasy

A sales analyst wants to show total sales by product category, with each category's contribution to the total. Which chart type is best?

A.Scatter plot
B.Box plot
C.Stacked bar chart
D.Line chart
AnswerC

A stacked bar chart segments each bar into category components, so segment heights show individual product contributions while the full bar gives the total. This simultaneously satisfies both requirements: per-category sales and each category's share of overall sales.

Why this answer

A stacked bar chart shows each product category as a bar segment whose height represents its sales contribution, and the total bar height represents overall sales — simultaneously displaying both individual contributions and the whole. This directly satisfies the requirement to show total sales by category with each category's contribution.

Exam trap

DA0-002 often tests the pie-chart-versus-stacked-bar confusion — candidates pick pie chart for 'contribution to total,' but pie charts can't show totals across multiple categories or support multi-level composition the way a stacked bar can.

How to eliminate wrong answers

Option A is wrong because a scatter plot shows the relationship between two continuous variables (correlation), not categorical contributions to a total. Option B is wrong because a box plot summarizes a distribution (median, quartiles, outliers) for a numeric variable, not compositional totals. Option D is wrong because a line chart shows trends over a continuous dimension like time, not parts-of-a-whole composition across categories.

759
Multi-Selecthard

Which TWO of the following are primary benefits of implementing a data governance program?

Select 2 answers
A.Faster data processing speed
B.Increased data volume
C.Improved data quality and consistency
D.Lower storage costs
E.Reduced data redundancy
AnswersC, E

Governance establishes standards, ownership and controls across the data lifecycle, which directly raises data quality and consistency. These are the primary benefits cited, as governance enforces uniform definitions and validation rather than merely storing or visualising data.

Why this answer

Option C is correct because a core purpose of data governance is establishing policies, standards, and stewardship roles that enforce data quality dimensions such as accuracy, completeness, and consistency across systems. Option E is correct because governance defines authoritative data sources, master data management, and ownership rules that eliminate duplicate and conflicting copies of data across the enterprise. The remaining options do not belong: A (faster data processing speed) is an infrastructure or query-optimization outcome, B (increased data volume) is a byproduct of data accumulation rather than a governance benefit, and D (lower storage costs) is a cost-optimization result typically achieved through tiering, deduplication, or archiving rather than governance itself.

Exam trap

The trap here is that candidates may confuse data governance with data management or data engineering tasks, mistakenly thinking it directly improves performance or reduces costs, when its core value is in quality, consistency, and compliance.

760
MCQeasy

A data team needs to extract data from a legacy system that only supports flat file exports. Which data acquisition method is most appropriate?

A.Database replication
B.API call
C.Web scraping
D.File transfer via SFTP
AnswerD

SFTP transfers the flat files the legacy system exports, satisfying the constraint that no direct database or API access exists. It preserves file integrity over an encrypted channel, so scheduled batch extraction works reliably without custom connectors.

Why this answer

The legacy system only supports flat file exports, meaning it cannot provide direct database or API access. SFTP (SSH File Transfer Protocol) is the most appropriate method because it securely transfers flat files over a network, aligning with the system's export capabilities while ensuring data integrity and encryption during transit.

Exam trap

The trap here is that candidates may confuse 'flat file exports' with a need for real-time or API-based methods, overlooking that SFTP is the standard secure file transfer protocol for batch-oriented legacy systems.

How to eliminate wrong answers

Option A is wrong because database replication requires the source system to support a database engine with replication features (e.g., transactional logs or CDC), which a legacy flat-file-only system lacks. Option B is wrong because an API call requires the legacy system to expose a programmatic interface (e.g., REST or SOAP), which is not available if it only supports flat file exports. Option C is wrong because web scraping is used to extract data from web pages via HTTP, not from a legacy system that exports flat files via a file transfer protocol.

761
MCQeasy

Refer to the exhibit. Which data quality dimension is compromised by the missing value for Charlie's salary?

A.Uniqueness
B.Completeness
C.Timeliness
D.Accuracy
AnswerB

Charlie's salary field is absent entirely, so the record lacks a required attribute. Completeness measures whether all expected values are present, and the missing value directly violates it, satisfying the stem's data quality dimension question.

Why this answer

Completeness measures whether all required data is present. Charlie's missing salary value means the record is incomplete, directly violating this dimension. In data quality frameworks, completeness is assessed by the proportion of non-null values in a field, and a null salary here fails that check.

Exam trap

CompTIA often tests the distinction between 'missing' (completeness) and 'wrong' (accuracy), leading candidates to confuse a null value with an incorrect value.

How to eliminate wrong answers

Option A is wrong because uniqueness refers to the absence of duplicate records or values, not missing data; a missing salary does not create a duplicate. Option C is wrong because timeliness concerns whether data is up-to-date or available when needed, not whether a value is present or absent. Option D is wrong because accuracy measures correctness of values against a reference source; a missing value is not an inaccurate value—it is an absent one.

762
MCQeasy

A dashboard designer wants to highlight the sales performance of individual sales representatives compared to team averages. Which chart type is most suitable for this comparison?

A.Scatter plot
B.Bar chart with average line
C.Pie chart
D.Line chart
AnswerB

A bar chart plots each representative's sales as discrete bars, while an average line provides a shared reference against which individual performance is instantly compared. This directly satisfies the requirement to highlight individual results relative to the team average.

Why this answer

A bar chart with an average line is the standard visualization for comparing individual categorical values (each sales rep) against a team benchmark. Each rep gets a bar whose height shows their sales, and a horizontal reference line drawn at the team average lets viewers instantly see who is above or below the mean. This combines categorical comparison with a statistical reference, which no other listed chart does.

Exam trap

The trap here is confusing 'comparison of individuals to a group benchmark' with 'comparison of two numeric variables,' which lures candidates toward scatter plots or line charts.

How to eliminate wrong answers

Option A is wrong because a scatter plot displays the relationship/correlation between two numeric variables (e.g., calls vs. revenue), not a categorical comparison of reps against a single average. Option C is wrong because a pie chart shows parts-of-a-whole proportions and cannot meaningfully display an average benchmark or compare many individual values. Option D is wrong because a line chart is designed for trends over a continuous dimension such as time, not for comparing discrete individuals to a team average.

763
MCQeasy

An analyst is presenting a recommendation to increase marketing spend. Which statement best follows the data-driven recommendation structure (evidence → insight → recommendation → expected impact)?

A.A 10% increase in marketing spend is recommended because we have budget.
B.The data shows sales are up, so we should spend more on marketing.
C.Based on a 5% lift in sales from previous campaigns, we recommend a 10% increase in marketing spend, expecting a 7% revenue growth.
D.We should increase marketing spend by 10% because it might boost sales.
AnswerC

The statement chains evidence (5% lift from prior campaigns), insight (spend drives sales), recommendation (10% increase) and expected impact (7% revenue growth), matching the required structure. The quantified forecast makes the recommendation testable rather than assertive.

Why this answer

The correct structure provides evidence, insight derived from it, a recommendation, and the expected impact.

764
Multi-Selecthard

Which THREE are best practices for designing a dashboard for executive consumption?

Select 3 answers
A.Include detailed raw data tables for transparency
B.Ensure the dashboard is responsive for mobile devices
C.Use a separate chart for each metric to avoid clutter
D.Provide interactive filters for time periods and regions
E.Display the most critical KPIs at the top
AnswersB, D, E

Executives frequently review metrics away from their desks, so responsive layout ensures the dashboard renders legibly on phones and tablets. This satisfies the executive-consumption constraint by preserving readability across the devices that audience actually uses.

Why this answer

Option B is correct because executives frequently consume dashboards on tablets and phones, so a responsive layout ensures KPIs remain legible and usable across screen sizes. Option D is correct because interactive filters for time periods and regions let executives drill into the specific slice of data relevant to their decision without needing a new report. Option E is correct because placing the most critical KPIs at the top follows the principle of prioritized information hierarchy, ensuring the most important metrics are seen first.

Option A is not ideal because detailed raw data tables overwhelm executive viewers and obscure high-level trends. Option C is not ideal because fragmenting every metric into its own chart creates visual clutter and prevents at-a-glance comparison of related KPIs.

Exam trap

The trap is that 'include detailed raw data' sounds like transparency and 'separate chart per metric' sounds thorough, but the exam tests whether you know executive dashboards prioritize brevity, consolidation, and interactivity.

765
MCQeasy

A data analyst is importing a fixed-width text file into a relational database. The file has no header row, and fields are separated by a single space, but some fields contain trailing spaces of varying lengths that shift the apparent column boundaries. The analyst must load the data reliably into the correct columns. Which approach is most appropriate?

A.Split each line on the single space delimiter and assign fields sequentially to columns.
B.Parse the file using fixed character positions derived from a documented layout specification rather than splitting on spaces.
C.Import the entire line into a single column and rely on downstream views to extract fields.
D.Replace all spaces with commas and then load the file as comma-separated values.
AnswerB

A fixed-width file defines each field by character position, so parsing by documented offsets preserves values even when trailing spaces vary. Splitting on spaces would misalign columns because multiple spaces are treated inconsistently. Using the layout specification is the reliable, repeatable way to load this format correctly.

Why this answer

Fixed-width files require position-based parsing because field boundaries are defined by character offsets, not delimiters. Documented offsets remain stable regardless of trailing spaces, whereas splitting on spaces or blindly converting delimiters misaligns columns. Loading the whole line or substituting commas defers or worsens the problem, so position-based parsing is the correct choice.

Exam trap

The trap here is assuming any whitespace-separated text is delimited data, when varying trailing spaces mean the file is actually fixed-width and must be parsed by position.

766
MCQhard

A data analyst needs to perform stratified sampling on a customer database to ensure proportional representation across three regions: North (40%), South (30%), and West (30%). The total sample size required is 1,000. How many customers should be sampled from the North region?

A.333
B.500
C.300
D.400
AnswerD

Stratified sampling allocates sample counts proportionally to each stratum's share of the population. North represents 40% of customers, so its allocation is 0.40 × 1,000 = 400. This satisfies the stem's proportional-representation constraint directly, giving North exactly its 40% share of the sample.

Why this answer

Stratified sampling allocates sample size proportionally to each stratum's share of the population. North represents 40% of the population, so 40% of the 1,000-customer sample — 400 customers — should be drawn from North.

Exam trap

DA0-002 often tests whether candidates can apply the proportional allocation formula — the trap is confusing the region's percentage with the sample count or mixing up which region gets which share.

How to eliminate wrong answers

Option A is wrong because 333 corresponds to roughly one-third (33.3%), which would be the allocation if the three regions were equal — but North is 40%, not 33%. Option B is wrong because 500 represents 50% of the sample, which would apply only if North were half the population. Option C is wrong because 300 represents 30%, which is the correct allocation for South or West, not North.

767
Multi-Selectmedium

Which TWO of the following are appropriate uses of min-max normalisation?

Select 2 answers
A.Transforming data to have mean 0 and standard deviation 1
B.Scaling features to a range of 0 to 1
C.Preparing data for linear regression with normally distributed residuals
D.Preparing data for k-nearest neighbours algorithm
E.Handling missing values
AnswersB, D

Min-max normalisation linearly transforms each feature to a fixed 0–1 range using the minimum and maximum values, satisfying the requirement to bound features within a common scale. This suits algorithms sensitive to magnitude, such as k-nearest neighbours or neural networks, where unbounded inputs distort distance calculations.

Why this answer

Option B is correct because min-max normalisation rescales each feature to a fixed range, typically [0, 1], using the formula (x − min)/(max − min), which is exactly its defining purpose. Option D is correct because k-nearest neighbours relies on distance calculations (e.g., Euclidean distance), so features on different scales would dominate the distance metric; min-max normalisation puts all features on a comparable 0–1 scale, improving the algorithm's behaviour. Option A is not a use of min-max normalisation but of standardisation (z-score scaling), which produces mean 0 and standard deviation 1.

Option C is not specific to min-max normalisation; linear regression with normally distributed residuals concerns the error distribution, not feature scaling, and standardisation is more commonly associated with such assumptions. Option E is incorrect because missing values are handled by imputation or deletion techniques, not by min-max normalisation, which requires complete numeric data.

Exam trap

DA0-002 often tests the confusion between min-max normalisation and z-score standardisation; candidates may incorrectly select the z-score description as a use of min-max.

768
MCQhard

A data analyst is examining the relationship between advertising spend and sales revenue across 50 regions. The analyst calculates a Pearson correlation coefficient of 0.85. Which of the following conclusions is most appropriate?

A.Increasing advertising spend causes an increase in sales revenue.
B.The relationship is not statistically significant because the correlation is less than 0.90.
C.Advertising spend explains 85% of the variation in sales revenue.
D.There is a strong positive linear relationship between advertising spend and sales revenue.
AnswerD

A Pearson correlation of 0.85 indicates a strong positive linear association. This means that as advertising spend increases, sales revenue tends to increase linearly. It does not imply causation, but it does summarize the strength and direction of the linear relationship. This is the most appropriate conclusion based solely on the correlation coefficient.

Why this answer

A Pearson correlation of 0.85 indicates a strong positive linear relationship between the two variables. It does not establish causation, nor does it directly state the proportion of variance explained. The correlation coefficient must be squared to get the coefficient of determination, and significance requires a hypothesis test.

Thus, the appropriate conclusion is the strong positive linear association.

Exam trap

The trap here is interpreting a high correlation as evidence of causation or confusing the correlation coefficient with the coefficient of determination.

769
Multi-Selecthard

A data analyst is designing a dashboard for a hospital's emergency department. The dashboard must display real-time patient wait times, current bed occupancy, and the number of patients waiting to be seen. The primary users are charge nurses and physicians who need to make quick staffing decisions. Which two design principles are most critical to ensure the dashboard is effective for this scenario? (Choose two.)

Select 2 answers
A.Include detailed patient demographic information to provide context for each waiting patient.
B.Use a complex, multi-layered drill-down structure to allow users to explore data from multiple angles.
C.Use a consistent color scheme with red indicating critical thresholds and green indicating normal status.
D.Display all available historical data to provide full context for current conditions.
E.Ensure the dashboard automatically refreshes frequently, such as every minute, to reflect current conditions.
AnswersC, E

In a high-pressure emergency department, rapid interpretation is essential. A consistent color scheme with red for critical and green for normal allows staff to instantly assess status without reading numbers. This reduces cognitive load and speeds decision-making, which is critical when every minute counts. It aligns with best practices for operational dashboards where alerts must be immediately visible.

Why this answer

For a real-time operational dashboard used in a high-pressure environment, immediate interpretability and data freshness are paramount. A consistent color scheme for thresholds allows quick status assessment, and frequent automatic refreshes ensure the data reflects current conditions. Both features directly support rapid, informed staffing decisions.

Exam trap

The trap here is assuming that more data and interactivity always improve a dashboard, when in fact simplicity and timeliness are more critical for real-time operational use.

770
MCQeasy

A junior analyst is asked to show how total monthly support ticket volume changed over the past 24 months so leadership can spot seasonal peaks. Which visualization is most appropriate?

A.A stacked bar chart of ticket categories per month.
B.A line chart with months on the horizontal axis and ticket volume on the vertical axis.
C.A scatter plot of ticket volume against month index.
D.A pie chart with one slice per month.
AnswerB

A line chart places time on the horizontal axis and connects ordered points, so rising and falling ticket volume reads as a continuous trajectory. That makes recurring seasonal peaks immediately visible as repeating humps across the 24-month window, which is exactly the pattern leadership needs. Lines also scale well to two dozen periods, unlike categorical charts that crowd as categories multiply.

Why this answer

Monthly ticket volume across 24 periods is a time series, and a line chart is the canonical encoding for ordered continuous change. Connecting the points lets readers perceive slope, peaks, and repeating seasonal cycles at a glance. The horizontal axis preserves chronological order, which is what distinguishes a trend display from a categorical comparison.

Exam trap

The trap here is reaching for a chart that shows composition or correlation when the question is fundamentally about change over an ordered time axis.

771
MCQhard

A data analyst needs to create a recursive CTE to traverse a hierarchical employee-manager table. Which of the following is a key requirement for a recursive CTE?

A.The CTE must include a WHERE clause in the recursive member
B.The CTE must use the RECURSIVE keyword in the WITH clause
C.The recursive CTE must have at least one anchor member that does not reference the CTE
D.The recursive member must use UNION instead of UNION ALL
AnswerC

A recursive CTE requires an anchor member returning the base rows plus a recursive member referencing the CTE itself, combined by UNION ALL. The anchor must not reference the CTE, otherwise the query cannot establish its initial result set.

Why this answer

A recursive CTE is defined by two parts: an anchor member that produces the base result set without referencing the CTE itself, and a recursive member that references the CTE and is combined with the anchor via UNION ALL (or UNION). The anchor member is mandatory — without it, the recursion has no starting point and the query fails. This is the defining structural requirement of a recursive CTE.

Exam trap

DA0-002 often tests the misconception that the RECURSIVE keyword or a WHERE clause is mandatory; the true structural requirement is the anchor member, which candidates frequently overlook.

How to eliminate wrong answers

Option A is wrong because a WHERE clause is not a structural requirement of the recursive member; filtering is common but optional, and recursion terminates via the join condition or an explicit depth limit, not a mandatory WHERE. Option B is wrong because the RECURSIVE keyword is optional in many engines — PostgreSQL, SQL Server, and Oracle accept WITH RECURSIVE or plain WITH depending on the dialect, and MySQL requires RECURSIVE, so it is not a universal requirement. Option D is wrong because UNION ALL is actually the more common and often required form; UNION (which deduplicates) is allowed but not mandatory, and using UNION ALL is standard practice for performance.

772
MCQhard

A data analyst discovers that a dataset contains multiple records for the same customer with different spellings (e.g., 'Jon' vs 'John'). Which data preparation step should be applied first?

A.Merge all records into one per customer.
B.Remove duplicates based on exact match.
C.Standardize text fields using a lookup table.
D.Flag records for manual review.
AnswerC

Variant spellings of the same customer must be reconciled before deduplication or matching. A lookup table maps each variant to a canonical value, standardising the text field first so subsequent joins and dedupe logic treat 'Jon' and 'John' as one entity.

Why this answer

The first step when dealing with inconsistent text values (like 'Jon' vs 'John') is to standardize the data using a lookup table or reference mapping. This ensures that all variations are normalized to a canonical form before any merging or deduplication is attempted, preventing data loss and preserving referential integrity.

Exam trap

The trap here is that candidates often jump to 'remove duplicates' (Option B) because they think of exact-match deduplication, but the question specifically tests the understanding that data quality issues like inconsistent spellings must be resolved through standardization before any deduplication logic can be applied.

How to eliminate wrong answers

Option A is wrong because merging records before standardizing spellings would combine data based on non-uniform keys, likely creating erroneous composite records or losing the ability to correctly identify which records belong to the same customer. Option B is wrong because removing duplicates based on exact match would treat 'Jon' and 'John' as different records, failing to identify them as the same customer and leaving the inconsistency unresolved. Option D is wrong because flagging records for manual review is a downstream action that should only be taken after automated standardization has been attempted; skipping standardization first would result in an unnecessarily large and inefficient manual review workload.

773
Multi-Selectmedium

Which THREE of the following are common characteristics of unstructured data?

Select 3 answers
A.Easily queried using SQL
B.Often stored in NoSQL databases or data lakes
C.Can include text, images, and video
D.Stored in relational tables
E.Lacks a predefined schema
AnswersB, C, E

NoSQL databases and data lakes both accept data without a predefined relational schema, so they accommodate raw text, media and logs. This satisfies the stem's requirement by naming the storage platforms that handle schema-less content, unlike warehouses requiring structure before loading.

Why this answer

Option B is correct because unstructured data, such as documents, media files, and logs, is commonly stored in NoSQL databases (e.g., MongoDB, Cassandra) or data lakes (e.g., Amazon S3, Azure Data Lake) that do not require a fixed relational schema. Option C is correct because unstructured data encompasses heterogeneous formats including free text, images, audio, and video, which cannot be easily decomposed into rows and columns. Option E is correct because unstructured data by definition lacks a predefined schema, meaning there is no fixed data model or rigid structure enforced at write time.

Option A is incorrect because SQL querying relies on structured, tabular schemas, which unstructured data does not provide. Option D is incorrect because relational tables are the storage model for structured data, not unstructured data.

Exam trap

The trap here is confusing 'unstructured' with 'semi-structured' or assuming that any data stored in a database must be queryable via SQL; candidates often pick 'Easily queried using SQL' because they conflate storage with query capability.

774
MCQeasy

A data analyst is tasked with creating a report that shows the proportion of total sales contributed by each product category. The analyst wants to use a chart that clearly displays the parts of a whole. Which chart type is most appropriate?

A.Histogram
B.Scatter plot
C.Line chart
D.Pie chart
AnswerD

A pie chart is specifically designed to show how individual categories contribute to a total, with each slice representing a proportion. It is ideal for displaying the percentage of total sales by product category, as it visually communicates the parts of a whole in an intuitive way. This makes it the most appropriate choice for the scenario.

Why this answer

A pie chart is the most appropriate because it visually represents each product category's share of total sales as slices of a whole. This directly addresses the requirement to show proportions, making it easy for the audience to compare contributions at a glance.

Exam trap

The trap here is selecting a chart that is familiar but mismatched to the data type, such as a line chart for categorical proportions.

775
MCQeasy

A data analyst is using SQL to extract data. The analyst wants to retrieve all records from a table named 'sales' where the 'amount' column is greater than 100. Which SQL clause should be used?

A.WHERE
B.ORDER BY
C.GROUP BY
D.HAVING
AnswerA

WHERE filters individual rows before any grouping or aggregation, returning only those sales records whose amount exceeds 100. It satisfies the stem's constraint of retrieving all matching rows from the sales table, unlike HAVING, which filters grouped results after aggregation and cannot reference non-aggregated row values in this way.

Why this answer

The WHERE clause in SQL is used to filter records based on a specified condition, such as 'amount > 100'. It is applied directly to the rows in the 'sales' table before any grouping or ordering, making it the correct choice for retrieving only records where the amount exceeds 100.

Exam trap

The trap here is that candidates often confuse HAVING with WHERE, thinking both can filter rows, but HAVING is only valid after GROUP BY and for aggregate conditions, while WHERE filters individual rows before any grouping.

How to eliminate wrong answers

Option B (ORDER BY) is wrong because it is used to sort the result set by one or more columns, not to filter rows based on a condition. Option C (GROUP BY) is wrong because it groups rows that have the same values in specified columns into summary rows, often for use with aggregate functions, and does not filter individual records. Option D (HAVING) is wrong because it is used to filter groups after the GROUP BY clause has been applied, typically with aggregate functions, and cannot be used to filter individual rows before grouping.

776
MCQeasy

A data analyst needs to show the relationship between advertising spend (in dollars) and the number of website visits. Both variables are continuous. Which chart type is most suitable?

A.Line chart
B.Scatter plot
C.Box plot
D.Bar chart
AnswerB

A scatter plot maps advertising spend on the x-axis and website visits on the y-axis, plotting each observation as an individual point to reveal the correlation between two continuous variables. This directly satisfies the stem’s constraint that both variables are continuous, as scatter plots are the standard chart type for visualising relationships between two quantitative measures.

Why this answer

A scatter plot is the correct choice for visualizing the relationship between two continuous variables, such as advertising spend and website visits. It plots each observation as a point on a two-dimensional plane, allowing the analyst to observe correlation, trends, clusters, and outliers. This makes it ideal for bivariate continuous data analysis.

Exam trap

The trap is confusing charts for time-series (line chart) with charts for bivariate relationships; candidates may pick line chart because both variables are continuous, but the exam expects recognition that scatter plots are for relationships between two continuous variables.

How to eliminate wrong answers

Option A is wrong because a line chart is typically used to show trends over time or ordered categories, not the relationship between two independent continuous variables. Option C is wrong because a box plot summarizes the distribution of a single continuous variable (or compares distributions across categories), not the relationship between two continuous variables. Option D is wrong because a bar chart is used for comparing categorical data or aggregated values, not for showing correlation between two continuous variables.

777
MCQeasy

Which data quality dimension ensures that data represents the real-world scenario correctly and without errors?

A.Completeness
B.Consistency
C.Accuracy
D.Timeliness
AnswerC

Accuracy verifies that values correctly reflect the real-world entities or events they describe, directly satisfying the stem's requirement for data representing reality without errors. Unlike completeness, which checks for missing values, or consistency, which checks formatting agreement across sources, accuracy measures correctness against an authoritative reference.

Why this answer

Accuracy is the data quality dimension that measures whether data correctly reflects the real-world object or event it describes, free from errors. If a customer's address is stored incorrectly or a transaction amount is wrong, the data is inaccurate even if it is complete and consistent. Completeness, consistency, and timeliness address different aspects and do not capture correctness of representation.

Exam trap

DA0-002 often tests the overlap between accuracy and completeness, tricking candidates into choosing completeness when the scenario describes incorrect values rather than missing values.

How to eliminate wrong answers

Option A is wrong because completeness measures whether all required data is present — a record can be complete but still contain incorrect values. Option B is wrong because consistency measures whether data agrees across systems or formats — data can be consistently wrong in multiple places. Option D is wrong because timeliness measures whether data is available when needed and up to date — stale data can still be accurate for the period it represents.

778
MCQmedium

A data analyst is analyzing survey responses where respondents rated satisfaction on a scale of 1-5. The analyst wants to visualize the distribution of responses. Which chart type is most appropriate?

A.Box plot
B.Scatter plot
C.Line chart
D.Histogram
AnswerD

A histogram bins the numeric 1–5 ratings into adjacent intervals and plots frequency per bin, directly revealing the distribution's shape, spread and skew. Bar charts suit categorical counts, not ordered numeric scales, so they obscure distributional patterns the analyst needs.

Why this answer

A histogram is the most appropriate chart for visualizing the distribution of a single discrete variable, such as satisfaction ratings on a 1-5 scale. It groups the responses into bins (each rating value) and displays the frequency of each bin using bars, clearly showing the shape, central tendency, and spread of the data.

Exam trap

The trap here is that candidates often confuse a histogram with a bar chart, but the key distinction is that a histogram is used for quantitative (ordinal or continuous) data where bin order matters, while a bar chart is for categorical (nominal) data with no inherent order.

How to eliminate wrong answers

Option A is wrong because a box plot summarizes data using five-number statistics (min, Q1, median, Q3, max) and is better for comparing distributions across groups, not for showing the detailed frequency distribution of a single ordinal variable. Option B is wrong because a scatter plot is used to visualize the relationship between two continuous variables, not the distribution of a single categorical or ordinal variable. Option C is wrong because a line chart is typically used to display trends over time or sequential data, not the frequency distribution of discrete survey responses.

779
MCQhard

In A/B testing, which factor is increased by having a larger sample size?

A.P-value
B.Effect size
C.Type I error rate
D.Statistical power
AnswerD

Larger samples shrink the standard error of the estimated effect, which directly raises statistical power — the probability of detecting a true difference when one exists. This satisfies the scenario's need to distinguish genuine treatment effects from random variation.

Why this answer

Statistical power is the probability of correctly rejecting a false null hypothesis (i.e., detecting a true effect). Increasing sample size reduces the standard error, making it easier to detect a true effect and thus increasing power. This is a fundamental principle in hypothesis testing and A/B testing.

Exam trap

The trap is conflating p-value with power or thinking that larger samples increase effect size; candidates often pick p-value because they associate it with significance, but the exam expects understanding that power is the probability of detecting a true effect and is directly boosted by sample size.

How to eliminate wrong answers

Option A is wrong because the p-value is the probability of observing the data given that the null hypothesis is true; it is not directly increased by sample size—in fact, with larger samples, smaller effects can become statistically significant, but the p-value itself is a calculated value from the test, not a factor that increases. Option B is wrong because effect size is a measure of the magnitude of the phenomenon (e.g., Cohen's d) and is independent of sample size; it is a property of the population, not the study design. Option C is wrong because the Type I error rate (alpha) is the probability of rejecting a true null hypothesis and is set by the researcher (typically 0.05); it is not affected by sample size.

780
MCQeasy

A marketing analyst must combine two datasets: a CRM extract with one row per customer and a transactions extract with many rows per customer. The analyst wants every customer from the CRM to appear in the output, even customers with no matching transactions. Which join type should be used?

A.LEFT JOIN
B.CROSS JOIN
C.FULL OUTER JOIN
D.INNER JOIN
AnswerA

A left join keeps all rows from the left table (the CRM customers) and attaches matching transaction rows where they exist, producing NULLs for customers without transactions. This satisfies the requirement that every customer appears while still enriching those who have purchase activity. It is the standard pattern for preserving a master list during an enrichment join.

Why this answer

A left join anchors the result on the CRM customer list and enriches it with transaction data, retaining customers who never purchased. Unmatched customers receive NULL transaction fields, which is exactly the desired behavior for a master-list enrichment. Inner, full outer, and cross joins either drop customers, add unwanted orphans, or fabricate pairings.

Exam trap

The trap here is defaulting to an inner join out of habit and silently dropping the non-purchasing customers the analyst explicitly needs.

781
MCQhard

A financial analyst is creating a quarterly report compliant with SOX. Which requirement is most critical for the report's audit trail?

A.Row-level security
B.Audit trail of data changes
C.Data lineage
D.Data dictionary
AnswerB

SOX requires verifiable traceability of who changed financial data, when, and what the prior value was. An audit trail of data changes provides this immutable record, satisfying the compliance requirement for the quarterly report's audit trail.

Why this answer

SOX compliance requires that financial reports be traceable and that any change to reported data be attributable to a specific user at a specific time. An audit trail of data changes (B) provides exactly this: a tamper-evident record of who changed what, when, and why, which is the core control auditors test for financial reporting integrity. Without it, the report cannot be defended as accurate or complete during an audit.

Exam trap

DA0-002 often tests the confusion between lineage (where data came from) and audit trail (who changed it and when) — candidates pick lineage because it sounds like traceability, but SOX specifically demands change attribution.

How to eliminate wrong answers

Option A is wrong because row-level security controls who can see which rows — it is an access control mechanism, not a record of changes, and does nothing to prove data integrity over time. Option C is wrong because data lineage shows where data originated and how it flowed through transformations, which supports provenance but does not capture the who/when/why of individual changes that SOX auditors require. Option D is wrong because a data dictionary merely documents field definitions and metadata; it is descriptive reference material, not an evidentiary change record.

782
MCQmedium

A data analyst wants to assign a unique sequential integer to each row in a result set, starting at 1, based on the order of the 'sales_amount' column descending. Which window function should be used?

A.DENSE_RANK() OVER (ORDER BY sales_amount DESC)
B.RANK() OVER (ORDER BY sales_amount DESC)
C.NTILE(1) OVER (ORDER BY sales_amount DESC)
D.ROW_NUMBER() OVER (ORDER BY sales_amount DESC)
AnswerD

ROW_NUMBER() assigns a unique sequential integer to every row, starting at 1, with no ties or gaps. Ordering by sales_amount DESC satisfies the stem's ranking constraint, so the highest sale receives 1. RANK() and DENSE RANK() would repeat values for tied amounts, breaking the uniqueness requirement.

Why this answer

ROW_NUMBER() assigns a unique sequential integer to each row starting at 1, with no ties and no gaps, based on the ORDER BY clause. Since the requirement is a unique sequential integer per row ordered by sales_amount DESC, ROW_NUMBER() is the correct function.

Exam trap

DA0-002 often tests the difference between ROW_NUMBER(), RANK(), and DENSE_RANK() — candidates pick RANK() or DENSE_RANK() forgetting that only ROW_NUMBER() guarantees unique sequential integers.

How to eliminate wrong answers

Option A is wrong because DENSE_RANK() assigns the same rank to tied values and leaves no gaps (e.g., 1,1,2), so it does not guarantee a unique integer per row. Option B is wrong because RANK() also assigns the same rank to ties and leaves gaps (e.g., 1,1,3), violating the 'unique sequential integer' requirement. Option C is wrong because NTILE(1) divides the result set into 1 bucket, assigning 1 to every row — it does not produce sequential integers.

783
MCQhard

A multinational retailer stores customer transaction data in a cloud data warehouse hosted in the European Union. A new analytics team in Brazil needs to analyze this data, but legal counsel warns that transferring personal data out of the EU may violate GDPR. The company decides to keep the data in the EU region and give the Brazil team remote access through a governed query layer. Which data governance concept does this decision primarily address?

A.Data virtualization
B.Data partitioning
C.Data deduplication
D.Data sovereignty
AnswerD

Data sovereignty holds that data is subject to the laws and governance of the jurisdiction where it is stored or collected. Keeping EU customer data within the EU region and granting remote access rather than transferring it addresses the legal restriction on moving personal data across borders, which is the core of this concept.

Why this answer

Data sovereignty means data is governed by the laws of the jurisdiction where it resides. Because GDPR restricts transferring EU personal data to Brazil, keeping the data in the EU and providing governed remote access respects that legal boundary. The decision is fundamentally about jurisdiction over data, not about storage optimization or query performance.

Exam trap

The trap here is choosing data virtualization because it describes the access mechanism, while the governance concept that actually drives the decision is legal jurisdiction over the stored data.

784
MCQeasy

Which type of chart is most suitable for showing the distribution of a single continuous variable?

A.Histogram
B.Line chart
C.Scatter plot
D.Bar chart
AnswerA

A histogram bins a continuous variable into intervals along the x-axis and plots frequency on the y-axis, revealing shape, spread and modality. No other standard chart type represents a single continuous variable's distribution as directly.

Why this answer

A histogram is the most suitable chart for showing the distribution of a single continuous variable because it groups the data into bins (intervals) along the x-axis and displays the frequency or count of observations within each bin on the y-axis. This allows you to see the shape, spread, and central tendency of the data, such as whether it is normally distributed, skewed, or has multiple modes. For a single continuous variable, a histogram directly visualizes the probability density or frequency distribution, which is its primary purpose in data analysis.

Exam trap

The trap here is that candidates often confuse a histogram with a bar chart, assuming both are interchangeable for showing frequencies, but a histogram is specifically for continuous data with no gaps between bars, while a bar chart is for discrete categories with gaps.

How to eliminate wrong answers

Option B is wrong because a line chart is designed to show trends over time or ordered categories by connecting data points with lines, not to display the distribution of a single continuous variable. Option C is wrong because a scatter plot is used to show the relationship between two continuous variables by plotting points on a Cartesian plane, not to visualize the distribution of a single variable. Option D is wrong because a bar chart is used for categorical data, where each bar represents a distinct category, and it does not handle continuous variables properly since it would require arbitrary binning without the statistical properties of a histogram.

785
MCQhard

An organization has multiple systems that store customer information inconsistently. To create a single authoritative view of customer data, they implement a process that identifies and merges duplicate records. This is an example of which data management discipline?

A.Data governance
B.Data warehousing
C.Data quality
D.Master Data Management (MDM)
AnswerD

Master Data Management creates the single authoritative golden record by matching and merging duplicates across systems, exactly the discipline described. It governs the organisation's core customer entities, unlike data quality or integration, which address different concerns.

Why this answer

Master Data Management (MDM) is the discipline focused on creating and maintaining a single, authoritative, consistent view of core business entities such as customers, products, and suppliers. Identifying and merging duplicate records to produce a trusted golden record is a core MDM activity. The scenario describes exactly this consolidation of inconsistent customer data across systems.

Exam trap

DA0-002 often tests the overlap between data governance, data quality, and MDM, causing candidates to choose governance or quality when the scenario specifically requires creating a single authoritative master record.

How to eliminate wrong answers

Option A is wrong because data governance defines policies, standards, and stewardship for data, but does not itself perform record matching and merging to create a golden record. Option B is wrong because data warehousing is about consolidating data for analytics and reporting, not about establishing an authoritative operational master record. Option C is wrong because data quality focuses on accuracy, completeness, and consistency of data, but the specific goal of a single authoritative customer view through deduplication and merging is MDM.

786
Multi-Selectmedium

A data analyst is using a decision tree to classify loan applications as approved or denied. To avoid overfitting, the analyst plans to prune the tree. Which TWO of the following are valid pruning techniques? (Choose two.)

Select 2 answers
A.Post-pruning by removing branches that do not improve validation performance.
B.Adding more training data to the dataset.
C.Pre-pruning by limiting the maximum depth of the tree.
D.Increasing the number of features used at each split.
E.Using a random forest ensemble instead of a single decision tree.
AnswersA, C

Post-pruning involves growing the full tree and then removing branches that provide little improvement in validation performance. This technique, such as cost-complexity pruning, simplifies the tree and reduces overfitting. It is effective because it uses validation data to guide the pruning process, ensuring that only branches that contribute to predictive accuracy are retained.

Why this answer

Pruning techniques aim to simplify a decision tree to reduce overfitting. Pre-pruning stops tree growth early, for example by limiting maximum depth. Post-pruning removes branches after the tree is fully grown, using validation performance to guide removal.

Both are standard methods. Increasing features, adding data, or using ensembles are not pruning techniques for a single tree.

Exam trap

The trap here is confusing ensemble methods or data augmentation with pruning, which specifically modifies the tree structure.

787
Multi-Selecteasy

A data analyst needs to retrieve the top 5 most expensive products from a 'products' table sorted by price descending. Which TWO SQL clauses are required to achieve this? (Select TWO).

Select 2 answers
A.HAVING COUNT(*) > 1
B.WHERE price > 100
C.ORDER BY price DESC
D.GROUP BY price
E.LIMIT 5
AnswersC, E

ORDER BY price DESC sorts the result set by the price column in descending order, placing the most expensive products first. This directly satisfies the stem's requirement to rank products by price from highest to lowest, which the LIMIT clause then truncates to the top 5 rows.

Why this answer

ORDER BY price DESC (C) is required because it sorts the result set by the price column in descending order, placing the most expensive products first. LIMIT 5 (E) is required because it restricts the result set to only the first five rows returned after sorting, giving the top 5 most expensive products. Together, ORDER BY price DESC followed by LIMIT 5 produce exactly the requested output.

The other options do not belong: HAVING COUNT(*) > 1 (A) filters grouped aggregate results, WHERE price > 100 (B) applies a fixed numeric filter rather than selecting the top 5, and GROUP BY price (D) aggregates rows by price instead of simply sorting and limiting them.

788
MCQmedium

A table named Orders has columns OrderID, CustomerID, OrderDate, and TotalAmount. Which column should be the primary key to uniquely identify each order?

A.OrderDate
B.OrderID
C.TotalAmount
D.CustomerID
AnswerB

OrderID is a surrogate identifier assigned one distinct value per order row, so it guarantees uniqueness and rejects nulls, satisfying the requirement to identify each order individually. CustomerID repeats across a customer's orders, OrderDate collides on same-day purchases, and TotalAmount duplicates readily, so none can enforce entity integrity as the primary key.

Why this answer

The OrderID column is the correct choice for the primary key because it contains unique values for each order, ensuring that each row can be uniquely identified. A primary key must be unique, non-null, and stable; OrderID satisfies all these requirements, whereas the other columns do not guarantee uniqueness or are subject to change.

Exam trap

The trap here is that candidates may confuse a column that is frequently used for filtering or grouping (like CustomerID or OrderDate) with one that guarantees uniqueness, overlooking the fundamental primary key requirement of uniqueness and non-nullability.

How to eliminate wrong answers

Option A is wrong because OrderDate is not unique; multiple orders can occur on the same date, and it can also be null, violating primary key constraints. Option C is wrong because TotalAmount can have duplicate values (e.g., two orders with the same total) and is not inherently unique or stable. Option D is wrong because CustomerID is not unique per order; a single customer can place many orders, so it cannot uniquely identify each order row.

789
Multi-Selecteasy

Which TWO of the following are examples of data transformation? (Choose TWO.)

Select 2 answers
A.Normalizing data to eliminate redundancy
B.Creating a backup of the database
C.Converting string dates to date format
D.Generating summary statistics
E.Removing duplicate records
AnswersA, C

Normalization is a transformation.

Why this answer

Data normalization is a transformation process that reorganizes data to reduce redundancy and improve integrity, typically by decomposing tables into smaller, related tables (e.g., achieving 3NF in relational databases). This changes the structure and representation of the data, which is a core example of data transformation.

Exam trap

CompTIA often tests the distinction between data transformation (changing format/structure) and data cleansing (removing errors/duplicates) or data analysis (generating summaries), leading candidates to mistakenly select removal of duplicates or summary statistics as transformations.

790
MCQeasy

A data analyst is presenting findings to a non-technical audience. Which data visualization technique is most appropriate?

A.Box plot
B.Heat map
C.Scatter plot
D.Bar chart
AnswerD

Bar charts encode magnitude by length, which people compare accurately without training, unlike pie slices or scatter plots. This satisfies the stem's non-technical audience constraint by using the most universally readable encoding for categorical comparison.

Why this answer

Bar charts are the most appropriate visualization for a non-technical audience because they present categorical comparisons in a simple, intuitive format that requires no statistical background to interpret. Unlike box plots, heat maps, or scatter plots, bar charts clearly show relative magnitudes with minimal cognitive load, making them ideal for communicating data insights to stakeholders who may not be familiar with data analysis conventions.

Exam trap

The trap here is that candidates often choose scatter plots or box plots because they seem more 'analytical,' but the question specifically targets audience comprehension, not analytical depth — CompTIA Data+ tests whether you prioritize clarity over complexity when the audience lacks technical background.

How to eliminate wrong answers

Option A is wrong because box plots display data distribution through quartiles and outliers, which require understanding of statistical concepts like median, interquartile range, and whiskers — concepts that are not intuitive for a non-technical audience. Option B is wrong because heat maps use color gradients to represent data density or correlation, which can be ambiguous without a legend and are better suited for identifying patterns in large matrices rather than straightforward comparisons. Option C is wrong because scatter plots show relationships between two continuous variables and rely on the audience's ability to interpret correlation or clustering, which is not immediately accessible to non-technical viewers.

791
MCQeasy

A marketing team wants to use customer email addresses collected for newsletter subscriptions to send third-party promotional offers. The data governance policy states that data must be used only for the purpose for which it was collected. Which action should the data governance team take?

A.Deny the use because email addresses are considered sensitive personal data under all regulations.
B.Allow the use if the third-party offers are related to the newsletter content.
C.Deny the use because it violates the purpose limitation principle unless new consent is obtained.
D.Allow the use because the email addresses are already in the company's possession.
AnswerC

Purpose limitation requires that personal data be used only for the purposes specified at collection. Using email addresses for third-party promotions exceeds the original newsletter subscription purpose. The governance team should deny the request unless the individuals provide explicit consent for the new purpose, which may also require updating the privacy notice.

Why this answer

Purpose limitation is a core data governance principle: personal data should be collected for specified, explicit, and legitimate purposes and not further processed in a manner incompatible with those purposes. Using newsletter subscription emails for third-party promotions is a new purpose that requires a new lawful basis, typically explicit consent. The governance team should deny the request unless proper consent is obtained and the privacy notice is updated.

Exam trap

The trap here is thinking that because the company already holds the data, it can use it for any purpose without additional consent.

792
MCQhard

A data analyst discovers a data discrepancy in a report that has already been distributed to executives. The discrepancy is due to a known bug that will be fixed next month. What is the most appropriate action?

A.Send a corrected report immediately with an apology
B.Wait for the next scheduled update
C.Ignore it since it will be fixed soon
D.Inform the executives of the error and its impact
AnswerD

Executives are making decisions on flawed figures, so proactively disclosing the error, its cause, and the expected fix date preserves trust and lets them correct course. Waiting until next month's fix leaves the inaccurate report unchallenged.

Why this answer

The most appropriate action is to inform the executives of the error and its impact because transparency and accountability are core ethical responsibilities for data analysts, especially when decisions may have been made based on flawed data. Executives need to understand the nature of the discrepancy, its potential consequences, and the timeline for a fix so they can assess risks and take corrective action. Simply sending a corrected report without context (Option A) may cause confusion or fail to address the impact, while waiting or ignoring (Options B and C) violates professional integrity and could lead to poor business decisions.

Exam trap

DA0-002 often tests the distinction between technical fixes and professional communication, trapping candidates who focus solely on correcting the data (Option A) rather than addressing the broader need for transparency and impact analysis (Option D).

How to eliminate wrong answers

Option A is wrong because sending a corrected report immediately with an apology, while well-intentioned, does not adequately explain the error's impact or the reason for the discrepancy, and may not give executives the full context needed to assess the situation. Option B is wrong because waiting for the next scheduled update delays critical information, allowing executives to continue using flawed data for decisions, which could compound the problem. Option C is wrong because ignoring the error, even if it will be fixed soon, is unethical and irresponsible; it undermines trust and could lead to significant negative consequences if the discrepancy affects key decisions.

793
MCQmedium

A data analyst is working with a dataset that contains customer names and addresses. Some records have missing state codes. Which data quality issue is this?

A.Duplication
B.Incompleteness
C.Outliers
D.Inconsistency
AnswerB

Incompleteness captures missing values within otherwise present records, which matches the absent state codes precisely. Unlike inaccuracy, which concerns incorrect values, or duplication, incompleteness addresses the null entries the analyst must resolve before geographic analysis. This satisfies the stem's constraint of records lacking required state data.

Why this answer

Incompleteness is the correct answer because missing state codes in customer address records represent a lack of required data. This is a classic example of incomplete data, where fields that should contain values are left null or blank, reducing the dataset's usability for analysis.

Exam trap

The trap here is that candidates may confuse incompleteness with inconsistency, but incompleteness is about missing data (nulls), while inconsistency is about contradictory data across records.

How to eliminate wrong answers

Option A is wrong because duplication refers to duplicate records (e.g., same customer appearing multiple times), not missing values. Option C is wrong because outliers are data points that deviate significantly from the norm (e.g., an unusually high age), not absent data. Option D is wrong because inconsistency involves contradictory or conflicting data (e.g., same customer with different state codes in different records), not missing values.

794
MCQmedium

A financial services firm stores customer records in a cloud data warehouse. A new analyst requests read access to a table containing account balances and Social Security numbers. The data governance team must decide how to grant access while following the principle of least privilege. Which action best satisfies this requirement?

A.Grant the analyst read access to the full table but require the analyst to sign an acceptable use policy.
B.Grant the analyst read access to the full table because the analyst is a full-time employee.
C.Grant the analyst read access only to a view that excludes Social Security numbers and contains account balances.
D.Grant the analyst write access to the table so they can correct any data errors they find.
AnswerC

This follows least privilege by providing only the columns required for the analyst's task while masking sensitive identifiers. A view enforces column-level security consistently and prevents direct access to the underlying table. It also supports data minimization and reduces the risk of accidental exposure.

Why this answer

Least privilege means granting only the minimum access necessary to perform a task. Since the analyst needs account balances but not Social Security numbers, a view that excludes the sensitive column is the appropriate control. This approach enforces column-level security, supports data minimization, and reduces the risk of exposing personally identifiable information while still enabling the required analysis.

Exam trap

The trap here is assuming that a signed policy or employee status can substitute for technical access controls that enforce least privilege.

795
Multi-Selectmedium

Which TWO chart types are appropriate for visualizing the distribution of a continuous variable? (Select two.)

Select 2 answers
A.Bar chart
B.Box plot
C.Line chart
D.Histogram
E.Pie chart
AnswersB, D

A box plot summarises a continuous variable's distribution through its median, quartiles and whiskers, directly exposing spread, skew and outliers. This satisfies the stem's requirement to visualise distribution rather than compare categories or show trends over time.

Why this answer

A box plot (B) is appropriate for visualizing the distribution of a continuous variable because it displays the median, quartiles, and potential outliers, providing a five-number summary of the data. A histogram (D) is also appropriate as it groups continuous data into bins and shows the frequency distribution, revealing the shape, central tendency, and spread of the variable.

Exam trap

The trap here is that candidates often confuse a bar chart with a histogram, thinking both can show distribution, but a bar chart is for categorical data while a histogram is specifically for continuous data with no gaps between bars.

796
MCQhard

A data engineer is ingesting a 40 GB JSON event log into a columnar analytics platform. The file contains deeply nested arrays of user actions, and queries only ever filter on three top-level fields: event_id, event_type, and event_timestamp. The ingestion is currently slow and queries scan excessive data. Which preparation approach is MOST appropriate?

A.Load the JSON as a single string column and rely on the query engine to parse it at read time.
B.Flatten the top-level fields into typed columns, extract the nested arrays into a separate child table, and partition or cluster on event_timestamp.
C.Store the file in a row-oriented relational table with indexes on all nested array paths.
D.Normalize every nested array element into its own row in a fully relational schema with foreign keys.
AnswerB

Extracting the frequently filtered scalars into native typed columns lets the columnar engine read only those columns and apply partition pruning on event_timestamp. Moving the rarely queried nested arrays into a child table keeps the main fact table narrow and fast. This matches the access pattern precisely and reduces both ingestion cost and scan volume.

Why this answer

Matching the storage layout to the access pattern is the decisive factor. Promoting the three filtered scalars to typed columns enables column pruning and partition pruning, while relocating the unused nested arrays keeps the hot table narrow. Parsing strings or fully normalizing both impose costs the workload does not justify.

Exam trap

The trap here is treating JSON flattening as all-or-nothing and either parsing at query time or fully normalizing, instead of selectively promoting the columns the workload actually filters on.

797
MCQeasy

A retail company wants to analyze customer purchase patterns to identify products frequently bought together. Which data mining technique is most appropriate?

A.Classification
B.Clustering
C.Regression
D.Association rules
AnswerD

Association rules discover co-occurrence relationships between items in transactional data, directly satisfying the requirement to identify products frequently bought together. Algorithms such as Apriori and FP-Growth generate rules like {bread} → {butter}, quantified by support, confidence and lift, which is precisely the market-basket analysis this retail scenario demands.

Why this answer

Association rules are specifically designed to uncover relationships between items in transactional datasets, such as 'customers who buy X also buy Y.' This technique generates rules like {bread, butter} → {milk} with metrics such as support, confidence, and lift, directly answering the question of which products are frequently bought together. Classification, clustering, and regression serve different purposes: they predict labels, group similar instances, or model continuous relationships, respectively. Therefore, association rules are the most appropriate choice for market basket analysis.

Exam trap

The trap here is confusing association rules with clustering because both are unsupervised and used for pattern discovery, but clustering groups similar items while association rules find co-occurrence relationships between items in transactions.

How to eliminate wrong answers

Option A is wrong because classification is a supervised learning technique used to assign predefined labels to records (e.g., spam/not spam), not to discover co-occurrence patterns among items. Option B is wrong because clustering is an unsupervised technique that groups similar data points based on distance metrics, but it does not identify 'if-then' relationships between products in transactions. Option C is wrong because regression is used to predict a continuous numeric value (e.g., sales amount) based on independent variables, not to find frequent itemsets or association rules.

798
MCQhard

A data analyst is preparing to acquire clickstream data from a web analytics platform via its REST API. The API returns paginated results with a maximum of 500 records per page and issues a short-lived bearer token that expires after one hour. The analyst needs to backfill six months of event data reliably. Which acquisition design is most appropriate?

A.Use the API's default page size and restart the entire backfill from the first page whenever the token expires.
B.Request the maximum page size and loop through pages, refreshing the bearer token before expiry and persisting a cursor or timestamp checkpoint after each successful page.
C.Request one record per page to avoid hitting rate limits, and store the token in the script so it can be reused across runs.
D.Pull all data in a single long-running request and write the response to disk when the connection closes.
AnswerB

Pagination with the maximum page size minimizes request count, token refresh prevents mid-run authentication failures, and a persisted checkpoint allows the job to resume without re-pulling data after an interruption. This design is resilient to the two stated constraints and avoids duplicate or missing records. Checkpointing on a timestamp or cursor also supports incremental runs beyond the initial backfill.

Why this answer

The API imposes two constraints: pagination and short-lived authentication. The sound design respects both by using the largest allowed page, proactively refreshing the token, and checkpointing after each page so progress survives interruptions. This yields efficient, resumable, and duplicate-avoiding acquisition.

Approaches that use tiny pages, single long requests, or full restarts on token expiry either trigger rate limits, cannot complete, or repeatedly re-download data.

Exam trap

The trap here is treating token expiry and pagination as separate concerns, when a robust backfill must handle both together with checkpointing.

799
MCQeasy

A data analyst needs to retrieve only unique job titles from the 'employees' table. Which SQL keyword should be used in the SELECT clause?

A.TOP
B.DISTINCT
C.UNIQUE
D.FILTER
AnswerB

DISTINCT in the SELECT clause removes duplicate rows from the result set, so repeated job titles collapse to a single occurrence. This satisfies the requirement to retrieve only unique job titles from the employees table.

Why this answer

The DISTINCT keyword in the SELECT clause removes duplicate rows from the result set, so 'SELECT DISTINCT job_title FROM employees' returns each unique job title only once. It is the standard SQL keyword defined in the SQL standard for deduplication of result rows. DISTINCT applies to the entire selected column list, not just one column.

Exam trap

The trap is confusing the UNIQUE constraint (a DDL keyword) with DISTINCT (a DML query keyword), leading candidates to pick UNIQUE for deduplication in a SELECT statement.

How to eliminate wrong answers

Option A (TOP) is wrong because TOP limits the number of rows returned (e.g., TOP 10) and does not remove duplicates. Option C (UNIQUE) is wrong because UNIQUE is a constraint used in table definitions (CREATE TABLE / ALTER TABLE) to enforce uniqueness, not a SELECT clause keyword for deduplication. Option D (FILTER) is wrong because FILTER is used with aggregate functions (e.g., COUNT(*) FILTER (WHERE ...)) to conditionally include rows in an aggregate, not to deduplicate result rows.

800
MCQeasy

A data analyst is creating a visualization to show the relationship between two continuous variables: advertising spend and sales revenue. The analyst wants to display the strength and direction of the linear relationship. Which type of chart should the analyst use?

A.Pie chart
B.Bar chart
C.Scatter plot
D.Histogram
AnswerC

A scatter plot displays individual data points on a two-dimensional plane, with one variable on the x-axis and the other on the y-axis. It allows the analyst to visually assess the strength, direction, and form of the relationship between two continuous variables. It is the standard tool for examining correlation and identifying patterns or outliers.

Why this answer

A scatter plot is the appropriate visualization for examining the relationship between two continuous variables. It plots each observation as a point, allowing the analyst to see if there is a positive, negative, or no correlation, and to spot outliers. Bar charts, pie charts, and histograms are designed for categorical or univariate data and cannot effectively display bivariate relationships.

Exam trap

The trap here is confusing a histogram, which shows one variable's distribution, with a scatter plot, which shows the relationship between two variables.

801
MCQmedium

A retail analytics team needs to load a 40 GB CSV file of point-of-sale transactions into a cloud data warehouse nightly. The file is generated as a single object by an upstream system, and the team must minimize load time. Which approach best addresses the load performance bottleneck?

A.Split the file into multiple smaller files and load them in parallel.
B.Convert the CSV to JSON and load it as a semi-structured format.
C.Compress the file using gzip and load the single compressed file.
D.Load the file into a staging table using row-by-row inserts.
AnswerA

Splitting the large CSV into multiple smaller files allows the data warehouse to ingest them concurrently, dramatically reducing overall load time. Parallel ingestion is a standard best practice for bulk loading large datasets, as it leverages distributed compute resources and avoids single-threaded bottlenecks. This directly addresses the performance issue without altering the data content.

Why this answer

The correct approach is to split the large file into multiple smaller files and load them in parallel. This leverages the distributed architecture of modern data warehouses, enabling concurrent ingestion and significantly reducing overall load time. Other options either do not address the parallelism bottleneck or introduce additional processing overhead that worsens performance.

Exam trap

The trap here is assuming that compressing the file alone will solve the load time issue, but compression only reduces transfer size and does not enable parallel processing.

802
Multi-Selectmedium

Which THREE of the following are common steps in data cleaning?

Select 3 answers
A.Removing outliers without justification
B.Imputing missing values
C.Standardizing data formats
D.Removing duplicate records
E.Increasing sample size
AnswersB, C, D

Missing values are often imputed to maintain dataset completeness.

Why this answer

Imputing missing values is a common data cleaning step because real-world datasets often have gaps due to data collection errors or system failures. Techniques like mean/median imputation, regression imputation, or using algorithms like k-NN help preserve sample size and avoid bias that would result from simply dropping rows. This ensures the dataset remains usable for analysis without introducing significant distortion.

Exam trap

CompTIA often tests the distinction between data cleaning steps and data collection or preprocessing steps, so the trap here is confusing 'increasing sample size' (a data augmentation or collection activity) with actual cleaning tasks like imputation, standardization, and deduplication.

803
MCQmedium

A data analyst is designing a dashboard and wants to maximize the data-ink ratio. Which action aligns with this principle?

A.Removing unnecessary gridlines
B.Using 3D effects for bars
C.Adding a background image
D.Using bright colors for all elements
AnswerA

Gridlines add non-data ink that consumes pixels without conveying values, so removing them raises the proportion of ink devoted to actual data. This directly increases the data-ink ratio, satisfying the dashboard design constraint stated in the stem.

Why this answer

The data-ink ratio, popularized by Edward Tufte, measures the proportion of ink used to convey actual data versus decorative or redundant elements. Removing unnecessary gridlines increases the ratio because gridlines add visual clutter without conveying data values, letting the data marks dominate the visual field. This is a core principle in dashboard design for clarity and cognitive efficiency.

Exam trap

The trap here is confusing 'minimalism' with 'data-ink ratio' — candidates may pick any visually simple option, but the principle specifically targets ink that does not encode data, so removing gridlines is correct while removing data labels would not be.

How to eliminate wrong answers

Option B is wrong because 3D effects add perspective distortion and decorative shading that reduce the data-ink ratio and can mislead viewers about relative magnitudes. Option C is wrong because a background image is pure decoration that competes with the data for visual attention and lowers the ratio. Option D is wrong because using bright colors for all elements creates visual noise and eliminates the ability to use color as an encoding channel, which reduces effective data-ink rather than maximizing it.

804
MCQhard

A data team is preparing a dashboard for executives. The team wants to highlight key performance indicators (KPIs) that are below target. Which of the following visualization techniques would most effectively draw attention to underperforming metrics without causing confusion?

A.Remove underperforming KPIs from the dashboard to avoid confusion.
B.Use a scatter plot to show the relationship between KPIs.
C.Apply conditional formatting to turn KPI values red when below target.
D.Use a pie chart showing the proportion of each KPI.
AnswerC

Conditional formatting applies a rule that automatically renders KPI values red when they fall below target, drawing attention through colour without altering the underlying data or adding extra chart elements. This satisfies the constraint of highlighting underperformance clearly while avoiding visual confusion.

Why this answer

Conditional formatting that turns KPI values red when below target is the most effective technique because it leverages pre-attentive visual processing — the human eye naturally notices color changes (especially red) before other visual elements. This allows executives to instantly identify underperforming metrics without needing to interpret complex chart types, reducing cognitive load and confusion.

Exam trap

CompTIA often tests the misconception that removing or hiding negative data simplifies communication, when in fact effective dashboards must highlight both successes and failures to drive data-driven decisions.

How to eliminate wrong answers

Option A is wrong because removing underperforming KPIs from the dashboard hides critical negative information, violating the principle of transparency in data communication and preventing executives from taking corrective action. Option B is wrong because a scatter plot is designed to show relationships or correlations between two variables, not to highlight individual KPI performance against a target; it would introduce unnecessary complexity and fail to draw immediate attention to underperforming metrics. Option D is wrong because a pie chart shows proportions of a whole, which is irrelevant for comparing KPIs against targets; it cannot effectively indicate which metrics are below target and often leads to misinterpretation due to difficulty in comparing angles and areas.

805
MCQhard

A data analyst is building a dashboard that will be used by both executives and operational managers. The executives need high-level summaries, while managers need granular details. Which dashboard design principle should the analyst apply?

A.Use a single chart that shows both summary and detail simultaneously.
B.Display all available data on one page for transparency.
C.Design the dashboard with drill-down capabilities from summary to detail.
D.Create two separate dashboards for each audience.
AnswerC

Drill-down lets one dashboard serve both audiences: executives view aggregated summaries, then navigate to granular detail on demand. This satisfies the stem's dual requirement without duplicating reports, since a single semantic layer filters progressively by dimension. Static summary-only or detail-only designs would fail one group.

Why this answer

Drill-down capabilities allow users to start with a high-level summary (e.g., total revenue by region) and then interactively navigate to granular details (e.g., individual transactions) without overwhelming either audience. This design principle supports both executive and operational manager needs within a single dashboard, maintaining clarity and performance by loading only the required level of detail on demand.

Exam trap

The trap here is that candidates often confuse 'drill-down' with 'separate dashboards' (Option D), thinking that different audiences require entirely different views, when in fact a single interactive dashboard with drill-down capabilities is the standard best practice for serving both summary and detail needs efficiently.

How to eliminate wrong answers

Option A is wrong because a single chart attempting to show both summary and detail simultaneously would likely become cluttered, violate data-ink ratio principles, and fail to provide clear high-level insights for executives or actionable granularity for managers. Option B is wrong because displaying all available data on one page creates information overload, reduces readability, and violates the principle of progressive disclosure, making it difficult for both audiences to quickly find relevant insights. Option D is wrong because creating two separate dashboards introduces maintenance overhead, potential data inconsistency, and forces users to switch contexts, whereas a single dashboard with drill-down capabilities is more efficient and user-friendly.

806
Matchingmedium

Match each data security concept to its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Convert data into code to prevent unauthorized access

Replace sensitive data with non-sensitive placeholder

Hide parts of data (e.g., show only last 4 digits)

Restrict data access based on user permissions

Log of all data access and modifications

Why these pairings

The correct matches are: Encryption - encoding data, Authentication - verifying identity, Authorization - access control, Non-repudiation - denying actions. Common confusions arise from swapping Encryption and Authentication.

807
MCQhard

A company is designing a data lake to store raw sensor data from IoT devices. The data arrives as JSON objects with varying schemas. Which storage approach is most appropriate?

A.Ingest into a relational database with a predefined schema
B.Store each JSON object as a separate file in a compressed columnar format
C.Convert all JSON to Avro with a fixed schema before storing
D.Store raw JSON files in a distributed file system and apply schema-on-read
AnswerD

Schema-on-read defers parsing until query time, so each JSON object's varying structure is interpreted individually rather than forced into a fixed table. This directly satisfies the stem's requirement to store raw sensor data whose schemas differ across IoT devices, avoiding ingestion-time transformation failures.

Why this answer

A data lake is designed to store raw data in its native format, and IoT sensor data with varying schemas is best handled by storing raw JSON files in a distributed file system (e.g., HDFS or Amazon S3). This approach leverages schema-on-read, where the schema is applied at query time rather than at write time, allowing flexibility for heterogeneous JSON objects without data loss or transformation overhead.

Exam trap

The trap here is that candidates confuse 'schema-on-read' with 'schema-on-write' and assume that converting to a structured format like Avro or columnar storage is always better for performance, ignoring the requirement to store raw, varying-schema data as-is.

How to eliminate wrong answers

Option A is wrong because relational databases require a predefined schema and enforce ACID constraints, which cannot accommodate JSON objects with varying schemas without costly schema migrations or data loss. Option B is wrong because storing each JSON object as a separate file in a compressed columnar format (e.g., Parquet or ORC) is inefficient for small, variable-schema records; columnar formats are optimized for analytical queries on large, homogeneous datasets, not for raw ingestion of many small, schema-varying JSON objects. Option C is wrong because converting all JSON to Avro with a fixed schema before storing defeats the purpose of a data lake, which is to preserve raw data; Avro requires a predefined schema at write time, and forcing a fixed schema on varying JSON objects would either lose data or require complex schema evolution management.

808
MCQeasy

Which data quality dimension is violated if a customer record has a missing phone number?

A.Consistency
B.Accuracy
C.Completeness
D.Validity
AnswerC

Completeness measures whether all required values are present in a record. A missing phone number means the attribute holds no value, directly violating this dimension. Other dimensions, such as accuracy or consistency, assume a value exists but is wrong or conflicting; here nothing is recorded at all.

Why this answer

Completeness refers to the extent to which data is not missing.

809
MCQmedium

A data analyst needs to combine rows from two tables based on a related column, but only wants rows that have matching values in both tables. Which join type should the analyst use?

A.RIGHT JOIN
B.INNER JOIN
C.FULL OUTER JOIN
D.LEFT JOIN
AnswerB

An INNER JOIN returns only rows where the join key exists in both tables, discarding unmatched rows from either side. This directly satisfies the requirement for rows with matching values in both tables, unlike outer joins that retain unmatched rows.

Why this answer

An INNER JOIN returns only the rows where the join key exists in both tables, which is exactly what the analyst needs when they want matched records only. Rows with no counterpart in either table are excluded from the result set. This is the default and most common join type in SQL.

Exam trap

DA0-002 often tests the difference between INNER JOIN (matches only) and OUTER JOINs (include unmatched rows with NULLs), so candidates who default to LEFT JOIN for 'combining tables' pick the wrong answer.

How to eliminate wrong answers

Option A is wrong because a RIGHT JOIN returns all rows from the right table plus matching rows from the left, including unmatched right-side rows with NULLs — it does not restrict to matches only. Option C is wrong because a FULL OUTER JOIN returns all rows from both tables, filling unmatched sides with NULLs, which is the opposite of 'matching values in both tables.' Option D is wrong because a LEFT JOIN returns all rows from the left table plus matches from the right, again including unmatched left-side rows.

810
MCQhard

A large e-commerce company has a data warehouse that stores daily sales transactions. The data analytics team creates a weekly report for the executive team that includes total revenue, number of orders, and average order value. Recently, the report showed a sudden 20% drop in total revenue compared to the previous week. Upon investigation, the data analyst discovers that a major promotional event ended the previous week, and the current week's data includes a new product launch with lower prices. The analyst needs to present this data to the executives in a way that accurately reflects the business situation. The analyst must also provide recommendations for actions. Which of the following is the best course of action?

A.Remove the new product launch data from the report to show a clearer comparison.
B.Present only the data from the new product launch to show growth potential.
C.Present the data with annotations explaining the promotion end and new product launch, and recommend analyzing customer retention and cross-selling opportunities.
D.Present the drop as-is and recommend cutting costs to compensate for the revenue decline.
AnswerC

Annotating the promotion end and product launch preserves the revenue drop's context, satisfying the need to reflect the business situation accurately rather than implying performance failure. Recommending retention and cross-selling analysis addresses the underlying cause: lower-priced launch items diluting average order value, so executives receive both explanation and actionable next steps.

Why this answer

It adheres to the principle of transparent data communication: the analyst must present the full dataset with contextual annotations (promotion end, new product launch) to avoid misleading executives. This approach enables informed decision-making, and the recommendation to analyze customer retention and cross-selling opportunities directly addresses the underlying business dynamics—understanding whether the revenue drop is temporary or signals a shift in customer behavior. In a data warehouse context, this ensures that the weekly report reflects the true state of the business rather than a manipulated or incomplete view.

Exam trap

The trap here is that candidates may think removing or cherry-picking data (Options A and B) simplifies the story for executives, but the exam tests the ethical and professional standard of presenting complete, contextualized data—even when it shows a negative trend—while offering actionable recommendations based on the full picture.

How to eliminate wrong answers

Option A is wrong because removing the new product launch data constitutes data manipulation that violates the integrity of the report; it would hide a significant business event and prevent executives from understanding the full revenue picture. Option B is wrong because presenting only the new product launch data is a form of selection bias that ignores the overall revenue decline and the impact of the promotion end, leading to an overly optimistic and incomplete analysis. Option D is wrong because presenting the drop as-is without context or analysis fails to provide actionable insights; recommending cost cuts without understanding the root cause (promotion end and new product pricing) could lead to misguided strategic decisions.

811
MCQhard

A financial analyst publishes a monthly report where a gauge chart shows the current debt-to-equity ratio against a target band. Executives repeatedly misread the gauge, assuming the needle's position alone indicates good or bad performance without checking the target thresholds. Which change would BEST improve accurate interpretation?

A.Change the gauge's color zones so that the entire arc turns red whenever the ratio exceeds the target, regardless of how far.
B.Replace the gauge with a pie chart showing the debt and equity components as two slices of total capital.
C.Replace the gauge with a bullet chart showing the current value as a bar, the target as a marker, and qualitative ranges as background bands.
D.Keep the gauge but enlarge it and add more tick labels around the arc so the needle's position is easier to read precisely.
AnswerC

A bullet chart encodes the same three elements more compactly and unambiguously: the bar gives the actual value, the target marker gives the benchmark, and the shaded bands give qualitative ranges. Because the value and target share one linear scale, executives can judge performance directly instead of inferring it from a needle's angle, which is what caused the misreading.

Why this answer

The executives are misreading the gauge because a needle's angle conveys no inherent good or bad without the reader mentally consulting the target, which they are skipping. A bullet chart solves this by placing the actual value, the target marker, and qualitative bands on one shared linear scale, so the comparison is explicit rather than inferred. Bigger gauges, binary color states, and pie charts each fail to make the target comparison legible.

Exam trap

The trap here is assuming the gauge is misread only because it is small or imprecise, when the real failure is that a needle angle never encodes the target benchmark the audience ignores.

812
MCQmedium

A data engineer needs to store logs from web servers that have varying fields. The logs are in JSON format. Which data type describes this JSON data?

A.Binary data
B.Structured data
C.Semi-structured data
D.Unstructured data
AnswerC

JSON logs have no fixed schema, yet carry tags and hierarchy, which defines semi-structured data. Unlike structured data in rigid relational tables, it permits varying fields per record while remaining machine-parseable, matching the web server logs described.

Why this answer

JSON data with varying fields is classified as semi-structured data because it has organizational properties (key-value pairs, nested structures) but does not conform to a rigid schema like a relational table. The logs from web servers may have different fields per record, which is a hallmark of semi-structured data, as it allows flexibility while still being self-describing.

Exam trap

The trap here is that candidates confuse 'structured' with any data that has a format, but JSON's lack of a fixed schema and varying fields disqualifies it from being structured data, which requires a rigid, predefined schema like a relational database table.

How to eliminate wrong answers

Option A is wrong because binary data refers to raw bytes or encoded formats (e.g., images, executables) that lack any inherent structure or human-readable format, whereas JSON is text-based and has explicit key-value organization. Option B is wrong because structured data requires a fixed schema with predefined fields and data types (e.g., rows in a SQL table), but JSON logs with varying fields violate this strict schema requirement. Option D is wrong because unstructured data has no predefined format or organization (e.g., plain text, video files), while JSON has a defined syntax with keys, values, and nesting, providing a clear structure.

813
MCQhard

A data analyst is presenting a story about declining sales. The narrative arc should include which three elements in order?

A.Resolution → Complication → Situation
B.Situation → Complication → Resolution
C.Situation → Resolution → Complication
D.Complication → Situation → Resolution
AnswerB

This three-part arc first establishes the stable business context, then introduces the sales decline as the disrupting complication, then closes with the recommended resolution, giving the audience a logical causal progression from problem to action.

Why this answer

A compelling data story about declining sales follows the classic narrative arc: Situation (establish context, e.g., 'Sales were steady in Q1'), Complication (introduce the conflict, e.g., 'Then a 20% drop occurred in Q2'), and Resolution (present the insight or action, e.g., 'We identified the cause and implemented a new pricing strategy'). This order mirrors the 'Situation-Complication-Resolution' framework used in data storytelling to guide the audience logically from context to problem to solution.

Exam trap

The trap here is that candidates often confuse the narrative order with a simple 'problem-solution' structure, mistakenly placing Complication first (Option D) or skipping the Situation entirely, but the exam requires the full Situation → Complication → Resolution sequence to ensure a complete and logical data story.

How to eliminate wrong answers

Option A is wrong because starting with Resolution (the solution) before establishing the Situation or Complication confuses the audience; they need context first to understand why the resolution matters. Option C is wrong because placing Resolution before Complication skips the core conflict that drives the narrative, making the story feel incomplete and the resolution unsupported. Option D is wrong because beginning with Complication without first setting the Situation leaves the audience without necessary background, making the problem seem arbitrary or disconnected from the data.

814
MCQeasy

A data analyst needs to present the distribution of customer ages to a non-technical audience. Which type of chart would be most appropriate?

A.Scatter plot
B.Histogram
C.Pie chart
D.Line chart
AnswerB

A histogram bins continuous numeric values such as ages into intervals, revealing the distribution's shape, peaks and spread. This suits a non-technical audience because the adjacent bars show frequency patterns that summary statistics alone would obscure.

Why this answer

A histogram is the most appropriate chart for displaying the distribution of a single continuous variable, such as customer ages, to a non-technical audience. It groups ages into bins and shows the frequency of customers within each bin, making the shape, center, and spread of the distribution immediately visible. This aligns with the DA0-001 objective of selecting visualizations that clearly communicate data insights to stakeholders.

Exam trap

The trap here is that candidates often confuse a histogram with a bar chart, thinking a bar chart can show distribution, but a bar chart is for categorical data, not continuous numeric data grouped into bins.

How to eliminate wrong answers

Option A is wrong because a scatter plot is used to show the relationship between two continuous variables, not the distribution of a single variable like age. Option C is wrong because a pie chart is designed for parts-of-a-whole categorical data, not for displaying the frequency distribution of a continuous numeric variable. Option D is wrong because a line chart is typically used to show trends over time or ordered categories, not the distribution of a single continuous variable.

815
MCQmedium

A data analyst is creating a dashboard and wants to maximize the data-ink ratio. Which action supports this principle?

A.Including detailed data tables alongside charts.
B.Removing gridlines and reducing chart borders.
C.Using 3D effects to make bars stand out.
D.Adding a background image to make the dashboard visually appealing.
AnswerB

Removing gridlines and shrinking chart borders eliminates non-data pixels, directly raising the data-ink ratio. Tufte's principle counts every drop of ink that does not encode a data value as waste, so stripping decorative framing and background rules satisfies the stem's constraint of maximising data-ink without altering the plotted values themselves.

Why this answer

Removing unnecessary gridlines and decorative elements reduces non-data ink, thereby increasing the data-ink ratio. The data-ink ratio is the proportion of ink used to display data versus total ink used in the chart.

816
MCQhard

A financial analyst is preparing a dataset of stock transactions for a machine learning model. The 'transaction_amount' column has a highly skewed distribution with a few extremely large values. The analyst decides to apply a logarithmic transformation to this column. Which statement best describes the effect of this transformation?

A.It eliminates all outliers from the dataset.
B.It reduces the impact of outliers by compressing the scale of large values.
C.It normalizes the data to a mean of 0 and standard deviation of 1.
D.It converts the data from continuous to categorical.
AnswerB

A logarithmic transformation compresses the range of large values more than small ones, reducing skewness and the influence of outliers. This makes the distribution more symmetric and can improve the performance of models that assume normality. It is a common technique for handling skewed data.

Why this answer

The logarithmic transformation is used to reduce right skewness by compressing the scale of large values. This lessens the leverage of extreme outliers on statistical models without removing them. The transformation is monotonic, so it preserves the order of data points.

It does not standardize the data or change its type; it simply alters the distribution shape to be more symmetric.

Exam trap

The trap here is confusing logarithmic transformation with standardization or outlier removal, which serve different purposes.

817
MCQmedium

A data analyst is working with a dataset that includes a categorical variable 'education_level' with four categories: High School, Bachelor's, Master's, and PhD. The analyst wants to include this variable in a linear regression model. Which encoding method should the analyst use to avoid the dummy variable trap?

A.Label encoding
B.Binary encoding
C.One-hot encoding with three categories (dropping one)
D.One-hot encoding with all four categories
AnswerC

One-hot encoding with k-1 categories (here, three) avoids the dummy variable trap by preventing perfect multicollinearity. The dropped category becomes the reference level, and the coefficients for the other categories represent the difference from that reference. This is the standard approach for including nominal categorical variables in linear regression.

Why this answer

The dummy variable trap occurs when all dummy variables are included in a model with an intercept, causing perfect multicollinearity. To avoid it, one category is omitted as the reference. One-hot encoding with three categories (k-1) achieves this.

Label encoding and binary encoding impose ordinality, which is inappropriate for nominal data. Including all four categories would cause multicollinearity.

Exam trap

The trap here is thinking that one-hot encoding all categories is fine; it actually creates perfect multicollinearity and unstable estimates.

818
Multi-Selecteasy

Which TWO of the following are valid SQL clauses used to filter and sort data?

Select 2 answers
A.DELETE
B.WHERE
C.ORDER BY
D.UPDATE
E.INSERT
AnswersB, C

WHERE filters rows before grouping or aggregation, applying a predicate to each row and returning only those satisfying the condition. This directly satisfies the stem's requirement for a valid SQL filtering clause, distinct from sorting clauses such as ORDER BY. It cannot sort results, but filtering alone qualifies it as one of the two valid answers.

Why this answer

Option B, WHERE, is correct because it is the SQL clause that filters rows by applying a Boolean predicate to each row before it is returned, as in SELECT ... FROM table WHERE condition. Option C, ORDER BY, is correct because it is the SQL clause that sorts the result set by one or more columns, optionally with ASC or DESC, as in SELECT ...

FROM table ORDER BY column. The remaining options are not filtering or sorting clauses: DELETE (A) is a DML statement that removes rows, UPDATE (D) is a DML statement that modifies existing rows, and INSERT (E) is a DML statement that adds new rows.

Exam trap

CompTIA often tests the distinction between SQL DML statements (DELETE, UPDATE, INSERT) and query clauses (WHERE, ORDER BY), trapping candidates who confuse data manipulation commands with data retrieval or sorting operations.

819
Multi-Selecthard

Which THREE factors should be considered when choosing a chart type for a dataset?

Select 3 answers
A.The animation capabilities of the software
B.The data types (categorical, numerical, time series)
C.The number of variables to display
D.The key insight or message to convey
E.The color scheme of the company logo
AnswersB, C, D

Data type dictates which encodings are valid: categorical fields suit bars, numerical fields suit histograms or scatter plots, and time series demand a continuous axis. Choosing a chart without matching the data type produces misleading visuals.

Why this answer

Option B is correct because the data type—categorical, numerical, or time series—directly determines which chart families are valid: categorical comparisons suit bar charts, time series suit line charts, and numerical distributions suit histograms or scatter plots. Option C is correct because the number of variables to display dictates dimensionality, so one variable may use a histogram, two variables a scatter plot, and three or more may require bubble charts, small multiples, or faceting. Option D is correct because the key insight or message to convey should drive the choice—showing trend, comparison, composition, correlation, or distribution each calls for a different chart type.

Option A is not a primary factor because animation is a presentation enhancement, not a determinant of whether a chart correctly encodes the data. Option E is not a primary factor because brand color schemes affect styling and accessibility, not the fundamental selection of chart type.

Exam trap

The trap here is that candidates often confuse aesthetic or software-specific features (like animation or branding) with the fundamental data characteristics that dictate chart appropriateness, leading them to select options that are about polish rather than analytical correctness.

820
MCQmedium

A marketing team wants to segment customers into groups based on purchasing behavior without prior labels. Which algorithm should the data analyst use?

A.K-means clustering
B.K-nearest neighbors
C.Linear regression
D.Decision tree
AnswerA

K-means clustering partitions unlabelled data into k groups by minimising within-cluster variance, directly satisfying the stem's requirement for segmentation without prior labels. Unlike supervised methods such as logistic regression or decision trees, it needs no target variable, making it the appropriate choice for discovering behavioural customer segments.

Why this answer

K-means clustering is the correct choice because it is an unsupervised learning algorithm that groups unlabeled data into clusters based on feature similarity. Since the marketing team has no prior labels for customer segments, K-means can partition customers by purchasing behavior patterns, such as frequency and monetary value, without needing predefined categories.

Exam trap

The trap here is that candidates often confuse unsupervised clustering (K-means) with supervised classification (K-nearest neighbors) because both involve 'K' and grouping, but KNN requires labeled data and predicts labels, while K-means discovers inherent structures without labels.

How to eliminate wrong answers

Option B is wrong because K-nearest neighbors is a supervised learning algorithm that requires labeled training data to classify or predict outcomes, making it unsuitable for unlabeled segmentation. Option C is wrong because linear regression is a supervised regression algorithm used to predict a continuous target variable, not to discover hidden groupings in unlabeled data. Option D is wrong because decision trees are typically used for supervised classification or regression tasks, relying on labeled data to split on features, and cannot perform unsupervised clustering without prior labels.

821
MCQmedium

A data analyst is working with a dataset that contains a column 'region' with values such as 'North', 'South', 'East', 'West', and 'N/A'. The analyst needs to prepare this column for a machine learning model. Which of the following is the most appropriate approach to handle the 'N/A' values?

A.Replace 'N/A' with the mode of the 'region' column.
B.Treat 'N/A' as a separate category and encode it as its own level.
C.Encode 'N/A' as a numeric value of 0 and other regions as 1-4.
D.Remove all rows where 'region' is 'N/A'.
AnswerB

Treating 'N/A' as a distinct category preserves the information that the region is missing or not applicable, which can be predictive. This approach avoids introducing false assumptions and allows the model to learn any signal associated with missingness. It is a common and valid strategy for categorical variables.

Why this answer

Treating 'N/A' as a separate category is the best approach because it retains the information that the region is missing or not applicable, which could be informative. Imputing with the mode or removing rows can introduce bias or lose data, and numeric encoding creates a false ordinal relationship.

Exam trap

The trap here is assuming that missing values must be imputed or removed, when sometimes they represent a valid category that should be preserved.

822
MCQmedium

A data analyst needs to combine customer information from a CRM table and order information from an orders table, returning only customers who have placed at least one order. Which type of join should the analyst use?

A.RIGHT JOIN
B.INNER JOIN
C.FULL OUTER JOIN
D.LEFT JOIN
AnswerB

An INNER JOIN returns only rows where the join key matches in both tables, so customers without orders are excluded and only customers with at least one order remain. A LEFT JOIN would retain unmatched customers, contradicting the stated requirement.

Why this answer

An INNER JOIN between the CRM table and the orders table returns only rows where there is a match in both tables based on the join key (e.g., customer ID). This satisfies the requirement to return only customers who have placed at least one order, because any customer without an order in the orders table will be excluded from the result set.

Exam trap

The CompTIA Data+ exam often tests the misconception that a LEFT JOIN will include only customers with orders because it 'joins' the tables, but the trap is that a LEFT JOIN preserves all rows from the left table, including those with no matches, so it does not filter out customers without orders.

How to eliminate wrong answers

Option A (RIGHT JOIN) is wrong because it returns all rows from the orders table and matching rows from the CRM table, which could include orders without a matching customer (if referential integrity is not enforced) and would not limit results to only customers with orders. Option C (FULL OUTER JOIN) is wrong because it returns all rows from both tables, including customers without orders and orders without customers, which violates the requirement to return only customers who have placed at least one order. Option D (LEFT JOIN) is wrong because it returns all rows from the CRM table and matching rows from the orders table, which would include customers with zero orders (where the orders columns are NULL), failing to filter out customers without orders.

823
MCQhard

A data scientist creates a box plot of employee salaries and notices many outliers above the upper whisker. What action should be taken to best understand the salary distribution?

A.Replace the box plot with a histogram of the salaries
B.Remove all outliers to create a more typical box plot
C.Trim the top 5% of salaries and recreate the box plot
D.Investigate the outliers to determine if they are data entry errors or valid extremes
AnswerD

Outliers above the upper whisker may represent genuine senior or executive salaries or data entry mistakes. Investigating each one distinguishes valid extremes from errors, giving an accurate picture of the salary distribution before deciding whether to exclude or retain them.

Why this answer

Outliers in a box plot represent data points that fall outside the typical range (beyond 1.5×IQR). They may be legitimate extreme values or errors. Investigating them is essential to understand whether they are valid (e.g., highly compensated executives) or mistakes (e.g., data entry errors).

Removing or trimming them without investigation could distort the analysis and hide important insights.

Exam trap

The trap here is assuming that outliers should always be removed or that changing the visualization will solve the problem. Candidates may think that a histogram or trimming will 'fix' the box plot, but the key is to investigate the cause of outliers first.

How to eliminate wrong answers

Option A is wrong because a histogram shows the distribution but does not help identify whether outliers are errors or valid; it may obscure individual extreme values. Option B is wrong because removing outliers without investigation can bias results and discard valid data. Option C is wrong because trimming the top 5% arbitrarily removes data and does not address the root cause of the outliers.

824
MCQhard

A data scientist is working with a dataset containing 1000 features and 500 samples. The goal is to build a predictive model. Which technique should be used to reduce the number of features while retaining most of the variance?

A.Ridge regression
B.Forward selection
C.Principal Component Analysis (PCA)
D.Lasso regression
AnswerC

PCA projects the 1000 features onto orthogonal principal components ordered by explained variance, letting you keep the top components that retain most variance while discarding the rest, reducing dimensionality despite having fewer samples than features.

Why this answer

Principal Component Analysis (PCA) is an unsupervised dimensionality reduction technique that transforms the original features into a set of orthogonal components, ordered by the variance they capture. Given 1000 features and only 500 samples, PCA is ideal because it reduces the feature space while retaining the maximum variance, helping to avoid overfitting and the curse of dimensionality.

Exam trap

CompTIA often tests the distinction between supervised feature selection (Lasso, Forward selection) and unsupervised dimensionality reduction (PCA), trapping candidates who confuse regularization with variance-based reduction.

How to eliminate wrong answers

Option A is wrong because Ridge regression is a regularization technique that shrinks coefficients but does not reduce the number of features; it retains all features with penalized weights. Option B is wrong because Forward selection is a supervised feature selection method that selects features based on their predictive power, not on variance retention, and it can be computationally expensive with 1000 features. Option D is wrong because Lasso regression performs feature selection by shrinking some coefficients to zero, but it is a supervised method that selects features based on target correlation, not on maximizing variance retention, and may not be optimal for unsupervised dimensionality reduction.

825
Multi-Selecthard

A data analyst is using Python pandas to perform exploratory data analysis. Which THREE methods are commonly used to assess data quality and distributions?

Select 3 answers
A.df.transpose()
B.df.describe()
C.df.info()
D.df.sort_values()
E.df.value_counts()
AnswersB, C, E

df.describe() returns count, mean, standard deviation, minimum, quartiles and maximum for numeric columns, exposing outliers, skew and missing values. This single call summarises distribution shape and completeness, making it a standard first step in data-quality assessment.

Why this answer

describe() gives summary statistics, info() shows data types and non-null counts, and value_counts() shows frequency distributions.

Page 10

Page 11 of 14

Page 12