Courseiva

CompTIA Data+ (DA0-002) (DA0-002) — Questions 676–750

1004 questions total · 14pages · All types, answers revealed

Page 9

Page 10 of 14

Page 11
676
MCQmedium

A marketing analyst is combining two datasets: one containing campaign IDs and spend, and another containing campaign IDs and impressions. The first dataset has 1,200 rows and the second has 950 rows. After an inner join on campaign_id, the result has 1,050 rows. Which statement best explains this result?

A.The inner join retained only rows where campaign_id exists in both datasets, so campaigns present in only one dataset were excluded.
B.The inner join produced a Cartesian product because campaign_id was not unique in one of the datasets.
C.The result indicates referential integrity was enforced, so orphaned campaign IDs were automatically repaired.
D.The join failed to include unmatched rows because the analyst should have used a cross join to preserve all campaigns.
AnswerA

An inner join returns only the intersection of keys from both tables. With 1,200 and 950 rows, the result of 1,050 means many campaigns matched but some from each side did not. This is the expected behavior of an inner join and explains why the output is smaller than the larger input while still substantial, reflecting partial overlap between the two campaign lists.

Why this answer

An inner join returns only rows whose join key appears in both inputs. Given 1,200 and 950 source rows, an output of 1,050 indicates substantial but incomplete overlap, which is exactly what an inner join produces. A Cartesian product, cross join, or referential integrity enforcement would not yield this count.

The analyst observed normal inner join behavior reflecting partial key matching between the two campaign datasets.

Exam trap

The trap here is assuming any row count change during a join indicates an error, when an inner join legitimately reduces rows to the matching key intersection.

677
MCQmedium

When communicating uncertainty in a report, which of the following is the most appropriate way to convey the reliability of a survey result showing 75% customer satisfaction?

A."The satisfaction rate might be lower or higher."
B."75% of customers are satisfied."
C."We are 95% confident that the true satisfaction rate is between 72% and 78%."
D."The margin of error is 3%."
AnswerC

A confidence interval quantifies sampling uncertainty, directly satisfying the stem's demand to convey reliability. Stating 95% confidence that the true rate lies between 72% and 78% gives the audience a precise, bounded estimate rather than a bare point figure, which would overstate certainty.

Why this answer

A confidence interval communicates both the point estimate and the uncertainty around it. Stating '95% confident the true rate is between 72% and 78%' gives the reader the estimate (75%), the precision (the interval width), and the confidence level, which is the most complete and honest way to convey reliability. The other options either omit the uncertainty or state it incompletely.

Exam trap

DA0-002 often tests the misconception that a margin of error alone conveys reliability, when the confidence level and interval bounds are needed for a complete statement of uncertainty.

How to eliminate wrong answers

Option A is wrong because 'might be lower or higher' is vague and provides no quantified uncertainty, so the reader cannot judge reliability. Option B is wrong because stating '75% of customers are satisfied' presents the sample statistic as if it were the population truth, ignoring sampling error entirely. Option D is wrong because 'the margin of error is 3%' gives only half the picture — it omits the confidence level and the resulting interval, so the reader cannot interpret the precision correctly.

678
MCQeasy

A marketing team wants to analyze customer sentiment from social media posts. Which data acquisition method is most appropriate?

A.Internal database query
B.Physical sensor data
C.Web scraping from public social media APIs
D.Survey questionnaire
AnswerC

Public social media APIs expose sentiment-bearing posts as structured JSON, letting the team acquire text at scale without breaching platform terms. This satisfies the requirement to analyse customer sentiment from social media, since scraping public APIs yields the raw opinion data the model needs.

Why this answer

Web scraping from public social media APIs is the most appropriate method for analyzing customer sentiment from social media posts because it directly collects the unstructured text data (posts, comments, tweets) that contains sentiment. Social media platforms provide APIs (e.g., Twitter API, Facebook Graph API) that allow programmatic access to public posts, enabling large-scale data acquisition. This method is real-time, scalable, and captures the authentic voice of customers, which is essential for sentiment analysis.

Internal databases and surveys do not capture social media data, and physical sensors are irrelevant to text-based sentiment.

Exam trap

The trap here is confusing data acquisition methods: candidates might think internal databases contain social media data or that surveys can capture unsolicited sentiment, but the key is recognizing that social media posts are external, unstructured text best obtained via APIs or scraping.

How to eliminate wrong answers

Option A is wrong because internal database queries only access data already stored within the organization, which does not include external social media posts. Option B is wrong because physical sensor data measures environmental or physical phenomena (e.g., temperature, motion), not text-based social media content. Option D is wrong because survey questionnaires are structured, solicited responses that may not reflect spontaneous social media sentiment and are limited in scale and timeliness.

679
MCQhard

An e-commerce company is merging customer data from three legacy systems. Two systems use email as unique identifier, but one system allows multiple customers per email. The third uses phone number. To create a unified customer view, the analyst should first:

A.Request the IT team to modify the legacy system
B.Build a customer matching rule that uses multiple attributes (email, phone, name) with a confidence score
C.Use email as primary key and ignore conflicts
D.Assign new unique IDs and discard existing identifiers
AnswerB

Email alone cannot be the match key because one legacy system permits duplicate customers per email, and phone alone is equally unreliable. A multi-attribute rule with a confidence score resolves this by scoring combined agreement across email, phone and name, letting the analyst merge only above a chosen threshold.

Why this answer

Merging data from systems with different identifier schemas requires a probabilistic matching approach. Using multiple attributes (email, phone, name) with a confidence score allows the analyst to resolve conflicts where email is not unique and phone numbers may be missing or formatted differently, creating a unified customer view without forcing a single key.

Exam trap

The trap here is that candidates assume a single unique identifier (email) can be forced as a primary key, ignoring the real-world data quality issue of non-unique emails, which the question explicitly states.

How to eliminate wrong answers

Option A is wrong because modifying legacy systems is often impractical, costly, and outside the analyst's scope; the question asks what the analyst should do first, not a long-term IT project. Option C is wrong because using email as primary key and ignoring conflicts would lose data integrity when one email maps to multiple customers, violating the goal of a unified view. Option D is wrong because assigning new unique IDs and discarding existing identifiers eliminates the ability to link records back to source systems and loses valuable matching context, making deduplication impossible.

680
MCQmedium

A data governance team is implementing a program to ensure consistent definitions and quality of customer data across the organization. They assign a senior manager to be accountable for the data asset. Which role does this manager fulfill?

A.Data analyst
B.Data custodian
C.Data owner
D.Data steward
AnswerC

Data owner is accountable for a specific data domain.

Why this answer

The data owner is the senior manager accountable for a specific data asset, including its quality, definition, and compliance. In the DA0-001 context, the data owner has ultimate responsibility for the data, not just day-to-day management. This role ensures consistent definitions and quality across the organization, aligning with the governance team's objectives.

Exam trap

The trap here is confusing the data owner's accountability with the data steward's operational duties, leading candidates to pick 'Data steward' because they associate governance with hands-on management rather than executive responsibility.

How to eliminate wrong answers

Option A is wrong because a data analyst focuses on analyzing and interpreting data, not on accountability for data definitions or quality. Option B is wrong because a data custodian is responsible for the technical environment and security of data, not for defining or governing its meaning. Option D is wrong because a data steward handles day-to-day data governance tasks like metadata management and quality monitoring, but does not hold the ultimate accountability that a senior manager does.

681
MCQmedium

A data analyst at an online retailer is examining a dataset of customer orders. The 'order_total' column has a mean of $85 and a median of $62. The analyst wants to describe the typical order amount for a presentation to the marketing team. Which measure of central tendency is most appropriate to report as the typical value, and why?

A.The mean, because it uses all data points and is the standard measure of central tendency.
B.The median, because it is resistant to extreme values and better represents the typical order in a skewed distribution.
C.The mode, because it identifies the most frequent order amount and is easy to understand.
D.The range, because it shows the spread of order amounts and highlights variability.
AnswerB

The median is the middle value when data is ordered, so it is not influenced by a few extremely large orders. In a right-skewed distribution where mean > median, the median better reflects the typical order amount for the majority of customers, making it the appropriate measure to report.

Why this answer

When a distribution is skewed, the mean is pulled toward the tail, while the median remains a robust measure of the center. Here, the mean ($85) is substantially higher than the median ($62), signaling right skew. The median therefore better represents the typical order amount for most customers.

Exam trap

The trap here is assuming the mean is always the best measure of central tendency, overlooking the impact of outliers in skewed data.

682
MCQeasy

A data analyst is preparing a presentation for executive leadership. The analyst wants to highlight the correlation between marketing spend and revenue over the past year. Which visualization type is most appropriate for showing this relationship?

A.Scatter plot
B.Pie chart
C.Bar chart
D.Histogram
AnswerA

A scatter plot places marketing spend on one axis and revenue on the other, so each point represents a paired observation and the relationship's direction, strength and outliers are visible. This satisfies the requirement to highlight correlation between two continuous measures, which bar or line charts cannot show as clearly.

Why this answer

A scatter plot is the most appropriate visualization for showing the relationship between two continuous variables—marketing spend and revenue—because it plots individual data points on an X-Y axis, allowing the analyst to visually assess correlation, trends, and outliers. This directly supports the goal of highlighting correlation, as the pattern of points (e.g., upward slope) indicates the strength and direction of the relationship.

Exam trap

The trap here is that candidates often confuse a bar chart or histogram with a scatter plot because they think any chart with axes can show relationships, but only a scatter plot directly plots paired continuous data to reveal correlation without aggregation.

How to eliminate wrong answers

Option B (Pie chart) is wrong because pie charts are designed to show parts of a whole (proportions) for categorical data, not the relationship between two continuous variables. Option C (Bar chart) is wrong because bar charts compare discrete categories or aggregated values, not the correlation between two continuous metrics; they would require binning or summarizing the data, losing the granularity needed for correlation analysis. Option D (Histogram) is wrong because histograms display the distribution of a single continuous variable (e.g., frequency of revenue values), not the relationship between two variables.

683
MCQeasy

In SQL, which string function would you use to remove leading and trailing spaces from a column named 'city'?

A.TRIM
B.RTRIM
C.LTRIM
D.CLEAN
AnswerA

TRIM removes both leading and trailing spaces from a string, returning the cleaned value for the city column. It precisely matches the requirement to strip spaces from both ends, unlike LTRIM or RTRIM which handle only one side.

Why this answer

The SQL TRIM function removes both leading and trailing spaces (or specified characters) from a string. Applied to a column like 'city', TRIM(city) returns the value with all surrounding whitespace stripped, which is exactly what the question asks for.

Exam trap

The trap is that LTRIM and RTRIM sound like they handle 'trimming' generically — candidates who don't recall that TRIM handles both sides may pick one of the directional variants.

How to eliminate wrong answers

Option B is wrong because RTRIM only removes trailing (right-side) spaces, leaving leading spaces intact. Option C is wrong because LTRIM only removes leading (left-side) spaces, leaving trailing spaces intact. Option D is wrong because CLEAN is not a standard SQL string function — it is not part of ANSI SQL and does not exist in major databases like PostgreSQL, MySQL, or SQL Server for this purpose.

684
MCQmedium

A data analyst at a pharmaceutical company is asked to create a report on clinical trial adverse events. The report will be shared with the FDA. The analyst must ensure that the data is accurate, complete, and traceable to source documents. Which data governance practice is most critical to meet these requirements?

A.Establishing data lineage that tracks the origin and transformation of each data element from source to report.
B.Implementing data versioning to maintain historical copies of the report and underlying datasets.
C.Conducting a manual review of the report by a second analyst before submission.
D.Applying data masking to protect patient identifiers in the report.
AnswerA

Data lineage provides a complete audit trail showing how data flows from source documents to the final report, including transformations. This traceability is essential for regulatory submissions like FDA reports, where accuracy and completeness must be verifiable. It allows auditors to confirm that no unauthorized changes occurred and that data integrity is maintained.

Why this answer

Data lineage is the most critical practice because it provides an auditable trail from source documents to the final report, ensuring that every data element can be traced and verified. This directly supports the FDA's requirements for accuracy, completeness, and traceability, making it indispensable for regulatory submissions.

Exam trap

The trap here is selecting data masking or versioning as the primary governance practice, when the FDA's core requirement is traceability, which lineage uniquely provides.

685
MCQhard

A data engineer is designing storage for a fraud detection system that ingests millions of transaction events per second. The system must store each event with its timestamp and support fast writes without predefined schema enforcement, while allowing later analytical queries over semi-structured payloads. Which storage approach best fits these requirements?

A.A document-oriented NoSQL database with schema-on-read
B.A relational OLTP database with third normal form tables
C.A graph database optimized for relationship traversal
D.A columnar data warehouse with strict schema-on-write
AnswerA

Document-oriented NoSQL databases accept semi-structured payloads without predefined schema enforcement, support high write throughput through horizontal scaling, and apply schema-on-read during queries. These characteristics align with ingesting millions of transaction events per second and later analyzing flexible JSON-like documents, making this approach the best fit.

Why this answer

The requirements combine high-velocity ingestion, absence of predefined schema, and support for semi-structured payloads with later analytical querying. Document-oriented NoSQL databases provide schema-on-read flexibility and horizontal write scalability that match these needs. Relational, columnar, and graph systems each impose constraints or optimizations that conflict with one or more of the stated requirements.

Exam trap

The trap here is assuming that any database capable of analytical queries must be a columnar warehouse, overlooking that schema-on-read NoSQL stores can also support analytics over semi-structured data.

686
MCQhard

A data analyst is evaluating a classification model that predicts whether customers will churn. The dataset is highly imbalanced, with only 5% of customers churning. The analyst wants to choose a metric that focuses on the model's ability to correctly identify actual churners. Which metric should the analyst prioritize?

A.Specificity
B.Recall
C.Precision
D.Accuracy
AnswerB

Recall, also known as sensitivity or true positive rate, measures the proportion of actual positives that are correctly identified. In this scenario, the analyst wants to correctly identify actual churners, so recall directly addresses that goal. High recall means the model captures most churners, which is crucial when the cost of missing a churner is high.

Why this answer

Recall is the metric that quantifies the model's ability to find all relevant instances of the positive class. Since the analyst wants to correctly identify actual churners, recall is the most appropriate metric. Accuracy and specificity focus on the negative class or overall correctness, and precision does not capture missed positives.

Exam trap

The trap here is defaulting to accuracy as the primary metric; in imbalanced datasets, it can be deceptively high even when the model fails to predict the minority class.

687
MCQhard

A sales dashboard shows a map with many overlapping markers in the same city, making it hard to read. What is the best improvement?

A.Add tooltips to show details on hover
B.Aggregate the data by region and use a choropleth map
C.Use a bubble chart instead of a map
D.Use different marker colors for each store
AnswerB

Choropleth mapping replaces individual point markers with shaded regions, eliminating the overlapping-marker clutter entirely. Aggregating to region level also matches the density comparison the map is meant to convey, satisfying the readability constraint in the stem.

Why this answer

Aggregating sales data by region and using a choropleth map eliminates visual clutter from overlapping markers by shading entire geographic areas based on a metric (e.g., total sales). This approach leverages spatial aggregation to provide a clear, high-level view of regional performance, which is the best practice when individual point markers become unreadable due to density.

Exam trap

The trap here is that candidates may choose tooltips (Option A) thinking interactivity solves the problem, but the question asks for the 'best improvement' to readability, and tooltips do not address the fundamental issue of overlapping markers obscuring the visualization.

How to eliminate wrong answers

Option A is wrong because tooltips only provide details on hover and do not solve the core problem of overlapping markers obscuring data; they add interactivity but do not reduce visual density. Option C is wrong because a bubble chart, while useful for comparing values, is not a map-based visualization and would lose the geographic context that the dashboard intends to convey. Option D is wrong because using different marker colors for each store does not address overlapping markers; it only adds visual differentiation without reducing clutter, and in dense areas, colored markers still overlap and remain unreadable.

688
MCQhard

Refer to the exhibit. A database administrator notices that queries filtering on both CustomerID and OrderDate are slow. Which single change would most likely improve performance for such queries?

A.Partition the table by OrderDate
B.Convert TotalAmount to VARCHAR
C.Add a composite index on (CustomerID, OrderDate)
D.Remove the primary key constraint
AnswerC

A composite index on (CustomerID, OrderDate) matches the query's equality filter followed by its range or sort column, letting the engine seek directly to the relevant CustomerID entries already ordered by OrderDate. This avoids scanning and sorting, unlike separate single-column indexes.

Why this answer

A composite index on (CustomerID, OrderDate) allows the database to satisfy queries that filter on both columns using a single index seek, avoiding a full table scan or multiple index lookups. The column order matters: CustomerID first supports equality filtering, and OrderDate second supports range or sort operations within each customer. This is the most direct performance improvement for the described query pattern.

Exam trap

The trap is choosing partitioning by OrderDate because it sounds like a performance fix — candidates overlook that the query filters on CustomerID first, and partitioning on the wrong column does not help and may hurt.

How to eliminate wrong answers

Option A is wrong because partitioning by OrderDate alone does not help queries that filter primarily by CustomerID — partition pruning would not eliminate partitions effectively, and it could even hurt queries that span many dates. Option B is wrong because converting TotalAmount to VARCHAR changes the data type and would break numeric comparisons and aggregations, with no benefit to filtering on CustomerID and OrderDate. Option D is wrong because removing the primary key constraint eliminates uniqueness enforcement and the clustered index, degrading performance and data integrity rather than improving it.

689
MCQmedium

A data quality report shows that 95% of records have all required fields completed, but 20% of the completed fields contain values that are outside valid ranges. Which data quality dimension is most affected?

A.Consistency
B.Accuracy
C.Timeliness
D.Completeness
AnswerB

Accuracy measures whether values correctly represent reality and fall within valid domains. Completeness is already high at 95%, but out-of-range values breach validity and correctness, so accuracy is the dimension most affected by this defect.

Why this answer

Accuracy measures how well data reflects real-world values or a defined standard. Here, 20% of completed fields contain values outside valid ranges, meaning the data is present but incorrect, directly degrading accuracy. Completeness (95% filled) is high, but the core issue is that the values themselves are wrong, not missing or late.

Exam trap

The trap here is that candidates see '95% of records have all required fields completed' and immediately think 'Completeness is high, so that dimension is fine,' but then incorrectly assume the 20% out-of-range values also affect Completeness, when in fact Accuracy is the dimension that suffers when present data is invalid.

How to eliminate wrong answers

Option A (Consistency) is wrong because consistency checks for logical coherence across datasets or over time (e.g., same customer ID format in two tables), not whether individual field values fall within valid ranges. Option C (Timeliness) is wrong because timeliness concerns whether data is available when needed or within a required time window, not the correctness of values. Option D (Completeness) is wrong because completeness measures the presence of data (95% of records have all required fields), which is high; the problem is with the quality of the present data, not its absence.

690
MCQeasy

Which chart type is best for showing the number of website visitors at each stage of a conversion funnel, from initial visit to purchase?

A.Stacked bar chart
B.Treemap
C.Funnel chart
D.Waterfall chart
AnswerC

A funnel chart plots sequential stages as progressively narrowing bars, directly representing drop-off between visit and purchase. This matches the requirement to show visitor counts at each conversion stage, which bar, line or pie charts cannot express as clearly.

Why this answer

A funnel chart is specifically designed to visualize the progressive reduction in volume across stages of a linear process, such as a conversion funnel. It clearly shows the number of visitors at each stage (e.g., initial visit, product view, add to cart, purchase) and the drop-off between them, making it the optimal choice for this scenario.

Exam trap

The trap here is that candidates often confuse a funnel chart with a waterfall chart because both show sequential steps, but a waterfall chart is for cumulative changes (additions/subtractions), not for displaying the count at each stage of a funnel.

How to eliminate wrong answers

Option A is wrong because a stacked bar chart is used to compare parts of a whole across categories, not to show the sequential reduction in a funnel; it would obscure the drop-off between stages. Option B is wrong because a treemap displays hierarchical data as nested rectangles based on proportion, which is not suitable for a linear, sequential process like a conversion funnel. Option D is wrong because a waterfall chart is designed to show the cumulative effect of sequential positive and negative values (e.g., financial statements), not the simple count of visitors at each stage of a funnel.

691
MCQmedium

A stock analyst is analyzing monthly sales data for a retail company and observes a consistent pattern of high sales every December. This pattern is most likely an example of which time series component?

A.Irregular
B.Cyclical
C.Seasonality
D.Trend
AnswerC

Seasonality describes a repeating pattern that recurs at fixed intervals shorter than a year, such as peak sales every December. This matches the stem's consistent annual December spike, distinguishing it from trend, which is a long-term directional movement.

Why this answer

A consistent spike every December is a calendar-driven repeating pattern, which is the definition of seasonality in time series analysis. The fixed annual period (12 months) and predictable timing make this a classic seasonal component.

Exam trap

The trap is confusing seasonality with trend — a December spike repeats every year (seasonal), whereas trend would be a steady multi-year rise or fall in sales, and the exam tests whether you anchor on the calendar periodicity.

How to eliminate wrong answers

Option A is wrong because irregular components are random shocks (e.g., a one-off supply disruption) that do not repeat predictably each December. Option B is wrong because cyclical patterns are multi-year, irregular fluctuations tied to business cycles, not a reliable annual December spike. Option D is wrong because trend captures the long-term direction of sales over many years, not the recurring within-year December peak.

692
MCQhard

A data team is preparing a quarterly business review for the CEO. The report must include both high-level summaries and the ability for the CEO to drill down into specific departments. Which reporting technique best meets this requirement?

A.A slide deck with one slide per department.
B.An interactive dashboard with drill-down capabilities.
C.A static PDF with a summary page and appendices.
D.A data dump in Excel with filters.
AnswerB

An interactive dashboard presents summary metrics initially, then lets the CEO filter or click through to department-level detail on demand. This satisfies the dual requirement of high-level summaries plus drill-down, which a static report or slide deck cannot deliver within a single artefact.

Why this answer

An interactive dashboard with drill-down capabilities (Option B) is the correct choice because it directly addresses the requirement for both high-level summaries and the ability to explore specific departments. Dashboards allow the CEO to view aggregated KPIs at a glance and then click through to detailed views for each department, providing a seamless, user-driven exploration experience without switching between separate reports or slides.

Exam trap

The trap here is that candidates often choose a static PDF (Option C) or a slide deck (Option A) because they associate 'report' with printed or presentation materials, but the question explicitly requires 'drill-down' capability, which is a hallmark of interactive business intelligence tools, not static documents.

How to eliminate wrong answers

Option A is wrong because a slide deck with one slide per department forces a linear, static presentation; the CEO cannot dynamically drill down from a summary view into a specific department without manually navigating slides, which breaks the requirement for interactive drill-down. Option C is wrong because a static PDF with a summary page and appendices is non-interactive; the CEO would have to jump to appendix pages manually, which is not a true drill-down capability and lacks the real-time filtering or cross-filtering that an interactive dashboard provides. Option D is wrong because a data dump in Excel with filters is a raw data file that requires the CEO to understand the data structure and apply filters manually; it does not offer a curated high-level summary or a guided drill-down path, and it risks overwhelming the user with granular data without pre-built aggregations.

693
MCQeasy

A data analyst is profiling a dataset and notices that the 'age' column contains negative values and values exceeding 120. The analyst needs to address these anomalies. Which data preparation technique is most appropriate?

A.Data aggregation
B.Imputation
C.Normalization
D.Outlier detection and treatment
AnswerD

Negative ages and ages above 120 are outliers that likely represent data entry errors. Outlier detection identifies such values, and treatment (e.g., removal, capping, or correction) addresses them. This technique is specifically designed to handle values that fall outside a plausible range, ensuring data quality.

Why this answer

The most appropriate technique is outlier detection and treatment because negative ages and ages over 120 are statistically implausible and likely errors. This method identifies and corrects or removes such values, ensuring the dataset's integrity. Other techniques like imputation or normalization do not address the underlying invalidity of the data.

Exam trap

The trap here is confusing outlier treatment with imputation, which is for missing values, or normalization, which only rescales data without fixing errors.

694
Multi-Selectmedium

Which THREE are best practices for data profiling during acquisition? (Choose three.)

Select 3 answers
A.Immediately normalize data
B.Check for completeness
C.Assess data types
D.Identify outliers
E.Skip validation for trusted sources
AnswersB, C, D

Ensuring all required fields are populated is essential.

Why this answer

Checking for completeness (Option B) is a best practice during data acquisition because it ensures that all required fields and records are present before further processing. Incomplete data can lead to incorrect analysis or failed transformations, so profiling for missing values or nulls is a fundamental validation step.

Exam trap

The trap here is that candidates confuse 'best practices for acquisition' with 'best practices for transformation,' leading them to select normalization (Option A) as an immediate step rather than a later processing stage.

695
MCQeasy

In simple linear regression, the coefficient of determination R² measures:

A.The probability that the slope is zero
B.The slope of the regression line
C.The proportion of variance in the dependent variable explained by the independent variable
D.The strength and direction of the linear relationship
AnswerC

R² quantifies the share of total variance in the dependent variable accounted for by the fitted regression on the independent variable, expressed as a proportion between 0 and 1. It therefore directly satisfies the stem's requirement to measure explained variance rather than correlation strength or slope.

Why this answer

R², the coefficient of determination, quantifies the proportion of variance in the dependent variable that is explained by the independent variable(s) in the regression model. It ranges from 0 to 1, where 1 means the model explains all variance and 0 means it explains none.

Exam trap

The trap is confusing R² with Pearson's r — R² measures explained variance (0 to 1) while r measures strength and direction (−1 to +1), and the exam offers both as plausible-sounding options.

How to eliminate wrong answers

Option A is wrong because the probability that the slope is zero is assessed via the p-value of the slope coefficient's t-test, not R². Option B is wrong because the slope of the regression line is the coefficient β₁ (or b), which represents the change in Y per unit change in X — R² is a goodness-of-fit measure, not a slope. Option D is wrong because the strength and direction of a linear relationship is measured by the correlation coefficient r (Pearson's r), which ranges from −1 to +1; R² is r squared and loses direction information.

696
MCQhard

A data analyst is using a recursive CTE to traverse an organizational hierarchy. What is the purpose of the anchor member in the recursive CTE?

A.It provides the initial seed or starting rows for the recursion.
B.It filters the final output of the recursive CTE.
C.It specifies how to join the CTE with itself recursively.
D.It defines the termination condition for the recursion.
AnswerA

The anchor member supplies the seed rows that begin the recursion, typically the hierarchy's root nodes. Without it, the recursive member has no starting set to iterate from, so the CTE cannot traverse the organisational hierarchy. It executes once, then the recursive member repeatedly joins against its output until no rows remain.

Why this answer

The anchor member initializes the recursion with the base result set.

697
Multi-Selecthard

A data analyst is evaluating data quality issues in a customer database. Which TWO actions are best practices for ensuring data consistency?

Select 2 answers
A.Allowing null values for foreign keys
B.Standardizing date formats across all tables
C.Implementing referential integrity constraints
D.Enabling cascading updates on primary keys
E.Using data profiling to identify duplicate records
AnswersB, C

Correct: Uniform formats ensure consistency in temporal data.

Why this answer

Standardizing date formats across all tables (Option B) ensures that date values are stored and interpreted uniformly, eliminating inconsistencies that arise from mixed formats (e.g., MM/DD/YYYY vs. DD-MM-YY). This practice directly supports data consistency by enforcing a single representation, which is critical for accurate querying, reporting, and integration across systems.

Exam trap

CompTIA often tests the distinction between data quality dimensions (e.g., consistency vs. accuracy), leading candidates to confuse data profiling (which identifies duplicates) with a direct method for enforcing consistency.

698
MCQeasy

In a regression analysis, the coefficient of determination (R²) is 0.85. How should this value be interpreted?

A.85% of the data points lie on the regression line
B.The slope of the regression line is 0.85
C.85% of the variance in the dependent variable is explained by the model
D.85% of the independent variables are significant
AnswerC

R² measures the proportion of variance in the dependent variable accounted for by the regression model. A value of 0.85 means the model explains 85% of that variance, leaving 15% attributable to other factors or random error.

Why this answer

R², the coefficient of determination, measures the proportion of variance in the dependent variable that is explained by the independent variables in the regression model. An R² of 0.85 means 85% of the variance in the dependent variable is accounted for by the model, indicating a strong fit.

Exam trap

DA0-002 often tests the misconception that R² represents the percentage of points on the line or the slope — candidates must remember it is the proportion of variance explained.

How to eliminate wrong answers

Option A is wrong because R² does not measure the percentage of data points lying on the regression line — that would be a different concept, and in most real regressions very few points lie exactly on the line. Option B is wrong because R² is not the slope; the slope is a separate coefficient (β₁) in the regression equation. Option D is wrong because R² says nothing about the significance of individual independent variables — that is assessed via t-tests and p-values for each coefficient.

699
Multi-Selecthard

A data analyst is preparing a report on customer satisfaction scores. To comply with GDPR, which THREE actions must be taken? (Select THREE.)

Select 3 answers
A.Retain data indefinitely for analysis
B.Ensure aggregates do not identify individuals
C.Include customer names for context
D.Anonymize personally identifiable information
E.Establish data retention periods for the report data
AnswersB, D, E

Aggregation must not allow re-identification, so ensuring aggregates do not identify individuals upholds GDPR's anonymisation principle. It satisfies the compliance constraint by preventing small cell sizes or unique combinations from exposing a single data subject within the satisfaction report.

Why this answer

Option B is correct because GDPR's data minimization and purpose-limitation principles require that aggregated statistics used for reporting must not allow re-identification of any individual, so ensuring aggregates do not identify individuals protects data subjects' privacy. Option D is correct because anonymizing personally identifiable information (PII) such as names, email addresses, and account numbers removes the personal data from scope of GDPR processing, satisfying the regulation's requirement to protect identifiable data. Option E is correct because GDPR Article 5(1)(e) mandates storage limitation, meaning the analyst must establish defined data retention periods for the report data rather than keeping it longer than necessary.

Option A is incorrect because retaining data indefinitely violates the GDPR storage-limitation principle. Option C is incorrect because including customer names for context unnecessarily introduces identifiable personal data, contradicting data minimization and the anonymization requirement.

Exam trap

DA0-002 often tests the misconception that 'aggregated data is automatically anonymous' or that retaining data indefinitely is acceptable for analytics — both violate GDPR's storage limitation and anonymization standards.

700
Matchingmedium

Match each data analysis technique to its primary purpose.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Model relationships between variables

Group similar data points without labels

Analyze data points collected over time

Compare means across multiple groups

Test association between categorical variables

Why these pairings

The correct matches are: Regression with predicting continuous outcomes, Clustering with grouping similar data, Classification with assigning categories, and PCA with reducing dimensionality. Common confusions include swapping regression and clustering definitions.

701
MCQmedium

A data architect is designing a storage solution for a healthcare provider. The system must store patient records with varying structures, including unstructured clinical notes and semi-structured JSON data from wearables. The solution must scale horizontally and handle high write throughput. Which storage type is most appropriate?

A.Data warehouse
B.Data lake
C.OLTP database
D.Relational database
AnswerB

A data lake stores raw data in any format, including unstructured and semi-structured, and can scale horizontally using distributed storage. It supports high write throughput and schema-on-read, making it ideal for diverse healthcare data like clinical notes and JSON from wearables.

Why this answer

The correct answer is Data lake because it can store raw data in any format, scale horizontally, and handle high write throughput. Relational databases, data warehouses, and OLTP databases are optimized for structured data and transactional workloads, making them less suitable for diverse, high-volume healthcare data.

Exam trap

The trap here is assuming that a data warehouse or relational database can handle unstructured data; however, they require schema-on-write and are not designed for raw, varied data formats.

702
Multi-Selectmedium

A data analyst is profiling a newly acquired customer dataset before loading it into a warehouse. The analyst notices that the 'country' column contains values such as 'USA', 'United States', 'U.S.A.', and 'US'. Which TWO actions are appropriate for standardizing this column during data preparation? (Choose two.)

Select 2 answers
A.Delete rows that do not exactly match the most common value for the country column.
B.Use an ISO 3166 country code list as the reference and map each variant to its corresponding alpha-2 or alpha-3 code.
C.Apply a fuzzy matching algorithm to cluster similar strings and assign the most frequent value as the canonical form.
D.Convert all values to uppercase and remove periods, then treat the results as distinct categories.
E.Create a mapping table that translates each variant to a single canonical country name or code.
AnswersB, E

ISO 3166 provides an authoritative, unambiguous set of country codes, so mapping variants to alpha-2 or alpha-3 codes yields a stable, internationally recognized canonical form. It avoids ambiguity in names and supports reliable joins and aggregation. This is a best practice for country data standardization and complements a mapping table with an external standard.

Why this answer

Standardizing a country column with known variants requires deterministic translation to a canonical form. A mapping table documents each variant explicitly, and ISO 3166 codes provide an authoritative reference that removes name ambiguity. Together they yield consistent values suitable for joins and aggregation.

Normalization alone leaves synonyms unresolved, fuzzy clustering risks merging distinct entities, and deleting rows sacrifices valid data rather than cleaning it.

Exam trap

The trap here is believing that simple case and punctuation normalization is sufficient standardization, when synonyms like 'United States' still remain distinct without a mapping or reference standard.

703
MCQeasy

A hospital's data team needs to choose a storage model for its new electronic health record (EHR) system. Patient records contain many nested, variable attributes such as a list of allergies, multiple insurance policies, and a changing set of lab results per visit. The schema changes frequently as new clinical fields are added, and the team wants to avoid costly migrations. Which data model should the team select?

A.Columnar model
B.Relational model
C.Document model
D.Key-value model
AnswerC

A document model stores each patient record as a self-contained document that can hold nested arrays and variable fields, such as an allergy list or a set of lab results. New clinical fields can be added without altering a global schema or performing migrations, which directly matches the team's requirement for flexible, fast-changing structures.

Why this answer

The team needs a schema-flexible structure that can hold nested and variable attributes, such as allergy lists and changing lab results, while allowing new fields to be added without migrations. A document model provides exactly this by storing each patient record as a self-contained, nested document that evolves with clinical requirements.

Exam trap

The trap here is assuming that any non-relational database automatically supports flexible nested schemas, when key-value stores actually treat values as opaque and offer no structure for nested attributes.

704
MCQmedium

A data analyst is creating a presentation for executives to explain why customer churn has increased over the last quarter. The analyst wants to present the story in a compelling way. Which narrative structure is most appropriate?

A.Problem, Hypothesis, Test
B.Background, Analysis, Recommendation
C.Situation, Complication, Resolution
D.Data, Visualization, Conclusion
AnswerC

Situation, Complication, Resolution frames the stable baseline, introduces the churn increase as the complicating disruption, then presents the recommended response. This structure gives executives a clear causal narrative explaining why churn rose and what to do, matching the persuasive intent.

Why this answer

The Situation-Complication-Resolution structure is ideal for executive presentations because it first establishes the context (situation), then introduces the problem (complication—increased churn), and finally proposes a solution (resolution). This narrative arc aligns with how executives process strategic issues, making the data story compelling and actionable. In contrast, other structures are better suited for technical reports or hypothesis testing, not high-level storytelling.

Exam trap

The CompTIA Data+ exam often tests the distinction between narrative structures for different audiences; the trap here is that candidates mistake 'Background, Analysis, Recommendation' (a common technical report format) as appropriate for executives, when in fact it lacks the persuasive arc needed for strategic decision-making.

How to eliminate wrong answers

Option A is wrong because 'Problem, Hypothesis, Test' is a scientific method structure used for experimental validation, not for presenting a business narrative to executives. Option B is wrong because 'Background, Analysis, Recommendation' is a linear report format that lacks the dramatic tension needed to engage an executive audience on a problem like churn. Option D is wrong because 'Data, Visualization, Conclusion' is a data-centric sequence that prioritizes outputs over storytelling, failing to frame the business impact and resolution in a compelling way.

705
MCQmedium

A data analyst needs to compare the salary distribution across five departments. Which visualization is most appropriate?

A.Line chart
B.Side-by-side box plot
C.Scatter plot
D.Stacked bar chart
AnswerB

A side-by-side box plot displays median, quartiles and outliers for each department on a shared salary axis, enabling direct distribution comparison across all five groups. This satisfies the stem's requirement to compare distributions rather than single summary values, which a bar chart of averages could not reveal.

Why this answer

A side-by-side box plot (option B) is the most appropriate visualization for comparing salary distributions across multiple departments because it displays the median, quartiles, and potential outliers for each group simultaneously. This allows the analyst to assess central tendency, spread, and skewness across all five departments in a single, compact chart.

Exam trap

The trap here is that candidates often confuse 'comparing distributions' with 'showing trends' or 'showing relationships,' leading them to incorrectly select a line chart or scatter plot instead of recognizing that a box plot is purpose-built for distribution comparison across groups.

How to eliminate wrong answers

Option A is wrong because a line chart is designed to show trends over a continuous interval (e.g., time series) and is not suitable for comparing distributions of categorical groups like departments. Option C is wrong because a scatter plot visualizes the relationship between two continuous variables, not the distribution of a single variable across categories. Option D is wrong because a stacked bar chart is used to show the composition of parts to a whole across categories, not the distribution (e.g., quartiles, outliers) of a continuous variable like salary.

706
MCQeasy

A data analyst wants to compare the sales revenue of five different product categories for the current month. Which chart type is most suitable for this comparison?

A.Bar chart
B.Histogram
C.Pie chart
D.Line chart
AnswerA

A bar chart encodes each product category as a separate bar with length proportional to revenue, letting viewers compare five discrete categories side by side. This suits the nominal, low-cardinality comparison; line charts imply continuity and pie charts obscure precise magnitude differences.

Why this answer

A bar chart encodes each product category as a discrete bar whose length maps to revenue, making side-by-side comparison of five categorical values immediate and precise. Categorical comparisons are the canonical use case for bar charts because the human eye judges length differences far more accurately than angles or slopes.

Exam trap

DA0-002 often tests the bar-chart-versus-histogram confusion — candidates see 'compare values across categories' and pick histogram, forgetting that a histogram requires a continuous numeric variable binned into ranges, not discrete category labels.

How to eliminate wrong answers

Option B is wrong because a histogram bins continuous numeric data into ranges to show a frequency distribution — product categories are discrete labels, not a continuous variable. Option C is wrong because a pie chart shows parts-of-a-whole proportions and becomes hard to read with five slices; it also doesn't support precise magnitude comparison as well as bars. Option D is wrong because a line chart implies a continuous trend over an ordered dimension (usually time), which is meaningless for unordered categories.

707
MCQeasy

A data analyst is creating a report to summarize customer satisfaction survey results. The survey asked customers to rate their satisfaction on a scale from 1 (very dissatisfied) to 5 (very satisfied). The analyst wants to show the number of respondents for each rating level. Which visualization is most appropriate?

A.Line chart
B.Bar chart
C.Pie chart
D.Histogram
AnswerB

A bar chart is ideal for displaying the frequency or count of categorical data, such as satisfaction ratings. Each bar represents a rating level, and the height shows the number of respondents. This makes it easy to compare counts across categories. It accurately represents the ordinal nature of the ratings without implying continuity or trend.

Why this answer

A bar chart is the most appropriate visualization for showing the frequency of categorical data like satisfaction ratings. It clearly displays the count for each rating level, allowing easy comparison. Other chart types either imply a trend, assume continuous data, or make comparison difficult.

Exam trap

The trap here is confusing a bar chart with a histogram, especially when the data is numerical but discrete.

708
Multi-Selectmedium

An analyst wants to use Python (pandas) to compute the average sales amount per region from a DataFrame 'df' with columns 'region' and 'sales'. Which TWO pandas operations are needed? (Select TWO).

Select 2 answers
A.df.fillna(0)
B.df.pivot_table(index='region', values='sales', aggfunc='mean')
C.df['sales'].apply(np.sqrt)
D.df.merge(df2, on='region')
E.df.groupby('region')['sales'].mean()
AnswersB, E

`pivot_table` groups rows by the `region` column and applies `aggfunc='mean'` to the `sales` values, producing one averaged figure per region. This directly satisfies the stem's requirement to compute average sales amount per region, collapsing many rows into a single aggregated result keyed by region.

Why this answer

Option B, df.pivot_table(index='region', values='sales', aggfunc='mean'), is correct because pivot_table with index='region' groups rows by region, selects the 'sales' column via values='sales', and applies the mean aggregation through aggfunc='mean', directly producing the average sales per region. Option E, df.groupby('region')['sales'].mean(), is correct because groupby('region') splits the DataFrame by region, ['sales'] selects the sales column, and .mean() computes the arithmetic average of sales within each group, yielding the same per-region averages. The other options do not compute grouped averages: A (df.fillna(0)) only replaces missing values with zero, C (df['sales'].apply(np.sqrt)) applies a square-root transformation element-wise, and D (df.merge(df2, on='region')) joins two DataFrames on the region key without any aggregation.

Exam trap

The trap is overcomplicating the question — candidates may look for a merge or a fillna step, but the core operation is simply group-and-aggregate, which both groupby().mean() and pivot_table() accomplish.

709
MCQmedium

An analyst receives a dataset of website sessions where the session_duration_seconds column contains several negative values and a few values exceeding 86,400 seconds. The analyst must prepare this data for analysis of average session length. Which action best addresses this data quality issue while preserving analytical integrity?

A.Replace all out-of-range values with the column mean so the average remains stable.
B.Take the absolute value of negative durations and cap all values at 86,400 seconds.
C.Leave the values unchanged because outliers are a natural part of real-world data.
D.Flag the out-of-range values, investigate their source, and correct or exclude them using a documented rule before computing the average.
AnswerD

Negative durations are logically impossible and durations beyond a day are implausible for a single session, so they indicate measurement or pipeline errors. Flagging and investigating them, then applying a documented correction or exclusion rule, preserves integrity and makes the average meaningful. Documenting the rule also keeps the process auditable and repeatable.

Why this answer

Negative and implausibly large session durations are invalid values that would distort an average. The sound approach is to flag them, investigate the cause, and apply a documented correction or exclusion rule. Transforming them into plausible values or ignoring them hides the defect, while retaining them uncritically treats impossible data as natural variation.

Exam trap

The trap here is treating invalid values such as negative durations as ordinary outliers, when they are logically impossible and must be investigated rather than simply retained or smoothed.

710
MCQmedium

A data analyst is working with a dataset that contains a column 'purchase_date' stored as a string in the format 'YYYY-MM-DD'. The analyst needs to calculate the number of days between each purchase and the current date. Which data transformation is required?

A.Apply one-hot encoding to the string.
B.Normalize the string values to a 0-1 range.
C.Convert the string to an integer data type.
D.Convert the string to a date data type.
AnswerD

To calculate the difference between dates, the string must be converted to a date data type so that date arithmetic can be performed. Many tools and languages provide functions to parse strings into dates, enabling subtraction to yield the number of days.

Why this answer

Calculating the number of days between dates requires date arithmetic, which is only possible if the values are in a date data type. Converting the string to a date allows the use of built-in functions to compute differences. Other transformations like integer conversion or encoding do not support date calculations.

Exam trap

The trap here is thinking that string manipulation alone can compute date differences, but without a proper date type, calculations are error-prone and complex.

711
MCQhard

In a table with columns 'employee_id' and 'manager_id', a data analyst needs to retrieve the hierarchy level of each employee, where the top manager has manager_id NULL. Which SQL feature is best suited?

A.A window function with ROW_NUMBER()
B.A recursive CTE
C.A GROUP BY clause with aggregation
D.A self-join with a LEFT JOIN
AnswerB

A recursive CTE references its own result set, walking manager_id links upward from each employee until the NULL root is reached, producing hierarchy levels. Self-joins need a known depth, and window functions cannot traverse variable-length parent-child chains.

Why this answer

Recursive CTE can traverse hierarchical data to compute levels.

712
MCQmedium

A data analyst is reviewing sales data and wants to find orders where the order total is between $100 and $500, inclusive. Which WHERE clause is correct?

A.total > 100 AND total < 500
B.total IN (100, 500)
C.total BETWEEN 100 AND 500
D.total >= 100 OR total <= 500
AnswerC

BETWEEN is inclusive at both bounds, so total BETWEEN 100 AND 500 returns rows where the total equals 100 or 500 as well as every value in between, satisfying the stem's inclusive requirement without needing separate >= and <= comparisons.

Why this answer

The BETWEEN operator is inclusive of both endpoints, so 'total BETWEEN 100 AND 500' returns rows where total is greater than or equal to 100 and less than or equal to 500 — exactly the inclusive range requested. It is equivalent to 'total >= 100 AND total <= 500' but more concise and readable.

Exam trap

DA0-002 often tests the misconception that BETWEEN is exclusive of its endpoints, when it is actually inclusive — and also tests the AND vs OR confusion in range conditions.

How to eliminate wrong answers

Option A is wrong because 'total > 100 AND total < 500' uses strict inequalities, excluding orders exactly equal to $100 or $500, which violates the inclusive requirement. Option B is wrong because 'total IN (100, 500)' matches only the two exact values 100 and 500, not the range between them. Option D is wrong because 'total >= 100 OR total <= 500' uses OR instead of AND, which is always true for any value (every number is either >= 100 or <= 500), returning all rows.

713
MCQeasy

A sales VP wants a quick summary of last month's revenue change and key drivers. Which report section is most relevant?

A.Executive summary
B.Data dictionary
C.Methodology notes
D.Row-level data
AnswerA

An executive summary condenses the period's revenue movement and its principal drivers into a brief narrative, matching the VP's request for a quick overview. Detailed transactional or variance sections supply supporting depth but not the immediate high-level answer.

Why this answer

Executive summaries provide high-level numbers and context for quick decision-making.

714
MCQhard

A data analyst is building a binary classification model to predict customer churn. The dataset is imbalanced, with only 10% churners. The analyst wants to evaluate model performance with a focus on correctly identifying churners. Which metric is most appropriate?

A.Recall (sensitivity)
B.F1-score
C.Precision
D.Accuracy
AnswerA

Recall measures the proportion of actual churners correctly identified, directly matching the stated focus on catching churners. With only 10% churners, accuracy is misleading because predicting 'no churn' everywhere scores 90%, whereas recall exposes missed churners.

Why this answer

Recall (sensitivity) is the most appropriate metric because it measures the proportion of actual churners correctly identified by the model. Since the dataset is imbalanced (only 10% churners) and the analyst's focus is on correctly identifying churners, recall directly addresses the cost of missing positive cases (false negatives). Accuracy would be misleading due to class imbalance, while precision and F1-score prioritize different trade-offs.

Exam trap

The trap here is that candidates often default to accuracy as the default metric, failing to recognize that class imbalance renders accuracy misleading, and that the question's explicit focus on 'correctly identifying churners' points directly to recall, not precision or F1-score.

How to eliminate wrong answers

Option B (F1-score) is wrong because it balances precision and recall, but the analyst's primary goal is to maximize identification of churners, not to balance false positives and false negatives; F1-score would penalize a model that achieves high recall at the expense of precision, which may be acceptable in this scenario. Option C (Precision) is wrong because it measures the proportion of predicted churners that are actual churners, focusing on false positives rather than false negatives; the analyst wants to minimize missed churners, not necessarily avoid false alarms. Option D (Accuracy) is wrong because with only 10% churners, a naive model predicting all non-churners would achieve 90% accuracy, masking poor performance on the minority class; accuracy is inappropriate for imbalanced classification problems.

715
MCQeasy

A hospital's data team is cataloging its data assets. They need to classify the blood type field stored for each patient (e.g., A+, O-). Which data type classification best describes this field?

A.Nominal
B.Interval
C.Ordinal
D.Ratio
AnswerA

Blood type values such as A+, O-, and AB+ are labels that identify categories with no inherent order or ranking. Nominal data classifies items into distinct groups where no category is greater or lesser than another. Because blood types cannot be meaningfully ordered or mathematically averaged, classifying them as nominal is correct for this hospital scenario.

Why this answer

Blood type is a categorical label that identifies which of the recognized blood groups a patient belongs to. There is no inherent order among A+, O-, AB+, and B-, and no arithmetic can be meaningfully performed on these values. Nominal classification is the only correct choice because it groups data into unordered, distinct categories.

Exam trap

The trap here is assuming that because blood types contain letters and symbols, they represent a ranked or measurable scale rather than unordered labels.

716
MCQmedium

A retail company wants to predict future sales based on historical data. Which modeling approach is most appropriate if the data shows a clear seasonal pattern?

A.Linear regression
B.Time series analysis
C.K-means clustering
D.Logistic regression
AnswerB

Time series analysis explicitly models sequential dependence and seasonality through components such as trend, seasonal, and residual terms, satisfying the stem's clear seasonal pattern requirement. Unlike regression, it uses autocorrelation and prior-period values, so forecasts of future sales account for recurring cycles rather than treating observations as independent.

Why this answer

Time series analysis is specifically designed to model data points indexed in time order, making it ideal for capturing and forecasting seasonal patterns. Unlike regression models, it accounts for autocorrelation, trends, and seasonality components, which are critical for accurate sales prediction from historical data.

Exam trap

The trap here is that candidates see 'predict future sales' and mistakenly choose linear regression, overlooking that time series methods are required when data has temporal dependencies and seasonality.

How to eliminate wrong answers

Option A is wrong because linear regression assumes independence of observations and cannot model time-dependent structures like seasonality or autocorrelation. Option C is wrong because K-means clustering is an unsupervised learning method used for grouping similar data points, not for forecasting future values. Option D is wrong because logistic regression is used for binary classification problems, not for predicting continuous numeric sales figures.

717
MCQhard

A data governance team is establishing policies to ensure data quality. They define rules for data accuracy, completeness, and consistency. Which data governance function is primarily responsible for defining and enforcing these rules?

A.Data stewardship
B.Data ownership
C.Data quality management
D.Master data management
AnswerC

Data quality management directly defines and enforces accuracy, completeness, and consistency rules, satisfying the stem's requirement for a governance function owning those three dimensions. It operationalises policy through profiling, validation, cleansing, and monitoring, unlike master data management, which governs shared reference entities rather than quality rule enforcement.

Why this answer

Data quality management is the governance function specifically responsible for defining data quality dimensions (accuracy, completeness, consistency, timeliness, validity, uniqueness) and implementing the rules, profiling, and monitoring to enforce them. It translates governance policy into measurable quality controls and remediation workflows.

Exam trap

DA0-002 often tests the boundary between data stewardship (who applies rules) and data quality management (who defines and enforces them) — candidates pick 'stewardship' because it sounds like the hands-on quality role.

How to eliminate wrong answers

Option A is wrong because data stewardship is about the day-to-day custodianship of data assets — stewards apply and monitor the rules, but the function that defines the quality dimensions and rules is data quality management. Option B is wrong because data ownership is about accountability and decision rights for a data domain (who approves access, defines policy), not the operational enforcement of quality rules. Option D is wrong because master data management focuses on creating a single trusted golden record for core entities (customer, product), which is a related but distinct discipline from defining quality rules across all data.

718
Multi-Selectmedium

A retail company wants to segment its customers based on purchase history. Which THREE methods are appropriate for customer segmentation?

Select 3 answers
A.RFM analysis
B.Linear regression
C.K-means clustering
D.t-test
E.Hierarchical clustering
AnswersA, C, E

RFM analysis scores customers on recency, frequency and monetary value derived from purchase history, directly satisfying the segmentation requirement. It groups customers into actionable tiers such as best, loyal or at-risk based on transactional behaviour.

Why this answer

RFM analysis (A) is correct because it segments customers by Recency, Frequency, and Monetary value of purchases, directly using purchase history to group customers into meaningful tiers. K-means clustering (C) is correct because it partitions customers into k groups based on feature similarity such as purchase behavior, making it a standard unsupervised segmentation technique. Hierarchical clustering (E) is correct because it builds a dendrogram of nested customer clusters, allowing segmentation at different granularity levels without pre-specifying the number of clusters.

Linear regression (B) is not appropriate because it predicts a continuous outcome rather than assigning customers to segments. The t-test (D) is not appropriate because it is a hypothesis test comparing means between two groups, not a segmentation method.

Exam trap

The trap is that regression and hypothesis-testing methods sound analytical and data-driven, so candidates may select them for segmentation even though they predict values or test differences rather than grouping customers.

719
MCQmedium

The exhibit shows an SQL query executed on an 'orders' table that contains 'order_id', 'customer_id', and 'order_date'. What is the purpose of this query?

A.Count total orders per customer regardless of date
B.Calculate average order count per customer for 2023
C.Find products with more than 5 orders in 2023
D.Identify customers who placed more than 5 orders in 2023
AnswerD

The query groups orders by customer_id, filters order_date to the 2023 range, and applies a HAVING count greater than five. This returns customers whose 2023 order count exceeds five, satisfying the stated purpose of identifying high-frequency customers.

Why this answer

The query groups orders by customer_id and filters using a HAVING clause with COUNT(*) > 5, which counts the number of orders per customer. The WHERE clause restricts orders to those placed in 2023, so the result identifies customers who placed more than 5 orders in that year. This matches option D exactly.

Exam trap

CompTIA often tests the distinction between WHERE and HAVING, and the trap here is confusing a count of orders per customer with a count of products or an average, leading candidates to pick option B or C.

How to eliminate wrong answers

Option A is wrong because the WHERE clause filters for order_date in 2023, so the count is not regardless of date. Option B is wrong because the query counts orders per customer, not the average order count per customer. Option C is wrong because the query operates on an 'orders' table with no product-related column; it counts orders per customer, not products.

720
MCQmedium

An analyst is reviewing the above SQL query used to acquire data. What does this query retrieve?

A.Customers who placed more than 5 orders in 2023
B.All customers who placed at least 5 orders in 2023
C.The total number of orders per customer in 2023
D.Customers who placed exactly 5 orders in 2023
AnswerA

The query groups orders by customer, filters transactions to the 2023 calendar year, applies a HAVING clause counting orders per customer, and returns only those exceeding five. This yields the set of customers meeting that order-frequency threshold within the specified period.

Why this answer

The SQL query uses a HAVING clause with COUNT(*) > 5 to filter customers who placed more than 5 orders in 2023. The WHERE clause restricts records to the year 2023, and the GROUP BY customer_id aggregates orders per customer. The condition '> 5' explicitly excludes customers with exactly 5 or fewer orders, making option A correct.

Exam trap

The trap here is confusing the comparison operator '>' with '>=', leading candidates to mistakenly include customers with exactly 5 orders when the query explicitly excludes them.

How to eliminate wrong answers

Option B is wrong because 'at least 5 orders' would require the condition COUNT(*) >= 5, not > 5. Option C is wrong because the query returns customer IDs, not the total number of orders per customer; the COUNT is used only for filtering, not as a selected column. Option D is wrong because 'exactly 5 orders' would require COUNT(*) = 5, not > 5.

721
MCQmedium

A data analyst is building a dashboard for executives and wants to ensure the most important metric, total revenue, is immediately visible. Which design principle should the analyst apply?

A.Appropriate precision
B.Consistent color coding
C.Visual hierarchy
D.Data-ink ratio
AnswerC

Visual hierarchy deliberately sizes, positions and contrasts elements so the eye lands first on the most critical metric. Placing total revenue at the top-left with dominant scale and colour ensures executives see it immediately, satisfying the stem's requirement for instant visibility.

Why this answer

Visual hierarchy is the design principle of arranging elements so the most important information draws the eye first, typically through size, position, color contrast, or whitespace. Placing total revenue prominently (e.g., top-left, largest font, high-contrast KPI card) ensures executives see it immediately. This directly addresses the requirement to make the key metric immediately visible.

Exam trap

The trap is confusing visual hierarchy (prioritizing what the eye sees first) with related but distinct principles like data-ink ratio (minimizing clutter) or consistent color coding (semantic color use), which do not by themselves elevate one metric above others.

How to eliminate wrong answers

Option A is wrong because appropriate precision concerns rounding and decimal places (e.g., showing $1.2M instead of $1,234,567.89), not prominence or placement. Option B is wrong because consistent color coding ensures the same meaning for a color across visuals; it aids comprehension but does not prioritize one metric over others. Option D is wrong because data-ink ratio (Tufte's principle) is about removing non-data ink like gridlines and 3D effects to reduce clutter, not about emphasizing a specific metric.

722
Multi-Selectmedium

An analyst is preparing data for an A/B test and wants to ensure valid results. Which TWO of the following should be considered when calculating the required sample size?

Select 2 answers
A.Data dimensionality
B.Desired effect size
C.Skewness of data
D.Number of features
E.Statistical power
AnswersB, E

Effect size quantifies the minimum difference the test must detect between variants. Smaller effects demand larger samples, so specifying it directly drives the sample size calculation and satisfies the validity constraint, preventing an underpowered test that misses genuine differences.

Why this answer

Sample size calculation depends on desired effect size and statistical power, among other factors like significance level.

723
Multi-Selectmedium

A data analyst is extracting data from a web page using web scraping techniques. The data will be used for market research. Which TWO of the following are common challenges associated with web scraping?

Select 2 answers
A.Limited API rate limits
B.Legal and ethical restrictions
C.Website structure changes
D.High latency of data transfer
E.Inconsistent data formatting
AnswersB, C

Many websites prohibit scraping in their terms of service, and legal issues may arise.

Why this answer

Web scraping often involves accessing data that may be protected by copyright, terms of service, or privacy regulations such as GDPR or the Computer Fraud and Abuse Act (CFAA). Even if data is publicly accessible, repurposing it for market research without permission can lead to legal liability or ethical violations, making this a fundamental challenge.

Exam trap

CompTIA Data+ often tests the distinction between API-related challenges (rate limits, authentication) and web-scraping-specific challenges (structure changes, legal/ethical issues), so candidates mistakenly select 'Limited API rate limits' because they confuse web scraping with API consumption.

724
MCQmedium

A dataset contains sales transactions with columns 'order_date', 'amount', and 'region'. The analyst wants to calculate the total sales per region for orders placed in 2023, but only include regions where total sales exceed $10,000. Which SQL clause should be used to filter the aggregated results?

A.HAVING
B.WHERE
C.GROUP BY
D.FILTER
AnswerA

HAVING filters groups after aggregation, unlike WHERE, which filters individual rows before grouping. Since the $10,000 threshold applies to SUM(amount) per region, not to individual transactions, HAVING is the only clause that can evaluate the aggregated total and exclude regions failing it.

Why this answer

HAVING is the SQL clause that filters rows after aggregation, so it can reference aggregate functions like SUM(amount) and apply conditions such as SUM(amount) > 10000. WHERE runs before GROUP BY and cannot reference aggregates, so it cannot filter on total sales per region.

Exam trap

DA0-002 often tests the WHERE vs. HAVING distinction — candidates pick WHERE because it 'filters,' forgetting that WHERE cannot reference aggregate functions.

How to eliminate wrong answers

Option B is wrong because WHERE filters individual rows before grouping and aggregation occur, so it cannot reference aggregate results like SUM(amount) — attempting to do so raises an error in standard SQL. Option C is wrong because GROUP BY is the clause that forms the groups (by region) but does not itself filter; it's a prerequisite for HAVING, not a filter. Option D is wrong because FILTER is not a standard SQL clause for post-aggregation filtering — it exists only as a modifier on aggregate functions in some dialects (e.g., PostgreSQL's FILTER (WHERE ...)), not as a standalone clause.

725
MCQeasy

A data analyst wants to compare the means of three different training methods on employee productivity. Which statistical test is most appropriate?

A.Correlation analysis
B.ANOVA
C.Chi-square test
D.t-test
AnswerB

ANOVA compares means across three or more independent groups in one test, controlling the family-wise error rate that repeated t-tests would inflate. With three training methods as the grouping factor, it fits the scenario directly.

Why this answer

ANOVA (Analysis of Variance) is used to compare means of three or more groups.

726
MCQeasy

A marketing team wants to explore the relationship between advertising spend (in dollars) and resulting revenue. Which chart type is most suitable?

A.Line chart
B.Table
C.Pie chart
D.Scatter plot
AnswerD

A scatter plot positions each observation by two continuous variables, plotting advertising spend on one axis against revenue on the other. This directly satisfies the stem's requirement to explore the relationship between two numeric measures, revealing correlation, clustering and outliers that category-based charts cannot show.

Why this answer

A scatter plot is designed to display the relationship between two quantitative variables, with one variable on the x-axis and the other on the y-axis. Plotting advertising spend against revenue allows the marketing team to visually assess correlation, trends, and outliers. This is the standard chart type for bivariate numerical analysis.

Exam trap

DA0-002 often tests the distinction between chart types for different analytical purposes, so candidates may pick a line chart out of habit for 'trends' when the question asks about the relationship between two variables, which requires a scatter plot.

How to eliminate wrong answers

Option A is wrong because a line chart is typically used to show trends over time (e.g., revenue over months), not the relationship between two independent variables. Option B is wrong because a table lists data but does not visually reveal relationships or correlations. Option C is wrong because a pie chart shows parts of a whole (proportions) for a single categorical variable, not the relationship between two numerical variables.

727
MCQhard

A data analyst at a retail company is building a dashboard for store managers to track sales performance. The data comes from three sources: point-of-sale (POS) systems, inventory, and customer loyalty. The POS table contains columns transaction_id, store_id, date, product_id, quantity, and price. The inventory table has product_id, store_id, stock_level, and reorder_point. The loyalty table has customer_id, transaction_id, and points_earned. The analyst creates a star schema with a sales_fact fact table containing all rows from POS, dimension tables for store, product, date, and customer. To calculate average transaction value, the analyst uses the formula SUM(quantity * price) / COUNT(*). Store managers report that the average transaction value appears too low, especially for stores with multiple registers. The analyst realizes that because each product sold in a transaction creates a separate row in sales_fact, a single transaction with multiple items contributes multiple rows. The current calculation divides by the number of rows rather than the number of distinct transactions. Which of the following is the best course of action to correct the average transaction value metric? (Choose one.)

A.Use the MEDIAN function instead of AVG
B.Aggregate the data at the transaction level before calculating the average
C.Use a different data model that denormalizes transaction totals into a new fact table
D.Create a calculated field that sums sales per transaction (quantity * price) and then averages across distinct transaction IDs
AnswerD

The inflated row count from multi-item transactions skews the divisor. Summing quantity times price per transaction and averaging across distinct transaction IDs restores the correct denominator, giving the true average transaction value store managers expect.

Why this answer

The metric is wrong because the denominator counts fact rows (one per product line), not transactions. The correct approach is to compute the transaction-level total (SUM of quantity * price grouped by transaction_id) and then average those distinct transaction totals. Option D captures this exactly: sum sales per transaction, then average across distinct transaction IDs, which yields the true average transaction value.

Exam trap

The trap here is confusing row-level aggregation with entity-level aggregation — candidates see 'average' and reach for AVG without checking the fact table grain, missing that COUNT(*) counts line items, not transactions.

How to eliminate wrong answers

Option A is wrong because switching to MEDIAN does not fix the denominator problem — the row-level grain still inflates the count and the median would be computed over product-line rows, not transactions. Option B is wrong because 'aggregate at the transaction level before calculating the average' is vague and, as written, does not specify that the average must be taken over distinct transaction IDs; it could still produce a row-weighted result and does not explicitly address the grain mismatch. Option C is wrong because denormalizing transaction totals into a new fact table is unnecessary and adds ETL complexity — the existing star schema already supports the correct calculation via a grouped subquery or a transaction-level aggregate.

728
MCQmedium

Refer to the exhibit. Which type of data is the field "region"?

A.Qualitative
B.Continuous
C.Quantitative
D.Discrete
AnswerA

Region labels categories such as "North" or "EMEA", so it is qualitative (nominal) data. It cannot be measured or averaged, unlike quantitative fields. This satisfies the stem's request to classify the field by its data type.

Why this answer

The field 'region' contains categorical labels (e.g., 'North', 'South', 'East', 'West') that represent distinct groups or categories, not numerical measurements. Qualitative data (also called categorical data) describes attributes or characteristics that can be named but not meaningfully ordered or measured on a numeric scale. Since 'region' assigns a name to a geographic area without any inherent numeric value or order, it is a classic example of qualitative data.

Exam trap

The trap here is that candidates may confuse 'region' with a numeric code (e.g., region ID 1, 2, 3) and incorrectly classify it as discrete quantitative data, but the field 'region' as shown contains text labels, making it qualitative.

How to eliminate wrong answers

Option B is wrong because continuous data represents measurements that can take any value within a range (e.g., temperature, time), but 'region' consists of discrete labels with no numeric continuum. Option C is wrong because quantitative data involves numerical values that can be counted or measured (e.g., sales amount, age), whereas 'region' is a non-numeric category. Option D is wrong because discrete data is a subset of quantitative data that takes countable integer values (e.g., number of customers), but 'region' is not numeric at all.

729
MCQmedium

A data analyst wants to show the relationship between advertising spend and sales revenue for 50 stores. Which chart type is most appropriate?

A.Line chart
B.Scatter plot
C.Bar chart
D.Pie chart
AnswerB

Advertising spend and sales revenue are two continuous numeric variables, and the analyst wants their relationship across 50 stores. A scatter plot places one variable on each axis, revealing correlation, clusters and outliers that bar or line charts cannot show.

Why this answer

A scatter plot is the most appropriate chart for showing the relationship between two continuous variables—advertising spend and sales revenue—across 50 stores. Each point on the plot represents one store, allowing the analyst to visually assess correlation, trends, or outliers. This aligns with the DA0-001 objective of selecting visualizations that best represent bivariate relationships.

Exam trap

The trap here is that candidates often confuse a line chart (which connects points in sequence) with a scatter plot (which treats points as independent observations), leading them to incorrectly choose a line chart when no temporal or ordered dimension exists.

How to eliminate wrong answers

Option A is wrong because a line chart is typically used to display trends over time or ordered categories, not to show the relationship between two independent continuous variables like advertising spend and sales revenue. Option C is wrong because a bar chart compares discrete categories or groups, not the correlation between two continuous metrics across 50 individual stores. Option D is wrong because a pie chart shows proportions of a whole for categorical data, which is irrelevant for analyzing the relationship between two numerical variables.

730
MCQmedium

A data engineer is designing a schema for a new application. The application will store user profiles with fields such as UserID, Email, and SignupDate. The team expects the user base to grow rapidly, and the schema may evolve to include new profile attributes like phone number or social media handles. They want to minimize downtime and avoid complex migrations when adding fields. Which data model should the engineer choose?

A.Relational
B.Graph
C.Key-value
D.Document
AnswerD

Document databases store data in flexible, JSON-like documents where each record can have its own structure. New attributes such as phone number or social media handles can be added to individual documents without altering a global schema or performing migrations. This flexibility aligns perfectly with the requirement to minimize downtime and accommodate evolving user profiles.

Why this answer

A document database provides schema flexibility by storing each user profile as a self-contained document, allowing new attributes to be added without global schema changes or downtime. Relational databases require migrations, key-value stores lack query flexibility, and graph databases are optimized for relationships rather than evolving profile attributes.

Exam trap

The trap here is assuming that a relational database's structured schema is always preferable, ignoring the need for frequent schema changes without downtime.

731
Multi-Selectmedium

An analyst is choosing a chart to show the correlation between two continuous variables. Which TWO chart types could be used? (Select two.)

Select 2 answers
A.Bubble chart
B.Scatter plot
C.Pie chart
D.Waterfall chart
E.Histogram
AnswersA, B

Bubble charts are scatter plots with a third variable; they also show correlation.

Why this answer

A bubble chart is an extension of a scatter plot that can show the correlation between two continuous variables on the x- and y-axes, while a third variable is represented by the size of the bubbles. For the specific purpose of showing correlation between exactly two continuous variables, the bubble chart is valid because the bubble size is optional and does not interfere with the primary x-y relationship. This makes it a correct choice for visualizing the relationship between two continuous variables.

Exam trap

A common misconception in data visualization is that a histogram can show relationships between two variables, but it only displays the frequency distribution of a single continuous variable. For the Data+ exam, remember that scatter plots and bubble charts are appropriate for showing correlation between two continuous variables.

732
MCQeasy

A data analyst is examining the relationship between two continuous variables: temperature and ice cream sales. The analyst wants to quantify the strength and direction of their linear association. Which statistical measure should the analyst use?

A.Covariance
B.Spearman rank correlation
C.Pearson correlation coefficient
D.Chi-square test of independence
AnswerC

The Pearson correlation coefficient measures the strength and direction of a linear relationship between two continuous variables. Temperature and ice cream sales are both continuous, and the analyst seeks a linear association, making Pearson correlation the appropriate measure. It ranges from -1 to 1, indicating perfect negative to perfect positive linear relationships.

Why this answer

Pearson correlation coefficient is the standard measure for quantifying the strength and direction of a linear relationship between two continuous variables. It is scale-independent, ranging from -1 to 1, and directly addresses the analyst's goal. Other options either measure different types of association or are not suitable for continuous data.

Exam trap

The trap here is selecting covariance because it also measures linear relationship, but covariance lacks standardization and does not convey strength on a fixed scale.

733
Multi-Selectmedium

Which TWO of the following are examples of semi-structured data?

Select 2 answers
A.XML document
B.JSON object
C.Relational table
D.Plain text file
E.CSV file
AnswersA, B

XML documents carry self-describing tags that impose hierarchy and meaning without a rigid relational schema, satisfying the stem's semi-structured requirement. Unlike flat CSV rows or fully schema-bound tables, XML's nested elements and attributes vary between records, so structure exists but is not fixed in advance.

Why this answer

Semi-structured data has some organizational structure (tags, keys, or markers) but does not conform to a rigid tabular schema, and both A (XML document) and B (JSON object) fit this definition because they use self-describing tags or key-value pairs to organize data hierarchically without requiring a fixed relational schema. XML documents are explicitly semi-structured since elements and attributes define structure while allowing flexible, nested, and optional fields. JSON objects are likewise semi-structured because their key-value pairs and nested arrays/objects provide structure without enforcing a strict table format.

The unmarked options do not belong: C (relational table) is structured data with a fixed schema, E (CSV file) is typically treated as structured tabular data with rows and columns, and D (plain text file) is unstructured data lacking any formal organizational markers.

734
MCQmedium

A data analyst creates a scatter plot showing the relationship between advertising spend and revenue. The plot shows a strong positive correlation. Which of the following should the analyst include in the report to ensure accurate communication?

A.Include a note that correlation does not imply causation.
B.Replace the scatter plot with a bar chart.
C.Remove any outliers from the plot.
D.Add a trend line to the scatter plot.
AnswerA

A strong positive correlation between advertising spend and revenue does not establish that spending caused the revenue change; a confounding variable could drive both. Stating this caveat satisfies the accurate-communication requirement, preventing readers from inferring a causal relationship the scatter plot cannot support.

Why this answer

Correlation does not imply causation. It is important for the analyst to include this caveat to prevent misinterpretation of the relationship. Option B is incorrect because a bar chart is not appropriate for showing the relationship between two continuous variables.

Option C is incorrect because removing outliers without justification can distort the analysis; outliers should be handled transparently. Option D is incorrect because while a trend line could be added, it is not mandatory and the question specifically asks for ensuring accurate communication, which is primarily about the correlation-causation warning.

735
MCQmedium

A data analyst uses a CTE to simplify a complex query. Which keyword is used to define a CTE?

A.DEFINE
B.CTE
C.DECLARE
D.WITH
AnswerD

The WITH keyword introduces a common table expression, defining a named temporary result set that exists only for the duration of a single statement. It satisfies the stem's requirement for simplifying a complex query by letting the analyst break logic into readable, reusable blocks referenced later in the main SELECT.

Why this answer

A Common Table Expression (CTE) is defined using the WITH keyword followed by a name and an AS clause containing the subquery. WITH precedes the main SELECT and can define one or more named CTEs separated by commas.

Exam trap

DA0-002 often tests basic SQL syntax recall — candidates confuse WITH (CTE) with DECLARE (procedural variables) or assume 'CTE' itself is a keyword.

How to eliminate wrong answers

Option A is wrong because DEFINE is not a SQL keyword for CTEs — it appears in some procedural languages (e.g., Snowflake scripting) but is not the standard CTE introducer. Option B is wrong because CTE is the concept's name, not a keyword — you never write 'CTE name AS (...)'. Option C is wrong because DECLARE is used for variables, cursors, and handlers in procedural SQL (T-SQL, PL/pgSQL), not for defining CTEs.

736
Multi-Selectmedium

An analyst is planning an A/B test to compare two website designs. Which TWO factors should be considered when calculating the required sample size?

Select 2 answers
A.Data type of the outcome variable
B.Desired effect size
C.Statistical power
D.Color scheme of the designs
E.Number of missing values
AnswersB, C

The desired effect size is the minimum detectable difference between designs that the test must reliably detect. Smaller effects require substantially larger samples, making it a core input to any sample size calculation for the A/B test.

Why this answer

The required sample size for an A/B test depends directly on the desired effect size (B), because smaller effects are harder to detect and demand more observations to distinguish a real difference from noise. It also depends on statistical power (C), conventionally set at 0.80, since higher power (lower Type II error risk) requires a larger sample to reliably detect the effect when it truly exists. The data type of the outcome variable (A) affects the choice of statistical test, not the sample-size formula's core inputs, and the color scheme (D) is merely the design variation being tested, not a computational factor.

The number of missing values (E) is a data-quality issue handled during cleaning or imputation and is not a standard parameter in sample-size calculation.

Exam trap

DA0-002 often tests the four inputs to sample size (α, power, effect size, variance) — candidates pick data type or missing values because they sound statistical, but only effect size and power are among the core parameters.

737
MCQeasy

Which stage of the data lifecycle involves converting raw data into a usable format, such as cleaning or validating?

A.Archival
B.Processing
C.Ingestion
D.Storage
AnswerB

Processing transforms raw data into a usable format through cleaning, validation, and standardisation, converting it into a reliable dataset. It sits between collection and storage or analysis, directly matching the stem's description of converting raw data.

Why this answer

Processing is the stage where raw data is transformed into a usable format through cleaning, validation, normalization, or aggregation. This step ensures data quality and consistency before analysis or storage, directly matching the question's description.

Exam trap

The trap here is confusing ingestion (data arrival) with processing (data transformation), as both occur early in the lifecycle but serve distinct purposes.

How to eliminate wrong answers

Option A is wrong because archival refers to moving data to long-term storage for compliance or historical purposes, not cleaning or validating. Option C is wrong because ingestion is the initial capture or import of raw data from sources, not its transformation. Option D is wrong because storage is the persistent retention of data in databases or filesystems, not the conversion into a usable format.

738
MCQmedium

A data analyst wants to retrieve data from a REST API that returns JSON. Which step is part of the data lifecycle for this activity?

A.Data archival
B.Data sharing
C.Data deletion
D.Data ingestion
AnswerD

Fetching JSON from a REST API is data acquisition into the analytics pipeline, which is the ingestion stage of the data lifecycle. The analyst is moving raw data from an external source into a repository for subsequent storage, transformation and analysis, so ingestion is the lifecycle step being performed.

Why this answer

Retrieving data from a REST API that returns JSON is a classic example of data ingestion, where data is acquired from an external source and brought into a system for processing or storage. This step is part of the data lifecycle's initial phase, often called data acquisition or ingestion. The other options—archival, sharing, and deletion—occur later in the lifecycle after data has been ingested and processed.

Exam trap

The trap here is confusing data ingestion with other lifecycle stages like data sharing or archival, especially when the question mentions retrieving data from an API, which might sound like sharing. Candidates must remember that ingestion is about bringing data in, not sending it out or storing it long-term.

How to eliminate wrong answers

Option A is wrong because data archival refers to the long-term storage of data that is no longer actively used, which happens after ingestion and processing. Option B is wrong because data sharing involves making data available to other users or systems, which is a later stage. Option C is wrong because data deletion is the final stage of the lifecycle, where data is removed, not retrieved.

739
MCQeasy

A hospital wants to analyze patient readmission rates. The data contains daily patient visits. What is the level of granularity?

A.Patient
B.Visit
C.Day
D.Hospital
AnswerB

Granularity describes the finest level at which data is captured. Each row represents a single daily patient visit, so the visit is the unit of analysis, not the patient, admission or hospital, enabling readmission calculations from visit-level records.

Why this answer

The level of granularity refers to the finest detail captured in the dataset. Since the data contains daily patient visits, each record represents a single visit event, not the patient or the day itself. Therefore, 'Visit' is the correct granularity because each row corresponds to one visit occurrence.

Exam trap

The trap here is confusing the subject of analysis (patient readmission rates) with the actual data granularity (each row is a visit), leading candidates to incorrectly select 'Patient' instead of 'Visit'.

How to eliminate wrong answers

Option A is wrong because 'Patient' would be the granularity if the data summarized all visits per patient (e.g., one row per patient with aggregated readmission counts), but here each visit is a separate record. Option C is wrong because 'Day' would be the granularity if the data aggregated all visits per day (e.g., total visits per day), but the data contains individual visit records, not daily summaries. Option D is wrong because 'Hospital' would be the granularity if the data aggregated across the entire hospital (e.g., total readmission rate for the hospital), but the data is at the individual visit level.

740
MCQhard

A data analyst is cleaning a dataset with missing values in a time series of daily temperatures. The missing values occur sporadically. Which imputation method is most appropriate to maintain the temporal trend?

A.Forward-fill
B.Mean imputation
C.Median imputation
D.Interpolation
AnswerD

Interpolation estimates missing points from neighbouring known values along the time axis, preserving the daily temperature trend and seasonality. Mean or median imputation flattens local variation, whereas interpolation maintains temporal continuity, satisfying the requirement to keep the series' trend intact.

Why this answer

Interpolation estimates missing values by using the values immediately before and after the gap, producing a smooth curve that preserves the temporal trend of a time series. For sporadically missing daily temperatures, linear interpolation between neighboring days is the most faithful reconstruction. It respects the ordered, continuous nature of the data better than simple fill methods.

Exam trap

The trap is that mean and median imputation are the most commonly taught missing-value fixes, so candidates default to them without recognizing that time-ordered data requires methods that respect temporal continuity.

How to eliminate wrong answers

Option A is wrong because forward-fill copies the last observed value into the gap, creating flat plateaus that distort the trend and are especially poor when values change steadily over time. Option B is wrong because mean imputation replaces missing values with the overall average, ignoring time order and flattening seasonal or trend variation. Option C is wrong because median imputation has the same flaw as mean imputation — it ignores temporal position and injects a constant that does not reflect local conditions.

741
MCQmedium

A retail company analyzes customer purchase data to improve inventory management. They store daily transaction records in a relational database and monthly aggregate reports in a data warehouse. Which difference between these storage methods best explains why the warehouse is more suitable for trend analysis?

A.The database uses a star schema while the warehouse uses a normalized schema.
B.The database enforces ACID transactions, while the warehouse uses eventual consistency.
C.The database is optimized for write-heavy OLTP, while the warehouse is optimized for read-heavy OLAP.
D.The database stores only current data, while the warehouse stores historical data.
AnswerC

The relational database handles write-heavy OLTP workloads with frequent row-level inserts and updates, while the warehouse uses columnar, read-heavy OLAP designs optimised for scanning and aggregating large historical datasets. That architectural difference is what makes the warehouse suitable for trend analysis.

Why this answer

OLTP databases are optimized for high-frequency write operations (INSERT/UPDATE/DELETE) and ACID compliance, making them ideal for transaction processing but poor for complex analytical queries. In contrast, a data warehouse is optimized for read-heavy OLAP workloads, using columnar storage, pre-aggregated tables, and indexing strategies that enable fast aggregation and trend analysis over large historical datasets. This architectural difference directly supports the retail company's need to analyze purchase trends over time.

Exam trap

CompTIA often tests the misconception that 'data warehouses only store historical data' (Option D) as the primary reason for trend analysis suitability, but the real differentiator is the workload optimization (OLTP vs. OLAP), not merely the presence of history.

How to eliminate wrong answers

Option A is wrong because a star schema (with fact and dimension tables) is actually typical of data warehouses for analytical queries, while OLTP databases usually use normalized schemas to reduce redundancy and maintain data integrity. Option B is wrong because data warehouses often support ACID or snapshot isolation for consistency, and eventual consistency is more characteristic of NoSQL systems, not traditional data warehouses. Option D is wrong because relational databases can store historical data as well; the key difference is not the presence of history but the optimization for read-heavy analytical queries versus write-heavy transactional processing.

742
MCQeasy

A marketing team conducted a customer satisfaction survey for five different departments (Sales, Support, Billing, Shipping, Returns). The survey asked customers to rate their satisfaction on a scale of 1 (Very Dissatisfied) to 5 (Very Satisfied). The data is ordinal and the team wants to visualize the distribution of responses for each department to quickly see which department has the most 'Very Satisfied' customers and which has the most 'Very Dissatisfied'. They also want to compare the spread of responses across departments. Which chart type should they use?

A.Stacked bar chart with departments on x-axis and counts of each rating stacked
B.Line chart with departments on x-axis and average rating on y-axis
C.Box plot for each department
D.Scatter plot with department as category and satisfaction score as value
AnswerA

Stacked bars place departments on the x-axis with rating counts stacked, so each department's proportion of Very Satisfied and Very Dissatisfied is directly comparable. Stacking preserves the ordinal categories while showing spread across departments, satisfying both stem requirements.

Why this answer

A stacked bar chart with departments on the x-axis and counts of each rating stacked shows the full distribution of ordinal responses per department, making it easy to compare which department has the most 'Very Satisfied' (top stack) and 'Very Dissatisfied' (bottom stack) responses. It also allows visual comparison of the spread across departments by comparing stack heights and segment proportions.

Exam trap

DA0-002 often tests whether candidates choose a visualization that preserves the full distribution of ordinal data rather than one that reduces it to a summary statistic like the mean or median.

How to eliminate wrong answers

Option B is wrong because a line chart of average rating collapses ordinal distribution into a single mean, hiding the spread and the counts of extreme responses. Option C is wrong because a box plot summarizes median, quartiles, and outliers but does not directly show counts of each rating category, making it harder to see which department has the most 'Very Satisfied' or 'Very Dissatisfied' responses. Option D is wrong because a scatter plot with department as a category and satisfaction score as value does not effectively show distribution counts or spread across ordinal categories; it is better for correlation between two continuous variables.

743
MCQhard

An analyst writes a SQL query that uses a window function: SELECT employee_id, salary, LAG(salary, 1) OVER (ORDER BY salary DESC) AS prev_salary FROM employees. What does the LAG function return for the row with the highest salary?

A.The same salary value
B.NULL
C.The next highest salary
D.Zero
AnswerB

LAG(salary, 1) OVER (ORDER BY salary DESC) fetches the preceding row's salary in descending order. The highest salary occupies the first row, which has no predecessor, so the function returns NULL rather than a wrapped or default value.

Why this answer

LAG(salary, 1) OVER (ORDER BY salary DESC) returns the salary value from the previous row in the window ordering. For the row with the highest salary, there is no preceding row, so LAG returns NULL. This is the defined behavior of LAG when the offset goes beyond the partition boundary.

Exam trap

DA0-002 often tests the boundary behavior of LAG/LEAD — candidates assume the first row returns the current value or zero, when the correct answer is NULL because no preceding row exists.

How to eliminate wrong answers

Option A is wrong because LAG does not return the current row's value — that would be the behavior of a self-reference or FIRST_VALUE, not LAG. Option C is wrong because the next highest salary is the value of the following row, which is what LEAD would return, not LAG. Option D is wrong because LAG returns NULL, not zero, when there is no preceding row; zero would only appear if the underlying column value were zero or if COALESCE were applied.

744
MCQmedium

An analyst is sampling a large customer database to estimate the average purchase amount. To ensure that the sample proportionally represents different customer segments (e.g., age groups), which sampling method should be used?

A.Systematic sampling
B.Simple random sampling
C.Cluster sampling
D.Stratified sampling
AnswerD

Stratified sampling divides the population into distinct strata — here, age groups — then draws proportionally from each, guaranteeing every segment is represented in the sample. This directly satisfies the stem's requirement for proportional representation across customer segments, unlike simple random sampling, which could under-represent smaller groups.

Why this answer

Stratified sampling divides the population into homogeneous subgroups (strata) based on a characteristic like age group, then samples proportionally from each stratum. This guarantees representation of every segment in the sample, which is exactly what the analyst needs to estimate the average purchase amount accurately across customer segments.

Exam trap

DA0-002 often tests sampling method recognition — candidates confuse stratified (proportional representation of known subgroups) with cluster (sampling whole groups) because both involve dividing the population.

How to eliminate wrong answers

Option A is wrong because systematic sampling selects every k-th element from an ordered list — it doesn't guarantee proportional representation of subgroups and can introduce periodicity bias if the list has a repeating pattern. Option B is wrong because simple random sampling gives every individual an equal chance but can, by chance, under- or over-represent small segments, especially in small samples. Option C is wrong because cluster sampling divides the population into clusters (e.g., stores, cities) and samples entire clusters — it reduces cost but increases variance and does not ensure segment proportionality.

745
MCQhard

A data engineer is designing a system to store raw sensor data from thousands of IoT devices. The data will be used later for various analytics projects, but the schema is not yet defined. Which storage solution is most appropriate?

A.Data lake
B.Data mart
C.Data warehouse
D.Relational database
AnswerA

A data lake stores raw data in its native format without requiring a predefined schema, satisfying the undefined-schema constraint. It handles high-volume ingestion from thousands of IoT devices and supports diverse downstream analytics, including structured, semi-structured and unstructured data, which a schema-on-write warehouse cannot accommodate.

Why this answer

A data lake is designed to store raw, unprocessed data in its native format without requiring a predefined schema, which is exactly what is needed for IoT sensor data whose schema is not yet defined. It supports schema-on-read, allowing analytics projects to interpret the data later as requirements evolve. This makes it the most appropriate choice for storing diverse, high-volume raw data for future analytics.

Exam trap

DA0-002 often tests the distinction between schema-on-write (data warehouse) and schema-on-read (data lake), so candidates who pick a data warehouse because it 'stores data for analytics' miss the requirement that the schema is not yet defined.

How to eliminate wrong answers

Option B is wrong because a data mart is a subset of a data warehouse focused on a specific business line, and it requires a defined schema and transformed data, which is not suitable for raw, schema-less IoT data. Option C is wrong because a data warehouse stores structured, processed data with a predefined schema (schema-on-write), which contradicts the requirement that the schema is not yet defined. Option D is wrong because a relational database requires a fixed schema and is not designed for the volume and variety of raw IoT sensor data.

746
Multi-Selectmedium

A company is acquiring social media data via a public API. Which TWO considerations are important for ensuring ethical and legal compliance?

Select 2 answers
A.Share raw data with third parties for additional insights
B.Use the data for any internal analysis without restrictions
C.Anonymize personal identifiable information (PII) before storage
D.Cache data indefinitely to avoid repeated API calls
E.Comply with the platform's terms of service
AnswersC, E

Social media data routinely contains personal identifiers, so anonymising PII before storage limits re-identification risk and supports privacy regulation compliance. Removing or masking identifiers at ingestion satisfies the ethical and legal requirement to protect data subjects throughout the acquisition pipeline.

Why this answer

Option C is correct because anonymizing personally identifiable information (PII) before storage reduces privacy risk and helps satisfy data-protection regulations such as GDPR and CCPA, which require limiting the processing and retention of identifiable personal data. Option E is correct because a public API is governed by the platform's terms of service, and using the data outside those contractual permissions—such as for prohibited purposes or beyond rate limits—can constitute a legal and ethical violation. Options A and B are not appropriate because sharing raw data with third parties or using it for unrestricted internal analysis can breach privacy obligations and the API's usage terms.

Option D is also incorrect because caching data indefinitely conflicts with data-minimization and retention-limitation principles and may violate the platform's terms.

Exam trap

The trap here is that candidates may confuse 'caching for efficiency' (Option D) with ethical compliance, overlooking that indefinite storage violates data minimization principles and platform terms, while 'internal analysis' (Option B) seems harmless but ignores explicit usage restrictions in the API's terms of service.

747
MCQmedium

An OLTP system processes thousands of transactions per second. Which property ensures that a transaction is fully completed or fully rolled back, preventing partial updates?

A.Isolation
B.Durability
C.Atomicity
D.Consistency
AnswerC

Atomicity guarantees each transaction executes as a single indivisible unit, so either every operation commits or none does. This directly satisfies the stem's requirement to prevent partial updates under high-volume OLTP load, where a mid-transaction failure would otherwise leave inconsistent data committed.

Why this answer

Atomicity guarantees that a transaction is treated as a single unit, completed entirely or not at all.

748
MCQmedium

A data analyst is reviewing a dataset containing house prices. The mean price is $350,000 and the median is $280,000. Which of the following best describes the distribution of house prices?

A.The distribution is right-skewed.
B.The distribution is symmetric.
C.The distribution is left-skewed.
D.The distribution is bimodal.
AnswerA

A mean above the median indicates right-skew, where high-value outliers pull the mean upward while the median stays resistant. This satisfies the stem's constraint of mean ($350,000) exceeding median ($280,000), confirming a positively skewed distribution of house prices.

Why this answer

When the mean ($350,000) is greater than the median ($280,000), the distribution is right-skewed (positively skewed). In a right-skewed distribution, a long tail of high values pulls the mean upward above the median. This is common with house prices, where a few very expensive homes inflate the mean.

Exam trap

DA0-002 often tests the mean-vs-median relationship: candidates who forget that mean > median indicates right skew (not left) pick the wrong direction.

How to eliminate wrong answers

Option B is wrong because in a symmetric distribution the mean and median are approximately equal, but here they differ by $70,000. Option C is wrong because left-skewness (negative skew) produces a mean lower than the median, the opposite of what is observed. Option D is wrong because bimodality refers to two distinct peaks in the distribution and cannot be inferred from mean and median alone.

749
MCQeasy

A data analyst is creating a report for a marketing campaign. The campaign data includes customer names, email addresses, and purchase history. Which of the following best describes the 'customer name' data type?

A.Nominal
B.Quantitative
C.Ordinal
D.Discrete
AnswerA

Customer names are labels with no inherent order or numeric meaning, so they are categorical. Nominal is the correct level of measurement because values merely distinguish one customer from another, unlike ordinal, interval or ratio data.

Why this answer

Customer names are categorical labels that identify individuals without any inherent order or numerical value. This fits the definition of nominal data, which is used for naming or classifying variables. In data analysis, nominal data can be stored as strings and used for grouping or filtering, but arithmetic operations are meaningless.

Exam trap

CompTIA often tests the distinction between nominal and ordinal data by presenting a label that could be mistaken for having an order (e.g., 'customer name' might be confused with 'rank' or 'tier'), but the trap here is that names are purely categorical with no intrinsic ranking.

How to eliminate wrong answers

Option B is wrong because quantitative data represents numerical measurements or counts (e.g., purchase amount), not text labels like names. Option C is wrong because ordinal data has a meaningful order or rank (e.g., customer satisfaction rating), but customer names have no inherent sequence. Option D is wrong because discrete data consists of countable numerical values (e.g., number of purchases), whereas customer names are non-numeric categories.

750
MCQmedium

A logistics company uses a key-value store to cache real-time shipment tracking updates. The data model has no fixed schema, and each shipment record may contain different attributes. Which data storage approach is being used?

A.Relational database
B.Data warehouse
C.Data lake
D.NoSQL key-value store
AnswerD

A key-value store maps unique keys to arbitrary values without requiring a fixed schema. Each shipment record can be stored under its shipment ID as the key, and the value can contain any set of attributes. This matches the description of schema-less, flexible storage for real-time caching of tracking updates.

Why this answer

The scenario describes storing shipment records under unique identifiers with varying attributes and no fixed schema, which is the defining characteristic of a key-value store. Key-value stores excel at fast, simple lookups by key, making them ideal for caching real-time tracking updates. Relational databases, warehouses, and lakes impose structure or latency that conflict with these requirements.

Exam trap

The trap here is equating any schema-less storage with a data lake, when the key-based, low-latency access pattern points specifically to a key-value store.

Page 9

Page 10 of 14

Page 11