Courseiva

CompTIA Data+ (DA0-002) (DA0-002) — Questions 601–675

1004 questions total · 14pages · All types, answers revealed

Page 8

Page 9 of 14

Page 10
601
Multi-Selectmedium

A data analyst is building a supervised learning model to predict customer churn. The target variable is binary (churn = yes/no). Which TWO modeling techniques are appropriate for this task? (Select two.)

Select 2 answers
A.K-means clustering
B.Linear regression
C.Logistic regression
D.Decision trees
E.Apriori algorithm
AnswersC, D

Logistic regression models the probability of a binary outcome by applying the logistic (sigmoid) function to a linear combination of predictors, bounding output between 0 and 1. This directly suits the churn yes/no target, unlike ordinary linear regression.

Why this answer

Logistic regression (C) is correct because it is a supervised classification technique that models the probability of a binary outcome (churn yes/no) using a sigmoid function, making it a standard choice for binary targets. Decision trees (D) are also correct because they are supervised classifiers that recursively split features to predict categorical class labels, and they handle binary targets naturally while offering interpretability. K-means clustering (A) is wrong because it is an unsupervised algorithm that groups unlabeled data and cannot predict a labeled binary target.

Linear regression (B) is wrong because it predicts continuous numeric values rather than class probabilities or discrete categories, so it is unsuitable for binary classification. The Apriori algorithm (E) is wrong because it is an unsupervised association-rule mining method for finding frequent itemsets, not a predictive classification model.

Exam trap

The trap here is that candidates may confuse unsupervised clustering (K-means) or association rule mining (Apriori) with supervised classification, or mistakenly think linear regression can be adapted for binary outcomes without transformation.

602
MCQmedium

A company wants to collect real-time clickstream data from its website. Which acquisition method is most suitable?

A.Streaming API
B.Web scraping
C.Batch processing nightly
D.Manual entry
AnswerA

A streaming API ingests events continuously as they occur, satisfying the real-time clickstream requirement. Unlike batch extraction, which introduces latency by collecting data at scheduled intervals, streaming delivers each click immediately for processing. This makes it the suitable acquisition method when low-latency, continuous event capture is the constraint.

Why this answer

A streaming API is the most suitable method for collecting real-time clickstream data because it enables continuous, low-latency ingestion of events as they occur. Unlike batch or manual methods, a streaming API (e.g., using WebSockets or HTTP/2 Server-Sent Events) pushes each click event immediately to the data pipeline, satisfying the real-time requirement.

Exam trap

CompTIA often tests the distinction between 'real-time' and 'near-real-time' or 'batch' methods, and the trap here is that candidates may confuse web scraping (which can be automated frequently) with true streaming, not realizing that scraping is still a pull-based, scheduled operation that cannot match the push-based immediacy of a streaming API.

How to eliminate wrong answers

Option B (Web scraping) is wrong because it is a pull-based technique that typically retrieves static HTML pages at intervals, not real-time event streams, and is inefficient for high-frequency click data. Option C (Batch processing nightly) is wrong because it introduces a delay of up to 24 hours, failing the real-time requirement. Option D (Manual entry) is wrong because it is error-prone, non-scalable, and cannot capture high-velocity clickstream data in real time.

603
MCQmedium

In a customer database, each row represents a customer with columns: CustomerID, Name, Address, Phone. What does the column "Name" represent?

A.Instance
B.Entity
C.Attribute
D.Record
AnswerC

Name is an attribute: a column describing a characteristic of the customer entity, whose rows are instances. It satisfies the stem's requirement by identifying the property recorded for each CustomerID, distinct from the entity itself.

Why this answer

In the context of a relational database, a column represents an attribute of an entity. The 'Name' column stores a specific characteristic (the customer's name) for each row, making it an attribute. This aligns with the data modeling concept where attributes define the properties of an entity.

Exam trap

The trap here is that candidates confuse 'attribute' with 'record' because they think of a row as containing all attributes, but the question specifically asks what a single column represents, not the row itself.

How to eliminate wrong answers

Option A is wrong because an instance refers to a single occurrence of an entity (e.g., a specific customer row), not a column. Option B is wrong because an entity is a table-level concept representing a real-world object (e.g., the Customer table), not a column within it. Option D is wrong because a record is a row in the table, which contains values for all attributes, not a single column like 'Name'.

604
MCQhard

A data analyst is preparing a dataset for a machine learning model. The dataset contains a categorical column 'color' with values 'red', 'green', 'blue', and 'yellow'. The analyst needs to transform this column into a numerical format suitable for the model. Which technique should be used?

A.One-hot encoding
B.Binary encoding
C.Hashing
D.Label encoding
AnswerA

One-hot encoding creates binary columns for each category, representing the presence of each color. This avoids implying any ordinal relationship, making it ideal for nominal categorical data like colors. The model can then treat each color as a separate feature without assuming order, which is essential for accurate learning.

Why this answer

One-hot encoding is the correct technique because it creates separate binary features for each color, eliminating any implied order. This is crucial for nominal data where categories have no inherent ranking. Other encoding methods like label encoding introduce ordinality, which can degrade model performance by suggesting false relationships.

Exam trap

The trap here is choosing label encoding for its simplicity, but it incorrectly imposes an order on nominal categories, which can mislead the model.

605
Multi-Selectmedium

A data governance team is establishing policies for data quality. Which THREE of the following are common dimensions of data quality? (Select 3)

Select 3 answers
A.Consistency
B.Completeness
C.Accuracy
D.Velocity
E.Volume
AnswersA, B, C

Data is uniform across systems.

Why this answer

Consistency is a common dimension of data quality because it ensures that data values are uniform across different datasets or systems, preventing contradictions. For example, if a customer's address is stored as '123 Main St' in one database and '123 Main Street' in another, consistency rules would flag this discrepancy. This dimension is critical for reliable reporting and integration.

Exam trap

The trap here is that candidates confuse the characteristics of big data (velocity, volume, variety) with the dimensions of data quality, leading them to select velocity or volume instead of the correct quality-focused options.

606
MCQhard

A financial services company is migrating its customer data from a legacy on-premises relational database to a cloud-based data warehouse. The legacy database uses a denormalized schema with a single table 'customer_master' that contains all customer attributes, including repeated groups for multiple accounts per customer (account1_type, account1_balance, account2_type, account2_balance, etc.). The data warehouse team wants to implement a normalized star schema with separate dimension and fact tables. During the ETL process, the team encounters an error: 'Data truncation: string data right truncation' when loading account_type values into the dim_account table. The account_type column in dim_account is defined as VARCHAR(10), but the source data contains account types like 'SavingsPlus' (11 characters) and 'CheckingPremium' (15 characters). The team must resolve this issue without losing data. Which course of action should the team take?

A.Truncate the account_type values to 10 characters during ETL.
B.Change the data type of dim_account.account_type to TEXT.
C.Ignore the error and continue loading with NULL values for truncated rows.
D.Increase the VARCHAR length of dim_account.account_type to accommodate the longest account type.
AnswerD

The truncation error occurs because dim_account.account_type is VARCHAR(10) while source values reach 15 characters. Widening the column to the longest value preserves every account type during load, satisfying the no-data-loss constraint without altering source records.

Why this answer

Increasing the VARCHAR length of dim_account.account_type to accommodate the longest account type (e.g., VARCHAR(15) for 'CheckingPremium') resolves the data truncation error without data loss. This aligns with the star schema design principle of preserving source data integrity while ensuring the column definition matches the actual data length. The team must avoid truncation or NULL insertion to maintain accurate dimensional attributes for analytics.

Exam trap

The trap here is that candidates may choose truncation (Option A) or NULL insertion (Option C) as quick fixes, overlooking the requirement to preserve data integrity, or mistakenly think TEXT (Option B) is a safe catch-all without considering performance implications in a data warehouse context.

How to eliminate wrong answers

Option A is wrong because truncating account_type values to 10 characters would lose data, violating the requirement to resolve the issue without data loss. Option B is wrong because changing the data type to TEXT is unnecessary and can introduce performance overhead in indexing and querying, as TEXT is a large object type not optimized for VARCHAR-like operations in a data warehouse. Option C is wrong because ignoring the error and loading NULL values for truncated rows would discard valid account_type data, breaking referential integrity and analytics accuracy.

607
Multi-Selecthard

A data analyst is building a narrative around a quarterly sales decline. The story should follow a narrative arc. Which THREE elements should be included in the story?

Select 3 answers
A.Random color scheme for every chart
B.Resolution: By diversifying suppliers, sales recovered in November.
C.Situation: Q3 sales were on track to meet targets.
D.Detailed description of each product's sales breakdown
E.Complication: A supply chain disruption caused a drop in October.
AnswersB, C, E

The resolution closes the narrative arc by stating the outcome and remedy, here supplier diversification restoring sales in November. It satisfies the stem's requirement that the story follow a narrative arc, giving the audience closure after the decline is described.

Why this answer

The question asks for the three elements of a narrative arc, which are Situation, Complication, and Resolution. Option C is correct because it establishes the Situation, the baseline context that Q3 sales were on track to meet targets, which sets the stage for the story. Option E is correct because it introduces the Complication, the supply chain disruption that caused the October drop, creating the conflict or turning point.

Option B is correct because it provides the Resolution, explaining how diversifying suppliers led to recovery in November, which closes the narrative arc. Option A is not correct because a random color scheme for every chart harms visual consistency and does not contribute to a narrative arc. Option D is not correct because a detailed description of each product's sales breakdown is granular data detail, not a structural element of the narrative arc.

Exam trap

The trap is selecting 'detailed breakdown' (Option D) as a narrative element — candidates confuse supporting data with narrative structure, but a narrative arc requires Situation, Complication, and Resolution, not exhaustive detail.

608
MCQmedium

A data analyst needs to compare sales data from the company's internal CRM with public demographic data from a government census. Which data concept best describes this scenario?

A.Internal vs. External data
B.Primary vs. Secondary data
C.Structured vs. Unstructured data
D.Quantitative vs. Qualitative data
AnswerA

The CRM data originates within the organisation, making it internal, while the census data comes from an outside government body, making it external. Combining both satisfies the stem's requirement to compare proprietary sales figures against public demographic information, which is precisely the internal versus external data distinction.

Why this answer

The scenario involves comparing internal CRM data (generated and owned by the company) with external government census data (publicly sourced from outside the organization). This directly maps to the Internal vs. External data concept, where internal data is collected within the enterprise (e.g., sales transactions, customer records) and external data is acquired from third-party sources (e.g., census bureaus, market research firms).

The key distinction is the data's origin and ownership, not its structure, collection method, or measurement type.

Exam trap

CompTIA often tests the Internal vs. External data concept by presenting a scenario where the key differentiator is the data's source (inside vs. outside the organization), tempting candidates to confuse it with Primary vs. Secondary data, which focuses on whether the data was collected firsthand or repurposed.

How to eliminate wrong answers

Option B (Primary vs. Secondary data) is wrong because both datasets could be primary (collected firsthand by the CRM or census) or secondary (repurposed from another source), but the question focuses on the origin relative to the organization, not the collection method. Option C (Structured vs.

Unstructured data) is wrong because both CRM sales data and census demographic data are typically structured (e.g., tables with rows and columns), so the contrast is not about format but about source. Option D (Quantitative vs. Qualitative data) is wrong because both datasets contain quantitative values (e.g., sales figures, population counts) and possibly qualitative labels (e.g., region names), but the core distinction in the scenario is internal versus external sourcing, not measurement scale.

609
MCQhard

A data analyst at an e-commerce company is profiling a transactions table with ten million rows. The 'order_total' column is stored as text and contains values like '$1,299.00', a few entries of 'N/A', and several negative amounts that represent refunds. The analyst must produce a numeric column suitable for summation in a revenue report. Which sequence of operations best prepares the column?

A.Strip the currency symbol and thousands separators, convert non-numeric entries to null or a flagged value, preserve negative signs for refunds, then cast to decimal.
B.Cast the column directly to a decimal type, letting the database coerce invalid strings to nulls.
C.Replace all non-numeric characters with zeros, then sum the column as text.
D.Take the absolute value of every entry first, then remove the dollar signs and cast to integer.
AnswerA

This sequence removes formatting characters that block conversion, explicitly handles the 'N/A' entries rather than letting them fail silently, keeps negative amounts so refunds reduce revenue correctly, and performs a deliberate cast to a numeric type. It also leaves a clear audit trail of which rows were invalid, which matters when the summed total feeds a financial report.

Why this answer

The column mixes formatting characters, non-numeric placeholders, and meaningful negative values. A robust cleanup removes currency symbols and separators, explicitly maps invalid entries to nulls or a flagged value, retains the negative sign for refunds, and then casts to a decimal type. Direct casting and character-stripping shortcuts either lose rows silently or corrupt the sign and precision of the amounts.

Exam trap

The trap here is trusting an implicit cast to clean currency-formatted text, when the dollar signs, commas, and 'N/A' entries actually cause errors or silent nulls that understate revenue.

610
Multi-Selectmedium

A hospital analytics team is preparing to share a de-identified patient outcomes dataset with an external research partner. The privacy officer requires that the dataset satisfy governance requirements for lawful secondary use while preserving analytical utility. Which two practices should the team apply? (Choose two.)

Select 2 answers
A.Retain the partner's access indefinitely so future research questions can reuse the dataset.
B.Share the dataset over an unencrypted file transfer so the partner can download it faster.
C.Remove or generalize all direct identifiers such as names, medical record numbers, and full dates of birth.
D.Execute a data-sharing agreement that defines permitted use, retention limits, and breach obligations.
E.Publish the dataset on a public website to streamline distribution to the research community.
AnswersC, D

Removing or generalizing direct identifiers reduces re-identification risk and aligns with de-identification expectations for secondary use. Direct identifiers like names and medical record numbers let anyone link records to a person, so they must be stripped or generalized before sharing. This step preserves analytical value in the remaining clinical variables while satisfying the privacy officer's governance requirement.

Why this answer

Lawful secondary use of health data combines de-identification of direct identifiers with a binding data-sharing agreement that scopes purpose, retention, and breach duties. Together they reduce re-identification risk and create enforceable accountability. Unencrypted transfer, indefinite access, and public release each undermine those controls and fail the privacy officer's governance requirements.

Exam trap

The trap here is treating de-identification as sufficient on its own, when governance also requires contractual controls that limit purpose, retention, and onward sharing.

611
MCQmedium

An analyst needs to show the part-to-whole relationship of market share among four competitors. Which chart type is most appropriate, considering best practices?

A.Pie chart
B.Scatter plot
C.Histogram
D.Box plot
AnswerA

A pie chart maps each competitor's slice to the whole market, so four segments clearly convey part-to-whole proportions. With only four categories, the slice count stays readable, satisfying best practise for showing market share composition without overwhelming the viewer.

Why this answer

A pie or donut chart is suitable for part-to-whole with few categories (5-7 slices max). Here, four competitors fit well.

612
MCQmedium

A data analyst wants to ensure a sample proportionally represents different regions in a population. Which sampling method should be used?

A.Simple random sampling
B.Cluster sampling
C.Systematic sampling
D.Stratified sampling
AnswerD

Stratified sampling divides the population into distinct regions (strata) and draws samples from each in proportion to its size, directly satisfying the requirement for proportional regional representation. Unlike simple random sampling, it guarantees every region appears at its correct weight, eliminating the under-representation that random chance can produce.

Why this answer

Stratified sampling divides the population into distinct subgroups (strata) based on a shared characteristic — here, region — and then draws a proportional random sample from each stratum. This guarantees that each region is represented in the sample in proportion to its size in the population, which is exactly what the analyst requires. Simple random sampling could, by chance, under- or over-represent certain regions.

Exam trap

DA0-002 often tests the distinction between stratified sampling (proportional representation of known subgroups) and cluster sampling (sampling whole naturally occurring groups), which candidates frequently confuse because both involve dividing the population.

How to eliminate wrong answers

Option A is wrong because simple random sampling selects individuals from the entire population without regard to region, so small regions may be underrepresented or missed entirely by chance. Option B is wrong because cluster sampling divides the population into clusters and samples entire clusters, which reduces cost but does not guarantee proportional representation of each region. Option C is wrong because systematic sampling picks every kth element from a list, which can introduce periodicity bias and does not ensure proportional regional representation.

613
MCQhard

In time series decomposition, a pattern that repeats at regular intervals (e.g., weekly, yearly) is called:

A.Cyclical
B.Irregular
C.Trend
D.Seasonality
AnswerD

Seasonality describes a repeating pattern occurring at fixed intervals, such as weekly or yearly cycles, distinct from trend and irregular components. This matches the stem's definition of a pattern repeating at regular intervals within time series decomposition.

Why this answer

Seasonality refers to a pattern that repeats at fixed, calendar-based intervals such as weekly, monthly, quarterly, or yearly. Because the question specifies 'regular intervals' like weekly or yearly, this is the defining characteristic of the seasonal component in time series decomposition. Seasonality is distinct from trend (long-term direction) and cyclical (irregular multi-year fluctuations).

Exam trap

The trap here is confusing cyclical with seasonal — both are repeating patterns, but cyclical has an irregular, multi-year period while seasonality has a fixed calendar period, and the exam relies on candidates glossing over that distinction.

How to eliminate wrong answers

Option A is wrong because cyclical patterns occur over longer, irregular periods (typically 2+ years) tied to economic or business cycles, not fixed calendar intervals. Option B is wrong because the irregular component represents random, unpredictable residual noise after trend, seasonal, and cyclical components are removed. Option C is wrong because trend describes the long-term upward or downward movement of the series over time, not a repeating pattern at fixed intervals.

614
MCQmedium

A data analyst is cleaning a dataset and finds that some cells in the 'email' column contain leading spaces. Which string function should be used to remove these spaces?

A.TRIM
B.LTRIM
C.REPLACE
D.SUBSTRING
AnswerA

TRIM removes leading and trailing whitespace characters from a string, directly eliminating the leading spaces found in the email column cells. Other string functions such as REPLACE or LTRIM target different patterns or only one side, whereas TRIM satisfies the exact cleaning requirement stated.

Why this answer

The TRIM function removes leading and trailing spaces (and sometimes other whitespace) from a string, which directly addresses the leading spaces in the email column. In most SQL dialects and data tools, TRIM is the standard function for stripping both leading and trailing spaces, making it the correct choice for cleaning the data.

Exam trap

The trap is selecting LTRIM because the question mentions 'leading spaces,' but TRIM is the more complete and standard function for removing spaces from both ends, and exams often expect the general-purpose function.

How to eliminate wrong answers

Option B is wrong because LTRIM only removes leading spaces, not trailing spaces; while it would fix the leading spaces, TRIM is more comprehensive and is the standard answer for removing spaces from both ends. Option C is wrong because REPLACE substitutes all occurrences of a specified substring, which could remove internal spaces in emails and is not targeted at leading spaces. Option D is wrong because SUBSTRING extracts a portion of a string based on position and length; it does not remove spaces.

615
MCQeasy

A data analyst creates a bar chart showing monthly sales for the current year. The CEO asks to see the same data but grouped by region. Which chart type is most appropriate for this new requirement?

A.Grouped bar chart
B.Pie chart
C.Stacked bar chart
D.Line chart
AnswerA

A grouped bar chart plots a second categorical dimension—region—as clustered bars within each monthly group, so both month and region remain visible. A simple bar chart cannot encode the added region breakdown, and a stacked chart obscures per-region comparison.

Why this answer

A grouped bar chart is the most appropriate choice because it allows the CEO to compare monthly sales across different regions side by side within each month. This chart type effectively displays two categorical variables (month and region) and one quantitative variable (sales), making it easy to see both regional performance and monthly trends simultaneously.

Exam trap

CompTIA often tests the distinction between grouped and stacked bar charts, trapping candidates who think any multi-category bar chart is equivalent, when in fact grouped bars are for comparing individual category values and stacked bars are for comparing totals and proportions.

How to eliminate wrong answers

Option B (Pie chart) is wrong because pie charts are designed to show parts of a whole for a single categorical variable at a single point in time, not to compare multiple categories (regions) across a time series (months). Option C (Stacked bar chart) is wrong because while it can show regions within months, it emphasizes the total sales per month and the proportion each region contributes, rather than allowing direct comparison of individual region sales across months. Option D (Line chart) is wrong because line charts are best for showing continuous trends over time for one or more series, but they do not effectively compare discrete categories like regions within each month; a grouped bar chart provides clearer categorical comparison.

616
MCQhard

A financial analytics team is building a data warehouse to support complex analytical queries on historical stock trades. The data volume is in terabytes, and queries frequently join multiple large tables and perform aggregations. The team needs a storage model that minimizes query latency for these read-heavy analytical workloads. Which data modeling approach is most appropriate?

A.Highly normalized OLTP schema
B.Dimensional star schema
C.Entity-attribute-value (EAV) model
D.Flat file with no indexing
AnswerB

A dimensional star schema organizes data into fact tables (e.g., trades) and denormalized dimension tables (e.g., date, stock, broker). This design reduces the number of joins, enables efficient aggregations, and is optimized for read-heavy analytical queries. It also supports columnar storage and partitioning, making it ideal for terabyte-scale historical trade analysis with minimal query latency.

Why this answer

A dimensional star schema is designed for analytical workloads, using fact and dimension tables to minimize joins and optimize aggregations. It supports columnar storage and partitioning, which are critical for terabyte-scale read-heavy queries. Normalized OLTP, EAV, and flat files each introduce performance bottlenecks or lack the necessary optimizations for complex analytical queries on historical stock trades.

Exam trap

The trap here is assuming that a normalized schema is always best for data integrity, overlooking that analytical workloads require denormalized dimensional models for performance.

617
Multi-Selecthard

An analyst is using SQL to analyze employee data. Which THREE of the following are valid uses of the WHERE clause? (Select three.)

Select 3 answers
A.Sort the result set by hire_date
B.Filter groups after aggregation using HAVING
C.Filter rows where manager_id is NULL using IS NULL
D.Filter rows where the name starts with 'J' using LIKE
E.Filter rows where salary is between 50,000 and 70,000 using BETWEEN
AnswersC, D, E

IS NULL tests for the absence of a value, correctly filtering rows where manager_id holds no data. Equality operators cannot match NULL because SQL treats it as unknown, so IS NULL is the only valid predicate for this filter.

Why this answer

Option C is correct because the WHERE clause can test for NULL values with the IS NULL predicate, so filtering rows where manager_id is NULL is a valid row-level filter. Option D is correct because WHERE supports the LIKE operator for pattern matching, so filtering names that start with 'J' (e.g., name LIKE 'J%') is a valid use. Option E is correct because WHERE supports the BETWEEN operator for range comparisons, so filtering salaries between 50,000 and 70,000 is a valid row-level filter.

Option A is incorrect because sorting the result set is done with the ORDER BY clause, not WHERE. Option B is incorrect because filtering groups after aggregation is performed with the HAVING clause, not WHERE, which filters rows before grouping.

Exam trap

DA0-002 often tests the WHERE vs HAVING distinction — candidates incorrectly use WHERE for aggregate filtering or confuse ORDER BY (sorting) with WHERE (filtering).

618
MCQhard

In Power BI, a developer needs to create a measure that calculates total sales for the same period last year. Which DAX function should be used?

A.TOTALYTD
B.PARALLELPERIOD
C.DATEADD
D.SAMEPERIODLASTYEAR
AnswerD

SAMEPERIODLASTYEAR returns a table of dates shifted back exactly one year, which the measure then evaluates against, satisfying the year-over-year comparison. It requires a marked date table, unlike DATEADD, which shifts by an arbitrary interval.

Why this answer

SAMEPERIODLASTYEAR is a DAX time intelligence function that returns a table of dates shifted back exactly one year, ideal for year-over-year comparisons. It is specifically designed for calculating metrics like total sales for the same period last year when used with a date table. This function is simple and directly addresses the requirement.

Exam trap

DA0-002 often tests the confusion between time intelligence functions, tricking candidates into choosing TOTALYTD or DATEADD when the question specifically asks for same period last year.

How to eliminate wrong answers

Option A is wrong because TOTALYTD calculates year-to-date totals for the current year, not the same period last year. Option B is wrong because PARALLELPERIOD shifts by a specified number of intervals (e.g., -1 year) but is more flexible and often used for non-standard periods; it can work but is not the most direct function for same period last year. Option C is wrong because DATEADD shifts dates by a specified interval but requires a date column and is more general; it can achieve the result but SAMEPERIODLASTYEAR is the purpose-built function.

619
MCQmedium

A scatter plot of advertising spend vs. revenue shows no clear correlation, but the analyst suspects a relationship exists. Which addition to the plot could help reveal a hidden trend?

A.Change to a bar chart
B.Increase the marker size
C.Remove data points with low spend
D.Add a trendline
AnswerD

A trendline fits a regression line through the scatter plot, summarising the overall direction and strength of the relationship between advertising spend and revenue. This can expose a subtle or non-linear pattern that individual points obscure, revealing the hidden trend.

Why this answer

A trendline (regression line) added to a scatter plot can reveal a relationship that is not obvious from raw points alone, such as a nonlinear or weak correlation. It summarizes the overall direction and strength of the relationship, helping the analyst detect hidden trends. Changing chart type, marker size, or removing data points does not reveal underlying correlation.

Exam trap

The trap is thinking that changing visual properties (marker size, chart type, filtering) reveals trends — only adding a trendline or fitting a model exposes the underlying relationship.

How to eliminate wrong answers

Option A is wrong because a bar chart is for categorical comparisons, not for showing correlation between two continuous variables — it would obscure the relationship further. Option B is wrong because increasing marker size only changes visual emphasis, not the analytical insight; it does not reveal correlation. Option C is wrong because removing low-spend data points biases the dataset and can hide or distort the true relationship rather than reveal it.

620
MCQmedium

A data analyst at a retail chain is importing a CSV file into a database. The file contains a 'transaction_date' column with values like '2023-13-01' and '2023-02-30'. The target column is defined as DATE. The analyst needs to ensure that invalid dates are flagged and not loaded. Which approach best handles this data quality issue during acquisition?

A.Use a Python script with the pandas library to read the CSV and apply the to_datetime function with errors='coerce' before loading into the database.
B.Set the database column to accept NULL values and load all rows, allowing the database to automatically convert invalid dates to NULL.
C.Load the data into a staging table with the column as VARCHAR, then use a function like TRY_CAST or TO_DATE with error handling to identify invalid dates.
D.Use a regular expression to validate the date format and reject rows that do not match the pattern.
AnswerC

Loading into a staging table with a string column allows the data to be ingested without conversion errors. Then, using a function like TRY_CAST (SQL Server) or TO_DATE with error handling (Oracle) can attempt conversion and return NULL or an error for invalid dates. This isolates invalid rows for review, ensuring only valid dates move to the final DATE column.

Why this answer

The correct approach is to stage the data as strings and then use database functions that attempt conversion with error handling. This allows invalid dates to be identified without failing the entire load, and only valid dates are moved to the final DATE column. This method is robust, scalable, and integrates with SQL-based ETL processes.

Exam trap

The trap here is assuming that format validation (like regex) is sufficient to catch invalid dates, but it cannot detect semantic errors such as month 13 or February 30.

621
MCQhard

A data analyst is preparing a quarterly business review report for a retail company. The report includes a bar chart showing sales by region. The analyst notices that the chart is very wide and the region labels are overlapping, making it hard to read. The analyst wants to improve readability without losing any regions. Which action is most appropriate?

A.Rotate the region labels 90 degrees so they fit vertically under each bar.
B.Increase the width of the chart to provide more space between bars.
C.Convert the bar chart to a horizontal bar chart, placing regions on the y-axis.
D.Use a pie chart instead, as it naturally avoids label overlap by placing labels outside the slices.
AnswerC

A horizontal bar chart gives each region label ample horizontal space, eliminating overlap and making labels easy to read. It also accommodates long region names without rotation. This is a standard best practice when category labels are long or numerous. It preserves all regions and improves readability significantly.

Why this answer

Converting to a horizontal bar chart is the most effective solution because it provides ample space for region labels on the y-axis, eliminating overlap and improving readability. It maintains the ability to compare sales across regions. The other options either partially mitigate the issue or introduce new problems such as reduced comparability or awkward label orientation.

Exam trap

The trap here is assuming that rotating labels or widening the chart is sufficient, when the root cause is insufficient horizontal space for labels on the x-axis.

622
MCQeasy

A data analyst wants to visualize the monthly sales trend for the past year. Which chart type is most appropriate?

A.Bar chart
B.Pie chart
C.Line chart
D.Scatter plot
AnswerC

A line chart plots a continuous metric against a time axis, making month-to-month sales movement and trend direction immediately visible across the twelve periods. Categorical charts such as bar or pie cannot show the temporal progression the analyst needs.

Why this answer

Line charts are best for showing trends over time.

623
MCQmedium

A database administrator wants to ensure that every value in a column matches values in a primary key column of another table. Which constraint enforces this rule?

A.Unique constraint
B.Primary key
C.Check constraint
D.Foreign key
AnswerD

A foreign key constrains each value in the child column to exist in the referenced primary key column of the parent table, enforcing referential integrity and rejecting inserts or updates that would orphan the row.

Why this answer

A foreign key constraint enforces referential integrity by requiring that values in a column match values in the primary key (or unique key) column of another table. This is exactly the rule described — ensuring every value in one table's column corresponds to a primary key value in another table.

Exam trap

DA0-002 often tests the confusion between unique constraints and foreign keys, since both involve uniqueness — candidates must remember that only a foreign key enforces a cross-table reference to a primary key.

How to eliminate wrong answers

Option A is wrong because a unique constraint only ensures values within a single column (or set of columns) are distinct; it does not reference another table. Option B is wrong because a primary key uniquely identifies rows in its own table and does not enforce cross-table references. Option C is wrong because a check constraint validates a condition on column values (e.g., range or format) within the same table, not a relationship to another table.

624
Multi-Selecthard

A data analyst is working with a dataset that contains a column for 'Order Date' stored as a string in the format 'YYYY-MM-DD'. The analyst needs to perform time-series analysis, such as calculating monthly sales trends. Which two actions should the analyst take to prepare the data for this analysis? (Choose two.)

Select 2 answers
A.Create a separate column for the month and year.
B.Convert the string to a date data type.
C.Encode the date as a numeric timestamp.
D.Normalize the date by subtracting the mean date.
E.Replace missing dates with the average date.
AnswersA, B

Creating separate columns for month and year facilitates grouping and aggregation for monthly trends. This derived column allows the analyst to easily group sales by month across years or by year-month combinations. It is a common data preparation step for time-series reporting when the tool does not support date functions directly.

Why this answer

To perform time-series analysis on a date string, the analyst must first convert it to a proper date data type so that date functions can be applied. Additionally, creating separate month and year columns simplifies grouping and aggregation for monthly trends. These two steps ensure the data is in a usable format for temporal analysis.

Exam trap

The trap here is assuming that dates can be normalized or averaged like numerical data, which is not appropriate for temporal analysis.

625
MCQhard

An analyst is reviewing a table that stores customer orders. The table contains columns: OrderID, CustomerName, Product1, Product1Qty, Product2, Product2Qty. This design violates which normal form?

A.No violation
B.Third normal form (3NF)
C.Second normal form (2NF)
D.First normal form (1NF)
AnswerD

Repeating Product1/Product2 column groups makes the table non-atomic, violating 1NF's requirement that each column hold a single value and no repeating groups exist. Splitting into an OrderItems table (OrderID, Product, Quantity) satisfies 1NF before addressing higher normal forms.

Why this answer

The table violates First Normal Form (1NF) because it contains repeating groups (Product1, Product1Qty, Product2, Product2Qty) instead of storing each product in a separate row. 1NF requires that each column contains atomic values and that there are no repeating groups or arrays. The presence of multiple product columns for a single order breaks this atomicity and normalization rule.

Exam trap

The trap here is that candidates often think the table is already in 1NF because it has a primary key (OrderID), but they overlook the repeating group columns that violate the atomicity requirement of 1NF.

How to eliminate wrong answers

Option A is wrong because the table clearly violates normalization rules due to repeating groups, so a violation exists. Option B is wrong because Third Normal Form (3NF) requires that the table already be in 2NF and have no transitive dependencies; the immediate violation is at the 1NF level, not 3NF. Option C is wrong because Second Normal Form (2NF) requires that the table first satisfy 1NF and then have no partial dependencies; since the table fails 1NF, it cannot be evaluated for 2NF.

626
MCQmedium

A simple linear regression model predicts sales (y) from advertising spend (x). The equation is y = 2.5x + 10, and R² = 0.81. Which interpretation is correct?

A.The correlation between sales and advertising is 0.81.
B.When advertising is $0, sales are $2.5.
C.81% of the variation in sales is explained by advertising spend.
D.For every $1 increase in advertising, sales increase by $10 on average.
AnswerC

R-squared is the coefficient of determination: 0.81 means 81% of the variance in the dependent variable (sales) is accounted for by the independent variable (advertising spend). The remaining 19% is unexplained by the model, so this interpretation matches the statistic precisely.

Why this answer

Slope indicates that each unit increase in x increases y by 2.5 units. R² of 0.81 means 81% of variance in y is explained by x.

627
MCQeasy

A data analyst is reviewing a dataset and notices that the 'ProductCategory' column contains inconsistent entries such as 'Electronics', 'electronics', and 'ELECTRONICS'. Which data cleaning technique should the analyst apply to standardize these values?

A.Standardization
B.Normalization
C.Case normalization
D.One-hot encoding
AnswerC

Case normalization involves converting all text to a consistent case, such as lowercase or uppercase. This resolves inconsistencies like 'Electronics', 'electronics', and 'ELECTRONICS' by making them identical. It is a simple and effective data cleaning step for standardizing categorical text values. After case normalization, the entries would all become 'electronics', for example.

Why this answer

Case normalization is the appropriate technique to standardize text values that differ only by case. By converting all entries to a consistent case, the analyst ensures that 'Electronics', 'electronics', and 'ELECTRONICS' are treated as the same category. Other techniques like normalization, standardization, and one-hot encoding are not designed to fix case inconsistencies in text data.

Exam trap

The trap here is confusing data normalization or standardization with text case normalization, applying numerical scaling methods to a categorical text issue.

628
MCQmedium

In a time series analysis, a retail analyst observes consistent peaks in sales every December and troughs every February. This pattern repeats annually. Which component of time series does this represent?

A.Irregular
B.Seasonality
C.Trend
D.Cyclical
AnswerB

Seasonality captures repeating fluctuations within a fixed period shorter than a year, so the annual December peaks and February troughs satisfy that constraint precisely. Unlike trend, which shows long-term direction, or cyclical variation, which lacks a fixed period, seasonality is defined by calendar-linked regularity.

Why this answer

Seasonality refers to regular patterns that repeat over fixed periods, such as months or quarters.

629
MCQhard

Refer to the exhibit. A data analyst runs this query to identify high-value customers. However, the result does not include customers with exactly 5 orders. Which data concept does the HAVING clause illustrate?

A.Data sorting with ORDER BY
B.Data joining with INNER JOIN
C.Data aggregation with filtering on aggregated values
D.Data filtering on row-level conditions
AnswerC

The HAVING clause filters groups produced by GROUP BY, applying a predicate to the aggregated COUNT rather than to individual rows, so customers with exactly five orders are excluded when the condition is a strict inequality.

Why this answer

The HAVING clause filters groups after aggregation, so it operates on aggregated values like COUNT(*), SUM(), or AVG(). In the exhibit, the query likely uses HAVING COUNT(order_id) > 5, which excludes customers with exactly 5 orders because the condition is strictly greater than 5. This illustrates aggregation with post-aggregation filtering, distinct from WHERE, which filters rows before grouping.

Exam trap

DA0-002 often tests the WHERE vs HAVING distinction, and candidates frequently miss strict inequality boundaries (e.g., > 5 excludes exactly 5) or mistakenly think HAVING filters rows rather than groups.

How to eliminate wrong answers

Option A is wrong because ORDER BY only sorts the result set; it does not filter groups or affect which rows are returned. Option B is wrong because INNER JOIN combines rows from two tables based on a join condition; it has nothing to do with filtering aggregated groups. Option D is wrong because row-level filtering is the job of the WHERE clause, which executes before GROUP BY, not HAVING, which executes after aggregation.

630
Multi-Selecteasy

A data analyst is preparing a presentation to share findings with non-technical stakeholders. Which TWO practices should the analyst follow to effectively communicate data insights? (Choose two.)

Select 2 answers
A.Tailor the message to the audience
B.Provide raw data tables
C.Use technical jargon to demonstrate expertise
D.Include all data anomalies and outliers
E.Focus on actionable insights
AnswersA, E

Tailoring the message to the audience satisfies the stem's non-technical stakeholder constraint by translating statistical findings into business-relevant language, avoiding jargon and technical notation. This ensures comprehension and drives decision-making, since the audience lacks the analytical background assumed in specialist reporting.

Why this answer

Option A, 'Tailor the message to the audience,' is correct because effective data communication requires framing findings in terms and context that non-technical stakeholders can readily understand, such as using plain language, relevant business metrics, and analogies instead of assuming technical background. Option E, 'Focus on actionable insights,' is correct because the primary purpose of presenting to stakeholders is to drive decisions, so the analyst should emphasize what the data means and what actions should follow rather than merely describing the data. Option B, 'Provide raw data tables,' is not appropriate because dumping unprocessed tables overwhelms non-technical audiences and obscures the key takeaways.

Option C, 'Use technical jargon to demonstrate expertise,' is wrong because jargon creates barriers and alienates the audience rather than clarifying insights. Option D, 'Include all data anomalies and outliers,' is not suitable because exhaustive detail distracts from the main message; anomalies should only be mentioned if they materially affect the conclusions.

Exam trap

The trap is selecting options that seem thorough (raw data, all outliers) or impressive (technical jargon) when they actually hinder communication; the exam tests understanding that audience-centric, action-oriented communication is key.

631
Multi-Selectmedium

A dataset contains outliers in a feature that will be used for linear regression. Which two outlier treatment methods are appropriate? (Choose TWO)

Select 2 answers
A.Cap the outliers at a percentile (e.g., 99th percentile)
B.Use min-max normalization
C.Increase the sample size
D.Remove the outlier rows
E.Replace outliers with the mean
AnswersA, D

Capping at the 99th percentile winsorises extreme values to a boundary, retaining the observation while limiting its leverage on the regression coefficients. This preserves sample size and satisfies the need to reduce outlier influence without discarding data.

Why this answer

Option A (Cap the outliers at a percentile, e.g., 99th percentile) is correct because winsorizing/capping replaces extreme values with a boundary value at a chosen percentile, which preserves the data points while reducing the leverage that extreme values exert on the least-squares fit, thereby limiting their disproportionate influence on the regression coefficients. Option D (Remove the outlier rows) is correct because deleting the offending observations eliminates their influence entirely, which is a standard and valid treatment when outliers are identified as erroneous or unrepresentative and the remaining sample is still adequate for fitting the model. Option B (Use min-max normalization) is not appropriate because rescaling features to a [0,1] range is a transformation that preserves the relative positions and leverage of outliers, so it does nothing to mitigate their effect on the regression.

Option C (Increase the sample size) does not treat the outliers themselves; adding more data does not reduce the influence of existing extreme points and may even introduce more outliers. Option E (Replace outliers with the mean) is generally inappropriate because substituting the mean distorts the distribution, artificially reduces variance, and biases the regression estimates rather than properly handling the extreme values.

Exam trap

DA0-002 often tests the confusion that normalization or mean-imputation 'handles' outliers, when in fact only capping, trimming, or robust methods (e.g., Huber regression) actually mitigate their influence.

632
MCQmedium

Refer to the exhibit. If the date column is stored as a string in 'MM/DD/YYYY' format, what will be the result?

A.Incorrect results because string comparison is lexicographic.
B.NULL values
C.Error because DATE type is expected.
D.Correct results because string comparison works for dates.
AnswerA

The different format causes lexicographic comparison to fail.

Why this answer

When dates are stored as strings in 'MM/DD/YYYY' format, string comparison is lexicographic (character-by-character). This means that '01/02/2023' (January 2) would be considered greater than '12/31/2022' because '0' > '1' at the first character, leading to incorrect chronological ordering. The comparison does not interpret the string as a date value.

Exam trap

CompTIA often tests the misconception that string comparison of dates in 'MM/DD/YYYY' format will yield correct chronological order, but the trap is that lexicographic comparison compares month first, not year, leading to incorrect results.

How to eliminate wrong answers

Option B is wrong because string comparison does not produce NULL values; it simply compares strings lexicographically and returns a valid boolean result. Option C is wrong because no error occurs; the database or application will perform string comparison without expecting a DATE type, as the column is defined as a string. Option D is wrong because string comparison does not work correctly for dates in this format; lexicographic order does not match chronological order for 'MM/DD/YYYY' strings.

633
MCQhard

A marketing analyst wants to segment customers based on purchasing behavior and demographics. The dataset includes continuous variables (spending amount, frequency) and categorical variables (region, gender). The analyst decides to use k-means clustering. What should the analyst do to prepare the data?

A.Use raw data because k-means works with mixed types
B.Standardize continuous variables and one-hot encode categorical variables
C.Apply PCA first to reduce dimensionality
D.Remove categorical variables entirely
AnswerB

K-means computes Euclidean distance, so unstandardised continuous variables with larger ranges dominate. Standardising spending and frequency equalises their influence, while one-hot encoding converts region and gender into numeric binary columns, since k-means cannot process categorical labels directly.

Why this answer

K-means clustering relies on Euclidean distance, which is sensitive to the scale of features. Standardizing continuous variables (e.g., spending amount, frequency) ensures they contribute equally to distance calculations, while one-hot encoding categorical variables (e.g., region, gender) converts them into numerical form without implying ordinal relationships, allowing k-means to process mixed data types correctly.

Exam trap

The trap here is that candidates assume k-means can natively handle mixed data types because it is a common clustering algorithm, but it strictly requires numerical input and scale normalization to avoid skewed distance calculations.

How to eliminate wrong answers

Option A is wrong because k-means cannot directly handle categorical variables; it requires numerical input and assumes continuous features, so using raw mixed-type data would produce meaningless distance calculations. Option C is wrong because PCA is a dimensionality reduction technique applied after preprocessing, not a substitute for standardizing and encoding; it may be used optionally but is not the required preparation step. Option D is wrong because removing categorical variables discards valuable demographic information that could improve segmentation, and k-means can incorporate them after proper encoding.

634
MCQeasy

A data analyst wants to compare the sales performance of four different stores over the same time period. Which chart type is most suitable?

A.Line chart with multiple lines
B.Grouped bar chart
C.Stacked bar chart
D.Pie chart with multiple pies
AnswerB

A grouped bar chart places the four stores side by side within each time period on a shared categorical axis, enabling direct comparison of their sales values. This satisfies the requirement to compare multiple discrete entities across the same period, which a single-series chart or pie chart cannot do.

Why this answer

A grouped bar chart is the most suitable choice because it allows direct comparison of discrete categories (four stores) across a common time period, with each group representing a time interval and individual bars representing each store's sales. This chart type excels at side-by-side comparisons of multiple entities over the same categorical axis, making it ideal for the analyst's goal.

Exam trap

The trap here is that candidates often choose a line chart (Option A) because they associate 'time period' with trend analysis, but the question explicitly asks to compare sales performance over the same time period, not to show changes over time, making the grouped bar chart the correct choice for direct comparison of discrete entities.

How to eliminate wrong answers

Option A is wrong because a line chart with multiple lines is designed to show trends over continuous time, but here the time period is fixed and the focus is on comparing discrete stores, not tracking changes over time. Option C is wrong because a stacked bar chart shows part-to-whole relationships within each time period, which obscures individual store performance and makes direct comparisons between stores difficult. Option D is wrong because a pie chart with multiple pies is meant to show proportions of a whole for each store separately, but comparing sales across stores requires a common baseline, which pie charts do not provide.

635
MCQeasy

A data analyst wants to predict customer churn based on categorical features like region and plan type, and continuous features like usage and tenure. Which regression type should be used?

A.Logistic regression
B.Ridge regression
C.Linear regression
D.Lasso regression
AnswerA

Logistic regression predicts a binary outcome, churn versus no churn, satisfying the categorical target requirement. It handles mixed predictors: region and plan type via dummy encoding, usage and tenure as continuous terms. Linear regression would model a continuous response, which churn is not, so logistic regression fits this classification scenario.

Why this answer

Logistic regression is the correct choice because the target variable, customer churn, is binary (churn vs. no churn). Logistic regression models the probability of a binary outcome using a sigmoid function, making it suitable for classification tasks with both categorical and continuous predictors.

Exam trap

CompTIA often tests the misconception that 'regression' in the option name implies it is only for continuous outcomes, leading candidates to overlook logistic regression as a valid classification technique.

How to eliminate wrong answers

Option B (Ridge regression) is wrong because it is a regularized form of linear regression used for continuous outcomes, not binary classification. Option C (Linear regression) is wrong because it predicts a continuous value and is inappropriate for a binary dependent variable; it can produce probabilities outside [0,1] and violates the assumption of normally distributed errors. Option D (Lasso regression) is wrong because, like Ridge, it is a regularized linear regression for continuous targets and performs feature selection via L1 penalty, but it does not handle binary classification.

636
MCQeasy

During data exploration, an analyst notices that the target variable has a heavily right-skewed distribution. Which data transformation would be most appropriate to make the distribution more symmetric?

A.Log transformation
B.Reciprocal transformation
C.No transformation needed
D.Square root transformation
AnswerA

Log transformation compresses the long right tail by rescaling values on a multiplicative scale, pulling extreme high observations closer to the bulk. This reduces positive skewness, making the target distribution more symmetric and better suited to techniques assuming normality.

Why this answer

Log transformation is appropriate for heavily right-skewed distributions because it compresses the high values and spreads out the low values, making the distribution more symmetric. Square root transformation is better for moderate skew, and reciprocal transformation is for severe skew. Therefore, option A (Log transformation) is correct.

637
MCQmedium

A company’s marketing team wants to segment customers based on purchase history, demographics, and website behavior. The data includes both numeric and categorical variables. Which clustering algorithm is best suited for handling mixed data types?

A.Hierarchical clustering with Gower distance
B.K-modes clustering
C.DBSCAN with Euclidean distance
D.K-means clustering
AnswerA

Hierarchical clustering with Gower distance computes pairwise dissimilarity across numeric and categorical attributes simultaneously, so mixed-type records can be segmented without arbitrary encoding. It satisfies the stem's mixed data constraint, unlike k-means, which relies on Euclidean distance and requires numeric, scaled input.

Why this answer

Hierarchical clustering with Gower distance is best suited for mixed data types because Gower distance computes a dissimilarity measure that handles both numeric and categorical variables by normalizing numeric differences and using a simple matching coefficient for categorical ones. This allows the algorithm to create a distance matrix that equally weights all variable types, making it ideal for segmenting customers with purchase history, demographics, and website behavior data.

Exam trap

The trap here is that candidates often assume K-means or DBSCAN can handle mixed data by simply encoding categorical variables, but they overlook that Euclidean distance on encoded data distorts the geometry and fails to preserve the natural dissimilarity structure of categorical variables.

How to eliminate wrong answers

Option B (K-modes clustering) is wrong because it is designed exclusively for categorical data and cannot handle numeric variables like purchase history or website behavior metrics. Option C (DBSCAN with Euclidean distance) is wrong because Euclidean distance is only meaningful for numeric data and cannot properly measure dissimilarity between categorical variables, leading to distorted clusters. Option D (K-means clustering) is wrong because it relies on Euclidean distance and assumes numeric, continuous data; it cannot directly incorporate categorical variables without encoding, and even with encoding, it is sensitive to scaling and does not naturally handle mixed types.

638
Multi-Selecthard

A data analyst is preparing a dataset for predictive modeling and must handle missing values in several numeric and categorical columns. The team needs defensible, documented choices rather than ad hoc deletion. Which TWO actions are appropriate for handling missing data in this scenario? (Choose two.)

Select 2 answers
A.Create an explicit missing-value indicator column alongside an imputed value so the model can learn from the missingness pattern.
B.Impute numeric missing values with the column mean and categorical missing values with the mode, without further review.
C.Document the missingness rate per column and investigate whether values are missing at random before choosing a treatment.
D.Replace all missing numeric values with zero so the column contains no nulls and requires no further processing.
E.Drop every row that contains any missing value across all columns to guarantee a complete dataset.
AnswersA, C

An indicator preserves the information that a value was absent, which is valuable when missingness correlates with the target. Pairing it with an imputed value keeps the record usable while letting the model distinguish imputed from observed cases. This is a documented, reproducible technique that satisfies the demand for defensible handling.

Why this answer

Defensible missing-data handling starts with diagnosing the extent and pattern of missingness, then applies a treatment matched to that pattern. Combining an indicator column with an imputed value preserves both usability and the signal contained in absence. Blind mean or mode substitution, blanket row deletion, and zero-filling all introduce bias without documentation.

Exam trap

The trap here is treating missing-data handling as a mechanical fill step and overlooking that the pattern of missingness itself can be informative and must be diagnosed first.

639
MCQeasy

A data analyst needs to merge two customer tables from different sources. One table uses 'CUST_ID' as the primary key, the other uses 'CustomerID'. To ensure accurate merging, the analyst should first:

A.Perform a fuzzy match on names
B.Normalize the key column names to a common format
C.Remove duplicate rows from both tables
D.Aggregate data by region
AnswerB

Mismatched key names such as 'CUST_ID' and 'CustomerID' cause the join to fail or produce a cartesian result. Standardising both columns to one common name and format lets the merge match rows correctly, satisfying the requirement for accurate joining across sources.

Why this answer

Normalizing key column names to a common format (Option B) is the correct first step because the merge operation requires a consistent join key. Without aligning 'CUST_ID' and 'CustomerID' to a single name and data type, the database or ETL tool will treat them as different columns, resulting in a cross join or an error. This step ensures referential integrity and enables an accurate inner or outer join based on the primary key.

Exam trap

The trap here is that candidates assume deduplication (Option C) is the most critical first step, but without first standardizing the join keys, any deduplication logic would operate on mismatched or incomplete data, leading to incorrect results.

How to eliminate wrong answers

Option A is wrong because performing a fuzzy match on names is an advanced, resource-intensive technique used only when exact key values are unavailable or inconsistent; it is unnecessary when the tables already have primary key columns that can be standardized. Option C is wrong because removing duplicate rows before aligning key names could inadvertently delete legitimate records that only appear duplicated due to key naming differences, and deduplication should occur after the merge or as a separate quality step. Option D is wrong because aggregating data by region is a post-merge analytical operation that has no bearing on resolving key column mismatches and would corrupt the granularity needed for accurate joining.

640
Multi-Selectmedium

A data analyst is creating a dashboard to monitor website traffic. The dashboard will include metrics such as page views, unique visitors, and average session duration. The analyst wants to ensure the dashboard effectively communicates performance to the marketing team. Which TWO design elements should the analyst prioritize? (Choose two.)

Select 2 answers
A.Use 3D charts to make the data more visually appealing.
B.Include as many metrics as possible to provide a comprehensive view.
C.Place the most important metrics in the top-left area of the dashboard.
D.Include a detailed data table beneath each chart for reference.
E.Use a consistent color scheme to indicate positive and negative trends.
AnswersC, E

In Western cultures, people tend to scan from top-left to bottom-right, so placing critical metrics in the top-left ensures they are seen first. This aligns with visual hierarchy principles, helping the marketing team quickly identify the most important information. It is a simple yet effective way to guide attention and improve dashboard usability.

Why this answer

Prioritizing a consistent color scheme for trends and placing key metrics in the top-left area enhances quick comprehension and guides the viewer's attention. These elements align with best practices for dashboard design, ensuring the marketing team can efficiently monitor website traffic performance.

Exam trap

The trap here is assuming that more data or decorative elements like 3D charts improve a dashboard, when in fact they often hinder clarity and speed of insight.

641
MCQeasy

An analyst creates a dashboard with multiple visualizations. Which feature allows users to change the data displayed across all charts simultaneously?

A.Linked chart
B.Drill-down
C.Filter or slicer
D.Data segmentation
AnswerC

A filter or slicer applies a single selection across every visual on the dashboard, so changing it updates all charts at once. This directly satisfies the stem's requirement for simultaneous, dashboard-wide data changes, unlike per-visual filters that affect only one chart.

Why this answer

A filter or slicer is an interactive control that, when changed, applies the same selection criterion to every visualization on the dashboard, enabling simultaneous cross-chart updates. This is the standard mechanism in BI tools (Tableau, Power BI, Looker) for synchronized data selection.

Exam trap

The trap is confusing 'linked chart' (a specific cross-highlighting interaction) with the general filter/slicer control that drives all visuals; the exam tests precise terminology.

How to eliminate wrong answers

Option A is wrong because a linked chart highlights or filters based on selection in one chart, but it is not the general control that changes data across all charts; it is a specific interaction pattern. Option B is wrong because drill-down navigates from a summary to more detailed data within a single visualization, not across all charts. Option D is wrong because 'data segmentation' is a general analysis concept (grouping data), not a dashboard interactivity feature that updates all charts.

642
MCQeasy

A data analyst is working with a dataset that contains a column for 'customer satisfaction rating' on a scale from 1 to 5, where 1 is very dissatisfied and 5 is very satisfied. The analyst wants to summarize the central tendency of this data. Which measure of central tendency is most appropriate for this ordinal data?

A.Range
B.Mode
C.Mean
D.Median
AnswerD

The median is the middle value when data is ordered, making it appropriate for ordinal data because it does not assume equal intervals. For satisfaction ratings, the median represents the central rating without implying that the difference between ratings is consistent. It is robust to outliers and skewed distributions, providing a more accurate summary of the typical response.

Why this answer

For ordinal data such as satisfaction ratings, the median is the most appropriate measure of central tendency because it does not assume equal intervals between categories. It identifies the middle value when the data is ordered, providing a robust summary that is not influenced by extreme ratings. The mean assumes equal intervals, and the mode may not represent the center, while the range measures spread.

Exam trap

The trap here is assuming that the mean is always the best measure of central tendency, but for ordinal data, the median is more appropriate due to unequal intervals.

643
MCQmedium

A sales manager receives a daily report at 8 AM via email showing yesterday's sales by region. This is an example of which report type?

A.Ad hoc report
B.Scheduled report
C.Self-service report
D.Operational report
AnswerB

A scheduled report runs automatically at a defined time and delivers results, here daily at 8 AM by email. This matches the recurring, time-triggered delivery described, distinguishing it from ad hoc, real-time, or dashboard report types.

Why this answer

A scheduled report is generated automatically at a predefined time and delivered to recipients without manual intervention. The 8 AM daily email with yesterday's sales by region matches this pattern exactly — a recurring, time-triggered distribution. This is the defining characteristic of scheduled reporting in BI tools like Power BI, Tableau, or SSRS subscriptions.

Exam trap

The trap here is confusing 'automated delivery' with 'operational' — candidates see 'daily sales' and pick operational report, but the defining trait is the scheduled trigger, not the content domain.

How to eliminate wrong answers

Option A is wrong because ad hoc reports are generated on demand in response to a specific, one-off user request, not delivered automatically on a fixed schedule. Option C is wrong because self-service reporting refers to business users building their own reports via tools, not automated delivery of a fixed report. Option D is wrong because operational reports support day-to-day transactional monitoring (e.g., real-time inventory), not a fixed daily summary delivered by email.

644
MCQeasy

A data analyst at a university is asked to classify the variable 'Student Classification' with possible values Freshman, Sophomore, Junior, Senior. Which measurement scale best describes this variable?

A.Ratio
B.Nominal
C.Ordinal
D.Interval
AnswerC

Ordinal scale applies to categorical data with a meaningful order but no defined difference between ranks. The classifications Freshman, Sophomore, Junior, Senior have a clear sequence, but the gap between each is not quantified. Thus, ordinal is the correct measurement scale for this variable.

Why this answer

The correct answer is Ordinal because the classifications have a natural order (Freshman, Sophomore, Junior, Senior) but the differences between them are not equal or measurable. Nominal would ignore the order, while interval and ratio require quantitative properties that are absent here.

Exam trap

The trap here is assuming that because the categories have a clear order, they must be interval or ratio; however, without equal intervals or a true zero, ordinal is the correct scale.

645
MCQhard

A data analyst is creating a report to compare the performance of sales regions across multiple years. The report will be used by regional managers to identify trends. Which visualization approach best supports this?

A.A single line chart with all regions overlaid
B.A bar chart with years on x-axis and regions as grouped bars
C.A stacked area chart with all regions
D.Small multiples of line charts, one per region
AnswerD

Small multiples of line charts, one per region, satisfy the multi-year trend comparison constraint: each region gets an identical axis scale, so managers compare slopes directly without overlapping lines obscuring patterns. Line charts encode temporal continuity, and faceting by region removes the clutter that a single multi-series chart would create across many regions.

Why this answer

Small multiples of line charts, one per region, allow regional managers to easily compare trends across years without visual clutter. Each region gets its own chart with consistent axes, making patterns and outliers clear. This approach supports the goal of identifying trends per region while enabling cross-region comparison.

Exam trap

DA0-002 often tests the confusion between charts that show trends (line charts) and those that compare categories (bar charts), and candidates may overlook that small multiples are a form of line chart that reduces clutter.

How to eliminate wrong answers

Option A is wrong because overlaying all regions on a single line chart can create a spaghetti plot, making it hard to distinguish individual regional trends, especially with many regions. Option B is wrong because a grouped bar chart with years on x-axis and regions as grouped bars is better for comparing discrete values, not for showing trends over time; it emphasizes comparison at each year rather than the trajectory. Option C is wrong because a stacked area chart shows cumulative totals and part-to-whole relationships, which obscures individual regional trends and makes it difficult to compare regions directly.

646
MCQhard

A data analyst is working with a dataset containing house prices. After building a multiple linear regression model, the analyst observes that the model performs well on training data but poorly on validation data. Which technique is most appropriate to address this issue?

A.Decrease the training data size
B.Use a polynomial transformation
C.Increase the number of features
D.Apply L2 regularization (Ridge)
AnswerD

L2 regularization (Ridge) adds a penalty proportional to the squared magnitude of coefficients, shrinking them towards zero and reducing variance without eliminating features. This directly counters the overfitting causing the train-validation performance gap, satisfying the scenario's need to improve generalisation on unseen house price data.

Why this answer

The model is overfitting the training data, as evidenced by high performance on training data but poor performance on validation data. L2 regularization (Ridge) adds a penalty term proportional to the square of the coefficients, which shrinks them and reduces model complexity, thereby improving generalization to unseen data.

Exam trap

CompTIA often tests the distinction between overfitting and underfitting, and candidates mistakenly choose polynomial transformation or adding features thinking they will improve fit, when in fact they increase model complexity and worsen overfitting.

How to eliminate wrong answers

Option A is wrong because decreasing the training data size would exacerbate overfitting by providing the model with even less information to learn generalizable patterns. Option B is wrong because polynomial transformation increases model complexity and feature interactions, which typically worsens overfitting rather than addressing it. Option C is wrong because increasing the number of features adds more predictors, which increases the risk of overfitting and does not directly penalize large coefficients.

647
MCQmedium

A data analyst is working with a dataset containing a categorical variable 'Education Level' with categories: High School, Bachelor's, Master's, PhD. The analyst wants to include this variable in a regression model. Which encoding technique should the analyst use?

A.Binary encoding
B.Label encoding
C.Ordinal encoding
D.One-hot encoding
AnswerD

One-hot encoding creates binary columns for each category, avoiding implying any ordinal relationship. For regression, it is appropriate for nominal categorical variables like Education Level, which has no inherent order. It allows the model to estimate a separate coefficient for each category, capturing differences without assuming a linear progression.

Why this answer

One-hot encoding is the preferred method for nominal categorical variables in regression because it avoids imposing an artificial order or equal spacing. It creates separate binary indicators, allowing the model to estimate distinct effects for each education level. Label, binary, and ordinal encoding all introduce numerical relationships that may misrepresent the categorical nature of the data.

Exam trap

The trap here is using label encoding for a nominal variable, which forces the model to treat categories as ordered numbers and can lead to incorrect coefficient estimates.

648
MCQeasy

A data analyst is tasked with gathering data from a legacy system that only exports CSV files. The files contain headers but no data types. Which tool would best facilitate initial data exploration?

A.Hadoop
B.Tableau
C.SQL database
D.Python pandas
AnswerD

Python pandas reads CSV files directly with `read_csv`, inferring column data types automatically despite the header-only source. This satisfies the stem's constraint of untyped legacy exports, enabling immediate exploration through `head()`, `info()` and `describe()` without prior schema definition or manual type assignment.

Why this answer

Python pandas is the best tool for initial data exploration of CSV files because its read_csv() function automatically infers data types, handles headers, and provides immediate exploratory methods like .info(), .describe(), and .head(). It requires no schema definition upfront, making it ideal for legacy exports with unknown types.

Exam trap

The trap is choosing a tool that requires predefined schemas (SQL) or is meant for downstream visualization (Tableau) instead of recognizing pandas as the flexible, schema-inferring exploration tool for raw CSVs.

How to eliminate wrong answers

Option A (Hadoop) is wrong because Hadoop is a distributed storage and processing framework designed for large-scale batch workloads, not quick interactive exploration of a single CSV. Option B (Tableau) is wrong because it is a visualization tool that connects to data sources but is not optimized for programmatic type inference and statistical profiling of raw CSVs. Option C (SQL database) is wrong because loading a CSV into SQL requires defining a schema and data types first, which contradicts the scenario where types are unknown.

649
Multi-Selectmedium

A data team must implement a data retention policy to reduce storage costs while meeting legal requirements. Which TWO actions best achieve this?

Select 2 answers
A.Set data retention limits with automated deletion
B.Use data compression
C.Increase primary storage capacity
D.Implement data deduplication
E.Archive historical data to tape or cloud archive
AnswersA, E

Automated deletion enforces retention limits without manual intervention, directly satisfying the legal requirement to purge data once its mandated retention period expires. By removing data systematically at the defined threshold, it also curtails ongoing storage consumption, which is the cost-reduction constraint the stem specifies.

Why this answer

Option A is correct because setting data retention limits with automated deletion enforces the legal retention policy by removing data once its required retention period expires, directly reducing stored volume and storage costs. Option E is correct because archiving historical data to tape or cloud archive tiers moves infrequently accessed data to much cheaper storage media while still preserving it for legal and compliance requirements. Together, these two actions address both cost reduction and legal retention.

Option B (data compression) reduces the size of stored data but does not enforce retention or remove data that has exceeded its legal retention period. Option C (increasing primary storage capacity) raises costs rather than reducing them and does nothing to meet retention requirements. Option D (data deduplication) eliminates redundant copies but does not implement a retention policy or remove expired data.

Exam trap

DA0-002 often tests the confusion between storage optimization techniques (compression, deduplication) and retention policy enforcement (automated deletion, archiving), leading candidates to pick efficiency features instead of compliance-driven actions.

650
MCQmedium

During EDA, an analyst calculates the Z-score for each data point in a dataset. A data point with a Z-score of 3.5 is identified. What does this indicate?

A.The data point has a high frequency
B.The data point is exactly at the mean
C.The data point is likely an outlier
D.The data point is within the interquartile range
AnswerC

A Z-score of 3.5 lies beyond three standard deviations from the mean, satisfying the stem's outlier criterion. Under a normal distribution roughly 99.7% of values fall within three standard deviations, so such an extreme standardised deviation is statistically improbable and warrants flagging as a likely outlier during EDA.

Why this answer

A Z-score of 3.5 means the data point lies 3.5 standard deviations above the mean. In most distributions, values beyond ±3 standard deviations are statistically rare (about 0.3% of a normal distribution) and are commonly flagged as outliers. This is the standard EDA heuristic for outlier detection.

Exam trap

DA0-002 often tests whether candidates confuse Z-score (standard deviations from mean) with IQR-based outlier detection or with frequency counts, luring them to pick 'high frequency' or 'within IQR' answers.

How to eliminate wrong answers

Option A is wrong because Z-score measures distance from the mean in standard deviation units, not frequency or count of occurrences. Option B is wrong because a Z-score of 0 indicates the point equals the mean; 3.5 is far from it. Option D is wrong because the interquartile range (IQR) is a separate outlier method (typically 1.5×IQR rule); Z-score does not describe IQR membership.

651
MCQhard

A data analyst is reviewing a data lake access policy that controls access to a data lake. The analyst wants to communicate that the current policy restricts data downloads to a specific IP range. Which of the following best describes the policy's effect?

A.The policy allows all users to download data.
B.The policy denies data downloads to all users.
C.The policy restricts data downloads to a specific IP range.
D.The policy allows data uploads from a specific IP range.
AnswerC

The policy's stated effect is exactly that downloads are limited to a defined IP range, so the description matches the access control as written. This satisfies the analyst's need to communicate the restriction accurately, confirming downloads outside that range are blocked.

Why this answer

The access policy explicitly restricts data downloads to a specific IP range by using a condition in the policy statement, such as a source IP address with a CIDR block. This means only requests originating from IP addresses within that range are allowed to perform the download action, effectively limiting downloads to authorized network locations.

Exam trap

The trap here is that candidates may confuse the policy's effect with a blanket denial (Option B) or misinterpret the restriction as applying to uploads (Option D), when the policy specifically targets downloads and allows them only from a defined IP range.

How to eliminate wrong answers

Option A is wrong because it describes an allow-all effect, which contradicts the explicit IP-based restriction in the policy; a policy that restricts to a specific IP range does not allow all users. Option B is wrong because it states the policy denies downloads to all users, but the policy actually allows downloads from the specified IP range while implicitly or explicitly denying others—it does not universally deny. Option D is wrong because it refers to data uploads (e.g., `s3:PutObject`), but the question specifically addresses downloads (`s3:GetObject`), and the policy effect is about restricting downloads, not uploads.

652
Drag & Dropmedium

Drag and drop the steps to perform a data audit in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Data audit begins with inventory, quality assessment, compliance check, documentation, and recommendations.

653
MCQeasy

A data analyst at a healthcare clinic is organizing patient records. The analyst needs to categorize each patient's blood pressure reading as 'Low', 'Normal', 'Elevated', or 'High' based on clinical thresholds. The categories have a clear order from lowest to highest risk. Which measurement scale best describes this classification?

A.Interval
B.Nominal
C.Ordinal
D.Ratio
AnswerC

Ordinal data has categories with a meaningful order but unequal or undefined intervals between them. The blood pressure classifications progress from low to high risk in a defined sequence, yet the difference between 'Normal' and 'Elevated' is not a fixed numeric quantity. This ordered categorization fits the ordinal scale precisely.

Why this answer

The blood pressure categories have a clear rank order from lowest to highest risk but lack equal intervals or a true zero, which is the defining characteristic of ordinal data. Nominal data would ignore the ranking, while interval and ratio scales require numeric properties that these qualitative labels do not possess.

Exam trap

The trap here is assuming that any set of categories is nominal, overlooking the meaningful order that elevates this classification to ordinal.

654
Multi-Selectmedium

Which TWO of the following are leading indicators that can help predict future performance?

Select 2 answers
A.Net profit margin
B.Monthly revenue
C.Website traffic
D.Number of qualified leads
E.Customer churn rate
AnswersC, D

Website traffic measures current visitor activity, which precedes conversion and revenue, so rising traffic signals future sales potential. It is a leading indicator because it changes before the outcome it predicts, unlike lagging metrics such as revenue.

Why this answer

Website traffic (C) is a leading indicator because it measures top-of-funnel visitor activity that precedes and predicts future conversions, revenue, and customer acquisition, rather than reporting results that have already occurred. Number of qualified leads (D) is also a leading indicator because qualified leads represent prospects who have shown buying intent and are likely to convert into customers in future periods, directly forecasting upcoming sales performance. By contrast, net profit margin (A), monthly revenue (B), and customer churn rate (E) are lagging indicators: they report outcomes that have already happened (profitability, realized sales, and lost customers) and therefore confirm past performance rather than predict future performance.

Exam trap

DA0-002 often tests the confusion between leading and lagging indicators, tricking candidates into selecting financial outcomes (revenue, profit) as leading when they are lagging.

655
MCQmedium

A report uses red and green colors to indicate performance; a stakeholder who is colorblind complains. What is the BEST remedy?

A.Remove color entirely
B.Change colors to blue and orange
C.Add patterns or labels in addition to color
D.Use a different chart type
AnswerC

Red and green differ only along the hue axis, which deuteranopia and protanopia collapse into similar tones. Adding patterns or text labels encodes the same performance status through a second, non-colour channel, preserving the report's meaning for colourblind stakeholders without removing the original palette.

Why this answer

Adding patterns or labels alongside color ensures that the data encoding is accessible to colorblind viewers without removing the informational value of color for those who can see it. This approach follows WCAG (Web Content Accessibility Guidelines) 2.1 Success Criterion 1.4.1, which requires that color is not the sole means of conveying information. In data visualization tools like Tableau or Power BI, this is implemented by adding shape markers or text annotations to series, preserving the original color scheme for non-colorblind users while providing redundant encoding.

Exam trap

The trap here is that candidates assume the best remedy is to simply change colors to a more colorblind-friendly palette (like blue and orange), but CompTIA often tests the deeper principle that color alone is never sufficient for accessibility—redundant encoding via patterns or labels is required to meet WCAG standards.

How to eliminate wrong answers

Option A is wrong because removing color entirely eliminates a key visual encoding channel, which can reduce the chart's effectiveness for the majority of viewers and may violate best practices for data visualization that leverage color for quick pattern recognition. Option B is wrong because while blue and orange are more colorblind-friendly than red and green, they still rely solely on color differentiation, which fails for users with complete color blindness (achromatopsia) or other deficiencies; the WCAG requirement is to not use color as the only means of distinction. Option D is wrong because changing the chart type does not address the core issue of color-dependent encoding; a different chart type (e.g., bar chart instead of pie chart) still uses color to differentiate categories, so the accessibility problem persists unless patterns or labels are added.

656
MCQeasy

A data analyst is designing a report that will be printed in black and white for a monthly management meeting. The report includes a bar chart comparing sales across five regions. To ensure the chart is easily interpretable in grayscale, which design choice should the analyst make?

A.Apply a gradient fill to the bars, with darker shades for higher sales.
B.Use a distinct color for each bar to differentiate regions.
C.Use a single color for all bars and rely on the x-axis labels for region identification.
D.Use a 3D bar chart to make the bars stand out.
AnswerC

Using a single color for all bars ensures that the bars are visually uniform and avoids any confusion from grayscale conversion. The x-axis labels clearly identify each region, so viewers can easily associate bars with regions. This approach is simple, clean, and effective for black and white printing, as it relies on position and labels rather than color.

Why this answer

For a black and white printed report, color cannot be used to differentiate categories. Using a single color for all bars and relying on x-axis labels ensures that each bar is clearly associated with its region without relying on color or shading. This approach is clean, avoids grayscale confusion, and maintains the integrity of the data presentation.

Exam trap

The trap here is assuming that color differentiation is always necessary, but in grayscale printing, color can become indistinguishable, so simpler designs are more effective.

657
MCQhard

A data analyst is evaluating a binary classification model for loan default prediction. The model achieves 98% accuracy on the test set, but the analyst notices that only 2% of loans in the dataset actually defaulted. The analyst is concerned that accuracy is misleading. Which metric should the analyst examine to better assess the model's ability to identify actual defaulters?

A.Mean Absolute Error (MAE)
B.Adjusted R-squared
C.Root Mean Squared Error (RMSE)
D.Recall
AnswerD

Recall, also called sensitivity or true positive rate, measures the proportion of actual defaulters that the model correctly identifies. With only 2% defaults, a model predicting 'no default' for all cases would achieve 98% accuracy but zero recall. Therefore, recall is the appropriate metric to assess the model's ability to catch actual defaulters, directly addressing the analyst's concern.

Why this answer

In imbalanced classification, accuracy can be misleading because a model that always predicts the majority class achieves high accuracy but fails to identify the minority class. Recall focuses on the minority class by measuring how many actual positives were correctly identified. For loan default prediction, missing a defaulter is costly, so recall is a key metric.

The other metrics are for regression or model fit, not classification performance.

Exam trap

The trap here is assuming that high accuracy always indicates a good model, even when the data is imbalanced and the cost of missing the minority class is high.

658
MCQhard

A data analyst has a time series of monthly sales data. They observe that sales are consistently higher every December and lower every January. Which component of time series does this pattern represent?

A.Irregular
B.Cyclical
C.Seasonality
D.Trend
AnswerC

Seasonality describes repeating, calendar-linked fluctuations within a fixed period, such as December peaks and January troughs each year. The consistent annual recurrence in the monthly sales series identifies this component rather than trend or cyclical variation.

Why this answer

A pattern that repeats every December (high) and every January (low) is tied to the calendar year, which is the defining characteristic of seasonality. Seasonal components repeat at fixed, known periods — here, an annual cycle with monthly data.

Exam trap

The trap is picking cyclical because both involve repetition — but cyclical spans irregular multi-year periods while seasonality is locked to a fixed calendar period like December, and the exam expects you to spot the calendar anchor.

How to eliminate wrong answers

Option A is wrong because the irregular component is random, unpredictable noise that cannot be attributed to a repeating calendar pattern. Option B is wrong because cyclical patterns span multiple years with irregular timing tied to economic cycles, not a fixed annual December/January rhythm. Option D is wrong because trend represents the long-term direction of the series (growth or decline over years), not a recurring within-year pattern.

659
MCQmedium

A data analyst is creating a dashboard in Looker Studio for an e-commerce company. They want to display the average order value by product category, and also allow users to filter by date range. Which combination of elements should be used?

A.Dimension: Order Value; Metric: Product Category; Control: Date Range
B.Dimension: Date; Metric: Product Category; Control: None
C.Dimension: Average Order Value; Metric: Product Category; Control: Date Range
D.Dimension: Product Category; Metric: Average Order Value; Control: Date Range
AnswerD

Product Category supplies the grouping dimension, Average Order Value the aggregated metric, and a Date Range control lets users filter the timeframe. This combination satisfies both the category breakdown and the interactive date filtering requirement.

Why this answer

In Looker Studio, the field you group by is the dimension and the field you aggregate is the metric. To show average order value by product category, Product Category must be the dimension and Average Order Value the metric, with a Date Range control added for user filtering. This is the only option that assigns the roles correctly and includes the required control.

Exam trap

DA0-002 often tests the dimension-vs-metric distinction, so the trap is reversing them (putting the measure in the dimension slot) or forgetting that a Date Range control is required for user-driven date filtering.

How to eliminate wrong answers

Option A is wrong because it reverses the roles—Order Value as dimension and Product Category as metric—which would not produce an average per category. Option B is wrong because it uses Date as the dimension and omits the Date Range control, so it cannot show category-level averages or let users filter by date. Option C is wrong because it puts Average Order Value in the dimension slot and Product Category in the metric slot, inverting the aggregation logic.

660
MCQmedium

A data analyst is compiling a monthly sales dashboard. The dataset contains a 'Revenue' column with a highly right-skewed distribution due to a few large enterprise transactions. The analyst wants a measure of central tendency that is robust to these outliers. Which measure should the analyst use?

A.Mode
B.Mean
C.Median
D.Range
AnswerC

The median is the middle value when data is ordered, so it is not affected by extreme values. With right-skewed revenue data, the median remains representative of the typical transaction, making it the robust measure of central tendency. It accurately reflects the center without being distorted by the large enterprise transactions.

Why this answer

The median is the appropriate measure of central tendency when data is skewed or contains outliers because it depends only on the middle position. The mean would be inflated by the large enterprise transactions, while the mode and range do not serve as robust central measures. Thus, the median best represents the typical revenue.

Exam trap

The trap here is assuming the mean is always the best measure of central tendency, ignoring its sensitivity to outliers in skewed distributions.

661
MCQhard

Given the linear regression output, which independent variable has the strongest effect on price, based on standardized coefficients?

A.bathrooms
B.sqft_living
C.Intercept
D.bedrooms
AnswerB

Standardised coefficients express each predictor's effect in standard-deviation units, so they are directly comparable across variables measured on different scales. The variable with the largest absolute standardised coefficient, sqft_living, therefore exerts the strongest effect on price.

Why this answer

Standardized coefficients (beta weights) allow comparison of the relative strength of independent variables by measuring the number of standard deviations the dependent variable changes per one standard deviation change in the predictor. In the regression output, sqft_living has the highest absolute standardized coefficient, indicating it has the strongest effect on price. The intercept is not an independent variable and its coefficient is not standardized for comparison.

Exam trap

The trap here is that candidates mistakenly compare unstandardized coefficients or p-values instead of standardized coefficients, leading them to choose a variable like bathrooms or bedrooms that appears significant but has a weaker standardized effect.

How to eliminate wrong answers

Option A is wrong because bathrooms may have a statistically significant coefficient, but its standardized coefficient is smaller than that of sqft_living, meaning it has a weaker relative effect on price. Option C is wrong because the intercept is a constant term representing the predicted price when all independent variables are zero; it is not an independent variable and its coefficient is not standardized for effect comparison. Option D is wrong because bedrooms, while possibly significant, has a lower absolute standardized coefficient than sqft_living, indicating a weaker influence on price per standard deviation change.

662
Multi-Selectmedium

Which TWO chart types are best suited to show the proportion of total sales contributed by each product category? (Select exactly two.)

Select 2 answers
A.Histogram
B.Scatter plot
C.Stacked bar chart
D.Pie chart
E.Line chart
AnswersC, D

A stacked bar chart encodes each category as a segment within a single bar, so segment heights show each product category's share of total sales. This directly satisfies the proportion-of-total requirement while also permitting comparison across periods.

Why this answer

A stacked bar chart (C) is correct because it displays each product category as a segment of a single bar, so the relative size of every segment directly shows that category's share of total sales. A pie chart (D) is correct because it divides a circle into slices whose angles are proportional to each category's contribution to the whole, which is the classic way to show part-to-whole proportions. A histogram (A) is not appropriate because it bins continuous numeric data to show a frequency distribution, not category shares.

A scatter plot (B) is not appropriate because it shows the relationship or correlation between two numeric variables as points. A line chart (E) is not appropriate because it is designed to show trends or changes over a continuous dimension such as time, not proportions of a total.

Exam trap

The trap is that candidates may pick a histogram or line chart because they see 'sales' and think trend or distribution, missing the 'proportion of total' phrasing that signals part-to-whole.

663
MCQhard

Refer to the exhibit. What is the best course of action for the data analyst?

A.Change the default value to 'NA'
B.Delete rows with missing 'Region'
C.Investigate source data to understand why 'Region' is missing
D.Ignore the warning as it is handled
AnswerC

Missing 'Region' values indicate a data quality defect at source. Investigating the source data identifies whether the gap stems from ingestion, mapping or entry errors, enabling a proper fix rather than masking the problem downstream.

Why this answer

Missing data in a critical field like 'Region' indicates a potential data quality issue that requires root cause analysis. The data analyst should investigate the source system to determine whether the missing values result from a collection error, a transformation bug, or a legitimate null, rather than making assumptions by deleting rows or filling defaults. This aligns with the principle of data provenance and ensures that downstream insights are based on accurate, traceable data.

Exam trap

The trap here is that candidates assume missing data should be immediately cleaned by deletion or default values, but the CompTIA Data+ exam tests the understanding that a data analyst's first responsibility is to investigate the source of the anomaly before applying any transformation.

How to eliminate wrong answers

Option A is wrong because blindly changing the default value to 'NA' introduces artificial data that may skew aggregations or filter logic, and it does not address the underlying reason for the missing values. Option B is wrong because deleting rows with missing 'Region' could discard valid records if the missing data is due to a temporary system glitch or a mapping error, leading to biased analysis and loss of potentially important context. Option D is wrong because ignoring the warning assumes the missing data is harmless, but in practice, missing region information can break geographic segmentation, violate data integrity constraints, and cause incorrect reporting in dashboards or ETL pipelines.

664
MCQmedium

A company is building a data pipeline to ingest sensor data from IoT devices. The data arrives continuously in small batches and must be processed in real-time for monitoring. Which type of data source best describes this scenario?

A.Transactional database
B.Streaming data
C.Web scraping
D.Flat file
AnswerB

Streaming data matches the continuous, small-batch arrival described in the stem, where each event is processed as it occurs rather than stored and queried later. This satisfies the real-time monitoring constraint, since latency stays low and the pipeline reacts to sensor readings immediately instead of waiting for scheduled batch windows.

Why this answer

B is correct because the scenario describes data arriving continuously in small batches that must be processed in real-time for monitoring. This is the defining characteristic of streaming data, which is typically ingested via technologies like Apache Kafka, Amazon Kinesis, or MQTT brokers, enabling low-latency processing and immediate alerting.

Exam trap

The trap here is that candidates may confuse 'real-time' with 'fast batch processing' and incorrectly choose a transactional database, not recognizing that streaming data sources are specifically designed for continuous, unbounded data flows with sub-second latency requirements.

How to eliminate wrong answers

Option A is wrong because a transactional database (e.g., PostgreSQL, MySQL) is designed for ACID-compliant, query-based storage and retrieval, not for continuous real-time ingestion of sensor data; it would introduce latency and cannot handle unbounded streams efficiently. Option C is wrong because web scraping is a technique for extracting data from web pages via HTTP requests (e.g., using BeautifulSoup or Scrapy), which is batch-oriented and not suited for real-time IoT sensor data. Option D is wrong because a flat file (e.g., CSV, JSON file) is a static storage format that requires manual or scheduled batch loads, making it incapable of supporting real-time processing or continuous ingestion.

665
Multi-Selecthard

A data analyst is integrating data from two source systems into a single customer dataset. Source A uses a customer ID format like 'CUST-12345', while Source B uses '12345'. Additionally, Source A records dates in 'MM/DD/YYYY' format, while Source B uses 'YYYY-MM-DD'. Which two data preparation tasks are essential to ensure the integrated dataset is consistent and usable? (Choose two.)

Select 2 answers
A.Apply encryption to all customer ID fields.
B.Convert all date values to a single standardized format.
C.Standardize the customer ID format across both sources.
D.Remove all records with missing values.
E.Aggregate the data by customer to reduce row count.
AnswersB, C

Converting dates to a uniform format ensures temporal comparisons and calculations are valid. Different date formats can lead to misinterpretation, sorting errors, and failed date functions. Standardizing dates is essential for consistent time-based analysis and reporting across the integrated dataset.

Why this answer

The key challenges are inconsistent customer ID formats and date formats across sources. Standardizing both ensures records can be matched and time-based analysis is accurate. Encryption, missing value removal, and aggregation do not directly solve these format inconsistencies and may introduce other issues.

Exam trap

The trap here is focusing on data security or missing values instead of the format harmonization needed for integration.

666
MCQhard

A data engineer is designing a data pipeline to ingest streaming data from IoT sensors. The sensors send data every second, and the pipeline must handle bursts of up to 10,000 messages per second. Which approach is most appropriate for capturing this data before processing?

A.Directly write each message to a relational database
B.Load directly into a data warehouse
C.Use a message queue to buffer the incoming data
D.Store data in flat files and process in nightly batches
AnswerC

A message queue decouples producers from consumers, buffering bursts of up to 10,000 messages per second so the ingestion tier is not overwhelmed. This satisfies the stem's burst-handling constraint by absorbing spikes and letting downstream processing drain at its own rate.

Why this answer

A message queue (e.g., Apache Kafka, Amazon Kinesis, or RabbitMQ) provides an asynchronous buffer that decouples the high-velocity ingestion (up to 10,000 messages/second) from downstream processing. This allows the pipeline to absorb burst traffic without overwhelming the processing layer, ensures data durability, and supports replayability in case of failures.

Exam trap

CompTIA often tests the misconception that relational databases or data warehouses can handle real-time streaming ingestion at scale, when in fact they require a buffering layer like a message queue to absorb bursts and decouple ingestion from processing.

How to eliminate wrong answers

Option A is wrong because directly writing each message to a relational database (RDBMS) at 10,000 messages/second would cause severe write contention, lock contention, and I/O bottlenecks, leading to dropped data and unacceptable latency. Option B is wrong because loading directly into a data warehouse (e.g., Snowflake, Redshift) is designed for batch or micro-batch ingestion, not for real-time streaming at this scale; it would incur high costs and fail to handle bursty throughput without prior buffering. Option D is wrong because storing data in flat files and processing in nightly batches introduces unacceptable latency (up to 24 hours) for streaming IoT data, and the file system cannot reliably handle 10,000 writes per second without data loss or corruption.

667
Multi-Selecthard

Which THREE of the following are valid methods for handling missing data?

Select 3 answers
A.Using a placeholder like 'Unknown' for categorical data
B.Ignoring missing values and proceeding with analysis
C.Replacing missing values with the mean of the column
D.Sorting the data to bring missing values to the top
E.Deleting rows with missing values
AnswersA, C, E

Placeholder is a valid approach.

Why this answer

Using a placeholder like 'Unknown' for categorical missing data preserves the dataset's structure and allows analysis to proceed without introducing statistical bias. This method is particularly valid for nominal data where the missing category can be treated as a distinct value, enabling downstream operations like one-hot encoding or frequency analysis without distorting the original distribution.

Exam trap

The trap here is that candidates may confuse 'handling missing data' with 'preprocessing steps'—sorting (Option D) is a data organization technique, not a valid method for dealing with missing values, and ignoring missing data (Option B) is often mistakenly considered acceptable in quick analyses, but it violates best practices for robust data science workflows.

668
Multi-Selecthard

When presenting data insights, which three practices help ensure clarity? (Choose THREE.)

Select 3 answers
A.Include all data points in the chart
B.Overlay multiple chart types on one axis
C.Use consistent color coding
D.Label axes clearly
E.Provide context and benchmarks
AnswersC, D, E

Consistent colour coding assigns stable meaning to each category across every chart, so viewers learn the mapping once and apply it throughout. This satisfies the clarity constraint by removing the cognitive cost of reinterpreting colours slide by slide.

Why this answer

Option C is correct because consistent color coding lets viewers instantly associate a color with the same category, metric, or series across every chart, slide, or dashboard, reducing cognitive load and preventing misinterpretation. Option D is correct because clearly labeling axes—including units, scale, and time period—ensures the audience knows exactly what is being measured and how to read the values, which is essential for accurate interpretation. Option E is correct because providing context and benchmarks (such as targets, prior-period comparisons, or industry averages) turns raw numbers into meaningful insights and helps the audience judge whether a result is good, bad, or expected.

Option A does not belong because including every data point creates clutter and obscures the key message; effective visualizations simplify and highlight what matters. Option B does not belong because overlaying multiple chart types on a single axis is confusing and can distort scale relationships, so dual axes or separate charts are preferred when different measures must be shown.

Exam trap

The trap here is that candidates may think including all data points (Option A) ensures completeness, but the exam tests the understanding that clarity often requires simplification and focus, not exhaustive detail.

669
MCQeasy

A data analyst is importing a CSV file that contains a mixture of numeric and text fields. What is the most common issue when importing?

A.Duplicate rows
B.Missing header row
C.Data types being incorrectly inferred
D.File size limitation
AnswerC

CSV files carry no type metadata, so the import engine must guess each column's type from sampled values. Mixed numeric and text fields cause misinference — for example, leading-zero codes becoming integers or numeric-looking text converting to numbers — which is the most common CSV import problem.

Why this answer

When importing CSV files, the most common issue is that the import tool (e.g., Excel, pandas, SQL Server Import Wizard) automatically infers data types based on the first rows it reads. Mixed numeric and text fields often cause the tool to guess wrong — for example, treating a numeric column with a stray text value as text, or converting leading-zero codes (like ZIP codes) to integers and losing the zeros. This type inference mismatch is the classic CSV import pitfall.

Exam trap

DA0-002 often tests the misconception that CSV import issues are about file size or duplicates, when the real culprit is automatic data type inference on mixed columns.

How to eliminate wrong answers

Option A is wrong because duplicate rows are a data quality issue that may or may not exist, but they are not an inherent import problem caused by mixed types. Option B is wrong because a missing header row is a file structure issue, not a type inference issue, and many import tools can handle headerless files with manual configuration. Option D is wrong because file size limitations are a platform constraint (e.g., Excel's 1,048,576 row limit) and are unrelated to the mixture of numeric and text fields.

670
MCQmedium

An analyst is conducting an A/B test to compare two website designs. The null hypothesis is that there is no difference in conversion rates. The p-value obtained is 0.03, and the significance threshold is 0.05. What should the analyst conclude?

A.Reject the null hypothesis; there is a significant difference.
B.Accept the alternative hypothesis that the new design is better.
C.The test is inconclusive; need a larger sample size.
D.Fail to reject the null hypothesis; there is no significant difference.
AnswerA

A p-value of 0.03 falls below the 0.05 significance threshold, so the null hypothesis of equal conversion rates is rejected. The result is statistically significant, indicating the observed difference between the two website designs is unlikely to arise from chance alone.

Why this answer

In hypothesis testing, if the p-value (0.03) is less than the significance level (0.05), you reject the null hypothesis. This indicates that there is statistically significant evidence of a difference in conversion rates between the two designs. The correct conclusion is to reject the null hypothesis.

Exam trap

The trap is interpreting a significant p-value as proof that the alternative hypothesis is true in a specific direction (e.g., new design is better), when it only indicates a difference.

How to eliminate wrong answers

Option B is wrong because rejecting the null hypothesis does not automatically mean the new design is better; it only indicates a significant difference, not the direction. Option C is wrong because the test is not inconclusive; the p-value is below the threshold, so a conclusion can be drawn. Option D is wrong because failing to reject the null hypothesis would require a p-value greater than 0.05.

671
MCQeasy

An organization needs to store raw data from IoT sensors in its native format for future analysis. Which storage solution is best suited for this purpose?

A.Relational database
B.Data lake
C.Data mart
D.Data warehouse
AnswerB

A data lake stores raw data in its native format without transformation or schema enforcement, preserving fidelity for future, possibly unknown, analytical needs. This schema-on-read approach suits high-volume, varied IoT sensor output better than warehouses or relational stores that demand predefined structure.

Why this answer

A data lake is designed to store raw data in its native format, including unstructured and semi-structured data from IoT sensors, without requiring a predefined schema. This allows the organization to preserve the original data for future analysis, unlike traditional databases that enforce structure upon ingestion.

Exam trap

The trap here is that candidates often confuse a data warehouse with a data lake, assuming both are for storage, but a data warehouse requires ETL and structured schemas, making it unsuitable for raw, native-format IoT data.

How to eliminate wrong answers

Option A is wrong because a relational database requires a predefined schema and is optimized for structured data, not raw, native-format IoT sensor data. Option C is wrong because a data mart is a subset of a data warehouse focused on a specific business domain, not designed for storing raw, unprocessed data. Option D is wrong because a data warehouse stores processed, structured, and transformed data for analytical queries, not raw data in its native format.

672
Multi-Selectmedium

A data analyst is building a Power BI report to track KPIs for a retail chain. Which TWO of the following are considered leading indicators? (Choose two.)

Select 2 answers
A.Number of employees
B.Customer satisfaction score
C.Number of website visits
D.Profit margin
E.Total sales revenue
AnswersB, C

Customer satisfaction score qualifies as a leading indicator because it measures perceptions that precede and predict future purchase behaviour, satisfying the stem's requirement for forward-looking KPIs. Unlike lagging revenue or profit figures, it signals upcoming retention and sales trends, letting the retail chain act before financial outcomes materialise.

Why this answer

Customer satisfaction score (B) is a leading indicator because it measures how customers feel about the company's products or service, which tends to predict future repeat purchases, retention, and revenue before those financial outcomes appear. Number of website visits (C) is also a leading indicator because it reflects current interest and engagement that typically precedes conversions and sales, giving an early signal of future demand. By contrast, profit margin (D) and total sales revenue (E) are lagging indicators, since they report financial results that have already occurred.

Number of employees (A) is a structural or capacity measure rather than a predictive indicator of future performance, so it does not qualify as a leading indicator here.

Exam trap

The trap is that candidates see financial metrics like profit margin and revenue and assume they are important KPIs, but importance does not equal leading—the exam tests whether you understand the temporal direction of the indicator.

673
MCQeasy

Which of the following is an example of unstructured data?

A.A JSON file
B.An image file
C.A relational database table
D.A CSV file with rows and columns
AnswerB

An image file stores pixel data with no predefined schema, so it cannot be queried by fixed fields or columns. That absence of a structured, tabular model is what makes it unstructured, unlike records held in relational tables or delimited formats.

Why this answer

An image file is unstructured data because it has no predefined data model or schema — its content is raw pixel data that requires specialized processing (e.g., computer vision) to extract meaning. This contrasts with structured formats that organize data into rows, columns, or key-value pairs.

Exam trap

DA0-002 often tests the boundary between semi-structured and unstructured data, tempting candidates to misclassify JSON or XML as unstructured when they are actually semi-structured.

How to eliminate wrong answers

Option A is wrong because a JSON file is semi-structured data — it has a defined syntax with keys and values that can be parsed into a schema. Option C is wrong because a relational database table is the archetypal structured data, with rows, columns, and defined data types. Option D is wrong because a CSV file with rows and columns is structured data, easily mapped to a tabular schema.

674
MCQeasy

A data analyst calculates the mean, median, and mode of a dataset. Which of the following best describes how these measures are used in descriptive statistics?

A.To identify outliers using standard deviation
B.To test hypotheses about population parameters
C.To describe the central tendency of the data
D.To determine the probability of an event
AnswerC

Mean, median and mode each summarise a dataset with a single representative value, describing where data clusters. They are the standard measures of central tendency in descriptive statistics, distinct from dispersion measures such as range or standard deviation.

Why this answer

Mean, median, and mode are measures of central tendency, which are used in descriptive statistics to summarize the center point of a dataset. They describe where the data tends to cluster. The correct answer is C because these measures describe the central tendency.

Exam trap

The trap is mixing up descriptive statistics (summarizing data) with inferential statistics (drawing conclusions about populations), leading candidates to choose hypothesis testing or probability options.

How to eliminate wrong answers

Option A is wrong because identifying outliers using standard deviation is a measure of dispersion, not central tendency. Option B is wrong because testing hypotheses about population parameters is inferential statistics, not descriptive. Option D is wrong because determining the probability of an event is probability theory, not descriptive statistics.

675
MCQeasy

When creating a data story, what should be the FIRST step?

A.Design the layout of the report
B.Gather all available data
C.Choose visualization tools
D.Identify the key message
AnswerD

Starting with the key message ensures every subsequent choice of chart, metric and narrative supports one clear conclusion, rather than assembling visuals first and hoping a story emerges. This satisfies the stem's requirement for the first step in data storytelling.

Why this answer

The first step in creating a data story is to identify the key message because the entire narrative, data selection, and visualization choices must align with a central insight. Without a clear message, subsequent steps like gathering data or designing layouts risk being unfocused and failing to communicate actionable findings. This ensures the story is driven by the insight, not by the data or tools.

Exam trap

CompTIA often tests the misconception that data gathering or tool selection is the logical starting point, tempting candidates to skip the critical step of defining the core insight first.

How to eliminate wrong answers

Option A is wrong because designing the layout of the report should occur after the key message is defined; layout is a presentation concern, not the foundational step. Option B is wrong because gathering all available data without a guiding message leads to analysis paralysis and irrelevant information; data should be collected with the key message in mind. Option C is wrong because choosing visualization tools depends on the data and the story to be told; selecting tools prematurely can limit the narrative or force inappropriate chart types.

Page 8

Page 9 of 14

Page 10