Courseiva

CompTIA Data+ (DA0-002) (DA0-002) — Questions 826–900

1004 questions total · 14pages · All types, answers revealed

Page 11

Page 12 of 14

Page 13
826
MCQmedium

A national hotel chain stores guest records in a cloud data warehouse. Before a loyalty-analytics project begins, the governance team must document what each field means, how it is formatted, and who owns it. Which deliverable best satisfies this requirement?

A.A data lineage diagram tracing how guest records flow from the property management system into the warehouse
B.A data catalog entry containing field definitions, data types, and the assigned data steward
C.A data quality scorecard reporting completeness and validity rates for each guest attribute
D.An entity-relationship diagram showing primary and foreign key relationships between guest tables
AnswerB

A data catalog is the governance artifact that captures technical metadata such as field names, data types, and formats alongside business metadata such as definitions and the accountable data steward. It makes guest-record semantics discoverable and auditable before the loyalty project consumes the data, directly fulfilling the requirement to document meaning, format, and ownership.

Why this answer

The requirement combines business meaning, technical format, and accountability for each field — precisely the contents of a data catalog entry. Catalogs are purpose-built for discoverability of definitions and stewardship, so they let the loyalty team interpret guest records correctly and know whom to contact about changes. Diagrams, lineage, and scorecards each answer different governance questions and leave the definition-and-ownership gap unfilled.

Exam trap

The trap here is assuming that any documentation of the data, such as a diagram or lineage map, satisfies a metadata-definition requirement even though it omits meaning, format, and ownership.

827
MCQmedium

A data analyst is preparing a dataset for a machine learning model. The dataset contains a 'country' column with 150 unique values. To reduce dimensionality, the analyst wants to group less frequent countries into an 'Other' category. Which technique is being applied?

A.One-hot encoding
B.Normalization
C.Binning
D.Imputation
AnswerC

Binning (or bucketing) groups continuous or categorical values into a smaller number of bins. Here, grouping infrequent countries into 'Other' is a form of categorical binning. This reduces the number of unique categories, simplifies the feature, and can improve model performance by limiting noise from rare categories. It is a standard dimensionality reduction technique for high-cardinality categorical variables.

Why this answer

The technique is binning, specifically grouping infrequent categories into an 'Other' bin. This reduces the cardinality of the 'country' feature, which can help prevent overfitting and improve model training efficiency. Binning is a common approach for handling high-cardinality categorical variables by consolidating rare levels into a single category.

Exam trap

The trap here is confusing binning with one-hot encoding, but one-hot encoding expands the feature space while binning reduces it.

828
MCQmedium

A retail company generates a daily PDF report showing the previous day's sales by region and product category. The report is automatically emailed to store managers at 6:00 AM. Which type of report is this?

A.Scheduled report
B.Self-service report
C.Ad hoc report
D.Operational report
AnswerA

Automated generation and emailing at a fixed 6:00 AM time defines a scheduled report, driven by a recurring time trigger rather than user request. This matches the daily delivery requirement, distinguishing it from ad hoc or on-demand reporting.

Why this answer

A scheduled report is generated automatically at a predefined time (6:00 AM daily) and delivered to recipients (store managers) via email. This matches the definition of a scheduled report, which runs on a recurring basis without manual intervention.

Exam trap

The trap is confusing scheduled reports with ad hoc or self-service reports; the key differentiator is the automatic, recurring delivery at a set time.

How to eliminate wrong answers

Option B is wrong because a self-service report is created and run by end-users on demand, not automatically delivered. Option C is wrong because an ad hoc report is a one-time, on-demand query created for a specific need, not a recurring scheduled delivery. Option D is wrong because an operational report typically refers to real-time or near-real-time monitoring reports for day-to-day operations, not a daily summary emailed to managers.

829
MCQmedium

A logistics company has data on delivery times (continuous) and distance traveled (continuous). They want to visualize the relationship between these two variables. Which chart type is most appropriate?

A.Histogram
B.Bar chart
C.Line chart
D.Scatter plot
AnswerD

A scatter plot places distance on one axis and delivery time on the other, plotting each shipment as a point so correlation, clustering and outliers between two continuous variables become visible. This directly satisfies the stem's relationship-visualisation requirement.

Why this answer

A scatter plot is the correct choice because it displays individual data points on a two-dimensional plane, with one continuous variable on the x-axis (distance traveled) and the other on the y-axis (delivery time). This allows analysts to visually assess the strength, direction, and form of the relationship between two continuous variables, including spotting correlations, clusters, or outliers. Unlike other charts, scatter plots are specifically designed for bivariate continuous data, making them ideal for this logistics use case.

Exam trap

The trap here is confusing charts that display distributions (histogram) or trends over time (line chart) with those that show relationships between two continuous variables, leading candidates to overlook the scatter plot as the only appropriate choice for bivariate continuous data.

How to eliminate wrong answers

Option A is wrong because a histogram is used to show the distribution of a single continuous variable by binning values into intervals, not to visualize the relationship between two variables. Option B is wrong because a bar chart compares categorical data or discrete groups using rectangular bars, and it cannot effectively represent two continuous variables simultaneously. Option C is wrong because a line chart is typically used to display trends over time or ordered sequences, connecting points with lines, which implies a sequential relationship that does not exist between distance and delivery time.

830
Multi-Selecthard

A data governance committee is defining metadata standards for a new data lake. They need to distinguish between technical metadata and business metadata to ensure proper data management. Which two of the following are examples of technical metadata? (Choose two.)

Select 2 answers
A.Data type of a column (e.g., VARCHAR, INTEGER)
B.Business definition of a 'customer'
C.Data quality threshold for completeness
D.Data owner responsible for a dataset
E.Number of rows in a table
AnswersA, E

Technical metadata describes the structure and format of data, such as data types, lengths, and constraints. The data type of a column directly defines how data is stored and processed, making it essential for schema management, ETL development, and query optimization. This is a classic example of technical metadata used by data engineers and database administrators.

Why this answer

Technical metadata describes the structure, format, and operational characteristics of data, such as data types and row counts. Business metadata provides semantic context, definitions, ownership, and quality rules. The data type of a column and the number of rows in a table are technical metadata, while business definitions, data owners, and quality thresholds are business metadata.

Exam trap

The trap here is conflating business governance roles and definitions with technical metadata, which strictly concerns the physical and structural attributes of data systems.

831
MCQhard

A data analyst is presenting a weekly sales report to a diverse audience that includes sales managers, marketing team, and finance. The analyst wants to ensure the data is understood correctly. Which communication strategy is most effective?

A.Use a single complex chart to show all data at once.
B.Create a slide deck with key insights and contextual explanations.
C.Email the raw data table to all stakeholders.
D.Use technical terms such as 'variance' and 'standard deviation'.
AnswerB

A slide deck pairing key insights with contextual explanations tailors the message for mixed audiences, letting sales, marketing and finance each interpret the figures correctly. This satisfies the need for shared understanding across differing backgrounds rather than raw data dumps.

Why this answer

A slide deck with key insights and contextual explanations tailors the presentation to a diverse audience, ensuring that sales managers, marketing, and finance can each grasp the relevant data without being overwhelmed. This approach aligns with best practices for data communication, where visual summaries and narrative context improve comprehension over raw data or overly technical jargon.

Exam trap

The trap here is that candidates may overvalue technical precision (Option D) or data completeness (Option C), forgetting that the primary goal is audience comprehension, not data exhaustiveness.

How to eliminate wrong answers

Option A is wrong because a single complex chart (e.g., a multi-series radar or stacked area chart) can obscure individual trends and overload viewers, leading to misinterpretation across roles with different analytical backgrounds. Option C is wrong because emailing raw data tables forces stakeholders to derive insights themselves, risking errors and inefficiency, especially for non-technical teams like marketing. Option D is wrong because using technical terms like 'variance' and 'standard deviation' without explanation alienates audience members who lack statistical training, violating the principle of audience-appropriate communication.

832
MCQhard

A financial analyst is preparing a report that must comply with the Sarbanes-Oxley (SOX) Act. What is the most critical requirement for this report?

A.Row-level security
B.Data anonymization
C.Data dictionary
D.Audit trail
AnswerD

An audit trail provides an immutable, timestamped record of who accessed or altered financial data, which directly satisfies SOX's requirement for verifiable internal controls over financial reporting. Unlike confidentiality or availability controls, this traceability is what auditors examine to evidence compliance, making it the critical requirement here.

Why this answer

SOX requires audit trails to ensure the integrity and traceability of financial data for compliance.

833
MCQhard

A data analyst is presenting a recommendation to reduce inventory costs. The evidence shows that overstocking occurs in 30% of warehouses. Which of the following best structures the recommendation?

A."Implement a just-in-time system."
B."Our inventory costs are too high. We need to fix this."
C."Overstocking occurs in 30% of warehouses. This indicates poor demand forecasting. We recommend implementing a just-in-time system to reduce inventory holding costs by 15%."
D."We should reduce inventory by 20%. This will save costs."
AnswerC

This structure states the evidence, interprets it as a cause, then gives a quantified recommendation, forming a logical evidence-to-action chain. It satisfies the stem's requirement for a well-structured recommendation by linking the 30% overstocking finding to a just-in-time proposal.

Why this answer

A proper data-driven recommendation follows the sequence: evidence → insight → recommendation → expected impact.

834
MCQhard

A database table has columns: OrderID (primary key), ProductID, CustomerID, CustomerName, OrderDate, ProductName. All products are purchased only by the customer who placed the order. Which normal form violation exists if CustomerName depends on CustomerID?

A.Boyce-Codd normal form (BCNF)
B.Third normal form (3NF)
C.Second normal form (2NF)
D.First normal form (1NF)
AnswerB

CustomerName depends on CustomerID, which is not a candidate key, creating a transitive dependency and violating 3NF.

Why this answer

The table violates Third Normal Form (3NF) because CustomerName depends on CustomerID, which is not a candidate key (the primary key is OrderID). 3NF requires that every non-key attribute be non-transitively dependent on the primary key; here, CustomerName is transitively dependent on OrderID via CustomerID. Since CustomerID is a non-key attribute (it is not part of the primary key), this transitive dependency breaks 3NF.

Exam trap

The trap here is that candidates often confuse transitive dependencies (3NF violation) with partial dependencies (2NF violation) or think that any dependency on a non-key attribute automatically violates BCNF, but the specific scenario of CustomerName depending on CustomerID is a textbook transitive dependency that breaks 3NF first.

How to eliminate wrong answers

Option A is wrong because Boyce-Codd Normal Form (BCNF) is a stricter version of 3NF that requires every determinant to be a candidate key; while this table also violates BCNF, the question asks which normal form violation exists, and the dependency described is a classic 3NF violation (transitive dependency), not a BCNF-specific one. Option C is wrong because Second Normal Form (2NF) is violated only when a non-key attribute depends on a proper subset of a composite primary key; here the primary key is a single column (OrderID), so no partial dependency exists, and 2NF is satisfied. Option D is wrong because First Normal Form (1NF) is violated only if there are repeating groups or non-atomic values; the table as described has atomic columns and no repeating groups, so 1NF is satisfied.

835
MCQhard

A data analyst is building a dashboard that tracks daily active users. The source system records one event row per user session, and a user can have multiple sessions in a day. The analyst needs to display the count of unique users per day. Which approach produces the correct daily active user count?

A.Sum the maximum session count per user across all dates
B.Average the number of session rows per user for each date
C.Count all session rows grouped by date
D.Count distinct user identifiers grouped by date
AnswerD

Applying a distinct count to the user identifier within each date collapses multiple sessions for the same user into one, producing the true number of unique users active that day. This directly matches the definition of daily active users and prevents session volume from distorting the metric.

Why this answer

Daily active users is defined as the number of distinct users who take an action on a given day. Because the source grain is one row per session, the analyst must deduplicate users within each date before counting. A distinct count of the user identifier grouped by date performs exactly that deduplication and yields the correct metric, whereas counting rows or averaging sessions measures activity volume instead of reach.

Exam trap

The trap here is treating the row grain of the source table as the unit of analysis, when the business question asks about distinct people rather than events.

836
MCQmedium

A data analyst at a retail company is profiling a newly acquired customer table. They observe that the 'last_purchase_date' column contains values such as '2023-13-45', '0000-00-00', and '2023-02-30'. Which data quality dimension is primarily violated?

A.Uniqueness
B.Completeness
C.Validity
D.Consistency
AnswerC

Validity checks whether data conforms to defined formats, types, or business rules. Here, the date values are not valid calendar dates (month 13, day 45, day 30 in February), so they violate the validity dimension. The analyst must apply date validation rules to flag or correct these entries before using the data for analysis.

Why this answer

The date values shown are not valid calendar dates, meaning they fail format and range checks. Validity ensures data adheres to defined rules, such as a date being a real date. Completeness, consistency, and uniqueness address different aspects and do not capture the core problem of impossible dates.

Exam trap

The trap here is confusing malformed values with missing values, leading to a completeness answer instead of validity.

837
Drag & Dropmedium

Drag and drop the steps to perform a data backup using the 3-2-1 rule in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

The 3-2-1 rule involves multiple copies, different media, offsite storage, and regular testing.

838
Multi-Selectmedium

A data analyst is profiling a new dataset and needs to assess data quality. Which two metrics are most appropriate for evaluating the completeness and consistency of the data? (Choose two.)

Select 2 answers
A.Data type mismatch count between source and target
B.Frequency of values violating defined business rules
C.Number of distinct values in a column
D.Percentage of missing values per column
E.Average record length in bytes
AnswersB, D

The frequency of values violating business rules measures consistency and validity. Business rules define acceptable ranges or formats, and violations indicate data that does not conform. This metric is crucial for identifying data quality issues that could lead to incorrect analysis or decisions.

Why this answer

Completeness is assessed by the percentage of missing values, which shows how much data is absent. Consistency is assessed by the frequency of business rule violations, which reveals non-conforming data. Together, these metrics provide a clear picture of data quality issues that need addressing before analysis.

Exam trap

The trap here is selecting metrics that seem related to data quality but do not directly measure completeness or consistency, such as distinct value counts or record length.

839
MCQmedium

A data analyst is compiling a quarterly report for a retail chain and needs to show how sales revenue is distributed across five product categories. The analyst wants a single visualization that makes it easy to compare each category's share of total revenue and also includes a separate small panel displaying total revenue for the quarter. Which visualization approach best meets these requirements?

A.A dashboard with a pie chart showing category shares and a separate KPI card displaying total revenue.
B.A stacked bar chart showing revenue by category, with the total revenue represented by the height of the entire stack.
C.A donut chart with each category's share, plus a small table in a separate panel listing the total revenue.
D.A pie chart showing each category's percentage of total revenue, with the total revenue displayed as the chart title.
AnswerA

A dashboard allows multiple visual elements to be combined. A pie chart clearly shows each category's proportion of total revenue, while a separate KPI card prominently displays the overall total. This directly satisfies both requirements: comparing shares and having a distinct panel for the total. It is the most effective and standard approach for this type of reporting need.

Why this answer

The requirement is to compare each category's share of total revenue and to have a separate small panel for total revenue. A dashboard with a pie chart and a KPI card meets both: the pie chart shows proportions, and the KPI card is a distinct panel for the total. Other options either embed the total in the chart or use less effective visuals for comparing shares.

Exam trap

The trap here is assuming that displaying the total in a chart title or as part of a stacked bar satisfies the need for a separate panel, when the scenario explicitly calls for a distinct panel.

840
MCQhard

A data modeler is designing a dimensional model for a sales analytics system. The fact table contains sales transactions, and the dimension tables include product, customer, and time. To reduce data redundancy, the modeler normalizes the dimension tables into multiple related tables. Which schema is being implemented?

A.Vault schema
B.Star schema
C.Galaxy schema
D.Snowflake schema
AnswerD

Normalising dimensions into multiple related tables produces a snowflake schema: each dimension is decomposed into its own hierarchy of linked tables rather than one flat denormalised table. This directly satisfies the stated goal of reducing redundancy in the product, customer and time dimensions.

Why this answer

The snowflake schema is a dimensional model where dimension tables are normalized into multiple related tables to reduce data redundancy. In this scenario, the product, customer, and time dimensions are split into sub-dimensions (e.g., product category, customer geography, time hierarchy), which is the defining characteristic of a snowflake schema. This contrasts with a star schema where dimensions remain denormalized.

Exam trap

CompTIA often tests the distinction between star and snowflake schemas by emphasizing normalization of dimensions; the trap here is that candidates may confuse 'normalized dimensions' with a star schema, which actually uses denormalized dimensions for simplicity and performance.

How to eliminate wrong answers

Option A is wrong because a vault schema (Data Vault) is a hybrid modeling approach focused on auditability and flexibility using hubs, links, and satellites, not on normalizing dimension tables for a sales analytics fact table. Option B is wrong because a star schema keeps dimension tables denormalized (single table per dimension) to optimize query performance, which directly contradicts the normalization described in the question. Option C is wrong because a galaxy schema (also called a fact constellation) contains multiple fact tables sharing dimension tables, not the normalization of a single fact table’s dimensions.

841
MCQeasy

The exhibit shows log entries. A data analyst wants to visualize the frequency of each error type over time. Which chart type is most appropriate?

A.Time series line chart
B.Bar chart of error types
C.Scatter plot of timestamp vs error code
D.Pie chart of error types
AnswerA

A time series line chart plots each error type's frequency against a continuous time axis, directly satisfying the requirement to visualise frequency over time. Unlike categorical charts such as bar or pie charts, it preserves temporal ordering and reveals trends, spikes and seasonality across error types.

Why this answer

A time series line chart plots error frequency on the y-axis against time on the x-axis, directly showing how each error type's occurrence changes over time — exactly what the analyst needs. Line charts are ideal for trend analysis and can display multiple error types as separate lines for comparison. This makes temporal patterns (spikes, trends, seasonality) immediately visible.

Exam trap

DA0-002 often tests the confusion between categorical comparison charts (bar, pie) and temporal trend charts (line), tricking candidates into choosing a chart that shows distribution but not change over time.

How to eliminate wrong answers

Option B is wrong because a bar chart of error types shows total counts per category but collapses the time dimension, hiding when errors occurred. Option C is wrong because a scatter plot of timestamp vs error code would show individual events but not aggregated frequency trends, making it hard to read patterns. Option D is wrong because a pie chart shows proportional distribution of error types at a single point in time, with no temporal axis at all.

842
Multi-Selecthard

A data analyst is preparing a dataset for analysis and needs to address data quality issues. The dataset contains missing values, outliers, and inconsistent formats. Which two techniques are appropriate for handling missing values? (Choose two.)

Select 2 answers
A.Mean imputation
B.Winsorizing
C.Listwise deletion
D.One-hot encoding
E.Z-score normalization
AnswersA, C

Mean imputation replaces missing numerical values with the mean of the observed values. It is a simple method that preserves the overall mean of the variable, though it can reduce variance. It is appropriate when missingness is random and the variable is numeric, making it a valid technique for this scenario.

Why this answer

The correct answers are Mean imputation and Listwise deletion because both directly address missing values. Mean imputation fills gaps with a central tendency measure, while listwise deletion removes incomplete records. The other techniques are for scaling, encoding, or outlier handling, not for missing data.

Exam trap

The trap here is confusing data preprocessing techniques; for example, assuming that normalization or encoding also handles missing values, when they do not.

843
Multi-Selecthard

An analyst is creating an executive summary for a quarterly business review. Which THREE components are essential for an effective executive summary?

Select 3 answers
A.Actionable recommendations
B.Context (e.g., compared to prior quarter)
C.Detailed data tables
D.SQL queries used to extract data
E.Headline number (e.g., revenue growth of 15%)
AnswersA, B, E

Actionable recommendations translate analytical findings into specific next steps, satisfying the executive summary's purpose of guiding decision-making rather than merely reporting data. Executives need to know what to do with the insights, so recommendations grounded in the analysis ensure the summary drives business action instead of leaving conclusions implicit.

Why this answer

An executive summary should start with a headline number, provide context, and include actionable recommendations.

844
MCQmedium

A logistics company stores shipment records in a relational database. The ShipmentID column uniquely identifies every shipment, and the CarrierName column stores the name of the carrier that moved each shipment, such as "Northwind Freight" or "Acme Logistics". Analysts frequently group shipments by carrier to compare on-time performance. Which statement correctly describes the relationship between these two columns in this table?

A.CarrierName is a candidate key because analysts group by it frequently
B.ShipmentID and CarrierName together form a composite primary key
C.ShipmentID is the primary key of the table, and CarrierName is a non-key descriptive attribute
D.ShipmentID is a foreign key that references CarrierName as its parent key
AnswerC

ShipmentID uniquely identifies each row, so it qualifies as the primary key and enforces entity integrity. CarrierName holds descriptive text about the carrier and repeats across many shipments, so it is a non-key attribute rather than an identifier. This structure supports grouping shipments by carrier for on-time performance comparisons.

Why this answer

Because ShipmentID uniquely identifies each shipment row, it serves as the primary key, while CarrierName simply describes an attribute of the shipment that can repeat and supports grouping. A foreign key must point to a unique parent key in another table, a candidate key must be unique, and a composite key is unnecessary when a single column already guarantees uniqueness. The described schema therefore fits the primary key plus descriptive attribute model.

Exam trap

The trap here is confusing a frequently grouped, repeating column with a key, when key status depends on uniqueness rather than on how often the column is used in queries.

845
Multi-Selecteasy

Which TWO of the following are characteristics of OLTP systems? (Select 2)

Select 2 answers
A.Typically uses a denormalized schema
B.Optimized for complex analytical queries
C.Stores historical data for trend analysis
D.Designed for high transaction throughput
E.Supports ACID transactions
AnswersD, E

OLTP handles many concurrent transactions.

Why this answer

OLTP systems are designed for high transaction throughput, handling large volumes of short, atomic transactions efficiently. They prioritize fast data processing and immediate consistency, making option D correct.

Exam trap

The trap here is that candidates often confuse OLTP with OLAP, mistakenly selecting denormalized schemas or analytical optimization as OLTP characteristics, when in fact OLTP emphasizes normalized schemas and high transaction throughput with ACID compliance.

846
MCQmedium

A company has a large data warehouse running on Snowflake. They receive daily CSV files from multiple sources and load them directly into the warehouse, then run SQL transformations to clean and aggregate the data. Which data integration approach does this describe?

A.ELT
B.Data streaming
C.ETL
D.CDC
AnswerA

ELT fits because raw CSV files load straight into Snowflake before any cleansing, letting the warehouse's own compute run the SQL transformations. This satisfies the stem's constraint of loading directly, then transforming in place — the defining reversal of extract-load-transform order versus ETL, where transformation precedes loading.

Why this answer

This describes ELT (Extract, Load, Transform) because the raw CSV files are first loaded directly into Snowflake, and then SQL transformations are applied within the warehouse. Unlike ETL, where data is transformed before loading, ELT leverages Snowflake's compute power to perform transformations after ingestion, which is efficient for large-scale batch processing.

Exam trap

The trap here is that candidates confuse ELT with ETL because both involve transformations, but the key distinction is the order of loading versus transforming; CompTIA often tests this by describing the sequence of operations to see if you recognize that loading raw data first is the hallmark of ELT.

How to eliminate wrong answers

Option B is wrong because data streaming involves continuous, real-time ingestion (e.g., using Kafka or Kinesis), not daily batch CSV file loads. Option C is wrong because ETL would transform the data before loading into Snowflake, but the question states raw CSV files are loaded directly and then transformed afterward. Option D is wrong because CDC (Change Data Capture) captures incremental changes from source databases (e.g., via Debezium or Oracle GoldenGate), not daily full-file CSV imports.

847
MCQeasy

A data engineer needs to extract data from a REST API and load it into a data warehouse. The data is received in JSON format. Which data type best describes JSON?

A.Transactional
B.Semi-structured
C.Unstructured
D.Structured
AnswerB

JSON is semi-structured: it encodes hierarchical key-value pairs and arrays without a fixed relational schema, yet retains tags and nesting that allow parsing. This distinguishes it from fully unstructured text and from rigidly structured tabular formats.

Why this answer

JSON (JavaScript Object Notation) is classified as a semi-structured data type because it uses a flexible, self-describing schema with key-value pairs and nested structures, but does not enforce a rigid tabular schema like relational databases. In the context of extracting data from a REST API, JSON allows for varying fields and hierarchical data, which aligns with the semi-structured category.

Exam trap

The trap here is that candidates confuse the presence of structure (keys and values) with being fully structured, overlooking that JSON lacks a fixed schema and allows variability, which places it in the semi-structured category.

How to eliminate wrong answers

Option A is wrong because transactional data refers to records of business transactions (e.g., sales, orders) typically stored in structured formats with ACID properties, not to the format of the data itself. Option C is wrong because unstructured data lacks any predefined structure or schema (e.g., raw text, images, video), whereas JSON has a defined syntax with keys, values, and nesting. Option D is wrong because structured data requires a fixed schema (e.g., rows and columns in a relational table), while JSON allows optional fields and varying data types, making it semi-structured.

848
MCQmedium

An analyst recommends a pricing change based on data showing price elasticity. The recommendation includes expected revenue impact. What is this an example of?

A.Data-driven recommendation
B.Uncertainty communication
C.Self-service analysis
D.Executive summary
AnswerA

Basing a pricing decision on elasticity analysis and quantifying the expected revenue impact is a data-driven recommendation: the choice is justified by evidence rather than intuition. This directly matches the scenario's constraint of recommending action supported by analytical findings.

Why this answer

A data-driven recommendation is a decision or course of action grounded in analysis of data — here, price elasticity analysis — and it includes a quantified expected outcome such as revenue impact. The analyst is not merely presenting data; they are translating analysis into a recommended action with a projected business result. This is the defining characteristic of a data-driven recommendation.

Exam trap

DA0-002 often tests the difference between presenting data and making a recommendation — candidates may pick 'executive summary' because the recommendation is communicated in that format, missing that the substance is the data-driven recommendation itself.

How to eliminate wrong answers

Option B is wrong because uncertainty communication focuses on expressing confidence intervals, assumptions, and limitations of the analysis, which is not the primary description here. Option C is wrong because self-service analysis refers to business users accessing and analyzing data themselves with governed tools, not to an analyst delivering a recommendation. Option D is wrong because an executive summary is a format for presenting findings concisely, not the substance of a data-driven recommendation.

849
MCQmedium

A data analyst is examining a dataset of customer transactions and notices that the distribution of transaction amounts is highly skewed to the right. The analyst wants to apply a transformation to make the distribution more symmetric for use in a linear regression model. Which of the following transformations is most appropriate?

A.Box-Cox transformation with lambda = 2
B.Square root transformation
C.Reciprocal transformation
D.Logarithmic transformation
AnswerD

The logarithmic transformation is highly effective for right-skewed data with positive values, as it compresses large values more than small ones, reducing skewness and stabilizing variance. It is commonly used for monetary amounts and can make the distribution more symmetric, improving the performance of linear models. It also helps with heteroscedasticity.

Why this answer

For highly right-skewed positive data like transaction amounts, the logarithmic transformation is the standard choice to reduce skewness and make the distribution more symmetric. It compresses the upper tail and can linearize relationships, which benefits linear regression. Other transformations like square root are weaker, while reciprocal and squaring are either too strong or counterproductive.

Exam trap

The trap here is selecting a transformation that is either too weak or that increases skew, such as squaring, instead of the logarithmic transformation suited for right-skewed monetary data.

850
MCQeasy

In a simple linear regression model y = 2.5 + 1.2x, what is the predicted value of y when x = 10?

A.12.0
B.13.7
C.14.5
D.10.0
AnswerC

Substituting x = 10 into the fitted equation gives y = 2.5 + 1.2(10) = 2.5 + 12 = 14.5. The slope of 1.2 contributes twelve units across the ten-unit increase, and the intercept of 2.5 adds the baseline, yielding the predicted response.

Why this answer

Plug x=10: y = 2.5 + 1.2*10 = 2.5 + 12 = 14.5.

851
MCQeasy

A dashboard needs to show sales trends for each of five regions over the past year. The intended audience wants to compare trends easily. Which chart type is best?

A.Line chart with multiple lines
B.Pie chart
C.Stacked bar chart
D.Area chart
AnswerA

A line chart with multiple lines plots each region's sales against a shared time axis, so trends across all five regions can be compared directly within one view. This satisfies the requirement to show trends over the past year while making regional comparison easy, since position and slope are read simultaneously.

Why this answer

A line chart with multiple lines is the canonical choice for showing trends over time across several categories because the x-axis represents the continuous time dimension and each line tracks one region's trajectory. It allows the audience to compare slopes, inflection points, and relative performance at a glance. Pie, stacked bar, and area charts obscure per-series trends or make cross-series comparison harder.

Exam trap

The trap is confusing 'compare trends' with 'compare composition' — candidates who focus on the five regions may wrongly pick a pie or stacked bar, which show proportions rather than trends over time.

How to eliminate wrong answers

Option B is wrong because a pie chart shows part-to-whole composition at a single point in time and cannot represent a time series across twelve months. Option C is wrong because a stacked bar chart emphasizes cumulative totals and makes it difficult to compare individual region trends when segments shift. Option D is wrong because an area chart, especially stacked, also emphasizes cumulative magnitude and can hide the trajectory of lower series behind upper ones.

852
MCQmedium

A financial services firm stores customer account records in a relational database. An analyst must assign a classification label to a new table containing Social Security numbers, account balances, and transaction histories. The firm's data governance policy defines four tiers: Public, Internal, Confidential, and Restricted. Which classification tier is most appropriate for this table?

A.Confidential
B.Internal
C.Public
D.Restricted
AnswerD

Restricted is the highest classification tier, reserved for data whose exposure causes severe legal, financial, or reputational harm. Social Security numbers are regulated personally identifiable information, and combining them with account balances and transaction histories amplifies risk. This table therefore requires the strictest access controls, encryption, and monitoring that the Restricted tier mandates.

Why this answer

Social Security numbers combined with account balances and transaction histories represent the most sensitive category of data the firm handles. The highest classification tier exists precisely for regulated identifiers and financial records whose disclosure triggers legal penalties and identity theft risk. Assigning Restricted ensures the table receives maximum access controls, encryption, and audit monitoring consistent with the governance policy.

Exam trap

The trap here is assuming that because the data belongs to the firm and stays internal, the Internal tier suffices, ignoring that regulated personal identifiers demand the highest protection level.

853
Multi-Selecthard

A data analyst is validating a regression model that predicts monthly churn probability for 50,000 subscribers. The analyst wants to detect whether the model is overfitting before deploying it. (Choose two.)

Select 2 answers
A.Remove all records with missing values from the validation set only
B.Increase the number of predictor variables until training accuracy reaches 100 percent
C.Apply a log transformation to the target churn probability before evaluation
D.Compare the model's error on the training set with its error on a held-out validation set
E.Use k-fold cross-validation to estimate performance across multiple data splits
AnswersD, E

A large gap between training error and validation error is the classic symptom of overfitting, because the model has memorized training noise that does not generalize. Comparing the two error values directly quantifies that gap and tells the analyst whether the model's complexity is justified, making this a core diagnostic step before deployment.

Why this answer

Overfitting is detected by measuring how well a model performs on data it did not train on. Holding out a validation set and comparing its error with training error reveals the generalization gap, while k-fold cross-validation repeats that comparison across multiple partitions to produce a stable estimate. Together they show whether the churn model has learned signal versus memorized noise, which is exactly what must be established before deployment.

Exam trap

The trap here is confusing actions that increase model complexity with techniques that actually diagnose overfitting.

854
MCQhard

In the data lifecycle, which phase involves converting raw data into a usable format for analysis?

A.Ingestion
B.Analysis
C.Archival
D.Processing
AnswerD

Processing transforms raw data into a usable, analysis-ready format through cleaning, validation, aggregation and enrichment. It sits between collection and analysis in the lifecycle, directly matching the stem's conversion of raw data into usable form.

Why this answer

The processing phase in the data lifecycle is specifically where raw data is cleaned, transformed, and structured into a usable format for analysis. This includes operations such as parsing, normalization, deduplication, and conversion into formats like Parquet or Avro, which are optimized for query engines like Apache Spark or Presto.

Exam trap

The trap here is that candidates often confuse 'ingestion' with 'processing' because both involve moving data, but ingestion is about raw data capture, while processing is about transformation and cleaning before analysis.

How to eliminate wrong answers

Option A is wrong because ingestion refers to the initial collection and import of raw data from sources (e.g., via Apache Kafka or Flume) into a storage system, not its transformation into a usable format. Option B is wrong because analysis is the phase where processed data is queried, visualized, or modeled to derive insights, not where raw data is converted. Option C is wrong because archival involves moving older or infrequently accessed data to long-term storage (e.g., Amazon S3 Glacier or tape) for compliance or cost savings, not for preparing data for analysis.

855
Multi-Selectmedium

A data analyst wants to segment customers based on purchasing behavior such as frequency, monetary value, and recency. Which TWO clustering evaluation methods can help determine the optimal number of clusters? (Select two.)

Select 2 answers
A.Correlation coefficient
B.ANOVA
C.Silhouette score
D.t-test
E.Elbow method
AnswersC, E

The silhouette score measures how similar each point is to its own cluster versus the nearest other cluster, producing a coefficient between -1 and 1. The cluster count with the highest average silhouette indicates the best-separated segmentation.

Why this answer

The elbow method uses within-cluster sum of squares, and the silhouette score measures cohesion and separation. Both help choose k. Correlation coefficient is for association, not clustering.

ANOVA and t-test are for hypothesis testing.

856
MCQhard

A data analyst is asked to compare the average sales across three different store locations. The data is normally distributed and variances are approximately equal. Which statistical test is most appropriate?

A.ANOVA
B.Pearson correlation
C.Chi-square test
D.Two-sample t-test
AnswerA

ANOVA compares means across three or more groups simultaneously while controlling family-wise error. With normal distributions and roughly equal variances, its assumptions hold, making it correct for comparing average sales across the three store locations.

Why this answer

ANOVA (Analysis of Variance) is the correct test because it compares the means of three or more independent groups (store locations) while controlling the Type I error rate. The assumptions of normality and equal variances are satisfied, which are key prerequisites for ANOVA. A t-test would be inappropriate for more than two groups as it would require multiple pairwise comparisons, inflating the risk of false positives.

Exam trap

The trap here is confusing the number of groups: candidates might select a t-test for comparing means, but t-tests are limited to two groups; ANOVA is required for three or more.

How to eliminate wrong answers

Option B is wrong because Pearson correlation measures the linear relationship between two continuous variables, not differences in means across groups. Option C is wrong because the chi-square test is used for categorical data to test independence or goodness-of-fit, not for comparing means of continuous variables. Option D is wrong because a two-sample t-test only compares two groups; using it for three groups would require multiple tests, increasing the chance of a Type I error.

857
Multi-Selectmedium

Which TWO of the following are best practices for designing a data dashboard?

Select 2 answers
A.Include animated transitions between data views.
B.Use consistent color schemes to indicate performance levels.
C.Use 3D effects to make charts more visually appealing.
D.Place the most important KPIs at the top of the dashboard.
E.Include as many charts as possible to provide comprehensive data.
AnswersB, D

Consistent colour encoding lets viewers map a hue directly to a performance band, so thresholds are read pre-attentively without consulting a legend each time. This satisfies the stem's design best-practise requirement by reducing cognitive load and preventing misinterpretation when comparing metrics across panels.

Why this answer

Option B is correct because consistent color schemes (for example, a fixed red/amber/green semantic palette) let users instantly interpret performance levels without re-learning the meaning of colors on each view, which is a core dashboard design principle. Option D is correct because placing the most important KPIs at the top follows visual hierarchy and the F-pattern reading behavior, ensuring decision-makers see critical metrics first without scrolling. Option A is not a best practice because animated transitions add latency and distraction rather than improving data comprehension.

Option C is wrong because 3D effects distort proportions and make values harder to compare accurately. Option E is wrong because cramming in as many charts as possible creates clutter and cognitive overload, obscuring the key insights a dashboard should deliver.

Exam trap

The trap here is that candidates confuse 'visually appealing' with 'effective communication' — CompTIA often tests that decorative elements like 3D effects and animations reduce data accuracy and user comprehension, even though they may look impressive.

858
MCQeasy

A healthcare database stores patient records. Each patient has a unique patient_id, and the database includes a table 'visits' with visit_id, patient_id, visit_date, and diagnosis_code. To ensure data integrity, which constraint should be applied to the patient_id column in the 'visits' table?

A.Unique constraint
B.Foreign key
C.Primary key
D.Check constraint
AnswerB

A foreign key on patient_id in the visits table references the primary key in the patients table, ensuring every visit maps to an existing patient. This enforces referential integrity, preventing orphan visit records, which is the integrity requirement stated in the stem.

Why this answer

A foreign key constraint on patient_id in the visits table enforces referential integrity by ensuring every patient_id value in visits matches an existing patient_id in the patients table. This prevents orphaned visit records and maintains consistency between the two tables.

Exam trap

DA0-002 often tests the confusion between primary key, unique, and foreign key constraints, where candidates incorrectly apply uniqueness to a column that should allow duplicates but reference another table.

How to eliminate wrong answers

Option A is wrong because a unique constraint would prevent duplicate patient_id values in the visits table, which is incorrect since a patient can have multiple visits. Option C is wrong because a primary key on patient_id in visits would also enforce uniqueness and not-null, which is inappropriate for a foreign key column that can repeat. Option D is wrong because a check constraint validates values against a condition (e.g., range), not referential integrity with another table.

859
MCQhard

An analyst is performing a logistic regression to predict customer churn (yes/no). The model outputs a probability of 0.75 for a particular customer. Which of the following best describes the interpretation?

A.The model predicts that the customer will not churn
B.There is a 75% chance that the customer will churn
C.The customer will definitely churn because the probability is above 0.5
D.The odds of churning are 0.75 to 1
AnswerB

Logistic regression outputs a calibrated probability of the positive class, so 0.75 means a 75% estimated likelihood of churn for that customer. This satisfies the stem's interpretation requirement, distinguishing probabilistic output from a deterministic classification decision.

Why this answer

In logistic regression, the model outputs a probability between 0 and 1 that the positive class (churn = yes) occurs. A value of 0.75 means the model estimates a 75% probability that the customer will churn, given the input features. This is a probabilistic prediction, not a deterministic one, and it does not mean the customer will definitely churn.

Exam trap

DA0-002 often tests the confusion between probability and odds, and between probabilistic and deterministic predictions, causing candidates to misinterpret the output of logistic regression.

How to eliminate wrong answers

Option A is wrong because 0.75 is greater than 0.5, so the model predicts the positive class (churn), not the negative class. Option C is wrong because logistic regression outputs probabilities, not certainties; a 0.75 probability still implies a 25% chance of not churning, so 'definitely' is incorrect. Option D is wrong because 0.75 is a probability, not odds; odds would be calculated as p/(1-p) = 0.75/0.25 = 3, meaning 3 to 1 odds of churning, not 0.75 to 1.

860
MCQmedium

A data analyst needs to create a dashboard that updates automatically every hour. The data source is a large database. Which approach minimizes performance impact?

A.Query the entire database each time
B.Use incremental refresh only for new or changed data
C.Export the data to Excel and import
D.Create a static report monthly
AnswerB

Incremental refresh loads only rows added or modified since the last run, using a partition filter on a date or watermark column. This avoids reprocessing the entire large database each hour, directly minimising the performance impact the scenario demands.

Why this answer

Incremental refresh queries only new or changed records since the last refresh, drastically reducing data transfer and processing load on the large database. This approach uses change-tracking mechanisms (e.g., timestamps, CDC) to avoid full-table scans, minimizing performance impact while maintaining near-real-time updates.

Exam trap

CompTIA often tests the misconception that 'more data is better' or that full refreshes are simpler and equally acceptable, but the trap here is ignoring the performance cost of full database scans on large datasets in favor of the more efficient incremental approach.

How to eliminate wrong answers

Option A is wrong because querying the entire database each hour performs a full table scan on a large database, causing excessive I/O, CPU, and memory usage that degrades performance for all users. Option C is wrong because exporting the entire database to Excel and importing it adds unnecessary data transformation overhead, loses real-time capability, and still requires a full data pull. Option D is wrong because a static monthly report does not meet the requirement for automatic hourly updates and provides stale data, making it functionally incorrect for the use case.

861
Multi-Selectmedium

An analyst is conducting an A/B test on a new checkout process. To calculate sample size, which THREE factors must be considered?

Select 3 answers
A.Number of control groups
B.Desired effect size
C.Significance level (alpha)
D.Statistical power
E.Population standard deviation
AnswersB, C, D

Desired effect size is the minimum lift worth detecting between control and variant checkout processes. Smaller effects need larger samples, so effect size is a required input when calculating sample size for the A/B test.

Why this answer

Statistical power, significance level (alpha), and desired effect size (minimum detectable effect) are essential for sample size calculation.

862
MCQhard

Refer to the exhibit. A data analyst is reviewing a data quality report. Which of the following actions should the analyst take first?

A.Delete the 1200 records with null emails.
B.Fill null emails with a placeholder.
C.Investigate the source system to understand why emails are missing.
D.Ignore the nulls as they are not critical.
AnswerC

Missing emails indicate a data quality defect whose origin must be established before remediation. Investigating the source system reveals whether the cause is an optional field, a validation rule, or an extraction fault, satisfying the need to diagnose root cause first.

Why this answer

The first step in data quality remediation is root cause analysis. Without understanding why 1200 records have null emails (e.g., a source system bug, a failed ETL join, or a missing required field), any corrective action like deletion or placeholder insertion risks introducing bias or masking a systemic issue. Investigating the source system aligns with the data governance principle of 'fix the source, not the symptom.'

Exam trap

CompTIA often tests the principle that 'fix the source, not the symptom'—the trap here is that candidates jump to data cleansing actions (delete, fill, ignore) without first diagnosing why the nulls exist, which is a classic data quality management mistake.

How to eliminate wrong answers

Option A is wrong because deleting 1200 records with null emails reduces dataset size and may discard valid records if the nulls are due to a temporary system glitch, not actual missing data. Option B is wrong because filling null emails with a placeholder (e.g., 'unknown@domain.com') introduces false data that can skew analysis, violate email format constraints, and mislead downstream processes. Option D is wrong because ignoring nulls assumes they are non-critical without verification; in many contexts (e.g., customer communications, deduplication), missing emails are critical and can lead to incomplete insights or compliance issues.

863
Multi-Selectmedium

A data analyst wants to retrieve the top 5 highest-paid employees from the 'employees' table. Which SQL clauses could be used to achieve this? (Select TWO.)

Select 2 answers
A.ORDER BY salary DESC
B.HAVING salary
C.ORDER BY salary ASC
D.LIMIT 5
E.GROUP BY salary
AnswersA, D

ORDER BY salary DESC sorts the result set by the salary column in descending order, placing the highest-paid employees first. Combined with a limiting clause such as TOP 5 or FETCH FIRST 5 ROWS ONLY, it satisfies the stem's requirement to retrieve exactly the top five earners.

Why this answer

Option A, ORDER BY salary DESC, is correct because it sorts the result set by the salary column in descending order, placing the highest-paid employees first so the top earners can be identified. Option D, LIMIT 5, is correct because it restricts the result set to only the first 5 rows, which after the descending sort yields exactly the top 5 highest-paid employees. Together, ORDER BY salary DESC LIMIT 5 produces the desired result.

Option C, ORDER BY salary ASC, is wrong because ascending order returns the lowest-paid employees first, the opposite of what is needed. Option B, HAVING salary, is wrong because HAVING filters groups after aggregation and is not used to sort or limit rows. Option E, GROUP BY salary, is wrong because grouping by salary aggregates rows by distinct salary values rather than selecting the top 5 individual employees.

Exam trap

The trap here is confusing the roles of ORDER BY and LIMIT: candidates might think HAVING or GROUP BY can be used to get top N, or they might choose ascending order instead of descending.

864
Multi-Selecthard

A data analyst is preparing a presentation for the executive team to explain why quarterly revenue fell short of targets. They want to use storytelling with data. Which THREE elements should be included in the narrative arc? (Choose three.)

Select 3 answers
A.Raw data tables with every transaction
B.Resolution: the recommended action or outcome
C.Situation: background and context
D.Detailed explanation of data cleaning steps
E.Complication: the problem or challenge
AnswersB, C, E

Resolution closes the narrative arc by prescribing the action executives should take after seeing why revenue fell short, satisfying the storytelling-with-data requirement for a clear outcome. Without it, the presentation identifies the problem but offers no decision or next step, leaving the executive audience without direction.

Why this answer

The narrative arc for storytelling with data follows a three-part structure: Situation, Complication, and Resolution. Option C (Situation: background and context) is correct because it establishes the setting and baseline — here, the quarterly revenue targets and prior performance — so the executive audience understands the starting point. Option E (Complication: the problem or challenge) is correct because it introduces the tension, namely that revenue fell short of targets, which is the core issue the presentation must explain.

Option B (Resolution: the recommended action or outcome) is correct because it closes the arc by proposing what should be done next, giving executives a clear takeaway rather than leaving them with an unresolved problem. Option A (raw data tables with every transaction) is not part of the narrative arc; exhaustive transaction-level detail overwhelms an executive audience and belongs in an appendix, not the story. Option D (detailed explanation of data cleaning steps) is also excluded because data preparation methodology is a technical process detail, not a narrative element that advances the executive storyline.

Exam trap

DA0-002 often tests whether candidates can distinguish narrative elements from analytical methodology — the trap is selecting 'data cleaning steps' or 'raw data tables' because they seem thorough, when storytelling explicitly excludes them from the arc.

865
MCQmedium

A data analyst is designing a dashboard for executives. Which best practice should be followed?

A.Use 3D effects to make charts more engaging
B.Include every data point in the dashboard
C.Minimize clutter and use clear visual hierarchy
D.Use rainbow color palette to highlight all data points
AnswerC

Executive dashboards demand rapid comprehension, so minimising clutter and applying clear visual hierarchy directs attention to the most decision-relevant metrics first. This satisfies the stem's executive-audience constraint, where cognitive load must stay low and key trends must be immediately apparent.

Why this answer

Executive dashboards should minimize clutter and use clear visual hierarchy so decision-makers can grasp key metrics at a glance. Executives need high-level signals, not exhaustive detail, so prioritizing the most important KPIs with clean layout and consistent visual encoding is the recommended practice.

Exam trap

DA0-002 often tests whether candidates confuse 'engaging' visuals (3D, rainbow colors) with effective ones — the exam rewards clarity and accessibility over decoration.

How to eliminate wrong answers

Option A is wrong because 3D effects distort proportions and make values harder to compare accurately — they are widely discouraged in data visualization best practice. Option B is wrong because including every data point overwhelms the audience and buries the signal in noise; executives need curated, aggregated views. Option D is wrong because rainbow palettes lack semantic ordering and are inaccessible to colorblind viewers; they also imply false categories rather than highlighting meaning.

866
MCQmedium

A data analyst is profiling a new dataset containing customer information. When assessing data quality, which metric would be most appropriate to determine if the 'email' column contains valid email addresses?

A.Pattern analysis
B.Null count
C.Cardinality
D.Row count
AnswerA

Pattern analysis examines whether values conform to an expected format, such as the structure of a valid email address. Applying it to the 'email' column reveals entries that deviate from that pattern, directly satisfying the requirement to assess validity of email addresses.

Why this answer

Pattern analysis is the correct metric because it validates data against a defined format or regular expression — exactly what's needed to confirm that values in the 'email' column conform to the structure of a valid email address (e.g., user@domain.tld). Data profiling tools use pattern/format analysis to detect values that deviate from expected structures, making it the appropriate quality dimension for format validation.

Exam trap

The trap here is confusing data quality dimensions — candidates often pick cardinality or null count because they sound like 'profiling' metrics, but the question specifically asks about validating format, which only pattern analysis addresses.

How to eliminate wrong answers

Option B is wrong because null count only measures missing values and says nothing about whether populated values are structurally valid email addresses. Option C is wrong because cardinality measures the number of distinct values in a column, which is useful for detecting duplicates or uniqueness but not format validity. Option D is wrong because row count simply reports the total number of records and provides no information about the correctness or format of any individual value.

867
MCQhard

A data architect is designing a system for a subscription streaming service. The service must record every play, pause, and skip event from millions of concurrent viewers with very low write latency, and it must later support analytical queries over months of event history. The architect wants a single storage layer that handles both needs without a separate transformation pipeline. Which data architecture should the architect choose?

A.A message queue that retains events for a fixed retention period and serves analytical queries directly
B.A traditional data warehouse that ingests events only through nightly ETL batches
C.A lakehouse that combines open table formats with ACID transactions and query engines over the same storage
D.A data lake that stores raw event files and requires a separate batch job to load them into a warehouse
AnswerC

A lakehouse uses open table formats such as Delta Lake or Apache Iceberg to add ACID transactions and schema enforcement directly on object storage, so streaming writes and analytical reads share one layer. It removes the need for a separate transformation pipeline while supporting low-latency ingestion and historical queries over the same data.

Why this answer

A lakehouse unifies streaming ingestion and analytical querying on one storage layer by adding ACID transactions and schema management to open table formats on object storage. This satisfies the requirement for low-latency event writes plus months of queryable history without a separate transformation pipeline.

Exam trap

The trap here is treating a message queue as a storage layer for long-term analytics, when its retention limits mean historical queries still require another system.

868
MCQmedium

An analyst is presenting findings to a non-technical audience. The data shows a 20% increase in customer churn after a price change. Which presentation approach is BEST?

A.Explain the p-value
B.Provide the raw data table
C.Use a simple bar chart comparing churn before and after
D.Show a complex statistical model
AnswerC

A simple bar chart comparing churn before and after the price change translates the 20% increase into an immediate visual comparison, avoiding statistical jargon. This satisfies the non-technical audience constraint by making the magnitude and direction of change obvious at a glance.

Why this answer

A simple bar chart visually and intuitively communicates the 20% increase in churn to a non-technical audience without requiring statistical literacy. This approach aligns with best practices for presenting data insights to stakeholders who need clear, actionable takeaways rather than technical details.

Exam trap

The trap here is that candidates often overcomplicate the presentation by choosing technical options (like p-values or models) to demonstrate rigor, forgetting that the exam prioritizes audience-appropriate communication over statistical depth.

How to eliminate wrong answers

Option A is wrong because explaining a p-value introduces statistical significance testing, which is unnecessary and confusing for a non-technical audience that only needs to understand the magnitude of the change. Option B is wrong because providing the raw data table overwhelms the audience with numbers and fails to highlight the key insight (the 20% increase) effectively. Option D is wrong because showing a complex statistical model is inappropriate for a non-technical audience, as it obscures the simple before-and-after comparison and may lead to misinterpretation or disengagement.

869
MCQhard

A financial services firm classifies its datasets into public, internal, confidential, and restricted tiers. A data engineer requests access to a restricted customer transaction dataset to build a fraud model. Which combination of controls best enforces the governance policy for this request?

A.Grant permanent read access after the engineer signs an acceptable use policy acknowledgment
B.Apply role-based access with least privilege, time-bound approval, and query-level audit logging
C.Mask all customer identifiers in the dataset and allow self-service access to the masked copy
D.Require manager approval, then grant access through a role shared by the entire analytics department
AnswerB

This combination enforces the restricted tier through layered controls: role-based access scopes permissions to the fraud-modeling function, least privilege limits the data to what the task needs, time-bound approval prevents standing access, and audit logging provides attributable accountability. Together they satisfy both the protective and evidentiary expectations that governance policy places on the most sensitive classification.

Why this answer

Restricted data demands defense in depth: authorization tied to a legitimate purpose, permissions scoped to the individual, a defined expiry, and records that attribute every query. Role-based access with least privilege supplies the scoping, time-bound approval removes standing access after the project, and audit logging makes usage reviewable. Controls that only document intent, share broad roles, or open masked copies each leave an enforcement gap for the highest sensitivity tier.

Exam trap

The trap here is treating a single control, such as masking or a signed acknowledgment, as sufficient enforcement for restricted data when the classification actually requires layered, attributable access controls.

870
MCQhard

A data analyst is using the IQR method to identify outliers in a dataset. The first quartile (Q1) is 25 and the third quartile (Q3) is 45. What is the upper bound for identifying outliers?

A.85
B.65
C.75
D.55
AnswerC

The IQR is Q3 minus Q1, giving 20. The upper outlier bound is Q3 plus 1.5 times the IQR, so 45 + 30 = 75. This satisfies the stem's request for the upper bound using the IQR method.

Why this answer

The IQR method defines outliers as values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR. Here, Q1=25, Q3=45, so IQR = 45-25 = 20. The upper bound is Q3 + 1.5*IQR = 45 + 1.5*20 = 45 + 30 = 75.

Therefore, any value above 75 is considered an outlier.

Exam trap

DA0-002 often tests the IQR calculation. Candidates might forget the 1.5 multiplier or miscalculate the IQR, leading to wrong bounds. The trap is using the wrong formula or arithmetic error.

How to eliminate wrong answers

Option A is wrong because 85 is Q3 + 2*IQR (45+40), which is not the standard IQR method. Option B is wrong because 65 is Q3 + 1*IQR (45+20), which is not the correct multiplier. Option D is wrong because 55 is Q3 + 0.5*IQR (45+10), which is also incorrect.

871
Multi-Selecthard

A data analyst is designing a dashboard for hospital administrators to monitor emergency department (ED) wait times. The dashboard will be viewed on large wall-mounted displays in the operations center and must be readable from a distance. The analyst wants to ensure the dashboard effectively communicates urgent situations. Which two design choices are most appropriate? (Choose two.)

Select 2 answers
A.Use a red-yellow-green color scale for wait-time thresholds, with red indicating waits exceeding the target.
B.Use large, bold fonts for key metrics and ensure high contrast between text and background.
C.Add a live scrolling ticker at the bottom showing individual patient wait times to provide full transparency.
D.Include a detailed data table below each chart showing the last 24 hours of wait times in 15-minute increments.
E.Use a 3D pie chart to show the proportion of patients in each triage category, as it is visually engaging.
AnswersA, B

A red-yellow-green scale leverages universal associations (red for danger, green for good) to quickly convey urgency. In a wall-mounted display viewed from a distance, this supports rapid situational awareness. It is appropriate as long as thresholds are clearly defined and the design accounts for color-blind users by also using text or icons.

Why this answer

The red-yellow-green color scale and large, high-contrast fonts directly address the need for rapid, distance-readable communication of urgent wait times. The other options introduce clutter, privacy risks, or visual distortions that hinder quick comprehension. Effective operational dashboards prioritize clarity and immediate signal detection.

Exam trap

The trap here is focusing on visual appeal or data volume rather than the primary use case: quick, clear communication of urgent status from a distance.

872
Multi-Selecthard

A data analyst is performing data cleaning. Which THREE steps are part of this process? (Choose three.)

Select 3 answers
A.Correcting inconsistent data
B.Normalization
C.Handling missing values
D.Feature engineering
E.Removing duplicate records
AnswersA, C, E

Correcting inconsistent data resolves conflicting formats, units and values so records agree across sources. This is a core data cleaning activity, satisfying the stem's requirement to identify steps that standardise and reconcile raw data before analysis.

Why this answer

Data cleaning is the process of detecting and correcting (or removing) corrupt or inaccurate records from a dataset, so option A (Correcting inconsistent data) is correct because fixing mismatched formats, units, or values (e.g., 'NY' vs 'New York') is a core cleaning task. Option C (Handling missing values) is correct because cleaning must address nulls or blanks through imputation, deletion, or flagging so downstream analysis is not skewed. Option E (Removing duplicate records) is correct because duplicate rows inflate counts and distort aggregates, and deduplication is a standard cleaning step.

Option B (Normalization) is not part of cleaning; it is a data transformation/scaling technique (e.g., min-max or z-score) typically applied during preprocessing/modeling. Option D (Feature engineering) is also not cleaning; it is the creation of new derived variables for modeling, which occurs after cleaning.

Exam trap

The trap here is that candidates confuse data cleaning with data transformation or feature engineering, leading them to select normalization or feature engineering as cleaning steps, when in fact cleaning strictly addresses data quality issues like consistency, completeness, and uniqueness.

873
MCQhard

A multinational corporation is implementing a data governance framework. The legal team requires that data stored in the European Union not be transferred to the United States without adequate safeguards. The data engineering team uses a cloud-based data warehouse that replicates data across regions for performance. Which mechanism should the governance team prioritize to ensure compliance with cross-border data transfer requirements?

A.Using standard contractual clauses (SCCs) for all data transfers, including intra-company transfers to the US.
B.Relying on the EU-US Privacy Framework certification of the cloud provider to cover all transfers.
C.Encrypting all data with customer-managed keys so that US-based staff cannot decrypt EU data.
D.Implementing data residency controls that restrict storage and processing of EU personal data to EU regions only.
AnswerD

Data residency controls ensure that EU personal data remains within EU regions, preventing unauthorized cross-border transfers. This directly addresses the legal requirement and is a fundamental governance mechanism for complying with data transfer restrictions. It also simplifies compliance by avoiding the need for complex legal safeguards for transfers.

Why this answer

Data residency controls are the most direct and preventive mechanism because they restrict EU personal data to EU regions, eliminating unauthorized cross-border transfers at the architectural level. This aligns with the legal requirement and avoids reliance on complex legal safeguards that may not fully address automatic replication in a cloud data warehouse.

Exam trap

The trap here is assuming that legal frameworks like SCCs or the EU-US Privacy Framework alone are sufficient, when the cloud architecture may still replicate data across regions without technical residency controls.

874
Multi-Selectmedium

Which THREE of the following are characteristics of a relational database?

Select 3 answers
A.Enforces referential integrity through foreign keys
B.Stores data in key-value pairs
C.Supports NoSQL document storage
D.Uses Structured Query Language (SQL) for data manipulation
E.Data is organized into tables with rows and columns
AnswersA, D, E

Referential integrity ensures relationships.

Why this answer

Relational databases enforce referential integrity through foreign keys, which ensure that relationships between tables remain consistent. A foreign key in a child table must match a primary key value in the parent table, preventing orphaned records and maintaining data integrity.

Exam trap

The trap here is that candidates may confuse key-value stores or document databases with relational databases, especially when they hear terms like 'keys' or 'documents' in other contexts, but relational databases strictly use tables, rows, columns, and SQL.

875
Drag & Dropmedium

Drag and drop the steps to resolve data integration conflicts in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Conflict resolution starts with identification, analysis, standardization, transformation, and merging.

876
MCQmedium

A company uses an OLTP system for processing customer transactions. Which characteristic is most important for this system to ensure that each transaction is processed reliably, even if multiple users access the system simultaneously?

A.It uses a columnar storage format
B.It stores data in a denormalized schema
C.It supports complex analytical queries
D.It follows ACID properties
AnswerD

ACID properties guarantee atomicity, consistency, isolation and durability, so concurrent OLTP transactions commit reliably without partial writes or interference. Isolation specifically handles simultaneous users, while durability ensures committed transactions survive failure, matching the stem's reliability requirement.

Why this answer

ACID properties (Atomicity, Consistency, Isolation, Durability) are the foundational guarantees that ensure each transaction in an OLTP system is processed reliably, even when multiple users access the system concurrently. Atomicity ensures all-or-nothing execution, Consistency preserves database invariants, Isolation prevents concurrent transactions from interfering with each other, and Durability guarantees committed transactions survive failures. These properties directly address the requirement for reliable transaction processing under simultaneous access.

Exam trap

The trap here is confusing OLTP with OLAP characteristics: candidates might pick columnar storage or complex analytical queries because they associate databases with analytics, but the question specifically asks for reliability under concurrent access, which points to ACID properties.

How to eliminate wrong answers

Option A is wrong because columnar storage is optimized for analytical workloads (OLAP) that scan large volumes of data, not for high-concurrency transactional processing; it typically degrades write performance and does not provide transactional guarantees. Option B is wrong because a denormalized schema is used in data warehousing to reduce joins and improve read performance for analytics, but it can introduce data redundancy and update anomalies that undermine transactional consistency. Option C is wrong because supporting complex analytical queries is a characteristic of OLAP systems, not OLTP; OLTP systems prioritize fast, simple transactions and would suffer performance degradation if burdened with complex analytical queries.

877
MCQmedium

A data analyst wants to visualize the distribution of employee salaries across departments and identify any outliers. Which chart type would best show quartiles, median, and potential outliers?

A.Heat map
B.Histogram
C.Box plot
D.Scatter plot
AnswerC

A box plot encodes the median, lower and upper quartiles, and whiskers, with points beyond the whiskers flagged as outliers. That five-number summary directly satisfies the requirement to show salary distribution spread and identify extreme values across departments.

Why this answer

A box plot is purpose-built to display the five-number summary — minimum, first quartile (Q1), median, third quartile (Q3), and maximum — with whiskers extending to 1.5×IQR and individual points plotted beyond that as outliers. This makes it the ideal chart for comparing salary distributions across departments while simultaneously surfacing outliers. No other listed chart encodes quartiles and outliers in a single view.

Exam trap

DA0-002 often tests the confusion between histogram and box plot — candidates pick histogram because it 'shows distribution,' but only the box plot explicitly encodes quartiles, median, and outliers in one view.

How to eliminate wrong answers

Option A is wrong because a heat map encodes values as color intensity across a matrix (e.g., correlation or density), and does not display quartiles, median, or outliers for a continuous variable. Option B is wrong because a histogram shows the frequency distribution of a single variable in bins, revealing shape and modality but not quartiles or labeled outliers directly. Option D is wrong because a scatter plot shows the relationship between two continuous variables, not the distributional summary of one variable across categories.

878
Multi-Selectmedium

A data analyst is designing a dashboard to monitor real-time website traffic. The dashboard will be used by the operations team to quickly identify anomalies and take immediate action. Which two design elements are most critical to include? (Choose two.)

Select 2 answers
A.Clear visual alerts (e.g., color-coded thresholds) for metrics exceeding normal ranges.
B.A detailed data table showing all raw traffic data for the current day.
C.A historical trend line showing traffic over the past year for context.
D.Drill-down capabilities to explore individual user sessions.
E.Automatic refresh with a timestamp indicating the last update.
AnswersA, E

Color-coded alerts or thresholds enable the operations team to instantly identify anomalies without having to interpret raw numbers. This supports rapid response, which is essential for real-time monitoring. Visual cues like red for critical, yellow for warning, and green for normal are preattentive and reduce cognitive load. This is a critical design element for anomaly detection.

Why this answer

The two most critical elements are automatic refresh with a timestamp and clear visual alerts for anomalies. Real-time monitoring demands that data is current and that deviations from normal are immediately visible. A timestamp builds trust in data freshness, while color-coded thresholds enable split-second recognition of issues.

Together, they empower the operations team to act quickly and confidently.

Exam trap

The trap here is assuming that more data or interactivity is always better, when in fact real-time dashboards should prioritize immediate, actionable signals over detailed exploration.

879
Multi-Selectmedium

A data analyst is preparing a report to present to a mixed audience of technical and non-technical stakeholders. Which THREE techniques should the analyst use to ensure effective communication? (Choose three.)

Select 3 answers
A.Tailor the narrative to address different concerns
B.Use only one chart type for consistency
C.Use technical jargon to demonstrate expertise
D.Provide high-level summaries for non-technical audience
E.Include detailed technical appendices for those interested
AnswersA, D, E

Tailoring the narrative lets the analyst connect findings to what each group cares about, so technical rigour and business impact are both addressed. This satisfies the mixed-audience constraint by preventing the single generic message that would lose either the specialists or the non-specialists.

Why this answer

Option A is correct because tailoring the narrative to address different concerns lets the analyst speak to both technical and non-technical stakeholders in terms they care about, keeping the presentation relevant and engaging for everyone. Option D is correct because providing high-level summaries gives non-technical audience members an accessible overview of the key findings without requiring them to parse technical details. Option E is correct because including detailed technical appendices satisfies the technical stakeholders who want deeper evidence and methodology, without cluttering the main presentation for others.

Option B does not belong because using only one chart type for consistency can obscure data that is better represented by a different visualization, reducing clarity. Option C does not belong because using technical jargon to demonstrate expertise alienates non-technical stakeholders and undermines effective communication rather than supporting it.

Exam trap

The trap here is that candidates often confuse 'consistency' with 'clarity,' mistakenly believing that using a single chart type (Option B) simplifies the message, when in fact it can hide critical patterns that require different visual encodings.

880
MCQhard

An analyst creates a scatter plot with three variables: X, Y, and a third variable represented by the size of the markers. This chart is called a:

A.Heat map
B.Treemap
C.Bubble chart
D.Waterfall chart
AnswerC

A bubble chart extends the scatter plot by encoding a third variable as marker size, so each point conveys three dimensions simultaneously. X and Y position the data point while the bubble's magnitude represents the additional measure, matching the described visual encoding exactly.

Why this answer

A bubble chart is a variation of a scatter plot where a third numeric variable is encoded by the size (area) of the markers. This allows the visualization of three dimensions of data simultaneously on a two-dimensional plane, making it the correct choice for the described chart.

Exam trap

The trap here is that candidates confuse a bubble chart with a heat map because both can represent three variables, but the heat map uses color gradients on a grid, not marker size on a scatter plot.

How to eliminate wrong answers

Option A is wrong because a heat map uses color intensity to represent the magnitude of a third variable across two categorical axes, not marker size. Option B is wrong because a treemap uses nested rectangles (tiles) to display hierarchical data, with area encoding a quantitative value, and does not use X/Y coordinates or marker sizes. Option D is wrong because a waterfall chart shows cumulative effects of sequential positive or negative values, typically in a financial context, and does not involve scatter plot markers or a third variable encoded by size.

881
MCQeasy

Which statistical test should be used to determine if there is a significant association between two categorical variables, such as gender and product preference?

A.ANOVA
B.Chi-square test
C.Pearson correlation
D.t-test
AnswerB

The chi-square test of independence compares observed versus expected frequencies across a contingency table, determining whether two categorical variables are associated. Gender and product preference are both nominal categories, so this test satisfies the requirement for assessing significant association between them.

Why this answer

The chi-square test is specifically designed to test for independence between two categorical variables, such as gender (male/female) and product preference (e.g., product A/B/C). It compares observed frequencies in each category combination to the frequencies expected if the variables were independent, using the chi-square statistic. A significant result indicates that the variables are associated, not independent.

Exam trap

The trap here is confusing tests for means (t-test, ANOVA) with tests for association between categorical variables, or mixing up correlation (for continuous variables) with chi-square (for categorical variables).

How to eliminate wrong answers

Option A is wrong because ANOVA compares means of a continuous dependent variable across three or more groups defined by a categorical independent variable, not associations between two categorical variables. Option C is wrong because Pearson correlation measures the linear relationship between two continuous variables, not categorical ones. Option D is wrong because a t-test compares means between two groups (or one group against a known value) for a continuous outcome, not categorical associations.

882
MCQmedium

A data analyst uses the following query: SELECT department, AVG(salary) AS avg_salary FROM employees GROUP BY department HAVING AVG(salary) > 50000. What is the purpose of the HAVING clause in this query?

A.To ensure only departments with more than 50000 employees are shown
B.To filter individual employee records before grouping
C.To sort departments by average salary
D.To filter groups (departments) based on the average salary
AnswerD

HAVING filters aggregated results after GROUP BY executes, unlike WHERE which filters rows beforehand. Here it discards department groups whose computed average salary does not exceed 50000, restricting output to qualifying groups based on the aggregate.

Why this answer

The HAVING clause filters groups after aggregation, so in this query it keeps only departments whose average salary exceeds 50000. It operates on the result of GROUP BY, which is why it can reference aggregate functions like AVG(salary) that WHERE cannot.

Exam trap

DA0-002 often tests the WHERE vs HAVING distinction — candidates pick 'filter individual employee records before grouping' because they conflate the two clauses, missing that HAVING operates on aggregated groups and is the only clause that can filter on aggregate results.

How to eliminate wrong answers

Option A is wrong because HAVING AVG(salary) > 50000 compares the average salary value, not a count of employees — to filter by employee count you would write HAVING COUNT(*) > 50000. Option B is wrong because filtering individual rows before grouping is the job of the WHERE clause, which runs before GROUP BY; HAVING runs after aggregation. Option C is wrong because sorting is done with ORDER BY, not HAVING — HAVING only includes or excludes groups, it does not order them.

883
Multi-Selecthard

After merging two datasets, an analyst finds that the resulting dataset has many null values in some columns. Which TWO steps should the analyst take to address this? (Select two.)

Select 2 answers
A.Ignore nulls and proceed.
B.Impute nulls with the median.
C.Remove all rows with nulls.
D.Replace nulls with a placeholder value like 'Unknown'.
E.Investigate the cause of nulls.
AnswersB, E

Median imputation fills missing numeric values with the column's central value, which is robust to outliers and preserves the row count needed for analysis. It satisfies the requirement to handle nulls introduced by the merge without discarding records, though it reduces variance and should follow cause investigation.

Why this answer

Option B is correct because imputing nulls with the median is a robust statistical technique that fills missing numeric values with the column's central tendency, reducing data loss while limiting the influence of outliers compared to the mean. Option E is correct because investigating the cause of nulls is an essential diagnostic step: nulls introduced by a dataset merge often indicate unmatched join keys, schema mismatches, or missing source records, and understanding the root cause determines whether imputation, key correction, or source remediation is appropriate. Options A, C, and D are not among the correct answers: ignoring nulls (A) can bias analysis and break models, removing all rows with nulls (C) can discard substantial valid data and skew distributions, and replacing nulls with a placeholder like 'Unknown' (D) is only suitable for categorical data and can corrupt numeric columns or mislead downstream processing.

Exam trap

The trap here is that candidates may think 'Ignore nulls and proceed' is acceptable, but the exam tests the understanding that nulls must be actively handled to ensure data quality and model validity, not simply overlooked.

884
MCQhard

A table Orders has OrderID (primary key), CustomerID, and CustomerEmail. During analysis, it is found that CustomerID uniquely identifies CustomerEmail. Which normal form is violated if both CustomerID and CustomerEmail are stored in this table?

A.Second normal form (2NF)
B.Third normal form (3NF)
C.No violation
D.First normal form (1NF)
AnswerB

CustomerEmail depends on CustomerID, which is a non-key attribute, creating a transitive dependency violating 3NF.

Why this answer

The table violates Third Normal Form (3NF) because CustomerEmail is transitively dependent on CustomerID, which is not a candidate key. In 3NF, every non-key attribute must depend only on the primary key (OrderID), not on another non-key attribute. Since CustomerID uniquely identifies CustomerEmail, CustomerEmail depends on CustomerID, not directly on OrderID, creating a transitive dependency.

Exam trap

The trap here is that candidates often confuse transitive dependencies with partial dependencies, mistakenly thinking that because CustomerID is not part of the primary key, the violation is 2NF rather than 3NF.

How to eliminate wrong answers

Option A is wrong because Second Normal Form (2NF) requires that all non-key attributes are fully functionally dependent on the entire primary key; here, the primary key is a single column (OrderID), so there is no partial dependency, and 2NF is satisfied. Option C is wrong because a violation does exist — the transitive dependency between CustomerID and CustomerEmail breaks 3NF. Option D is wrong because First Normal Form (1NF) is not violated; the table has atomic values and a primary key, so it meets 1NF requirements.

885
MCQeasy

A data analyst runs the following query: SELECT DISTINCT city FROM customers. What is the primary purpose of using the DISTINCT keyword in this query?

A.To sort the cities alphabetically
B.To count the number of cities
C.To filter cities that start with a specific letter
D.To remove duplicate city names
AnswerD

DISTINCT collapses duplicate rows in the result set, so each city name appears once. Applied to the single selected column, it returns the unique list of cities, eliminating repeated values that would otherwise appear for every matching customer record.

Why this answer

The DISTINCT keyword in a SELECT statement eliminates duplicate rows from the result set. In 'SELECT DISTINCT city FROM customers', if multiple customers live in the same city, that city name appears only once in the output, producing a list of unique city values. DISTINCT operates on the entire row of selected columns, so if multiple columns are selected, the combination must be unique to be retained.

Exam trap

The trap is confusing DISTINCT with ORDER BY or COUNT, causing candidates to select an answer that describes sorting or counting rather than deduplication.

How to eliminate wrong answers

Option A is wrong because sorting is performed by the ORDER BY clause, not DISTINCT — DISTINCT does not guarantee any particular order, and the result set order is undefined without an explicit ORDER BY. Option B is wrong because counting rows requires the COUNT() aggregate function (e.g., 'SELECT COUNT(DISTINCT city) FROM customers'), not the DISTINCT keyword alone. Option C is wrong because filtering rows based on a condition requires the WHERE clause (e.g., 'WHERE city LIKE 'A%''), not DISTINCT, which only removes duplicates after the rows are selected.

886
MCQmedium

A dashboard designer is creating a KPI dashboard for executives. Which of the following is a leading indicator?

A.Number of qualified leads
B.Net profit margin
C.Monthly revenue
D.Customer churn rate
AnswerA

Qualified leads measure pipeline activity that precedes future revenue, making them a leading indicator executives can act on. Lagging indicators such as closed revenue or profit report outcomes already realised, so they cannot forecast upcoming performance.

Why this answer

Number of qualified leads is a leading indicator as it predicts future sales, while revenue and customer churn are lagging indicators.

887
MCQhard

A data scientist runs a linear regression model to predict customer spending based on income. The R-squared value is 0.45 and the p-value for the slope coefficient is 0.03. At a significance level of α=0.05, which of the following conclusions is correct?

A.The slope is not statistically significant, and the model explains 55% of the variance.
B.The slope is statistically significant, and the model explains 45% of the variance.
C.The slope is statistically significant, and the model explains 55% of the variance.
D.The slope is not statistically significant, and the model explains 45% of the variance.
AnswerB

With α=0.05, the slope p-value of 0.03 falls below the threshold, so the slope is statistically significant. R-squared of 0.45 means income explains 45% of the variance in spending. Both conditions in the stem are satisfied simultaneously.

Why this answer

The p-value (0.03) is less than α (0.05), so the slope is statistically significant. R²=0.45 means the model explains 45% of the variance.

888
Multi-Selectmedium

Which TWO of the following are examples of unstructured data? (Select 2)

Select 2 answers
A.MP4 video
B.CSV file
C.XML file
D.JPEG image
E.JSON document
AnswersA, D

Video files are unstructured.

Why this answer

A is correct because MP4 video files contain binary data that lacks a predefined schema or tabular structure, making them a classic example of unstructured data. Unlike structured data, MP4 files store audiovisual content in a container format that cannot be easily queried or analyzed without specialized processing.

Exam trap

The trap here is that candidates often confuse semi-structured data (XML, JSON, CSV) with unstructured data, forgetting that semi-structured data still has a defined schema or metadata, unlike raw binary or free-form text.

889
MCQmedium

A data analyst at a subscription streaming service is designing a dashboard for the customer success team. The team needs to monitor, side by side, the current month's churn rate against the same month last year and against a target churn rate of 4.5%. The analyst wants viewers to instantly see whether the current rate is above or below target. Which design approach best supports this?

A.A detailed table listing monthly churn rates for the last 24 months with conditional formatting on the target column
B.A KPI card showing the current churn rate, the prior-year value, and the target, with a conditional indicator that changes when the target is exceeded
C.A gauge chart showing the current churn rate with a needle and colored zones for acceptable and unacceptable ranges
D.A stacked area chart showing churn rate components over the last 12 months
AnswerB

A KPI card presents the current value alongside a comparison and a target, and a conditional indicator turns the target comparison into an instant visual signal. This directly answers whether churn is above or below 4.5% while keeping the prior-year context visible, matching the customer success team's monitoring need efficiently.

Why this answer

A KPI card with the current churn rate, the prior-year comparison, and the target, plus a conditional indicator, packages exactly the three values the team needs and converts the target comparison into an immediate visual cue. Gauges, tables, and stacked area charts either hide the target relationship or require manual interpretation, slowing the team's ability to act when churn drifts above 4.5%.

Exam trap

The trap here is equating a visually elaborate chart, such as a gauge, with a more effective at-a-glance indicator than a simple labeled KPI card.

890
Multi-Selecthard

Which THREE are considered best practices in dashboard design? (Select three.)

Select 3 answers
A.Using heat maps to visualize correlation
B.Using 3D charts to add depth
C.Maximizing the data-ink ratio
D.Providing interactive filters for exploration
E.Including every data point in the dashboard
AnswersA, C, D

Heat maps encode correlation coefficients as colour intensity across a matrix of variable pairs, letting viewers spot strong positive or negative relationships instantly. This satisfies dashboard design best practice by compressing dense statistical relationships into a single glanceable visual.

Why this answer

Option A is correct because heat maps encode values through color intensity across a matrix, making correlation and density patterns between two dimensions immediately visible without requiring users to read individual numbers. Option C is correct because maximizing the data-ink ratio, a principle from Edward Tufte, means removing non-data decoration so that most pixels convey actual information, which improves clarity and reduces cognitive load. Option D is correct because interactive filters let users drill into relevant subsets and explore the data themselves, turning a static report into an analytical tool that answers multiple questions.

Option B is not a best practice because 3D charts distort proportions and depth perception, making values harder to compare accurately than their 2D equivalents. Option E is not a best practice because cramming every data point into a dashboard creates clutter and obscures the key signals, whereas dashboards should surface aggregated or prioritized metrics.

Exam trap

CompTIA often tests the misconception that adding visual flair (like 3D effects) or exhaustive data improves a dashboard, when in reality these choices degrade readability and violate core principles of effective data visualization.

891
MCQeasy

A data analyst wants to compare the revenue across five different product categories. Which chart type is best suited?

A.Scatter plot
B.Pie chart
C.Line chart
D.Bar chart
AnswerD

A bar chart encodes each category as a separate bar on a categorical axis, letting viewers compare five discrete revenue values accurately through position and length. This suits nominal comparison better than pie or line charts, which obscure precise differences between individual categories.

Why this answer

A bar chart is the best choice for comparing a categorical variable (five product categories) against a quantitative measure (revenue). Each category gets its own bar, making relative magnitudes easy to compare side by side. Bar charts handle a small number of discrete categories cleanly, which matches this scenario exactly.

Exam trap

The trap here is choosing a pie chart because 'comparing categories' sounds like parts of a whole; CompTIA expects you to recognize that bar charts are superior for magnitude comparison across discrete categories.

How to eliminate wrong answers

Option A is wrong because a scatter plot displays the relationship between two continuous variables, not comparisons across discrete categories. Option B is wrong because a pie chart shows parts of a whole as proportions and becomes hard to read with five categories; it is also poor for precise comparisons of revenue magnitudes. Option C is wrong because a line chart is designed for trends over a continuous dimension such as time, not for comparing discrete categories.

892
MCQhard

A data analyst is using Power BI to create a report that shows sales by region. The data includes duplicate rows for some transactions due to a data entry error. The analyst needs to count only unique transactions. Which DAX function should be used to create a measure for unique count?

A.COUNT
B.DISTINCTCOUNT
C.COUNTROWS
D.SUMX
AnswerB

DISTINCTCOUNT counts each unique value in a column once, so duplicate transaction rows collapse to a single count. This directly satisfies the requirement to count only unique transactions, unlike COUNTROWS, which would tally every row including the erroneous duplicates.

Why this answer

DISTINCTCOUNT is the DAX function specifically designed to count the number of distinct (unique) values in a column, ignoring duplicates. In this scenario, the analyst needs to count unique transactions despite duplicate rows, so DISTINCTCOUNT on the transaction ID column will return the correct count.

Exam trap

DA0-002 often tests the confusion between COUNT/COUNTROWS (which include duplicates) and DISTINCTCOUNT (which excludes duplicates), leading candidates to pick a function that overcounts.

How to eliminate wrong answers

Option A is wrong because COUNT counts all non-blank values, including duplicates, so it would overcount. Option C is wrong because COUNTROWS counts all rows in a table, including duplicates, which would also overcount. Option D is wrong because SUMX is an iterator that sums an expression over a table, not a counting function; it is used for calculations like total sales, not unique counts.

893
MCQeasy

A data analyst is building a linear regression model to predict sales based on advertising spend. The analyst notices that the residuals are not normally distributed and have a non‑constant variance. Which of the following transformations is most appropriate to apply to the dependent variable?

A.Standardization (z-score)
B.Normalization (min-max scaling)
C.Logarithmic transformation
D.Square root transformation
AnswerC

Non-constant variance and non-normal residuals violate linear regression assumptions. A logarithmic transformation of the dependent variable compresses the scale of large values, stabilising variance and pulling the residual distribution towards normality, which is the standard remedy for this pattern of heteroscedasticity.

Why this answer

The logarithmic transformation is the most appropriate choice because it stabilizes non‑constant variance (heteroscedasticity) and helps make the residuals more normally distributed, which are key assumptions for linear regression. By compressing the scale of the dependent variable (sales), it reduces the impact of large values and often linearizes multiplicative relationships, such as diminishing returns from advertising spend.

Exam trap

CompTIA often tests the misconception that any scaling technique (standardization or normalization) can fix heteroscedasticity or non‑normality, but these methods only change the range or center of the data, not the shape of the residual distribution or the variance structure.

How to eliminate wrong answers

Option A is wrong because standardization (z-score) centers and scales the data to mean 0 and standard deviation 1, but it does not address heteroscedasticity or non‑normal residuals; it merely changes the units of the dependent variable without altering the shape of the distribution. Option B is wrong because normalization (min-max scaling) rescales the data to a fixed range (e.g., 0 to 1), which also fails to correct non‑constant variance or non‑normality; it is primarily used for feature scaling in algorithms like neural networks, not for satisfying regression assumptions. Option D is wrong because the square root transformation is typically used for count data (e.g., Poisson-distributed outcomes) to stabilize variance, but it is less effective than the log transformation when the variance increases proportionally with the mean, which is common in sales data; the log transformation is the standard choice for multiplicative relationships and heteroscedasticity.

894
MCQeasy

A data architect needs to store raw data from various sources, including social media feeds and log files, for future analysis. The data may be used for machine learning and ad-hoc queries. Which storage solution is most appropriate for storing raw data in its native format?

A.Data lake
B.Data mart
C.Relational database
D.Data warehouse
AnswerA

A data lake stores raw data in its native format without schema enforcement, accommodating social media feeds and log files. This satisfies the requirement for future machine learning and ad-hoc queries, unlike warehouses that demand predefined structure.

Why this answer

A data lake is designed to store raw data in its native format, including unstructured and semi-structured data from sources like social media feeds and log files. It supports schema-on-read, making it ideal for future machine learning and ad-hoc queries without requiring upfront transformation. This aligns directly with the requirement to preserve raw data for flexible analysis.

Exam trap

The trap here is that candidates confuse a data lake with a data warehouse, assuming both are for analytics, but the key distinction is that a data warehouse requires structured, transformed data while a data lake preserves raw, native-format data.

How to eliminate wrong answers

Option B is wrong because a data mart is a subset of a data warehouse optimized for a specific business domain, not for storing raw, diverse data in native format. Option C is wrong because a relational database enforces a rigid schema and ACID constraints, making it unsuitable for unstructured data like social media feeds and log files. Option D is wrong because a data warehouse stores processed, structured data optimized for reporting and BI, not raw data in its native format.

895
Multi-Selecthard

A national retailer is consolidating data from 40 regional stores into a central analytics platform. Each region uses different codes for the same product categories, and store managers report sales in local currencies. Before loading the data, the integration team must resolve these inconsistencies. Which two activities are appropriate steps to standardize the data? (Choose two.)

Select 2 answers
A.Convert all monetary values to a single reporting currency using a defined exchange rate table.
B.Build a mapping table that translates each region's product category codes into a single enterprise code set.
C.Delete records from regions whose currency differs from the headquarters currency.
D.Store each region's data in a separate database and report from each database independently.
E.Allow each region to keep its own category codes and currencies, and resolve differences during reporting.
AnswersA, B

Converting local currency amounts to one reporting currency using a governed exchange rate table makes financial figures directly comparable across regions. The rate table documents which rate applies to which period, preserving auditability. This standardization step addresses the unit inconsistency in monetary values and is essential before aggregating sales at the enterprise level.

Why this answer

Standardizing inconsistent data requires resolving both the category code mismatch and the currency unit mismatch. A crosswalk mapping table unifies product categories under one enterprise code set, while conversion using a governed exchange rate table unifies monetary values into one reporting currency. Deleting records, isolating databases, or deferring translation to report time all leave the inconsistencies unresolved or shift the burden downstream.

Exam trap

The trap here is thinking that consolidation only means moving data into one place, when it also requires resolving semantic and unit mismatches during transformation.

896
MCQmedium

A retail company is integrating sales data from three regional databases into a central data warehouse. The 'product_id' column is defined as an integer in two databases but as a variable-length string in the third. During the ETL process, the analyst must ensure that product_id values are consistent for joining with the product dimension table. Which data transformation should the analyst perform?

A.Apply data normalization to product_id
B.Convert all product_id values to string
C.Convert all product_id values to integer
D.Use data aggregation on product_id
AnswerB

Converting all product_id values to a string data type ensures consistency and avoids loss of leading zeros or non-numeric characters. String representation is the most flexible for joining across sources, as it accommodates numeric and alphanumeric identifiers. This transformation aligns the data types for the join operation.

Why this answer

The product_id column must have a consistent data type across all sources to perform joins. Converting all values to string is the safest choice because it preserves leading zeros and any alphanumeric characters. Integer conversion could fail or lose information if the string column contains non-numeric values.

String conversion ensures compatibility without data loss.

Exam trap

The trap here is assuming that numeric identifiers should always be stored as integers, ignoring the possibility of alphanumeric or leading-zero values.

897
MCQeasy

A data analyst is tasked with collecting data from multiple spreadsheets provided by different departments. Each spreadsheet has different column names and formats. What is the best first step?

A.Develop a data dictionary and standardize column names
B.Discard any mismatched data
C.Use a machine learning model to clean data
D.Immediately load all data into a database
AnswerA

Differing column names and formats across departmental spreadsheets create schema conflicts that must be resolved before any merging or analysis. A data dictionary documents each source field and its meaning, enabling analysts to map and standardise names consistently, satisfying the stem's requirement for a first step addressing heterogeneous source structures.

Why this answer

Developing a data dictionary and standardizing column names ensures consistency across all data sources before loading, reducing errors and facilitating integration. Immediately loading data can cause inconsistencies. Discarding mismatched data loses potentially valuable information.

Using a machine learning model is an unnecessary and complex first step.

898
Multi-Selectmedium

A data analyst is preparing a dashboard that will be embedded in a public-facing web page and refreshed nightly. The source system contains customer names and account numbers, but the dashboard only needs aggregated counts by region and product line. Which two practices should the analyst apply to reduce disclosure risk while keeping the dashboard functional? (Choose two.)

Select 2 answers
A.Restrict the dashboard to authenticated internal users so the public page shows a login prompt instead of the report.
B.Embed the full source table in the page and rely on the chart configuration to display only aggregated values.
C.Obscure the identifiers by applying a simple substitution cipher to the name and account number columns before publishing.
D.Suppress or combine aggregate cells whose counts fall below a minimum group size threshold.
E.Remove the customer name and account number fields from the dataset feeding the dashboard, since only aggregates are required.
AnswersD, E

Small-count cells are the classic re-identification vector: a region and product line with one or two customers can effectively name those individuals even without a name column. Enforcing a minimum group size, or collapsing sparse cells into an 'other' bucket, preserves the analytic value of large groups while removing the disclosure risk. This complements field removal because it protects against inference from the aggregates themselves.

Why this answer

Minimizing the dataset to only the fields the dashboard consumes removes direct identifiers at the source, and enforcing a minimum group size on aggregate cells blocks the inference route where a tiny count can single out an individual. Together they let regional and product-line counts be published on a public page while removing the two most common re-identification paths.

Exam trap

The trap here is assuming that hiding identifier columns in the chart configuration is equivalent to removing them from the data, when the fields still travel to the client.

899
MCQeasy

A data analyst notices that customer addresses in the database contain invalid ZIP codes. Which data quality dimension is being violated?

A.Validity
B.Timeliness
C.Consistency
D.Completeness
AnswerA

Validity checks whether values conform to defined formats and permissible ranges. A ZIP code that does not match the required pattern breaches that rule, so the dimension violated is validity rather than accuracy, completeness or consistency.

Why this answer

A is correct because validity refers to the degree to which data conforms to its defined format, rules, or constraints. Invalid ZIP codes (e.g., a five-digit code containing letters or a non-existent postal code) directly violate the format and domain rules expected for that field, making this a validity issue.

Exam trap

The trap here is that candidates confuse 'validity' with 'completeness' or 'consistency,' mistakenly thinking a missing or mismatched ZIP code is a completeness or consistency issue, when in fact the violation is about the data not conforming to the required format or rule set.

How to eliminate wrong answers

Option B (Timeliness) is wrong because timeliness concerns whether data is available when needed, not whether individual values match expected formats. Option C (Consistency) is wrong because consistency checks for logical coherence across related data sets or fields (e.g., ZIP code matching city/state), not the intrinsic correctness of a single value. Option D (Completeness) is wrong because completeness measures whether all required data is present (e.g., missing ZIP codes), not whether present data is correctly formatted.

900
MCQmedium

A data analyst uses a CTE to find employees who earn more than the average salary in their department. Which SQL clause is used to define the CTE?

A.DECLARE
B.WITH
C.DEFINE
D.CTE
AnswerB

The WITH clause introduces a named common table expression, letting the analyst define the per-department average salary subquery once and reference it in the main SELECT. This satisfies the stem's requirement for computing employees earning above their department's average without repeating the aggregate logic.

Why this answer

The WITH clause is the standard SQL syntax that introduces a Common Table Expression (CTE), allowing the analyst to define a named temporary result set (e.g., WITH dept_avg AS (SELECT dept, AVG(salary) ...)) that can then be referenced in the main query. This is defined in the SQL standard and supported by PostgreSQL, SQL Server, Oracle, and MySQL 8+. The CTE exists only for the duration of the single statement that follows it.

Exam trap

The trap here is confusing the conceptual name 'CTE' with the actual SQL keyword — candidates who know the term but not the syntax may select 'CTE' or 'DEFINE' instead of the correct WITH clause.

How to eliminate wrong answers

Option A is wrong because DECLARE is used in T-SQL/PL-SQL to declare variables, cursors, or table variables — not to define a CTE. Option C is wrong because DEFINE is not a SQL keyword for query result sets; it appears in other contexts (e.g., Snowflake's DEFINE or Oracle SQL*Plus) but never introduces a CTE. Option D is wrong because CTE is a conceptual term for 'Common Table Expression,' not an actual SQL keyword — no SQL dialect uses 'CTE' as a clause.

Page 11

Page 12 of 14

Page 13