CompTIA · Free Practice Questions · Last reviewed May 2026
30real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
22% of exam · 6 sample questions below
A data analyst is pulling data from a production database for a report. The database contains customer orders with a column 'order_date'. The analyst notices that some orders have dates in the future. Which data quality issue does this represent?
Invalid data type
Inconsistent data
Missing data
Violation of business rules
Future orders are not valid per business rules, indicating a data quality issue.
A data engineer is designing a data pipeline to ingest streaming data from IoT sensors. The sensors send data every second, and the pipeline must handle bursts of up to 10,000 messages per second. Which approach is most appropriate for capturing this data before processing?
Directly write each message to a relational database
Load directly into a data warehouse
Use a message queue to buffer the incoming data
A message queue decouples producers from consumers, buffering bursts of up to 10,000 messages per second so the ingestion tier is not overwhelmed. This satisfies the stem's burst-handling constraint by absorbing spikes and letting downstream processing drain at its own rate.
Store data in flat files and process in nightly batches
A data analyst needs to combine two datasets: one contains customer information (customer_id, name, address) and the other contains order information (order_id, customer_id, order_date). The analyst wants to include all customers, even those who have not placed orders. Which type of join should be used?
FULL OUTER JOIN
INNER JOIN
LEFT JOIN
A LEFT JOIN returns every row from the customer table plus matching order rows, emitting NULLs where no orders exist. This satisfies the stem's requirement to retain customers without orders, unlike an INNER JOIN which would exclude them.
RIGHT JOIN
A data analyst is tasked with extracting data from a legacy system that outputs fixed-width text files. The analyst needs to parse these files into a structured format. Which tool or method is most appropriate for this task?
A spreadsheet application
An ETL tool with a graphical interface
A scripting language such as Python
A scripting language such as Python parses fixed-width files by slicing each line at defined column positions, handling the absence of delimiters. This directly addresses the fixed-width constraint, unlike delimiter-based tools that cannot infer field boundaries.
SQL
A company is merging two databases from different departments. In Database A, customer IDs are integers. In Database B, customer IDs are alphanumeric strings. To merge, the data analyst must reconcile these differences. Which step should be taken first?
Drop the ID column and use a surrogate key
Convert all IDs to integers using CAST
Perform data profiling to understand the ID formats and relationships
Profiling first reveals each database's ID format, value patterns and relationships before any mapping or conversion. Reconciling integer versus alphanumeric IDs requires knowing actual contents and referential links, so profiling satisfies the stem's requirement to understand formats and relationships.
Create a mapping table based on the first character
A data analyst is building a dataset from multiple sources and needs to ensure data quality. During the data acquisition phase, which activity is most important to perform?
Data visualization
Data cleaning
Data profiling
Profiling computes null counts, distinct values, ranges and formats across each source column, exposing anomalies before data enters the warehouse. This satisfies the data quality requirement by revealing inconsistencies at acquisition time, when they are cheapest to correct.
Data modeling
Want more Data Acquisition and Preparation practice?
Practice this domain20% of exam · 6 sample questions below
A retail company stores customer purchase history in a relational database. The database contains a table 'transactions' with columns: transaction_id, customer_id, product_id, quantity, price, and transaction_date. A data analyst needs to create a report that shows total revenue per customer for the last quarter. Which data concept describes the relationship between customer_id and total revenue?
Foreign key
Composite attribute
Derived attribute
Total revenue is not stored but computed by aggregating quantity multiplied by price grouped by customer_id. Because it is calculated from existing stored columns rather than persisted, it is a derived attribute, satisfying the report's need for per-customer revenue.
Atomic attribute
A healthcare database stores patient records. Each patient has a unique patient_id, and the database includes a table 'visits' with visit_id, patient_id, visit_date, and diagnosis_code. To ensure data integrity, which constraint should be applied to the patient_id column in the 'visits' table?
Unique constraint
Foreign key
A foreign key on patient_id in the visits table references the primary key in the patients table, ensuring every visit maps to an existing patient. This enforces referential integrity, preventing orphan visit records, which is the integrity requirement stated in the stem.
Primary key
Check constraint
A data engineer is designing a data warehouse for a multinational corporation. The company has sales data from different regions with varying currencies and date formats. To ensure consistency, which data concept should be applied to standardize the data before loading into the warehouse?
Data cleansing
Data transformation
Data transformation converts source values into a consistent format, normalising currencies and date formats during ETL before loading. This satisfies the stem's standardisation constraint by ensuring multinational sales records are comparable and queryable within the warehouse.
Data profiling
Data masking
An e-commerce company uses a star schema for its data warehouse. The fact table 'sales_fact' contains foreign keys to dimension tables: customer_dim, product_dim, time_dim, and store_dim. A business user wants to know the total sales for each product category in the last month. Which join operation is required to retrieve this data?
Self-join on the fact table
Cross join between fact and dimension tables
Inner join between fact table and dimension tables
An inner join matches each fact row to its related dimension rows via the foreign keys, letting the query group sales by product category. Because every sales_fact row has valid dimension references, inner joins return all required combinations without dropping matching data.
Left outer join between fact and dimension tables
A data scientist is building a machine learning model to predict customer churn. The dataset includes both numerical features (age, income) and categorical features (gender, marital status). Which data concept describes the process of converting categorical features into numerical values that can be used by the algorithm?
Data sampling
Encoding
Encoding maps categorical values such as gender and marital status into numeric representations, for example one-hot or ordinal vectors, which the algorithm can process. Numerical features like age and income need no conversion, so encoding is the concept that addresses the categorical constraint.
Feature scaling
Dimensionality reduction
A company's database has a table 'orders' with columns: order_id, customer_id, order_date, and total_amount. A data analyst needs to identify customers who have placed more than 5 orders in the past year. Which data concept should be used to group orders by customer and count them?
Joining with other tables
Filtering with WHERE clause
Sorting with ORDER BY
Aggregation with GROUP BY
Aggregation with GROUP BY satisfies the requirement to group orders by customer_id and count them. GROUP BY collapses rows sharing a customer_id into single groups, then COUNT() tallies each customer's orders. Filtering with HAVING COUNT(*) > 5 after grouping identifies those exceeding five orders in the past year, which WHERE cannot do on aggregates.
Want more Data Concepts and Environments practice?
Practice this domainA data analyst needs to identify the most frequently occurring value in a dataset. Which measure of central tendency should they use?
Mode
The mode identifies the value appearing most frequently, directly satisfying the requirement to find the most frequently occurring value. Unlike the mean, which averages all values, or the median, which finds the middle position, the mode alone measures frequency of occurrence, making it the appropriate measure of central tendency for this scenario.
Standard deviation
Median
Mean
A retail company wants to predict future sales based on historical data. Which modeling approach is most appropriate if the data shows a clear seasonal pattern?
Linear regression
Time series analysis
Time series analysis explicitly models sequential dependence and seasonality through components such as trend, seasonal, and residual terms, satisfying the stem's clear seasonal pattern requirement. Unlike regression, it uses autocorrelation and prior-period values, so forecasts of future sales account for recurring cycles rather than treating observations as independent.
K-means clustering
Logistic regression
A data analyst is building a model to predict customer churn. The dataset has 10,000 records with 500 churned customers. The model predicts churn with 95% accuracy, but only identifies 10% of actual churners. Which metric best highlights this issue?
Accuracy
F1 score
Recall
Recall measures the proportion of actual churners correctly identified, so 10% recall exposes the model's failure to catch churn despite 95% accuracy. Accuracy is misleading here because the 500 churners are a small minority of 10,000 records.
Precision
A data analyst is performing a hypothesis test with a significance level of 0.05. The p-value obtained is 0.03. What should the analyst conclude?
Reject the null hypothesis
The p-value of 0.03 falls below the 0.05 significance level, meaning the observed result is unlikely under the null hypothesis. The analyst therefore rejects the null hypothesis in favour of the alternative, accepting the result as statistically significant.
Fail to reject the null hypothesis
Accept the null hypothesis
The result is practically significant
A data scientist trains a regression model and observes high variance with low bias. Which technique is most appropriate to reduce variance?
Apply Ridge regularization
High variance with low bias indicates overfitting, where coefficients fit training noise. Ridge adds an L2 penalty on squared coefficient magnitudes, shrinking them and constraining model complexity, which reduces variance at the cost of slightly increased bias.
Increase polynomial features
Use a smaller training set
Remove correlated features
A data analyst is cleaning a dataset and finds missing values in a categorical variable representing customer region. Which imputation method is most appropriate?
Drop rows with missing values
Mode imputation
Mode imputation fills missing categorical entries with the variable's most frequent region, preserving the existing distribution and requiring no numeric assumptions. It suits categorical data, where mean or median imputation is meaningless, satisfying the requirement to handle missing region values sensibly.
Mean imputation
Median imputation
Want more Data Analysis practice?
Practice this domain20% of exam · 6 sample questions below
A data analyst creates a bubble chart showing country GDP (x-axis), life expectancy (y-axis), and population (bubble size). However, large bubbles overlap and obscure many data points. Which corrective action should the analyst take?
Increase the chart canvas size
Set bubble opacity to 70%
Setting bubble opacity to 70% lets overlapping marks remain partially visible, so obscured data points stay readable while bubble size still encodes population. This directly addresses the stem's constraint that large bubbles overlap and hide points, without altering the GDP or life expectancy axes.
Reduce all bubble sizes uniformly
Remove outlier countries with large populations
An analyst wants to show the distribution of test scores for 500 students. Which visualization type is best for understanding the shape of the distribution?
Line chart
Pie chart
Scatter plot
Histogram
A histogram bins the 500 continuous test scores into intervals and plots frequency per bin, directly revealing the distribution's shape, centre, spread and skew. Bar charts suit categorical counts, and box plots summarise rather than display the full distribution.
A dashboard shows sales by region using a map with color intensity. Users complain that two regions with very different sales appear nearly the same color. What is the most likely cause?
The map projection is distorted
The color scale uses a sequential palette with insufficient contrast
A sequential palette maps values onto one hue's lightness ramp, so two regions with very different sales can land on similar shades when the scale's contrast is too low or its range poorly fitted to the data. Widening the lightness range or switching to a diverging scale restores the visible difference.
The monitor resolution is too low
Users are color blind
An analyst creates a stacked bar chart showing quarterly sales by product category. The chart becomes hard to read because some categories have very small contributions. Which redesign is most effective?
Combine small categories into an 'Other' group
Combining small categories into an 'Other' group reduces the number of segments competing for limited bar height, so each remaining category occupies a larger, readable portion. This directly addresses the stem's constraint: tiny contributions that become illegible in a stacked bar. Aggregation preserves the total while restoring visual clarity.
Change to a pie chart for each quarter
Increase the width of each bar
Switch to a 3D stacked column chart
A data analyst needs to present the correlation between advertising spend and website traffic. Which chart type is most appropriate?
Bar chart
Line chart
Pie chart
Scatter plot
A scatter plot encodes each observation as a point on two continuous axes, revealing the strength, direction, and shape of the relationship between advertising spend and website traffic. Correlation is a pairwise association, so this two-variable encoding satisfies the requirement precisely.
Which TWO actions are best practices for creating effective data visualizations?
Avoid using more than five slices in a pie chart
Capping pie slices at five preserves accurate part-to-whole comparison, since human perception struggles to judge angles beyond a handful of wedges. This satisfies the best-practice constraint by preventing cluttered, unreadable charts, so viewers can distinguish each category's proportion at a glance rather than squinting at slivers.
Maximize data-ink ratio by removing all whitespace
Always include gridlines with high contrast
Use 3D effects to make charts look professional
Use color to represent data values consistently
Consistent colour mapping lets viewers compare values across charts without relearning the encoding each time, directly satisfying the best-practice requirement for effective visualisations. Assigning the same hue to identical categories or magnitudes reduces cognitive load and prevents misreading trends.
Want more Visualization and Reporting practice?
Practice this domainA data team is preparing a dashboard for executives. The team wants to highlight key performance indicators (KPIs) that are below target. Which of the following visualization techniques would most effectively draw attention to underperforming metrics without causing confusion?
Remove underperforming KPIs from the dashboard to avoid confusion.
Use a scatter plot to show the relationship between KPIs.
Apply conditional formatting to turn KPI values red when below target.
Conditional formatting applies a rule that automatically renders KPI values red when they fall below target, drawing attention through colour without altering the underlying data or adding extra chart elements. This satisfies the constraint of highlighting underperformance clearly while avoiding visual confusion.
Use a pie chart showing the proportion of each KPI.
A data analyst needs to present the distribution of customer ages to a non-technical audience. Which type of chart would be most appropriate?
Scatter plot
Histogram
A histogram bins continuous numeric values such as ages into intervals, revealing the distribution's shape, peaks and spread. This suits a non-technical audience because the adjacent bars show frequency patterns that summary statistics alone would obscure.
Pie chart
Line chart
A data analyst creates a report showing sales by product category. The analyst notices that one category has a very high sales figure due to a one-time bulk order. Which of the following is the best way to communicate this insight to stakeholders?
Delete the bulk order from the dataset.
Add a note to the chart explaining the bulk order.
Annotating the chart identifies the one-time bulk order as the cause of the spike, preventing stakeholders from misreading it as sustained demand. This preserves the accurate sales figure while supplying the context needed for sound decisions.
Remove the category with the bulk order from the report.
Use a pie chart to show the proportion of each category.
A data analyst is building a dashboard that will be used by both executives and operational managers. The executives need high-level summaries, while managers need granular details. Which dashboard design principle should the analyst apply?
Use a single chart that shows both summary and detail simultaneously.
Display all available data on one page for transparency.
Design the dashboard with drill-down capabilities from summary to detail.
Drill-down lets one dashboard serve both audiences: executives view aggregated summaries, then navigate to granular detail on demand. This satisfies the stem's dual requirement without duplicating reports, since a single semantic layer filters progressively by dimension. Static summary-only or detail-only designs would fail one group.
Create two separate dashboards for each audience.
A data analyst wants to compare the sales performance of four different stores over the same time period. Which chart type is most suitable?
Line chart with multiple lines
Grouped bar chart
A grouped bar chart places the four stores side by side within each time period on a shared categorical axis, enabling direct comparison of their sales values. This satisfies the requirement to compare multiple discrete entities across the same period, which a single-series chart or pie chart cannot do.
Stacked bar chart
Pie chart with multiple pies
A data analyst creates a scatter plot showing the relationship between advertising spend and revenue. The plot shows a strong positive correlation. Which of the following should the analyst include in the report to ensure accurate communication?
Include a note that correlation does not imply causation.
A strong positive correlation between advertising spend and revenue does not establish that spending caused the revenue change; a confounding variable could drive both. Stating this caveat satisfies the accurate-communication requirement, preventing readers from inferring a causal relationship the scatter plot cannot support.
Replace the scatter plot with a bar chart.
Remove any outliers from the plot.
Add a trend line to the scatter plot.
Want more Data Governance practice?
Practice this domainThe DA0-002 exam has 90 questions and must be completed in 90 minutes. The passing score is 675/1000.
Multiple-choice and performance-based questions covering IT security, networking, and operations. Some questions are performance-based (PBQs), asking you to complete tasks in a simulated environment.
The exam covers 5 domains: Data Acquisition and Preparation, Data Concepts and Environments, Data Analysis, Visualization and Reporting, Data Governance. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official CompTIA DA0-002 exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.