Courseiva
← Back to CompTIA Data+ (DA0-002) questions

Scenario-based practice

Hard Difficulty Questions

Practise CompTIA Data+ (DA0-002) practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

20
scenario questions
DA0-002
exam code
CompTIA
vendor

Scenario guide

How to approach hard difficulty questions

These are the questions most candidates get wrong. They require connecting multiple concepts, reading tricky output, or knowing edge-case behaviour that isn't on most study cards. Practising them trains you to operate under uncertainty — a necessary skill on the real exam.

Quick answer

Hard Difficulty Questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Related practice questions

Related DA0-002 topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmultiple choice
Full question →

During an ETL process, a data quality check fails due to duplicate customer IDs. Which data quality dimension is violated?

Question 2hardmultiple choice
Full question →

An analyst creates a dashboard with a large amount of decorative graphics and background images. According to CompTIA Data+, which design principle is being violated?

Question 3hardmultiple choice
Full question →

An analyst is fitting a polynomial regression model and wants to choose the degree that minimizes overfitting. Which technique should the analyst use?

Question 4hardmultiple choice
Full question →

Refer to the exhibit. An analyst sees this log during data acquisition. What action should be taken first?

Exhibit

Refer to the exhibit.
```
INFO: Starting extraction from source 'SalesDB'
INFO: Connected successfully.
WARN: Data type mismatch in column 'Price': expected NUMERIC, found VARCHAR
ERROR: Row 1523: failed to parse date '2024-02-30'
INFO: Extraction completed with 1 warning and 1 error.
```
Question 5hardmultiple choice
Full question →

During a data mining project, an analyst discovers that a significant number of records have a negative value for the age field. What is the most appropriate first step?

Question 6hardmultiple choice
Full question →

A financial institution is merging transaction data from two different systems. System A stores currency amounts as integers in cents, and System B stores as decimals in dollars. What is the best way to integrate the data?

Question 7hardmultiple choice
Full question →

A data engineer is tasked with acquiring data from a third-party vendor that provides daily file drops via SFTP. The files are large (10 GB each). The pipeline must load data into a data warehouse. Which approach optimizes for speed and reliability?

Question 8hardmultiple choice
Full question →

A healthcare data analyst is presenting findings on patient readmission rates to a group of hospital administrators. The analysis reveals a 15% increase in readmissions over the past quarter for patients aged 65+ from a specific zip code. However, the administrators are skeptical because previous quarterly reports showed no such trend, and they suspect data quality issues. The analyst must communicate this insight effectively while maintaining credibility. Which of the following approaches should the analyst take?

Question 9hardmultiple choice
Full question →

An IT operations team monitors 200 servers. Each server reports CPU utilization (0-100%) every five minutes for the past year. The team wants to visualize the data to identify servers that are consistently over 80% utilization and detect any unusual spikes. They have a large dataset with 100,000+ records per server. The current visualization is a single scatter plot with CPU utilization on the y-axis, time on the x-axis, and each server as a different colored point. The chart is extremely cluttered, with points overlapping and colors indistinguishable. What should the team do to improve the visualization?

Question 10hardmultiple choice
Full question →

A data governance team is drafting a policy for handling personally identifiable information (PII). According to data governance best practices, which document should define the classification levels and handling procedures?

Question 11hardmultiple choice
Full question →

A data analyst needs to visualize sales per capita across U.S. states. States with small populations but high sales (e.g., Delaware) appear too prominent on a choropleth map. Which technique best addresses this issue?

Question 12hardmultiple choice
Full question →

A data analyst is presenting results from a customer segmentation analysis to a non-technical audience. The segmentation was performed using k-means clustering, and the analyst wants to explain the characteristics of each cluster. Which visualization is most appropriate for this audience?

Question 13hardmultiple choice
Full question →

In time series decomposition, a pattern that repeats at regular intervals (e.g., weekly, yearly) is called:

Question 14hardmultiple choice
Full question →

A data scientist trains a regression model and observes high variance with low bias. Which technique is most appropriate to reduce variance?

Question 15hardmulti select
Full question →

Which THREE of the following are properties of ratio data? (Choose THREE.)

Question 16hardmultiple choice
Full question →

An analyst is reviewing a table that stores customer orders. The table contains columns: OrderID, CustomerName, Product1, Product1Qty, Product2, Product2Qty. This design violates which normal form?

Question 17hardmulti select
Full question →

A company is designing a data pipeline to process streaming data from social media feeds. Which THREE of the following are characteristics of streaming data? (Select THREE).

Question 18hardmultiple choice
Full question →

A data analyst is using a public API to collect historical weather data. The API has a rate limit of 100 requests per minute, but the analyst needs to retrieve 10,000 records as quickly as possible. What strategy should be used?

Question 19hardmultiple choice
Full question →

A healthcare analytics team is building a predictive model to identify patients at high risk of readmission within 30 days of discharge. The dataset includes 50,000 patient records with 200 features, including demographics, vital signs, lab results, and historical admissions. The target variable is binary (readmitted or not). The team uses a logistic regression model and achieves an AUC of 0.72 on the test set. However, the model's calibration is poor: for patients predicted to have a 70% risk, the actual readmission rate is only 40%. The team wants to improve calibration without significantly reducing discrimination (AUC). The data scientist suggests applying Platt scaling. However, the team lead is concerned that Platt scaling may reduce the model's ability to rank patients correctly. Which of the following is the best course of action?

Question 20hardmultiple choice
Full question →

A data analyst is reviewing the error log from a nightly batch load. What is the most likely cause of the error?

Exhibit

Refer to the exhibit.

Error log from a data pipeline:

[2025-03-15 10:32:14] ERROR: Duplicate key value violates unique constraint 'order_pkey'
[2025-03-15 10:32:14] Detail: Key (order_id)=(12345) already exists.
[2025-03-15 10:32:15] WARNING: Batch load incomplete. 4999 of 5000 rows inserted.

These DA0-002 practice questions are part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style DA0-002 questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.

CompTIA Data+ (DA0-002) Practice Questions | Courseiva