DA0-002 Data Acquisition and Preparation Practice Question
A data engineer is designing an ETL pipeline to extract sales data from a legacy on-premise database and load it into a cloud data warehouse. The database is slow and queries during business hours affect performance. Which extraction strategy minimizes impact?
⚠ Common exam trap
Candidates often confuse 'log shipping' (a high-availability technique) with 'Change Data Capture' (an extraction method), or assume that any periodic query (like hourly SELECT *) is acceptable without considering the cumulative performance impact on a slow legacy database.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Incremental extraction using Change Data Capture (CDC)
Incremental extraction using Change Data Capture (CDC) minimizes impact on the legacy on-premise database by reading only the changed rows (inserts, updates, deletes) from transaction logs or change tables, rather than issuing heavy SELECT queries. This avoids full table scans or frequent queries during business hours, preserving database performance for operational workloads.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Query the database with SELECT * every hour
Why it's wrong here
Hourly SELECT * re-reads the entire table twenty-four times a day, multiplying load on the already slow database and worsening the business-hours impact. It is tempting because frequent polling gives near-real-time freshness, and it would suit a small, lightly used source where full refreshes are cheap and change data capture is unavailable.
- ✓
Incremental extraction using Change Data Capture (CDC)
Why this is correct
Incremental extraction using Change Data Capture reads only changed rows from the transaction log, avoiding full-table scans that degrade the legacy database during business hours. This directly satisfies the stem's constraint of minimising performance impact on a slow on-premise source, while reducing data volume transferred to the cloud warehouse.
- ✗
Full table extraction nightly
Why it's wrong here
A nightly full table scan reads every row from the slow legacy database, so the heavy load still lands on it, just later. It is tempting because full extraction is the simplest way to guarantee completeness when no reliable change-tracking column exists, and off-peak timing does reduce contention with daytime users.
- ✗
Use a database log shipping
Why it's wrong here
Log shipping copies and restores backup logs on a schedule, so it still reads the production database and lags behind. It is tempting as a low-impact replication technique, but it would suit standby database copies, not extracting changed sales rows into a cloud warehouse.
About these practice questions
Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.