DA0-002 Data Acquisition and Preparation Practice Question
A retail analytics team needs to load a 40 GB CSV file of point-of-sale transactions into a cloud data warehouse nightly. The file is generated as a single object by an upstream system, and the team must minimize load time. Which approach best addresses the load performance bottleneck?
⚠ Common exam trap
The trap here is assuming that compressing the file alone will solve the load time issue, but compression only reduces transfer size and does not enable parallel processing.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Split the file into multiple smaller files and load them in parallel.
The correct approach is to split the large file into multiple smaller files and load them in parallel. This leverages the distributed architecture of modern data warehouses, enabling concurrent ingestion and significantly reducing overall load time. Other options either do not address the parallelism bottleneck or introduce additional processing overhead that worsens performance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Split the file into multiple smaller files and load them in parallel.
Why this is correct
Splitting the large CSV into multiple smaller files allows the data warehouse to ingest them concurrently, dramatically reducing overall load time. Parallel ingestion is a standard best practice for bulk loading large datasets, as it leverages distributed compute resources and avoids single-threaded bottlenecks. This directly addresses the performance issue without altering the data content.
- ✗
Convert the CSV to JSON and load it as a semi-structured format.
Why it's wrong here
Converting to JSON increases file size and parsing complexity, often slowing down loads. Semi-structured formats are useful for schema flexibility, but they do not inherently improve load performance for a large batch file. This approach would likely worsen the bottleneck rather than alleviate it, making it a poor choice for the stated goal.
- ✗
Compress the file using gzip and load the single compressed file.
Why it's wrong here
While compression reduces network transfer time, the warehouse still processes a single object, which limits parallelism. The decompression step may also add overhead. This does not solve the core bottleneck of a single large file being loaded sequentially. Compression is beneficial but insufficient when the goal is to minimize load time for a monolithic file.
- ✗
Load the file into a staging table using row-by-row inserts.
Why it's wrong here
Row-by-row inserts are extremely slow for large datasets and would drastically increase load time. Bulk loading mechanisms are designed to handle large volumes efficiently. This method contradicts the requirement to minimize load time and would create a severe performance penalty, making it unsuitable for a 40 GB nightly load.
About these practice questions
One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.