DEA-C02 Data Movement Practice Question
A data engineer is loading a large number of small JSON files into a Snowflake table using the COPY command. The process is taking much longer than expected. Which action is most likely to resolve the performance bottleneck?
⚠ Common exam trap
Candidates often try to optimize the COPY command parameters or virtual warehouse size, overlooking the fact that small files fundamentally cripple metadata throughput regardless of compute power.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Aggregating the small JSON files into fewer, larger files before initiating the COPY command.
Small files create significant metadata overhead for Snowflake, as every file requires a separate request and tracking entry. Consolidating small files into larger batches (100MB-250MB) before loading significantly improves throughput. This reduces the number of I/O operations and allows the virtual warehouse to spend more time processing data rather than managing file-level metadata.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increasing the size of the virtual warehouse from Medium to 2X-Large.
Why it's wrong here
While a larger warehouse provides more threads, it cannot overcome the overhead of processing thousands of tiny files. The bottleneck is likely at the metadata and network request layer. Simply adding more compute will not efficiently solve the problem and will lead to significantly higher costs without a proportional performance gain.
- ✓
Aggregating the small JSON files into fewer, larger files before initiating the COPY command.
Why this is correct
This is the most effective way to optimize ingestion. By reducing the total number of files, you decrease the overhead associated with listing and opening files in cloud storage. Larger files allow Snowflake to utilize its parallel processing capabilities more effectively, leading to much faster and more cost-efficient data movement.
- ✗
Changing the target table's data retention period to 0 days during the load.
Why it's wrong here
Data retention (Time Travel) affects storage and the ability to recover data, but it has no impact on the performance of the COPY command itself. Disabling it will not reduce the time it takes to parse JSON files or manage the metadata overhead of a large number of small files.
- ✗
Using the 'STRIP_OUTER_ARRAY = TRUE' file format option in the COPY command.
Why it's wrong here
This option is used for parsing JSON files that are wrapped in a top-level array. While necessary for correct parsing of certain JSON structures, it does not improve performance for a large number of small files. It is a functional setting rather than a performance optimization for bulk data movement.
Visual reference
About these practice questions
This DEA-C02 question is part of Courseiva's 229-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Snowflake exam blueprint
This DEA-C02 practice question is part of Courseiva's free Snowflake certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C02 exam.