DEA-C02 Performance Optimization Practice Question
A data engineer runs a transformation that loads a 50 GB staged file into a target table using a COPY INTO statement with a single large file. The warehouse is a MEDIUM multi-cluster warehouse. The load takes far longer than expected because one node processes the entire file. Which change most directly improves load throughput?
⚠ Common exam trap
The trap here is reaching for a larger warehouse or multi-cluster scaling when the real constraint is that a single file cannot be parallelized.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Split the staged file into multiple smaller files of roughly 100-250 MB compressed and load them in the same COPY INTO statement.
COPY INTO distributes work per file, so throughput scales with the number of files rather than the size of the warehouse. Breaking a single large file into many moderately sized compressed files enables parallel processing across nodes, which is the direct remedy for a load dominated by one file.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Split the staged file into multiple smaller files of roughly 100-250 MB compressed and load them in the same COPY INTO statement.
Why this is correct
COPY INTO parallelizes work across the files in a stage, so a single monolithic file forces one thread to process the entire load. Splitting the data into many smaller compressed files lets the warehouse distribute file processing across nodes, dramatically increasing throughput. The 100-250 MB compressed target size is the recommended range that balances parallelism against per-file overhead, making this the most direct fix for the described bottleneck.
- ✗
Set the ON_ERROR option to CONTINUE and re-run the COPY INTO statement.
Why it's wrong here
ON_ERROR controls how the load handles malformed rows, such as skipping them instead of aborting. It has no effect on how data files are distributed to threads or how quickly valid rows are ingested. Re-running with this option would still process the same single large file through one thread, so load duration would not meaningfully improve for this scenario.
- ✗
Enable multi-cluster scaling by setting the maximum cluster count to four on the warehouse.
Why it's wrong here
Multi-cluster warehouses help by adding clusters when many concurrent queries queue, which is a concurrency problem. A single COPY INTO statement is one query, so additional clusters do not split its work and would not be engaged. The load remains constrained by the one file being processed by one thread, so this change does not improve throughput for the described job.
- ✗
Increase the warehouse size from MEDIUM to 4X-LARGE before running the same COPY INTO statement.
Why it's wrong here
Resizing the warehouse adds compute nodes, but a single file can only be read by one thread, so the extra capacity sits idle during the load. The bottleneck is file-level parallelism, not node count. Scaling up would consume substantially more credits without addressing the root cause, so it is not the most direct or economical improvement for this scenario.
About these practice questions
One of 229 original DEA-C02 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Snowflake exam blueprint
This DEA-C02 practice question is part of Courseiva's free Snowflake certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C02 exam.