Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is loading data from Amazon S3 into an Amazon Redshift cluster using the COPY command. The S3 bucket contains 500 Parquet files, each about 200 MB, in a single prefix. The COPY job is running slowly and consuming excessive cluster resources. The engineer wants to improve performance without changing the data format or the cluster size. Which action should the engineer take?

⚠ Common exam trap

The trap here is assuming that COPY always parallelizes perfectly regardless of file layout, when in fact a single prefix with many files can create a leader-node and slice-distribution bottleneck.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Split the data into multiple prefixes and run multiple COPY commands in parallel, or use a manifest file to distribute the load across slices.

Redshift COPY achieves high throughput by having multiple slices read files in parallel. When hundreds of files sit in one prefix, the leader node may serialize listing and assignment, and some slices may read more files than others, causing skew and resource contention. Distributing files across prefixes or using a manifest gives COPY finer-grained control over parallel reads, balancing work across slices and improving load speed without altering the format or cluster size.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use the COPY command with the PARALLEL OFF option to reduce the number of slices used.

    Why it's wrong here

    PARALLEL OFF forces COPY to load data serially using a single slice, which drastically reduces throughput and increases load time. It is intended for small loads or when you need to preserve order, not for optimizing a large Parquet load. Disabling parallelism would make the slow performance worse, not better.

  • ✓

    Split the data into multiple prefixes and run multiple COPY commands in parallel, or use a manifest file to distribute the load across slices.

    Why this is correct

    Redshift COPY parallelizes across slices by reading multiple files concurrently. When many files are in one prefix, the leader node can become a bottleneck and slice distribution may be uneven. Splitting into multiple prefixes or using a manifest file lets COPY distribute files more evenly across slices, improving throughput and reducing resource contention without changing the data format or cluster size.

  • ✗

    Convert the Parquet files to CSV and load them with the CSV option to enable faster parsing.

    Why it's wrong here

    Parquet is a columnar, compressed format that Redshift can load efficiently and is generally faster than CSV for analytics workloads. Converting to CSV would increase file size, require more parsing, and likely worsen performance. The scenario explicitly asks to improve performance without changing the data format, so this approach is counterproductive.

  • ✗

    Add the COMPUPDATE OFF and STATUPDATE OFF options to the COPY command to skip compression and statistics updates.

    Why it's wrong here

    COMPUPDATE and STATUPDATE control compression encoding and statistics updates, not file distribution. Turning them off may slightly reduce CPU usage but does not address the underlying parallelization bottleneck caused by many files in a single prefix. The slow performance and resource consumption would largely remain.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.