Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer must load a 2 GB uncompressed CSV file from Amazon S3 into Amazon Redshift using the COPY command. The cluster is a two-node ra3.xlplus cluster, and the load is running far slower than expected. The engineer wants the fastest reliable improvement without changing the cluster. What should the engineer do?

⚠ Common exam trap

The trap here is thinking that more cluster nodes or compression will speed up a single-file COPY, when the real limiter is that one file maps to one slice.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Split the single CSV file into multiple files and load them in parallel with a single COPY command.

Redshift distributes a COPY workload across slices, but each input file is read by only one slice. A single 2 GB CSV therefore runs on one slice while the rest idle. Splitting the file into several smaller files, ideally a multiple of the slice count, allows all slices to read concurrently, which is the fastest reliable fix that does not alter the cluster.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert the CSV to GZIP before loading and keep it as one file.

    Why it's wrong here

    GZIP reduces the bytes transferred from S3 but keeps the data in a single file, so the load still runs on one slice. Compression helps I/O volume but does not create parallelism. The serialization bottleneck remains, and the load stays much slower than a parallel multi-file load.

  • ✗

    Add the COMPUPDATE OFF and STATUPDATE OFF parameters to the COPY command.

    Why it's wrong here

    COMPUPDATE and STATUPDATE control whether Redshift analyzes compression and updates statistics during or after the load. Turning them off can speed up small loads but does nothing to parallelize reading of a single large file, which is the actual bottleneck. The load remains slice-serialized and slow.

  • ✗

    Increase the cluster to four nodes so more slices can read the same file simultaneously.

    Why it's wrong here

    Even with more slices, Redshift assigns a single input file to a single slice, so additional nodes cannot read the same file concurrently. Adding nodes increases cost without fixing the serialization. The engineer also wants a fix that avoids changing the cluster, so this violates the constraint.

  • ✓

    Split the single CSV file into multiple files and load them in parallel with a single COPY command.

    Why this is correct

    Amazon Redshift parallelizes COPY across slices, but a single file can only be read by one slice at a time, so a lone 2 GB file serializes the load. Splitting it into multiple files lets each slice read its own file concurrently, dramatically improving throughput. This is the standard remedy and requires no cluster change.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.