Courseiva
Data Store Management →mediumMultiple Select

DEA-C01 Data Store Management Practice Question

A data engineer is building an Amazon Redshift data warehouse that ingests large staged files from Amazon S3 using the COPY command. The team wants to maximize load performance and minimize the time spent on ingestion. Which TWO practices should the engineer apply? (Choose two.)

⚠ Common exam trap

The trap here is equating query-performance tuning such as sort keys and VACUUM with load-performance tuning, when COPY speed is driven primarily by parallel file distribution and compressed columnar formats.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Load compressed columnar files such as Parquet in a single COPY statement

COPY performance in Redshift depends on parallelism and data volume. Splitting input into multiple similarly sized files lets every slice read concurrently, and using compressed columnar formats such as Parquet reduces bytes moved and enables column pruning. Together these practices shorten load time. Sort keys, single-file loads, and post-load VACUUM operations affect query performance or add overhead rather than accelerating ingestion, so they do not belong in this optimization.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Load compressed columnar files such as Parquet in a single COPY statement

    Why this is correct

    COPY supports columnar formats like Parquet and ORC, and columnar compression reduces the bytes transferred from S3 while allowing column-level pruning during load. Loading multiple Parquet files in one COPY statement lets Redshift distribute them across slices in parallel. Combining columnar compression with parallel file distribution reduces I/O and CPU work, directly improving ingestion throughput and shortening load windows.

  • ✗

    Run VACUUM SORT ONLY on the target table immediately after each COPY

    Why it's wrong here

    VACUUM SORT ONLY re-sorts rows to restore sort order after loads, which consumes cluster resources and time but does not improve the COPY operation itself. It addresses query performance degradation from unsorted data, not ingestion speed. Running it after every load adds overhead to the pipeline and competes with concurrent workloads. This practice does not help meet the stated goal of maximizing load performance.

  • ✗

    Use a single large compressed file to reduce the number of S3 GET requests

    Why it's wrong here

    A single large file is read by one slice, so it cannot exploit Redshift's parallel load architecture. The reduction in S3 GET requests does not compensate for the loss of parallelism, and load time scales with file size on one slice. Compression is beneficial, but it must be combined with multiple files. This option directly undermines the parallel-load goal, making it a poor choice.

  • ✓

    Split the input data into multiple files sized roughly equal and load them in parallel

    Why this is correct

    Redshift parallelizes COPY across slices, and each slice reads a subset of the input files. Using multiple files of similar size lets all slices work concurrently rather than having one large file processed by a single slice. This dramatically shortens load time for large datasets. A single monolithic file forces serial reads and underutilizes the cluster, so splitting into comparable files is a core performance practice for COPY.

  • ✗

    Add a sort key to the target table on the timestamp column before loading

    Why it's wrong here

    A sort key improves query performance by enabling zone maps and ordered scans, but it does not accelerate the COPY ingestion itself. In fact, loading unsorted data into a sorted table causes Redshift to sort during the load, which can add overhead. Defining a sort key is a query-optimization decision, not a load-throughput practice, so it does not satisfy the goal of minimizing ingestion time here.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.