Courseiva
Data Store Management →mediumMultiple Select

DEA-C01 Data Store Management Practice Question

A data engineer is configuring an Amazon Redshift cluster for a workload that runs large nightly ELT jobs loading data from Amazon S3 and then executes complex analytical queries. The team wants to improve query performance and reduce the time the cluster spends on data loading. Which TWO configuration choices should the engineer make? (Choose two.)

⚠ Common exam trap

The trap here is treating maintenance commands like VACUUM and broad distribution styles like ALL as universal performance fixes, when they can lengthen loads and inflate storage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the COPY command with a manifest file and parallelism to load data from multiple S3 objects.

Bulk loading with COPY and a manifest, combined with parallelism, uses Redshift's MPP architecture to shorten load time. Choosing an appropriate distribution key on large fact tables reduces inter-slice data movement during complex joins, improving query performance. Row-by-row inserts, restricting maintenance to the load window, and applying ALL distribution to every table all degrade performance or increase load time.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the COPY command with a manifest file and parallelism to load data from multiple S3 objects.

    Why this is correct

    COPY is the native, massively parallel load path in Redshift and automatically splits work across slices when multiple files or a manifest are provided. Using a manifest ensures a consistent, complete set of files is loaded, and parallelism exploits the cluster's MPP architecture to shorten load time. This directly addresses the goal of reducing time spent loading data.

  • ✗

    Enable automatic vacuum and analyze only during the nightly load window.

    Why it's wrong here

    VACUUM and ANALYZE are necessary maintenance operations, but restricting them to the load window competes with the ELT jobs and can extend the window rather than shorten it. Redshift can run automatic vacuum and analyze in the background, so pinning them to the load window is counterproductive. This choice does not improve query performance or reduce load time.

  • ✓

    Define distribution keys on the largest fact tables to co-locate join rows on the same slice.

    Why this is correct

    Choosing the right distribution key (for example, the common join column) causes matching rows to land on the same slice, letting joins execute locally without network redistribution. This reduces data movement during complex analytical queries and is a core Redshift tuning technique for large fact tables. It directly improves query performance for the described workload.

  • ✗

    Load data using single-row INSERT statements in a loop from the application.

    Why it's wrong here

    Row-by-row INSERT is extremely slow on Redshift because each statement is a separate transaction that the leader node must coordinate, and it prevents the parallel load path. For nightly ELT of large datasets it would dramatically increase load time. It is the opposite of the recommended bulk-load approach and should be avoided.

  • ✗

    Set the cluster's distribution style to ALL for every table to avoid joins.

    Why it's wrong here

    Distribution style ALL replicates the entire table to every slice, which speeds up joins for small dimension tables but is disastrous for large fact tables because it multiplies storage and load time. Applying ALL to every table would increase load duration and consume excessive space, working against both stated goals. It is a misapplication of a technique intended for small tables.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.