DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift cluster. The data is in Parquet format and is partitioned by date. The engineer wants to load only the latest partition into Redshift and ensure the load is efficient. Which method should the engineer use?
⚠ Common exam trap
The trap here is thinking there is a PARTITION parameter in the COPY command, or that you must process the entire bucket to filter later.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the Redshift COPY command specifying the S3 path for the latest partition.
The Redshift COPY command is the most efficient way to load Parquet data from S3. By specifying the S3 path to the latest partition, you load only the required data. Other methods like using Glue or Spectrum involve extra processing or are less efficient for this specific task.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use the Redshift COPY command specifying the S3 path for the latest partition.
Why this is correct
The COPY command can load data directly from a specific S3 prefix. By specifying the path to the latest partition (e.g., s3://bucket/data/date=2023-10-01/), you load only that partition efficiently. This is the most direct and performant method for Parquet data.
- ✗
Use Amazon Redshift Spectrum to query the latest partition and insert into a Redshift table.
Why it's wrong here
Redshift Spectrum allows querying external data in S3, but inserting the results into a Redshift table would require a CREATE TABLE AS or INSERT INTO statement, which is less efficient than a direct COPY for bulk loading. It also adds complexity and may not be as fast for large datasets.
- ✗
Use AWS Glue to read the entire S3 bucket and write to Redshift, filtering by date.
Why it's wrong here
Reading the entire bucket and filtering in Glue is inefficient because it processes all partitions, not just the latest. This increases runtime and cost. The requirement is to load only the latest partition efficiently, so this approach is suboptimal.
- ✗
Use the Redshift COPY command with the PARTITION option.
Why it's wrong here
The Redshift COPY command does not have a PARTITION option. You can specify a prefix in the S3 path to load only certain partitions, but there is no PARTITION parameter. This option is invalid and would result in an error.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.