DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue to read a large dataset from Amazon S3 and write it to Amazon Redshift. The job intermittently fails with 'Communication link failure' errors during the write phase. The dataset is several hundred gigabytes and the Redshift cluster is under heavy query load. Which change is MOST likely to resolve the failures while preserving data integrity?
⚠ Common exam trap
The trap here is treating the communication failure as a Glue-side timeout or capacity problem, when the real fix is to change the write path to use S3 staging plus Redshift COPY.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Write the transformed data to Amazon S3 in Parquet, then use a Redshift COPY command from S3 instead of a direct JDBC write.
Staging data in S3 and using the Redshift COPY command is the recommended pattern for bulk loads. It minimizes open connections, leverages Redshift's parallel load capability, and provides transactional integrity. Increasing workers, extending timeouts, or switching to Spectrum does not resolve the underlying contention or integrity concerns during direct writes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch the write operation to use Redshift Spectrum and query the S3 data directly without loading.
Why it's wrong here
Redshift Spectrum queries external tables in S3 but does not load data into Redshift tables. If the requirement is to persist transformed data in Redshift for downstream analytics, Spectrum does not satisfy it. It also does not address the write-phase failure because no write occurs. This changes the architecture rather than fixing the ingestion issue.
- ✓
Write the transformed data to Amazon S3 in Parquet, then use a Redshift COPY command from S3 instead of a direct JDBC write.
Why this is correct
The Glue Redshift connector can stage data in S3 and issue a COPY command, which is far more resilient and efficient than row-by-row JDBC inserts. COPY is atomic per transaction, handles large volumes well, and reduces the number of open connections to Redshift. This approach preserves data integrity and avoids the communication link failures caused by direct writes under load.
- ✗
Increase the Glue job's number of workers and enable auto-scaling.
Why it's wrong here
Adding workers increases parallelism, which can overwhelm the Redshift cluster further and worsen connection failures. The error stems from write-phase contention, not insufficient Glue compute. More concurrent connections to Redshift during heavy query load would likely increase the failure rate rather than resolve it.
- ✗
Configure the Glue connection to use a Redshift JDBC URL with a longer socketTimeout and retry the write.
Why it's wrong here
Extending socketTimeout may mask the symptom temporarily, but it does not address the underlying cause of the communication failure, which is typically resource contention or a dropped connection during COPY. It also does not guarantee data integrity if a partial write occurs. This is a superficial fix rather than a robust solution.
Visual reference
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.