DEA-C01 Data Operations and Support Practice Question
A data engineer is designing a data pipeline that ingests data from multiple sources into Amazon S3, then processes it with AWS Glue and loads it into Amazon Redshift. Which THREE practices should be implemented to ensure data quality?
⚠ Common exam trap
DEA-C01 often tests the confusion between cost optimisation (compression) and data quality practices, so candidates pick compression or manual sampling instead of automated validation, profiling, and monitoring.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement data validation checks at the ingestion stage
Option A is correct because validating data at ingestion (for example, checking types, nulls, ranges, and referential integrity before writing to Amazon S3) prevents corrupt or malformed records from propagating downstream into Glue ETL jobs and Redshift tables. Option B is correct because AWS Glue DataBrew provides visual data profiling, column statistics, and built-in transformations/rules that can enforce schema consistency and detect anomalies before loading into Redshift. Option E is correct because Amazon CloudWatch alarms on Glue job metrics, Redshift load events, and custom anomaly metrics enable proactive detection of pipeline failures and unexpected data patterns, which is essential for ongoing data quality assurance. Option C is not correct because compressing files reduces storage and I/O costs but does not validate or improve data accuracy, completeness, or consistency. Option D is not correct because manual sampling is ad hoc, non-scalable, and not a reliable or automated data quality control compared with systematic validation, profiling, and monitoring.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Implement data validation checks at the ingestion stage
Why this is correct
Validating records as they land in Amazon S3 catches malformed, incomplete or out-of-range data before Glue transforms it, preventing corrupt rows from propagating into Redshift. This satisfies the requirement to ensure data quality across the multi-source ingestion pipeline at its earliest stage.
- ✓
Use AWS Glue DataBrew for data profiling and schema enforcement
Why this is correct
Glue DataBrew profiles datasets and enforces schemas, detecting anomalies, duplicates and type mismatches before transformation. Running it between S3 ingestion and Redshift loading satisfies the data quality requirement by standardising structure and flagging deviations early in the pipeline.
- ✗
Compress data files to reduce storage costs
Why it's wrong here
Compression reduces storage and transfer cost, not data quality; it changes file size, not values, completeness or validity. It is tempting because compression is a genuine best practice for S3 ingestion pipelines, and would be correct if the question asked about cost optimisation or query performance rather than data quality.
- ✗
Use manual sampling to check data quality periodically
Why it's wrong here
Manual sampling is periodic and human-driven, so it misses defects between checks and cannot scale across multiple sources; automated validation rules at ingestion and transformation catch issues continuously. It is tempting because sampling is a real quality technique, and would suit one-off exploratory profiling rather than a production pipeline.
- ✓
Set up Amazon CloudWatch alarms for pipeline failures and data anomalies
Why this is correct
CloudWatch alarms detect pipeline failures and anomalous metrics, enabling proactive response before bad data propagates to Redshift. This satisfies the data quality requirement by providing continuous monitoring and alerting across the Glue and S3 stages of the pipeline.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.