Courseiva

DEA-C01 Data Operations and Support Practice Question

A data engineer is designing a data pipeline that ingests data from multiple sources into Amazon S3, then processes it with AWS Glue and loads it into Amazon Redshift. Which THREE practices should be implemented to ensure data quality?

⚠ Common exam trap

DEA-C01 often tests the confusion between cost optimisation (compression) and data quality practices, so candidates pick compression or manual sampling instead of automated validation, profiling, and monitoring.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implement data validation checks at the ingestion stage

Option A is correct because validating data at ingestion (for example, checking types, nulls, ranges, and referential integrity before writing to Amazon S3) prevents corrupt or malformed records from propagating downstream into Glue ETL jobs and Redshift tables. Option B is correct because AWS Glue DataBrew provides visual data profiling, column statistics, and built-in transformations/rules that can enforce schema consistency and detect anomalies before loading into Redshift. Option E is correct because Amazon CloudWatch alarms on Glue job metrics, Redshift load events, and custom anomaly metrics enable proactive detection of pipeline failures and unexpected data patterns, which is essential for ongoing data quality assurance. Option C is not correct because compressing files reduces storage and I/O costs but does not validate or improve data accuracy, completeness, or consistency. Option D is not correct because manual sampling is ad hoc, non-scalable, and not a reliable or automated data quality control compared with systematic validation, profiling, and monitoring.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Implement data validation checks at the ingestion stage

    Why this is correct

    Validating records as they land in Amazon S3 catches malformed, incomplete or out-of-range data before Glue transforms it, preventing corrupt rows from propagating into Redshift. This satisfies the requirement to ensure data quality across the multi-source ingestion pipeline at its earliest stage.

  • ✓

    Use AWS Glue DataBrew for data profiling and schema enforcement

    Why this is correct

    Glue DataBrew profiles datasets and enforces schemas, detecting anomalies, duplicates and type mismatches before transformation. Running it between S3 ingestion and Redshift loading satisfies the data quality requirement by standardising structure and flagging deviations early in the pipeline.

  • ✗

    Compress data files to reduce storage costs

    Why it's wrong here

    Compression reduces storage and transfer cost, not data quality; it changes file size, not values, completeness or validity. It is tempting because compression is a genuine best practice for S3 ingestion pipelines, and would be correct if the question asked about cost optimisation or query performance rather than data quality.

  • ✗

    Use manual sampling to check data quality periodically

    Why it's wrong here

    Manual sampling is periodic and human-driven, so it misses defects between checks and cannot scale across multiple sources; automated validation rules at ingestion and transformation catch issues continuously. It is tempting because sampling is a real quality technique, and would suit one-off exploratory profiling rather than a production pipeline.

  • ✓

    Set up Amazon CloudWatch alarms for pipeline failures and data anomalies

    Why this is correct

    CloudWatch alarms detect pipeline failures and anomalous metrics, enabling proactive response before bad data propagates to Redshift. This satisfies the data quality requirement by providing continuous monitoring and alerting across the Glue and S3 stages of the pipeline.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.