DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket containing sensitive customer records. The job must remove all columns flagged as PII before writing to a target S3 location. The PII columns are not known in advance and vary by file. Which approach should the engineer use to ensure the PII is removed dynamically?
⚠ Common exam trap
The trap here is assuming AWS Glue has a built-in PII detection transform or that the Data Catalog automatically identifies PII columns.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a custom PySpark script within the Glue job that scans column names and data samples against a regex pattern for PII, then drops matching columns.
The requirement is to dynamically remove PII columns that vary by file. AWS Glue does not have a native PII detection transform, and the Data Catalog does not store PII classifications unless manually added. A custom PySpark script within the Glue job can inspect schema and data, apply regex patterns, and drop matching columns. This provides the necessary flexibility and scalability.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use AWS Glue's `DetectPII` transform to automatically identify and redact PII columns based on built-in patterns.
Why it's wrong here
AWS Glue does not provide a built-in `DetectPII` transform. While Glue has transforms like `Filter`, `Map`, and `ApplyMapping`, there is no native PII detection transform. The engineer would need to implement custom logic or use another service like Amazon Macie for PII detection. This option describes a non-existent feature.
- ✗
Use the AWS Glue DynamicFrame's `drop_fields` method with a list of PII column names derived from the Glue Data Catalog metadata.
Why it's wrong here
The Glue Data Catalog stores technical metadata such as column names and data types, but it does not automatically identify which columns contain PII. Unless the catalog is manually annotated with custom classification, deriving PII column names from it is not possible. This approach would fail to dynamically detect PII columns that vary by file.
- ✗
Use AWS Glue DataBrew to create a recipe that identifies PII columns using pattern matching and apply it as a transformation step in the Glue job.
Why it's wrong here
DataBrew recipes can identify PII using pattern matching, but integrating a DataBrew recipe directly into a Glue ETL job is not a supported native operation. DataBrew is a separate visual data preparation service, and its recipes are not executed within Glue jobs. This approach would require exporting the recipe logic or using DataBrew independently, which does not meet the requirement of dynamic removal within the Glue job.
- ✓
Use a custom PySpark script within the Glue job that scans column names and data samples against a regex pattern for PII, then drops matching columns.
Why this is correct
A custom PySpark script can dynamically inspect each DataFrame's schema and sample data to apply regex patterns that identify PII-like content (e.g., email, SSN). It can then drop those columns before writing. This approach is flexible, scales with Glue's distributed processing, and handles varying PII columns per file, satisfying the requirement without relying on non-existent built-in transforms.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.