DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue Studio to build an ETL job that reads semi-structured JSON from Amazon S3. The source files contain nested arrays and inconsistent keys, and the engineer wants the job to automatically infer the schema at runtime without a Data Catalog table. Which transform or configuration should the engineer use to read the data most reliably?
⚠ Common exam trap
The trap here is assuming a Glue crawler or a hard-coded Spark schema is required to read JSON into Glue, when DynamicFrame.from_options with recurse handles runtime inference directly.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a DynamicFrame with the from_options method and set the format to "json" with the recurse parameter enabled.
Reading semi-structured JSON from S3 without a Data Catalog entry is best done with DynamicFrame.from_options, specifying the JSON format and enabling recursion so nested arrays and inconsistent keys are handled during inference. This avoids the rigidity of hard-coded schemas and the catalog dependency of crawler-based reads.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Create a DynamicFrame with the from_options method and set the format to "json" with the recurse parameter enabled.
Why this is correct
DynamicFrame.from_options with format "json" and recurse=true handles nested and inconsistent JSON by flattening and inferring the schema dynamically, which suits semi-structured data without a catalog table. It is the standard programmatic approach in AWS Glue for reading JSON from S3 while accommodating schema variability at runtime.
- ✗
Use a Spark DataFrame read with spark.read.schema() and a hard-coded StructType that matches all nested fields.
Why it's wrong here
Hard-coding a StructType requires knowing the schema in advance, which conflicts with the goal of runtime inference for inconsistent keys. If new fields appear, the job fails or silently drops them. This is the opposite of the dynamic inference behavior the engineer is seeking.
- ✗
Configure the job to use the ApplyMapping transform with a manually defined mapping for every nested attribute before reading.
Why it's wrong here
ApplyMapping operates on an already-read DynamicFrame; it does not handle reading or schema inference. Using it before reading is not possible in the Glue job flow, and manually defining mappings contradicts the requirement to infer schema at runtime for inconsistent JSON structures.
- ✗
Run an AWS Glue crawler against the S3 prefix and then use the catalog table in the job with a from_catalog read.
Why it's wrong here
A crawler-based catalog read is viable, but the scenario explicitly states the engineer wants schema inference at runtime without a Data Catalog table. Crawlers also struggle with inconsistent keys and nested arrays, often producing a schema that drops or misclassifies fields, so this approach does not meet the stated requirement.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.