Courseiva
Data Operations and Support →mediumMultiple Choice

DEA-C01 Data Operations and Support Practice Question

A data engineer maintains an AWS Glue ETL job that processes JSON files from Amazon S3 and writes Parquet to another S3 location. The job has been running successfully for months. Recently, the job started failing intermittently with the error 'Unable to infer schema for JSON'. The engineer confirms the source bucket contains valid JSON files. Which action should the engineer take to resolve the failure?

⚠ Common exam trap

The trap here is assuming that increasing compute resources or changing file formats will fix schema inference errors, when the real solution is to provide an explicit schema.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Define an explicit schema using a Glue Data Catalog table or a DynamicFrame with a specified schema, instead of relying on schema inference.

The error 'Unable to infer schema for JSON' occurs when AWS Glue cannot automatically determine a consistent schema from the source files. This often happens when JSON records have varying fields or data types. Defining an explicit schema through the Glue Data Catalog or programmatically ensures the job can parse the data without relying on inference, resolving the failure while maintaining data integrity.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable job bookmarks to track previously processed files and avoid reprocessing.

    Why it's wrong here

    Job bookmarks help with incremental processing by tracking which files have been processed, but they do not resolve schema inference errors. The failure occurs because Glue cannot determine the schema from the JSON content, not because it is reprocessing old files. Bookmarks would not prevent the error from occurring on new files with inconsistent schemas.

  • ✗

    Increase the number of AWS Glue DPUs allocated to the job to handle larger file sizes.

    Why it's wrong here

    Adding DPUs increases compute capacity but does not address schema inference failures. The error indicates Glue cannot determine the schema from the JSON data, often due to inconsistent or malformed records. More DPUs would not fix a parsing issue and could increase cost unnecessarily. The engineer should first inspect the data and define the schema explicitly.

  • ✓

    Define an explicit schema using a Glue Data Catalog table or a DynamicFrame with a specified schema, instead of relying on schema inference.

    Why this is correct

    When JSON files have inconsistent structures, Glue's automatic schema inference can fail. Providing an explicit schema via a Data Catalog table or by passing a schema to the DynamicFrame resolves the ambiguity. This ensures the job can parse the data reliably regardless of variations, and is the recommended approach for production workloads with evolving or irregular JSON.

  • ✗

    Convert the source JSON files to CSV format before running the job, then update the job to read CSV.

    Why it's wrong here

    Converting to CSV might work but it is a significant change to the data pipeline and may lose nested structures. It does not directly address the root cause of schema inference failure. The error is about JSON parsing, not format preference. An explicit schema is a more targeted and less disruptive fix that preserves the original data format.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.