Courseiva
Data Operations and Support →mediumMultiple Choice

DEA-C01 Data Operations and Support Practice Question

A data engineer runs an AWS Glue ETL job that reads CSV files from an Amazon S3 bucket, applies transformations, and writes Parquet output to another S3 bucket. The job fails with the error 'AnalysisException: Unable to infer schema for CSV. It must be specified manually.' The CSV files are stored with a header row, and the job's script uses the default Glue DynamicFrame reader without specifying format options. What is the MOST likely cause of the failure?

⚠ Common exam trap

The trap here is assuming that a schema inference error is caused by permissions or compression, when it actually stems from data formatting or missing header configuration.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The CSV files have inconsistent or malformed data that prevents Glue from inferring a schema, or the header option is not set correctly.

The error 'Unable to infer schema for CSV' occurs when AWS Glue cannot determine column names and types from the source data. This often happens when the header option is not enabled or when the CSV data has inconsistencies such as varying column counts or mixed data types. Ensuring the reader is configured with 'withHeader' set to true and that data is well-formed allows Glue to infer the schema correctly.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The CSV files are compressed with gzip, and Glue cannot infer schema from compressed CSV files.

    Why it's wrong here

    Glue can infer schemas from compressed CSV files, including gzip, if the compression format is specified or automatically detected. Compression does not prevent schema inference; it only affects how data is read. The error message does not mention compression, and the scenario does not state that files are compressed. Therefore, compression is not the likely cause.

  • ✓

    The CSV files have inconsistent or malformed data that prevents Glue from inferring a schema, or the header option is not set correctly.

    Why this is correct

    Glue's schema inference can fail if CSV files have inconsistent columns, missing headers, or if the 'withHeader' option is not set to true. In this scenario, the header row exists but the default reader may not treat it as a header, leading to type conflicts. Setting 'withHeader' to true and ensuring consistent data resolves the error.

  • ✗

    The S3 bucket containing the CSV files does not have the correct bucket policy allowing AWS Glue to read objects.

    Why it's wrong here

    A missing bucket policy would produce an AccessDenied error when Glue attempts to list or read objects, not a schema inference failure. The AnalysisException indicates that Glue could access the data but could not determine the column structure. The error specifically points to schema inference, so permissions are not the root cause here.

  • ✗

    The Glue job's IAM role lacks permissions to read the CSV files from the S3 bucket.

    Why it's wrong here

    Insufficient IAM permissions would result in an access denied error during the read operation, not a schema inference exception. The job would fail before attempting to parse the data. The error clearly indicates that schema inference was attempted but unsuccessful, implying that the files were accessible. Thus, IAM permissions are not the issue.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.