Courseiva
Data Ingestion and TransformationmediumMultiple ChoiceObjective-mapped

DEA-C01 Data Ingestion and Transformation Practice Question

Exhibit

Refer to the exhibit.

2019-11-15T10:00:00Z ERROR: Task failed: 'NoneType' object has no attribute 'read'

Refer to the exhibit. A data engineer is running an AWS Glue job that reads data from an S3 source. The job fails with the error shown. What is the MOST likely cause?

⚠ Common exam trap

Watch out — candidates often assume permission errors (Option A) are the default cause of any S3-related failure, but the specific error message (NullPointerException) points to data corruption or empty files, not access control issues.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

One of the source files is empty or corrupted.

The error message indicates that the Glue job encountered a 'NullPointerException' or similar parsing failure when reading from S3. This typically occurs when a source file is empty or corrupted, causing the Spark DataFrame reader to fail during schema inference or data parsing. AWS Glue jobs rely on Spark's ability to read files; an empty or malformed file triggers a runtime error because Spark cannot extract any records or infer a valid schema from it.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • The IAM role does not have s3:GetObject permission.

    Why it's wrong here

    Would cause access denied.

  • One of the source files is empty or corrupted.

    Why this is correct

    Empty file can return None when read, causing 'NoneType' has no attribute 'read'.

  • The file is in JSON format but the schema expects Parquet.

    Why it's wrong here

    Would cause a parsing error.

  • The Glue job has insufficient memory allocated.

    Why it's wrong here

    Would cause out-of-memory error.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,711 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.