DEA-C01 Data Operations and Support Practice Question
A data engineer is troubleshooting a failed AWS Glue ETL job that reads from an S3 bucket. The job logs show the following error: 'java.lang.RuntimeException: java.lang.ClassNotFoundException: Class org.apache.hadoop.fs.s3a.S3AFileSystem not found'. Which TWO actions will resolve this issue?
⚠ Common exam trap
DEA-C01 often tests the confusion between classpath errors (missing JAR) and permission errors (IAM), leading candidates to choose IAM or VPC options when the error is clearly a ClassNotFoundException.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Include the hadoop-aws jar as an extra jar in the Glue job configuration.
The error 'ClassNotFoundException: org.apache.hadoop.fs.s3a.S3AFileSystem' means the S3A filesystem implementation class is not on the Glue job's classpath, so the fix must supply that library. Option B is correct because adding the hadoop-aws jar (which contains org.apache.hadoop.fs.s3a.S3AFileSystem) as an extra jar via the Glue job's --extra-jars configuration puts the missing S3A class on the classpath. Option D is correct because newer AWS Glue versions (Glue 3.0 and later, built on Spark 3.x/Hadoop 3.x) bundle the S3A filesystem library, so running the job on such a version provides the class natively. Option A is wrong because a VPC S3 endpoint only fixes network routing to S3, not a missing Java class. Option C is wrong because IAM permissions govern authorization, and a ClassNotFoundException is a classpath problem, not an access-denied error. Option E is wrong because EMRFS is an EMR-specific filesystem and is not a valid S3 access mode to switch to within AWS Glue.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable VPC S3 endpoint for the Glue job.
Why it's wrong here
A VPC S3 endpoint provides private network routing to S3; it does not supply the missing S3A Hadoop connector class on the Glue worker classpath. It is tempting because S3 connectivity errors often implicate networking, and an endpoint would be correct for private-subnet access failures, not ClassNotFoundException.
- ✓
Include the hadoop-aws jar as an extra jar in the Glue job configuration.
Why this is correct
The ClassNotFoundException means the S3A filesystem implementation is absent from the job's classpath. Adding the hadoop-aws jar as an extra jar supplies the missing org.apache.hadoop.fs.s3a.S3AFileSystem class, letting the Glue job resolve and read from S3.
- ✗
Update the IAM role to allow access to S3.
Why it's wrong here
The ClassNotFoundException is a classpath problem: the S3A filesystem JAR is absent from the job's dependencies. IAM permissions govern authorisation, not class loading, so a broader role cannot make a missing class appear. Granting S3 access is correct when a job fails with AccessDenied errors instead.
- ✓
Use a Glue version that includes the S3A filesystem library (e.g., Glue 3.0 or later).
Why this is correct
Glue 3.0 and later bundle the Hadoop S3A connector, so `org.apache.hadoop.fs.s3a.S3AFileSystem` resolves on the classpath without extra configuration. Earlier Glue versions omit this library, which is precisely why the job throws ClassNotFoundException when reading from S3. Upgrading satisfies the missing-dependency constraint directly.
- ✗
Change the S3 access mode from S3A to EMRFS.
Why it's wrong here
Glue's S3 connection uses the EMRFS connector natively; the S3AFileSystem class belongs to Hadoop's s3a:// scheme, which Glue does not bundle. Switching schemes masks the missing class rather than supplying it. EMRFS is the right choice on EMR clusters, where it is the default filesystem implementation.
Visual reference
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.