DEA-C01 Data Security and Governance Practice Question
A data engineer stores raw customer records in an Amazon S3 bucket and runs an AWS Glue job that writes curated Parquet files to a second bucket. The governance team requires that the curated data carry a verifiable record of which job run produced it and that any modification to a curated file be detectable. The engineer must also prove that the curated dataset has not been altered since a nightly baseline. Which combination of AWS features should the engineer use?
⚠ Common exam trap
The trap here is treating S3 Versioning or CloudTrail logging as proof of integrity, when only a content digest compared against a trusted baseline can reveal that file bytes were altered.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Compute an SHA-256 checksum per curated file, store the checksums and Glue job run IDs in a manifest, and use that manifest to verify the files against the nightly baseline.
Provenance and tamper evidence require linking each curated file to the job run that created it and being able to prove the bytes have not changed. A manifest that pairs a Glue job run ID with a per-file SHA-256 checksum delivers both. Comparing the nightly recomputed checksums against the baseline manifest surfaces any alteration, while versioning, CloudTrail, and Object Lock address adjacent concerns but not content integrity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable AWS CloudTrail data events on the curated bucket and use AWS Glue job bookmarks to record the last processed run in the Data Catalog.
Why it's wrong here
CloudTrail data events log S3 API activity such as PutObject and GetObject, which supports auditing who acted, but they do not store a content hash that proves a file is unchanged. Glue job bookmarks track incremental processing state, not file integrity, so together they cannot detect silent byte-level modification of curated output.
- ✓
Compute an SHA-256 checksum per curated file, store the checksums and Glue job run IDs in a manifest, and use that manifest to verify the files against the nightly baseline.
Why this is correct
A per-file SHA-256 checksum recorded alongside the Glue job run ID gives both provenance and tamper evidence. Recomputing checksums and comparing them with the stored manifest detects any byte-level change to a curated file, and the run ID ties each file to the job execution that produced it, which is exactly what the governance team requires.
- ✗
Enable S3 Versioning on the curated bucket and configure an S3 event notification to write each version ID to an Amazon DynamoDB table keyed by the Glue job run ID.
Why it's wrong here
Versioning preserves prior object versions and event notifications can record version IDs, but neither produces a cryptographic digest of file contents. An attacker or faulty process could alter an object and create a new version, and the DynamoDB record would not reveal that the bytes changed, so alteration detection and tamper evidence are not achieved.
- ✗
Enable S3 Object Lock in governance mode on the curated bucket and attach a retention period equal to the nightly baseline interval.
Why it's wrong here
S3 Object Lock governance mode prevents deletion or overwrite for a retention period, which protects immutability but does not create a verifiable provenance record of which Glue job run wrote a file, nor does it provide a digest to detect alteration. It also would block the nightly rewrite the pipeline depends on, so it does not satisfy the provenance requirement.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.