DEA-C01 Data Store Management Practice Question
A data engineering team is building a data lake on Amazon S3. They need to catalog data and make it queryable by Amazon Athena and Amazon Redshift Spectrum. The data arrives in multiple formats and the schema evolves frequently. Which TWO actions should the team take to support schema evolution and efficient querying? (Choose two.)
⚠ Common exam trap
The trap here is assuming that a single format or manual bucket separation solves schema evolution, when the Data Catalog plus columnar, partitioned storage is what enables both evolution and efficient querying.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Define tables in the AWS Glue Data Catalog and use AWS Glue crawlers to infer and update schemas as new data arrives.
The AWS Glue Data Catalog with crawlers keeps table and partition metadata current as schemas evolve, and both Athena and Redshift Spectrum read from it. Using open columnar formats such as Parquet or ORC with partitioned prefixes enables column projection and partition pruning, which cut scanned bytes and cost. Together, these two actions support evolving schemas and efficient querying across both engines.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Define tables in the AWS Glue Data Catalog and use AWS Glue crawlers to infer and update schemas as new data arrives.
Why this is correct
The AWS Glue Data Catalog is the central metastore used by Athena and Redshift Spectrum. Crawlers inspect data, infer schemas, and update table definitions, so new columns and partitions are reflected automatically. This supports frequent schema evolution without manual DDL, and both query engines can immediately use the updated metadata.
- ✗
Store all data as uncompressed CSV to maximize compatibility with all query engines.
Why it's wrong here
Uncompressed CSV is row-oriented and forces engines to read every column and row, increasing scan volume and cost. It also does not handle nested structures or evolving schemas well. This choice conflicts with efficient querying and is unnecessary since Athena and Spectrum support columnar formats.
- ✗
Create a separate S3 bucket for each schema version and require analysts to query the correct bucket manually.
Why it's wrong here
This fragments the data lake and forces analysts to know which bucket holds the current schema, which is error-prone and does not integrate with the Data Catalog. It also complicates partition management and increases storage duplication. Schema evolution is better handled in place with catalog updates and compatible formats.
- ✓
Use open columnar formats such as Apache Parquet or ORC with partition prefixes, and register partitions in the Data Catalog.
Why this is correct
Parquet and ORC are columnar, compressed, and self-describing, enabling column projection and predicate pushdown that reduce scanned bytes. Storing data under partition prefixes and registering those partitions in the Data Catalog allows Athena and Redshift Spectrum to prune irrelevant data. This combination supports efficient querying as schemas and partitions evolve.
- ✗
Enable S3 Object Lock on the data lake bucket to preserve schema versions.
Why it's wrong here
S3 Object Lock prevents deletion or overwriting of objects for a retention period; it does not manage or version schemas. It would also block legitimate updates to table metadata if applied incorrectly. Schema evolution is handled by the Data Catalog and format choices, not by object immutability.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.