MLA-C01 Data Preparation for Machine Learning Practice Question
A machine learning team stores training data in an Amazon S3 bucket and wants to catalog it so that Amazon Athena and Amazon SageMaker Feature Store can discover the schema. The data is partitioned by year, month, and day in Hive-style prefixes, and new partitions are added daily. A data engineer must ensure new partitions are automatically discoverable without manual intervention. Which solution meets these requirements?
⚠ Common exam trap
The trap here is reaching for Athena's partition repair command, which works only on demand and still leaves the catalog stale between runs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create an AWS Glue crawler with a daily schedule that points at the S3 prefix and updates the Data Catalog.
A scheduled AWS Glue crawler is the standard mechanism for automatically discovering new Hive-style partitions in S3 and refreshing the Data Catalog. Because it runs on a daily schedule, new year/month/day partitions appear in the catalog without any manual command, and Athena and other catalog-aware services can query them immediately.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable S3 Inventory and point Athena at the inventory report to enumerate partitions.
Why it's wrong here
S3 Inventory produces a periodic CSV or Parquet listing of objects, but it is a reporting feature, not a Data Catalog updater. Athena cannot automatically translate an inventory manifest into table partitions, so the schema and partition metadata would still be missing from the Data Catalog for the consuming services.
- ✗
Run MSCK REPAIR TABLE on an Athena table each time new data arrives.
Why it's wrong here
MSCK REPAIR TABLE adds partitions to an existing Athena table definition, but it must be invoked after each new partition lands and requires the table to already exist with a matching schema. It is a manual, event-driven operation rather than an automated discovery mechanism, so it does not meet the no-manual-intervention requirement.
- ✗
Register the S3 prefix as a SageMaker Feature Store offline store and let it infer the partitions.
Why it's wrong here
SageMaker Feature Store requires data to be ingested through its feature group APIs, which write to a defined offline store layout. It does not crawl an arbitrary partitioned S3 prefix to infer Hive partitions, so this approach neither catalogs the existing data nor produces the Data Catalog entries Athena needs for discovery.
- ✓
Create an AWS Glue crawler with a daily schedule that points at the S3 prefix and updates the Data Catalog.
Why this is correct
A scheduled Glue crawler scans the S3 location, detects the Hive-style year/month/day prefixes as partition columns, and updates the Data Catalog tables automatically each day. This gives Athena and other integrated services an up-to-date schema and partition list with no manual steps, satisfying the automation requirement.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.