Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A company ingests daily log data into an S3 bucket. They need to update the existing ML training dataset with new data without reprocessing the entire history. Which two strategies should they adopt? (Choose two.)

⚠ Common exam trap

AWS often tests the misconception that S3 supports append operations or that simply copying new files to the same bucket constitutes an incremental update strategy, when in reality S3 objects are immutable and a proper processing framework like AWS Glue with job bookmarks is required.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use AWS Glue to incrementally process new partitions

Option B is correct because AWS Glue can perform incremental processing by reading only newly added partitions (using job bookmarks and partition predicates) rather than reprocessing the full history, which directly satisfies the requirement to update the dataset without re-running the entire pipeline. Option C is correct because partitioning the S3 data by a key such as date lets new daily log data land in a new partition, so downstream training jobs or Glue ETL can scan only the new partition and append it to the existing dataset. Option A is incorrect because S3 objects are immutable and cannot be appended to in place; a single large file would require rewriting the whole object. Option D is incorrect because manually copying files into the same bucket does not provide incremental processing logic or partition awareness, and is error-prone. Option E is incorrect because overwriting the entire dataset forces full reprocessing, which is exactly what the scenario wants to avoid.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Store all data in a single large file and use append operations

    Why it's wrong here

    S3 objects are immutable, so append operations are impossible; each update rewrites the entire large file, which is exactly the full reprocessing the scenario forbids. It is tempting because a single file simplifies downstream reads, and appending would be correct for a mutable store such as a database or log file.

  • ✓

    Use AWS Glue to incrementally process new partitions

    Why this is correct

    AWS Glue incrementally processes only newly arrived S3 partitions, satisfying the requirement to update the training dataset without reprocessing the entire history. By tracking partition metadata in the Data Catalog, Glue reads just the fresh daily logs, cutting compute cost and runtime compared with full-dataset reprocessing.

  • ✓

    Use a partition key such as date to add new partitions

    Why this is correct

    Partitioning S3 data by date lets downstream jobs target only the newest partition prefix. New daily logs land in their own partition, so the existing training dataset is extended without scanning or reprocessing historical data.

  • ✗

    Manually copy new files to the same S3 bucket

    Why it's wrong here

    Copying files manually into the same prefix gives no incremental ingestion mechanism, so the training pipeline cannot detect which objects are new and must rescan the whole bucket. It is tempting because S3 is the correct storage layer, but manual copying suits one-off uploads, not scheduled incremental dataset refreshes.

  • ✗

    Overwrite the entire existing dataset with the new data

    Why it's wrong here

    Overwriting discards all historical data, so the model loses prior examples and cannot learn from the full history. It is tempting when only the newest data matters, but incremental strategies such as appending new partitions or merging with existing datasets preserve history without full reprocessing.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.