Courseiva
easyMultiple Choice

MLA-C01 Practice Question: A data scientist needs to label a large dataset…

A data scientist needs to label a large dataset of product images for a classification model. They want to reduce labeling costs by prioritizing uncertain samples. Which Amazon SageMaker Ground Truth feature should they use?

⚠ Common exam trap

It's easy for candidates to confuse 'active learning' with 'automated data labeling' — candidates often think automated labeling reduces costs by skipping humans entirely, but active learning specifically reduces costs by selectively using humans only on uncertain samples.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Active learning

Active learning in Amazon SageMaker Ground Truth automatically selects the most uncertain or informative samples from the unlabeled dataset and sends them to human annotators. This prioritization reduces labeling costs by focusing budget on the samples that will most improve model performance, rather than labeling all data indiscriminately.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Pre-built worker templates

    Why it's wrong here

    Pre-built worker templates only supply the task UI for common labelling jobs such as bounding boxes or semantic segmentation; they do not score model uncertainty or select which images get labelled. They would be the right choice when the labelling task type is already known and you simply need a ready-made interface.

  • ✓

    Active learning

    Why this is correct

    Active learning selects the samples the model is least confident about and sends only those for human labelling, so annotation effort concentrates on uncertain images rather than the whole dataset. This directly satisfies the requirement to reduce labelling costs by prioritising uncertain samples.

  • ✗

    Consolidated labeling

    Why it's wrong here

    Consolidated labeling merges annotations from multiple workers into one output; it does not select which samples get labelled. Active learning is the feature that queries uncertain or low-confidence samples, and it would be correct when the aim is cutting labelling cost through selective sampling.

  • ✗

    Automated data labeling

    Why it's wrong here

    Automated data labeling uses model predictions to label data, but it does not prioritise uncertain samples for human review. It suits reducing cost on high-confidence data, whereas active learning targets the uncertain samples the stem requires.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.