Courseiva
Fundamentals of AI and ML →mediumMultiple Select

SageMaker Ground Truth Capabilities: Workflows, Workforce & Active Learning

Which THREE statements about Amazon SageMaker Ground Truth are correct? (Choose three.)

Quick Answer

The answer is that Amazon SageMaker Ground Truth integrates with Amazon SageMaker to use the labeled data for training, which is a core capability tested on the AWS Certified AI Practitioner AIF-C01 exam. This is correct because Ground Truth is designed as a data labeling service that feeds directly into SageMaker’s model training pipelines; once a labeling job is complete, the output dataset is automatically stored in an S3 bucket and can be immediately referenced by a SageMaker training job without manual data transfer. On the exam, this question tests your understanding of how Ground Truth’s built-in workflows—such as those for image classification and object detection—simplify the labeling process by providing pre-built UI templates, while a common trap is confusing Ground Truth with a standalone labeling tool that does not integrate with SageMaker. To remember this, think of Ground Truth as the “labeling engine” that powers SageMaker’s training: it creates the high-quality labeled data that SageMaker consumes, not just a separate annotation service.

⚠ Common exam trap

AWS often tests the misconception that Ground Truth is limited to text data or only supports public workforces, while in reality it handles multiple data modalities and offers flexible workforce options including private and vendor-managed.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

It provides built-in workflows for image classification and object detection.

Option B is correct because Amazon SageMaker Ground Truth ships with built-in labeling task templates and workflows for common job types, including image classification and object detection (as well as text, semantic segmentation, and bounding boxes). Option C is correct because Ground Truth supports automated data labeling, which uses active learning to have a model label high-confidence data while sending only low-confidence samples to human labelers, reducing cost and effort. Option D is correct because Ground Truth integrates directly with Amazon SageMaker, writing labeled datasets to Amazon S3 in augmented manifest format that SageMaker training jobs can consume. Option A is wrong because Ground Truth handles images, video, text, and 3D point clouds, not just text. Option E is wrong because Ground Truth supports multiple workforces, including private workforces, vendor-managed workforces, and Amazon Mechanical Turk, so it is not limited to a public Mechanical Turk workforce.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    It can only be used for text data.

    Why it's wrong here

    Ground Truth handles images, video, audio, text and 3D point clouds through built-in labelling task types, so a text-only limitation is false. A text-only tool would fit narrow NLP labelling projects, but Ground Truth's purpose is multi-modal dataset annotation at scale.

  • ✓

    It provides built-in workflows for image classification and object detection.

    Why this is correct

    Ground Truth ships managed labelling interfaces and task templates for image classification and object detection, so teams avoid building custom annotation tooling. This built-in workflow support satisfies the stem's requirement that the statement describe a genuine Ground Truth capability.

  • ✓

    It supports automated data labeling using active learning.

    Why this is correct

    Ground Truth uses active learning: a model labels the confident examples automatically and routes only low-confidence items to human labellers, cutting cost. This automated labelling mechanism is a genuine Ground Truth feature, satisfying the stem's requirement.

  • ✓

    It integrates with Amazon SageMaker to use the labeled data for training.

    Why this is correct

    Ground Truth writes labelled datasets directly into Amazon S3 in augmented manifest format, which SageMaker training jobs consume natively. This tight integration lets the labelled output feed model training without custom transformation, satisfying the stem's requirement.

  • ✗

    It can only use a public workforce from Amazon Mechanical Turk.

    Why it's wrong here

    Ground Truth supports private workforces, its own vendor-managed workforce, and third-party vendors alongside Mechanical Turk, so restricting it to the public MTurk pool is false. MTurk alone would suit only low-sensitivity, crowdsourced labelling where data confidentiality is not a constraint.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on AIF-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company wants to use Amazon SageMaker Ground Truth to build a labeled dataset for a custom object detection model. Which TWO labeling strategies are available? (Choose two.)

medium
  • ✓ A.Private workforce labeling (company employees)
  • ✓ B.Crowd-based labeling using Amazon Mechanical Turk
  • C.Automated labeling using pre-trained models
  • D.Active learning with manual verification
  • E.Fully automated labeling via AWS Lambda

Why A: Option A is correct because SageMaker Ground Truth supports a private workforce, where the company's own employees (or a private vendor workforce) perform the labeling tasks through the Ground Truth console, which is a standard strategy for sensitive or proprietary datasets. Option B is correct because Ground Truth also integrates with Amazon Mechanical Turk, allowing crowd-based labeling by a large, distributed public workforce for scalable annotation. Option C is not one of the two primary workforce-based labeling strategies asked for here; automated labeling in Ground Truth is a feature (auto-labeling) but the question targets the available labeling workforce strategies. Option D is not a distinct labeling strategy offered as a standalone choice; active learning is a Ground Truth feature that routes uncertain items to humans, not a separate labeling strategy option in this context. Option E is incorrect because fully automated labeling via AWS Lambda is not a Ground Truth labeling strategy; Lambda can be used for custom pre/post-processing, but it does not replace the human labeling workforce options.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.