Courseiva

AIF-C01 Fundamentals of AI and ML Practice Question

Which TWO services can be used to preprocess data for machine learning in AWS? (Choose two.)

⚠ Common exam trap

The AIF-C01 exam often tests the distinction between data querying services (like Athena) and data preprocessing services, leading candidates to mistakenly choose Athena because it can 'process' data via SQL, but it lacks the ML-specific transformation capabilities required for preprocessing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

AWS Glue

AWS Glue (A) is correct because it is a fully managed, serverless ETL service that can run Apache Spark jobs and Glue Studio/DynamicFrames to clean, transform, join, and normalize raw data into feature-ready datasets for ML training. Amazon SageMaker Data Wrangler (C) is correct because it is purpose-built for ML data preparation, providing a visual interface with 300+ built-in transforms, data quality/insight reports, and one-click export to SageMaker pipelines or feature stores. Amazon Athena (B) is a serverless interactive query service over S3 using SQL, suited for ad hoc analytics rather than building reusable preprocessing pipelines. Amazon Redshift (D) is a data warehouse for analytics and SQL workloads, not a dedicated ML preprocessing service. AWS Lambda (E) is a general-purpose serverless compute service that could run small transformation code but is not a data preprocessing service for ML.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    AWS Glue

    Why this is correct

    AWS Glue is a serverless ETL service that cleanses, transforms and catalogues data at scale using Spark jobs or Glue Studio, producing prepared datasets for ML training. It satisfies the preprocessing requirement by handling transformation and data quality tasks before model training.

  • ✗

    Amazon Athena

    Why it's wrong here

    Athena queries data in place across S3 using SQL, returning results rather than performing reusable transformation and feature-engineering steps for ML datasets. It is tempting because SQL can reshape data, but Athena is an interactive query service, not a preprocessing pipeline.

  • ✓

    Amazon SageMaker Data Wrangler

    Why this is correct

    SageMaker Data Wrangler provides visual, low-code data preparation with over 300 built-in transforms, plus bias and anomaly analysis, exporting directly to SageMaker pipelines. It satisfies the preprocessing requirement by importing, transforming and featurising data before training.

  • ✗

    Amazon Redshift

    Why it's wrong here

    Redshift is a petabyte-scale data warehouse for storing and analysing structured data, not a preprocessing engine for preparing ML datasets. It is tempting because warehouses often hold the source data, but transformation and feature preparation occur in services such as SageMaker Data Wrangler or Glue.

  • ✗

    AWS Lambda

    Why it's wrong here

    Lambda runs event-driven code and can transform records, but it is compute rather than a purpose-built data preparation service with built-in transforms and catalogued datasets. It is tempting because Lambda commonly triggers preprocessing pipelines, yet the question asks for services that perform the preprocessing itself.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.