AIF-C01 Fundamentals of AI and ML Practice Question
Which TWO services can be used to preprocess data for machine learning in AWS? (Choose two.)
⚠ Common exam trap
The AIF-C01 exam often tests the distinction between data querying services (like Athena) and data preprocessing services, leading candidates to mistakenly choose Athena because it can 'process' data via SQL, but it lacks the ML-specific transformation capabilities required for preprocessing.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
AWS Glue
AWS Glue (A) is correct because it is a fully managed, serverless ETL service that can run Apache Spark jobs and Glue Studio/DynamicFrames to clean, transform, join, and normalize raw data into feature-ready datasets for ML training. Amazon SageMaker Data Wrangler (C) is correct because it is purpose-built for ML data preparation, providing a visual interface with 300+ built-in transforms, data quality/insight reports, and one-click export to SageMaker pipelines or feature stores. Amazon Athena (B) is a serverless interactive query service over S3 using SQL, suited for ad hoc analytics rather than building reusable preprocessing pipelines. Amazon Redshift (D) is a data warehouse for analytics and SQL workloads, not a dedicated ML preprocessing service. AWS Lambda (E) is a general-purpose serverless compute service that could run small transformation code but is not a data preprocessing service for ML.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
AWS Glue
Why this is correct
AWS Glue is a serverless ETL service that cleanses, transforms and catalogues data at scale using Spark jobs or Glue Studio, producing prepared datasets for ML training. It satisfies the preprocessing requirement by handling transformation and data quality tasks before model training.
- ✗
Amazon Athena
Why it's wrong here
Athena queries data in place across S3 using SQL, returning results rather than performing reusable transformation and feature-engineering steps for ML datasets. It is tempting because SQL can reshape data, but Athena is an interactive query service, not a preprocessing pipeline.
- ✓
Amazon SageMaker Data Wrangler
Why this is correct
SageMaker Data Wrangler provides visual, low-code data preparation with over 300 built-in transforms, plus bias and anomaly analysis, exporting directly to SageMaker pipelines. It satisfies the preprocessing requirement by importing, transforming and featurising data before training.
- ✗
Amazon Redshift
Why it's wrong here
Redshift is a petabyte-scale data warehouse for storing and analysing structured data, not a preprocessing engine for preparing ML datasets. It is tempting because warehouses often hold the source data, but transformation and feature preparation occur in services such as SageMaker Data Wrangler or Glue.
- ✗
AWS Lambda
Why it's wrong here
Lambda runs event-driven code and can transform records, but it is compute rather than a purpose-built data preparation service with built-in transforms and catalogued datasets. It is tempting because Lambda commonly triggers preprocessing pipelines, yet the question asks for services that perform the preprocessing itself.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.