AIF-C01 Fundamentals of AI and ML Practice Question
An e-commerce company stores user interaction logs in Amazon S3. They want to use machine learning to segment users based on purchasing behavior. Which unsupervised learning algorithm is most appropriate?
⚠ Common exam trap
The AIF-C01 exam often tests the distinction between supervised and unsupervised learning by presenting a clustering problem and including supervised algorithms as distractors, leading candidates to mistakenly pick a familiar algorithm like random forest or linear regression without recognizing the lack of labeled data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
K-means clustering
K-means clustering is the most appropriate unsupervised learning algorithm for segmenting users based on purchasing behavior because it groups data points into clusters based on feature similarity without requiring labeled training data. The e-commerce scenario involves discovering natural groupings (segments) in user interaction logs, which is a classic clustering task, and K-means efficiently partitions users into K distinct segments by minimizing within-cluster variance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Linear regression
Why it's wrong here
Linear regression predicts a continuous numeric target from labelled data, so it cannot discover the behavioural groupings segmentation requires. It would be correct for forecasting a value such as spend, not for clustering users by purchasing behaviour without labels.
- ✗
Random forest
Why it's wrong here
Random forest is a supervised classifier or regressor requiring labelled outcomes; it cannot uncover latent segments in unlabelled interaction logs. It would be the right choice for predicting a known label, such as churn, rather than clustering users by behaviour.
- ✓
K-means clustering
Why this is correct
K-means clustering partitions unlabelled interaction logs into k groups by minimising within-cluster variance, directly satisfying the requirement to segment users by purchasing behaviour without predefined labels. Unlike supervised methods, it needs no target variable, making it appropriate for discovering behavioural cohorts in the S3-stored data.
- ✗
Neural network
Why it's wrong here
Can be used but typically supervised; overkill for simple segmentation.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.