easyMultiple Choice
AIF-C01 Is an example of reinforcement learning? Practice Question
Which of the following is an example of reinforcement learning?
⚠ Common exam trap
The AWS AI Practitioner exam often tests the distinction between supervised, unsupervised, and reinforcement learning by presenting tasks that involve feedback (like rewards) versus tasks that use labeled data or no labels, and the trap here is that candidates may confuse any task with a 'goal' or 'outcome' as reinforcement learning, even when it uses pre-labeled examples or historical data without an interactive environment.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A robot learning to navigate a maze by receiving rewards for reaching the goal
Reinforcement learning involves an agent learning to make decisions by interacting with an environment and receiving rewards or penalties for its actions. Option D describes a robot learning to navigate a maze by receiving rewards for reaching the goal, which is a classic example of an agent optimizing its policy through trial and error to maximize cumulative reward.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Predicting house prices based on historical data
Why it's wrong here
Predicting house prices fits a regression model to historical labelled data, which is supervised learning with no agent, action or reward. It is tempting because predictions improve as more data arrives, and it would be correct if the question asked for a supervised regression example.
- ✗
Grouping customers into segments based on purchasing behavior
Why it's wrong here
Grouping customers into segments is unsupervised clustering, which learns structure from unlabelled data with no reward signal, agent or environment. It is tempting because segmentation also improves through iteration, and it would be the correct answer if the question asked for an unsupervised learning example.
- ✗
Detecting spam emails using labeled examples
Why it's wrong here
Spam detection trains a classifier on labelled examples, which is supervised learning, not trial-and-error reward maximisation. It is tempting because spam filters do adapt over time, and it would be the correct choice if the question asked for a supervised classification example.
- ✓
A robot learning to navigate a maze by receiving rewards for reaching the goal
Why this is correct
A robot learning to navigate a maze by receiving rewards for reaching the goal exemplifies reinforcement learning: an agent takes actions within an environment and learns a policy from reward signals, rather than labelled examples. This satisfies the stem's requirement for trial-and-error learning driven by cumulative reward maximisation.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.