easyMultiple Choice
AIF-C01 Practice Question: Which type of machine learning is used when a…
Which type of machine learning is used when a model learns to play a game by receiving rewards or penalties for its actions?
⚠ Common exam trap
AWS often tests the misconception that reinforcement learning is a type of supervised learning because both involve feedback, but the key trap is that supervised learning uses immediate, correct labels while reinforcement learning uses delayed, evaluative rewards or penalties without explicit correct answers.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Reinforcement learning
Reinforcement learning is correct because it involves an agent learning to make decisions by interacting with an environment, receiving rewards for desirable actions and penalties for undesirable ones. This trial-and-error approach, often modeled using Markov Decision Processes (MDPs), is specifically designed for sequential decision-making tasks like game playing, where the agent must maximize cumulative reward over time.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Reinforcement learning
Why this is correct
Reinforcement learning trains an agent through interaction with an environment, using rewards and penalties to shape behaviour. The agent learns a policy that maximises cumulative reward, which is precisely how game-playing models improve their actions over time.
- ✗
Semi-supervised learning
Why it's wrong here
Semi-supervised learning combines a small labelled set with a large unlabelled one to improve classification, typically via pseudo-labelling or consistency regularisation. Rewards and penalties are neither labels nor unlabelled examples, so this paradigm cannot represent the sequential feedback loop the scenario describes.
- ✗
Supervised learning
Why it's wrong here
Supervised learning maps inputs to known target labels supplied in advance, such as classifying images or predicting prices. A game agent receives no pre-labelled correct move; it must discover which actions yield reward through trial and error, which is reinforcement learning's credit-assignment problem.
- ✗
Unsupervised learning
Why it's wrong here
Unsupervised learning finds structure in unlabelled data through clustering, dimensionality reduction or density estimation; it has no reward signal, so an agent cannot learn which moves pay off. It would be the right choice for segmenting customers or compressing features, not for training a game-playing policy.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.