Courseiva
easyMultiple Choice

AIF-C01 Is an example of reinforcement learning? Practice Question

Which of the following is an example of reinforcement learning?

⚠ Common exam trap

The AWS AI Practitioner exam often tests the distinction between supervised, unsupervised, and reinforcement learning by presenting tasks that involve feedback (like rewards) versus tasks that use labeled data or no labels, and the trap here is that candidates may confuse any task with a 'goal' or 'outcome' as reinforcement learning, even when it uses pre-labeled examples or historical data without an interactive environment.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

A robot learning to navigate a maze by receiving rewards for reaching the goal

Reinforcement learning involves an agent learning to make decisions by interacting with an environment and receiving rewards or penalties for its actions. Option D describes a robot learning to navigate a maze by receiving rewards for reaching the goal, which is a classic example of an agent optimizing its policy through trial and error to maximize cumulative reward.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Predicting house prices based on historical data

    Why it's wrong here

    Predicting house prices fits a regression model to historical labelled data, which is supervised learning with no agent, action or reward. It is tempting because predictions improve as more data arrives, and it would be correct if the question asked for a supervised regression example.

  • ✗

    Grouping customers into segments based on purchasing behavior

    Why it's wrong here

    Grouping customers into segments is unsupervised clustering, which learns structure from unlabelled data with no reward signal, agent or environment. It is tempting because segmentation also improves through iteration, and it would be the correct answer if the question asked for an unsupervised learning example.

  • ✗

    Detecting spam emails using labeled examples

    Why it's wrong here

    Spam detection trains a classifier on labelled examples, which is supervised learning, not trial-and-error reward maximisation. It is tempting because spam filters do adapt over time, and it would be the correct choice if the question asked for a supervised classification example.

  • ✓

    A robot learning to navigate a maze by receiving rewards for reaching the goal

    Why this is correct

    A robot learning to navigate a maze by receiving rewards for reaching the goal exemplifies reinforcement learning: an agent takes actions within an environment and learns a policy from reward signals, rather than labelled examples. This satisfies the stem's requirement for trial-and-error learning driven by cumulative reward maximisation.

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.