Courseiva

AI0-001 AI Concepts and Techniques Practice Question

Which machine learning paradigm involves training an agent to make decisions by interacting with an environment and receiving rewards or penalties based on its actions?

⚠ Common exam trap

CompTIA AI often tests the distinction between reinforcement learning and supervised learning by phrasing the question to emphasize 'rewards or penalties' — candidates mistakenly think supervised learning uses penalties (like loss functions) and confuse it with RL's delayed reward signals from an environment.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Reinforcement learning

Reinforcement learning (RL) is the correct paradigm because it explicitly involves an agent learning a policy through trial-and-error interactions with an environment, receiving scalar reward signals (positive or negative) to maximize cumulative reward. This matches the question's description of making decisions based on rewards or penalties, which is the defining characteristic of RL, as opposed to learning from labeled data or discovering hidden patterns without feedback.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Unsupervised learning

    Why it's wrong here

    Unsupervised learning finds structure in unlabelled data through clustering or dimensionality reduction, with no environment, actions or reward signal involved. It is tempting because it also learns without labelled targets, and it would be correct for segmenting customers or detecting anomalies, but reinforcement learning's agent-environment feedback loop is absent here.

  • ✓

    Reinforcement learning

    Why this is correct

    Reinforcement learning trains an agent through trial-and-error interaction with an environment, using reward and penalty signals to shape a policy. This directly matches the stem's requirement for decisions driven by rewards or penalties, unlike supervised or unsupervised paradigms that learn from static labelled or unlabelled datasets.

  • ✗

    Supervised learning

    Why it's wrong here

    Supervised learning maps labelled input–output pairs to predict known targets, so it cannot learn from reward signals generated by an agent's own actions. It is tempting because it excels at classification and regression where historical labelled data exists, such as predicting customer churn, but reinforcement learning is required when no correct action labels are available.

  • ✗

    Self-supervised learning

    Why it's wrong here

    Self-supervised learning generates supervisory labels from the data's own structure, such as predicting masked tokens, so it never involves an agent, actions, or reward signals from an environment. It is tempting because it also learns without manual labels, and it would be the right choice for pre-training representations on large unlabelled text or image corpora.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.