Courseiva

AI0-001 AI Concepts and Foundations Practice Question

A media company uses a reinforcement learning agent to schedule promotional banners on its homepage. The agent receives a reward when users click a banner, and it has learned to show the same sensational headline repeatedly because it historically generated high clicks. The editorial team is concerned that this harms long-term user trust. Which modification best aligns the agent's objective with long-term user satisfaction?

⚠ Common exam trap

The trap here is assuming that changing exploration or discount parameters will fix a misaligned objective, when the core issue is that the reward function itself rewards repetitive clickbait.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Add a reward penalty for showing the same banner repeatedly and include a signal for user retention or satisfaction.

The agent's behavior stems from a reward function that values clicks without regard for long-term consequences. Adding a repetition penalty and a retention or satisfaction signal reshapes the objective so the agent learns to avoid overexposing sensational content and to prioritize sustainable user engagement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the discount factor so future rewards are weighted more heavily.

    Why it's wrong here

    A higher discount factor makes the agent value future rewards more, but if the reward signal still measures only clicks, the agent will optimize for repeated sensational headlines even further into the future. The problem is the reward definition, not the horizon. This change alone does not incorporate trust or satisfaction, so it fails to address the editorial concern.

  • ✗

    Reduce the size of the experience replay buffer so the agent forgets older click patterns.

    Why it's wrong here

    Shrinking the replay buffer makes the agent rely on recent experiences, which can increase variance and may even reinforce the latest sensational headline. It does not change the reward function, so the agent still pursues clicks. This is a stability and sample-efficiency tuning change, not an objective alignment fix, and it could worsen the repetition problem.

  • ✗

    Switch from an epsilon-greedy policy to a softmax exploration policy.

    Why it's wrong here

    Exploration strategy affects how the agent discovers actions, but the underlying reward remains click-based, so the agent will still favor sensational headlines once it learns they yield clicks. Softmax exploration may vary banner selection during learning, yet the converged policy will reflect the same misaligned objective. This does not address long-term user trust.

  • ✓

    Add a reward penalty for showing the same banner repeatedly and include a signal for user retention or satisfaction.

    Why this is correct

    Redefining the reward to penalize repetition and include retention or satisfaction directly changes what the agent optimizes. This aligns the objective with long-term trust rather than short-term clicks. By incorporating both a repetition penalty and a satisfaction signal, the agent learns to balance engagement with sustainable user experience, which is exactly what the editorial team needs.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.