Courseiva

AI0-001 Implementing AI Solutions Practice Question

During testing of a customer service chatbot, the team notices that the model sometimes generates plausible-sounding but factually incorrect answers about company policies. Which evaluation approach is BEST to systematically detect and quantify this issue?

⚠ Common exam trap

AI0-001 often tests the confusion between general software testing types (unit, integration, regression) and AI-specific evaluation metrics, tricking candidates into choosing familiar testing terminology over purpose-built LLM evaluation approaches.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Evaluation framework with faithfulness and answer relevancy metrics on a held-out test set

An evaluation framework with faithfulness and answer relevancy metrics on a held-out test set is the best approach because it directly measures whether the model's output is grounded in the provided source (faithfulness) and whether it actually addresses the user's question (answer relevancy). These are the standard RAG/LLM evaluation metrics designed to detect hallucinations and off-topic answers systematically. A held-out test set ensures the measurement is repeatable and quantifiable across model versions.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Regression testing comparing old and new model outputs

    Why it's wrong here

    Regression testing compares old and new model outputs, so it detects changed behaviour, not whether either output is factually correct. It suits guarding against unintended drift after a model update. Quantifying hallucinated policy claims needs scoring against a curated ground-truth dataset, which regression diffs cannot supply.

  • ✗

    Unit tests on the data pipeline

    Why it's wrong here

    Unit tests on the data pipeline check ingestion, transformation and formatting logic, never the semantic truth of generated answers. They are correct for validating that training or retrieval data is parsed and shaped as expected. Detecting fabricated policy statements requires evaluating model outputs against authoritative reference answers.

  • ✗

    Integration tests for API calls

    Why it's wrong here

    Integration tests verify that API calls connect and return responses; they assert plumbing, not factual accuracy of generated policy text. They would be correct for confirming a chatbot's endpoints, authentication and error handling work end to end. Detecting plausible-sounding falsehoods requires grounded evaluation against reference policy answers.

  • ✓

    Evaluation framework with faithfulness and answer relevancy metrics on a held-out test set

    Why this is correct

    Faithfulness metrics measure whether each claim in the generated answer is grounded in the retrieved context, while answer relevancy scores how well it addresses the question. Running both over a held-out test set systematically quantifies hallucinated policy statements.

About these practice questions

One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.