Courseiva

AI0-001 Implementing AI Solutions Practice Question

A media company is deploying an AI system that generates short news summaries from full articles. Before launch, the responsible AI review board asks the team to define monitoring that will detect harmful or degraded behavior in production. Which TWO monitoring practices should the team implement? (Choose two.)

⚠ Common exam trap

The trap here is selecting infrastructure or output-length metrics that feel like monitoring but do not actually detect harmful or factually degraded generated content.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Track hallucination and factual-consistency rates by automatically comparing generated summaries against source articles using a grounded entailment check.

Responsible deployment of a generative summarization system requires monitoring both automated factual grounding and structured human review. Grounding checks quantify hallucinations against source articles, while rubric-based human sampling catches bias, tone, and framing issues that metrics miss. Together they provide the safety and quality signals the review board needs. Infrastructure and length metrics do not measure harmful or degraded content behavior.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Record the GPU utilization of the inference cluster to confirm the model is running within expected compute bounds.

    Why it's wrong here

    GPU utilization is an infrastructure metric that supports cost and capacity management, not content safety. It cannot detect hallucinations, bias, or tone problems in generated summaries. Relying on it would leave the review board without any signal about whether the AI is producing harmful or inaccurate news content in production.

  • ✗

    Monitor the average token length of generated summaries to ensure outputs stay within the expected range.

    Why it's wrong here

    Token length is a weak proxy for quality and does not detect hallucinations, bias, or unsafe content. A summary can be the right length while being factually wrong or toxic. While length monitoring may catch truncation bugs, it does not satisfy the review board's request for detecting harmful or degraded behavior in a news summarization context.

  • ✗

    Track the number of API calls per minute to detect traffic spikes that could indicate abuse or system overload.

    Why it's wrong here

    Throughput monitoring is valuable for capacity and abuse detection but says nothing about the quality, safety, or factual accuracy of generated summaries. A system can be operating at normal traffic while producing harmful output. This metric does not address the review board's concern about detecting degraded or harmful behavior in the generated content itself.

  • ✓

    Track hallucination and factual-consistency rates by automatically comparing generated summaries against source articles using a grounded entailment check.

    Why this is correct

    Automated grounding checks compare each generated claim against the source article and flag unsupported statements, which directly measures hallucination risk in a summarization system. Tracking this rate over time reveals model or data drift, prompt regressions, and changes in source content that degrade factual fidelity. It gives the review board a concrete, ongoing safety signal rather than a one-time evaluation.

  • ✓

    Log and sample generated summaries for human review, with a rubric covering factual accuracy, bias, and tone, and route flagged items to an escalation queue.

    Why this is correct

    Structured human review with a defined rubric catches harms that automated metrics miss, including subtle bias, misleading framing, and tone issues. Sampling keeps the process tractable while escalation queues ensure flagged outputs are investigated and remediated. This creates an accountability loop that complements automated checks and gives the review board evidence of ongoing oversight.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.