AI0-001 AI Implementation and Operations Practice Question
A hospital's clinical decision support model was validated at 94 percent accuracy on a held-out set. After go-live, clinicians report that the model's suggestions are frequently irrelevant for elderly patients, even though overall accuracy in the monitoring dashboard has barely moved. Which monitoring practice would have surfaced this problem?
⚠ Common exam trap
The trap here is trusting a stable top-line accuracy number as proof of health, when a small or distinct subgroup can fail badly while the average barely moves.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Tracking performance metrics disaggregated by patient demographic and clinical subgroups, with alerts when any subgroup falls below its own threshold.
The clue is that aggregate accuracy held steady while a specific patient group received poor suggestions, which is the classic signature of a subgroup masked by population averaging. Disaggregated performance monitoring with per-segment thresholds is the practice that reveals it. Aggregate accuracy alerts, latency metrics, and input drift detection each observe a different dimension and cannot confirm that elderly patients specifically are being served badly.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Measuring inference latency and throughput to confirm the model responds within clinical workflow time limits.
Why it's wrong here
Latency and throughput describe operational health, not predictive correctness. A model can answer instantly and still give irrelevant suggestions. These metrics would be relevant to a performance incident, but they say nothing about whether predictions are accurate for a particular patient group.
- ✓
Tracking performance metrics disaggregated by patient demographic and clinical subgroups, with alerts when any subgroup falls below its own threshold.
Why this is correct
A subgroup that is a small fraction of the population can degrade sharply while the aggregate metric stays flat, which matches the reported pattern. Segmenting metrics by age and other clinical attributes makes that hidden weakness visible. Subgroup thresholds then trigger action before clinicians lose trust in the system.
- ✗
Monitoring overall model accuracy and alerting when it drops more than two percentage points from the validation baseline.
Why it's wrong here
This is precisely the aggregate monitoring the scenario says did not detect the problem. A decline confined to elderly patients is diluted by the larger population and stays under the alert threshold. The metric is measured correctly but at the wrong granularity to expose subgroup failure.
- ✗
Comparing the live distribution of input features against the training distribution to detect covariate shift.
Why it's wrong here
Input drift monitoring can raise a flag if elderly patient records differ from training data, but it does not measure whether predictions for them are correct. Shift can occur without any accuracy loss, and accuracy can degrade with no input shift. It is an indirect indicator rather than the outcome-based check this situation requires.
About these practice questions
One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.