AI0-001 AI Models and Data Engineering Practice Question
A medical imaging team is developing an AI model to detect tumors from CT scans. They have 10,000 labeled scans, but the labels were created by a semi-automated process with an estimated 20% error rate (mislabeled tumor vs. no tumor). The team trains a convolutional neural network (CNN) and achieves 90% accuracy on a held-out test set that was carefully validated by an expert radiologist. However, when deployed to a new hospital's patient population, the accuracy drops to 70%. The team suspects domain shift and label noise. Which strategy is most likely to improve model robustness for the new hospital?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use active learning to select the most uncertain predictions from the new hospital's data, then have an expert radiologist correct those labels
Active learning selects the most uncertain predictions from the new hospital's data, allowing an expert radiologist to efficiently correct the most informative labels. This directly addresses both label noise (by correcting mislabeled examples) and domain shift (by focusing on samples where the model is uncertain in the new domain). Option B is wrong because random selection may not target the most impactful errors, wasting expert effort. Option C is wrong because adding more noisy labels from the same flawed process will amplify label noise without correcting the domain shift. Option D is wrong because reducing model complexity and dropout are regularization techniques that do not fix label noise or domain shift.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use active learning to select the most uncertain predictions from the new hospital's data, then have an expert radiologist correct those labels
Why this is correct
Active learning targets the new hospital's own distribution, and expert correction removes label noise in exactly those uncertain cases. This adapts the model to the shifted domain while fixing the 20% labelling error, addressing both suspected causes.
- ✗
Randomly select 1,000 scans from the new hospital and have them re-labeled by the radiologist
Why it's wrong here
Re-labelling 1,000 new-hospital scans corrects the target-domain labels but leaves the 20% noise in the 10,000 training scans, so the model still learns corrupted patterns. It is tempting because domain adaptation via local labels helps, yet label-noise correction requires cleaning or re-labelling the training set itself.
- ✗
Collect 20,000 more scans with the same semi-automated labeling process
Why it's wrong here
Adding 20,000 scans through the same semi-automated pipeline amplifies the 20% mislabelling rate and introduces no new-hospital images, leaving both label noise and domain shift untouched. This approach suits scaling a dataset whose labels are already reliable and drawn from the target domain.
- ✗
Reduce the CNN's number of layers and apply dropout to combat overfitting
Why it's wrong here
Shrinking the CNN and adding dropout addresses variance within the training distribution, not the covariate shift between hospitals or the 20% label noise. Dropout is a regularisation choice for overfitting on clean, same-domain data; it cannot correct systematically corrupted targets or scanner differences.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.