A startup is developing a mobile app that uses facial recognition to verify user identity for account access. The app is intended for a global audience, but the training data predominantly includes images of light-skinned individuals. During beta testing, users with darker skin tones report frequent verification failures, while light-skinned users have a high success rate. The startup wants to release the app soon and needs to address this fairness issue without delaying the launch too much. The team has limited resources. Which approach should they take to most effectively mitigate the bias while meeting the launch timeline?
Retraining on augmented, demographically diverse data corrects the underlying representation imbalance causing disparate error rates across skin tones. This directly addresses the root cause of the fairness failure while remaining feasible within the startup's limited resources and launch timeline.
Why this answer
The root cause of the bias is a skewed training dataset that underrepresents darker skin tones. Collecting more diverse data and augmenting the existing dataset directly addresses the data imbalance, allowing the facial recognition model to learn robust features for all skin tones. Retraining the model on this enriched dataset is the most effective long-term fix that aligns with responsible AI principles, and with focused effort it can be completed within a reasonable timeline without introducing the risks of post-hoc patches.
Exam trap
AWS often tests the misconception that a quick operational fix (like adjusting thresholds or adding manual review) can effectively solve algorithmic bias, when in fact the only principled solution is to address the data imbalance at the source.
How to eliminate wrong answers
Option A is wrong because applying a post-processing rule to artificially increase acceptance rates for darker-skinned users does not fix the underlying model bias; it merely masks the problem and can lead to higher false acceptance rates, undermining security. Option B is wrong because lowering the similarity threshold for all users would increase false positives across the board, reducing the overall security of the verification system without addressing the specific failure mode for darker skin tones. Option C is wrong because deferring verification to manual human review for darker-skinned users creates a separate, slower process that introduces user friction, scales poorly with limited resources, and does not resolve the model's inherent bias.