How to Mitigate Bias in Machine Learning Models Responsibly
Exhibit
Refer to the exhibit.
```
{
"ModelName": "credit-risk-v1",
"InputName": "features",
"JobName": "bias-report-20240101",
"ProcessingJob": {
"ProcessingResources": {
"ClusterConfig": {
"InstanceCount": 1,
"InstanceType": "ml.m5.large"
}
}
},
"AppSpecification": {
"ImageUri": "683313688378.dkr.ecr.us-west-2.amazonaws.com/sagemaker-clarify-processing:1.0"
},
"Config": {
"BiasConfig": {
"Label": "approved",
"Facet": ["gender"],
"GroupVariable": ["age_group"]
}
},
"OutputConfig": {
"S3OutputPath": "s3://my-bucket/bias-reports/"
}
}A data scientist runs the SageMaker Clarify job shown in the exhibit for a credit risk model. After reviewing the results, they find a high bias metric for the gender facet. Which action is most consistent with responsible AI?
Quick Answer
The answer is to investigate the root cause and retrain with balanced data. This is correct because responsible AI demands mitigating bias at its source rather than superficially masking it; SageMaker Clarify detects statistical bias, but only a root-cause analysis—checking for skewed data collection, labeling errors, or proxy features like zip codes—allows you to retrain with balanced data that genuinely reduces the bias metric. On the AWS Certified AI Practitioner AIF-C01 exam, this scenario tests your understanding that fairness is a core principle of responsible AI, and a common trap is assuming removing a sensitive attribute like gender eliminates bias, when in reality correlated proxies can perpetuate it. Remember the mnemonic: “Don’t delete the feature, dig for the source.”
⚠ Common exam trap
The AIF-C01 exam often tests the misconception that simply removing a sensitive attribute (like gender) is sufficient to eliminate bias, but the trap here is that proxy features can still encode the same bias, making root-cause investigation and balanced retraining the only responsible action.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Investigate the root cause and retrain with balanced data
Responsible AI requires understanding and mitigating bias at its source, not just masking it. Investigating the root cause (e.g., data collection bias, labeling bias, or proxy features) and retraining with balanced data directly addresses the high bias metric detected by SageMaker Clarify, aligning with AWS's principle of fairness. Simply removing the gender attribute may not eliminate bias if other features act as proxies, and increasing the threshold does not fix the underlying model bias.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Proceed with deployment because the model is already in production
Why it's wrong here
Deploying despite a high bias metric ignores the finding and perpetuates disparate impact on the gender facet. Responsible AI practise requires investigating and mitigating the bias, for example by rebalancing training data or adjusting thresholds, before releasing the model.
- ✗
Remove the gender attribute from the training data and retrain
Why it's wrong here
Dropping the gender attribute does not remove bias, since correlated proxies such as postcode or occupation still encode it, and the Clarify job would simply report a different facet. It is tempting because removing a protected attribute feels like anonymisation, and it would be reasonable where the attribute is genuinely irrelevant to the target and no proxy correlation exists.
- ✓
Investigate the root cause and retrain with balanced data
Why this is correct
SageMaker Clarify's bias metric indicates disparate impact across gender, so investigating the root cause and retraining with balanced data addresses the underlying skew. This satisfies responsible AI by remediating the model rather than merely documenting or ignoring the measured bias.
- ✗
Increase the acceptance threshold for the model
Why it's wrong here
Raising the acceptance threshold shifts the decision boundary to cut approvals, altering error rates without addressing the disparate impact measured across the gender facet. It is tempting because threshold tuning is a legitimate lever for balancing precision and recall, and would be correct if the goal were purely to reduce false positives rather than mitigate bias.
Go deeper
Related to this question
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
Same concept, more angles
3 more ways this is tested on AIF-C01
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. Refer to the exhibit. An AWS customer runs SageMaker Clarify to evaluate bias in their training data. The report shows multiple metrics with status 'violated'. What should the customer do next?
medium- ✓ A.Use data augmentation to balance the dataset
- B.Reduce the number of features
- C.Retrain the model with more data
- D.Ignore the metrics because thresholds are too strict
Why A: SageMaker Clarify bias metrics such as Class Imbalance (CI) or Difference in Positive Proportions in Labels (DPPL) flag potential bias in the training data or model predictions. When a report shows violations, the next step is to review the findings and apply a targeted mitigation. Among the options, balancing the dataset through data augmentation directly addresses the representative imbalance; reducing features or simply adding more data does not target the demographic imbalance, and ignoring the metrics is not appropriate.
Variation 2. A data science team is building a resume screening model and wants to ensure it does not exhibit gender bias. Which TWO actions are most effective for mitigating bias? (Choose TWO.)
easy- ✓ A.Apply adversarial debiasing techniques during training.
- B.Use a more complex deep learning model.
- C.Remove the gender attribute and all correlated features from the dataset.
- ✓ D.Regularly audit model predictions for disparate impact across genders.
- E.Ensure the training dataset has equal numbers of male and female candidates.
Why A: Option A is correct because adversarial debiasing trains a predictor alongside an adversary that tries to detect the protected attribute (gender) from the predictor's outputs, forcing the model to learn representations that cannot discriminate by gender, which directly mitigates bias during training. Option D is correct because regularly auditing model predictions for disparate impact across genders (e.g., using metrics like demographic parity, equal opportunity, or the 80% rule) detects bias that may persist or emerge after deployment and enables corrective action. Option B is not correct because increasing model complexity with deep learning does not inherently reduce bias and can even amplify it by fitting spurious correlations in the data. Option C is not correct because simply removing the gender attribute and correlated features does not eliminate bias, since proxy variables and historical patterns can still encode gender information. Option E is not correct because equal representation of male and female candidates in the training set does not guarantee fairness, as bias can arise from label imbalance, feature correlations, or unequal outcomes despite balanced group sizes.
Variation 3. A healthcare organization is developing a clinical decision support system using Amazon Bedrock with a large language model (LLM) to analyze patient symptoms and suggest potential diagnoses. The system must comply with HIPAA and internal responsible AI guidelines. During testing, the model occasionally generates diagnoses that are inconsistent with established medical guidelines and shows a tendency to recommend more aggressive treatments for patients from certain demographic groups. The team has already implemented data encryption, access controls, and basic content filtering. They need to further reduce biased and unsafe outputs without delaying the deployment timeline. What should the team do next?
hard- A.Increase the logging of all model inputs and outputs to Amazon CloudWatch and set up alarms for any mentions of protected attributes.
- B.Replace the current LLM with a different pre-trained model that has been benchmarked for lower bias on medical datasets.
- ✓ C.Fine-tune the model using a curated dataset of anonymized patient records that is balanced across demographic groups and aligned with clinical guidelines.
- D.Apply stronger content filtering rules using Amazon Comprehend Medical to block any diagnosis that contains demographic-related terms.
Why C: Fine-tuning the model with a balanced, curated dataset directly addresses both the bias and clinical accuracy issues at the model level, which is the most effective approach for reducing biased and unsafe outputs without delaying deployment. This method adjusts the model's internal weights to align with established medical guidelines and demographic fairness, rather than relying on post-processing filters or logging that do not fix the root cause. Since the team has already implemented basic content filtering, fine-tuning provides a targeted, efficient solution that can be completed within a reasonable timeline.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.