1
Guidelines for Responsible AI
hard
A hospital uses an Amazon SageMaker model to predict sepsis risk from electronic health records and displays a risk score to clinicians. An internal review finds that the model was trained on data from a single urban hospital and performs worse for patients from rural clinics. The review board asks the data science team to quantify and document this performance gap across patient subgroups before the model is expanded. Which approach should the team take?