A data scientist is building a text classification model using a pre-trained BERT model from the Hugging Face library on SageMaker. The scientist wants to fine-tune the model on a custom dataset. Which TWO steps are necessary to set up the fine-tuning job? (Select TWO.)
The HuggingFace estimator is the SageMaker-provided class that packages the training container, script and hyperparameters for Hugging Face models. It satisfies the requirement to set up a fine-tuning job for the pre-trained BERT model on a custom dataset.
Why this answer
Option A is correct because the HuggingFace estimator from the SageMaker Python SDK is the purpose-built, supported interface for launching Hugging Face training jobs on SageMaker; it handles the training container, script entry point, and hyperparameters needed to fine-tune a pre-trained BERT model. Option D is correct because the HuggingFace estimator lets you pin the exact framework and library versions via the transformers_version and pytorch_version (or tensorflow_version) arguments, which is essential for reproducibility and compatibility with the pre-trained BERT checkpoint. Option B is not needed because SageMaker Clarify provides bias and explainability analysis, not a prerequisite for fine-tuning.
Option C is unnecessary since the HuggingFace estimator already supplies a managed container with PyTorch and Transformers, so building a custom Docker image is only required for unsupported dependencies. Option E is not required because data preprocessing can be done inside the training script or beforehand; SageMaker Processing is optional and not a mandatory setup step for the fine-tuning job.
Exam trap
AWS often tests the misconception that custom Docker containers are required for any non-standard framework, but the HuggingFace estimator eliminates that need by providing a managed environment with version control.