A company is training a deep learning model on Amazon SageMaker using a custom Docker container. The training job fails with the error 'CannotStartContainerError: API error (500): failed to create shim task'. The team verifies that the container image is compatible with the selected instance type. What is the most likely cause of this error?
Trap 1: The training data is stored in the wrong S3 bucket
S3 permissions or path issues do not cause container start errors.
Trap 2: The container image does not have the correct entry point
Missing entry point would cause a different error.
Trap 3: The GPU drivers are outdated
Outdated GPU drivers would cause a different error, but SageMaker manages drivers.
- A
The instance type does not have enough memory for the container
Insufficient memory is a common cause of container startup failures.
- B
The training data is stored in the wrong S3 bucket
Why it fails: S3 permissions or path issues do not cause container start errors.
- C
The container image does not have the correct entry point
Why it fails: Missing entry point would cause a different error.
- D
The GPU drivers are outdated
Why it fails: Outdated GPU drivers would cause a different error, but SageMaker manages drivers.