Courseiva
ML Model Development →easyMultiple Choice

MLA-C01 ML Model Development Practice Question

A machine learning engineer has trained a model in SageMaker and wants to deploy it to a real-time endpoint for low-latency inference. The model artifacts are stored in Amazon S3, and the engineer needs to create the endpoint with the least operational effort. Which sequence of actions should the engineer take?

⚠ Common exam trap

The trap here is conflating model registration or batch transform with real-time hosting, when a real-time endpoint requires the model, configuration, and endpoint resources.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Create a model, create an endpoint configuration, and create an endpoint using the SageMaker API or the AWS SDK for Python (Boto3).

A SageMaker real-time endpoint is composed of a model, an endpoint configuration, and an endpoint. Creating these three resources through the SageMaker API or Boto3 is the direct, low-effort way to deploy artifacts from Amazon S3 for low-latency inference. Batch transform, model registry pipelines, and self-managed EC2 hosting either do not provide real-time serving or add unnecessary operational burden.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Create a model, create an endpoint configuration, and create an endpoint using the SageMaker API or the AWS SDK for Python (Boto3).

    Why this is correct

    Deploying to a real-time endpoint in SageMaker requires three resources: a model that points to the artifacts and image, an endpoint configuration that defines the instance type and count, and an endpoint that provisions the compute. Using the SageMaker API or Boto3 performs these steps directly and is the standard low-effort path for a single model deployment.

  • ✗

    Package the model into a Docker image, push it to Amazon ECR, and run it on an Amazon EC2 instance behind an Application Load Balancer.

    Why it's wrong here

    This approach requires managing the container image, EC2 instances, scaling, and load balancing manually. SageMaker real-time endpoints handle provisioning and scaling of the inference compute. Building a custom stack on EC2 is significantly more operational effort and does not use SageMaker's managed hosting capabilities.

  • ✗

    Create a batch transform job that reads from Amazon S3 and writes predictions to Amazon S3, then expose the output as a REST API.

    Why it's wrong here

    Batch transform is designed for offline, high-throughput inference over a dataset, not for low-latency real-time requests. It does not provision a persistent endpoint and would require additional components to serve REST traffic. This approach adds operational effort and does not meet the low-latency requirement.

  • ✗

    Register the model in the SageMaker model registry and deploy it through a SageMaker pipeline with a manual approval step.

    Why it's wrong here

    The model registry and pipelines support governance and automation across multiple models, which is more operational overhead than needed for a single endpoint. A manual approval step also delays deployment. While these features are useful at scale, they do not represent the least-effort path for this simple real-time deployment.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.