Courseiva

MLA-C01 Deployment and Orchestration of ML Workflows Practice Question

A machine learning engineer needs to deploy a TensorFlow model that requires a custom inference environment with specific system libraries. The model will be used in a real-time application with variable traffic. They want to minimize cold start latency. Which SageMaker hosting option should they choose?

⚠ Common exam trap

Watch out — candidates often confuse 'minimizing cold start latency' with 'scaling to zero' and incorrectly choose Serverless Inference, failing to recognize that Serverless inherently introduces cold starts on first request after idle periods.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

SageMaker real-time endpoint with a custom container

SageMaker real-time endpoints with a custom container are the correct choice because they provide persistent, always-on infrastructure that eliminates cold start latency. By packaging the TensorFlow model with required system libraries in a custom Docker image, the engineer ensures the inference environment is ready immediately, and the endpoint can scale to handle variable traffic with minimal delay.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    SageMaker real-time endpoint with a custom container

    Why this is correct

    A real-time endpoint with a custom container packages the required system libraries while keeping an instance continuously provisioned, so requests are served without container startup delay. This directly minimises cold start latency for variable traffic, unlike serverless inference which scales to zero.

  • ✗

    SageMaker Serverless Inference with a custom container

    Why it's wrong here

    Serverless Inference scales to zero, so every request after idle time triggers a full container initialisation, worsening cold start latency for a custom TensorFlow image. It suits sporadic, latency-tolerant workloads. Real-time variable traffic with a custom container needs a provisioned endpoint with autoscaling, which keeps instances warm.

  • ✗

    SageMaker Multi-Model Endpoint with a custom container

    Why it's wrong here

    Multi-Model Endpoints load models on demand from shared storage, adding cold start latency and offering no per-model custom inference environment. They are tempting because they are the right choice when hosting many models behind one endpoint to cut hosting cost, not for latency-sensitive custom environments.

  • ✗

    SageMaker Asynchronous Inference with a custom container

    Why it's wrong here

    Asynchronous Inference queues requests and returns responses via Amazon S3, designed for large payloads and long processing, not sub-second real-time responses. It cannot minimise cold start latency for an interactive application. A provisioned real-time endpoint with autoscaling and a custom container is the correct hosting option here.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.