Courseiva

MLA-C01 Deployment and Orchestration of ML Workflows Practice Question

A company has 200 small PyTorch models that are each used infrequently but need to be available for real-time inference. To minimize costs, they want to host all models on a single endpoint. Which SageMaker feature should they use?

⚠ Common exam trap

MLA-C01 often tests the confusion between multi-model endpoints (many models, one container, dynamic load) and multi-container endpoints (few containers, different frameworks, static), so candidates who see 'multiple models' and pick multi-container get it wrong.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Multi-model endpoint (MME)

SageMaker multi-model endpoints (MME) let a single endpoint host hundreds or thousands of models behind one container, loading each model into memory or disk on demand and unloading idle ones. This is purpose-built for the scenario of many small, infrequently used models that must still serve real-time inference, because you pay for one endpoint's worth of instances rather than one endpoint per model. The SageMaker SDK/API passes a TargetModel parameter on each InvokeEndpoint call so the endpoint knows which model to load and serve.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Multi-model endpoint (MME)

    Why this is correct

    Multi-model endpoint hosts many models on one endpoint, loading each from S3 on invocation and caching it in memory. Because the PyTorch models are infrequent, this satisfies the single-endpoint and cost-minimisation constraints without provisioning per-model hosting.

  • ✗

    Multi-container endpoint

    Why it's wrong here

    Multi-container endpoints run several containers behind one endpoint but each container serves its own model on a shared instance, so 200 models cannot share one container's memory. It suits hosting a handful of distinct frameworks together, not many small models loaded dynamically on demand.

  • ✗

    Batch Transform job

    Why it's wrong here

    Batch Transform processes an entire dataset offline in a job, returning no persistent real-time endpoint for on-demand inference. It is tempting because it is cost-effective for infrequent work, and would be correct if predictions could be generated in bulk ahead of time rather than served interactively.

  • ✗

    Asynchronous inference endpoint

    Why it's wrong here

    Asynchronous inference queues requests and returns results via Amazon S3, so callers cannot receive an immediate synchronous response. It is tempting for large payloads or long processing times, and would be correct when near-real-time queued inference with extended runtimes is acceptable.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on MLA-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company has 50 small PyTorch models that are used infrequently for inference. They want to minimize costs while maintaining the ability to serve all models from a single endpoint. Which SageMaker feature should they use?

easy
  • A.Multi-container endpoint
  • B.Batch transform job
  • C.Real-time endpoint with 50 production variants
  • ✓ D.Multi-model endpoint

Why D: SageMaker multi-model endpoints allow hosting multiple models on a single endpoint, loading them on demand and sharing the same serving container. This is cost-effective for many infrequently used models because you only pay for the endpoint instance and not for separate endpoints per model. It supports hundreds of models and dynamically loads them from S3.

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.