Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: A company deploys a model on Amazon SageMaker for…

A company deploys a model on Amazon SageMaker for real-time inference. The inference latency is too high. The model is a large deep learning model. The company wants to reduce latency without significantly impacting accuracy. Which approach should the company consider?

⚠ Common exam trap

AWS often tests the misconception that reducing instance size or increasing batch size directly reduces latency, when in fact these actions typically increase latency or degrade throughput for real-time inference.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use SageMaker Inference Recommender to test different instance types and optimizations.

SageMaker Inference Recommender is designed specifically to automate load testing and benchmarking across various instance types and model optimizations (e.g., Elastic Inference, GPU acceleration, serialization formats). It provides latency and throughput metrics to identify the optimal configuration for reducing inference latency while maintaining accuracy, making it the correct choice for a large deep learning model with high latency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the batch size for inference.

    Why it's wrong here

    Increasing batch size groups requests, improving throughput but delaying each individual response until the batch fills, which raises per-request latency for real-time inference. Batching suits offline or high-throughput asynchronous scoring. Real-time endpoints need single-request latency reduced through model or hardware optimisation.

  • ✗

    Use a smaller instance type to reduce inference time.

    Why it's wrong here

    A smaller instance type provides less CPU, GPU memory and compute, so a large deep learning model's forward pass runs slower or spills, increasing latency. Smaller instances suit lightweight models or cost reduction where latency is not the constraint. The scenario needs accelerated compute or model optimisation instead.

  • ✓

    Use SageMaker Inference Recommender to test different instance types and optimizations.

    Why this is correct

    Inference Recommender benchmarks the large deep learning model across instance types, batch sizes and optimisations such as compilation, identifying configurations that cut latency. This empirical comparison avoids guesswork and preserves accuracy, since it tests rather than alters the model's weights.

  • ✗

    Enable SageMaker Model Monitor to detect performance issues.

    Why it's wrong here

    Model Monitor observes data drift, bias and quality metrics after deployment; it detects degradation rather than altering inference execution, so latency is unchanged. It is the right choice for ongoing production quality assurance. Reducing latency requires optimising the model or the serving instance instead.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.