Courseiva
easyMultiple Choice

MLA-C01 Practice Question: A company uses Amazon SageMaker to deploy a…

A company uses Amazon SageMaker to deploy a real-time inference endpoint. They notice increased latency in predictions during peak hours. Which should they investigate first to address the issue?

⚠ Common exam trap

Many exam-takers confuse training infrastructure (instance type, artifact size) with inference infrastructure, or assume that data labeling quality affects inference speed, when the immediate cause of peak-hour latency is almost always insufficient endpoint capacity due to misconfigured auto-scaling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Review the endpoint auto-scaling policy

Increased latency during peak hours is a classic symptom of insufficient compute capacity to handle the request volume. The first step is to review the endpoint's auto-scaling policy to ensure it is configured to scale out instances proactively or reactively based on a relevant metric like 'SageMakerVariantInvocationsPerInstance'. If the policy has a high cooldown period or a low target metric value, it may not add instances quickly enough, causing requests to queue and latency to spike.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Review the endpoint auto-scaling policy

    Why this is correct

    Peak-hour latency typically indicates insufficient capacity, so reviewing the endpoint auto-scaling policy reveals whether instance counts and target-tracking thresholds keep pace with demand. This satisfies the requirement to investigate the most likely cause first, before examining model artefacts or payload sizes.

  • ✗

    Check the data labeling job status

    Why it's wrong here

    Data labelling job status concerns preparing training datasets and has no bearing on the runtime latency of a deployed inference endpoint. It is tempting because labelling is part of the SageMaker workflow, but it would be the correct area to check when investigating training data quality or a stalled ground-truth job, not endpoint performance.

  • ✗

    Modify the training instance type

    Why it's wrong here

    Training instance type affects model fitting cost and duration, not the latency of an already-deployed real-time endpoint during peak traffic. It is tempting because instance sizing influences performance generally, but the correct choice here is endpoint-side scaling or instance selection, which directly governs inference response times under load.

  • ✗

    Increase the model artifact size

    Why it's wrong here

    A larger model artifact increases download and load time and typically raises per-inference latency, so enlarging it worsens the symptom rather than addressing peak-hour latency. Artifact size matters when optimising cold-start or container image pull times, not for handling concurrent real-time inference traffic.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.