Courseiva

AIF-C01 Fundamentals of AI and ML Practice Question

A company wants to deploy a real-time inference endpoint for a custom model on SageMaker. The model has high latency (100ms) and they need to handle variable traffic with spikes. Which deployment strategy is most cost-effective?

⚠ Common exam trap

Test-takers frequently confuse multi-model endpoints with cost-effective scaling for a single model, not realizing that multi-model endpoints are designed for hosting many models, not for handling variable traffic for one model with high latency.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy on a single SageMaker endpoint with automatic scaling

A single SageMaker endpoint with automatic scaling allows the endpoint to dynamically adjust the number of instances based on traffic patterns, handling variable traffic and spikes cost-effectively. For a model with 100ms latency, automatic scaling can add instances during spikes and remove them during low traffic, ensuring you only pay for the compute resources you use while maintaining low inference latency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy on a SageMaker multi-model endpoint

    Why it's wrong here

    Multi-model endpoints load models on demand from S3 and unload idle ones, adding cold-start latency and suiting many low-traffic models sharing an instance, not one latency-sensitive model absorbing spikes. It is tempting because it cuts hosting cost, and it would be correct when hosting hundreds of infrequently invoked models.

  • ✗

    Use batch transform

    Why it's wrong here

    Batch transform processes a complete dataset as an asynchronous job and has no persistent endpoint, so it cannot serve real-time requests or absorb traffic spikes. It is tempting because it is cheap for bulk scoring, and it would be correct for offline inference over stored data rather than interactive low-latency predictions.

  • ✓

    Deploy on a single SageMaker endpoint with automatic scaling

    Why this is correct

    Automatic scaling adjusts instance count to match traffic spikes while a single endpoint avoids paying for idle capacity during troughs. This satisfies the variable-traffic constraint cost-effectively, unlike always-on multi-instance deployments that over-provision for peak load.

  • ✗

    Use a single large instance type

    Why it's wrong here

    A single large instance provisions fixed capacity sized for peak load, so it sits idle between spikes and cannot scale out when traffic surges. It is tempting because one instance is simple to configure, and it would be correct for steady, predictable workloads where autoscaling adds no benefit.

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.