Courseiva
Fundamentals of AI and ML →mediumMultiple Choice

AIF-C01 Fundamentals of AI and ML Practice Question

An ML team is deploying a real-time inference endpoint for a computer vision model using Amazon SageMaker. The model requires GPU acceleration for low latency. Which instance type should the team choose to minimize cost while meeting the GPU requirement?

⚠ Common exam trap

Candidates may be tempted to choose ml.p3.2xlarge due to its reputation for ML training, but for inference, ml.g5 instances often provide better price-performance. The ml.p3.2xlarge, while GPU-equipped, is more expensive per hour and may be over-provisioned for typical inference workloads, leading to unnecessary costs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

ml.g5.xlarge

(ml.g5.xlarge) is correct because it provides a GPU (NVIDIA A10G Tensor Core GPU) necessary for low-latency GPU acceleration in computer vision inference, while being the most cost-effective GPU instance among the options. The ml.g5.xlarge offers sufficient GPU compute for real-time inference at a lower hourly cost compared to ml.p3.2xlarge, making it the optimal choice for minimizing cost while meeting the GPU requirement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    ml.g5.xlarge

    Why this is correct

    ml.g5.xlarge pairs an NVIDIA A10G GPU with four vCPUs, providing the required GPU acceleration at the lowest cost within the G5 family. Smaller GPU instances lack sufficient acceleration, while larger G5 sizes exceed the stated requirement.

  • ✗

    ml.c5.xlarge

    Why it's wrong here

    ml.c5.xlarge is a compute-optimised CPU instance with no GPU, so it cannot satisfy the GPU acceleration requirement at all. It is tempting because it is the cheapest option in the list and suits CPU-bound inference, but the stem explicitly mandates GPU hardware for low-latency computer vision.

  • ✗

    ml.p3.2xlarge

    Why it's wrong here

    ml.p3.2xlarge provides a single NVIDIA V100 GPU, but ml.g4dn.xlarge offers GPU acceleration at lower cost for inference workloads. It is tempting because p3 is a well-known GPU family, yet it targets training and heavier compute rather than cost-minimised real-time inference.

  • ✗

    ml.p4d.24xlarge

    Why it's wrong here

    ml.p4d.24xlarge carries eight A100 GPUs, far exceeding the single-GPU need and inflating cost dramatically. It is tempting because it is the most powerful GPU instance, suited to large-scale distributed training, but the stem asks to minimise cost while merely meeting GPU acceleration.

About these practice questions

One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.