Courseiva

MLA-C01 Deployment and Orchestration of ML Workflows Practice Question

A machine learning team needs to deploy a PyTorch model that has been compiled with SageMaker Neo to improve inference performance on edge devices. Which TWO statements about SageMaker Neo are correct? (Select TWO.)

⚠ Common exam trap

MLA-C01 often tests the misconception that SageMaker Neo is tied to SageMaker training or built-in algorithms, when in fact it is a standalone compilation service that works with models from any source.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Neo reduces model inference latency through optimization techniques

Option A is correct because SageMaker Neo applies compiler-level optimizations such as operator fusion, constant folding, and quantization-aware graph rewrites that reduce inference latency and model size on the target hardware. Option C is correct because Neo is a compilation service that takes a trained model plus a target hardware specification (for example, an Intel x86 CPU with a specific instruction set, an ARM Cortex-A SoC, or an NVIDIA Jetson GPU) and produces a hardware-optimized executable, so the compiled artifact is tied to that target. Option B is incorrect because Neo can compile models trained anywhere, including on-premises or in other clouds, as long as the model is provided in a supported framework format such as PyTorch, TensorFlow, MXNet, or ONNX. Option D is incorrect because Neo supports custom models and frameworks, not just SageMaker built-in algorithms. Option E is incorrect because endpoint auto scaling is handled by Application Auto Scaling and SageMaker endpoint scaling policies, not by Neo, which only handles model compilation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Neo reduces model inference latency through optimization techniques

    Why this is correct

    SageMaker Neo applies operator fusion, constant folding and quantisation-aware compilation to shrink the model graph, directly satisfying the stem's requirement to improve inference performance on constrained edge hardware. This optimisation lowers per-inference latency rather than merely packaging the PyTorch artefacts for deployment.

  • ✗

    Neo requires the model to be trained on SageMaker

    Why it's wrong here

    Neo compiles models from frameworks including PyTorch, TensorFlow and XGBoost regardless of training location; it consumes model artefacts from Amazon S3. Tempting because SageMaker-native workflows are common, and Neo would be correct where the model was trained in SageMaker, but that is not a prerequisite.

  • ✓

    Neo compiles models for a specific hardware target, such as Intel or ARM

    Why this is correct

    Neo compiles the trained model into an optimised binary for a declared target architecture, such as Intel x86 or ARM, rather than running the framework graph directly. This satisfies the edge deployment requirement because the compiled artefact matches the device's processor.

  • ✗

    Neo can only compile models trained with SageMaker built-in algorithms

    Why it's wrong here

    Neo accepts models from any framework it supports, including PyTorch, TensorFlow, MXNet and XGBoost, trained anywhere. Tempting because built-in algorithms integrate smoothly with Neo, and it would be correct only if the question asked about models produced by SageMaker built-in algorithms, not the general constraint.

  • ✗

    Neo automatically scales SageMaker endpoints based on demand

    Why it's wrong here

    Neo is a compilation tool that converts trained models into optimised executables for target hardware; endpoint autoscaling is handled by Application Auto Scaling on SageMaker endpoints. Tempting because both concern deployment performance, and autoscaling would be correct for a scenario needing dynamic capacity adjustment under variable inference load.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.