Courseiva

MLA-C01 Deployment and Orchestration of ML Workflows Practice Question

A team needs to deploy a PyTorch model that uses custom CUDA kernels. They want to use NVIDIA Triton Inference Server on SageMaker for high-performance serving. Which SageMaker configuration is required to use Triton?

⚠ Common exam trap

AWS often tests the misconception that custom containers are always required for custom code, but the trap here is that SageMaker's pre-built Triton container fully supports custom CUDA kernels, making option A a redundant and incorrect choice.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the SageMaker pre-built Triton Inference Server container available in Amazon ECR

SageMaker provides a pre-built Triton Inference Server container in Amazon ECR that is optimized for high-performance serving of models, including those with custom CUDA kernels. This container eliminates the need to build a custom image from scratch, ensuring compatibility with SageMaker's deployment infrastructure and reducing operational overhead.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Create a custom container from scratch with Triton and deploy on SageMaker

    Why it's wrong here

    Creating a custom container from scratch is not the specific SageMaker configuration required for Triton, as SageMaker provides native integration that handles the complex setup. This native support offers pre-built images and manages the intricate deployment of Triton, CUDA, and PyTorch, which is essential for models using custom CUDA kernels. Building from scratch would necessitate manual configuration of all these components, bypassing the intended SageMaker simplification. This option is tempting because custom containers offer maximum flexibility for deploying inference servers or models not natively supported by SageMaker, or when highly specific, non-standard customisation is needed.

  • ✓

    Use the SageMaker pre-built Triton Inference Server container available in Amazon ECR

    Why this is correct

    SageMaker provides a pre-built container with Triton, ready for deployment.

  • ✗

    Use a Multi-Model Endpoint with Triton

    Why it's wrong here

    Multi-Model Endpoints let one endpoint host several models behind shared serving infrastructure; they do not provide the CUDA kernel execution environment Triton requires. MME is the right choice when consolidating many low-traffic models onto one endpoint, not for enabling custom CUDA kernels.

  • ✗

    Attach an Amazon Elastic Inference accelerator to the endpoint

    Why it's wrong here

    Elastic Inference accelerators attach to supported frameworks via a separate inference library and do not expose the CUDA runtime Triton needs to load custom kernels. EI suits cost-reduced inference for standard TensorFlow, PyTorch or MXNet models, not custom CUDA kernel execution.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.