Courseiva

CCNA Mla Deployment Orchestration Questions

19 of 94 questions · Page 2/2 · Mla Deployment Orchestration topic · Answers revealed

76
MCQeasy

A company uses SageMaker Pipelines to automate their ML workflow. They notice that the pipeline reruns all steps even when the input data has not changed. Which feature should they enable to avoid unnecessary recomputation?

A.Enable pipeline caching
B.Use a Lambda step to check input changes
C.Use a Conditional step to skip steps
D.Set the pipeline execution mode to 'Parallel'
AnswerA

Pipeline caching stores step outputs keyed on input signatures, so unchanged data and code let SageMaker skip re-execution and reuse prior artefacts. This directly removes the unnecessary recomputation the company observes, cutting cost and runtime without altering pipeline structure.

Why this answer

Pipeline caching in SageMaker Pipelines automatically reuses the output of a step if its inputs (including parameters, data, and code) have not changed since the last successful execution. This avoids recomputation by comparing a hash of the step's dependencies against previous runs, making it the correct feature to prevent unnecessary reruns when input data remains identical.

Exam trap

The trap here is that candidates confuse caching with conditional branching or parallel execution, assuming that skipping steps via conditions or running steps in parallel will avoid recomputation, when in fact only caching directly reuses prior outputs based on input immutability.

How to eliminate wrong answers

Option B is wrong because a Lambda step is used for custom processing or integration (e.g., invoking external APIs), not for detecting input changes or caching step outputs; it would add complexity without solving the core caching requirement. Option C is wrong because a Conditional step evaluates a condition to branch the pipeline (e.g., skip a step based on a metric), but it does not automatically detect unchanged inputs or cache results; it requires manual logic and still incurs overhead for the condition check. Option D is wrong because setting the pipeline execution mode to 'Parallel' controls whether steps run sequentially or concurrently, but it does not prevent recomputation of steps whose inputs have not changed; it only affects execution order, not caching.

77
Multi-Selectmedium

A company uses SageMaker to deploy a model and wants to perform A/B testing by splitting traffic between two model variants. Which TWO actions should they take? (Select TWO.)

Select 2 answers
A.Configure two production variants on the endpoint, each with an initial weight
B.Use SageMaker Model Registry to approve both variants
C.Use the UpdateEndpointWeightsAndCapacity API to adjust traffic after analysis
D.Deploy each variant to a separate endpoint and use Route53 weighted routing
E.Enable shadow testing on the endpoint
AnswersA, C

Configuring two production variants on a single SageMaker endpoint, each assigned an initial weight, directly enables traffic splitting for A/B testing. The endpoint distributes inference requests across variants according to those weights, satisfying the stem's requirement to compare two model variants without deploying separate endpoints.

Why this answer

Option A is correct because SageMaker A/B testing is implemented by configuring a single endpoint with two production variants, each assigned an InitialVariantWeight so the endpoint can split inference traffic between them. Option C is correct because after analyzing the variants' performance, you can shift traffic by calling UpdateEndpointWeightsAndCapacity to change each variant's VariantWeight and DesiredInstanceCount without redeploying the endpoint. Option B is not required: Model Registry approval is a governance step and does not itself create or split endpoint traffic.

Option D is wrong because deploying separate endpoints with Route53 weighted routing is a DNS-level workaround, not SageMaker's native A/B testing mechanism. Option E is wrong because shadow testing mirrors live traffic to a new variant for evaluation without serving its responses, so it does not split production traffic between two variants.

Exam trap

The trap here is that candidates confuse shadow testing (Option E) with A/B testing, but shadow testing does not split live traffic—it only mirrors requests for offline analysis, while A/B testing requires actual traffic distribution between variants.

78
MCQmedium

A team uses MLflow on SageMaker for experiment tracking. They want to automate the retraining of a model when new training data arrives in an S3 bucket. Which combination of services should they use?

A.SageMaker Pipelines scheduled trigger every hour
B.EventBridge -> Lambda -> SageMaker Training Job
C.S3 Event Notifications -> SQS -> SageMaker Training Job
D.AWS Step Functions with S3 poller
AnswerB

EventBridge detects the new S3 object and triggers a Lambda function that starts a SageMaker Training Job, giving event-driven retraining. This satisfies the automation requirement, since S3 alone cannot invoke training and Lambda supplies the orchestration glue.

Why this answer

EventBridge can capture S3 events (via CloudTrail or S3 event notifications routed to EventBridge) and trigger a Lambda function that starts a SageMaker Training Job. This event-driven pattern is the standard AWS-native way to automate retraining when new data lands in S3. It avoids polling and integrates cleanly with MLflow tracking on SageMaker.

Exam trap

MLA-C01 often tests event-driven vs scheduled automation, tricking candidates into choosing polling or scheduled solutions when the requirement is to react to new data arrival.

How to eliminate wrong answers

Option A is wrong because a scheduled hourly trigger is time-based, not event-driven, so it does not react to new data arrival and may waste resources or delay retraining. Option C is wrong because SQS is a queue, not a compute trigger — you still need a consumer (Lambda or similar) to start the training job, and S3 event notifications to SQS alone do not invoke SageMaker. Option D is wrong because Step Functions does not natively poll S3; you would still need EventBridge or S3 notifications to drive it, making this combination incomplete.

79
MCQmedium

A machine learning engineer needs to deploy a TensorFlow model that requires a custom inference environment with specific system libraries. The model will be used in a real-time application with variable traffic. They want to minimize cold start latency. Which SageMaker hosting option should they choose?

A.SageMaker real-time endpoint with a custom container
B.SageMaker Serverless Inference with a custom container
C.SageMaker Multi-Model Endpoint with a custom container
D.SageMaker Asynchronous Inference with a custom container
AnswerA

A real-time endpoint with a custom container packages the required system libraries while keeping an instance continuously provisioned, so requests are served without container startup delay. This directly minimises cold start latency for variable traffic, unlike serverless inference which scales to zero.

Why this answer

SageMaker real-time endpoints with a custom container are the correct choice because they provide persistent, always-on infrastructure that eliminates cold start latency. By packaging the TensorFlow model with required system libraries in a custom Docker image, the engineer ensures the inference environment is ready immediately, and the endpoint can scale to handle variable traffic with minimal delay.

Exam trap

The trap here is that candidates often confuse 'minimizing cold start latency' with 'scaling to zero' and incorrectly choose Serverless Inference, failing to recognize that Serverless inherently introduces cold starts on first request after idle periods.

How to eliminate wrong answers

Option B is wrong because SageMaker Serverless Inference automatically scales to zero when idle, incurring cold start latency (typically 5–10 seconds) when traffic resumes, which contradicts the requirement to minimize cold start latency. Option C is wrong because SageMaker Multi-Model Endpoints are designed to host multiple models on a single container, but they still require a pre-configured inference environment; they do not inherently reduce cold start latency for a single custom model. Option D is wrong because SageMaker Asynchronous Inference is intended for non-real-time workloads with larger payloads and queuing, and it also experiences cold starts when scaling from zero, making it unsuitable for real-time applications with variable traffic.

80
MCQeasy

A company wants to deploy a model using a serverless inference endpoint that can automatically scale to zero when not in use and has a configurable maximum concurrency. Which SageMaker inference option meets these requirements?

A.Serverless inference
B.Real-time endpoint with auto-scaling
C.Batch transform
D.Asynchronous inference
AnswerA

Serverless inference provisions compute on demand and scales to zero during idle periods, eliminating charges when unused. Its configurable maximum concurrency caps simultaneous invocations, matching the stem's scaling and concurrency constraints. Provisioned endpoints cannot scale to zero, and asynchronous inference targets queued payloads rather than interactive low-latency requests.

Why this answer

SageMaker Serverless Inference is the correct choice because it automatically scales to zero when the endpoint is idle, eliminating costs during periods of no traffic, and it allows you to configure a maximum concurrency limit per endpoint to control throughput. This fully managed, pay-per-invoke option is designed for workloads with intermittent or unpredictable traffic patterns, meeting both requirements precisely.

Exam trap

The trap here is that candidates confuse 'auto-scaling' with 'scaling to zero' and incorrectly choose the real-time endpoint with auto-scaling, not realizing that auto-scaling maintains a minimum instance count and cannot reduce to zero.

How to eliminate wrong answers

Option B is wrong because a real-time endpoint with auto-scaling can scale down to a minimum number of instances (e.g., 1) but cannot scale to zero; it always keeps at least one instance running, incurring base costs. Option C is wrong because batch transform is not a real-time inference endpoint; it processes entire datasets offline and does not support automatic scaling to zero or configurable concurrency for live requests. Option D is wrong because asynchronous inference endpoints can scale to zero when idle, but they do not support a configurable maximum concurrency; concurrency is managed internally based on the payload size and queue depth, not directly set by the user.

81
MCQeasy

A machine learning engineer wants to automatically trigger a retraining pipeline whenever new training data arrives in an S3 bucket. The pipeline uses SageMaker Pipelines. Which AWS service should be used to detect the S3 event and start the pipeline?

A.SageMaker Pipelines native S3 trigger
B.AWS Step Functions
C.Amazon CloudWatch Logs
D.Amazon EventBridge
AnswerD

Amazon EventBridge natively receives S3 bucket notifications, matches them against a rule, and starts the SageMaker pipeline as a target. This satisfies the detect-and-trigger requirement without polling, unlike Lambda alone which would still need EventBridge or S3 notification wiring.

Why this answer

Amazon EventBridge can be configured to listen for S3 events (e.g., PutObject) and then invoke a Lambda function that starts the SageMaker Pipeline execution. Step Functions could orchestrate the pipeline but is not needed to trigger on S3 events. SageMaker Pipelines does not natively listen to S3 events.

CloudWatch Events is the older name for EventBridge.

82
MCQmedium

A team built a SageMaker Pipeline that includes a training step and a model evaluation step. They want to automatically register a model in SageMaker Model Registry only if the evaluation metric (accuracy) exceeds 0.9. Which pipeline step should be used to implement this conditional logic?

A.RegisterModel step
B.Condition step
C.Processing step
D.Transform step
AnswerB

A condition step evaluates a property such as the evaluation metric against a threshold and gates downstream steps accordingly. Placing the register step inside its true branch means registration occurs only when accuracy exceeds 0.9, exactly as the stem requires.

Why this answer

The Condition step in SageMaker Pipelines allows you to add conditional branching logic, such as evaluating a metric and proceeding only if a condition is met. In this scenario, you would use a ConditionStep to check if the accuracy metric from the evaluation step exceeds 0.9, and then conditionally execute a RegisterModel step to register the model in SageMaker Model Registry. A common misconception is that the RegisterModel step itself can conditionally register a model based on metrics, but in SageMaker Pipelines, conditional logic must be implemented explicitly with a ConditionStep.

Exam trap

A common trap is thinking the RegisterModel step itself can conditionally register a model based on metrics, but in SageMaker Pipelines, conditional logic must be implemented explicitly with a ConditionStep.

How to eliminate wrong answers

Option A is wrong because the RegisterModel step is used to register a model in the Model Registry, but it does not have built-in conditional logic; it would register the model unconditionally unless placed inside a ConditionStep. Option C is wrong because a Processing step is used for data processing, feature engineering, or evaluation tasks, not for implementing conditional branching logic. Option D is wrong because a Transform step is used for batch inference or model serving, not for conditional evaluation or registration decisions.

83
Multi-Selecthard

An ML engineer is designing a SageMaker Pipeline for model training and registration. They need to ensure that the pipeline can be re-run with different datasets without manual intervention, and that the steps are only re-executed if inputs have changed. Which THREE features should they configure? (Select THREE.)

Select 3 answers
A.Add a Condition step to manually check for data changes
B.Enable step caching to reuse outputs when inputs are unchanged
C.Configure lineage tracking to record the origin of models
D.Use Parameterized execution to pass different values at runtime
E.Define pipeline parameters for dataset location and hyperparameters
AnswersB, D, E

Step caching stores each step's outputs keyed by its input signature; when a re-run supplies unchanged inputs, SageMaker skips execution and reuses the cached output. This directly satisfies the requirement that steps are only re-executed if inputs have changed.

Why this answer

Option B is correct because SageMaker Pipelines step caching stores each step's output keyed by a hash of the step's inputs (code, arguments, and input artifacts), so when a pipeline re-runs and a step's inputs are unchanged, the cached output is reused instead of re-executing the step. Option D is correct because parameterized execution lets the pipeline be invoked with different parameter values at runtime (for example via StartPipelineExecution with PipelineParameters), enabling re-runs on new datasets without editing or manually rebuilding the pipeline definition. Option E is correct because defining pipeline parameters for dataset location and hyperparameters is the mechanism that makes the pipeline dynamic and reusable, so the same pipeline definition can be pointed at different data and training configurations.

Option A is not appropriate because a Condition step performs branching logic on property values during execution; it does not detect data changes or prevent re-execution, and it would require manual logic rather than automatic cache-based skipping. Option C is not appropriate because lineage tracking records the provenance of artifacts (data, models, jobs) for audit and traceability; it does not control whether steps re-execute or enable parameterized re-runs.

Exam trap

MLA-C01 often tests the distinction between features that enable reusability and those that provide auditing or conditional logic. Candidates might incorrectly select lineage tracking or condition steps, but the key is to focus on automation and parameterization for re-runs with different datasets.

84
Multi-Selectmedium

A team is migrating their ML infrastructure to AWS and wants to use infrastructure as code to manage SageMaker Studio domains, user profiles, and associated resources. Which services can they use for this purpose? (Select THREE.)

Select 3 answers
A.AWS CDK (Cloud Development Kit)
B.SageMaker Python SDK
C.Boto3
D.Terraform by HashiCorp
E.AWS CloudFormation
AnswersA, D, E

AWS CDK provisions SageMaker Studio domains and user profiles as infrastructure as code, synthesising CloudFormation templates from familiar programming languages. It satisfies the migration's IaC requirement by managing these resources declaratively and repeatably, unlike manual console configuration. CDK constructs also handle associated resources such as IAM roles and VPC settings within the same stack.

Why this answer

AWS CDK (A) is correct because it is an infrastructure-as-code framework that synthesizes CloudFormation templates, and it provides L2 constructs such as CfnDomain, CfnUserProfile, and CfnApp for managing SageMaker Studio domains, user profiles, and associated resources declaratively. Terraform by HashiCorp (D) is correct because it is a third-party IaC tool with dedicated AWS provider resources like aws_sagemaker_domain, aws_sagemaker_user_profile, and aws_sagemaker_app that let teams manage the same SageMaker Studio resources in HCL state files. AWS CloudFormation (E) is correct because it is AWS's native IaC service and directly supports the AWS::SageMaker::Domain, AWS::SageMaker::UserProfile, and AWS::SageMaker::App resource types for provisioning and updating these resources.

The SageMaker Python SDK (B) is not an infrastructure-as-code service; it is a high-level library for training, deploying, and interacting with models and endpoints, not for declaratively managing Studio domain infrastructure. Boto3 (C) is the AWS SDK for Python and can call SageMaker APIs imperatively, but it is not an infrastructure-as-code tool with declarative templates or state management, so it does not fit the IaC requirement.

Exam trap

The trap here is that candidates often confuse the SageMaker Python SDK (used for ML workflows) or Boto3 (used for general AWS API calls) with infrastructure as code tools, but neither provides declarative, state-managed provisioning of SageMaker Studio resources like CloudFormation, CDK, or Terraform do.

85
MCQhard

A team needs to deploy a PyTorch model that uses custom CUDA kernels. They want to use NVIDIA Triton Inference Server on SageMaker for high-performance serving. Which SageMaker configuration is required to use Triton?

A.Create a custom container from scratch with Triton and deploy on SageMaker
B.Use the SageMaker pre-built Triton Inference Server container available in Amazon ECR
C.Use a Multi-Model Endpoint with Triton
D.Attach an Amazon Elastic Inference accelerator to the endpoint
AnswerB

SageMaker provides a pre-built container with Triton, ready for deployment.

Why this answer

SageMaker provides a pre-built Triton Inference Server container in Amazon ECR that is optimized for high-performance serving of models, including those with custom CUDA kernels. This container eliminates the need to build a custom image from scratch, ensuring compatibility with SageMaker's deployment infrastructure and reducing operational overhead.

Exam trap

AWS often tests the misconception that custom containers are always required for custom code, but the trap here is that SageMaker's pre-built Triton container fully supports custom CUDA kernels, making option A a redundant and incorrect choice.

How to eliminate wrong answers

Option A is wrong because creating a custom container from scratch is unnecessary and error-prone; SageMaker already offers a pre-built Triton container that handles the integration with SageMaker's hosting environment, including health checks and model loading. Option C is wrong because Multi-Model Endpoints are designed to host multiple models on a single container, but they do not inherently support Triton's specific features like dynamic batching and model pipelines; Triton requires its own server process, which is not compatible with the Multi-Model Endpoint architecture. Option D is wrong because Amazon Elastic Inference accelerators are deprecated and do not support custom CUDA kernels or Triton; they are limited to specific frameworks like TensorFlow and PyTorch without custom ops, and they cannot accelerate custom CUDA code.

86
MCQhard

A machine learning engineer is using SageMaker Pipelines to orchestrate a training workflow. The pipeline includes a processing step that outputs a dataset, which is then used by a training step. The engineer notices that the processing step runs every time the pipeline executes, even when the input data has not changed. The engineer wants to avoid re-running the processing step if the input data and code are unchanged, while ensuring that downstream steps still execute if the processing step is skipped. Which approach should the engineer take?

A.Use a condition step to check if the input data has changed, and if not, skip the processing step and directly pass the previous output to the training step.
B.Configure the processing step to use a spot instance, which will reduce cost but not affect whether the step runs.
C.Set the pipeline's execution mode to 'Reprocess' and manually skip the processing step by editing the pipeline definition before each run.
D.Enable caching on the processing step by setting the cache policy to CacheConfig with a time-to-live (TTL) and a cache key that includes the input data S3 URI and the processing script's S3 URI.
AnswerD

SageMaker Pipelines caching allows a step to be skipped if the cache key (which includes input artifacts and parameters) matches a previous successful run. By including the input data S3 URI and the processing script's S3 URI in the cache key, the step will be skipped when neither has changed. Downstream steps will still run because the processing step's outputs are retrieved from cache.

Why this answer

SageMaker Pipelines supports step caching, where the cache key is derived from the step's input artifacts and parameters. By configuring a cache policy with a TTL on the processing step and ensuring the cache key includes the input data URI and script URI, the step is skipped when these inputs are unchanged. The cached outputs are then used by downstream steps, maintaining pipeline integrity.

Exam trap

The trap here is assuming that condition steps can automatically skip steps and reuse outputs, when in fact caching is the native mechanism for avoiding redundant step execution.

87
MCQmedium

A team has 200 small ML models that need to be served via HTTPS endpoints. Each model is used infrequently, and the team wants to minimize hosting costs. Which SageMaker deployment approach is MOST cost-effective?

A.Use SageMaker Serverless Inference for each model
B.Deploy each model on a separate real-time endpoint
C.Use Batch Transform for all models
D.Use a single multi-model endpoint (MME)
AnswerD

A multi-model endpoint loads multiple models behind one HTTPS endpoint, sharing the underlying instance and loading models on demand. With 200 infrequently used models, this avoids provisioning 200 always-on endpoints, directly satisfying the stem's cost-minimisation constraint.

Why this answer

A single multi-model endpoint (MME) hosts many models behind one HTTPS endpoint and loads them on demand into shared compute, which is ideal for 200 infrequently used small models because you pay for one endpoint's underlying instances rather than 200 separate ones. SageMaker dynamically loads and unloads models from S3 as requests arrive, dramatically reducing hosting cost while still providing real-time HTTPS inference.

Exam trap

The trap is assuming Serverless Inference is always cheapest for infrequent use; for many models behind one endpoint, MME's shared hosting is more cost-effective than many serverless endpoints.

How to eliminate wrong answers

Option A is wrong because Serverless Inference, while cost-effective for spiky traffic, is billed per invocation and per compute duration and is not designed to host 200 distinct models behind a single endpoint; managing 200 serverless endpoints adds overhead and can cost more than one MME. Option B is wrong because 200 separate real-time endpoints each incur continuous instance-hour charges, which is the most expensive option for infrequent use. Option C is wrong because Batch Transform is for offline, asynchronous batch scoring of large datasets, not for serving live HTTPS requests.

88
MCQmedium

A machine learning team uses SageMaker Pipelines to automate retraining. They want to avoid re-running data processing steps if the data has not changed since the last successful pipeline run. Which built-in feature should they enable?

A.Pipeline caching
B.Model lineage tracking
C.Parameterized pipeline executions
D.Step parallelism
AnswerA

Pipeline caching skips steps whose inputs, code and parameters are unchanged since the last successful run, reusing prior outputs. This satisfies the requirement to avoid re-running data processing when data has not changed, cutting cost and execution time.

Why this answer

Pipeline caching is the correct choice because SageMaker Pipelines can cache the outputs of each step based on a hash of the step's input parameters, configuration, and code. If the hash matches a previous successful run, the cached output is reused, avoiding redundant execution of data processing steps when the underlying data hasn't changed.

Exam trap

The trap here is that candidates confuse lineage tracking (Option B) with caching, assuming that tracking data versions automatically prevents re-execution, when in fact lineage only records history without affecting pipeline execution behavior.

How to eliminate wrong answers

Option B is wrong because model lineage tracking (via SageMaker ML Lineage Tracking) records the relationships between data, models, and training jobs, but it does not prevent re-running steps; it only provides auditability and provenance. Option C is wrong because parameterized pipeline executions allow you to pass different input values at runtime, but they do not automatically skip unchanged steps—caching is required for that. Option D is wrong because step parallelism controls the concurrency of step execution within a pipeline, not the reuse of previous outputs.

89
Multi-Selectmedium

A company is using AWS Step Functions to orchestrate their ML retraining pipeline. They want to trigger retraining when new data arrives, but only if the model's performance has degraded below a threshold. Which THREE AWS services should they use together to achieve this? (Choose three.)

Select 3 answers
A.AWS Step Functions
B.AWS Lambda
C.Amazon EventBridge
D.Amazon CloudWatch Logs
E.SageMaker Model Registry
AnswersA, B, C

Step Functions provides the orchestration layer, chaining the data-arrival event, the degradation check and the retraining workflow into one state machine, which satisfies the requirement to trigger retraining conditionally rather than on every new object.

Why this answer

Amazon EventBridge (C) is the correct event-routing service to detect the arrival of new data (e.g., an S3 PutObject event) and trigger the pipeline, since it can match events and invoke targets such as Step Functions. AWS Step Functions (A) is correct because it orchestrates the multi-step ML retraining workflow, including the conditional logic that checks whether model performance has degraded below the threshold before proceeding. AWS Lambda (B) is correct because it provides the serverless compute to evaluate the performance metric against the threshold and return a decision that Step Functions uses in its Choice state to branch into retraining or stop.

Amazon CloudWatch Logs (D) is not correct because it is a logging/monitoring service, not an event trigger or orchestration component for this pipeline. SageMaker Model Registry (E) is not correct because, while it tracks model versions and approval status, it does not itself trigger retraining based on incoming data or performance degradation.

Exam trap

MLA-C01 often tests the confusion between orchestration services (Step Functions), compute/glue logic (Lambda), and event routing (EventBridge), catching candidates who include Model Registry or CloudWatch Logs as part of the trigger/evaluate/orchestrate trio.

90
Multi-Selectmedium

A company uses SageMaker Pipelines to automate their ML workflow. They need to add model versioning and approval workflow. Which THREE steps should they include in their pipeline to achieve this? (Choose THREE.)

Select 3 answers
A.RegisterModel step
B.Training step
C.Condition step
D.Processing step for evaluation
E.Transform step
AnswersA, C, D

RegisterModel publishes a trained model artefact to the SageMaker Model Registry, creating a numbered model package version each run. This supplies the versioning half of the requirement, letting the pipeline track distinct model iterations and their approval status.

Why this answer

Option A (RegisterModel step) is correct because it creates a model package in the SageMaker Model Registry, which is the core mechanism for versioning models and tracking their approval status. Option C (Condition step) is correct because it enables the pipeline to branch based on evaluation results or approval status, allowing the workflow to proceed only when the model meets criteria or is approved. Option D (Processing step for evaluation) is correct because it runs the evaluation logic that produces the metrics used by the Condition step to decide whether the model should be registered or approved.

Option B (Training step) is not one of the three required steps here since training alone does not provide versioning or approval; it is typically a prerequisite but not the mechanism for versioning/approval. Option E (Transform step) is not correct because batch transform is for inference, not for model versioning or approval workflow.

Exam trap

The trap here is that candidates may think the Training step alone suffices for versioning, but AWS explicitly separates model training from model registration, requiring the RegisterModel step for registry integration.

91
Multi-Selecteasy

A company wants to trigger a model retraining pipeline whenever new training data arrives in an S3 bucket. They also need to send a notification to a Slack channel when the retraining completes. Which TWO AWS services should they use to implement this event-driven workflow? (Select TWO.)

Select 2 answers
A.Amazon SQS
B.AWS Lambda
C.AWS CloudTrail
D.Amazon EventBridge
E.SageMaker Model Registry
AnswersB, D

AWS Lambda provides the compute layer that reacts to the S3 event notification, invoking code to start the retraining pipeline without polling. It satisfies the event-driven trigger requirement, and the same function can publish to Slack via a webhook once retraining completes, meeting the notification constraint.

Why this answer

AWS Lambda is correct because it can be triggered directly by S3 events (e.g., s3:ObjectCreated) to invoke the model retraining pipeline. Amazon EventBridge is correct because it can capture completion events from the retraining pipeline (e.g., SageMaker training job state changes) and route them to a target like a Slack webhook via Lambda or SNS, enabling the notification workflow.

Exam trap

A common pitfall is confusing event-trigger services (Lambda, EventBridge) with storage/audit services (SQS, CloudTrail, Model Registry). Candidates often select SQS for decoupling or CloudTrail for monitoring, but these do not directly trigger the retraining or send notifications in an event-driven manner.

92
MCQeasy

A company has 50 small PyTorch models that are used infrequently for inference. They want to minimize costs while maintaining the ability to serve all models from a single endpoint. Which SageMaker feature should they use?

A.Multi-container endpoint
B.Batch transform job
C.Real-time endpoint with 50 production variants
D.Multi-model endpoint
AnswerD

A multi-model endpoint hosts many models behind one endpoint, loading each into memory on invocation and unloading idle ones. This shares the instance across all 50 infrequently used PyTorch models, satisfying the cost-minimisation and single-endpoint constraints.

Why this answer

SageMaker multi-model endpoints allow hosting multiple models on a single endpoint, loading them on demand and sharing the same serving container. This is cost-effective for many infrequently used models because you only pay for the endpoint instance and not for separate endpoints per model. It supports hundreds of models and dynamically loads them from S3.

Exam trap

The trap is confusing multi-model endpoints with multi-container endpoints; multi-container is for different frameworks or models that need separate containers, while multi-model is for many models sharing the same container and is more cost-effective for infrequent use.

How to eliminate wrong answers

Option A is wrong because multi-container endpoints are for hosting multiple containers that each serve a different model, but they are limited to a small number (up to 15) and are not designed for 50 models; they also require all containers to be running, increasing cost. Option B is wrong because batch transform is for offline, batch inference, not for serving from a single endpoint in real-time. Option C is wrong because real-time endpoint with 50 production variants would require 50 separate models deployed on the same endpoint, which is not cost-effective and has limits (initially 10 variants per endpoint, though can be increased, but still expensive).

93
Multi-Selectmedium

A machine learning team needs to deploy a PyTorch model that has been compiled with SageMaker Neo to improve inference performance on edge devices. Which TWO statements about SageMaker Neo are correct? (Select TWO.)

Select 2 answers
A.Neo reduces model inference latency through optimization techniques
B.Neo requires the model to be trained on SageMaker
C.Neo compiles models for a specific hardware target, such as Intel or ARM
D.Neo can only compile models trained with SageMaker built-in algorithms
E.Neo automatically scales SageMaker endpoints based on demand
AnswersA, C

SageMaker Neo applies operator fusion, constant folding and quantisation-aware compilation to shrink the model graph, directly satisfying the stem's requirement to improve inference performance on constrained edge hardware. This optimisation lowers per-inference latency rather than merely packaging the PyTorch artefacts for deployment.

Why this answer

Option A is correct because SageMaker Neo applies compiler-level optimizations such as operator fusion, constant folding, and quantization-aware graph rewrites that reduce inference latency and model size on the target hardware. Option C is correct because Neo is a compilation service that takes a trained model plus a target hardware specification (for example, an Intel x86 CPU with a specific instruction set, an ARM Cortex-A SoC, or an NVIDIA Jetson GPU) and produces a hardware-optimized executable, so the compiled artifact is tied to that target. Option B is incorrect because Neo can compile models trained anywhere, including on-premises or in other clouds, as long as the model is provided in a supported framework format such as PyTorch, TensorFlow, MXNet, or ONNX.

Option D is incorrect because Neo supports custom models and frameworks, not just SageMaker built-in algorithms. Option E is incorrect because endpoint auto scaling is handled by Application Auto Scaling and SageMaker endpoint scaling policies, not by Neo, which only handles model compilation.

Exam trap

MLA-C01 often tests the misconception that SageMaker Neo is tied to SageMaker training or built-in algorithms, when in fact it is a standalone compilation service that works with models from any source.

94
MCQhard

A financial services company needs to deploy a machine learning model for real-time fraud detection. The model must be highly available across multiple Availability Zones and must support automatic scaling based on request volume. The company also needs to perform canary deployments to test new model versions with a small percentage of traffic before full rollout. Which SageMaker feature should they use?

A.SageMaker real-time endpoint with production variants
B.SageMaker Multi-Model Endpoint
C.SageMaker Batch Transform
D.SageMaker Serverless Inference
AnswerA

A SageMaker real-time endpoint with production variants hosts multiple model versions behind one endpoint and supports weighted traffic splitting, enabling canary deployments. Multi-AZ deployment and automatic scaling satisfy the availability and request-volume scaling constraints in the stem.

Why this answer

SageMaker real-time endpoints with production variants enable canary deployments by routing a small percentage of traffic to a new model version while the majority goes to the current version. This feature also supports multi-AZ deployment for high availability and automatic scaling based on request volume via Application Auto Scaling, meeting all the stated requirements.

Exam trap

AWS often tests the distinction between real-time endpoints with production variants and Multi-Model Endpoints, where candidates mistakenly think Multi-Model Endpoints support canary deployments because they can host multiple models, but they lack traffic splitting and weighted routing capabilities.

How to eliminate wrong answers

Option B is wrong because SageMaker Multi-Model Endpoint hosts multiple models on the same endpoint but does not support canary deployments or traffic shifting between model versions; it is designed for cost-efficient hosting of many models, not staged rollouts. Option C is wrong because SageMaker Batch Transform is for offline, asynchronous inference on large datasets, not real-time fraud detection with low latency and automatic scaling. Option D is wrong because SageMaker Serverless Inference automatically scales to zero and has a cold start latency that is unsuitable for real-time fraud detection requiring consistent sub-second response times, and it does not support canary deployments with traffic splitting.

← PreviousPage 2 of 2 · 94 questions total

Ready to test yourself?

Try a timed practice session using only Mla Deployment Orchestration questions.