Courseiva
ModelinghardMultiple ChoiceObjective-mapped

Reducing SageMaker Inference Latency with Feature Importance

A company is deploying a real-time fraud detection system using a gradient boosting model on AWS SageMaker. The model uses 200 features and is trained on 50 GB of data. The inference latency requirement is under 10 ms per request. During load testing, the endpoint shows average latency of 15 ms. Which change is MOST likely to reduce latency below 10 ms?

Quick Answer

The core relationship this question is testing is that inference latency for a tree-based model like gradient boosting scales with how much work each prediction requires, and feature count is a direct driver of that work. Every feature a gradient boosting model considers adds potential split points for the trees to evaluate at inference time, so trimming 200 features down to the 50 most important ones, selected using the model's own feature importance scores, reduces the computation each request has to do without touching the underlying infrastructure. This is why it's the most direct lever available: it attacks the actual source of the 15 ms latency rather than working around it with more hardware, and because it's a modeling change rather than an infrastructure change, it doesn't introduce the cost or complexity that scaling up compute would. Feature reduction based on importance also has the advantage of typically preserving most of the model's predictive power, since low-importance features contribute little to the model's decisions in the first place, so latency comes down without a proportional loss in accuracy. When a scenario gives you a latency target that's just barely missed and highlights a large feature count as part of the setup, that's usually a signal that trimming the feature set, rather than upgrading hardware, is the intended, cost-effective fix.

⚠ Common exam trap

It's easy for candidates to assume GPU instances universally speed up inference, but for tree-based models like gradient boosting, the bottleneck is sequential tree traversal, not parallel computation, so feature reduction is the correct optimization.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Reduce the number of features to the top 50 based on feature importance

Reducing the number of features from 200 to the top 50 directly decreases the amount of data each inference request must process, which lowers both feature engineering overhead and model evaluation time. For gradient boosting models on SageMaker, fewer features mean fewer decision tree splits to traverse per prediction, which can significantly reduce latency without requiring hardware changes. This is the most direct and cost-effective way to meet the 10 ms requirement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Switch to a GPU-based instance type

    Why it's wrong here

    Gradient boosting models typically run faster on CPUs; GPUs are beneficial for neural networks.

  • Reduce the number of features to the top 50 based on feature importance

    Why this is correct

    Fewer features reduce inference computation time, directly lowering latency.

  • Increase the number of trees in the model

    Why it's wrong here

    More trees increase computation and latency.

  • Use a larger batch size for inference

    Why it's wrong here

    Batch size affects throughput, not per-request latency for real-time endpoints.

About these practice questions

One of 1,672 original MLS-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

3 more ways this is tested on MLS-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company has deployed a real-time inference endpoint using SageMaker for a fraud detection model. The model uses a Random Forest classifier. The endpoint receives predictions but the latency is too high. The metric shows p99 latency of 500ms, but the requirement is under 200ms. The team has already optimized the instance type to the maximum allowed by their budget. The data scientist suggests: A) Reducing the number of trees in the Random Forest model. B) Switching to a linear model like Logistic Regression. C) Enabling SageMaker's batch transform instead of real-time endpoint. D) Adding more instances to the endpoint behind a load balancer. Which option will MOST effectively reduce latency while maintaining acceptable accuracy?

easy
  • A.Switch to a linear model like Logistic Regression
  • B.Reduce the number of trees in the Random Forest model
  • C.Enable SageMaker's batch transform
  • D.Add more instances to the endpoint

Why B: (Reducing the number of trees) is the most effective method to reduce latency while maintaining acceptable accuracy. Fewer trees directly decrease inference time of the Random Forest model, although it may slightly impact accuracy. Switching to a linear model (Option A) would reduce latency but likely result in significant accuracy loss. Batch transform (Option C) is not suitable for real-time inference. Adding more instances (Option D) improves throughput but not per-request latency.

Variation 2. A company is deploying a real-time fraud detection model using Amazon SageMaker. The model must make predictions in under 100 milliseconds. The data scientist uses a pre-trained XGBoost model and deploys it to a SageMaker endpoint with an ml.c5.xlarge instance. After load testing, the average latency is 150 ms. Which action should the data scientist take to reduce latency?

medium
  • A.Reduce the number of trees in the XGBoost model
  • B.Deploy multiple instances behind a load balancer
  • C.Enable SageMaker Neo to compile the model for the target instance
  • D.Use a larger instance type to increase compute capacity

Why C: SageMaker Neo optimizes trained models for the target hardware platform by compiling them into an efficient runtime. This reduces inference latency without changing the model architecture, making it ideal for meeting the sub-100ms requirement when the current latency is 150ms on an ml.c5.xlarge instance.

Variation 3. A company is deploying a machine learning model for real-time fraud detection. The model must have low latency (under 100 ms) and high throughput. The data scientist trains a gradient boosting model and deploys it to a SageMaker endpoint with a single ml.c5.xlarge instance. During load testing, the endpoint exceeds the latency threshold. Which change is MOST likely to reduce latency?

hard
  • A.Replace the model with a simpler model, such as logistic regression
  • B.Use a larger instance type, such as ml.c5.4xlarge
  • C.Switch to batch transform for inference
  • D.Enable automatic scaling on the endpoint

Why A: Replacing the gradient boosting model with a simpler model like logistic regression reduces the computational complexity per inference. Gradient boosting involves traversing many decision trees, each requiring multiple conditional checks and arithmetic operations, while logistic regression is a single linear transformation. This directly lowers CPU utilization per request, reducing latency under the same instance resources.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.