MLA-C01 ML Model Development Practice Question
A machine learning engineer is training a tabular regression model using the SageMaker built-in XGBoost algorithm. They want to reduce overfitting and improve generalization without changing the algorithm. Which SageMaker hyperparameter should they tune to control the fraction of features randomly sampled per tree?
⚠ Common exam trap
A common mix-up: candidates confuse row subsampling (subsample) with column subsampling (colsample_bytree) when the scenario explicitly asks for feature sampling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
colsample_bytree
The SageMaker built-in XGBoost algorithm exposes colsample_bytree to control column subsampling per tree. Setting it below 1.0 introduces feature-level randomness that combats overfitting and can improve generalization on tabular data. Other hyperparameters such as subsample, eta, and max_depth affect different aspects of training and do not implement the requested feature-sampling behavior.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
colsample_bytree
Why this is correct
In the SageMaker built-in XGBoost algorithm, colsample_bytree specifies the subsample ratio of columns when constructing each tree. Lowering it introduces feature-level randomness, which reduces overfitting and often improves generalization on tabular regression tasks. It is a native XGBoost hyperparameter exposed by the SageMaker estimator, so tuning it directly addresses the scenario without changing the algorithm.
- ✗
subsample
Why it's wrong here
subsample controls the fraction of training rows sampled per boosting round, not the fraction of features. While it also adds randomness and can reduce overfitting, the scenario specifically asks for feature sampling. Using subsample would alter row sampling behavior, which is a different regularization axis and would not satisfy the stated requirement of controlling features per tree.
- ✗
max_depth
Why it's wrong here
max_depth limits tree depth and thus model complexity, which can reduce overfitting, but it does not perform feature sampling. The requirement is specifically about the fraction of features randomly chosen per tree. Adjusting depth changes the tree structure rather than injecting column subsampling randomness, so it would not meet the stated goal in the SageMaker built-in XGBoost algorithm.
- ✗
eta
Why it's wrong here
eta, or learning rate, scales each boosting step's contribution. Reducing eta can slow overfitting but does not control feature subsampling; it changes optimization dynamics. The engineer asked for the fraction of features randomly sampled per tree, which is not governed by eta. Tuning eta alone would not directly implement the requested feature-level regularization mechanism in the SageMaker XGBoost container.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.