Amazon Web Services · Free Practice Questions · Last reviewed May 2026
24real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
A data science team uses Amazon SageMaker to train models on a large dataset stored in S3. The dataset is 500 GB in CSV format and is updated daily. The team wants to optimize data loading for training jobs to reduce I/O wait time. Which data ingestion strategy is MOST effective?
Use SageMaker File input mode and increase the EBS volume size to 1 TB.
Use SageMaker Pipe input mode to stream data directly from S3.
Pipe mode streams training data directly from S3 into the algorithm container, eliminating the download-to-local-storage phase and reducing I/O wait and startup time for large datasets. File mode would copy the full 500 GB first, which is slower and consumes substantial disk.
Convert the CSV files to Parquet format and use File input mode.
Load the data into an Amazon EFS file system and mount it to the training instance.
A data engineer needs to transform large CSV files stored in S3 into Parquet format and load them into a data warehouse for analysis. The transformation must be cost-effective and serverless. Which AWS service should be used?
Amazon Athena
Amazon EMR with Spark
AWS Glue
AWS Glue is serverless and pay-per-use, satisfying the cost-effective, no-infrastructure constraint. Its Spark-based ETL jobs read CSV from S3, transform schemas, and write Parquet directly, with the Glue Data Catalog and crawlers handling metadata for warehouse loading without managing clusters.
AWS Data Pipeline
A company uses Amazon Kinesis Data Firehose to deliver streaming data to an S3 bucket. The data is JSON and must be partitioned by year, month, and day. The delivery stream is configured with a buffer interval of 60 seconds and buffer size of 5 MB. The data producer sends about 1 MB per second. The data is arriving in S3 but the partitions are not being created as expected. What is the MOST likely reason?
The data is encrypted with AWS KMS and Firehose cannot write to encrypted buckets.
The delivery stream does not have dynamic partitioning enabled with the appropriate custom prefix.
Dynamic partitioning is what evaluates the JSON keys and derives the year/month/day S3 prefix per record; without it enabled and a matching custom prefix, Firehose writes to a flat default location. Buffer settings only affect flush timing, not partition paths, so they cannot explain the missing partitions.
The buffer interval is too short for the data volume, causing incomplete records.
The S3 bucket has versioning enabled, which prevents partitioning.
A data scientist uses Amazon SageMaker to train a model. The training dataset is 10 GB and stored in S3. The training job uses a ml.m5.large instance. The data must be available on the local file system during training. Which input mode should be used?
Local input mode
Batch input mode
File input mode
File mode downloads data to the local file system, making it available for training.
Pipe input mode
A company uses Amazon Redshift as a data warehouse. They need to load 50 TB of clickstream data from S3 into Redshift daily. The data arrives in 5-minute intervals as gzipped CSV files. The target table has a sort key and a distribution key. The load must complete within 2 hours. Which approach is MOST efficient?
Use AWS Glue to transform the data and write to Redshift using JDBC.
Use a staging table and then merge using a stored procedure.
Use a series of INSERT statements from a Lambda function.
Use the COPY command with a manifest file and gzip compression.
COPY is optimized for bulk loading from S3.
A machine learning engineer needs to process a large dataset that does not fit on a single Amazon SageMaker notebook instance's EBS volume. The data is stored in S3. What is the MOST efficient way to access the data from the notebook?
Increase the EBS volume size to 5 TB.
Mount the S3 bucket as a file system using s3fs.
Read the data directly from S3 using the boto3 library.
Streaming objects directly from S3 via boto3 avoids copying the dataset onto the notebook's EBS volume, so the size constraint is bypassed entirely. The instance reads only needed objects on demand, which is more efficient than downloading or resizing storage.
Use SageMaker File input mode in the notebook.
Want more Data Engineering practice?
Practice this domain20% of exam · 6 sample questions below
A company is using Amazon SageMaker to train a model with a custom algorithm. The training script reads data from an S3 bucket using boto3. The training job fails with an 'AccessDenied' error when trying to access the S3 bucket. The IAM role attached to the SageMaker notebook instance has full S3 access. What is the most likely cause?
The S3 bucket has a bucket policy that denies access from the SageMaker service.
The SageMaker execution role used for the training job does not have S3 access permissions.
SageMaker training jobs run under a separate execution role, not the notebook instance role. The notebook's full S3 access is irrelevant; without S3 permissions on the execution role, boto3 calls inside the training container return AccessDenied.
The training script is using an incorrect S3 bucket name.
The SageMaker training job is not configured to use the S3 VPC endpoint.
An IAM policy is attached to a SageMaker execution role. A data scientist tries to create a training job using a custom algorithm stored in an ECR repository. The training job fails with an 'AccessDenied' error when pulling the Docker image from ECR. What is the missing permission?
ecr:GetDownloadUrlForLayer and ecr:BatchGetImage on the ECR repository
SageMaker pulls the image using the execution role, which requires both ecr:GetDownloadUrlForLayer and ecr:BatchGetImage on the repository; BatchGetImage alone returns metadata, while GetDownloadUrlForLayer supplies the layer download URL. Granting both on the ECR repository resolves the AccessDenied failure during image pull.
ecr:PutImage on the ECR repository
s3:GetObject on the ECR repository
sagemaker:CreateTrainingJob on the ECR resource
A DevOps engineer created a SageMaker notebook instance using the Terraform configuration shown. The notebook instance is in a VPC with a public subnet. However, the notebook instance cannot access the internet. What is the most likely cause?
The role_arn is incorrect or missing permissions.
The instance type ml.t2.medium does not support internet access.
The subnet does not have a route to an internet gateway.
Without a route to an internet gateway, the notebook cannot access the internet despite the setting.
The direct_internet_access parameter is set to 'Enabled' but should be 'Disabled'.
A company is using Amazon SageMaker to train a XGBoost model on a large dataset. The training job is taking a long time. The data scientist wants to reduce training time without sacrificing model accuracy. The dataset is 100 GB in CSV format stored in S3. What is the most effective approach?
Reduce the number of instances to avoid communication overhead.
Use Pipe mode to stream data from S3 instead of downloading it first.
Pipe mode streams training data directly from S3 into the algorithm, avoiding the full download and disk conversion that File mode performs. This reduces startup and epoch time on the 100 GB CSV dataset while preserving accuracy, since all records are still consumed.
Use random sampling to reduce the dataset size to 10 GB.
Use SageMaker Managed Spot Training to reduce cost, but training time may increase due to interruptions.
A data scientist is deploying a model using Amazon SageMaker for real-time inference. The model is memory-intensive and requires a GPU. Which instance type should be selected for the endpoint?
i3.2xlarge
c5.2xlarge
r5.2xlarge
p3.2xlarge
The p3.2xlarge instance provides an NVIDIA Tesla V100 GPU with substantial memory, matching the endpoint's GPU and memory-intensive inference requirement. CPU-based instances such as m5 would lack GPU acceleration, making real-time inference impractically slow for this model.
A data scientist is deploying a model on Amazon SageMaker for real-time inference. The model is a PyTorch model that requires custom inference code. The data scientist needs to handle variable-length inputs and optimize inference latency. Which TWO steps should the data scientist take? (Choose TWO.)
Enable SageMaker batch transform to process requests in batches.
Use the SageMaker PyTorch container without any modifications.
Set the endpoint to use multiple variants for A/B testing.
Use TorchScript to compile the model for optimized inference.
TorchScript traces or scripts the PyTorch model into a serialised, optimised representation that SageMaker serves without the Python interpreter overhead, cutting inference latency. It also handles dynamic, variable-length inputs through scripting, satisfying the custom-code and latency constraints.
Provide a custom inference script (inference.py) that defines how to load the model and process requests.
SageMaker real-time endpoints require an inference.py entry point defining model_fn, input_fn, predict_fn and output_fn, letting the data scientist load the PyTorch model and handle variable-length inputs. This satisfies the stem's custom inference code and latency-optimisation requirement.
Want more Machine Learning Implementation and Operations practice?
Practice this domainA team is building a product recommendation system using matrix factorization in Amazon SageMaker. They notice that the model's training loss decreases steadily but validation loss starts increasing after 5 epochs. What is the most likely cause?
Underfitting
Not enough training data
Learning rate too high
Overfitting
Diverging losses — training loss falling while validation loss rises after five epochs — is the classic signature of overfitting: the model memorises training data and generalises poorly. Regularisation, early stopping, or more data would counter it.
A company is using Amazon SageMaker to train a deep learning model on a large dataset. The training job is taking too long. The team wants to reduce training time without changing the model architecture. Which action should they take?
Increase the learning rate by a factor of 10
Use SageMaker's distributed training with multiple instances
Distributed training partitions the workload across multiple instances, so gradient computation and data throughput are parallelised rather than bound to one machine. This shortens wall-clock training time while leaving the model architecture unchanged, as the question requires.
Reduce the number of epochs
Reduce the batch size
A data scientist is deploying a regression model in Amazon SageMaker that predicts housing prices. The model shows high bias (underfitting). Which action is most likely to reduce bias?
Reduce the amount of training data
Increase regularization strength
Use a simpler model
Add more features or increase model complexity
High bias means the model is too simple to capture the underlying relationship, so adding informative features or increasing model complexity raises its capacity to fit the data, directly reducing underfitting. Regularisation would worsen bias, not help.
A data scientist is training a deep learning model using Amazon SageMaker. The training loss is decreasing, but the validation loss starts increasing after 10 epochs. The model is overfitting. Which TWO actions should the data scientist take to reduce overfitting? (Choose 2.)
Increase the number of layers
Remove L2 regularization
Increase the number of training steps
Add dropout layers
Dropout layers randomly deactivate a proportion of neurons during each training iteration, forcing the network to learn more robust, distributed representations rather than memorising training samples. This directly counteracts the overfitting indicated by rising validation loss after epoch ten.
Add early stopping based on validation loss
Early stopping halts training once validation loss stops improving, directly countering the overfitting pattern described. It restores the best-performing weights rather than the final epoch's, so the model retains generalisation instead of memorising the training set.
An e-commerce company uses a linear regression model to predict customer lifetime value (LTV). The model shows high variance on the test set, with training RMSE much lower than test RMSE. Which of the following is the MOST effective approach to reduce overfitting?
Apply L2 regularization (Ridge regression)
L2 regularization adds a penalty proportional to the squared magnitude of the coefficients to the loss function, shrinking weights and reducing variance. This directly addresses the large train-to-test RMSE gap, the stem's stated overfitting constraint, more reliably than gathering more data.
Use a polynomial kernel in a support vector regressor
Add more features, including interaction terms
Increase training data size by duplicating existing samples
A company wants to use Amazon SageMaker to train a deep learning model using a custom TensorFlow script. The data is stored in an S3 bucket. Which SageMaker API operation should be used to launch the training job?
CreateHyperParameterTuningJob
CreateEndpoint
CreateTransformJob
CreateTrainingJob
CreateTrainingJob is the SageMaker API that launches a training job, accepting the Docker image, S3 input channels, instance type, and hyperparameters. It provisions the specified compute, runs the custom TensorFlow script, and writes model artefacts to S3.
Want more Modeling practice?
Practice this domain24% of exam · 6 sample questions below
A team is building a regression model to predict house prices. The dataset includes a column 'zip_code' with 100 unique values. The data scientist one-hot encodes this column, resulting in 100 new binary columns. The model shows poor performance on a validation set. What is the most likely cause?
One-hot encoding introduced multicollinearity among the binary columns.
One-hot encoding reduced the number of features, causing underfitting.
The one-hot encoding introduced high variance, but the validation set has low variance.
The model suffers from the curse of dimensionality due to the large number of features.
One-hot encoding expands a single categorical column into 100 sparse binary features, inflating the feature space so observations become sparse and distances lose meaning. This sparsity-driven curse of dimensionality degrades the regression model's validation performance, rather than any encoding error or label leakage.
During exploratory data analysis, a data scientist plots the distribution of a numerical feature and observes a heavy right skew. The feature has many outliers at the high end. Which transformation is most appropriate to reduce skewness?
Apply a log transformation to the feature.
A log transformation compresses the large values that cause the heavy right skew, pulling high-end outliers closer to the bulk of the distribution and producing a more symmetric, near-normal shape. It directly satisfies the stem's requirement to reduce skewness, and is the standard remedy for positively skewed numerical features.
Apply z-score normalization.
Apply one-hot encoding.
Apply min-max scaling.
A data scientist is analyzing a dataset with missing values in 30% of the rows for the 'age' column. The data scientist decides to impute the missing values with the median of the observed 'age' values. What is a potential drawback of this approach?
The imputation will introduce bias if the missing values are not random.
Imputation using median is computationally expensive for large datasets.
The imputed values may reduce the variance of the 'age' distribution.
Replacing 30% of ages with a single median collapses those rows onto the centre, so the spread of the imputed column shrinks relative to the true distribution. This understates variance and weakens any model relying on age dispersion.
The imputed values will increase the variance of the feature, leading to overfitting.
A data scientist is exploring a dataset with 500 features and 10,000 samples. The data scientist computes the pairwise correlation matrix and finds that many features have correlations above 0.9. The data scientist wants to reduce the dataset to 50 features while preserving as much variance as possible. Which technique should be used?
Remove all but one feature from each group of highly correlated features.
Apply Principal Component Analysis (PCA) and keep the top 50 principal components.
PCA is an unsupervised linear transformation that projects the 500 correlated features onto orthogonal components ordered by explained variance. Keeping the top 50 components satisfies the 50-feature constraint while maximising retained variance, unlike correlation-based filtering.
Use Linear Discriminant Analysis (LDA) to project to 50 dimensions.
Use t-Distributed Stochastic Neighbor Embedding (t-SNE) to reduce to 50 dimensions.
During EDA, a data scientist notices that a feature has a high proportion of missing values (e.g., 70%). The feature is continuous and expected to be important based on domain knowledge. What is the best approach to handle this?
Remove the feature entirely to avoid bias.
Create a binary indicator for missingness and impute the continuous values with the median.
The missingness indicator preserves the information that a value was absent, which may itself correlate with the target, while median imputation keeps the continuous feature usable despite 70% gaps. This satisfies the constraint that the feature is important and continuous.
Impute missing values with -1 since it is out of range.
Drop all rows with missing values in that feature.
A data scientist is performing EDA on a dataset with 1,000 features and 10,000 rows. The target variable is binary. After checking for multicollinearity, the scientist finds many pairs of features with correlation > 0.95. Which action should be taken to prepare the data for modeling?
Apply PCA to all features to decorrelate them.
Standardize all features using StandardScaler.
For each highly correlated pair, remove one feature based on domain knowledge or higher correlation with target.
Dropping one feature from each pair correlated above 0.95 removes redundancy while retaining predictive signal, since the survivor is chosen by domain knowledge or stronger target correlation. This addresses the multicollinearity constraint without discarding information wholesale.
Randomly drop half of the correlated features.
Want more Exploratory Data Analysis practice?
Practice this domainThe MLS-C01 exam has 65 questions and must be completed in 180 minutes. The passing score is 750/1000.
Scenario-based questions covering exam objectives with detailed answer explanations.
The exam covers 4 domains: Data Engineering, Machine Learning Implementation and Operations, Modeling, Exploratory Data Analysis. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official Amazon Web Services MLS-C01 exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.