Courseiva

CCNA Data Preparation for Machine Learning Questions

69 questions · Data Preparation for Machine Learning · All types, answers revealed

1
MCQmedium

A machine learning engineer is using Amazon SageMaker Feature Store to manage features for a fraud detection model. The engineer needs to ensure that the feature group can serve both batch and real-time predictions. The feature group is configured with an online store enabled. Which additional configuration is required to support batch predictions?

A.Increase the online store's read capacity to handle batch loads.
B.Set the feature group's event time to the ingestion time to allow batch queries.
C.Enable an offline store for the feature group and specify an S3 bucket for storage.
D.Configure the feature group to use a custom KMS key for encryption, which enables batch access.
AnswerC

To support batch predictions, the feature group must have an offline store, which stores historical feature data in Amazon S3. The online store is optimized for low-latency real-time serving, but batch predictions require access to historical data. Enabling the offline store and specifying an S3 bucket allows batch retrieval of features for training and batch inference.

Why this answer

Feature Store's online store is for real-time serving, while the offline store is for batch serving and training. To support batch predictions, the feature group must have an offline store enabled, which stores feature data in S3. This allows the engineer to retrieve historical features in bulk for batch inference jobs.

Without an offline store, only real-time serving is possible.

Exam trap

The trap here is assuming that the online store can handle batch predictions or that encryption or event time settings enable batch access, when in fact an offline store is required.

2
MCQhard

A company uses AWS Glue ETL jobs to transform data for machine learning. They have a dataset with a column 'income' that is heavily right-skewed. Which transformation should be applied to make the distribution more Gaussian-like?

A.Log transformation (natural log)
B.Standardization (z-score)
C.Min-max scaling to [0,1]
D.Equal-width binning
AnswerA

A natural log transformation compresses the long right tail of a heavily skewed variable, pulling extreme high values closer to the bulk of the data. This reduces skewness and stabilises variance, making the 'income' distribution approximately Gaussian-like, which many ML algorithms assume.

Why this answer

A log transformation is appropriate for heavily right-skewed data because it compresses the long tail by applying a concave function, pulling extreme values closer to the mean and making the distribution more symmetric. In AWS Glue ETL, you can apply this using Spark SQL's `LOG` function or a Python UDF with `numpy.log`, which directly addresses the skewness to better approximate a Gaussian distribution for downstream ML models.

Exam trap

The trap here is that candidates confuse scaling (standardization or min-max) with shape-changing transformations, assuming any normalization makes data Gaussian, when in fact only non-linear transformations like log or Box-Cox address skewness.

How to eliminate wrong answers

Option B is wrong because standardization (z-score) centers and scales data to have mean 0 and standard deviation 1, but it does not change the shape of the distribution—it only rescales, so right-skewness remains. Option C is wrong because min-max scaling to [0,1] linearly compresses the data into a fixed range, which preserves the relative distances and does not alter skewness or make the distribution Gaussian-like. Option D is wrong because equal-width binning discretizes the continuous 'income' column into fixed intervals, which loses granularity and does not transform the distribution toward Gaussian—it creates a categorical or ordinal feature instead.

3
MCQeasy

A company has 10 TB of log data in compressed JSON format stored in Amazon S3. The data needs to be processed and transformed into a structured format for machine learning. The processing requires complex transformations, including parsing nested JSON and joining with a reference table. The company wants to minimize infrastructure management. Which approach should the company use?

A.Use SageMaker Processing jobs to run custom scripts.
B.Use Amazon Athena to query and transform the data.
C.Use Amazon EMR with Apache Spark.
D.Use AWS Glue ETL with PySpark.
AnswerD

AWS Glue ETL with PySpark handles nested JSON parsing and reference-table joins at scale, while remaining serverless. This satisfies the 10 TB dataset and the requirement to minimise infrastructure management, since Glue provisions and scales the Spark environment automatically.

Why this answer

AWS Glue ETL with PySpark (Option D) is the best choice because it provides a fully serverless environment, minimizing infrastructure management. Glue can handle complex transformations like parsing nested JSON and joining with reference tables using PySpark, and it scales automatically for large datasets (10 TB). Amazon EMR (Option C) requires cluster management and provisioning, which contradicts the goal of minimizing management overhead.

Exam trap

Candidates may assume that large-scale data (10 TB) requires a provisioned cluster like EMR, but AWS Glue can scale to petabyte-scale workloads and is fully serverless, aligning with the goal of minimizing infrastructure management.

How to eliminate wrong answers

Option A is wrong because SageMaker Processing jobs are optimized for ML-specific tasks like training data preprocessing, not for general-purpose ETL on 10 TB of data; they lack native support for complex joins and nested JSON parsing at scale. Option B is wrong because Amazon Athena is a serverless query engine that excels at ad-hoc SQL queries but struggles with complex transformations like parsing deeply nested JSON and joining large reference tables due to its per-query pricing and lack of native procedural logic. Option D is wrong because AWS Glue ETL with PySpark is a valid alternative for ETL, but it is less performant and more expensive than EMR for large-scale (10 TB) data processing due to its auto-scaling overhead and limited tuning capabilities; EMR provides finer control over cluster configuration and cost optimization for batch jobs.

4
MCQmedium

A data engineer is preparing a large training dataset stored in Amazon S3 as many small Parquet files, and a SageMaker training job that reads directly from S3 is spending most of its time on the input channel rather than on model computation. The engineer needs to improve the input throughput without changing the model code or the training algorithm. Which action should the engineer take?

A.Enable SageMaker Training Compiler and set the framework to accelerated mode.
B.Use the SageMaker File System Input with Amazon FSx for Lustre linked to the S3 bucket.
C.Increase the number of records per S3 GET request by enabling S3 Transfer Acceleration on the bucket.
D.Convert the output to TFRecord format and use Pipe mode with the SageMaker TensorFlow estimator.
AnswerB

FSx for Lustre linked to the S3 bucket exposes the dataset as a high-throughput, low-latency POSIX file system that SageMaker training jobs can mount, so the many small Parquet files are read far faster than repeated S3 GET requests. It improves input throughput without changing model code or the training algorithm, satisfying the requirement.

Why this answer

Feeding a training job from a high-performance shared file system removes the request-per-small-object overhead that dominates when many tiny Parquet files are read directly from S3. Amazon FSx for Lustre linked to the S3 bucket presents the data as files that the training container mounts and reads at high throughput, improving the input channel without altering model code or the training algorithm, which is exactly what the scenario requires.

Exam trap

The trap here is assuming that any SageMaker input-mode or acceleration feature will fix slow data loading, when the real issue is per-object overhead from many small files rather than raw bandwidth.

5
MCQhard

A data scientist is preparing a large dataset (50 GB) for training a TensorFlow model on SageMaker. The dataset consists of many small CSV files. Training is slow due to I/O bottlenecks. Which data preparation strategy most effectively accelerates training?

A.Convert the dataset to TFRecord format and use tf.data pipeline with prefetching
B.Convert the dataset to Parquet format and use Apache Arrow for loading
C.Compress the CSV files and decompress during data loading
D.Use a larger instance type with more vCPUs
AnswerA

TFRecord stores records in a compact binary format, eliminating per-file parsing overhead from thousands of small CSVs. The tf.data pipeline with prefetching overlaps data loading with GPU computation, directly resolving the I/O bottleneck constraining training throughput. This satisfies the scenario's requirement to accelerate training on a 50 GB dataset.

Why this answer

TFRecord format stores data in a binary, row-oriented format that TensorFlow's tf.data API can read efficiently, especially with prefetching to overlap data loading with model computation. This eliminates the per-file open/parse overhead of many small CSV files, which is the primary cause of I/O bottlenecks in this scenario.

Exam trap

The trap here is that candidates often choose larger instances (Option D) as a brute-force fix, failing to recognize that the root cause is the small-file I/O pattern, which requires a format change (TFRecord) rather than more compute resources.

How to eliminate wrong answers

Option B is wrong because Parquet is a columnar storage format optimized for analytical queries and selective column reads, not for sequential row-by-row training loops typical in deep learning; Apache Arrow adds overhead without solving the small-file problem. Option C is wrong because compressing CSV files reduces storage size but increases CPU load during decompression, often worsening I/O bottlenecks due to the many small files still requiring individual decompression. Option D is wrong because increasing vCPUs does not fix the fundamental I/O bottleneck caused by many small files; it may even exacerbate contention on shared storage without addressing the file access pattern.

6
MCQeasy

A data engineer needs to prepare a large dataset for machine learning. The data is stored in an Amazon RDS MySQL database and needs to be transformed and moved to an S3 bucket in Parquet format for use with SageMaker. Which AWS service is most suitable for this extraction, transformation, and loading (ETL) task?

A.Use AWS Glue ETL jobs with PySpark to read from RDS, apply transformations, and write to S3 as Parquet.
B.Use Amazon Athena CTAS statements to copy data from RDS to S3.
C.Use SageMaker Data Wrangler to connect to RDS and export transformed data to S3.
D.Use Amazon EMR with Spark to read from RDS, transform, and write to S3.
AnswerA

AWS Glue ETL jobs with PySpark natively connect to Amazon RDS MySQL through JDBC, apply distributed transformations, and write Parquet directly to S3, satisfying the required format conversion. Glue's serverless Spark engine handles the large dataset scale without managing infrastructure, and its built-in SageMaker integration streamlines downstream machine learning consumption.

Why this answer

AWS Glue ETL jobs with PySpark are the most suitable service for this task because Glue is a fully managed, serverless ETL service that can natively connect to Amazon RDS MySQL via JDBC, apply transformations using PySpark, and write the output directly to S3 in Parquet format. This aligns perfectly with the requirement to extract, transform, and load a large dataset into a machine-learning-ready format without managing infrastructure.

Exam trap

The trap here is that candidates may confuse SageMaker Data Wrangler's ability to connect to RDS and export data with a full ETL capability, overlooking that it is an interactive tool for data preparation within SageMaker Studio rather than a serverless batch ETL service like AWS Glue.

How to eliminate wrong answers

Option B is wrong because Amazon Athena CTAS statements cannot read directly from Amazon RDS; Athena only queries data already in S3 or other data sources via federated queries, but CTAS itself requires the source to be in S3 or a cataloged table, not a live RDS database. Option C is wrong because SageMaker Data Wrangler is designed for interactive data preparation and feature engineering within SageMaker Studio, not for running serverless ETL jobs at scale; it can import data from RDS but lacks the native ability to schedule or run large-scale batch transformations and write to S3 as Parquet without additional infrastructure. Option D is wrong because while Amazon EMR with Spark can technically perform this task, it requires provisioning and managing a cluster, which adds operational overhead; AWS Glue is more suitable as a serverless, cost-effective alternative for this specific ETL workload without the need to manage EC2 instances or cluster lifecycle.

7
Multi-Selecthard

A team is preparing text data for a natural language processing (NLP) model. They have a corpus of customer reviews. Which THREE preprocessing steps are essential to reduce noise and improve model performance?

Select 3 answers
A.Apply one-hot encoding to each word
B.Remove punctuation and special characters
C.Compute TF-IDF vectors
D.Perform stemming or lemmatization
E.Convert all text to lowercase
AnswersB, D, E

Removing punctuation and special characters strips non-linguistic tokens that inflate vocabulary size and dilute token frequency statistics, directly reducing noise in the customer review corpus. This normalisation step ensures the NLP model learns from meaningful word patterns rather than artefacts like commas, hashtags or emojis, satisfying the stem's noise-reduction requirement.

Why this answer

Option B is correct because removing punctuation and special characters eliminates non-linguistic symbols that add noise and are typically not useful features for NLP models, helping normalize the token stream. Option D is correct because stemming or lemmatization reduces inflected words to their base or dictionary form (e.g., 'running' to 'run'), decreasing vocabulary size and helping the model generalize across morphological variants. Option E is correct because converting all text to lowercase ensures that words like 'Review' and 'review' are treated as the same token, preventing spurious vocabulary duplication and improving consistency.

Option A is not a noise-reduction preprocessing step; one-hot encoding is a feature representation technique that actually increases dimensionality and does not clean the text. Option C is also not a preprocessing cleaning step; TF-IDF is a numerical vectorization/weighting method applied after text has already been normalized and tokenized.

Exam trap

AWS often tests the distinction between preprocessing steps (cleaning) and feature engineering steps (vectorization), so the trap here is that candidates mistake TF-IDF or one-hot encoding as essential preprocessing for noise reduction when they are actually downstream representation techniques.

8
MCQeasy

A data scientist is preparing a dataset stored in Amazon S3 for a SageMaker training job. The dataset contains missing values in several columns. The scientist wants to impute missing values with the mean of each column. Which SageMaker built-in algorithm or processing method should be used to perform this imputation efficiently?

A.Use the SageMaker 'PCA' algorithm to fill missing values.
B.Use the SageMaker 'Linear Learner' algorithm to predict missing values.
C.Use the SageMaker 'Impute' transform in Data Wrangler.
D.Use the SageMaker 'XGBoost' algorithm to predict missing values.
AnswerC

SageMaker Data Wrangler offers an 'Impute' transform that allows replacing missing values with various strategies, including mean, median, or mode. This is a straightforward and efficient way to handle missing numerical data within the Data Wrangler interface, which can then be exported to a processing job or pipeline. It directly fulfills the requirement of mean imputation.

Why this answer

Mean imputation is a common preprocessing step for numerical features. In SageMaker Data Wrangler, the 'Impute' transform provides a user-friendly way to replace missing values with the column mean, median, or mode. This transform can be applied to multiple columns and is part of the data preparation flow.

It ensures that the dataset is complete before feeding it into a training algorithm, which is essential for algorithms that cannot handle missing values natively.

Exam trap

The trap here is assuming that built-in algorithms like XGBoost handle missing values, but the question specifically asks for mean imputation, which requires a dedicated preprocessing transform.

9
MCQmedium

A data engineer needs to prepare a dataset for a fraud detection model. The dataset contains a highly skewed numerical feature with extreme outliers. The engineer decides to apply a logarithmic transformation to this feature before training. Which SageMaker Data Wrangler transform should be used to apply the logarithmic transformation?

A.Use the 'Standard Scaler' transform in Data Wrangler.
B.Use the 'One-Hot Encoding' transform in Data Wrangler.
C.Use the 'Log Transform' transform in Data Wrangler.
D.Use the 'Min-Max Scaler' transform in Data Wrangler.
AnswerC

The 'Log Transform' transform in SageMaker Data Wrangler applies a natural logarithm (base e) to the selected numeric column. This is specifically designed to reduce right skewness and mitigate the impact of extreme outliers, making the feature more suitable for models that assume normality or are sensitive to scale. It directly addresses the scenario's need for a logarithmic transformation.

Why this answer

The logarithmic transformation is a common technique to reduce right skewness and stabilize variance in numerical data. In SageMaker Data Wrangler, the 'Log Transform' transform applies a natural logarithm to the selected column, effectively compressing the scale of large values and making the distribution more symmetric. This helps models that are sensitive to feature distributions, such as linear models, to perform better.

Other transforms like scaling or encoding do not change the distribution shape.

Exam trap

The trap here is confusing scaling transforms with distribution-changing transforms; scaling does not alter skewness or outliers.

10
MCQmedium

A team is using Amazon SageMaker for feature engineering. They have a dataset with a column 'TransactionDate' in string format (e.g., '2023-01-15 10:30:00'). They need to create features: year, month, day, hour, and day_of_week. What is the most efficient way to do this in a SageMaker processing job?

A.Use pandas datetime functions and then split
B.Use SageMaker built-in first party algorithms
C.Use AWS Glue for transformation
D.Use SQL query in Athena on S3 data
AnswerA

Pandas' `pd.to_datetime` parses the string column into datetime64 in one vectorised pass, then `.dt.year`, `.dt.month`, `.dt.day`, `.dt.hour` and `.dt.dayofweek` extract all five features directly. This satisfies the efficiency constraint by avoiding per-row Python loops or repeated parsing inside the SageMaker processing job.

Why this answer

Using pandas datetime functions within a SageMaker processing job is the most efficient approach for this task. SageMaker processing jobs run custom Python scripts, and pandas provides vectorized operations (e.g., `pd.to_datetime()`, `.dt.year`, `.dt.month`, `.dt.day`, `.dt.hour`, `.dt.dayofweek`) that parse the string column and extract all required features in a single pass without external dependencies or data movement.

Exam trap

AWS often tests the misconception that SageMaker built-in algorithms can handle feature engineering, but they are strictly for training and inference, not data preprocessing — the trap here is assuming 'first-party algorithms' include data transformation capabilities.

How to eliminate wrong answers

Option B is wrong because SageMaker built-in first-party algorithms (e.g., XGBoost, Linear Learner) are designed for model training, not for feature engineering or data transformation tasks like datetime parsing. Option C is wrong because AWS Glue is an ETL service that introduces additional overhead (e.g., Spark cluster startup, schema inference) and is less efficient for a simple in-memory pandas operation within a SageMaker processing job. Option D is wrong because using SQL in Athena on S3 data requires querying the raw data from S3, which incurs scan costs and latency, and Athena's SQL functions for datetime extraction (e.g., `EXTRACT`) are less flexible and slower than pandas for this specific transformation.

11
MCQhard

A data engineer is using Amazon SageMaker Processing to run a data preprocessing script on a dataset with 500 million rows. The script runs out of memory on a single ml.r5.24xlarge instance. The engineer needs to modify the processing job to handle the dataset size. Which approach is most cost-effective and scalable?

A.Configure the Processing job with multiple instances and use ShardedByS3Key for data splitting.
B.Write the script to process data in chunks and write intermediate results to local ephemeral storage.
C.Increase the instance type to a larger one like ml.p3dn.24xlarge with more memory.
D.Reduce the number of instances to one and increase the volume size for swap space.
AnswerA

ShardedByS3Key distributes objects across multiple instances, so each processes a subset in parallel and memory per instance stays bounded. Scaling horizontally on smaller instances costs less than one oversized ml.r5.24xlarge and handles 500 million rows.

Why this answer

SageMaker Processing with ShardedByS3Key splits the input dataset by S3 object boundaries across multiple instances, allowing distributed processing of the 500 million rows without exceeding memory on any single instance. This approach is cost-effective as it uses multiple smaller instances (e.g., ml.r5.xlarge) rather than a single oversized instance, and scales linearly with data size.

Exam trap

AWS often tests the misconception that increasing instance size or using swap space is the primary solution for memory issues, whereas the correct approach is to distribute the workload horizontally using SageMaker's built-in data sharding feature.

How to eliminate wrong answers

Option B is wrong because writing intermediate results to local ephemeral storage does not solve the out-of-memory issue; the script still loads the entire dataset into memory before chunking, and local storage is limited and not designed for large-scale intermediate data. Option C is wrong because increasing to a larger instance like ml.p3dn.24xlarge (which has 192 GB memory vs. ml.r5.24xlarge's 768 GB) actually reduces memory, and GPU instances are not optimized for memory-intensive preprocessing; this approach is neither cost-effective nor scalable. Option D is wrong because reducing to a single instance and increasing volume size for swap space relies on disk-based swapping, which is orders of magnitude slower than RAM and will cause severe performance degradation or job failure due to I/O bottlenecks.

12
MCQhard

A machine learning engineer is using Amazon SageMaker Processing to preprocess a large dataset. The processing job runs a custom Python script that uses the pandas library to read multiple CSV files from an S3 input prefix. The script must write the processed output to a different S3 prefix. Which configuration of the ProcessingInput and ProcessingOutput parameters is correct for this scenario?

A.ProcessingInput with source set to the S3 input prefix and destination set to the S3 output prefix; ProcessingOutput with source set to the S3 input prefix and destination set to '/opt/ml/processing/output'.
B.ProcessingInput with source set to '/opt/ml/processing/input' and destination set to '/opt/ml/processing/output'; ProcessingOutput with source set to the S3 input prefix and destination set to the S3 output prefix.
C.ProcessingInput with source set to the S3 input prefix and destination set to '/opt/ml/processing/input'; ProcessingOutput with source set to '/opt/ml/processing/output' and destination set to the S3 output prefix.
D.ProcessingInput with source set to '/opt/ml/processing/input' and destination set to the S3 input prefix; ProcessingOutput with source set to the S3 output prefix and destination set to '/opt/ml/processing/output'.
AnswerC

This configuration correctly maps the S3 input prefix to a local directory inside the processing container and the local output directory back to S3. The source for ProcessingInput is the S3 URI, and destination is the local path where the data will be available. For ProcessingOutput, source is the local path where the script writes outputs, and destination is the S3 URI. This is the standard pattern for SageMaker Processing jobs.

Why this answer

In SageMaker Processing, ProcessingInput specifies the S3 source and the local destination path where the data will be mounted in the container. ProcessingOutput specifies the local source path where the script writes output and the S3 destination where it will be uploaded. This bidirectional mapping ensures that the container can access input data and persist output data to S3.

The correct configuration uses S3 URIs for external locations and local paths for container locations.

Exam trap

The trap here is mixing up the source and destination fields for ProcessingInput and ProcessingOutput; source always refers to the origin (S3 for input, local for output) and destination to the target (local for input, S3 for output).

13
MCQmedium

A data engineer must prepare a 4 TB Parquet dataset stored in Amazon S3 for a SageMaker training job that runs on 8 ml.p4d.24xlarge instances. The engineer wants the fastest possible data throughput during training while minimizing per-epoch I/O overhead. The dataset is immutable for the duration of the training run. Which approach BEST meets these requirements?

A.Copy the dataset once into an Amazon FSx for Lustre file system linked to the S3 bucket, and point the training job's channel at the FSx mount.
B.Use SageMaker FastFile mode, which streams objects from S3 on demand with a local cache and exposes a POSIX-like interface.
C.Use SageMaker File mode, which copies the dataset to the attached Amazon EBS volume before training begins.
D.Use SageMaker Pipe mode with the dataset converted to protobuf RecordIO and streamed directly to the training algorithm.
AnswerA

FSx for Lustre provides a high-throughput, low-latency POSIX file system that can lazily load data from the linked S3 bucket on first access. For a large immutable dataset read repeatedly across epochs, this eliminates repeated S3 GET overhead, gives shared high-bandwidth access to all training instances, and is the recommended pattern for accelerating SageMaker training on large datasets.

Why this answer

FSx for Lustre linked to S3 offers a shared, high-throughput file system that materializes data on first access and serves subsequent epochs at much lower latency than repeated S3 reads. Because the dataset is immutable during training, the lazy-load-then-cache behavior is safe and efficient, and all eight instances read from the same fast mount rather than each pulling from S3.

Exam trap

The trap here is assuming Pipe mode or FastFile mode always beats a file system for very large datasets, when in fact a linked FSx for Lustre mount removes repeated object-store latency for multi-epoch immutable data.

14
Multi-Selectmedium

A machine learning engineer is preparing a dataset for a binary classification model. The dataset has 10,000 rows and 200 features, with 5% positive class. The engineer suspects class imbalance may affect model performance. Which TWO actions should the engineer take to mitigate imbalance? (Choose 2.)

Select 2 answers
A.Perform PCA to reduce dimensions
B.Remove features with low variance
C.Use k-fold cross-validation
D.Apply SMOTE only to training data
E.Use class weights in the algorithm
AnswersD, E

SMOTE synthesises new minority-class observations by interpolating between nearest minority neighbours. Restricting it to the training split keeps those synthetic points out of validation and test data, preventing the leakage and inflated metrics that applying it before splitting would produce.

Why this answer

Option D is correct because SMOTE (Synthetic Minority Over-sampling Technique) generates synthetic samples of the minority class and must be applied only to the training data to avoid data leakage into validation/test folds, directly addressing the 5% positive-class imbalance. Option E is correct because setting class weights in the algorithm (e.g., class_weight='balanced' in scikit-learn) penalizes misclassification of the minority class more heavily, which mitigates imbalance without altering the dataset. Option A is incorrect because PCA is a dimensionality-reduction technique for the 200 features and does nothing to change the 5% class ratio.

Option B is incorrect because removing low-variance features is a feature-selection step unrelated to class distribution. Option C is incorrect because k-fold cross-validation is an evaluation/resampling strategy that provides more reliable performance estimates but does not itself rebalance classes.

Exam trap

The trap here is that candidates may confuse techniques for handling class imbalance with general data preprocessing or evaluation methods, leading them to select PCA or cross-validation as solutions, when in fact only resampling (SMOTE) and cost-sensitive learning (class weights) directly address the imbalance problem.

15
MCQeasy

A data scientist has a 40 GB CSV dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with pandas, and the scientist wants to reduce both storage cost and training-time I/O without changing the logical schema. Which data preparation action should be taken?

A.Move the dataset to Amazon EFS and mount it to the training container.
B.Convert the dataset to JSON Lines and enable S3 Transfer Acceleration.
C.Split the CSV into many smaller CSV files and keep the same column layout.
D.Convert the dataset to Parquet with Snappy compression and keep the same columns.
AnswerD

Parquet is a columnar format, so the training script reads only the columns it needs, and Snappy compression shrinks the on-disk footprint substantially compared with text CSV. The logical schema is preserved because column names and types remain the same. This directly reduces S3 storage cost and the bytes transferred during training, satisfying both requirements without changing the data model.

Why this answer

Parquet stores data by column and compresses each column efficiently, so a training script that needs a subset of columns reads far fewer bytes than it would from CSV. Snappy compression further reduces the S3 object size, lowering storage cost. Because the column names and types stay the same, the logical schema is unchanged and the pandas-based script can read the Parquet files with minimal modification.

Exam trap

The trap here is assuming that splitting or relocating files reduces cost, when only changing the storage format and compression actually shrinks the bytes stored and read.

16
MCQhard

A social media company is processing a real-time stream of user activity data from Amazon Kinesis Data Streams to train a machine learning model for content recommendation. The raw data includes user ID, timestamp, content ID, interaction type (like, share, comment), and device type. The data scientists need to aggregate features per user over a sliding window of 7 days, including counts of interaction types, unique content IDs engaged, and a moving average of interaction timestamps. The aggregated data will be used to update a user embedding model. The streaming data volume is approximately 500 records per second, and the company uses an AWS Glue streaming ETL job for transformation. However, the Glue job is failing frequently with high latency and checkpoint errors. The team needs a more robust solution to prepare the streaming data features. Which approach should the team take?

A.Increase the DPU count on the Glue streaming ETL job and reduce the checkpoint interval to improve performance.
B.Use Amazon Kinesis Data Analytics for Apache Flink to perform the sliding window aggregations with built-in state management and exactly-once processing, then write the features to S3 and DynamoDB.
C.Use AWS Lambda functions to process records from Kinesis, store intermediate aggregation results in Amazon DynamoDB, and read them back to compute windowed features.
D.Use Amazon SageMaker Processing jobs that run periodically every hour to read data from S3 (landing from Kinesis Firehose) and perform the aggregations batch-wise.
AnswerB

Apache Flink on Kinesis Data Analytics provides native event-time sliding windows with managed keyed state and exactly-once checkpointing, eliminating the Glue job's checkpoint failures and latency at 500 records per second. It satisfies the 7-day per-user aggregation constraint, then sinks features to S3 and DynamoDB.

Why this answer

Amazon Kinesis Data Analytics for Apache Flink provides native support for sliding window aggregations with managed state and exactly-once processing semantics, which directly addresses the high latency and checkpoint errors seen in the Glue streaming ETL job. Flink's checkpointing mechanism ensures fault-tolerant state management for the 7-day sliding window, while Glue's Spark Streaming engine struggles with long-running stateful operations at 500 records/sec due to its micro-batch architecture and checkpoint overhead.

Exam trap

The trap here is that candidates assume increasing resources (DPU) on Glue streaming ETL will fix performance issues, but the root cause is Spark's micro-batch architecture's inability to efficiently manage long-running stateful sliding windows, which Flink's native streaming engine is designed for.

How to eliminate wrong answers

Option A is wrong because increasing DPU count and reducing checkpoint interval on a Glue streaming ETL job exacerbates checkpoint errors and latency due to Spark's micro-batch overhead and lack of native long-lived state management for sliding windows. Option C is wrong because AWS Lambda functions have a maximum execution timeout of 15 minutes and no built-in state management, making them unsuitable for maintaining 7-day sliding window aggregations across 500 records/sec without external state stores that introduce eventual consistency and latency. Option D is wrong because using hourly SageMaker Processing jobs on S3 data from Kinesis Firehose introduces a minimum 1-hour delay, which violates the real-time requirement for updating a user embedding model with sliding window features.

17
Multi-Selectmedium

A data engineer is preparing a dataset in Amazon SageMaker Data Wrangler for a binary classification model. The dataset contains missing values in several numeric columns, and the engineer wants a reusable, reproducible transformation that can be applied identically to the training data and to future inference data. The engineer plans to export the transformation and integrate it into a SageMaker Pipeline. Which TWO actions should the engineer take to ensure the imputation is consistent between training and inference? (Choose two.)

Select 2 answers
A.Configure the imputation transform in the Data Wrangler flow so it learns the fill statistic (such as the mean) from the training data during the flow run.
B.Replace missing numeric values with zero in both the training and inference datasets using a custom Python script.
C.Export the Data Wrangler flow to a SageMaker Pipeline and include the generated processing step so it runs on both training and inference inputs.
D.Compute the mean of each numeric column on the full dataset including inference data, then hard-code those values into the pipeline.
E.Delete all rows that contain missing values from the training dataset before training.
AnswersA, C

Learning the fill statistic from the training data ensures the imputation parameters are derived from the correct distribution and not from inference data, which would leak information. Data Wrangler captures these learned parameters as part of the flow, making the transformation reproducible. This is the foundation for applying the same imputation consistently at inference time.

Why this answer

Consistency between training and inference requires that the imputation statistic be learned from training data and that the same transformation be reused at inference. Configuring Data Wrangler to learn the fill statistic from training data captures the parameters, and exporting the flow into a SageMaker Pipeline processing step ensures those exact parameters and logic are applied to both training and inference inputs.

Exam trap

The trap here is treating imputation as a one-time data cleanup rather than a parameterized transformation whose learned statistics must be persisted and reapplied at inference.

18
MCQeasy

A data scientist has a 200 GB Parquet dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with the SageMaker training toolkit's File mode, and the job currently spends a long time downloading before training begins. The team wants to reduce startup time without changing the training algorithm. Which change should the data scientist make?

A.Switch the input channel to FastFile mode so the training container streams data instead of downloading the full dataset first.
B.Enable Pipe mode and reformat the Parquet files into the recordIO-protobuf format.
C.Add more instances to the training cluster and enable distributed data parallel.
D.Increase the volume size of the training instance so the full dataset fits in local storage.
AnswerA

SageMaker FastFile mode exposes S3 objects to the training container through a POSIX-like interface that streams bytes on demand, so the job can begin reading immediately without a full upfront download. For large Parquet datasets this substantially reduces the time before training starts while leaving the algorithm unchanged.

Why this answer

FastFile mode removes the blocking full download by presenting S3 objects as a streamed filesystem inside the container, so the training process starts as soon as it reads the first bytes. This directly shortens the pre-training wait for a large Parquet dataset and requires no change to the training algorithm or data format.

Exam trap

The trap here is conflating capacity with latency, adding disk space when the real delay is waiting for a complete copy of the data.

19
MCQhard

A machine learning engineer runs a SageMaker Processing job that must load a 200 GB dataset from S3, compute statistics, and write a small summary to S3. The job repeatedly fails with an out-of-disk-space error on the processing instance. Which change is MOST likely to resolve the failure?

A.Add a lifecycle configuration script that deletes files from the input directory after each epoch.
B.Raise the max_runtime_in_seconds parameter and enable network isolation on the processing job.
C.Increase the processing instance's volume size and use ShardedByS3Key so each instance downloads only its share of objects.
D.Switch the input channel to Pipe mode so the container streams records instead of downloading files.
AnswerC

The failure is caused by the local volume filling with downloaded input data. Enlarging the attached EBS volume provides headroom, and ShardedByS3Key distributes the S3 objects across instances so no single instance downloads the entire dataset. Together these directly address the root cause rather than merely retrying the job.

Why this answer

Processing jobs download input channels to local storage, so a 200 GB dataset can exhaust the default volume. Enlarging the volume gives the container room, and ShardedByS3Key spreads objects across multiple instances so each holds only a fraction. Pipe mode is not applicable to Processing jobs, and the other options adjust runtime or cleanup behavior without changing disk consumption.

Exam trap

The trap here is assuming Pipe mode applies to every SageMaker job type, when streaming input is a training-job feature and Processing jobs always mount files locally.

20
MCQmedium

A data engineer is building a data pipeline for a machine learning model that requires both structured and unstructured data. The structured data (customer demographics) is in Amazon RDS, and the unstructured data (customer support chat logs) is in Amazon S3 as JSON files. The engineer needs to combine these datasets into a single training dataset stored in S3 in Parquet format. They must also perform feature engineering such as text vectorization on the chat logs. The pipeline should be serverless and cost-effective. Which approach should they use?

A.Use a SageMaker Processing job with a custom Python script that reads from both sources and writes to S3.
B.Use Amazon Athena to join the data from RDS and S3, then export the results as Parquet.
C.Use AWS Glue ETL with a Spark script that reads from RDS (via JDBC) and S3, performs transformations, and writes Parquet.
D.Use Amazon Kinesis Data Analytics to read from RDS and S3 and produce a continuous stream of processed data.
AnswerC

AWS Glue ETL provides a serverless Spark environment that reads RDS through JDBC and S3 JSON concurrently, then applies ML Transform for text vectorisation. It writes Parquet directly to S3, satisfying the serverless, cost-effective constraint without managing clusters or provisioning infrastructure.

Why this answer

AWS Glue ETL with a Spark script is the correct choice because it natively supports reading from both Amazon RDS (via JDBC) and Amazon S3 (JSON), performing complex transformations like text vectorization, and writing the output as Parquet. Glue is serverless, cost-effective (pay per DPU-hour), and fully managed, making it ideal for batch ETL pipelines that combine structured and unstructured data for ML training.

Exam trap

The trap here is that candidates often choose SageMaker Processing (Option A) because it is associated with ML, but they overlook that Glue ETL is the designated AWS service for serverless data preparation and transformation, especially when combining disparate data sources like RDS and S3.

How to eliminate wrong answers

Option A is wrong because SageMaker Processing jobs are designed for ML-specific tasks like training or inference, not general-purpose ETL; they lack native JDBC connectors for RDS and require custom networking setup, increasing complexity and cost. Option B is wrong because Amazon Athena cannot perform feature engineering like text vectorization; it is an interactive query service for SQL-on-data, not a transformation engine, and cannot write Parquet with custom logic. Option D is wrong because Kinesis Data Analytics is for real-time stream processing, not batch ETL; it would introduce unnecessary latency and cost for a one-time or scheduled training dataset generation, and it cannot directly write Parquet to S3 without additional sinks.

21
MCQmedium

A data scientist needs to split a dataset into training, validation, and test sets. The dataset has a categorical target variable with imbalanced class distribution. Which splitting technique ensures that each subset has a similar proportion of each class?

A.K-fold cross-validation split
B.Chronological split
C.Stratified split
D.Random split
AnswerC

Stratified splitting preserves the target class proportions within each training, validation, and test subset by sampling per class rather than randomly across the whole dataset. This directly satisfies the stem's imbalanced categorical target constraint, preventing minority classes from being underrepresented or absent in any split.

Why this answer

Stratified splitting preserves the original class proportions in each subset (training, validation, test) by sampling each class independently. This is critical for imbalanced datasets to avoid skewed distributions that could bias model evaluation or training.

Exam trap

AWS often tests the distinction between data splitting techniques and model evaluation methods, so the trap here is that candidates confuse k-fold cross-validation (a validation strategy) with a static split technique, leading them to select option A.

How to eliminate wrong answers

Option A is wrong because k-fold cross-validation is a resampling technique for model evaluation, not a method for creating a single static split into training, validation, and test sets. Option B is wrong because chronological split orders data by time, which is irrelevant for a categorical target with imbalanced classes and does not guarantee proportional class representation. Option D is wrong because random split does not account for class distribution; with imbalanced data, it can produce subsets with significantly different class proportions, especially for rare classes.

22
MCQmedium

A data engineer is preparing a dataset for a SageMaker training job. The dataset contains a timestamp column and is stored in Amazon S3 as CSV files. The engineer needs to ensure that the training job reads the data efficiently and that the data is partitioned by date to improve query performance in Amazon Athena. Which action should the engineer take?

A.Convert the CSV files to Parquet format and store them in a single prefix without partitioning.
B.Convert the CSV files to Parquet format and partition them by date in S3 using a Hive-style partition scheme (e.g., year=2023/month=01/day=01/).
C.Keep the CSV format but partition the files by date using a flat prefix structure (e.g., 2023-01-01/).
D.Use AWS Glue to crawl the CSV files and create a table in the AWS Glue Data Catalog, then use Athena to query the data directly without converting or partitioning.
AnswerB

Converting to Parquet improves compression and columnar read performance, which benefits SageMaker training jobs. Partitioning by date using Hive-style prefixes allows Athena to prune partitions and reduce query scan, improving performance. This meets both efficiency and query performance requirements.

Why this answer

Converting CSV to Parquet reduces storage size and enables columnar reads, which speeds up SageMaker training jobs. Partitioning by date with Hive-style prefixes allows Athena to prune partitions, reducing query cost and latency. Together, these actions satisfy both the training efficiency and Athena query performance requirements.

Exam trap

The trap here is thinking that simply cataloging CSV data or partitioning without converting format is sufficient for both SageMaker efficiency and Athena performance.

23
Multi-Selecteasy

A data engineer is using AWS Glue to prepare a dataset for ML. The engineer wants to split the dataset into training and testing sets while preserving the distribution of the target variable. Which TWO methods achieve this goal? (Select TWO)

Select 2 answers
A.Use Amazon Athena to create views with random sampling
B.Use the `train_test_split` function from scikit-learn in a SageMaker notebook
C.Use AWS Glue's built-in random split transform
D.Use a custom Spark script with stratified sampling
E.Use Amazon SageMaker's built-in SplitType parameter in a Processing Job
AnswersB, D

The stratify parameter maintains class proportions.

Why this answer

The `train_test_split` function from scikit-learn supports the `stratify` parameter, which preserves the distribution of the target variable when splitting a dataset into training and testing sets. This is a standard, reliable method for stratified splitting in Python-based ML workflows, and it can be used directly in a SageMaker notebook.

Exam trap

The trap here is that candidates often confuse random splitting (which is available in many tools like Glue and Athena) with stratified splitting, assuming that any 'random' operation preserves distribution, but only stratified methods explicitly maintain class proportions.

24
Multi-Selectmedium

A data scientist is using SageMaker Data Wrangler to prepare a dataset for a binary classification model. The dataset contains a mix of numerical and categorical features. The scientist wants to perform feature engineering to improve model performance. Which TWO actions are appropriate for handling categorical features in Data Wrangler? (Choose two.)

Select 2 answers
A.Apply one-hot encoding to categorical features with low cardinality.
B.Apply target encoding to categorical features with high cardinality.
C.Apply PCA to categorical features for dimensionality reduction.
D.Apply log transformation to categorical features.
E.Apply min-max scaling to categorical features.
AnswersA, B

One-hot encoding is a standard technique for converting categorical variables into a numerical format suitable for many machine learning algorithms. It creates binary columns for each category. For low-cardinality features (few unique values), this is efficient and avoids the curse of dimensionality. Data Wrangler provides a 'One-Hot Encoding' transform that can be applied directly.

Why this answer

Categorical features need to be encoded into numerical representations before training. One-hot encoding is effective for low-cardinality features, creating binary indicators. Target encoding is suitable for high-cardinality features, replacing categories with target statistics and reducing dimensionality.

Both are available as transforms in SageMaker Data Wrangler. Scaling, PCA, and log transformations are designed for numerical features and are not applicable to categorical data.

Exam trap

The trap here is assuming that numerical transforms like scaling or PCA can be applied to categorical features without encoding; they cannot and will cause errors.

25
MCQmedium

A team is building a recommendation system and wants to store and serve features for online and offline models. The features include user statistics (updated daily) and movie metadata (static). The team needs low-latency inference for real-time recommendations and wants to reuse features across multiple models. Which AWS service should the team use to store, manage, and serve these features?

A.Amazon DynamoDB with TTL.
B.AWS Glue Data Catalog.
C.SageMaker Feature Store.
D.Amazon S3 with AWS Lambda for serving.
AnswerC

SageMaker Feature Store provides an online store for low-latency real-time inference and an offline store for training, with feature groups reusable across models. It directly satisfies the daily-updated user statistics, static movie metadata and cross-model reuse requirements.

Why this answer

Amazon SageMaker Feature Store is purpose-built for storing, managing, and serving ML features with low-latency retrieval for online inference and batch serving for offline training. It supports feature reuse across multiple models by providing a centralized feature registry, consistent feature definitions, and both online (low-latency) and offline (S3-based) stores, which directly matches the team's requirements for real-time recommendations and cross-model reuse.

Exam trap

The trap here is that candidates often confuse a general-purpose database (DynamoDB) or a data catalog (Glue) with a purpose-built ML feature store, overlooking the need for feature-specific capabilities like online/offline consistency, feature versioning, and reuse across models.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB with TTL is a key-value and document database that can store features but lacks built-in feature management capabilities such as feature versioning, point-in-time consistency across online/offline stores, and a feature registry; TTL only handles data expiration, not the orchestration needed for ML feature reuse. Option B is wrong because AWS Glue Data Catalog is a metadata repository for data assets (tables, schemas) and does not provide a low-latency online serving endpoint or feature-specific storage; it is used for data discovery and ETL, not for serving features in real-time inference. Option D is wrong because Amazon S3 with AWS Lambda for serving introduces high latency due to Lambda cold starts and S3 GET request overhead, making it unsuitable for low-latency real-time recommendations; additionally, it lacks feature store capabilities like consistent feature definitions, offline/online synchronization, and feature reuse across models.

26
MCQhard

A company uses Amazon SageMaker Data Wrangler to prepare data for ML. The dataset contains a timestamp column and sensor readings from IoT devices. The data scientist needs to create features such as moving averages and rolling statistics over time windows. Which Data Wrangler transformation type should be selected?

A.Join
B.Custom Python script
C.Group by and aggregate
D.Window function
AnswerD

Window functions compute rolling aggregates across ordered partitions, so moving averages and rolling statistics over the timestamp column are produced directly within Data Wrangler. This satisfies the requirement for time-windowed feature engineering, unlike row-level transforms such as numeric or categorical encodings, which cannot aggregate across neighbouring rows.

Why this answer

Window functions in Amazon SageMaker Data Wrangler allow you to compute moving averages, rolling statistics, and other time-window-based aggregations over ordered partitions of data. This is the correct transformation type because it directly supports operations like `SUM() OVER (ORDER BY timestamp ROWS BETWEEN 2 PRECEDING AND CURRENT ROW)` without requiring custom code or losing row-level granularity.

Exam trap

The trap here is that candidates confuse 'Group by and aggregate' with 'Window function' because both involve aggregation, but Group by reduces rows while Window functions preserve row-level detail, which is essential for rolling statistics.

How to eliminate wrong answers

Option A is wrong because Join is used to combine datasets based on a common key, not to compute rolling statistics over a time window. Option B is wrong because while a Custom Python script could technically implement moving averages, Data Wrangler provides a native Window function transformation that is more efficient, easier to maintain, and avoids the overhead of writing and debugging custom code. Option C is wrong because Group by and aggregate collapses rows into summary statistics per group, which loses the individual row-level detail needed for rolling window calculations.

27
MCQeasy

A data scientist is working with a dataset that contains missing values in several numeric features. The data scientist wants to impute the missing values with the median of each feature. Which Amazon SageMaker Data Wrangler transformation should be used?

A.Replace missing with constant
B.Custom transform with Python
C.Drop missing rows
D.Handle missing values (with median strategy)
AnswerD

The Handle missing values transformation with the median strategy computes each numeric feature's median and substitutes it for nulls, satisfying the requirement to impute with the median rather than mean or mode. It operates directly on the dataset within Data Wrangler's flow.

Why this answer

Amazon SageMaker Data Wrangler includes a built-in 'Handle missing values' transformation that supports imputation with the median strategy. This directly matches the requirement to replace missing numeric values with the median of each feature without writing custom code.

Exam trap

The trap here is that candidates may confuse the 'Replace missing with constant' option (which uses a fixed value) with the median strategy, or they may overcomplicate the solution by choosing a custom Python transform when a built-in option exists.

How to eliminate wrong answers

Option A is wrong because 'Replace missing with constant' imputes a user-specified constant value (e.g., 0 or a fixed number), not the median of the feature. Option B is wrong because 'Custom transform with Python' would require writing custom Python code to compute and apply the median, which is unnecessary when a built-in transformation exists. Option C is wrong because 'Drop missing rows' removes entire rows with missing values, discarding potentially valuable data instead of imputing the missing values.

28
MCQmedium

A team is using Amazon SageMaker Processing for data preprocessing. They have a Parquet dataset in Amazon S3. Which configuration will provide the most efficient reading of the dataset during processing?

A.Read the Parquet files as text using SparkContext.textFile
B.Split the dataset into many small Parquet files (e.g., 1 MB each)
C.Convert the Parquet files to CSV before processing
D.Read the Parquet files directly using SparkSession.read.parquet
AnswerD

SparkSession.read.parquet reads Parquet's columnar, compressed format directly, enabling predicate pushdown and column pruning so SageMaker Processing scans only needed columns and row groups from Amazon S3, avoiding full-dataset deserialisation and delivering the most efficient read.

Why this answer

SageMaker Processing natively integrates with Apache Spark, and reading Parquet files directly via `SparkSession.read.parquet` leverages columnar storage, predicate pushdown, and compression (e.g., Snappy) to minimize I/O and deserialization overhead. This approach is far more efficient than text-based or format-conversion methods, as Parquet is optimized for analytical workloads and preserves schema information.

Exam trap

AWS often tests the misconception that many small files improve parallelism, but in distributed systems like Spark on SageMaker, small files increase S3 API call overhead and scheduler latency, making larger Parquet files (e.g., 128 MB–1 GB) far more efficient for reading.

How to eliminate wrong answers

Option A is wrong because `SparkContext.textFile` reads data as plain text lines, which is incompatible with binary Parquet format and would result in corrupted data or require manual parsing, losing all columnar optimization. Option B is wrong because splitting the dataset into many small 1 MB Parquet files increases S3 LIST and GET request overhead, causing task scheduling delays and poor I/O throughput due to excessive file metadata operations. Option C is wrong because converting Parquet to CSV before processing introduces unnecessary serialization/deserialization costs, increases data size (CSV lacks compression and columnar storage), and discards schema and type information, leading to slower read performance.

29
MCQeasy

A data scientist is preparing a CSV dataset in Amazon S3 for a SageMaker training job. Several rows contain missing values in numeric feature columns, and the chosen algorithm cannot handle NaNs. The scientist wants a repeatable, code-based transformation that runs inside a SageMaker Processing job before training. Which step is the MOST appropriate?

A.Use an S3 Lifecycle rule to expire objects containing missing values so only clean data remains in the bucket.
B.Open the CSV in a SageMaker notebook, manually edit the missing cells, and save the file back to S3.
C.Configure the SageMaker training job's input channel with a content type that tells the algorithm to ignore NaN values automatically.
D.Write a preprocessing script that uses pandas to impute or drop missing values and run it in a SageMaker Processing job with a scikit-learn container.
AnswerD

SageMaker Processing jobs run a containerized script against data in S3 and write outputs back to S3, which makes the transformation repeatable, versionable, and independent of notebook state. Using pandas for imputation or row removal inside the scikit-learn container is a standard, well-supported approach that produces a clean dataset the training job can consume directly.

Why this answer

Running a pandas-based cleaning script in a SageMaker Processing job makes the imputation or row removal reproducible and code-driven, and the job reads from and writes to S3 so the cleaned dataset is available to training. This is the canonical pattern for repeatable preprocessing that must run before a training job and be re-executed as data changes.

Exam trap

The trap here is treating missing-value handling as an algorithm setting or a storage-lifecycle concern, when it is really a preprocessing step that belongs in a scripted Processing job.

30
MCQmedium

A machine learning engineer is building a pipeline to preprocess text data for a sentiment analysis model. The data consists of customer reviews. The engineer wants to convert the text into numerical features while preserving the semantic meaning of words. Which technique should be used?

A.One-hot encoding of each word
B.Bag-of-words with TF-IDF
C.Hashing vectorizer
D.Word embeddings (e.g., Word2Vec or GloVe)
AnswerD

Word embeddings map tokens to dense vectors whose geometric relationships encode semantic similarity, so reviews with comparable meaning produce comparable features. This preserves semantic meaning, unlike bag-of-words or TF-IDF, which treat terms as independent and lose contextual relationships.

Why this answer

Word embeddings (like Word2Vec or GloVe) are dense vector representations that capture semantic relationships between words based on their context in a large corpus. For sentiment analysis, preserving semantic meaning (e.g., 'good' and 'excellent' having similar vectors) is critical, and embeddings directly encode this, unlike sparse or count-based methods.

Exam trap

The trap here is that candidates often choose TF-IDF (Option B) because it is a common text preprocessing technique, but they overlook the explicit requirement to 'preserve semantic meaning,' which only dense embeddings can achieve.

How to eliminate wrong answers

Option A is wrong because one-hot encoding treats each word as an independent binary feature with no semantic similarity—vectors for 'good' and 'excellent' are orthogonal, losing all contextual meaning. Option B is wrong because bag-of-words with TF-IDF produces sparse, high-dimensional vectors based on word frequency and inverse document frequency, which ignore word order and context, failing to capture semantic relationships. Option C is wrong because a hashing vectorizer uses a hash function to map words to fixed-size indices, which can cause collisions and still produces sparse, frequency-based features without any semantic understanding.

31
MCQmedium

A data engineer is preparing a dataset for a time series forecasting model. The dataset contains a timestamp column and a target variable. The engineer wants to create additional features such as lag values and rolling averages. Which SageMaker Data Wrangler transform should be used to generate these time series features?

A.Use the 'Handle Outliers' transform to create lag features.
B.Use the 'Balance Data' transform to generate rolling averages.
C.Use the 'Time Series' transform to create lag and rolling window features.
D.Use the 'Featurize Text' transform to extract date parts.
AnswerC

SageMaker Data Wrangler includes a 'Time Series' transform that can generate lag features, rolling statistics (like mean, sum), and other time-based features from a timestamp column. This transform is specifically designed for time series data preparation, allowing the engineer to create the required lag values and rolling averages efficiently.

Why this answer

Time series feature engineering often involves creating lag features (past values) and rolling statistics (e.g., moving averages) to capture temporal patterns. In SageMaker Data Wrangler, the 'Time Series' transform provides a dedicated set of operations for these tasks, including lag, rolling window, and date part extraction. This transform simplifies the process and ensures correct handling of time order.

Other transforms like text featurization or outlier handling serve different purposes and cannot produce these features.

Exam trap

The trap here is assuming any transform that manipulates numerical data can create lag features; only the Time Series transform is designed for temporal feature engineering.

32
MCQmedium

A data engineer needs to prepare a large dataset (10 TB) stored in Amazon S3 for a training job on SageMaker. The data is in CSV format, but the training algorithm expects Parquet for performance. The engineer must transform the data with minimal cost and without writing custom code. Which service should be used?

A.Use AWS Glue to create a crawler and ETL job that converts CSV to Parquet.
B.Use SageMaker Processing with a TensorFlow script to read CSV and write Parquet.
C.Use Amazon S3 Select to convert the data to Parquet during retrieval.
D.Use Amazon EMR with a Spark job to convert the files.
AnswerA

AWS Glue's crawler infers the CSV schema and its serverless Spark ETL writes Parquet directly to S3, satisfying the no-custom-code and minimal-cost constraints for the 10 TB conversion. Unlike Lambda, which caps runtime and memory, Glue scales to terabyte workloads without cluster management.

Why this answer

AWS Glue is the correct choice because it provides a serverless, pay-per-use ETL service that can automatically convert CSV to Parquet without writing custom code. The Glue crawler infers the schema, and the ETL job uses built-in transforms to efficiently handle 10 TB of data with minimal cost, as it only charges for the resources consumed during the job execution.

Exam trap

The trap here is that candidates often confuse Amazon S3 Select's ability to filter data with the ability to transform data formats, but S3 Select only returns filtered results in the original format and cannot perform format conversion like CSV to Parquet.

How to eliminate wrong answers

Option B is wrong because SageMaker Processing with a TensorFlow script requires writing custom code, which violates the 'without writing custom code' requirement. Option C is wrong because Amazon S3 Select only supports filtering data using SQL queries on CSV or JSON objects; it cannot convert data to Parquet format. Option D is wrong because Amazon EMR with a Spark job requires provisioning and managing a cluster, incurring higher costs and operational overhead compared to the serverless Glue approach.

33
MCQhard

A data scientist is training a binary classifier on a highly imbalanced dataset (1:100 class ratio). The dataset contains 500,000 rows and 30 features. The data is stored in S3 in Parquet format. The data scientist wants to use SageMaker's built-in XGBoost algorithm. Which data preparation technique should the data scientist apply to best address the class imbalance without causing data leakage?

A.Undersample the majority class to create a balanced dataset, then split.
B.Use the scale_pos_weight parameter in XGBoost to assign higher weight to the minority class.
C.Oversample the minority class using SMOTE on the entire dataset before splitting into train/validation sets.
D.Randomly oversample the minority class by duplicating rows, then perform stratified train/test split.
AnswerB

scale_pos_weight multiplies the minority class's gradient contribution during XGBoost training, countering the 1:100 skew. Because it adjusts only the loss weighting rather than duplicating or synthesising rows, no information crosses between train and validation splits, avoiding the leakage that resampling before splitting would cause.

Why this answer

The scale_pos_weight parameter in XGBoost directly adjusts the loss function to penalize misclassifications of the minority class more heavily, effectively handling class imbalance without modifying the dataset. This avoids data leakage because the weighting is applied during training only, not during preprocessing, and does not involve any synthetic data generation or resampling that could inadvertently expose test information.

Exam trap

AWS often tests the misconception that resampling techniques (like SMOTE or random oversampling) are always safe, when in fact applying them before splitting introduces data leakage, whereas built-in parameters like scale_pos_weight avoid this pitfall.

How to eliminate wrong answers

Option A is wrong because undersampling the majority class reduces the dataset size significantly (from 500,000 rows to ~10,000 rows), discarding valuable information and potentially degrading model performance, and it does not inherently prevent data leakage if done before splitting. Option C is wrong because applying SMOTE on the entire dataset before splitting causes data leakage: synthetic samples generated from the full dataset can incorporate information from the test set, leading to overly optimistic validation metrics. Option D is wrong because randomly oversampling the minority class by duplicating rows before splitting can cause data leakage if duplicates of the same row appear in both training and validation sets, and it does not introduce new variance, leading to overfitting.

34
MCQeasy

An organization stores raw data in Amazon S3 as CSV files. They need to perform serverless data transformation and convert the data to Parquet format for efficient ML training. Which AWS service is most appropriate?

A.AWS Glue
B.Amazon EMR
C.Amazon Athena
D.Amazon Redshift
AnswerA

AWS Glue provides serverless Spark-based ETL that reads CSV from Amazon S3 and writes Parquet, satisfying the stem's transformation and format-conversion requirement. Its crawler and job model needs no cluster management, and Parquet's columnar layout accelerates downstream ML training.

Why this answer

AWS Glue is the most appropriate service because it is a fully managed, serverless ETL service designed specifically for data transformation tasks like converting CSV to Parquet. It automatically handles schema inference, data partitioning, and optimization for ML training workloads without requiring infrastructure management.

Exam trap

The trap here is that candidates often confuse Amazon Athena's ability to query Parquet data with the ability to transform data into Parquet, but Athena is a query engine, not an ETL transformation service.

How to eliminate wrong answers

Option B (Amazon EMR) is wrong because it requires provisioning and managing clusters, which contradicts the 'serverless' requirement; it is better suited for large-scale big data processing with frameworks like Spark or Hadoop, not simple serverless transformations. Option C (Amazon Athena) is wrong because it is an interactive query service for analyzing data directly in S3 using SQL, not a transformation engine; it cannot convert file formats like CSV to Parquet. Option D (Amazon Redshift) is wrong because it is a data warehouse for analytics and SQL-based querying, not a serverless transformation service; it requires loading data into a cluster and does not natively convert CSV to Parquet in S3.

35
MCQmedium

A company runs an online retail business and wants to build a product recommendation system. They have a dataset of customer purchases stored in Amazon S3 as CSV files. The dataset includes columns: 'customer_id', 'product_id', 'purchase_date', 'quantity', 'price', and 'category'. The data science team plans to use Amazon SageMaker to train a factorization machines model. During data exploration, they discover that the 'category' column has 1,200 unique values, and many categories appear only a few times. The 'product_id' column has 50,000 unique values. They want to include both features in the model. The team is concerned about the high cardinality of these features. Which approach should they take to prepare these features for the factorization machines model?

A.Apply one-hot encoding to both 'product_id' and 'category' columns.
B.Drop the 'category' column and only use 'product_id' since it has more granularity.
C.Encode both columns as integer indices and feed them directly to the factorization machines algorithm as categorical features.
D.Apply principal component analysis (PCA) to reduce the dimensionality of the categorical features.
AnswerC

Factorization machines handle high-cardinality categoricals by learning latent factor vectors per index, so integer-encoding category and product_id directly satisfies the stem's concern. One-hot encoding would explode dimensionality; index encoding lets the algorithm capture interactions between sparse features efficiently.

Why this answer

Amazon SageMaker's factorization machines algorithm natively supports categorical features encoded as integer indices (0-based). This avoids the explosion of features from one-hot encoding (which would create 51,200 columns) and leverages the algorithm's ability to learn interactions between high-cardinality features via factorized parameters, making it both memory-efficient and effective for sparse data.

Exam trap

The trap here is that candidates default to one-hot encoding (Option A) as the standard categorical encoding technique, not realizing that factorization machines are specifically designed to avoid that explosion by accepting raw integer indices as categorical features.

How to eliminate wrong answers

Option A is wrong because one-hot encoding 1,200 categories and 50,000 products would create 51,200 binary columns, causing extreme sparsity and memory blowup, which undermines the factorization machine's efficiency and can lead to poor generalization. Option B is wrong because dropping the 'category' column discards valuable hierarchical information (e.g., product type) that could improve recommendation quality; factorization machines are designed to handle high-cardinality features, so there is no need to drop it. Option D is wrong because PCA is a linear dimensionality reduction technique for continuous features, not suitable for categorical data; applying PCA to one-hot encoded categories would destroy the interpretability of interactions and is not a standard preprocessing step for factorization machines.

36
Multi-Selecteasy

A company ingests daily log data into an S3 bucket. They need to update the existing ML training dataset with new data without reprocessing the entire history. Which two strategies should they adopt? (Choose two.)

Select 2 answers
A.Store all data in a single large file and use append operations
B.Use AWS Glue to incrementally process new partitions
C.Use a partition key such as date to add new partitions
D.Manually copy new files to the same S3 bucket
E.Overwrite the entire existing dataset with the new data
AnswersB, C

AWS Glue incrementally processes only newly arrived S3 partitions, satisfying the requirement to update the training dataset without reprocessing the entire history. By tracking partition metadata in the Data Catalog, Glue reads just the fresh daily logs, cutting compute cost and runtime compared with full-dataset reprocessing.

Why this answer

Option B is correct because AWS Glue can perform incremental processing by reading only newly added partitions (using job bookmarks and partition predicates) rather than reprocessing the full history, which directly satisfies the requirement to update the dataset without re-running the entire pipeline. Option C is correct because partitioning the S3 data by a key such as date lets new daily log data land in a new partition, so downstream training jobs or Glue ETL can scan only the new partition and append it to the existing dataset. Option A is incorrect because S3 objects are immutable and cannot be appended to in place; a single large file would require rewriting the whole object.

Option D is incorrect because manually copying files into the same bucket does not provide incremental processing logic or partition awareness, and is error-prone. Option E is incorrect because overwriting the entire dataset forces full reprocessing, which is exactly what the scenario wants to avoid.

Exam trap

AWS often tests the misconception that S3 supports append operations or that simply copying new files to the same bucket constitutes an incremental update strategy, when in reality S3 objects are immutable and a proper processing framework like AWS Glue with job bookmarks is required.

37
MCQeasy

Refer to the exhibit. A data scientist is trying to use AWS Glue to read data from the S3 bucket `ml-data-bucket`. The Glue job fails with an access denied error. What is the most likely cause?

A.The policy allows s3:PutObject but the job only reads
B.The policy does not specify the bucket ARN without /*
C.The Glue job role does not have the required permissions
D.The policy does not include s3:ListBucket permission on the bucket
AnswerD

Without s3:ListBucket, Glue cannot enumerate the bucket's objects, so its read attempt fails with AccessDenied even when s3:GetObject is granted. The stem's constraint is a Glue job reading from `ml-data-bucket`; Glue's crawler and job bootstrap require ListBucket to resolve prefixes before fetching any object.

Why this answer

The error occurs because the IAM policy attached to the Glue job role grants s3:GetObject on the bucket objects (via the `arn:aws:s3:::ml-data-bucket/*` resource) but does not include the s3:ListBucket permission on the bucket itself (`arn:aws:s3:::ml-data-bucket`). When AWS Glue reads data from S3, it first performs a ListBucket operation to enumerate objects in the bucket or prefix, and without that permission, the request is denied even if GetObject is allowed.

Exam trap

AWS often tests the subtle distinction between bucket-level permissions (like s3:ListBucket) and object-level permissions (like s3:GetObject), where candidates assume that granting GetObject on objects is sufficient for reading data, forgetting that listing the bucket is a prerequisite for discovering those objects.

How to eliminate wrong answers

Option A is wrong because the error is an access denied on a read operation, not a write operation; s3:PutObject is irrelevant to reading data. Option B is wrong because the policy does specify the bucket ARN without `/*` for the s3:ListBucket permission (as required), but the issue is that the s3:ListBucket permission itself is missing entirely. Option C is wrong because the Glue job role does have some permissions (as shown in the exhibit), but the specific missing permission is s3:ListBucket, not a general lack of permissions.

38
MCQeasy

A company has a dataset of 2 billion records stored as text files in Amazon S3. The data is partitioned by year and month. The data science team wants to read only the last 6 months of data for model training using SageMaker. To minimize data scanned and reduce costs, which approach should the team use?

A.Use S3 Select to retrieve only the last 6 months of data by applying an SQL expression on each object.
B.Use AWS Glue to create a catalog table with partitions, then query with Athena to create a filtered dataset in S3.
C.Use SageMaker Processing with a script that lists all objects in the bucket and reads only those with the desired prefixes.
D.Use SageMaker Processing with Input Mode 'File' and specify the S3 prefix for the last 6 months.
AnswerB

Partition pruning is the mechanism: the Glue catalog exposes year and month partitions, so Athena reads only the six relevant partitions rather than scanning the full 2 billion records, directly minimising bytes scanned and cost.

Why this answer

AWS Glue can crawl the S3 data to create a catalog table with partitions by year and month. Athena can then query only the partitions corresponding to the last 6 months, scanning minimal data and writing the filtered results back to S3 for SageMaker training. This approach leverages partition pruning to reduce costs and avoids loading or processing the full 2 billion records.

Exam trap

AWS often tests the misconception that SageMaker's Input Mode 'File' or S3 Select can efficiently filter partitioned data, but the key trap is that partition pruning requires a catalog service (like Glue) and a query engine (like Athena) to avoid scanning all objects or listing the entire bucket.

How to eliminate wrong answers

Option A is wrong because S3 Select operates on a single object at a time and cannot filter across multiple objects or partitions; applying it to 2 billion records would require iterating over all objects, negating cost savings. Option C is wrong because listing all objects in the bucket and reading only those with desired prefixes still requires enumerating the entire bucket, which incurs significant API costs and does not minimize data scanned (the script must still list all objects). Option D is wrong because SageMaker Processing with Input Mode 'File' downloads the entire dataset to the training instance; specifying a prefix for the last 6 months would still download all files under that prefix, but the data is partitioned by year and month, so using the prefix alone does not guarantee partition pruning—the team would need to explicitly list only the relevant prefixes, which is inefficient compared to Glue+Athena.

39
Multi-Selecthard

A company is building a real-time inference pipeline for an ML model. The raw data arrives in JSON format via Amazon Kinesis Data Streams. Before invoking the SageMaker endpoint, the data must be preprocessed to match the training data format. Which THREE steps should be included in the preprocessing function? (Select THREE)

Select 3 answers
A.Ensure that missing values are handled consistently with the training phase
B.Convert the data to a CSV string for model input
C.Apply the same feature engineering transformations (e.g., scaling, encoding) that were used during training
D.Re-train the model periodically using new data
E.Parse the JSON payload
AnswersA, C, E

Handling missing values identically to training preserves the feature distribution the model learned, preventing inference-time skew. Because Kinesis delivers raw JSON that may contain nulls, this step satisfies the requirement that preprocessed data match the training format before the SageMaker endpoint is invoked.

Why this answer

The preprocessing function must first parse the JSON payload (E), because the raw records arrive from Kinesis Data Streams in JSON format and the individual feature fields cannot be accessed until the JSON is decoded into a usable structure. It must also apply the same feature engineering transformations used during training (C), such as scaling and encoding, since the SageMaker endpoint expects inputs in the exact distribution and representation the model learned; using different transformations would cause training-serving skew and degrade predictions. Missing values must be handled consistently with the training phase (A), because imputation or drop logic applied at inference must mirror training-time behavior to keep the feature semantics identical.

Option B is not required because the model's input format is not stated to be CSV, and forcing a CSV string could conflict with the actual serialization the endpoint expects. Option D is incorrect because retraining the model is a separate MLOps activity, not part of the per-record preprocessing function invoked before calling the endpoint.

Exam trap

The trap here is that candidates confuse the preprocessing function's scope with broader MLOps tasks like model retraining, or assume a specific serialization format like CSV is required when JSON is natively supported by SageMaker endpoints.

40
MCQmedium

A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset. The dataset contains a column with date strings in the format 'YYYY-MM-DD'. The data scientist wants to extract the year, month, and day as separate features. Which Data Wrangler transform should be used?

A.Encode categorical transform.
B.Scale values transform.
C.Parse date transform.
D.Handle missing transform.
AnswerC

The Parse date transform interprets the 'YYYY-MM-DD' string as a datetime type, from which Data Wrangler can derive year, month and day components. This satisfies the requirement to extract those three separate features without custom code.

Why this answer

The 'Parse date' transform in Amazon SageMaker Data Wrangler is specifically designed to convert date strings into structured datetime components. By applying this transform to the 'YYYY-MM-DD' column, the data scientist can automatically extract year, month, and day as separate features, enabling downstream feature engineering without manual string parsing.

Exam trap

The trap here is that candidates may confuse 'Parse date' with 'Encode categorical' because dates can be treated as categorical features, but the question specifically asks for extracting year, month, and day as separate features, which requires parsing the date string into its components, not encoding the entire date as a category.

How to eliminate wrong answers

Option A is wrong because 'Encode categorical' transform is used to convert categorical variables into numerical representations (e.g., one-hot encoding), not to parse date strings. Option B is wrong because 'Scale values' transform normalizes or standardizes numerical features (e.g., min-max scaling, z-score), which is irrelevant for extracting date components. Option D is wrong because 'Handle missing' transform addresses null or missing values through imputation or deletion, not date parsing.

41
MCQhard

A machine learning engineer is using Amazon SageMaker Data Wrangler to prepare a dataset for a regression model. The dataset contains a categorical feature with high cardinality (over 10,000 unique values). The engineer wants to encode this feature efficiently without creating thousands of binary columns, which would explode the feature space. Which encoding technique should the engineer use in Data Wrangler?

A.Ordinal encoding
B.Binary encoding
C.Target encoding
D.One-hot encoding
AnswerC

Target encoding replaces each category with the mean of the target variable for that category, producing a single numeric feature. This handles high cardinality without expanding the feature space. Data Wrangler supports target encoding, and it is effective for regression tasks when properly regularized to avoid overfitting.

Why this answer

Target encoding is designed for high-cardinality categorical features, replacing each category with a statistic of the target variable. This produces a single feature, avoiding the curse of dimensionality. In SageMaker Data Wrangler, target encoding is available as a transformation and can be configured with smoothing to reduce overfitting, making it the appropriate choice for this regression scenario.

Exam trap

The trap here is assuming that binary encoding is sufficient for high cardinality, but it still creates multiple columns and may not capture the target relationship as effectively as target encoding.

42
Multi-Selecthard

A machine learning engineer is preparing data for a SageMaker training job and needs to split a large dataset into training, validation, and test sets while avoiding data leakage from the same entity appearing in multiple splits. The dataset contains multiple rows per customer, and the target is customer churn. Which TWO strategies are appropriate? (Choose two.)

Select 2 answers
A.Split the data by a hash of the customer ID so that all rows for a given customer land in exactly one split.
B.Duplicate rows for customers in the minority class so that each split contains examples of every customer.
C.Sort the dataset by timestamp and take the first 80 percent of rows for training and the last 20 percent for validation.
D.Use GroupShuffleSplit from scikit-learn with the customer ID as the group parameter to partition customers across splits.
E.Use a random row-level split with a fixed seed, because a fixed seed makes the split reproducible and therefore leakage-free.
AnswersA, D

Hashing the customer ID and assigning splits by hash range guarantees that every row for a customer goes to the same split, which prevents the same entity from appearing in both training and validation. This is a standard grouped-split technique for entity-level leakage and is deterministic and reproducible across reruns when the same hash function and boundaries are used.

Why this answer

Entity-level leakage occurs when rows from the same customer appear in more than one split, letting the model memorize customer-specific patterns and inflating validation metrics. Hashing the customer ID into split buckets and using GroupShuffleSplit with the customer ID as the group both ensure that all rows for a customer stay together, which is the correct way to partition this churn dataset.

Exam trap

The trap here is confusing reproducibility with leakage prevention, since a fixed random seed produces repeatable splits but still scatters a customer's rows across training and validation.

43
MCQeasy

A machine learning engineer is preparing a dataset for training a model on Amazon SageMaker. The dataset contains numerical features with varying scales, and the engineer wants to ensure that all features contribute equally during training. Which data preparation step should the engineer take?

A.Apply one-hot encoding to all numerical features.
B.Apply feature scaling, such as standardization or normalization, to the numerical features.
C.Remove outliers from the dataset.
D.Apply principal component analysis (PCA) to reduce dimensionality.
AnswerB

Feature scaling transforms numerical features to a similar scale, preventing features with larger magnitudes from dominating the learning process. Standardization (z-score) or normalization (min-max) are common methods. This ensures all features contribute equally, which is the goal stated in the scenario.

Why this answer

Feature scaling adjusts the range or distribution of numerical features so that they are on a similar scale. This is crucial for algorithms that rely on distance calculations or gradient descent, as unscaled features can lead to slow convergence or biased results. Standardization or normalization are appropriate methods to achieve equal contribution from all features.

Exam trap

The trap here is confusing feature scaling with other preprocessing steps like encoding or dimensionality reduction, which do not address varying scales.

44
MCQmedium

A data scientist is using SageMaker Data Wrangler to prepare a large dataset. The data contains duplicate rows, which could bias the model. Which built-in step in Data Wrangler can automatically detect and remove duplicates?

A.Amazon QuickSight duplicate detection
B.Handle Duplicates transform in Data Wrangler
C.AWS Glue Studio FindDuplicates transform
D.Amazon DataZone catalog
AnswerB

The Handle Duplicates transform operates directly on the imported dataset within Data Wrangler, detecting and removing duplicate rows without external code. It satisfies the stem's requirement for a built-in step that automatically eliminates duplicates, preventing the bias they would introduce during model training.

Why this answer

The Handle Duplicates transform is a built-in step in SageMaker Data Wrangler specifically designed to detect and remove duplicate rows from a dataset. It provides configurable options such as selecting a subset of columns for duplicate detection and choosing whether to keep the first or last occurrence, directly addressing the bias risk from duplicate rows in ML training data.

Exam trap

The trap here is that candidates confuse AWS Glue Studio transforms (like FindDuplicates) with SageMaker Data Wrangler's built-in steps, as both are AWS data preparation services but operate in different environments and have distinct feature sets.

How to eliminate wrong answers

Option A is wrong because Amazon QuickSight is a business intelligence (BI) service for visualization and dashboards, not a data preparation tool with built-in duplicate detection for ML pipelines. Option C is wrong because AWS Glue Studio FindDuplicates is a transform available in AWS Glue Studio (a separate ETL service), not within SageMaker Data Wrangler's interface or step library. Option D is wrong because Amazon DataZone is a data catalog and governance service for managing data assets across an organization, not a data preparation tool that detects or removes duplicates.

45
MCQhard

A financial services company is developing a fraud detection model using Amazon SageMaker. They have a dataset with 10 million transactions, each with 300 features. The dataset is highly imbalanced (0.1% fraud). They have performed feature engineering and now need to split the data for training, validation, and test sets. The data is stored in CSV files in Amazon S3. They plan to use SageMaker's built-in XGBoost algorithm. To ensure proper evaluation and avoid data leakage, which data splitting strategy should they use?

A.Randomly shuffle the entire dataset and then split into 80% training, 10% validation, 10% test.
B.Use k-fold cross-validation on the entire dataset and average the results.
C.Perform a stratified split on the target variable to ensure each set has the same fraud ratio.
D.Apply SMOTE to balance the dataset first, then split randomly into training, validation, and test sets.
AnswerC

A stratified split preserves the 0.1% fraud ratio across training, validation and test sets, preventing the minority class from being absent or severely under-represented in any split. This satisfies the stem's requirement for proper evaluation of the imbalanced target, since random splitting could yield validation sets with too few fraud cases to assess model performance reliably.

Why this answer

A stratified split preserves the original 0.1% fraud ratio across training, validation, and test sets, which is critical for imbalanced datasets. This ensures each subset is representative of the population, allowing SageMaker's XGBoost to be evaluated fairly without data leakage. Random splits (Option A) could accidentally create a validation or test set with zero fraud cases, making evaluation meaningless.

Exam trap

The trap here is that candidates often choose random splitting (Option A) out of habit, forgetting that imbalanced datasets require stratified sampling to avoid evaluation sets with zero positive cases, which would render metrics like precision and recall undefined.

How to eliminate wrong answers

Option A is wrong because random shuffling and splitting an imbalanced dataset (0.1% fraud) risks producing validation or test sets with no fraud examples, leading to misleading accuracy metrics and inability to detect model overfitting. Option B is wrong because k-fold cross-validation on the entire dataset would leak information from future folds into training when used for final model selection, and it does not provide a held-out test set for unbiased final evaluation. Option D is wrong because applying SMOTE before splitting introduces synthetic data that can leak information across the split boundaries, causing data leakage and overly optimistic performance estimates; SMOTE should only be applied to the training set after splitting.

46
MCQhard

A machine learning team is building a model to predict customer churn. They have historical data that includes customer activity logs, each with a timestamp. The team wants to ensure that the training data does not contain any data leakage from the future. Which approach should they take when preparing the training and validation datasets?

A.Use stratified sampling based on churn label
B.Randomly split the data 80/20 for training and validation
C.Use k-fold cross-validation with shuffling
D.Split the data by time, using data before a certain date for training and after for validation
AnswerD

Time-based split ensures no future data influences training.

Why this answer

Splitting by time (chronological split) prevents data leakage by ensuring that the validation set contains only future data relative to the training set. In time-series or timestamped data, random splits can allow the model to learn from future patterns, artificially inflating performance. This approach respects the temporal dependency inherent in customer churn prediction.

Exam trap

AWS often tests the concept of data leakage in time-series contexts, where candidates mistakenly choose random splits or cross-validation with shuffling, overlooking that temporal order must be preserved to avoid future data leaking into training.

How to eliminate wrong answers

Option A is wrong because stratified sampling based on churn label preserves class distribution but does not address temporal leakage; it can still mix future and past data. Option B is wrong because random splitting ignores the timestamp order, allowing future data to leak into the training set and causing the model to learn from events that haven't occurred yet. Option C is wrong because k-fold cross-validation with shuffling randomly reorders the data, which breaks the time sequence and introduces future information into training folds.

47
MCQeasy

A retail company is building a machine learning model to predict customer churn. The data engineering team has extracted customer transaction data from Amazon Aurora and stored it as CSV files in Amazon S3. The data includes customer IDs, transaction amounts, timestamps, and product categories. A data scientist discovers that the dataset contains several missing values in the 'transaction_amount' column for about 15% of the records. The data scientist also notices that the 'customer_id' column has some duplicate entries. The team wants to prepare the data for training a churn model using Amazon SageMaker. The data is approximately 50 GB in size. What should the data scientist do to handle the missing values and duplicates efficiently while preparing the data for training?

A.Use a SageMaker notebook instance with Pandas to load the entire dataset into memory, fill missing values with the median, and drop duplicate customer IDs.
B.Use an AWS Glue ETL job to read the data from S3, apply transformations to fill missing values with the mean or median, and drop duplicate customer IDs, then write the cleaned data back to S3.
C.Drop all records with missing values in the transaction_amount column and remove duplicate customer IDs using an Athena SQL query, then store the result in S3.
D.Use an Amazon EMR cluster with Spark to read the CSV files, impute missing transaction amounts with the mean or median, and remove duplicate customers.
AnswerB

AWS Glue handles the 50 GB scale serverlessly, and its ETL transforms can impute missing transaction_amount values and drop duplicate customer IDs in one pass, writing cleaned output back to S3 for SageMaker training. This satisfies the efficiency constraint that single-node pandas processing would struggle with.

Why this answer

AWS Glue ETL jobs are serverless and designed to handle large-scale data transformations (like 50 GB) without requiring manual cluster management. Glue can read CSV files from S3, apply transformations to impute missing values with the mean or median, drop duplicate customer IDs, and write the cleaned data back to S3, all while scaling automatically to handle the data volume efficiently.

Exam trap

The trap here is that candidates often choose Option A (Pandas in a notebook) because it seems simple, but they overlook the memory limitations of a single-instance notebook when processing 50 GB of data, which is a classic 'scale vs. simplicity' trick in the MLA-C01 exam.

How to eliminate wrong answers

Option A is wrong because loading a 50 GB dataset into memory using Pandas in a SageMaker notebook instance is inefficient and likely to cause out-of-memory errors, as Pandas is single-threaded and not designed for distributed processing of large datasets. Option C is wrong because dropping all records with missing values (15% of data) would discard a significant portion of the dataset, potentially biasing the model, and Athena SQL queries do not natively support imputation of missing values with mean or median without complex workarounds. Option D is wrong because while Amazon EMR with Spark could handle the task, it requires provisioning and managing a cluster, which is more complex and less cost-effective than the serverless AWS Glue approach for this specific data preparation task.

48
MCQeasy

A data scientist is preparing a large dataset for training a machine learning model. The dataset contains missing values in several columns. Which approach is the MOST efficient for handling missing values in a large dataset using AWS services?

A.Use AWS Glue ETL to write a custom Python script that imputes missing values with the mean.
B.Use Amazon SageMaker Data Wrangler to impute missing values using built-in transforms.
C.Use pandas in a SageMaker notebook to impute missing values with the median.
D.Remove all rows with missing values from the dataset.
AnswerB

Data Wrangler provides built-in imputation transforms that run as scalable Spark processing, avoiding custom code for a large dataset. This satisfies the efficiency constraint by handling missing values across many columns in one visual flow, with results exportable directly to SageMaker training.

Why this answer

Amazon SageMaker Data Wrangler provides a visual interface and built-in transforms for handling missing values efficiently at scale, without writing custom code. Glue ETL is more code-heavy, and imputation with pandas is not scalable for large datasets. Removing all rows with missing values is not always optimal and may not be efficient.

49
MCQeasy

An ML engineer needs to split a dataset into training, validation, and test sets. The dataset has a time-based column that should not be leaked. Which split method is most appropriate?

A.Stratified split based on target
B.Temporal split based on date
C.Random split with 70/20/10
D.K-fold cross-validation
AnswerB

A temporal split partitions rows by date, so training uses earlier records and validation/test use later ones. This preserves chronological order and prevents future information leaking into training, directly satisfying the stem's constraint that the time-based column must not be leaked. Random or stratified splits would mix periods and leak future data.

Why this answer

A temporal split ensures that the time-based column is not leaked by preserving the chronological order of the data. This method uses the date column to assign earlier records to the training set and later records to the validation and test sets, preventing future information from influencing the model during training.

Exam trap

AWS often tests the concept of data leakage by presenting random or stratified splits as viable options, trapping candidates who overlook the time-based column and assume standard splitting methods are always safe.

How to eliminate wrong answers

Option A is wrong because a stratified split based on the target variable preserves class proportions but does not account for time order, leading to potential data leakage when time-dependent patterns exist. Option C is wrong because a random split ignores the temporal structure entirely, allowing future data points to appear in the training set and causing leakage. Option D is wrong because K-fold cross-validation shuffles data randomly across folds, which breaks the time sequence and introduces leakage; it is unsuitable for time-series or time-sensitive data.

50
MCQmedium

A company is building a fraud detection model on an imbalanced dataset (99% legitimate, 1% fraudulent). To improve recall on the minority class, they want to resample data. Which combination of techniques should they use?

A.SMOTE on entire dataset before train/test split
B.Random oversampling of minority class before train/test split
C.Random undersampling of majority class
D.SMOTE on training set only
AnswerD

SMOTE synthesises new minority-class examples by interpolating between existing fraudulent cases, directly raising recall on the 1% class. Applying it only to the training set preserves the genuine 99:1 distribution in validation and test data, preventing the inflated performance estimates that leakage from resampled holdout data would cause.

Why this answer

SMOTE (Synthetic Minority Over-sampling Technique) generates synthetic minority-class samples by interpolating between existing minority instances. Applying SMOTE only to the training set prevents synthetic samples from leaking into the test set, which would inflate evaluation metrics and produce an overly optimistic model. This is the correct resampling approach for improving recall on an imbalanced fraud dataset.

Exam trap

MLA-C01 often tests whether candidates apply resampling before the train/test split — the trap is forgetting that SMOTE or oversampling on the full dataset leaks synthetic information into the test set and invalidates evaluation.

How to eliminate wrong answers

Option A is wrong because applying SMOTE before the train/test split causes data leakage — synthetic samples derived from test-set instances contaminate training, producing misleadingly high evaluation metrics. Option B is wrong because random oversampling before the split duplicates minority instances across train and test, again causing leakage and overfitting. Option C is wrong because random undersampling of the majority class discards potentially useful legitimate-transaction data and, used alone, does not address the recall goal as effectively as SMOTE on the training set.

51
MCQmedium

A company is building a time series forecasting model using SageMaker DeepAR. The raw data is a CSV with columns: timestamp, item_id, and value. What is the correct data format required for DeepAR training?

A.JSON Lines files with 'start', 'target', and optional fields per time series
B.A wide-format CSV where each column is a different time series
C.Parquet files with a schema containing timestamp, item_id, and value
D.A single CSV file with columns: timestamp, item_id, value
AnswerA

DeepAR consumes JSON Lines, where each line holds one time series with a 'start' timestamp and a 'target' array of values. The stem's CSV columns must therefore be pivoted into that per-series structure; CSV is not accepted directly by the algorithm.

Why this answer

DeepAR requires time series data to be provided in JSON Lines format, where each line represents a single time series with a 'start' timestamp (in ISO 8601 format), a 'target' array of values, and optional fields like 'cat' for categorical features. This structured format allows DeepAR to handle variable-length sequences and missing values natively, which is not possible with simple CSV or wide-format data.

Exam trap

The trap here is that candidates assume DeepAR can accept raw CSV data like other SageMaker built-in algorithms (e.g., XGBoost), but DeepAR is a specialized time series algorithm that requires a specific JSON Lines structure with 'start' and 'target' fields, not a simple tabular format.

How to eliminate wrong answers

Option B is wrong because wide-format CSV (each column as a separate time series) is not supported by DeepAR; it expects each time series to be a separate JSON object, not columns. Option C is wrong because Parquet files are not a native input format for DeepAR; the built-in algorithm specifically requires JSON Lines or RecordIO-protobuf format. Option D is wrong because a single CSV with timestamp, item_id, and value columns does not provide the 'start' and 'target' structure DeepAR needs; it would require significant preprocessing to group by item_id and convert to the required JSON Lines format.

52
MCQeasy

A data scientist is preparing a dataset for a machine learning model that predicts customer churn. The dataset contains a column 'CustomerID' that is a unique identifier. What should the data scientist do with this column before training the model?

A.Keep the column as a feature because it uniquely identifies each customer.
B.Use the column as the target variable.
C.Remove the column from the feature set.
D.Encode the column using one-hot encoding.
AnswerC

CustomerID is a unique identifier carrying no generalisable signal; each value appears once, so the model would memorise rather than learn. Removing it from the feature set prevents noise and spurious splits, satisfying the requirement to prepare the churn dataset correctly.

Why this answer

'CustomerID' is a unique identifier with no predictive power for churn. Including it as a feature would cause the model to memorize individual customers rather than learn generalizable patterns, leading to overfitting and poor performance on unseen data. In machine learning, such columns should be removed during data preparation to ensure the model learns from meaningful features.

Exam trap

The trap here is that candidates may think unique identifiers are useful for tracking or that they can be encoded as categorical features, but the exam tests the principle that identifiers with no predictive relationship to the target must be removed to avoid overfitting and data leakage.

How to eliminate wrong answers

Option A is wrong because keeping 'CustomerID' as a feature introduces a high-cardinality categorical variable with no correlation to the target, which can cause overfitting and degrade model generalization. Option B is wrong because the target variable for churn prediction should be a binary or categorical label indicating churn status, not a unique identifier that has no relationship to the outcome. Option D is wrong because one-hot encoding a unique identifier like 'CustomerID' would create thousands of sparse binary columns, dramatically increasing dimensionality without adding any predictive value, and is computationally wasteful.

53
MCQeasy

A machine learning engineer needs to split a dataset into training, validation, and test sets for a SageMaker training job. The dataset is stored in Amazon S3 as a single CSV file. The engineer wants to ensure that the splits are reproducible and that the test set is never used during training or hyperparameter tuning. Which approach should the engineer use?

A.Manually split the CSV file using a Python script with a fixed random seed, write the three splits to separate S3 prefixes, and use the training and validation prefixes in the training job, keeping the test prefix for final evaluation.
B.Use the SageMaker training job's built-in data splitting feature by specifying a validation split percentage in the hyperparameters.
C.Use SageMaker Data Wrangler to split the data into three parts and export them to S3, then use all three parts in the training job with different channels.
D.Use Amazon Athena to run a query that randomly assigns rows to three splits, and store the results in S3. Then use all three splits in the training job.
AnswerA

Using a Python script with a fixed random seed ensures reproducibility. Writing the splits to separate S3 prefixes allows the training job to access only the training and validation data, while the test set remains untouched for final evaluation. This meets all requirements.

Why this answer

To ensure reproducibility and prevent data leakage, the engineer should split the data with a fixed random seed and store each split in separate S3 locations. The training job should only consume the training and validation sets, while the test set is reserved for final model evaluation. This is a standard best practice for ML workflows.

Exam trap

The trap here is assuming that SageMaker training jobs automatically split data or that using all splits in training is acceptable.

54
MCQmedium

A machine learning team stores training data in an Amazon S3 bucket and wants to catalog it so that Amazon Athena and Amazon SageMaker Feature Store can discover the schema. The data is partitioned by year, month, and day in Hive-style prefixes, and new partitions are added daily. A data engineer must ensure new partitions are automatically discoverable without manual intervention. Which solution meets these requirements?

A.Enable S3 Inventory and point Athena at the inventory report to enumerate partitions.
B.Run MSCK REPAIR TABLE on an Athena table each time new data arrives.
C.Register the S3 prefix as a SageMaker Feature Store offline store and let it infer the partitions.
D.Create an AWS Glue crawler with a daily schedule that points at the S3 prefix and updates the Data Catalog.
AnswerD

A scheduled Glue crawler scans the S3 location, detects the Hive-style year/month/day prefixes as partition columns, and updates the Data Catalog tables automatically each day. This gives Athena and other integrated services an up-to-date schema and partition list with no manual steps, satisfying the automation requirement.

Why this answer

A scheduled AWS Glue crawler is the standard mechanism for automatically discovering new Hive-style partitions in S3 and refreshing the Data Catalog. Because it runs on a daily schedule, new year/month/day partitions appear in the catalog without any manual command, and Athena and other catalog-aware services can query them immediately.

Exam trap

The trap here is reaching for Athena's partition repair command, which works only on demand and still leaves the catalog stale between runs.

55
MCQeasy

A data engineer stores raw ML training data in Amazon S3 and needs to catalog the schema, track partition changes, and make the data queryable by Amazon Athena without running ETL. Which AWS service should the engineer use?

A.Amazon SageMaker Feature Store with an offline store backed by S3.
B.AWS Glue ETL job with a Python shell script.
C.AWS Glue Data Catalog with an AWS Glue crawler.
D.Amazon Redshift Spectrum with an external schema.
AnswerC

A Glue crawler inspects data in S3, infers schema and partitions, and registers tables in the Glue Data Catalog, which Athena uses as its metastore for querying S3 data in place. This satisfies cataloging, partition tracking, and query access without moving or transforming the data, matching the stated requirement exactly.

Why this answer

The Glue Data Catalog is the central metadata repository that stores table definitions, schemas, and partition information for data in S3, and Athena natively uses it as its metastore. A crawler automates schema inference and partition discovery, so no ETL is required. The other services either transform data, serve features, or consume the catalog rather than create it.

Exam trap

The trap here is confusing a service that consumes the Glue Data Catalog, such as Athena or Redshift Spectrum, with the service that actually builds and maintains the catalog entries.

56
MCQeasy

A data engineer needs to convert a JSON dataset to Parquet format for efficient querying with Amazon Athena. The JSON files are in an S3 bucket. Which service can perform this conversion with minimal coding?

A.Amazon SageMaker Processing
B.Amazon EMR
C.AWS Lambda
D.AWS Glue Studio with a visual job
AnswerD

AWS Glue Studio's visual job editor generates the PySpark ETL code that reads JSON from S3, applies a schema, and writes Parquet back to S3, satisfying the minimal-coding constraint. Its built-in transforms and crawler-derived schemas remove hand-written conversion logic, and the output is directly queryable by Athena.

Why this answer

AWS Glue Studio with a visual job is the correct choice because it provides a no-code, drag-and-drop interface to create ETL jobs that can read JSON from S3 and write it as Parquet, with built-in schema inference and transformation capabilities. This minimizes coding effort while leveraging Glue's serverless Spark engine for efficient conversion, making it ideal for preparing data for Athena queries.

Exam trap

The trap here is that candidates often confuse AWS Glue Studio with AWS Glue DataBrew or assume that any AWS service with 'processing' in its name (like SageMaker Processing) is suitable for simple ETL tasks, overlooking the specific no-code visual job capability of Glue Studio.

How to eliminate wrong answers

Option A is wrong because Amazon SageMaker Processing is designed for data preprocessing and model training workflows within the ML pipeline, not for simple file format conversion; it requires writing custom processing scripts and managing infrastructure, which adds unnecessary complexity. Option B is wrong because Amazon EMR is a managed Hadoop/Spark cluster that can perform the conversion, but it requires provisioning and configuring a cluster, writing Spark or Hive code, and managing lifecycle, which is far more coding and operational overhead than a visual job. Option C is wrong because AWS Lambda has a maximum execution time of 15 minutes and a deployment package size limit, making it impractical for converting large JSON datasets to Parquet; it also requires custom Python code with libraries like PyArrow or Pandas, which is not minimal coding.

57
Multi-Selecteasy

Which TWO actions are recommended best practices when preparing training data for a machine learning model in AWS? (Choose two.)

Select 2 answers
A.Remove all outliers from the dataset.
B.Train the model on the entire dataset to maximize data usage.
C.Check for and handle missing values appropriately.
D.Split the data into training, validation, and test sets.
E.Always normalize all features to a [0,1] range.
AnswersC, D

Missing values bias model training and can cause failures in algorithms that reject nulls. Detecting and handling them—via imputation, removal, or indicator flags—preserves data integrity, satisfying the best-practise requirement for robust training data preparation in AWS pipelines such as SageMaker Data Wrangler or Glue.

Why this answer

Option C is correct because missing values can bias or break training algorithms, so best practice is to detect them (e.g., with pandas isnull() or Amazon SageMaker Data Wrangler) and handle them via imputation, removal, or model-native handling. Option D is correct because splitting data into training, validation, and test sets lets you fit parameters, tune hyperparameters, and estimate generalization performance on unseen data, avoiding overfitting and data leakage. Option A is not recommended because not all outliers are errors; blindly removing them can discard legitimate signal and distort the distribution.

Option B is wrong because training on the entire dataset leaves no held-out data for validation or unbiased testing. Option E is wrong because normalization is not always required and [0,1] scaling is only one of several techniques (e.g., standardization, log transforms) chosen based on the algorithm and feature distribution.

Exam trap

The trap here is that candidates assume all outliers must be removed (Option A) or that normalization is always required (Option E), but the exam tests nuanced understanding that these steps depend on the algorithm and data characteristics, not blanket rules.

58
MCQeasy

A data engineer is preparing a large dataset of 10 TB for ML training on Amazon SageMaker. The data is stored in Amazon S3 as CSV files. To reduce training time and cost, the engineer wants to use a columnar format that is optimized for analytical queries. Which format should the engineer convert the data to?

A.XML
B.Parquet
C.ORC
D.JSON Lines
AnswerB

Parquet stores data column-wise with compression and encoding, so SageMaker reads only the columns each training job needs rather than scanning every field. That columnar layout and smaller footprint cut I/O and cost across the 10 TB CSV dataset, unlike row-based formats such as CSV or JSON.

Why this answer

Parquet is a columnar storage format that is highly optimized for analytical queries and is natively supported by Amazon SageMaker for efficient data loading. By converting the 10 TB of CSV data to Parquet, the data engineer can reduce I/O and storage costs because columnar formats allow SageMaker to read only the columns needed for training, rather than scanning entire rows. This directly addresses the goal of reducing training time and cost for ML workloads.

Exam trap

AWS often tests the distinction between columnar formats (Parquet vs. ORC) by making both appear correct, but the trap here is that ORC is tightly coupled with Hive and less commonly used with SageMaker, while Parquet is the de facto standard for AWS-native ML and analytics services.

How to eliminate wrong answers

Option A (XML) is wrong because XML is a verbose, row-oriented text format that is not optimized for analytical queries; it would increase storage size and I/O overhead, making training slower and more expensive. Option C (ORC) is also a columnar format optimized for analytical queries, but it is primarily designed for and tightly integrated with the Apache Hive ecosystem, whereas Parquet is the more universally supported and recommended format for Amazon SageMaker and AWS analytics services. Option D (JSON Lines) is wrong because it is a row-oriented, text-based format that lacks the compression and columnar pruning benefits of Parquet, leading to higher storage costs and slower data access for ML training.

59
Multi-Selectmedium

A data engineer is using Amazon SageMaker Data Wrangler to prepare a dataset for a machine learning model. The engineer needs to ensure that the data preparation flow is reproducible and can be integrated into an MLOps pipeline. Which TWO actions should the engineer take to achieve this? (Choose two.)

Select 2 answers
A.Manually document each transformation step in a text file and apply them using a Jupyter notebook.
B.Use the Data Wrangler interface to directly train a model without exporting the flow.
C.Use Data Wrangler's built-in scheduling feature to run the flow daily.
D.Export the Data Wrangler flow to a SageMaker Pipeline as a processing step.
E.Export the Data Wrangler flow as a Python script and commit it to a Git repository.
AnswersD, E

Exporting the Data Wrangler flow to a SageMaker Pipeline creates a processing step that can be versioned, scheduled, and executed as part of an MLOps workflow. This ensures reproducibility because the pipeline definition captures the entire data preparation logic, and it can be rerun with different inputs or parameters.

Why this answer

To ensure reproducibility and MLOps integration, the Data Wrangler flow should be exported either as a SageMaker Pipeline processing step or as a Python script committed to version control. Both methods capture the data preparation logic in a versionable, executable form that can be automated and tracked, meeting the requirements.

Exam trap

The trap here is assuming that Data Wrangler has a built-in scheduling feature or that manual documentation is sufficient for reproducibility.

60
Multi-Selecthard

A data engineer is building an AWS Glue ETL job that reads raw JSON clickstream events from Amazon S3, flattens nested structures, and writes Parquet to a curated S3 prefix for SageMaker training. The job must run daily on only the newly arrived files and must keep the Glue Data Catalog table current so Athena and SageMaker can query it. Which TWO configurations should the engineer apply to meet these requirements? (Choose two.)

Select 2 answers
A.Convert the input JSON to CSV before processing to improve Glue read performance.
B.Enable job bookmarks on the Glue ETL job.
C.Set the job's maximum capacity to the smallest possible value to reduce cost.
D.Enable the Glue Data Catalog update option and configure crawlers or partition projection for the output prefix.
E.Use a Development endpoint to run the daily job interactively.
AnswersB, D

Job bookmarks persist state about which S3 objects and partitions have already been processed, so each daily run reads only newly arrived files. Without bookmarks, every run reprocesses the entire input prefix, which wastes DPU hours and can duplicate curated output. Bookmarks are the standard Glue mechanism for incremental processing and directly satisfy the requirement to handle only new data.

Why this answer

Job bookmarks give Glue the persistent state needed to process only newly arrived S3 objects, which is exactly the incremental requirement. Keeping the Data Catalog current, either through the job's catalog update option or a crawler over the curated prefix, ensures Athena and SageMaker see the latest schema and partitions. Together these two settings make the daily pipeline both efficient and queryable.

Exam trap

The trap here is treating Glue job bookmarks as a performance tuning knob rather than the stateful mechanism that actually enables incremental reads of new S3 objects.

61
MCQhard

A data scientist creates a feature group as shown in the exhibit. When ingesting data with an 'age' column of integer values, the ingestion fails. What is the most likely cause?

A.The role does not have permissions to write to the feature store.
B.The `age` feature type should be `Integral`, not `String`.
C.The `OnlineStoreConfig` must include a `SecurityConfig`.
D.The `EventTimeFeatureName` is incorrectly spelled.
AnswerB

Feature Store validates each value against its declared feature type. An integer age ingested into a feature defined as String violates that schema and the ingestion job fails. Redeclaring the feature as Integral aligns the type with the data.

Why this answer

The feature group definition specifies the 'age' column as a `String` type, but the ingested data contains integer values. Amazon SageMaker Feature Store requires that the data types of ingested records match the schema defined in the feature group. When a mismatch occurs, such as providing an integer for a string field, the ingestion fails with a type conversion error.

Exam trap

AWS often tests the distinction between schema definition and actual data types, trapping candidates who overlook that the feature group schema must exactly match the ingested data's types, not just the column names.

How to eliminate wrong answers

Option A is wrong because the question states the ingestion fails specifically due to a data type mismatch, not a permissions issue; a permissions error would typically occur at the API call level, not during data parsing. Option C is wrong because `SecurityConfig` is not a required field in `OnlineStoreConfig`; the online store configuration only requires an `EnableOnlineStore` boolean and optionally a `SecurityGroupIdList` and `SubnetIdList` for VPC settings. Option D is wrong because the `EventTimeFeatureName` is spelled correctly as 'EventTime' in the exhibit, and a misspelling would cause a different error (e.g., 'InvalidParameterValue') rather than a data type mismatch.

62
MCQmedium

A company collects sensor data from IoT devices. The data arrives with missing timestamps due to network issues. For anomaly detection, the engineer needs to create features that capture rolling statistics over fixed windows. Which data preprocessing step is essential before feature generation?

A.Remove missing timestamps
B.Resample data to a fixed frequency
C.Sort data by device ID
D.Impute missing values with forward fill
AnswerB

Rolling statistics require evenly spaced observations; irregular timestamps from dropped network packets make window boundaries ambiguous. Resampling to a fixed frequency creates uniform intervals, so rolling means and variances are computed over consistent, comparable windows.

Why this answer

Resampling the data to a fixed frequency is essential because rolling window statistics require a consistent time index to compute accurate aggregations over fixed windows. Without a uniform timestamp grid, the window boundaries become ambiguous and the resulting features will be misaligned or incomplete, undermining the anomaly detection model.

Exam trap

AWS often tests the distinction between handling missing values (imputation) and handling irregular timestamps (resampling), leading candidates to confuse forward-fill as a solution for time alignment when it only addresses missing data points, not the underlying time index irregularity.

How to eliminate wrong answers

Option A is wrong because simply removing missing timestamps discards valuable data and does not address the need for a consistent time index; the remaining timestamps remain irregularly spaced. Option C is wrong because sorting by device ID organizes data by device but does not fix the irregular timestamp spacing required for fixed-window rolling statistics. Option D is wrong because forward-fill imputation fills missing values but does not create a uniform time grid; the timestamps themselves remain irregular, so rolling windows cannot be applied consistently.

63
Multi-Selecteasy

A data engineer is using AWS Glue to prepare a dataset for machine learning. The dataset has several columns with outliers. The engineer wants to detect and handle outliers in a scalable manner. Which TWO approaches should the engineer consider? (Select TWO.)

Select 2 answers
A.Manually remove outliers by inspecting the data in Amazon S3.
B.Train a neural network to identify anomalies and remove them.
C.Use pandas in a SageMaker notebook to calculate z-scores and filter outliers.
D.Use AWS Glue DynamicFrame with Apache Spark to compute interquartile range (IQR) and filter outliers.
E.Use Amazon SageMaker Data Wrangler to apply an outlier detection transform.
AnswersD, E

Computing IQR per column on a DynamicFrame lets Spark calculate quartiles and filter rows outside the fences in a distributed, scalable manner. This satisfies the outlier-detection constraint across the full dataset without collecting data to a single node.

Why this answer

Option D is correct because AWS Glue DynamicFrames run on Apache Spark, which scales horizontally across a cluster, and computing the interquartile range (IQR) with Spark SQL/DataFrame operations lets the engineer detect and filter outliers (e.g., values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) across large datasets efficiently. Option E is correct because Amazon SageMaker Data Wrangler provides built-in outlier detection transforms (such as Robust Standard Deviation, Standard Deviation, and Quantile-based detection) that can be applied to a dataset at scale and exported into a processing or training pipeline, fitting the ML preparation workflow. Option A is not appropriate because manually inspecting data in Amazon S3 does not scale and is error-prone for large datasets.

Option B is not appropriate because training a neural network just to remove outliers is overkill, costly, and unnecessary when statistical methods suffice. Option C is not appropriate because pandas in a SageMaker notebook operates in memory on a single instance, so it does not scale to large datasets the way Spark or Data Wrangler do.

Exam trap

The trap here is that candidates may assume that only a single AWS service can handle outlier detection at scale, but the question requires selecting two approaches, and both Glue DynamicFrames and SageMaker Data Wrangler are valid, scalable, and managed AWS solutions for this task.

64
MCQmedium

A data scientist must join a 50 GB transactional table with a small 5 MB lookup table in AWS Glue before writing Parquet output for SageMaker training. The join currently shuffles the large table across the cluster and the job runs slowly. Which optimization should the data scientist apply?

A.Configure a broadcast join by passing the small table to the join as a broadcast hint.
B.Convert the small lookup table to a Glue DynamicFrame and apply a relationalize transform.
C.Repartition both tables on the join key before the join operation.
D.Increase the number of DPUs allocated to the Glue job so more executors perform the shuffle.
AnswerA

When one side of a join is small enough to fit in memory, broadcasting it to every executor avoids shuffling the large table across the network. In AWS Glue with Spark, this is expressed through a broadcast hint or by relying on the automatic broadcast threshold, and it directly removes the shuffle that is slowing the job.

Why this answer

A broadcast join sends the small lookup table to every executor so the large table can be streamed locally without a network shuffle. Because the lookup table is only a few megabytes, it fits comfortably in executor memory. Increasing cluster size, repartitioning the large table, or reshaping nested data all leave the expensive shuffle in place, so none of them addresses the actual bottleneck.

Exam trap

The trap here is responding to a slow Spark join by adding cluster capacity, when the real fix is changing the join strategy so the shuffle disappears entirely.

65
Multi-Selecthard

A data engineer is optimizing Amazon Athena queries on large datasets stored in S3 for machine learning data preparation. Which THREE practices improve query performance?

Select 3 answers
A.Partition the data by a frequently filtered column, such as date
B.Use uncompressed CSV files for simplicity
C.Partition the data by every column to maximize filtering
D.Store data in columnar formats like Parquet or ORC
E.Compress the data with Snappy or gzip
AnswersA, D, E

Partitioning by a frequently filtered column such as date lets Athena prune irrelevant partitions, scanning only the data each query needs. This directly reduces bytes read from S3, which is the dominant cost and latency factor for large datasets.

Why this answer

Option A is correct because partitioning the S3 data by a frequently filtered column such as date lets Athena prune partitions and scan only the relevant prefixes, drastically reducing the amount of data read and thus improving performance and lowering cost. Option D is correct because columnar formats like Parquet or ORC allow Athena to read only the columns referenced by a query and provide better compression and predicate pushdown, minimizing I/O compared to row-based formats. Option E is correct because compressing data with Snappy or gzip reduces the bytes transferred from S3 and read by Athena, and these splittable/columnar-friendly codecs work well with Parquet/ORC to further speed up scans.

Option B is incorrect because uncompressed CSV is a row-based, non-splittable-friendly format that forces Athena to scan all columns and more bytes, hurting performance. Option C is incorrect because partitioning by every column creates an excessive number of small partitions and files, which increases metadata overhead and can degrade rather than improve query performance.

Exam trap

AWS often tests the misconception that more partitions always improve performance, but in reality, over-partitioning leads to metastore overhead and small file problems that degrade query performance.

66
MCQeasy

A data engineer stores a training dataset in Amazon S3 across a prefix hierarchy organized by year, month, and day, and analysts query it with Amazon Athena. The engineer wants Athena to scan only the relevant date range instead of the entire dataset. Which action should the engineer take?

A.Enable S3 Transfer Acceleration on the bucket to speed up object retrieval during queries.
B.Convert the dataset to JSON and rely on Athena schema-on-read to skip irrelevant records.
C.Store the data as a single large uncompressed CSV object so Athena reads it sequentially.
D.Define the S3 prefix structure as partition columns in the table and query with a WHERE clause on those columns.
AnswerD

Athena prunes partitions when the query filters on partition columns that match the S3 prefix layout, so only the objects in the matching year, month, and day prefixes are read. Registering the hierarchy as partition columns and filtering on them in the WHERE clause directly reduces the bytes scanned and the query cost for the date-range queries.

Why this answer

Athena reduces cost and latency by reading only the partitions that satisfy the query predicate. When the S3 layout already encodes year, month, and day, exposing those levels as partition columns in the table definition lets a WHERE filter on those columns prune the scan to the matching prefixes. The engineer should register the partitions and query against the partition keys rather than scanning the whole dataset.

Exam trap

The trap here is focusing on making individual object reads faster when the real win is not reading the irrelevant objects at all through partition pruning.

67
MCQeasy

A data scientist is preparing a dataset for training a binary classification model. The dataset has 100,000 rows and 50 features. The target variable is imbalanced, with only 5% positive cases. Which technique should the data scientist apply to address the class imbalance BEFORE training?

A.Principal Component Analysis (PCA) dimensionality reduction
B.Random oversampling of the minority class
C.Standard scaling of numerical features
D.One-hot encoding of categorical variables
AnswerB

Random oversampling duplicates minority-class rows, raising the 5% positive rate toward parity so the algorithm no longer biases toward the majority class. It satisfies the pre-training constraint by rebalancing class distribution before the model sees the data, unlike threshold tuning applied afterwards.

Why this answer

Random oversampling of the minority class (Option B) directly addresses the class imbalance by duplicating examples from the positive class until the class distribution is more balanced. This prevents the binary classification model from being biased toward the majority class, which is critical when only 5% of the 100,000 rows are positive cases. Oversampling is applied before training to ensure the model sees sufficient minority examples during learning.

Exam trap

AWS often tests whether candidates confuse data preprocessing techniques (scaling, encoding, dimensionality reduction) with methods that directly modify the class distribution, leading them to pick a plausible but irrelevant option like PCA or scaling.

How to eliminate wrong answers

Option A is wrong because PCA dimensionality reduction reduces the number of features but does not alter the class distribution; it would not fix the 5% imbalance and could even discard variance useful for separating the minority class. Option C is wrong because standard scaling normalizes numerical feature ranges but has no effect on the ratio of positive to negative samples; it addresses feature magnitude, not class imbalance. Option D is wrong because one-hot encoding converts categorical variables into binary columns but does not change the target variable's distribution; it is a preprocessing step for feature representation, not for balancing classes.

68
MCQmedium

A healthcare company is building a model to predict patient readmission rates. The dataset contains a mix of numeric features (age, blood pressure, lab test results) and categorical features (gender, diagnosis code, hospital department). The dataset has 2 million rows. The data is stored in an Amazon S3 bucket, and they use AWS Glue to catalog and preprocess the data. The data scientist notices that the 'diagnosis_code' column has 10,000 unique codes, and 20% of the rows have missing values for 'blood_pressure'. They plan to use a SageMaker built-in XGBoost model. For optimal model performance, which preprocessing steps should they apply using AWS Glue ETL?

A.Impute missing 'blood_pressure' with the mean, and apply label encoding to 'diagnosis_code'.
B.Impute missing 'blood_pressure' with median, and apply integer encoding to 'diagnosis_code'.
C.Replace missing 'blood_pressure' with -1 and apply one-hot encoding to 'diagnosis_code' after grouping rare codes into 'other'.
D.Apply one-hot encoding to 'diagnosis_code' and drop rows with missing 'blood_pressure'.
AnswerB

Median is robust; integer encoding is sufficient for tree-based models like XGBoost.

Why this answer

XGBoost handles missing values natively, so median imputation for 'blood_pressure' is robust to outliers and preserves data distribution, while integer encoding (label encoding) for 'diagnosis_code' with 10,000 unique values is efficient and avoids the dimensionality explosion of one-hot encoding. AWS Glue ETL can apply these transformations using built-in functions like `Imputer` and `StringIndexer` without excessive memory overhead.

Exam trap

The trap here is that candidates overestimate the need for one-hot encoding with high-cardinality categorical features, forgetting that tree-based models like XGBoost can effectively use integer encoding, and they may also default to mean imputation without considering outlier sensitivity.

How to eliminate wrong answers

Option A is wrong because mean imputation for 'blood_pressure' is sensitive to outliers, which can skew the model, and label encoding is a form of integer encoding but the term 'label encoding' often implies ordinal mapping that may introduce unintended ordinal relationships; however, the primary flaw is the mean imputation choice. Option C is wrong because replacing missing 'blood_pressure' with -1 introduces an arbitrary value that XGBoost may misinterpret as a valid numeric pattern, and one-hot encoding 'diagnosis_code' with 10,000 categories (even after grouping rare codes) still creates a very high-dimensional sparse matrix that degrades performance and increases memory usage in Glue ETL. Option D is wrong because dropping 20% of rows with missing 'blood_pressure' leads to significant data loss and potential bias, and one-hot encoding 'diagnosis_code' with 10,000 categories is computationally prohibitive and unnecessary for tree-based models like XGBoost.

69
MCQhard

A team is using AWS Glue to process streaming data from Amazon Kinesis. The streaming data contains both structured and semi-structured fields. The team needs to flatten the semi-structured fields into columns for downstream ML training. Which Glue feature is BEST suited?

A.Relationalize transform
B.Spigot transform
C.ResolveChoice transform
D.ApplyMapping transform
AnswerA

Relationalize flattens nested semi-structured data into separate relational tables that Glue can join, converting arrays and structs into columns. This directly satisfies the requirement to expose semi-structured Kinesis fields as flat columns for downstream ML training.

Why this answer

The Relationalize transform is specifically designed to flatten nested JSON or semi-structured fields into a relational structure, making it ideal for converting complex streaming data from Kinesis into flat columns for ML training. It automatically handles arrays and structs by creating separate tables or columns, which is exactly what the team needs for downstream processing.

Exam trap

The trap here is that candidates confuse 'flattening semi-structured data' with simple schema operations like type resolution or column mapping, leading them to choose ResolveChoice or ApplyMapping instead of the specialized Relationalize transform.

How to eliminate wrong answers

Option B is wrong because the Spigot transform is used to sample or write a subset of data to a specified location for debugging or testing, not for flattening semi-structured fields. Option C is wrong because the ResolveChoice transform resolves ambiguity when a column has multiple data types (e.g., string vs. int) by casting to a chosen type, but it does not flatten nested structures. Option D is wrong because the ApplyMapping transform renames, casts, or drops columns based on a mapping specification, but it cannot flatten nested JSON or semi-structured data into separate columns.

Ready to test yourself?

Try a timed practice session using only Data Preparation for Machine Learning questions.