Be able to select and sequence the right AWS services to ingest, transform, and store training data in S3, then apply the correct SageMaker transform for duplicates, imbalance, or mixed feature types. The most important thing is matching the tool to the source and target format without leaking data across splits.
Start practicing
Data Preparation for Machine Learning — choose a session length
Free · No account required
Domain overview
This domain covers getting data into and ready for SageMaker: ingestion from RDS, S3, and other sources; transformation with Data Wrangler, Glue, and Processing jobs; and handling imbalance, duplicates, and mixed feature types. Questions are scenario-based, asking you to pick the right AWS service, format, or built-in transform for a stated constraint.
Exam objectives
Choosing AWS Glue or SageMaker Processing to transform RDS data into Parquet on S3
Using SageMaker Data Wrangler built-in transforms to detect and remove duplicate rows
Applying SageMaker Processing or imbalanced-learn techniques to rebalance skewed class datasets
Handling mixed numeric and categorical features with encoding, scaling, and SageMaker Feature Store
Assuming Data Wrangler alone moves data from RDS to S3 without a Glue or Processing job
Forgetting that Parquet conversion requires a schema-aware engine, not a plain S3 copy
Resampling before the train/test split, leaking synthetic samples into validation data
Click any question to see the full explanation and answer options, or start a focused practice session above.
A data scientist is preparing a large dataset for training a machine learning model. The dataset contains missing values in several columns. Which approach is the MOST efficient for handling missing values in a large dataset using AWS services?
2A data scientist is preparing a dataset for a machine learning model that predicts customer churn. The dataset contains a column 'CustomerID' that is a unique identifier. What should the data scientist do with this column before training the model?
3A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset. The dataset contains a column with date strings in the format 'YYYY-MM-DD'. The data scientist wants to extract the year, month, and day as separate features. Which Data Wrangler transform should be used?
4A data engineer is using AWS Glue to prepare a dataset for machine learning. The dataset has several columns with outliers. The engineer wants to detect and handle outliers in a scalable manner. Which TWO approaches should the engineer consider? (Select TWO.)
5A company uses AWS Glue ETL jobs to transform data for machine learning. They have a dataset with a column 'income' that is heavily right-skewed. Which transformation should be applied to make the distribution more Gaussian-like?
6A team is using Amazon SageMaker Processing for data preprocessing. They have a Parquet dataset in Amazon S3. Which configuration will provide the most efficient reading of the dataset during processing?
7A team is preparing text data for a natural language processing (NLP) model. They have a corpus of customer reviews. Which THREE preprocessing steps are essential to reduce noise and improve model performance?
8A data scientist is preparing a dataset for training a binary classification model. The dataset has 100,000 rows and 50 features. The target variable is imbalanced, with only 5% positive cases. Which technique should the data scientist apply to address the class imbalance BEFORE training?
9A machine learning engineer is building a pipeline to preprocess text data for a sentiment analysis model. The data consists of customer reviews. The engineer wants to convert the text into numerical features while preserving the semantic meaning of words. Which technique should be used?
10A company uses Amazon SageMaker Data Wrangler to prepare data for ML. The dataset contains a timestamp column and sensor readings from IoT devices. The data scientist needs to create features such as moving averages and rolling statistics over time windows. Which Data Wrangler transformation type should be selected?
11A data engineer is preparing a large dataset of 10 TB for ML training on Amazon SageMaker. The data is stored in Amazon S3 as CSV files. To reduce training time and cost, the engineer wants to use a columnar format that is optimized for analytical queries. Which format should the engineer convert the data to?
12A machine learning team is building a model to predict customer churn. They have historical data that includes customer activity logs, each with a timestamp. The team wants to ensure that the training data does not contain any data leakage from the future. Which approach should they take when preparing the training and validation datasets?
13A data scientist is working with a dataset that contains missing values in several numeric features. The data scientist wants to impute the missing values with the median of each feature. Which Amazon SageMaker Data Wrangler transformation should be used?
14A company is building a real-time inference pipeline for an ML model. The raw data arrives in JSON format via Amazon Kinesis Data Streams. Before invoking the SageMaker endpoint, the data must be preprocessed to match the training data format. Which THREE steps should be included in the preprocessing function? (Select THREE)
15A data engineer is using AWS Glue to prepare a dataset for ML. The engineer wants to split the dataset into training and testing sets while preserving the distribution of the target variable. Which TWO methods achieve this goal? (Select TWO)
16A data scientist creates a feature group as shown in the exhibit. When ingesting data with an 'age' column of integer values, the ingestion fails. What is the most likely cause?
17An organization stores raw data in Amazon S3 as CSV files. They need to perform serverless data transformation and convert the data to Parquet format for efficient ML training. Which AWS service is most appropriate?
18A data scientist is using SageMaker Data Wrangler to prepare a large dataset. The data contains duplicate rows, which could bias the model. Which built-in step in Data Wrangler can automatically detect and remove duplicates?
19A data scientist needs to split a dataset into training, validation, and test sets. The dataset has a categorical target variable with imbalanced class distribution. Which splitting technique ensures that each subset has a similar proportion of each class?
20A company ingests daily log data into an S3 bucket. They need to update the existing ML training dataset with new data without reprocessing the entire history. Which two strategies should they adopt? (Choose two.)
21A data scientist is training a binary classifier on a highly imbalanced dataset (1:100 class ratio). The dataset contains 500,000 rows and 30 features. The data is stored in S3 in Parquet format. The data scientist wants to use SageMaker's built-in XGBoost algorithm. Which data preparation technique should the data scientist apply to best address the class imbalance without causing data leakage?
22A data engineer needs to prepare a large dataset for machine learning. The data is stored in an Amazon RDS MySQL database and needs to be transformed and moved to an S3 bucket in Parquet format for use with SageMaker. Which AWS service is most suitable for this extraction, transformation, and loading (ETL) task?
23A team is building a recommendation system and wants to store and serve features for online and offline models. The features include user statistics (updated daily) and movie metadata (static). The team needs low-latency inference for real-time recommendations and wants to reuse features across multiple models. Which AWS service should the team use to store, manage, and serve these features?
24A company has 10 TB of log data in compressed JSON format stored in Amazon S3. The data needs to be processed and transformed into a structured format for machine learning. The processing requires complex transformations, including parsing nested JSON and joining with a reference table. The company wants to minimize infrastructure management. Which approach should the company use?
25A company has a dataset of 2 billion records stored as text files in Amazon S3. The data is partitioned by year and month. The data science team wants to read only the last 6 months of data for model training using SageMaker. To minimize data scanned and reduce costs, which approach should the team use?
26A company is building a fraud detection model on an imbalanced dataset (99% legitimate, 1% fraudulent). To improve recall on the minority class, they want to resample data. Which combination of techniques should they use?
27A team is using Amazon SageMaker for feature engineering. They have a dataset with a column 'TransactionDate' in string format (e.g., '2023-01-15 10:30:00'). They need to create features: year, month, day, hour, and day_of_week. What is the most efficient way to do this in a SageMaker processing job?
28An ML engineer needs to split a dataset into training, validation, and test sets. The dataset has a time-based column that should not be leaked. Which split method is most appropriate?
29A company collects sensor data from IoT devices. The data arrives with missing timestamps due to network issues. For anomaly detection, the engineer needs to create features that capture rolling statistics over fixed windows. Which data preprocessing step is essential before feature generation?
30A machine learning engineer is preparing a dataset for a binary classification model. The dataset has 10,000 rows and 200 features, with 5% positive class. The engineer suspects class imbalance may affect model performance. Which TWO actions should the engineer take to mitigate imbalance? (Choose 2.)
31A data engineer needs to convert a JSON dataset to Parquet format for efficient querying with Amazon Athena. The JSON files are in an S3 bucket. Which service can perform this conversion with minimal coding?
32A team is using AWS Glue to process streaming data from Amazon Kinesis. The streaming data contains both structured and semi-structured fields. The team needs to flatten the semi-structured fields into columns for downstream ML training. Which Glue feature is BEST suited?
33Refer to the exhibit. A data scientist is trying to use AWS Glue to read data from the S3 bucket `ml-data-bucket`. The Glue job fails with an access denied error. What is the most likely cause?
34A data engineer is optimizing Amazon Athena queries on large datasets stored in S3 for machine learning data preparation. Which THREE practices improve query performance?
35A data scientist is preparing a large dataset (50 GB) for training a TensorFlow model on SageMaker. The dataset consists of many small CSV files. Training is slow due to I/O bottlenecks. Which data preparation strategy most effectively accelerates training?
36A company is building a time series forecasting model using SageMaker DeepAR. The raw data is a CSV with columns: timestamp, item_id, and value. What is the correct data format required for DeepAR training?
37A data engineer needs to prepare a large dataset (10 TB) stored in Amazon S3 for a training job on SageMaker. The data is in CSV format, but the training algorithm expects Parquet for performance. The engineer must transform the data with minimal cost and without writing custom code. Which service should be used?
38A data engineer is using Amazon SageMaker Processing to run a data preprocessing script on a dataset with 500 million rows. The script runs out of memory on a single ml.r5.24xlarge instance. The engineer needs to modify the processing job to handle the dataset size. Which approach is most cost-effective and scalable?
39Which TWO actions are recommended best practices when preparing training data for a machine learning model in AWS? (Choose two.)
40A company runs an online retail business and wants to build a product recommendation system. They have a dataset of customer purchases stored in Amazon S3 as CSV files. The dataset includes columns: 'customer_id', 'product_id', 'purchase_date', 'quantity', 'price', and 'category'. The data science team plans to use Amazon SageMaker to train a factorization machines model. During data exploration, they discover that the 'category' column has 1,200 unique values, and many categories appear only a few times. The 'product_id' column has 50,000 unique values. They want to include both features in the model. The team is concerned about the high cardinality of these features. Which approach should they take to prepare these features for the factorization machines model?
41A healthcare company is building a model to predict patient readmission rates. The dataset contains a mix of numeric features (age, blood pressure, lab test results) and categorical features (gender, diagnosis code, hospital department). The dataset has 2 million rows. The data is stored in an Amazon S3 bucket, and they use AWS Glue to catalog and preprocess the data. The data scientist notices that the 'diagnosis_code' column has 10,000 unique codes, and 20% of the rows have missing values for 'blood_pressure'. They plan to use a SageMaker built-in XGBoost model. For optimal model performance, which preprocessing steps should they apply using AWS Glue ETL?
42A financial services company is developing a fraud detection model using Amazon SageMaker. They have a dataset with 10 million transactions, each with 300 features. The dataset is highly imbalanced (0.1% fraud). They have performed feature engineering and now need to split the data for training, validation, and test sets. The data is stored in CSV files in Amazon S3. They plan to use SageMaker's built-in XGBoost algorithm. To ensure proper evaluation and avoid data leakage, which data splitting strategy should they use?
43A data engineer is building a data pipeline for a machine learning model that requires both structured and unstructured data. The structured data (customer demographics) is in Amazon RDS, and the unstructured data (customer support chat logs) is in Amazon S3 as JSON files. The engineer needs to combine these datasets into a single training dataset stored in S3 in Parquet format. They must also perform feature engineering such as text vectorization on the chat logs. The pipeline should be serverless and cost-effective. Which approach should they use?
44A retail company is building a machine learning model to predict customer churn. The data engineering team has extracted customer transaction data from Amazon Aurora and stored it as CSV files in Amazon S3. The data includes customer IDs, transaction amounts, timestamps, and product categories. A data scientist discovers that the dataset contains several missing values in the 'transaction_amount' column for about 15% of the records. The data scientist also notices that the 'customer_id' column has some duplicate entries. The team wants to prepare the data for training a churn model using Amazon SageMaker. The data is approximately 50 GB in size. What should the data scientist do to handle the missing values and duplicates efficiently while preparing the data for training?
45A social media company is processing a real-time stream of user activity data from Amazon Kinesis Data Streams to train a machine learning model for content recommendation. The raw data includes user ID, timestamp, content ID, interaction type (like, share, comment), and device type. The data scientists need to aggregate features per user over a sliding window of 7 days, including counts of interaction types, unique content IDs engaged, and a moving average of interaction timestamps. The aggregated data will be used to update a user embedding model. The streaming data volume is approximately 500 records per second, and the company uses an AWS Glue streaming ETL job for transformation. However, the Glue job is failing frequently with high latency and checkpoint errors. The team needs a more robust solution to prepare the streaming data features. Which approach should the team take?
46A data engineer is building an AWS Glue ETL job that reads raw JSON clickstream events from Amazon S3, flattens nested structures, and writes Parquet to a curated S3 prefix for SageMaker training. The job must run daily on only the newly arrived files and must keep the Glue Data Catalog table current so Athena and SageMaker can query it. Which TWO configurations should the engineer apply to meet these requirements? (Choose two.)
47A data scientist has a 40 GB CSV dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with pandas, and the scientist wants to reduce both storage cost and training-time I/O without changing the logical schema. Which data preparation action should be taken?
48A data engineer needs to prepare a dataset for a fraud detection model. The dataset contains a highly skewed numerical feature with extreme outliers. The engineer decides to apply a logarithmic transformation to this feature before training. Which SageMaker Data Wrangler transform should be used to apply the logarithmic transformation?
49A data scientist is preparing a dataset stored in Amazon S3 for a SageMaker training job. The dataset contains missing values in several columns. The scientist wants to impute missing values with the mean of each column. Which SageMaker built-in algorithm or processing method should be used to perform this imputation efficiently?
50A data engineer must prepare a 4 TB Parquet dataset stored in Amazon S3 for a SageMaker training job that runs on 8 ml.p4d.24xlarge instances. The engineer wants the fastest possible data throughput during training while minimizing per-epoch I/O overhead. The dataset is immutable for the duration of the training run. Which approach BEST meets these requirements?
51A machine learning engineer is using Amazon SageMaker Processing to preprocess a large dataset. The processing job runs a custom Python script that uses the pandas library to read multiple CSV files from an S3 input prefix. The script must write the processed output to a different S3 prefix. Which configuration of the ProcessingInput and ProcessingOutput parameters is correct for this scenario?
52A machine learning engineer is using Amazon SageMaker Data Wrangler to prepare a dataset for a regression model. The dataset contains a categorical feature with high cardinality (over 10,000 unique values). The engineer wants to encode this feature efficiently without creating thousands of binary columns, which would explode the feature space. Which encoding technique should the engineer use in Data Wrangler?
53A data scientist is using SageMaker Data Wrangler to prepare a dataset for a binary classification model. The dataset contains a mix of numerical and categorical features. The scientist wants to perform feature engineering to improve model performance. Which TWO actions are appropriate for handling categorical features in Data Wrangler? (Choose two.)
54A data scientist is preparing a CSV dataset in Amazon S3 for a SageMaker training job. Several rows contain missing values in numeric feature columns, and the chosen algorithm cannot handle NaNs. The scientist wants a repeatable, code-based transformation that runs inside a SageMaker Processing job before training. Which step is the MOST appropriate?
55A data engineer is preparing a dataset in Amazon SageMaker Data Wrangler for a binary classification model. The dataset contains missing values in several numeric columns, and the engineer wants a reusable, reproducible transformation that can be applied identically to the training data and to future inference data. The engineer plans to export the transformation and integrate it into a SageMaker Pipeline. Which TWO actions should the engineer take to ensure the imputation is consistent between training and inference? (Choose two.)
56A machine learning engineer is preparing a dataset for training a model on Amazon SageMaker. The dataset contains numerical features with varying scales, and the engineer wants to ensure that all features contribute equally during training. Which data preparation step should the engineer take?
57A data engineer is preparing a dataset for a time series forecasting model. The dataset contains a timestamp column and a target variable. The engineer wants to create additional features such as lag values and rolling averages. Which SageMaker Data Wrangler transform should be used to generate these time series features?
58A data engineer is preparing a large training dataset stored in Amazon S3 as many small Parquet files, and a SageMaker training job that reads directly from S3 is spending most of its time on the input channel rather than on model computation. The engineer needs to improve the input throughput without changing the model code or the training algorithm. Which action should the engineer take?
59A machine learning engineer is using Amazon SageMaker Feature Store to manage features for a fraud detection model. The engineer needs to ensure that the feature group can serve both batch and real-time predictions. The feature group is configured with an online store enabled. Which additional configuration is required to support batch predictions?
60A data engineer is using Amazon SageMaker Data Wrangler to prepare a dataset for a machine learning model. The engineer needs to ensure that the data preparation flow is reproducible and can be integrated into an MLOps pipeline. Which TWO actions should the engineer take to achieve this? (Choose two.)
61A machine learning team stores training data in an Amazon S3 bucket and wants to catalog it so that Amazon Athena and Amazon SageMaker Feature Store can discover the schema. The data is partitioned by year, month, and day in Hive-style prefixes, and new partitions are added daily. A data engineer must ensure new partitions are automatically discoverable without manual intervention. Which solution meets these requirements?
62A machine learning engineer is preparing data for a SageMaker training job and needs to split a large dataset into training, validation, and test sets while avoiding data leakage from the same entity appearing in multiple splits. The dataset contains multiple rows per customer, and the target is customer churn. Which TWO strategies are appropriate? (Choose two.)
63A machine learning engineer needs to split a dataset into training, validation, and test sets for a SageMaker training job. The dataset is stored in Amazon S3 as a single CSV file. The engineer wants to ensure that the splits are reproducible and that the test set is never used during training or hyperparameter tuning. Which approach should the engineer use?
64A data engineer is preparing a dataset for a SageMaker training job. The dataset contains a timestamp column and is stored in Amazon S3 as CSV files. The engineer needs to ensure that the training job reads the data efficiently and that the data is partitioned by date to improve query performance in Amazon Athena. Which action should the engineer take?
65A data scientist has a 200 GB Parquet dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with the SageMaker training toolkit's File mode, and the job currently spends a long time downloading before training begins. The team wants to reduce startup time without changing the training algorithm. Which change should the data scientist make?
66A data engineer stores raw ML training data in Amazon S3 and needs to catalog the schema, track partition changes, and make the data queryable by Amazon Athena without running ETL. Which AWS service should the engineer use?
67A machine learning engineer runs a SageMaker Processing job that must load a 200 GB dataset from S3, compute statistics, and write a small summary to S3. The job repeatedly fails with an out-of-disk-space error on the processing instance. Which change is MOST likely to resolve the failure?
68A data scientist must join a 50 GB transactional table with a small 5 MB lookup table in AWS Glue before writing Parquet output for SageMaker training. The join currently shuffles the large table across the cluster and the job runs slowly. Which optimization should the data scientist apply?
69A data engineer stores a training dataset in Amazon S3 across a prefix hierarchy organized by year, month, and day, and analysts query it with Amazon Athena. The engineer wants Athena to scan only the relevant date range instead of the entire dataset. Which action should the engineer take?
Deep-dive questions
The most-searched questions in this domain — detailed explanations, worked examples, full answer breakdowns.
Be able to select and sequence the right AWS services to ingest, transform, and store training data in S3, then apply the correct SageMaker transform for duplicates, imbalance, or mixed feature types. The most important thing is matching the tool to the source and target format without leaking data across splits.
The Courseiva MLA-C01 question bank contains 69 questions in the Data Preparation for Machine Learning domain, covering the 28% of the exam attributed to this domain in the official Amazon Web Services blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Data Preparation for Machine Learning domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included