MLA-C01 · domain
Data Preparation for Machine Learning
This domain covers getting data into and ready for SageMaker: ingestion from RDS, S3, and other sources; transformation with Data Wrangler, Glue, and Processing jobs; and handling imbalance, duplicates, and mixed feature types. Questions are scenario-based, asking you to pick the right AWS service, format, or built-in transform for a stated constraint.
Focused practice
Practice Data Preparation for Machine Learning questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Data Preparation for Machine Learning
Be able to select and sequence the right AWS services to ingest, transform, and store training data in S3, then apply the correct SageMaker transform for duplicates, imbalance, or mixed feature types. The most important thing is matching the tool to the source and target format without leaking data across splits.
Choosing AWS Glue or SageMaker Processing to transform RDS data into Parquet on S3
Using SageMaker Data Wrangler built-in transforms to detect and remove duplicate rows
Applying SageMaker Processing or imbalanced-learn techniques to rebalance skewed class datasets
Handling mixed numeric and categorical features with encoding, scaling, and SageMaker Feature Store
Watch out for
Common Data Preparation for Machine Learning exam traps
- ▸Assuming Data Wrangler alone moves data from RDS to S3 without a Glue or Processing job
- ▸Forgetting that Parquet conversion requires a schema-aware engine, not a plain S3 copy
- ▸Resampling before the train/test split, leaking synthetic samples into validation data
Question index
All Data Preparation for Machine Learning questions (69)
Click any question to see the full explanation, or start a practice session above.
A machine learning engineer is using Amazon SageMaker Feature Store to manage features for a fraud detection model. The engineer needs to ensure that the feature group can serve both batch and real-time predictions. The feature group is configured with an online store enabled. Which additional configuration is required to support batch predictions?
Medium2A company uses AWS Glue ETL jobs to transform data for machine learning. They have a dataset with a column 'income' that is heavily right-skewed. Which transformation should be applied to make the distribution more Gaussian-like?
Hard3A company has 10 TB of log data in compressed JSON format stored in Amazon S3. The data needs to be processed and transformed into a structured format for machine learning. The processing requires complex transformations, including parsing nested JSON and joining with a reference table. The company wants to minimize infrastructure management. Which approach should the company use?
Easy4A data engineer is preparing a large training dataset stored in Amazon S3 as many small Parquet files, and a SageMaker training job that reads directly from S3 is spending most of its time on the input channel rather than on model computation. The engineer needs to improve the input throughput without changing the model code or the training algorithm. Which action should the engineer take?
Medium5A data scientist is preparing a large dataset (50 GB) for training a TensorFlow model on SageMaker. The dataset consists of many small CSV files. Training is slow due to I/O bottlenecks. Which data preparation strategy most effectively accelerates training?
Hard6A data engineer needs to prepare a large dataset for machine learning. The data is stored in an Amazon RDS MySQL database and needs to be transformed and moved to an S3 bucket in Parquet format for use with SageMaker. Which AWS service is most suitable for this extraction, transformation, and loading (ETL) task?
Easy7A team is preparing text data for a natural language processing (NLP) model. They have a corpus of customer reviews. Which THREE preprocessing steps are essential to reduce noise and improve model performance?
Hard8A data scientist is preparing a dataset stored in Amazon S3 for a SageMaker training job. The dataset contains missing values in several columns. The scientist wants to impute missing values with the mean of each column. Which SageMaker built-in algorithm or processing method should be used to perform this imputation efficiently?
Easy9A data engineer needs to prepare a dataset for a fraud detection model. The dataset contains a highly skewed numerical feature with extreme outliers. The engineer decides to apply a logarithmic transformation to this feature before training. Which SageMaker Data Wrangler transform should be used to apply the logarithmic transformation?
Medium10A team is using Amazon SageMaker for feature engineering. They have a dataset with a column 'TransactionDate' in string format (e.g., '2023-01-15 10:30:00'). They need to create features: year, month, day, hour, and day_of_week. What is the most efficient way to do this in a SageMaker processing job?
Medium11A data engineer is using Amazon SageMaker Processing to run a data preprocessing script on a dataset with 500 million rows. The script runs out of memory on a single ml.r5.24xlarge instance. The engineer needs to modify the processing job to handle the dataset size. Which approach is most cost-effective and scalable?
Hard12A machine learning engineer is using Amazon SageMaker Processing to preprocess a large dataset. The processing job runs a custom Python script that uses the pandas library to read multiple CSV files from an S3 input prefix. The script must write the processed output to a different S3 prefix. Which configuration of the ProcessingInput and ProcessingOutput parameters is correct for this scenario?
Hard13A data engineer must prepare a 4 TB Parquet dataset stored in Amazon S3 for a SageMaker training job that runs on 8 ml.p4d.24xlarge instances. The engineer wants the fastest possible data throughput during training while minimizing per-epoch I/O overhead. The dataset is immutable for the duration of the training run. Which approach BEST meets these requirements?
Medium14A machine learning engineer is preparing a dataset for a binary classification model. The dataset has 10,000 rows and 200 features, with 5% positive class. The engineer suspects class imbalance may affect model performance. Which TWO actions should the engineer take to mitigate imbalance? (Choose 2.)
Medium15A data scientist has a 40 GB CSV dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with pandas, and the scientist wants to reduce both storage cost and training-time I/O without changing the logical schema. Which data preparation action should be taken?
Easy16A social media company is processing a real-time stream of user activity data from Amazon Kinesis Data Streams to train a machine learning model for content recommendation. The raw data includes user ID, timestamp, content ID, interaction type (like, share, comment), and device type. The data scientists need to aggregate features per user over a sliding window of 7 days, including counts of interaction types, unique content IDs engaged, and a moving average of interaction timestamps. The aggregated data will be used to update a user embedding model. The streaming data volume is approximately 500 records per second, and the company uses an AWS Glue streaming ETL job for transformation. However, the Glue job is failing frequently with high latency and checkpoint errors. The team needs a more robust solution to prepare the streaming data features. Which approach should the team take?
Hard17A data engineer is preparing a dataset in Amazon SageMaker Data Wrangler for a binary classification model. The dataset contains missing values in several numeric columns, and the engineer wants a reusable, reproducible transformation that can be applied identically to the training data and to future inference data. The engineer plans to export the transformation and integrate it into a SageMaker Pipeline. Which TWO actions should the engineer take to ensure the imputation is consistent between training and inference? (Choose two.)
Medium18A data scientist has a 200 GB Parquet dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with the SageMaker training toolkit's File mode, and the job currently spends a long time downloading before training begins. The team wants to reduce startup time without changing the training algorithm. Which change should the data scientist make?
Easy19A machine learning engineer runs a SageMaker Processing job that must load a 200 GB dataset from S3, compute statistics, and write a small summary to S3. The job repeatedly fails with an out-of-disk-space error on the processing instance. Which change is MOST likely to resolve the failure?
Hard20A data engineer is building a data pipeline for a machine learning model that requires both structured and unstructured data. The structured data (customer demographics) is in Amazon RDS, and the unstructured data (customer support chat logs) is in Amazon S3 as JSON files. The engineer needs to combine these datasets into a single training dataset stored in S3 in Parquet format. They must also perform feature engineering such as text vectorization on the chat logs. The pipeline should be serverless and cost-effective. Which approach should they use?
Medium21A data scientist needs to split a dataset into training, validation, and test sets. The dataset has a categorical target variable with imbalanced class distribution. Which splitting technique ensures that each subset has a similar proportion of each class?
Medium22A data engineer is preparing a dataset for a SageMaker training job. The dataset contains a timestamp column and is stored in Amazon S3 as CSV files. The engineer needs to ensure that the training job reads the data efficiently and that the data is partitioned by date to improve query performance in Amazon Athena. Which action should the engineer take?
Medium23A data engineer is using AWS Glue to prepare a dataset for ML. The engineer wants to split the dataset into training and testing sets while preserving the distribution of the target variable. Which TWO methods achieve this goal? (Select TWO)
Easy24A data scientist is using SageMaker Data Wrangler to prepare a dataset for a binary classification model. The dataset contains a mix of numerical and categorical features. The scientist wants to perform feature engineering to improve model performance. Which TWO actions are appropriate for handling categorical features in Data Wrangler? (Choose two.)
Medium25A team is building a recommendation system and wants to store and serve features for online and offline models. The features include user statistics (updated daily) and movie metadata (static). The team needs low-latency inference for real-time recommendations and wants to reuse features across multiple models. Which AWS service should the team use to store, manage, and serve these features?
Medium26A company uses Amazon SageMaker Data Wrangler to prepare data for ML. The dataset contains a timestamp column and sensor readings from IoT devices. The data scientist needs to create features such as moving averages and rolling statistics over time windows. Which Data Wrangler transformation type should be selected?
Hard27A data scientist is working with a dataset that contains missing values in several numeric features. The data scientist wants to impute the missing values with the median of each feature. Which Amazon SageMaker Data Wrangler transformation should be used?
Easy28A team is using Amazon SageMaker Processing for data preprocessing. They have a Parquet dataset in Amazon S3. Which configuration will provide the most efficient reading of the dataset during processing?
Medium29A data scientist is preparing a CSV dataset in Amazon S3 for a SageMaker training job. Several rows contain missing values in numeric feature columns, and the chosen algorithm cannot handle NaNs. The scientist wants a repeatable, code-based transformation that runs inside a SageMaker Processing job before training. Which step is the MOST appropriate?
Easy30A machine learning engineer is building a pipeline to preprocess text data for a sentiment analysis model. The data consists of customer reviews. The engineer wants to convert the text into numerical features while preserving the semantic meaning of words. Which technique should be used?
Medium31A data engineer is preparing a dataset for a time series forecasting model. The dataset contains a timestamp column and a target variable. The engineer wants to create additional features such as lag values and rolling averages. Which SageMaker Data Wrangler transform should be used to generate these time series features?
Medium32A data engineer needs to prepare a large dataset (10 TB) stored in Amazon S3 for a training job on SageMaker. The data is in CSV format, but the training algorithm expects Parquet for performance. The engineer must transform the data with minimal cost and without writing custom code. Which service should be used?
Medium33A data scientist is training a binary classifier on a highly imbalanced dataset (1:100 class ratio). The dataset contains 500,000 rows and 30 features. The data is stored in S3 in Parquet format. The data scientist wants to use SageMaker's built-in XGBoost algorithm. Which data preparation technique should the data scientist apply to best address the class imbalance without causing data leakage?
Hard34An organization stores raw data in Amazon S3 as CSV files. They need to perform serverless data transformation and convert the data to Parquet format for efficient ML training. Which AWS service is most appropriate?
Easy35A company runs an online retail business and wants to build a product recommendation system. They have a dataset of customer purchases stored in Amazon S3 as CSV files. The dataset includes columns: 'customer_id', 'product_id', 'purchase_date', 'quantity', 'price', and 'category'. The data science team plans to use Amazon SageMaker to train a factorization machines model. During data exploration, they discover that the 'category' column has 1,200 unique values, and many categories appear only a few times. The 'product_id' column has 50,000 unique values. They want to include both features in the model. The team is concerned about the high cardinality of these features. Which approach should they take to prepare these features for the factorization machines model?
Medium36A company ingests daily log data into an S3 bucket. They need to update the existing ML training dataset with new data without reprocessing the entire history. Which two strategies should they adopt? (Choose two.)
Easy37Refer to the exhibit. A data scientist is trying to use AWS Glue to read data from the S3 bucket `ml-data-bucket`. The Glue job fails with an access denied error. What is the most likely cause?
Easy38A company has a dataset of 2 billion records stored as text files in Amazon S3. The data is partitioned by year and month. The data science team wants to read only the last 6 months of data for model training using SageMaker. To minimize data scanned and reduce costs, which approach should the team use?
Easy39A company is building a real-time inference pipeline for an ML model. The raw data arrives in JSON format via Amazon Kinesis Data Streams. Before invoking the SageMaker endpoint, the data must be preprocessed to match the training data format. Which THREE steps should be included in the preprocessing function? (Select THREE)
Hard40A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset. The dataset contains a column with date strings in the format 'YYYY-MM-DD'. The data scientist wants to extract the year, month, and day as separate features. Which Data Wrangler transform should be used?
Medium41A machine learning engineer is using Amazon SageMaker Data Wrangler to prepare a dataset for a regression model. The dataset contains a categorical feature with high cardinality (over 10,000 unique values). The engineer wants to encode this feature efficiently without creating thousands of binary columns, which would explode the feature space. Which encoding technique should the engineer use in Data Wrangler?
Hard42A machine learning engineer is preparing data for a SageMaker training job and needs to split a large dataset into training, validation, and test sets while avoiding data leakage from the same entity appearing in multiple splits. The dataset contains multiple rows per customer, and the target is customer churn. Which TWO strategies are appropriate? (Choose two.)
Hard43A machine learning engineer is preparing a dataset for training a model on Amazon SageMaker. The dataset contains numerical features with varying scales, and the engineer wants to ensure that all features contribute equally during training. Which data preparation step should the engineer take?
Easy44A data scientist is using SageMaker Data Wrangler to prepare a large dataset. The data contains duplicate rows, which could bias the model. Which built-in step in Data Wrangler can automatically detect and remove duplicates?
Medium45A financial services company is developing a fraud detection model using Amazon SageMaker. They have a dataset with 10 million transactions, each with 300 features. The dataset is highly imbalanced (0.1% fraud). They have performed feature engineering and now need to split the data for training, validation, and test sets. The data is stored in CSV files in Amazon S3. They plan to use SageMaker's built-in XGBoost algorithm. To ensure proper evaluation and avoid data leakage, which data splitting strategy should they use?
Hard46A machine learning team is building a model to predict customer churn. They have historical data that includes customer activity logs, each with a timestamp. The team wants to ensure that the training data does not contain any data leakage from the future. Which approach should they take when preparing the training and validation datasets?
Hard47A retail company is building a machine learning model to predict customer churn. The data engineering team has extracted customer transaction data from Amazon Aurora and stored it as CSV files in Amazon S3. The data includes customer IDs, transaction amounts, timestamps, and product categories. A data scientist discovers that the dataset contains several missing values in the 'transaction_amount' column for about 15% of the records. The data scientist also notices that the 'customer_id' column has some duplicate entries. The team wants to prepare the data for training a churn model using Amazon SageMaker. The data is approximately 50 GB in size. What should the data scientist do to handle the missing values and duplicates efficiently while preparing the data for training?
Easy48A data scientist is preparing a large dataset for training a machine learning model. The dataset contains missing values in several columns. Which approach is the MOST efficient for handling missing values in a large dataset using AWS services?
Easy49An ML engineer needs to split a dataset into training, validation, and test sets. The dataset has a time-based column that should not be leaked. Which split method is most appropriate?
Easy50A company is building a fraud detection model on an imbalanced dataset (99% legitimate, 1% fraudulent). To improve recall on the minority class, they want to resample data. Which combination of techniques should they use?
Medium51A company is building a time series forecasting model using SageMaker DeepAR. The raw data is a CSV with columns: timestamp, item_id, and value. What is the correct data format required for DeepAR training?
Medium52A data scientist is preparing a dataset for a machine learning model that predicts customer churn. The dataset contains a column 'CustomerID' that is a unique identifier. What should the data scientist do with this column before training the model?
Easy53A machine learning engineer needs to split a dataset into training, validation, and test sets for a SageMaker training job. The dataset is stored in Amazon S3 as a single CSV file. The engineer wants to ensure that the splits are reproducible and that the test set is never used during training or hyperparameter tuning. Which approach should the engineer use?
Easy54A machine learning team stores training data in an Amazon S3 bucket and wants to catalog it so that Amazon Athena and Amazon SageMaker Feature Store can discover the schema. The data is partitioned by year, month, and day in Hive-style prefixes, and new partitions are added daily. A data engineer must ensure new partitions are automatically discoverable without manual intervention. Which solution meets these requirements?
Medium55A data engineer stores raw ML training data in Amazon S3 and needs to catalog the schema, track partition changes, and make the data queryable by Amazon Athena without running ETL. Which AWS service should the engineer use?
Easy56A data engineer needs to convert a JSON dataset to Parquet format for efficient querying with Amazon Athena. The JSON files are in an S3 bucket. Which service can perform this conversion with minimal coding?
Easy57Which TWO actions are recommended best practices when preparing training data for a machine learning model in AWS? (Choose two.)
Easy58A data engineer is preparing a large dataset of 10 TB for ML training on Amazon SageMaker. The data is stored in Amazon S3 as CSV files. To reduce training time and cost, the engineer wants to use a columnar format that is optimized for analytical queries. Which format should the engineer convert the data to?
Easy59A data engineer is using Amazon SageMaker Data Wrangler to prepare a dataset for a machine learning model. The engineer needs to ensure that the data preparation flow is reproducible and can be integrated into an MLOps pipeline. Which TWO actions should the engineer take to achieve this? (Choose two.)
Medium60A data engineer is building an AWS Glue ETL job that reads raw JSON clickstream events from Amazon S3, flattens nested structures, and writes Parquet to a curated S3 prefix for SageMaker training. The job must run daily on only the newly arrived files and must keep the Glue Data Catalog table current so Athena and SageMaker can query it. Which TWO configurations should the engineer apply to meet these requirements? (Choose two.)
Hard61A data scientist creates a feature group as shown in the exhibit. When ingesting data with an 'age' column of integer values, the ingestion fails. What is the most likely cause?
Hard62A company collects sensor data from IoT devices. The data arrives with missing timestamps due to network issues. For anomaly detection, the engineer needs to create features that capture rolling statistics over fixed windows. Which data preprocessing step is essential before feature generation?
Medium63A data engineer is using AWS Glue to prepare a dataset for machine learning. The dataset has several columns with outliers. The engineer wants to detect and handle outliers in a scalable manner. Which TWO approaches should the engineer consider? (Select TWO.)
Easy64A data scientist must join a 50 GB transactional table with a small 5 MB lookup table in AWS Glue before writing Parquet output for SageMaker training. The join currently shuffles the large table across the cluster and the job runs slowly. Which optimization should the data scientist apply?
Medium65A data engineer is optimizing Amazon Athena queries on large datasets stored in S3 for machine learning data preparation. Which THREE practices improve query performance?
Hard66A data engineer stores a training dataset in Amazon S3 across a prefix hierarchy organized by year, month, and day, and analysts query it with Amazon Athena. The engineer wants Athena to scan only the relevant date range instead of the entire dataset. Which action should the engineer take?
Easy67A data scientist is preparing a dataset for training a binary classification model. The dataset has 100,000 rows and 50 features. The target variable is imbalanced, with only 5% positive cases. Which technique should the data scientist apply to address the class imbalance BEFORE training?
Easy68A healthcare company is building a model to predict patient readmission rates. The dataset contains a mix of numeric features (age, blood pressure, lab test results) and categorical features (gender, diagnosis code, hospital department). The dataset has 2 million rows. The data is stored in an Amazon S3 bucket, and they use AWS Glue to catalog and preprocess the data. The data scientist notices that the 'diagnosis_code' column has 10,000 unique codes, and 20% of the rows have missing values for 'blood_pressure'. They plan to use a SageMaker built-in XGBoost model. For optimal model performance, which preprocessing steps should they apply using AWS Glue ETL?
Medium69A team is using AWS Glue to process streaming data from Amazon Kinesis. The streaming data contains both structured and semi-structured fields. The team needs to flatten the semi-structured fields into columns for downstream ML training. Which Glue feature is BEST suited?
HardOther domains
All MLA-C01 exam domains
Frequently asked questions
- What does the Data Preparation for Machine Learning domain cover on the MLA-C01 exam?
- Be able to select and sequence the right AWS services to ingest, transform, and store training data in S3, then apply the correct SageMaker transform for duplicates, imbalance, or mixed feature types. The most important thing is matching the tool to the source and target format without leaking data across splits.
- How many questions are in this domain?
- This page lists all 69 Data Preparation for Machine Learning questions in the MLA-C01 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Data Preparation for Machine Learning questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.