Courseiva

CCNA Fundamentals of AI and ML Questions

75 of 90 questions · Page 1/2 · Fundamentals of AI and ML · Answers revealed

1
MCQeasy

A hospital wants to build a model that predicts whether a patient has a specific disease based on labeled historical medical records where each record is marked either positive or negative. Which type of machine learning problem does this represent?

A.Reinforcement learning using reward signals
B.Supervised learning using regression
C.Unsupervised learning using clustering
D.Supervised learning using classification
AnswerD

The records carry known labels of positive or negative, which is exactly the supervision signal classification uses to learn a decision boundary. Because the target is a discrete category rather than a continuous number, the task is classification. The model can then predict the disease status for new patients, matching the hospital's goal.

Why this answer

Because each historical record already includes a known positive or negative label and the target is a discrete category, the task is supervised classification. Regression would predict a continuous value, clustering ignores the labels, and reinforcement learning requires rewards from interaction. Classification directly models the disease outcome the hospital wants to predict.

Exam trap

The trap here is confusing labeled binary prediction with regression simply because both are supervised learning tasks.

2
MCQhard

A data scientist trains a model to predict whether a loan applicant will default. After deployment, the model performs well on applicants similar to the training data but poorly on applicants from a newly added geographic region that was underrepresented in training. Which statement best describes the underlying problem?

A.The training data does not represent the new region, so the model cannot generalize to that subpopulation.
B.The model has overfit the training set because it memorized noise in the original regions.
C.The model requires more training epochs to converge on the new region's data.
D.The evaluation metric used is inappropriate and should be replaced with accuracy.
AnswerA

When a subpopulation is underrepresented or absent in training, the model learns patterns that do not transfer to it, producing poor predictions for that group. This is a data coverage and distribution shift problem, and it matches the observed failure on the newly added geographic region exactly.

Why this answer

The model fails specifically on a subpopulation that was underrepresented in training, which is a data coverage and distribution shift problem. Good performance on familiar applicants confirms the algorithm works, but without representative training examples from the new region the model cannot generalize there. The remedy is additional representative data or techniques that address distribution shift.

Exam trap

The trap here is labeling any train-versus-deployment gap as overfitting, when the failure is isolated to a subpopulation missing from the training distribution.

3
MCQhard

A company is building a model to detect fraudulent transactions. The dataset has 1,000,000 transactions, of which only 1,000 are fraudulent. The team wants to evaluate the model's performance. Which metric is most appropriate to use as the primary evaluation metric?

A.F1 score
B.Recall
C.Accuracy
D.Precision
AnswerA

F1 score is the harmonic mean of precision and recall, providing a balance between the two. In fraud detection, both false positives and false negatives have costs, so a balanced metric is essential. F1 score is robust to class imbalance and is the most appropriate primary metric here.

Why this answer

For highly imbalanced datasets like fraud detection, accuracy is misleading. The F1 score balances precision and recall, making it suitable when both false positives and false negatives matter. It provides a single metric that reflects the model's ability to correctly identify fraud without being overwhelmed by the majority class.

Exam trap

The trap here is choosing accuracy because it is commonly used, but it is deceptive when classes are heavily imbalanced.

4
MCQhard

A machine learning team notices their model performs excellently on the training dataset but poorly on new, unseen data. They want to reduce this gap without collecting more data. Which action most directly addresses the problem?

A.Apply regularization techniques such as L2 penalty or dropout
B.Increase model complexity by adding more layers and parameters
C.Train for many more epochs until training loss approaches zero
D.Remove the validation dataset and evaluate only on training data
AnswerA

Regularization constrains the model so it cannot fit training noise as easily, which improves generalization to unseen data. L2 penalties shrink weights and dropout randomly deactivates units during training, both discouraging memorization. This directly targets the overfitting gap described without requiring additional data collection.

Why this answer

The described pattern is classic overfitting: strong training performance with weak performance on unseen data. Regularization such as L2 penalties or dropout constrains the model and improves generalization. Adding capacity, training longer, or discarding validation all worsen or conceal the issue instead of correcting it.

Exam trap

The trap here is assuming that a model performing better on training data will automatically perform better on new data.

5
MCQhard

A financial services firm wants an internal assistant that answers employee questions about company travel policy in natural language. The policy documents change frequently, and the firm requires answers to cite the specific policy section used. Which approach BEST meets these requirements?

A.Deploy a keyword search engine that returns matching policy paragraphs and rely on employees to interpret the results themselves.
B.Fine-tune a large language model on the full policy corpus so it memorizes the documents and answers from memory.
C.Use retrieval-augmented generation, where relevant policy passages are retrieved and supplied to the model to ground its answer and source citation.
D.Train a small classification model that maps each employee question to one of the fixed policy section titles.
AnswerC

Retrieval-augmented generation fetches the most relevant passages from the current policy store at query time and passes them to the language model as context. The model then answers based on those passages and can reference the specific section, satisfying the citation requirement. Because the knowledge lives in the document store, updating policies requires only re-indexing the changed documents, not retraining.

Why this answer

Retrieval-augmented generation grounds each answer in freshly retrieved policy passages, so responses reflect the current documents and can point to the exact section used. Because updates only require re-indexing changed content rather than retraining a model, it also handles the frequently changing policy corpus efficiently.

Exam trap

The trap here is assuming that fine-tuning a model on documents is the way to make it answer about those documents, when grounding through retrieval is what enables current content and citations.

6
MCQmedium

A data scientist is using SageMaker to train a model on a dataset with many features. They suspect some features are redundant. Which feature engineering technique would help?

A.Feature scaling
B.One-hot encoding
C.Principal Component Analysis (PCA)
D.Polynomial features
AnswerC

Principal Component Analysis projects the many correlated features onto a smaller set of orthogonal components capturing most variance, eliminating redundancy. This directly addresses the stem's suspicion of redundant features by reducing dimensionality while retaining information.

Why this answer

Principal Component Analysis (PCA) is a dimensionality reduction technique that transforms the original correlated features into a smaller set of uncorrelated principal components, effectively removing redundancy while preserving most of the variance in the data. In SageMaker, PCA can be applied via the built-in PCA algorithm or as a preprocessing step in a scikit-learn container to reduce feature space and eliminate multicollinearity.

Exam trap

The AIF-C01 exam often tests the distinction between feature reduction (PCA) and feature transformation (scaling, encoding, polynomial expansion) to see if candidates confuse techniques that change feature count versus those that only change feature values.

How to eliminate wrong answers

Option A is wrong because feature scaling (e.g., StandardScaler, MinMaxScaler) normalizes the range of features but does not remove redundant or correlated features; it only changes the scale. Option B is wrong because one-hot encoding is used to convert categorical variables into numerical format, not to address feature redundancy among many continuous or numerical features. Option D is wrong because polynomial features create interaction and higher-order terms, which actually increase the number of features and can introduce more redundancy, not reduce it.

7
MCQmedium

A data science team needs to choose a machine learning approach for a project that requires predicting customer churn based on historical data. The team has a labeled dataset with 10,000 records and needs to interpret the model's decisions to provide business insights. Which machine learning technique should the team prioritize?

A.Random forest.
B.K-means clustering.
C.Linear regression.
D.Deep neural network with multiple hidden layers.
AnswerA

Random forest is supervised and handles labelled tabular data well, and its ensemble of decision trees exposes feature importance values. Those importances give the interpretable business insight the team needs, unlike opaque deep learning or unsupervised clustering.

Why this answer

Random forest is the best choice because it handles the classification task of predicting churn (a binary outcome) from a labeled dataset, provides feature importance scores for interpretability, and works well with 10,000 records without overfitting due to its ensemble of decision trees. Its built-in ability to rank input features directly supports the team's need to derive business insights from model decisions.

Exam trap

The AWS AI Practitioner exam often tests the distinction between supervised and unsupervised learning, and the trap here is that candidates might choose a powerful but opaque model like a deep neural network, overlooking the explicit requirement for interpretability and the modest dataset size that favors simpler, more explainable ensemble methods.

How to eliminate wrong answers

Option B (K-means clustering) is wrong because it is an unsupervised learning algorithm used for grouping unlabeled data, not for predicting a labeled target like churn. Option C (Linear regression) is wrong because it is designed for regression tasks predicting continuous values, not for binary classification like churn prediction. Option D (Deep neural network with multiple hidden layers) is wrong because, while it can handle classification, it requires large datasets (typically >100,000 records) to avoid overfitting and offers poor interpretability, making it unsuitable for deriving clear business insights from a 10,000-record dataset.

8
MCQmedium

A financial institution wants to predict whether a loan applicant will default. They have a historical dataset with loan outcomes (default or no default) and various applicant features. Which type of machine learning should they use?

A.Reinforcement learning
B.Unsupervised learning
C.Clustering
D.Supervised learning
AnswerD

Supervised learning uses labeled data to train a model that maps input features to a target output. Here, the target is binary (default or no default), and historical labeled data is available. This allows training a classification model to predict default for new applicants, making supervised learning the correct choice.

Why this answer

Supervised learning is ideal when historical labeled data is available and the goal is to predict a known outcome. The loan default prediction task is a binary classification problem, which supervised learning handles by learning from past examples to classify new applicants.

Exam trap

The trap here is confusing clustering with classification because both group data, but clustering is unsupervised and cannot predict a labeled outcome.

9
MCQmedium

An ML team is deploying a real-time inference endpoint for a computer vision model using Amazon SageMaker. The model requires GPU acceleration for low latency. Which instance type should the team choose to minimize cost while meeting the GPU requirement?

A.ml.g5.xlarge
B.ml.c5.xlarge
C.ml.p3.2xlarge
D.ml.p4d.24xlarge
AnswerA

ml.g5.xlarge pairs an NVIDIA A10G GPU with four vCPUs, providing the required GPU acceleration at the lowest cost within the G5 family. Smaller GPU instances lack sufficient acceleration, while larger G5 sizes exceed the stated requirement.

Why this answer

(ml.g5.xlarge) is correct because it provides a GPU (NVIDIA A10G Tensor Core GPU) necessary for low-latency GPU acceleration in computer vision inference, while being the most cost-effective GPU instance among the options. The ml.g5.xlarge offers sufficient GPU compute for real-time inference at a lower hourly cost compared to ml.p3.2xlarge, making it the optimal choice for minimizing cost while meeting the GPU requirement.

Exam trap

Candidates may be tempted to choose ml.p3.2xlarge due to its reputation for ML training, but for inference, ml.g5 instances often provide better price-performance. The ml.p3.2xlarge, while GPU-equipped, is more expensive per hour and may be over-provisioned for typical inference workloads, leading to unnecessary costs.

How to eliminate wrong answers

Option A (ml.g5.xlarge) is wrong because it uses a GPU (NVIDIA A10G) that is more powerful and expensive than needed for this use case, leading to higher cost; however, it does meet the GPU requirement, so the primary issue is cost inefficiency, not technical incompatibility. Option B (ml.c5.xlarge) is wrong because it is a CPU-only instance (based on Intel Xeon Scalable processors) and lacks any GPU, failing to meet the explicit GPU acceleration requirement for low-latency computer vision inference. Option D (ml.p4d.24xlarge) is wrong because it provides 8 NVIDIA A100 GPUs, which is massively over-provisioned for a single real-time inference endpoint, resulting in significantly higher cost without any benefit for this workload.

10
Multi-Selecthard

Which TWO are best practices for model monitoring in production on AWS?

Select 2 answers
A.Disable logging to reduce latency
B.Use only CPU instances
C.Monitor input data drift
D.Retrain model daily
E.Monitor prediction drift
AnswersC, E

Input data drift monitoring detects shifts between production inference data and the training distribution, such as changing customer demographics or feature ranges. This satisfies the best-practice requirement by triggering alerts or retraining before degraded inputs silently corrupt prediction quality in the deployed AWS model.

Why this answer

Option C (Monitor input data drift) is correct because production models degrade when the statistical distribution of incoming features diverges from training data, so tracking data drift with tools like SageMaker Model Monitor detects this covariate shift early. Option E (Monitor prediction drift) is correct because shifts in the distribution of model outputs (e.g., changing class proportions or score distributions) signal concept drift or upstream data issues that require investigation or retraining. Together, input and prediction drift monitoring are the two core best practices for maintaining model quality in production.

Option A is wrong because disabling logging removes the observability needed to detect drift, errors, and bias, and logging overhead is typically negligible relative to inference. Option B is wrong because instance type (CPU vs. GPU) is a performance/cost choice, not a monitoring best practice.

Option D is wrong because retraining daily is not a best practice by itself; retraining should be triggered by monitored drift or performance degradation, and daily retraining can be costly and destabilizing.

Exam trap

The AIF-C01 exam often tests the misconception that retraining on a fixed schedule (e.g., daily) is a best practice, when in reality it should be event-driven based on drift or performance metrics.

11
Multi-Selectmedium

A data science team is preparing a dataset for training a machine learning model. They need to perform data preprocessing to improve model performance. Which TWO of the following are common data preprocessing techniques? (Choose two.)

Select 2 answers
A.One-hot encoding
B.Feature selection
C.Cross-validation
D.Normalization
E.Gradient descent
AnswersA, D

One-hot encoding converts categorical variables into a binary vector representation, allowing machine learning algorithms to handle non-numeric data. It is a fundamental preprocessing step for categorical features, ensuring they can be used in models that require numerical input. This is a common technique.

Why this answer

Normalization and one-hot encoding are standard preprocessing techniques. Normalization scales numerical features, while one-hot encoding transforms categorical variables into a numerical format. Both are applied to the data before training to ensure the model can effectively learn from all features.

Exam trap

The trap here is confusing model training techniques like gradient descent or evaluation methods like cross-validation with data preprocessing steps.

12
MCQeasy

Refer to the exhibit. A data scientist ran a training job on Amazon SageMaker. The job failed with the error shown. What is the most likely cause?

A.The S3 input path is incorrect
B.The IAM role does not have permission to access S3
C.The training code has a syntax error
D.The batch size is too large for the instance's GPU memory
AnswerD

A batch size exceeding GPU memory triggers an out-of-memory failure during the forward or backward pass, since activations for the whole batch must be held simultaneously. Reducing the batch size, or using gradient accumulation, directly addresses this constraint rather than altering the model architecture.

Why this answer

The error message indicates a CUDA out-of-memory error, which occurs when the GPU memory is insufficient for the requested batch size. Option D is correct because increasing the batch size beyond the GPU's memory capacity causes the training job to fail with this specific error.

Exam trap

AWS often tests the distinction between infrastructure errors (S3, IAM) and runtime errors (CUDA memory), where candidates mistakenly attribute a GPU memory error to a misconfiguration in data access or code syntax.

How to eliminate wrong answers

Option A is wrong because an incorrect S3 input path would result in a 'NoSuchKey' or '404' error, not a CUDA out-of-memory error. Option B is wrong because an IAM role lacking S3 permissions would produce an 'AccessDenied' error, not a GPU memory error. Option C is wrong because a syntax error in the training code would raise a Python exception (e.g., SyntaxError) before any GPU operations, not a CUDA memory error.

13
Multi-Selecteasy

A data scientist wants to deploy a custom model built with TensorFlow to Amazon SageMaker for real-time inference. Which TWO steps are required? (Choose two.)

Select 2 answers
A.Create an Amazon ECR repository for the inference container
B.Upload the model artifacts to an S3 bucket
C.Submit a training job to SageMaker
D.Create a SageMaker endpoint configuration
E.Convert the model to ONNX format
AnswersB, D

SageMaker real-time endpoints pull model artefacts from Amazon S3, so the trained TensorFlow files must reside in a bucket the execution role can read. This satisfies the stem's deployment requirement: without artefacts in S3, CreateModel cannot locate the model data, and endpoint creation fails.

Why this answer

Option B is correct because SageMaker requires model artifacts (the trained TensorFlow SavedModel or model.tar.gz) to be stored in an Amazon S3 bucket, which is then referenced by the SageMaker model when deploying. Option D is correct because deploying to a real-time endpoint requires creating an endpoint configuration that specifies the production variant, instance type, and initial instance count, which is then used to create the endpoint. Option A is not required because the TensorFlow inference container is already provided and maintained by SageMaker as a prebuilt Docker image, so no custom ECR repository is needed unless using a custom container.

Option C is not required because the model is already built with TensorFlow; a SageMaker training job is only needed if training within SageMaker. Option E is not required because SageMaker's TensorFlow container supports native TensorFlow SavedModel format, so ONNX conversion is unnecessary.

Exam trap

The trap here is that candidates often think they must build a custom container (Option A) or convert the model (Option E), but SageMaker's pre-built TensorFlow containers eliminate those steps, and the key requirements are simply uploading artifacts to S3 and creating the endpoint configuration.

14
MCQeasy

A startup wants to build a product recommendation engine for their e-commerce platform. They have user purchase history and item metadata. They want a fully managed solution that can automatically train and deploy a recommendation model without needing to manage the underlying ML lifecycle. The solution should provide personalized recommendations based on collaborative filtering. Which AWS service should they use?

A.Use Amazon Kendra
B.Use Amazon Lex
C.Use Amazon Personalize
D.Use Amazon SageMaker built-in Factorization Machines algorithm
AnswerC

Amazon Personalize is fully managed, handling training, tuning and deployment of collaborative-filtering recommenders without ML lifecycle management. It ingests purchase history and item metadata to produce personalised recommendations, meeting the startup's requirement for a managed solution requiring no underlying ML operations.

Why this answer

Amazon Personalize is a fully managed service that enables you to build and deploy recommendation models without managing the underlying ML lifecycle. It supports collaborative filtering out of the box, using user purchase history and item metadata to generate personalized recommendations, which directly matches the startup's requirements.

Exam trap

The trap here is that candidates may confuse Amazon SageMaker's built-in algorithms (like Factorization Machines) with a fully managed recommendation service, overlooking the requirement for automatic lifecycle management and instead focusing only on the algorithm capability.

How to eliminate wrong answers

Option A is wrong because Amazon Kendra is an intelligent search service that uses natural language processing to answer questions and retrieve documents, not a recommendation engine for collaborative filtering. Option B is wrong because Amazon Lex is a service for building conversational interfaces (chatbots) using speech and text, not for generating product recommendations. Option D is wrong because while Amazon SageMaker's built-in Factorization Machines algorithm can perform collaborative filtering, it requires you to manage the ML lifecycle (data preparation, training, deployment, scaling), which contradicts the requirement for a fully managed solution that automatically handles these tasks.

15
MCQmedium

A data scientist is training a binary classification model to predict customer churn. The dataset has 10,000 records with 9,500 non-churners and 500 churners. After training a logistic regression model, the model achieves 95% accuracy on the test set. However, the business team reports that the model is not useful because it predicts almost all customers as non-churners. Which metric should the data scientist use to evaluate the model's performance in this scenario?

A.Accuracy
B.R-squared
C.Precision
D.Recall
AnswerD

Recall measures the proportion of actual churners correctly identified, directly exposing the model's failure on the 500-positive minority class. With 95% accuracy achievable by predicting all non-churners, recall reveals that sensitivity to churn is near zero, satisfying the need for a metric robust to this class imbalance.

Why this answer

(Recall) is correct because in this highly imbalanced dataset (95% non-churners vs 5% churners), the model's 95% accuracy is misleading—it can achieve this by simply predicting the majority class (non-churner) for all samples. Recall measures the proportion of actual churners correctly identified (True Positives / (True Positives + False Negatives)), directly addressing the business need to detect churn. A high recall ensures the model captures most churners, even at the cost of some false positives.

Exam trap

The AIF-C01 exam often tests the misconception that high accuracy always indicates a good model, especially in imbalanced datasets, leading candidates to overlook metrics like recall or precision that better reflect model utility for the specific business problem.

How to eliminate wrong answers

Option A is wrong because accuracy is a poor metric for imbalanced datasets; a model that predicts all samples as the majority class can achieve high accuracy (95% here) while failing to identify any churners, making it useless for the business goal. Option B is wrong because R-squared is a metric for regression models, measuring the proportion of variance explained by the independent variables, and is not applicable to binary classification tasks like churn prediction. Option C is wrong because precision (True Positives / (True Positives + False Positives)) focuses on the correctness of positive predictions; while important, it does not capture the model's ability to find all churners—a model with high precision but low recall might still miss most churners, which is the core issue reported by the business team.

16
Multi-Selecthard

A fintech startup is preparing its first machine learning project to detect fraudulent card transactions. The team must decide which characteristics make a problem well suited to supervised learning. Which TWO characteristics indicate that supervised learning is appropriate? (Choose two.)

Select 2 answers
A.The team wants the algorithm to discover hidden structure without any predefined output.
B.Historical transactions exist that were confirmed as fraudulent or legitimate by investigators.
C.The team prefers a model that explains each prediction with feature importance values.
D.The goal is to predict a discrete outcome for each new transaction.
E.The dataset contains millions of unlabeled transactions with no investigator feedback.
AnswersB, D

Confirmed historical outcomes are exactly the labeled data supervised learning requires. Each transaction carries a target the model can learn from, allowing it to map transaction features to a fraud or legitimate decision. Without such labels, a supervised classifier could not be trained, so this characteristic is a defining indicator that the approach fits.

Why this answer

Supervised learning needs labeled examples and a defined target to predict. Confirmed fraudulent or legitimate transactions supply those labels, and a discrete fraud-or-not decision supplies a suitable classification target. Unlabeled data, label-free exploration goals, and explainability preferences do not establish the labeled input-output pairing that supervised learning fundamentally depends on.

Exam trap

The trap here is assuming that a large transaction volume or a desire for explainability alone makes a problem supervised, when labeled outcomes and a defined target are what actually matter.

17
MCQhard

A media company wants to build a system that generates short promotional descriptions for its articles. The team has no labeled dataset of article-summary pairs but has a large corpus of published articles and descriptions. They want to leverage a pretrained foundation model and adapt it to their domain with minimal labeling effort. Which approach best fits this scenario?

A.Use reinforcement learning with reader click rewards to train a description generator from scratch.
B.Fine-tune a pretrained foundation model on the company's article and description corpus for text generation.
C.Train a supervised classification model that assigns each article to a predefined topic category.
D.Apply unsupervised clustering to group articles by topic and use cluster IDs as descriptions.
AnswerB

Foundation models are pretrained on broad text and can be adapted to a specific domain through fine-tuning on in-domain examples. The company's corpus of articles paired with descriptions provides exactly the material to steer the model toward generating promotional text in its own style. This leverages pretrained capability while requiring far less labeling than training a model from scratch.

Why this answer

The company needs generated text and already holds a domain corpus of articles paired with descriptions. Fine-tuning a pretrained foundation model adapts broad language ability to the company's style with modest additional data, which matches the stated preference for minimal labeling. Classification, clustering, and from-scratch reinforcement learning either produce the wrong output type or discard the pretrained model advantage.

Exam trap

The trap here is assuming any use of a foundation model must be prompt-only, overlooking that fine-tuning on in-domain pairs is a valid adaptation path.

18
MCQmedium

A data scientist is training a model using Amazon SageMaker and notices the training loss is decreasing but validation loss starts increasing after a few epochs. Which technique should they apply to address this?

A.Increase batch size
B.Increase the learning rate
C.Add more training data
D.Add regularization (e.g., L1 or L2)
AnswerD

Rising validation loss while training loss falls signals overfitting. L1 or L2 regularization penalises large weights, constraining model complexity so it generalises better to unseen data, directly addressing the divergence between the two loss curves.

Why this answer

The scenario describes overfitting, where the model memorizes training data but fails to generalize to validation data. Adding regularization (L1 or L2) penalizes large weights, reducing model complexity and improving generalization. This is a standard technique in SageMaker training jobs, often configured via the `regularizer` hyperparameter in frameworks like TensorFlow or MXNet.

Exam trap

The trap here is that candidates confuse overfitting with underfitting or optimization issues, and incorrectly choose to increase learning rate or batch size, not recognizing that rising validation loss with falling training loss is the classic signature of overfitting.

How to eliminate wrong answers

Option A is wrong because increasing batch size typically stabilizes gradient estimates but does not directly address overfitting; it may even reduce generalization by sharpening minima. Option B is wrong because increasing the learning rate can cause divergence or overshooting of the loss minimum, worsening both training and validation loss. Option C is wrong because adding more training data can help generalization but is not a direct fix for overfitting when validation loss increases; it may not be feasible or sufficient, and regularization is the immediate corrective action.

19
MCQhard

An e-commerce company stores user interaction logs in Amazon S3. They want to use machine learning to segment users based on purchasing behavior. Which unsupervised learning algorithm is most appropriate?

A.Linear regression
B.Random forest
C.K-means clustering
D.Neural network
AnswerC

K-means clustering partitions unlabelled interaction logs into k groups by minimising within-cluster variance, directly satisfying the requirement to segment users by purchasing behaviour without predefined labels. Unlike supervised methods, it needs no target variable, making it appropriate for discovering behavioural cohorts in the S3-stored data.

Why this answer

K-means clustering is the most appropriate unsupervised learning algorithm for segmenting users based on purchasing behavior because it groups data points into clusters based on feature similarity without requiring labeled training data. The e-commerce scenario involves discovering natural groupings (segments) in user interaction logs, which is a classic clustering task, and K-means efficiently partitions users into K distinct segments by minimizing within-cluster variance.

Exam trap

The AIF-C01 exam often tests the distinction between supervised and unsupervised learning by presenting a clustering problem and including supervised algorithms as distractors, leading candidates to mistakenly pick a familiar algorithm like random forest or linear regression without recognizing the lack of labeled data.

How to eliminate wrong answers

Option A is wrong because linear regression is a supervised learning algorithm used for predicting continuous numeric values (e.g., sales amount) from labeled data, not for discovering unlabeled user segments. Option B is wrong because random forest is a supervised ensemble learning method used for classification or regression on labeled datasets, and it cannot perform unsupervised segmentation without target labels. Option D is wrong because neural networks are typically used in supervised or reinforcement learning contexts; while they can be adapted for unsupervised tasks (e.g., autoencoders), they are not the most straightforward or appropriate choice for simple user segmentation compared to K-means clustering.

20
MCQmedium

A media company has millions of customer support emails but no labels indicating topic or sentiment. The company wants to discover natural groupings of emails and reduce dimensionality before further analysis. Which approach should they use?

A.Use reinforcement learning to reward correct email topic assignments
B.Perform regression on email length to predict customer churn
C.Apply unsupervised learning techniques such as clustering and dimensionality reduction
D.Train a supervised classifier on the emails using sentiment labels
AnswerC

With no labels, unsupervised learning is the appropriate paradigm. Clustering algorithms can group emails by textual similarity, and dimensionality reduction techniques such as principal component analysis or t-SNE can compress the feature space for visualization and downstream analysis. This directly satisfies the goal of discovering groupings without predefined categories.

Why this answer

The absence of labels rules out supervised methods and points to unsupervised learning. Clustering reveals natural groupings among the emails, while dimensionality reduction techniques help compress and visualize the feature space. Together they satisfy both stated goals without requiring any annotation effort.

Exam trap

The trap here is reaching for a familiar supervised classifier even though the data has no labels to train on.

21
MCQmedium

A financial services company needs to ensure that the machine learning models used for loan approval are explainable and meet regulatory compliance. Which AWS feature can help explain model predictions?

A.SageMaker Ground Truth
B.SageMaker Clarify
C.SageMaker Automatic Model Tuning
D.SageMaker Model Monitor
AnswerB

SageMaker Clarify computes feature attribution values, such as SHAP, showing how much each input contributed to an individual prediction. This satisfies the explainability and regulatory compliance constraint by documenting the reasoning behind each loan approval decision.

Why this answer

SageMaker Clarify is the correct AWS service for explaining model predictions because it provides feature attribution and bias detection capabilities. It uses SHAP (SHapley Additive exPlanations) to generate explainability reports, which are essential for meeting regulatory compliance in financial services like loan approval.

Exam trap

The trap here is confusing monitoring (Model Monitor) with explainability (Clarify), as both relate to model governance but serve fundamentally different purposes—monitoring tracks performance over time, while Clarify explains individual predictions.

How to eliminate wrong answers

Option A is wrong because SageMaker Ground Truth is a data labeling service for creating training datasets, not for explaining model predictions. Option C is wrong because SageMaker Automatic Model Tuning (hyperparameter optimization) adjusts model parameters to improve performance, but does not provide explainability or feature attribution. Option D is wrong because SageMaker Model Monitor detects data drift and model quality degradation over time, but does not generate explanations for individual predictions.

22
MCQeasy

A media company stores 40 TB of raw video footage in Amazon S3 and wants to automatically detect scene boundaries, identify on-screen text, and flag unsafe frames without building custom computer vision models. Which AWS service should they use?

A.Amazon Comprehend
B.Amazon Polly
C.Amazon Rekognition Video
D.Amazon Transcribe
AnswerC

Amazon Rekognition Video is a managed computer vision service that analyzes stored and streaming video for scene detection, text detection, content moderation, and activity recognition, so the media company can process S3-hosted footage without training any models. It exposes these capabilities through a single API call, which matches the requirement to detect scene boundaries, on-screen text, and unsafe frames with no custom ML development.

Why this answer

Amazon Rekognition Video is purpose-built for analyzing video stored in S3 or streamed in real time, offering scene detection, text detection, content moderation, and activity recognition through managed APIs. Because the company wants those visual capabilities without training custom models, the managed video analysis service is the correct fit. The other services address speech, text, or audio generation rather than visual video understanding.

Exam trap

The trap here is assuming that a general-purpose language or speech service can analyze visual video content.

23
MCQmedium

A company wants to build a model to forecast monthly sales. The data is a time series with trend and seasonality. Which SageMaker algorithm is most appropriate?

A.XGBoost
B.K-Means
C.Linear Learner
D.DeepAR
AnswerD

DeepAR is a supervised recurrent neural network designed for time series forecasting. It learns from many related series and natively models trend, seasonality and uncertainty, directly satisfying the stem's requirement for monthly sales data exhibiting both trend and seasonality.

Why this answer

DeepAR is the most appropriate algorithm because it is specifically designed for time series forecasting, handling both trend and seasonality through autoregressive recurrent neural networks. It learns from multiple related time series and produces probabilistic forecasts, making it ideal for monthly sales prediction.

Exam trap

The trap here is that candidates often choose XGBoost or Linear Learner because they are familiar with regression tasks, but fail to recognize that time series forecasting requires algorithms that explicitly model temporal dependencies and seasonality, which DeepAR is built for.

How to eliminate wrong answers

Option A is wrong because XGBoost is a gradient boosting algorithm for tabular data, not designed to capture temporal dependencies or seasonality in time series without extensive feature engineering. Option B is wrong because K-Means is an unsupervised clustering algorithm that groups data points by similarity, with no capability for forecasting sequential data. Option C is wrong because Linear Learner is a linear regression model that assumes independence of observations and cannot model complex time series patterns like seasonality or long-term trends.

24
MCQmedium

A team is training a binary classification model using Amazon SageMaker. They notice that the training accuracy is 99% but the test accuracy is only 70%. Which technique should they apply first to address this?

A.Reduce training data
B.Apply regularization
C.Increase learning rate
D.Increase model complexity
AnswerB

Regularization adds penalty for large weights, helping to reduce overfitting.

Why this answer

The high training accuracy (99%) paired with significantly lower test accuracy (70%) is a classic symptom of overfitting, where the model memorizes the training data instead of learning generalizable patterns. Regularization (Option B) is the first-line technique to combat overfitting by adding a penalty to the loss function (e.g., L1 or L2 regularization), which discourages overly complex decision boundaries. In Amazon SageMaker, this can be implemented via hyperparameters like `l1` or `l2` in built-in algorithms or by adding dropout layers in a custom framework.

Exam trap

AWS often tests the misconception that overfitting is solved by increasing data or model complexity, when in fact the first step should be regularization to penalize overly complex models.

How to eliminate wrong answers

Option A is wrong because reducing training data would worsen overfitting by providing the model with fewer examples to learn from, making it even more prone to memorization. Option C is wrong because increasing the learning rate can cause the model to overshoot optimal weights during training, leading to divergence or poor convergence, but it does not directly address the variance problem of overfitting. Option D is wrong because increasing model complexity (e.g., adding more layers or parameters) would exacerbate overfitting by giving the model more capacity to memorize noise in the training data.

25
MCQmedium

A company wants to use Amazon SageMaker to train a model using a custom Docker container that has specific dependencies. The training code is stored in an S3 bucket. Which steps must be taken to run the training job?

A.Install dependencies via SageMaker's lifecycle configuration instead of a custom container
B.Push the custom container to Amazon ECR and create a training job with the container URI
C.Use SageMaker's built-in framework container and override the entry point
D.Upload the container to S3 and reference it in the training job
AnswerB

Pushing the image to Amazon ECR gives SageMaker a registry URI it can pull from, satisfying the custom-dependency requirement. The training job then references that image URI alongside the S3 code location, so SageMaker runs your container rather than a built-in algorithm image.

Why this answer

Amazon SageMaker requires custom Docker containers to be stored in Amazon Elastic Container Registry (ECR) to run training jobs. The container URI from ECR is specified in the `AlgorithmSpecification` parameter of the `CreateTrainingJob` API call, allowing SageMaker to pull and execute the container with the training code from S3. Option B correctly describes this mandatory workflow.

Exam trap

AWS often tests the misconception that any S3-uploaded artifact (including Docker images) can be directly referenced in a training job, but SageMaker strictly requires container images to be stored in ECR, not S3.

How to eliminate wrong answers

Option A is wrong because lifecycle configurations are used to customize notebook instances (e.g., install packages on Jupyter kernels), not to provide dependencies for training jobs; training jobs run in ephemeral containers that do not use lifecycle configurations. Option C is wrong because overriding the entry point of a built-in framework container only works if the container already includes the required dependencies; if custom dependencies are needed, a custom container must be built and pushed to ECR. Option D is wrong because SageMaker does not accept Docker containers stored in S3; containers must be registered in ECR and referenced by their URI.

26
Multi-Selectmedium

A data scientist is evaluating different AWS services for building a machine learning pipeline. Which THREE components are part of Amazon SageMaker? (Select THREE.)

Select 3 answers
A.AWS Glue
B.Notebook instances
C.Ground Truth
D.Model registry
E.Amazon Athena
AnswersB, C, D

Notebook instances are a core SageMaker component, providing managed Jupyter environments for data exploration, preprocessing and model development within the pipeline. They satisfy the stem's requirement to identify genuine SageMaker features, unlike unrelated AWS analytics or storage services that lack integrated notebook tooling.

Why this answer

Amazon SageMaker is a fully managed ML platform, and three of the listed items are native SageMaker capabilities. B (Notebook instances) is correct because SageMaker provides managed Jupyter notebook instances for data exploration, preprocessing, and model development. C (Ground Truth) is correct because SageMaker Ground Truth is the built-in data labeling service used to create high-quality training datasets, including with automated labeling and human review workflows.

D (Model registry) is correct because SageMaker includes a model registry for cataloging, versioning, and managing trained models through approval and deployment stages. A (AWS Glue) is a separate ETL and data catalog service, and E (Amazon Athena) is a separate serverless query service for S3 data, so neither is a component of SageMaker.

Exam trap

The trap here is that candidates often confuse AWS Glue (a separate ETL service) as part of SageMaker because both are used in ML pipelines, but Glue is not a SageMaker component.

27
MCQeasy

A company is using Amazon Comprehend for sentiment analysis on customer reviews. They notice that the sentiment is often incorrect for negative reviews with sarcasm. What is the likely cause?

A.The model is not fine-tuned for the domain
B.The pre-trained model cannot handle sarcasm well
C.Insufficient training data
D.The input text is too long
AnswerB

Amazon Comprehend's pre-trained sentiment model learns patterns from literal text and lacks the contextual reasoning needed to detect sarcasm, where stated sentiment inverts intended meaning. Sarcastic negative reviews therefore get classified by surface wording, producing incorrect sentiment labels.

Why this answer

Amazon Comprehend's pre-trained sentiment analysis models are trained on general text corpora and lack the ability to detect sarcasm, which relies on contextual cues, tone, and figurative language. Sarcasm often inverts the literal sentiment (e.g., 'Great job, as always' for a failure), and standard NLP models without explicit sarcasm detection or fine-tuning cannot reliably interpret this inversion. Therefore, the likely cause is that the pre-trained model cannot handle sarcasm well.

Exam trap

The AIF-C01 exam often tests the misconception that 'fine-tuning' or 'more data' can fix any NLP issue, but here the trap is that sarcasm is a distinct linguistic challenge that pre-trained models inherently fail at, regardless of domain or data volume, unless specifically addressed with sarcasm-aware training or custom classifiers.

How to eliminate wrong answers

Option A is wrong because while fine-tuning can improve domain-specific accuracy, the core issue here is not domain mismatch but the model's inherent inability to detect sarcasm—a linguistic phenomenon that even domain-tuned models struggle with unless specifically trained on sarcastic examples. Option C is wrong because insufficient training data is not the primary cause; Amazon Comprehend's pre-trained model is trained on vast datasets, but sarcasm detection requires specialized training data and architectures (e.g., contrastive learning) that the base model lacks. Option D is wrong because input text length is not the issue; Comprehend handles up to 5,000 UTF-8 characters per request, and sarcasm is a semantic problem, not a truncation or length-related one.

28
MCQmedium

An organization wants to detect anomalies in real-time streaming data from IoT devices. The data includes sensor readings, and the team plans to use a machine learning model. Which AWS service should be used to build and deploy the model with minimal operational overhead?

A.Amazon SageMaker
B.AWS Glue
C.Amazon QuickSight
D.Amazon Kinesis Data Analytics
AnswerA

Amazon SageMaker provides fully managed infrastructure for building, training and deploying models, with built-in algorithms suited to streaming anomaly detection. It satisfies the stem's minimal operational overhead constraint by handling provisioning, scaling and endpoint hosting, letting the team focus on the model rather than servers.

Why this answer

Amazon SageMaker is the correct choice because it provides a fully managed environment for building, training, and deploying machine learning models at scale. For real-time anomaly detection on streaming IoT data, SageMaker can host a trained model as a real-time endpoint that processes incoming sensor readings via Amazon Kinesis Data Streams or AWS Lambda, minimizing operational overhead by handling infrastructure, scaling, and monitoring automatically.

Exam trap

AWS often tests the misconception that Amazon Kinesis Data Analytics can build and deploy custom ML models, when in fact it only supports built-in ML functions for simple anomaly detection and cannot train or host custom models.

How to eliminate wrong answers

Option B (AWS Glue) is wrong because it is a serverless data integration and ETL service for preparing and transforming batch data, not for building or deploying machine learning models for real-time anomaly detection. Option C (Amazon QuickSight) is wrong because it is a business intelligence (BI) service for visualizing and analyzing data, not for building or deploying ML models. Option D (Amazon Kinesis Data Analytics) is wrong because it is designed for real-time stream processing using SQL or Apache Flink, but it does not provide the capability to build, train, or deploy custom machine learning models; it is limited to built-in ML functions like anomaly detection on simple metrics, not custom model deployment.

29
MCQmedium

A data scientist is using Amazon SageMaker to train a deep learning model for image classification. The training job is taking too long. The dataset consists of 100,000 images stored in Amazon S3. Which action can the data scientist take to reduce training time without modifying the model architecture?

A.Convert images to CSV format before training.
B.Use a GPU instance type for training.
C.Enable checkpointing to save intermediate models.
D.Reduce the number of training epochs.
AnswerB

GPU instances accelerate the matrix operations in deep neural network training far more than CPU instances, cutting training time without altering the architecture. The 100,000 images in Amazon S3 already stream efficiently, so compute acceleration is the binding constraint.

Why this answer

GPU instances are specifically designed for parallel processing of matrix operations, which are fundamental to deep learning training. By switching to a GPU instance type (e.g., p3 or p4d families) in SageMaker, the data scientist can significantly accelerate the training of the image classification model without altering the model architecture, as the dataset of 100,000 images benefits from GPU's massive parallelism for forward and backward passes.

Exam trap

The trap here is that candidates may confuse checkpointing (which helps with recovery, not speed) or reducing epochs (which changes training duration but also model performance) with legitimate performance optimizations, while overlooking that GPU acceleration directly addresses the computational bottleneck without altering the model or dataset.

How to eliminate wrong answers

Option A is wrong because converting images to CSV format would increase data size, lose spatial structure, and introduce unnecessary serialization overhead, making training slower, not faster. Option C is wrong because checkpointing saves intermediate model states for fault tolerance or resumption, but it does not reduce training time; it may even add overhead due to I/O operations. Option D is wrong because reducing the number of training epochs would change the training process and likely degrade model accuracy, which violates the constraint of not modifying the model architecture (epochs are a hyperparameter, not part of architecture, but the question implies no changes that affect training duration by reducing work).

30
MCQeasy

A data scientist at a retail company is tasked with building a model to predict customer churn. The dataset contains 100,000 records with features such as age, purchase history, customer support interactions, and a binary label indicating whether the customer churned in the past. The team needs a model that can be deployed for real-time inference with low latency. They have limited time and want to use a built-in algorithm from Amazon SageMaker that is optimized for classification tasks. Which approach should they take?

A.Use Amazon SageMaker PCA algorithm
B.Use Amazon SageMaker XGBoost algorithm
C.Use Amazon SageMaker K-Means algorithm
D.Use Amazon SageMaker BlazingText algorithm
AnswerB

SageMaker's built-in XGBoost is optimised for tabular binary classification and serves low-latency real-time inference from a managed endpoint, meeting the deployment constraint. It handles the 100,000-record dataset without custom training code, saving the limited time available.

Why this answer

Amazon SageMaker's built-in XGBoost algorithm is optimized for classification tasks like binary churn prediction, supports real-time inference with low latency via SageMaker endpoints, and can handle the dataset size of 100,000 records efficiently. It is a supervised learning algorithm that directly uses the binary label for training, making it the correct choice for this scenario.

Exam trap

The trap here is that candidates may confuse unsupervised algorithms (PCA, K-Means) or domain-specific algorithms (BlazingText for text) with general-purpose supervised classification algorithms, overlooking that XGBoost is the only built-in SageMaker algorithm among the options designed for tabular classification with real-time inference needs.

How to eliminate wrong answers

Option A is wrong because PCA (Principal Component Analysis) is an unsupervised dimensionality reduction algorithm, not a classification algorithm, and cannot predict churn from a binary label. Option C is wrong because K-Means is an unsupervised clustering algorithm used for grouping data, not for supervised classification tasks like churn prediction. Option D is wrong because BlazingText is optimized for text classification and word embeddings, not for tabular data with features like age and purchase history.

31
MCQeasy

A retail company wants to build a system that predicts next month's sales for each of its 500 stores based on historical sales, local holidays, and marketing spend. The target values are continuous dollar amounts, and the company has labeled historical data for every store. Which type of machine learning problem does this represent?

A.Unsupervised learning, because the model discovers hidden groupings of stores on its own.
B.Supervised regression, because the model learns from labeled data to predict a continuous numeric value.
C.Supervised classification, because the model assigns each store to a predicted category.
D.Reinforcement learning, because the model improves by receiving rewards after each prediction.
AnswerB

The historical records pair input features such as past sales, holidays, and marketing spend with a known numeric target, next month's sales. Predicting a continuous quantity from labeled examples is exactly regression within supervised learning. This matches both the availability of labels and the continuous nature of the dollar amount the company needs forecasted.

Why this answer

The company holds labeled historical data and needs a continuous numeric output, next month's sales in dollars. That combination defines supervised regression, where the model learns a mapping from input features to a numeric target. Classification, unsupervised learning, and reinforcement learning all misalign with either the presence of labels or the continuous nature of the predicted value.

Exam trap

The trap here is assuming any business forecasting problem must be classification just because it involves predicting a future outcome.

32
Multi-Selectmedium

Which TWO of the following are examples of supervised learning tasks that can be performed using Amazon SageMaker built-in algorithms?

Select 2 answers
A.Principal Component Analysis (PCA)
B.XGBoost
C.Linear Learner
D.Latent Dirichlet Allocation (LDA)
E.K-Means
AnswersB, C

XGBoost is a gradient-boosted decision tree algorithm that trains on labelled data to predict a target, making it a supervised task. SageMaker provides it as a built-in algorithm, satisfying the stem's requirement for supervised learning examples.

Why this answer

XGBoost (B) is correct because Amazon SageMaker's built-in XGBoost algorithm is a supervised gradient-boosted trees implementation used for classification and regression on labeled data. Linear Learner (C) is correct because SageMaker's built-in Linear Learner algorithm trains supervised linear models for classification and regression using labeled datasets. PCA (A) is not correct because it is an unsupervised dimensionality-reduction technique that does not use target labels.

LDA (D) is not correct because it is an unsupervised topic-modeling algorithm. K-Means (E) is not correct because it is an unsupervised clustering algorithm that groups unlabeled data.

Exam trap

The AIF-C01 exam often tests the distinction between supervised and unsupervised learning by listing algorithms like PCA, LDA, and K-Means alongside supervised ones, trapping candidates who recognize the algorithm names but forget their learning paradigm.

33
MCQhard

A healthcare provider wants to predict which patients are likely to be readmitted within 30 days. The historical dataset has 12,000 admissions and includes age, diagnosis codes, length of stay, and prior admissions. The team has limited machine learning experience and needs an explainable model that shows which factors drove each prediction. Which AWS approach is most appropriate?

A.Use Amazon Personalize to rank patients by readmission risk based on their historical interactions.
B.Use Amazon Forecast to predict readmission counts and rely on its built-in accuracy metrics for explanation.
C.Use Amazon Lex to build a conversational interface that asks clinicians to estimate readmission probability.
D.Use Amazon SageMaker Autopilot to train candidate models on the labeled dataset, then review the model explainability report to understand feature influence.
AnswerD

SageMaker Autopilot automates feature engineering, algorithm selection, and hyperparameter tuning for tabular supervised problems, which suits a team with limited ML experience. It also produces a model explainability report showing how each feature contributes to predictions, addressing the explainability requirement. The labeled readmission outcome makes this a supervised binary classification task that Autopilot handles directly.

Why this answer

The task is supervised binary classification on tabular patient data, which SageMaker Autopilot handles by automatically exploring preprocessing, algorithms, and hyperparameters. Its model explainability report attributes predictions to input features, satisfying the need for transparency. Forecasting, recommendation, and conversational services do not perform per-patient clinical risk prediction with feature-level explanations.

Exam trap

The trap here is selecting a forecasting or recommendation service because it also produces scores, when the actual task is tabular binary classification with explainability.

34
MCQhard

A company wants to deploy a real-time inference endpoint for a custom model on SageMaker. The model has high latency (100ms) and they need to handle variable traffic with spikes. Which deployment strategy is most cost-effective?

A.Deploy on a SageMaker multi-model endpoint
B.Use batch transform
C.Deploy on a single SageMaker endpoint with automatic scaling
D.Use a single large instance type
AnswerC

Automatic scaling adjusts instance count to match traffic spikes while a single endpoint avoids paying for idle capacity during troughs. This satisfies the variable-traffic constraint cost-effectively, unlike always-on multi-instance deployments that over-provision for peak load.

Why this answer

A single SageMaker endpoint with automatic scaling allows the endpoint to dynamically adjust the number of instances based on traffic patterns, handling variable traffic and spikes cost-effectively. For a model with 100ms latency, automatic scaling can add instances during spikes and remove them during low traffic, ensuring you only pay for the compute resources you use while maintaining low inference latency.

Exam trap

The trap here is that candidates often confuse multi-model endpoints with cost-effective scaling for a single model, not realizing that multi-model endpoints are designed for hosting many models, not for handling variable traffic for one model with high latency.

How to eliminate wrong answers

Option A is wrong because a multi-model endpoint is designed to host multiple models on a shared instance to reduce costs, but it does not inherently handle high-latency models (100ms) well under variable traffic spikes, as it can lead to resource contention and increased latency. Option B is wrong because batch transform is an asynchronous, offline inference method for processing large datasets in batches, not suitable for real-time inference endpoints that require immediate responses. Option D is wrong because using a single large instance type is not cost-effective for variable traffic with spikes; you would over-provision for peak traffic and pay for idle capacity during low traffic, whereas automatic scaling adjusts resources dynamically.

35
MCQmedium

During model training, the loss decreases rapidly for the first few epochs and then plateaus. The validation loss starts increasing after some epochs. What should the team do to improve generalization?

A.Increase learning rate
B.Early stopping
C.Increase training epochs
D.Add more layers
AnswerB

Rising validation loss while training loss plateaus signals overfitting beyond the optimal point. Early stopping halts training when validation loss stops improving, restoring the best weights and preventing further memorisation, thereby improving generalisation without changing the model architecture.

Why this answer

The validation loss increasing while training loss continues to decrease is a classic sign of overfitting. Early stopping (Option B) halts training when validation performance stops improving, preventing the model from memorizing noise in the training data and thereby improving generalization.

Exam trap

The AIF-C01 exam often tests the misconception that overfitting is solved by increasing model complexity or training longer, when in fact the opposite is true—early stopping or regularization techniques are required to curb overfitting.

How to eliminate wrong answers

Option A is wrong because increasing the learning rate would cause larger weight updates, potentially overshooting minima and destabilizing training, which does not address overfitting. Option C is wrong because increasing training epochs would allow the model to continue fitting the training data even more closely, exacerbating overfitting rather than improving generalization. Option D is wrong because adding more layers increases model capacity, making it more prone to overfitting on the training data, not less.

36
MCQmedium

A retail company uses a machine learning model to forecast daily product demand. The model is a time series model that uses historical sales data. The model has been performing well, but recently the forecasts have been consistently too low, leading to stockouts. The data scientist notices that the model was trained on data up to last year, and the company has since launched a successful marketing campaign that increased sales by 20%. The data scientist needs to update the model to reflect the new sales patterns. Which approach should the data scientist take?

A.Add a feature for the marketing campaign and continue using the old model.
B.Switch to a different model type, such as ARIMA, without retraining.
C.Multiply the model's predictions by 1.2 to account for the marketing campaign.
D.Retrain the model on the most recent data that includes the sales from the marketing campaign.
AnswerD

The marketing campaign shifted the underlying sales distribution, so the stale training data no longer represents current demand. Retraining on recent data that includes campaign-period sales lets the model learn the new pattern, correcting the systematic under-forecasting.

Why this answer

Retraining the model on the most recent data that includes the sales from the marketing campaign allows the model to learn the new underlying pattern in the time series. Since the model is a time series model, it relies on historical patterns to make forecasts; retraining on data that captures the 20% sales lift ensures the model adapts to the new demand level, reducing the persistent underforecasting.

Exam trap

The trap here is that candidates may think a simple multiplicative adjustment (Option C) is sufficient, but the AWS exam tests the understanding that time series models must be retrained on the new distribution to maintain forecast accuracy, as static adjustments ignore changes in the underlying data-generating process.

How to eliminate wrong answers

Option A is wrong because simply adding a feature for the marketing campaign without retraining the model does not update the model's learned parameters; the old model's weights are still based on pre-campaign data, so it cannot properly incorporate the new pattern. Option B is wrong because switching to a different model type like ARIMA without retraining means the new model has no learned parameters from the data, and ARIMA requires fitting to the specific time series to estimate its autoregressive and moving average components. Option C is wrong because multiplying predictions by a constant factor (1.2) assumes a uniform multiplicative effect that may not hold across all products or time periods, and it does not address potential changes in seasonality, trend, or other dynamics introduced by the campaign.

37
MCQhard

A data scientist wants to perform automatic model tuning (hyperparameter optimization) on SageMaker. They need to find the best hyperparameters for a gradient boosting model. Which strategy is BEST for this task?

A.Random search
B.Grid search
C.Exhaustive search
D.Bayesian optimization
AnswerD

Bayesian optimisation builds a probabilistic surrogate model of the objective function, then uses an acquisition function to pick each next hyperparameter combination, converging on strong configurations in far fewer training jobs than random or grid search. This directly satisfies the stem's requirement for efficient automatic model tuning on SageMaker.

Why this answer

Bayesian optimization is the best strategy for automatic model tuning on SageMaker because it builds a probabilistic model of the objective function and uses it to select the most promising hyperparameters to evaluate next. This approach is far more sample-efficient than random or grid search, making it ideal for expensive-to-evaluate models like gradient boosting, where each training run consumes significant time and compute resources.

Exam trap

The AIF-C01 exam often tests the misconception that exhaustive or grid search is the most thorough and therefore best approach, but the trap is that they ignore the practical constraints of compute cost and time, making Bayesian optimization the superior choice for automatic model tuning in SageMaker.

How to eliminate wrong answers

Option A is wrong because random search, while better than grid search for high-dimensional spaces, does not use past evaluation results to inform future trials, making it less efficient than Bayesian optimization for finding optimal hyperparameters. Option B is wrong because grid search exhaustively evaluates all combinations of a predefined set of hyperparameter values, which is computationally prohibitive for gradient boosting models with many continuous hyperparameters and does not scale well. Option C is wrong because exhaustive search is essentially a synonym for grid search and suffers from the same curse of dimensionality, making it impractical for hyperparameter optimization in SageMaker's automatic model tuning context.

38
MCQeasy

A startup is building a recommendation engine for their e-commerce platform. They need a fully managed service that can generate personalized product recommendations based on user behavior. Which AWS service should they use?

A.Amazon Personalize
B.Amazon Rekognition
C.Amazon Forecast
D.Amazon Comprehend
AnswerA

Amazon Personalize is purpose-built for recommendations, using user-interaction data to train custom models without managing infrastructure. It satisfies the fully managed constraint, unlike SageMaker, which requires the startup to build, train and host the recommender themselves.

Why this answer

Amazon Personalize is a fully managed machine learning service specifically designed to generate real-time personalized product recommendations by processing user behavior data (e.g., clicks, purchases, views) and item metadata. It uses the same technology that powers Amazon.com's recommendation engine, making it the correct choice for this e-commerce use case.

Exam trap

The AIF-C01 exam often tests the distinction between AWS AI services by presenting a use case that sounds like 'forecasting' or 'analysis' but actually requires personalization, leading candidates to confuse Amazon Forecast (time-series) with Amazon Personalize (recommendations).

How to eliminate wrong answers

Option B (Amazon Rekognition) is wrong because it is a computer vision service for image and video analysis (e.g., object detection, facial recognition), not for generating product recommendations. Option C (Amazon Forecast) is wrong because it is a time-series forecasting service for predicting future metrics (e.g., demand, sales), not for personalized recommendations based on user behavior. Option D (Amazon Comprehend) is wrong because it is a natural language processing (NLP) service for extracting insights from text (e.g., sentiment, entities), not for recommendation generation.

39
MCQhard

A data scientist needs to preprocess categorical data with high cardinality (e.g., zip code with 50,000 unique values). Which technique is most appropriate?

A.Target encoding
B.Label encoding
C.Ordinal encoding
D.One-hot encoding
AnswerA

Target encoding replaces each category with a statistic of the target, producing a single numeric column regardless of cardinality. This suits zip codes with 50,000 unique values, where one-hot encoding would create 50,000 sparse columns and degrade model performance and memory use.

Why this answer

Target encoding is the most appropriate technique for high-cardinality categorical features like zip codes because it replaces each category with the mean of the target variable for that category, effectively compressing 50,000 unique values into a single numeric feature while retaining predictive signal. This avoids the dimensionality explosion of one-hot encoding and the arbitrary ordering of label or ordinal encoding, which can mislead models.

Exam trap

AWS often tests the trap that candidates assume one-hot encoding is always the safest choice for categorical data, ignoring the practical infeasibility of high cardinality, and fail to recognize target encoding as the standard solution for such cases.

How to eliminate wrong answers

Option B (Label encoding) is wrong because it assigns arbitrary integer labels to categories, implying an ordinal relationship that does not exist for nominal data like zip codes, which can degrade model performance. Option C (Ordinal encoding) is wrong because it also imposes a false order on categories, suitable only for ordinal data, not for high-cardinality nominal features. Option D (One-hot encoding) is wrong because it creates 50,000 binary columns, leading to extreme dimensionality, memory exhaustion, and the curse of dimensionality, making it computationally infeasible.

40
Multi-Selecteasy

Which TWO services can be used to preprocess data for machine learning in AWS? (Choose two.)

Select 2 answers
A.AWS Glue
B.Amazon Athena
C.Amazon SageMaker Data Wrangler
D.Amazon Redshift
E.AWS Lambda
AnswersA, C

AWS Glue is a serverless ETL service that cleanses, transforms and catalogues data at scale using Spark jobs or Glue Studio, producing prepared datasets for ML training. It satisfies the preprocessing requirement by handling transformation and data quality tasks before model training.

Why this answer

AWS Glue (A) is correct because it is a fully managed, serverless ETL service that can run Apache Spark jobs and Glue Studio/DynamicFrames to clean, transform, join, and normalize raw data into feature-ready datasets for ML training. Amazon SageMaker Data Wrangler (C) is correct because it is purpose-built for ML data preparation, providing a visual interface with 300+ built-in transforms, data quality/insight reports, and one-click export to SageMaker pipelines or feature stores. Amazon Athena (B) is a serverless interactive query service over S3 using SQL, suited for ad hoc analytics rather than building reusable preprocessing pipelines.

Amazon Redshift (D) is a data warehouse for analytics and SQL workloads, not a dedicated ML preprocessing service. AWS Lambda (E) is a general-purpose serverless compute service that could run small transformation code but is not a data preprocessing service for ML.

Exam trap

The AIF-C01 exam often tests the distinction between data querying services (like Athena) and data preprocessing services, leading candidates to mistakenly choose Athena because it can 'process' data via SQL, but it lacks the ML-specific transformation capabilities required for preprocessing.

41
MCQmedium

A company is using Amazon Rekognition to detect objects in images. They need to detect custom objects that are specific to their domain. What should they do?

A.Use Amazon Rekognition's built-in labels
B.Use Amazon SageMaker Object Detection algorithm
C.Use Amazon Rekognition Custom Labels
D.Use Amazon Comprehend
AnswerC

Rekognition Custom Labels trains a bespoke model on your own labelled images, so it detects domain-specific objects that the general-purpose Rekognition detector cannot recognise. This directly satisfies the requirement for custom, domain-specific object detection without building a model from scratch.

Why this answer

Amazon Rekognition Custom Labels allows you to train a custom model using your own labeled images to detect domain-specific objects that are not covered by Rekognition's built-in labels. This is the correct service for custom object detection without needing to build a model from scratch.

Exam trap

The trap here is that candidates may confuse Amazon Rekognition Custom Labels with SageMaker Object Detection, not realizing that Custom Labels is a managed service specifically designed for custom image analysis without requiring ML expertise.

How to eliminate wrong answers

Option A is wrong because built-in labels are pre-trained on general categories and cannot detect custom domain-specific objects. Option B is wrong because Amazon SageMaker Object Detection algorithm requires you to build, train, and deploy a custom model from scratch, which is more complex and not the recommended approach when Rekognition Custom Labels can handle the task with less effort. Option D is wrong because Amazon Comprehend is a natural language processing (NLP) service for text analysis, not for image object detection.

42
MCQeasy

A company needs to store large amounts of unstructured training data (images, videos) in a cost-effective manner while ensuring low-latency retrieval for training jobs running on Amazon SageMaker. Which storage solution should be used?

A.Amazon EFS
B.Amazon S3
C.Amazon RDS
D.Amazon EBS
AnswerB

Amazon S3 provides durable, low-cost object storage for unstructured images and video, and SageMaker training jobs read directly from S3 with high throughput and low latency. This satisfies both the cost-effectiveness and low-latency retrieval constraints in the stem.

Why this answer

Amazon S3 is the correct choice because it is designed for cost-effective, scalable storage of unstructured data (images, videos) and integrates natively with Amazon SageMaker for low-latency data retrieval during training jobs. S3 provides high throughput and can be accessed directly from SageMaker training instances without the need for file system mounting, making it ideal for large-scale ML workloads.

Exam trap

The trap here is that candidates often confuse the need for low-latency retrieval with the need for a mounted file system (EFS or EBS), not realizing that S3's direct integration with SageMaker provides both low latency and high throughput for training workloads without the cost and complexity of file storage.

How to eliminate wrong answers

Option A is wrong because Amazon EFS is a file system that provides shared access for EC2 instances but is not optimized for the high-throughput, cost-effective storage of large unstructured datasets like images and videos; it also incurs higher costs per GB compared to S3 and can introduce latency overhead when used with SageMaker. Option C is wrong because Amazon RDS is a relational database service designed for structured data with SQL queries, not for storing unstructured training data such as images and videos, and it would be prohibitively expensive and inefficient for large-scale blob storage. Option D is wrong because Amazon EBS provides block-level storage volumes attached to a single EC2 instance, which is not suitable for sharing large datasets across multiple SageMaker training jobs and lacks the cost efficiency and scalability of object storage for unstructured data.

43
MCQeasy

A company wants to build a system that automatically categorizes customer support tickets into predefined categories (e.g., billing, technical, account). The team has a large dataset of historical tickets with their category labels. Which type of machine learning problem is this?

A.Regression
B.Binary classification
C.Multi-class classification
D.Clustering
AnswerC

Historical tickets carry one label from a set of more than two predefined categories, so the model must assign each ticket to exactly one of several mutually exclusive classes. That maps to multi-class classification, not binary, regression or clustering.

Why this answer

This is a multi-class classification problem because the model must assign each support ticket to one of three or more predefined categories (e.g., billing, technical, account). The dataset provides labeled historical tickets, making it a supervised learning task, and the output is a discrete class label from a set of more than two categories, which distinguishes it from binary classification.

Exam trap

The AIF-C01 exam often tests the distinction between binary and multi-class classification by presenting a scenario with multiple categories but implying a simple yes/no decision, leading candidates to mistakenly choose binary classification when the number of classes exceeds two.

How to eliminate wrong answers

Option A is wrong because regression predicts continuous numerical values (e.g., ticket resolution time), not discrete categories. Option B is wrong because binary classification only handles two classes (e.g., spam vs. not spam), whereas this problem involves three or more categories. Option D is wrong because clustering is an unsupervised learning technique that groups data without using predefined labels, while this problem uses labeled historical data for supervised learning.

44
MCQhard

A healthcare company is using Amazon SageMaker to deploy a model that makes predictions on patient data. They need to ensure that the model's predictions are explainable to comply with regulations. Which approach should they take?

A.Use SageMaker Model Monitor to track predictions
B.Use SageMaker Experiments to log model parameters
C.Use SageMaker Clarify to generate feature importance and explanations
D.Use SageMaker Debugger to analyze training gradients
AnswerC

SageMaker Clarify computes feature attribution and produces explanations for individual predictions, exposing which input features drove each output. This satisfies the regulatory requirement for explainable predictions on patient data, unlike accuracy-focused or purely operational measures.

Why this answer

SageMaker Clarify is specifically designed to provide model explainability, including feature importance and SHAP-based explanations, which are essential for regulatory compliance in healthcare. It helps stakeholders understand why a model made a particular prediction, addressing transparency requirements.

Exam trap

The AIF-C01 exam often tests the distinction between monitoring (Model Monitor), tracking (Experiments), debugging (Debugger), and explainability (Clarify), so the trap here is confusing operational monitoring with the need for interpretable explanations required by compliance frameworks.

How to eliminate wrong answers

Option A is wrong because SageMaker Model Monitor is used for detecting data drift and model quality degradation over time, not for generating per-prediction explanations. Option B is wrong because SageMaker Experiments tracks and organizes model training runs and parameters, but does not produce explainability reports for individual predictions. Option D is wrong because SageMaker Debugger monitors training metrics and gradients to debug training issues, not to explain model predictions post-deployment.

45
MCQmedium

A marketing agency wants to analyze customer feedback from social media posts to gauge sentiment. They have no labeled data and limited ML expertise. The team needs a managed service that provides pre-trained models for sentiment analysis without requiring them to train or manage infrastructure. They also need to process text in multiple languages. Which AWS service should they use?

A.Use Amazon Comprehend with its default sentiment analysis model
B.Use Amazon SageMaker to train a custom sentiment analysis model
C.Use AWS Glue to build a custom NLP pipeline
D.Use Amazon Rekognition for text analysis
AnswerA

Amazon Comprehend provides pre-trained sentiment analysis via a managed API, requiring no labelled data, training, or infrastructure management. Its built-in multilingual support handles social media text across languages, matching the agency's lack of ML expertise and need for immediate sentiment detection.

Why this answer

Amazon Comprehend is a fully managed natural language processing (NLP) service that provides pre-trained models for sentiment analysis, key phrase extraction, and language detection. It requires no labeled data, no model training, and no infrastructure management, making it ideal for teams with limited ML expertise. Comprehend natively supports multiple languages, including Spanish, French, German, and many others, directly addressing the requirement to process text in multiple languages.

Exam trap

The trap here is that candidates may confuse Amazon Rekognition (image/video analysis) with text analysis services, or assume that any AWS ML service (like SageMaker or Glue) can handle NLP tasks without recognizing the specific managed service designed for unstructured text.

How to eliminate wrong answers

Option B is wrong because Amazon SageMaker is a platform for building, training, and deploying custom ML models, which requires labeled data, ML expertise, and infrastructure management—contradicting the requirements for a pre-trained, managed service with no training. Option C is wrong because AWS Glue is a serverless data integration and ETL service, not an NLP service; it cannot perform sentiment analysis or provide pre-trained models. Option D is wrong because Amazon Rekognition is a computer vision service for analyzing images and videos, not for text analysis or sentiment detection.

46
Multi-Selectmedium

Which THREE are SageMaker built-in algorithms suitable for regression tasks?

Select 3 answers
A.Linear Learner
B.K-Means
C.PCA
D.DeepAR
E.XGBoost
AnswersA, D, E

Linear Learner trains a linear function to predict a continuous target, supporting both regression and classification. It is a SageMaker built-in algorithm explicitly designed for regression, satisfying the question's requirement for a suitable built-in option.

Why this answer

Linear Learner (A) is a SageMaker built-in algorithm that supports both classification and regression; in regression mode it learns a linear function minimizing a chosen loss (e.g., squared error) on continuous targets. DeepAR (D) is a built-in supervised time-series forecasting algorithm that predicts continuous numeric values, so it is used for regression-style forecasting tasks. XGBoost (E) is a built-in gradient-boosted trees algorithm whose objective can be set to reg:squarederror or reg:logistic, making it suitable for regression on tabular data.

K-Means (B) is an unsupervised clustering algorithm that groups data points and does not predict a continuous target, and PCA (C) is an unsupervised dimensionality-reduction algorithm, so neither belongs to regression tasks.

Exam trap

The AIF-C01 exam often tests the distinction between supervised and unsupervised algorithms, and the trap here is that candidates may confuse dimensionality reduction (PCA) or clustering (K-Means) with regression tasks, assuming any algorithm that processes numeric data can perform regression.

47
MCQeasy

A data scientist wants to host a pre-trained model on Amazon SageMaker for real-time inference with minimal latency. Which approach should they use?

A.Run inference using AWS Lambda with the model packaged as a container
B.Use SageMaker batch transform
C.Create a SageMaker asynchronous inference endpoint
D.Deploy the model on a SageMaker real-time endpoint
AnswerD

A SageMaker real-time endpoint keeps model artefacts loaded on dedicated instances behind a persistent HTTPS endpoint, returning predictions synchronously with low latency. This satisfies the minimal-latency requirement for real-time inference, unlike batch transform or asynchronous inference, which introduce queuing and storage round-trips.

Why this answer

SageMaker real-time endpoints are designed for low-latency, synchronous inference. They keep the model loaded and ready to respond to individual requests, making them ideal for real-time applications where minimal latency is critical.

Exam trap

The AIF-C01 exam often tests the distinction between synchronous (real-time) and asynchronous inference patterns, and the trap here is that candidates may confuse 'asynchronous inference' with 'real-time' because both can handle requests, but only real-time endpoints guarantee minimal latency for individual predictions.

How to eliminate wrong answers

Option A is wrong because AWS Lambda has a maximum execution timeout of 15 minutes and limited memory (up to 10 GB), making it unsuitable for hosting large pre-trained models that require persistent, low-latency inference. Option B is wrong because SageMaker batch transform is an asynchronous, offline process for processing large datasets in batches, not for real-time inference with minimal latency. Option C is wrong because SageMaker asynchronous inference endpoints are designed for requests with large payloads and longer processing times, where immediate response is not required; they introduce queuing and processing delays that are incompatible with minimal latency requirements.

48
MCQmedium

A retail bank has millions of unlabeled customer transaction records and wants to discover natural groupings of spending behavior without defining any categories in advance. The data science team plans to use an unsupervised learning approach. Which technique is designed for this goal?

A.Linear regression, which models the relationship between a continuous target and one or more input features.
B.Logistic regression, which estimates the probability that a record belongs to a particular class.
C.A decision tree classifier, which splits data using feature thresholds to predict a labeled category.
D.K-means clustering, which partitions records into groups so that members of a group are more similar to each other.
AnswerD

K-means is an unsupervised algorithm that groups unlabeled records into a chosen number of clusters based on feature similarity. It directly matches the bank's goal of discovering natural spending-behavior segments without predefined categories. The team can then profile each cluster, for example frequent travelers versus everyday local spenders, to inform marketing or risk decisions.

Why this answer

The bank has no predefined categories and wants to uncover structure in unlabeled data, which is the defining use case for unsupervised learning. K-means clustering groups similar records together so the team can inspect and name the resulting segments afterward, turning raw transaction behavior into actionable customer groupings.

Exam trap

The trap here is reaching for a familiar supervised algorithm such as regression or a classifier even though the data has no labels and no target to predict.

49
Multi-Selecthard

A data engineer is using Amazon SageMaker Data Wrangler to prepare tabular data for ML. Which THREE data transformations are natively supported? (Choose three.)

Select 3 answers
A.One-hot encoding for categorical features
B.Audio feature extraction
C.Text vectorization using TF-IDF
D.Custom Python code via Pandas or Spark
E.Image resizing and normalization
AnswersA, C, D

Data Wrangler includes a built-in one-hot encoding transform that converts categorical columns into binary indicator features, selectable directly in the transformation list. It satisfies the native transformation requirement without custom code, unlike techniques requiring external libraries or manual scripting.

Why this answer

SageMaker Data Wrangler natively supports one-hot encoding as a categorical-encoding transform, which converts categorical features into binary indicator columns suitable for ML models, so option A is correct. It also provides a text vectorization transform using TF-IDF (and other text transforms like bag-of-words and n-gram) to convert text fields into numeric feature vectors, making option C correct. Data Wrangler additionally allows custom transformations through the Custom Transform node, where users can write their own Pandas or PySpark code, so option D is correct.

Options B and E are not native Data Wrangler tabular transforms: audio feature extraction and image resizing/normalization are handled by other AWS services or frameworks (for example, SageMaker Processing with librosa or image libraries), not by Data Wrangler's built-in transform list.

Exam trap

AWS often tests the distinction between natively supported transformations in SageMaker Data Wrangler versus those requiring external services or custom scripts, leading candidates to mistakenly select audio or image processing options that are not part of Data Wrangler's built-in capabilities.

50
Multi-Selectmedium

Which AWS services can be used to build, train, and deploy custom machine learning models? (Choose two.)

Select 2 answers
A.Amazon Polly
B.Amazon Lex
C.AWS Deep Learning AMIs
D.Amazon Rekognition
E.Amazon SageMaker
AnswersC, E

AWS Deep Learning AMIs provide pre-configured Amazon EC2 images bundling frameworks such as TensorFlow and PyTorch with GPU drivers, satisfying the build-and-train requirement. They supply the compute environment for custom model development, though deployment needs a separate service. This addresses the training half of the stem directly.

Why this answer

AWS Deep Learning AMIs (C) are correct because they provide pre-configured Amazon EC2 images with popular deep learning frameworks (TensorFlow, PyTorch, MXNet) and GPU drivers, giving you the full environment needed to build and train custom models on your own infrastructure. Amazon SageMaker (E) is correct because it is a fully managed platform whose built-in capabilities (notebooks, training jobs, automatic model tuning, and hosting endpoints) let you build, train, and deploy custom ML models end to end. The other options are managed AI services with pre-trained models rather than tools for creating custom models: Amazon Polly (A) only converts text to lifelike speech, Amazon Lex (B) only builds conversational chatbots using NLU, and Amazon Rekognition (D) only provides pre-trained image and video analysis.

Exam trap

The trap here is that candidates confuse pre-built AI services (Polly, Lex, Rekognition) with platforms that allow custom model development, leading them to select services that only consume pre-trained models rather than build and train custom ones.

51
Multi-Selecthard

A financial services company is preparing to train a machine learning model on customer transaction data. The data science team must address data quality concerns before training, because poor data directly harms model performance. Which TWO practices best improve the quality of the training data? (Choose two.)

Select 2 answers
A.Add more duplicate rows to increase the effective dataset size
B.Increase the learning rate so the model converges faster on noisy data
C.Handle missing values by imputing them or removing affected records after analysis
D.Remove all features that have any correlation with the target variable
E.Normalize or standardize numeric features so they share a comparable scale
AnswersC, E

Missing values can bias or destabilize training if left untreated. Analyzing the pattern of missingness and then imputing sensible values or removing affected records produces a cleaner dataset that better represents the population. This is a foundational data quality practice directly tied to the goal of improving model performance before training begins.

Why this answer

Improving data quality before training centers on correcting defects in the data itself. Handling missing values prevents bias and instability, while scaling numeric features ensures algorithms treat inputs comparably. Hyperparameter changes, row duplication, and deleting predictive features do not repair data defects and in several cases worsen results.

Exam trap

The trap here is treating model tuning knobs or dataset inflation as data quality fixes when quality work targets the data itself.

52
Multi-Selectmedium

Which TWO techniques are commonly used to prevent overfitting in machine learning models? (Select TWO.)

Select 2 answers
A.Add more irrelevant features
B.Use cross-validation
C.Increase model complexity
D.Reduce the amount of training data
E.Use regularization
AnswersB, E

Cross-validation partitions the data into folds, training and validating on different subsets, so performance estimates reflect generalisation rather than memorisation of one split. It detects overfitting by exposing the gap between training and held-out fold error.

Why this answer

Option B (Use cross-validation) is correct because techniques like k-fold cross-validation estimate how well a model generalizes to unseen data by training and validating on different data splits, which helps detect and mitigate overfitting during model selection and hyperparameter tuning. Option E (Use regularization) is correct because methods such as L1 (Lasso) and L2 (Ridge) regularization add a penalty term to the loss function that constrains the magnitude of model weights, reducing variance and preventing the model from fitting noise in the training data. The unmarked options do not belong: adding more irrelevant features (A) increases dimensionality and noise, making overfitting worse; increasing model complexity (C) raises variance and typically worsens overfitting; and reducing the amount of training data (D) gives the model less information to learn generalizable patterns, also promoting overfitting.

Exam trap

AWS often tests the misconception that adding more data or features always helps model performance, when in fact irrelevant features or reducing training data can worsen overfitting, and candidates may incorrectly associate 'more complexity' with better generalization.

53
MCQhard

An ML engineer wants to store training data in a format optimized for linear data scanning and columnar access in SageMaker. Which format is most appropriate?

A.JSON
B.Image (JPEG/PNG)
C.Parquet
D.CSV
AnswerC

Parquet is a columnar format, so SageMaker reads only the columns a query needs and scans them linearly, cutting I/O versus row-based formats. This matches the stated requirement for columnar access and efficient linear scanning of training data.

Why this answer

Parquet is a columnar storage format optimized for both linear data scanning and columnar access, making it ideal for training data in SageMaker. It reduces I/O by storing data by columns rather than rows, enabling efficient retrieval of specific features during model training.

Exam trap

AWS often tests the misconception that CSV is the most efficient format for training data, but Parquet's columnar storage and compression provide superior performance for linear scanning and columnar access in distributed ML pipelines.

How to eliminate wrong answers

Option A is wrong because JSON is a row-oriented text format that requires full parsing for columnar access, leading to high I/O overhead and slower linear scans. Option B is wrong because image formats like JPEG/PNG are binary and designed for visual data, not structured tabular data, and lack columnar access capabilities. Option D is wrong because CSV is a row-oriented text format that, while simple, requires scanning entire rows to access specific columns and lacks compression and schema optimization.

54
MCQeasy

A hospital wants to build a system that automatically assigns a specialty department (for example, Cardiology, Neurology, or Orthopedics) to each free-text patient referral note. The hospital has a large archive of past referral notes that were already labeled by clinicians with the correct department. Which type of machine learning problem does this scenario describe?

A.Supervised learning, because the model learns from labeled historical examples to predict a categorical outcome.
B.Generative AI, because the model must create a new department name for every referral note it processes.
C.Unsupervised learning, because the model must discover hidden structure within the referral notes without guidance.
D.Reinforcement learning, because the model improves through rewards received after each department assignment.
AnswerA

The archive of referral notes already carries clinician-assigned department labels, so the model has ground-truth targets during training and predicts one category from a fixed set. That is the definition of supervised classification, where the algorithm learns a mapping from input text to a discrete label and can then generalize to new, unseen referral notes.

Why this answer

Because every historical referral note already has a clinician-assigned department, the training data is fully labeled and the target is one of several discrete categories. That combination defines supervised classification, in which a model learns the relationship between note text and department and then predicts the department for new notes.

Exam trap

The trap here is assuming that working with text automatically makes a problem unsupervised or generative, when the presence of labeled outcomes is what determines the learning type.

55
MCQeasy

Which metric is most appropriate for evaluating a classification model when false positives are costly?

A.Precision
B.F1 score
C.Recall
D.Accuracy
AnswerA

Precision measures the proportion of positive predictions that are actually correct, so it directly penalises false positives. When false positives carry high cost, maximising precision minimises those costly incorrect positive classifications, unlike recall or accuracy.

Why this answer

Precision is the most appropriate metric when false positives are costly because it measures the proportion of true positive predictions among all positive predictions (TP / (TP + FP)). A high precision indicates that when the model predicts a positive class, it is very likely correct, minimizing the number of false positives. This directly aligns with the business requirement to avoid costly false alarms.

Exam trap

The AIF-C01 exam often tests the distinction between precision and recall by framing a cost scenario, and the trap here is that candidates confuse 'costly false positives' with 'costly false negatives' and incorrectly choose recall or F1 score without analyzing which error type is being penalized.

How to eliminate wrong answers

Option B (F1 score) is wrong because it is the harmonic mean of precision and recall, balancing both false positives and false negatives; it does not specifically penalize false positives more heavily. Option C (Recall) is wrong because it measures the proportion of actual positives correctly identified (TP / (TP + FN)), which is useful when false negatives are costly, not false positives. Option D (Accuracy) is wrong because it considers overall correct predictions (TP + TN) divided by total predictions, which can be misleading in imbalanced datasets and does not isolate the cost of false positives.

56
MCQeasy

A junior data scientist is building a model to classify incoming customer support tickets into one of eight predefined categories such as Billing, Shipping, or Returns. Historical tickets already have correct category labels. Which type of machine learning is being used?

A.Generative AI fine-tuning
B.Reinforcement learning
C.Unsupervised learning
D.Supervised learning
AnswerD

Supervised learning trains on labeled examples, and here every historical ticket already carries the correct category label, so the model learns a mapping from ticket text to one of the eight predefined classes. This is textbook multi-class classification, a supervised task, which is exactly what the scenario describes.

Why this answer

The presence of correct historical labels plus a fixed set of eight output categories makes this multi-class classification, which belongs to supervised learning. The model is taught from known input-output pairs, then predicts the category for new tickets. Unsupervised, reinforcement, and generative approaches all lack the labeled input-output mapping central to this task.

Exam trap

The trap here is assuming any text-related task must involve generative AI or unsupervised clustering, when predefined labels and fixed categories make it supervised classification.

57
MCQeasy

A company wants to use AI to automatically transcribe customer service calls into text. Which AWS service is most suitable?

A.Amazon Transcribe
B.Amazon Comprehend
C.Amazon Polly
D.Amazon Rekognition
AnswerA

Amazon Transcribe is a fully managed automatic speech recognition service that converts audio to text, supporting batch and streaming transcription with speaker diarisation. It directly satisfies the requirement to transcribe customer service calls without building custom acoustic models.

Why this answer

Amazon Transcribe is the correct choice because it is a fully managed automatic speech recognition (ASR) service designed specifically to convert speech into text. It can handle real-time streaming or batch processing of audio files, making it ideal for transcribing customer service calls into searchable text.

Exam trap

The trap here is that candidates often confuse Amazon Transcribe (speech-to-text) with Amazon Polly (text-to-speech) or assume Amazon Comprehend can process audio directly, when in fact Comprehend only works on text input.

How to eliminate wrong answers

Option B is wrong because Amazon Comprehend is a natural language processing (NLP) service used for extracting insights like sentiment, entities, and key phrases from text, not for transcribing audio. Option C is wrong because Amazon Polly is a text-to-speech (TTS) service that converts text into lifelike speech, the opposite of the required speech-to-text functionality. Option D is wrong because Amazon Rekognition is a computer vision service for analyzing images and videos, such as object detection and facial recognition, and has no capability to process audio or transcribe speech.

58
Multi-Selecthard

A company is using Amazon Fraud Detector to detect fraudulent transactions. Which TWO actions can be taken to improve model accuracy? (Select TWO.)

Select 2 answers
A.Increase the volume of event data
B.Deploy the model to multiple endpoints
C.Use a different detector type
D.Use a different model version
E.Select event variables that are more predictive
AnswersA, E

Fraud Detector's models learn from historical event data, so a larger volume of labelled fraud and legitimate events gives the algorithm more signal to distinguish patterns. This directly addresses the accuracy constraint by reducing overfitting to sparse samples.

Why this answer

Option A is correct because Amazon Fraud Detector's model accuracy improves with more historical event data — a larger volume of labeled fraud/legitimate events gives the training algorithm more examples to learn patterns from, reducing overfitting and improving predictions. Option E is correct because model accuracy depends heavily on feature quality; choosing event variables with stronger predictive power (e.g., customer age, order price, IP address, email domain) directly improves the model's ability to distinguish fraudulent from legitimate transactions. Option B is incorrect because deploying a model to multiple endpoints only affects availability/scalability of predictions, not the underlying model's accuracy.

Option C is incorrect because changing the detector type (e.g., online fraud, transaction fraud) alters the use case rather than inherently improving accuracy for the existing scenario. Option D is incorrect because switching to a different model version simply selects an already-trained model; it does not by itself improve accuracy unless retraining with better data or variables occurs.

Exam trap

The AIF-C01 exam often tests the misconception that changing model versions or detector types alone improves accuracy, when in reality accuracy improvements require data or feature enhancements.

59
MCQeasy

A company wants to automatically detect anomalies in their AWS CloudTrail logs to identify potential security threats. Which AWS service is specifically designed for this purpose?

A.Amazon Macie
B.AWS Config
C.Amazon GuardDuty
D.Amazon Inspector
AnswerC

GuardDuty is a managed threat-detection service that continuously analyses CloudTrail management events, VPC Flow Logs and DNS logs using machine learning and threat intelligence, surfacing anomalous or malicious activity without agents — precisely the automated anomaly detection the company needs.

Why this answer

Amazon GuardDuty is a threat detection service that continuously monitors AWS accounts and workloads using machine learning, anomaly detection, and integrated threat intelligence. It specifically analyzes CloudTrail management and data events, VPC Flow Logs, and DNS logs to identify unauthorized behavior or potential security threats, making it the correct choice for automatically detecting anomalies in CloudTrail logs.

Exam trap

The AIF-C01 exam often tests the distinction between services that detect threats (GuardDuty) versus services that protect data (Macie), assess vulnerabilities (Inspector), or track configuration compliance (Config), leading candidates to confuse their primary use cases.

How to eliminate wrong answers

Option A is wrong because Amazon Macie is a data security and data privacy service that uses machine learning to discover, classify, and protect sensitive data stored in Amazon S3, not to analyze CloudTrail logs for security threats. Option B is wrong because AWS Config is a service that evaluates and records resource configurations and compliance against desired policies, not designed for real-time anomaly detection in log data. Option D is wrong because Amazon Inspector is a vulnerability management service that scans EC2 instances and container images for software vulnerabilities and unintended network exposure, not for analyzing CloudTrail logs.

60
Multi-Selecteasy

Which TWO of the following are types of feature scaling?

Select 2 answers
A.One-hot encoding
B.Principal Component Analysis (PCA)
C.Standardization
D.Binning
E.Normalization (Min-Max)
AnswersC, E

Standardization rescales each feature to zero mean and unit variance by subtracting the mean and dividing by the standard deviation, a distinct scaling technique alongside normalization, which instead maps values into a fixed range such as 0 to 1.

Why this answer

Standardization (C) is a feature-scaling technique that transforms each feature to have a mean of 0 and a standard deviation of 1 (z-score, (x − μ)/σ), which is exactly what the question asks for. Normalization (Min-Max) (E) is also a feature-scaling method that rescales values into a fixed range, typically [0, 1], via (x − min)/(max − min). One-hot encoding (A) is a categorical-encoding technique that creates binary columns, not a scaling method.

Principal Component Analysis (B) is a dimensionality-reduction technique that projects data onto principal components, not a scaling method. Binning (D) is a discretization technique that groups continuous values into bins, not a scaling method.

Exam trap

AWS often tests the distinction between feature scaling (changing the numeric range of features) and data transformation techniques like encoding or dimensionality reduction, leading candidates to confuse one-hot encoding or PCA with scaling methods.

61
MCQeasy

A team trained a deep learning model that achieves 99% accuracy on training data but only 70% on validation data. What is the most likely issue?

A.Underfitting
B.Overfitting
C.Data leakage
D.Feature scaling
AnswerB

The large gap between 99% training accuracy and 70% validation accuracy shows the model memorised training data rather than learning generalisable patterns. Overfitting is the specific condition producing this divergence, matching the stem's reported metrics exactly.

Why this answer

The model performs exceptionally well on training data (99% accuracy) but significantly worse on validation data (70% accuracy). This large gap indicates the model has memorized the training data, including noise and irrelevant patterns, rather than learning generalizable features — a classic symptom of overfitting.

Exam trap

The AIF-C01 exam often tests the distinction between overfitting and underfitting by presenting a scenario where training accuracy is high but validation accuracy is low, tempting candidates to incorrectly choose underfitting if they focus only on the low validation score.

How to eliminate wrong answers

Option A is wrong because underfitting would show poor performance on both training and validation data, not high training accuracy with low validation accuracy. Option C is wrong because data leakage typically causes both training and validation accuracy to be artificially high, not a large gap between them. Option D is wrong because feature scaling issues would generally affect model convergence or performance uniformly across datasets, not create a specific training-validation accuracy disparity.

62
MCQhard

A company is using Amazon SageMaker to train a large language model with hundreds of billions of parameters. The model does not fit into the memory of a single GPU. Which approach should they use to train the model efficiently?

A.Use a larger instance with more GPU memory, such as p4d.24xlarge
B.Use SageMaker's data parallelism strategy
C.Use SageMaker's model parallelism strategy with the SageMaker distributed training library
D.Reduce the model size by pruning layers until it fits into memory
AnswerC

Model parallelism shards the model's layers and parameters across multiple GPUs, so no single device must hold the full hundreds-of-billions-parameter model. This directly resolves the stem's constraint that the model does not fit into one GPU's memory.

Why this answer

SageMaker's model parallelism strategy with the SageMaker distributed training library is specifically designed for training large models that do not fit into the memory of a single GPU. It partitions the model layers across multiple GPUs, enabling efficient training of models with hundreds of billions of parameters by overlapping computation and communication.

Exam trap

The AIF-C01 exam often tests the distinction between data parallelism and model parallelism, and the trap here is that candidates may confuse data parallelism (which splits data, not the model) as a solution for models that don't fit in memory, when in fact model parallelism is required for such cases.

How to eliminate wrong answers

Option A is wrong because even the largest GPU instances like p4d.24xlarge have limited GPU memory (40 GB per A100 GPU), which is insufficient for a model with hundreds of billions of parameters; scaling vertically is not feasible for such large models. Option B is wrong because SageMaker's data parallelism strategy replicates the entire model on each GPU and splits the data across GPUs, which requires the model to fit into a single GPU's memory; it does not solve the memory constraint issue. Option D is wrong because pruning layers to reduce model size would degrade model quality and is not a practical or efficient approach for training large language models; the goal is to train the full model, not a smaller version.

63
MCQeasy

A data scientist wants to quickly build a supervised learning model for binary classification on a tabular dataset with 10,000 rows and 200 features. The dataset has some missing values and requires minimal code. Which AWS service should the data scientist use?

A.Amazon SageMaker Studio Lab
B.Amazon SageMaker Clarify
C.Amazon SageMaker Autopilot
D.Amazon SageMaker JumpStart
AnswerC

SageMaker Autopilot automates algorithm selection, feature engineering and hyperparameter tuning for tabular classification, handling missing values and returning an explainable model with minimal code. It directly meets the binary classification requirement on the 10,000-row, 200-feature dataset.

Why this answer

Amazon SageMaker Autopilot is the correct choice because it automatically performs data preprocessing (including handling missing values), feature engineering, model selection, and hyperparameter tuning for supervised learning tasks like binary classification. It requires minimal code—users can simply point to a tabular dataset in Amazon S3 and specify the target column, and Autopilot will automatically train and evaluate multiple candidate models, making it ideal for quickly building a binary classifier on a 10,000-row, 200-feature dataset with missing values.

Exam trap

The AIF-C01 exam often tests the distinction between automated ML services (Autopilot) and model hosting or development environments (Studio Lab, JumpStart), so the trap here is that candidates may confuse SageMaker Autopilot with SageMaker JumpStart, thinking JumpStart also automates model building, when in fact JumpStart only provides pre-built models and requires manual configuration.

How to eliminate wrong answers

Option A is wrong because Amazon SageMaker Studio Lab is a free, no-code ML development environment that provides JupyterLab notebooks and limited compute resources, but it does not automate model building or handle missing values—it requires the user to write all code manually. Option B is wrong because Amazon SageMaker Clarify is designed for bias detection, model explainability, and fairness analysis, not for building or training supervised learning models; it cannot handle missing values or perform automated model selection. Option D is wrong because Amazon SageMaker JumpStart provides pre-built models and solutions for transfer learning and fine-tuning, but it does not automatically preprocess missing values or perform automated model selection for tabular binary classification—it requires the user to select and configure a model manually.

64
MCQmedium

A company wants to automatically detect anomalies in server metrics. Which algorithm is most appropriate?

A.XGBoost
B.One-class SVM
C.Linear SVM
D.K-Means
AnswerB

One-class SVM learns a decision boundary describing normal behaviour from only normal samples, flagging deviations as anomalies. Server metrics rarely contain labelled anomalies, so this unsupervised approach suits the scenario better than supervised classifiers requiring labelled attack examples.

Why this answer

One-class SVM is specifically designed for anomaly detection, as it learns a boundary around the normal data points in the feature space and identifies any point falling outside this boundary as an anomaly. This makes it ideal for detecting unusual patterns in server metrics without requiring labeled anomaly examples.

Exam trap

The AIF-C01 exam often tests the distinction between supervised and unsupervised learning, and the trap here is that candidates may choose XGBoost or Linear SVM because they are familiar with them for classification, forgetting that anomaly detection typically requires a one-class approach when only normal data is available.

How to eliminate wrong answers

Option A is wrong because XGBoost is a supervised ensemble learning algorithm used for classification and regression, not for unsupervised anomaly detection; it requires labeled training data and is not designed to identify outliers without prior examples. Option C is wrong because Linear SVM is a supervised binary classifier that separates data into two classes using a hyperplane, and it cannot perform one-class anomaly detection without negative samples. Option D is wrong because K-Means is an unsupervised clustering algorithm that partitions data into clusters based on distance, but it does not inherently detect anomalies; while outliers can be inferred from cluster distances, it is not a dedicated anomaly detection method and lacks the statistical boundary learning of one-class SVM.

65
MCQmedium

A company is using Amazon Rekognition to detect objects in images. They find that the service sometimes mislabels objects. What is the best way to improve accuracy for their specific use case?

A.Use a larger image size
B.Contact AWS support
C.Increase the confidence threshold
D.Use Amazon SageMaker to build a custom model
AnswerD

Rekognition's pre-trained labels cannot be tuned to niche classes, so mislabelling persists. SageMaker lets you train a custom model on your own annotated images, matching the exact object categories and visual conditions of your use case, which directly addresses the accuracy constraint.

Why this answer

Amazon Rekognition is a pre-trained service that may not perform optimally for specialized or domain-specific use cases. By using Amazon SageMaker to build a custom model, you can train a model on your own labeled dataset, which directly addresses the mislabeling issue by tailoring the model to your specific images and objects.

Exam trap

The trap here is that candidates often assume increasing the confidence threshold is a universal fix for accuracy issues, but the AIF-C01 exam tests the understanding that pre-trained services have limitations and that custom training (via SageMaker) is required for domain-specific improvements.

How to eliminate wrong answers

Option A is wrong because using a larger image size does not inherently improve Rekognition's detection accuracy; the service already resizes images to a standard input size, and larger images may only increase processing time without correcting mislabeling. Option B is wrong because contacting AWS support will not modify the underlying pre-trained model or improve its accuracy for your specific use case; support can only assist with service configuration or bugs, not model retraining. Option C is wrong because increasing the confidence threshold reduces false positives but does not fix systematic mislabeling; it may cause the service to return fewer results, potentially missing correct detections, without addressing the root cause of incorrect object identification.

66
MCQmedium

A data science team is using Amazon SageMaker to train multiple models with different hyperparameters. They want to track metrics, compare runs, and reproduce the best result. Which SageMaker feature should they use?

A.SageMaker Model Registry
B.SageMaker Debugger
C.SageMaker Autopilot
D.SageMaker Experiments
AnswerD

SageMaker Experiments groups training runs into experiments and trials, automatically logging hyperparameters, metrics and artefacts for each job. This directly satisfies the team's need to track metrics across runs, compare them side by side, and reproduce the best result by retrieving its exact configuration.

Why this answer

SageMaker Experiments is the correct feature because it is specifically designed to track, organize, and compare machine learning training runs (trials) with different hyperparameters and metrics. It allows data scientists to log parameters, metrics, and artifacts for each run, compare results across runs, and retrieve the exact configuration needed to reproduce the best-performing model.

Exam trap

The trap here is that candidates often confuse SageMaker Experiments with SageMaker Model Registry, mistakenly thinking that model versioning and run tracking are the same feature, when in fact Experiments focuses on the iterative training process and Registry focuses on the final model lifecycle.

How to eliminate wrong answers

Option A is wrong because SageMaker Model Registry is a catalog for managing and versioning trained models, not for tracking and comparing individual training runs or hyperparameter experiments. Option B is wrong because SageMaker Debugger monitors training jobs in real time for issues like vanishing gradients or overfitting, but it does not provide a structured way to log, compare, or reproduce runs with different hyperparameters. Option C is wrong because SageMaker Autopilot automatically explores different algorithms and hyperparameters to find the best model, but it does not give the team the ability to manually track, compare, and reproduce their own custom runs with specific hyperparameters.

67
MCQmedium

A retailer wants to group its customers into distinct behavioral segments for targeted marketing, but it has no predefined segment labels and no historical outcomes to learn from. Which machine learning approach should the retailer use?

A.Clustering
B.Reinforcement learning
C.Regression
D.Supervised classification
AnswerA

Clustering is an unsupervised technique that groups similar records without predefined labels. The retailer only has customer attributes and behavior, and wants to discover natural segments, so an algorithm such as k-means can partition customers by similarity. This directly matches the goal of finding structure in unlabeled data.

Why this answer

Because no segment labels exist and the objective is to discover natural groupings among customers, this is an unsupervised learning problem best solved with clustering. Classification, regression, and reinforcement learning all depend on labels, targets, or reward signals that the retailer does not have, so clustering is the only approach that fits the described situation.

Exam trap

The trap here is reaching for classification because the business wants 'segments', when the absence of predefined labels is precisely what makes this an unsupervised clustering problem.

68
MCQmedium

A retail company wants to forecast daily product demand for the next quarter. They have three years of historical sales data that includes seasonal spikes and promotional periods, and they want a fully managed AWS service that can automatically train and tune a forecasting model without writing deep learning code. Which AWS service best fits this requirement?

A.Amazon Comprehend
B.Amazon Polly
C.Amazon Rekognition
D.Amazon Forecast
AnswerD

Amazon Forecast is a fully managed time-series forecasting service that ingests historical demand data with related features such as promotions and holidays, then automatically selects and tunes algorithms like DeepAR+ and CNN-QR. It directly matches the requirement of no custom deep learning code and handles seasonality, making it the correct fit for quarterly demand forecasting.

Why this answer

The scenario calls for a managed service that learns from historical numeric sales data containing seasonality and promotions. Amazon Forecast is purpose-built for this, automatically training and tuning time-series models and exposing forecasts without custom deep learning code. The other services address computer vision, speech, and text analytics, none of which produce demand forecasts.

Exam trap

The trap here is assuming any fully managed AI service can handle forecasting, when forecasting requires a dedicated time-series service.

69
Multi-Selectmedium

A data scientist is preparing data for a classification task. Which TWO techniques are commonly used for handling missing values? (Choose two.)

Select 2 answers
A.Label encoding
B.Normalization
C.Imputing with mean
D.Dropping rows with any missing values
E.One-hot encoding
AnswersC, D

Imputing with mean replaces missing numerical entries with the column average, preserving dataset size and avoiding dropped rows. This directly satisfies the stem's requirement for a common missing-value technique, since mean imputation is a standard preprocessing method in classification workflows.

Why this answer

Option C (Imputing with mean) is correct because replacing missing numeric entries with the column mean is a standard, simple imputation technique that preserves the dataset size and avoids discarding information. Option D (Dropping rows with any missing values) is also correct because listwise deletion is a common, straightforward approach to handling missing data, especially when the missingness is minimal or random. The other options do not address missing values: A (Label encoding) converts categorical labels into integer codes, B (Normalization) rescales numeric feature ranges, and E (One-hot encoding) creates binary indicator columns for categories—all are preprocessing steps for existing values, not missing-data handling.

Exam trap

The AIF-C01 exam often tests the distinction between data preprocessing techniques (e.g., encoding, scaling) and missing value handling, so candidates mistakenly select label encoding or normalization because they are common preprocessing steps, even though they do not address missing data.

70
Multi-Selectmedium

A logistics company is planning its first machine learning project and the leadership team asks which statements correctly describe fundamental machine learning concepts. Which TWO statements are accurate? (Choose two.)

Select 2 answers
A.In supervised learning, the model learns from input-output pairs where the desired output is provided during training.
B.A machine learning model's predictions are deterministic rules written by engineers rather than patterns inferred from data.
C.Reinforcement learning requires a fully labeled dataset of correct actions for every possible situation before training can begin.
D.Model training always improves accuracy on unseen data as more epochs are run, so training should continue until loss reaches zero.
E.Unsupervised learning can identify patterns or groupings in data that has no predefined labels.
AnswersA, E

Supervised learning depends on labeled examples that pair each input with the correct output, allowing the algorithm to adjust its parameters to minimize error against those known targets. This is exactly how tasks such as delivery-time prediction or package damage classification are trained. Without the provided outputs, the model would have no error signal to learn from, so this statement correctly captures the core of supervised learning.

Why this answer

Supervised learning is defined by training on input-output pairs with known targets, while unsupervised learning finds structure in data without any labels. Both statements describe core, vendor-neutral machine learning concepts that a logistics team should understand before scoping its first project.

Exam trap

The trap here is accepting the common myth that more training always helps, or confusing reinforcement learning with supervised learning that needs labeled correct actions.

71
MCQmedium

A company uses Amazon SageMaker to train a model. The training job fails with 'InsufficientInstanceCapacity' error. What is the most likely cause?

A.The request rate is too high.
B.The dataset size exceeds the instance storage limit.
C.The requested instance type is not available in the specified region.
D.The training image is not compatible with the instance type.
AnswerC

SageMaker provisions the requested instance type from capacity within the chosen region; when that pool is exhausted, the training job fails with InsufficientInstanceCapacity. The constraint is regional instance availability, not IAM permissions, data format or hyperparameters, so switching instance type or region resolves it.

Why this answer

The 'InsufficientInstanceCapacity' error in Amazon SageMaker indicates that AWS does not currently have enough available capacity for the requested instance type in the specified region or Availability Zone. This is a common transient error when demand for a particular instance type exceeds supply, and it is not related to request rate, dataset size, or image compatibility.

Exam trap

The AIF-C01 exam often tests the distinction between capacity errors and throttling errors, so the trap here is confusing 'InsufficientInstanceCapacity' with a rate-limiting or quota error, leading candidates to incorrectly select Option A.

How to eliminate wrong answers

Option A is wrong because 'InsufficientInstanceCapacity' is a capacity error, not a throttling error; throttling (e.g., from high request rate) would return a 'ThrottlingException' or 'RequestLimitExceeded' error. Option B is wrong because dataset size exceeding instance storage limits would cause an 'OutOfMemory' or 'DiskFull' error, not a capacity error. Option D is wrong because image compatibility issues would result in a 'ClientError' or 'ImageNotFoundException', not an instance capacity error.

72
MCQhard

A financial services company uses a machine learning model to approve loan applications. The model is a gradient boosting classifier trained on historical loan data. Recently, the company noticed that the model's approval rate for applicants from a certain demographic group is significantly lower than for other groups, even though the model's overall accuracy remains high. The data science team has been asked to address this potential bias while minimizing the impact on overall model performance. The team has access to the training data and the trained model. They have limited time and budget. Which course of action should the team take first?

A.Remove the sensitive attribute from the training data and retrain the model.
B.Collect more data from the under-represented demographic group and retrain the model.
C.Analyze the training data for bias and retrain the model using bias mitigation techniques such as reweighting.
D.Adjust the model's decision threshold for the affected group after deployment.
AnswerC

Inspecting the training data for demographic imbalance or proxy features reveals the bias source before any modelling change. Reweighting or resampling then corrects the disparity at its root, which is cheaper and faster than post-hoc threshold tuning or full model replacement.

Why this answer

The first step should be to analyze the training data for bias and then retrain the model using bias mitigation techniques such as reweighting, because this addresses the root cause of the disparity while aiming to preserve overall performance. Reweighting adjusts sample weights to reduce the influence of biased historical patterns, and it can be applied with limited time and budget since it works with existing data and model. This is a targeted, cost-effective first action before considering more expensive data collection or post-deployment adjustments.

Exam trap

AIF-C01 often tests the misconception that removing the sensitive attribute eliminates bias; candidates must recognize that proxy variables and historical bias persist, and that data analysis plus mitigation techniques is the correct first step.

How to eliminate wrong answers

Option A is wrong because simply removing the sensitive attribute does not eliminate bias — the model can still learn proxy features correlated with the sensitive attribute, and it may reduce accuracy without fixing the disparity. Option B is wrong because collecting more data from the under-represented group is expensive and time-consuming, and the question states the team has limited time and budget. Option D is wrong because adjusting the decision threshold after deployment is a post-hoc fix that can mask bias but does not address the underlying model or data issues, and it may not be sustainable or compliant.

73
Multi-Selecthard

A company is training a deep learning model for image classification. Which THREE practices help reduce overfitting? (Choose three.)

Select 3 answers
A.L2 regularization
B.Increasing model depth
C.Increasing learning rate
D.Dropout
E.Data augmentation
AnswersA, D, E

L2 regularization penalises large weights by adding a squared-magnitude term to the loss function, constraining the model's capacity to memorise training noise. This directly satisfies the stem's requirement to reduce overfitting in deep image classification, since simpler weight distributions generalise better to unseen images.

Why this answer

L2 regularization (option A) is correct because it adds a penalty proportional to the squared magnitude of the weights to the loss function, discouraging large weights and thus reducing the model's ability to memorize training noise. Dropout (option D) is correct because it randomly deactivates a fraction of neurons during each training step, forcing the network to learn redundant, more robust feature representations instead of relying on specific units. Data augmentation (option E) is correct because it synthetically expands the training set by applying label-preserving transformations (e.g., flips, rotations, crops) to images, exposing the model to more varied inputs and improving generalization.

Increasing model depth (option B) is not correct here because adding layers increases capacity, which typically makes overfitting worse rather than better. Increasing the learning rate (option C) is not correct because it only affects optimization speed and stability, and an excessively high rate can cause divergence or poor convergence, not reduced overfitting.

Exam trap

The AIF-C01 exam often tests the misconception that increasing model complexity (depth) or tuning the learning rate can mitigate overfitting, when in fact these changes either exacerbate the problem or address unrelated training dynamics.

74
MCQmedium

A team has built a regression model to predict house prices. The RMSE is 50,000 on the test set. Which action is most appropriate to improve model performance?

A.Remove outliers from training data
B.Apply feature scaling
C.Add more relevant features
D.Use a different evaluation metric
AnswerC

An RMSE of 50,000 indicates underfitting or missing signal, so adding relevant features gives the regression model additional predictive information. This addresses the performance gap more directly than hyperparameter tuning or collecting more rows of the same variables.

Why this answer

Adding more relevant features can provide the model with additional predictive signals, potentially reducing bias and lowering RMSE if the new features have genuine correlation with house prices. Since RMSE is already 50,000, the model may be underfitting due to insufficient input variables, and enriching the feature set is a direct way to capture more variance in the target variable.

Exam trap

The AWS AI Practitioner exam often tests the misconception that data preprocessing steps like scaling or outlier removal are universal fixes for high error, when in fact the most appropriate first step for a high RMSE in regression is to improve the feature set to address underfitting.

How to eliminate wrong answers

Option A is wrong because removing outliers from training data can reduce variance but may also discard valuable extreme cases that reflect real market conditions, and it does not address the core issue of underfitting or missing predictive signals. Option B is wrong because feature scaling (e.g., normalization or standardization) is important for gradient-based optimization and distance-based algorithms, but it does not inherently improve model accuracy for tree-based or linear regression models when RMSE is already computed on unscaled data; it affects convergence speed, not predictive power. Option D is wrong because changing the evaluation metric (e.g., from RMSE to MAE or R²) does not improve the model's actual predictive performance; it only changes how performance is measured, potentially masking the same underlying error.

75
MCQeasy

A company wants to predict customer churn. They have historical data with features like usage minutes, support tickets, contract length. The target is binary: churn/not churn. Which ML algorithm is best suited?

A.Logistic regression
B.Principal Component Analysis (PCA)
C.Linear regression
D.K-means clustering
AnswerA

Logistic regression outputs a probability between 0 and 1 via the sigmoid function, fitting the binary churn/not-churn target directly. It also yields interpretable coefficients across usage minutes, support tickets and contract length, so the company can see which features drive churn.

Why this answer

Logistic regression is the best choice because it is specifically designed for binary classification tasks like predicting churn (churn/not churn). It models the probability of the target class using a logistic (sigmoid) function, making it interpretable and efficient for this type of supervised learning problem with a categorical outcome.

Exam trap

The AIF-C01 exam often tests the distinction between supervised and unsupervised learning, and the trap here is that candidates may confuse dimensionality reduction (PCA) or clustering (K-means) with classification, or mistakenly apply linear regression to a binary outcome without recognizing the need for a logistic function.

How to eliminate wrong answers

Option B is wrong because Principal Component Analysis (PCA) is an unsupervised dimensionality reduction technique, not a classification algorithm; it reduces feature space but does not predict a binary target. Option C is wrong because linear regression predicts a continuous numeric output, not a binary class; using it for classification would violate the assumption of normally distributed errors and produce unbounded predictions. Option D is wrong because K-means clustering is an unsupervised learning algorithm used for grouping unlabeled data into clusters, not for predicting a known binary target variable.

Page 1 of 2 · 90 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Fundamentals of AI and ML questions.