Courseiva

AIF-C01 · domain

Fundamentals of AI and ML

This domain covers core AI, ML, and generative AI concepts plus the AWS services that implement them. Questions ask you to classify learning types, pick the right AWS AI service for a task, choose feature engineering or data preparation techniques, and match SageMaker tools to a described scenario. Expect scenario-based multiple choice and multi-select.

90 questions27 easy39 medium24 hard

Focused practice

Practice Fundamentals of AI and ML questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Fundamentals of AI and ML

You must map a described business problem to the correct AWS AI/ML service and the correct data or feature technique. The single most important thing: read the modality and data shape first, because they eliminate most wrong answers immediately.

Selecting Amazon Comprehend, Transcribe, Polly, Translate, or Rekognition for a stated NLP or vision task

Choosing SageMaker built-in algorithms, Autopilot, or JumpStart for tabular classification with minimal code

Applying feature engineering: PCA for redundancy, target or frequency encoding for high-cardinality categoricals

Distinguishing supervised, unsupervised, reinforcement, and generative learning and their AWS use cases

Watch out for

Common Fundamentals of AI and ML exam traps

  • ▸Confusing Comprehend (NLP text analysis) with Transcribe (speech-to-text) or Polly (text-to-speech) when the scenario names the modality.
  • ▸Picking one-hot encoding for a categorical feature with tens of thousands of unique values, which explodes dimensionality.
  • ▸Assuming SageMaker Autopilot or JumpStart trains custom deep learning models when the task is simple tabular binary classification.

Question index

All Fundamentals of AI and ML questions (90)

Click any question to see the full explanation, or start a practice session above.

1

A hospital wants to build a model that predicts whether a patient has a specific disease based on labeled historical medical records where each record is marked either positive or negative. Which type of machine learning problem does this represent?

Easy
2

A data scientist trains a model to predict whether a loan applicant will default. After deployment, the model performs well on applicants similar to the training data but poorly on applicants from a newly added geographic region that was underrepresented in training. Which statement best describes the underlying problem?

Hard
3

A company is building a model to detect fraudulent transactions. The dataset has 1,000,000 transactions, of which only 1,000 are fraudulent. The team wants to evaluate the model's performance. Which metric is most appropriate to use as the primary evaluation metric?

Hard
4

A machine learning team notices their model performs excellently on the training dataset but poorly on new, unseen data. They want to reduce this gap without collecting more data. Which action most directly addresses the problem?

Hard
5

A financial services firm wants an internal assistant that answers employee questions about company travel policy in natural language. The policy documents change frequently, and the firm requires answers to cite the specific policy section used. Which approach BEST meets these requirements?

Hard
6

A data scientist is using SageMaker to train a model on a dataset with many features. They suspect some features are redundant. Which feature engineering technique would help?

Medium
7

A data science team needs to choose a machine learning approach for a project that requires predicting customer churn based on historical data. The team has a labeled dataset with 10,000 records and needs to interpret the model's decisions to provide business insights. Which machine learning technique should the team prioritize?

Medium
8

A financial institution wants to predict whether a loan applicant will default. They have a historical dataset with loan outcomes (default or no default) and various applicant features. Which type of machine learning should they use?

Medium
9

An ML team is deploying a real-time inference endpoint for a computer vision model using Amazon SageMaker. The model requires GPU acceleration for low latency. Which instance type should the team choose to minimize cost while meeting the GPU requirement?

Medium
10

Which TWO are best practices for model monitoring in production on AWS?

Hard
11

A data science team is preparing a dataset for training a machine learning model. They need to perform data preprocessing to improve model performance. Which TWO of the following are common data preprocessing techniques? (Choose two.)

Medium
12

Refer to the exhibit. A data scientist ran a training job on Amazon SageMaker. The job failed with the error shown. What is the most likely cause?

Easy
13

A data scientist wants to deploy a custom model built with TensorFlow to Amazon SageMaker for real-time inference. Which TWO steps are required? (Choose two.)

Easy
14

A startup wants to build a product recommendation engine for their e-commerce platform. They have user purchase history and item metadata. They want a fully managed solution that can automatically train and deploy a recommendation model without needing to manage the underlying ML lifecycle. The solution should provide personalized recommendations based on collaborative filtering. Which AWS service should they use?

Easy
15

A data scientist is training a binary classification model to predict customer churn. The dataset has 10,000 records with 9,500 non-churners and 500 churners. After training a logistic regression model, the model achieves 95% accuracy on the test set. However, the business team reports that the model is not useful because it predicts almost all customers as non-churners. Which metric should the data scientist use to evaluate the model's performance in this scenario?

Medium
16

A fintech startup is preparing its first machine learning project to detect fraudulent card transactions. The team must decide which characteristics make a problem well suited to supervised learning. Which TWO characteristics indicate that supervised learning is appropriate? (Choose two.)

Hard
17

A media company wants to build a system that generates short promotional descriptions for its articles. The team has no labeled dataset of article-summary pairs but has a large corpus of published articles and descriptions. They want to leverage a pretrained foundation model and adapt it to their domain with minimal labeling effort. Which approach best fits this scenario?

Hard
18

A data scientist is training a model using Amazon SageMaker and notices the training loss is decreasing but validation loss starts increasing after a few epochs. Which technique should they apply to address this?

Medium
19

An e-commerce company stores user interaction logs in Amazon S3. They want to use machine learning to segment users based on purchasing behavior. Which unsupervised learning algorithm is most appropriate?

Hard
20

A media company has millions of customer support emails but no labels indicating topic or sentiment. The company wants to discover natural groupings of emails and reduce dimensionality before further analysis. Which approach should they use?

Medium
21

A financial services company needs to ensure that the machine learning models used for loan approval are explainable and meet regulatory compliance. Which AWS feature can help explain model predictions?

Medium
22

A media company stores 40 TB of raw video footage in Amazon S3 and wants to automatically detect scene boundaries, identify on-screen text, and flag unsafe frames without building custom computer vision models. Which AWS service should they use?

Easy
23

A company wants to build a model to forecast monthly sales. The data is a time series with trend and seasonality. Which SageMaker algorithm is most appropriate?

Medium
24

A team is training a binary classification model using Amazon SageMaker. They notice that the training accuracy is 99% but the test accuracy is only 70%. Which technique should they apply first to address this?

Medium
25

A company wants to use Amazon SageMaker to train a model using a custom Docker container that has specific dependencies. The training code is stored in an S3 bucket. Which steps must be taken to run the training job?

Medium
26

A data scientist is evaluating different AWS services for building a machine learning pipeline. Which THREE components are part of Amazon SageMaker? (Select THREE.)

Medium
27

A company is using Amazon Comprehend for sentiment analysis on customer reviews. They notice that the sentiment is often incorrect for negative reviews with sarcasm. What is the likely cause?

Easy
28

An organization wants to detect anomalies in real-time streaming data from IoT devices. The data includes sensor readings, and the team plans to use a machine learning model. Which AWS service should be used to build and deploy the model with minimal operational overhead?

Medium
29

A data scientist is using Amazon SageMaker to train a deep learning model for image classification. The training job is taking too long. The dataset consists of 100,000 images stored in Amazon S3. Which action can the data scientist take to reduce training time without modifying the model architecture?

Medium
30

A data scientist at a retail company is tasked with building a model to predict customer churn. The dataset contains 100,000 records with features such as age, purchase history, customer support interactions, and a binary label indicating whether the customer churned in the past. The team needs a model that can be deployed for real-time inference with low latency. They have limited time and want to use a built-in algorithm from Amazon SageMaker that is optimized for classification tasks. Which approach should they take?

Easy
31

A retail company wants to build a system that predicts next month's sales for each of its 500 stores based on historical sales, local holidays, and marketing spend. The target values are continuous dollar amounts, and the company has labeled historical data for every store. Which type of machine learning problem does this represent?

Easy
32

Which TWO of the following are examples of supervised learning tasks that can be performed using Amazon SageMaker built-in algorithms?

Medium
33

A healthcare provider wants to predict which patients are likely to be readmitted within 30 days. The historical dataset has 12,000 admissions and includes age, diagnosis codes, length of stay, and prior admissions. The team has limited machine learning experience and needs an explainable model that shows which factors drove each prediction. Which AWS approach is most appropriate?

Hard
34

A company wants to deploy a real-time inference endpoint for a custom model on SageMaker. The model has high latency (100ms) and they need to handle variable traffic with spikes. Which deployment strategy is most cost-effective?

Hard
35

During model training, the loss decreases rapidly for the first few epochs and then plateaus. The validation loss starts increasing after some epochs. What should the team do to improve generalization?

Medium
36

A retail company uses a machine learning model to forecast daily product demand. The model is a time series model that uses historical sales data. The model has been performing well, but recently the forecasts have been consistently too low, leading to stockouts. The data scientist notices that the model was trained on data up to last year, and the company has since launched a successful marketing campaign that increased sales by 20%. The data scientist needs to update the model to reflect the new sales patterns. Which approach should the data scientist take?

Medium
37

A data scientist wants to perform automatic model tuning (hyperparameter optimization) on SageMaker. They need to find the best hyperparameters for a gradient boosting model. Which strategy is BEST for this task?

Hard
38

A startup is building a recommendation engine for their e-commerce platform. They need a fully managed service that can generate personalized product recommendations based on user behavior. Which AWS service should they use?

Easy
39

A data scientist needs to preprocess categorical data with high cardinality (e.g., zip code with 50,000 unique values). Which technique is most appropriate?

Hard
40

Which TWO services can be used to preprocess data for machine learning in AWS? (Choose two.)

Easy
41

A company is using Amazon Rekognition to detect objects in images. They need to detect custom objects that are specific to their domain. What should they do?

Medium
42

A company needs to store large amounts of unstructured training data (images, videos) in a cost-effective manner while ensuring low-latency retrieval for training jobs running on Amazon SageMaker. Which storage solution should be used?

Easy
43

A company wants to build a system that automatically categorizes customer support tickets into predefined categories (e.g., billing, technical, account). The team has a large dataset of historical tickets with their category labels. Which type of machine learning problem is this?

Easy
44

A healthcare company is using Amazon SageMaker to deploy a model that makes predictions on patient data. They need to ensure that the model's predictions are explainable to comply with regulations. Which approach should they take?

Hard
45

A marketing agency wants to analyze customer feedback from social media posts to gauge sentiment. They have no labeled data and limited ML expertise. The team needs a managed service that provides pre-trained models for sentiment analysis without requiring them to train or manage infrastructure. They also need to process text in multiple languages. Which AWS service should they use?

Medium
46

Which THREE are SageMaker built-in algorithms suitable for regression tasks?

Medium
47

A data scientist wants to host a pre-trained model on Amazon SageMaker for real-time inference with minimal latency. Which approach should they use?

Easy
48

A retail bank has millions of unlabeled customer transaction records and wants to discover natural groupings of spending behavior without defining any categories in advance. The data science team plans to use an unsupervised learning approach. Which technique is designed for this goal?

Medium
49

A data engineer is using Amazon SageMaker Data Wrangler to prepare tabular data for ML. Which THREE data transformations are natively supported? (Choose three.)

Hard
50

Which AWS services can be used to build, train, and deploy custom machine learning models? (Choose two.)

Medium
51

A financial services company is preparing to train a machine learning model on customer transaction data. The data science team must address data quality concerns before training, because poor data directly harms model performance. Which TWO practices best improve the quality of the training data? (Choose two.)

Hard
52

Which TWO techniques are commonly used to prevent overfitting in machine learning models? (Select TWO.)

Medium
53

An ML engineer wants to store training data in a format optimized for linear data scanning and columnar access in SageMaker. Which format is most appropriate?

Hard
54

A hospital wants to build a system that automatically assigns a specialty department (for example, Cardiology, Neurology, or Orthopedics) to each free-text patient referral note. The hospital has a large archive of past referral notes that were already labeled by clinicians with the correct department. Which type of machine learning problem does this scenario describe?

Easy
55

Which metric is most appropriate for evaluating a classification model when false positives are costly?

Easy
56

A junior data scientist is building a model to classify incoming customer support tickets into one of eight predefined categories such as Billing, Shipping, or Returns. Historical tickets already have correct category labels. Which type of machine learning is being used?

Easy
57

A company wants to use AI to automatically transcribe customer service calls into text. Which AWS service is most suitable?

Easy
58

A company is using Amazon Fraud Detector to detect fraudulent transactions. Which TWO actions can be taken to improve model accuracy? (Select TWO.)

Hard
59

A company wants to automatically detect anomalies in their AWS CloudTrail logs to identify potential security threats. Which AWS service is specifically designed for this purpose?

Easy
60

Which TWO of the following are types of feature scaling?

Easy
61

A team trained a deep learning model that achieves 99% accuracy on training data but only 70% on validation data. What is the most likely issue?

Easy
62

A company is using Amazon SageMaker to train a large language model with hundreds of billions of parameters. The model does not fit into the memory of a single GPU. Which approach should they use to train the model efficiently?

Hard
63

A data scientist wants to quickly build a supervised learning model for binary classification on a tabular dataset with 10,000 rows and 200 features. The dataset has some missing values and requires minimal code. Which AWS service should the data scientist use?

Easy
64

A company wants to automatically detect anomalies in server metrics. Which algorithm is most appropriate?

Medium
65

A company is using Amazon Rekognition to detect objects in images. They find that the service sometimes mislabels objects. What is the best way to improve accuracy for their specific use case?

Medium
66

A data science team is using Amazon SageMaker to train multiple models with different hyperparameters. They want to track metrics, compare runs, and reproduce the best result. Which SageMaker feature should they use?

Medium
67

A retailer wants to group its customers into distinct behavioral segments for targeted marketing, but it has no predefined segment labels and no historical outcomes to learn from. Which machine learning approach should the retailer use?

Medium
68

A retail company wants to forecast daily product demand for the next quarter. They have three years of historical sales data that includes seasonal spikes and promotional periods, and they want a fully managed AWS service that can automatically train and tune a forecasting model without writing deep learning code. Which AWS service best fits this requirement?

Medium
69

A data scientist is preparing data for a classification task. Which TWO techniques are commonly used for handling missing values? (Choose two.)

Medium
70

A logistics company is planning its first machine learning project and the leadership team asks which statements correctly describe fundamental machine learning concepts. Which TWO statements are accurate? (Choose two.)

Medium
71

A company uses Amazon SageMaker to train a model. The training job fails with 'InsufficientInstanceCapacity' error. What is the most likely cause?

Medium
72

A financial services company uses a machine learning model to approve loan applications. The model is a gradient boosting classifier trained on historical loan data. Recently, the company noticed that the model's approval rate for applicants from a certain demographic group is significantly lower than for other groups, even though the model's overall accuracy remains high. The data science team has been asked to address this potential bias while minimizing the impact on overall model performance. The team has access to the training data and the trained model. They have limited time and budget. Which course of action should the team take first?

Hard
73

A company is training a deep learning model for image classification. Which THREE practices help reduce overfitting? (Choose three.)

Hard
74

A team has built a regression model to predict house prices. The RMSE is 50,000 on the test set. Which action is most appropriate to improve model performance?

Medium
75

A company wants to predict customer churn. They have historical data with features like usage minutes, support tickets, contract length. The target is binary: churn/not churn. Which ML algorithm is best suited?

Easy
76

A financial services firm must build a model that flags potentially fraudulent card transactions in under 200 milliseconds while keeping all data inside its own Amazon VPC. The fraud team has thousands of labeled historical transactions and the pattern changes slowly over months. Which approach best balances latency, data residency, and the need for periodic retraining?

Medium
77

A startup needs to predict customer churn based on historical data containing labels (churned or not). Which type of machine learning should they use?

Easy
78

A logistics company wants to use machine learning to predict delivery times. A data scientist is preparing the project and must identify which characteristics describe supervised learning rather than unsupervised learning. (Choose two.)

Medium
79

A social media company needs to automatically detect and flag toxic comments in multiple languages. They have a large stream of user comments and require real-time moderation. Which AWS service is best suited for this task?

Easy
80

A company is deploying a machine learning model for real-time fraud detection. The model must have latency under 100ms. Which infrastructure choice is most appropriate?

Hard
81

A hospital wants an AI system that reviews chest X-ray images and flags those likely to contain pneumonia so radiologists can prioritize their queue. The hospital has 40,000 historical X-rays, each already labeled by a radiologist as pneumonia present or absent. Which learning approach best fits this scenario?

Medium
82

A financial institution is deploying a fraud detection model using Amazon SageMaker. The model must be able to handle sudden spikes in inference requests during promotional events while keeping costs low. The team wants to use a serverless architecture to avoid provisioning idle capacity and to scale automatically from zero. However, the inference latency requirement is under 5 seconds for each request. Which SageMaker inference option should they choose?

Hard
83

A team is evaluating a classification model. The confusion matrix shows: TP=80, FN=20, FP=10, TN=90. What is the precision?

Easy
84

A startup is preparing a dataset to train a supervised machine learning model and wants to follow sound data preparation practices. Which TWO activities are appropriate parts of preparing training data? (Choose two.)

Medium
85

A startup with limited ML expertise wants to quickly prototype a binary classification model using a small customer dataset. They need a managed environment to run Jupyter notebooks and access pre-built algorithms. Which AWS service should they choose?

Easy
86

A hospital wants to detect pneumonia from chest X-ray images. Radiologists have already labeled thousands of past X-rays as either 'pneumonia' or 'no pneumonia'. The hospital wants a model that generalizes to new X-rays. Which type of machine learning task is this?

Medium
87

A media company stores thousands of hours of unlabeled video footage and wants to build a searchable index that lets editors retrieve clips by describing their content in natural language. The team has no annotated dataset and no budget to label one. Which machine learning approach is the MOST appropriate starting point?

Hard
88

Which TWO of the following are best practices for preparing training data for a machine learning model?

Hard
89

A company wants to use AWS services to process natural language text. Which TWO AWS services provide natural language processing (NLP) capabilities? (Select TWO.)

Easy
90

Which THREE statements about Amazon SageMaker Ground Truth are correct? (Choose three.)

Medium

Frequently asked questions

What does the Fundamentals of AI and ML domain cover on the AIF-C01 exam?
You must map a described business problem to the correct AWS AI/ML service and the correct data or feature technique. The single most important thing: read the modality and data shape first, because they eliminate most wrong answers immediately.
How many questions are in this domain?
This page lists all 90 Fundamentals of AI and ML questions in the AIF-C01 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Fundamentals of AI and ML questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
aws-ai-practitioner AWS-AI-PRACTITIONER ai ml fundamentals Practice Questions