Courseiva

AIF-C01 · topic practice

Fundamentals of AI and ML practice questions

This domain covers core AI, ML, and generative AI concepts plus the AWS services that implement them. Questions ask you to classify learning types, pick the right AWS AI service for a task, choose feature engineering or data preparation techniques, and match SageMaker tools to a described scenario. Expect scenario-based multiple choice and multi-select.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Fundamentals of AI and ML

What the exam tests

What to know about Fundamentals of AI and ML

You must map a described business problem to the correct AWS AI/ML service and the correct data or feature technique. The single most important thing: read the modality and data shape first, because they eliminate most wrong answers immediately.

Selecting Amazon Comprehend, Transcribe, Polly, Translate, or Rekognition for a stated NLP or vision task

Choosing SageMaker built-in algorithms, Autopilot, or JumpStart for tabular classification with minimal code

Applying feature engineering: PCA for redundancy, target or frequency encoding for high-cardinality categoricals

Distinguishing supervised, unsupervised, reinforcement, and generative learning and their AWS use cases

Watch out for

Common Fundamentals of AI and ML exam traps

  • ▸Confusing Comprehend (NLP text analysis) with Transcribe (speech-to-text) or Polly (text-to-speech) when the scenario names the modality.
  • ▸Picking one-hot encoding for a categorical feature with tens of thousands of unique values, which explodes dimensionality.
  • ▸Assuming SageMaker Autopilot or JumpStart trains custom deep learning models when the task is simple tabular binary classification.

Practice set

Fundamentals of AI and ML questions

20 questions · select your answer, then reveal the explanation

A company is training a deep learning model on Amazon SageMaker using a custom Docker container. The training job fails with the error 'CannotStartContainerError: API error (500): failed to create shim task'. The team verifies that the container image is compatible with the selected instance type. What is the most likely cause of this error?

A machine learning engineer is using Amazon SageMaker to train a model and wants to automatically stop the training job if the loss does not improve for 10 consecutive epochs. Which SageMaker feature should be used?

During a SageMaker training job, the data scientist observes that the loss is not decreasing after the initial few epochs. The model is a deep neural network with ReLU activations. Which hyperparameter adjustment is most likely to help?

Which TWO factors should be considered when choosing between a CPU-based instance and a GPU-based instance for training a machine learning model on Amazon SageMaker? (Choose two.)

Refer to the exhibit. A data scientist attaches the above IAM policy to a SageMaker notebook instance role. The notebook is in the same AWS account as the S3 bucket. When trying to read a file from 's3://my-bucket/training/data.csv', the data scientist gets an Access Denied error. What is the most likely cause?

Exhibit

Refer to the exhibit.

```json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "s3:GetObject",
        "s3:PutObject"
      ],
      "Resource": "arn:aws:s3:::my-bucket/training/*"
    }
  ]
}
```

Refer to the exhibit. A data scientist is training a neural network model on SageMaker. The training log shows the loss values per epoch. Which issue is most likely occurring?

Exhibit

Refer to the exhibit.

```
2023-09-15 10:15:30,123 INFO     - Training job started
2023-09-15 10:15:35,456 INFO     - Epoch 1/10: loss=2.3456, accuracy=0.1234
2023-09-15 10:15:40,789 INFO     - Epoch 2/10: loss=2.3001, accuracy=0.1300
2023-09-15 10:15:46,012 INFO     - Epoch 3/10: loss=2.2800, accuracy=0.1350
2023-09-15 10:15:51,234 INFO     - Epoch 4/10: loss=2.3100, accuracy=0.1280
2023-09-15 10:15:56,456 WARNING - Loss increased from 2.2800 to 2.3100
2023-09-15 10:16:01,678 INFO     - Epoch 5/10: loss=2.3500, accuracy=0.1200
```

A company is deploying a machine learning model for real-time fraud detection. The model must make predictions with latency under 10 milliseconds. The data scientist trained a gradient boosting model that achieves high accuracy but has inference latency of 50 milliseconds. The team has access to a larger instance type with more CPU cores. Which approach should the data scientist take to reduce inference latency while maintaining accuracy?

A data scientist is preparing data for a machine learning model. What is the purpose of splitting the data into training, validation, and test sets?

A team trains a model using Amazon SageMaker built-in XGBoost. After training, they want to evaluate feature importance. Which SageMaker feature allows them to view this?

A data scientist is working with a dataset that contains both numerical and categorical features. Which algorithm is commonly used for regression tasks in AWS SageMaker?

A team is training a deep learning model using Horovod distributed training on SageMaker. They observe that the loss stops decreasing after a few epochs. Which technique should they implement to reduce overfitting?

A company is using Amazon SageMaker to train a model. They want to automatically stop training if the model performance stops improving on a validation dataset. Which SageMaker feature should they enable?

A company is training a deep learning model on Amazon SageMaker using a large dataset stored in S3. Training jobs are frequently failing with 'OutOfMemoryError'. The training algorithm uses PyTorch. How should the data scientist solve this without reducing model accuracy?

A financial services company needs to deploy a real-time fraud detection model with sub-100ms inference latency. The model is a large ensemble requiring 8 GB of memory per request. The workload has bursty traffic. Which Amazon SageMaker deployment strategy best meets these requirements?

A company wants to use Amazon SageMaker Ground Truth to build a labeled dataset for a custom object detection model. Which TWO labeling strategies are available? (Choose two.)

Refer to the exhibit. A SageMaker training job fails with an 'AccessDenied' error when trying to read files from the S3 bucket 'my-training-data'. The IAM role used by the training job has the policy shown. What is the most likely reason for the failure?

Exhibit

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "s3:GetObject",
      "Resource": "arn:aws:s3:::my-training-data/*"
    }
  ]
}

Refer to the exhibit. A SageMaker real-time endpoint is experiencing increasing latency and memory errors after running for a few hours. What is the most likely cause and recommended fix?

Exhibit

Refer to the exhibit: CloudWatch Logs excerpt from a SageMaker real-time endpoint with instance type ml.m5.large. Average inference time is 500ms but increases over time. Logs show repeated 'MemoryError: cannot allocate memory' after several hours of operation.

A data scientist is building a binary classification model for fraud detection. The dataset is highly imbalanced (99% legitimate, 1% fraud). Which metric is most appropriate to evaluate model performance?

A data scientist is using Amazon SageMaker to train a model. The training job is taking longer than expected. Which change would most likely reduce training time?

Which TWO of the following are best practices for data preprocessing in machine learning? (Select TWO.)

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Fundamentals of AI and ML sessions

Start a Fundamentals of AI and ML only practice session

Every question in these sessions is drawn from the Fundamentals of AI and ML domain — nothing else.

Related practice questions

Related AIF-C01 topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the AIF-C01 exam test about Fundamentals of AI and ML?
You must map a described business problem to the correct AWS AI/ML service and the correct data or feature technique. The single most important thing: read the modality and data shape first, because they eliminate most wrong answers immediately.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Fundamentals of AI and ML questions in a focused session?
Yes — the session launcher on this page draws every question from the Fundamentals of AI and ML domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other AIF-C01 topics?
Use the topic links above to move to related areas, or go back to the AIF-C01 question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the AIF-C01 exam covers. They are not copied from any real exam or dump site.