Courseiva

CCNA Fundamentals of AI and ML Questions

15 of 90 questions · Page 2/2 · Fundamentals of AI and ML · Answers revealed

76
MCQmedium

A financial services firm must build a model that flags potentially fraudulent card transactions in under 200 milliseconds while keeping all data inside its own Amazon VPC. The fraud team has thousands of labeled historical transactions and the pattern changes slowly over months. Which approach best balances latency, data residency, and the need for periodic retraining?

A.Train a supervised classification model with Amazon SageMaker, deploy it to a real-time endpoint inside the VPC, and schedule periodic retraining jobs.
B.Deploy a pre-trained foundation model from Amazon Bedrock and prompt it to classify each transaction.
C.Build an unsupervised anomaly detection model on unlabeled data and run batch transform once per day.
D.Use Amazon Fraud Detector in evaluation mode and export predictions to an S3 bucket outside the VPC for scoring.
AnswerA

Fraud flagging with thousands of labeled transactions is a supervised classification problem, and a SageMaker real-time endpoint keeps inference within the VPC at low millisecond latency. Scheduled retraining jobs let the model adapt as fraud patterns drift over months. This combination satisfies the latency, residency, and retraining requirements directly without introducing external data movement.

Why this answer

A supervised classification model trained on the labeled history and deployed to a SageMaker real-time endpoint inside the VPC meets the latency and residency constraints, while scheduled retraining handles gradual fraud pattern drift. Purpose-built supervised learning uses the available labels, and in-VPC endpoints keep transaction data within the controlled network. Batch, unsupervised, or general foundation-model approaches each miss at least one hard requirement.

Exam trap

The trap here is treating a managed fraud service or foundation model as automatically better than a purpose-built supervised model that meets the stated latency and residency constraints.

77
MCQeasy

A startup needs to predict customer churn based on historical data containing labels (churned or not). Which type of machine learning should they use?

A.Reinforcement learning
B.Unsupervised learning
C.Supervised learning
D.Semi-supervised learning
AnswerC

Historical data already contains the churn outcome as a label, so the model learns a mapping from input features to that known target. Supervised learning is defined by training on labelled examples, satisfying the labelled churned/not-churned constraint.

Why this answer

The startup has labeled historical data (churned or not), which is the defining characteristic of supervised learning. The goal is to learn a mapping from input features to the known output labels to predict churn for new customers. This is a classic classification problem, making supervised learning the correct choice.

Exam trap

The AIF-C01 exam often tests the distinction between supervised and unsupervised learning by presenting a scenario with labeled data, where candidates might mistakenly choose unsupervised learning if they overlook the presence of labels.

How to eliminate wrong answers

Option A is wrong because reinforcement learning involves an agent learning through trial-and-error interactions with an environment to maximize cumulative reward, not from labeled historical data. Option B is wrong because unsupervised learning finds hidden patterns or structures in unlabeled data, but here the labels (churned/not) are explicitly provided. Option D is wrong because semi-supervised learning uses a small amount of labeled data with a large amount of unlabeled data, but the problem states the historical data contains labels, implying fully labeled data is available.

78
Multi-Selectmedium

A logistics company wants to use machine learning to predict delivery times. A data scientist is preparing the project and must identify which characteristics describe supervised learning rather than unsupervised learning. (Choose two.)

Select 2 answers
A.The training process requires no historical outcomes and relies only on feature similarity.
B.The goal is to learn a mapping from input features to a known output so the model can predict on new data.
C.The algorithm discovers hidden structure in data without any predefined output values.
D.The training dataset includes a target label for each example, such as actual delivery duration.
E.The model is evaluated primarily by how well it separates data into groups with no ground truth.
AnswersB, D

Supervised learning explicitly learns a function that maps input features to a known output, enabling predictions on unseen examples. For delivery times, the model learns from features such as distance, traffic, and time of day to predict duration for new orders, which is the core objective of supervised learning.

Why this answer

Supervised learning is defined by training on labeled examples and learning a mapping from inputs to a known output so the model can predict on new data. In the delivery-time scenario, historical trips with actual durations provide those labels. The remaining characteristics describe unsupervised learning, which discovers structure or groups without predefined outputs or ground truth.

Exam trap

The trap here is confusing clustering-style structure discovery with supervised prediction because both can operate on the same features.

79
MCQeasy

A social media company needs to automatically detect and flag toxic comments in multiple languages. They have a large stream of user comments and require real-time moderation. Which AWS service is best suited for this task?

A.Amazon Lex
B.Amazon Comprehend
C.Amazon Rekognition
D.Amazon Translate
AnswerB

Amazon Comprehend provides pre-trained natural language processing, including toxicity detection and multi-language support, and processes streaming text via real-time endpoints. This satisfies the stem's need for automatic, low-latency flagging of toxic comments across multiple languages.

Why this answer

Amazon Comprehend is the correct choice because it is a natural language processing (NLP) service that can perform real-time toxicity detection across multiple languages using its built-in content moderation and custom classification capabilities. It analyzes text streams to identify toxic comments (e.g., hate speech, threats) and integrates with AWS streaming services like Amazon Kinesis for real-time processing.

Exam trap

The trap here is that candidates may confuse Amazon Comprehend's NLP capabilities with Amazon Lex's conversational AI or Amazon Translate's language translation, assuming any language-related service can detect toxicity, but only Comprehend provides the specific text analysis APIs for content moderation.

How to eliminate wrong answers

Option A is wrong because Amazon Lex is a service for building conversational interfaces (chatbots) using automatic speech recognition (ASR) and natural language understanding (NLU), not for analyzing text for toxicity. Option C is wrong because Amazon Rekognition is designed for image and video analysis (e.g., object detection, facial recognition), not for processing text comments. Option D is wrong because Amazon Translate is a machine translation service that converts text between languages but does not perform toxicity detection or content moderation.

80
MCQhard

A company is deploying a machine learning model for real-time fraud detection. The model must have latency under 100ms. Which infrastructure choice is most appropriate?

A.Amazon SageMaker real-time endpoints
B.Amazon EC2 with Deep Learning AMI
C.Amazon SageMaker batch transform
D.Amazon SageMaker notebook instance
AnswerA

SageMaker real-time endpoints keep the model loaded on persistent instances and return synchronous predictions with consistently low latency, satisfying the sub-100ms requirement. Batch transform and asynchronous inference introduce queuing or storage overheads unsuitable for immediate fraud decisions.

Why this answer

Amazon SageMaker real-time endpoints are designed for low-latency inference, typically in the tens of milliseconds, making them suitable for real-time fraud detection where latency must be under 100ms. They deploy a model behind a persistent HTTPS endpoint that auto-scales to handle incoming requests with minimal delay.

Exam trap

The trap here is that candidates often confuse batch transform with real-time inference, assuming that any SageMaker inference capability can serve low-latency requests, but batch transform is explicitly asynchronous and designed for high-throughput, not low-latency.

How to eliminate wrong answers

Option B is wrong because Amazon EC2 with Deep Learning AMI requires manual setup of the inference server, scaling, and load balancing, which introduces operational overhead and cannot guarantee sub-100ms latency without significant custom engineering. Option C is wrong because Amazon SageMaker batch transform is designed for asynchronous, offline inference on large datasets, not for real-time, low-latency predictions. Option D is wrong because Amazon SageMaker notebook instance is an interactive development environment for building and testing models, not a production inference endpoint.

81
MCQmedium

A hospital wants an AI system that reviews chest X-ray images and flags those likely to contain pneumonia so radiologists can prioritize their queue. The hospital has 40,000 historical X-rays, each already labeled by a radiologist as pneumonia present or absent. Which learning approach best fits this scenario?

A.Supervised learning using a classification model trained on the labeled images.
B.Reinforcement learning where the model is rewarded for each correct flag.
C.Generative AI using a foundation model to synthesize new X-ray images for training.
D.Unsupervised learning using clustering to group similar X-rays together.
AnswerA

Each X-ray is paired with a radiologist-provided label indicating pneumonia present or absent, which is exactly the labeled data supervised classification needs. The model learns to map image features to one of two discrete classes, matching the flagging task. This approach directly uses the existing annotations and produces the binary output the radiologists require for triage.

Why this answer

The scenario provides thousands of images already annotated by radiologists with a binary outcome. Supervised classification consumes those labels to learn a decision boundary between pneumonia present and absent, producing the flagging behavior the hospital wants. Unsupervised, reinforcement, and generative approaches either ignore the labels or address a different problem than binary image triage.

Exam trap

The trap here is treating medical imaging as inherently unsupervised because images look complex, even though expert labels are already available.

82
MCQhard

A financial institution is deploying a fraud detection model using Amazon SageMaker. The model must be able to handle sudden spikes in inference requests during promotional events while keeping costs low. The team wants to use a serverless architecture to avoid provisioning idle capacity and to scale automatically from zero. However, the inference latency requirement is under 5 seconds for each request. Which SageMaker inference option should they choose?

A.Use Amazon SageMaker Serverless Inference
B.Use Amazon SageMaker Multi-Model Endpoints
C.Use Amazon SageMaker real-time endpoints with auto-scaling
D.Use Amazon SageMaker Asynchronous Inference
AnswerA

Serverless Inference provisions compute on demand, scales automatically from zero and charges only per request, eliminating idle capacity costs. It meets the under-five-second latency requirement, since cold-start and invocation latency stay within that threshold for typical payloads.

Why this answer

Amazon SageMaker Serverless Inference is the correct choice because it automatically scales from zero to handle sudden spikes in inference requests, aligning with the requirement to avoid provisioning idle capacity. It also meets the sub-5-second latency requirement for fraud detection, as it is designed for low-latency, on-demand inference without managing underlying infrastructure.

Exam trap

AWS often tests the misconception that serverless inference cannot meet low-latency requirements, but SageMaker Serverless Inference is specifically designed for sub-second to few-second latency, making it suitable for real-time fraud detection scenarios.

How to eliminate wrong answers

Option B is wrong because Multi-Model Endpoints require provisioned instances and do not scale from zero; they are designed to host multiple models on a single endpoint but still incur costs for idle capacity. Option C is wrong because real-time endpoints with auto-scaling still require a baseline of provisioned instances, which can lead to idle capacity costs during low-traffic periods, and they do not scale from zero. Option D is wrong because Asynchronous Inference is intended for large payloads and longer processing times (typically minutes), not for sub-5-second latency requirements, and it queues requests rather than providing real-time responses.

83
MCQeasy

A team is evaluating a classification model. The confusion matrix shows: TP=80, FN=20, FP=10, TN=90. What is the precision?

A.0.89
B.0.75
C.0.80
D.0.90
AnswerA

Precision measures the proportion of positive predictions that are truly positive, calculated as TP/(TP+FP). With TP=80 and FP=10, precision is 80/90 = 0.889, rounding to 0.89. This satisfies the stem's requirement to derive precision from the given confusion matrix values rather than recall or accuracy.

Why this answer

Precision is calculated as TP / (TP + FP). Here, TP=80 and FP=10, so precision = 80 / (80 + 10) = 80 / 90 = 0.888..., which rounds to 0.89. This metric measures the proportion of positive identifications that were actually correct.

Exam trap

The AIF-C01 exam often tests the distinction between precision and recall by providing confusion matrix values that make one metric easy to miscalculate if you confuse the denominator (TP+FP vs TP+FN).

How to eliminate wrong answers

Option B (0.75) is wrong because it incorrectly uses FN in the denominator, likely confusing precision with recall (TP / (TP + FN)). Option C (0.80) is wrong because it uses only TP divided by the total number of actual positives (TP + FN), which is recall, not precision. Option D (0.90) is wrong because it uses TN in the denominator or calculates accuracy (TP + TN) / total, which is not precision.

84
Multi-Selectmedium

A startup is preparing a dataset to train a supervised machine learning model and wants to follow sound data preparation practices. Which TWO activities are appropriate parts of preparing training data? (Choose two.)

Select 2 answers
A.Splitting the dataset into training, validation, and test subsets before tuning the model.
B.Adding synthetic copies of the same training examples until the dataset reaches a target size.
C.Removing duplicate records and correcting inconsistent label values across the dataset.
D.Tuning the model's hyperparameters repeatedly against the test set until the score improves.
E.Deleting all records that contain any missing values to guarantee a clean dataset.
AnswersA, C

Holding out validation and test subsets lets the team tune hyperparameters on validation data and estimate final performance on untouched test data. Without this separation, performance estimates become optimistically biased because the model has effectively seen the evaluation data during development, so this is a core preparation practice.

Why this answer

Sound preparation centers on protecting evaluation integrity and improving data quality: separating train, validation, and test subsets keeps performance estimates honest, while removing duplicates and fixing inconsistent labels prevents leakage and contradictory signals. Tuning on the test set, deleting all incomplete rows, and duplicating examples either leak information, discard useful data, or add no new information.

Exam trap

The trap here is treating 'cleaner and bigger' as always better, so deleting every incomplete row or duplicating records looks reasonable even though both harm data quality or evaluation validity.

85
MCQeasy

A startup with limited ML expertise wants to quickly prototype a binary classification model using a small customer dataset. They need a managed environment to run Jupyter notebooks and access pre-built algorithms. Which AWS service should they choose?

A.AWS Lambda
B.Amazon SageMaker
C.Amazon EMR
D.AWS Glue
AnswerB

Amazon SageMaker provides managed Jupyter notebook instances and built-in algorithms, letting users with limited ML expertise train and deploy a binary classifier quickly. It removes infrastructure management, satisfying the need for a managed environment with pre-built algorithms for prototyping.

Why this answer

Amazon SageMaker is the correct choice because it provides a fully managed environment for Jupyter notebooks and includes built-in, pre-built algorithms for binary classification. This allows the startup to quickly prototype without deep ML expertise, as SageMaker handles infrastructure, scaling, and model training.

Exam trap

The AIF-C01 exam often tests the distinction between managed ML platforms (SageMaker) and general-purpose compute or data processing services (Lambda, EMR, Glue), leading candidates to pick a service that can run code but lacks the specific notebook and pre-built algorithm capabilities required.

How to eliminate wrong answers

Option A is wrong because AWS Lambda is a serverless compute service for running code in response to events, not a managed environment for Jupyter notebooks or pre-built ML algorithms. Option C is wrong because Amazon EMR is a big data processing service using frameworks like Apache Spark and Hadoop, not designed for interactive Jupyter notebook-based ML prototyping with pre-built algorithms. Option D is wrong because AWS Glue is a serverless data integration and ETL service, not a platform for running Jupyter notebooks or accessing pre-built ML models.

86
MCQmedium

A hospital wants to detect pneumonia from chest X-ray images. Radiologists have already labeled thousands of past X-rays as either 'pneumonia' or 'no pneumonia'. The hospital wants a model that generalizes to new X-rays. Which type of machine learning task is this?

A.Reinforcement learning
B.Supervised image classification
C.Unsupervised anomaly detection
D.Clustering
AnswerB

The labeled X-rays provide input images paired with correct diagnoses, which is exactly the training data supervised classification needs. A convolutional neural network can learn visual patterns distinguishing pneumonia from normal lungs and then predict on unseen X-rays, matching the hospital's goal of generalization.

Why this answer

Labeled images with known diagnoses are the defining input for supervised classification. The model learns from pneumonia and no-pneumonia examples and predicts the class for new X-rays, which is precisely the generalization the hospital wants. Unsupervised, clustering, and reinforcement approaches either ignore the labels or lack the required structure.

Exam trap

The trap here is assuming medical imaging must use unsupervised anomaly detection, when the presence of radiologist-assigned labels makes it supervised classification.

87
MCQhard

A media company stores thousands of hours of unlabeled video footage and wants to build a searchable index that lets editors retrieve clips by describing their content in natural language. The team has no annotated dataset and no budget to label one. Which machine learning approach is the MOST appropriate starting point?

A.Use a pretrained multimodal embedding model to encode video frames and text into a shared vector space, then retrieve clips by similarity.
B.Train a supervised video classifier from scratch using the footage, treating each video file as its own class.
C.Apply k-means clustering to raw pixel values of every frame to produce searchable cluster identifiers.
D.Build a reinforcement learning agent that learns to select the clip most likely to satisfy an editor through trial and error.
AnswerA

A pretrained multimodal model maps both video content and text descriptions into a common embedding space, so a natural-language query can be compared directly against indexed clips without any custom labels. This leverages existing pretrained knowledge, requires no annotation budget, and supports flexible search. It is the standard foundation for semantic video retrieval systems.

Why this answer

With no labels and a need to search footage using free-form language, a pretrained multimodal embedding model is the most practical foundation. It places video and text in a shared vector space so that descriptive queries match relevant clips by similarity, enabling semantic retrieval without any custom annotation effort.

Exam trap

The trap here is assuming that a large unlabeled video collection must be turned into a supervised classification dataset, when pretrained multimodal embeddings can support retrieval directly.

88
Multi-Selecthard

Which TWO of the following are best practices for preparing training data for a machine learning model?

Select 2 answers
A.Handle missing values by imputing or removing them.
B.Split the data into training, validation, and test sets.
C.Remove all outliers to improve model robustness.
D.Use the entire dataset for training to maximize data usage.
E.Avoid shuffling the data to preserve original order.
AnswersA, B

Missing values break many algorithms and bias estimates, so imputing sensible substitutes or removing affected rows prevents errors and skewed learning. This is a standard preparation step ensuring the training set is complete and representative before model fitting.

Why this answer

Option A is correct because missing values can bias or break many ML algorithms, so best practice is to address them explicitly—either by imputation (e.g., mean/median/mode or model-based) or by removing affected rows/columns when appropriate. Option B is correct because splitting data into training, validation, and test sets enables unbiased model fitting, hyperparameter tuning, and final performance estimation, preventing data leakage and overfitting. Option C is not a best practice because outliers may be legitimate signal; blindly removing them can distort the distribution and reduce robustness, so they should be investigated and handled contextually.

Option D is not a best practice because training on the entire dataset leaves no held-out data for validation or testing, making it impossible to reliably estimate generalization. Option E is not a best practice because shuffling is typically needed to remove ordering bias and ensure representative mini-batches, especially when data is sorted by class or time.

Exam trap

The AIF-C01 exam often tests the misconception that removing all outliers is always beneficial, when in fact domain knowledge is required to distinguish between noise and legitimate extreme values that may be critical for model accuracy.

89
Multi-Selecteasy

A company wants to use AWS services to process natural language text. Which TWO AWS services provide natural language processing (NLP) capabilities? (Select TWO.)

Select 2 answers
A.Amazon Translate
B.Amazon Rekognition
C.Amazon Comprehend
D.Amazon Polly
E.Amazon Lex
AnswersC, E

Amazon Comprehend is a fully managed NLP service providing entity recognition, sentiment analysis, key phrase extraction and language detection on text. It directly satisfies the stem's requirement for an AWS service delivering natural language processing capabilities.

Why this answer

Amazon Comprehend (C) is correct because it is AWS's fully managed NLP service that uses machine learning to extract insights from text, including sentiment analysis, entity recognition, key phrase extraction, language detection, and topic modeling. Amazon Lex (E) is correct because it provides NLP capabilities for building conversational interfaces (chatbots and voice bots) using automatic speech recognition and natural language understanding to interpret user intent and slot values. Amazon Translate (A) is not marked correct because it performs language translation rather than general NLP analysis, even though it is a text-processing service.

Amazon Rekognition (B) is not marked correct because it is a computer vision service for image and video analysis, not natural language text. Amazon Polly (D) is not marked correct because it is a text-to-speech service that converts written text into lifelike speech, which is speech synthesis rather than NLP understanding.

Exam trap

The trap here is that candidates often confuse text-to-speech (Polly) or translation (Translate) with NLP, but these services do not perform language understanding or analysis—they only convert or generate speech/translation without extracting meaning.

90
Multi-Selectmedium

Which THREE statements about Amazon SageMaker Ground Truth are correct? (Choose three.)

Select 3 answers
A.It can only be used for text data.
B.It provides built-in workflows for image classification and object detection.
C.It supports automated data labeling using active learning.
D.It integrates with Amazon SageMaker to use the labeled data for training.
E.It can only use a public workforce from Amazon Mechanical Turk.
AnswersB, C, D

Ground Truth ships managed labelling interfaces and task templates for image classification and object detection, so teams avoid building custom annotation tooling. This built-in workflow support satisfies the stem's requirement that the statement describe a genuine Ground Truth capability.

Why this answer

Option B is correct because Amazon SageMaker Ground Truth ships with built-in labeling task templates and workflows for common job types, including image classification and object detection (as well as text, semantic segmentation, and bounding boxes). Option C is correct because Ground Truth supports automated data labeling, which uses active learning to have a model label high-confidence data while sending only low-confidence samples to human labelers, reducing cost and effort. Option D is correct because Ground Truth integrates directly with Amazon SageMaker, writing labeled datasets to Amazon S3 in augmented manifest format that SageMaker training jobs can consume.

Option A is wrong because Ground Truth handles images, video, text, and 3D point clouds, not just text. Option E is wrong because Ground Truth supports multiple workforces, including private workforces, vendor-managed workforces, and Amazon Mechanical Turk, so it is not limited to a public Mechanical Turk workforce.

Exam trap

AWS often tests the misconception that Ground Truth is limited to text data or only supports public workforces, while in reality it handles multiple data modalities and offers flexible workforce options including private and vendor-managed.

← PreviousPage 2 of 2 · 90 questions total

Ready to test yourself?

Try a timed practice session using only Fundamentals of AI and ML questions.