Courseiva

AWS Certified AI Practitioner AIF-C01 (AIF-C01) — Questions 601–675

862 questions total · 12pages · All types, answers revealed

Page 8

Page 9 of 12

Page 10
601
MCQmedium

A team is building a classifier to detect fraudulent transactions. The dataset has 99.9% legitimate transactions and 0.1% fraudulent. Which evaluation metric is most appropriate?

A.Mean Absolute Error (MAE)
B.Root Mean Squared Error (RMSE)
C.AUC-ROC
D.Accuracy
AnswerC

With only 0.1% fraud, accuracy is misleading because predicting all legitimate scores 99.9%. AUC-ROC measures ranking performance across all thresholds, remaining informative under such extreme class imbalance, so it distinguishes fraudulent from legitimate transactions regardless of the skewed prior.

Why this answer

AUC-ROC is the most appropriate metric for this highly imbalanced classification problem because it evaluates the model's ability to distinguish between fraudulent (positive) and legitimate (negative) transactions across all classification thresholds, without being biased by the overwhelming majority of legitimate transactions. Unlike accuracy, AUC-ROC focuses on the trade-off between true positive rate and false positive rate, making it robust to class imbalance where 99.9% of data is negative.

Exam trap

The AWS AI Practitioner exam often tests the misconception that accuracy is always a good metric, but the trap here is that accuracy is dangerously misleading for imbalanced datasets, and candidates must recognize that AUC-ROC or precision-recall metrics are required for such scenarios.

How to eliminate wrong answers

Option A is wrong because Mean Absolute Error (MAE) is a regression metric that measures average absolute differences between predicted and actual continuous values, not suitable for binary classification tasks like fraud detection. Option B is wrong because Root Mean Squared Error (RMSE) is also a regression metric that penalizes larger errors more heavily, and it cannot evaluate classification performance on imbalanced datasets. Option D is wrong because accuracy would be 99.9% even if the model predicts all transactions as legitimate, completely failing to detect any fraud; it is misleading for imbalanced datasets where the minority class is the one of interest.

602
MCQmedium

Refer to the exhibit. A data scientist created this endpoint config for a foundation model in Amazon SageMaker. However, the endpoint fails to scale under load. What is the most likely reason?

A.Missing AutoScaling configuration
B.Variant weight is 1.0
C.Instance type is too small
D.InitialInstanceCount is 1
AnswerA

SageMaker endpoints do not scale automatically unless an autoscaling policy is attached to the endpoint variant. Without Application Auto Scaling configured against a target metric such as InvocationsPerInstance, the endpoint keeps fixed instance count and cannot absorb rising load.

Why this answer

The endpoint fails to scale under load because the endpoint configuration shown lacks an AutoScaling policy. Without AutoScaling, SageMaker will not automatically adjust the number of instances based on traffic, so even if the initial instance count is 1, the endpoint cannot add more instances to handle increased load. AutoScaling must be explicitly configured via Application Auto Scaling to define scaling policies and target tracking metrics.

Exam trap

AWS often tests the misconception that setting a higher InitialInstanceCount or choosing a larger instance type alone enables scaling, when in fact AutoScaling must be explicitly configured as a separate step.

How to eliminate wrong answers

Option B is wrong because a variant weight of 1.0 is the default and does not prevent scaling; it simply means all traffic is routed to that variant. Option C is wrong because the instance type being 'too small' would cause performance issues or throttling, but it does not prevent the endpoint from scaling out; scaling is controlled by AutoScaling, not instance size. Option D is wrong because an InitialInstanceCount of 1 is a valid starting point; the endpoint can still scale out if AutoScaling is configured, so a single initial instance does not inherently block scaling.

603
MCQmedium

A company is using Amazon Bedrock to summarize long documents. They notice that the summary sometimes omits key details. What is the most likely cause?

A.The model is overfitted
B.The prompt lacks examples
C.The model's context window is too small
D.The temperature parameter is too high
AnswerC

Summarisation requires the whole document plus the prompt to fit within the model's context window. When the input exceeds that token limit, content is truncated before inference, so later sections never reach the model and their details are omitted from the summary.

Why this answer

When summarizing long documents with Amazon Bedrock, the model's context window determines the maximum amount of text it can process at once. If the document exceeds this limit, the model truncates or ignores portions, leading to omitted key details. This is the most likely cause because summarization requires the model to attend to the entire input, and a small context window directly prevents full coverage.

Exam trap

AWS often tests the distinction between model capacity limits (context window) and output quality parameters (temperature, prompt engineering), leading candidates to incorrectly attribute omission errors to randomness or lack of examples rather than the fundamental constraint of input size.

How to eliminate wrong answers

Option A is wrong because overfitting refers to a model memorizing training data and failing to generalize, which is unrelated to truncation or omission of details in a summarization task. Option B is wrong because while prompt examples (few-shot learning) can improve output quality, their absence does not cause the model to physically ignore parts of the input; the core issue is capacity, not prompting technique. Option D is wrong because a high temperature parameter increases randomness and creativity in generation, which might cause irrelevant or verbose output, but it does not cause the model to skip or omit key details from the input.

604
MCQmedium

A financial services company is building an application on Amazon Bedrock that generates personalized investment summaries. The compliance team requires that every generated summary includes an exact, verifiable citation from the company's approved regulatory documents, and that the model must not fabricate any citation. The company has a large corpus of approved PDF documents stored in Amazon S3. Which approach should the company use to meet these requirements?

A.Use Amazon Bedrock Knowledge Bases to ingest the S3 documents and configure the model to generate responses grounded in retrieved passages, then verify the returned source citations.
B.Enable model invocation logging in Amazon Bedrock and audit the logs to confirm that any citation the model produces exists in the S3 corpus.
C.Increase the model's temperature setting so the model explores more of its training data and is more likely to recall the exact regulatory text.
D.Fine-tune the foundation model on the approved regulatory documents using Amazon Bedrock custom models, then rely on the fine-tuned model to reproduce exact citations from memory.
AnswerA

Amazon Bedrock Knowledge Bases performs managed retrieval-augmented generation by ingesting documents from Amazon S3 into a vector store, retrieving relevant passages at query time, and returning responses with source attribution. This directly satisfies the requirement for grounded, verifiable citations while reducing fabrication, because the model conditions its output on retrieved approved content rather than only on its pretrained parameters.

Why this answer

Grounding the model in an approved document corpus is the reliable way to produce verifiable citations. Amazon Bedrock Knowledge Bases ingests the S3 documents, retrieves relevant passages per query, and returns source attributions, so generated summaries reference real approved text instead of relying on parametric memory. Sampling settings, fine-tuning, and logging do not provide retrieval-backed citation integrity.

Exam trap

The trap here is assuming that fine-tuning or a higher temperature makes a model recall exact source text, when only retrieval-based grounding reliably ties output to specific approved documents.

605
MCQeasy

Which AWS service can be used to convert a recorded speech file into text for further analysis?

A.Amazon Polly
B.Amazon Translate
C.Amazon Rekognition
D.Amazon Transcribe
AnswerD

Amazon Transcribe applies automatic speech recognition to audio files, returning accurate transcripts for downstream analysis. It directly satisfies the stem's requirement to convert recorded speech into text, unlike Amazon Polly, which performs the reverse text-to-speech conversion, or Amazon Translate, which handles language translation rather than speech transcription.

Why this answer

Amazon Transcribe is an automatic speech recognition (ASR) service that converts audio speech into text. It is specifically designed for this task, supporting both real-time streaming and batch processing of recorded speech files, making it the correct choice for converting a recorded speech file into text for further analysis.

Exam trap

The trap here is confusing Amazon Transcribe (speech-to-text) with Amazon Polly (text-to-speech), as both deal with speech but in opposite directions, leading candidates to select Polly when the question asks for converting speech into text.

How to eliminate wrong answers

Option A is wrong because Amazon Polly is a text-to-speech (TTS) service that converts text into lifelike speech, not speech into text. Option B is wrong because Amazon Translate is a neural machine translation service that translates text between languages, it does not process audio or perform speech recognition. Option C is wrong because Amazon Rekognition is a computer vision service for image and video analysis, such as object detection, facial recognition, and content moderation, and it does not handle audio transcription.

606
MCQmedium

A support organization notices its generative AI assistant sometimes invents policy details that do not exist. Leadership wants a measurable way to track whether this behavior improves over time. Which action should the team take?

A.Switch to a model with a larger parameter count and assume the fabrication problem disappears.
B.Build a labeled evaluation set of questions with verified answers and score responses for factual consistency on a recurring basis.
C.Increase the model's maximum output token limit so it has more room to explain itself.
D.Rely on customer complaints as the primary signal for whether fabricated policy details are decreasing.
AnswerB

A curated evaluation set with known-correct answers turns an anecdotal complaint into a repeatable metric, allowing the team to compare prompts, retrieval strategies, or models over time. This directly addresses the request for a measurable way to track improvement in factual accuracy and supports regression testing as the system evolves.

Why this answer

Turning an observed quality problem into a tracked metric requires a fixed evaluation set with verified answers and a scoring procedure that can be rerun after each change. This makes comparisons across prompts, retrieval configurations, and models meaningful, and it establishes a regression baseline so the organization can prove whether fabrication is actually declining.

Exam trap

The trap here is assuming that newer, larger models or more output space automatically fix hallucination, when without a measurement framework no one can tell whether the problem improved.

607
MCQmedium

A financial services company wants to deploy a generative AI assistant that answers questions using its internal policy documents. The team is concerned that the model may invent policies that do not exist. Which approach best reduces this risk while keeping the model's answers tied to the company's actual documents?

A.Switch to a larger foundation model with more parameters to improve factual recall.
B.Use retrieval augmented generation to fetch relevant document passages and include them in the prompt.
C.Lower the maximum token limit so the model produces shorter answers with fewer chances to err.
D.Increase the model's temperature setting so it explores more diverse phrasings.
AnswerB

Retrieval augmented generation retrieves relevant passages from the company's own document store and supplies them as context, so the model generates answers conditioned on real policy text. This grounds responses in verifiable source material and sharply reduces fabricated content, directly addressing the concern about invented policies while keeping answers tied to actual documents.

Why this answer

Retrieval augmented generation grounds the model by supplying relevant passages from the company's own document repository at inference time. Because the model conditions its answer on retrieved, verifiable content, it is far less likely to invent policies. This directly satisfies the requirement that answers stay tied to actual internal documents rather than relying on the model's parametric memory.

Exam trap

The trap here is assuming that a bigger model or shorter answers will fix fabrication, when the real fix is supplying authoritative source content at inference time.

608
MCQeasy

A company wants to extract text and data from scanned PDF invoices for automated processing. Which AWS service is MOST appropriate for this task?

A.Amazon Rekognition
B.Amazon Transcribe
C.Amazon Comprehend
D.Amazon Textract
AnswerD

Amazon Textract uses optical character recognition with layout and table understanding, extracting text and structured data such as line items and totals from scanned invoices. This directly satisfies the automated processing requirement, unlike services limited to plain OCR or general document storage.

Why this answer

Amazon Textract is specifically designed to extract text, handwriting, and data from scanned documents like PDF invoices. It uses machine learning to recognize and extract key-value pairs, tables, and form data, making it the most appropriate service for automated invoice processing.

Exam trap

The trap here is that candidates may confuse Amazon Textract with Amazon Rekognition, since both can process images, but Rekognition lacks the specialized document layout analysis and key-value extraction capabilities required for invoice processing.

How to eliminate wrong answers

Option A is wrong because Amazon Rekognition is optimized for image and video analysis (e.g., object detection, facial recognition), not for extracting structured text and data from scanned documents. Option B is wrong because Amazon Transcribe is a speech-to-text service for converting audio to text, not for processing scanned PDFs. Option C is wrong because Amazon Comprehend is a natural language processing (NLP) service for analyzing text sentiment, entities, and key phrases, but it cannot extract text from scanned images or PDFs—it requires text input.

609
MCQmedium

A retail company wants to forecast daily product demand for the next quarter. They have three years of historical sales data that includes seasonal spikes and promotional periods, and they want a fully managed AWS service that can automatically train and tune a forecasting model without writing deep learning code. Which AWS service best fits this requirement?

A.Amazon Comprehend
B.Amazon Polly
C.Amazon Rekognition
D.Amazon Forecast
AnswerD

Amazon Forecast is a fully managed time-series forecasting service that ingests historical demand data with related features such as promotions and holidays, then automatically selects and tunes algorithms like DeepAR+ and CNN-QR. It directly matches the requirement of no custom deep learning code and handles seasonality, making it the correct fit for quarterly demand forecasting.

Why this answer

The scenario calls for a managed service that learns from historical numeric sales data containing seasonality and promotions. Amazon Forecast is purpose-built for this, automatically training and tuning time-series models and exposing forecasts without custom deep learning code. The other services address computer vision, speech, and text analytics, none of which produce demand forecasts.

Exam trap

The trap here is assuming any fully managed AI service can handle forecasting, when forecasting requires a dedicated time-series service.

610
MCQhard

An enterprise is building a RAG solution with Amazon Bedrock and needs to ensure that the retrieved documents are from authorised sources only. They also must prevent the model from generating responses that contain personally identifiable information (PII). Which two Bedrock features combined address these requirements?

A.Amazon Bedrock Knowledge Bases and Bedrock Guardrails
B.Amazon Bedrock Model Evaluation and Bedrock Studio
C.Amazon Bedrock Agents and Bedrock Guardrails
D.Amazon Bedrock Playground and prompt engineering
AnswerA

Bedrock Knowledge Bases restrict retrieval to indexed, approved data sources, satisfying the authorised-sources constraint. Guardrails then apply PII filters and deny topics to both input prompts and generated output, blocking PII leakage. Together they cover retrieval provenance and response safety, the two distinct requirements in the stem.

Why this answer

Amazon Bedrock Knowledge Bases lets you ground responses in your own curated data sources (e.g., S3, OpenSearch), so retrieval is limited to the documents you configure and index. Amazon Bedrock Guardrails provides configurable safeguards, including sensitive information filters that can detect and block PII in model responses. Combining Knowledge Bases for controlled retrieval with Guardrails for PII filtering addresses both requirements.

Exam trap

AWS often tests the distinction between features that enable retrieval (Knowledge Bases) versus those that enforce safety (Guardrails), and candidates may mistakenly think Agents alone can handle both authorization and filtering, or that prompt engineering is sufficient for PII control.

How to eliminate wrong answers

Option B is wrong because Amazon Bedrock Model Evaluation is used for assessing model performance, not for controlling document source authorization or PII filtering, and Bedrock Studio is a collaborative development environment without these specific controls. Option C is wrong because Amazon Bedrock Agents are for orchestrating multi-step tasks and can use Knowledge Bases, but they do not inherently provide the PII filtering capability; Guardrails is needed for that, but the combination of Agents and Guardrails alone does not address the authorized source requirement as directly as Knowledge Bases does. Option D is wrong because Amazon Bedrock Playground is an interactive testing interface and prompt engineering is a technique, not a feature; neither provides the necessary access control or automated PII filtering required.

611
MCQmedium

A bank wants to use Amazon Augmented AI (A2I) to review high-value loan applications that require human judgment. Which workflow best implements human-in-the-loop review for these predictions?

A.Set up an A2I workflow with a confidence threshold, so low-confidence predictions are sent to human reviewers
B.Route all loan applications to human reviewers for approval
C.Use Amazon Mechanical Turk to review all predictions in real time
D.Configure a private workforce in Amazon SageMaker Ground Truth
AnswerA

Amazon A2I triggers human review when inference confidence falls below a configured threshold, routing only uncertain predictions to reviewers. This satisfies the bank's need for human judgment on high-value loan applications without reviewing every prediction, since A2I integrates this condition directly into the workflow.

Why this answer

Amazon A2I enables human review of low-confidence predictions or specific conditions. The best practice is to set a confidence threshold; predictions below that threshold are sent to human reviewers.

612
MCQmedium

A media company uses Amazon Bedrock to generate personalized news summaries. They notice that summaries sometimes include details not present in the source articles. They want to reduce these hallucinations without retraining the model. Which approach should they use?

A.Decrease the maximum token length in the InvokeModel request.
B.Switch to a larger foundation model with more parameters.
C.Use Retrieval Augmented Generation (RAG) by retrieving relevant passages from the source articles and including them in the prompt.
D.Increase the model's temperature setting to encourage more creative outputs.
AnswerC

RAG grounds the model's response by supplying relevant, authoritative content from the source articles directly in the prompt. This reduces hallucination because the model conditions its output on the retrieved text rather than relying solely on parametric knowledge. It requires no retraining and integrates with Amazon Bedrock Knowledge Bases or custom retrieval, making it suitable for dynamic news content.

Why this answer

Retrieval Augmented Generation supplies the model with relevant excerpts from the source articles at inference time, so the generated summary is conditioned on actual content rather than the model's internal knowledge. This reduces fabricated details without retraining. Other options either increase randomness, rely on model size, or truncate output, none of which ground the response in the provided material.

Exam trap

The trap here is assuming that a larger or more capable model will automatically stop hallucinating, when grounding the prompt with retrieved source content is what actually constrains the output.

613
Multi-Selectmedium

A machine learning engineer is implementing a RAG system using Amazon Bedrock and a vector database. They need to chunk a large set of PDF documents before embedding. Which THREE considerations are important for chunking strategy? (Select THREE.)

Select 3 answers
A.Overlap between consecutive chunks
B.The method used to split chunks (e.g., by sentence, paragraph, or fixed token count)
C.Chunk size (e.g., number of tokens)
D.The embedding model's maximum input length
E.The metadata associated with each document
AnswersA, B, C

Overlap between consecutive chunks preserves context across boundaries, so sentences split mid-thought still appear intact in at least one chunk. This improves retrieval quality in the RAG pipeline by preventing meaning loss at chunk edges.

Why this answer

Options A, B, and C are the three core design parameters of any chunking strategy. A is correct because adding overlap between consecutive chunks preserves context that would otherwise be lost at chunk boundaries, improving retrieval quality when a relevant passage spans two chunks. B is correct because the splitting method (sentence, paragraph, semantic, or fixed token count) determines how coherently meaning is preserved within each chunk and directly affects embedding quality.

C is correct because chunk size in tokens controls the trade-off between retrieval granularity and the amount of context each embedding captures; too small loses context, too large dilutes relevance. Option D is not a chunking-strategy consideration per se — the embedding model's maximum input length is a hard constraint that caps chunk size, but it is a model property rather than a strategic choice. Option E is about document metadata and filtering, which relates to retrieval and indexing, not to how documents are split into chunks.

Exam trap

AWS AI Practitioner exams often test the distinction between constraints (like model limits) and active strategy choices (like chunk size and overlap), so candidates mistakenly select the embedding model's maximum input length as a chunking strategy consideration instead of recognizing it as a boundary condition.

614
Multi-Selectmedium

A data scientist is preparing data for a classification task. Which TWO techniques are commonly used for handling missing values? (Choose two.)

Select 2 answers
A.Label encoding
B.Normalization
C.Imputing with mean
D.Dropping rows with any missing values
E.One-hot encoding
AnswersC, D

Imputing with mean replaces missing numerical entries with the column average, preserving dataset size and avoiding dropped rows. This directly satisfies the stem's requirement for a common missing-value technique, since mean imputation is a standard preprocessing method in classification workflows.

Why this answer

Option C (Imputing with mean) is correct because replacing missing numeric entries with the column mean is a standard, simple imputation technique that preserves the dataset size and avoids discarding information. Option D (Dropping rows with any missing values) is also correct because listwise deletion is a common, straightforward approach to handling missing data, especially when the missingness is minimal or random. The other options do not address missing values: A (Label encoding) converts categorical labels into integer codes, B (Normalization) rescales numeric feature ranges, and E (One-hot encoding) creates binary indicator columns for categories—all are preprocessing steps for existing values, not missing-data handling.

Exam trap

The AIF-C01 exam often tests the distinction between data preprocessing techniques (e.g., encoding, scaling) and missing value handling, so candidates mistakenly select label encoding or normalization because they are common preprocessing steps, even though they do not address missing data.

615
Multi-Selecteasy

Which TWO statements about the bias-variance tradeoff are correct? (Choose TWO.)

Select 2 answers
A.High bias typically leads to overfitting
B.High variance models are insensitive to changes in training data
C.High bias typically leads to underfitting
D.Increasing model complexity typically decreases variance
E.High variance typically leads to overfitting
AnswersC, E

High bias means the model's assumptions are too rigid for the underlying data, so it cannot capture the true relationship. Consequently it performs poorly on both training and test data — the defining symptom of underfitting.

Why this answer

Option C is correct because high bias means the model makes overly simplistic assumptions about the data, so it fails to capture the underlying pattern and systematically underfits both training and test data. Option E is correct because high variance means the model is overly sensitive to the particular training samples, fitting noise as well as signal, which produces excellent training performance but poor generalization—the hallmark of overfitting. Option A is wrong because high bias leads to underfitting, not overfitting; overfitting is associated with high variance.

Option B is wrong because high variance models are highly sensitive to changes in training data, not insensitive. Option D is wrong because increasing model complexity typically increases variance (and decreases bias), not decreases variance.

Exam trap

The trap is mixing up the symptoms: candidates often think high bias causes overfitting (it's the opposite) or that more complexity reduces variance (it increases it).

616
MCQhard

A legal firm uses Amazon Bedrock to generate contract summaries. They want to evaluate the quality of summaries against human-written reference summaries. The evaluation should capture both the overlap of n-grams and the semantic similarity. Which combination of automated metrics is MOST appropriate?

A.Exact match and F1 score
B.ROUGE and BLEU
C.ROUGE and BERTScore
D.BLEU and BERTScore
AnswerC

ROUGE measures n-gram overlap with the reference summaries, satisfying the lexical requirement, while BERTScore uses contextual embeddings to capture semantic similarity beyond exact wording. Together they cover both constraints in the stem, unlike metrics addressing only one dimension.

Why this answer

ROUGE measures n-gram overlap (recall-oriented) for summarization, while BERTScore uses contextual embeddings to capture semantic similarity. BLEU is for translation. Combining ROUGE and BERTScore gives both lexical and semantic evaluation.

617
MCQmedium

A financial services company uses Amazon Bedrock with the Anthropic Claude 3 Sonnet model to answer employee questions about internal policies. The policy documents are updated frequently, and the model occasionally provides outdated or incorrect policy details. The company wants the model to base its answers on the most current authoritative documents without retraining the model. Which approach should they use?

A.Switch to a larger foundation model with a higher parameter count.
B.Use Retrieval Augmented Generation (RAG) by storing policy documents in an Amazon Bedrock knowledge base and retrieving relevant passages at inference time.
C.Fine-tune the foundation model on the updated policy documents.
D.Increase the model's temperature setting to encourage more creative and up-to-date answers.
AnswerB

RAG with an Amazon Bedrock knowledge base indexes the policy documents and retrieves the most relevant passages for each query, then passes them to the model as context. This grounds responses in the latest authoritative content without retraining, and updating the knowledge base data source keeps answers current. It directly addresses outdated or incorrect policy details by supplying fresh source text at inference time.

Why this answer

Retrieval Augmented Generation with an Amazon Bedrock knowledge base retrieves relevant, current policy passages and supplies them to the model as context, so answers are grounded in authoritative documents. It avoids retraining costs and keeps pace with frequent updates, directly solving the outdated or incorrect policy detail problem.

Exam trap

The trap here is assuming that a larger or fine-tuned model automatically knows frequently updated internal documents, when grounding requires retrieval of current source content at inference time.

618
MCQmedium

A financial services firm needs to ensure that all calls to Amazon Bedrock APIs are logged for audit purposes. Which AWS service should they enable to capture API calls?

A.AWS CloudTrail
B.Amazon S3 server access logs
C.Amazon CloudWatch Logs
D.AWS Config
AnswerA

CloudTrail records API activity as management and data events, capturing every Bedrock API call with caller identity, timestamp and source IP. Enabling a trail delivers the audit logging the firm requires, satisfying the compliance constraint without altering application code.

Why this answer

AWS CloudTrail records API activity in AWS accounts, including Bedrock API calls, providing audit logs.

619
Multi-Selectmedium

A data science team uses Amazon SageMaker to train models. To comply with SOC 2, they must ensure that access to training data is logged, that the data is encrypted at rest, and that model training jobs are isolated from each other. Which THREE actions should they take? (Choose three.)

Select 3 answers
A.Enable Amazon Inspector to scan training instances for vulnerabilities.
B.Enable server-side encryption on the S3 bucket containing training data using SSE-KMS.
C.Use SageMaker Debugger to monitor training jobs.
D.Enable AWS CloudTrail to capture SageMaker API calls.
E.Use SageMaker VPC mode to launch training jobs in a private subnet.
AnswersB, D, E

SSE-KMS encrypts data at rest.

Why this answer

Enabling server-side encryption on the S3 bucket containing training data using SSE-KMS ensures data at rest is encrypted, which is a direct requirement for SOC 2 compliance. SSE-KMS provides envelope encryption with a customer-managed AWS KMS key, allowing fine-grained access control and audit trails for the encryption keys.

Exam trap

The trap here is that candidates may confuse Amazon Inspector with a logging or encryption service, or think SageMaker Debugger provides security logging, when in fact Inspector only scans for vulnerabilities and Debugger only monitors model training metrics.

620
MCQmedium

During model training, a data scientist notices that the model performs very well on the training data but poorly on the test data. The scientist suspects high variance. Which technique is MOST likely to reduce the variance and improve test performance?

A.Lower the learning rate
B.Increase the number of features
C.Apply regularization (e.g., L1 or L2)
D.Decrease the amount of training data
AnswerC

High variance means the model memorises training noise, so it generalises poorly. L1 or L2 regularization penalises large weights, constraining model complexity and reducing that variance, which directly addresses the gap between strong training and weak test performance described in the stem.

Why this answer

High variance indicates the model is overfitting to the training data, capturing noise rather than the underlying pattern. Regularization (L1/L2) adds a penalty to the loss function for large coefficients, effectively constraining the model complexity and reducing variance, which improves generalization to unseen test data.

Exam trap

AWS often tests the misconception that lowering the learning rate or reducing data fixes overfitting, but the correct approach is to reduce model complexity through regularization or feature selection.

How to eliminate wrong answers

Option A is wrong because lowering the learning rate controls the step size during gradient descent and affects convergence speed, not model complexity or variance. Option B is wrong because increasing the number of features typically adds more dimensions, which can exacerbate overfitting and increase variance, not reduce it. Option D is wrong because decreasing the amount of training data usually worsens overfitting and increases variance, as the model has fewer examples to learn the true underlying distribution.

621
MCQeasy

A data scientist wants to fine-tune a foundation model on a specific domain dataset using Amazon SageMaker. Which built-in SageMaker feature can simplify the training process?

A.SageMaker Neo
B.SageMaker Canvas
C.SageMaker JumpStart
D.SageMaker Ground Truth
AnswerC

SageMaker JumpStart provides pre-trained foundation models with ready-made notebooks and one-click deployment, removing much of the manual configuration involved in fine-tuning. This simplifies domain-specific training by supplying the model artefacts and training scaffolding out of the box.

Why this answer

SageMaker JumpStart provides pre-trained foundation models and built-in training scripts that simplify fine-tuning on custom datasets. It handles the underlying infrastructure, hyperparameter configurations, and model deployment, allowing the data scientist to focus on the domain-specific data rather than writing custom training loops or managing SageMaker training jobs manually.

Exam trap

The trap in this question is that candidates may confuse SageMaker services that are related to aspects of ML workflows but not specifically designed for fine-tuning foundation models. SageMaker Neo handles model optimization for deployment, SageMaker Canvas is a no-code tool for building ML models without writing code, and SageMaker Ground Truth is for data labeling. Only SageMaker JumpStart provides pre-trained models and built-in training scripts suitable for fine-tuning.

How to eliminate wrong answers

Option A is wrong because SageMaker Neo is a model optimization and compilation service for deploying models on edge devices, not a tool for fine-tuning foundation models. Option B is wrong because SageMaker Canvas is a no-code visual interface for building ML models using pre-built algorithms and does not support fine-tuning of foundation models. Option D is wrong because SageMaker Ground Truth is a data labeling service used to create high-quality training datasets, not a feature for training or fine-tuning models.

622
Multi-Selectmedium

A healthcare organization is deploying an AI system to assist in diagnosing diseases from medical images. They need to ensure the system is robust, safe, and subject to human oversight. Which TWO actions align with responsible AI guidelines? (Select TWO.)

Select 2 answers
A.Monitor model performance over time to detect data drift
B.Deploy the model directly to production without testing
C.Use Amazon Augmented AI (A2I) to set up human review for uncertain predictions
D.Permanently disable human review to speed up diagnoses
E.Use only one source of training data
AnswersA, C

Continuous performance monitoring detects data drift, where incoming medical images diverge from training distribution and degrade accuracy. This satisfies the robustness requirement by surfacing degradation before unsafe predictions reach clinicians, enabling retraining or intervention rather than silent failure.

Why this answer

Option A is correct because continuously monitoring model performance over time to detect data drift is essential for maintaining robustness and safety in a clinical AI system, since changes in patient demographics, imaging equipment, or disease prevalence can degrade accuracy and introduce risk. Option C is correct because Amazon Augmented AI (A2I) enables human review workflows for low-confidence or uncertain predictions, directly satisfying the requirement for meaningful human oversight in high-stakes diagnostic decisions. Options B, D, and E do not belong: deploying directly to production without testing violates safety and validation practices, permanently disabling human review removes the required oversight, and relying on only one source of training data undermines robustness and generalizability across diverse patient populations.

Exam trap

AIF-C01 often tests whether candidates recognize that human oversight and continuous monitoring are mandatory for high-risk AI — the trap is selecting options that prioritize speed or simplicity (disable review, skip testing, single data source) because they sound operationally efficient, when they actually violate responsible AI principles.

623
MCQeasy

A developer is using Amazon Bedrock to create a chatbot. They want to ensure the bot does not generate toxic or offensive content. Which feature should they enable?

A.Use careful prompt engineering to avoid toxic responses.
B.Fine-tune the model on a dataset of safe responses.
C.Enable content filtering on the Bedrock model.
D.Implement external response validation using a third-party API.
AnswerC

Content filtering applies configurable thresholds that block or mask harmful categories such as hate, violence and sexual content in both prompts and responses. Enabling it on the Bedrock model directly satisfies the requirement to prevent toxic output from the chatbot.

Why this answer

Amazon Bedrock provides built-in content filtering capabilities that can be enabled at the model invocation level to automatically detect and block toxic or offensive content in both input prompts and generated responses. This feature uses predefined safety filters (e.g., hate, insults, sexual content, violence) and is the most direct and managed way to prevent harmful outputs without requiring custom development.

Exam trap

A common misconception is that prompt engineering alone is sufficient for safety, when in fact Bedrock's content filtering is the explicit, managed feature designed to enforce content policies at runtime.

How to eliminate wrong answers

Option A is wrong because careful prompt engineering can reduce but not guarantee the elimination of toxic responses, as the underlying model may still generate harmful content due to its training data or adversarial inputs. Option B is wrong because fine-tuning the model on a dataset of safe responses requires significant data preparation, cost, and expertise, and it does not provide a runtime guard against all toxic outputs, especially for edge cases. Option D is wrong because implementing external response validation using a third-party API adds latency, complexity, and potential cost, and it is not a native Bedrock feature; Bedrock already offers content filtering as a first-class, integrated service.

624
MCQmedium

A company is developing an AI system that transcribes medical consultations. To ensure privacy and security, they need to implement controls that protect patient health information (PHI). Which AWS service can help anonymize data before it is used for model training?

A.Amazon SageMaker Ground Truth
B.AWS Lake Formation
C.Amazon Rekognition
D.Amazon Comprehend Medical
AnswerD

Amazon Comprehend Medical uses natural language processing to detect and redact protected health information such as names, dates and medical record numbers within clinical text, satisfying the requirement to anonymise PHI before it reaches model training.

Why this answer

Amazon Comprehend Medical is a HIPAA-eligible NLP service purpose-built for extracting and detecting PHI from unstructured medical text such as clinical notes and transcriptions. It includes a DetectPHI API that identifies 18+ PHI categories (names, dates, IDs, contact info) so they can be redacted or anonymized before the data is used for model training.

Exam trap

AIF-C01 often tests whether candidates confuse general-purpose AI services (Rekognition, Comprehend) with the medical-specialized variant (Comprehend Medical), so picking plain Comprehend or Rekognition misses the PHI-specific capability.

How to eliminate wrong answers

Option A is wrong because SageMaker Ground Truth is a data labeling service for building training datasets — it does not detect or anonymize PHI. Option B is wrong because Lake Formation is a data lake governance service for access control and fine-grained permissions, not content-level PHI detection. Option C is wrong because Amazon Rekognition performs image and video analysis (faces, objects, moderation), not medical text PHI anonymization.

625
MCQmedium

A company is using Amazon Bedrock to power a chatbot that provides customer support. The security team wants to ensure that the chatbot does not generate responses that include profanity, hate speech, or prompts that attempt to bypass safety filters (jailbreak attempts). They also want to log any blocked interactions for review. Which AWS service or feature should be used to meet these requirements?

A.Amazon Bedrock Guardrails with content filters and prompt attack detection, and enable model invocation logging to capture blocked interactions.
B.Amazon Comprehend for sentiment analysis and Amazon Macie for detecting sensitive data, with alerts sent to Amazon SNS.
C.Amazon SageMaker Clarify for bias detection and Amazon Augmented AI (A2I) for human review, with results logged to Amazon S3.
D.AWS WAF with managed rules for bot control and a custom Lambda function to inspect responses for profanity, logging to Amazon CloudWatch.
AnswerA

Amazon Bedrock Guardrails provides configurable content filters to block hate speech, profanity, and other harmful content. It also includes prompt attack detection to identify jailbreak attempts. When model invocation logging is enabled, Bedrock logs the full request and response, including blocked interactions, allowing for review. This native integration meets all requirements without custom code.

Why this answer

The requirements are to block harmful content and jailbreak attempts in a Bedrock-powered chatbot, and to log blocked interactions. Amazon Bedrock Guardrails offers content filters for profanity, hate speech, and more, plus prompt attack detection. Enabling model invocation logging captures all interactions, including those blocked by Guardrails, for later review.

This is a managed solution that directly addresses the needs without custom development.

Exam trap

The trap here is thinking that general-purpose security services like AWS WAF or content analysis services like Comprehend can moderate generative AI outputs, when they lack the specific filters for profanity and jailbreak detection.

626
MCQeasy

A company wants to track API calls made to Amazon SageMaker for audit purposes. Which AWS service should they enable?

A.AWS CloudTrail
B.Amazon Macie
C.AWS Config
D.Amazon CloudWatch Logs
AnswerA

AWS CloudTrail records API activity in the account, including SageMaker InvokeEndpoint calls, capturing identity, time and source. Enabling it provides the audit trail the company requires, unlike CloudWatch metrics, which show performance rather than who called what.

Why this answer

AWS CloudTrail is the correct service because it records API activity across AWS services, including Amazon SageMaker. By enabling CloudTrail, the company can capture all SageMaker API calls (e.g., CreateModel, InvokeEndpoint) for audit, compliance, and security analysis. CloudTrail logs provide details such as the identity of the caller, the time of the call, and the request parameters, which are essential for auditing.

Exam trap

The trap here is that candidates often confuse AWS CloudTrail (for API auditing) with Amazon CloudWatch Logs (for log monitoring), mistakenly thinking CloudWatch Logs is the primary service for tracking API calls, but CloudTrail is the dedicated service for recording API activity across AWS.

How to eliminate wrong answers

Option B (Amazon Macie) is wrong because Macie is a data security service that uses machine learning to discover, classify, and protect sensitive data in Amazon S3, not to track API calls. Option C (AWS Config) is wrong because Config evaluates and records resource configuration changes (e.g., SageMaker endpoint configuration), not API call activity. Option D (Amazon CloudWatch Logs) is wrong because CloudWatch Logs is for monitoring, storing, and accessing log files from applications and AWS services, but it does not natively capture API calls; it can ingest CloudTrail logs but is not the primary service for API auditing.

627
MCQmedium

A developer is trying to invoke the Claude v2 model in Amazon Bedrock from a Lambda function. The Lambda function's IAM role has the following policy attached: { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "bedrock:InvokeModel", "Resource": "*" } ] } When the Lambda function runs, it receives the error shown in the exhibit. Which additional step is most likely needed to resolve this issue?

A.Change the AWS region to one where Claude v2 is available.
B.Use a different model ID such as 'anthropic.claude-v1'.
C.Request access to the Anthropic Claude model through the Amazon Bedrock console.
D.Add a condition to the IAM policy to specify the model ARN.
AnswerC

Model access in Amazon Bedrock is granted separately from IAM permissions. Even with bedrock:InvokeModel allowed, the account must first request access to the Anthropic Claude model in the Bedrock console, otherwise invocation returns an access-denied error despite the permissive policy.

Why this answer

Amazon Bedrock requires explicit user access approval for third-party foundation models like Anthropic Claude before they can be invoked. Even with a valid IAM policy allowing bedrock:InvokeModel on all resources, the model itself must be granted access via the Bedrock console's 'Model access' section. Without this step, the API returns an access denied error regardless of IAM permissions.

Exam trap

The trap here is that candidates assume a wildcard IAM policy (Resource: '*') grants full access, but Bedrock requires an additional explicit model access approval that is independent of IAM, causing them to incorrectly focus on policy or region changes.

How to eliminate wrong answers

Option A is wrong because the error is not due to regional availability; Claude v2 is available in multiple regions (e.g., us-west-2, us-east-1) and the region choice does not bypass the model access requirement. Option B is wrong because using a different model ID like 'anthropic.claude-v1' would still require explicit model access for that specific model; the error is about access, not model version. Option D is wrong because adding a condition to specify the model ARN does not resolve the missing model access approval; the IAM policy already allows all resources, and the issue is at the service level, not the policy level.

628
MCQmedium

A healthcare startup is using Amazon SageMaker to train a model on patient data. They need to ensure that the training data does not contain any personally identifiable information (PII) before being used. Which AWS service can automatically detect and report PII in the data stored in S3?

A.Amazon Rekognition
B.AWS Glue DataBrew
C.Amazon Comprehend Medical
D.Amazon Macie
AnswerD

Amazon Macie uses machine learning and pattern matching to automatically discover, classify and report sensitive data such as PII in S3. It continuously evaluates buckets and raises findings, directly satisfying the requirement to detect and report PII before SageMaker training begins.

Why this answer

Amazon Macie is a fully managed data security and privacy service that uses machine learning to automatically discover, classify, and protect sensitive data stored in Amazon S3. It can detect personally identifiable information (PII) such as names, addresses, and credit card numbers, and provides detailed reports. This directly addresses the requirement to detect and report PII in S3 before using it for training.

Exam trap

AIF-C01 often tests the confusion between Macie and Comprehend Medical — candidates may think Comprehend Medical detects all PII, but Macie is the dedicated service for PII discovery in S3, while Comprehend Medical is for clinical text extraction.

How to eliminate wrong answers

Option A is wrong because Amazon Rekognition is a computer vision service for image and video analysis, not for detecting PII in text data stored in S3. Option B is wrong because AWS Glue DataBrew is a data preparation tool that can profile data but does not automatically detect PII with the same specialized focus as Macie. Option C is wrong because Amazon Comprehend Medical is designed for extracting medical information from unstructured text, not for general PII detection in S3.

629
MCQhard

A data scientist is evaluating two different foundation models for a summarization task. They want to compare the quality of summaries generated by each model against a set of human-written reference summaries. Which set of metrics is most appropriate for this automated evaluation?

A.Accuracy, precision, recall, and F1-score
B.ROUGE, BLEU, and BERTScore
C.Latency and throughput
D.Mean squared error (MSE) and R-squared
AnswerB

ROUGE, BLEU and BERTScore compare generated text against reference summaries using n-gram overlap and semantic similarity. This directly satisfies the stem's constraint of automated evaluation against human-written references, unlike classification metrics such as accuracy or F1.

Why this answer

ROUGE, BLEU, and BERTScore are specifically designed for evaluating text generation quality against reference summaries. ROUGE measures n-gram overlap (recall-oriented), BLEU measures precision of n-gram matches, and BERTScore uses contextual embeddings from BERT to capture semantic similarity, making them ideal for summarization tasks.

Exam trap

AWS often tests the distinction between evaluation metrics for classification (accuracy, F1) versus generation (ROUGE, BLEU, BERTScore), and candidates mistakenly apply classification metrics to summarization tasks.

How to eliminate wrong answers

Option A is wrong because accuracy, precision, recall, and F1-score are classification metrics used for tasks like spam detection or sentiment analysis, not for evaluating the quality of generated text against references. Option C is wrong because latency and throughput are performance metrics for system efficiency (e.g., inference speed), not for assessing summary quality. Option D is wrong because mean squared error (MSE) and R-squared are regression metrics for continuous value prediction, not for comparing text outputs.

630
MCQeasy

A startup wants to generate product descriptions from a few keywords using a foundation model. They need a fully managed serverless solution that requires no infrastructure setup. Which AWS service should they use?

A.Amazon SageMaker
B.Amazon Comprehend
C.AWS Lambda
D.Amazon Bedrock
AnswerD

Amazon Bedrock provides serverless access to foundation models through a single API, so the startup generates product descriptions from keywords without provisioning or managing any infrastructure. It satisfies the stem's fully managed, no-setup constraint directly, unlike self-hosted alternatives requiring instance or endpoint management.

Why this answer

Amazon Bedrock is a fully managed serverless service that provides access to foundation models (FMs) from leading AI providers via a simple API, making it ideal for generating product descriptions from keywords without any infrastructure management. It directly supports generative AI tasks like text generation, unlike other AWS services that focus on different ML or NLP capabilities.

Exam trap

The trap here is that candidates may confuse Amazon SageMaker's managed ML capabilities with a serverless generative AI service, overlooking that SageMaker requires explicit infrastructure setup for model hosting, while Bedrock is purpose-built for serverless access to foundation models.

How to eliminate wrong answers

Option A is wrong because Amazon SageMaker is a fully managed machine learning platform that requires setting up training jobs, endpoints, and infrastructure for custom models, not a serverless solution for directly using pre-built foundation models. Option B is wrong because Amazon Comprehend is a natural language processing (NLP) service for tasks like sentiment analysis and entity extraction, not for generative text creation from keywords. Option C is wrong because AWS Lambda is a serverless compute service that runs custom code but does not natively provide access to foundation models; you would need to integrate it with another service like Bedrock to generate descriptions, making it not a standalone solution for this use case.

631
MCQmedium

A company wants to reduce the cost of running a large number of inference requests for a text classification task. The responses can tolerate a slight delay. Which cost optimization strategy should they implement?

A.Store the training data in a vector store for faster retrieval
B.Use batch inference to process multiple requests asynchronously
C.Use a larger, more accurate model to reduce the number of retries
D.Implement model caching to reuse results for identical prompts
AnswerB

Batch inference submits many prompts as a single asynchronous job, removing the per-token real-time pricing premium and improving throughput per compute unit. Because the stem states responses tolerate slight delay, latency is not a constraint, so batching delivers the required cost reduction.

Why this answer

Batch inference processes multiple requests together, reducing cost per request at the expense of higher latency. Model caching can help if there are many repeated requests, but batch inference is more general. Switching to a larger model increases cost.

Using a vector store is unrelated.

632
MCQhard

A company using Amazon Bedrock needs to redact personally identifiable information (PII) from user inputs before sending them to the foundation model. Which Bedrock Guardrails component should be configured?

A.Content filters
B.Topic restrictions
C.Word filters
D.Sensitive information filters
AnswerD

Sensitive information filters detect and redact PII such as names, addresses and credit card numbers in prompts, directly satisfying the requirement to remove personal data before it reaches the foundation model. Configuring these filters with appropriate PII entities and action set to Mask achieves the redaction.

Why this answer

PII redaction in Guardrails automatically detects and redacts PII in user inputs or model responses.

633
MCQeasy

A developer is testing different prompts for a text generation model on Amazon Bedrock. Which parameter controls the randomness of the model's output?

A.top_p
B.stop_sequences
C.temperature
D.max_tokens
AnswerC

Temperature directly scales the sampling distribution's entropy: low values sharpen probabilities toward the highest-scoring token, high values flatten them, increasing output variability. This is precisely the randomness control the developer needs when testing prompts on Amazon Bedrock, unlike top-p, which truncates the candidate token set rather than rescaling probabilities.

Why this answer

Temperature directly controls the randomness of the model's output by scaling the logits before applying the softmax function. A higher temperature (e.g., 1.5) increases randomness and creativity, while a lower temperature (e.g., 0.1) makes the output more deterministic and focused. In Amazon Bedrock, this parameter is a core setting for text generation models like Anthropic Claude and Amazon Titan.

Exam trap

Candidates often confuse top_p (nucleus sampling) with randomness control, when temperature is the primary parameter for that purpose. Both affect output diversity but through different mechanisms.

How to eliminate wrong answers

Option A is wrong because top_p controls nucleus sampling, which limits the cumulative probability of token choices, not the overall randomness of the output. Option B is wrong because stop_sequences define specific strings that halt generation, not randomness. Option D is wrong because max_tokens sets the maximum number of tokens in the generated response, not the randomness of the output.

634
MCQhard

An organization is using Amazon Bedrock to power a customer service chatbot. They notice that the chatbot occasionally generates hallucinated information about product specifications. Which strategy should be implemented to reduce hallucinations?

A.Fine-tune the model on a dataset of product specification conversations.
B.Integrate a Retrieval Augmented Generation (RAG) system with the product catalog.
C.Use more detailed prompts with explicit instructions to avoid speculation.
D.Increase the temperature parameter to make outputs more conservative.
AnswerB

Integrating RAG grounds responses in retrieved product-catalogue content, so the model conditions on factual specifications rather than relying solely on parametric memory. This directly targets the hallucination source by supplying authoritative context at inference time, satisfying the requirement to reduce fabricated product details without retraining the foundation model.

Why this answer

Retrieval Augmented Generation (RAG) grounds the model's responses in authoritative, up-to-date product catalog data, directly reducing hallucinations by ensuring the chatbot references verified facts rather than relying solely on its parametric memory. This is the most effective strategy because it provides a retrieval-based factual foundation that fine-tuning or prompt engineering alone cannot guarantee.

Exam trap

The AIF-C01 exam often tests the misconception that prompt engineering or fine-tuning alone can solve hallucination problems, when in fact they lack the dynamic, verifiable grounding that RAG provides.

How to eliminate wrong answers

Option A is wrong because fine-tuning on product specification conversations may reinforce patterns from the training data but does not prevent the model from generating plausible-sounding but incorrect details when faced with queries outside the fine-tuned distribution; it also cannot dynamically incorporate real-time catalog updates. Option C is wrong because while more detailed prompts can reduce speculation, they do not provide the model with access to external, authoritative data—hallucinations can still occur when the model's internal knowledge is incomplete or outdated. Option D is wrong because increasing the temperature parameter makes outputs more random and creative, not more conservative; decreasing temperature would make outputs more deterministic and less prone to hallucination, but even low temperature cannot eliminate hallucinations without a retrieval mechanism.

635
MCQhard

A company uses Amazon Bedrock to deploy a foundation model for a real-time chat application. Users report that responses are slow. Which optimization is MOST likely to reduce latency without degrading quality?

A.Use a larger model variant to improve inference speed
B.Increase the temperature to make the model generate faster
C.Enable response streaming using the Converse API or InvokeModelWithResponseStream
D.Switch from a text generation model to an embedding model
AnswerC

Streaming returns tokens incrementally as the model generates them, so users see the first words almost immediately rather than waiting for the full completion. This directly addresses the real-time chat latency constraint while leaving model quality untouched, since the same foundation model and parameters are used.

Why this answer

Enabling response streaming with the Converse API or InvokeModelWithResponseStream allows the model to send tokens to the client as they are generated, rather than waiting for the full response. This reduces the user's perceived latency because the first token appears much sooner, even though the total generation time remains similar. This optimization directly addresses the real-time chat requirement without degrading output quality.

Exam trap

The trap here is that candidates confuse 'perceived latency' with 'total generation time' and assume that only model-level changes (like model size or parameters) can affect speed, overlooking the architectural optimization of streaming.

How to eliminate wrong answers

Option A is wrong because using a larger model variant typically increases inference time and latency, not reduces it, due to higher computational requirements. Option B is wrong because temperature controls the randomness of token selection, not the speed of generation; increasing it does not make the model generate faster and can degrade response quality. Option D is wrong because switching from a text generation model to an embedding model would fundamentally change the application's capability—embedding models produce vector representations, not conversational text, so the chat application would no longer function.

636
MCQeasy

A company is using Amazon Bedrock to build a generative AI application. The company wants to prevent the model from generating toxic or harmful content while still allowing creative responses. Which feature should the company enable?

A.Amazon Bedrock Guardrails with content filters.
B.AWS Key Management Service (KMS) to encrypt model responses.
C.AWS Identity and Access Management (IAM) policies to restrict model output.
D.Amazon CloudWatch Logs to monitor and block harmful content.
AnswerA

Guardrails applies configurable content filters that block toxic categories such as hate, violence and insults on both prompts and responses, while threshold tuning preserves creative output. This satisfies the requirement to prevent harmful generation without suppressing legitimate creative responses.

Why this answer

Amazon Bedrock Guardrails with content filters is the correct feature because it allows the company to define and enforce policies that block toxic or harmful content in model inputs and outputs, while still permitting creative responses within safe boundaries. This feature provides configurable thresholds for content categories like hate, insults, and sexual content, enabling precise control over model behavior without restricting overall creativity.

Exam trap

The trap here is that candidates may confuse security services (like KMS for encryption or IAM for access control) with content moderation capabilities, assuming any AWS security service can filter model outputs, when in fact only Bedrock Guardrails provides purpose-built content filters for generative AI.

How to eliminate wrong answers

Option B is wrong because AWS KMS encrypts data at rest and in transit but does not inspect or filter model responses for toxic content; encryption ensures confidentiality, not content safety. Option C is wrong because IAM policies control access to AWS resources and actions (e.g., who can invoke a model) but cannot restrict the actual text output of a model; they are for authorization, not content moderation. Option D is wrong because Amazon CloudWatch Logs can monitor and store logs for analysis but cannot actively block harmful content in real-time; it is a logging and monitoring service, not a content filter.

637
MCQeasy

A startup is deploying a foundation model on Amazon SageMaker for real-time inference. They notice high latency (over 2 seconds per request). Which action is most likely to reduce latency?

A.Enable auto-scaling on the SageMaker endpoint to handle more concurrent requests.
B.Switch to a smaller, distilled version of the model.
C.Deploy the model on a CPU-based instance instead of GPU.
D.Increase the batch size parameter in the inference request.
AnswerB

A smaller, distilled model reduces the number of parameters and floating-point operations per inference, directly cutting compute time on the SageMaker endpoint. This satisfies the stem's real-time latency constraint, since inference duration scales with model size; distillation preserves much of the original accuracy while delivering sub-second responses.

Why this answer

Using a smaller, distilled version of the model directly reduces the computational complexity per inference request. Distillation compresses the model by training a smaller student network to mimic a larger teacher model, resulting in fewer parameters and faster forward passes. This is the most direct way to cut latency when the model size is the bottleneck, as it reduces the number of floating-point operations (FLOPs) required per request.

Exam trap

AWS often tests the distinction between latency (time per single request) and throughput (requests per second), so candidates mistakenly choose auto-scaling or batch size increases, which improve throughput but not per-request latency.

How to eliminate wrong answers

Option A is wrong because enabling auto-scaling adds more endpoint instances to handle higher concurrency, but it does not reduce the latency of a single inference request; it only improves throughput under load. Option C is wrong because CPU-based instances are generally slower for deep learning inference than GPU instances, especially for large foundation models, so switching to CPU would increase latency, not reduce it. Option D is wrong because increasing the batch size in the inference request means processing multiple inputs together, which increases the time to first byte for each individual request and does not reduce per-request latency; it is a throughput optimization, not a latency reduction technique.

638
Multi-Selecteasy

A company needs to build a system that converts text-based customer reviews into audio files for accessibility. Which AWS service should be used?

Select 1 answer
A.Amazon Transcribe
B.Amazon Translate
C.Amazon Rekognition
D.Amazon Polly
E.Amazon Comprehend
AnswersD

Amazon Polly is a text-to-speech service that converts text into lifelike speech, making it the correct choice for generating audio files from customer reviews.

Why this answer

Amazon Polly is a text-to-speech (TTS) service that converts text into lifelike speech, making it the correct choice for generating audio files from text-based customer reviews. It supports multiple languages and voices, and can output audio in formats like MP3 or OGG, which is ideal for accessibility use cases. The other listed services do not convert text to audio: Amazon Transcribe converts speech to text, Amazon Translate translates text, Amazon Rekognition analyzes images and video, and Amazon Comprehend performs natural language processing.

Therefore, only one correct answer exists among the options.

Exam trap

The trap here is that candidates often confuse Amazon Transcribe (speech-to-text) with Amazon Polly (text-to-speech), or mistakenly think Amazon Comprehend can generate audio because it processes text, but neither service produces audio output.

639
Multi-Selectmedium

A company is using Amazon Bedrock Agents to build a travel booking assistant that can search for flights, book hotels, and answer questions about travel policies. Which TWO components are required to enable the agent to call external services? (Select TWO.)

Select 2 answers
A.Action group with API schema
B.AWS Lambda function
C.Amazon DynamoDB table
D.Amazon SageMaker endpoint
E.Bedrock Knowledge Base
AnswersA, B

An action group with an API schema defines the operations the agent may invoke and maps them to external endpoints, satisfying the requirement to call services such as flight search and hotel booking. Amazon Bedrock Agents uses the OpenAPI schema to determine parameters and orchestrate each API call during task execution.

Why this answer

Option A (Action group with API schema) is correct because an action group defines the set of tasks the Bedrock Agent can perform and includes an OpenAPI schema that describes the external APIs the agent can invoke, which is the mechanism for calling external services. Option B (AWS Lambda function) is correct because Bedrock Agents use Lambda functions as the compute layer to execute the business logic for action groups, receiving the agent's request and returning the API response. Option C (Amazon DynamoDB table) is not required; it is only a possible data store for the Lambda function's implementation, not a component that enables external service calls.

Option D (Amazon SageMaker endpoint) is not required; it is used for hosting custom ML models, not for enabling agent action groups. Option E (Bedrock Knowledge Base) is not required; it provides retrieval-augmented generation over data sources for answering questions, not for calling external services.

640
MCQeasy

What does the temperature parameter control in a text generation model?

A.The number of candidate tokens considered at each step
B.The degree of randomness in the generated output
C.The similarity to the training data distribution
D.The maximum number of tokens to generate
AnswerB

Temperature scales the probability distribution over the model's vocabulary before sampling, so higher values flatten it and increase randomness while lower values sharpen it toward the most likely tokens. This directly governs output variability, satisfying the question's focus on how stochastic versus deterministic the generated text becomes.

Why this answer

Temperature is a scaling factor applied to the logits before the softmax function, controlling the sharpness of the resulting probability distribution. A low temperature (e.g., 0.1) sharpens the distribution, making the highest-probability token overwhelmingly likely and producing deterministic, focused output. A high temperature (e.g., 1.5) flattens the distribution, giving lower-probability tokens more chance of being sampled, producing more diverse and creative — but potentially incoherent — output.

Exam trap

The trap is conflating temperature with other sampling controls — candidates who don't distinguish temperature (probability sharpening) from top-k/top-p (candidate restriction) or max_tokens (length limit) pick A or D.

How to eliminate wrong answers

Option A is wrong because the number of candidate tokens considered at each step is controlled by top-k or top-p (nucleus) sampling, not temperature — temperature reshapes probabilities across all tokens but doesn't restrict the candidate set. Option C is wrong because similarity to the training data distribution is an emergent property of the model's weights and training, not something temperature tunes at inference time. Option D is wrong because the maximum output length is controlled by the max_tokens (or max_new_tokens) parameter, which is entirely separate from the sampling distribution.

641
MCQeasy

A developer wants to quickly test different prompts and models for a text summarization task without writing any code. Which AWS service should they use?

A.Amazon Bedrock Playground
B.Amazon SageMaker Studio
C.Amazon Bedrock Agents
D.Amazon Bedrock Knowledge Bases
AnswerA

Amazon Bedrock Playground provides a no-code console for experimenting with foundation models, letting the developer adjust prompts and switch between models interactively. This directly satisfies the stem's constraints: rapid prompt and model testing for summarisation without writing any code, unlike API-based or SDK-driven approaches.

Why this answer

Amazon Bedrock Playground is a no-code interface within the AWS Management Console that allows developers to interactively test different foundation models and prompts for tasks like text summarization. It provides a chat-like environment where you can select models, adjust parameters (e.g., temperature, max tokens), and see responses immediately without writing any code. This makes it the ideal choice for rapid experimentation and prototyping.

Exam trap

Candidates often confuse the Bedrock Playground (no-code) with Amazon SageMaker Studio (requires coding). The trick is that SageMaker Studio is for building, training, and deploying ML models, not for quick no-code prompt testing. The Playground is the correct answer for rapid experimentation without code.

How to eliminate wrong answers

Option B is wrong because Amazon SageMaker Studio is a full-featured integrated development environment (IDE) for building, training, and deploying machine learning models, which requires writing code (e.g., Python notebooks) and is not designed for no-code prompt testing. Option C is wrong because Amazon Bedrock Agents is a service for creating autonomous agents that can orchestrate multi-step tasks using foundation models, APIs, and knowledge bases, not for quick no-code prompt experimentation. Option D is wrong because Amazon Bedrock Knowledge Bases is a service for creating a retrievable knowledge store that foundation models can query to augment responses (RAG), not a tool for directly testing prompts and models without code.

642
MCQeasy

A startup wants to add image generation to its design tool. The team does not want to manage GPU infrastructure, train models, or host endpoints, and they want to call a managed API for a text-to-image foundation model available in Amazon Bedrock. Which action should they take?

A.Use an Amazon Bedrock image generation model through the InvokeModel API with a text prompt.
B.Use Amazon Polly to convert text prompts into images for the design tool.
C.Use Amazon Rekognition to synthesize new design images from text descriptions.
D.Deploy an open-source diffusion model on Amazon SageMaker endpoints and manage scaling themselves.
AnswerA

Amazon Bedrock provides serverless access to image generation foundation models, including Amazon Titan Image Generator, through the InvokeModel API. The startup sends a text prompt and receives generated images without provisioning GPUs or managing endpoints. This matches the requirement for a managed text-to-image API with no infrastructure ownership.

Why this answer

Amazon Bedrock exposes image generation foundation models such as Amazon Titan Image Generator as managed, serverless APIs invoked with a text prompt. This delivers text-to-image capability without GPU provisioning, training, or endpoint management. SageMaker self-hosting adds operational work, while Rekognition analyzes images and Polly synthesizes speech, so neither can generate design images from text.

Exam trap

The trap here is matching the word image to Amazon Rekognition, which analyzes existing images rather than generating new ones from text.

643
Multi-Selectmedium

A startup is building a semantic search system over their product catalog using Amazon Bedrock. They want to convert product descriptions into vector embeddings and store them in a vector database for similarity search. Which TWO actions should they take? (Select TWO.)

Select 2 answers
A.Store the embeddings in a vector database such as Amazon OpenSearch Serverless with the k-NN plugin
B.Use Amazon Titan Image Generator to create embeddings for each product image
C.Use Amazon Bedrock InvokeModel with Anthropic Claude to generate embeddings
D.Store the embedding vectors in an Amazon DynamoDB table as binary attributes
E.Use Amazon Titan Embeddings to generate vector embeddings from product descriptions
AnswersA, E

Amazon OpenSearch Serverless with the k-NN plugin provides the vector index that stores embeddings and performs approximate nearest-neighbour similarity search, satisfying the stem's requirement for a vector database. Bedrock generates the embeddings; this service persists and queries them, returning semantically similar product descriptions by distance in vector space.

Why this answer

Option E is correct because Amazon Titan Embeddings (e.g., amazon.titan-embed-text-v1/v2) is a purpose-built Bedrock text embedding model that converts product descriptions into dense vector embeddings suitable for semantic similarity search. Option A is correct because Amazon OpenSearch Serverless with the k-NN plugin is a managed vector database that supports storing embeddings and running approximate nearest-neighbor (ANN) similarity queries, which is exactly what the semantic search system requires. Option B is wrong because Titan Image Generator produces images from prompts, not embeddings, and it would not generate text-description embeddings.

Option C is wrong because Anthropic Claude is a text generation/reasoning model on Bedrock and does not expose an embeddings API via InvokeModel. Option D is wrong because DynamoDB is a key-value/document store without native vector similarity search, so storing vectors as binary attributes would not support k-NN retrieval.

Exam trap

AWS often tests the distinction between generative models (like Claude) and embedding models (like Titan Embeddings), so candidates mistakenly think any Bedrock model can produce embeddings via InvokeModel, but only specific embedding models output vectors.

644
Multi-Selecthard

Which TWO of the following are valid methods to reduce the risk of foundation models generating harmful or biased content?

Select 2 answers
A.Use a smaller model
B.Use a content filter
C.Apply prompt engineering to guide output
D.Fine-tune the model on a biased dataset
E.Disable all logging
AnswersB, C

Content filters intercept prompts and completions at runtime, blocking harmful or biased output before it reaches users. This directly satisfies the stem's requirement to reduce risk from foundation models, since filtering operates on the model's generated content itself rather than on training data or access controls.

Why this answer

Option B (Use a content filter) is correct because content filters act as a post-processing guardrail that screens both prompts and model completions for harmful, violent, hateful, or otherwise policy-violating content, blocking or redacting it before it reaches users. Option C (Apply prompt engineering to guide output) is correct because carefully crafted system prompts, few-shot examples, and instructions can steer the foundation model toward safe, neutral, and on-topic responses, reducing the likelihood of biased or harmful generations. Option A (Use a smaller model) is not a valid mitigation because model size does not determine safety or bias; smaller models can still produce harmful or biased content.

Option D (Fine-tune the model on a biased dataset) would actually increase the risk by reinforcing biased patterns in the model's outputs. Option E (Disable all logging) does not reduce harmful content generation and instead removes the audit trail needed to detect, monitor, and remediate such issues.

Exam trap

AWS often tests the misconception that simply using a smaller model or disabling logging can reduce bias, when in fact these actions either have no effect or worsen the problem, whereas content filters and prompt engineering are direct, effective mitigation strategies.

645
MCQmedium

An e-commerce company uses Amazon Bedrock to generate product descriptions from keywords. Some descriptions contain inaccurate details about product specifications. Which approach should the company take to reduce factual errors?

A.Increase the maxTokens parameter to allow more detailed descriptions.
B.Use a different foundation model from Bedrock for each product category.
C.Deploy the model to a SageMaker endpoint and use human-in-the-loop validation.
D.Include the product specifications in the prompt and instruct the model to base the description on the provided data.
AnswerD

Supplying the specifications directly in the prompt grounds generation in authoritative source data, so the model conditions its output on those facts rather than relying on parametric knowledge that may be outdated or hallucinated. This directly satisfies the stem's constraint of reducing inaccurate specification details, since the model is instructed to base descriptions solely on the provided data.

Why this answer

Providing the product specifications directly in the prompt and instructing the model to base the description on that data grounds the generation in factual information, reducing hallucinations. This technique, known as prompt engineering with in-context learning, ensures the model uses the given data rather than relying on its training data, which may contain inaccuracies.

Exam trap

AWS often tests the misconception that increasing model parameters or changing models alone improves factual accuracy, when in fact prompt engineering with grounded data is the most effective and efficient method to reduce hallucinations.

How to eliminate wrong answers

Option A is wrong because increasing maxTokens only allows longer outputs but does not improve factual accuracy; it may even increase the chance of hallucinations by generating more unverified content. Option B is wrong because using a different foundation model for each category does not inherently reduce factual errors; all models can hallucinate, and this approach adds complexity without addressing the root cause of inaccurate specifications. Option C is wrong because deploying to a SageMaker endpoint with human-in-the-loop validation is an operational pattern for custom models, but it is overkill and inefficient for this use case; prompt engineering (Option D) is a simpler, more direct solution that avoids the latency and cost of human review for every generation.

646
MCQeasy

Which AWS service is BEST suited for extracting text from scanned PDF documents, such as invoices and receipts?

A.Amazon Comprehend
B.Amazon Textract
C.Amazon Rekognition
D.Amazon Transcribe
AnswerB

Amazon Textract uses optical character recognition with layout analysis to extract text, forms and tables from scanned documents such as invoices and receipts. This directly satisfies the requirement for extracting text from scanned PDFs, unlike generic storage or compute services.

Why this answer

Amazon Textract is specifically designed to extract text, handwriting, and data from scanned documents like invoices and receipts. It uses machine learning to recognize and extract printed text, forms, and tables from document images, going beyond simple OCR by understanding the structure of the document.

Exam trap

The trap is that candidates often confuse Amazon Rekognition's text-in-image capability with Amazon Textract's document-specific extraction. However, Rekognition lacks the ability to extract text from structured documents like invoices and receipts, as it is designed for image analysis rather than document processing.

How to eliminate wrong answers

Option A is wrong because Amazon Comprehend is a natural language processing (NLP) service that extracts insights and relationships from text, but it cannot extract text from scanned images or PDFs. Option C is wrong because Amazon Rekognition is primarily for image and video analysis, including object detection and facial recognition, and while it can detect text in images, it is not optimized for extracting structured data from dense documents like invoices. Option D is wrong because Amazon Transcribe is an automatic speech recognition (ASR) service that converts audio to text, and it has no capability to process scanned PDFs or images.

647
MCQeasy

A developer wants to test different foundation models quickly without setting up infrastructure. Which AWS service allows interactive prompting and comparison of multiple models?

A.Amazon Comprehend
B.Amazon Bedrock Playground
C.Amazon Lex
D.Amazon SageMaker Studio
AnswerB

Amazon Bedrock Playground provides a console interface for interactive prompting and side-by-side comparison of multiple foundation models without provisioning any infrastructure. This directly satisfies the stem's requirement to test different models quickly, since no servers, endpoints or deployment configuration are needed before experimentation.

Why this answer

Amazon Bedrock Playground is a feature within Amazon Bedrock that provides a web-based interface for interactive prompting and side-by-side comparison of multiple foundation models (FMs). It allows developers to test different models quickly without provisioning any infrastructure, making it ideal for rapid experimentation and evaluation.

Exam trap

The trap here is that candidates may confuse Amazon Bedrock Playground with SageMaker Studio, assuming both are for model experimentation, but SageMaker Studio requires infrastructure setup and lacks the built-in multi-model comparison interface that Bedrock Playground provides.

How to eliminate wrong answers

Option A is wrong because Amazon Comprehend is a natural language processing (NLP) service for extracting insights like sentiment, entities, and key phrases from text; it does not support interactive prompting or comparison of foundation models. Option C is wrong because Amazon Lex is a service for building conversational interfaces (chatbots) using automatic speech recognition (ASR) and natural language understanding (NLU), not for testing or comparing foundation models. Option D is wrong because Amazon SageMaker Studio is an integrated development environment (IDE) for building, training, and deploying machine learning models, but it requires setting up infrastructure (e.g., instances, kernels) and does not provide a built-in interactive playground for comparing multiple foundation models.

648
MCQeasy

A startup is building a customer support chatbot using Amazon Bedrock with the Claude foundation model. The chatbot needs to answer questions based on a knowledge base of frequently asked questions (FAQs) stored in an Amazon S3 bucket. The team wants to implement Retrieval Augmented Generation (RAG) to provide accurate and context-aware responses. They are evaluating different approaches to integrate the knowledge base. What is the most efficient way to implement RAG with Bedrock?

A.Use AWS Lambda to fetch documents from S3 and inject them into the prompt.
B.Manually extract all FAQs and include them in the prompt each time the chatbot responds.
C.Fine-tune the Claude model on the FAQs so the model memorizes the knowledge base.
D.Use Amazon Bedrock Knowledge Bases to directly connect the S3 bucket and retrieve relevant documents for the prompt.
AnswerD

Bedrock Knowledge Bases handles the full RAG pipeline natively: it ingests the S3 bucket, chunks and embeds documents, stores vectors, and retrieves relevant passages into the prompt. This removes the need to build custom chunking, embedding, and vector-store orchestration, making it the most efficient route.

Why this answer

Amazon Bedrock Knowledge Bases is a fully managed RAG capability that natively connects to Amazon S3 (and other data sources), automatically chunks and embeds documents into a vector store, and retrieves the most relevant passages at query time to augment the prompt. This eliminates the need to build custom retrieval pipelines, manage embeddings, or manually inject documents. It is the most efficient, purpose-built approach for RAG on Bedrock.

Exam trap

AIF-C01 often tests the misconception that fine-tuning is the way to inject factual knowledge into a model, when in fact RAG (via Bedrock Knowledge Bases) is the correct pattern for dynamic, grounded question answering.

How to eliminate wrong answers

Option A is wrong because using Lambda to fetch documents from S3 and inject them into the prompt is a custom, manual retrieval approach that lacks semantic search, chunking, and vector indexing — it is inefficient and unscalable. Option B is wrong because manually including all FAQs in every prompt is impractical, exceeds context window limits, increases cost, and does not scale as the knowledge base grows. Option C is wrong because fine-tuning teaches the model style and patterns, not factual retrieval; it cannot reliably memorize a dynamic knowledge base and risks hallucination and stale information.

649
MCQeasy

A company needs to audit all API calls made to Amazon Bedrock, including model invocations and guardrail evaluations. Which AWS service should they enable to capture these API calls for compliance?

A.AWS CloudTrail
B.Amazon Macie
C.AWS Config
D.Amazon GuardDuty
AnswerA

AWS CloudTrail records Bedrock control-plane and data-plane API activity, including model invocations and guardrail evaluations, as audit events. It satisfies the compliance requirement to capture every API call by logging requests, identities, timestamps and source IPs to an immutable trail, which can then be queried or exported for auditing.

Why this answer

AWS CloudTrail records API calls for all AWS services, including Bedrock. It captures the caller identity, API, parameters, and response elements, which can be used for auditing and compliance.

650
Multi-Selecthard

A company is deploying a machine learning model to detect fraudulent transactions. The dataset is highly imbalanced (1% fraud). The team needs to evaluate model performance and minimize false positives while maintaining high recall. Which TWO metrics should they focus on? (Select TWO.)

Select 2 answers
A.Accuracy
B.Recall
C.Precision
D.F1 score
E.AUC-ROC
AnswersB, C

Recall measures the proportion of actual frauds correctly identified, directly addressing the requirement to maintain high recall on the 1% minority class. It exposes missed frauds (false negatives), which are the costly errors when the positive class is rare.

Why this answer

Option B (Recall) is correct because with only 1% fraud, recall measures the proportion of actual fraudulent transactions the model successfully catches, which directly supports the requirement to maintain high recall and avoid missing fraud cases. Option C (Precision) is correct because minimizing false positives means reducing legitimate transactions incorrectly flagged as fraud, and precision is exactly the ratio of true positives to all positive predictions (TP / (TP + FP)), so higher precision directly reflects fewer false alarms. Accuracy (A) is misleading on a 1% imbalanced dataset, since a model predicting 'no fraud' for everything would score 99% accuracy while catching zero fraud.

F1 score (D) balances precision and recall but does not by itself target the specific goal of minimizing false positives while keeping recall high. AUC-ROC (E) summarizes ranking performance across thresholds but does not directly express the false-positive/recall trade-off the team must optimize.

Exam trap

A common misconception is that accuracy is a reliable metric for imbalanced datasets, or that F1 score alone is sufficient when the question explicitly asks for two separate metrics to independently control false positives and recall.

651
Multi-Selectmedium

A logistics company is planning its first machine learning project and the leadership team asks which statements correctly describe fundamental machine learning concepts. Which TWO statements are accurate? (Choose two.)

Select 2 answers
A.In supervised learning, the model learns from input-output pairs where the desired output is provided during training.
B.A machine learning model's predictions are deterministic rules written by engineers rather than patterns inferred from data.
C.Reinforcement learning requires a fully labeled dataset of correct actions for every possible situation before training can begin.
D.Model training always improves accuracy on unseen data as more epochs are run, so training should continue until loss reaches zero.
E.Unsupervised learning can identify patterns or groupings in data that has no predefined labels.
AnswersA, E

Supervised learning depends on labeled examples that pair each input with the correct output, allowing the algorithm to adjust its parameters to minimize error against those known targets. This is exactly how tasks such as delivery-time prediction or package damage classification are trained. Without the provided outputs, the model would have no error signal to learn from, so this statement correctly captures the core of supervised learning.

Why this answer

Supervised learning is defined by training on input-output pairs with known targets, while unsupervised learning finds structure in data without any labels. Both statements describe core, vendor-neutral machine learning concepts that a logistics team should understand before scoping its first project.

Exam trap

The trap here is accepting the common myth that more training always helps, or confusing reinforcement learning with supervised learning that needs labeled correct actions.

652
MCQmedium

A financial services company has deployed a machine learning model that approves or denies loan applications in real time. The compliance team requires that any applicant who is denied must receive a meaningful explanation of the decision, and the company must be able to prove which model version and input features produced each decision for audit purposes. Which AWS service should the company use to capture the model's feature attributions and store them for each inference request?

A.Amazon SageMaker Experiments to track the training runs
B.Amazon SageMaker Model Monitor with a data quality baseline
C.Amazon SageMaker Clarify with online explainability enabled on the endpoint
D.AWS CloudTrail data events on the SageMaker endpoint
AnswerC

SageMaker Clarify online explainability runs within the endpoint and returns a feature attribution for each individual request, so the company can attach a per-applicant explanation to each denial decision. Combined with endpoint data capture writing to Amazon S3, this produces the per-inference audit record the compliance team requires, tying each decision to the model version and the input features that drove it.

Why this answer

Per-decision explainability requires a capability that computes feature attributions at inference time and persists them for audit. SageMaker Clarify online explainability does exactly this inside the endpoint, returning a SHAP-based attribution for each request, and endpoint data capture stores the request and response in Amazon S3. Together they satisfy both the applicant-facing explanation requirement and the internal audit trail obligation.

Exam trap

The trap here is assuming that a monitoring service such as SageMaker Model Monitor produces per-request explanations, when it only aggregates drift and quality statistics across traffic.

653
Multi-Selectmedium

Which TWO factors are most important when selecting a foundation model in Amazon Bedrock for a text summarization task with strict latency requirements?

Select 2 answers
A.Average response latency per request.
B.Model size in billions of parameters.
C.Maximum input token limit.
D.Output quality and token efficiency for summarization tasks.
E.Availability of fine-tuning capability for domain adaptation.
AnswersA, D

Average response latency per request directly measures whether a candidate model can satisfy the stem's strict latency requirement, since Bedrock models vary widely in inference speed. Selecting on this metric ensures the summarisation workload meets its response-time constraint rather than optimising quality alone.

Why this answer

Option A is correct because with strict latency requirements the deciding metric is the model's average response latency per request, which directly determines whether the summarization workload can meet its service-level objective. Option D is correct because output quality and token efficiency determine whether the summary is usable and how many tokens must be generated, and fewer generated tokens translate directly into lower end-to-end latency. Option B is not decisive on its own: parameter count correlates loosely with latency but does not guarantee it, since architecture, hardware, and serving optimizations also matter.

Option C, the maximum input token limit, matters for handling long documents but does not address the strict latency constraint. Option E, fine-tuning availability, supports domain adaptation and accuracy rather than latency, so it is not one of the two most important factors here.

Exam trap

A common misconception is that model size (parameters) is the primary driver of latency, but in practice, latency depends on inference optimization, model quantization, and hardware, not just parameter count.

654
MCQhard

A software company is deploying a generative AI assistant on Amazon Bedrock. They need the assistant to include citations to source documents in its answers and to avoid answering when no supporting document is found. Which configuration should they use?

A.Enable model invocation logging and parse the logs to extract source references.
B.Use a higher temperature setting so the model explores more possible answers, including citations.
C.Configure an Amazon Bedrock knowledge base and enable citation generation, and instruct the model to answer only from retrieved context.
D.Increase the model's maximum token count so it can include citations in its response.
AnswerC

Amazon Bedrock knowledge bases can return citations that link generated statements to the retrieved source chunks, and a system prompt can instruct the model to answer only when supporting context exists. This satisfies both requirements: source citations and abstention when no document is found. It leverages built-in retrieval and citation features rather than custom post-processing.

Why this answer

An Amazon Bedrock knowledge base with citation generation returns responses that reference the retrieved source chunks, and a prompt that restricts answers to retrieved context makes the assistant abstain when no supporting document exists. This combination meets both the attribution and the no-support behavior requirements.

Exam trap

The trap here is assuming that logging or sampling parameters can produce citations, when attribution and abstention depend on retrieval configuration and prompt constraints.

655
MCQmedium

A company wants to build a customer service chatbot that answers questions about their internal policy documents. The documents are updated monthly, and the team cannot afford to retrain a model each time. Which approach is MOST appropriate?

A.Use a larger foundation model with a longer context window and paste all documents into each prompt
B.Train a custom model from scratch on the policy documents each month
C.Fine-tune a base LLM on the policy documents monthly
D.Use Retrieval-Augmented Generation (RAG) with the policy documents indexed in a vector store
AnswerD

RAG retrieves relevant passages from the vector store at query time and supplies them as context, so monthly document updates only require re-indexing, not retraining. This satisfies the constraint that the team cannot afford to retrain the model each time.

Why this answer

Retrieval-Augmented Generation (RAG) is the most appropriate approach because it allows the chatbot to answer questions based on the latest policy documents without retraining the model. RAG retrieves relevant document chunks from a vector store at inference time and injects them into the prompt, enabling the model to ground its responses in the most current information. This avoids the cost and latency of monthly fine-tuning or retraining, and it scales efficiently as documents are updated.

Exam trap

The AWS AI Practitioner exam often tests the misconception that fine-tuning is the only way to adapt a model to new data. The trap here is that RAG provides a cheaper, faster, and more maintainable solution for frequently updated knowledge bases without retraining.

How to eliminate wrong answers

Option A is wrong because pasting all documents into each prompt is impractical due to context window limits (typically 4K–32K tokens) and rapidly increasing costs; policy documents can easily exceed these limits, causing truncation and loss of relevant information. Option B is wrong because training a model from scratch each month is prohibitively expensive (requiring thousands of GPU hours and massive datasets) and unnecessary for a task that only needs to reference a changing document set. Option C is wrong because fine-tuning a base LLM monthly on the policy documents would still require significant compute resources and time, and it risks catastrophic forgetting of general knowledge; moreover, fine-tuning does not efficiently handle monthly updates because the model must be retrained on the entire corpus each time.

656
Multi-Selecteasy

Which TWO actions can help reduce the likelihood of hallucinations in a generative AI model used for question answering?

Select 2 answers
A.Increase the maximum token count to allow more complete answers.
B.Use Retrieval Augmented Generation (RAG) with a trusted knowledge base.
C.Fine-tune the model on the training data used for the application.
D.Set a lower temperature parameter (e.g., 0.1) to reduce randomness.
E.Use a larger foundation model with more parameters.
AnswersB, D

RAG grounds generation in retrieved passages from a trusted knowledge base, so answers are conditioned on verifiable source text rather than parametric memory alone. This directly reduces fabrication because the model cites retrieved evidence, satisfying the requirement to lower hallucination likelihood in question answering.

Why this answer

Option B is correct because Retrieval Augmented Generation (RAG) grounds the model's responses in documents retrieved from a trusted knowledge base, so answers are conditioned on verifiable source content rather than the model's parametric memory, which substantially reduces fabricated or unsupported statements. Option D is correct because lowering the temperature parameter (e.g., to 0.1) sharpens the next-token probability distribution, making the model select high-probability, more deterministic tokens instead of sampling low-probability alternatives that often produce invented facts. Option A does not belong because increasing the maximum token count only allows longer outputs; it does not improve factual grounding and can even give hallucinations more room to expand.

Option C does not belong because fine-tuning on the application's training data can reinforce patterns and biases in that data and does not guarantee factual accuracy, and may even increase confident hallucination. Option E does not belong because a larger foundation model with more parameters may be more fluent but is not inherently less prone to hallucination, since scale alone does not provide source grounding or reduce sampling randomness.

Exam trap

AWS often tests the misconception that simply increasing model size or output length improves answer quality, when in fact grounding through RAG and controlling randomness via temperature are the direct mechanisms to reduce hallucinations.

657
MCQhard

A company uses an LLM to summarize medical research papers. They are concerned about hallucinations. Which combination of techniques would most effectively reduce hallucinations in this context?

A.Increase temperature and top-p sampling parameters
B.Use a smaller model with less capacity
C.Few-shot prompting and fine-tuning on more data
D.Retrieval-Augmented Generation (RAG) and Bedrock Guardrails
AnswerD

Retrieval-Augmented Generation grounds each summary in retrieved source passages, so the model cites actual paper content rather than relying on parametric memory. Bedrock Guardrails then applies contextual grounding checks, filtering responses unsupported by those passages. Together they directly target the factual-accuracy constraint, reducing fabricated medical claims.

Why this answer

Retrieval-Augmented Generation (RAG) grounds the model in retrieved documents, and Bedrock Guardrails can enforce content policies and factuality checks, together reducing hallucinations.

658
MCQeasy

A company wants to automatically summarize customer support tickets into a short paragraph. Which AWS service is MOST appropriate for this task?

A.Amazon Bedrock
B.Amazon Rekognition
C.Amazon Polly
D.Amazon Comprehend
AnswerA

Amazon Bedrock offers access to foundation models that perform abstractive summarisation, condensing ticket text into a short paragraph. It provides a managed, serverless API without infrastructure management, making it the most appropriate service for this natural language generation task.

Why this answer

Amazon Bedrock provides access to foundation models (FMs) from providers like Anthropic and AI21 that excel at natural language generation tasks, including summarization. By invoking a model such as Claude or Jurassic-2 via Bedrock's API, you can pass the customer support ticket text and receive a concise paragraph summary. This makes Bedrock the most appropriate service for generative summarization.

Exam trap

The trap here is that candidates may confuse Amazon Comprehend's text analysis features (like key phrase extraction) with generative summarization, but Comprehend cannot produce new text; it only extracts or classifies existing content.

How to eliminate wrong answers

Option B (Amazon Rekognition) is wrong because it is designed for image and video analysis, not text summarization. Option C (Amazon Polly) is wrong because it converts text to speech, not text summarization. Option D (Amazon Comprehend) is wrong because it performs natural language processing tasks like entity extraction and sentiment analysis, but it does not generate new text summaries; it lacks generative capabilities.

659
MCQeasy

A retail company uses a recommendation system that occasionally suggests inappropriate products to minors. Which responsible AI practice should be applied?

A.Implement human review of flagged recommendations
B.Rely solely on user feedback to improve
C.Disable the recommendation system entirely
D.Increase the volume of training data
AnswerA

Human review of flagged recommendations directly satisfies the need to catch inappropriate outputs before minors see them. Automated filters alone cannot reliably judge context, so routing borderline cases to reviewers provides the oversight layer that mitigates harm, matching the responsible AI principle of accountability and safety in this retail scenario.

Why this answer

The correct practice is to implement human review of flagged recommendations. This aligns with the responsible AI principle of accountability, where automated systems must have oversight mechanisms to catch and correct inappropriate outputs, especially when minors are involved. Human-in-the-loop (HITL) validation ensures that edge cases or subtle context (e.g., age-inappropriate product suggestions) are caught before they reach end users, rather than relying solely on automated filters or feedback loops.

Exam trap

AWS often tests the misconception that more data or automation alone can solve fairness and safety issues, when in fact responsible AI requires explicit governance mechanisms like human oversight for high-stakes or vulnerable-user scenarios.

How to eliminate wrong answers

Option B is wrong because relying solely on user feedback to improve is reactive and can expose minors to harm before any corrective action is taken; feedback loops are slow and may not capture subtle or rare inappropriate recommendations. Option C is wrong because disabling the recommendation system entirely is an extreme, non-scalable response that eliminates business value and does not teach the system to behave responsibly; responsible AI aims to mitigate harm, not abandon functionality. Option D is wrong because increasing the volume of training data does not inherently address the problem of inappropriate recommendations; if the training data itself contains biased or unlabeled age-sensitive content, more data can amplify the issue rather than fix it.

660
MCQmedium

A company uses Amazon SageMaker to train a model. The training job fails with 'InsufficientInstanceCapacity' error. What is the most likely cause?

A.The request rate is too high.
B.The dataset size exceeds the instance storage limit.
C.The requested instance type is not available in the specified region.
D.The training image is not compatible with the instance type.
AnswerC

SageMaker provisions the requested instance type from capacity within the chosen region; when that pool is exhausted, the training job fails with InsufficientInstanceCapacity. The constraint is regional instance availability, not IAM permissions, data format or hyperparameters, so switching instance type or region resolves it.

Why this answer

The 'InsufficientInstanceCapacity' error in Amazon SageMaker indicates that AWS does not currently have enough available capacity for the requested instance type in the specified region or Availability Zone. This is a common transient error when demand for a particular instance type exceeds supply, and it is not related to request rate, dataset size, or image compatibility.

Exam trap

The AIF-C01 exam often tests the distinction between capacity errors and throttling errors, so the trap here is confusing 'InsufficientInstanceCapacity' with a rate-limiting or quota error, leading candidates to incorrectly select Option A.

How to eliminate wrong answers

Option A is wrong because 'InsufficientInstanceCapacity' is a capacity error, not a throttling error; throttling (e.g., from high request rate) would return a 'ThrottlingException' or 'RequestLimitExceeded' error. Option B is wrong because dataset size exceeding instance storage limits would cause an 'OutOfMemory' or 'DiskFull' error, not a capacity error. Option D is wrong because image compatibility issues would result in a 'ClientError' or 'ImageNotFoundException', not an instance capacity error.

661
MCQhard

A financial services company uses a machine learning model to approve loan applications. The model is a gradient boosting classifier trained on historical loan data. Recently, the company noticed that the model's approval rate for applicants from a certain demographic group is significantly lower than for other groups, even though the model's overall accuracy remains high. The data science team has been asked to address this potential bias while minimizing the impact on overall model performance. The team has access to the training data and the trained model. They have limited time and budget. Which course of action should the team take first?

A.Remove the sensitive attribute from the training data and retrain the model.
B.Collect more data from the under-represented demographic group and retrain the model.
C.Analyze the training data for bias and retrain the model using bias mitigation techniques such as reweighting.
D.Adjust the model's decision threshold for the affected group after deployment.
AnswerC

Inspecting the training data for demographic imbalance or proxy features reveals the bias source before any modelling change. Reweighting or resampling then corrects the disparity at its root, which is cheaper and faster than post-hoc threshold tuning or full model replacement.

Why this answer

The first step should be to analyze the training data for bias and then retrain the model using bias mitigation techniques such as reweighting, because this addresses the root cause of the disparity while aiming to preserve overall performance. Reweighting adjusts sample weights to reduce the influence of biased historical patterns, and it can be applied with limited time and budget since it works with existing data and model. This is a targeted, cost-effective first action before considering more expensive data collection or post-deployment adjustments.

Exam trap

AIF-C01 often tests the misconception that removing the sensitive attribute eliminates bias; candidates must recognize that proxy variables and historical bias persist, and that data analysis plus mitigation techniques is the correct first step.

How to eliminate wrong answers

Option A is wrong because simply removing the sensitive attribute does not eliminate bias — the model can still learn proxy features correlated with the sensitive attribute, and it may reduce accuracy without fixing the disparity. Option B is wrong because collecting more data from the under-represented group is expensive and time-consuming, and the question states the team has limited time and budget. Option D is wrong because adjusting the decision threshold after deployment is a post-hoc fix that can mask bias but does not address the underlying model or data issues, and it may not be sustainable or compliant.

662
Multi-Selecthard

A company is training a deep learning model for image classification. Which THREE practices help reduce overfitting? (Choose three.)

Select 3 answers
A.L2 regularization
B.Increasing model depth
C.Increasing learning rate
D.Dropout
E.Data augmentation
AnswersA, D, E

L2 regularization penalises large weights by adding a squared-magnitude term to the loss function, constraining the model's capacity to memorise training noise. This directly satisfies the stem's requirement to reduce overfitting in deep image classification, since simpler weight distributions generalise better to unseen images.

Why this answer

L2 regularization (option A) is correct because it adds a penalty proportional to the squared magnitude of the weights to the loss function, discouraging large weights and thus reducing the model's ability to memorize training noise. Dropout (option D) is correct because it randomly deactivates a fraction of neurons during each training step, forcing the network to learn redundant, more robust feature representations instead of relying on specific units. Data augmentation (option E) is correct because it synthetically expands the training set by applying label-preserving transformations (e.g., flips, rotations, crops) to images, exposing the model to more varied inputs and improving generalization.

Increasing model depth (option B) is not correct here because adding layers increases capacity, which typically makes overfitting worse rather than better. Increasing the learning rate (option C) is not correct because it only affects optimization speed and stability, and an excessively high rate can cause divergence or poor convergence, not reduced overfitting.

Exam trap

The AIF-C01 exam often tests the misconception that increasing model complexity (depth) or tuning the learning rate can mitigate overfitting, when in fact these changes either exacerbate the problem or address unrelated training dynamics.

663
MCQeasy

Which AWS service allows teams to collaborate on building generative AI applications with a visual interface, share prompts, and manage prompt versions?

A.Amazon Bedrock Playground
B.Amazon Bedrock Studio
C.AWS CodeCommit
D.Amazon SageMaker Studio
AnswerB

Amazon Bedrock Studio provides a browser-based visual workspace where teams build generative AI applications, experiment with and share prompts, and manage prompt versions collaboratively. This directly satisfies the stem's requirements for a visual interface, prompt sharing and version management.

Why this answer

Amazon Bedrock Studio is the correct service because it provides a visual interface specifically designed for teams to collaborate on building generative AI applications. It allows users to share prompts, manage prompt versions, and iterate on foundation model configurations in a team-based environment, directly matching the question's requirements.

Exam trap

AWS exams often test the distinction between a single-user testing tool (Playground) and a collaborative development environment (Studio), leading candidates to mistakenly choose Amazon Bedrock Playground because they overlook the 'team collaboration' and 'version management' keywords.

How to eliminate wrong answers

Option A is wrong because Amazon Bedrock Playground is a single-user, experimental interface for testing foundation models with individual prompts, lacking team collaboration features, prompt sharing, or version management. Option C is wrong because AWS CodeCommit is a source control service for storing Git repositories, not a visual interface for building generative AI applications or managing prompts. Option D is wrong because Amazon SageMaker Studio is an integrated development environment (IDE) for machine learning workflows, including model training and deployment, but it does not provide a dedicated visual interface for generative AI prompt collaboration or version management.

664
MCQeasy

A company wants to prevent an Amazon Bedrock chatbot from discussing specific prohibited topics like competitor pricing. Which Bedrock feature should they configure?

A.Bedrock Knowledge Base
B.Bedrock Guardrails – topic denial
C.Bedrock Agents
D.Bedrock Model Evaluation
AnswerB

Bedrock Guardrails topic denial lets you define denied topics with natural-language descriptions, then blocks prompts and responses matching them. This directly satisfies the requirement to prevent discussion of competitor pricing, filtering both input and output at inference time without retraining the foundation model.

Why this answer

Amazon Bedrock Guardrails with topic denial is the correct feature because it allows administrators to define a set of prohibited topics (e.g., competitor pricing) and configure the model to refuse to engage in conversations about them. When the model detects a user input or generates a response that falls within a denied topic, Guardrails intercepts the interaction and returns a predefined denial message, effectively blocking the unwanted discussion. This is a content moderation capability specifically designed to enforce policy-based restrictions on model behavior, making it the ideal choice for preventing a chatbot from discussing specific prohibited topics.

Exam trap

The trap here is that candidates often confuse Bedrock Guardrails with Bedrock Agents or Knowledge Bases, assuming that agents or knowledge bases can inherently filter topics, when in fact Guardrails is the dedicated service for content moderation and policy enforcement.

How to eliminate wrong answers

Option A is wrong because Bedrock Knowledge Base is a feature for connecting the model to a company's private data sources (e.g., documents, databases) to enable retrieval-augmented generation (RAG), not for blocking or denying topics. Option C is wrong because Bedrock Agents are used to orchestrate multi-step tasks by invoking APIs and knowledge bases, but they do not natively provide topic-level denial or content filtering; they rely on Guardrails for such safety controls. Option D is wrong because Bedrock Model Evaluation is a tool for assessing model performance (e.g., accuracy, toxicity) on benchmark datasets, not for enforcing runtime content restrictions or blocking specific topics.

665
Multi-Selectmedium

A data science team is evaluating the output quality of a text generation model. They want to use both automated metrics and human judgment. Which THREE approaches should they include in their evaluation strategy? (Select THREE.)

Select 3 answers
A.Conduct human evaluation with a rating rubric
B.Measure inference latency under load
C.Compute perplexity on the generated text
D.Use ROUGE or BLEU scores against reference texts
E.Run the model on a task‑specific benchmark dataset
AnswersA, D, E

Human evaluation with a rating rubric captures subjective qualities automated metrics miss, such as coherence, relevance and fluency, satisfying the stem's requirement for human judgment alongside automated metrics. A structured rubric ensures consistent, reproducible scores across annotators, making it a valid component of the three-approach evaluation strategy.

Why this answer

Option A is correct because human evaluation with a rating rubric directly captures subjective quality dimensions such as fluency, coherence, relevance, and factual consistency that automated metrics cannot fully assess, which is essential when the team explicitly wants human judgment. Option D is correct because ROUGE (recall-oriented n-gram/LCS overlap) and BLEU (precision-oriented n-gram overlap with brevity penalty) are standard automated reference-based metrics for comparing generated text against reference texts, matching the team's goal of using automated metrics. Option E is correct because running the model on a task-specific benchmark dataset provides standardized, objective, task-relevant performance measurement (e.g., accuracy, F1, exact match) that complements both human ratings and text-overlap metrics.

Option B does not belong because inference latency under load measures operational performance and scalability, not output quality of the generated text. Option C does not belong because perplexity is typically computed on a language model's token likelihoods over a corpus and is not a reliable measure of the quality of generated text, especially for open-ended generation.

Exam trap

AWS often tests the distinction between performance metrics (e.g., latency, throughput) and quality metrics (e.g., human evaluation, automated scores), so candidates may mistakenly include inference latency as a quality measure.

666
MCQeasy

Which of the following is a primary benefit of using Bedrock Agents for building generative AI applications?

A.Automatically fine-tune the underlying model on new data
B.Orchestrate multi-step tasks and call external APIs via action groups
C.Optimize prompts automatically without any manual tuning
D.Guarantee that the model's responses are factually correct
AnswerB

Action groups let Bedrock Agents invoke external APIs and orchestrate multi-step workflows, decomposing a request into sequenced actions with API calls between steps. This satisfies the requirement for building applications that automate complex tasks rather than single-turn text generation.

Why this answer

The primary benefit of using Bedrock Agents is their ability to orchestrate multi-step tasks and call external APIs through action groups. This allows the agent to break down complex user requests into a sequence of actions, invoke APIs, and return results, enabling sophisticated automation.

Exam trap

AIF-C01 often tests the capabilities of Bedrock Agents, and candidates may overestimate their abilities, such as assuming they automatically fine-tune models or guarantee correctness, rather than focusing on orchestration and API calls.

How to eliminate wrong answers

Option A is wrong because Bedrock Agents do not automatically fine-tune the underlying model; fine-tuning is a separate process that requires training data and is not a core feature of Agents. Option C is wrong because Bedrock Agents do not optimize prompts automatically; prompt engineering is still a manual process, although Agents can use the model to generate prompts for actions. Option D is wrong because no generative AI model can guarantee factually correct responses; Bedrock Agents can improve accuracy by grounding in data, but they do not guarantee correctness.

667
MCQmedium

A data scientist is evaluating a logistic regression model for a binary classification task. The model's AUC-ROC score is 0.95 on the training set and 0.51 on the test set. What is the MOST likely issue?

A.The model is overfitting the training data
B.The test set is too small
C.The learning rate is too low
D.The model is underfitting the training data
AnswerA

The AUC-ROC gap between training (0.95) and test (0.51) shows the model memorised training data rather than learning generalisable patterns. Near-random test performance confirms overfitting, satisfying the scenario's requirement to identify the most likely issue for this binary classification model.

Why this answer

A large gap between training and test AUC-ROC indicates overfitting — the model memorizes training data but fails to generalize. Cross-validation can help detect and mitigate this. Data leakage or class imbalance could also contribute, but overfitting is the primary symptom.

668
MCQhard

An insurance company uses an Amazon SageMaker model to set premium discounts. An internal audit finds that the model's error rates are substantially higher for one demographic group than another, even though overall accuracy is strong. The data science team must investigate and quantify this disparity before the model can be re-approved. Which approach should they take first?

A.Increase the model's complexity by adding more layers to improve fit on the training data
B.Retrain the model on a larger dataset and compare overall accuracy before and after
C.Deploy the model with Amazon SageMaker Model Monitor to track data drift in production
D.Run a SageMaker Clarify bias analysis that computes subgroup performance metrics such as accuracy difference and disparate impact across the demographic facets
AnswerD

SageMaker Clarify bias analysis computes pre-training and post-training metrics broken out by configured facets, including accuracy difference, recall difference, and disparate impact. That directly quantifies the higher error rate observed for one demographic group and produces the evidence the audit requires. It is the correct first step because measurement must precede any mitigation decision.

Why this answer

Before any remediation, the team must quantify the disparity by computing fairness metrics broken out by demographic facet. SageMaker Clarify bias analysis produces exactly these subgroup metrics, such as accuracy difference and disparate impact, giving the audit the evidence it needs. Retraining, monitoring, and added model capacity all act without first measuring the gap, so they cannot satisfy the re-approval requirement.

Exam trap

The trap here is jumping to a fix such as retraining or monitoring before measuring the disparity, when the audit first requires quantified subgroup fairness metrics.

669
Multi-Selecthard

Which THREE are best practices for building a secure and scalable generative AI application using Amazon Bedrock? (Choose 3)

Select 3 answers
A.Implement guardrails to filter harmful content
B.Deploy models on EC2 instances for better control
C.Store API keys in source code for easy access
D.Use AWS KMS to encrypt data and model artifacts
E.Use foundation models from multiple providers via Bedrock
AnswersA, D, E

Guardrails in Amazon Bedrock apply configurable policies that filter harmful or inappropriate content in prompts and responses, denying unsafe interactions. This directly enforces the security posture the stem requires for a secure generative AI application.

Why this answer

Option A is correct because Amazon Bedrock Guardrails let you define denied topics, content filters, word filters, and PII redaction policies that are applied to both prompts and model responses, which is a core best practice for filtering harmful or non-compliant content in a generative AI application. Option D is correct because AWS KMS provides customer-managed keys and envelope encryption for data at rest, including model artifacts, knowledge bases, and logs, satisfying security and compliance requirements for protecting sensitive data. Option E is correct because Bedrock offers a unified, serverless API to access foundation models from multiple providers such as Anthropic, Meta, Mistral, Cohere, and Amazon, enabling model choice, fallback, and cost/performance optimization without managing infrastructure, which supports scalability.

Option B is not a Bedrock best practice because deploying models on self-managed EC2 instances moves you away from Bedrock's serverless, managed scaling and shifts patching, scaling, and security responsibilities to you. Option C is incorrect because hardcoding API keys in source code exposes credentials to leakage via repositories and logs; you should use AWS Secrets Manager or IAM roles instead.

Exam trap

Often, candidates mistakenly think that 'more control' (like EC2) is always better for security, when in fact managed services like Bedrock reduce attack surface and operational burden, making them the recommended approach for secure and scalable generative AI applications.

670
MCQmedium

A data scientist is building a model to predict housing prices. The dataset contains features such as square footage, number of bedrooms, and location. After training a linear regression model, the RMSE on the test set is significantly higher than on the training set. What is the MOST likely cause?

A.Data leakage from test to training set
B.Overfitting
C.Underfitting
D.Insufficient training data
AnswerB

A large gap between low training error and high test RMSE indicates the model has memorised training noise rather than the underlying relationship. Overfitting is therefore the most likely cause, since the model fits training data well but fails to generalise to unseen housing records.

Why this answer

Overfitting. When RMSE on the test set is significantly higher than on the training set, it indicates the model has learned noise and specific patterns from the training data that do not generalize to unseen data. This is the classic symptom of overfitting in linear regression, where the model may have too many features or excessive complexity relative to the data.

Exam trap

In the AWS AI Practitioner exam, this scenario tests the distinction between overfitting and underfitting. A large gap between test and training RMSE indicates overfitting, not data leakage or insufficient data.

How to eliminate wrong answers

Option A is wrong because data leakage from test to training set would typically cause both training and test errors to be artificially low, not a large gap between them. Option C is wrong because underfitting would result in high RMSE on both training and test sets, not a significant difference. Option D is wrong because insufficient training data can contribute to overfitting but is not the most likely direct cause; the primary issue described is the performance gap, which is the hallmark of overfitting.

671
Multi-Selectmedium

A data scientist is evaluating whether to use fine-tuning or Retrieval-Augmented Generation (RAG) for a legal document analysis application. Which TWO statements correctly describe when to use each approach?

Select 2 answers
A.Use RAG to reduce model hallucinations without any external data
B.Use fine-tuning to teach the model a specific output format or style
C.Use fine-tuning to enable the model to access real-time information
D.Use fine-tuning to inject new factual knowledge into the model
E.Use RAG when the knowledge base changes frequently
AnswersB, E

Fine-tuning adjusts model weights on labelled examples, embedding a consistent output structure or tone into the model itself. This suits the legal application when a fixed response format or style is required, rather than retrieving external document content at inference time.

Why this answer

Option B is correct because fine-tuning adjusts the model's weights on labeled examples, which is the right technique for teaching a consistent output format, tone, or style (e.g., always returning structured legal citations in a fixed schema). Option E is correct because RAG retrieves documents from an external knowledge base at inference time, so when the underlying corpus changes frequently, you simply update the index/vector store rather than retraining the model. Option A is wrong because RAG reduces hallucinations by grounding responses in retrieved external data, not by working 'without any external data.' Option C is wrong because real-time information access is achieved through retrieval (RAG) or tool calls, not fine-tuning, which bakes in static weights at training time.

Option D is wrong because fine-tuning is unreliable for injecting new factual knowledge; RAG is the preferred approach for adding or updating facts.

Exam trap

AWS often tests the misconception that fine-tuning can inject new factual knowledge or enable real-time updates, when in fact RAG is the appropriate technique for dynamic or external knowledge integration.

672
Multi-Selecthard

Which THREE are benefits of using Amazon Bedrock over self-managing foundation models on EC2? (Choose THREE.)

Select 3 answers
A.Built-in integration with AWS services such as AWS CloudWatch and AWS CloudTrail.
B.Lower data transfer costs between cloud regions.
C.Access to a curated set of foundation models from different providers.
D.Managed infrastructure for model hosting and scaling.
E.Greater control over model fine-tuning and customization.
AnswersA, C, D

Amazon Bedrock natively emits metrics to CloudWatch and logs API activity to CloudTrail without custom instrumentation, satisfying the operational visibility requirement that self-managed EC2 inference stacks must build and maintain themselves. This removes undifferentiated monitoring plumbing, letting teams focus on model usage rather than telemetry infrastructure.

Why this answer

Option A is correct because Amazon Bedrock natively integrates with AWS monitoring and auditing services like CloudWatch (for metrics/logs) and CloudTrail (for API activity logging), giving visibility and compliance without custom instrumentation. Option C is correct because Bedrock provides a curated catalog of foundation models from multiple providers (e.g., Anthropic, AI21, Stability AI, Amazon Titan, Meta) through a single API, which self-managing on EC2 would require sourcing and integrating individually. Option D is correct because Bedrock is a fully managed service that handles model hosting, provisioning, and automatic scaling of inference capacity, removing the operational burden of managing EC2 instances, AMIs, autoscaling groups, and GPU capacity.

Option B is not correct because Bedrock does not inherently reduce inter-region data transfer costs; those are governed by standard AWS data transfer pricing and are unrelated to the choice between Bedrock and self-managed EC2. Option E is not correct because self-managing foundation models on EC2 generally offers greater control over fine-tuning and customization, whereas Bedrock provides more constrained, managed customization options.

Exam trap

The trap here is that candidates may confuse 'managed infrastructure' with 'greater control'—Bedrock simplifies operations but reduces customization flexibility, so option E is a common distractor for those who think managed services offer more control than self-managed solutions.

673
MCQeasy

An e-commerce company uses an LLM to generate product descriptions. They observe that occasionally the model outputs factually incorrect information about products. What is the term for this phenomenon?

A.Bias amplification
B.Overfitting
C.Hallucination
D.Concept drift
AnswerC

Hallucination describes an LLM generating fluent but factually incorrect output, such as false product details. The model produces plausible text without grounding in verified data, which is precisely the phenomenon observed in the generated descriptions.

Why this answer

Hallucination is the term for an LLM generating fluent, confident output that is factually incorrect or unsupported by its training data — exactly what happens when the model invents product specifications, prices, or features. It occurs because the model predicts statistically plausible tokens rather than retrieving verified facts. This is a well-known limitation of generative AI and the reason techniques like RAG and grounding are used to reduce it.

Exam trap

The trap is conflating hallucination with bias or drift — candidates who see 'incorrect information' may pick bias amplification, but the AIF-C01 exam distinguishes factual fabrication (hallucination) from fairness issues (bias) and temporal degradation (drift).

How to eliminate wrong answers

Option A is wrong because bias amplification refers to a model disproportionately reinforcing existing societal or data biases in its outputs, not fabricating facts. Option B is wrong because overfitting describes a model memorizing training data too closely and performing poorly on new data — it does not describe confident false statements at inference time. Option D is wrong because concept drift is the degradation of model accuracy over time as real-world data distributions change, which is a monitoring concern, not a per-output factual error.

674
MCQeasy

Which AWS service provides human review workflows to handle low-confidence predictions or high-risk decisions in an AI system?

A.AWS Lambda
B.Amazon Augmented AI (A2I)
C.Amazon Bedrock
D.Amazon SageMaker Ground Truth
AnswerB

Amazon Augmented AI provides built-in human review workflows, routing low-confidence inferences to human reviewers via configurable task types. It directly satisfies the stem's requirement to handle low-confidence predictions or high-risk decisions, unlike Rekognition or SageMaker Ground Truth alone.

Why this answer

Amazon Augmented AI (A2I) provides human review workflows that integrate with ML predictions, allowing organizations to route low-confidence predictions or high-risk decisions to human reviewers. It supports both Amazon-built integrations (e.g., Rekognition, Textract) and custom ML models, and it manages the review UI, workforce, and result consolidation. This is exactly the service designed for human-in-the-loop review of AI decisions.

Exam trap

The trap is confusing Amazon Augmented AI (A2I) with SageMaker Ground Truth — both involve humans, but A2I is for reviewing live predictions (human-in-the-loop), while Ground Truth is for labeling training data.

How to eliminate wrong answers

Option A is wrong because AWS Lambda is a serverless compute service for running code in response to events — it can orchestrate review logic but does not provide human review workflows out of the box. Option C is wrong because Amazon Bedrock is a managed service for accessing foundation models via APIs; it does not provide human review workflows. Option D is wrong because SageMaker Ground Truth is a data labeling service for building training datasets, not a runtime human-review workflow for live predictions.

675
MCQhard

A financial services company uses Bedrock Agents to automate a multi-step loan approval process. The agent needs to call an external credit scoring API and a compliance database, then combine results. The agent currently fails when the API returns a 503 error. How should the practitioner address this?

A.Adjust the Bedrock Agent's granularity setting to 'high'
B.Reduce the agent's trace truncation limit to shorten the context
C.Implement a custom Lambda function for the action group to handle errors and retries
D.Use Bedrock Guardrails with a topic denial rule for error messages
AnswerC

A custom Lambda backing the action group lets the practitioner catch the 503, apply retry with backoff, and return a structured response the agent can act on. Bedrock Agents' built-in invocation offers no such error handling, so this satisfies the resilience constraint.

Why this answer

Lambda functions integrated with action groups can implement retry logic, error handling, and custom processing. Granularity tuning or trace truncation does not handle errors. Guardrails are for content filtering, not API error handling.

Page 8

Page 9 of 12

Page 10