Courseiva

AWS Certified AI Practitioner AIF-C01 (AIF-C01) — Questions 826–862

862 questions total · 12pages · All types, answers revealed

Page 11

Page 12 of 12

826
MCQhard

A media company stores thousands of hours of unlabeled video footage and wants to build a searchable index that lets editors retrieve clips by describing their content in natural language. The team has no annotated dataset and no budget to label one. Which machine learning approach is the MOST appropriate starting point?

A.Use a pretrained multimodal embedding model to encode video frames and text into a shared vector space, then retrieve clips by similarity.
B.Train a supervised video classifier from scratch using the footage, treating each video file as its own class.
C.Apply k-means clustering to raw pixel values of every frame to produce searchable cluster identifiers.
D.Build a reinforcement learning agent that learns to select the clip most likely to satisfy an editor through trial and error.
AnswerA

A pretrained multimodal model maps both video content and text descriptions into a common embedding space, so a natural-language query can be compared directly against indexed clips without any custom labels. This leverages existing pretrained knowledge, requires no annotation budget, and supports flexible search. It is the standard foundation for semantic video retrieval systems.

Why this answer

With no labels and a need to search footage using free-form language, a pretrained multimodal embedding model is the most practical foundation. It places video and text in a shared vector space so that descriptive queries match relevant clips by similarity, enabling semantic retrieval without any custom annotation effort.

Exam trap

The trap here is assuming that a large unlabeled video collection must be turned into a supervised classification dataset, when pretrained multimodal embeddings can support retrieval directly.

827
Multi-Selecthard

Which TWO of the following are best practices for preparing training data for a machine learning model?

Select 2 answers
A.Handle missing values by imputing or removing them.
B.Split the data into training, validation, and test sets.
C.Remove all outliers to improve model robustness.
D.Use the entire dataset for training to maximize data usage.
E.Avoid shuffling the data to preserve original order.
AnswersA, B

Missing values break many algorithms and bias estimates, so imputing sensible substitutes or removing affected rows prevents errors and skewed learning. This is a standard preparation step ensuring the training set is complete and representative before model fitting.

Why this answer

Option A is correct because missing values can bias or break many ML algorithms, so best practice is to address them explicitly—either by imputation (e.g., mean/median/mode or model-based) or by removing affected rows/columns when appropriate. Option B is correct because splitting data into training, validation, and test sets enables unbiased model fitting, hyperparameter tuning, and final performance estimation, preventing data leakage and overfitting. Option C is not a best practice because outliers may be legitimate signal; blindly removing them can distort the distribution and reduce robustness, so they should be investigated and handled contextually.

Option D is not a best practice because training on the entire dataset leaves no held-out data for validation or testing, making it impossible to reliably estimate generalization. Option E is not a best practice because shuffling is typically needed to remove ordering bias and ensure representative mini-batches, especially when data is sorted by class or time.

Exam trap

The AIF-C01 exam often tests the misconception that removing all outliers is always beneficial, when in fact domain knowledge is required to distinguish between noise and legitimate extreme values that may be critical for model accuracy.

828
Multi-Selectmedium

A data science team wants to document and share their model's intended use, performance, and limitations with stakeholders. They also need to track the model's version and deployment history. Which TWO AWS services or features should they use?

Select 2 answers
A.Amazon SageMaker Pipelines
B.Amazon SageMaker Clarify
C.Amazon SageMaker Studio
D.Amazon SageMaker Model Cards
E.Amazon SageMaker Model Registry
AnswersD, E

Amazon SageMaker Model Cards capture intended use, performance metrics, and limitations in a structured document, directly satisfying the stakeholder documentation requirement. Model Registry separately handles versioning and deployment history, so Model Cards alone address only the documentation half of the stem's two-part need.

Why this answer

SageMaker Model Cards provide standardized documentation for transparency. SageMaker Model Registry tracks model versions, deployment stages, and metadata. SageMaker Pipelines is for ML workflows, not documentation or version tracking.

SageMaker Studio is an IDE. SageMaker Clarify is for bias and explainability.

829
Multi-Selectmedium

A team is using Amazon SageMaker to build a text classification model. They have raw text data in CSV files stored in Amazon S3. Before training, they need to perform feature engineering. Which THREE actions should they take? (Select THREE.)

Select 3 answers
A.Normalize numerical features to have zero mean and unit variance
B.Tokenize the text into words or subwords
C.Remove punctuation and stop words from the text
D.Convert the text into numerical features using techniques like TF-IDF or word embeddings
E.Encode categorical variables using one-hot encoding
AnswersB, C, D

Tokenisation splits raw text into words or subwords, producing discrete units that downstream vectorisers such as TF-IDF or embedding layers require. Without this step, the CSV text strings cannot be mapped to any numerical representation, so the feature engineering pipeline for the classification model cannot proceed.

Why this answer

For text classification, the raw text must be transformed into a machine-readable representation, so option B is correct because tokenizing the text into words or subwords is the essential first step that splits raw strings into discrete units a model can process. Option C is correct because removing punctuation and stop words reduces noise and dimensionality, helping the model focus on meaningful tokens rather than common, low-information words. Option D is correct because converting the tokenized text into numerical features via TF-IDF or word embeddings is required, since SageMaker training algorithms and models operate on numeric vectors, not raw strings.

Option A is not appropriate here because normalizing numerical features to zero mean and unit variance applies to numeric columns, not raw text data. Option E is also not appropriate because one-hot encoding is used for categorical variables, which is not the primary transformation needed for free-form text in this scenario.

Exam trap

In the AWS AI Practitioner exam, candidates often confuse feature engineering steps specific to text data with general preprocessing for numerical data, leading them to incorrectly select normalization or one-hot encoding.

830
MCQeasy

Refer to the exhibit. A developer runs this command but gets an error: 'An error occurred (AccessDeniedException) when calling the ListFoundationModels operation'. What is the most likely cause?

A.The IAM role does not have bedrock:ListFoundationModels permission
B.The AWS CLI version is outdated
C.The foundation model is not available in us-west-2
D.The region us-west-2 does not support Bedrock
AnswerA

ListFoundationModels is a Bedrock control-plane action, so the caller's identity must hold bedrock:ListFoundationModels in its attached IAM policy. AccessDeniedException is returned by IAM authorisation evaluation, not by model availability or Region configuration, making a missing or explicitly denied permission the direct cause.

Why this answer

The error 'AccessDeniedException' when calling ListFoundationModels indicates that the IAM role or user executing the AWS CLI command lacks the required permission to list foundation models in Amazon Bedrock. The specific permission needed is bedrock:ListFoundationModels, which must be attached to the IAM identity via a policy. Without this permission, the API call is denied regardless of other factors like region or CLI version.

Exam trap

AWS often tests the distinction between service availability errors (e.g., region not supported) and IAM permission errors, where candidates mistakenly attribute an AccessDeniedException to regional or model availability issues rather than missing IAM permissions.

How to eliminate wrong answers

Option B is wrong because an outdated AWS CLI version would typically produce a different error (e.g., 'InvalidClientTokenId' or 'UnrecognizedClientException'), not an AccessDeniedException, and the ListFoundationModels API is available in recent CLI versions. Option C is wrong because the error is an access denial, not a model availability issue; if a model were unavailable, the error would be something like 'ValidationException' or 'ResourceNotFoundException' when trying to use that specific model. Option D is wrong because us-west-2 (Oregon) fully supports Amazon Bedrock and its APIs; the error is explicitly an IAM permissions issue, not a regional unsupported service error.

831
MCQeasy

What is the primary advantage of the transformer architecture over previous RNN-based architectures for natural language processing tasks?

A.It processes tokens sequentially to maintain the order of the sequence
B.It requires less training data because it uses convolutional layers
C.It uses a gated recurrent unit to selectively forget information
D.It relies on a self-attention mechanism that allows parallel processing of all tokens
AnswerD

Self-attention computes relationships between every token pair simultaneously, so the model processes all positions in parallel rather than sequentially as RNNs must. This removes the sequential dependency that prevented parallelisation, directly satisfying the stem's demand for the primary advantage: dramatically faster training on long sequences.

Why this answer

The transformer architecture's primary advantage is its self-attention mechanism, which allows the model to process all tokens in the input sequence in parallel rather than sequentially. This parallelization dramatically reduces training time and enables the model to capture long-range dependencies more effectively than RNNs, which must process tokens one by one and suffer from vanishing gradients over long sequences.

Exam trap

In the AWS AI Practitioner exam, a common trap is to confuse the sequential processing of RNNs with a desirable property for maintaining order, but transformers actually use positional encodings to preserve sequence order while still enabling parallel processing.

How to eliminate wrong answers

Option A is wrong because processing tokens sequentially is a characteristic of RNNs, not a primary advantage of transformers; transformers process tokens in parallel. Option B is wrong because transformers do not use convolutional layers; they rely on self-attention and feed-forward networks, and they typically require more training data, not less, to achieve their performance. Option C is wrong because gated recurrent units (GRUs) are a type of RNN architecture, not a feature of transformers; transformers use self-attention and do not have gating mechanisms for forgetting information.

832
MCQhard

A company is using Amazon Bedrock to build a sentiment analysis application for customer reviews. They need to evaluate the model's performance against a labeled test dataset. They want to use a metric that compares the model's predicted sentiment (positive, negative, neutral) to the ground truth labels. Which metric is MOST appropriate?

A.ROUGE-1
B.BERTScore
C.BLEU
D.Accuracy
AnswerD

Accuracy directly measures the proportion of predicted sentiment labels matching ground truth across the three classes, giving a single comparable figure for the labelled test set. It satisfies the requirement to compare predictions against ground truth without weighting any class.

Why this answer

Accuracy is the most appropriate metric for this multi-class classification task because it directly measures the proportion of correctly predicted sentiments (positive, negative, neutral) out of total predictions, given a labeled test dataset. It is simple, interpretable, and well-suited for evaluating categorical outputs where class distribution is balanced or the cost of misclassification is uniform.

Exam trap

AWS often tests the misconception that advanced NLP metrics like BLEU or ROUGE are suitable for classification tasks, when in fact they are designed for generation tasks, leading candidates to overcomplicate the answer.

How to eliminate wrong answers

Option A is wrong because ROUGE-1 is an n-gram overlap metric designed for evaluating text summarization and generation tasks, not for classification accuracy of discrete sentiment labels. Option B is wrong because BERTScore computes semantic similarity between generated and reference texts using contextual embeddings, which is overkill and inappropriate for comparing categorical labels like sentiment classes. Option C is wrong because BLEU measures precision of n-gram matches in machine translation output, not the correctness of classification predictions.

833
MCQmedium

A company uses Amazon Bedrock Guardrails to filter harmful content. They want to ensure that the model does not generate responses containing specific keywords related to their internal project names. Which Guardrails component should they configure?

A.Harmful content filters
B.Topic restrictions
C.Word filters
D.Grounding checks
AnswerC

Word filters block or mask prompts and responses containing configured custom terms, so adding internal project names stops the model emitting them. This is distinct from managed word filters, which target profanity, and from denied topics, which cover broader subject areas.

Why this answer

Word filters in Amazon Bedrock Guardrails are specifically designed to block or mask responses containing custom keywords or phrases. By configuring word filters, the company can ensure that internal project names are not generated in model responses. This is the precise component for keyword-based filtering.

Exam trap

The trap is confusing word filters with topic restrictions; candidates may think topic restrictions handle specific keywords, but they are for broader thematic blocking, while word filters are for exact terms.

How to eliminate wrong answers

Option A is wrong because harmful content filters target categories like hate, violence, and sexual content, not specific custom keywords. Option B is wrong because topic restrictions block entire topics based on natural language descriptions, not exact keyword matches. Option D is wrong because grounding checks verify that responses are grounded in source documents, not filter specific words.

834
Multi-Selecteasy

A company wants to use AWS services to process natural language text. Which TWO AWS services provide natural language processing (NLP) capabilities? (Select TWO.)

Select 2 answers
A.Amazon Translate
B.Amazon Rekognition
C.Amazon Comprehend
D.Amazon Polly
E.Amazon Lex
AnswersC, E

Amazon Comprehend is a fully managed NLP service providing entity recognition, sentiment analysis, key phrase extraction and language detection on text. It directly satisfies the stem's requirement for an AWS service delivering natural language processing capabilities.

Why this answer

Amazon Comprehend (C) is correct because it is AWS's fully managed NLP service that uses machine learning to extract insights from text, including sentiment analysis, entity recognition, key phrase extraction, language detection, and topic modeling. Amazon Lex (E) is correct because it provides NLP capabilities for building conversational interfaces (chatbots and voice bots) using automatic speech recognition and natural language understanding to interpret user intent and slot values. Amazon Translate (A) is not marked correct because it performs language translation rather than general NLP analysis, even though it is a text-processing service.

Amazon Rekognition (B) is not marked correct because it is a computer vision service for image and video analysis, not natural language text. Amazon Polly (D) is not marked correct because it is a text-to-speech service that converts written text into lifelike speech, which is speech synthesis rather than NLP understanding.

Exam trap

The trap here is that candidates often confuse text-to-speech (Polly) or translation (Translate) with NLP, but these services do not perform language understanding or analysis—they only convert or generate speech/translation without extracting meaning.

835
MCQhard

A team trained a binary classifier to detect fraudulent transactions. The dataset is highly imbalanced (1% fraud). The model achieves 99% accuracy but only catches 5% of actual fraud cases. Which metric should the team primarily optimize?

A.Recall
B.F1 score
C.Accuracy
D.Precision
AnswerA

Recall measures the proportion of actual fraud cases detected. With only 5% caught, recall exposes the model's failure despite 99% accuracy, which is misleading under 1% class imbalance. Optimising recall directly targets catching more fraud.

Why this answer

In fraud detection with 1% fraud rate, 99% accuracy is misleading because the model can achieve it by simply predicting 'not fraud' for all transactions. The model catches only 5% of actual fraud cases, meaning recall (true positive rate) is critically low at 5%. Optimizing recall directly increases the proportion of actual fraud cases correctly identified, which is the primary business requirement in fraud detection where missing a fraud (false negative) is far more costly than a false alarm.

Exam trap

In AWS AI Practitioner exams, note that accuracy can be misleading in imbalanced datasets, and recall is often the key metric for fraud detection because missing a fraud (false negative) is far more costly than a false alarm.

How to eliminate wrong answers

Option B (F1 score) is wrong because F1 score is the harmonic mean of precision and recall; while it balances both, the question specifically states the model catches only 5% of fraud cases, so the immediate priority is to improve recall first before balancing with precision. Option C (Accuracy) is wrong because accuracy is misleading in imbalanced datasets — a model that predicts 'not fraud' for every transaction achieves 99% accuracy but fails at the actual task of detecting fraud. Option D (Precision) is wrong because precision measures how many of the predicted fraud cases are actually fraud; optimizing precision would further reduce the number of fraud predictions, making the recall even worse, which is the opposite of what is needed.

836
Multi-Selecthard

Which THREE considerations are essential for ensuring responsible AI in a model that predicts employee performance? (Choose 3)

Select 3 answers
A.Minimize the number of features to reduce cost
B.Publish the model's predictions publicly for transparency
C.Incorporate human review before final decisions
D.Ensure employee data privacy and consent
E.Test for bias across demographic groups
AnswersC, D, E

Human review before final decisions satisfies the accountability constraint: a performance prediction can carry employment consequences, so a person must evaluate context the model cannot capture and challenge biased or erroneous outputs. This keeps a human answerable for outcomes, rather than deferring to automated scoring, as responsible AI practise requires.

Why this answer

Option C is correct because responsible AI in an HR context requires human-in-the-loop oversight: performance predictions should inform, not automate, decisions about employees, so a human reviewer can catch errors and provide context before any action is taken. Option D is correct because employee performance data is personal and often sensitive, so the model must comply with privacy and consent requirements (e.g., GDPR lawful basis, purpose limitation, and data minimization) to be ethically and legally sound. Option E is correct because bias testing across demographic groups (e.g., comparing error rates, selection rates, and disparate impact metrics by gender, age, or ethnicity) is essential to detect and mitigate discriminatory outcomes in performance predictions.

Option A does not belong because minimizing features to cut cost is an efficiency concern, not a responsible-AI safeguard, and can even harm fairness if it removes relevant variables. Option B does not belong because publishing individual employees' predicted performance publicly would violate privacy and confidentiality rather than constitute meaningful transparency, which is better served by model documentation and explainability to authorized stakeholders.

Exam trap

The AIF-C01 exam often tests the misconception that transparency means public disclosure of all model outputs, whereas in responsible AI, transparency refers to explainability and auditability of the model's logic, not exposing sensitive predictions.

837
MCQmedium

A developer is using the Amazon Bedrock Converse API to build a chat application. They want the model to maintain context across multiple turns. Which parameter should they set to ensure the conversation history is included?

A.inferenceConfig
B.messages
C.additionalModelRequestFields
D.system
AnswerB

The messages parameter carries the ordered array of user and assistant turns sent to the model. Including prior turns in messages supplies the conversation history, enabling the model to maintain context across multiple turns as the stem requires.

Why this answer

The Converse API's `messages` parameter accepts an array of message objects, each with a `role` (user or assistant) and `content`. By including previous turns in this array, the model receives the full conversation history as context, enabling coherent multi-turn dialogue. Without it, each API call is stateless and the model has no memory of prior exchanges.

Exam trap

AWS often tests the distinction between stateless API calls and stateful conversation management, where candidates mistakenly think `system` or `inferenceConfig` handles history, but only `messages` carries the turn-by-turn dialogue.

How to eliminate wrong answers

Option A is wrong because `inferenceConfig` controls generation parameters like temperature and max tokens, not conversation history. Option C is wrong because `additionalModelRequestFields` passes extra inference parameters specific to certain models (e.g., Cohere's return_likelihoods), not dialogue context. Option D is wrong because `system` sets a system prompt or persona for the model, but does not include the user-assistant message history.

838
Multi-Selecthard

A research team is using Amazon SageMaker to fine-tune a large language model. They want to optimize training cost and time without sacrificing model quality. Which THREE strategies should they implement? (Choose 3)

Select 3 answers
A.Use a larger instance type with more GPUs.
B.Apply parameter-efficient fine-tuning (PEFT) techniques like LoRA.
C.Increase the batch size to the maximum that fits in GPU memory.
D.Use managed spot training with checkpointing.
E.Enable mixed precision training (FP16).
AnswersB, D, E

Parameter-efficient fine-tuning such as LoRA freezes the base model weights and trains only small low-rank adapter matrices, cutting trainable parameters and GPU memory dramatically. This directly satisfies the stem's constraint of reducing training cost and time while preserving model quality, since the pretrained knowledge remains intact.

Why this answer

Option B is correct because parameter-efficient fine-tuning techniques such as LoRA freeze most of the pretrained model weights and train only small low-rank adapter matrices, which drastically reduces GPU memory consumption, training time, and compute cost while preserving model quality. Option D is correct because SageMaker managed spot training can cut training costs by up to 90% versus on-demand instances, and pairing it with checkpointing to Amazon S3 lets training resume from the last checkpoint after a spot interruption rather than restarting from scratch. Option E is correct because mixed precision training with FP16 reduces memory footprint and leverages GPU tensor cores for faster matrix math, speeding up training and allowing larger effective batch sizes with minimal impact on model accuracy.

Option A is not correct because simply moving to a larger multi-GPU instance raises cost and does not by itself optimize time or cost efficiency, and may be unnecessary if PEFT and FP16 already fit the workload. Option C is not correct because maximizing batch size to the GPU memory limit is not universally beneficial; overly large batches can hurt convergence and model quality and may require extensive hyperparameter retuning, so it is not a reliable cost-and-time optimization strategy.

Exam trap

The AIF-C01 exam often tests the misconception that simply scaling up hardware (larger instances) or maximizing batch size is the best optimization strategy, when in fact algorithmic efficiency (PEFT, mixed precision) and cost-saving infrastructure (spot instances) are the correct approaches for balancing cost, time, and quality.

839
Multi-Selecthard

A company is deploying a chatbot using Amazon Bedrock and wants to ensure that the model does not generate offensive or inappropriate content. Which THREE measures can they apply?

Select 3 answers
A.Use a system prompt to define ethical guidelines and constraints
B.Implement a human-in-the-loop review process for flagged responses
C.Increase the temperature to make outputs more creative and less likely to repeat offensive phrases
D.Enable content filtering via Bedrock's guardrails or built-in filters
E.Fine-tune the model on a dataset of safe conversations only
AnswersA, B, D

A system prompt sets persistent instructions that steer the model's tone and refusals across every turn, so offensive outputs are discouraged at generation time. It complements, rather than replaces, output filtering, satisfying the requirement to constrain inappropriate content.

Why this answer

Option A is correct because a system prompt sets the model's behavioral boundaries and can explicitly instruct it to refuse or avoid offensive, harmful, or inappropriate content, making it a first-line control for responsible generation. Option B is correct because a human-in-the-loop review process catches flagged or borderline responses that automated filters miss, providing an additional safety layer before content reaches users. Option D is correct because Amazon Bedrock Guardrails (and built-in content filters) apply configurable policies such as denied topics, content filters for hate/violence/sexual/insults, and PII redaction to block or mask inappropriate model outputs at inference time.

Option C is not appropriate because raising temperature increases randomness and creativity, which generally makes offensive or unpredictable output more likely, not less. Option E is not a reliable standalone measure because fine-tuning on safe conversations does not guarantee the model will never generate offensive content and does not provide runtime enforcement, unlike guardrails or review processes.

Exam trap

AWS often tests the misconception that increasing temperature or fine-tuning on safe data alone can prevent harmful outputs, when in reality these methods do not provide runtime content filtering and can even increase risk or be impractical for rapid deployment.

840
Multi-Selectmedium

A company uses Amazon Macie to discover sensitive data in an S3 bucket containing training datasets. The bucket policy currently prohibits access from external accounts. Which TWO steps are necessary to allow a cross-account SageMaker training job to access this bucket while maintaining security?

Select 2 answers
A.Configure Macie to automatically grant access to the SageMaker execution role
B.Create a VPC endpoint for S3 and associate it with the SageMaker VPC
C.Add a bucket policy that grants the SageMaker execution role from the other account s3:GetObject and s3:ListBucket permissions
D.Attach an IAM policy to the SageMaker execution role that allows s3:GetObject and s3:ListBucket on the source bucket
E.Remove the bucket policy that prohibits external access
AnswersC, D

The bucket policy is the resource-based control that explicitly overrides the current prohibition on external accounts. Granting the cross-account SageMaker execution role s3:GetObject and s3:ListBucket authorises that principal directly, satisfying the cross-account access requirement while keeping the bucket closed to everyone else.

Why this answer

Option C is correct because the S3 bucket policy is the resource-based policy that must explicitly grant the cross-account SageMaker execution role access; since the bucket currently prohibits external accounts, you must add a statement allowing that role principal s3:GetObject and s3:ListBucket on the bucket and its objects. Option D is correct because the SageMaker execution role in the other account also needs an identity-based IAM policy permitting s3:GetObject and s3:ListBucket on the source bucket, since cross-account access requires both the identity policy and the resource policy to allow the action. Option A is wrong because Macie is a data-discovery and classification service and cannot grant access to a SageMaker execution role.

Option B is wrong because an S3 VPC endpoint only affects private network routing for traffic within a VPC and does not by itself authorize cross-account access. Option E is wrong because removing the bucket policy's external-access prohibition would broadly open the bucket rather than securely scoping access to the SageMaker execution role.

Exam trap

AIF-C01 often tests the misconception that a single policy (either identity-based or resource-based) suffices for cross-account access, when in fact both sides must explicitly grant permission.

841
Multi-Selectmedium

A developer is building an agent using Amazon Bedrock Agents to handle customer support inquiries. The agent needs to look up order status from a database and escalate complex issues to a human. Which THREE components are essential for this agent?

Select 3 answers
A.A Lambda function that connects to the database and returns results
B.A fine-tuned foundation model specific to the support domain
C.A custom model trained on historical support tickets
D.A knowledge base containing troubleshooting documentation
E.An action group with an OpenAPI schema for the database query
AnswersA, D, E

A Lambda function provides the executable logic that queries the order database and returns structured results to the agent. Without it, the action group has no backend to fulfil the order-status lookup the scenario requires.

Why this answer

Option A is correct because Amazon Bedrock Agents invoke AWS Lambda functions as action executors to run custom business logic, such as querying an order-status database and returning structured results to the agent. Option E is correct because an action group defines the agent's callable APIs, and its OpenAPI schema tells the agent the available operations, parameters, and request/response format needed to query the database. Option D is correct because a knowledge base (backed by Amazon Bedrock Knowledge Bases and typically Amazon OpenSearch Serverless or another vector store) lets the agent retrieve troubleshooting documentation via RAG so it can answer support questions and decide when to escalate.

Option B is not needed because Bedrock Agents work with existing foundation models and do not require a domain-specific fine-tuned model. Option C is not needed because training a custom model on historical tickets is not a required component; the agent uses foundation models plus action groups and knowledge bases.

Exam trap

AIF-C01 often tests the misconception that fine-tuning or custom models are required for Bedrock Agents, but the essential components are action groups, Lambda functions, and knowledge bases.

842
MCQeasy

Refer to the exhibit. A developer wants to choose a model that can generate text (not just embeddings) and has the lowest cost. Based on the exhibit, which model should they select?

A.Titan Embed Text
B.Titan Text Express
C.Titan Text Lite
D.Need more information
AnswerC

Titan Text Lite is a text-generation model, satisfying the requirement to generate text rather than embeddings, and carries the lowest cost among the exhibit's generative options. Embeddings models such as Titan Embeddings cannot produce text, so they fail the primary constraint regardless of price.

Why this answer

Titan Text Lite is the correct choice because it is designed for text generation tasks and is explicitly positioned as the lowest-cost option among Amazon's Titan text generation models. Unlike Titan Embed Text, which only produces embeddings and cannot generate text, Titan Text Lite offers a cost-efficient solution for generating text while meeting the requirement of lowest cost.

Exam trap

The trap here is that candidates may confuse Titan Embed Text as a text generation model due to its 'Text' name, or assume that 'Express' implies lower cost, when in fact the naming convention indicates performance tier rather than cost efficiency.

How to eliminate wrong answers

Option A is wrong because Titan Embed Text is an embeddings model that converts text into numerical vectors and cannot generate text, failing the core requirement. Option B is wrong because Titan Text Express, while capable of text generation, is a higher-cost model compared to Titan Text Lite, making it not the lowest-cost option. Option D is wrong because the exhibit provides sufficient information to determine that Titan Text Lite is the correct model based on the stated requirements of text generation capability and lowest cost.

843
MCQeasy

Which AWS service provides access to a wide variety of foundation models from different providers through a single API, without managing underlying infrastructure?

A.Amazon SageMaker
B.AWS Lambda
C.Amazon Rekognition
D.Amazon Bedrock
AnswerD

Amazon Bedrock exposes foundation models from multiple providers — including Anthropic, Meta and Amazon — through one unified API, removing the need to provision or manage servers. This directly satisfies the stem's requirement for single-API, multi-provider model access without underlying infrastructure management.

Why this answer

Amazon Bedrock is a fully managed service that provides access to a wide variety of foundation models (FMs) from leading AI providers such as Anthropic, Meta, Stability AI, and Amazon itself through a single API. It eliminates the need to manage underlying infrastructure, such as GPU clusters or model hosting, allowing developers to integrate generative AI capabilities directly into applications.

Exam trap

The trap here is that candidates may confuse Amazon Bedrock with Amazon SageMaker, assuming SageMaker's JumpStart or built-in algorithms provide similar multi-model API access, but SageMaker requires explicit model deployment and infrastructure management, whereas Bedrock is purpose-built for serverless foundation model access.

How to eliminate wrong answers

Option A is wrong because Amazon SageMaker is a machine learning platform focused on building, training, and deploying custom models, not a service that provides pre-built foundation models from multiple providers via a single API; it requires managing infrastructure for model hosting. Option B is wrong because AWS Lambda is a serverless compute service for running code in response to events, not a service for accessing foundation models or managing model inference APIs. Option C is wrong because Amazon Rekognition is a computer vision service for image and video analysis (e.g., facial recognition, object detection), not a generative AI service that provides access to large language models or foundation models.

844
MCQhard

A healthcare startup uses Amazon SageMaker to train a model predicting patient readmission. They need to ensure the model's predictions do not discriminate based on protected attributes like age or race. Which SageMaker feature allows them to monitor and mitigate bias during training?

A.SageMaker Model Monitor
B.SageMaker Autopilot
C.SageMaker Debugger
D.SageMaker Clarify
AnswerD

SageMaker Clarify computes bias metrics such as disparate impact and demographic parity on training data and model outputs, detecting imbalance across protected attributes like age and race. It also provides SHAP-based feature attribution, satisfying the requirement to monitor and mitigate bias during training.

Why this answer

SageMaker Clarify is the correct choice because it is specifically designed to detect and mitigate bias in machine learning models. It provides built-in capabilities to analyze training data and model predictions for bias against protected attributes such as age or race, and can generate bias reports and suggest mitigation strategies during training.

Exam trap

The trap here is that candidates often confuse SageMaker Clarify with SageMaker Model Monitor, assuming that monitoring for data drift also covers bias detection, but Clarify is the dedicated service for bias detection and mitigation.

How to eliminate wrong answers

Option A is wrong because SageMaker Model Monitor is used to detect data drift and model quality degradation in production, not to analyze or mitigate bias during training. Option B is wrong because SageMaker Autopilot automates the process of building, training, and tuning models, but it does not include built-in bias detection or mitigation features. Option C is wrong because SageMaker Debugger is designed to monitor training jobs for issues like vanishing gradients or overfitting by capturing tensors and metrics, not for bias detection or mitigation.

845
MCQeasy

A company wants to use a pre-trained foundation model for sentiment analysis without any customization. Which Amazon Machine Learning service provides access to foundation models via API?

A.Amazon Bedrock
B.Amazon Textract
C.Amazon Comprehend
D.Amazon Rekognition
AnswerA

Amazon Bedrock supplies serverless API access to pre-trained foundation models from providers such as Anthropic and Amazon, so no model training or hosting is needed. This directly satisfies the stem's constraint of using a foundation model for sentiment analysis without customisation, unlike services requiring you to build or fine-tune your own model.

Why this answer

Amazon Bedrock is a fully managed service that provides access to a wide range of pre-trained foundation models (FMs) from leading AI providers like Anthropic, Meta, and Amazon via a unified API. For sentiment analysis, you can invoke an FM such as Anthropic Claude or Meta Llama directly through the Bedrock API without any customization, making it the correct choice for this use case.

Exam trap

AWS often tests the distinction between a service that provides access to foundation models (Bedrock) and a service that offers a pre-built, non-customizable ML capability (Comprehend), leading candidates to mistakenly choose Comprehend because it also handles sentiment analysis.

How to eliminate wrong answers

Option B (Amazon Textract) is wrong because it is a document analysis service designed to extract text, handwriting, and data from scanned documents, not a service that provides access to foundation models via API. Option C (Amazon Comprehend) is wrong because it is a natural language processing (NLP) service that offers pre-built sentiment analysis, but it does not provide access to foundation models; it uses its own proprietary models and cannot be used to invoke third-party FMs. Option D (Amazon Rekognition) is wrong because it is an image and video analysis service for tasks like object detection and facial recognition, not a service for accessing foundation models or performing text-based sentiment analysis.

846
Multi-Selecteasy

A company is using Amazon Bedrock to generate content for a marketing application. The company wants to ensure that the model does not generate content that violates the company's brand guidelines, which prohibit certain keywords and tones. Which TWO features should the company use to enforce these guidelines? (Choose two.)

Select 2 answers
A.Enable Amazon CloudWatch Logs to capture model output and manually review.
B.Create a prompt template that instructs the model to adhere to brand guidelines and avoid prohibited keywords.
C.Configure Amazon Bedrock Guardrails with custom deny topics and content filters.
D.Use AWS IAM policies to restrict the model's output to only approved words.
E.Encrypt the model responses using AWS KMS to prevent unauthorized viewing.
AnswersB, C

A prompt template shapes model output through instruction, embedding brand rules and prohibited keywords directly in the request. It satisfies the guideline requirement by steering generation at inference time, though it relies on the model's compliance rather than hard enforcement.

Why this answer

Prompt engineering allows the company to embed brand guidelines directly into the instruction given to the model, effectively steering the output away from prohibited keywords and tones. Option C is correct because Amazon Bedrock Guardrails provides a managed, policy-based mechanism to define custom deny topics and content filters that can block or mask unwanted content at inference time, enforcing brand guidelines without manual intervention.

Exam trap

The trap here is that candidates often confuse IAM policies with content moderation, mistakenly believing that IAM can restrict model output vocabulary, when in fact IAM only governs API-level permissions and has no awareness of the semantic content of model responses.

847
MCQmedium

A company uses Amazon Titan Text Express for a real-time chat application. Users report that responses are too slow. The application uses the InvokeModel API with default settings. Which change is MOST likely to reduce latency?

A.Use the Converse API instead of InvokeModel
B.Reduce the temperature parameter to 0
C.Switch from Titan Text Express to Titan Text Lite
D.Increase the maxTokens parameter to allow longer responses
AnswerC

Titan Text Lite is the smaller, faster variant in the Titan Text family, trading some capability for lower per-token latency. For a real-time chat workload where speed is the reported problem, this substitution directly reduces response time.

Why this answer

Titan Text Lite is a smaller, faster model optimized for low-latency use cases like real-time chat, whereas Titan Text Express prioritizes higher quality and throughput at the cost of speed. Switching to Titan Text Lite directly reduces inference time because it has fewer parameters and lower computational overhead per request, making it the most effective change to reduce latency with default InvokeModel API settings.

Exam trap

The AWS exam often tests the misconception that all models in a family have identical performance characteristics, but Titan Text Lite and Express are deliberately differentiated by speed versus quality, and candidates may overlook that model selection is the primary latency lever.

How to eliminate wrong answers

Option A is wrong because the Converse API is a higher-level abstraction for multi-turn conversations that adds additional processing overhead, not reducing latency. Option B is wrong because reducing the temperature parameter affects output randomness, not inference speed; latency is determined by model size and infrastructure, not sampling parameters. Option D is wrong because increasing maxTokens forces the model to generate longer sequences, which increases latency proportionally to the number of tokens generated.

848
MCQmedium

A healthcare organization uses an ML model to predict patient readmission risk. The model performs well overall but has significantly higher false negative rates for elderly patients. The team needs to mitigate this bias. Which step should they take FIRST?

A.Retrain the model with fairness constraints using SageMaker Clarify
B.Remove the age feature from the model
C.Collect more data for elderly patients to reduce representation bias
D.Use SageMaker Clarify to compute bias metrics and identify the extent of the disparity
AnswerD

SageMaker Clarify quantifies bias through metrics such as disparate impact and conditional demographic disparity across the elderly subgroup, establishing the baseline magnitude before any mitigation. Measuring the disparity first satisfies the stem's requirement to identify the extent of bias, since remediation techniques cannot be selected or validated without that evidence.

Why this answer

The first step is to identify and understand the bias through measurement. SageMaker Clarify can compute fairness metrics like false negative rate difference to quantify the disparity. Retraining with fairness constraints or reweighting are mitigation steps that come after measurement.

849
MCQhard

A startup is choosing between two foundation models for a summarization feature. Model X is a large general-purpose model with strong benchmark scores; Model Y is a smaller domain-specialized model with lower general benchmarks but excellent results on the startup's own document samples. Cost per token for Model Y is roughly one third of Model X. Which evaluation practice should the team follow?

A.Select Model Y solely because its cost per token is lower, since summarization quality is equivalent across models.
B.Deploy both models and let end users vote on which summaries they prefer before any offline evaluation.
C.Select Model X because higher public benchmark scores reliably predict performance on every downstream task.
D.Evaluate both models on a held-out set of representative documents, then weigh measured quality against cost and latency.
AnswerD

Task-specific evaluation on representative, held-out data is the most reliable signal of downstream performance, and combining it with cost and latency reflects the real trade-offs of production. This approach uses the startup's own evidence rather than generic benchmarks or price alone, which is the sound engineering practice for model selection.

Why this answer

Model selection should be driven by measured performance on data that resembles the production workload, combined with operational constraints such as cost and latency. Public benchmarks and price alone are weak proxies, while a held-out evaluation set gives direct evidence of whether the cheaper specialized model meets the quality bar for summarization.

Exam trap

The trap here is equating leaderboard rankings or lowest price with fitness for a specific task, when only task-specific evaluation reveals the real trade-off.

850
Multi-Selectmedium

A developer is building a chatbot that must refuse to answer questions about internal financial data. They also need to filter out any offensive language from user inputs. Which TWO Bedrock features should they use? (Choose TWO.)

Select 2 answers
A.Bedrock Agent
B.Bedrock Guardrails – topic denial
C.Bedrock Guardrails – content filtering
D.Prompt engineering with negative instructions
E.Bedrock Knowledge Base
AnswersB, C

Bedrock Guardrails topic denial lets you define denied topics, so the chatbot refuses questions about internal financial data. Guardrails also apply content filters that block offensive language in user inputs, satisfying both requirements in one configuration. This directly addresses the stem's dual constraints: refusing financial topics and filtering profanity.

Why this answer

Option B, Bedrock Guardrails – topic denial, is correct because it lets you define denied topics (such as internal financial data) that the model must refuse to discuss, directly satisfying the requirement to block questions about internal financial data. Option C, Bedrock Guardrails – content filtering, is correct because it detects and filters harmful or offensive content like hate speech, insults, and profanity in both inputs and outputs, which addresses the need to filter offensive language from user inputs. Option A, Bedrock Agent, is not correct because it orchestrates tasks and API calls for action-taking workflows, not content refusal or profanity filtering.

Option D, prompt engineering with negative instructions, is not correct because prompt-level instructions are not a reliable enforcement mechanism and lack the managed policy controls of Guardrails. Option E, Bedrock Knowledge Base, is not correct because it is a RAG feature for retrieving and grounding answers in data sources, not for denying topics or filtering offensive language.

851
MCQeasy

A company wants to use Amazon Bedrock to generate product descriptions for an e-commerce catalog. They need to process 100,000 product records efficiently and cost-effectively. Which inference option should they choose?

A.On-demand inference
B.Batch inference
C.Model caching
D.Provisioned throughput
AnswerB

Batch inference processes large volumes of records asynchronously through a single job, which suits the 100,000-record catalogue without maintaining real-time throughput. Amazon Bedrock batch jobs typically complete within 24 hours at roughly 50% lower cost than on-demand inference, directly satisfying the stem's efficiency and cost-effectiveness constraints.

Why this answer

Batch inference is the correct choice because it is designed for asynchronous, high-volume processing of large datasets like 100,000 product records. It processes requests in bulk, significantly reducing per-record cost compared to real-time options, and is ideal for non-latency-sensitive workloads such as generating product descriptions for an entire catalog.

Exam trap

A common pitfall is assuming that on-demand inference is suitable for all use cases. For large-scale, non-real-time batch jobs like generating 100,000 product descriptions, Batch inference in Amazon Bedrock is the most cost-effective and efficient choice. Candidates often overlook the asynchronous processing and cost savings of batch inference.

How to eliminate wrong answers

Option A is wrong because on-demand inference is optimized for real-time, low-latency requests and charges per-token, making it cost-prohibitive for processing 100,000 records in bulk. Option C is wrong because model caching reduces latency for repeated inference calls by caching model weights or intermediate results, but it does not provide a cost-effective batch processing mechanism for large-scale offline tasks. Option D is wrong because provisioned throughput guarantees a fixed level of throughput for consistent, low-latency inference, but it incurs hourly costs regardless of usage, making it overkill and expensive for a one-time or infrequent batch job.

852
Multi-Selectmedium

A company wants to build a multi-language customer support chatbot using Amazon Bedrock. The chatbot should support English, Spanish, and French. The team needs to translate user queries into English before processing and then translate responses back. Which TWO approaches could achieve this? (Choose TWO)

Select 2 answers
A.Use a separate model fine-tuned for translation for each language pair
B.Use a Bedrock Agent with an action group that calls a translation Lambda function
C.Instruct the foundation model via prompt engineering to translate the query and response
D.Configure Bedrock Guardrails to translate automatically
E.Store translated documents in a Bedrock Knowledge Base
AnswersB, C

A Bedrock Agent action group invoking a Lambda function lets the agent call deterministic translation code for queries and responses, rather than relying on model output alone. This satisfies the requirement to translate between English, Spanish and French reliably.

Why this answer

Option B is correct because a Bedrock Agent can invoke an action group backed by a Lambda function that calls a translation service (e.g., Amazon Translate) to convert the user query into English and translate the response back, giving a deterministic, programmatic translation step within the agent workflow. Option C is correct because a foundation model on Bedrock can be instructed through prompt engineering to perform the translation as part of the same inference call, e.g., asking it to translate the Spanish or French input to English, answer, then translate the answer back, which is a valid no-extra-service approach. Option A is not appropriate because fine-tuning a separate model per language pair is costly and unnecessary when Bedrock models or a translation service can handle translation directly.

Option D is wrong because Bedrock Guardrails are for content filtering, PII redaction, and topic denial, not translation. Option E is wrong because a Knowledge Base is for retrieval-augmented generation over stored documents, not for translating queries and responses.

Exam trap

AIF-C01 often tests whether candidates confuse Guardrails or Knowledge Bases with translation capabilities, or overlook the simplicity of prompt engineering for multilingual tasks.

853
Multi-Selectmedium

A company wants to build a system that automatically routes support tickets to the appropriate department based on the ticket text. The system must handle new categories that emerge over time without retraining. Which TWO approaches should the company combine to achieve this? (Select TWO.)

Select 2 answers
A.Use Amazon Forecast to predict ticket volume
B.Use Amazon Comprehend for custom classification
C.Use Amazon Comprehend for topic modeling and entity detection
D.Use Amazon Rekognition to analyze ticket images
E.Use Amazon Kendra to index support documents and perform intelligent search
AnswersB, E

Custom classification requires retraining for new categories, which does not meet the requirement.

Why this answer

Option B is correct because Amazon Comprehend custom classification can categorize ticket text into departments, providing the routing decision. Option E is correct because Amazon Kendra indexes support documents and uses semantic search to match ticket text to the most relevant department knowledge base, allowing new or emerging categories to be handled through indexed content without retraining the classifier. Together, Comprehend custom classification and Kendra enable routing that adapts to new categories by using Kendra to surface relevant department knowledge for unfamiliar tickets.

Option A is incorrect because Amazon Forecast predicts time-series values like ticket volume, not ticket content or routing. Option C is incorrect because Comprehend topic modeling and entity detection discover topics and entities but do not directly route tickets to departments. Option D is incorrect because Amazon Rekognition analyzes images and video, not support ticket text.

Exam trap

The trap is that candidates may select Comprehend topic modeling (C) because it sounds like it handles new categories without retraining, but topic modeling does not perform routing to departments; the correct approach combines Comprehend custom classification (B) with Kendra (E) for semantic search-based routing.

854
MCQmedium

A developer is using Bedrock Agents to create a travel planning assistant. The agent needs to call a hotel booking API and a flight API. What is the correct way to define these external API calls in Bedrock Agents?

A.Create action groups with the API schemas and Lambda functions
B.Hardcode the API endpoints in the agent's prompt
C.Configure the APIs as guardrails content filters
D.Define the APIs as knowledge base data sources
AnswerA

Action groups bind an OpenAPI schema describing each API's operations to a Lambda function that executes the call. Defining the hotel and flight APIs this way lets the Bedrock agent invoke them during orchestration, satisfying the requirement for external API calls.

Why this answer

Action groups in Bedrock Agents define the tools (APIs) that the agent can call, including the OpenAPI schema and Lambda function integration. Knowledge Bases are for RAG, not API calls. Guardrails are for content filtering.

The playground is for testing prompts.

855
MCQmedium

A company is building a chatbot using Amazon Bedrock to answer customer questions about their product catalog. The chatbot should only use information from the company's internal knowledge base and should not generate answers based on the model's pre-training data. Which feature should be enabled?

A.Use prompt engineering to instruct the model to only use the knowledge base
B.Configure a knowledge base with Retrieval Augmented Generation (RAG)
C.Enable model invocation logging to review responses
D.Fine-tune the model on the product catalog data
AnswerB

RAG grounds responses in the supplied knowledge base by retrieving relevant documents and injecting them into the prompt context, so the model answers from company data rather than its pre-training weights. This directly satisfies the constraint that answers must come only from the internal catalogue.

Why this answer

Configuring a knowledge base with Retrieval Augmented Generation (RAG) allows the chatbot to retrieve relevant documents from the company's internal knowledge base and use them as context for generating answers. This ensures the model's responses are grounded solely in the provided data, preventing reliance on its pre-training knowledge.

Exam trap

The trap here is that candidates often confuse fine-tuning with RAG, assuming fine-tuning alone can restrict the model to a specific knowledge domain, when in fact fine-tuning does not prevent the model from using its pre-training data and can still produce off-topic responses.

How to eliminate wrong answers

Option A is wrong because prompt engineering alone cannot reliably prevent the model from using its pre-training data; it only provides instructions that the model may still override with its internal knowledge. Option C is wrong because model invocation logging only records responses for auditing and debugging, it does not constrain the model's source of information. Option D is wrong because fine-tuning adapts the model to the product catalog but does not guarantee that the model will ignore its pre-training data; it can still generate answers from its original training corpus.

856
MCQmedium

A financial services company has built an internal assistant on Amazon Bedrock using Anthropic Claude 3 Sonnet. Employees ask questions that require retrieving the latest internal policy documents, which are updated frequently and stored in Amazon S3. The company wants the assistant to answer with accurate, up-to-date citations without retraining the model. Which approach should they implement?

A.Increase the temperature parameter to encourage the model to generate more detailed and specific policy answers.
B.Enable model invocation logging in Amazon Bedrock to capture requests and responses for later review.
C.Fine-tune the Claude 3 Sonnet model on the policy documents using Amazon Bedrock custom models.
D.Use Retrieval Augmented Generation (RAG) by creating a knowledge base in Amazon Bedrock that indexes the S3 documents.
AnswerD

RAG with an Amazon Bedrock knowledge base retrieves relevant document chunks from the S3 data source at query time and passes them to the model as context. This grounds responses in current content and can return citations, all without retraining or fine-tuning the underlying foundation model.

Why this answer

Grounding a foundation model in frequently changing internal documents is best achieved with Retrieval Augmented Generation. An Amazon Bedrock knowledge base ingests the S3 content, chunks and embeds it, and retrieves relevant passages at inference time so the model can answer with current, citable information. This avoids retraining and keeps responses accurate as policies evolve.

Exam trap

The trap here is assuming that fine-tuning is the way to add new knowledge, when RAG is the appropriate pattern for frequently updated, citable source material.

857
Multi-Selectmedium

A company wants to use AWS Lake Formation to govern access to data used for AI training. They need to ensure that only approved columns of sensitive tables are visible to data scientists. Which THREE steps should they implement? (Choose THREE)

Select 3 answers
A.Use AWS Glue crawlers to catalog the data and populate the Data Catalog
B.Create a SageMaker notebook instance and attach an IAM role
C.Define column-level permissions in Lake Formation to grant access to specific columns for the data scientist role
D.Enable S3 versioning on the training data bucket
E.Register the S3 bucket containing the training data with Lake Formation
AnswersA, C, E

AWS Glue crawlers scan the underlying data, infer schemas and register tables in the Data Catalog. Lake Formation permissions are granted against those catalogued tables and columns, so cataloguing is the prerequisite step enabling column-level access control.

Why this answer

The scenario requires governing access to AI training data with Lake Formation and restricting visibility to approved columns, so the solution must include cataloging, registering the data location, and applying column-level grants. Option A is correct because AWS Glue crawlers scan the S3 data, infer schemas, and populate the AWS Glue Data Catalog, which Lake Formation relies on to define and enforce permissions on tables and columns. Option C is correct because Lake Formation supports column-level permissions, allowing you to grant SELECT on only specific columns of a table to the data scientist role, which directly satisfies the requirement that only approved columns be visible.

Option E is correct because the S3 bucket containing the training data must be registered with Lake Formation so that Lake Formation manages access to the underlying data and can enforce its table and column permissions. Option B is not required because a SageMaker notebook instance with an IAM role is a compute/access mechanism, not a governance step for restricting column visibility. Option D is not relevant because S3 versioning provides object version retention and recovery, not fine-grained column access control.

Exam trap

AIF-C01 often tests the steps for Lake Formation governance, and candidates may overlook the need to register the S3 bucket or confuse it with other services like SageMaker, leading to incorrect selections.

858
MCQeasy

Which of the following is an example of reinforcement learning?

A.Predicting house prices based on historical data
B.Grouping customers into segments based on purchasing behavior
C.Detecting spam emails using labeled examples
D.A robot learning to navigate a maze by receiving rewards for reaching the goal
AnswerD

A robot learning to navigate a maze by receiving rewards for reaching the goal exemplifies reinforcement learning: an agent takes actions within an environment and learns a policy from reward signals, rather than labelled examples. This satisfies the stem's requirement for trial-and-error learning driven by cumulative reward maximisation.

Why this answer

Reinforcement learning involves an agent learning to make decisions by interacting with an environment and receiving rewards or penalties for its actions. Option D describes a robot learning to navigate a maze by receiving rewards for reaching the goal, which is a classic example of an agent optimizing its policy through trial and error to maximize cumulative reward.

Exam trap

The AWS AI Practitioner exam often tests the distinction between supervised, unsupervised, and reinforcement learning by presenting tasks that involve feedback (like rewards) versus tasks that use labeled data or no labels, and the trap here is that candidates may confuse any task with a 'goal' or 'outcome' as reinforcement learning, even when it uses pre-labeled examples or historical data without an interactive environment.

How to eliminate wrong answers

Option A is wrong because predicting house prices based on historical data is a supervised learning regression task, where the model learns from labeled input-output pairs (features and continuous target values), not from rewards or interactions. Option B is wrong because grouping customers into segments based on purchasing behavior is an unsupervised learning clustering task (e.g., using k-means), which finds patterns in unlabeled data without any reward signal. Option C is wrong because detecting spam emails using labeled examples is a supervised learning classification task, where the model is trained on pre-labeled examples (spam/not spam) to predict labels, not through environmental feedback.

859
MCQhard

An enterprise deploys a foundation model on Amazon Bedrock with a knowledge base. Users report that the model is returning outdated information. What is the most likely cause?

A.The model was fine-tuned
B.The model is not the latest version
C.The knowledge base data source is not refreshed
D.The inference parameters are incorrect
AnswerC

Stale source content propagates directly into retrieval: Amazon Bedrock knowledge bases sync from the configured data source, so unchanged documents keep returning outdated chunks regardless of model capability. Refreshing or re-syncing the data source satisfies the freshness constraint the stem describes.

Why this answer

When a knowledge base is attached to a foundation model on Amazon Bedrock, the model retrieves information from the data source to augment its responses. If the data source is not refreshed, the model will return outdated information even if the model itself is current. Option C directly addresses this by identifying the stale data source as the root cause.

Exam trap

The trap here is that candidates may confuse model versioning (Option B) with data freshness, but the question specifically ties the symptom to the knowledge base, making the refresh cycle the critical factor.

How to eliminate wrong answers

Option A is wrong because fine-tuning adjusts the model's weights on a specific dataset, which does not inherently cause outdated information; in fact, fine-tuning could update the model with newer data. Option B is wrong because using an older model version might affect performance or capabilities, but the question specifically states the model is returning outdated information, which points to the knowledge base content, not the model version. Option D is wrong because inference parameters (e.g., temperature, top_p) control randomness and creativity of responses, not the freshness or accuracy of the information retrieved from the knowledge base.

860
Multi-Selectmedium

Which THREE statements about Amazon SageMaker Ground Truth are correct? (Choose three.)

Select 3 answers
A.It can only be used for text data.
B.It provides built-in workflows for image classification and object detection.
C.It supports automated data labeling using active learning.
D.It integrates with Amazon SageMaker to use the labeled data for training.
E.It can only use a public workforce from Amazon Mechanical Turk.
AnswersB, C, D

Ground Truth ships managed labelling interfaces and task templates for image classification and object detection, so teams avoid building custom annotation tooling. This built-in workflow support satisfies the stem's requirement that the statement describe a genuine Ground Truth capability.

Why this answer

Option B is correct because Amazon SageMaker Ground Truth ships with built-in labeling task templates and workflows for common job types, including image classification and object detection (as well as text, semantic segmentation, and bounding boxes). Option C is correct because Ground Truth supports automated data labeling, which uses active learning to have a model label high-confidence data while sending only low-confidence samples to human labelers, reducing cost and effort. Option D is correct because Ground Truth integrates directly with Amazon SageMaker, writing labeled datasets to Amazon S3 in augmented manifest format that SageMaker training jobs can consume.

Option A is wrong because Ground Truth handles images, video, text, and 3D point clouds, not just text. Option E is wrong because Ground Truth supports multiple workforces, including private workforces, vendor-managed workforces, and Amazon Mechanical Turk, so it is not limited to a public Mechanical Turk workforce.

Exam trap

AWS often tests the misconception that Ground Truth is limited to text data or only supports public workforces, while in reality it handles multiple data modalities and offers flexible workforce options including private and vendor-managed.

861
MCQhard

A developer is using the Amazon Bedrock InvokeModel API with a model that has a context window of 8,000 tokens. The developer sends a prompt that is 7,500 tokens long and expects a response of about 1,000 tokens. The API call fails with an error indicating the input exceeds the model's context window. Why did this happen?

A.The API automatically reserves 2,000 tokens for output, reducing available input capacity
B.The model requires a minimum of 512 tokens for internal processing
C.The context window includes both input and output tokens, so the total of 7,500 + 1,000 = 8,500 exceeds the 8,000 limit
D.The model's context window counts only input tokens; output tokens are separate
AnswerC

The context window caps combined input and output tokens, not input alone. With 7,500 prompt tokens plus an expected 1,000 response tokens, the total of 8,500 exceeds the model's 8,000-token limit, so Bedrock rejects the request before generation begins.

Why this answer

The context window of a foundation model includes both the input prompt and the generated output tokens. The developer's prompt of 7,500 tokens plus the expected 1,000-token response totals 8,500 tokens, which exceeds the model's 8,000-token limit. The InvokeModel API enforces this combined limit, so even though the prompt alone is under the limit, the total token count causes the error.

Exam trap

The trap here is that candidates often assume the context window applies only to the input prompt, ignoring that the output response also consumes tokens from the same window, leading them to incorrectly select option D.

How to eliminate wrong answers

Option A is wrong because the API does not automatically reserve a fixed number of tokens for output; the context window is shared dynamically between input and output, and the error occurs only when the total exceeds the limit. Option B is wrong because there is no minimum token requirement for internal processing; the model can handle prompts of any length up to the context window limit. Option D is wrong because the context window counts both input and output tokens together, not input tokens separately; output tokens consume part of the same window.

862
MCQeasy

Which of the following is NOT one of the core principles of responsible AI as defined by AWS?

A.Transparency
B.Fairness
C.Profitability
D.Robustness
AnswerC

AWS responsible AI centres on fairness, explainability, privacy, security, safety, controllability and governance. Profitability is a commercial objective, not a responsible AI principle, so it is correctly excluded, satisfying the question's requirement to identify the option that is NOT a core principle.

Why this answer

Profitability is not one of the core principles of responsible AI as defined by AWS. AWS defines six core principles: Fairness, Transparency, Robustness, Privacy, Security, and Explainability. Profitability is a business objective, not a principle for ensuring ethical and trustworthy AI systems.

Exam trap

The trap here is that candidates may confuse business objectives like profitability or cost-efficiency with the ethical and technical governance principles that AWS specifically defines for responsible AI.

How to eliminate wrong answers

Option A is wrong because Transparency is a core AWS responsible AI principle that requires AI systems to be open about their capabilities, limitations, and decision-making processes. Option B is wrong because Fairness is a core principle that mandates AI systems should treat all groups equitably and avoid bias. Option D is wrong because Robustness is a core principle that ensures AI systems perform reliably under varying conditions and resist adversarial inputs.

Page 11

Page 12 of 12