Courseiva

AWS Certified AI Practitioner AIF-C01 (AIF-C01) — Questions 676–750

862 questions total · 12pages · All types, answers revealed

Page 9

Page 10 of 12

Page 11
676
MCQmedium

A developer is building a retrieval-augmented generation (RAG) assistant on Amazon Bedrock. The assistant must answer questions about internal policy documents that change frequently, and answers must cite the source passages. The developer wants a managed capability that handles chunking, embedding, and retrieval so the application code stays minimal. Which Amazon Bedrock feature should the developer use?

A.Amazon Bedrock model evaluation
B.Amazon Bedrock knowledge bases
C.Amazon Bedrock Agents
D.Amazon Bedrock Guardrails
AnswerB

Amazon Bedrock knowledge bases provide a managed RAG workflow: they ingest and chunk source documents, generate embeddings, store them in a vector store, and retrieve relevant passages at query time. The model can then generate answers grounded in those retrieved passages. This removes the need for the developer to orchestrate chunking, embedding, and retrieval manually, matching the requirement for minimal application code.

Why this answer

Amazon Bedrock knowledge bases are the managed RAG capability that handles ingestion, chunking, embedding, vector storage, and retrieval, enabling grounded answers with source citations while keeping application code small. Evaluation, Guardrails, and Agents address quality measurement, content safety, and task orchestration respectively, none of which deliver the required retrieval pipeline on their own.

Exam trap

The trap here is conflating orchestration or safety features with retrieval, when only knowledge bases perform the managed chunking, embedding, and retrieval work.

677
MCQmedium

A financial services company is deploying a foundation model to analyze customer sentiment from call transcripts. The model outputs must be consistent and deterministic for auditing purposes. Which parameter configuration should the company use?

A.Set temperature to 0.1 and top_p to 0.9.
B.Set temperature to 0.7 and top_p to 1.0.
C.Set temperature to 0.5 and top_p to 0.5.
D.Set temperature to 0 and top_p to 1.
AnswerD

Setting temperature to 0 makes the model select the highest-probability token at each step, eliminating the random sampling that produces run-to-run variation. Keeping top_p at 1 disables nucleus filtering, so it cannot reintroduce randomness. This satisfies the audit requirement for consistent, deterministic sentiment outputs from identical transcripts.

Why this answer

Setting temperature to 0 and top_p to 1 forces the model to always select the highest-probability token at each step, producing deterministic and repeatable outputs. This is essential for auditing and compliance in financial services, where consistency is required. Any nonzero temperature introduces randomness, which undermines determinism.

Exam trap

AWS often tests the misconception that low temperature (e.g., 0.1) is 'deterministic enough,' but only temperature exactly 0 guarantees deterministic outputs, and top_p must be 1 to avoid interfering with the argmax selection.

How to eliminate wrong answers

Option A is wrong because temperature 0.1 still introduces slight randomness, making outputs non-deterministic and unsuitable for auditing. Option B is wrong because temperature 0.7 introduces significant randomness, and top_p 1.0 does not constrain it, leading to high variability. Option C is wrong because temperature 0.5 introduces randomness, and top_p 0.5 further restricts token sampling but does not eliminate the stochastic behavior from the nonzero temperature.

678
MCQmedium

A team has built a regression model to predict house prices. The RMSE is 50,000 on the test set. Which action is most appropriate to improve model performance?

A.Remove outliers from training data
B.Apply feature scaling
C.Add more relevant features
D.Use a different evaluation metric
AnswerC

An RMSE of 50,000 indicates underfitting or missing signal, so adding relevant features gives the regression model additional predictive information. This addresses the performance gap more directly than hyperparameter tuning or collecting more rows of the same variables.

Why this answer

Adding more relevant features can provide the model with additional predictive signals, potentially reducing bias and lowering RMSE if the new features have genuine correlation with house prices. Since RMSE is already 50,000, the model may be underfitting due to insufficient input variables, and enriching the feature set is a direct way to capture more variance in the target variable.

Exam trap

The AWS AI Practitioner exam often tests the misconception that data preprocessing steps like scaling or outlier removal are universal fixes for high error, when in fact the most appropriate first step for a high RMSE in regression is to improve the feature set to address underfitting.

How to eliminate wrong answers

Option A is wrong because removing outliers from training data can reduce variance but may also discard valuable extreme cases that reflect real market conditions, and it does not address the core issue of underfitting or missing predictive signals. Option B is wrong because feature scaling (e.g., normalization or standardization) is important for gradient-based optimization and distance-based algorithms, but it does not inherently improve model accuracy for tree-based or linear regression models when RMSE is already computed on unscaled data; it affects convergence speed, not predictive power. Option D is wrong because changing the evaluation metric (e.g., from RMSE to MAE or R²) does not improve the model's actual predictive performance; it only changes how performance is measured, potentially masking the same underlying error.

679
Multi-Selecthard

A team is designing a RAG system on Amazon Bedrock. They need to chunk a large set of PDF documents into smaller pieces for embedding. Which THREE considerations should guide their chunking strategy? (Choose three.)

Select 3 answers
A.All chunks must be exactly the same length for optimal performance
B.Chunk size should be small enough to fit multiple chunks within the model's context window after including the query
C.Consider the embedding model's maximum input token limit
D.Overlapping chunks should be avoided to reduce redundancy
E.Chunks should align with natural semantic boundaries (e.g., paragraphs, sections)
AnswersB, C, E

Chunk size must leave room for the query and generated response inside the model's context window, so several retrieved chunks can be concatenated without truncation. This satisfies the stem's embedding-and-retrieval constraint: oversized chunks consume the window, reducing how many passages ground the answer.

Why this answer

Option B is correct because chunk size must be small enough that several retrieved chunks plus the user query and prompt fit within the foundation model's context window; otherwise retrieved content will be truncated or the request will fail. Option C is correct because each chunk is passed to the embedding model, so it must not exceed that model's maximum input token limit (for example, Amazon Titan Text Embeddings V2 supports up to 8,192 tokens), or the chunk will be rejected or silently truncated. Option E is correct because splitting on natural semantic boundaries such as paragraphs, headings, and sections preserves coherent meaning, which produces more accurate embeddings and better retrieval relevance than arbitrary fixed-length cuts.

Option A is not required: uniform chunk length is not a performance requirement, and forcing equal sizes often breaks semantic boundaries. Option D is not correct: overlap is commonly recommended because it preserves context across chunk boundaries and prevents information loss at split points, even though it adds some redundancy.

Exam trap

AWS often tests the misconception that uniform chunk sizes are optimal, when in fact semantic alignment and context window constraints are more critical for RAG performance.

680
MCQhard

An AI practitioner is fine-tuning an Amazon Titan Text model on a dataset of customer support conversations to improve response accuracy. After training, the model's perplexity on the validation set is low, but during inference, the model frequently generates off-topic or nonsensical responses to real customer queries. What is the most likely cause?

A.The model's context window is too small for the inference queries
B.The temperature during inference is set too low
C.The fine-tuning dataset is too small, causing overfitting
D.The validation set does not represent the distribution of real customer queries
AnswerD

Low validation perplexity with nonsensical live responses indicates the validation set shares the training distribution but not the real query distribution, so the model never learned the actual input patterns. The gap is distributional mismatch, not overfitting to the validation set itself.

Why this answer

The core issue is a mismatch between the validation set and real-world inference data. Low perplexity on the validation set indicates the model fits that specific distribution well, but if the validation set does not reflect the diversity, phrasing, or intent of actual customer queries, the model will generate off-topic or nonsensical responses during inference. This is a classic case of distribution shift, where the model has not generalized to the true target domain.

Exam trap

The AIF-C01 exam often tests the distinction between overfitting (Option C) and distribution mismatch (Option D), where candidates mistakenly attribute low validation perplexity with good generalization, but the trap is that overfitting would still show high perplexity on a representative validation set, whereas here the validation set itself is the problem.

How to eliminate wrong answers

Option A is wrong because a small context window would truncate input, not cause off-topic or nonsensical responses; it would more likely cause incomplete or irrelevant answers due to missing context, but the model would still stay on-topic within its truncated window. Option B is wrong because a temperature set too low makes the model deterministic and repetitive, reducing creativity but not causing off-topic or nonsensical outputs; low temperature actually forces the model to pick the most likely token, which should keep responses coherent and on-topic if the model is well-trained. Option C is wrong because a small fine-tuning dataset causing overfitting would lead to low perplexity on the validation set (which matches the training distribution) but high perplexity on unseen data; however, the question states perplexity is low on the validation set, and overfitting would typically cause poor performance on any out-of-distribution data, not specifically off-topic or nonsensical responses—this is a subtle but important distinction, as overfitting often produces memorized or repetitive outputs, not random nonsense.

681
MCQeasy

A data scientist trains a linear regression model to predict housing prices. The model achieves a low training error but a high test error. Which concept does this BEST illustrate?

A.Bias-variance tradeoff
B.Regularization
C.Underfitting
D.Overfitting
AnswerD

Low training error with high test error means the model has memorised training noise rather than learning generalisable patterns, so it fails on unseen data. Overfitting is the specific term for this train-test performance gap, which the stem describes directly.

Why this answer

The scenario describes a model that performs well on training data but poorly on unseen test data, which is the classic definition of overfitting. Overfitting occurs when the model learns noise and specific patterns in the training set rather than the underlying generalizable relationship, leading to high variance and poor test performance.

Exam trap

AWS often tests the distinction between overfitting and the bias-variance tradeoff, where candidates mistakenly select 'bias-variance tradeoff' because they recognize high variance, but the question explicitly asks for the concept best illustrated by the specific error pattern, which is overfitting.

How to eliminate wrong answers

Option A is wrong because the bias-variance tradeoff is a broader concept that describes the balance between underfitting (high bias) and overfitting (high variance); while overfitting is a manifestation of high variance, the question specifically asks for the concept best illustrated by low training error and high test error, which is directly overfitting, not the tradeoff itself. Option B is wrong because regularization is a technique used to reduce overfitting by adding a penalty term (e.g., L1 or L2) to the loss function, not the problem being illustrated. Option C is wrong because underfitting is characterized by high training error and high test error, which is the opposite of the given scenario.

682
MCQeasy

What is the primary role of the self-attention mechanism in the Transformer architecture?

A.To allow each token to attend to all other tokens in the sequence, capturing long-range dependencies
B.To process tokens in parallel by alternating attention and feed-forward layers
C.To reduce the vocabulary size by mapping tokens to embeddings
D.To generate the next token one at a time in an autoregressive manner
AnswerA

Self-attention computes pairwise relevance scores between every token and all others, so each position aggregates context across the entire sequence in one layer. This directly satisfies the stem's requirement to capture long-range dependencies, unlike recurrent architectures whose signal degrades over distance.

Why this answer

The self-attention mechanism allows each token in the input sequence to directly attend to every other token, computing a weighted sum of all token representations. This enables the model to capture long-range dependencies and contextual relationships regardless of distance, which is the fundamental innovation of the Transformer architecture over recurrent or convolutional models.

Exam trap

AWS often tests the distinction between the specific function of a component (self-attention's role in capturing dependencies) and the broader architectural or procedural behavior (parallel processing, embedding, or autoregressive generation), leading candidates to confuse the mechanism with its effects or surrounding architecture.

How to eliminate wrong answers

Option B is wrong because processing tokens in parallel by alternating attention and feed-forward layers describes the overall Transformer architecture, not the specific role of self-attention; self-attention is the component that enables parallelization by removing sequential recurrence, but its primary role is dependency capture. Option C is wrong because reducing vocabulary size by mapping tokens to embeddings is the function of the embedding layer (e.g., token embedding or word embedding), not self-attention; self-attention operates on the embedded representations to model relationships. Option D is wrong because generating the next token one at a time in an autoregressive manner describes the decoding process (e.g., in GPT-style models), not the role of self-attention; self-attention is used within both encoder and decoder to compute context, but autoregressive generation is a separate inference strategy.

683
Multi-Selectmedium

A company is using Amazon SageMaker to train machine learning models. The security team wants to ensure that the training data is encrypted at rest and that the SageMaker notebook instances cannot access the internet. Which TWO actions should the company take? (Choose TWO.)

Select 2 answers
A.Enable S3 server-side encryption with AWS KMS (SSE-KMS) for the training data bucket
B.Create an AWS CloudTrail trail to log all S3 data events
C.Enable encryption at rest for the SageMaker endpoint using the AWS Management Console
D.Disable internet access for the SageMaker notebook instance by placing it in a VPC without a NAT gateway or internet gateway
E.Use AWS Security Token Service (STS) to generate temporary credentials for the notebook instance
AnswersA, D

SSE-KMS encrypts objects at rest using KMS keys.

Why this answer

Enabling S3 server-side encryption with AWS KMS (SSE-KMS) ensures that the training data stored in the S3 bucket is encrypted at rest. This satisfies the security team's requirement for data encryption at rest, as SSE-KMS provides envelope encryption with a customer-managed or AWS-managed KMS key, giving the company control over the encryption keys and auditability via AWS CloudTrail.

Exam trap

The trap here is that candidates often confuse encryption at rest for the endpoint (Option C) with encryption of the training data in S3, or they mistakenly think that CloudTrail logging (Option B) or STS credentials (Option E) provide encryption, when in fact they address auditing and access control, not data encryption.

684
MCQmedium

A financial services firm uses Amazon Bedrock to generate investment summaries. They need to prevent the model from generating content containing personally identifiable information (PII) such as social security numbers. Which feature should they configure in Bedrock Guardrails?

A.Contextual grounding checks
B.Content filtering with category-based harmful content filters
C.Sensitive information filters with PII redaction
D.Word filters with a custom list of terms
AnswerC

Sensitive information filters detect and redact personally identifiable information, matching the stem's requirement to block social security numbers in generated summaries. Bedrock Guardrails applies these filters to both prompts and responses, so PII is caught before reaching users. Unlike content filters, which target harmful categories such as hate or violence, this mechanism specifically addresses data privacy.

Why this answer

Bedrock Guardrails include a PII redaction filter that can detect and block or mask PII in model inputs and outputs.

685
Multi-Selectmedium

A product team wants its generative AI assistant to answer questions about internal policy documents accurately rather than from the model's general training knowledge. Which TWO techniques should they apply? (Choose two.)

Select 2 answers
A.Increase the temperature setting so the model explores a wider range of possible answers.
B.Shorten the system prompt to a single sentence to reduce token consumption.
C.Instruct the model to answer only from the provided context and to state when the context is insufficient.
D.Ask the model to produce the longest possible answer for every question.
E.Retrieve relevant passages from the policy corpus and include them in the prompt before generation.
AnswersC, E

An explicit instruction that constrains the model to the supplied context, plus a fallback behavior for missing information, prevents the model from filling gaps with plausible but invented policy details. This complements retrieval by defining how the model should behave when retrieved evidence is absent or incomplete.

Why this answer

Grounding combines retrieval of authoritative source passages with an instruction that limits the model to that evidence and requires an explicit fallback when evidence is missing. Together these reduce reliance on parametric knowledge and make answers traceable to internal documents, which is what accurate policy question answering demands.

Exam trap

The trap here is treating generation settings such as temperature or prompt length as accuracy controls, when grounding depends on supplying evidence and constraining the model to it.

686
MCQmedium

An organization is required to provide transparency about AI-generated content. Which of the following is the best practice to comply with transparency requirements?

A.Clearly label AI-generated content with a disclosure statement
B.Store metadata but not display it to users
C.Use a watermark that is invisible to users
D.Only disclose AI generation if the content is inaccurate
AnswerA

Labelling AI-generated content with a disclosure statement directly satisfies the transparency requirement by informing users of the content's synthetic origin. Unlike watermarking, which embeds provenance signals in the media itself, or audit logging, which records system activity internally, disclosure operates at the point of consumption, giving the audience explicit notice.

Why this answer

Clearly labeling AI-generated content with a disclosure statement is the best practice to provide transparency. This directly informs users that the content is AI-generated, fulfilling transparency requirements. It is a straightforward and effective method that does not rely on hidden mechanisms.

Exam trap

AIF-C01 often tests... the difference between transparency and other concepts like explainability or accountability, and candidates might choose technical solutions like watermarking over direct disclosure.

How to eliminate wrong answers

Option B is wrong because storing metadata without displaying it does not provide transparency to users; transparency requires disclosure. Option C is wrong because an invisible watermark does not inform users at the point of consumption; it may be useful for provenance but not for immediate transparency. Option D is wrong because disclosing only when content is inaccurate is insufficient; transparency should apply regardless of accuracy to build trust.

687
MCQmedium

A company uses an AI system to screen job applications. The system was trained on resumes from previous hires, which predominantly came from a specific demographic. As a result, the system may unfairly filter out qualified candidates from other backgrounds. Which responsible AI practice should the company implement?

A.Implement bias detection metrics and monitor outcomes by demographic groups
B.Focus solely on improving the model's precision and recall
C.Defer all screening decisions to a human recruiter
D.Increase the size of the training dataset without regard to demographic composition
AnswerA

Training data skewed toward one demographic produces disparate impact, so measuring outcomes across demographic groups and applying bias detection metrics exposes and quantifies that skew. This satisfies the stem's requirement to address unfair filtering of qualified candidates from other backgrounds.

Why this answer

Implementing bias detection metrics and monitoring outcomes by demographic groups directly addresses the risk of unfair filtering. This practice aligns with the responsible AI principle of fairness, requiring continuous evaluation of model outputs across protected groups to identify and mitigate demographic disparities in hiring decisions.

Exam trap

AWS often tests the misconception that improving model accuracy or simply adding more data automatically fixes bias, when in fact biased training data requires targeted fairness interventions like reweighting, resampling, or adversarial debiasing.

How to eliminate wrong answers

Option B is wrong because focusing solely on precision and recall ignores fairness and can amplify existing biases if the training data is skewed. Option C is wrong because deferring all screening decisions to a human recruiter is impractical at scale and does not address the root cause of bias in the AI system; it also introduces human bias. Option D is wrong because increasing the training dataset size without regard to demographic composition may perpetuate or even worsen the existing demographic imbalance, failing to ensure representativeness.

688
Multi-Selectmedium

A company uses Amazon Bedrock Agents to automate a multi-step customer support workflow. The agent needs to query a customer database and update a ticket system. Which TWO components are required to enable the agent to interact with these external systems?

Select 2 answers
A.A vector store like Amazon OpenSearch Serverless
B.Bedrock Guardrails
C.Bedrock Knowledge Bases
D.Lambda functions that implement the business logic for each action
E.Action groups that define the APIs or database operations
AnswersD, E

Lambda functions supply the executable business logic behind each Bedrock Agent action group, performing the actual database queries and ticket-system updates the agent invokes. They satisfy the stem's requirement to interact with external systems, since action groups only define the API schema; the Lambda function executes the call.

Why this answer

Option E is correct because action groups are the Bedrock Agents construct that defines the APIs, operations, and parameters the agent can invoke to interact with external systems such as the customer database and ticket system. Option D is correct because each action group is backed by a Lambda function that contains the business logic to actually query the database and update the ticket system when the agent calls the action. Together, action groups describe what the agent can do and Lambda functions execute it, which is exactly what is needed for these external interactions.

Option A is not required because a vector store like Amazon OpenSearch Serverless is used for retrieval-augmented generation over document embeddings, not for transactional database queries or ticket updates. Option B is not required because Bedrock Guardrails filter and constrain content for safety and compliance, not enable external system integration. Option C is not required because Bedrock Knowledge Bases provide managed RAG over data sources for informational retrieval, not the action-based API or database operations needed here.

Exam trap

AIF-C01 often tests the components required for Bedrock Agents, and candidates may confuse knowledge bases with action groups, or overlook the need for Lambda functions to implement the business logic.

689
MCQeasy

An application needs to store and search vector embeddings of 10 million documents for a RAG system. Which Amazon vector store is a fully managed, serverless option that integrates natively with Amazon Bedrock Knowledge Bases?

A.Amazon Aurora PostgreSQL with pgvector
B.MongoDB Atlas
C.Amazon OpenSearch Serverless
D.Pinecone
AnswerC

Amazon OpenSearch Serverless provides a fully managed, serverless vector engine that scales without capacity planning and integrates natively with Amazon Bedrock Knowledge Bases as a supported vector store, satisfying both the serverless constraint and the ten-million-document scale.

Why this answer

Amazon OpenSearch Serverless is a fully managed, serverless vector store that integrates natively with Amazon Bedrock Knowledge Bases. It supports vector search for embeddings and automatically scales compute and storage capacity, making it ideal for RAG workloads with 10 million documents without requiring infrastructure management.

Exam trap

A common pitfall is choosing Amazon Aurora PostgreSQL with pgvector because it is a managed database service, but it requires manual scaling and lacks native integration with Amazon Bedrock Knowledge Bases. Amazon OpenSearch Serverless is the fully managed, serverless option that automatically scales and natively integrates for RAG workloads.

How to eliminate wrong answers

Option A is wrong because Amazon Aurora PostgreSQL with pgvector is a relational database with a vector extension, not a fully managed serverless vector store; it requires manual scaling and does not integrate natively with Amazon Bedrock Knowledge Bases. Option B is wrong because MongoDB Atlas is a third-party, non-AWS managed document database that offers vector search but is not a fully managed serverless option within AWS and lacks native integration with Amazon Bedrock Knowledge Bases. Option D is wrong because Pinecone is a third-party, standalone vector database that is not an AWS service and does not integrate natively with Amazon Bedrock Knowledge Bases.

690
MCQeasy

A company wants to predict customer churn. They have historical data with features like usage minutes, support tickets, contract length. The target is binary: churn/not churn. Which ML algorithm is best suited?

A.Logistic regression
B.Principal Component Analysis (PCA)
C.Linear regression
D.K-means clustering
AnswerA

Logistic regression outputs a probability between 0 and 1 via the sigmoid function, fitting the binary churn/not-churn target directly. It also yields interpretable coefficients across usage minutes, support tickets and contract length, so the company can see which features drive churn.

Why this answer

Logistic regression is the best choice because it is specifically designed for binary classification tasks like predicting churn (churn/not churn). It models the probability of the target class using a logistic (sigmoid) function, making it interpretable and efficient for this type of supervised learning problem with a categorical outcome.

Exam trap

The AIF-C01 exam often tests the distinction between supervised and unsupervised learning, and the trap here is that candidates may confuse dimensionality reduction (PCA) or clustering (K-means) with classification, or mistakenly apply linear regression to a binary outcome without recognizing the need for a logistic function.

How to eliminate wrong answers

Option B is wrong because Principal Component Analysis (PCA) is an unsupervised dimensionality reduction technique, not a classification algorithm; it reduces feature space but does not predict a binary target. Option C is wrong because linear regression predicts a continuous numeric output, not a binary class; using it for classification would violate the assumption of normally distributed errors and produce unbounded predictions. Option D is wrong because K-means clustering is an unsupervised learning algorithm used for grouping unlabeled data into clusters, not for predicting a known binary target variable.

691
MCQmedium

A company is using Amazon Bedrock to build an application that requires very low latency responses (under 100ms). They are currently using a large model but need faster inference. Which model selection strategy is MOST appropriate?

A.Increase the temperature parameter to speed up generation
B.Use a larger model with more parameters for better accuracy
C.Switch to a smaller model that can meet the latency requirement while still providing acceptable quality
D.Use batch processing instead of real-time streaming
AnswerC

Smaller models have fewer parameters, so each forward pass requires less computation and memory bandwidth, cutting inference latency. This satisfies the sub-100ms constraint while retaining acceptable output quality, unlike prompt tuning or provisioned throughput, which do not reduce per-token compute.

Why this answer

Smaller models have fewer parameters, which reduces the computational cost per inference, directly lowering latency. For latency-sensitive applications requiring under 100ms responses, a smaller model can often provide acceptable quality while meeting the strict timing requirement, whereas a larger model would introduce higher inference latency due to increased matrix operations and memory bandwidth demands.

Exam trap

A common misconception in AWS exams is that adjusting hyperparameters like temperature can affect inference speed, but temperature only influences output diversity and randomness.

How to eliminate wrong answers

Option A is wrong because the temperature parameter controls the randomness of token sampling, not the speed of generation; increasing temperature does not reduce inference time. Option B is wrong because using a larger model with more parameters increases computational complexity and latency, which is the opposite of what is needed for sub-100ms responses. Option D is wrong because batch processing aggregates multiple requests for parallel processing, which increases overall throughput but introduces higher per-request latency due to queuing and batching delays, making it unsuitable for real-time low-latency requirements.

692
MCQeasy

A financial services company uses Amazon Rekognition to verify customer identities. To ensure responsible AI practices, which measure should the company prioritize?

A.Use only black-box models to protect intellectual property
B.Increase model complexity to improve accuracy
C.Minimize the amount of training data collected
D.Regularly audit the model for demographic bias
AnswerD

Regular demographic bias audits directly satisfy responsible AI's fairness requirement by measuring whether Rekognition's identity verification accuracy differs across groups. This detects disparate error rates before they cause harm, which the other measures do not address.

Why this answer

Regularly auditing the model for demographic bias is a core responsible AI practice, especially for identity verification systems where biased outcomes could lead to unfair treatment of certain customer groups. Amazon Rekognition's facial analysis and comparison features must be tested across diverse demographics to ensure equitable performance, as bias can arise from imbalanced training data or algorithmic artifacts.

Exam trap

The trap here is that candidates may confuse 'responsible AI' with generic model optimization (like increasing accuracy or reducing data), but the exam specifically tests the principle of fairness through bias auditing and transparency.

How to eliminate wrong answers

Option A is wrong because using only black-box models contradicts responsible AI principles; explainability and transparency are critical for auditing bias and ensuring fairness, and black-box models obscure how decisions are made, making it harder to detect issues. Option B is wrong because increasing model complexity does not inherently improve accuracy and can amplify bias or reduce interpretability; responsible AI prioritizes balanced performance and fairness over raw accuracy. Option C is wrong because minimizing training data can exacerbate bias by underrepresenting certain demographic groups, leading to poor generalization and unfair outcomes; responsible AI requires diverse, representative datasets.

693
Multi-Selectmedium

A company is building a content generation application using Amazon Bedrock. They need to ensure that the model does not generate offensive content and also avoids discussing certain prohibited topics. Which TWO Bedrock features should be combined to achieve this?

Select 2 answers
A.Bedrock Model Evaluation
B.Bedrock Guardrails topic denial
C.Bedrock Guardrails content filters
D.Bedrock Knowledge Bases
E.Bedrock Agents
AnswersB, C

Topic denial defines prohibited subjects and blocks model responses that stray into them, enforcing the stem's requirement to avoid certain banned topics. It complements content filters, which handle offensiveness, giving the two-feature combination the scenario demands.

Why this answer

Bedrock Guardrails content filters (C) are the right choice for blocking offensive content because they let you configure thresholds across categories such as hate, violence, sexual, insults, and misconduct, filtering both prompts and model responses. Bedrock Guardrails topic denial (B) is also correct because it lets you define specific prohibited topics with natural-language descriptions and optional example phrases, so the model refuses to discuss them. Together, content filters handle offensiveness while topic denial handles restricted subject matter, which is exactly the two-part requirement.

Bedrock Model Evaluation (A) only measures and compares model quality or metrics; it does not block content at runtime. Bedrock Knowledge Bases (D) is for retrieval-augmented generation over your own data, and Bedrock Agents (E) orchestrates multi-step tasks and API calls, neither of which enforces content or topic restrictions.

694
MCQhard

Refer to the exhibit. A developer sees this error when calling Amazon Bedrock for inference. What is the MOST likely cause and recommended solution?

A.The model ID is incorrect; use a different model
B.The prompt is too long; reduce the number of tokens in the prompt
C.The request rate exceeds the model's throughput limit; implement retries with exponential backoff
D.Increase the max_tokens_to_sample value
AnswerC

ThrottlingException indicates the account's requests per minute or tokens per minute exceed the model's allocated throughput. Retrying with exponential backoff and jitter spreads retries, letting transient capacity recover instead of compounding the overload with immediate repeat calls.

Why this answer

The error indicates a throttling exception from Amazon Bedrock, which occurs when the request rate exceeds the model's throughput limit. The recommended solution is to implement retries with exponential backoff to handle transient rate limits gracefully, as this aligns with AWS best practices for managing API call limits.

Exam trap

The trap here is that candidates may confuse a throttling error with a model ID or prompt length issue, because the error message may not explicitly state 'throttling' and instead show a generic 'ServiceUnavailable' or 'TooManyRequests' response, leading them to incorrectly modify the model or prompt instead of implementing retry logic.

How to eliminate wrong answers

Option A is wrong because a model ID error would produce a different error (e.g., 'ValidationException' or 'ResourceNotFoundException'), not a throttling-related error. Option B is wrong because a prompt that is too long would cause a 'ValidationException' regarding token limits, not a throttling error. Option D is wrong because increasing max_tokens_to_sample would increase the output length, potentially worsening throttling or causing a different error, but it does not address the rate limit issue.

695
Multi-Selecthard

Which TWO of the following are key components of a responsible AI governance framework?

Select 2 answers
A.Develop and enforce AI ethics policies and standards
B.Focus solely on compliance with legal regulations
C.Minimize human involvement in AI lifecycle decisions
D.Conduct regular bias and fairness impact assessments
E.Deploy AI models as black boxes to avoid scrutiny
AnswersA, D

Policies provide the foundation for governance.

Why this answer

A responsible AI governance framework must include the development and enforcement of AI ethics policies and standards to ensure alignment with societal values, fairness, and accountability. These policies guide the design, deployment, and monitoring of AI systems, embedding ethical principles such as transparency, privacy, and non-discrimination into the AI lifecycle. Without such policies, organizations risk deploying AI that violates ethical norms or regulatory expectations.

Exam trap

The AIF-C01 exam often tests the distinction between mere legal compliance and comprehensive ethical governance, trapping candidates who think that meeting regulatory requirements alone constitutes responsible AI, while ignoring proactive fairness and transparency measures.

696
MCQmedium

An AI practitioner is evaluating a text generation model and notices that the model sometimes produces plausible-sounding but factually incorrect statements. What is this phenomenon called?

A.Hallucination
B.Catastrophic forgetting
C.Bias amplification
D.Overfitting
AnswerA

Hallucination describes a model generating fluent, plausible output that is factually wrong or unsupported by its training data. It directly matches the stem's constraint: text that sounds credible yet is incorrect. This arises because generative models predict likely token sequences rather than verifying truth against a source.

Why this answer

Hallucination in LLMs refers to generating content that is not grounded in the training data or provided context. It is a known challenge for generative models.

697
MCQmedium

A data scientist is using Amazon Bedrock to build a question-answering system over a large corpus of technical manuals. They want to ensure that the model's answers are grounded in the retrieved documents and that the model does not hallucinate. Which feature should they enable?

A.Embedding model with higher dimensionality
B.Larger chunk sizes in the knowledge base
C.Bedrock Guardrails with grounding support
D.Bedrock Agents with a multi-step reasoning prompt
AnswerC

Bedrock Guardrails with grounding support validates responses against the retrieved source documents, applying a grounding threshold that filters claims unsupported by the reference material. This directly satisfies the requirement that answers remain grounded in the technical manuals and prevents hallucinated content, unlike guardrails focused solely on harmful categories or denied topics.

Why this answer

Bedrock Knowledge Bases provides source attribution, and when combined with model inference, the model can be instructed to answer only from the retrieved chunks. However, Bedrock Guardrails' grounding check specifically verifies that the model's response is supported by the retrieved context, reducing hallucination.

698
MCQeasy

A developer is using Amazon Bedrock's Claude model to summarize long documents. The developer notices that the summaries sometimes miss key points. Which parameter adjustment is most likely to improve summary completeness?

A.Increase the max_tokens parameter.
B.Increase the top_k parameter.
C.Increase the temperature parameter.
D.Increase the top_p parameter.
AnswerA

Truncation is the mechanism: if max_tokens is too low, the summary is cut off before covering all key points. Raising it allows the model to emit a complete summary, directly addressing the missed key points.

Why this answer

Increasing max_tokens allows the model to generate longer outputs, which is essential when summarizing long documents because the summary may need more tokens to capture all key points. If max_tokens is too low, the model truncates the response, potentially omitting important details. This directly addresses the issue of missing key points by providing sufficient output length for a complete summary.

Exam trap

The AIF-C01 exam often tests the misconception that parameters controlling randomness (temperature, top_k, top_p) affect output length or completeness, when in fact they only influence token selection diversity and creativity.

How to eliminate wrong answers

Option B is wrong because increasing top_k controls the number of highest-probability tokens considered during sampling, which affects randomness and diversity, not the length or completeness of the output. Option C is wrong because increasing temperature increases randomness in token selection, which can lead to more creative but less focused summaries, potentially worsening completeness. Option D is wrong because increasing top_p (nucleus sampling) also controls randomness by selecting tokens with cumulative probability, and does not extend the output length or guarantee inclusion of key points.

699
MCQhard

A company uses Amazon Bedrock with a Knowledge Base for RAG. Users report that the assistant gives incorrect answers for questions that require understanding of data tables. After reviewing, the team suspects the chunking strategy is breaking table structures. Which change would BEST preserve the integrity of tabular data?

A.Use a different vector store like Pinecone
B.Switch from fixed-size chunking to semantic chunking that respects table boundaries
C.Increase the chunk overlap to 50%
D.Decrease the chunk size to 100 tokens
AnswerB

Semantic chunking splits on meaning rather than character count, so table rows and headers stay within one chunk. Fixed-size splitting severs rows from headers, destroying the relationships the model needs, whereas boundary-aware chunking preserves tabular structure for retrieval.

Why this answer

Semantic chunking that respects table boundaries preserves the logical structure of tabular data by ensuring that rows, columns, and headers remain intact within a single chunk. Fixed-size chunking can split a table mid-row or mid-column, causing the knowledge base to retrieve incomplete or misaligned data, which leads to incorrect RAG answers. This approach directly addresses the root cause—broken table structures—without changing the vector store or overlap settings.

Exam trap

The trap here is that candidates often confuse vector store selection with chunking strategy, assuming a different database will magically fix data integrity issues, when in fact the chunking method directly controls how tabular data is preserved.

How to eliminate wrong answers

Option A is wrong because changing the vector store (e.g., to Pinecone) does not affect how chunks are created; the chunking strategy is independent of the vector database used. Option C is wrong because increasing chunk overlap to 50% still uses fixed-size boundaries and may duplicate broken table fragments, but does not prevent the initial splitting of table structures. Option D is wrong because decreasing chunk size to 100 tokens makes it more likely that tables are fragmented into even smaller, meaningless pieces, worsening the problem.

700
MCQhard

A healthcare company uses Amazon Bedrock with a foundation model to generate patient education materials. They must ensure that the model does not include protected health information (PHI) in its responses, even if it appears in the prompt. Which Amazon Bedrock feature should they configure?

A.Amazon Bedrock Agents with an action group to query a patient database
B.Amazon Bedrock provisioned throughput
C.Guardrails for Amazon Bedrock with sensitive information filters
D.Amazon Bedrock model evaluation with toxicity metrics
AnswerC

Guardrails for Amazon Bedrock includes sensitive information filters that can detect and block or mask personally identifiable information (PII) and custom regex patterns. By configuring these filters, the company can prevent PHI from appearing in model responses. This directly addresses the requirement to avoid including PHI in generated patient education materials.

Why this answer

Guardrails for Amazon Bedrock provides sensitive information filters that can detect and block or mask PII and custom patterns in both prompts and responses. Configuring these filters prevents PHI from appearing in generated materials. Model evaluation, provisioned throughput, and Agents do not offer runtime content redaction, so they cannot meet the compliance requirement.

Exam trap

The trap here is assuming that model evaluation or agent-based data retrieval can enforce privacy, when only runtime guardrail filters can detect and redact sensitive information in responses.

701
MCQmedium

A company is building a document summarization application using Amazon Bedrock. They want to prototype the application quickly by testing different models and prompts interactively. Which AWS service or feature should they use?

A.Amazon SageMaker Studio
B.Amazon Bedrock Playground
C.Amazon Bedrock Studio
D.AWS Cloud9
AnswerB

Amazon Bedrock Playground provides an interactive console for selecting foundation models, adjusting inference parameters and iterating on prompts without writing code, which directly satisfies the requirement to prototype rapidly and compare models and prompts interactively.

Why this answer

Amazon Bedrock Playground is the correct choice because it provides an interactive, no-code environment within the AWS Management Console specifically designed for experimenting with different foundation models and prompts. It allows you to quickly test and compare model outputs, adjust parameters, and iterate on prompts without writing any code or setting up infrastructure, making it ideal for rapid prototyping of a document summarization application.

Exam trap

Candidates may confuse 'Amazon Bedrock Studio', a collaborative development environment, with 'Amazon Bedrock Playground', which is the interactive, no-code console for testing models and prompts. Additionally, some might incorrectly choose general-purpose tools like SageMaker Studio or Cloud9, not realizing that the Playground is purpose-built for rapid prototyping with no setup required.

How to eliminate wrong answers

Option A is wrong because Amazon SageMaker Studio is a full-featured IDE for machine learning development, requiring you to write code and manage infrastructure, which is overkill for quick interactive prompt testing and not the purpose-built tool for Bedrock model experimentation. Option C is wrong because Amazon Bedrock Studio is not a real AWS service; the correct interactive environment is called Amazon Bedrock Playground. Option D is wrong because AWS Cloud9 is a cloud-based IDE for writing, running, and debugging code, and it does not provide the specialized, no-code interface for testing Bedrock models and prompts interactively.

702
MCQhard

A research team is using Amazon Bedrock to analyze scientific papers. They want the model to generate answers based only on papers published after 2023. Which approach should they use?

A.Fine-tune the model on a dataset of post-2023 papers and deploy it.
B.Set the maxTokens to a low value to force the model to rely on recent context.
C.Include a system prompt instructing the model to ignore data before 2023.
D.Use Amazon Bedrock Knowledge Bases with a metadata filter to retrieve only papers published after 2023, and generate responses based on retrieved content.
AnswerD

Amazon Bedrock Knowledge Bases with a metadata filter satisfies the post-2023 constraint by restricting vector retrieval to documents whose publication-date metadata matches the filter, so the foundation model generates answers grounded solely in that retrieved subset rather than its training data. This enforces the temporal restriction at retrieval time, preventing older papers from entering the context window.

Why this answer

Amazon Bedrock Knowledge Bases with a metadata filter allows you to restrict retrieval to only documents that match specific metadata criteria, such as publication year. By filtering the vector search to only include papers published after 2023, the model generates responses based solely on that retrieved content, ensuring it does not rely on pre-2023 data. This approach is the only one that guarantees the model's answers are grounded exclusively in the specified time range.

Exam trap

AWS often tests the misconception that a system prompt or fine-tuning can reliably restrict a model's knowledge to a specific time period, when in fact only a retrieval-based approach with metadata filtering can enforce such temporal constraints.

How to eliminate wrong answers

Option A is wrong because fine-tuning the model on a dataset of post-2023 papers does not prevent the model from using its pre-existing training data (which includes pre-2023 knowledge) during inference; fine-tuning adjusts weights but does not erase prior knowledge, so the model could still generate answers based on older information. Option B is wrong because setting maxTokens to a low value limits the length of the generated response but does not control the temporal scope of the model's knowledge; the model can still draw on pre-2023 training data regardless of token count. Option C is wrong because a system prompt instructing the model to ignore data before 2023 is merely a suggestion and not a technical enforcement; the model has no inherent mechanism to filter its own training data by date, so it may still generate answers based on pre-2023 information, especially if the prompt is not strictly followed.

703
MCQmedium

A financial services firm wants to deploy a generative AI application that answers customer questions about account balances and recent transactions. The firm has strict latency requirements (responses under 2 seconds) and wants to minimize costs. Which strategy for model selection and deployment is MOST appropriate?

A.Select a smaller, faster foundation model (e.g., Amazon Titan Text Lite) and use on-demand inference
B.Use the largest available foundation model via on-demand inference for highest accuracy
C.Fine-tune a large model specifically on account data and deploy on a dedicated endpoint
D.Deploy a large model using Provisioned Throughput to guarantee low latency
AnswerA

A smaller foundation model such as Amazon Titan Text Lite generates tokens faster, directly satisfying the sub-two-second latency constraint. On-demand inference avoids provisioning charges, minimising cost for variable customer traffic. Larger models would add latency and expense without improving simple balance and transaction lookups.

Why this answer

For latency-sensitive and cost-conscious applications, selecting a smaller, faster model is preferable over a large model or custom deployment. Provisioned Throughput for dedicated capacity would increase cost and may not be needed if the base model performs adequately.

704
Multi-Selectmedium

A company is using an LLM to generate customer support responses. They want to reduce hallucinations and improve the accuracy of the responses. Which TWO approaches are most effective? (Select TWO.)

Select 2 answers
A.Increase the temperature parameter to 1.0
B.Use a smaller model to reduce complexity
C.Remove all system prompts
D.Apply Bedrock Guardrails with contextual grounding check
E.Use Retrieval-Augmented Generation (RAG) to retrieve relevant documents
AnswersD, E

Contextual grounding checks compare each generated response against the retrieved source passages, filtering or blocking output that is unsupported by that reference material. This directly targets hallucination by enforcing evidential grounding, satisfying the requirement to improve response accuracy.

Why this answer

RAG grounds the model in retrieved facts, while Bedrock Guardrails with contextual grounding check validates responses against sources. Both are proven techniques to reduce hallucinations.

705
MCQhard

A company is building a RAG application that indexes thousands of PDF documents. They notice that some documents are very long (hundreds of pages) and the vector search often returns irrelevant chunks. Which configuration change would MOST improve retrieval relevance?

A.Switch from Amazon OpenSearch Serverless to Pinecone
B.Increase the embedding dimension from 1024 to 4096
C.Use a larger, more capable foundation model for response generation
D.Adjust the chunk size and overlap to better capture context from the documents
AnswerD

Long PDFs produce chunks that mix unrelated topics, so embeddings drift and retrieval returns irrelevant passages. Reducing chunk size with sensible overlap keeps each vector semantically focused, directly improving relevance without changing the embedding model or index.

Why this answer

Chunk size and overlap directly control how much semantic context each embedded vector carries. When documents are hundreds of pages long, a fixed chunk size (e.g., 512 tokens) can split related sentences across chunks, producing vectors that represent fragments rather than complete ideas. Tuning chunk size and overlap ensures each chunk is semantically self-contained, which is the single most impactful lever for retrieval precision in a RAG pipeline.

Exam trap

AIF-C01 often tests the misconception that retrieval quality is improved by upgrading the model or vector database, when the actual root cause is usually data preparation — chunking, embedding model choice, or metadata filtering.

How to eliminate wrong answers

Option A is wrong because switching vector stores (OpenSearch Serverless to Pinecone) changes the infrastructure, not the embedding quality — the same poorly-chunked vectors would be indexed either way. Option B is wrong because increasing embedding dimensions from 1024 to 4096 does not fix semantic fragmentation; it only changes vector granularity and increases cost, and most models are trained for a fixed dimension. Option C is wrong because the foundation model only generates the final answer from retrieved context — it has no effect on which chunks are retrieved in the first place.

706
MCQmedium

A machine learning team wants to detect bias in a deployed model's predictions on new data. They use Amazon SageMaker. Which service should they use to generate bias reports after deployment?

A.Amazon SageMaker Clarify
B.Amazon SageMaker Debugger
C.Amazon SageMaker Model Monitor
D.Amazon SageMaker Role Manager
AnswerA

Amazon SageMaker Clarify generates post-deployment bias reports by monitoring live endpoint traffic and computing metrics such as disparate impact against baseline data. This satisfies the stem's requirement to detect bias in predictions on new data after deployment, which pre-training Clarify analysis alone cannot cover.

Why this answer

Amazon SageMaker Clarify provides bias detection and explainability for ML models, both during training and after deployment. SageMaker Model Monitor detects data drift but not bias. SageMaker Debugger is for training debugging.

SageMaker Role Manager is for managing IAM roles.

707
MCQeasy

Which AWS service provides a serverless API for accessing foundation models with per-token pricing?

A.Amazon Bedrock
B.Amazon API Gateway
C.AWS Lambda
D.Amazon SageMaker
AnswerA

Amazon Bedrock offers a serverless, API-driven way to invoke foundation models from multiple providers, with usage billed per input and output token. That matches the stem's serverless API and per-token pricing constraints without managing infrastructure.

Why this answer

Amazon Bedrock is a fully managed service that provides a serverless API for accessing foundation models (FMs) from providers like AI21 Labs, Anthropic, Cohere, Meta, and Stability AI. It offers per-token pricing, meaning you pay only for the number of tokens processed in both input and output, with no upfront commitments or infrastructure management required.

Exam trap

The trap here is that candidates often confuse Amazon API Gateway (a serverless API front-end) with Bedrock's serverless model inference API, or mistakenly think AWS Lambda provides built-in FM access, when in fact Lambda is just compute and requires explicit integration with a model service.

How to eliminate wrong answers

Option B is wrong because Amazon API Gateway is a managed service for creating, publishing, and securing RESTful and WebSocket APIs, but it does not provide access to foundation models or per-token pricing; it is a front-end API layer that would need to be integrated with a backend service like Bedrock or Lambda. Option C is wrong because AWS Lambda is a serverless compute service that runs code in response to events, but it does not natively provide access to foundation models or per-token pricing; you would need to write custom code to call an FM API, and you pay per invocation and duration, not per token. Option D is wrong because Amazon SageMaker is a fully managed machine learning platform for building, training, and deploying custom models, but it is not a serverless API for foundation models with per-token pricing; it typically involves provisioning instances and paying for compute time, not per-token consumption.

708
MCQhard

A security engineer is configuring logging for Amazon Bedrock model invocations. They need to capture both the input and output of all API calls for compliance audits. Which set of steps should they take?

A.Enable Bedrock model invocation logging and specify an S3 bucket and optionally CloudWatch Logs as the destination
B.Enable CloudTrail for the Bedrock API and configure S3 event notifications
C.Use VPC Flow Logs to capture network traffic and reconstruct model inputs from packet data
D.Enable AWS Config rules for Bedrock and stream logs to Amazon Kinesis
AnswerA

Enabling model invocation logging in Bedrock and designating an S3 bucket (with optional CloudWatch Logs) captures both the request input and the model output for every API call, directly meeting the compliance audit requirement to record full invocation content.

Why this answer

Bedrock model invocation logging captures inputs and outputs to S3 and/or CloudWatch Logs. CloudTrail records API calls but not the model inputs/outputs.

709
MCQmedium

An IAM policy allows creation of SageMaker training jobs only if they use a specific VPC security group. A user tries to create a training job without specifying that security group. What will happen?

A.The request will succeed but SageMaker will ignore the condition
B.The request will succeed because the condition is optional
C.The request will be denied because the training job resource ARN is invalid
D.The request will be denied with an AccessDenied error
AnswerD

The IAM condition is not satisfied, so the request is denied.

Why this answer

IAM policies are evaluated before any AWS API action is executed. If the policy includes a condition that requires a specific VPC security group for SageMaker training jobs, and the user's request does not include that security group, the condition is not met, resulting in an explicit deny (AccessDenied error). AWS IAM denies the request by default if the condition in a policy is not satisfied, regardless of whether the condition is marked as optional in the API.

Exam trap

The trap here is that candidates assume an optional API parameter means the IAM condition is also optional, but IAM conditions are strictly enforced regardless of whether the parameter is required by the API.

How to eliminate wrong answers

Option A is wrong because IAM policies do not ignore conditions; if a condition is not met, the request is denied, not silently ignored. Option B is wrong because the condition is not optional from an IAM perspective; even if the API parameter is optional, the IAM policy condition must be satisfied for the request to be allowed. Option C is wrong because the training job resource ARN is not invalid; the request is denied due to the policy condition, not due to an ARN format issue.

710
MCQhard

A company has built a RAG application using Amazon Bedrock Knowledge Bases. Users report that answers are sometimes based on irrelevant or incorrect document chunks. The team has verified that the embedding model is appropriate and the documents are correctly indexed. What is the MOST likely cause of the poor retrieval quality?

A.The foundation model is too small
B.The prompt template is missing instructions
C.The vector store is too slow
D.The chunking strategy is suboptimal
AnswerD

Suboptimal chunking splits documents so that semantically related content is fragmented or unrelated text is merged, degrading embedding quality and returning irrelevant chunks. Since embeddings and indexing are verified correct, chunk boundaries are the remaining retrieval-quality cause.

Why this answer

Chunking strategy directly affects retrieval relevance. If chunks are too large, they may contain irrelevant information; if too small, they may miss context. Overlap size also matters.

Optimizing chunking often fixes relevance issues when embeddings are correct.

711
MCQeasy

Refer to the exhibit. A developer is reviewing CloudWatch Logs for a deployed model and notices the same input appears multiple times with slightly different probabilities. What responsible AI concern does this pattern suggest?

A.The model is overfitting to the training data.
B.The model is not robust; it produces inconsistent predictions for the same input.
C.The model is exhibiting bias against a demographic group.
D.The input data is drifting from the training distribution.
AnswerB

Inconsistent probabilities for identical inputs indicate the model lacks robustness, directly matching the stem's repeated-input pattern. Non-determinism at inference, whether from sampling, dropout, or unstable weights, breaches the reliability principle of responsible AI, which expects stable, reproducible outputs for the same input.

Why this answer

The pattern of the same input producing slightly different probabilities indicates that the model's predictions are not deterministic for identical inputs. This violates the principle of robustness in responsible AI, which requires that a model should produce consistent outputs for the same input under the same conditions. Inconsistent predictions for identical inputs undermine trust and reliability, making the model non-robust.

Exam trap

The AWS AI Practitioner exam often tests the distinction between robustness (consistency for the same input) and other AI concerns like bias or drift, so the trap here is confusing non-deterministic output with data drift or overfitting.

How to eliminate wrong answers

Option A is wrong because overfitting refers to a model memorizing training data noise and performing poorly on unseen data, not to inconsistent predictions for the same input. Option C is wrong because bias against a demographic group would manifest as systematic differences in predictions across groups, not as random variation for identical inputs. Option D is wrong because input data drift describes a change in the distribution of incoming features over time, which would affect predictions for different inputs, not cause inconsistent outputs for the same input.

712
MCQmedium

A company uses Amazon Bedrock to generate marketing copy. They want to measure the quality of generated text compared to reference text. Which metric is most appropriate?

A.F1 score
B.BLEU
C.RMSE
D.Accuracy
AnswerB

BLEU compares generated text against reference text using n-gram overlap, producing a precision-based score. It suits marketing copy evaluation where reference outputs exist. Perplexity, by contrast, measures language model likelihood without references, so it cannot assess similarity to target text.

Why this answer

BLEU (Bilingual Evaluation Understudy) is the most appropriate metric for evaluating the quality of generated text against reference text in tasks like machine translation and text generation. It measures n-gram precision between the generated and reference texts, making it ideal for assessing marketing copy generated by Amazon Bedrock.

Exam trap

AWS often tests the distinction between classification/regression metrics and text generation metrics, leading candidates to mistakenly apply F1 score or accuracy to evaluate generated text quality instead of using BLEU or similar sequence-based metrics.

How to eliminate wrong answers

Option A is wrong because F1 score is a classification metric that measures harmonic mean of precision and recall, not suitable for evaluating text generation quality against reference text. Option C is wrong because RMSE (Root Mean Square Error) is a regression metric used for continuous numerical predictions, not for text or sequence evaluation. Option D is wrong because Accuracy is a classification metric that measures the proportion of correct predictions, which does not account for the sequential and linguistic nuances of generated text.

713
MCQmedium

A financial services firm must build a model that flags potentially fraudulent card transactions in under 200 milliseconds while keeping all data inside its own Amazon VPC. The fraud team has thousands of labeled historical transactions and the pattern changes slowly over months. Which approach best balances latency, data residency, and the need for periodic retraining?

A.Train a supervised classification model with Amazon SageMaker, deploy it to a real-time endpoint inside the VPC, and schedule periodic retraining jobs.
B.Deploy a pre-trained foundation model from Amazon Bedrock and prompt it to classify each transaction.
C.Build an unsupervised anomaly detection model on unlabeled data and run batch transform once per day.
D.Use Amazon Fraud Detector in evaluation mode and export predictions to an S3 bucket outside the VPC for scoring.
AnswerA

Fraud flagging with thousands of labeled transactions is a supervised classification problem, and a SageMaker real-time endpoint keeps inference within the VPC at low millisecond latency. Scheduled retraining jobs let the model adapt as fraud patterns drift over months. This combination satisfies the latency, residency, and retraining requirements directly without introducing external data movement.

Why this answer

A supervised classification model trained on the labeled history and deployed to a SageMaker real-time endpoint inside the VPC meets the latency and residency constraints, while scheduled retraining handles gradual fraud pattern drift. Purpose-built supervised learning uses the available labels, and in-VPC endpoints keep transaction data within the controlled network. Batch, unsupervised, or general foundation-model approaches each miss at least one hard requirement.

Exam trap

The trap here is treating a managed fraud service or foundation model as automatically better than a purpose-built supervised model that meets the stated latency and residency constraints.

714
MCQhard

A data scientist is using Amazon Bedrock to generate product descriptions. They notice the output is often repetitive and lacks creativity. Which combination of parameter adjustments is MOST likely to produce more diverse and less repetitive output?

A.Decrease temperature and decrease top-p
B.Decrease temperature and increase top-p
C.Increase temperature and decrease top-p
D.Increase temperature and increase top-p
AnswerD

Temperature scales the sampling distribution's randomness, while top-p restricts sampling to the smallest token set whose cumulative probability exceeds p. Raising both widens the candidate pool and flattens selection, producing more diverse, less repetitive product descriptions.

Why this answer

Increasing temperature raises the probability of sampling lower-probability tokens, which increases randomness and diversity. Increasing top-p (nucleus sampling) expands the set of tokens considered for sampling, further reducing repetitiveness. Together, these adjustments encourage the model to explore a wider range of possible continuations, producing more creative and less repetitive output.

Exam trap

AWS often tests the misconception that increasing temperature alone is sufficient for diversity, but candidates forget that top-p must also be increased to avoid the model repeatedly sampling from a narrow set of high-probability tokens.

How to eliminate wrong answers

Option A is wrong because decreasing both temperature and top-p makes the model more deterministic and focused on the highest-probability tokens, which increases repetitiveness and reduces creativity. Option B is wrong because decreasing temperature while increasing top-p partially counteracts the effect: lower temperature narrows token probabilities, so even with a larger top-p set, the model still tends to pick the same high-probability tokens, limiting diversity. Option C is wrong because increasing temperature but decreasing top-p restricts the sampling pool to only the most probable tokens, which can still lead to repetitive patterns despite higher randomness within that narrow set.

715
MCQhard

A company wants to forecast product demand across thousands of SKUs with different demand patterns. They have 3 years of historical sales data, plus external factors like holidays and promotions. Which combination of AWS services and approach would deliver the most accurate forecasts with minimal manual effort?

A.Use Amazon Comprehend to analyze customer reviews and correlate with sales
B.Upload data to Amazon QuickSight and use its built-in forecasting widget
C.Use Amazon Forecast with the DeepAR+ algorithm and provide item metadata, holiday calendars, and promotion data
D.Train individual ARIMA models for each SKU using Amazon SageMaker built-in algorithms
AnswerC

DeepAR+ is a supervised recurrent neural network that learns from many related time series simultaneously, so thousands of SKUs train one model. Built-in holiday calendars and promotion metadata let it capture external demand drivers, while Amazon Forecast automates training and tuning, minimising manual effort.

Why this answer

Amazon Forecast is purpose-built for time-series forecasting and automatically handles multiple SKUs, holidays, and promotions. SageMaker would require building custom models from scratch. Comprehend is NLP, and QuickSight is for visualization.

716
Multi-Selecthard

A company wants to evaluate the performance of a generative AI model before deployment. Which TWO metrics are most relevant for measuring model quality? (Select two.)

Select 2 answers
A.BLEU score
B.Response time
C.Perplexity
D.Model size
E.CPU utilization
AnswersA, C

BLEU compares n-gram overlap between generated and reference text, directly quantifying translation and summarisation fidelity. It satisfies the stem's need to measure generative output quality against expected answers, unlike latency or cost metrics that assess operational rather than quality performance.

Why this answer

BLEU score (A) is correct because it measures the quality of generated text by comparing n-gram overlap between model output and reference text, making it a standard metric for evaluating generative AI models such as translation and summarization systems. Perplexity (C) is correct because it quantifies how well a language model predicts a sample of text, with lower perplexity indicating better model confidence and language modeling quality. Response time (B) is not a quality metric but a latency/performance measure, model size (D) reflects resource footprint rather than output quality, and CPU utilization (E) is an infrastructure efficiency metric unrelated to the model's generative quality.

Exam trap

AWS exams often test the distinction between model quality metrics (like BLEU and perplexity) and operational or performance metrics (like response time or resource utilization), leading candidates to mistakenly select speed or size as relevant for quality assessment.

717
Multi-Selectmedium

A retail bank is preparing to launch an Amazon Bedrock-based assistant that recommends credit card products to customers. The responsible AI review board requires the team to document how the system's outputs can be explained to regulators and customers. Which TWO actions best support explainability for this deployment? (Choose two.)

Select 2 answers
A.Increase the model's temperature setting so recommendations show more variety across customers
B.Enable verbose AWS CloudTrail logging of every InvokeModel API call for the assistant
C.Store the retrieved source documents and the model's cited references alongside each generated recommendation for later review
D.Reduce the number of credit card products in the catalog so the model has fewer options to choose from
E.Provide a customer-facing disclosure that the recommendation is generated by AI and how to request a human review
AnswersC, E

Capturing the grounding documents and citations creates an auditable trail showing why a specific product was recommended. When a regulator or customer asks how a recommendation was produced, the team can reproduce the evidence the model used. This directly supports explainability because the reasoning basis is preserved rather than lost after inference, and it aligns with transparency expectations for consequential financial recommendations.

Why this answer

Explainability for a regulated recommendation assistant requires preserving the evidence behind each output and communicating the automated nature of the decision to affected customers. Retaining retrieved sources and citations makes reasoning reproducible during audits, while AI disclosure with a human-review path gives customers transparency and contestability. Logging and catalog changes do not expose reasoning.

Exam trap

The trap here is equating operational logging or output variety with explainability, when explainability specifically requires capturing and communicating the basis for a decision.

718
MCQhard

A bank uses an AI system to detect fraudulent transactions. The model has high precision but low recall for small transactions, potentially missing fraud. Which approach aligns with responsible AI?

A.Send all flagged transactions to customers for confirmation
B.Focus only on precision to minimize false positives
C.Tune the model to achieve an acceptable balance between recall and precision
D.Increase the detection threshold to reduce false positives
AnswerC

Tuning to balance recall and precision directly addresses the low recall on small transactions, reducing missed fraud while keeping false positives acceptable. This aligns with responsible AI by mitigating harm from undetected fraud rather than optimising one metric alone.

Why this answer

Responsible AI requires balancing competing objectives like precision and recall to align with ethical principles and business needs. In fraud detection, high precision with low recall means many fraudulent transactions are missed, which can lead to significant financial losses and erode customer trust. Tuning the model to achieve an acceptable trade-off ensures that the system is both effective and fair, minimizing harm while maintaining operational viability.

Exam trap

The AIF-C01 exam often tests the misconception that increasing the detection threshold improves model performance overall, when in fact it only reduces false positives at the cost of lowering recall, which can be detrimental in high-stakes applications like fraud detection.

How to eliminate wrong answers

Option A is wrong because sending all flagged transactions to customers for confirmation shifts the burden to users, degrades user experience, and may not be scalable or timely for real-time fraud detection, nor does it address the underlying model imbalance. Option B is wrong because focusing only on precision ignores the critical need to catch actual fraud (recall), which can result in substantial financial losses and violates the responsible AI principle of beneficence. Option D is wrong because increasing the detection threshold reduces false positives but further lowers recall, worsening the problem of missed fraud and contradicting the goal of responsible AI.

719
MCQmedium

A retail company runs a product-question answering feature on Amazon Bedrock. During peak hours, requests intermittently fail with a ThrottlingException even though average usage is well within quota. The team needs a solution that smooths bursty traffic, retries failed calls, and avoids overwhelming the model endpoint, with minimal application code changes. Which approach should they take?

A.Introduce client-side retries with exponential backoff and jitter around the Amazon Bedrock InvokeModel calls.
B.Switch the application to call the model through Amazon Bedrock Provisioned Throughput and remove all retry logic.
C.Place an Amazon SQS queue between the application and Amazon Bedrock, and have a worker invoke the model as messages arrive.
D.Increase the maximum token count in each request so fewer total requests are needed during peak periods.
AnswerA

Exponential backoff with jitter retries throttled requests after progressively longer, randomized delays, which spreads retries out and prevents synchronized retry storms. It is a small, well-understood code change that directly addresses transient ThrottlingException errors during bursts. AWS SDKs and the Bedrock runtime support configurable retry behaviour, making this the lowest-effort effective fix.

Why this answer

Bursty traffic that briefly exceeds per-account or per-model quotas produces throttling that is transient by nature. Retrying with exponential backoff and jitter lets the client back off and spread retries, converting failures into eventual successes without architectural change. Provisioned Throughput, queues, and larger token limits either cost more, add latency, or do not address retry behaviour, so backoff and jitter is the targeted remedy.

Exam trap

The trap here is treating throttling as a hard capacity shortage that requires Provisioned Throughput, when bursty transient throttling is usually resolved with backoff and jitter retries.

720
MCQmedium

An organization needs to generate high-quality images from text prompts for a marketing campaign. They require the ability to edit specific regions of an image (inpainting) and extend images beyond their original boundaries (outpainting). Which AWS service or model should they choose?

A.Stable Diffusion XL via Amazon Bedrock
B.Amazon Rekognition
C.Amazon Titan Image Generator
D.Amazon SageMaker JumpStart with a custom GAN
AnswerC

Titan Image Generator includes features for inpainting and outpainting.

Why this answer

Amazon Titan Image Generator is the correct choice because it natively supports both inpainting (editing specific regions of an image) and outpainting (extending images beyond their original boundaries) through its image conditioning capabilities. This service is specifically designed for generative image tasks like text-to-image generation, inpainting, and outpainting, making it ideal for the marketing campaign's requirements.

Exam trap

The trap is that candidates may confuse Amazon Rekognition (an image analysis service) with a generative AI service, or assume that any text-to-image model like Stable Diffusion XL inherently supports inpainting/outpainting, when in fact Amazon Titan Image Generator is the AWS-managed service that offers these features.

How to eliminate wrong answers

Option A is wrong because Stable Diffusion XL via Amazon Bedrock is a text-to-image model that does not natively support inpainting or outpainting; while it can be used for image generation, it lacks the built-in region-specific editing and boundary extension features that Amazon Titan Image Generator provides. Option B is wrong because Amazon Rekognition is a computer vision service for image and video analysis (e.g., object detection, facial recognition, content moderation), not a generative AI model for creating or editing images from text prompts. Option D is wrong because Amazon SageMaker JumpStart with a custom GAN requires significant custom development and training to implement inpainting and outpainting capabilities, whereas Amazon Titan Image Generator offers these features out-of-the-box without the need for custom model building.

721
MCQeasy

A company wants to use a foundation model to automatically summarize lengthy documents. Which capability of foundation models is being utilized?

A.Text generation
B.Sentiment analysis
C.Text classification
D.Machine translation
AnswerA

Summarisation is a text-generation task: the model consumes the source document as input and autoregressively produces a condensed natural-language output. This directly satisfies the stem's requirement to automatically summarise lengthy documents, since the capability being exercised is generating new text rather than classification, embedding or retrieval.

Why this answer

Summarization is a text generation task where the model produces a concise version of the original content. Foundation models (e.g., GPT, Claude) are pre-trained on vast corpora and can generate coherent summaries by predicting the next tokens conditioned on the input document. This directly utilizes the text generation capability, not classification or translation.

Exam trap

The AIF-C01 exam often tests the distinction between text generation and text classification, so the trap here is that candidates may confuse summarization (a generative task) with classification or analysis tasks, especially when the question emphasizes 'understanding' the document rather than 'producing' new text.

How to eliminate wrong answers

Option B (Sentiment analysis) is wrong because it involves classifying the emotional tone of text (positive, negative, neutral), not generating a summary. Option C (Text classification) is wrong because it assigns predefined labels or categories to text, whereas summarization requires generating new text. Option D (Machine translation) is wrong because it converts text from one language to another, not condensing content within the same language.

722
MCQmedium

Refer to the exhibit. An AWS CloudTrail log shows the creation of an IAM policy for a SageMaker execution role. Which responsible AI concern does this configuration raise?

A.Insufficient training data
B.Lack of least privilege access control
C.Violation of data residency requirements
D.Absence of model monitoring
AnswerB

The created IAM policy grants broader permissions than the SageMaker execution role requires, violating least privilege. Over-scoped role permissions expand the blast radius if the role is compromised, which is the responsible AI access-control concern raised by this CloudTrail configuration.

Why this answer

The CloudTrail log shows the creation of an IAM policy for a SageMaker execution role. If this policy grants overly broad permissions (e.g., `s3:*` or `iam:PassRole` to all resources), it violates the principle of least privilege, which is a core responsible AI concern. Overly permissive roles can lead to unauthorized access to training data, models, or other AWS resources, undermining security and governance.

Exam trap

AWS often tests the distinction between operational security concerns (like least privilege) and other responsible AI pillars (like data governance or model monitoring), so candidates may confuse a broad IAM policy with a data residency or monitoring issue.

How to eliminate wrong answers

Option A is wrong because insufficient training data is a data quality or quantity issue, not a security or access control concern raised by an IAM policy creation event. Option C is wrong because violation of data residency requirements relates to where data is stored or processed (e.g., cross-region transfers), not to the permissions granted in an IAM policy. Option D is wrong because absence of model monitoring refers to the lack of ongoing tracking of model performance or bias, which is not directly indicated by the creation of an IAM policy for an execution role.

723
Multi-Selecteasy

Which TWO of the following are types of unsupervised learning? (Select TWO.)

Select 2 answers
A.Classification
B.Dimensionality reduction
C.Clustering
D.Reinforcement learning
E.Regression
AnswersB, C

Dimensionality reduction is unsupervised because it finds structure in unlabelled data, projecting high-dimensional inputs onto fewer components without target labels. Techniques such as principal component analysis learn patterns purely from feature distributions, unlike supervised classification or regression.

Why this answer

Dimensionality reduction (B) is a type of unsupervised learning because it seeks to compress or transform unlabeled data into a lower-dimensional representation (e.g., via PCA or t-SNE) without any target labels. Clustering (C) is also unsupervised, as algorithms like k-means or DBSCAN group unlabeled data points by similarity without predefined classes. Classification (A) and regression (E) are supervised learning tasks, since they require labeled training data with discrete or continuous target values, respectively.

Reinforcement learning (D) is a separate paradigm in which an agent learns from reward signals via interaction with an environment, not from unlabeled data in the unsupervised sense.

Exam trap

The AWS AI Practitioner exam often tests the distinction between supervised and unsupervised learning by presenting classification and regression as plausible unsupervised options, exploiting the common misconception that any 'grouping' or 'reduction' task is unsupervised, while in fact classification and regression require labeled data.

724
MCQeasy

A startup needs to predict customer churn based on historical data containing labels (churned or not). Which type of machine learning should they use?

A.Reinforcement learning
B.Unsupervised learning
C.Supervised learning
D.Semi-supervised learning
AnswerC

Historical data already contains the churn outcome as a label, so the model learns a mapping from input features to that known target. Supervised learning is defined by training on labelled examples, satisfying the labelled churned/not-churned constraint.

Why this answer

The startup has labeled historical data (churned or not), which is the defining characteristic of supervised learning. The goal is to learn a mapping from input features to the known output labels to predict churn for new customers. This is a classic classification problem, making supervised learning the correct choice.

Exam trap

The AIF-C01 exam often tests the distinction between supervised and unsupervised learning by presenting a scenario with labeled data, where candidates might mistakenly choose unsupervised learning if they overlook the presence of labels.

How to eliminate wrong answers

Option A is wrong because reinforcement learning involves an agent learning through trial-and-error interactions with an environment to maximize cumulative reward, not from labeled historical data. Option B is wrong because unsupervised learning finds hidden patterns or structures in unlabeled data, but here the labels (churned/not) are explicitly provided. Option D is wrong because semi-supervised learning uses a small amount of labeled data with a large amount of unlabeled data, but the problem states the historical data contains labels, implying fully labeled data is available.

725
MCQhard

A data science team is comparing two approaches for customizing a foundation model in Amazon Bedrock for a domain-specific classification task. Approach one is providing a small set of labeled examples directly in the prompt for each request. Approach two is fine-tuning the model on a larger labeled dataset. The team wants the lowest operational overhead and the fastest way to start, and their label set is small and changes frequently. Which statement best describes the trade-off they should consider?

A.Prompt-based examples offer lower operational overhead and adapt quickly to changing labels, while fine-tuning can improve consistency when a larger stable dataset is available.
B.Prompt-based examples require a dedicated training job in Amazon Bedrock, so they carry the same operational overhead as fine-tuning.
C.Fine-tuning eliminates the need for any labeled data because the model learns the task from unlabeled inputs automatically.
D.Fine-tuning is preferred because it always produces higher accuracy than prompt-based examples regardless of dataset size.
AnswerA

Supplying labeled examples in the prompt, often called few-shot prompting, requires no training job and can be updated instantly as labels change. Fine-tuning is better suited when a larger, stable labeled dataset exists and consistent behavior matters. This statement correctly frames the trade-off given the team's small, changing label set.

Why this answer

The team's constraints are low operational overhead, fast start, and a small label set that changes often. Few-shot prompting embeds labeled examples in each request, needs no training job, and updates instantly. Fine-tuning is more appropriate for larger, stable datasets where consistent behavior is worth the training and maintenance effort, so it is not the best fit here.

Exam trap

The trap here is assuming fine-tuning is always superior for customization, when small or rapidly changing label sets make prompt-based examples the lower-overhead and more adaptable choice.

726
MCQmedium

A company is building a search application that retrieves relevant documents based on semantic meaning rather than exact keyword matches. Which combination of services would BEST enable this capability?

A.Amazon Titan Embeddings model and a vector database
B.Amazon Bedrock and Amazon Kendra
C.Amazon Lex and Amazon ElastiCache
D.Amazon Comprehend and Amazon DynamoDB
AnswerA

Amazon Titan Embeddings converts documents and queries into dense vectors, and a vector database performs approximate nearest-neighbour similarity search over them. This pairing satisfies the stem's semantic-meaning requirement, retrieving conceptually related documents even when no keywords overlap, unlike lexical matching such as BM25 or OpenSearch keyword queries.

Why this answer

Amazon Titan Embeddings converts text into dense vector representations that capture semantic meaning, and storing these vectors in a vector database (e.g., Amazon OpenSearch Serverless with vector engine or Aurora PostgreSQL with pgvector) enables similarity search based on cosine distance or Euclidean distance. This combination retrieves documents by semantic proximity rather than exact keyword matches, which is the core requirement for semantic search.

Exam trap

The trap here is that candidates may confuse Amazon Kendra's built-in semantic search capabilities with the need to generate and query custom vector embeddings, or assume that any NLP service like Comprehend can produce embeddings suitable for similarity search, when in fact only embedding models and vector databases provide the precise semantic retrieval pipeline described.

How to eliminate wrong answers

Option B is wrong because Amazon Bedrock is a service for accessing foundation models, and Amazon Kendra is an intelligent search service that uses keyword and semantic search but does not natively support storing and querying custom vector embeddings for user-defined semantic retrieval; it is designed for out-of-the-box enterprise search with its own indexing. Option C is wrong because Amazon Lex is a conversational AI service for building chatbots, and Amazon ElastiCache is an in-memory caching service; neither provides vector embedding generation or vector similarity search for semantic document retrieval. Option D is wrong because Amazon Comprehend is a natural language processing service for extracting entities and sentiment, not for generating dense vector embeddings, and Amazon DynamoDB is a NoSQL key-value and document database that lacks native vector similarity search capabilities (unless combined with a separate vector index, which is not the intended combination).

727
Multi-Selectmedium

A data scientist has trained a random forest model that achieves 92% accuracy on the training set but only 75% on the test set. The dataset has 1000 samples and 20 features. Which THREE actions could help improve the model's generalization? (Select THREE.)

Select 3 answers
A.Increase the number of trees in the forest
B.Increase the maximum depth of each tree
C.Reduce the maximum depth of each tree
D.Decrease the number of trees in the forest
E.Increase the minimum number of samples required to split an internal node
AnswersA, C, E

Adding more trees reduces variance in the ensemble's averaged predictions, which typically improves generalisation on unseen data. This addresses the 92% versus 75% train-test gap, satisfying the scenario's requirement for an action that improves the random forest model's generalisation.

Why this answer

The scenario describes overfitting: a large gap between training accuracy (92%) and test accuracy (75%). Option A (increase the number of trees in the forest) is correct because adding trees to a random forest averages more decorrelated trees, reducing variance without increasing overfitting, which typically improves test-set generalization. Option C (reduce the maximum depth of each tree) is correct because shallower trees have lower variance and cannot memorize training noise as easily, directly countering the overfitting.

Option E (increase the minimum number of samples required to split an internal node) is correct because requiring more samples per split (e.g., raising min_samples_split) forces splits to be based on more evidence, producing simpler, more generalizable trees. Option B (increase the maximum depth) is wrong because deeper trees increase variance and worsen overfitting. Option D (decrease the number of trees) is wrong because fewer trees reduce the averaging benefit and can increase variance, hurting generalization.

Exam trap

A common mistake is assuming that increasing model complexity always improves accuracy, but overfitting requires regularization such as reducing tree depth or increasing minimum split samples.

728
MCQeasy

A media company uses Amazon Rekognition to automatically moderate user-uploaded images on its platform. The moderation team reports that some images containing nudity are being approved, while harmless images of sculptures are being rejected. The company wants to review the specific labels and confidence scores that Rekognition assigned to each image before deciding whether to appeal. Which action should the team take to obtain this information?

A.Enable AWS CloudTrail data events on the Amazon Rekognition API to capture label details
B.Use Amazon Augmented AI (A2I) to route every image to human reviewers before moderation
C.Configure Amazon Rekognition Custom Labels to train a new moderation model on the rejected images
D.Enable Amazon Rekognition content moderation and inspect the ModerationLabels returned in the DetectModerationLabels response
AnswerD

Amazon Rekognition's DetectModerationLabels API returns ModerationLabels with a name and a confidence score for each detected category, such as Explicit Nudity or Suggestive. Reviewing these labels and scores lets the team see exactly why an image was approved or rejected and supports a documented appeal process. This directly addresses the need to inspect per-image classification details.

Why this answer

DetectModerationLabels returns a ModerationLabels list where each entry includes a category name and a confidence score, giving the moderation team the exact evidence behind each decision. This supports appeals and threshold tuning. Custom Labels would replace the classifier, A2I adds human review rather than exposing existing scores, and CloudTrail logs API activity without capturing label content.

Exam trap

The trap here is confusing API activity logging in AWS CloudTrail with the actual moderation label output, leading teams to expect CloudTrail to contain confidence scores it never records.

729
Multi-Selectmedium

A company is building a RAG solution using Amazon Bedrock Knowledge Bases. Which TWO steps are essential in the document ingestion pipeline? (Select TWO.)

Select 2 answers
A.Setting up a model endpoint for real-time inference
B.Creating a Bedrock Guardrail for the documents
C.Generating embeddings for each chunk
D.Chunking documents into smaller segments
E.Fine-tuning the model on the documents
AnswersC, D

Embeddings convert each chunk into vectors that capture semantic meaning, enabling similarity search against the user query. Without this vectorisation step, Amazon Bedrock Knowledge Bases cannot retrieve relevant passages, so the RAG pipeline fails at its core retrieval stage.

Why this answer

Option D is correct because chunking documents into smaller segments is a required preprocessing step in a Bedrock Knowledge Bases ingestion pipeline; splitting source documents into manageable chunks allows the system to process and index content effectively. Option C is correct because generating embeddings for each chunk is essential: Bedrock invokes an embedding model (for example, Amazon Titan Embeddings or Cohere Embed) to convert each chunk into a vector, which is then stored in the vector store for semantic retrieval during RAG queries. Option A is not essential to the ingestion pipeline because a model endpoint for real-time inference relates to query-time generation, not document ingestion.

Option B is incorrect because Bedrock Guardrails apply content filtering and safety policies at inference time and are not a required ingestion step. Option E is incorrect because fine-tuning the model on the documents is a separate model-customization activity and is not part of the Knowledge Bases ingestion pipeline, which relies on embeddings and vector storage rather than fine-tuning.

Exam trap

AWS often tests the distinction between the ingestion pipeline (chunking and embedding) and the inference pipeline (model endpoints, guardrails, and fine-tuning), so candidates mistakenly select options that belong to the query or model customization phase.

730
MCQeasy

A media company wants to create an internal tool that drafts short news summaries from press releases. They have no machine learning engineers, no labeled training data, and need a working prototype within two weeks. Which approach best describes how they should build this capability using generative AI concepts on AWS?

A.Use a pretrained foundation model through an API and steer its behavior with prompt engineering and few-shot examples.
B.Build a rules-based extractive summarizer that selects the first and last sentence of each press release.
C.Deploy an Amazon Rekognition custom labels project and post-process its output into summary sentences.
D.Train a small transformer model from scratch on the company's archive of press releases using Amazon SageMaker training jobs.
AnswerA

A foundation model already encodes broad language understanding from large-scale pretraining, so the team only needs to supply task instructions and a handful of example summaries in the prompt. This delivers a working prototype quickly without labeled data, GPUs, or model training, which matches the two-week constraint and the absence of ML staff.

Why this answer

Foundation models are pretrained on broad corpora and can be adapted to a new task through prompting rather than retraining. Supplying instructions plus a few example summaries in context lets a team with no ML specialists and no labeled dataset produce a functional summarizer within days. Training from scratch or using a vision service does not fit the constraint or the modality.

Exam trap

The trap here is assuming that any useful generative AI application requires custom model training, when in-context prompting of a pretrained foundation model is usually sufficient for a first prototype.

731
Multi-Selecteasy

A company wants to log all model invocation requests in Amazon Bedrock for audit and troubleshooting. Which TWO destinations can they configure for invocation logging? (Choose 2)

Select 2 answers
A.Amazon CloudWatch Logs
B.Amazon S3
C.Amazon DynamoDB
D.AWS CloudTrail
E.Amazon Kinesis Data Firehose
AnswersA, B

Amazon CloudWatch Logs is a supported destination for Bedrock model invocation logging, satisfying the audit and troubleshooting requirement. It captures text, image, and embedding invocation data as log events, enabling near-real-time querying and retention policies for compliance reviews.

Why this answer

Amazon Bedrock invocation logging supports two destination types: Amazon CloudWatch Logs (option A) and Amazon S3 (option B). Option A is correct because Bedrock can deliver text and image invocation logs to a CloudWatch Logs log group, enabling real-time monitoring and troubleshooting via CloudWatch. Option B is correct because Bedrock can also deliver invocation logs to an S3 bucket for durable, long-term audit storage and later analysis.

Option C is incorrect because DynamoDB is not a supported invocation logging destination for Bedrock. Option D is incorrect because CloudTrail records API activity/management events, not the model invocation request/response payloads that invocation logging captures. Option E is incorrect because Kinesis Data Firehose is not a configurable destination for Bedrock invocation logging.

732
Multi-Selectmedium

A logistics company wants to use machine learning to predict delivery times. A data scientist is preparing the project and must identify which characteristics describe supervised learning rather than unsupervised learning. (Choose two.)

Select 2 answers
A.The training process requires no historical outcomes and relies only on feature similarity.
B.The goal is to learn a mapping from input features to a known output so the model can predict on new data.
C.The algorithm discovers hidden structure in data without any predefined output values.
D.The training dataset includes a target label for each example, such as actual delivery duration.
E.The model is evaluated primarily by how well it separates data into groups with no ground truth.
AnswersB, D

Supervised learning explicitly learns a function that maps input features to a known output, enabling predictions on unseen examples. For delivery times, the model learns from features such as distance, traffic, and time of day to predict duration for new orders, which is the core objective of supervised learning.

Why this answer

Supervised learning is defined by training on labeled examples and learning a mapping from inputs to a known output so the model can predict on new data. In the delivery-time scenario, historical trips with actual durations provide those labels. The remaining characteristics describe unsupervised learning, which discovers structure or groups without predefined outputs or ground truth.

Exam trap

The trap here is confusing clustering-style structure discovery with supervised prediction because both can operate on the same features.

733
MCQhard

A data scientist is using Amazon Bedrock to generate product descriptions. The current prompt produces inconsistent results; sometimes the descriptions are too verbose, other times too short. The scientist wants to reduce output variability and set a consistent tone. Which combination of parameters should be adjusted?

A.Set temperature to 0 and top-p to 0.5
B.Increase max tokens to 500 and set stop sequences
C.Lower temperature to 0.1 and set top-p to 0.9
D.Increase temperature to 0.9 and set top-p to 1.0
AnswerC

Temperature scales the sampling distribution's randomness, so lowering it to 0.1 sharply reduces variability in length and tone. Top-p at 0.9 restricts sampling to the smallest token set whose cumulative probability reaches 90%, trimming unlikely tokens while retaining fluency, giving consistent, controlled descriptions.

Why this answer

Setting a low temperature (0.1) makes the model more deterministic and less random, reducing output variability, while a high top-p (0.9) allows the model to consider a broad but controlled set of likely tokens, ensuring consistent tone without being overly restrictive. This combination balances creativity and consistency, directly addressing the problem of inconsistent verbosity.

Exam trap

AWS often tests the misconception that lowering temperature alone is sufficient for consistency, but the trap here is that candidates overlook how top-p must be set high enough to avoid overly deterministic outputs, or they mistakenly pair low temperature with low top-p (as in option A), which can cause repetitive or unnatural text.

How to eliminate wrong answers

Option A is wrong because setting temperature to 0 makes the model fully deterministic, which can lead to repetitive or overly rigid outputs, and top-p to 0.5 further restricts token selection, potentially causing unnatural descriptions; this combination does not allow for the controlled variability needed for a consistent tone. Option B is wrong because increasing max tokens to 500 only controls output length, not variability or tone, and setting stop sequences only terminates generation at specific points, neither of which reduces randomness or enforces a consistent style. Option D is wrong because increasing temperature to 0.9 increases randomness and output variability, and setting top-p to 1.0 includes all tokens, which exacerbates inconsistency, the opposite of what is needed.

734
Multi-Selecthard

A retail bank is deploying an Amazon SageMaker model that recommends credit limit increases to existing cardholders. The bank's responsible AI review board requires that the model's decisions be explainable to customers who request an adverse action notice, and that the team be able to detect whether any single input feature is disproportionately driving predictions. Which TWO capabilities should the team implement to meet these requirements? (Choose two.)

Select 2 answers
A.Enable SageMaker Debugger to capture tensor values during training
B.Use SageMaker Clarify explainability to generate SHAP-based feature attributions for individual predictions
C.Run SageMaker Clarify bias detection to measure feature importance and disparate impact across groups
D.Deploy the model behind an Application Load Balancer with AWS WAF for request filtering
E.Configure SageMaker Model Monitor to track feature drift against a baseline distribution
AnswersB, C

SageMaker Clarify explainability computes SHAP values that quantify how much each input feature contributed to a specific prediction. For adverse action notices, this gives the bank concrete per-customer reasons, such as utilization rate or recent delinquencies, that drove the credit limit decision. It directly supports the explainability requirement at the individual prediction level.

Why this answer

SageMaker Clarify provides both SHAP-based per-prediction explanations and bias metrics such as disparate impact and feature importance. Together they let the bank produce customer-facing adverse action reasons and detect whether a particular feature disproportionately drives decisions. Model Monitor, Debugger, and load balancer controls address drift, training internals, and traffic security rather than explainability or fairness.

Exam trap

The trap here is treating SageMaker Model Monitor as a fairness tool, when it only detects input or output drift and does not compute feature attributions or bias metrics.

735
MCQmedium

A product team wants to compare two foundation models on their own customer-support transcripts before choosing one for a chatbot. They need a quantitative measure of how well each model's answers match reference answers. Which evaluation approach fits this need?

A.Human evaluation where reviewers rate each answer on a five-point scale
B.Automated evaluation using similarity metrics such as ROUGE or BERTScore between model outputs and reference answers
C.Reviewing each model's published model card and benchmark leaderboard rankings
D.Measuring inference latency and cost per thousand tokens for each model
AnswerB

Automated metrics compare generated text against reference answers and produce numeric scores, enabling objective side-by-side comparison across models on the same dataset. This suits the requirement for a quantitative measure over the team's own transcripts, and it scales to many examples without manual grading.

Why this answer

Automated text-similarity metrics score generated answers against reference answers and yield numbers that can be compared across models on the same transcripts. Human ratings are qualitative and costly, latency and cost measure operations rather than quality, and public benchmarks do not reflect the company's domain data. Automated evaluation on in-domain data is the fit.

Exam trap

The trap here is confusing operational metrics such as latency and cost with quality metrics that actually measure answer correctness.

736
MCQeasy

Which component of the Transformer architecture allows the model to weigh the importance of different tokens in the input sequence when generating each output token?

A.Layer normalization
B.Feed-forward network
C.Positional encoding
D.Self-attention mechanism
AnswerD

Self-attention computes query, key and value projections for every token, then scores each token against all others so their weighted contributions shape the current output. This weighting is the mechanism the stem requires for judging token importance, unlike feed-forward layers, which process positions independently.

Why this answer

The self-attention mechanism (option D) is the core component of the Transformer that computes attention scores between every pair of tokens in the input sequence, allowing the model to dynamically weigh the importance of each token when generating an output token. This enables the model to capture long-range dependencies and contextual relationships without the sequential constraints of RNNs.

Exam trap

AWS often tests the misconception that positional encoding or feed-forward networks handle token relationships, but the self-attention mechanism is the only component that directly computes pairwise token importance.

How to eliminate wrong answers

Option A is wrong because layer normalization stabilizes training by normalizing activations across features, but it does not weigh token importance or model token relationships. Option B is wrong because the feed-forward network applies a non-linear transformation to each token independently after attention, but it has no mechanism to compare or weigh tokens against each other. Option C is wrong because positional encoding injects information about token order into the input embeddings, but it does not compute relative importance between tokens; that is the role of self-attention.

737
MCQmedium

A developer is building an agent using Amazon Bedrock Agents. The agent needs to call an external API to retrieve weather data. What must the developer define to enable this capability?

A.A Knowledge Base with weather documents
B.An action group with an OpenAPI schema and a Lambda function
C.A Bedrock Guardrail to filter the weather data
D.A prompt flow in Bedrock Studio
AnswerB

An action group with an OpenAPI schema and a Lambda function satisfies the requirement to call an external API. The OpenAPI schema defines the API operations the agent can invoke, while the Lambda function executes the actual weather data retrieval, letting Amazon Bedrock Agents orchestrate the call through defined action group parameters.

Why this answer

To enable an Amazon Bedrock Agent to call an external API, the developer must define an action group that includes an OpenAPI schema describing the API's operations and a Lambda function to execute the API call and return results. The action group bridges the agent's intent recognition with the actual API invocation, making option B the correct choice.

Exam trap

The AWS AI Practitioner exam often tests the distinction between action groups (for API calls) and knowledge bases (for document retrieval), leading candidates to mistakenly choose a knowledge base when the question involves external API integration.

How to eliminate wrong answers

Option A is wrong because a Knowledge Base stores and retrieves documents for RAG, not for executing API calls. Option C is wrong because Bedrock Guardrails enforce content policies and safety filters, not API integration. Option D is wrong because prompt flows in Bedrock Studio orchestrate multi-step prompts but do not directly enable external API calls from an agent.

738
MCQeasy

A small insurance firm is selecting an AI service to classify customer emails by topic. The compliance officer insists that the provider publish clear documentation on how the model was built, its intended use, and its limitations, and that the firm be able to see model version details. Which AWS resource should the firm consult to evaluate these transparency characteristics of Amazon Bedrock foundation models?

A.Amazon Bedrock model cards in the model catalog
B.Amazon CloudWatch Logs insights queries
C.AWS Service Health Dashboard
D.AWS Artifact reports
AnswerA

Bedrock model cards document each foundation model's intended use cases, training approach, limitations, and responsible AI considerations, and they identify model versions. Reviewing them gives the insurance firm the transparency evidence the compliance officer requested before selecting a model for email classification.

Why this answer

Transparency about how a model was built, its intended use, and its limits is documented in model cards. Amazon Bedrock publishes model cards in its model catalog, giving the insurer the version and suitability details the compliance officer wants. Compliance certifications, service health, and log queries describe infrastructure, availability, or runtime events, not model characteristics.

Exam trap

The trap here is confusing infrastructure compliance documentation such as AWS Artifact with model-level transparency artifacts such as Bedrock model cards.

739
MCQeasy

A company wants to monitor for malicious activity in their machine learning pipelines, such as unauthorized access to training data or model artifacts. Which AWS service can provide automated threat detection and continuous monitoring?

A.AWS Config
B.Amazon GuardDuty
C.AWS Shield
D.Amazon Inspector
AnswerB

GuardDuty continuously analyses CloudTrail, VPC flow logs and DNS logs using threat intelligence and machine learning to detect unauthorised access and anomalous behaviour. This satisfies the requirement for automated threat detection and continuous monitoring across AWS accounts, including access to S3-stored training data and model artefacts.

Why this answer

Amazon GuardDuty is a threat detection service that continuously monitors for malicious activity and unauthorized behavior across AWS workloads, including machine learning pipelines. It uses machine learning, anomaly detection, and integrated threat intelligence to identify threats such as unauthorized access to S3 buckets containing training data or model artifacts, without requiring manual intervention.

Exam trap

AWS often tests the distinction between services that monitor for security threats (GuardDuty) versus services that manage compliance (AWS Config), protect against DDoS (AWS Shield), or scan for vulnerabilities (Amazon Inspector), leading candidates to confuse configuration auditing with active threat detection.

How to eliminate wrong answers

Option A is wrong because AWS Config is a service for evaluating and auditing resource configurations against compliance rules, not for continuous threat detection or monitoring for malicious activity. Option C is wrong because AWS Shield is a managed Distributed Denial of Service (DDoS) protection service, designed to safeguard against network and transport layer attacks, not for detecting unauthorized access or malicious behavior in ML pipelines. Option D is wrong because Amazon Inspector is a vulnerability management service that scans for software vulnerabilities and unintended network exposure, not for real-time threat detection or monitoring of malicious activity.

740
MCQeasy

Which AWS service can convert a text document from one language to another?

A.Amazon Polly
B.Amazon Textract
C.Amazon Comprehend
D.Amazon Translate
AnswerD

Amazon Translate performs neural machine translation between supported languages, directly satisfying the stem's requirement to convert a text document from one language to another. Unlike Amazon Comprehend, which extracts entities and sentiment without translating, Translate outputs the same content in the target language, making it the precise fit for this scenario.

Why this answer

Amazon Translate is a neural machine translation service that delivers fast, high-quality, and customizable language translation. It uses deep learning models to translate text documents from a source language to a target language, making option D the correct choice for this task.

Exam trap

The trap here is that candidates may confuse Amazon Comprehend's language detection capability with translation, but Comprehend only identifies the language of text, it does not convert it to another language.

How to eliminate wrong answers

Option A is wrong because Amazon Polly is a text-to-speech service that converts text into lifelike speech, not a translation service. Option B is wrong because Amazon Textract is a document analysis service that extracts text, handwriting, and data from scanned documents using OCR, but it does not translate languages. Option C is wrong because Amazon Comprehend is a natural language processing (NLP) service that extracts insights like entities, sentiment, and key phrases from text, but it does not perform language translation.

741
MCQeasy

A social media company needs to automatically detect and flag toxic comments in multiple languages. They have a large stream of user comments and require real-time moderation. Which AWS service is best suited for this task?

A.Amazon Lex
B.Amazon Comprehend
C.Amazon Rekognition
D.Amazon Translate
AnswerB

Amazon Comprehend provides pre-trained natural language processing, including toxicity detection and multi-language support, and processes streaming text via real-time endpoints. This satisfies the stem's need for automatic, low-latency flagging of toxic comments across multiple languages.

Why this answer

Amazon Comprehend is the correct choice because it is a natural language processing (NLP) service that can perform real-time toxicity detection across multiple languages using its built-in content moderation and custom classification capabilities. It analyzes text streams to identify toxic comments (e.g., hate speech, threats) and integrates with AWS streaming services like Amazon Kinesis for real-time processing.

Exam trap

The trap here is that candidates may confuse Amazon Comprehend's NLP capabilities with Amazon Lex's conversational AI or Amazon Translate's language translation, assuming any language-related service can detect toxicity, but only Comprehend provides the specific text analysis APIs for content moderation.

How to eliminate wrong answers

Option A is wrong because Amazon Lex is a service for building conversational interfaces (chatbots) using automatic speech recognition (ASR) and natural language understanding (NLU), not for analyzing text for toxicity. Option C is wrong because Amazon Rekognition is designed for image and video analysis (e.g., object detection, facial recognition), not for processing text comments. Option D is wrong because Amazon Translate is a machine translation service that converts text between languages but does not perform toxicity detection or content moderation.

742
MCQmedium

A public sector agency is building a chatbot on Amazon Bedrock that answers citizen questions about benefits eligibility. The agency must ensure the chatbot never provides medical or legal advice, and that responses stay grounded in the agency's official policy documents rather than the foundation model's general knowledge. Which Amazon Bedrock Guardrails configuration should the team apply?

A.Configure a sensitive information filter to redact medical terms and legal citations, and enable the profanity filter at HIGH strength to keep responses professional.
B.Configure a denied topics policy for medical and legal advice, and use contextual grounding checks with the official policy documents supplied as the grounding source.
C.Configure automated reasoning checks against a formal policy of eligibility rules, and rely on the foundation model's built-in knowledge to answer general questions.
D.Configure a word filter containing a blocklist of medical and legal terms, and set the guardrail action to NONE so responses are flagged for human review instead of blocked.
AnswerB

A denied topics policy blocks prompts and responses that fall within the defined medical and legal advice subjects, directly enforcing the prohibition. Contextual grounding checks evaluate each response against the supplied policy documents and intervene when the answer is not supported by that source, which keeps the chatbot anchored to official agency content instead of the model's general knowledge.

Why this answer

Two constraints must be enforced: no medical or legal advice, and answers grounded in official policy documents. A denied topics policy names and blocks those advice categories at both input and output. Contextual grounding checks compare each response against the supplied policy documents and intervene when support is weak, which prevents the model from answering from its own general knowledge.

Exam trap

The trap here is using a word filter or sensitive information filter to enforce a topical prohibition, when those mechanisms match or redact specific strings rather than classifying whether a response constitutes prohibited advice.

743
Multi-Selectmedium

A retail company wants to build an application that uses a foundation model on Amazon Bedrock to answer customer questions about product availability. They need the model to access real-time inventory data from their internal database and perform actions such as reserving an item. Which TWO capabilities should they implement to achieve this? (Choose two.)

Select 2 answers
A.Enable model invocation logging to capture inventory queries.
B.Use Retrieval Augmented Generation with a static document store containing product manuals.
C.Define an action group in the agent that maps to an API schema for the inventory and reservation operations.
D.Use Amazon Bedrock Agents to orchestrate calls to an AWS Lambda function that queries the inventory database.
E.Increase the model's temperature to allow more creative problem-solving.
AnswersC, D

Action groups in Amazon Bedrock Agents use OpenAPI schemas to describe available operations. By defining an action group with the inventory query and reservation endpoints, the agent knows how and when to call them. This enables the model to perform the required actions through the agent's orchestration, making it a correct implementation step.

Why this answer

Amazon Bedrock Agents orchestrate tasks by invoking action groups, which are defined with OpenAPI schemas and backed by Lambda functions. This allows the agent to query real-time inventory and perform reservations. RAG over static documents, temperature changes, and logging do not provide live data access or transactional capabilities, so they cannot satisfy the requirements.

Exam trap

The trap here is assuming that RAG alone can provide real-time data and actions, when live database access and transactions require agent action groups backed by compute.

744
MCQhard

A developer is prompting a foundation model to classify customer feedback into categories. The model sometimes returns extra commentary along with the category label, breaking downstream parsing. The developer wants more deterministic, tightly formatted output without retraining the model. Which technique best addresses this?

A.Set the temperature to zero and constrain the output using a structured format specification.
B.Add more few-shot examples that include verbose explanations of each category.
C.Increase the maximum output token count to give the model more room to respond.
D.Increase the top-p sampling value to broaden the token selection pool.
AnswerA

Lowering temperature to zero makes token selection greedy and far more deterministic, while constraining output with a structured format such as a JSON schema enforces the exact shape of the response. Together these yield tightly formatted labels that downstream systems can parse reliably, without any retraining, directly solving the developer's problem.

Why this answer

Setting temperature to zero makes generation greedy and deterministic, while a structured output specification constrains the response to an exact schema. This combination eliminates extraneous commentary and guarantees a parseable label, achieving the desired formatting control without any model retraining. It directly targets both the randomness and the format-enforcement gaps causing the parsing failures.

Exam trap

The trap here is reaching for sampling knobs like top-p or token limits, which affect variety and length rather than enforcing a strict, deterministic output structure.

745
Multi-Selecteasy

A data science team is building a resume screening model and wants to ensure it does not exhibit gender bias. Which TWO actions are most effective for mitigating bias? (Choose TWO.)

Select 2 answers
A.Apply adversarial debiasing techniques during training.
B.Use a more complex deep learning model.
C.Remove the gender attribute and all correlated features from the dataset.
D.Regularly audit model predictions for disparate impact across genders.
E.Ensure the training dataset has equal numbers of male and female candidates.
AnswersA, D

Adversarial debiasing adds a competing objective that penalises the model whenever a classifier can predict gender from its internal representations, forcing those representations to become gender-invariant. This directly reduces the model's reliance on gender-correlated features, satisfying the requirement to mitigate bias during training.

Why this answer

Option A is correct because adversarial debiasing trains a predictor alongside an adversary that tries to detect the protected attribute (gender) from the predictor's outputs, forcing the model to learn representations that cannot discriminate by gender, which directly mitigates bias during training. Option D is correct because regularly auditing model predictions for disparate impact across genders (e.g., using metrics like demographic parity, equal opportunity, or the 80% rule) detects bias that may persist or emerge after deployment and enables corrective action. Option B is not correct because increasing model complexity with deep learning does not inherently reduce bias and can even amplify it by fitting spurious correlations in the data.

Option C is not correct because simply removing the gender attribute and correlated features does not eliminate bias, since proxy variables and historical patterns can still encode gender information. Option E is not correct because equal representation of male and female candidates in the training set does not guarantee fairness, as bias can arise from label imbalance, feature correlations, or unequal outcomes despite balanced group sizes.

Exam trap

A common misconception is that simply removing a protected attribute or balancing the dataset is sufficient to eliminate bias, when in reality, correlated features and model complexity can still encode bias, requiring more advanced debiasing techniques like adversarial training or regular auditing.

746
MCQmedium

A solutions architect must choose a foundation model for an application that summarizes lengthy internal audit reports. The reports average 60,000 tokens, and the summaries must reflect details from the beginning, middle, and end of each document. Cost per request matters, but recall of details is the top priority. Which model characteristic should drive the selection?

A.The model's inference latency percentile
B.The number of parameters in the model
C.The model's supported output modalities
D.The model's maximum context window length
AnswerD

A document of roughly 60,000 tokens must fit inside the model's context window along with the prompt and the generated summary. If the window is smaller, content must be truncated or chunked, which risks losing the details the team cares about most. Context window length is therefore the gating characteristic for this workload.

Why this answer

When a single request must contain an entire long document, the maximum context window is the hard constraint that decides feasibility. Parameter count, output modality, and latency influence quality or experience but do not determine whether 60,000 tokens can be processed without losing the details the task requires.

Exam trap

The trap here is equating a larger parameter count with the ability to handle longer inputs, when context window size is a separate and independently configured limit.

747
MCQmedium

A data scientist is using Amazon SageMaker to train a deep learning model. The training job fails with a 'ResourceLimitExceeded' error. What is the MOST likely cause of this error?

A.The account has reached the limit for concurrent training jobs or instance usage
B.The training data contains corrupted files
C.The model is too large for the chosen instance type
D.The training script has a syntax error
AnswerA

The ResourceLimitExceeded error is thrown by SageMaker when the account's service quota for concurrent training jobs or ml instance usage is exhausted. Since the stem describes a training job failing immediately, the constraint satisfied is the per-account quota ceiling, not data, code, or IAM permissions.

Why this answer

The 'ResourceLimitExceeded' error in Amazon SageMaker indicates that the AWS account has exceeded the service quota for concurrent training jobs or the total number of instances being used. This is a common limit enforced by AWS to prevent resource overconsumption, and it can be resolved by requesting a quota increase via the Service Quotas console or by reducing the number of parallel jobs.

Exam trap

AWS often tests the distinction between resource limits (service quotas) and runtime errors (data corruption, memory, syntax) to see if candidates understand that 'ResourceLimitExceeded' is an AWS infrastructure constraint, not a model or code issue.

How to eliminate wrong answers

Option B is wrong because corrupted training data typically causes data loading errors (e.g., 'Unable to read file' or 'EOFError'), not a 'ResourceLimitExceeded' error which is a service quota issue. Option C is wrong because a model that is too large for the chosen instance type results in an 'OutOfMemory' or 'InsufficientInstanceCapacity' error, not a resource limit error. Option D is wrong because a syntax error in the training script would produce a Python or framework-specific exception (e.g., 'SyntaxError' or 'NameError') during script execution, not a resource limit error.

748
Multi-Selecthard

A data scientist is fine-tuning a foundation model on Amazon Bedrock for a custom summarization task. Which THREE practices should they follow to optimize the fine-tuning process?

Select 3 answers
A.Start with a base model that is already strong in the domain.
B.Use the default hyperparameters without tuning.
C.Use a representative dataset that reflects the target task.
D.Monitor training loss and validation loss to avoid overfitting.
E.Train for as many epochs as possible.
AnswersA, C, D

Selecting a domain-strong base model reduces the volume of task-specific examples and training steps required, since the model already encodes relevant vocabulary and structure. This directly satisfies the stem's optimisation goal by lowering compute cost and convergence time during Bedrock fine-tuning, rather than compensating for weak domain representation through extra data.

Why this answer

Option A is correct because selecting a base model already strong in the target domain gives the fine-tuning process a better starting point, reducing the amount of task-specific data and compute needed to reach high summarization quality. Option C is correct because a representative dataset that reflects the target task ensures the model learns the desired summarization style, domain vocabulary, and input-output distribution, which directly improves fine-tuning effectiveness. Option D is correct because monitoring training loss and validation loss lets the data scientist detect overfitting early and apply mitigations such as early stopping, regularization, or more data, keeping generalization strong.

Option B is not correct because leaving hyperparameters at defaults ignores task-specific tuning of values like learning rate, batch size, and epoch count, which often materially affects fine-tuning results. Option E is not correct because training for as many epochs as possible typically causes overfitting, degrading validation performance rather than optimizing the process.

Exam trap

The AIF-C01 exam often tests the misconception that more epochs always improve model performance, when in fact excessive training leads to overfitting, and they expect candidates to recognize that monitoring loss curves and using early stopping are critical practices.

749
MCQmedium

A company is implementing a Retrieval-Augmented Generation (RAG) pipeline with Amazon Bedrock Knowledge Bases. They need to store vector embeddings for their documents. Which vector store options are natively supported by Bedrock Knowledge Bases?

A.Amazon OpenSearch Serverless, Pinecone, MongoDB Atlas, and Amazon Aurora pgvector
B.Only Pinecone and Amazon OpenSearch Serverless
C.Amazon DynamoDB and Amazon RDS for MySQL
D.Only Amazon OpenSearch Serverless
AnswerA

Bedrock Knowledge Bases integrates natively with these four vector stores, so embeddings can be written and queried without custom connectors. Amazon OpenSearch Serverless, Pinecone, MongoDB Atlas and Aurora pgvector are the supported backends, satisfying the requirement to store vector embeddings.

Why this answer

Amazon Bedrock Knowledge Bases natively support multiple vector stores, including Amazon OpenSearch Serverless, Pinecone, MongoDB Atlas, and Amazon Aurora with pgvector. These integrations allow customers to choose a vector store that fits their existing infrastructure and performance requirements. The service provides built-in connectors to these databases for storing and retrieving embeddings.

Exam trap

AIF-C01 often tests the specific list of supported vector stores; candidates may assume only AWS-native services are supported and overlook third-party options like Pinecone or MongoDB Atlas.

How to eliminate wrong answers

Option B is wrong because it incorrectly limits support to only Pinecone and OpenSearch Serverless, omitting MongoDB Atlas and Aurora pgvector. Option C is wrong because DynamoDB and RDS for MySQL are not natively supported vector stores for Bedrock Knowledge Bases; they lack the vector search capabilities required. Option D is wrong because it claims only OpenSearch Serverless is supported, which is false as other vector stores are also supported.

750
MCQhard

A media company must give an external analytics vendor temporary, auditable access to a curated S3 dataset used for fine-tuning a model. The vendor works from its own AWS account, and the company's security policy forbids sharing long-term IAM credentials. Which approach best meets the policy while keeping access auditable?

A.Generate a presigned URL for each S3 object and email the URLs to the vendor on a weekly schedule.
B.Configure a cross-account IAM role in the company account that the vendor assumes using AWS Security Token Service, with an external ID condition.
C.Enable S3 Block Public Access and place the dataset behind an Amazon CloudFront signed cookie distribution.
D.Create an IAM user in the company account for the vendor and rotate the access keys every 24 hours with a script.
AnswerB

A cross-account role lets the vendor's own principals call AssumeRole and receive temporary credentials, so no long-term secrets are shared. The external ID condition guards against the confused deputy problem, and because the vendor assumes a distinct role, CloudTrail records its activity separately. This satisfies both the no-shared-credentials policy and the auditability requirement.

Why this answer

Cross-account IAM roles with AWS Security Token Service temporary credentials are the standard way to grant a partner access without exchanging secrets. The external ID prevents the confused deputy problem, and separate role assumption produces clean CloudTrail attribution. Presigned URLs, rotated IAM user keys, and CDN-signed access all either share secrets or obscure who actually accessed the data.

Exam trap

The trap here is treating time-limited presigned URLs or rotated access keys as equivalent to temporary role-based credentials, when only role assumption avoids sharing secrets and preserves clear identity attribution.

Page 9

Page 10 of 12

Page 11