Courseiva

AWS Certified AI Practitioner AIF-C01 (AIF-C01) — Questions 1–75

862 questions total · 12pages · All types, answers revealed

Page 1 of 12

Page 2
1
MCQhard

A healthcare company must train a model on sensitive patient data while complying with privacy regulations. They want to add noise to the training process to prevent re-identification. Which technique should they implement?

A.Differential privacy
B.k-anonymity
C.Federated learning
D.Homomorphic encryption
AnswerA

Differential privacy injects calibrated statistical noise into training or query outputs, bounding any single patient's influence on the model. This mathematically limits re-identification risk while preserving aggregate utility, satisfying privacy regulations when training on sensitive patient records.

Why this answer

Differential privacy is the correct technique because it adds calibrated noise to the training process (e.g., via gradient clipping and noise injection in stochastic gradient descent) to ensure that the model's outputs do not reveal whether any individual's data was included in the training set. This provides a formal mathematical guarantee (ε-differential privacy) that limits the risk of re-identification, which is essential for complying with privacy regulations like HIPAA or GDPR when training on sensitive patient data.

Exam trap

AWS often tests the misconception that federated learning alone provides privacy guarantees, but the trap here is that federated learning only addresses data locality, not re-identification resistance, which requires a formal privacy technique like differential privacy.

How to eliminate wrong answers

Option B (k-anonymity) is wrong because it is a data anonymization technique applied to static datasets (e.g., generalizing quasi-identifiers in a table) rather than a training process technique; it does not add noise during model training and can be vulnerable to attacks like homogeneity or background knowledge attacks. Option C (Federated learning) is wrong because it is a distributed training approach that keeps data on local devices but does not inherently add noise to prevent re-identification; without differential privacy, model updates can still leak sensitive information. Option D (Homomorphic encryption) is wrong because it allows computation on encrypted data but does not add noise to the training process; it protects data in transit or at rest but does not prevent re-identification from model outputs.

2
MCQhard

A company is building a RAG application using Amazon Bedrock Knowledge Bases. They want to ensure that the model only answers based on the retrieved documents and does not use its internal knowledge. Which configuration should they use?

A.Set the model temperature to 0
B.Use a fine-tuned model that has only seen the company's documents
C.Enable the 'Only use retrieved context' option in the Knowledge Base configuration
D.Use a Guardrail with topic denial for out-of-scope topics
AnswerC

Constraining generation to retrieved passages is exactly what the 'Only use retrieved context' setting enforces, so the model cannot fall back on its pretrained parametric knowledge. This satisfies the stem's requirement that answers derive solely from the ingested documents.

Why this answer

Amazon Bedrock Knowledge Bases provides a specific configuration toggle called 'Only use retrieved context' that forces the foundation model to generate responses solely from the information in the retrieved document chunks, ignoring any internal knowledge the model may have been trained on. This ensures strict adherence to the retrieved context, which is critical for RAG applications requiring factual grounding.

Exam trap

The trap here is that candidates often confuse model temperature or fine-tuning with retrieval grounding, but neither mechanism restricts the model from using its internal knowledge; only the explicit 'Only use retrieved context' toggle in Amazon Bedrock Knowledge Bases enforces that constraint.

How to eliminate wrong answers

Option A is wrong because setting the model temperature to 0 reduces randomness but does not prevent the model from using its internal knowledge; it only makes outputs more deterministic. Option B is wrong because fine-tuning a model on the company's documents does not guarantee it will ignore its pre-trained knowledge; the model can still hallucinate or blend internal knowledge with retrieved context. Option D is wrong because Guardrails with topic denial can block out-of-scope topics but do not force the model to rely exclusively on retrieved documents; the model may still use its internal knowledge for in-scope queries.

3
Multi-Selecthard

Which THREE practices are recommended for responsible AI when deploying foundation models? (Choose three.)

Select 3 answers
A.Avoid collecting user feedback to reduce bias
B.Include human review for high-stakes decisions
C.Implement guardrails to filter harmful content
D.Continuously monitor model outputs for drift
E.Use a black box approach to keep model internals secret
AnswersB, C, D

Human review for high-stakes decisions satisfies the accountability constraint by keeping a person responsible for consequential outcomes, rather than delegating judgement to a probabilistic model. Foundation models can produce confident but wrong outputs, so oversight catches harmful errors before they affect people, aligning with responsible AI principles.

Why this answer

Option B is correct because responsible AI deployment requires human-in-the-loop oversight for consequential decisions (e.g., hiring, lending, medical triage), so that a qualified person can validate or override model output before it affects someone's rights or safety. Option C is correct because guardrails — input/output filters, content moderation classifiers, and policy enforcement layers — are a standard control to prevent foundation models from generating harmful, unsafe, or policy-violating content. Option D is correct because foundation models can degrade or shift behavior over time due to data drift, concept drift, or upstream model updates, so continuous monitoring of outputs (with metrics, alerts, and retraining triggers) is essential to maintain reliability and fairness.

Option A is not recommended: collecting user feedback is a valuable signal for identifying bias and improving models, so avoiding it would reduce accountability rather than enhance responsibility. Option E is not recommended: opaque 'black box' secrecy undermines transparency, explainability, and auditability, which are core responsible AI principles; model internals and documentation should be appropriately disclosed to stakeholders.

Exam trap

A common misconception is that avoiding user feedback reduces bias, when in fact it starves the system of data needed to detect and correct bias, making it a harmful anti-pattern. AWS recommends continuous feedback and monitoring for responsible AI.

4
MCQhard

A company uses Amazon Bedrock with a custom model deployed via Amazon SageMaker. They want to monitor for data drift in input prompts over time. Which AWS service is best suited for this?

A.Amazon CloudWatch
B.Amazon SageMaker Model Monitor
C.AWS CloudTrail
D.Amazon Athena
AnswerB

SageMaker Model Monitor captures incoming request data and compares it against a baseline, detecting drift in prompt feature distributions over time. This satisfies the requirement to monitor input prompts for data drift on a custom SageMaker-deployed model.

Why this answer

Amazon SageMaker Model Monitor is the correct choice because it is specifically designed to detect data drift in machine learning models, including input prompts for custom models deployed via SageMaker. It continuously monitors the distribution of input data against a baseline and alerts when drift occurs, which aligns with the requirement to monitor input prompts over time.

Exam trap

The trap here is that candidates often confuse general monitoring services like CloudWatch with specialized ML monitoring tools, assuming CloudWatch can handle data drift detection when it actually lacks the statistical analysis required for such tasks.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch is a monitoring service for AWS resources and applications (e.g., metrics, logs, alarms), but it does not have built-in capabilities to detect data drift in ML model inputs. Option C is wrong because AWS CloudTrail records API activity for auditing and governance, not for monitoring data drift in model inputs. Option D is wrong because Amazon Athena is an interactive query service for analyzing data in S3 using SQL, not a monitoring tool for data drift.

5
Multi-Selectmedium

Which THREE steps are typically involved in fine-tuning a foundation model? (Select THREE.)

Select 3 answers
A.Deploy the model immediately without additional training
B.Prepare a labeled dataset specific to the target domain
C.Train the model on the domain dataset with a lower learning rate
D.Select a pre-trained foundation model as the starting point
E.Choose a model architecture with more parameters than the base model
AnswersB, C, D

Fine-tuning adapts a foundation model's weights to a target task, which requires supervised examples the model can learn from. Preparing a labelled dataset specific to the target domain supplies those input-output pairs, satisfying the scenario's need for task-relevant training data rather than relying on the model's general pre-trained knowledge.

Why this answer

Fine-tuning a foundation model begins with selecting an appropriate pre-trained foundation model as the starting point (D), since the whole point of fine-tuning is to adapt existing general-purpose weights rather than train from scratch. Next, you must prepare a labeled dataset specific to the target domain (B), because supervised fine-tuning requires task-relevant input-output pairs to steer the model toward the desired behavior. Then you train the model on that domain dataset with a lower learning rate (C), which is standard practice to avoid catastrophic forgetting and to gently nudge the pre-trained weights instead of overwriting them.

Option A is incorrect because deploying the model immediately without additional training is the opposite of fine-tuning—it describes using the base model as-is. Option E is incorrect because fine-tuning does not require choosing an architecture with more parameters than the base model; you typically fine-tune the same pre-trained architecture, and increasing parameter count is not a fine-tuning step.

Exam trap

AWS often tests the distinction between fine-tuning and other adaptation methods (like prompt engineering or retrieval-augmented generation), and the trap here is that candidates might think fine-tuning requires a larger model or no additional data, when in fact it requires a labeled dataset and the same architecture.

6
MCQeasy

A financial services company is deploying a machine learning model to approve loans. They want to ensure that the model does not discriminate based on race or gender. Which AWS service or feature can help them detect bias in the model's predictions?

A.Amazon SageMaker Clarify
B.AWS CloudTrail
C.Amazon SageMaker Model Monitor
D.AWS Identity and Access Management (IAM)
AnswerA

SageMaker Clarify runs bias metrics such as disparate impact across protected attributes like race and gender, both pre-training and on model predictions. This directly satisfies the requirement to detect discrimination in loan approval outputs, unlike general monitoring or debugging tools.

Why this answer

Amazon SageMaker Clarify is purpose-built to detect bias in ML models and datasets. It computes bias metrics such as Disparate Impact (DI) and Difference in Conditional Acceptance (DCA) across protected attributes like race and gender, and also provides explainability via SHAP values. This directly addresses the requirement to detect discrimination in loan approval predictions.

Exam trap

AIF-C01 often tests the distinction between SageMaker Clarify (bias/fairness detection) and SageMaker Model Monitor (drift/quality monitoring) — candidates confuse the two because both are 'SageMaker monitoring' features.

How to eliminate wrong answers

Option B is wrong because AWS CloudTrail only logs API activity and account events for auditing — it has no capability to analyze model predictions for bias. Option C is wrong because SageMaker Model Monitor detects data drift, model quality degradation, and feature attribution drift in production, but it does not compute fairness/bias metrics against protected attributes. Option D is wrong because IAM controls authentication and authorization to AWS resources; it has no role in evaluating model fairness.

7
MCQhard

A company is building a multi-modal application that processes images and text to answer questions about product defects. Which foundation model approach is BEST?

A.Use an image captioning model and then analyze the caption text
B.Use a text-to-image generation model and analyze the generated image
C.Use a multi-modal foundation model that processes both images and text
D.Use a separate image analysis model and a text model, then combine outputs
AnswerC

A multi-modal foundation model encodes images and text into a shared embedding space, letting one model jointly reason over both modalities. This directly satisfies the stem's requirement to answer defect questions from combined image and text input, unlike text-only or vision-only models that cannot correlate the two.

Why this answer

Multi-modal foundation models (e.g., CLIP, Flamingo, GPT-4V) are specifically designed to jointly process and reason over images and text in a unified architecture. This allows the model to directly correlate visual defects with textual descriptions without intermediate lossy transformations, making it the most effective approach for a multi-modal QA task.

Exam trap

AWS often tests the misconception that combining two separate single-modal models (Option D) is equivalent to a true multi-modal model, but the trap is that late fusion lacks the joint embedding and cross-attention mechanisms needed for coherent multi-modal reasoning.

How to eliminate wrong answers

Option A is wrong because image captioning models convert the entire image into a single text caption, losing fine-grained spatial and defect-specific details that are critical for accurate defect analysis. Option B is wrong because text-to-image generation models create new images from text, which is the inverse of the required task and cannot analyze existing product images for defects. Option D is wrong because using separate models and combining outputs introduces a late-fusion bottleneck, where alignment between visual features and text is not learned end-to-end, leading to poorer performance on tasks requiring joint reasoning.

8
MCQeasy

A solutions architect needs to build a generative AI application that can invoke foundation models from Amazon and third-party providers through a single, unified API without managing any infrastructure. The architect wants the fastest path to a working prototype using AWS-native tooling. Which AWS service should the architect choose?

A.Amazon SageMaker AI
B.Amazon Bedrock
C.Amazon Polly
D.Amazon Comprehend
AnswerB

Amazon Bedrock is the fully managed AWS service that exposes foundation models from Amazon and multiple third-party providers through one unified API, so the architect can prototype quickly without provisioning servers. It handles model hosting and scaling, and supports capabilities such as knowledge bases and guardrails, making it the direct fit for a single-API, serverless generative AI prototype.

Why this answer

Amazon Bedrock is purpose-built for invoking foundation models from Amazon and third-party providers through one API, with no infrastructure to manage. SageMaker AI offers flexibility but adds operational overhead, while Comprehend and Polly serve narrow NLP and speech use cases rather than generative model invocation. The unified, serverless model access makes Bedrock the correct choice for a fast prototype.

Exam trap

The trap here is assuming that any AWS AI service can invoke foundation models, when only Amazon Bedrock provides unified, serverless access to Amazon and third-party foundation models.

9
MCQhard

A company runs a question-answering application on Amazon Bedrock that answers from a large knowledge base. Recently, users have reported that the model gives incomplete answers, often missing details from the middle of documents. The team suspects the chunking strategy is suboptimal. Which adjustment is MOST likely to improve completeness?

A.Change the embedding model to a more powerful one
B.Use overlapping chunks so that boundaries do not cut off meaningful content
C.Increase the chunk size to include more context per chunk
D.Decrease the chunk size to capture more granular details
AnswerB

Overlapping chunks repeat content across adjacent boundaries, so sentences or details split between chunks still appear intact in at least one chunk. This directly addresses incomplete answers caused by boundaries cutting off meaningful content from the middle of documents.

Why this answer

Overlapping chunks ensure that sentences or concepts that span chunk boundaries are not lost, improving retrieval completeness. Larger chunks may cause loss of precision, smaller chunks may lose context, and embedding model change is not directly related to chunking.

10
MCQeasy

A data scientist needs to convert text into numerical vectors for semantic search. Which type of foundation model should they use?

A.Embedding model
B.Text generation model
C.Multimodal model
D.Image generation model
AnswerA

Embedding models map text into dense numerical vectors positioned so semantically similar content sits close together in vector space. This directly satisfies the semantic search requirement, because the resulting vectors can be indexed and compared by similarity rather than generating or classifying text.

Why this answer

An embedding model is specifically designed to convert text (or other data) into dense numerical vectors that capture semantic meaning. These vectors enable efficient similarity comparisons in vector space, which is the core requirement for semantic search. Text generation models produce sequences of tokens, not fixed-length vector representations, making them unsuitable for this task.

Exam trap

AWS often tests the distinction between generative models (which produce new content) and representation models (which encode data into vectors), leading candidates to mistakenly choose a text generation model for tasks requiring vector embeddings.

How to eliminate wrong answers

Option B is wrong because text generation models (e.g., GPT, Claude) are autoregressive and output token sequences, not fixed-length embedding vectors; they cannot directly produce numerical vectors for semantic search. Option C is wrong because multimodal models process and generate multiple data types (e.g., text, images, audio) but are not optimized for converting text alone into semantic vectors; their primary purpose is cross-modal understanding, not embedding generation. Option D is wrong because image generation models (e.g., DALL-E, Stable Diffusion) create visual outputs from text prompts and have no mechanism to produce numerical vectors for text-based semantic search.

11
MCQmedium

A team is using Amazon SageMaker to train a deep learning model. They notice that the training loss decreases steadily but the validation loss starts increasing after 10 epochs. Which technique should they apply to address this issue?

A.Apply early stopping
B.Reduce the batch size
C.Increase the learning rate
D.Add more layers to the network
AnswerA

Early stopping halts training once validation loss stops improving, directly countering the overfitting that emerges after epoch 10. By restoring the best-performing checkpoint, it prevents the model memorising training data while validation performance degrades, satisfying the scenario's need to arrest the diverging loss curves.

Why this answer

The scenario describes overfitting, where the model memorizes training data but fails to generalize to validation data. Early stopping halts training when validation loss stops improving, preventing overfitting while preserving the best model weights. This is a standard regularization technique in SageMaker training jobs, configurable via the `use_early_stopping` parameter in the `Estimator` or `HyperparameterTuner`.

Exam trap

AWS often tests the misconception that increasing model complexity or adjusting batch size can fix overfitting, when the correct first-line approach is early stopping or other regularization techniques.

How to eliminate wrong answers

Option B is wrong because reducing batch size introduces noisier gradient estimates, which can actually worsen generalization and does not directly address overfitting. Option C is wrong because increasing the learning rate can cause divergence or overshooting of the loss minimum, exacerbating validation loss increase. Option D is wrong because adding more layers increases model capacity, which typically worsens overfitting rather than mitigating it.

12
Multi-Selecthard

A company is developing a generative AI application using Amazon Bedrock for code generation. They want to reduce costs without sacrificing throughput. Which THREE approaches can help achieve cost optimization?

Select 3 answers
A.Always use the largest available model for best results
B.Right-size model selection by choosing the smallest model that meets accuracy requirements
C.Use batch inference for non-real-time workloads
D.Fine-tune a large foundation model on the company's codebase
E.Enable model caching to reuse previously generated responses for similar inputs
AnswersB, C, E

Code generation accuracy requirements are often met by smaller models. Selecting the smallest model that satisfies them lowers per-token inference cost while maintaining throughput, directly addressing the cost constraint without sacrificing the application's output quality.

Why this answer

Option B is correct because right-sizing model selection—choosing the smallest Bedrock model that still meets the application's accuracy requirements—directly reduces per-token inference costs while maintaining acceptable throughput for code generation. Option C is correct because Bedrock batch inference processes large volumes of non-real-time requests at a significantly lower cost per token than on-demand synchronous invocation, which is ideal for workloads that don't need immediate responses. Option E is correct because enabling model caching (e.g., via prompt caching or an application-level cache) lets the application reuse previously generated responses for similar inputs, cutting the number of billable model invocations and thus reducing cost without lowering throughput.

Option A is incorrect because always using the largest model increases cost unnecessarily and does not optimize spend. Option D is incorrect because fine-tuning a large foundation model adds training and storage costs and does not by itself reduce inference cost or improve throughput.

Exam trap

AIF-C01 often tests the misconception that using the largest model or fine-tuning always improves results, but cost optimization requires right-sizing, batch processing, and caching instead.

13
MCQhard

A company is deploying a generative AI application that creates marketing copy. They want to ensure the outputs do not include harmful or inappropriate content. Which AWS service can enforce content policies and filter undesirable outputs?

A.AWS WAF
B.Amazon Comprehend
C.Amazon Bedrock Guardrails
D.AWS Shield
AnswerC

Amazon Bedrock Guardrails applies configurable content filters and denied-topic policies to model inputs and outputs, blocking harmful or inappropriate marketing copy. This directly satisfies the stem's requirement to enforce content policies and filter undesirable generative AI outputs.

Why this answer

Amazon Bedrock Guardrails lets you define content filters, denied topics, word filters, and sensitive information filters that are applied to both inputs and outputs of foundation models. It is purpose-built to enforce responsible AI policies and block harmful or inappropriate content in generative AI applications. This directly addresses the requirement to filter undesirable outputs from a marketing copy generator.

Exam trap

AIF-C01 often tests the confusion between general-purpose AWS security services (WAF, Shield) and AI-specific responsible AI tooling — candidates must recognize that content moderation for generative AI requires Bedrock Guardrails, not traditional network security services.

How to eliminate wrong answers

Option A is wrong because AWS WAF is a web application firewall that protects against HTTP-layer attacks like SQL injection and XSS — it has no understanding of generative AI content semantics and cannot filter model outputs. Option B is wrong because Amazon Comprehend is a natural language processing service for sentiment analysis, entity recognition, and topic modeling; it does not enforce content policies or block outputs in a generative AI pipeline. Option D is wrong because AWS Shield is a managed DDoS protection service, unrelated to content moderation or AI output filtering.

14
MCQeasy

A hospital wants to build a model that predicts whether a patient has a specific disease based on labeled historical medical records where each record is marked either positive or negative. Which type of machine learning problem does this represent?

A.Reinforcement learning using reward signals
B.Supervised learning using regression
C.Unsupervised learning using clustering
D.Supervised learning using classification
AnswerD

The records carry known labels of positive or negative, which is exactly the supervision signal classification uses to learn a decision boundary. Because the target is a discrete category rather than a continuous number, the task is classification. The model can then predict the disease status for new patients, matching the hospital's goal.

Why this answer

Because each historical record already includes a known positive or negative label and the target is a discrete category, the task is supervised classification. Regression would predict a continuous value, clustering ignores the labels, and reinforcement learning requires rewards from interaction. Classification directly models the disease outcome the hospital wants to predict.

Exam trap

The trap here is confusing labeled binary prediction with regression simply because both are supervised learning tasks.

15
MCQhard

Refer to the exhibit. A developer runs the CLI command to summarize text using Claude v2 in Bedrock. The output is shorter than expected. Which change should the developer make to allow a longer response?

A.Increase 'max_tokens_to_sample' to 1000
B.Change the prompt to include 'Write a long summary'
C.Set 'stop_reason' to 'none'
D.Use a different region like us-west-2
AnswerA

The response is truncated because max_tokens_to_sample caps generated output length. Raising it to 1000 permits Claude v2 to emit more tokens before stopping, directly satisfying the requirement for a longer summary in the Bedrock CLI invocation.

Why this answer

The 'max_tokens_to_sample' parameter in the Bedrock InvokeModel API directly controls the maximum number of tokens the model can generate in its response. By default, this value is often set low (e.g., 256 tokens), which truncates the output. Increasing it to 1000 allows Claude v2 to produce a longer summary up to that token limit.

Exam trap

A common mistake is to assume that simply asking the model to produce a longer response via prompt engineering will override the API parameter 'max_tokens_to_sample'. In AWS Bedrock, the 'max_tokens_to_sample' parameter is the definitive control for response length.

How to eliminate wrong answers

Option B is wrong because simply adding 'Write a long summary' to the prompt does not override the hard token limit set by 'max_tokens_to_sample'; the model will still stop generating once the token budget is exhausted, regardless of instruction phrasing. Option C is wrong because 'stop_reason' is an output field returned by the API indicating why generation stopped (e.g., 'max_tokens', 'stop_sequence'), not an input parameter; setting it to 'none' is invalid and would not affect response length. Option D is wrong because the region (e.g., us-west-2) does not impose a different default token limit for Claude v2; the 'max_tokens_to_sample' parameter is a model-level constraint independent of regional deployment.

16
MCQhard

A company is developing a fraud detection system using a neural network. The training loss decreases steadily but the validation loss begins to increase after a certain number of epochs. Which action should be taken to address this issue?

A.Implement early stopping to halt training when validation performance degrades
B.Add more layers to the neural network
C.Increase the amount of training data
D.Increase the learning rate
AnswerA

Early stopping monitors validation loss each epoch and halts training once it stops improving, directly countering the divergence described. It preserves the weights from the best-performing epoch, preventing the overfitting that causes validation loss to rise while training loss keeps falling. This satisfies the stem's requirement without altering the model architecture.

Why this answer

The described behavior—training loss decreasing while validation loss increases—is a classic sign of overfitting. Early stopping monitors the validation loss and halts training when it stops improving (or begins to degrade), preventing the model from memorizing noise in the training data. This directly addresses the overfitting issue without requiring architectural or data changes.

Exam trap

AWS AI Practitioner often tests the distinction between overfitting (validation loss increasing) and underfitting (both losses high), leading candidates to mistakenly choose more data or more layers when the correct immediate fix is early stopping.

How to eliminate wrong answers

Option B is wrong because adding more layers increases model capacity, which typically worsens overfitting by allowing the network to fit training data noise even more closely. Option C is wrong because while more training data can help generalization, it does not directly stop the ongoing overfitting during training; early stopping is the immediate corrective action. Option D is wrong because increasing the learning rate can cause the optimizer to overshoot minima, leading to unstable training and potentially faster divergence on the validation set, not a solution to overfitting.

17
MCQhard

A data scientist trains a model to predict whether a loan applicant will default. After deployment, the model performs well on applicants similar to the training data but poorly on applicants from a newly added geographic region that was underrepresented in training. Which statement best describes the underlying problem?

A.The training data does not represent the new region, so the model cannot generalize to that subpopulation.
B.The model has overfit the training set because it memorized noise in the original regions.
C.The model requires more training epochs to converge on the new region's data.
D.The evaluation metric used is inappropriate and should be replaced with accuracy.
AnswerA

When a subpopulation is underrepresented or absent in training, the model learns patterns that do not transfer to it, producing poor predictions for that group. This is a data coverage and distribution shift problem, and it matches the observed failure on the newly added geographic region exactly.

Why this answer

The model fails specifically on a subpopulation that was underrepresented in training, which is a data coverage and distribution shift problem. Good performance on familiar applicants confirms the algorithm works, but without representative training examples from the new region the model cannot generalize there. The remedy is additional representative data or techniques that address distribution shift.

Exam trap

The trap here is labeling any train-versus-deployment gap as overfitting, when the failure is isolated to a subpopulation missing from the training distribution.

18
MCQmedium

A developer is using the Amazon Bedrock API to generate text. They notice that the model sometimes returns harmful content despite setting safety parameters. What is the BEST way to add an additional layer of content filtering?

A.Fine-tune the model on a curated safe dataset
B.Configure content filters in Amazon Bedrock Guardrails
C.Improve prompt engineering with more specific instructions
D.Use AWS WAF to filter API responses
AnswerB

Amazon Bedrock Guardrails applies configurable content filters that evaluate both prompts and model responses independently of the model's own safety parameters, blocking harmful categories the model may still emit and adding a deterministic policy layer.

Why this answer

Amazon Bedrock Guardrails provides a dedicated, configurable content filtering layer that can block harmful content at inference time, independent of the model's built-in safety parameters. This allows developers to enforce custom policies (e.g., hate speech, violence) without modifying the model itself, making it the best additional safeguard.

Exam trap

The AIF-C01 exam often tests the misconception that fine-tuning or prompt engineering alone can fully prevent harmful outputs, when in fact a separate, configurable guardrail layer is the recommended approach for production-grade content filtering in Amazon Bedrock.

How to eliminate wrong answers

Option A is wrong because fine-tuning the model on a curated safe dataset adjusts the model's weights to reduce harmful outputs, but it does not guarantee filtering of all harmful content at inference and requires significant retraining effort; it is not an 'additional layer' but a model modification. Option C is wrong because improving prompt engineering with more specific instructions can guide the model's behavior but cannot reliably block harmful content that the model might generate despite instructions, as it lacks enforcement at the API response level. Option D is wrong because AWS WAF is a web application firewall designed to filter HTTP requests to web applications, not to inspect or filter the content of API responses from Bedrock; it operates at the network layer, not the application content layer.

19
MCQhard

A company is building a model to predict loan default. They have historical data with 5% default rate. The model must minimize false negatives (missed defaults) because each default costs $50,000. False positives (incorrectly flagged defaults) cost $500 in customer service time. The model currently has a recall of 0.70 and precision of 0.80. Which of the following actions would MOST likely reduce the total cost?

A.Increase the model's precision by raising the classification threshold
B.Use a different algorithm that trades off recall for precision
C.Add more features to the model without changing the threshold
D.Increase the model's recall by lowering the classification threshold
AnswerD

Lowering the threshold captures more positives, improving recall and reducing the most costly errors (false negatives).

Why this answer

The cost of a false negative ($50,000) is 100 times greater than a false positive ($500). Lowering the classification threshold increases recall (reduces false negatives) at the expense of precision (increases false positives). Given the extreme cost asymmetry, the net expected cost will decrease even if many more false positives occur, because each additional true positive saves $50,000 while each extra false positive costs only $500.

Option D directly increases recall, which is the correct lever for this cost structure.

Exam trap

AWS often tests the misconception that higher precision is always better, but in cost-sensitive scenarios with asymmetric costs, maximizing recall (even at the cost of precision) is the correct strategy to minimize total financial loss.

How to eliminate wrong answers

Option A is wrong because raising the threshold increases precision but decreases recall, which would increase false negatives—the most costly error type—and thus raise total cost. Option B is wrong because trading off recall for precision means reducing recall to gain precision, which again increases false negatives and total cost; the correct trade-off is the opposite. Option C is wrong because adding features without changing the threshold does not guarantee improved recall; it may improve overall model accuracy but does not directly target the reduction of false negatives, and could even degrade recall if the new features introduce noise.

20
Multi-Selecthard

A company is designing a RAG pipeline for a legal document review system. They need to ingest hundreds of documents, create embeddings, and store them for retrieval. Which THREE steps are essential in the ingestion phase of the RAG pipeline?

Select 3 answers
A.Configure Bedrock Guardrails for the application
B.Chunk the documents into smaller pieces
C.Fine-tune the base model on legal documents
D.Generate embeddings for each chunk using an embedding model
E.Store the embeddings in a vector store
AnswersB, D, E

Chunking splits lengthy legal documents into smaller passages before embedding, which is essential because embedding models have fixed token limits and retrieval precision degrades when a single vector represents an entire contract. Smaller chunks let the retriever return the specific clause relevant to a query, satisfying the ingestion requirement to prepare documents for accurate semantic search.

Why this answer

The ingestion phase of a RAG pipeline is responsible for preparing source documents so they can later be retrieved: option B (chunk the documents into smaller pieces) is essential because splitting long legal documents into appropriately sized chunks keeps each unit within the embedding model's token limit and improves retrieval granularity and relevance. Option D (generate embeddings for each chunk using an embedding model) is essential because RAG retrieval works by comparing the vector representation of a query against vector representations of the content, so every chunk must be converted into an embedding. Option E (store the embeddings in a vector store) is essential because the resulting vectors must be persisted in a vector database or index so that similarity search can be performed at query time.

Option A (Bedrock Guardrails) is a runtime safety and content-filtering concern applied to model inputs/outputs, not a data ingestion step, and option C (fine-tuning the base model) is a separate model-customization activity that is not required to ingest documents into a RAG pipeline.

Exam trap

AIF-C01 often tests the components of RAG, and candidates might include fine-tuning or guardrails as part of ingestion. The trap is confusing training-time activities with ingestion-time activities.

21
MCQeasy

Which Amazon Bedrock feature allows you to invoke a model and receive the response token by token as it is generated, reducing perceived latency for the end user?

A.Model invocation API
B.Provisioned throughput
C.Batch inference
D.Streaming responses
AnswerD

Streaming responses return generated tokens incrementally over a persistent connection rather than buffering the full completion, so the user sees output immediately. This directly addresses the perceived-latency constraint in the stem, even though total generation time is unchanged.

Why this answer

Streaming responses in Amazon Bedrock allow the model to send back partial results token by token as they are generated, rather than waiting for the entire response to be complete. This reduces perceived latency for end users by enabling them to see the output incrementally, which is especially important for real-time applications like chatbots or interactive assistants.

Exam trap

The trap here is that candidates confuse the standard synchronous Model invocation API (which returns the full response at once) with the streaming capability, assuming that 'invocation' inherently includes streaming, when in fact a separate API call (InvokeModelWithResponseStream) is required.

How to eliminate wrong answers

Option A is wrong because the Model invocation API is the general endpoint used to call a model synchronously, but it does not inherently provide token-by-token streaming; streaming requires a separate mechanism (e.g., InvokeModelWithResponseStream). Option B is wrong because Provisioned Throughput is a pricing and capacity feature that guarantees a certain number of inference tokens per minute, but it does not change how responses are delivered (streaming vs. non-streaming). Option C is wrong because Batch inference processes multiple requests asynchronously in bulk and returns results after all are complete, which is the opposite of real-time token-by-token delivery.

22
MCQmedium

A media company uses a foundation model on Amazon Bedrock to generate article summaries. The model occasionally omits important details. Which prompt engineering technique is most likely to improve completeness?

A.Use a lower temperature setting
B.Increase the max tokens limit
C.Include a list of required key points in the prompt
D.Add 'Be concise' to the prompt
AnswerC

Listing required key points in the prompt explicitly constrains the model's output coverage, directing it to address each specified element rather than relying on its own judgement of relevance. This directly counteracts the omission of important details by making completeness an explicit instruction.

Why this answer

Explicitly listing required key points in the prompt guides the foundation model to cover all specified elements, directly addressing the omission issue. This technique, often called 'constrained generation' or 'structured prompting,' forces the model to attend to each required detail, improving completeness without altering model parameters.

Exam trap

A common misconception is that adjusting model parameters like temperature or token limits can fix content quality issues like missing details. In practice, prompt structure and explicit instructions, such as listing required key points, are the primary tools for controlling output completeness and accuracy.

How to eliminate wrong answers

Option A is wrong because lowering temperature reduces randomness and creativity but does not guarantee inclusion of specific details; it may even make outputs more repetitive or conservative, potentially omitting important points. Option B is wrong because increasing max tokens only allows longer outputs but does not instruct the model what content to include; the model may still omit key details within the expanded token budget. Option D is wrong because adding 'Be concise' encourages brevity, which is counterproductive to completeness and may cause the model to omit even more details.

23
MCQeasy

A company wants to build a chatbot that responds to customer queries using a foundation model. They need low latency and want to avoid managing infrastructure. Which AWS service should they use?

A.Amazon EC2
B.AWS Lambda
C.Amazon Bedrock
D.Amazon SageMaker
AnswerC

Amazon Bedrock provides serverless access to foundation models through a single API, so the company avoids provisioning or scaling infrastructure while obtaining the low-latency inference the chatbot requires. No model hosting or GPU capacity management falls on the customer.

Why this answer

Amazon Bedrock is a fully managed service that provides access to foundation models (FMs) from leading AI providers via a simple API, eliminating the need to manage underlying infrastructure. It is designed for building generative AI applications like chatbots with low latency, as it handles model hosting, scaling, and inference optimization automatically. This makes it the ideal choice for the company's requirement of low-latency responses without infrastructure management.

Exam trap

AWS often tests the misconception that AWS Lambda can handle any serverless workload, but candidates must recognize that Lambda is unsuitable for large model inference due to its execution time, memory, and GPU limitations, whereas Bedrock is purpose-built for foundation model access.

How to eliminate wrong answers

Option A is wrong because Amazon EC2 requires you to provision, configure, and manage virtual servers, including installing and maintaining the foundation model and its dependencies, which contradicts the requirement to avoid managing infrastructure. Option B is wrong because AWS Lambda is a serverless compute service for running short-duration code (up to 15 minutes) and is not designed to host large foundation models; it lacks the GPU support and memory capacity needed for model inference. Option D is wrong because Amazon SageMaker is a machine learning platform that requires you to manage endpoints, instances, and scaling for model deployment, which still involves infrastructure management and does not provide the fully managed, API-based access to foundation models that Bedrock offers.

24
MCQmedium

A developer needs to ensure that a generative AI application on Amazon Bedrock does not produce harmful or inappropriate content. Which feature should they configure?

A.Provisioned Throughput
B.Model Invocation Logging
C.Knowledge Bases for Amazon Bedrock
D.Guardrails for Amazon Bedrock
AnswerD

Guardrails for Amazon Bedrock applies configurable content filters and denied-topic policies to both prompts and model responses, blocking harmful or inappropriate output. This directly satisfies the requirement that the generative AI application must not produce such content.

Why this answer

Guardrails for Amazon Bedrock is the correct feature because it allows developers to define policies that filter and block harmful or inappropriate content in both user inputs and model outputs. This includes configurable thresholds for topics, content filters (e.g., hate, insults, sexual, violence), and denied topics, ensuring the generative AI application adheres to safety and responsible AI requirements.

Exam trap

AWS often tests the distinction between monitoring/logging features (like Model Invocation Logging) and active content filtering features (like Guardrails), leading candidates to mistakenly choose logging as a safety mechanism when it only records data without blocking harmful content.

How to eliminate wrong answers

Option A is wrong because Provisioned Throughput is a pricing and capacity feature that reserves model inference capacity for consistent performance, not a content safety mechanism. Option B is wrong because Model Invocation Logging records API calls and responses for auditing and monitoring, but it does not actively filter or block harmful content. Option C is wrong because Knowledge Bases for Amazon Bedrock enables retrieval-augmented generation (RAG) by connecting to external data sources, but it does not provide content moderation or safety controls.

25
MCQhard

A company is building a model to detect fraudulent transactions. The dataset has 1,000,000 transactions, of which only 1,000 are fraudulent. The team wants to evaluate the model's performance. Which metric is most appropriate to use as the primary evaluation metric?

A.F1 score
B.Recall
C.Accuracy
D.Precision
AnswerA

F1 score is the harmonic mean of precision and recall, providing a balance between the two. In fraud detection, both false positives and false negatives have costs, so a balanced metric is essential. F1 score is robust to class imbalance and is the most appropriate primary metric here.

Why this answer

For highly imbalanced datasets like fraud detection, accuracy is misleading. The F1 score balances precision and recall, making it suitable when both false positives and false negatives matter. It provides a single metric that reflects the model's ability to correctly identify fraud without being overwhelmed by the majority class.

Exam trap

The trap here is choosing accuracy because it is commonly used, but it is deceptive when classes are heavily imbalanced.

26
MCQmedium

A company is using Amazon Bedrock to build a code generation assistant for internal developers. They want to reduce costs by processing batch requests during off-peak hours. Which Bedrock feature should they use?

A.Bedrock Guardrails
B.Bedrock Batch Inference
C.Provisioned Throughput
D.Bedrock Agents
AnswerB

Bedrock Batch Inference processes large volumes of requests asynchronously at reduced cost, ideal for non-urgent code generation jobs run during off-peak hours. It satisfies the cost-reduction constraint by trading immediate responses for lower-priced bulk processing.

Why this answer

B is correct because Amazon Bedrock Batch Inference is specifically designed for processing large volumes of inference requests asynchronously, which allows the company to schedule batch jobs during off-peak hours to reduce costs. This feature leverages lower pricing for asynchronous workloads and is ideal for non-real-time use cases like code generation for internal developers.

Exam trap

The AWS exam often tests the distinction between cost-optimization features (Batch Inference) and performance-guarantee features (Provisioned Throughput), so candidates may mistakenly choose Provisioned Throughput thinking it reduces costs, when it actually increases costs for guaranteed capacity.

How to eliminate wrong answers

Option A is wrong because Bedrock Guardrails is a safety feature for implementing content filters and data privacy controls, not for cost optimization or batch processing. Option C is wrong because Provisioned Throughput provides reserved capacity for consistent, low-latency inference at a fixed cost, which is more expensive and not designed for cost savings via off-peak batch processing. Option D is wrong because Bedrock Agents is a framework for building autonomous agents that orchestrate tasks and APIs, not a feature for batch inference or cost reduction.

27
MCQeasy

A hospital's IT team wants a generative AI assistant that answers patient-billing questions using only the hospital's internal policy documents, without retraining the model. Which approach should they use?

A.Fine-tuning the foundation model on the full set of policy documents
B.Increasing the model's temperature setting so it explores more of its pretrained knowledge
C.Retrieval Augmented Generation (RAG) by querying a vector store of the policy documents and passing retrieved passages to the model
D.Prepending a system instruction telling the model to answer only from hospital policy
AnswerC

RAG retrieves relevant passages from an external knowledge source at inference time and includes them in the prompt, so the model grounds its answer in the hospital's own policies without any weight updates. This directly satisfies the requirement to avoid retraining while keeping answers tied to authoritative internal documents.

Why this answer

Retrieval Augmented Generation supplies the model with relevant, current policy text at query time, so responses are grounded in the hospital's own documents. Fine-tuning changes model weights and needs retraining, temperature alters sampling randomness, and a system prompt alone provides no source content. Only retrieval of the policy passages gives accurate, updatable answers without retraining.

Exam trap

The trap here is assuming that instructing a model to stay on topic is equivalent to giving it the source material it needs to answer accurately.

28
MCQhard

A company operates in a region where Amazon Bedrock is not available. They want to use generative AI but must keep data within the country. Which solution should they consider?

A.Use Amazon SageMaker to host an open-source model in the local region.
B.Wait for Bedrock to become available in their region; there is no alternative.
C.Use Amazon Bedrock in the nearest available region with cross-region inference.
D.Use an API from a third-party generative AI provider with AWS PrivateLink.
AnswerA

SageMaker lets you host an open-source model on instances inside the local region, so training and inference data never leave the country. This satisfies the data-residency constraint that rules out Amazon Bedrock, which is unavailable there.

Why this answer

Amazon SageMaker allows you to host open-source models (e.g., Llama 2, Falcon) in any AWS region, including those where Bedrock is unavailable. This satisfies the data residency requirement because the model and data never leave the local region. SageMaker provides full control over the infrastructure, enabling compliance with local data sovereignty laws.

Exam trap

The trap here is that candidates assume Bedrock is the only AWS generative AI service, overlooking SageMaker's capability to host open-source models, which is a common misconception tested in the AIF-C01 exam.

How to eliminate wrong answers

Option B is wrong because waiting for Bedrock availability is unnecessary; SageMaker offers a viable alternative today. Option C is wrong because cross-region inference would send data outside the required country boundary, violating the data residency constraint. Option D is wrong because using a third-party API, even with AWS PrivateLink, still involves data leaving the AWS network to an external provider, which may not guarantee data remains within the country.

29
MCQmedium

A team is evaluating two different foundation models for a sentiment analysis task. They have a labeled test dataset. Which evaluation approach should they use to compare the models' performance on this task?

A.Use ROUGE scores to compare model outputs
B.Run a human evaluation study with 100 judges
C.Use BLEU scores to measure n-gram overlap with reference labels
D.Compute accuracy, precision, recall, and F1-score on the labeled test set
AnswerD

A labelled test set supports supervised classification metrics, so computing accuracy, precision, recall, and F1-score quantifies each model's predictions against ground-truth sentiment labels. This gives an objective, comparable measurement of both models on the same task.

Why this answer

Accuracy, precision, recall, and F1-score are standard classification metrics that directly measure how well a model's predicted sentiment labels match the ground truth labels in a labeled test dataset. These metrics provide a quantitative, reproducible comparison of model performance on a supervised sentiment analysis task, unlike text-generation metrics or subjective human evaluation.

Exam trap

AWS often tests the distinction between metrics for generative tasks (ROUGE, BLEU) versus classification tasks (accuracy, precision, recall, F1), leading candidates to mistakenly apply text-generation metrics to a classification problem.

How to eliminate wrong answers

Option A is wrong because ROUGE scores are designed for evaluating text summarization by measuring n-gram recall between generated and reference summaries, not for comparing classification outputs like sentiment labels. Option B is wrong because while human evaluation can assess subjective quality, it is not the most efficient or objective approach for comparing two models on a labeled test set; it introduces variability and cost, and the question explicitly asks for an evaluation approach using the labeled dataset. Option C is wrong because BLEU scores measure n-gram precision for machine translation and are not appropriate for sentiment analysis, which requires classification metrics rather than sequence overlap with reference texts.

30
MCQhard

A machine learning team notices their model performs excellently on the training dataset but poorly on new, unseen data. They want to reduce this gap without collecting more data. Which action most directly addresses the problem?

A.Apply regularization techniques such as L2 penalty or dropout
B.Increase model complexity by adding more layers and parameters
C.Train for many more epochs until training loss approaches zero
D.Remove the validation dataset and evaluate only on training data
AnswerA

Regularization constrains the model so it cannot fit training noise as easily, which improves generalization to unseen data. L2 penalties shrink weights and dropout randomly deactivates units during training, both discouraging memorization. This directly targets the overfitting gap described without requiring additional data collection.

Why this answer

The described pattern is classic overfitting: strong training performance with weak performance on unseen data. Regularization such as L2 penalties or dropout constrains the model and improves generalization. Adding capacity, training longer, or discarding validation all worsen or conceal the issue instead of correcting it.

Exam trap

The trap here is assuming that a model performing better on training data will automatically perform better on new data.

31
MCQhard

A company wants to adapt a foundation model for a custom domain with very limited labeled data and minimal cost. Which approach is most suitable?

A.Pre-training from scratch
B.Prompt engineering with few-shot examples
C.Reinforcement learning from human feedback
D.Full fine-tuning
AnswerB

Few-shot prompting supplies a handful of labelled examples within the prompt itself, letting the foundation model adapt to the custom domain without weight updates. This satisfies both constraints: very limited labelled data and minimal cost, since no fine-tuning compute is required.

Why this answer

Prompt engineering with few-shot examples is the most suitable approach because it allows the company to adapt a foundation model to a custom domain using very limited labeled data and minimal cost. By providing a few input-output examples directly in the prompt, the model can infer the desired task without any weight updates, making it efficient for low-resource scenarios.

Exam trap

AWS often tests the misconception that full fine-tuning is always the best way to adapt a model, but the trap here is that candidates overlook the cost and data requirements, failing to recognize that prompt engineering with few-shot examples is the most efficient when labeled data is scarce and budget is tight.

How to eliminate wrong answers

Option A is wrong because pre-training from scratch requires massive amounts of unlabeled data and significant computational resources, which contradicts the requirement for very limited labeled data and minimal cost. Option C is wrong because reinforcement learning from human feedback (RLHF) requires a large dataset of human preferences and multiple model training iterations, making it costly and data-intensive. Option D is wrong because full fine-tuning updates all model weights, which requires a substantial labeled dataset and significant compute, and is not suitable when labeled data is very limited.

32
MCQhard

A company is building a RAG application with Amazon Bedrock Knowledge Bases. They want to ensure that the retriever returns the most semantically relevant chunks. They are using a large document corpus with many similar passages. Which chunking strategy is MOST likely to improve retrieval accuracy?

A.Fixed‑size token chunking with no overlap
B.Overlapping fixed‑size chunks with 50% overlap
C.Semantic chunking that splits at paragraph or section boundaries
D.Very small chunks (50 tokens) to maximize granularity
AnswerC

Semantic chunking splits text where meaning shifts, using embedding similarity between adjacent sentences to detect paragraph or section boundaries. This preserves coherent, self-contained passages, so the retriever's vector comparison distinguishes between the many similar passages in the corpus, directly improving semantic relevance of returned chunks.

Why this answer

Semantic chunking splits documents at natural boundaries like paragraphs or sections, preserving the coherence of each chunk. This ensures that each chunk contains a complete, self-contained idea, which allows the retriever to match the semantic meaning of the query more accurately. In a corpus with many similar passages, this approach reduces noise and improves the relevance of retrieved chunks.

Exam trap

A common misconception is that more granularity (smaller chunks) always improves retrieval accuracy, but the trap here is that overly small chunks lose context and semantic completeness, which actually degrades relevance in a RAG system.

How to eliminate wrong answers

Option A is wrong because fixed-size token chunking with no overlap can cut sentences or ideas in half, breaking semantic coherence and causing the retriever to miss relevant context. Option B is wrong because overlapping fixed-size chunks still suffer from arbitrary boundaries that can split concepts, and the overlap introduces redundancy without guaranteeing semantic completeness. Option D is wrong because very small chunks (50 tokens) maximize granularity but often lack sufficient context to capture the full meaning of a passage, leading to poor semantic matching and increased retrieval of irrelevant fragments.

33
MCQhard

A company wants to use a rules-based approach to approve loan applications but finds that it cannot keep up with changing regulations. They have historical data with decisions. Which approach should they adopt?

A.Train a supervised machine learning model on historical decisions
B.Continue refining the rules-based system manually
C.Implement a deep neural network without feature engineering
D.Use unsupervised learning to cluster applications
AnswerA

A supervised model learns the mapping from applicant features to historical approve or decline decisions, capturing patterns that rules cannot express. Retraining on fresh data lets the model adapt as regulations shift, addressing the rigidity that made the rules-based approach unworkable.

Why this answer

Supervised machine learning can automatically learn patterns from historical loan decision data, adapting to changing regulations without manual rule updates. By training a model on labeled examples of approved and rejected applications, the system can generalize to new cases and adjust as the underlying regulatory logic shifts over time, which is precisely the limitation of a static rules-based approach.

Exam trap

AWS often tests the distinction between supervised and unsupervised learning by presenting a scenario where historical labels exist, tempting candidates to choose unsupervised clustering (Option D) because they confuse 'grouping similar applications' with 'predicting approval decisions.'

How to eliminate wrong answers

Option B is wrong because manually refining a rules-based system is exactly what the company finds unsustainable—it cannot keep up with changing regulations, and this approach does not leverage the historical data to automate adaptation. Option C is wrong because implementing a deep neural network without feature engineering is impractical for tabular loan data; deep learning typically requires large datasets and careful feature extraction, and it is overkill for a problem that can be solved with simpler supervised models. Option D is wrong because unsupervised learning clusters applications without using the historical approval labels, so it cannot predict whether a new application should be approved—it lacks the target variable needed for decision-making.

34
MCQmedium

A healthcare company is training a model on sensitive patient data using Amazon SageMaker. They need to ensure that individual patient data cannot be reverse-engineered from the model. Which technique should they implement during training?

A.Data encryption at rest
B.AWS Identity and Access Management (IAM) policies
C.Differential privacy
D.SageMaker Model Monitor
AnswerC

Differential privacy adds calibrated noise during training, bounding any single patient's influence on learned parameters. This satisfies the requirement that individual records cannot be reverse-engineered, providing a formal privacy guarantee rather than mere access control or encryption at rest.

Why this answer

Differential privacy is the technique specifically designed to prevent reverse-engineering of individual records from a trained model by injecting calibrated statistical noise during training. It provides a mathematical guarantee that the presence or absence of any single patient's data has a bounded effect on the model's output, directly addressing the requirement that individual patient data cannot be inferred. This is the only option that operates at the training algorithm level to protect individual records.

Exam trap

AIF-C01 often tests the distinction between data protection mechanisms, and candidates confuse encryption or access control (which protect data at rest/in transit) with differential privacy (which protects against inference from the model itself).

How to eliminate wrong answers

Option A is wrong because encryption at rest protects data on disk but does nothing to prevent a trained model from memorizing and leaking individual records through inference. Option B is wrong because IAM policies control who can access AWS resources, not whether the model itself leaks training data — access control is orthogonal to model privacy. Option D is wrong because SageMaker Model Monitor detects data drift and quality issues in deployed models; it does not provide any privacy guarantee against training data extraction.

35
Multi-Selecthard

A healthcare company is building an application on Amazon Bedrock that uses a foundation model to answer patient questions about medications. The company must reduce the risk of harmful or inaccurate medical advice. Which TWO strategies should they implement? (Choose two.)

Select 2 answers
A.Use Amazon Bedrock Guardrails to filter harmful content and define denied topics related to specific medical advice.
B.Disable all logging so patient interactions are not stored.
C.Fine-tune the model on a small set of patient questions and answers collected from previous chats.
D.Ground responses in approved clinical guidelines by using a knowledge base with Retrieval Augmented Generation.
E.Increase the model's temperature to make answers more diverse and comprehensive.
AnswersA, D

Amazon Bedrock Guardrails can block harmful content and enforce denied topics, which directly addresses the requirement to reduce harmful or inaccurate medical advice. By configuring topic denial and content filters, the application can prevent the model from responding to restricted medical queries. This provides a managed safety layer without changing the underlying model, and it can be applied consistently across invocations.

Why this answer

Guardrails provide a managed safety layer that filters harmful content and blocks denied medical topics, while RAG grounds responses in approved clinical guidelines so answers reflect authoritative sources. Together they reduce both harmful outputs and factual inaccuracies, which is essential for a patient-facing medication assistant.

Exam trap

The trap here is treating model tuning or temperature changes as safety controls, when the primary mitigations are content filtering and grounding in approved sources.

36
Multi-Selectmedium

A company is deploying a customer‑facing chatbot using Amazon Bedrock. They need to ensure the chatbot never reveals personally identifiable information (PII) and refuses to discuss the topic of 'employee salaries'. Which TWO Bedrock Guardrails features should they configure together? (Select TWO.)

Select 2 answers
A.Topic denial
B.Grounding check
C.PII detection and redaction
D.Bedrock Knowledge Base
E.Content filtering
AnswersA, C

Topic denial defines denied subjects and blocks any prompt or response touching them, so configuring it with PII handling makes the chatbot refuse salary discussions. It satisfies the explicit refusal requirement for the 'employee salaries' topic.

Why this answer

Option A (Topic denial) is correct because Amazon Bedrock Guardrails' denied topics feature lets you define a custom topic such as 'employee salaries' with a name, definition, and example phrases, and the guardrail will block user inputs and model responses that fall into that topic, directly satisfying the requirement to refuse discussing salaries. Option C (PII detection and redaction) is correct because Guardrails' sensitive information filters can detect and either block or mask PII entities (such as NAME, EMAIL, PHONE, SSN, and ADDRESS) in both prompts and completions, ensuring the chatbot never reveals personally identifiable information. Option B (Grounding check) is not correct here because grounding and relevance checks only validate whether responses are supported by and relevant to a provided source (RAG) context, which does not prevent PII leakage or off-topic salary discussions.

Option D (Bedrock Knowledge Base) is not correct because a knowledge base is a managed RAG resource for retrieving source content, not a guardrail feature for blocking topics or redacting PII. Option E (Content filtering) is not correct because content filters target harmful categories such as hate, violence, sexual, insults, and misconduct, not the specific business rules of PII redaction and salary-topic denial.

Exam trap

AIF-C01 often tests whether candidates can distinguish Bedrock Guardrails policies (topic denial, PII redaction, content filters, grounding) from adjacent Bedrock features like Knowledge Bases — picking a non-guardrail feature is the classic mistake.

37
Multi-Selectmedium

A company uses Amazon Bedrock with Anthropic Claude for a question-answering system. They want to reduce costs while maintaining acceptable latency. Which TWO actions would help achieve this? (Choose two.)

Select 2 answers
A.Use a smaller model variant (e.g., Claude Instant instead of Claude)
B.Switch from text generation to image generation
C.Enable prompt caching for frequently used system prompts
D.Increase the context window to include more documents
E.Use a vector database to pre-filter documents before sending to the model
AnswersA, C

Claude Instant processes tokens faster and costs less per token than larger Claude variants, directly lowering both inference spend and response latency. It satisfies the cost-reduction goal provided the smaller model still meets the question-answering quality the system requires.

Why this answer

Option A is correct because choosing a smaller, cheaper model variant such as Claude Instant instead of a larger Claude model directly lowers the per-token inference cost while still providing acceptable latency for question-answering workloads. Option C is correct because enabling prompt caching for frequently reused system prompts avoids re-processing the same input tokens on every request, reducing both token costs and latency for repeated prompt prefixes. Option B is wrong because switching to image generation changes the workload entirely and does not reduce the cost of a text-based question-answering system.

Option D is wrong because increasing the context window adds more input tokens, which raises cost and can increase latency rather than reduce them. Option E is wrong because using a vector database to pre-filter documents is a retrieval optimization that can improve relevance, but it does not by itself reduce Bedrock model inference costs or guarantee lower latency.

Exam trap

The trap is that candidates may think increasing context window (D) or using a vector database (E) are cost-effective, but both can increase latency. Vector database retrieval adds overhead, and prompt caching is often overlooked as a cost-saving technique.

38
MCQmedium

A company is developing a speech-to-text application for a diverse user base. To ensure inclusive design, they test the model with different accents and dialects. They find that error rates are higher for certain accents. Which responsible AI principle is most directly violated?

A.Robustness
B.Fairness
C.Veracity
D.Privacy and security
AnswerB

Higher error rates for particular accents constitute disparate performance across demographic groups, which is precisely the harm the fairness principle addresses. Inclusive design testing exists to surface such bias, and the stem's finding that accuracy varies by accent directly satisfies fairness's requirement of equitable outcomes regardless of accent or dialect.

Why this answer

Fairness in responsible AI refers to avoiding bias and ensuring equitable performance across different demographic groups. Higher error rates for certain accents indicate that the model performs unequally across user groups, which is a direct violation of fairness. This is a classic example of algorithmic bias in speech recognition systems.

Exam trap

AIF-C01 often tests the distinction between fairness and robustness — candidates may choose robustness because the issue involves error rates, but the key signal is the disparity across demographic groups, which is fairness.

How to eliminate wrong answers

Option A is wrong because robustness refers to a system's ability to maintain performance under adverse or unexpected conditions (e.g., noisy environments, adversarial inputs), not to equitable performance across demographic groups. Option C is wrong because veracity refers to truthfulness and accuracy of information — while error rates relate to accuracy, the disparity across accents is fundamentally a fairness issue, not a truthfulness issue. Option D is wrong because privacy and security concerns data protection and unauthorized access, not performance disparities across user groups.

39
MCQmedium

A data scientist is designing a RAG pipeline using Amazon Bedrock Knowledge Bases. They need to store embeddings of document chunks and perform similarity searches. Which vector store is a serverless option that integrates directly with Bedrock Knowledge Bases?

A.Amazon OpenSearch Serverless
B.Amazon Aurora with pgvector
C.MongoDB Atlas
D.Pinecone
AnswerA

Amazon OpenSearch Serverless provides a fully managed, serverless vector engine that Bedrock Knowledge Bases supports as a native vector store, so embeddings are indexed and queried without provisioning clusters. This satisfies the stem's serverless constraint, unlike self-managed OpenSearch or provisioned alternatives requiring capacity planning.

Why this answer

Amazon Bedrock Knowledge Bases natively integrates with Amazon OpenSearch Serverless as a vector store for storing embeddings and performing similarity searches. OpenSearch Serverless is a fully serverless option that automatically scales and requires no infrastructure management, making it the correct choice for a serverless vector store that works directly with Bedrock Knowledge Bases.

Exam trap

The trap here is that candidates may confuse 'serverless' with 'managed' and select a third-party option like Pinecone or MongoDB Atlas, which are managed but not natively integrated with Bedrock Knowledge Bases, or choose Aurora with pgvector thinking its serverless variant qualifies, but Bedrock Knowledge Bases does not support it as a direct vector store integration.

How to eliminate wrong answers

Option B is wrong because Amazon Aurora with pgvector is a relational database that supports vector embeddings but is not serverless (Aurora Serverless is available but not directly integrated with Bedrock Knowledge Bases as a vector store). Option C is wrong because MongoDB Atlas is a third-party managed database that requires separate configuration and is not natively integrated with Bedrock Knowledge Bases. Option D is wrong because Pinecone is a third-party vector database that is not serverless in the AWS context and does not have direct native integration with Bedrock Knowledge Bases.

40
MCQeasy

A machine learning engineer wants to detect if sensitive data, such as personally identifiable information (PII), exists in a training dataset stored in S3 before training a model. Which AWS service should they use?

A.Amazon Inspector
B.Amazon Macie
C.Amazon GuardDuty
D.AWS Config
AnswerB

Macie uses managed data identifiers and pattern matching to scan S3 objects and report PII such as names, addresses and credit card numbers. It satisfies the requirement to detect sensitive data in the training dataset before training begins, without modifying the data.

Why this answer

Amazon Macie uses machine learning to automatically discover, classify, and protect sensitive data in S3 buckets.

41
MCQeasy

A company is using Amazon Bedrock to generate marketing copy. They want to ensure the model's responses are factually accurate and grounded in their proprietary knowledge base. Which feature should they use?

A.Model customization
B.Fine-tuning
C.Retrieval Augmented Generation (RAG)
D.Prompt engineering
AnswerC

Retrieval Augmented Generation queries the proprietary knowledge base and injects relevant retrieved passages into the prompt, grounding responses in the company's own content. This satisfies the factual-accuracy requirement by supplying authoritative context the base model lacks, reducing hallucination.

Why this answer

Retrieval Augmented Generation (RAG) is the correct choice because it retrieves relevant documents from the company's proprietary knowledge base and provides them as context to the foundation model at inference time. This grounds the model's responses in factual, up-to-date information without modifying the underlying model weights, ensuring accuracy and reducing hallucinations.

Exam trap

A common misconception is that you must fine-tune or customize a model to incorporate proprietary knowledge. However, with Amazon Bedrock, RAG allows you to ground responses in your knowledge base without retraining, which is more cost-effective and keeps information current.

How to eliminate wrong answers

Option A is wrong because model customization (e.g., using Amazon Bedrock's Custom Model Import or training a new base model) changes the model's weights and behavior, but it does not inherently ground responses in a specific knowledge base; it requires large amounts of labeled data and can still hallucinate. Option B is wrong because fine-tuning adjusts model parameters on a dataset, which can embed knowledge but is static, expensive, and does not allow dynamic retrieval of proprietary information; it also risks overfitting or catastrophic forgetting. Option D is wrong because prompt engineering only modifies the input prompt to guide the model's output; it does not inject external factual data and cannot guarantee grounding in a proprietary knowledge base, as the model relies solely on its pre-trained parameters.

42
MCQhard

A financial services firm wants an internal assistant that answers employee questions about company travel policy in natural language. The policy documents change frequently, and the firm requires answers to cite the specific policy section used. Which approach BEST meets these requirements?

A.Deploy a keyword search engine that returns matching policy paragraphs and rely on employees to interpret the results themselves.
B.Fine-tune a large language model on the full policy corpus so it memorizes the documents and answers from memory.
C.Use retrieval-augmented generation, where relevant policy passages are retrieved and supplied to the model to ground its answer and source citation.
D.Train a small classification model that maps each employee question to one of the fixed policy section titles.
AnswerC

Retrieval-augmented generation fetches the most relevant passages from the current policy store at query time and passes them to the language model as context. The model then answers based on those passages and can reference the specific section, satisfying the citation requirement. Because the knowledge lives in the document store, updating policies requires only re-indexing the changed documents, not retraining.

Why this answer

Retrieval-augmented generation grounds each answer in freshly retrieved policy passages, so responses reflect the current documents and can point to the exact section used. Because updates only require re-indexing changed content rather than retraining a model, it also handles the frequently changing policy corpus efficiently.

Exam trap

The trap here is assuming that fine-tuning a model on documents is the way to make it answer about those documents, when grounding through retrieval is what enables current content and citations.

43
MCQeasy

A company is using Amazon Bedrock to generate responses for customer support. They want to ensure that the model does not expose personally identifiable information (PII) in its outputs. Which AWS feature can be configured to automatically redact PII from model responses?

A.Amazon Macie
B.Amazon SageMaker Model Monitor
C.Amazon Bedrock Guardrails
D.AWS CloudTrail
AnswerC

Amazon Bedrock Guardrails applies configurable sensitive-information filters that detect and automatically redact PII such as names, addresses and card numbers from model responses. This satisfies the requirement to prevent PII exposure in outputs without altering the underlying foundation model itself.

Why this answer

Amazon Bedrock Guardrails is the correct choice because it provides configurable policies that can automatically detect and redact personally identifiable information (PII) from model inputs and outputs. This feature is specifically designed for Amazon Bedrock to enforce content safety and compliance requirements, including PII redaction, without requiring custom code or external services.

Exam trap

The trap here is that candidates may confuse Amazon Macie (a data discovery service for S3) with a real-time content filtering capability, or assume that SageMaker Model Monitor can be applied to Bedrock, when in fact only Bedrock Guardrails provides native PII redaction for model responses.

How to eliminate wrong answers

Option A is wrong because Amazon Macie is a data security service that discovers and protects sensitive data in Amazon S3, not a feature for redacting PII from model responses in Amazon Bedrock. Option B is wrong because Amazon SageMaker Model Monitor detects data drift and model quality issues for SageMaker endpoints, not for Bedrock, and does not perform PII redaction. Option D is wrong because AWS CloudTrail records API activity for auditing and governance, not for modifying or filtering model responses.

44
MCQhard

A company uses a Bedrock Agent to handle customer support tickets. The agent needs to look up order status from a legacy API that requires authentication. The agent should also escalate to a human if the query is not supported. Which combination of components should the developer configure?

A.Configure a Guardrail to block unsupported queries and return a static response.
B.Define an action group with a Lambda function that calls the legacy API, and configure the agent instructions to escalate when it cannot answer.
C.Use a Bedrock Knowledge Base to store order status information and rely on RAG to answer.
D.Fine-tune the model on historical tickets so it can guess order status.
AnswerB

An action group with a Lambda function gives the agent a callable tool to query the legacy API, including authentication handled inside the function. Agent instructions then define the escalation path when no action matches the query.

Why this answer

Bedrock Agents use action groups to invoke external APIs; defining an action group with a Lambda function lets the agent call the legacy API with authentication handled inside Lambda. Configuring agent instructions to escalate when the agent cannot answer satisfies the human-handoff requirement. This combination directly addresses both the API integration and the escalation behavior.

Exam trap

AIF-C01 often tests the misconception that a Knowledge Base or fine-tuning can replace live API integration, when in fact real-time authenticated data requires an action group backed by Lambda.

How to eliminate wrong answers

Option A is wrong because Guardrails filter or block content based on policies; they do not provide API integration or dynamic escalation logic, and a static response does not handle order lookups. Option C is wrong because a Knowledge Base with RAG is for retrieving from indexed documents, not for calling a live authenticated API that returns real-time order status. Option D is wrong because fine-tuning teaches style and patterns, not live data retrieval; the model cannot 'guess' current order status accurately and fine-tuning does not provide API access.

45
MCQhard

A financial services firm must extract text from scanned loan application forms to automate data entry. The forms are in various languages. Which combination of AWS AI services should be used?

A.Amazon Transcribe and Amazon Polly
B.Amazon Kendra and Amazon Lex
C.Amazon Textract and Amazon Translate
D.Amazon Rekognition and Amazon Comprehend
AnswerC

Amazon Textract uses OCR to extract text and form data from scanned documents, while Amazon Translate handles the multilingual requirement by translating the extracted content. Together they satisfy both the scanned-form extraction and varied-language constraints.

Why this answer

Amazon Textract is designed to extract text and data from scanned documents, including forms, using OCR (Optical Character Recognition). Amazon Translate then converts the extracted text from various languages into a target language, enabling automated data entry across multilingual loan applications. Together, they directly address the requirement of extracting text from scanned forms in multiple languages.

Exam trap

AWS often tests the distinction between OCR-based extraction (Textract) and image analysis (Rekognition), leading candidates to mistakenly choose Rekognition for document text extraction when it is designed for object/face detection, not form data extraction.

How to eliminate wrong answers

Option A is wrong because Amazon Transcribe converts speech to text (not text from scanned documents) and Amazon Polly converts text to speech, neither of which extracts text from scanned forms. Option B is wrong because Amazon Kendra is an intelligent search service and Amazon Lex builds conversational chatbots; neither performs OCR or text extraction from scanned documents. Option D is wrong because Amazon Rekognition analyzes images and videos for objects/faces (not form text extraction) and Amazon Comprehend performs NLP on text (not extraction from scanned documents), so they cannot extract structured data from scanned forms.

46
MCQmedium

A data scientist is using SageMaker to train a model on a dataset with many features. They suspect some features are redundant. Which feature engineering technique would help?

A.Feature scaling
B.One-hot encoding
C.Principal Component Analysis (PCA)
D.Polynomial features
AnswerC

Principal Component Analysis projects the many correlated features onto a smaller set of orthogonal components capturing most variance, eliminating redundancy. This directly addresses the stem's suspicion of redundant features by reducing dimensionality while retaining information.

Why this answer

Principal Component Analysis (PCA) is a dimensionality reduction technique that transforms the original correlated features into a smaller set of uncorrelated principal components, effectively removing redundancy while preserving most of the variance in the data. In SageMaker, PCA can be applied via the built-in PCA algorithm or as a preprocessing step in a scikit-learn container to reduce feature space and eliminate multicollinearity.

Exam trap

The AIF-C01 exam often tests the distinction between feature reduction (PCA) and feature transformation (scaling, encoding, polynomial expansion) to see if candidates confuse techniques that change feature count versus those that only change feature values.

How to eliminate wrong answers

Option A is wrong because feature scaling (e.g., StandardScaler, MinMaxScaler) normalizes the range of features but does not remove redundant or correlated features; it only changes the scale. Option B is wrong because one-hot encoding is used to convert categorical variables into numerical format, not to address feature redundancy among many continuous or numerical features. Option D is wrong because polynomial features create interaction and higher-order terms, which actually increase the number of features and can introduce more redundancy, not reduce it.

47
MCQmedium

Refer to the exhibit. An AWS customer runs SageMaker Clarify to evaluate bias in their training data. The report shows multiple metrics with status 'violated'. What should the customer do next?

A.Use data augmentation to balance the dataset
B.Reduce the number of features
C.Retrain the model with more data
D.Ignore the metrics because thresholds are too strict
AnswerA

Data augmentation can balance representation.

Why this answer

SageMaker Clarify bias metrics such as Class Imbalance (CI) or Difference in Positive Proportions in Labels (DPPL) flag potential bias in the training data or model predictions. When a report shows violations, the next step is to review the findings and apply a targeted mitigation. Among the options, balancing the dataset through data augmentation directly addresses the representative imbalance; reducing features or simply adding more data does not target the demographic imbalance, and ignoring the metrics is not appropriate.

Exam trap

A common misconception is that adding more data will automatically reduce bias. Without addressing the specific imbalance or bias source, adding data can amplify existing disparities. The correct approach is to use Clarify's metrics to guide targeted mitigation, such as balancing the dataset.

How to eliminate wrong answers

Option B is wrong because reducing the number of features does not fix bias in the training data labels or target distribution; feature reduction may even remove attributes needed to detect bias, and bias violations typically stem from label imbalance, not feature count. Option C is wrong because simply adding more data without addressing the underlying imbalance (e.g., collecting more data from the majority class) can worsen bias metrics; the key is to ensure the new data is balanced across sensitive groups. Option D is wrong because ignoring violated metrics violates the core principle of responsible AI; SageMaker Clarify thresholds are configurable but should be set based on business and ethical requirements, not arbitrarily dismissed.

48
MCQhard

Refer to the exhibit. An IAM policy is attached to a user. Which models can the user invoke?

A.Only Claude v2
B.No models
C.Claude v2 and any model with a name containing 'claude'
D.Any model in the account
AnswerA

The policy's Allow statement names only the Claude v2 model resource, so invocation is scoped to that single model. Other foundation models lack an explicit Allow, and IAM denies by default, satisfying the exhibit's constraint that permissions are granted per model ARN rather than account-wide.

Why this answer

The IAM policy explicitly allows the `bedrock:InvokeModel` action only on the resource ARN `arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-v2`. This means the user can invoke only the Claude v2 model. No other models, including other Claude versions or any model with 'claude' in its name, are permitted because the resource ARN is exact and does not use wildcards.

Exam trap

The AIF-C01 exam often tests the distinction between an exact resource ARN and a wildcard pattern; candidates mistakenly assume that 'claude' in the model ID implies all Claude models are allowed, but without a wildcard, only the exact model specified is permitted.

How to eliminate wrong answers

Option B is wrong because the policy does allow invocation of Claude v2, so the user can invoke at least one model. Option C is wrong because the policy uses an exact resource ARN (`anthropic.claude-v2`), not a wildcard pattern like `*claude*`; models with 'claude' in their name but not exactly `claude-v2` are not allowed. Option D is wrong because the policy restricts invocation to a single specific model, not any model in the account.

49
MCQmedium

A data science team needs to choose a machine learning approach for a project that requires predicting customer churn based on historical data. The team has a labeled dataset with 10,000 records and needs to interpret the model's decisions to provide business insights. Which machine learning technique should the team prioritize?

A.Random forest.
B.K-means clustering.
C.Linear regression.
D.Deep neural network with multiple hidden layers.
AnswerA

Random forest is supervised and handles labelled tabular data well, and its ensemble of decision trees exposes feature importance values. Those importances give the interpretable business insight the team needs, unlike opaque deep learning or unsupervised clustering.

Why this answer

Random forest is the best choice because it handles the classification task of predicting churn (a binary outcome) from a labeled dataset, provides feature importance scores for interpretability, and works well with 10,000 records without overfitting due to its ensemble of decision trees. Its built-in ability to rank input features directly supports the team's need to derive business insights from model decisions.

Exam trap

The AWS AI Practitioner exam often tests the distinction between supervised and unsupervised learning, and the trap here is that candidates might choose a powerful but opaque model like a deep neural network, overlooking the explicit requirement for interpretability and the modest dataset size that favors simpler, more explainable ensemble methods.

How to eliminate wrong answers

Option B (K-means clustering) is wrong because it is an unsupervised learning algorithm used for grouping unlabeled data, not for predicting a labeled target like churn. Option C (Linear regression) is wrong because it is designed for regression tasks predicting continuous values, not for binary classification like churn prediction. Option D (Deep neural network with multiple hidden layers) is wrong because, while it can handle classification, it requires large datasets (typically >100,000 records) to avoid overfitting and offers poor interpretability, making it unsuitable for deriving clear business insights from a 10,000-record dataset.

50
MCQhard

A retail company uses Amazon Bedrock with Anthropic Claude to generate personalized marketing emails. They want to include dynamic content such as the customer's name and recent purchase history. Which API should they use to enable multi-turn conversations with context management?

A.Amazon Comprehend API
B.Converse API
C.InvokeModel API
D.Amazon Lex API
AnswerB

Converse API provides a unified, model-agnostic interface with built-in conversation history handling, so prior turns are passed as structured messages. This maintains context across turns, enabling dynamic personalisation such as customer names and purchase history in generated emails.

Why this answer

The Converse API (option B) is the correct choice because it is specifically designed for multi-turn conversations with context management in Amazon Bedrock. It automatically handles conversation history, allowing the model to maintain context across multiple exchanges, which is essential for generating personalized marketing emails that reference dynamic content like customer names and purchase history from previous interactions.

Exam trap

AWS often tests the distinction between low-level (InvokeModel) and high-level (Converse) APIs, where candidates mistakenly choose InvokeModel because they think it offers more control, but they overlook that Converse is purpose-built for multi-turn conversations with built-in context management.

How to eliminate wrong answers

Option A is wrong because Amazon Comprehend is a natural language processing (NLP) service for extracting insights like sentiment or entities from text, not for generating responses or managing multi-turn conversations. Option C is wrong because the InvokeModel API is a lower-level API that requires you to manually manage conversation history and context, making it less suitable for multi-turn interactions without additional overhead. Option D is wrong because Amazon Lex is a service for building conversational interfaces (chatbots) with its own intent and slot management, but it does not integrate with Bedrock's foundation models like Anthropic Claude for generating personalized email content.

51
MCQeasy

Refer to the exhibit. An ML team finds that their training data is stored in two subfolders under s3://my-bucket/train/. They need to ensure that the dataset is balanced for training a classification model. What should they do?

A.Use AWS Glue to create a balanced dataset
B.Use Amazon Rekognition custom labels
C.Count the number of files in each subfolder and resample
D.Enable versioning on the bucket
AnswerC

Class imbalance is the constraint: file counts per subfolder reveal each class's representation. Counting then resampling — oversampling the minority or undersampling the majority — equalises class distribution before training, directly satisfying the balance requirement. File counts act as a proxy for label frequency in this image-classification setup.

Why this answer

The core requirement is to balance the dataset by ensuring an equal number of samples from each subfolder (class). Counting the files in each subfolder and then resampling (e.g., undersampling the majority class or oversampling the minority class) directly addresses class imbalance. This is a standard data preprocessing step before training a classification model, and it does not require any additional AWS services beyond the existing S3 storage.

Exam trap

AWS exams often test the misconception that a fully managed AI service (like Rekognition or Glue) can automatically handle dataset imbalance, when in reality the responsibility for data preprocessing and balancing lies with the ML practitioner.

How to eliminate wrong answers

Option A is wrong because AWS Glue is a serverless data integration service for ETL (Extract, Transform, Load) jobs, not a tool specifically designed for dataset balancing or resampling; using Glue for this simple counting and resampling task would be overkill and inefficient. Option B is wrong because Amazon Rekognition Custom Labels is a managed service for training custom image classification models, but it does not provide a mechanism to balance an existing dataset stored in S3; it expects the user to provide a balanced dataset as input. Option D is wrong because enabling versioning on the S3 bucket only preserves, retrieves, and restores every version of every object; it does not affect the distribution of files across subfolders or help in balancing the dataset.

52
MCQhard

An organization uses Amazon Bedrock to generate content. They have implemented guardrails to block toxic content. However, some users are able to bypass the guardrails by encoding their prompts. What step should be taken to improve security?

A.Encode the prompts before sending to the model.
B.Enable prompt injection detection in the guardrail configuration.
C.Use a different foundation model that is less susceptible.
D.Restrict access to the model using IAM policies.
AnswerB

Prompt injection detection can identify and block encoded or malicious prompts.

Why this answer

Amazon Bedrock guardrails include a built-in prompt injection detection capability that can identify and block attempts to bypass content filters through encoded or obfuscated prompts. Enabling this feature specifically addresses the scenario where users encode their inputs to evade toxic content blocking, as it analyzes the decoded intent of the prompt rather than just the surface-level encoding.

Exam trap

The AIF-C01 exam often tests the misconception that encoding or encrypting inputs is a security measure, when in reality it is a common bypass technique that must be countered by content inspection mechanisms like prompt injection detection.

How to eliminate wrong answers

Option A is wrong because encoding the prompts before sending them to the model would not improve security; it would actually compound the problem by further obfuscating the input, making it harder for guardrails to detect toxic content. Option C is wrong because the susceptibility to encoded prompts is not a model-specific vulnerability; it is a function of the input processing layer, and switching foundation models would not prevent encoding-based bypasses. Option D is wrong because restricting access with IAM policies controls who can invoke the model but does not inspect or sanitize the content of prompts, so it cannot prevent users from submitting encoded toxic inputs.

53
MCQmedium

A machine learning practitioner is building a binary classifier for a medical diagnosis application. The cost of a false negative (missing a disease) is very high. Which evaluation metric should the team emphasize?

A.Accuracy
B.F1 score
C.Recall
D.Precision
AnswerC

Recall measures the proportion of actual positives correctly identified, directly minimising false negatives. In medical diagnosis, where missing a disease carries severe consequences, maximising recall ensures fewer diseased patients are wrongly classified as healthy. Precision would instead penalise false positives, which matter less here than the constraint of avoiding missed diagnoses.

Why this answer

Recall (sensitivity) measures the proportion of actual positives correctly identified, which is critical when the cost of false negatives is high, as in medical diagnosis where missing a disease could have severe consequences. By maximizing recall, the model minimizes false negatives, ensuring that most patients with the disease are detected, even at the expense of more false positives.

Exam trap

AWS often tests the distinction between recall and precision in high-stakes scenarios, and the trap here is that candidates may choose F1 score thinking it balances both metrics, not realizing that when false negatives are the primary concern, recall should be the emphasized metric over a balanced measure.

How to eliminate wrong answers

Option A is wrong because accuracy considers both true positives and true negatives, but in imbalanced medical datasets, a high accuracy can be achieved by simply predicting the majority class (e.g., healthy patients), which would miss many diseased cases (false negatives). Option B is wrong because the F1 score is the harmonic mean of precision and recall, balancing both; while it accounts for false negatives, it also penalizes false positives, which is not the primary concern when the cost of false negatives is extremely high. Option D is wrong because precision focuses on minimizing false positives (i.e., ensuring that predicted positives are correct), which is less relevant when the priority is to avoid missing actual positives (false negatives).

54
MCQmedium

A financial services company uses Amazon Bedrock to generate investment report summaries. They have strict compliance requirements that the model must not discuss certain topics like insider trading or unapproved financial advice. Which Bedrock feature should they use to deny these topics?

A.Prompt engineering with negative instructions
B.Bedrock Knowledge Bases
C.Bedrock Agents
D.Bedrock Guardrails
AnswerD

Bedrock Guardrails apply configurable denied-topics policies that block prompts and responses covering specified subjects such as insider trading or unapproved financial advice. This satisfies the compliance requirement by enforcing topic denial at inference time, independent of the underlying foundation model.

Why this answer

Bedrock Guardrails include topic denial policies to block specific subjects. Knowledge Bases, Agents, and prompt engineering alone cannot reliably enforce topic restrictions.

55
MCQhard

A practitioner is using Amazon Bedrock to invoke Anthropic Claude for a text generation task. They need the model to output a JSON object with specific keys, and they have observed that the model occasionally produces malformed JSON. Which parameter adjustment is MOST likely to improve JSON formatting consistency?

A.Increase the top_k parameter
B.Decrease the temperature parameter
C.Increase the max_tokens parameter
D.Add a stop sequence of '}'
AnswerB

Lowering temperature sharpens the probability distribution over tokens, so Claude samples higher-confidence continuations instead of exploring unlikely ones. That reduces the random drift producing malformed JSON, directly addressing the stem's inconsistent formatting. It does not guarantee schema compliance, but it is the parameter adjustment most likely to improve consistency without prompt changes.

Why this answer

Decreasing the temperature parameter reduces the randomness of the model's output, making it more deterministic and less likely to deviate from the expected JSON structure. Lower temperature values (e.g., 0.1–0.3) encourage the model to choose higher-probability tokens, which improves formatting consistency for structured outputs like JSON.

Exam trap

AWS often tests the misconception that increasing randomness parameters (top_k, temperature) improves output quality, when in fact reducing randomness is the key to enforcing strict formatting rules like JSON syntax.

How to eliminate wrong answers

Option A is wrong because increasing top_k expands the pool of candidate tokens the model can sample from, which increases randomness and can worsen JSON formatting issues. Option C is wrong because max_tokens controls the maximum length of the output, not the formatting or structure; it does not affect the model's tendency to produce malformed JSON. Option D is wrong because adding a stop sequence of '}' would prematurely terminate the output at the first closing brace, potentially cutting off nested JSON objects or arrays, and does not enforce correct JSON syntax throughout the generation.

56
MCQhard

A data scientist fine-tuned a large language model on Amazon SageMaker for financial report generation. The model produces responses that are too short and incomplete, often cutting off mid-sentence. What parameter should be adjusted first?

A.Increase the temperature parameter
B.Increase the top_p parameter
C.Increase the maximum token count
D.Switch to a different foundation model
AnswerC

Truncated, mid-sentence output indicates generation stopped at the configured output limit rather than the model finishing naturally. Raising the maximum token count lets the model complete longer financial reports, directly addressing the premature cutoff constraint in the stem.

Why this answer

The max tokens parameter limits the length of generated responses. Increasing it allows the model to produce longer completions. Temperature, top_p, and model change affect quality or diversity, but not the length cap.

57
MCQeasy

A data scientist needs to grant an IAM user access to a specific Amazon SageMaker notebook instance. The user should only be able to start and stop the notebook instance, but not delete it. Which IAM policy statement should be used?

A.{"Effect":"Allow","Action":["sagemaker:Start*","sagemaker:Stop*"],"Resource":"*"}
B.{"Effect":"Allow","Action":["sagemaker:StartNotebookInstance","sagemaker:StopNotebookInstance"],"Resource":"arn:aws:sagemaker:us-east-1:123456789012:notebook-instance/MyNotebook"}
C.{"Effect":"Allow","Action":"sagemaker:*","Resource":"*"}
D.{"Effect":"Allow","Action":"sagemaker:*","Resource":"arn:aws:sagemaker:us-east-1:123456789012:notebook-instance/MyNotebook"}
AnswerB

The policy grants only the sagemaker:StartNotebookInstance and sagemaker:StopNotebookInstance actions, scoped to the exact notebook-instance ARN. Because IAM denies any action absent from an Allow statement, DeleteNotebookInstance is implicitly denied, satisfying the requirement that the user start and stop the instance without deleting it.

Why this answer

It uses the specific actions `sagemaker:StartNotebookInstance` and `sagemaker:StopNotebookInstance` with a resource ARN that targets only the intended notebook instance. This grants the least privilege required to start and stop the instance while explicitly preventing deletion, as no delete action is included. The resource ARN restricts the policy to a single notebook instance, ensuring the user cannot affect other resources.

Exam trap

The trap here is that candidates often choose a wildcard action like `sagemaker:Start*` or `sagemaker:*` thinking it covers the needed actions, but they overlook that these patterns grant unintended permissions (e.g., delete or other start/stop actions on different resources), violating the principle of least privilege.

How to eliminate wrong answers

Option A is wrong because it uses wildcard actions `sagemaker:Start*` and `sagemaker:Stop*`, which could match unintended actions like `sagemaker:StartPipelineExecution` or `sagemaker:StopTrainingJob`, and the resource `*` grants access to all SageMaker resources, violating least privilege. Option C is wrong because `sagemaker:*` allows all SageMaker actions, including `sagemaker:DeleteNotebookInstance`, which the user should not have. Option D is wrong because `sagemaker:*` on a specific resource still grants all actions on that notebook instance, including deletion, which exceeds the required permissions.

58
MCQhard

A research lab is using Amazon SageMaker to fine-tune a large language model (LLM) for scientific text summarization. The training dataset contains 10 million documents, and the lab has a limited budget but needs to minimize training time. They have access to SageMaker Training with managed spot instances, which offer significant cost savings but are interruptible. The team is considering different training strategies to balance cost, time, and model quality. Which strategy should they use?

A.Use SageMaker's distributed training with data parallelism on multiple managed spot instances, and enable checkpointing.
B.Fine-tune only the last few layers of the model on a smaller subset of the data.
C.Use a single on-demand instance to avoid interruptions and maximize throughput.
D.Use a single large GPU instance to train the model from scratch.
AnswerA

Data parallelism shards the 10-million-document dataset across multiple managed spot instances, cutting wall-clock training time while spot pricing reduces cost. Checkpointing persists progress so interrupted instances resume rather than restart, preserving model quality within budget.

Why this answer

The best strategy. Using SageMaker's distributed training with data parallelism across multiple managed spot instances allows parallel processing of the 10-million-document dataset, significantly reducing training time. Spot instances offer cost savings of up to 90% compared to on-demand, and enabling checkpointing ensures that if any instance is interrupted, training can resume from the last checkpoint without losing progress.

Option B is incorrect because fine-tuning only the last few layers on a subset of data compromises model quality and does not leverage the full dataset. Option C is incorrect because a single on-demand instance is expensive and slower than distributed spot instances. Option D is incorrect because training from scratch on a single GPU is prohibitively slow and costly, especially given the budget constraints.

59
MCQmedium

A financial institution wants to predict whether a loan applicant will default. They have a historical dataset with loan outcomes (default or no default) and various applicant features. Which type of machine learning should they use?

A.Reinforcement learning
B.Unsupervised learning
C.Clustering
D.Supervised learning
AnswerD

Supervised learning uses labeled data to train a model that maps input features to a target output. Here, the target is binary (default or no default), and historical labeled data is available. This allows training a classification model to predict default for new applicants, making supervised learning the correct choice.

Why this answer

Supervised learning is ideal when historical labeled data is available and the goal is to predict a known outcome. The loan default prediction task is a binary classification problem, which supervised learning handles by learning from past examples to classify new applicants.

Exam trap

The trap here is confusing clustering with classification because both group data, but clustering is unsupervised and cannot predict a labeled outcome.

60
MCQmedium

A company wants to send personalized product recommendations to customers based on their browsing history and previous purchases. Which AWS service is BEST suited for this?

A.Amazon Personalize
B.Amazon SageMaker built-in factorization machines
C.Amazon Forecast
D.Amazon Rekognition
AnswerA

Amazon Personalize ingests browsing history and purchase events, then trains a custom recommendation model using the same algorithms behind Amazon.com. It satisfies the requirement for personalised product recommendations by generating real-time, user-specific suggestions through a deployed campaign, without needing machine learning expertise.

Why this answer

Amazon Personalize is a fully managed machine learning service specifically designed to build real-time personalized recommendation systems. It uses the same technology as Amazon.com's recommendation engine, processing user-item interaction data (browsing history, purchases) to generate tailored product suggestions. This makes it the ideal choice for the described use case.

Exam trap

The trap here is that candidates might confuse Amazon SageMaker's general ML capabilities with a specialized managed service like Amazon Personalize, assuming SageMaker's built-in algorithms are equally suited for recommendation tasks without considering the operational overhead and lack of pre-built recommendation pipelines.

How to eliminate wrong answers

Option B is wrong because Amazon SageMaker built-in factorization machines are a general-purpose algorithm for matrix factorization, requiring significant custom data preparation, model training, and deployment effort, whereas Amazon Personalize provides an end-to-end managed solution with built-in recommendation models. Option C is wrong because Amazon Forecast is designed for time-series forecasting (e.g., demand planning, sales predictions), not for generating personalized product recommendations based on user behavior. Option D is wrong because Amazon Rekognition is a computer vision service for image and video analysis (e.g., object detection, facial recognition), which is unrelated to recommendation systems.

61
MCQmedium

An organization wants to prototype a new generative AI application and allow multiple team members to collaborate on prompt engineering and model selection without writing code. Which tool should they use?

A.Amazon SageMaker Canvas
B.Amazon Bedrock Playground
C.Amazon CodeWhisperer
D.Amazon Bedrock Studio
AnswerD

Amazon Bedrock Studio provides a browser-based, no-code workspace where teams collaboratively build generative AI prototypes, covering exactly the stem's prompt engineering and model selection requirements. It supports shared projects with multiple collaborators, satisfying the no-code and multi-user collaboration constraints that a command-line or SDK-only approach would fail.

Why this answer

Amazon Bedrock Studio is a web-based, no-code environment specifically designed for teams to collaboratively prototype generative AI applications, experiment with prompts, compare foundation models, and share work — all without writing code. It provides a shared workspace with prompt engineering tools and model selection, which matches the requirement exactly.

Exam trap

The trap is confusing Bedrock Playground (single-user, quick test) with Bedrock Studio (collaborative, no-code project workspace); the exam tests whether you know which tool supports team collaboration without code.

How to eliminate wrong answers

Option A is wrong because SageMaker Canvas is a no-code machine learning tool for building predictive models (classification, regression, forecasting), not a collaborative generative AI prompt-engineering playground. Option B is wrong because Bedrock Playground is a single-user console for testing prompts and models; it lacks the team collaboration and project-sharing features that Bedrock Studio provides. Option C is wrong because Amazon CodeWhisperer (now Amazon Q Developer) is an AI coding assistant that generates code suggestions inside an IDE — it requires writing code and is not a no-code collaboration tool.

62
MCQmedium

An ML team is deploying a real-time inference endpoint for a computer vision model using Amazon SageMaker. The model requires GPU acceleration for low latency. Which instance type should the team choose to minimize cost while meeting the GPU requirement?

A.ml.g5.xlarge
B.ml.c5.xlarge
C.ml.p3.2xlarge
D.ml.p4d.24xlarge
AnswerA

ml.g5.xlarge pairs an NVIDIA A10G GPU with four vCPUs, providing the required GPU acceleration at the lowest cost within the G5 family. Smaller GPU instances lack sufficient acceleration, while larger G5 sizes exceed the stated requirement.

Why this answer

(ml.g5.xlarge) is correct because it provides a GPU (NVIDIA A10G Tensor Core GPU) necessary for low-latency GPU acceleration in computer vision inference, while being the most cost-effective GPU instance among the options. The ml.g5.xlarge offers sufficient GPU compute for real-time inference at a lower hourly cost compared to ml.p3.2xlarge, making it the optimal choice for minimizing cost while meeting the GPU requirement.

Exam trap

Candidates may be tempted to choose ml.p3.2xlarge due to its reputation for ML training, but for inference, ml.g5 instances often provide better price-performance. The ml.p3.2xlarge, while GPU-equipped, is more expensive per hour and may be over-provisioned for typical inference workloads, leading to unnecessary costs.

How to eliminate wrong answers

Option A (ml.g5.xlarge) is wrong because it uses a GPU (NVIDIA A10G) that is more powerful and expensive than needed for this use case, leading to higher cost; however, it does meet the GPU requirement, so the primary issue is cost inefficiency, not technical incompatibility. Option B (ml.c5.xlarge) is wrong because it is a CPU-only instance (based on Intel Xeon Scalable processors) and lacks any GPU, failing to meet the explicit GPU acceleration requirement for low-latency computer vision inference. Option D (ml.p4d.24xlarge) is wrong because it provides 8 NVIDIA A100 GPUs, which is massively over-provisioned for a single real-time inference endpoint, resulting in significantly higher cost without any benefit for this workload.

63
MCQmedium

A company wants to use Amazon Bedrock to generate responses grounded in their proprietary knowledge base. They need to minimize hallucinations and ensure responses are based on the provided documents. Which feature should they enable?

A.Topic restrictions
B.Word filters
C.Grounding check
D.Content filters
AnswerC

Grounding check measures whether each response is supported by the retrieved source passages, flagging or filtering ungrounded statements. This directly reduces hallucination by tying outputs to the proprietary knowledge base documents rather than model memory.

Why this answer

Grounding check in Bedrock Guardrails verifies that the model's response is supported by the source documents, reducing hallucinations.

64
MCQhard

A financial services company is deploying a foundation model on Amazon Bedrock to generate compliance reports from internal audit logs. The model must not output any personally identifiable information (PII). They have configured a Bedrock Guardrail with sensitive information filters set to the 'HIGH' sensitivity level. During testing in a staging environment, testers still observed PII being occasionally generated in the report outputs. The guardrail did not block these instances because the PII was embedded in a context that the guardrail's pattern matching did not catch (e.g., structured JSON data with embedded names). The company requires a solution that minimizes latency and cost, as they process thousands of reports daily. They cannot afford to increase inference time significantly due to strict SLAs. They also want to avoid re-engineering the entire solution. Which additional step should they take to effectively eliminate PII leakage while maintaining performance?

A.Add a prompt instruction to the model to never output PII, with few-shot examples of non-PII outputs.
B.Fine-tune the foundation model on a dataset that excludes PII.
C.Increase the guardrail sensitivity to 'MAXIMUM'.
D.Implement a post-processing Lambda function that uses Amazon Comprehend's PII detection to scan and redact any PII from the model output before returning it.
AnswerD

Amazon Comprehend's PII detection uses trained machine-learning models rather than regex patterns, so it recognises names embedded in structured JSON that Guardrail filters miss. Running it as a post-processing Lambda adds only milliseconds per report, satisfying the latency and cost constraints without re-engineering the Bedrock invocation path.

Why this answer

Amazon Comprehend's PII detection API can be invoked as a post-processing step to scan and redact PII from the model output without requiring any changes to the model or guardrail configuration. This approach adds minimal latency (typically under 100ms per request) and cost per API call is low, making it suitable for high-throughput scenarios. It directly addresses the guardrail's failure to catch PII embedded in structured contexts like JSON, as Comprehend uses machine learning models that can identify PII even when it's not in plain text patterns.

Exam trap

The trap here is that candidates assume increasing guardrail sensitivity or prompt engineering can solve all PII detection failures, but they overlook that guardrails rely on pattern matching and cannot handle contextually embedded PII, whereas a dedicated ML-based detection service like Amazon Comprehend is designed for that exact scenario.

How to eliminate wrong answers

Option A is wrong because prompt instructions and few-shot examples are not reliable for preventing PII leakage; the model may still generate PII due to its training data or contextual reasoning, and this approach adds no deterministic enforcement. Option B is wrong because fine-tuning a foundation model to exclude PII is expensive, time-consuming, and requires a large curated dataset; it also risks degrading model performance on the compliance reporting task and does not guarantee elimination of all PII. Option C is wrong because the guardrail's sensitivity levels (HIGH, MAXIMUM) only affect pattern-matching rules and confidence thresholds; they cannot detect PII embedded in non-standard formats like structured JSON, so increasing sensitivity does not solve the core issue.

65
Multi-Selecthard

Which TWO are best practices for model monitoring in production on AWS?

Select 2 answers
A.Disable logging to reduce latency
B.Use only CPU instances
C.Monitor input data drift
D.Retrain model daily
E.Monitor prediction drift
AnswersC, E

Input data drift monitoring detects shifts between production inference data and the training distribution, such as changing customer demographics or feature ranges. This satisfies the best-practice requirement by triggering alerts or retraining before degraded inputs silently corrupt prediction quality in the deployed AWS model.

Why this answer

Option C (Monitor input data drift) is correct because production models degrade when the statistical distribution of incoming features diverges from training data, so tracking data drift with tools like SageMaker Model Monitor detects this covariate shift early. Option E (Monitor prediction drift) is correct because shifts in the distribution of model outputs (e.g., changing class proportions or score distributions) signal concept drift or upstream data issues that require investigation or retraining. Together, input and prediction drift monitoring are the two core best practices for maintaining model quality in production.

Option A is wrong because disabling logging removes the observability needed to detect drift, errors, and bias, and logging overhead is typically negligible relative to inference. Option B is wrong because instance type (CPU vs. GPU) is a performance/cost choice, not a monitoring best practice.

Option D is wrong because retraining daily is not a best practice by itself; retraining should be triggered by monitored drift or performance degradation, and daily retraining can be costly and destabilizing.

Exam trap

The AIF-C01 exam often tests the misconception that retraining on a fixed schedule (e.g., daily) is a best practice, when in reality it should be event-driven based on drift or performance metrics.

66
MCQhard

A healthcare AI system predicts patient diagnoses. The data collection process primarily samples from urban hospitals, leading to underrepresentation of rural populations. Which type of bias is this, and what is the most effective mitigation strategy?

A.Representation bias; collect additional data from rural hospitals
B.Historical bias; reweight urban samples to reduce their influence
C.Aggregation bias; use regularization to simplify the model
D.Measurement bias; apply data augmentation to rural records
AnswerA

Sampling only urban hospitals underrepresents rural populations, producing representation bias in the training data. Collecting additional data from rural hospitals rebalances the dataset so the model learns patterns from the missing subgroup, directly addressing the sampling gap rather than adjusting outputs after training.

Why this answer

This is representation bias, which occurs when the training data underrepresents certain subgroups (rural populations) relative to the population the model will serve. The most effective mitigation is to collect additional data from the underrepresented group (rural hospitals) so the model learns their patterns and does not systematically misdiagnose them.

Exam trap

AIF-C01 often tests the distinction between bias types (representation vs. historical vs. measurement vs. aggregation), so candidates who see 'underrepresentation' but pick a mitigation that only reweights existing data choose the wrong answer.

How to eliminate wrong answers

Option B is wrong because historical bias refers to past inequities embedded in data (e.g., biased labeling or access), and reweighting urban samples does not fix the missing rural data. Option C is wrong because aggregation bias occurs when a single model is applied across heterogeneous subgroups that should be modeled separately; regularization does not address missing data. Option D is wrong because measurement bias relates to how features or labels are measured inconsistently across groups, and data augmentation of rural records cannot substitute for real rural data.

67
MCQeasy

Which AWS service provides a serverless experience for building and scaling generative AI applications with access to various foundation models?

A.Amazon Bedrock
B.Amazon SageMaker
C.Amazon Lex
D.AWS Lambda
AnswerA

Amazon Bedrock delivers a fully managed, serverless API that exposes multiple foundation models from providers such as Anthropic, Meta and Amazon, removing infrastructure provisioning. This directly satisfies the stem's serverless requirement for building and scaling generative AI applications, since you invoke models on demand without managing servers or capacity.

Why this answer

Amazon Bedrock is a fully managed service that provides a serverless experience for building and scaling generative AI applications. It offers access to a variety of foundation models (FMs) from providers like AI21 Labs, Anthropic, Cohere, Meta, Stability AI, and Amazon via a single API, without the need to manage underlying infrastructure.

Exam trap

The trap here is that candidates may confuse Amazon SageMaker's broad ML capabilities with the specific serverless, foundation-model-focused offering of Amazon Bedrock, or mistakenly think AWS Lambda alone provides generative AI model access when it is merely a compute trigger.

How to eliminate wrong answers

Option B is wrong because Amazon SageMaker is a comprehensive machine learning (ML) platform that requires users to manage the entire ML lifecycle, including provisioning instances, training custom models, and deploying endpoints; it is not a serverless service specifically designed for accessing pre-built foundation models. Option C is wrong because Amazon Lex is a service for building conversational interfaces (chatbots) using automatic speech recognition (ASR) and natural language understanding (NLU), and it does not provide access to foundation models for generative AI tasks. Option D is wrong because AWS Lambda is a serverless compute service that runs code in response to events, but it does not natively provide access to foundation models or a managed API for generative AI; it can be used as part of a solution but is not the primary service for building and scaling generative AI applications with FMs.

68
MCQmedium

A company needs to generate high-quality product images from textual descriptions for an e-commerce catalog. They want to use a foundation model on AWS that specializes in text-to-image generation. Which model provider should they use through Amazon Bedrock?

A.Stability AI
B.Meta
C.Anthropic
D.Cohere
AnswerA

Stability AI provides text-to-image foundation models such as Stable Diffusion through Amazon Bedrock, purpose-built for generating images from textual prompts. This matches the e-commerce catalogue requirement for a provider specialising in text-to-image generation rather than text or embedding models.

Why this answer

Stability AI is the correct provider because its Stable Diffusion models, available through Amazon Bedrock, are specifically designed and optimized for text-to-image generation. These models use a diffusion process to generate high-quality images from textual descriptions, making them ideal for the e-commerce catalog use case.

Exam trap

Candidates may confuse foundation models that excel in text generation (like Anthropic or Meta) with those specialized for multimodal tasks, assuming all major AI providers offer image generation capabilities through Amazon Bedrock.

How to eliminate wrong answers

Option B (Meta) is wrong because Meta's foundation models on Bedrock, such as Llama 2, are large language models (LLMs) focused on text generation and understanding, not image generation. Option C (Anthropic) is wrong because Anthropic's Claude models are conversational AI assistants designed for text-based tasks like summarization and Q&A, lacking image generation capabilities. Option D (Cohere) is wrong because Cohere's models specialize in natural language processing tasks such as text generation, classification, and embeddings, with no support for image synthesis.

69
MCQeasy

A retail company is preparing to launch a generative AI customer support assistant built on Amazon Bedrock. Before launch, the responsible AI review board asks the team to document the assistant's intended purpose, its known limitations, and the evaluation results from fairness and accuracy testing. Which AWS resource should the team produce to satisfy this request?

A.An Amazon CloudWatch dashboard showing invocation counts, latency percentiles, and error rates for the assistant.
B.An AWS Trusted Advisor check report listing cost optimization and security recommendations for the account.
C.An Amazon SageMaker Model Card that records intended use, risk rating, and evaluation results for the model.
D.An AWS Artifact report downloaded from the compliance portal covering the AWS services in use.
AnswerC

SageMaker Model Cards are the governance artifact designed for exactly this purpose: they capture intended use, out-of-scope uses, risk rating, training details, and evaluation metrics in a structured document. Producing a model card gives the review board a single, standardized record of purpose, limitations, and fairness and accuracy results, which is what the scenario asks the team to deliver.

Why this answer

A responsible AI review board wants structured documentation of purpose, limitations, and evaluation outcomes. SageMaker Model Cards are purpose-built for this, capturing intended use, out-of-scope applications, risk rating, and performance and fairness metrics in a consistent format. The other artifacts cover AWS compliance posture, runtime telemetry, or account advisories, none of which describe the model itself.

Exam trap

The trap here is confusing provider-side compliance documentation such as AWS Artifact with customer-side model governance documentation, which must describe the customer's own model, not AWS's certifications.

70
MCQmedium

A company needs to convert a large number of recorded customer service calls into text for analysis. Which AWS service should they use?

A.Amazon Textract
B.Amazon Comprehend
C.Amazon Transcribe
D.Amazon Polly
AnswerC

Amazon Transcribe is the managed automatic speech recognition service that converts audio recordings into text transcripts. It handles batch call recordings directly, satisfying the requirement to transform a large volume of customer service calls into analysable text.

Why this answer

Amazon Transcribe is the correct service because it is specifically designed to convert speech from audio files (such as recorded customer service calls) into text using automatic speech recognition (ASR). This directly matches the requirement to transcribe audio content for subsequent text analysis.

Exam trap

The trap here is confusing Amazon Transcribe (speech-to-text) with Amazon Polly (text-to-speech) or Amazon Textract (text from images), as candidates often misremember which AWS service handles audio transcription versus document text extraction.

How to eliminate wrong answers

Option A is wrong because Amazon Textract is designed to extract text and data from scanned documents and images (e.g., PDFs, forms), not from audio recordings. Option B is wrong because Amazon Comprehend is a natural language processing (NLP) service that analyzes text for insights like sentiment or entities, but it cannot convert speech to text. Option D is wrong because Amazon Polly is a text-to-speech service that converts text into lifelike speech, the opposite of the required speech-to-text conversion.

71
Multi-Selectmedium

A data science team is preparing a dataset for training a machine learning model. They need to perform data preprocessing to improve model performance. Which TWO of the following are common data preprocessing techniques? (Choose two.)

Select 2 answers
A.One-hot encoding
B.Feature selection
C.Cross-validation
D.Normalization
E.Gradient descent
AnswersA, D

One-hot encoding converts categorical variables into a binary vector representation, allowing machine learning algorithms to handle non-numeric data. It is a fundamental preprocessing step for categorical features, ensuring they can be used in models that require numerical input. This is a common technique.

Why this answer

Normalization and one-hot encoding are standard preprocessing techniques. Normalization scales numerical features, while one-hot encoding transforms categorical variables into a numerical format. Both are applied to the data before training to ensure the model can effectively learn from all features.

Exam trap

The trap here is confusing model training techniques like gradient descent or evaluation methods like cross-validation with data preprocessing steps.

72
MCQeasy

Which of the following correctly describes the purpose of pre-training in the context of large language models?

A.To evaluate the model's performance on a benchmark dataset
B.To reduce the model's size for deployment on edge devices
C.To adapt the model to a specific downstream task using labeled data
D.To learn general language representations from a large, unlabeled corpus
AnswerD

Pre-training learns broad statistical patterns and representations from massive unlabelled text, satisfying the stem's requirement to describe its purpose. This self-supervised phase precedes fine-tuning, which adapts those general representations to specific labelled tasks. It is distinct from inference-time prompting, which performs no weight updates.

Why this answer

Pre-training in large language models involves learning general language representations from a large, unlabeled corpus using self-supervised objectives like masked language modeling or next-token prediction. This phase builds a foundational understanding of syntax, semantics, and world knowledge without task-specific labels, which can later be fine-tuned for downstream tasks.

Exam trap

AWS often tests the distinction between pre-training (unsupervised on unlabeled data) and fine-tuning (supervised on labeled data), so the trap here is confusing the purpose of fine-tuning with pre-training.

How to eliminate wrong answers

Option A is wrong because evaluating performance on a benchmark dataset is a separate evaluation step, not the purpose of pre-training. Option B is wrong because reducing model size for edge deployment is achieved through techniques like pruning, quantization, or distillation, not pre-training. Option C is wrong because adapting a model to a specific downstream task using labeled data describes fine-tuning, not pre-training.

73
MCQeasy

Refer to the exhibit. A data scientist ran a training job on Amazon SageMaker. The job failed with the error shown. What is the most likely cause?

A.The S3 input path is incorrect
B.The IAM role does not have permission to access S3
C.The training code has a syntax error
D.The batch size is too large for the instance's GPU memory
AnswerD

A batch size exceeding GPU memory triggers an out-of-memory failure during the forward or backward pass, since activations for the whole batch must be held simultaneously. Reducing the batch size, or using gradient accumulation, directly addresses this constraint rather than altering the model architecture.

Why this answer

The error message indicates a CUDA out-of-memory error, which occurs when the GPU memory is insufficient for the requested batch size. Option D is correct because increasing the batch size beyond the GPU's memory capacity causes the training job to fail with this specific error.

Exam trap

AWS often tests the distinction between infrastructure errors (S3, IAM) and runtime errors (CUDA memory), where candidates mistakenly attribute a GPU memory error to a misconfiguration in data access or code syntax.

How to eliminate wrong answers

Option A is wrong because an incorrect S3 input path would result in a 'NoSuchKey' or '404' error, not a CUDA out-of-memory error. Option B is wrong because an IAM role lacking S3 permissions would produce an 'AccessDenied' error, not a GPU memory error. Option C is wrong because a syntax error in the training code would raise a Python exception (e.g., SyntaxError) before any GPU operations, not a CUDA memory error.

74
MCQeasy

A data scientist is using SageMaker Clarify to analyze a binary classification model for gender bias. The dataset has 80% male and 20% female applicants. The model predicts positive outcomes for 60% of males and 30% of females. Which fairness metric would directly capture this disparity in prediction rates?

A.Equalized odds
B.Disparate impact
C.Accuracy difference
D.Demographic parity
AnswerD

Demographic parity compares positive prediction rates across groups, so 60% male versus 30% female reveals a clear disparity. Other metrics such as equalised odds or disparate impact assess error rates or ratios, not the raw outcome-rate gap described.

Why this answer

Demographic parity (also called statistical parity) measures whether the probability of a positive prediction is equal across groups. Here, 60% of males versus 30% of females receive positive outcomes, a clear disparity in prediction rates, which is exactly what demographic parity captures. It compares selection rates directly without conditioning on actual labels.

Exam trap

AIF-C01 often tests the distinction between demographic parity (equal prediction rates) and equalized odds (equal error rates) — candidates who focus on 'bias' generically may pick equalized odds without noticing the question only provides prediction rates, not ground-truth labels.

How to eliminate wrong answers

Option A is wrong because equalized odds compares true positive rates and false positive rates across groups, which requires ground-truth labels and measures error-rate parity rather than raw prediction-rate parity. Option B is wrong because disparate impact is a legal/regulatory ratio (typically the 80% rule under EEOC guidelines) computed as the ratio of selection rates; while related, the question asks for the metric that directly captures the disparity in prediction rates, which is demographic parity, and disparate impact is a derived ratio rather than the underlying parity metric. Option C is wrong because accuracy difference compares overall model accuracy between groups, which does not isolate prediction-rate disparity and requires labels.

75
MCQmedium

An e-commerce company uses an Amazon Lex chatbot to handle customer inquiries. They want to implement human oversight for sensitive interactions, such as when the chatbot cannot provide a confident response. Which AWS service should they integrate?

A.Amazon Rekognition
B.Amazon Comprehend
C.Amazon Augmented AI (A2I)
D.Amazon SageMaker Ground Truth
AnswerC

Amazon Augmented AI (A2I) adds a human review workflow directly into the inference path, triggering reviewers when confidence scores fall below a threshold. This satisfies the stem's requirement for human oversight of low-confidence chatbot responses, unlike services that only monitor metrics or route notifications after the fact.

Why this answer

Amazon Augmented AI (A2I) is the correct service because it provides built-in human review workflows for ML predictions, allowing you to route low-confidence responses from Amazon Lex to human reviewers for oversight. This directly addresses the requirement for human oversight on sensitive interactions where the chatbot cannot provide a confident response.

Exam trap

In AWS exams, candidates often confuse data processing services (like Comprehend or Rekognition) with the human review service (A2I). They mistakenly choose a service that analyzes text or images rather than the one that orchestrates human oversight.

How to eliminate wrong answers

Option A is wrong because Amazon Rekognition is a computer vision service for image and video analysis, not designed for human review of chatbot interactions. Option B is wrong because Amazon Comprehend is a natural language processing (NLP) service for extracting insights from text, but it does not include a human-in-the-loop review mechanism. Option D is wrong because Amazon SageMaker Ground Truth is used for creating training data labels via human workers, not for real-time human oversight of production chatbot responses.

Page 1 of 12

Page 2