Courseiva

CCNA Foundation Model Applications Questions

53 of 128 questions · Page 2/2 · Foundation Model Applications topic · Answers revealed

76
MCQmedium

A retail company wants to compare the output quality of several foundation models available in Amazon Bedrock for a product description generation task. They need a repeatable, automated way to score responses against reference descriptions. Which AWS capability should they use?

A.Amazon Bedrock model evaluation jobs with automatic metrics such as BERTScore and ROUGE.
B.AWS Trusted Advisor to assess foundation model performance.
C.Amazon Bedrock provisioned throughput to benchmark model latency.
D.Amazon CloudWatch Logs to capture model responses and manually review them.
AnswerA

Amazon Bedrock supports model evaluation jobs that can automatically score model outputs against reference data using metrics like BERTScore and ROUGE. This provides a repeatable, automated comparison across multiple foundation models, which matches the requirement. It removes manual scoring and produces consistent results that can guide model selection for the product description task.

Why this answer

Amazon Bedrock model evaluation jobs run automated scoring against reference datasets using metrics such as BERTScore and ROUGE, enabling consistent comparison of multiple foundation models. This directly supports selecting the best model for product description generation with repeatable, quantitative results rather than subjective manual review.

Exam trap

The trap here is confusing operational monitoring or capacity features with model quality evaluation, which requires scoring outputs against references.

77
MCQeasy

A startup uses Amazon Bedrock with a provisioned throughput to generate product images. They now have unpredictable traffic and want to reduce costs. What should they do?

A.Switch to batch inference using Amazon Bedrock.
B.Keep the provisioned throughput but reduce the number of units.
C.Use a different model or service like Amazon SageMaker with spot instances.
D.Switch to on-demand mode in Amazon Bedrock.
AnswerD

On-demand mode bills per token or image, so cost scales directly with actual usage rather than a fixed provisioned commitment. This satisfies the stem's unpredictable traffic and cost-reduction constraint, since idle provisioned capacity would otherwise be paid for regardless of demand.

Why this answer

On-demand mode in Amazon Bedrock allows you to pay per inference request without committing to a provisioned throughput, making it ideal for unpredictable traffic patterns. This eliminates the cost of idle capacity while still providing access to the same foundation models. Option D directly addresses the need to reduce costs when traffic is variable.

Exam trap

The trap here is that candidates may assume provisioned throughput is always more cost-effective for any workload, overlooking that on-demand mode is specifically designed to eliminate idle costs for unpredictable traffic patterns.

How to eliminate wrong answers

Option A is wrong because batch inference is designed for processing large volumes of data asynchronously, not for handling unpredictable real-time traffic, and it still requires provisioning resources that may incur costs even when idle. Option B is wrong because reducing the number of provisioned throughput units still leaves you with committed capacity that must be paid for regardless of usage, which does not solve the cost issue for unpredictable traffic. Option C is wrong because switching to a different model or service like Amazon SageMaker with spot instances introduces additional complexity and does not leverage the native on-demand pricing model of Bedrock, which is specifically designed for variable workloads.

78
Multi-Selectmedium

A healthcare startup is using Amazon Bedrock to build a patient education chatbot. The chatbot must generate responses that are empathetic, accurate, and compliant with medical privacy regulations. The startup wants to implement safeguards to prevent the model from generating harmful or inappropriate content. Which TWO actions should the startup take to meet these requirements? (Choose two.)

Select 2 answers
A.Use Amazon Comprehend Medical to detect and redact protected health information (PHI) in user inputs and model outputs.
B.Configure Amazon Bedrock Guardrails to filter harmful content and define denied topics.
C.Enable Amazon Bedrock model evaluation to automatically monitor and block harmful outputs in real time.
D.Use a foundation model with a low temperature setting to reduce randomness.
E.Implement prompt engineering to instruct the model to be empathetic and avoid medical advice.
AnswersA, B

Amazon Comprehend Medical can identify protected health information (PHI) in text, enabling redaction before sending data to the model and before returning outputs to users. This helps comply with medical privacy regulations such as HIPAA. While Bedrock Guardrails also offer PII redaction, Comprehend Medical is specifically trained for medical entities. This action directly supports the privacy compliance requirement.

Why this answer

To prevent harmful content and ensure medical privacy compliance, the startup should use Amazon Bedrock Guardrails to filter harmful categories and denied topics, and Amazon Comprehend Medical to detect and redact PHI. Guardrails provide runtime content moderation, while Comprehend Medical specializes in medical data privacy. Together they address both safety and regulatory requirements.

Other options like low temperature or prompt engineering do not enforce safeguards, and model evaluation is not a real-time filter.

Exam trap

The trap here is thinking that prompt engineering or low temperature settings are sufficient safeguards, when actual enforcement requires dedicated filtering services like Guardrails and Comprehend Medical.

79
MCQeasy

A company uses Amazon Bedrock to build a chatbot. The chatbot needs to answer questions based on internal company documents. Which AWS service should be integrated with Bedrock to enable Retrieval Augmented Generation (RAG) without managing infrastructure?

A.Amazon OpenSearch Service
B.Amazon DynamoDB
C.Amazon RDS
D.Amazon Kendra
AnswerD

Amazon Kendra provides a fully managed, ML-powered enterprise search service that indexes internal documents and returns relevant passages, which Bedrock then feeds to the foundation model for grounded answers. It satisfies the RAG requirement without infrastructure management, unlike self-managed vector stores or OpenSearch clusters.

Why this answer

Amazon Kendra is a fully managed intelligent search service that can be directly integrated with Amazon Bedrock to implement Retrieval Augmented Generation (RAG) without any infrastructure management. It indexes internal company documents and retrieves relevant passages, which are then passed to the foundation model as context to generate accurate, grounded answers.

Exam trap

AWS often tests the distinction between managed services that require infrastructure management (like OpenSearch Service) and fully managed services (like Kendra) that abstract away all infrastructure concerns, making candidates incorrectly choose OpenSearch for its search capabilities.

How to eliminate wrong answers

Option A is wrong because Amazon OpenSearch Service requires you to manage clusters, configure indexing, and handle scaling — it is not a serverless, zero-infrastructure option. Option B is wrong because Amazon DynamoDB is a NoSQL key-value and document database designed for transactional workloads, not for semantic search or document retrieval needed in RAG. Option C is wrong because Amazon RDS is a relational database service that requires provisioning and managing database instances, and it lacks native semantic search capabilities for document retrieval.

80
MCQeasy

A retail analytics team wants an assistant that answers questions about last quarter's sales using data stored in an Amazon S3 bucket of PDF reports and CSV exports. They want the model to cite the underlying documents and avoid inventing figures. Which Amazon Bedrock capability should they use?

A.Guardrails for Amazon Bedrock configured with a denied topics policy.
B.Model evaluation jobs that score responses against a ground-truth dataset.
C.Provisioned Throughput purchased for the foundation model.
D.Knowledge Bases for Amazon Bedrock with the S3 bucket as a data source.
AnswerD

Knowledge Bases for Amazon Bedrock ingests documents from sources such as Amazon S3, chunks and embeds them into a vector store, and retrieves relevant passages at query time so the model grounds answers in those documents and can return citations. This directly satisfies the requirement to answer from the sales reports and avoid fabricated figures.

Why this answer

Grounding answers in private documents requires retrieval, not filtering or capacity. Knowledge Bases for Amazon Bedrock manages ingestion, chunking, embedding, and retrieval from the S3 source, so responses are generated from the sales reports and can include citations. Evaluation, Guardrails, and Provisioned Throughput address quality measurement, content safety, and capacity respectively, none of which supply the source data.

Exam trap

The trap here is confusing content filtering or capacity reservation with retrieval-augmented grounding of private data.

81
Multi-Selectmedium

A company is using Amazon Bedrock to build a conversational agent. They want to ensure the agent maintains context across multiple turns in a conversation. Which TWO strategies should the developer implement? (Choose two.)

Select 2 answers
A.Use Amazon Bedrock's session management feature to automatically persist conversation state.
B.Include the entire conversation history in each prompt sent to the model.
C.Enable the model's memory parameter to retain information across API calls.
D.Use a single API call with a streaming response to keep the connection open and maintain context.
E.Store conversation history in an external database and retrieve relevant turns to include in the prompt.
AnswersB, E

Including the entire conversation history in each prompt allows the model to see previous exchanges and maintain context. This is a common technique for multi-turn conversations with foundation models, as they are stateless and do not remember previous interactions. However, it increases token usage and may hit context length limits, so it should be managed carefully.

Why this answer

To maintain context across multiple turns, developers must include conversation history in each prompt. This can be done by sending the full history or by storing history externally and retrieving relevant parts. Foundation models are stateless, so context must be explicitly provided.

Options suggesting built-in memory or session management are incorrect.

Exam trap

The trap here is assuming that Amazon Bedrock or the foundation models automatically maintain conversation state, when in fact they are stateless and require manual context management.

82
MCQhard

A company uses Amazon Bedrock to generate code. They want to ensure the code follows security best practices and does not contain vulnerabilities. Which approach is most effective?

A.Implement a post-processing step using AWS WAF.
B.Use Amazon CodeGuru Security to review generated code.
C.Train a custom model on the company’s secure code.
D.Use a foundation model trained only on secure code.
AnswerB

Amazon CodeGuru Security applies static analysis with detector rules tuned to identify security vulnerabilities and deviations from coding best practices, directly satisfying the requirement that generated code be checked for vulnerabilities. Unlike generic content filters, it inspects code semantics, catching issues such as injection flaws and insecure data handling within the Bedrock output.

Why this answer

Amazon CodeGuru Security reviews code for security vulnerabilities and provides recommendations. Using a model trained on secure code may not be sufficient; WAF is for web traffic; training a custom model requires significant effort and may not catch all issues.

83
MCQhard

A financial services company is using Amazon Bedrock to generate investment summaries. They must ensure that the model does not provide personalized financial advice, which is a regulatory requirement. Which AWS feature should they use to block the model from generating such advice?

A.Amazon Bedrock Guardrails with a denied topics policy.
B.Amazon Bedrock model evaluation with a custom metric for regulatory compliance.
C.Amazon Bedrock Agents with an action group that checks for financial advice keywords.
D.Amazon Bedrock Provisioned Throughput with a custom model that has been fine-tuned to avoid financial advice.
AnswerA

Amazon Bedrock Guardrails allows you to define denied topics, which are subjects the model should avoid. By specifying 'personalized financial advice' as a denied topic, the guardrail will block the model from generating content on that topic. This is the appropriate feature to enforce regulatory compliance by preventing certain types of responses.

Why this answer

To block the model from generating personalized financial advice, the company should use Amazon Bedrock Guardrails with a denied topics policy. This feature allows defining topics that the model must avoid, effectively preventing non-compliant responses. Other options like model evaluation or fine-tuning do not provide real-time blocking.

Exam trap

The trap here is thinking that fine-tuning or model evaluation can reliably block specific content, when in fact Guardrails with denied topics is the purpose-built feature for real-time filtering.

84
MCQeasy

A company uses Amazon Bedrock to generate product descriptions. They notice that the output sometimes contains incorrect information. What should they do to improve accuracy?

A.Increase the temperature parameter.
B.Implement Retrieval-Augmented Generation (RAG).
C.Use a larger foundation model.
D.Use AWS WAF to filter outputs.
AnswerB

Retrieval-Augmented Generation grounds responses in authoritative source documents retrieved at inference time, so the model cites real product data instead of relying solely on parametric memory. This directly reduces fabricated or incorrect details in generated descriptions.

Why this answer

Retrieval-Augmented Generation (RAG) enhances the accuracy of foundation model outputs by grounding the generation in authoritative, up-to-date external knowledge sources. Instead of relying solely on the model's parametric memory, RAG retrieves relevant documents or data from a vector database (e.g., Amazon OpenSearch Serverless) and injects them into the prompt context, reducing hallucinations and incorrect information in product descriptions.

Exam trap

AWS often tests the misconception that simply using a larger or more powerful model (Option C) is the universal fix for accuracy issues, when in fact the root cause of hallucinations is often a lack of grounded, retrievable context that RAG specifically addresses.

How to eliminate wrong answers

Option A is wrong because increasing the temperature parameter makes the model's output more random and creative, which would likely increase, not decrease, the frequency of incorrect information. Option C is wrong because using a larger foundation model does not inherently fix factual accuracy; larger models can still hallucinate or produce outdated information without access to current or domain-specific data. Option D is wrong because AWS WAF is a web application firewall that filters HTTP traffic for security threats (e.g., SQL injection, XSS) and has no mechanism to validate or correct the factual accuracy of generated text.

85
MCQeasy

A media company wants to automatically generate short video captions from uploaded audio files using a foundation model on AWS. The solution should transcribe speech and then produce concise captions. Which combination of AWS services should they use?

A.Amazon Kinesis Data Streams to ingest audio, then Amazon Bedrock to generate captions directly from the audio stream.
B.Amazon Rekognition to analyze the audio, then Amazon Bedrock to generate captions.
C.Amazon Polly to convert audio to text, then Amazon Comprehend to generate captions.
D.Amazon Transcribe to convert audio to text, then Amazon Bedrock to generate captions from the transcript.
AnswerD

Amazon Transcribe performs automatic speech recognition to convert audio into text, and Amazon Bedrock provides access to foundation models that can summarize or rewrite that transcript into concise captions. This two-step pipeline matches the requirement: transcription followed by generative caption creation. Both are managed AWS services, so the company avoids building and operating its own speech and language models.

Why this answer

The task requires two capabilities: speech-to-text and generative text creation. Amazon Transcribe provides accurate automatic speech recognition, and Amazon Bedrock supplies foundation models that turn the transcript into concise captions. Together they form a managed, scalable pipeline that meets the scenario without custom model development.

Exam trap

The trap here is confusing text-to-speech with speech-to-text, or assuming a foundation model can accept raw audio input directly.

86
Multi-Selectmedium

Which TWO actions can help reduce bias in a foundation model’s outputs? (Choose two.)

Select 2 answers
A.Fine-tune the model on a balanced, representative dataset
B.Use careful prompt engineering with neutral wording
C.Restrict model access to a subset of users
D.Increase temperature to add randomness
E.Use a larger foundation model
AnswersA, B

Fine-tuning on a balanced, representative dataset corrects skewed statistical associations in the training data, so the model's outputs reflect fairer distributions across demographic groups. This directly targets the root cause of bias rather than merely masking its symptoms at inference time.

Why this answer

Option A is correct because fine-tuning on a balanced, representative dataset directly addresses bias by exposing the model to fair, diverse examples during training, which adjusts its learned weights and reduces skewed associations in its outputs. Option B is correct because careful prompt engineering with neutral wording avoids injecting biased framing or leading context into the input, so the model is less likely to amplify stereotypes or produce skewed responses. Option C is incorrect because restricting access to a subset of users is an access-control measure that does nothing to change the model's inherent biases.

Option D is incorrect because increasing temperature only adds randomness to token sampling, which can make outputs more erratic rather than less biased. Option E is incorrect because a larger foundation model does not inherently reduce bias and may even reproduce or amplify biases present in its larger training corpus.

Exam trap

AWS AI Practitioner exam often tests the misconception that increasing randomness (temperature) or model size can inherently fix bias, when in fact these changes do not address the underlying data or prompt-level causes of biased outputs.

87
Multi-Selectmedium

Which TWO actions are recommended for improving the factual accuracy of a foundation model's responses when using RAG?

Select 2 answers
A.Include relevant context from the knowledge base in the prompt
B.Increase the max_tokens parameter
C.Provide clear instructions in the system prompt
D.Use the largest foundation model available
E.Increase the temperature parameter
AnswersA, C

Injecting retrieved passages into the prompt gives the model authoritative source text to condition on, so answers cite supplied evidence rather than relying on parametric memory. This grounds generation in the knowledge base, directly improving factual accuracy.

Why this answer

Option A is correct because RAG's core mechanism is retrieving relevant passages from the knowledge base and injecting them into the prompt as grounding context, which directly supplies the model with the facts it needs and reduces hallucination. Option C is correct because clear system-prompt instructions (for example, telling the model to answer only from the provided context and to say it doesn't know when the context is insufficient) constrain the model's behavior and improve factual fidelity. Option B is not recommended because max_tokens only caps response length and does not affect factual grounding.

Option D is not recommended because a larger model does not guarantee factual accuracy and does not address retrieval quality. Option E is not recommended because raising temperature increases randomness and creativity, which typically worsens factual accuracy.

Exam trap

AWS often tests the misconception that larger models or higher randomness (temperature) inherently improve response quality, but in RAG, factual accuracy depends on retrieval quality and prompt engineering, not model size or creativity parameters.

88
Multi-Selectmedium

A data scientist is using a foundation model to summarize long documents. Which TWO of the following steps are most likely to improve the quality of the summaries?

Select 2 answers
A.Break the input document into chunks and summarize each chunk separately.
B.Use a high temperature parameter to increase creativity.
C.Provide few-shot examples of desired summaries in the prompt.
D.Use a low frequency penalty to reduce repetition.
E.Use a longer context length by increasing the max tokens parameter.
AnswersA, C

Chunking splits the long document into segments that fit the model's context window, letting each chunk be summarised without truncation. This satisfies the long-document constraint, since whole-document input would exceed context limits and lose detail, degrading summary quality.

Why this answer

Option A is correct because chunking a long document and summarizing each chunk separately keeps each request within the model's effective context window, avoiding truncation and the degraded recall that occurs when a foundation model must attend to very long inputs, and the chunk summaries can then be combined into a final summary. Option C is correct because few-shot prompting supplies concrete examples of the desired summary style, length, and level of detail, which steers the model's output distribution toward the target format more reliably than a bare instruction. Option B is not appropriate because a high temperature increases randomness and creativity, which harms factual fidelity and consistency in summarization.

Option D is not the priority here because frequency penalty only discourages repeated tokens and does not address the core problems of long-input context limits or output style alignment. Option E is not correct because increasing max tokens only raises the output length cap; it does not extend the model's usable context window for the input document and can even encourage overly long, unfocused summaries.

Exam trap

AWS often tests the misconception that increasing max tokens extends the model's input capacity, when in reality it only controls the output length, while the input is constrained by the model's inherent context window.

89
MCQmedium

A team deployed a text generation model on Amazon Bedrock. They want to monitor for toxic content in model outputs. Which evaluation approach is MOST effective?

A.Enable CloudWatch Logs and set a metric filter for toxic words
B.Use Amazon SageMaker Ground Truth for human annotation
C.Manually review a sample of outputs each week
D.Use Amazon Bedrock Model Evaluation with toxicity metrics
AnswerD

Bedrock Model Evaluation runs automatic toxicity metrics against model outputs, directly satisfying the requirement to monitor generated text for toxic content. It provides scored, repeatable assessment rather than ad-hoc inspection, making it the most effective ongoing evaluation approach for this deployment.

Why this answer

Amazon Bedrock Model Evaluation with toxicity metrics is the most effective approach because it provides automated, built-in evaluation of model outputs for toxic content using predefined metrics, directly integrated with the Bedrock service. This eliminates the need for manual effort or custom filtering, ensuring consistent and scalable monitoring of harmful content.

Exam trap

The trap here is that candidates may choose CloudWatch metric filters (Option A) because they associate monitoring with logs, but fail to recognize that toxicity detection requires semantic understanding beyond simple keyword matching.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs with a metric filter for toxic words is a simplistic, keyword-based approach that cannot detect nuanced or context-dependent toxicity, such as sarcasm or implicit hate speech, and requires manual setup of word lists. Option B is wrong because Amazon SageMaker Ground Truth for human annotation is designed for creating labeled datasets, not for real-time or automated monitoring of model outputs, and introduces latency and cost overhead. Option C is wrong because manually reviewing a sample of outputs each week is not scalable, introduces human bias, and fails to provide continuous or real-time monitoring, making it ineffective for production systems.

90
MCQmedium

A data scientist uses Amazon Bedrock. The model responses are too long. Which parameter should they adjust to limit the output length?

A.temperature
B.max_tokens
C.stop sequences
D.top_p
AnswerB

Setting max_tokens caps the number of tokens the model generates, directly satisfying the requirement to limit response length. In Amazon Bedrock, this parameter bounds output size independently of the prompt, so the data scientist can truncate verbose answers without altering the input or switching models.

Why this answer

The `max_tokens` parameter directly controls the maximum number of tokens (words or subwords) the model can generate in a single response. By reducing this value, the data scientist caps the output length, preventing overly long responses. Temperature and top_p affect randomness and diversity, not length, while stop sequences define when generation halts but do not enforce a hard token limit.

Exam trap

AWS often tests the distinction between parameters that control output length (`max_tokens`) versus those that control output randomness or diversity (`temperature`, `top_p`), leading candidates to confuse 'limiting length' with 'limiting creativity'.

How to eliminate wrong answers

Option A is wrong because temperature controls the randomness of token selection (higher values increase creativity, lower values make output more deterministic), not the length of the response. Option C is wrong because stop sequences are custom strings (e.g., '###' or 'END') that tell the model to cease generation when encountered, but they do not limit the total number of tokens generated before that point. Option D is wrong because top_p (nucleus sampling) limits the cumulative probability of token choices to a threshold (e.g., 0.9), affecting diversity, not the maximum output length.

91
MCQmedium

A company uses a foundation model for real-time translation in a chat application. The latency is high. Which optimization would reduce latency the most?

A.Increase batch size
B.Use model distillation to create a smaller model
C.Use a larger model
D.Use a CDN for model weights
AnswerB

Distillation trains a compact student model to mimic the larger teacher, cutting inference compute and memory per token. Fewer parameters mean faster forward passes, directly addressing the real-time chat latency constraint rather than merely tuning prompts or batching.

Why this answer

Model distillation reduces the size of the foundation model by training a smaller 'student' model to mimic the behavior of a larger 'teacher' model. This directly decreases inference latency because the smaller model requires fewer computational resources (FLOPs) per forward pass, which is critical for real-time translation in a chat application where low latency is paramount.

Exam trap

The AIF-C01 exam often tests the distinction between throughput optimization (batch size) and latency optimization (model size/distillation), leading candidates to mistakenly choose increasing batch size when the question explicitly asks for reducing latency.

How to eliminate wrong answers

Option A is wrong because increasing batch size improves throughput (more requests processed per unit time) but does not reduce per-request latency; in fact, it can increase latency for individual requests as the model must wait for the batch to fill. Option C is wrong because using a larger model increases the number of parameters and computational complexity, which would increase latency, not reduce it. Option D is wrong because a CDN for model weights only accelerates the initial download of the model to edge locations, not the inference latency of each translation request; once the model is loaded, inference speed is determined by the model architecture and hardware, not network delivery.

92
MCQhard

A media company uses Amazon Bedrock to generate article summaries. They notice that for long articles, the model sometimes ignores instructions placed at the beginning of the prompt. The company wants to improve the model's adherence to instructions without changing the model or increasing cost significantly. Which prompt engineering technique should they apply?

A.Use a higher temperature value to make the model more creative in following instructions.
B.Split the article into smaller chunks and summarize each chunk separately, then concatenate the summaries.
C.Increase the maxTokenCount parameter to allow the model to process more of the article.
D.Move the instructions to the end of the prompt and repeat them after the article text.
AnswerD

Placing instructions at the end of the prompt, after the long context, leverages the model's tendency to pay more attention to recent tokens. Repeating key instructions after the article text reinforces them, improving adherence without changing the model or adding significant cost. This is a known prompt engineering technique for long-context scenarios.

Why this answer

The model's tendency to overlook early instructions in long contexts is a known limitation. By moving instructions to the end of the prompt and repeating them after the article, the company places them where the model's attention is strongest. This prompt engineering technique improves adherence without retraining or increasing cost, making it the most effective solution.

Exam trap

The trap here is thinking that increasing maxTokenCount or temperature will fix instruction-following, when the real issue is prompt structure and attention placement.

93
MCQmedium

Refer to the exhibit. A data scientist created this endpoint config for a foundation model in Amazon SageMaker. However, the endpoint fails to scale under load. What is the most likely reason?

A.Missing AutoScaling configuration
B.Variant weight is 1.0
C.Instance type is too small
D.InitialInstanceCount is 1
AnswerA

SageMaker endpoints do not scale automatically unless an autoscaling policy is attached to the endpoint variant. Without Application Auto Scaling configured against a target metric such as InvocationsPerInstance, the endpoint keeps fixed instance count and cannot absorb rising load.

Why this answer

The endpoint fails to scale under load because the endpoint configuration shown lacks an AutoScaling policy. Without AutoScaling, SageMaker will not automatically adjust the number of instances based on traffic, so even if the initial instance count is 1, the endpoint cannot add more instances to handle increased load. AutoScaling must be explicitly configured via Application Auto Scaling to define scaling policies and target tracking metrics.

Exam trap

AWS often tests the misconception that setting a higher InitialInstanceCount or choosing a larger instance type alone enables scaling, when in fact AutoScaling must be explicitly configured as a separate step.

How to eliminate wrong answers

Option B is wrong because a variant weight of 1.0 is the default and does not prevent scaling; it simply means all traffic is routed to that variant. Option C is wrong because the instance type being 'too small' would cause performance issues or throttling, but it does not prevent the endpoint from scaling out; scaling is controlled by AutoScaling, not instance size. Option D is wrong because an InitialInstanceCount of 1 is a valid starting point; the endpoint can still scale out if AutoScaling is configured, so a single initial instance does not inherently block scaling.

94
MCQmedium

A financial services company is building an application on Amazon Bedrock that generates personalized investment summaries. The compliance team requires that every generated summary includes an exact, verifiable citation from the company's approved regulatory documents, and that the model must not fabricate any citation. The company has a large corpus of approved PDF documents stored in Amazon S3. Which approach should the company use to meet these requirements?

A.Use Amazon Bedrock Knowledge Bases to ingest the S3 documents and configure the model to generate responses grounded in retrieved passages, then verify the returned source citations.
B.Enable model invocation logging in Amazon Bedrock and audit the logs to confirm that any citation the model produces exists in the S3 corpus.
C.Increase the model's temperature setting so the model explores more of its training data and is more likely to recall the exact regulatory text.
D.Fine-tune the foundation model on the approved regulatory documents using Amazon Bedrock custom models, then rely on the fine-tuned model to reproduce exact citations from memory.
AnswerA

Amazon Bedrock Knowledge Bases performs managed retrieval-augmented generation by ingesting documents from Amazon S3 into a vector store, retrieving relevant passages at query time, and returning responses with source attribution. This directly satisfies the requirement for grounded, verifiable citations while reducing fabrication, because the model conditions its output on retrieved approved content rather than only on its pretrained parameters.

Why this answer

Grounding the model in an approved document corpus is the reliable way to produce verifiable citations. Amazon Bedrock Knowledge Bases ingests the S3 documents, retrieves relevant passages per query, and returns source attributions, so generated summaries reference real approved text instead of relying on parametric memory. Sampling settings, fine-tuning, and logging do not provide retrieval-backed citation integrity.

Exam trap

The trap here is assuming that fine-tuning or a higher temperature makes a model recall exact source text, when only retrieval-based grounding reliably ties output to specific approved documents.

95
MCQmedium

A media company uses Amazon Bedrock to generate personalized news summaries. They notice that summaries sometimes include details not present in the source articles. They want to reduce these hallucinations without retraining the model. Which approach should they use?

A.Decrease the maximum token length in the InvokeModel request.
B.Switch to a larger foundation model with more parameters.
C.Use Retrieval Augmented Generation (RAG) by retrieving relevant passages from the source articles and including them in the prompt.
D.Increase the model's temperature setting to encourage more creative outputs.
AnswerC

RAG grounds the model's response by supplying relevant, authoritative content from the source articles directly in the prompt. This reduces hallucination because the model conditions its output on the retrieved text rather than relying solely on parametric knowledge. It requires no retraining and integrates with Amazon Bedrock Knowledge Bases or custom retrieval, making it suitable for dynamic news content.

Why this answer

Retrieval Augmented Generation supplies the model with relevant excerpts from the source articles at inference time, so the generated summary is conditioned on actual content rather than the model's internal knowledge. This reduces fabricated details without retraining. Other options either increase randomness, rely on model size, or truncate output, none of which ground the response in the provided material.

Exam trap

The trap here is assuming that a larger or more capable model will automatically stop hallucinating, when grounding the prompt with retrieved source content is what actually constrains the output.

96
MCQmedium

A financial services company uses Amazon Bedrock with the Anthropic Claude 3 Sonnet model to answer employee questions about internal policies. The policy documents are updated frequently, and the model occasionally provides outdated or incorrect policy details. The company wants the model to base its answers on the most current authoritative documents without retraining the model. Which approach should they use?

A.Switch to a larger foundation model with a higher parameter count.
B.Use Retrieval Augmented Generation (RAG) by storing policy documents in an Amazon Bedrock knowledge base and retrieving relevant passages at inference time.
C.Fine-tune the foundation model on the updated policy documents.
D.Increase the model's temperature setting to encourage more creative and up-to-date answers.
AnswerB

RAG with an Amazon Bedrock knowledge base indexes the policy documents and retrieves the most relevant passages for each query, then passes them to the model as context. This grounds responses in the latest authoritative content without retraining, and updating the knowledge base data source keeps answers current. It directly addresses outdated or incorrect policy details by supplying fresh source text at inference time.

Why this answer

Retrieval Augmented Generation with an Amazon Bedrock knowledge base retrieves relevant, current policy passages and supplies them to the model as context, so answers are grounded in authoritative documents. It avoids retraining costs and keeps pace with frequent updates, directly solving the outdated or incorrect policy detail problem.

Exam trap

The trap here is assuming that a larger or fine-tuned model automatically knows frequently updated internal documents, when grounding requires retrieval of current source content at inference time.

97
MCQmedium

A developer is trying to invoke the Claude v2 model in Amazon Bedrock from a Lambda function. The Lambda function's IAM role has the following policy attached: { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "bedrock:InvokeModel", "Resource": "*" } ] } When the Lambda function runs, it receives the error shown in the exhibit. Which additional step is most likely needed to resolve this issue?

A.Change the AWS region to one where Claude v2 is available.
B.Use a different model ID such as 'anthropic.claude-v1'.
C.Request access to the Anthropic Claude model through the Amazon Bedrock console.
D.Add a condition to the IAM policy to specify the model ARN.
AnswerC

Model access in Amazon Bedrock is granted separately from IAM permissions. Even with bedrock:InvokeModel allowed, the account must first request access to the Anthropic Claude model in the Bedrock console, otherwise invocation returns an access-denied error despite the permissive policy.

Why this answer

Amazon Bedrock requires explicit user access approval for third-party foundation models like Anthropic Claude before they can be invoked. Even with a valid IAM policy allowing bedrock:InvokeModel on all resources, the model itself must be granted access via the Bedrock console's 'Model access' section. Without this step, the API returns an access denied error regardless of IAM permissions.

Exam trap

The trap here is that candidates assume a wildcard IAM policy (Resource: '*') grants full access, but Bedrock requires an additional explicit model access approval that is independent of IAM, causing them to incorrectly focus on policy or region changes.

How to eliminate wrong answers

Option A is wrong because the error is not due to regional availability; Claude v2 is available in multiple regions (e.g., us-west-2, us-east-1) and the region choice does not bypass the model access requirement. Option B is wrong because using a different model ID like 'anthropic.claude-v1' would still require explicit model access for that specific model; the error is about access, not model version. Option D is wrong because adding a condition to specify the model ARN does not resolve the missing model access approval; the IAM policy already allows all resources, and the issue is at the service level, not the policy level.

98
MCQeasy

A startup is deploying a foundation model on Amazon SageMaker for real-time inference. They notice high latency (over 2 seconds per request). Which action is most likely to reduce latency?

A.Enable auto-scaling on the SageMaker endpoint to handle more concurrent requests.
B.Switch to a smaller, distilled version of the model.
C.Deploy the model on a CPU-based instance instead of GPU.
D.Increase the batch size parameter in the inference request.
AnswerB

A smaller, distilled model reduces the number of parameters and floating-point operations per inference, directly cutting compute time on the SageMaker endpoint. This satisfies the stem's real-time latency constraint, since inference duration scales with model size; distillation preserves much of the original accuracy while delivering sub-second responses.

Why this answer

Using a smaller, distilled version of the model directly reduces the computational complexity per inference request. Distillation compresses the model by training a smaller student network to mimic a larger teacher model, resulting in fewer parameters and faster forward passes. This is the most direct way to cut latency when the model size is the bottleneck, as it reduces the number of floating-point operations (FLOPs) required per request.

Exam trap

AWS often tests the distinction between latency (time per single request) and throughput (requests per second), so candidates mistakenly choose auto-scaling or batch size increases, which improve throughput but not per-request latency.

How to eliminate wrong answers

Option A is wrong because enabling auto-scaling adds more endpoint instances to handle higher concurrency, but it does not reduce the latency of a single inference request; it only improves throughput under load. Option C is wrong because CPU-based instances are generally slower for deep learning inference than GPU instances, especially for large foundation models, so switching to CPU would increase latency, not reduce it. Option D is wrong because increasing the batch size in the inference request means processing multiple inputs together, which increases the time to first byte for each individual request and does not reduce per-request latency; it is a throughput optimization, not a latency reduction technique.

99
MCQeasy

A startup wants to add image generation to its design tool. The team does not want to manage GPU infrastructure, train models, or host endpoints, and they want to call a managed API for a text-to-image foundation model available in Amazon Bedrock. Which action should they take?

A.Use an Amazon Bedrock image generation model through the InvokeModel API with a text prompt.
B.Use Amazon Polly to convert text prompts into images for the design tool.
C.Use Amazon Rekognition to synthesize new design images from text descriptions.
D.Deploy an open-source diffusion model on Amazon SageMaker endpoints and manage scaling themselves.
AnswerA

Amazon Bedrock provides serverless access to image generation foundation models, including Amazon Titan Image Generator, through the InvokeModel API. The startup sends a text prompt and receives generated images without provisioning GPUs or managing endpoints. This matches the requirement for a managed text-to-image API with no infrastructure ownership.

Why this answer

Amazon Bedrock exposes image generation foundation models such as Amazon Titan Image Generator as managed, serverless APIs invoked with a text prompt. This delivers text-to-image capability without GPU provisioning, training, or endpoint management. SageMaker self-hosting adds operational work, while Rekognition analyzes images and Polly synthesizes speech, so neither can generate design images from text.

Exam trap

The trap here is matching the word image to Amazon Rekognition, which analyzes existing images rather than generating new ones from text.

100
Multi-Selecthard

Which TWO of the following are valid methods to reduce the risk of foundation models generating harmful or biased content?

Select 2 answers
A.Use a smaller model
B.Use a content filter
C.Apply prompt engineering to guide output
D.Fine-tune the model on a biased dataset
E.Disable all logging
AnswersB, C

Content filters intercept prompts and completions at runtime, blocking harmful or biased output before it reaches users. This directly satisfies the stem's requirement to reduce risk from foundation models, since filtering operates on the model's generated content itself rather than on training data or access controls.

Why this answer

Option B (Use a content filter) is correct because content filters act as a post-processing guardrail that screens both prompts and model completions for harmful, violent, hateful, or otherwise policy-violating content, blocking or redacting it before it reaches users. Option C (Apply prompt engineering to guide output) is correct because carefully crafted system prompts, few-shot examples, and instructions can steer the foundation model toward safe, neutral, and on-topic responses, reducing the likelihood of biased or harmful generations. Option A (Use a smaller model) is not a valid mitigation because model size does not determine safety or bias; smaller models can still produce harmful or biased content.

Option D (Fine-tune the model on a biased dataset) would actually increase the risk by reinforcing biased patterns in the model's outputs. Option E (Disable all logging) does not reduce harmful content generation and instead removes the audit trail needed to detect, monitor, and remediate such issues.

Exam trap

AWS often tests the misconception that simply using a smaller model or disabling logging can reduce bias, when in fact these actions either have no effect or worsen the problem, whereas content filters and prompt engineering are direct, effective mitigation strategies.

101
MCQmedium

An e-commerce company uses Amazon Bedrock to generate product descriptions from keywords. Some descriptions contain inaccurate details about product specifications. Which approach should the company take to reduce factual errors?

A.Increase the maxTokens parameter to allow more detailed descriptions.
B.Use a different foundation model from Bedrock for each product category.
C.Deploy the model to a SageMaker endpoint and use human-in-the-loop validation.
D.Include the product specifications in the prompt and instruct the model to base the description on the provided data.
AnswerD

Supplying the specifications directly in the prompt grounds generation in authoritative source data, so the model conditions its output on those facts rather than relying on parametric knowledge that may be outdated or hallucinated. This directly satisfies the stem's constraint of reducing inaccurate specification details, since the model is instructed to base descriptions solely on the provided data.

Why this answer

Providing the product specifications directly in the prompt and instructing the model to base the description on that data grounds the generation in factual information, reducing hallucinations. This technique, known as prompt engineering with in-context learning, ensures the model uses the given data rather than relying on its training data, which may contain inaccuracies.

Exam trap

AWS often tests the misconception that increasing model parameters or changing models alone improves factual accuracy, when in fact prompt engineering with grounded data is the most effective and efficient method to reduce hallucinations.

How to eliminate wrong answers

Option A is wrong because increasing maxTokens only allows longer outputs but does not improve factual accuracy; it may even increase the chance of hallucinations by generating more unverified content. Option B is wrong because using a different foundation model for each category does not inherently reduce factual errors; all models can hallucinate, and this approach adds complexity without addressing the root cause of inaccurate specifications. Option C is wrong because deploying to a SageMaker endpoint with human-in-the-loop validation is an operational pattern for custom models, but it is overkill and inefficient for this use case; prompt engineering (Option D) is a simpler, more direct solution that avoids the latency and cost of human review for every generation.

102
MCQhard

A software company is deploying a generative AI assistant on Amazon Bedrock. They need the assistant to include citations to source documents in its answers and to avoid answering when no supporting document is found. Which configuration should they use?

A.Enable model invocation logging and parse the logs to extract source references.
B.Use a higher temperature setting so the model explores more possible answers, including citations.
C.Configure an Amazon Bedrock knowledge base and enable citation generation, and instruct the model to answer only from retrieved context.
D.Increase the model's maximum token count so it can include citations in its response.
AnswerC

Amazon Bedrock knowledge bases can return citations that link generated statements to the retrieved source chunks, and a system prompt can instruct the model to answer only when supporting context exists. This satisfies both requirements: source citations and abstention when no document is found. It leverages built-in retrieval and citation features rather than custom post-processing.

Why this answer

An Amazon Bedrock knowledge base with citation generation returns responses that reference the retrieved source chunks, and a prompt that restricts answers to retrieved context makes the assistant abstain when no supporting document exists. This combination meets both the attribution and the no-support behavior requirements.

Exam trap

The trap here is assuming that logging or sampling parameters can produce citations, when attribution and abstention depend on retrieval configuration and prompt constraints.

103
MCQeasy

A company wants to automatically summarize customer support tickets into a short paragraph. Which AWS service is MOST appropriate for this task?

A.Amazon Bedrock
B.Amazon Rekognition
C.Amazon Polly
D.Amazon Comprehend
AnswerA

Amazon Bedrock offers access to foundation models that perform abstractive summarisation, condensing ticket text into a short paragraph. It provides a managed, serverless API without infrastructure management, making it the most appropriate service for this natural language generation task.

Why this answer

Amazon Bedrock provides access to foundation models (FMs) from providers like Anthropic and AI21 that excel at natural language generation tasks, including summarization. By invoking a model such as Claude or Jurassic-2 via Bedrock's API, you can pass the customer support ticket text and receive a concise paragraph summary. This makes Bedrock the most appropriate service for generative summarization.

Exam trap

The trap here is that candidates may confuse Amazon Comprehend's text analysis features (like key phrase extraction) with generative summarization, but Comprehend cannot produce new text; it only extracts or classifies existing content.

How to eliminate wrong answers

Option B (Amazon Rekognition) is wrong because it is designed for image and video analysis, not text summarization. Option C (Amazon Polly) is wrong because it converts text to speech, not text summarization. Option D (Amazon Comprehend) is wrong because it performs natural language processing tasks like entity extraction and sentiment analysis, but it does not generate new text summaries; it lacks generative capabilities.

104
Multi-Selecthard

Which THREE are benefits of using Amazon Bedrock over self-managing foundation models on EC2? (Choose THREE.)

Select 3 answers
A.Built-in integration with AWS services such as AWS CloudWatch and AWS CloudTrail.
B.Lower data transfer costs between cloud regions.
C.Access to a curated set of foundation models from different providers.
D.Managed infrastructure for model hosting and scaling.
E.Greater control over model fine-tuning and customization.
AnswersA, C, D

Amazon Bedrock natively emits metrics to CloudWatch and logs API activity to CloudTrail without custom instrumentation, satisfying the operational visibility requirement that self-managed EC2 inference stacks must build and maintain themselves. This removes undifferentiated monitoring plumbing, letting teams focus on model usage rather than telemetry infrastructure.

Why this answer

Option A is correct because Amazon Bedrock natively integrates with AWS monitoring and auditing services like CloudWatch (for metrics/logs) and CloudTrail (for API activity logging), giving visibility and compliance without custom instrumentation. Option C is correct because Bedrock provides a curated catalog of foundation models from multiple providers (e.g., Anthropic, AI21, Stability AI, Amazon Titan, Meta) through a single API, which self-managing on EC2 would require sourcing and integrating individually. Option D is correct because Bedrock is a fully managed service that handles model hosting, provisioning, and automatic scaling of inference capacity, removing the operational burden of managing EC2 instances, AMIs, autoscaling groups, and GPU capacity.

Option B is not correct because Bedrock does not inherently reduce inter-region data transfer costs; those are governed by standard AWS data transfer pricing and are unrelated to the choice between Bedrock and self-managed EC2. Option E is not correct because self-managing foundation models on EC2 generally offers greater control over fine-tuning and customization, whereas Bedrock provides more constrained, managed customization options.

Exam trap

The trap here is that candidates may confuse 'managed infrastructure' with 'greater control'—Bedrock simplifies operations but reduces customization flexibility, so option E is a common distractor for those who think managed services offer more control than self-managed solutions.

105
MCQmedium

A developer is building a retrieval-augmented generation (RAG) assistant on Amazon Bedrock. The assistant must answer questions about internal policy documents that change frequently, and answers must cite the source passages. The developer wants a managed capability that handles chunking, embedding, and retrieval so the application code stays minimal. Which Amazon Bedrock feature should the developer use?

A.Amazon Bedrock model evaluation
B.Amazon Bedrock knowledge bases
C.Amazon Bedrock Agents
D.Amazon Bedrock Guardrails
AnswerB

Amazon Bedrock knowledge bases provide a managed RAG workflow: they ingest and chunk source documents, generate embeddings, store them in a vector store, and retrieve relevant passages at query time. The model can then generate answers grounded in those retrieved passages. This removes the need for the developer to orchestrate chunking, embedding, and retrieval manually, matching the requirement for minimal application code.

Why this answer

Amazon Bedrock knowledge bases are the managed RAG capability that handles ingestion, chunking, embedding, vector storage, and retrieval, enabling grounded answers with source citations while keeping application code small. Evaluation, Guardrails, and Agents address quality measurement, content safety, and task orchestration respectively, none of which deliver the required retrieval pipeline on their own.

Exam trap

The trap here is conflating orchestration or safety features with retrieval, when only knowledge bases perform the managed chunking, embedding, and retrieval work.

106
MCQmedium

A financial services company is deploying a foundation model to analyze customer sentiment from call transcripts. The model outputs must be consistent and deterministic for auditing purposes. Which parameter configuration should the company use?

A.Set temperature to 0.1 and top_p to 0.9.
B.Set temperature to 0.7 and top_p to 1.0.
C.Set temperature to 0.5 and top_p to 0.5.
D.Set temperature to 0 and top_p to 1.
AnswerD

Setting temperature to 0 makes the model select the highest-probability token at each step, eliminating the random sampling that produces run-to-run variation. Keeping top_p at 1 disables nucleus filtering, so it cannot reintroduce randomness. This satisfies the audit requirement for consistent, deterministic sentiment outputs from identical transcripts.

Why this answer

Setting temperature to 0 and top_p to 1 forces the model to always select the highest-probability token at each step, producing deterministic and repeatable outputs. This is essential for auditing and compliance in financial services, where consistency is required. Any nonzero temperature introduces randomness, which undermines determinism.

Exam trap

AWS often tests the misconception that low temperature (e.g., 0.1) is 'deterministic enough,' but only temperature exactly 0 guarantees deterministic outputs, and top_p must be 1 to avoid interfering with the argmax selection.

How to eliminate wrong answers

Option A is wrong because temperature 0.1 still introduces slight randomness, making outputs non-deterministic and unsuitable for auditing. Option B is wrong because temperature 0.7 introduces significant randomness, and top_p 1.0 does not constrain it, leading to high variability. Option C is wrong because temperature 0.5 introduces randomness, and top_p 0.5 further restricts token sampling but does not eliminate the stochastic behavior from the nonzero temperature.

107
MCQhard

Refer to the exhibit. A developer sees this error when calling Amazon Bedrock for inference. What is the MOST likely cause and recommended solution?

A.The model ID is incorrect; use a different model
B.The prompt is too long; reduce the number of tokens in the prompt
C.The request rate exceeds the model's throughput limit; implement retries with exponential backoff
D.Increase the max_tokens_to_sample value
AnswerC

ThrottlingException indicates the account's requests per minute or tokens per minute exceed the model's allocated throughput. Retrying with exponential backoff and jitter spreads retries, letting transient capacity recover instead of compounding the overload with immediate repeat calls.

Why this answer

The error indicates a throttling exception from Amazon Bedrock, which occurs when the request rate exceeds the model's throughput limit. The recommended solution is to implement retries with exponential backoff to handle transient rate limits gracefully, as this aligns with AWS best practices for managing API call limits.

Exam trap

The trap here is that candidates may confuse a throttling error with a model ID or prompt length issue, because the error message may not explicitly state 'throttling' and instead show a generic 'ServiceUnavailable' or 'TooManyRequests' response, leading them to incorrectly modify the model or prompt instead of implementing retry logic.

How to eliminate wrong answers

Option A is wrong because a model ID error would produce a different error (e.g., 'ValidationException' or 'ResourceNotFoundException'), not a throttling-related error. Option B is wrong because a prompt that is too long would cause a 'ValidationException' regarding token limits, not a throttling error. Option D is wrong because increasing max_tokens_to_sample would increase the output length, potentially worsening throttling or causing a different error, but it does not address the rate limit issue.

108
MCQhard

A healthcare company uses Amazon Bedrock with a foundation model to generate patient education materials. They must ensure that the model does not include protected health information (PHI) in its responses, even if it appears in the prompt. Which Amazon Bedrock feature should they configure?

A.Amazon Bedrock Agents with an action group to query a patient database
B.Amazon Bedrock provisioned throughput
C.Guardrails for Amazon Bedrock with sensitive information filters
D.Amazon Bedrock model evaluation with toxicity metrics
AnswerC

Guardrails for Amazon Bedrock includes sensitive information filters that can detect and block or mask personally identifiable information (PII) and custom regex patterns. By configuring these filters, the company can prevent PHI from appearing in model responses. This directly addresses the requirement to avoid including PHI in generated patient education materials.

Why this answer

Guardrails for Amazon Bedrock provides sensitive information filters that can detect and block or mask PII and custom patterns in both prompts and responses. Configuring these filters prevents PHI from appearing in generated materials. Model evaluation, provisioned throughput, and Agents do not offer runtime content redaction, so they cannot meet the compliance requirement.

Exam trap

The trap here is assuming that model evaluation or agent-based data retrieval can enforce privacy, when only runtime guardrail filters can detect and redact sensitive information in responses.

109
MCQhard

A research team is using Amazon Bedrock to analyze scientific papers. They want the model to generate answers based only on papers published after 2023. Which approach should they use?

A.Fine-tune the model on a dataset of post-2023 papers and deploy it.
B.Set the maxTokens to a low value to force the model to rely on recent context.
C.Include a system prompt instructing the model to ignore data before 2023.
D.Use Amazon Bedrock Knowledge Bases with a metadata filter to retrieve only papers published after 2023, and generate responses based on retrieved content.
AnswerD

Amazon Bedrock Knowledge Bases with a metadata filter satisfies the post-2023 constraint by restricting vector retrieval to documents whose publication-date metadata matches the filter, so the foundation model generates answers grounded solely in that retrieved subset rather than its training data. This enforces the temporal restriction at retrieval time, preventing older papers from entering the context window.

Why this answer

Amazon Bedrock Knowledge Bases with a metadata filter allows you to restrict retrieval to only documents that match specific metadata criteria, such as publication year. By filtering the vector search to only include papers published after 2023, the model generates responses based solely on that retrieved content, ensuring it does not rely on pre-2023 data. This approach is the only one that guarantees the model's answers are grounded exclusively in the specified time range.

Exam trap

AWS often tests the misconception that a system prompt or fine-tuning can reliably restrict a model's knowledge to a specific time period, when in fact only a retrieval-based approach with metadata filtering can enforce such temporal constraints.

How to eliminate wrong answers

Option A is wrong because fine-tuning the model on a dataset of post-2023 papers does not prevent the model from using its pre-existing training data (which includes pre-2023 knowledge) during inference; fine-tuning adjusts weights but does not erase prior knowledge, so the model could still generate answers based on older information. Option B is wrong because setting maxTokens to a low value limits the length of the generated response but does not control the temporal scope of the model's knowledge; the model can still draw on pre-2023 training data regardless of token count. Option C is wrong because a system prompt instructing the model to ignore data before 2023 is merely a suggestion and not a technical enforcement; the model has no inherent mechanism to filter its own training data by date, so it may still generate answers based on pre-2023 information, especially if the prompt is not strictly followed.

110
MCQeasy

Which AWS service provides a serverless API for accessing foundation models with per-token pricing?

A.Amazon Bedrock
B.Amazon API Gateway
C.AWS Lambda
D.Amazon SageMaker
AnswerA

Amazon Bedrock offers a serverless, API-driven way to invoke foundation models from multiple providers, with usage billed per input and output token. That matches the stem's serverless API and per-token pricing constraints without managing infrastructure.

Why this answer

Amazon Bedrock is a fully managed service that provides a serverless API for accessing foundation models (FMs) from providers like AI21 Labs, Anthropic, Cohere, Meta, and Stability AI. It offers per-token pricing, meaning you pay only for the number of tokens processed in both input and output, with no upfront commitments or infrastructure management required.

Exam trap

The trap here is that candidates often confuse Amazon API Gateway (a serverless API front-end) with Bedrock's serverless model inference API, or mistakenly think AWS Lambda provides built-in FM access, when in fact Lambda is just compute and requires explicit integration with a model service.

How to eliminate wrong answers

Option B is wrong because Amazon API Gateway is a managed service for creating, publishing, and securing RESTful and WebSocket APIs, but it does not provide access to foundation models or per-token pricing; it is a front-end API layer that would need to be integrated with a backend service like Bedrock or Lambda. Option C is wrong because AWS Lambda is a serverless compute service that runs code in response to events, but it does not natively provide access to foundation models or per-token pricing; you would need to write custom code to call an FM API, and you pay per invocation and duration, not per token. Option D is wrong because Amazon SageMaker is a fully managed machine learning platform for building, training, and deploying custom models, but it is not a serverless API for foundation models with per-token pricing; it typically involves provisioning instances and paying for compute time, not per-token consumption.

111
MCQmedium

A company uses Amazon Bedrock to generate marketing copy. They want to measure the quality of generated text compared to reference text. Which metric is most appropriate?

A.F1 score
B.BLEU
C.RMSE
D.Accuracy
AnswerB

BLEU compares generated text against reference text using n-gram overlap, producing a precision-based score. It suits marketing copy evaluation where reference outputs exist. Perplexity, by contrast, measures language model likelihood without references, so it cannot assess similarity to target text.

Why this answer

BLEU (Bilingual Evaluation Understudy) is the most appropriate metric for evaluating the quality of generated text against reference text in tasks like machine translation and text generation. It measures n-gram precision between the generated and reference texts, making it ideal for assessing marketing copy generated by Amazon Bedrock.

Exam trap

AWS often tests the distinction between classification/regression metrics and text generation metrics, leading candidates to mistakenly apply F1 score or accuracy to evaluate generated text quality instead of using BLEU or similar sequence-based metrics.

How to eliminate wrong answers

Option A is wrong because F1 score is a classification metric that measures harmonic mean of precision and recall, not suitable for evaluating text generation quality against reference text. Option C is wrong because RMSE (Root Mean Square Error) is a regression metric used for continuous numerical predictions, not for text or sequence evaluation. Option D is wrong because Accuracy is a classification metric that measures the proportion of correct predictions, which does not account for the sequential and linguistic nuances of generated text.

112
MCQmedium

A retail company runs a product-question answering feature on Amazon Bedrock. During peak hours, requests intermittently fail with a ThrottlingException even though average usage is well within quota. The team needs a solution that smooths bursty traffic, retries failed calls, and avoids overwhelming the model endpoint, with minimal application code changes. Which approach should they take?

A.Introduce client-side retries with exponential backoff and jitter around the Amazon Bedrock InvokeModel calls.
B.Switch the application to call the model through Amazon Bedrock Provisioned Throughput and remove all retry logic.
C.Place an Amazon SQS queue between the application and Amazon Bedrock, and have a worker invoke the model as messages arrive.
D.Increase the maximum token count in each request so fewer total requests are needed during peak periods.
AnswerA

Exponential backoff with jitter retries throttled requests after progressively longer, randomized delays, which spreads retries out and prevents synchronized retry storms. It is a small, well-understood code change that directly addresses transient ThrottlingException errors during bursts. AWS SDKs and the Bedrock runtime support configurable retry behaviour, making this the lowest-effort effective fix.

Why this answer

Bursty traffic that briefly exceeds per-account or per-model quotas produces throttling that is transient by nature. Retrying with exponential backoff and jitter lets the client back off and spread retries, converting failures into eventual successes without architectural change. Provisioned Throughput, queues, and larger token limits either cost more, add latency, or do not address retry behaviour, so backoff and jitter is the targeted remedy.

Exam trap

The trap here is treating throttling as a hard capacity shortage that requires Provisioned Throughput, when bursty transient throttling is usually resolved with backoff and jitter retries.

113
MCQeasy

A company wants to use a foundation model to automatically summarize lengthy documents. Which capability of foundation models is being utilized?

A.Text generation
B.Sentiment analysis
C.Text classification
D.Machine translation
AnswerA

Summarisation is a text-generation task: the model consumes the source document as input and autoregressively produces a condensed natural-language output. This directly satisfies the stem's requirement to automatically summarise lengthy documents, since the capability being exercised is generating new text rather than classification, embedding or retrieval.

Why this answer

Summarization is a text generation task where the model produces a concise version of the original content. Foundation models (e.g., GPT, Claude) are pre-trained on vast corpora and can generate coherent summaries by predicting the next tokens conditioned on the input document. This directly utilizes the text generation capability, not classification or translation.

Exam trap

The AIF-C01 exam often tests the distinction between text generation and text classification, so the trap here is that candidates may confuse summarization (a generative task) with classification or analysis tasks, especially when the question emphasizes 'understanding' the document rather than 'producing' new text.

How to eliminate wrong answers

Option B (Sentiment analysis) is wrong because it involves classifying the emotional tone of text (positive, negative, neutral), not generating a summary. Option C (Text classification) is wrong because it assigns predefined labels or categories to text, whereas summarization requires generating new text. Option D (Machine translation) is wrong because it converts text from one language to another, not condensing content within the same language.

114
Multi-Selectmedium

A retail company wants to build an application that uses a foundation model on Amazon Bedrock to answer customer questions about product availability. They need the model to access real-time inventory data from their internal database and perform actions such as reserving an item. Which TWO capabilities should they implement to achieve this? (Choose two.)

Select 2 answers
A.Enable model invocation logging to capture inventory queries.
B.Use Retrieval Augmented Generation with a static document store containing product manuals.
C.Define an action group in the agent that maps to an API schema for the inventory and reservation operations.
D.Use Amazon Bedrock Agents to orchestrate calls to an AWS Lambda function that queries the inventory database.
E.Increase the model's temperature to allow more creative problem-solving.
AnswersC, D

Action groups in Amazon Bedrock Agents use OpenAPI schemas to describe available operations. By defining an action group with the inventory query and reservation endpoints, the agent knows how and when to call them. This enables the model to perform the required actions through the agent's orchestration, making it a correct implementation step.

Why this answer

Amazon Bedrock Agents orchestrate tasks by invoking action groups, which are defined with OpenAPI schemas and backed by Lambda functions. This allows the agent to query real-time inventory and perform reservations. RAG over static documents, temperature changes, and logging do not provide live data access or transactional capabilities, so they cannot satisfy the requirements.

Exam trap

The trap here is assuming that RAG alone can provide real-time data and actions, when live database access and transactions require agent action groups backed by compute.

115
Multi-Selecthard

A data scientist is fine-tuning a foundation model on Amazon Bedrock for a custom summarization task. Which THREE practices should they follow to optimize the fine-tuning process?

Select 3 answers
A.Start with a base model that is already strong in the domain.
B.Use the default hyperparameters without tuning.
C.Use a representative dataset that reflects the target task.
D.Monitor training loss and validation loss to avoid overfitting.
E.Train for as many epochs as possible.
AnswersA, C, D

Selecting a domain-strong base model reduces the volume of task-specific examples and training steps required, since the model already encodes relevant vocabulary and structure. This directly satisfies the stem's optimisation goal by lowering compute cost and convergence time during Bedrock fine-tuning, rather than compensating for weak domain representation through extra data.

Why this answer

Option A is correct because selecting a base model already strong in the target domain gives the fine-tuning process a better starting point, reducing the amount of task-specific data and compute needed to reach high summarization quality. Option C is correct because a representative dataset that reflects the target task ensures the model learns the desired summarization style, domain vocabulary, and input-output distribution, which directly improves fine-tuning effectiveness. Option D is correct because monitoring training loss and validation loss lets the data scientist detect overfitting early and apply mitigations such as early stopping, regularization, or more data, keeping generalization strong.

Option B is not correct because leaving hyperparameters at defaults ignores task-specific tuning of values like learning rate, batch size, and epoch count, which often materially affects fine-tuning results. Option E is not correct because training for as many epochs as possible typically causes overfitting, degrading validation performance rather than optimizing the process.

Exam trap

The AIF-C01 exam often tests the misconception that more epochs always improve model performance, when in fact excessive training leads to overfitting, and they expect candidates to recognize that monitoring loss curves and using early stopping are critical practices.

116
MCQeasy

A marketing agency uses a foundation model to generate images for social media campaigns. Some generated images have contained violent or inappropriate content, damaging the brand. The agency needs to prevent such content from being displayed automatically. They are using Amazon Bedrock for image generation with Stable Diffusion. What is the most effective way to filter out inappropriate images?

A.Use Amazon Rekognition to analyze images after generation.
B.Manually review all images before posting.
C.Restrict the prompt to avoid triggering keywords.
D.Enable the safety checker in Amazon Bedrock's image generation models.
AnswerD

Enabling the safety checker activates Bedrock's built-in content filters, which evaluate each generated image against violence and inappropriate-content thresholds before returning it. This directly satisfies the agency's need to block harmful output automatically at generation time, rather than relying on manual review or post-hoc moderation of images already displayed.

Why this answer

Amazon Bedrock's image generation models (including Stable Diffusion) have a built-in safety checker that automatically detects and blocks NSFW or inappropriate content during generation, preventing such images from being output without manual effort. Option A (Amazon Rekognition) adds cost and latency for post-generation analysis, while Option B (manual review) is not scalable. Option C (restricting the prompt) is unreliable as models can still generate inappropriate content from seemingly safe prompts.

117
MCQhard

An e-commerce company is using a foundation model to generate product descriptions. They want to reduce costs by caching frequently requested descriptions. Which AWS service should they use to implement a cache?

A.Amazon CloudFront
B.Amazon DynamoDB
C.Amazon S3
D.Amazon ElastiCache
AnswerD

Amazon ElastiCache provides an in-memory cache, delivering sub-millisecond latency for repeatedly requested product descriptions, which directly satisfies the stem's cost-reduction constraint by offloading duplicate foundation model inferences. Unlike persistent stores, its volatile, RAM-based architecture suits transient cached text, avoiding repeated compute charges.

Why this answer

Amazon ElastiCache is the correct choice because it provides an in-memory caching layer (using Redis or Memcached) that can store frequently requested product descriptions, reducing the need to invoke the foundation model repeatedly. This directly lowers inference costs and latency by serving cached responses instead of generating new ones each time.

Exam trap

The AIF-C01 exam often tests the distinction between caching at the application layer (ElastiCache) versus caching at the content delivery layer (CloudFront), leading candidates to mistakenly choose CloudFront for any caching need.

How to eliminate wrong answers

Option A is wrong because Amazon CloudFront is a content delivery network (CDN) that caches static and dynamic content at edge locations, but it is not designed for application-level caching of model-generated text; it caches HTTP responses, not arbitrary key-value data. Option B is wrong because Amazon DynamoDB is a fully managed NoSQL database optimized for high-throughput, low-latency reads and writes, but it is not a caching service; using it as a cache would incur higher costs and lack native TTL-based eviction policies for transient data. Option C is wrong because Amazon S3 is an object storage service for storing large amounts of unstructured data, not a low-latency cache; retrieving descriptions from S3 would introduce significant latency compared to an in-memory cache, defeating the purpose of cost reduction.

118
MCQmedium

A company uses Amazon Bedrock to generate summarizations of lengthy reports. Users report that the summaries are too verbose and include excessive detail. Which prompt engineering technique should the team apply to address this issue?

A.Reduce the input context length to limit available information.
B.Increase the maxTokens parameter in the inference request.
C.Include few-shot examples of desired outputs.
D.Add explicit constraints like 'Provide a concise summary in two sentences.'
AnswerD

Adding explicit length constraints directly controls the model's output verbosity, which is the stated problem. Unlike vague instructions such as "be brief", specifying "two sentences" gives Amazon Bedrock a concrete target the model can satisfy during generation, tightening the summary without altering the underlying report content or requiring model retraining.

Why this answer

Adding explicit constraints like 'Provide a concise summary in two sentences' directly instructs the model to limit verbosity and detail. This prompt engineering technique uses clear, specific instructions to control output length and style, which is the most effective way to address overly verbose summaries without altering model parameters or input data.

Exam trap

The trap here is that candidates confuse reducing input length (Option A) with controlling output length, or they mistakenly think increasing maxTokens (Option B) can somehow shorten output, when in fact it does the opposite.

How to eliminate wrong answers

Option A is wrong because reducing input context length does not guarantee concise output; the model may still generate verbose summaries from the remaining text, and it risks losing critical information needed for accurate summarization. Option B is wrong because increasing the maxTokens parameter actually allows the model to generate longer outputs, which would exacerbate the verbosity issue rather than solve it. Option C is wrong because few-shot examples can guide output format but are less direct and reliable than explicit constraints; they may not consistently enforce conciseness, especially if the examples themselves are not perfectly aligned with the desired brevity.

119
MCQmedium

A retail company is using Amazon Bedrock to generate personalized product recommendations in real time. The model sometimes produces recommendations that include products the company no longer sells. The company wants to ensure that only in-stock products are recommended. The product catalog is stored in an Amazon DynamoDB table that is updated continuously. Which solution should the company implement to meet this requirement with minimal latency?

A.Fine-tune the foundation model daily with the latest product catalog from DynamoDB.
B.Increase the model's top-p parameter to include more diverse product recommendations.
C.Store the product catalog in an Amazon Bedrock knowledge base and rely on the model to retrieve only in-stock items.
D.Use Amazon Bedrock agents with an action group that queries DynamoDB for in-stock products, and filter the model's output accordingly.
AnswerD

Amazon Bedrock agents can invoke an action group that calls DynamoDB to check current inventory. The agent can fetch in-stock products and either provide them as context to the model or filter the model's recommendations before returning them. This ensures real-time accuracy with minimal latency because DynamoDB provides fast, scalable reads. This is the correct pattern for dynamic inventory validation.

Why this answer

To ensure recommendations only include in-stock products with minimal latency, the company should use Amazon Bedrock agents with an action group that queries DynamoDB. The agent can validate inventory in real time and filter the model's output. DynamoDB offers low-latency reads, making this suitable for real-time recommendations.

Fine-tuning, top-p adjustments, or knowledge bases do not provide the necessary real-time inventory validation.

Exam trap

The trap here is assuming that a knowledge base automatically reflects real-time updates from DynamoDB, when in fact it requires a sync and does not perform live inventory checks.

120
Multi-Selecteasy

Which TWO AWS services can be used together to build a chatbot that leverages a foundation model for natural language understanding?

Select 2 answers
A.Amazon Rekognition
B.Amazon Lex
C.Amazon Polly
D.AWS Glue
E.Amazon Bedrock
AnswersB, E

Amazon Lex supplies the conversational interface, handling intent recognition, slot elicitation and dialogue flow, while Amazon Bedrock provides access to foundation models for natural language understanding. Together they satisfy the stem's requirement that two AWS services combine to build a foundation-model-powered chatbot.

Why this answer

Amazon Lex (B) is correct because it provides the conversational interface—intents, utterances, slots, and dialog management—needed to build a chatbot that interprets user input. Amazon Bedrock (E) is correct because it offers access to foundation models via a managed API, supplying the natural language understanding and generative responses the chatbot requires; Lex can invoke Bedrock through a Lambda fulfillment function to combine both capabilities. Amazon Rekognition (A) is a computer-vision service for image and video analysis, not natural language understanding, so it does not fit.

Amazon Polly (C) only converts text to speech for voice output and does not provide NLU or foundation-model reasoning. AWS Glue (D) is a serverless ETL and data-catalog service for analytics pipelines, unrelated to building conversational chatbots.

Exam trap

AWS often tests the distinction between services that handle conversational interfaces (Lex) versus those that provide generative AI models (Bedrock), tempting candidates to pick Polly (speech output) or Rekognition (vision) as part of a chatbot, when they are not core to NLU or FM integration.

121
MCQhard

A team is fine-tuning a foundation model using SageMaker. They want to minimize training time while keeping the model's original knowledge. Which technique is BEST suited?

A.Use Parameter Efficient Fine-Tuning (PEFT) such as LoRA
B.Use distributed training across multiple GPUs
C.Use prompt engineering instead of fine-tuning
D.Full fine-tuning on the new dataset
AnswerA

PEFT with LoRA freezes the base weights and trains small low-rank adapter matrices, cutting trainable parameters and GPU time substantially. Because original weights remain unchanged, the model's pretrained knowledge is preserved, satisfying both the speed and retention constraints.

Why this answer

Parameter Efficient Fine-Tuning (PEFT) methods like LoRA (Low-Rank Adaptation) are best suited because they freeze the pre-trained model weights and inject trainable low-rank matrices into specific layers, drastically reducing the number of trainable parameters. This minimizes training time and computational cost while preserving the model's original knowledge, as only a small fraction of parameters are updated during fine-tuning.

Exam trap

AWS often tests the distinction between techniques that modify the model (fine-tuning) versus those that only change the input (prompt engineering), and the trap here is that candidates may choose distributed training (Option B) thinking it reduces time, but it does not address parameter efficiency or knowledge preservation as directly as PEFT.

How to eliminate wrong answers

Option B is wrong because distributed training across multiple GPUs accelerates training but does not inherently preserve the model's original knowledge or reduce the number of updated parameters; it still requires full or partial parameter updates and does not address the goal of minimizing training time through parameter efficiency. Option C is wrong because prompt engineering is a zero-shot or few-shot inference technique that does not involve training at all, so it cannot be used to fine-tune the model on a new dataset. Option D is wrong because full fine-tuning updates all model parameters, which is computationally expensive, time-consuming, and risks catastrophic forgetting of the original knowledge, contrary to the goal of minimizing training time while preserving original knowledge.

122
MCQhard

A developer is using Amazon Bedrock with the Cohere Command model to generate summaries of technical documents. They need to control the length of the summaries and ensure they do not exceed a certain number of tokens. Which parameter should they use in the InvokeModel API call?

A.top_p
B.max_tokens
C.stop_sequences
D.temperature
AnswerB

The max_tokens parameter specifies the maximum number of tokens to generate in the response. By setting this parameter, the developer can control the length of the summary and ensure it does not exceed the desired token limit. It is a standard parameter supported by Cohere Command on Amazon Bedrock.

Why this answer

The max_tokens parameter directly controls the maximum number of tokens in the model's response. Setting it ensures the summary does not exceed the desired length. Other parameters like temperature, top_p, and stop_sequences influence creativity, diversity, or stopping conditions, but do not provide a precise token limit.

Exam trap

The trap here is confusing stop_sequences with a length control mechanism; stop_sequences can end output early but do not enforce a specific token count.

123
MCQeasy

A developer is building a prototype that needs to generate product descriptions from a few keywords. The developer wants to use a foundation model on Amazon Bedrock but has no labeled training data and wants to avoid managing infrastructure. Which approach should the developer use?

A.Deploy an open-source model on an Amazon EC2 instance and write custom inference code.
B.Train a custom model from scratch using Amazon SageMaker.
C.Fine-tune a foundation model with a small dataset of product descriptions.
D.Use the foundation model with a well-crafted prompt that includes the keywords and desired format.
AnswerD

Foundation models on Amazon Bedrock can generate product descriptions from keywords using zero-shot prompting. By crafting a clear prompt that specifies the keywords and desired output format, the developer can obtain useful results without any training data or infrastructure management. This is the simplest and most appropriate approach for a prototype with no labeled data.

Why this answer

For a prototype with no labeled data and a desire to avoid infrastructure management, using a foundation model on Amazon Bedrock with a well-designed prompt is ideal. Prompt engineering allows the model to generate product descriptions from keywords without any training. Fine-tuning, training from scratch, or self-managing an EC2 deployment all introduce unnecessary complexity and data requirements that the scenario does not support.

Exam trap

The trap here is thinking that any customization like fine-tuning is required to get useful output, when zero-shot prompting with a foundation model is often sufficient for simple generation tasks.

124
MCQeasy

Refer to the exhibit. A developer runs this command but gets an error: 'An error occurred (AccessDeniedException) when calling the ListFoundationModels operation'. What is the most likely cause?

A.The IAM role does not have bedrock:ListFoundationModels permission
B.The AWS CLI version is outdated
C.The foundation model is not available in us-west-2
D.The region us-west-2 does not support Bedrock
AnswerA

ListFoundationModels is a Bedrock control-plane action, so the caller's identity must hold bedrock:ListFoundationModels in its attached IAM policy. AccessDeniedException is returned by IAM authorisation evaluation, not by model availability or Region configuration, making a missing or explicitly denied permission the direct cause.

Why this answer

The error 'AccessDeniedException' when calling ListFoundationModels indicates that the IAM role or user executing the AWS CLI command lacks the required permission to list foundation models in Amazon Bedrock. The specific permission needed is bedrock:ListFoundationModels, which must be attached to the IAM identity via a policy. Without this permission, the API call is denied regardless of other factors like region or CLI version.

Exam trap

AWS often tests the distinction between service availability errors (e.g., region not supported) and IAM permission errors, where candidates mistakenly attribute an AccessDeniedException to regional or model availability issues rather than missing IAM permissions.

How to eliminate wrong answers

Option B is wrong because an outdated AWS CLI version would typically produce a different error (e.g., 'InvalidClientTokenId' or 'UnrecognizedClientException'), not an AccessDeniedException, and the ListFoundationModels API is available in recent CLI versions. Option C is wrong because the error is an access denial, not a model availability issue; if a model were unavailable, the error would be something like 'ValidationException' or 'ResourceNotFoundException' when trying to use that specific model. Option D is wrong because us-west-2 (Oregon) fully supports Amazon Bedrock and its APIs; the error is explicitly an IAM permissions issue, not a regional unsupported service error.

125
MCQeasy

A company wants to use a pre-trained foundation model for sentiment analysis without any customization. Which Amazon Machine Learning service provides access to foundation models via API?

A.Amazon Bedrock
B.Amazon Textract
C.Amazon Comprehend
D.Amazon Rekognition
AnswerA

Amazon Bedrock supplies serverless API access to pre-trained foundation models from providers such as Anthropic and Amazon, so no model training or hosting is needed. This directly satisfies the stem's constraint of using a foundation model for sentiment analysis without customisation, unlike services requiring you to build or fine-tune your own model.

Why this answer

Amazon Bedrock is a fully managed service that provides access to a wide range of pre-trained foundation models (FMs) from leading AI providers like Anthropic, Meta, and Amazon via a unified API. For sentiment analysis, you can invoke an FM such as Anthropic Claude or Meta Llama directly through the Bedrock API without any customization, making it the correct choice for this use case.

Exam trap

AWS often tests the distinction between a service that provides access to foundation models (Bedrock) and a service that offers a pre-built, non-customizable ML capability (Comprehend), leading candidates to mistakenly choose Comprehend because it also handles sentiment analysis.

How to eliminate wrong answers

Option B (Amazon Textract) is wrong because it is a document analysis service designed to extract text, handwriting, and data from scanned documents, not a service that provides access to foundation models via API. Option C (Amazon Comprehend) is wrong because it is a natural language processing (NLP) service that offers pre-built sentiment analysis, but it does not provide access to foundation models; it uses its own proprietary models and cannot be used to invoke third-party FMs. Option D (Amazon Rekognition) is wrong because it is an image and video analysis service for tasks like object detection and facial recognition, not a service for accessing foundation models or performing text-based sentiment analysis.

126
MCQmedium

A company is building a chatbot using Amazon Bedrock to answer customer questions about their product catalog. The chatbot should only use information from the company's internal knowledge base and should not generate answers based on the model's pre-training data. Which feature should be enabled?

A.Use prompt engineering to instruct the model to only use the knowledge base
B.Configure a knowledge base with Retrieval Augmented Generation (RAG)
C.Enable model invocation logging to review responses
D.Fine-tune the model on the product catalog data
AnswerB

RAG grounds responses in the supplied knowledge base by retrieving relevant documents and injecting them into the prompt context, so the model answers from company data rather than its pre-training weights. This directly satisfies the constraint that answers must come only from the internal catalogue.

Why this answer

Configuring a knowledge base with Retrieval Augmented Generation (RAG) allows the chatbot to retrieve relevant documents from the company's internal knowledge base and use them as context for generating answers. This ensures the model's responses are grounded solely in the provided data, preventing reliance on its pre-training knowledge.

Exam trap

The trap here is that candidates often confuse fine-tuning with RAG, assuming fine-tuning alone can restrict the model to a specific knowledge domain, when in fact fine-tuning does not prevent the model from using its pre-training data and can still produce off-topic responses.

How to eliminate wrong answers

Option A is wrong because prompt engineering alone cannot reliably prevent the model from using its pre-training data; it only provides instructions that the model may still override with its internal knowledge. Option C is wrong because model invocation logging only records responses for auditing and debugging, it does not constrain the model's source of information. Option D is wrong because fine-tuning adapts the model to the product catalog but does not guarantee that the model will ignore its pre-training data; it can still generate answers from its original training corpus.

127
MCQmedium

A financial services company has built an internal assistant on Amazon Bedrock using Anthropic Claude 3 Sonnet. Employees ask questions that require retrieving the latest internal policy documents, which are updated frequently and stored in Amazon S3. The company wants the assistant to answer with accurate, up-to-date citations without retraining the model. Which approach should they implement?

A.Increase the temperature parameter to encourage the model to generate more detailed and specific policy answers.
B.Enable model invocation logging in Amazon Bedrock to capture requests and responses for later review.
C.Fine-tune the Claude 3 Sonnet model on the policy documents using Amazon Bedrock custom models.
D.Use Retrieval Augmented Generation (RAG) by creating a knowledge base in Amazon Bedrock that indexes the S3 documents.
AnswerD

RAG with an Amazon Bedrock knowledge base retrieves relevant document chunks from the S3 data source at query time and passes them to the model as context. This grounds responses in current content and can return citations, all without retraining or fine-tuning the underlying foundation model.

Why this answer

Grounding a foundation model in frequently changing internal documents is best achieved with Retrieval Augmented Generation. An Amazon Bedrock knowledge base ingests the S3 content, chunks and embeds it, and retrieves relevant passages at inference time so the model can answer with current, citable information. This avoids retraining and keeps responses accurate as policies evolve.

Exam trap

The trap here is assuming that fine-tuning is the way to add new knowledge, when RAG is the appropriate pattern for frequently updated, citable source material.

128
MCQhard

An enterprise deploys a foundation model on Amazon Bedrock with a knowledge base. Users report that the model is returning outdated information. What is the most likely cause?

A.The model was fine-tuned
B.The model is not the latest version
C.The knowledge base data source is not refreshed
D.The inference parameters are incorrect
AnswerC

Stale source content propagates directly into retrieval: Amazon Bedrock knowledge bases sync from the configured data source, so unchanged documents keep returning outdated chunks regardless of model capability. Refreshing or re-syncing the data source satisfies the freshness constraint the stem describes.

Why this answer

When a knowledge base is attached to a foundation model on Amazon Bedrock, the model retrieves information from the data source to augment its responses. If the data source is not refreshed, the model will return outdated information even if the model itself is current. Option C directly addresses this by identifying the stale data source as the root cause.

Exam trap

The trap here is that candidates may confuse model versioning (Option B) with data freshness, but the question specifically ties the symptom to the knowledge base, making the refresh cycle the critical factor.

How to eliminate wrong answers

Option A is wrong because fine-tuning adjusts the model's weights on a specific dataset, which does not inherently cause outdated information; in fact, fine-tuning could update the model with newer data. Option B is wrong because using an older model version might affect performance or capabilities, but the question specifically states the model is returning outdated information, which points to the knowledge base content, not the model version. Option D is wrong because inference parameters (e.g., temperature, top_p) control randomness and creativity of responses, not the freshness or accuracy of the information retrieved from the knowledge base.

← PreviousPage 2 of 2 · 128 questions total

Ready to test yourself?

Try a timed practice session using only Foundation Model Applications questions.