Courseiva

CCNA Applications of Foundation Models Questions

75 of 128 questions · Page 1/2 · Applications of Foundation Models · Answers revealed

1
Multi-Selecthard

Which THREE practices are recommended for responsible AI when deploying foundation models? (Choose three.)

Select 3 answers
A.Avoid collecting user feedback to reduce bias
B.Include human review for high-stakes decisions
C.Implement guardrails to filter harmful content
D.Continuously monitor model outputs for drift
E.Use a black box approach to keep model internals secret
AnswersB, C, D

Human review for high-stakes decisions satisfies the accountability constraint by keeping a person responsible for consequential outcomes, rather than delegating judgement to a probabilistic model. Foundation models can produce confident but wrong outputs, so oversight catches harmful errors before they affect people, aligning with responsible AI principles.

Why this answer

Option B is correct because responsible AI deployment requires human-in-the-loop oversight for consequential decisions (e.g., hiring, lending, medical triage), so that a qualified person can validate or override model output before it affects someone's rights or safety. Option C is correct because guardrails — input/output filters, content moderation classifiers, and policy enforcement layers — are a standard control to prevent foundation models from generating harmful, unsafe, or policy-violating content. Option D is correct because foundation models can degrade or shift behavior over time due to data drift, concept drift, or upstream model updates, so continuous monitoring of outputs (with metrics, alerts, and retraining triggers) is essential to maintain reliability and fairness.

Option A is not recommended: collecting user feedback is a valuable signal for identifying bias and improving models, so avoiding it would reduce accountability rather than enhance responsibility. Option E is not recommended: opaque 'black box' secrecy undermines transparency, explainability, and auditability, which are core responsible AI principles; model internals and documentation should be appropriately disclosed to stakeholders.

Exam trap

A common misconception is that avoiding user feedback reduces bias, when in fact it starves the system of data needed to detect and correct bias, making it a harmful anti-pattern. AWS recommends continuous feedback and monitoring for responsible AI.

2
MCQhard

A company uses Amazon Bedrock with a custom model deployed via Amazon SageMaker. They want to monitor for data drift in input prompts over time. Which AWS service is best suited for this?

A.Amazon CloudWatch
B.Amazon SageMaker Model Monitor
C.AWS CloudTrail
D.Amazon Athena
AnswerB

SageMaker Model Monitor captures incoming request data and compares it against a baseline, detecting drift in prompt feature distributions over time. This satisfies the requirement to monitor input prompts for data drift on a custom SageMaker-deployed model.

Why this answer

Amazon SageMaker Model Monitor is the correct choice because it is specifically designed to detect data drift in machine learning models, including input prompts for custom models deployed via SageMaker. It continuously monitors the distribution of input data against a baseline and alerts when drift occurs, which aligns with the requirement to monitor input prompts over time.

Exam trap

The trap here is that candidates often confuse general monitoring services like CloudWatch with specialized ML monitoring tools, assuming CloudWatch can handle data drift detection when it actually lacks the statistical analysis required for such tasks.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch is a monitoring service for AWS resources and applications (e.g., metrics, logs, alarms), but it does not have built-in capabilities to detect data drift in ML model inputs. Option C is wrong because AWS CloudTrail records API activity for auditing and governance, not for monitoring data drift in model inputs. Option D is wrong because Amazon Athena is an interactive query service for analyzing data in S3 using SQL, not a monitoring tool for data drift.

3
MCQhard

A company is building a multi-modal application that processes images and text to answer questions about product defects. Which foundation model approach is BEST?

A.Use an image captioning model and then analyze the caption text
B.Use a text-to-image generation model and analyze the generated image
C.Use a multi-modal foundation model that processes both images and text
D.Use a separate image analysis model and a text model, then combine outputs
AnswerC

A multi-modal foundation model encodes images and text into a shared embedding space, letting one model jointly reason over both modalities. This directly satisfies the stem's requirement to answer defect questions from combined image and text input, unlike text-only or vision-only models that cannot correlate the two.

Why this answer

Multi-modal foundation models (e.g., CLIP, Flamingo, GPT-4V) are specifically designed to jointly process and reason over images and text in a unified architecture. This allows the model to directly correlate visual defects with textual descriptions without intermediate lossy transformations, making it the most effective approach for a multi-modal QA task.

Exam trap

AWS often tests the misconception that combining two separate single-modal models (Option D) is equivalent to a true multi-modal model, but the trap is that late fusion lacks the joint embedding and cross-attention mechanisms needed for coherent multi-modal reasoning.

How to eliminate wrong answers

Option A is wrong because image captioning models convert the entire image into a single text caption, losing fine-grained spatial and defect-specific details that are critical for accurate defect analysis. Option B is wrong because text-to-image generation models create new images from text, which is the inverse of the required task and cannot analyze existing product images for defects. Option D is wrong because using separate models and combining outputs introduces a late-fusion bottleneck, where alignment between visual features and text is not learned end-to-end, leading to poorer performance on tasks requiring joint reasoning.

4
MCQeasy

A solutions architect needs to build a generative AI application that can invoke foundation models from Amazon and third-party providers through a single, unified API without managing any infrastructure. The architect wants the fastest path to a working prototype using AWS-native tooling. Which AWS service should the architect choose?

A.Amazon SageMaker AI
B.Amazon Bedrock
C.Amazon Polly
D.Amazon Comprehend
AnswerB

Amazon Bedrock is the fully managed AWS service that exposes foundation models from Amazon and multiple third-party providers through one unified API, so the architect can prototype quickly without provisioning servers. It handles model hosting and scaling, and supports capabilities such as knowledge bases and guardrails, making it the direct fit for a single-API, serverless generative AI prototype.

Why this answer

Amazon Bedrock is purpose-built for invoking foundation models from Amazon and third-party providers through one API, with no infrastructure to manage. SageMaker AI offers flexibility but adds operational overhead, while Comprehend and Polly serve narrow NLP and speech use cases rather than generative model invocation. The unified, serverless model access makes Bedrock the correct choice for a fast prototype.

Exam trap

The trap here is assuming that any AWS AI service can invoke foundation models, when only Amazon Bedrock provides unified, serverless access to Amazon and third-party foundation models.

5
MCQeasy

A company wants to build a chatbot that responds to customer queries using a foundation model. They need low latency and want to avoid managing infrastructure. Which AWS service should they use?

A.Amazon EC2
B.AWS Lambda
C.Amazon Bedrock
D.Amazon SageMaker
AnswerC

Amazon Bedrock provides serverless access to foundation models through a single API, so the company avoids provisioning or scaling infrastructure while obtaining the low-latency inference the chatbot requires. No model hosting or GPU capacity management falls on the customer.

Why this answer

Amazon Bedrock is a fully managed service that provides access to foundation models (FMs) from leading AI providers via a simple API, eliminating the need to manage underlying infrastructure. It is designed for building generative AI applications like chatbots with low latency, as it handles model hosting, scaling, and inference optimization automatically. This makes it the ideal choice for the company's requirement of low-latency responses without infrastructure management.

Exam trap

AWS often tests the misconception that AWS Lambda can handle any serverless workload, but candidates must recognize that Lambda is unsuitable for large model inference due to its execution time, memory, and GPU limitations, whereas Bedrock is purpose-built for foundation model access.

How to eliminate wrong answers

Option A is wrong because Amazon EC2 requires you to provision, configure, and manage virtual servers, including installing and maintaining the foundation model and its dependencies, which contradicts the requirement to avoid managing infrastructure. Option B is wrong because AWS Lambda is a serverless compute service for running short-duration code (up to 15 minutes) and is not designed to host large foundation models; it lacks the GPU support and memory capacity needed for model inference. Option D is wrong because Amazon SageMaker is a machine learning platform that requires you to manage endpoints, instances, and scaling for model deployment, which still involves infrastructure management and does not provide the fully managed, API-based access to foundation models that Bedrock offers.

6
MCQhard

A company wants to adapt a foundation model for a custom domain with very limited labeled data and minimal cost. Which approach is most suitable?

A.Pre-training from scratch
B.Prompt engineering with few-shot examples
C.Reinforcement learning from human feedback
D.Full fine-tuning
AnswerB

Few-shot prompting supplies a handful of labelled examples within the prompt itself, letting the foundation model adapt to the custom domain without weight updates. This satisfies both constraints: very limited labelled data and minimal cost, since no fine-tuning compute is required.

Why this answer

Prompt engineering with few-shot examples is the most suitable approach because it allows the company to adapt a foundation model to a custom domain using very limited labeled data and minimal cost. By providing a few input-output examples directly in the prompt, the model can infer the desired task without any weight updates, making it efficient for low-resource scenarios.

Exam trap

AWS often tests the misconception that full fine-tuning is always the best way to adapt a model, but the trap here is that candidates overlook the cost and data requirements, failing to recognize that prompt engineering with few-shot examples is the most efficient when labeled data is scarce and budget is tight.

How to eliminate wrong answers

Option A is wrong because pre-training from scratch requires massive amounts of unlabeled data and significant computational resources, which contradicts the requirement for very limited labeled data and minimal cost. Option C is wrong because reinforcement learning from human feedback (RLHF) requires a large dataset of human preferences and multiple model training iterations, making it costly and data-intensive. Option D is wrong because full fine-tuning updates all model weights, which requires a substantial labeled dataset and significant compute, and is not suitable when labeled data is very limited.

7
Multi-Selecthard

A healthcare company is building an application on Amazon Bedrock that uses a foundation model to answer patient questions about medications. The company must reduce the risk of harmful or inaccurate medical advice. Which TWO strategies should they implement? (Choose two.)

Select 2 answers
A.Use Amazon Bedrock Guardrails to filter harmful content and define denied topics related to specific medical advice.
B.Disable all logging so patient interactions are not stored.
C.Fine-tune the model on a small set of patient questions and answers collected from previous chats.
D.Ground responses in approved clinical guidelines by using a knowledge base with Retrieval Augmented Generation.
E.Increase the model's temperature to make answers more diverse and comprehensive.
AnswersA, D

Amazon Bedrock Guardrails can block harmful content and enforce denied topics, which directly addresses the requirement to reduce harmful or inaccurate medical advice. By configuring topic denial and content filters, the application can prevent the model from responding to restricted medical queries. This provides a managed safety layer without changing the underlying model, and it can be applied consistently across invocations.

Why this answer

Guardrails provide a managed safety layer that filters harmful content and blocks denied medical topics, while RAG grounds responses in approved clinical guidelines so answers reflect authoritative sources. Together they reduce both harmful outputs and factual inaccuracies, which is essential for a patient-facing medication assistant.

Exam trap

The trap here is treating model tuning or temperature changes as safety controls, when the primary mitigations are content filtering and grounding in approved sources.

8
MCQhard

Refer to the exhibit. An IAM policy is attached to a user. Which models can the user invoke?

A.Only Claude v2
B.No models
C.Claude v2 and any model with a name containing 'claude'
D.Any model in the account
AnswerA

The policy's Allow statement names only the Claude v2 model resource, so invocation is scoped to that single model. Other foundation models lack an explicit Allow, and IAM denies by default, satisfying the exhibit's constraint that permissions are granted per model ARN rather than account-wide.

Why this answer

The IAM policy explicitly allows the `bedrock:InvokeModel` action only on the resource ARN `arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-v2`. This means the user can invoke only the Claude v2 model. No other models, including other Claude versions or any model with 'claude' in its name, are permitted because the resource ARN is exact and does not use wildcards.

Exam trap

The AIF-C01 exam often tests the distinction between an exact resource ARN and a wildcard pattern; candidates mistakenly assume that 'claude' in the model ID implies all Claude models are allowed, but without a wildcard, only the exact model specified is permitted.

How to eliminate wrong answers

Option B is wrong because the policy does allow invocation of Claude v2, so the user can invoke at least one model. Option C is wrong because the policy uses an exact resource ARN (`anthropic.claude-v2`), not a wildcard pattern like `*claude*`; models with 'claude' in their name but not exactly `claude-v2` are not allowed. Option D is wrong because the policy restricts invocation to a single specific model, not any model in the account.

9
MCQhard

A financial services company is deploying a foundation model on Amazon Bedrock to generate compliance reports from internal audit logs. The model must not output any personally identifiable information (PII). They have configured a Bedrock Guardrail with sensitive information filters set to the 'HIGH' sensitivity level. During testing in a staging environment, testers still observed PII being occasionally generated in the report outputs. The guardrail did not block these instances because the PII was embedded in a context that the guardrail's pattern matching did not catch (e.g., structured JSON data with embedded names). The company requires a solution that minimizes latency and cost, as they process thousands of reports daily. They cannot afford to increase inference time significantly due to strict SLAs. They also want to avoid re-engineering the entire solution. Which additional step should they take to effectively eliminate PII leakage while maintaining performance?

A.Add a prompt instruction to the model to never output PII, with few-shot examples of non-PII outputs.
B.Fine-tune the foundation model on a dataset that excludes PII.
C.Increase the guardrail sensitivity to 'MAXIMUM'.
D.Implement a post-processing Lambda function that uses Amazon Comprehend's PII detection to scan and redact any PII from the model output before returning it.
AnswerD

Amazon Comprehend's PII detection uses trained machine-learning models rather than regex patterns, so it recognises names embedded in structured JSON that Guardrail filters miss. Running it as a post-processing Lambda adds only milliseconds per report, satisfying the latency and cost constraints without re-engineering the Bedrock invocation path.

Why this answer

Amazon Comprehend's PII detection API can be invoked as a post-processing step to scan and redact PII from the model output without requiring any changes to the model or guardrail configuration. This approach adds minimal latency (typically under 100ms per request) and cost per API call is low, making it suitable for high-throughput scenarios. It directly addresses the guardrail's failure to catch PII embedded in structured contexts like JSON, as Comprehend uses machine learning models that can identify PII even when it's not in plain text patterns.

Exam trap

The trap here is that candidates assume increasing guardrail sensitivity or prompt engineering can solve all PII detection failures, but they overlook that guardrails rely on pattern matching and cannot handle contextually embedded PII, whereas a dedicated ML-based detection service like Amazon Comprehend is designed for that exact scenario.

How to eliminate wrong answers

Option A is wrong because prompt instructions and few-shot examples are not reliable for preventing PII leakage; the model may still generate PII due to its training data or contextual reasoning, and this approach adds no deterministic enforcement. Option B is wrong because fine-tuning a foundation model to exclude PII is expensive, time-consuming, and requires a large curated dataset; it also risks degrading model performance on the compliance reporting task and does not guarantee elimination of all PII. Option C is wrong because the guardrail's sensitivity levels (HIGH, MAXIMUM) only affect pattern-matching rules and confidence thresholds; they cannot detect PII embedded in non-standard formats like structured JSON, so increasing sensitivity does not solve the core issue.

10
MCQeasy

A company is using Amazon Bedrock to generate code snippets. Developers report that the generated code sometimes contains security vulnerabilities. Which action should the team take to mitigate this risk?

A.Deploy the model in a sandbox environment to limit its access to sensitive systems.
B.Implement a manual code review process after generation.
C.Add a system prompt that instructs the model to follow security best practices and avoid known vulnerabilities.
D.Reduce the temperature parameter to 0 to make the output deterministic.
AnswerC

A system prompt steers the foundation model at inference time, so instructing it to follow security best practices directly reduces vulnerable code output without retraining or infrastructure changes. This satisfies the scenario's constraint of mitigating risk within the existing Amazon Bedrock setup, though prompt-level guidance offers weaker assurance than automated code scanning.

Why this answer

Adding a system prompt that instructs the model to follow security best practices and avoid known vulnerabilities directly influences the model's output at inference time. Amazon Bedrock supports system prompts that act as high-level instructions to guide the foundation model's behavior, making this a proactive, scalable mitigation that does not require manual intervention or architectural changes.

Exam trap

AWS often tests the misconception that reducing temperature or isolating the environment can fix output quality issues, when in fact only prompt-level guidance directly addresses the model's generation behavior.

How to eliminate wrong answers

Option A is wrong because deploying the model in a sandbox environment limits access to sensitive systems but does not prevent the model from generating code with security vulnerabilities; the model's output itself remains unchanged. Option B is wrong because implementing a manual code review process after generation is a reactive measure that does not reduce the risk at the source; it adds human overhead and delays, and is not a mitigation that addresses the model's tendency to produce insecure code. Option D is wrong because reducing the temperature parameter to 0 makes the output deterministic but does not teach or enforce security best practices; it only reduces randomness, not the likelihood of generating vulnerable patterns.

11
MCQeasy

A company wants to use a foundation model to automatically moderate user-generated content. The model must filter out inappropriate content with high accuracy. Which Amazon service is best suited for this task?

A.Amazon Translate
B.Amazon Rekognition
C.Amazon Polly
D.Amazon Comprehend
AnswerD

Amazon Comprehend provides pre-trained content moderation that detects harmful or inappropriate text categories, returning confidence scores for filtering user-generated content. It delivers the required accuracy without building custom models, unlike general-purpose services such as Bedrock or Rekognition.

Why this answer

Amazon Comprehend is the correct choice because it is a natural language processing (NLP) service that can analyze text for sentiment, key phrases, and — critically — toxicity and inappropriate content using built-in or custom classifiers. This directly matches the requirement to moderate user-generated text with high accuracy, as it can detect hate speech, profanity, and other harmful language.

Exam trap

The trap here is that candidates may confuse Amazon Rekognition's ability to detect unsafe content in images with the need to moderate text, leading them to select Rekognition instead of recognizing that Comprehend is the NLP service for text analysis.

How to eliminate wrong answers

Option A is wrong because Amazon Translate is a machine translation service that converts text between languages; it has no capability to analyze or moderate content for appropriateness. Option B is wrong because Amazon Rekognition is designed for image and video analysis (e.g., object detection, facial recognition, unsafe content detection in visual media), not for moderating text-based user-generated content. Option C is wrong because Amazon Polly is a text-to-speech service that converts written text into lifelike speech; it performs no content moderation or analysis.

12
MCQeasy

A startup needs to build a real-time text translation feature for a customer chat application. Latency must be under 200 ms per request. Which AWS approach is BEST suited?

A.Use Amazon Comprehend for language detection and a custom translation model
B.Use Amazon Bedrock with a multilingual foundation model
C.Use Amazon Translate with real-time translation
D.Use Amazon Transcribe and then Amazon Bedrock
AnswerC

Amazon Translate provides synchronous, real-time translation via its TranslateText API, returning results within milliseconds, which satisfies the sub-200 ms latency constraint for live chat. Purpose-built neural machine translation avoids the overhead of provisioning or invoking general-purpose models, making it the appropriate managed service for low-latency text translation.

Why this answer

Amazon Translate's real-time translation API is purpose-built for low-latency text translation, typically achieving sub-200 ms response times for small payloads. It directly translates text without the overhead of running a large foundation model or performing intermediate steps like transcription, making it the best fit for this latency-sensitive chat application.

Exam trap

The trap here is that candidates may assume a large foundation model (e.g., via Bedrock) is always the best choice for multilingual tasks, overlooking that purpose-built services like Amazon Translate are specifically optimized for low-latency, high-throughput translation at a fraction of the cost and complexity.

How to eliminate wrong answers

Option A is wrong because Amazon Comprehend is designed for natural language processing (e.g., sentiment analysis, entity extraction), not real-time translation, and building a custom translation model would introduce significant latency and complexity. Option B is wrong because Amazon Bedrock with a multilingual foundation model introduces inference latency that often exceeds 200 ms for real-time requests, and it is not optimized for the single-purpose, high-throughput translation task required here. Option D is wrong because Amazon Transcribe is for speech-to-text, not text translation, and chaining it with Bedrock adds unnecessary latency and complexity for a text-only translation feature.

13
MCQhard

A media company uses Amazon Bedrock to generate short product descriptions. They notice that outputs vary in tone and sometimes include unwanted promotional claims. They want consistent, brand-aligned results while keeping the same foundation model. Which action best addresses this requirement?

A.Switch to a larger foundation model with more parameters to improve output quality and consistency.
B.Create a guardrail in Amazon Bedrock with a denied topics policy and a word filter for prohibited promotional terms.
C.Use a prompt template that specifies brand voice, required structure, and forbidden claims, and reuse it for every generation request.
D.Lower the top-p value to 0.1 so the model only considers the most probable tokens.
AnswerC

A well-crafted prompt template embeds the desired tone, format, and explicit exclusions directly into every request, giving the model clear instructions for consistent output. Reusing the template standardizes results across many generations without changing the model. This is the most direct way to enforce brand alignment for text style and claims.

Why this answer

Consistent brand-aligned generation is achieved by controlling the instructions the model receives. A reusable prompt template that defines voice, structure, and forbidden claims gives the model explicit guidance on every request, producing more uniform output. Guardrails, sampling parameters, and larger models can support quality but do not by themselves encode brand style and claim restrictions.

Exam trap

The trap here is treating guardrails or sampling settings as a substitute for clear prompt instructions, when style and claim control require explicit guidance in the prompt.

14
MCQhard

A financial services company uses Amazon Bedrock to power an internal assistant that answers employee questions about HR policies. The company must ensure that the assistant never reveals sensitive employee data. The HR policy documents are stored in an Amazon S3 bucket and are updated frequently. The company wants the assistant to cite the exact policy document and section for each answer. Which solution meets these requirements with the LEAST operational overhead?

A.Build a custom retrieval system using Amazon OpenSearch Service and write a Lambda function to call the model with retrieved passages.
B.Use the model's built-in knowledge by prompting it to answer from its training data, and add a guardrail to filter sensitive information.
C.Fine-tune a foundation model on the HR policy documents and deploy it with a guardrail that blocks sensitive data.
D.Create an Amazon Bedrock knowledge base backed by the S3 bucket, and use the RetrieveAndGenerate API with citations enabled.
AnswerD

Amazon Bedrock knowledge bases can be connected to an S3 bucket, automatically chunk and index documents, and support frequent updates. The RetrieveAndGenerate API can return generated answers with citations that point to the exact source documents and sections. This approach minimizes operational overhead because Bedrock manages the vector store and retrieval pipeline, and it meets both the citation and data freshness requirements.

Why this answer

The company needs accurate, cited answers from frequently updated HR documents with minimal operational overhead. Amazon Bedrock knowledge bases integrate with S3, automatically manage indexing and retrieval, and support citations via the RetrieveAndGenerate API. This managed solution avoids custom infrastructure and ensures answers are grounded in the latest policies.

Fine-tuning, custom retrieval, or relying on model training data all fail to meet the citation and maintenance requirements efficiently.

Exam trap

The trap here is assuming that fine-tuning a model on internal documents will provide citations, when in fact fine-tuning does not produce source references and requires retraining for updates.

15
MCQhard

A generative AI application occasionally produces factually incorrect responses. The team has already tried prompt engineering and increasing the temperature parameter. Which next step is MOST effective to improve factual accuracy?

A.Use a larger foundation model
B.Fine-tune the model on company data
C.Reduce the temperature to 0
D.Implement a Retrieval Augmented Generation (RAG) pipeline
AnswerD

RAG retrieves relevant documents from a knowledge base and injects them into the prompt, grounding generation in verifiable source content. Prompt engineering and temperature tuning cannot supply missing facts, so retrieval directly targets factual accuracy.

Why this answer

Retrieval Augmented Generation (RAG) is the most effective next step because it grounds the model's responses in a verified external knowledge base, directly addressing factual inaccuracies without requiring retraining. Unlike prompt engineering or temperature adjustments, RAG provides real-time access to authoritative documents, reducing hallucinations by constraining the model's output to retrieved evidence.

Exam trap

AWS often tests the misconception that reducing temperature or using a larger model directly fixes factual accuracy, when in fact these methods address output randomness and capacity, not the root cause of hallucination, which is lack of grounded knowledge.

How to eliminate wrong answers

Option A is wrong because using a larger foundation model does not inherently improve factual accuracy; larger models can still hallucinate and may even produce more confident but incorrect answers. Option B is wrong because fine-tuning on company data improves domain-specific performance but does not guarantee factual correctness for dynamic or external facts, and it requires substantial computational resources and labeled data. Option C is wrong because reducing temperature to 0 makes the model deterministic but does not prevent it from generating plausible-sounding but factually incorrect responses; it only reduces randomness, not hallucinations.

16
MCQmedium

A startup is using Amazon Bedrock to power a virtual assistant. They need to ensure that personally identifiable information (PII) is not included in the model's responses. Which feature should they enable?

A.Enable PII redaction in the Bedrock guardrails.
B.Enable model invocation logging.
C.Configure a VPC endpoint.
D.Enable data encryption at rest.
AnswerA

Bedrock guardrails apply PII redaction as a configurable filter that detects and masks sensitive data such as names, addresses and account numbers in both prompts and model responses. This directly satisfies the startup's requirement that personally identifiable information never appears in the virtual assistant's output, without retraining or altering the underlying foundation model.

Why this answer

Amazon Bedrock Guardrails provide a configurable content filtering and PII redaction feature that can automatically detect and mask personally identifiable information (PII) in model inputs and outputs. By enabling PII redaction within guardrails, the startup can ensure that sensitive data like names, addresses, or credit card numbers are removed or obfuscated before the virtual assistant's responses reach the user. This is the direct and intended mechanism for preventing PII leakage in model responses.

Exam trap

The trap here is that candidates often confuse data protection features like encryption or logging with content filtering, not realizing that PII redaction is a specific guardrail policy that actively modifies model outputs in real time.

How to eliminate wrong answers

Option B is wrong because model invocation logging captures metadata and request/response payloads for auditing and debugging, but it does not actively redact or filter PII from responses — it only records what was sent and received. Option C is wrong because configuring a VPC endpoint provides private network connectivity to Bedrock without traversing the public internet, but it has no capability to inspect or modify the content of model responses for PII. Option D is wrong because enabling data encryption at rest protects stored data (e.g., logs, model artifacts) from unauthorized access, but it does not perform real-time redaction of PII in model outputs during inference.

17
MCQeasy

A startup wants to build a mobile app that generates personalized workout plans using a foundation model. They need to minimize infrastructure management and pay only for what they use. Which AWS service should they use to access foundation models via a single API?

A.Amazon Polly
B.Amazon Comprehend
C.Amazon SageMaker
D.Amazon Bedrock
AnswerD

Amazon Bedrock is a fully managed service that provides access to multiple foundation models from different providers through a single API. It eliminates infrastructure management and offers pay-as-you-go pricing, making it ideal for a startup that wants to minimize operational overhead and control costs.

Why this answer

The startup needs a managed service that offers foundation models via a single API with pay-as-you-go pricing. Amazon Bedrock provides exactly that, allowing developers to experiment with and deploy models from multiple providers without managing infrastructure. Other services like SageMaker require more management, while Comprehend and Polly serve different purposes.

Exam trap

The trap here is confusing Amazon SageMaker with Amazon Bedrock; SageMaker is a broader ML platform that requires more infrastructure management and does not offer a unified API for multiple foundation models.

18
MCQmedium

Refer to the exhibit. You receive this response from Amazon Bedrock. What is the most likely cause of the incomplete information?

A.The max_tokens limit was reached
B.The prompt was too short
C.The temperature was too high
D.The model lacks knowledge about capitals
AnswerA

Generation halts once the max_tokens ceiling is hit, so the response ends mid-sentence with content omitted. The truncated, incomplete output in the exhibit matches that hard cutoff, since the model cannot emit further tokens beyond the configured limit.

Why this answer

The response from Amazon Bedrock shows an incomplete sentence that cuts off mid-thought, which is a classic symptom of hitting the max_tokens limit. When the generated output reaches the specified maximum number of tokens, the model stops generating immediately, resulting in truncated text. This is the most likely cause because the output is syntactically incomplete but otherwise coherent up to the cutoff point.

Exam trap

AWS often tests the distinction between output truncation (max_tokens) and output quality issues (temperature, prompt engineering), so the trap here is that candidates may incorrectly attribute a truncated response to model ignorance or randomness rather than the explicit token limit.

How to eliminate wrong answers

Option B is wrong because the prompt length does not directly cause incomplete output; a short prompt can still produce a complete response if the max_tokens limit is high enough. Option C is wrong because temperature controls randomness and creativity, not the length or truncation of the output; high temperature might produce less coherent text but would not cut off mid-sentence. Option D is wrong because the model's lack of knowledge about capitals would result in incorrect or hallucinated information, not a truncated or incomplete sentence.

19
Multi-Selectmedium

A company is deploying a foundation model on Amazon Bedrock to generate product descriptions. They want to ensure the model's output is factually consistent with the provided product specifications and avoids hallucinated features. Which TWO techniques should they use? (Choose two.)

Select 2 answers
A.Set the temperature to a high value to encourage the model to explore more creative descriptions.
B.Fine-tune the model on a dataset of existing product descriptions to improve its writing style.
C.Use Amazon Bedrock Guardrails to define a denied topic for features not present in the specifications.
D.Use retrieval-augmented generation (RAG) with Amazon Bedrock Knowledge Bases to ground the model in the product specifications.
E.Provide clear instructions and product specifications in the prompt, and ask the model to only use the given information.
AnswersD, E

RAG with Amazon Bedrock Knowledge Bases retrieves relevant product specifications and provides them to the model as context. This grounds the generation in factual data, reducing hallucinations and ensuring the output aligns with the actual features. It is a direct method to improve factual consistency.

Why this answer

To ensure factual consistency, the company should ground the model in the actual product specifications. Retrieval-augmented generation with Amazon Bedrock Knowledge Bases provides relevant context, while clear prompt instructions that restrict the model to given information further reduce hallucinations. Both techniques work together to keep output aligned with facts.

Exam trap

The trap here is thinking that Guardrails or fine-tuning can enforce factual consistency, when they address different concerns like content filtering or style adaptation, not grounding in specific data.

20
MCQeasy

A developer wants to compare the output quality of several foundation models available in Amazon Bedrock for a text summarization task. They need to evaluate responses side by side using the same prompt. Which Amazon Bedrock feature should they use?

A.Amazon Bedrock model evaluation
B.Amazon Bedrock provisioned throughput
C.Amazon Bedrock Agents
D.Amazon Bedrock Guardrails
AnswerA

Amazon Bedrock model evaluation allows you to compare foundation models using either automatic metrics or human evaluation. You can submit the same prompt to multiple models and assess summarization quality side by side. This directly supports the developer's goal of comparing output quality across models for a specific task, making it the appropriate feature.

Why this answer

Amazon Bedrock model evaluation provides both automatic and human-based evaluation workflows to compare foundation models on tasks like summarization. It lets you run the same prompt across models and review results side by side. Guardrails, Agents, and provisioned throughput serve different purposes such as content filtering, orchestration, and capacity reservation, so they cannot fulfill the comparison requirement.

Exam trap

The trap here is mixing up model evaluation with guardrails, since both relate to model outputs, but only evaluation is designed to compare and score quality across models.

21
Multi-Selecteasy

Which TWO of the following are benefits of using Amazon Bedrock for building applications with foundation models?

Select 2 answers
A.No infrastructure management
B.Automatic model fine-tuning
C.Access to multiple foundation models
D.Free tier for all models
E.Built-in image generation capability
AnswersA, C

Amazon Bedrock is fully managed and serverless, so teams consume foundation models through a single API without provisioning or scaling any compute, satisfying the stem's benefit of eliminating infrastructure management. This removes capacity planning and patching overhead, letting developers focus on application logic rather than hosting model endpoints.

Why this answer

Option A (No infrastructure management) is correct because Amazon Bedrock is a fully managed, serverless service: AWS handles provisioning, scaling, patching, and hosting of the underlying compute for the foundation models, so developers just call the API without managing any servers or clusters. Option C (Access to multiple foundation models) is correct because Bedrock provides a single unified API to choose from a range of foundation models from providers such as Anthropic, AI21 Labs, Cohere, Meta, Stability AI, and Amazon, letting you switch or compare models without integrating separate SDKs or endpoints. Option B is not a Bedrock benefit as stated, since fine-tuning is an optional capability you must explicitly configure (and not all models support it), not an automatic feature.

Option D is incorrect because Bedrock is not free for all models; usage is billed per input/output token or per image, and pricing varies by model. Option E is incorrect because image generation is only available through specific models (e.g., Stability AI or Amazon Titan Image Generator), not as a universal built-in capability of the service.

Exam trap

AWS often tests the misconception that Amazon Bedrock includes built-in capabilities like automatic fine-tuning or image generation, when in reality these are model-specific features that you must explicitly select and configure, not inherent service features.

22
Multi-Selectmedium

A logistics company is building an internal assistant on Amazon Bedrock that must answer operational questions using its private runbooks and must return citations so staff can verify answers. The team also needs to control cost by limiting how much source text is sent with each request. Which TWO capabilities should they combine to meet these requirements? (Choose two.)

Select 2 answers
A.Knowledge Bases for Amazon Bedrock to ingest the runbooks and retrieve relevant passages.
B.Guardrails for Amazon Bedrock with a profanity filter enabled.
C.Citations returned from retrieved source chunks in the Knowledge Bases response.
D.Provisioned Throughput to guarantee dedicated model capacity.
E.A larger maxTokens value to include more of each runbook.
AnswersA, C

Knowledge Bases for Amazon Bedrock ingests the runbooks from a supported data source, chunks and embeds them, and retrieves only the passages relevant to each question. That retrieval step both grounds answers in private content and keeps prompt size bounded, which directly serves the citation and cost-control requirements.

Why this answer

Retrieval is the mechanism that both grounds answers in private runbooks and bounds prompt size, because only the most relevant chunks are injected. Knowledge Bases for Amazon Bedrock performs that ingestion and retrieval, and it can return citations linking generated text to source chunks, enabling verification. Capacity reservation, output-length increases, and profanity filtering do not supply grounding, traceability, or cost control.

Exam trap

The trap here is assuming that sending more source text or reserving capacity improves answer quality, when targeted retrieval is what controls both accuracy and cost.

23
MCQmedium

A financial services company uses Amazon Bedrock with the Anthropic Claude 3 Haiku model to answer employee questions about internal policies. The knowledge base is updated weekly, and the company wants the model to cite the exact source document and page number in its responses. Which approach should the company use to meet these requirements?

A.Increase the temperature setting in the InvokeModel API call to encourage the model to include source references.
B.Fine-tune the Anthropic Claude 3 Haiku model on the internal policy documents and deploy the custom model.
C.Use Amazon Bedrock Guardrails to filter responses and automatically append document citations.
D.Use Amazon Bedrock Knowledge Bases with a vector store and enable citations in the RetrieveAndGenerate API call.
AnswerD

Amazon Bedrock Knowledge Bases supports retrieval-augmented generation and can return citations that identify the source documents and passages used to generate a response. By enabling citations in the RetrieveAndGenerate API call, the model's answers include references to the exact source, satisfying the requirement for document and page-level attribution.

Why this answer

The company needs answers grounded in specific internal documents with citations. Amazon Bedrock Knowledge Bases performs retrieval-augmented generation by fetching relevant passages from a connected data source and can return citations that point to the source document and page. This directly satisfies the requirement, whereas fine-tuning, temperature adjustments, or Guardrails do not provide source attribution.

Exam trap

The trap here is assuming that fine-tuning or prompt engineering can make a model cite sources, when citation requires a retrieval mechanism such as Amazon Bedrock Knowledge Bases.

24
MCQmedium

A developer is using Amazon Bedrock with the Claude model for text summarization. The output sometimes includes inaccurate information. What is the best practice to reduce hallucinations?

A.Use a larger model
B.Increase temperature
C.Use retrieval augmented generation
D.Decrease max tokens
AnswerC

Retrieval augmented generation grounds the model's response in documents fetched from a knowledge base, so summarisation draws on supplied source text rather than parametric memory alone. This constrains the model to verifiable content, directly reducing fabricated output in the summarisation task.

Why this answer

Retrieval Augmented Generation (RAG) grounds the model's output in external, authoritative knowledge sources by retrieving relevant documents and injecting them into the prompt context. This directly reduces hallucinations because the model generates summaries based on factual retrieved data rather than relying solely on its parametric memory, which is the primary source of inaccuracies in text summarization tasks.

Exam trap

AWS often tests the misconception that model size or output length adjustments are the primary levers for accuracy, when in fact grounding techniques like RAG are the standard solution for reducing hallucinations in production systems.

How to eliminate wrong answers

Option A is wrong because using a larger model (e.g., moving from Claude 3 Haiku to Sonnet) may improve general capabilities but does not inherently reduce hallucinations; larger models can still confidently fabricate information without external grounding. Option B is wrong because increasing temperature introduces more randomness into token selection, which actually increases the likelihood of hallucinated or nonsensical outputs rather than reducing them. Option D is wrong because decreasing max tokens limits the length of the output but does not address the root cause of hallucination—the model's lack of factual grounding—and may even truncate important context, leading to incomplete or misleading summaries.

25
MCQmedium

A developer is building a RAG-based Q&A bot with Amazon Bedrock Knowledge Bases. They need a managed vector store for document embeddings. Which service should they use?

A.Amazon OpenSearch Serverless
B.Amazon DynamoDB
C.Amazon RDS
D.Amazon S3
AnswerA

Amazon OpenSearch Serverless provides a fully managed vector engine that Amazon Bedrock Knowledge Bases can use natively as its vector store, removing server provisioning and cluster scaling work. It satisfies the stem's managed vector store constraint for document embeddings, unlike self-managed alternatives requiring infrastructure upkeep.

Why this answer

Amazon Bedrock Knowledge Bases requires a vector store to store and query document embeddings for Retrieval-Augmented Generation (RAG). Amazon OpenSearch Serverless provides a managed, scalable vector engine that supports k-NN (k-nearest neighbor) search, making it the correct choice for this use case. It integrates natively with Bedrock Knowledge Bases to handle embedding storage and similarity search without manual infrastructure management.

Exam trap

The trap here is that candidates may confuse Amazon DynamoDB or Amazon RDS as viable options because they can store data, but they lack native vector search capabilities required for RAG, leading to an incorrect choice.

How to eliminate wrong answers

Option B (Amazon DynamoDB) is wrong because it is a key-value and document database that does not natively support vector similarity search or k-NN indexing, making it unsuitable as a vector store for RAG. Option C (Amazon RDS) is wrong because it is a relational database service that lacks built-in vector search capabilities; while extensions like pgvector for PostgreSQL exist, Amazon RDS is not a managed vector store and would require custom implementation. Option D (Amazon S3) is wrong because it is an object storage service that cannot perform vector similarity queries; it can store raw documents but not embeddings in a searchable vector index.

26
MCQhard

A healthcare company needs to use a foundation model for analyzing medical records while complying with HIPAA. They plan to use Amazon Bedrock. What should they do to meet HIPAA requirements?

A.Use a model that is HIPAA eligible in a region that supports BAA
B.Implement access logging for all API calls
C.Encrypt data at rest and in transit
D.All of the above
AnswerD

Bedrock's HIPAA eligibility requires a signed AWS Business Associate Addendum, use of HIPAA-eligible models, and no PHI in non-eligible services. Selecting all of the above satisfies the stem's compliance requirement, since each measure is mandatory rather than optional.

Why this answer

HIPAA compliance in Amazon Bedrock requires a combination of controls: using a HIPAA-eligible model in a region where AWS offers a Business Associate Addendum (BAA), enabling access logging for auditability, and encrypting data at rest and in transit. None of the individual options alone satisfy all HIPAA requirements; only the full set of controls ensures compliance.

Exam trap

The trap here is that candidates often pick a single security control (like encryption or logging) thinking it alone ensures HIPAA compliance, but the exam tests that HIPAA requires a combination of administrative, physical, and technical safeguards, all of which must be addressed.

How to eliminate wrong answers

Option A is wrong because while using a HIPAA-eligible model in a BAA-supported region is necessary, it does not address audit logging or encryption requirements. Option B is wrong because access logging alone provides audit trails but does not ensure the model is HIPAA-eligible or that data encryption is enforced. Option C is wrong because encrypting data at rest and in transit is critical but does not cover the need for a BAA or access logging.

All three are required together.

27
MCQhard

A company uses Amazon Bedrock to generate product descriptions. They need to ensure outputs do not contain offensive language. Which service should they integrate to filter content?

A.Amazon Comprehend
B.Amazon Rekognition
C.Bedrock Guardrails
D.AWS WAF
AnswerC

Bedrock Guardrails applies configurable content filters that intercept both prompts and responses, blocking offensive language before it reaches users. This satisfies the requirement to filter outputs, unlike prompt engineering or post-processing, because filtering happens within the Bedrock inference path itself.

Why this answer

Amazon Bedrock Guardrails is the correct choice because it is specifically designed to enforce content policies for foundation model outputs, including filtering for offensive language, hate speech, and other harmful content. It integrates directly with Bedrock to apply customizable safety filters and deny topics without requiring additional services or custom code.

Exam trap

The trap here is that candidates often confuse Amazon Comprehend's text analysis capabilities (like sentiment detection) with real-time content filtering, but Comprehend lacks the policy enforcement and integration with Bedrock that Guardrails provides.

How to eliminate wrong answers

Option A is wrong because Amazon Comprehend is a natural language processing (NLP) service for extracting insights like sentiment, entities, and key phrases from text, but it does not provide real-time content filtering or policy enforcement for Bedrock outputs. Option B is wrong because Amazon Rekognition is an image and video analysis service that detects objects, faces, and text in visual media, not a text-based content filter for offensive language. Option D is wrong because AWS WAF is a web application firewall that protects HTTP/HTTPS endpoints from common web exploits like SQL injection and cross-site scripting, not a content moderation filter for LLM-generated text.

28
MCQeasy

A company is building a customer support chatbot using Amazon Bedrock. They need to store conversation history for context across sessions. Which AWS service is best suited for this purpose?

A.Amazon S3
B.Amazon DynamoDB
C.Amazon RDS
D.Amazon ElastiCache
AnswerB

Amazon DynamoDB stores conversation history as session-keyed items, giving the chatbot low-latency retrieval of prior turns across sessions. Its schema-flexible, key-value design satisfies the persistence requirement without a relational model, and it scales with request volume, unlike stateless compute or storage-only services.

Why this answer

Amazon DynamoDB is the best choice for storing conversation history because it is a fully managed NoSQL key-value and document database that provides single-digit millisecond latency at any scale. It supports flexible schema, which is ideal for storing variable-length chat sessions, and its Time to Live (TTL) feature can automatically expire old conversations to manage storage costs. DynamoDB also integrates natively with AWS Lambda and Amazon Bedrock for real-time retrieval and update of context across sessions.

Exam trap

The trap here is that candidates often confuse durability with performance, picking Amazon S3 for its low cost or Amazon ElastiCache for its speed, without recognizing that DynamoDB uniquely combines low latency, persistence, and flexible schema for session state management.

How to eliminate wrong answers

Option A is wrong because Amazon S3 is an object storage service designed for large, unstructured data like files and backups, not for low-latency, frequent read/write operations required for real-time conversation history retrieval. Option C is wrong because Amazon RDS is a relational database that requires a fixed schema and is overkill for simple key-value session storage; it also incurs higher operational overhead and latency compared to DynamoDB for this use case. Option D is wrong because Amazon ElastiCache is an in-memory caching service (Redis/Memcached) that is volatile and not designed for durable, persistent storage of conversation history across sessions, making it unsuitable for long-term context retention.

29
MCQhard

A law firm uses a foundation model to draft legal briefs. To ensure accuracy, they want to ground the model's outputs in authoritative legal sources. They have a large database of prior case law and statutes stored in Amazon S3. The firm's IT team must implement a solution that reduces hallucinations while being cost-effective. The solution should allow the model to retrieve relevant documents and generate responses based on them. Which approach should they take?

A.Fine-tune the model on the legal database.
B.Manually attach relevant documents to each prompt.
C.Use a larger foundation model with more parameters.
D.Use Amazon Bedrock Agents to create a RAG application.
AnswerD

Amazon Bedrock Agents orchestrate retrieval-augmented generation: documents from Amazon S3 are vectorised into a knowledge base, and relevant passages are retrieved and injected into the prompt, grounding outputs in authoritative case law and statutes while reducing hallucinations cost-effectively.

Why this answer

Amazon Bedrock Agents with a knowledge base can implement Retrieval-Augmented Generation (RAG): the agent retrieves relevant documents from S3 and uses them as context for the model, grounding responses and reducing hallucinations. Option A (fine-tuning) is expensive and does not guarantee grounding for all queries. Option B (manually attaching documents) is not scalable.

Option C (using a larger model) increases cost without solving hallucination.

30
MCQhard

A healthcare company is deploying a conversational AI using a foundation model on Amazon Bedrock for patient triage. The application must minimize hallucinations and ensure factual accuracy. Which combination of techniques should the team implement?

A.Implement Retrieval-Augmented Generation (RAG) using a knowledge base on Amazon Bedrock and a system prompt demanding factual responses.
B.Fine-tune the model on a large dataset of medical transcripts and deploy with default parameters.
C.Use reinforcement learning from human feedback (RLHF) on the deployed model.
D.Set the maxTokens to a low value to force shorter, more focused answers.
AnswerA

RAG grounds responses in a curated knowledge base, so the model cites retrieved clinical content rather than relying on parametric memory, directly minimising hallucination. The system prompt reinforces factual, non-speculative answers. Together they satisfy the accuracy constraint for patient triage.

Why this answer

Retrieval-Augmented Generation (RAG) with a knowledge base on Amazon Bedrock grounds the model's responses in authoritative, up-to-date documents, dramatically reducing hallucinations. Pairing RAG with a system prompt that demands factual, sourced answers further constrains the model. This combination directly addresses the requirement to minimize hallucinations and ensure factual accuracy for patient triage.

Exam trap

AIF-C01 often tests the misconception that fine-tuning or RLHF eliminates hallucinations, when RAG with grounded retrieval is the primary technique for factual accuracy.

How to eliminate wrong answers

Option B is wrong because fine-tuning on medical transcripts improves domain style but does not prevent hallucinations and can even amplify incorrect patterns; default parameters offer no grounding. Option C is wrong because RLHF is a training-time alignment technique not available as a runtime control on Bedrock and does not provide factual grounding. Option D is wrong because limiting maxTokens only shortens output; it does not improve factual accuracy and can truncate critical information.

31
MCQmedium

A developer uses Amazon Bedrock to generate code. Some outputs contain syntax errors. What is the most likely cause?

A.The prompt lacks constraints or examples
B.The max_tokens is too low
C.The temperature is too high
D.The model lacks knowledge of the language
AnswerA

Foundation models infer intent from prompt context, so an unconstrained prompt lets the model choose arbitrary syntax and libraries. Adding explicit language, style and formatting constraints plus few-shot code examples narrows the output distribution, which is the direct mechanism preventing the syntax errors described.

Why this answer

Syntax errors in generated code typically arise when the prompt lacks sufficient constraints or examples to guide the model toward producing syntactically valid output. Amazon Bedrock's foundation models rely on clear instructions and few-shot examples to adhere to language syntax rules; without them, the model may generate plausible-looking but incorrect code.

Exam trap

The AWS exam often tests the misconception that syntax errors are due to model limitations (e.g., lack of knowledge or parameter settings) rather than the more common cause of insufficient prompt engineering, such as missing constraints or examples.

How to eliminate wrong answers

Option B is wrong because max_tokens controls the length of the output, not the syntactic correctness; a low max_tokens might truncate code but does not cause syntax errors. Option C is wrong because temperature affects randomness and creativity, not syntax; a high temperature might produce more varied but still syntactically valid code. Option D is wrong because Bedrock's foundation models are trained on vast corpora including many programming languages and have sufficient knowledge of language syntax; the issue is prompt design, not model capability.

32
MCQeasy

Which pricing model does Amazon Bedrock use for foundation model inference?

A.Per-request
B.Per-hour instance
C.Per-GB storage
D.Per-token
AnswerD

Amazon Bedrock charges for foundation model inference by the number of input and output tokens processed, satisfying the stem's demand for its actual pricing model. This consumption-based approach bills each API call according to tokens consumed, unlike provisioned throughput's hourly commitment, making per-token the accurate answer.

Why this answer

Amazon Bedrock charges for foundation model inference based on the number of tokens processed, which includes both input and output tokens. Each model has a specific per-token price, and the total cost is calculated by multiplying the token count by the model's rate. This is the standard pricing model for generative AI inference services like Bedrock.

Exam trap

AWS often tests the distinction between on-demand inference (per-token) and provisioned throughput (per-hour instance), so candidates mistakenly select per-request thinking it covers all inference, but Bedrock's on-demand pricing is explicitly token-based.

How to eliminate wrong answers

Option A is wrong because Amazon Bedrock does not use a per-request pricing model; instead, it charges per token, which accounts for the variable length of each request. Option B is wrong because per-hour instance pricing applies to provisioned throughput or dedicated instances, not to on-demand inference, which is token-based. Option C is wrong because per-GB storage pricing is used for data storage services like Amazon S3 or EBS, not for inference compute in Bedrock.

33
MCQhard

A security engineer creates the above IAM policy to allow a user to invoke an Amazon Bedrock model. However, invocation fails. What is the issue?

A.The action should be "bedrock:InvokeModelWithResponseStream".
B.The resource ARN is missing the account ID.
C.The ARN should use "foundation-model" instead of "model".
D.The statement is missing a condition for the model ID.
AnswerC

Bedrock foundation models are addressed with the resource ARN segment "foundation-model", not "model". The policy's incorrect ARN means the Allow statement matches no real resource, so the invocation is implicitly denied despite otherwise valid permissions.

Why this answer

The IAM policy's resource ARN incorrectly uses 'model' in the path, but Amazon Bedrock requires 'foundation-model' to reference foundation models. The correct ARN format for invoking a Bedrock foundation model is 'arn:aws:bedrock:region::foundation-model/model-id'. Using 'model' instead of 'foundation-model' causes the policy to not match any valid Bedrock resource, resulting in an invocation failure.

Exam trap

AWS often tests the distinction between 'model' and 'foundation-model' in Bedrock ARNs, as candidates may assume all Bedrock models use the same resource type, overlooking that foundation models require a specific path.

How to eliminate wrong answers

Option A is wrong because 'bedrock:InvokeModelWithResponseStream' is a separate action for streaming responses, but the standard 'bedrock:InvokeModel' action is sufficient for non-streaming invocation; the failure is not due to the action name. Option B is wrong because the resource ARN for Bedrock foundation models does not require an account ID; the ARN format uses a double colon (::) in the account ID position, which is correct for service-owned resources. Option D is wrong because a condition for the model ID is optional and not required for invocation; the primary issue is the incorrect resource type in the ARN.

34
MCQhard

A developer is integrating an Amazon Bedrock foundation model into an application that must support multi-turn conversations, maintain chat history, and switch between different provider models with minimal code changes. The application should use a consistent request and response format. Which Amazon Bedrock API should the developer use?

A.The InvokeModel API
B.The Converse API
C.The ListFoundationModels API
D.The CreateModelCustomizationJob API
AnswerB

The Converse API provides a unified, model-agnostic interface for multi-turn conversations, accepting a structured messages array and returning a consistent response shape across supported models. It supports system prompts and tool use, and it reduces provider-specific code when switching models, directly satisfying the requirement for consistent formats and minimal changes.

Why this answer

The Converse API offers a unified conversational interface with a consistent message structure and response format across supported foundation models, simplifying multi-turn applications and reducing code changes when switching providers. InvokeModel requires provider-specific payloads, while model listing and customization job APIs serve discovery and training rather than conversational inference.

Exam trap

The trap here is choosing InvokeModel for portability, when its provider-specific payloads and responses require custom code for each model rather than a unified conversation format.

35
MCQeasy

A developer is using Amazon Bedrock to generate summaries of news articles. They notice that the model sometimes includes information not present in the original article. Which term describes this phenomenon?

A.Data leakage
B.Hallucination
C.Underfitting
D.Overfitting
AnswerB

Hallucination refers to the phenomenon where a foundation model generates content that is not grounded in the input data or factual reality. In this scenario, the model adds details not present in the original article, which is a classic example of hallucination. This is a common challenge when using generative AI and requires mitigation strategies such as grounding or retrieval augmentation.

Why this answer

The phenomenon where a foundation model generates information not present in the source material is called hallucination. It is a key challenge in generative AI, especially for summarization tasks. Mitigation includes grounding the model with retrieval-augmented generation or using guardrails to filter unsupported claims.

Exam trap

The trap here is confusing hallucination with overfitting or data leakage, which are unrelated to the model adding unsupported content during inference.

36
Multi-Selecteasy

Which TWO techniques can reduce the cost of running a fine-tuned foundation model on Amazon SageMaker? (Choose TWO.)

Select 2 answers
A.Implement structured pruning to remove less important model parameters.
B.Use larger instance types with more GPUs to speed up inference.
C.Apply model quantization to reduce precision from FP32 to FP16 or INT8.
D.Store the model parameters in FP32 to maintain accuracy during inference.
E.Increase the number of training epochs to achieve higher accuracy.
AnswersA, C

Structured pruning removes entire neurons, channels or attention heads, shrinking parameter count and memory footprint. The smaller model needs fewer SageMaker instance hours and less GPU memory, directly lowering inference hosting cost while retaining most accuracy.

Why this answer

Option A is correct because structured pruning removes less important weights, neurons, or channels from the fine-tuned model, producing a smaller model that requires fewer compute and memory resources during SageMaker inference, which directly lowers hosting cost. Option C is correct because quantization reduces numerical precision from FP32 to FP16 or INT8, shrinking model size and memory bandwidth needs and enabling faster, cheaper inference on SageMaker endpoints, especially with GPU instances that support lower-precision arithmetic. Option B is not correct because using larger GPU instances increases the hourly cost of the endpoint rather than reducing it.

Option D is not correct because keeping parameters in FP32 preserves accuracy but consumes more memory and compute, raising cost. Option E is not correct because increasing training epochs affects training time and accuracy, not the cost of running inference on the deployed model.

Exam trap

AWS often tests the distinction between techniques that reduce inference cost (pruning, quantization) versus those that improve training speed or accuracy, leading candidates to mistakenly select options that increase resource usage or are irrelevant to inference cost.

37
MCQeasy

A startup company is developing an e-commerce platform and wants to use Amazon Bedrock to generate product descriptions automatically. They have a small team of developers who are not machine learning experts. The product catalog is stored in a DynamoDB table, and each product has attributes like name, category, price, and a brief description. The company wants the generated descriptions to reflect the unique brand voice, which is documented in a few internal style guides stored as PDF files in Amazon S3. They need a solution that allows them to quickly test the approach without significant infrastructure changes or model training. The development team is familiar with AWS SDKs and want to minimize ongoing maintenance. The team has already set up a Bedrock foundation model (Claude) and can make API calls. They tested simple prompts but the output lacked the brand's informal yet professional tone. They want to incorporate examples from the style guides directly into the prompt without retraining. The team fears that including the entire style guide in each prompt would exceed token limits and increase costs. Which approach should they take to effectively incorporate the brand voice with minimal changes?

A.Fine-tune the foundation model using the style guides with Amazon Bedrock Custom Models.
B.Use Amazon Bedrock with a custom prompt template that includes a few representative examples from the style guides as few-shot examples in the system prompt.
C.Concatenate all style guide PDFs into a single text and include it in every prompt.
D.Use Amazon Comprehend to analyze the style guides and extract a list of keywords to include in the prompt.
AnswerB

Few-shot examples embedded in the system prompt steer Claude's tone without weight updates, satisfying the no-training constraint. Selecting only representative excerpts keeps token usage within limits, avoiding the cost and context-window problems of pasting whole style guides, and requires no infrastructure change beyond prompt edits.

Why this answer

Amazon Bedrock supports few-shot prompting, where you include a few representative examples in the prompt to guide the model's style without retraining. By extracting a few examples from the style guides and incorporating them into the system prompt, the model can learn the desired tone and apply it to new product descriptions. This approach requires minimal changes, no model training, and keeps token usage manageable by not including the entire style guide.

Exam trap

AIF-C01 often tests the misconception that fine-tuning is always necessary to adapt a model to a specific style, when in fact few-shot prompting can achieve similar results with less effort and cost.

How to eliminate wrong answers

Option A is wrong because fine-tuning requires significant data preparation, training time, and cost, and the team wants to avoid model training. Option C is wrong because concatenating all style guide PDFs into every prompt would likely exceed token limits and increase costs, as the team fears. Option D is wrong because Amazon Comprehend extracts keywords and topics, not writing style or tone, so it would not effectively capture the brand voice.

38
MCQmedium

A company wants its Amazon Bedrock application to always answer in a formal tone, never discuss competitors, and never reveal internal project codenames. The controls must apply consistently to every request and response without changing the underlying model. Which Amazon Bedrock capability should the company configure?

A.A higher temperature setting
B.Amazon Bedrock Guardrails
C.Provisioned Throughput
D.A larger context window
AnswerB

Guardrails let you define denied topics, content filters, word filters, and sensitive information filters that are evaluated on both prompts and responses, independent of the foundation model. A denied topic can block competitor discussion, a word filter can block codenames, and contextual grounding checks can enforce answer fidelity. This provides consistent, model-agnostic policy enforcement exactly as required.

Why this answer

Amazon Bedrock Guardrails enforce content policies such as denied topics, word filters, and sensitive information detection on inputs and outputs independently of the model, so the same rules apply across every request. Sampling parameters, throughput reservations, and context window size influence generation behavior or capacity but cannot block topics or enforce tone consistently.

Exam trap

The trap here is assuming that prompt engineering or sampling settings can enforce hard policy controls, when only Guardrails apply model-independent filters to requests and responses.

39
Multi-Selecthard

A marketing team is using a foundation model to generate marketing copy. Which THREE of the following should they consider to ensure responsible and cost-effective use?

Select 3 answers
A.Bias mitigation to avoid unfair stereotypes
B.Cost per token for the model
C.Model size (number of parameters)
D.Toxicity detection in generated content
E.Latency of model inference
AnswersA, B, D

Bias mitigation directly addresses fairness, a core responsible-AI dimension for generative marketing copy that could otherwise propagate stereotypes. It satisfies the stem's responsible-use constraint by reducing discriminatory outputs, complementing cost controls. Unlike accuracy or latency tuning, bias mitigation targets the ethical risk inherent in open-ended text generation.

Why this answer

Option A (Bias mitigation to avoid unfair stereotypes) is correct because foundation models can reproduce or amplify biases present in their training data, and marketing copy that relies on unfair stereotypes creates reputational, legal, and ethical risk, so teams should apply bias detection and mitigation techniques. Option B (Cost per token for the model) is correct because foundation models are typically billed by input and output tokens, so tracking cost per token directly supports cost-effective use and lets the team choose the most economical model or prompt design for the workload. Option D (Toxicity detection in generated content) is correct because generative models can produce offensive, harmful, or brand-damaging text, so toxicity detection and filtering are needed to keep marketing content responsible and safe to publish.

Option C (Model size, number of parameters) is not one of the required answers here because parameter count is only an indirect proxy for capability and cost, and by itself it does not ensure responsible or cost-effective use. Option E (Latency of model inference) is not one of the required answers because inference latency affects user experience and responsiveness, not the responsible-use or token-cost concerns the scenario asks about.

Exam trap

AWS often tests the misconception that model size (parameters) is a key cost driver, but in practice, cost is tied to token consumption and inference infrastructure, not just parameter count, and latency is a performance metric, not a cost or responsibility factor.

40
MCQmedium

A media company runs a daily news digest built on Amazon Bedrock with Anthropic Claude. Editors complain that summaries of long policy documents sometimes omit the final recommendations, even though the source text clearly contains them near the end. The requests currently pass only the document body and set a maximum output length of 300 tokens. Which change best addresses the truncation of the source content before the model reasons over it?

A.Chunk the document and apply retrieval or summarization so all sections reach the model.
B.Raise the temperature setting so the model explores more of the document.
C.Increase the maxTokens parameter so the digest can be longer.
D.Switch the model invocation to streaming responses.
AnswerA

When the source exceeds the model's context window, the trailing content is silently truncated, so recommendations at the end never reach inference. Splitting the document into chunks and either summarizing hierarchically or retrieving the most relevant passages ensures every section is represented, restoring the omitted recommendations without exceeding context limits.

Why this answer

The symptom points to input truncation: content beyond the context window is discarded before the model sees it, so the closing recommendations vanish. Increasing output length or changing sampling cannot recover text that never entered the prompt. Chunking the document, then summarizing or retrieving relevant chunks, ensures the full source is represented and the recommendations reach the model.

Exam trap

The trap here is assuming that a larger output token limit also expands how much source text the model can read.

41
MCQmedium

A company is using a foundation model on Amazon Bedrock to generate customer support responses. They notice that the model sometimes produces harmful or offensive content. Which approach is MOST effective to mitigate this issue?

A.Use prompt engineering to instruct the model to avoid harmful content
B.Enable model invocation logging to review and block responses
C.Fine-tune the model on a curated dataset of safe responses
D.Configure Amazon Bedrock Guardrails with content filters
AnswerD

Guardrails apply configurable content filters that evaluate both prompts and model responses, blocking harmful categories before they reach the user. This enforces safety at the Bedrock layer without retraining the foundation model, directly addressing the offensive-output constraint.

Why this answer

Amazon Bedrock Guardrails provides configurable content filters that can block harmful, offensive, or inappropriate content in both user inputs and model outputs. This is the most effective approach because it operates at the inference layer, applying safety policies consistently across all requests without requiring model retraining or manual review. Prompt engineering alone is unreliable, and fine-tuning may not generalize to all harmful content patterns.

Exam trap

The AIF-C01 exam often tests the misconception that prompt engineering or fine-tuning alone is sufficient for safety, when in fact a dedicated guardrail mechanism is required for reliable, policy-based content filtering at inference time.

How to eliminate wrong answers

Option A is wrong because prompt engineering can be easily bypassed by adversarial inputs or model drift, and it does not provide deterministic enforcement of safety policies. Option B is wrong because model invocation logging only records responses for auditing; it does not block harmful content in real time. Option C is wrong because fine-tuning on a curated dataset of safe responses reduces but does not eliminate the risk of generating harmful content, especially for edge cases or novel inputs not seen during training.

42
Multi-Selectmedium

A hospital is deploying an Amazon Bedrock-powered assistant that answers staff questions about internal policies. The compliance team requires that the assistant refuse requests for individual patient diagnoses and that any response containing protected health information be masked before it is returned. Which TWO capabilities should the team configure in Amazon Bedrock Guardrails to meet these requirements? (Choose two.)

Select 2 answers
A.Enable model invocation logging to capture all prompts and responses for later compliance review.
B.Create a Provisioned Throughput purchase for the model to guarantee response capacity for clinical staff.
C.Configure a sensitive information policy that detects and masks protected health information in model responses.
D.Set the model temperature to zero so responses are deterministic and cannot include patient details.
E.Define a denied topics policy that blocks requests and responses related to individual patient diagnosis.
AnswersC, E

Sensitive information filters in Guardrails use pattern matching and identifiers to detect categories such as personally identifiable information and can redact or mask matches. Configuring this for protected health information ensures detected data is masked before the response reaches the user, satisfying the masking requirement.

Why this answer

Guardrails enforce content policy on both inputs and outputs. A denied topics policy blocks requests and responses about individual patient diagnosis, while a sensitive information policy detects and masks protected health information in responses. Together they satisfy the refusal and masking requirements.

Logging, temperature settings, and Provisioned Throughput do not prevent or redact content at inference time.

Exam trap

The trap here is treating invocation logging as a compliance control, when logging only records interactions and performs no real-time blocking or masking.

43
Multi-Selecteasy

A company uses Amazon Bedrock to build a question-answering system. Which THREE features of Amazon Bedrock can improve answer accuracy? (Choose three.)

Select 3 answers
A.Retrieval Augmented Generation (RAG)
B.Auto-scaling of provisioned throughput
C.Model fine-tuning
D.Encryption at rest
E.Prompt engineering
AnswersA, C, E

Retrieval Augmented Generation grounds responses in your own data by retrieving relevant passages and injecting them into the prompt, so answers reflect actual source content rather than parametric guesses. This directly satisfies the accuracy requirement by reducing hallucination and stale knowledge, which a question-answering system over proprietary documents demands.

Why this answer

Retrieval Augmented Generation (RAG) (A) is correct because it grounds the model's responses in an external knowledge base retrieved from sources like Amazon OpenSearch Serverless or Aurora, injecting relevant, up-to-date context into the prompt so answers are factually accurate rather than hallucinated. Model fine-tuning (C) is correct because adapting a foundation model on domain-specific labeled data adjusts its weights to better match the company's terminology, style, and task patterns, directly raising answer quality for that use case. Prompt engineering (E) is correct because carefully designing instructions, few-shot examples, and output formatting in the prompt steers the model toward more precise, relevant, and consistent answers without retraining.

Auto-scaling of provisioned throughput (B) only affects performance and cost under load, not answer correctness, and encryption at rest (D) is a security control protecting stored data, neither of which improves the accuracy of generated answers.

Exam trap

AWS often tests the distinction between features that improve accuracy (RAG, fine-tuning, prompt engineering) versus features that improve operational aspects like scalability (auto-scaling) or security (encryption), leading candidates to mistakenly select non-accuracy-related options.

44
MCQhard

A marketing firm uses Amazon Bedrock to generate ad copy. They notice that the generated text often includes factual inaccuracies about their products. Which technique would most effectively reduce these inaccuracies?

A.Implement Retrieval-Augmented Generation (RAG) with a product knowledge base.
B.Use longer, more detailed prompts.
C.Increase the temperature parameter to 0.9.
D.Fine-tune the model on a dataset of previous ad copies.
AnswerA

Retrieval-Augmented Generation grounds each response in retrieved product facts, so the model conditions on authoritative content rather than relying solely on parametric memory. This directly targets the factual inaccuracies described, since the product knowledge base supplies verified details at inference time, satisfying the accuracy constraint in the stem.

Why this answer

Retrieval-Augmented Generation (RAG) grounds the model's output in a trusted, external knowledge base by retrieving relevant product documents before generating text. This directly addresses factual inaccuracies because the model references authoritative data rather than relying solely on its parametric memory, which may contain outdated or incorrect information.

Exam trap

The AIF-C01 exam often tests the misconception that fine-tuning or prompt engineering alone can fix factual accuracy issues, when in reality RAG is the standard solution for grounding model outputs in external, verifiable data.

How to eliminate wrong answers

Option B is wrong because longer prompts do not fix the underlying knowledge gap; they only provide more context but cannot inject new, accurate facts that the model lacks. Option C is wrong because increasing temperature to 0.9 increases randomness and creativity, which would likely worsen factual inaccuracies by encouraging more hallucinated or divergent outputs. Option D is wrong because fine-tuning on previous ad copies would reinforce existing patterns and biases, including any inaccuracies present in the training data, rather than introducing a reliable source of truth.

45
MCQhard

A media company is using Amazon Bedrock to generate marketing copy with a foundation model. They want to ensure the output adheres to brand voice guidelines (e.g., friendly, professional). Which prompt engineering strategy is most effective for this requirement?

A.Provide five example outputs in the prompt that match the desired tone.
B.Include instructions like 'Do not use technical jargon' in every user prompt.
C.Set the temperature parameter to a low value (e.g., 0.1) to reduce randomness.
D.Use a system prompt that explicitly describes the brand voice and expectations.
AnswerD

A system prompt sets persistent behavioural instructions that condition every response, so describing the brand voice there enforces friendly, professional tone across all generated copy without repeating guidance per request. Unlike few-shot examples, which shape format through demonstration, this directly constrains style, satisfying the adherence requirement.

Why this answer

Amazon Bedrock supports system prompts that set overarching context and behavioral guidelines for the model. By explicitly describing the brand voice (e.g., 'friendly, professional') in the system prompt, the model consistently applies these constraints across all user interactions, which is more effective than per-instruction tuning.

Exam trap

AWS often tests the misconception that parameter tuning (like temperature) or few-shot examples are sufficient for style control, when in fact system prompts provide the most direct and scalable mechanism for enforcing behavioral constraints in foundation models.

How to eliminate wrong answers

Option A is wrong because providing example outputs (few-shot prompting) can guide tone but is less reliable than a system prompt for consistent adherence across diverse inputs, and it consumes prompt token budget without guaranteeing the model internalizes the rule. Option B is wrong because including instructions like 'Do not use technical jargon' in every user prompt is redundant, inefficient, and can be overridden by the model's tendency to follow the most recent instruction, whereas a system prompt sets a persistent baseline. Option C is wrong because lowering the temperature parameter reduces randomness but does not enforce specific brand voice constraints; it only makes outputs more deterministic, which may still produce off-tone content if the model's training data lacks the desired style.

46
MCQmedium

A healthcare company uses Amazon Bedrock to generate patient summaries. They need to ensure no protected health information (PHI) is leaked in the output. Which AWS service can they use to detect and mask PHI in text?

A.Amazon Comprehend Medical
B.Amazon Macie
C.AWS Glue
D.Amazon Rekognition
AnswerA

Amazon Comprehend Medical applies natural language processing purpose-built for clinical text, detecting protected health information such as names, dates and medical record numbers, then masking it. This directly satisfies the requirement to prevent PHI leakage in generated patient summaries.

Why this answer

Amazon Comprehend Medical is specifically designed to extract and identify protected health information (PHI) from unstructured medical text using natural language processing (NLP). It can detect entities such as patient names, dates, medical conditions, and medications, and provides APIs to mask or redact that PHI before output. This makes it the correct choice for the healthcare company's requirement to prevent PHI leakage in patient summaries generated by Amazon Bedrock.

Exam trap

AWS often tests the distinction between general-purpose data protection services (like Macie) and domain-specific medical NLP services (like Comprehend Medical), leading candidates to choose Macie because it is associated with sensitive data discovery, even though it cannot perform inline text masking.

How to eliminate wrong answers

Option B (Amazon Macie) is wrong because Macie is a data security service that discovers and protects sensitive data stored in Amazon S3 using machine learning and pattern matching, but it does not provide real-time PHI detection or masking in text streams or API outputs. Option C (AWS Glue) is wrong because Glue is a serverless data integration service for ETL (extract, transform, load) jobs, not a text analysis or PHI detection service. Option D (Amazon Rekognition) is wrong because Rekognition is an image and video analysis service that can detect objects, faces, and text in media, but it is not designed to identify or mask PHI in textual data.

47
MCQeasy

A startup needs to generate product descriptions from bullet points using a foundation model. They want a fully managed serverless experience. Which AWS service should they use?

A.Amazon Comprehend
B.Amazon Bedrock
C.Amazon Polly
D.Amazon Lex
AnswerB

Amazon Bedrock provides serverless access to foundation models through a managed API, requiring no infrastructure provisioning or capacity management. This satisfies the startup's need to generate product descriptions from bullet points with a fully managed, pay-per-use experience.

Why this answer

Amazon Bedrock is a fully managed serverless service that provides access to foundation models (FMs) from leading AI providers via an API, making it ideal for generating product descriptions from bullet points. It eliminates infrastructure management while allowing you to invoke models like Anthropic Claude or Amazon Titan for text generation tasks.

Exam trap

The trap here is that candidates confuse Amazon Comprehend (a text analysis service) with a generative AI service, or assume Polly or Lex can generate text descriptions when they are specialized for speech and conversation, respectively.

How to eliminate wrong answers

Option A is wrong because Amazon Comprehend is a natural language processing (NLP) service for extracting insights (e.g., sentiment, entities) from text, not for generating new content from bullet points. Option C is wrong because Amazon Polly is a text-to-speech service that converts text into lifelike speech, not a foundation model for text generation. Option D is wrong because Amazon Lex is a service for building conversational interfaces (chatbots) using automatic speech recognition and natural language understanding, not for generating product descriptions from bullet points.

48
MCQmedium

A media company uses Amazon Bedrock to generate personalized news summaries for its subscribers. The model occasionally produces summaries that include outdated facts from its training data. The company wants the summaries to reflect only the most recent articles from its internal content management system (CMS). The CMS exposes a REST API that returns the latest articles. Which approach should the company take to ensure the generated summaries are grounded in the latest articles?

A.Fine-tune the foundation model using Amazon Bedrock continued pre-training with a small set of recent articles.
B.Retrain the foundation model on the CMS articles daily using Amazon Bedrock custom model training.
C.Use Amazon Bedrock agents with an action group that calls the CMS REST API to fetch the latest articles, and include those articles in the prompt.
D.Increase the model's temperature setting to encourage more creative and current outputs.
AnswerC

Amazon Bedrock agents can be configured with action groups that invoke external APIs, such as the CMS REST API, to retrieve up-to-date content. The retrieved articles can then be inserted into the prompt as context, grounding the model's output in the latest facts. This is the intended pattern for dynamic, real-time grounding without retraining or fine-tuning the model.

Why this answer

The requirement is to ground generated summaries in the most recent articles from a CMS. Amazon Bedrock agents with action groups can call the CMS REST API at inference time, fetch the latest articles, and pass them as context to the foundation model. This retrieval-augmented approach ensures factual recency without retraining.

Other options either alter generation parameters or involve static training methods that cannot dynamically incorporate fresh content.

Exam trap

The trap here is assuming that fine-tuning or retraining the model is the only way to update its knowledge, when in fact retrieval-augmented generation via agents or knowledge bases is the appropriate pattern for dynamic, up-to-date grounding.

49
MCQmedium

A startup prototypes a support assistant using Amazon Bedrock and needs to compare how three different foundation models handle the same set of 200 support tickets. They want objective quality scores, including accuracy against reference answers and robustness, before committing to one model. Which approach should they use?

A.Create an Amazon Bedrock model evaluation job with the tickets and reference answers.
B.Increase Provisioned Throughput for each model and measure response latency.
C.Enable Guardrails for Amazon Bedrock on each model and compare block rates.
D.Send the tickets through each model and have engineers rate the replies informally.
AnswerA

Amazon Bedrock model evaluation jobs run a dataset of prompts and reference responses against selected models and produce metrics such as accuracy, robustness, and toxicity. Running the 200 tickets through one job gives comparable, objective scores across the three models, which is exactly what the startup needs to choose a model before committing.

Why this answer

The startup needs standardized, repeatable quality metrics across multiple models. Amazon Bedrock model evaluation jobs accept a dataset with prompts and reference responses and compute metrics such as accuracy, robustness, and toxicity for each selected model, making results directly comparable. Guardrail block rates, informal human ratings, and latency measurements capture different properties and cannot substitute for objective quality scoring.

Exam trap

The trap here is substituting a safety-control signal or a performance metric for genuine answer-quality evaluation against reference responses.

50
MCQhard

A financial services firm uses Amazon Bedrock with a foundation model to draft client emails. Compliance requires that no personally identifiable information from the prompt ever appear in the response, and that the model refuse requests for investment guarantees. The team wants a managed, configurable layer rather than custom prompt engineering alone. Which combination should they apply?

A.A higher temperature combined with a system prompt instructing the model to be careful.
B.Provisioned Throughput plus a custom lexicon in the tokenizer.
C.Guardrails for Amazon Bedrock with sensitive information filters and a denied topics policy.
D.Amazon Bedrock model evaluation with a toxicity metric and automatic retries.
AnswerC

Guardrails for Amazon Bedrock provides configurable safeguards applied to both prompts and responses. Sensitive information filters detect and block or mask personally identifiable information such as names, addresses, and account numbers, while a denied topics policy makes the model refuse requests about investment guarantees. This managed layer meets both compliance requirements without relying solely on prompt wording.

Why this answer

The firm needs runtime enforcement on both input and output. Guardrails for Amazon Bedrock applies sensitive information filters that detect and mask or block PII, and a denied topics policy that makes the model decline investment-guarantee requests. These are managed, configurable controls evaluated per request, unlike prompt wording, capacity reservations, or offline evaluation, none of which prevent prohibited content from reaching clients.

Exam trap

The trap here is treating prompt instructions or offline evaluation scores as equivalent to enforced runtime safeguards.

51
MCQmedium

A multinational corporation uses a foundation model via Amazon Bedrock to translate internal communication documents from English to multiple languages. They notice that the translations often miss company-specific jargon and acronyms, leading to confusion. The company has a glossary of approved translations for terms like 'Project Atlas' and 'Operation Synergy.' They want to improve translation accuracy quickly and with minimal effort. What approach should they take?

A.Use prompt engineering to include the glossary in each translation request.
B.Use a larger foundation model that has better language understanding.
C.Fine-tune the foundation model on a corpus of bilingual company documents.
D.Switch to Amazon Translate with custom terminology.
AnswerA

Embedding the approved glossary directly in each prompt supplies the model with the exact company-specific terms and acronyms at inference time, requiring no training or fine-tuning. This satisfies the minimal-effort, quick-improvement constraint for jargon like 'Project Atlas'.

Why this answer

Prompt engineering allows the company to inject the glossary directly into the context window of the foundation model with each translation request. This approach requires no model retraining or infrastructure changes, enabling rapid improvement by simply appending the approved translations as instructions or few-shot examples. It is the quickest and least effortful method to enforce company-specific terminology without altering the underlying model.

Exam trap

The trap here is that candidates may overestimate the effort required for prompt engineering or underestimate the speed and simplicity of in-context learning, leading them to choose fine-tuning or a different service when the most direct and minimal-effort solution is to augment the prompt with the glossary.

How to eliminate wrong answers

Option B is wrong because simply using a larger foundation model does not guarantee it will learn or prioritize the company's specific jargon and acronyms; larger models have broader general knowledge but still lack domain-specific customizations without additional context or fine-tuning. Option C is wrong because fine-tuning requires preparing a labeled bilingual corpus of company documents, which is time-consuming and resource-intensive, contradicting the requirement for minimal effort and quick improvement. Option D is wrong because Amazon Translate is a different service from Amazon Bedrock; the question explicitly states they are using a foundation model via Bedrock, and switching to a separate service would involve architectural changes and additional integration effort, not a minimal-effort adjustment.

52
MCQmedium

A developer is using Amazon Bedrock to generate text summaries. The output sometimes includes irrelevant information. What is the most effective prompt engineering technique to improve relevance?

A.Add a negative prompt specifying what to avoid
B.Use few-shot examples with summaries
C.Increase max tokens
D.Decrease temperature
AnswerB

Few-shot examples supply the model with concrete input-output pairs demonstrating the desired summary style and length, steering generation toward relevant content. This conditioning is more effective than abstract instructions because the model infers the expected pattern directly from demonstrated summaries.

Why this answer

Few-shot examples provide the model with explicit patterns of desired output, directly guiding it to produce summaries that match the format and content of the examples. This technique is the most effective for improving relevance because it gives the model concrete reference points, reducing the likelihood of including irrelevant information.

Exam trap

AWS often tests the misconception that adjusting generation parameters (like temperature or token limits) can substitute for explicit prompt structure, when in fact few-shot examples directly teach the model the expected output format and content relevance.

How to eliminate wrong answers

Option A is wrong because negative prompts (e.g., 'avoid irrelevant details') are less reliable in foundation models; they can be ignored or misinterpreted, and they do not provide the structured guidance that few-shot examples offer. Option C is wrong because increasing max tokens only expands the output length, which can actually increase the chance of including irrelevant information rather than improving relevance. Option D is wrong because decreasing temperature reduces randomness but does not teach the model what relevant content looks like; it may still produce irrelevant information if the prompt lacks clear examples.

53
MCQhard

A financial services firm uses Amazon Bedrock with Anthropic Claude to analyze earnings call transcripts. They need the model to output results in a strict JSON schema for downstream processing. The model occasionally returns prose or invalid JSON. Which Amazon Bedrock feature should they use to enforce the output structure?

A.Increase the top_p value to 1.0 to make the output more deterministic.
B.Guardrails for Amazon Bedrock with a denied topics policy.
C.Converse API with a tool definition that specifies the JSON schema as input parameters.
D.Model invocation logging to Amazon CloudWatch.
AnswerC

The Converse API supports tool use, where you define a tool with a JSON schema for its input parameters. When the model decides to use the tool, it returns structured arguments matching that schema, effectively enforcing JSON output. This is the intended mechanism in Amazon Bedrock for reliable structured generation, and it works across supported models without custom parsing hacks.

Why this answer

The Converse API's tool use capability lets you define a tool whose input schema is the desired JSON structure. The model then returns arguments conforming to that schema when it invokes the tool, providing reliable structured output. Guardrails, logging, and sampling parameters do not enforce a schema, so they cannot guarantee valid JSON for downstream processing.

Exam trap

The trap here is confusing content filtering or logging features with output formatting controls, when only tool use with a defined schema actually constrains the model to produce structured JSON.

54
Multi-Selecthard

A data scientist is fine-tuning a foundation model on SageMaker. They want to prevent overfitting. Which THREE actions can help? (Select THREE.)

Select 3 answers
A.Apply dropout
B.Increase training data size
C.Increase the number of epochs
D.Use early stopping
E.Use a smaller learning rate
AnswersA, B, D

Dropout randomly deactivates neurons during each training step, forcing the network to learn redundant representations rather than memorising training samples. This regularisation directly counteracts the overfitting the data scientist observes when fine-tuning the foundation model on SageMaker.

Why this answer

Option A (Apply dropout) is correct because randomly deactivating neurons during training acts as a regularizer that reduces the model's reliance on specific weights, thereby curbing overfitting. Option B (Increase training data size) is correct because more diverse examples give the model a broader signal and reduce the chance it memorizes the limited training set. Option D (Use early stopping) is correct because monitoring validation loss and halting training when it stops improving prevents the model from continuing to fit noise in the training data.

Option C (Increase the number of epochs) is not correct because training longer typically worsens overfitting rather than preventing it. Option E (Use a smaller learning rate) is not correct because a smaller learning rate mainly affects optimization stability and convergence speed, not the model's tendency to overfit.

Exam trap

AWS often tests the misconception that increasing epochs or using a smaller learning rate directly prevents overfitting, when in fact these are hyperparameter tuning strategies that can exacerbate or fail to address overfitting without explicit regularization.

55
MCQhard

A financial services company uses a foundation model for document analysis. They need to ensure the model does not output sensitive customer information from its training data. What is the most effective mitigation?

A.Implement output filtering using an external service
B.Choose a model that has been fine-tuned on financial data
C.Apply data masking before sending input
D.Use a private endpoint
AnswerA

Output filtering inspects generated responses before delivery, blocking any text matching sensitive customer data patterns. Because training-data leakage cannot be fully prevented at the model level, an external guardrail enforces the constraint at the point where disclosure would occur.

Why this answer

Output filtering using an external service is the most effective mitigation because it acts as a post-processing layer that can detect and redact sensitive customer information (e.g., PII, account numbers) before the model's response is returned to the user. This approach does not rely on the model's internal training or input modifications, which can be bypassed or incomplete. It provides a robust, policy-driven control that can be updated independently of the model.

Exam trap

The trap here is that candidates often confuse input-side controls (like data masking or fine-tuning) with output-side controls, assuming that protecting the input or training the model on domain data is sufficient to prevent leakage of memorized sensitive information.

How to eliminate wrong answers

Option B is wrong because fine-tuning on financial data does not guarantee the model will not memorize or regurgitate sensitive customer information from its original training data; fine-tuning adjusts the model's behavior but does not erase existing memorized data. Option C is wrong because data masking before sending input only protects the input data, not the model's outputs; the model could still output sensitive information from its training data that was never masked. Option D is wrong because using a private endpoint secures the network connection and access control but does not prevent the model from generating outputs containing sensitive training data; it addresses data-in-transit security, not output content safety.

56
MCQhard

A media company uses Amazon Bedrock to generate short video scripts. Writers report that outputs are inconsistent in tone and sometimes ignore the required scene structure. The team wants a reusable, versioned artifact that enforces the role, tone, and output format across many invocations without retraining the model, and they want to track changes over time. Which Amazon Bedrock feature should they use?

A.Amazon Bedrock Prompt Management with prompt versions
B.Provisioned Throughput on the foundation model
C.Amazon Bedrock Guardrails with a denied topics policy
D.Amazon Bedrock model evaluation jobs
AnswerA

Prompt Management stores prompts as reusable, versioned resources in Amazon Bedrock. Teams can define the role, tone, and required scene structure once, reference the prompt by its ARN across invocations, and create immutable versions to track changes. This enforces consistent instructions without retraining and provides the version history the writers need.

Why this answer

A reusable, versioned instruction artifact is exactly what Amazon Bedrock Prompt Management provides. Prompts can be authored, stored, and referenced by ARN, with versions that capture changes over time, so every invocation uses the same role, tone, and structure guidance without touching model weights. Evaluation jobs, Guardrails, and Provisioned Throughput address quality measurement, output filtering, and capacity respectively, not prompt standardization.

Exam trap

The trap here is confusing Guardrails, which filters outputs against policies, with Prompt Management, which stores and versions the instructions that shape output.

57
Multi-Selecthard

Which THREE are best practices for ensuring generated content complies with corporate brand guidelines when using Amazon Bedrock?

Select 3 answers
A.Implement guardrails to restrict tone, topics, and language
B.Use prompt engineering to specify brand voice and style
C.Increase the temperature for more creative outputs
D.Use random prompts to test variability
E.Fine-tune the model on a dataset of brand-compliant content
AnswersA, B, E

Guardrails for Amazon Bedrock applies configurable policies that block or filter disallowed topics, words, and tones at inference time. This enforces brand tone and language constraints consistently across every request, independent of prompt wording, satisfying the requirement to restrict generated content to corporate guidelines.

Why this answer

Option A is correct because Amazon Bedrock Guardrails let you define denied topics, word filters, and content filters that constrain the model's tone, subject matter, and language so outputs stay within corporate brand boundaries. Option B is correct because prompt engineering—embedding explicit instructions about brand voice, style, and formatting in the prompt—directly steers the model toward compliant outputs without retraining. Option E is correct because fine-tuning a model on a curated dataset of brand-compliant content adapts the model's behavior to consistently reproduce the organization's approved tone and terminology.

Option C is not correct because raising the temperature increases randomness and creativity, which makes outputs less predictable and more likely to drift from brand guidelines. Option D is not correct because using random prompts to test variability does not enforce compliance; it merely measures output variation and does nothing to align content with brand standards.

Exam trap

AWS often tests the misconception that increasing temperature or using random prompts can help enforce brand guidelines, when in fact these actions increase variability and reduce control, directly opposing the goal of compliance.

58
MCQmedium

A company fine-tunes a foundation model on SageMaker using a custom dataset. They notice the training job takes too long. Which optimization technique is specifically designed to reduce training time for foundation models?

A.Distributed training using SageMaker Data Parallelism
B.Using a smaller instance type
C.Using Spot Instances
D.Reducing batch size
AnswerA

SageMaker distributed data parallelism shards the training dataset and gradients across multiple GPUs or instances, so each worker processes a subset and synchronises updates. This cuts wall-clock training time for large foundation models that cannot fit efficiently on a single device.

Why this answer

SageMaker Data Parallelism distributes the training workload across multiple GPUs or instances, splitting the data and synchronizing gradients using optimized all-reduce algorithms. This specifically reduces training time for large foundation models by enabling parallel computation, which is the most direct technique for accelerating training at scale.

Exam trap

AWS often tests the misconception that cost-saving techniques like Spot Instances or smaller instances also improve performance, but the question specifically asks for optimization to reduce training time, not cost.

How to eliminate wrong answers

Option B is wrong because using a smaller instance type reduces computational capacity, which would increase training time rather than reduce it. Option C is wrong because Spot Instances reduce cost by using spare AWS capacity, but they do not inherently speed up training; they may even cause interruptions that prolong total time. Option D is wrong because reducing batch size can actually slow convergence and increase the number of training steps, potentially increasing overall training time.

59
Multi-Selectmedium

A company is using Amazon Bedrock to generate marketing content. They want to evaluate the quality of the generated text. Which TWO metrics are most appropriate for evaluating text quality?

Select 2 answers
A.Precision
B.Perplexity
C.Accuracy
D.F1 score
E.BLEU (Bilingual Evaluation Understudy)
AnswersB, E

Perplexity quantifies how confidently a language model predicts a sample, so lower values indicate the generated marketing text is more fluent and statistically coherent. It directly measures the text-quality dimension the stem asks about, unlike task-specific overlap metrics.

Why this answer

Perplexity (B) is a standard intrinsic metric for generative language models that measures how well a probability distribution predicts a sample, with lower values indicating the model is less surprised by the text and thus produces more fluent, coherent output. BLEU (E) is a reference-based metric that compares n-gram overlap between generated text and one or more reference texts, making it well suited for evaluating the quality of machine-generated marketing copy against human-written examples. Precision (A), Accuracy (C), and F1 score (D) are classification metrics that require discrete predicted labels and ground-truth classes, so they do not apply to free-form text generation where the output is a sequence of tokens rather than a category.

Exam trap

AWS often tests the distinction between classification metrics (precision, accuracy, F1) and generation evaluation metrics (perplexity, BLEU), leading candidates to mistakenly apply classification concepts to text quality assessment.

60
MCQhard

A developer sends the above request to Amazon Bedrock with Anthropic Claude. The model returns a response that stops before reaching 500 tokens. What is the most likely reason?

A.The temperature is set too high
B.The model is not trained on this topic
C.The model reached a stop sequence
D.The token limit is exceeded
AnswerC

A stop sequence supplied in the request causes generation to halt immediately when that string is produced, regardless of remaining max token budget. The response therefore ends well before 500 tokens, which is expected behaviour rather than truncation or throttling.

Why this answer

The model stopped before reaching 500 tokens because the request likely included a stop sequence (e.g., `\n\nHuman:` or a custom stop token) that matched the generated output. When a stop sequence is encountered, Bedrock immediately halts generation, even if the token limit has not been reached. This is the most direct explanation for a premature stop.

Exam trap

AWS often tests the distinction between a stop sequence and a token limit; the trap here is that candidates confuse a premature stop with exceeding the token limit, but a stop sequence causes an early halt while a token limit would cause truncation at the limit.

How to eliminate wrong answers

Option A is wrong because a high temperature increases randomness and can cause the model to generate more tokens or diverge, not stop early. Option B is wrong because Bedrock's Claude models are trained on a broad corpus and can generate responses on any topic; lack of training would produce low-quality or repetitive text, not a stop before the token limit. Option D is wrong because if the token limit were exceeded, the model would truncate the response at the limit, not stop before reaching it.

61
Multi-Selecthard

A data science team is fine-tuning a foundation model on Amazon SageMaker. Which THREE steps are part of the best practice? (Choose three.)

Select 3 answers
A.Increase model size to improve performance.
B.Monitor for catastrophic forgetting during fine-tuning.
C.Use early stopping to prevent overfitting.
D.Deploy the model to production immediately after fine-tuning.
E.Use a diverse dataset representing various scenarios.
AnswersB, C, E

Fine-tuning can overwrite pretrained weights, causing catastrophic forgetting where the model loses general capabilities. Monitoring benchmark performance during training detects this degradation early, satisfying the best-practise requirement to preserve the foundation model's original knowledge while adapting it to the new task.

Why this answer

Option B is correct because fine-tuning a foundation model on a narrow dataset can cause catastrophic forgetting, where the model loses previously learned general capabilities; monitoring for this (e.g., via evaluation on held-out general benchmarks) is a recognized best practice. Option C is correct because early stopping halts training when validation performance stops improving, which prevents overfitting to the fine-tuning dataset and is a standard SageMaker technique (e.g., via the EarlyStopping callback or stopping criterion). Option E is correct because using a diverse dataset that represents various scenarios improves generalization and reduces bias, ensuring the fine-tuned model performs well across the range of real-world inputs it will encounter.

Option A is not correct because simply increasing model size does not guarantee better fine-tuning results and can increase cost, latency, and overfitting risk. Option D is not correct because deploying immediately after fine-tuning skips essential validation, evaluation, and testing steps before production release.

Exam trap

AWS often tests the misconception that fine-tuning always requires a larger model or immediate deployment, while the real best practices focus on validation, monitoring, and data diversity to maintain model robustness.

62
MCQmedium

An e-commerce company uses Amazon Bedrock to generate product descriptions. They notice the descriptions are too long and contain repetitive phrases. Which parameter adjustment can help?

A.Increase frequency penalty
B.Increase temperature
C.Increase top_p
D.Decrease presence penalty
AnswerA

Raising the frequency penalty directly penalises tokens proportional to how often they have already appeared, which suppresses the repetitive phrasing the stem describes. It does not shorten output, so pair it with a lower maximum length to address the excessive description length.

Why this answer

Increasing the frequency penalty reduces the likelihood of the model repeating the same phrases or tokens, directly addressing the issue of repetitive language in generated product descriptions. This parameter penalizes tokens that have already appeared in the text, encouraging more diverse output and naturally shortening overly long descriptions by avoiding redundant loops.

Exam trap

AWS often tests the distinction between frequency penalty and presence penalty, where candidates confuse 'penalizing repetition' with 'reducing randomness' and incorrectly choose temperature or top_p adjustments.

How to eliminate wrong answers

Option B is wrong because increasing temperature makes the model more random and creative, which could actually worsen verbosity and introduce more irrelevant phrases rather than reducing repetition. Option C is wrong because increasing top_p (nucleus sampling) expands the set of possible next tokens, which may increase diversity but does not specifically penalize repeated tokens and can still produce long, repetitive text. Option D is wrong because decreasing presence penalty would reduce the penalty for tokens that have already appeared, making the model more likely to repeat itself, which is the opposite of what is needed.

63
MCQmedium

A company runs a chatbot using a large language model on Amazon Bedrock. They notice high latency during peak hours. Which action would be MOST effective to reduce latency without degrading response quality?

A.Increase the number of concurrent invocations
B.Switch to a smaller model
C.Decrease the maxTokens parameter
D.Use Provisioned Throughput for model inference
AnswerD

Provisioned Throughput reserves dedicated model units for your Amazon Bedrock model, eliminating the queueing that causes peak-hour latency. Unlike on-demand inference, which shares capacity, it guarantees consistent throughput without altering the model, prompt, or parameters — so response quality stays identical while latency drops.

Why this answer

Provisioned Throughput on Amazon Bedrock reserves dedicated capacity for model inference, ensuring consistent low latency even during peak hours. This eliminates the variability caused by resource contention in the on-demand tier, directly addressing high latency without altering model size or output quality.

Exam trap

AWS often tests the misconception that reducing model size or output length is the primary way to reduce latency, but the real bottleneck in peak-hour scenarios is often infrastructure contention, which Provisioned Throughput resolves without sacrificing quality.

How to eliminate wrong answers

Option A is wrong because increasing concurrent invocations without dedicated capacity can exacerbate resource contention, leading to throttling and higher latency. Option B is wrong because switching to a smaller model reduces response quality (e.g., lower accuracy or coherence), which degrades the chatbot's performance. Option C is wrong because decreasing maxTokens truncates responses, degrading output quality by cutting off reasoning or context, and does not address the root cause of latency from infrastructure contention.

64
MCQeasy

A developer is using Amazon Bedrock to build a chatbot that answers customer queries. The chatbot must only respond based on the provided company documentation. Which approach best meets this requirement?

A.Use prompt engineering to instruct the model to only use documentation.
B.Use a RAG architecture with the company documentation as the knowledge base.
C.Fine-tune a foundation model on the company documentation.
D.Use a text classification model to filter responses.
AnswerB

RAG constrains responses by retrieving passages from the company documentation and supplying them as context, so answers derive from that corpus rather than the model's parametric knowledge. This satisfies the requirement to answer only from provided documentation.

Why this answer

Retrieval-Augmented Generation (RAG) architecture retrieves relevant chunks from the company documentation at query time and injects them into the prompt, ensuring the model's response is grounded solely in the provided documents. This approach prevents the model from relying on its internal training data or generating information outside the documentation, which is critical for a closed-domain chatbot.

Exam trap

The AIF-C01 exam often tests the distinction between prompt engineering and RAG, where candidates mistakenly believe a well-crafted prompt can fully control model behavior without a retrieval mechanism, overlooking the fact that foundation models inherently generate responses from their training data unless explicitly grounded via external knowledge retrieval.

How to eliminate wrong answers

Option A is wrong because prompt engineering alone cannot guarantee the model will ignore its pre-trained knowledge; the model may still hallucinate or use information not present in the documentation, as it has no mechanism to enforce retrieval of specific content. Option C is wrong because fine-tuning a foundation model on the company documentation embeds the data into the model's weights, which can lead to outdated or incomplete responses and does not allow dynamic retrieval of the latest documentation; it also risks overfitting and does not scale well with changing content. Option D is wrong because a text classification model filters responses after generation, but it cannot ensure the response is based on the documentation; it only labels or rejects outputs, which is insufficient for generating accurate, document-grounded answers.

65
MCQhard

A company operates a customer service platform that uses Amazon Bedrock with a foundation model to generate automated responses. The system has been in production for three months. Recently, customers have reported that responses are becoming repetitive and less relevant over time. The development team notices that the model's performance has degraded, especially for queries about newer products that were added after the initial deployment. The team currently uses a static prompt with a fixed knowledge base that was set up at launch. The model is invoked via the Bedrock API with standard settings. The team wants to improve response quality without incurring high costs or extensive re-engineering. What should the team do?

A.Increase the temperature parameter to 0.9 to introduce more randomness and reduce repetition.
B.Fine-tune the model every week on the latest customer interactions using Amazon SageMaker.
C.Switch to a larger foundation model to handle the increased complexity of new products.
D.Implement a feedback loop to periodically update the knowledge base with new product information and use a dynamic prompt that includes recent interactions.
AnswerD

Updating the knowledge base periodically addresses the stale fixed data, which is the root cause of degraded relevance for newer products. A dynamic prompt incorporating recent interactions supplies current context at inference time, countering repetition without retraining or re-engineering, keeping costs low as the stem requires.

Why this answer

The core issue is that the static knowledge base and prompt have become stale as new products were added. Implementing a feedback loop to periodically update the knowledge base with new product information and using a dynamic prompt that includes recent interactions directly addresses the root cause of degradation—outdated context—without requiring costly fine-tuning or model swaps. This approach leverages Amazon Bedrock's native capabilities for retrieval-augmented generation (RAG) to keep responses relevant and non-repetitive.

Exam trap

A common pitfall is assuming model performance degradation always requires tuning the model (e.g., temperature or fine-tuning), when in fact it is often a data freshness and context management issue that can be solved with RAG and dynamic prompting.

How to eliminate wrong answers

Option A is wrong because increasing the temperature parameter to 0.9 introduces more randomness, which may reduce repetition but will also make responses less coherent and less relevant, especially for factual queries about new products; it does not address the stale knowledge base. Option B is wrong because fine-tuning the model every week on the latest customer interactions using Amazon SageMaker would incur high costs and extensive re-engineering, contradicting the requirement to avoid high costs and extensive re-engineering; also, fine-tuning is overkill when the issue is outdated context, not model architecture. Option C is wrong because switching to a larger foundation model does not solve the problem of missing new product information in the knowledge base; it would increase latency and cost without guaranteeing relevance, and the model's performance degradation is due to static context, not model capacity.

66
MCQhard

A company uses Amazon Bedrock to generate product descriptions. They want to ensure the outputs consistently follow a specific brand tone (professional yet friendly). They have a small set of example descriptions (few-shot examples) but do not want to fine-tune the model. Which strategy best achieves consistent tone without modifying the base model?

A.Fine-tune the model on a dataset of product descriptions that exemplify the desired tone.
B.Use a system prompt that defines the brand tone and include few-shot examples in the prompt.
C.Implement prompt chaining by breaking the task into multiple steps, each with its own prompt.
D.Use retrieval-augmented generation (RAG) to pull example descriptions from a database and prepend them to the prompt.
AnswerB

A system prompt sets persistent behavioural instructions, while few-shot examples demonstrate the desired professional-yet-friendly phrasing in context. Together they steer generation at inference time, achieving consistent brand tone without altering weights or incurring fine-tuning cost.

Why this answer

Using a system prompt to define the brand tone and including few-shot examples in the prompt is the best strategy because it guides the model's behavior without modifying its weights. This leverages in-context learning, which is effective for style adaptation. Fine-tuning is unnecessary and more resource-intensive.

Exam trap

The trap is thinking that fine-tuning is required for consistent tone, but few-shot prompting with a system prompt is often sufficient and avoids model modification.

How to eliminate wrong answers

Option A is wrong because fine-tuning modifies the base model, which the company explicitly wants to avoid. Option C is wrong because prompt chaining breaks the task into steps but does not inherently enforce a consistent tone; it's more for complex reasoning. Option D is wrong because RAG is for incorporating external knowledge, not for style control; prepending examples is similar to few-shot but RAG typically retrieves relevant documents, not style examples.

67
MCQeasy

A company needs to summarize thousands of customer reviews daily using a foundation model. The solution must minimize latency and cost while handling variable traffic. Which AWS service should they use?

A.Amazon SageMaker real-time endpoint
B.Amazon Comprehend
C.Amazon Lex
D.Amazon Bedrock with on-demand mode
AnswerD

Bedrock on-demand mode bills per token processed and automatically scales with variable traffic, avoiding provisioned throughput charges during idle periods. This satisfies both the latency and cost constraints for summarising thousands of reviews daily without capacity planning.

Why this answer

Amazon Bedrock with on-demand mode is correct because it provides serverless access to foundation models (FMs) with pay-per-use pricing, which minimizes cost for variable traffic and eliminates the need to provision infrastructure. The on-demand mode handles thousands of daily summarization requests with low latency by leveraging AWS's scalable inference infrastructure, making it ideal for variable workloads without upfront commitments.

Exam trap

The trap here is that candidates often confuse Amazon Comprehend's pre-built NLP capabilities with foundation model summarization, failing to recognize that Comprehend cannot perform generative abstractive summarization and is limited to extractive tasks like key phrases and sentiment.

How to eliminate wrong answers

Option A is wrong because Amazon SageMaker real-time endpoints require provisioning and managing dedicated instances, leading to higher costs and latency overhead for variable traffic, and they are not optimized for foundation model inference without custom containers. Option B is wrong because Amazon Comprehend is a natural language processing (NLP) service for tasks like sentiment analysis and entity extraction, not a foundation model service for generative summarization; it lacks the ability to use customizable FMs for abstractive summarization. Option C is wrong because Amazon Lex is designed for building conversational chatbots using intent and slot models, not for summarizing text with foundation models; it does not provide FM inference capabilities.

68
MCQmedium

A company uses Amazon Bedrock to generate marketing copy. The summaries are too verbose. Which parameter should be decreased to directly limit the length of the output?

A.max_tokens
B.temperature
C.top_p
D.frequency_penalty
AnswerA

Lowering max_tokens caps the number of tokens the model may generate, directly truncating verbose marketing copy at the configured ceiling. It satisfies the stem's constraint of limiting output length, unlike temperature or top_p, which alter sampling randomness rather than imposing a hard generation limit.

Why this answer

The `max_tokens` parameter directly controls the maximum number of tokens (words or subwords) in the generated output. By decreasing this value, you explicitly cap the length of the marketing copy, making it less verbose. This is the most direct way to limit output length in Amazon Bedrock and other LLM APIs.

Exam trap

The trap here is that candidates confuse parameters that affect output style (temperature, top_p, frequency_penalty) with the one that directly controls output length (max_tokens), leading them to pick a parameter that changes how the model writes rather than how much it writes.

How to eliminate wrong answers

Option B (temperature) is wrong because it controls the randomness of token selection, not the length of the output; lowering temperature makes output more deterministic but does not shorten it. Option C (top_p) is wrong because it sets a cumulative probability threshold for nucleus sampling, affecting diversity of token choices, not the total number of tokens generated. Option D (frequency_penalty) is wrong because it penalizes tokens based on their frequency in the generated text, reducing repetition but not directly limiting the overall length of the response.

69
MCQeasy

Refer to the exhibit. This is an Amazon Bedrock invocation request for Claude. What is the purpose of the "stop_sequences" parameter?

A.It tells the model to stop generating when it encounters that sequence
B.It specifies a character sequence for the model to include in its response
C.It limits the number of tokens in the response
D.It controls the randomness of the response
AnswerA

The stop_sequences parameter supplies one or more strings that halt generation as soon as the model emits them, truncating output at that point. It satisfies the invocation request's need to control where Claude stops, rather than limiting token count or filtering content.

Why this answer

The 'stop_sequences' parameter in Amazon Bedrock's invocation request for Claude tells the model to halt generation as soon as it encounters a specified character sequence. This allows developers to control the output format, such as stopping at a newline or a custom delimiter, ensuring the response ends exactly where intended.

Exam trap

AWS often tests the distinction between parameters that control output length (max_tokens_to_sample) versus those that control output termination (stop_sequences), leading candidates to confuse token limits with stop sequences.

How to eliminate wrong answers

Option B is wrong because 'stop_sequences' does not instruct the model to include a sequence; it tells the model to stop when that sequence is generated, not to add it. Option C is wrong because token limits are controlled by the 'max_tokens_to_sample' parameter, not by stop sequences. Option D is wrong because randomness is controlled by the 'temperature' parameter, not by stop sequences.

70
Multi-Selecteasy

A company is using Amazon Bedrock to generate code snippets. They want to ensure the generated code is secure. Which TWO practices should they implement?

Select 2 answers
A.Increase the max token limit to generate longer code.
B.Use guardrails to block insecure code patterns.
C.Set the temperature to 0 for deterministic output.
D.Review and test all generated code before deployment.
E.Use a larger model for better accuracy.
AnswersB, D

Guardrails apply configurable content filters that intercept prompts and responses, blocking insecure code patterns before they reach developers. This satisfies the stem's security requirement by preventing vulnerable snippets at generation time rather than after deployment.

Why this answer

Option B is correct because Amazon Bedrock Guardrails let you define denied topics, content filters, and (critically for this scenario) sensitive-information and custom word/regex filters that can detect and block insecure code patterns such as hardcoded credentials or known dangerous constructs before the response is returned. Option D is correct because no generative model guarantees secure output; generated code must be treated as untrusted and go through human review plus SAST/DAST and unit testing in the CI/CD pipeline before deployment, which is the standard secure-SDLC control for AI-generated artifacts. Option A does not belong because raising the max token limit only affects output length, not security posture.

Option C does not belong because temperature 0 makes sampling more deterministic but does not make code secure. Option E does not belong because a larger model may improve accuracy or fluency, but accuracy is not a security control and does not prevent insecure code.

Exam trap

The AIF-C01 exam often tests the misconception that model parameters like temperature or token limits can substitute for explicit security controls, when in fact only guardrails and human review directly address code security.

71
Multi-Selecthard

A team is deploying a foundation model on Amazon Bedrock for a customer-facing assistant. They must reduce hallucinations and keep answers grounded in approved company content while controlling inference cost. Which TWO approaches should the team implement? (Choose two.)

Select 2 answers
A.Remove all system instructions so the model relies solely on its pretrained knowledge.
B.Switch to the largest available foundation model to guarantee factual accuracy.
C.Configure Amazon Bedrock Guardrails contextual grounding and relevance checks to filter ungrounded responses.
D.Increase the model temperature to encourage more creative and varied responses.
E.Use Amazon Bedrock knowledge bases to retrieve relevant approved passages and include them in the prompt before generation.
AnswersC, E

Guardrails contextual grounding evaluates whether a response is supported by the provided reference source and whether it is relevant to the user query, blocking or flagging outputs that fail thresholds. This adds a verification layer after generation, catching unsupported claims that retrieval alone may not prevent, and it strengthens grounding for customer-facing answers.

Why this answer

Grounding answers in approved content through Bedrock knowledge bases supplies factual context, while Guardrails contextual grounding and relevance checks verify that generated responses are supported by that context. Together they reduce hallucination and keep outputs aligned with company material. Raising temperature, removing instructions, or simply choosing a larger model does not provide grounding and can increase cost or fabrication risk.

Exam trap

The trap here is treating model size or creativity settings as accuracy controls, when grounding requires retrieved source content plus verification of the response against it.

72
MCQeasy

A developer invokes an Amazon Bedrock model and receives the above response. What does the 'stopReason' field indicate?

A.The model encountered an error.
B.The model reached a defined stop sequence.
C.The model hit the maximum token limit.
D.The model stopped due to a safety filter.
AnswerB

The stopReason field reports why generation halted. A value indicating a stop sequence means the model emitted one of the custom strings supplied in the request, terminating output early rather than hitting the token limit or natural end.

Why this answer

The 'stopReason' field in an Amazon Bedrock response indicates why the model stopped generating tokens. When set to 'stop', it means the model encountered a defined stop sequence (such as a special token like <|endoftext|> or a user-specified string) and halted generation normally. This is the expected behavior for a successful, complete response.

Exam trap

The trap here is that candidates confuse 'stop' (normal completion via stop sequence) with 'length' (token limit reached), as both end generation but have different implications for response completeness and cost.

How to eliminate wrong answers

Option A is wrong because a model error would typically result in an HTTP error code or a different field like 'error' or 'failure', not a 'stopReason' of 'stop'. Option C is wrong because hitting the maximum token limit would produce a 'stopReason' of 'length', not 'stop'. Option D is wrong because a safety filter intervention would produce a 'stopReason' of 'content_filtered' or similar, not 'stop'.

73
MCQeasy

A developer receives the above response from invoking a Bedrock model. Which field indicates that the model completed its response normally?

A.output
B.stop_reason
C.text
D.role
AnswerB

The `stop_reason` field reports why generation halted, returning `end_turn` when the model finished naturally rather than hitting the `max_tokens` limit or a stop sequence. This directly satisfies the stem's requirement to identify normal completion, distinguishing it from truncation or content filtering.

Why this answer

The `stop_reason` field in the Bedrock response indicates why the model stopped generating text. A value of `"stop"` or `"end_turn"` (depending on the model) signals that the model completed its response normally, as opposed to hitting a token limit, content filter, or other interruption.

Exam trap

The trap here is that candidates confuse the `output` container or the `text` field with the completion indicator, overlooking the dedicated `stop_reason` field that explicitly signals normal termination.

How to eliminate wrong answers

Option A is wrong because `output` is a container object that holds the generated content, not a field that indicates the completion status. Option C is wrong because `text` is a field within the output that contains the actual generated string, but it does not convey why generation stopped. Option D is wrong because `role` indicates the conversational role (e.g., user or assistant) in a multi-turn context, not the model's completion state.

74
MCQeasy

A company uses Amazon Bedrock to build a conversational AI. They want to enforce role-based access to the model. Which AWS service should they use?

A.AWS Config
B.AWS Identity and Access Management (IAM)
C.AWS CloudTrail
D.AWS Organizations
AnswerB

AWS IAM supplies the role-based access control the stem demands, letting you attach identity-based policies to roles and users that permit or deny specific Bedrock actions such as InvokeModel on particular model ARNs. Bedrock integrates with IAM for authorisation, so permissions are enforced per caller without application-level checks.

Why this answer

AWS Identity and Access Management (IAM) is the correct service because it enables fine-grained, role-based access control (RBAC) to Amazon Bedrock models. You can define IAM policies that specify which principals (users, groups, or roles) are allowed to invoke specific foundation models, ensuring that only authorized roles can interact with the conversational AI.

Exam trap

The trap here is that candidates often confuse AWS Config (which audits configurations) or CloudTrail (which logs actions) with IAM, mistakenly thinking that logging or compliance tools can enforce access control, when in fact only IAM provides the authorization layer for Bedrock model invocation.

How to eliminate wrong answers

Option A is wrong because AWS Config is a service for evaluating and auditing resource configurations against compliance rules, not for enforcing role-based access to Bedrock models. Option C is wrong because AWS CloudTrail records API activity for auditing and governance, but it does not control or enforce access permissions. Option D is wrong because AWS Organizations manages multi-account governance and policy inheritance across accounts, but it does not provide the granular, per-model role-based access control needed for Bedrock.

75
MCQhard

Which parameter controls the randomness of generated text in a foundation model?

A.top_p
B.stop sequences
C.max_tokens
D.temperature
AnswerD

Temperature directly scales the model's output logits before the softmax, flattening or sharpening the probability distribution over the next token. Lower values make high-probability tokens dominate, producing deterministic text; higher values spread probability mass, increasing randomness. This satisfies the stem's requirement for the parameter governing randomness in foundation models.

Why this answer

Temperature is the correct parameter because it directly controls the randomness of token sampling in a foundation model. A lower temperature (e.g., 0.1) makes the model more deterministic by concentrating probability mass on the most likely tokens, while a higher temperature (e.g., 1.5) flattens the probability distribution, increasing the likelihood of less probable tokens and thus generating more diverse or creative outputs.

Exam trap

AWS often tests the distinction between temperature (which reshapes the probability distribution) and top_p (which truncates the token set), leading candidates to confuse 'randomness control' with 'diversity via cumulative probability threshold'.

How to eliminate wrong answers

Option A is wrong because top_p (nucleus sampling) controls the cumulative probability threshold for token selection, not the randomness of the distribution itself; it dynamically chooses a set of tokens whose cumulative probability exceeds p, which is a different mechanism for diversity. Option B is wrong because stop sequences define specific strings that halt text generation (e.g., '\n\n' or a period), and they have no effect on the randomness or sampling behavior of the model. Option C is wrong because max_tokens sets a hard limit on the number of tokens generated in the output, controlling length rather than the stochasticity of token selection.

Page 1 of 2 · 128 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Applications of Foundation Models questions.