Courseiva

CCNA Fundamentals of Generative AI Questions

75 of 125 questions · Page 1/2 · Fundamentals of Generative AI · Answers revealed

1
Multi-Selectmedium

Which THREE steps are typically involved in fine-tuning a foundation model? (Select THREE.)

Select 3 answers
A.Deploy the model immediately without additional training
B.Prepare a labeled dataset specific to the target domain
C.Train the model on the domain dataset with a lower learning rate
D.Select a pre-trained foundation model as the starting point
E.Choose a model architecture with more parameters than the base model
AnswersB, C, D

Fine-tuning adapts a foundation model's weights to a target task, which requires supervised examples the model can learn from. Preparing a labelled dataset specific to the target domain supplies those input-output pairs, satisfying the scenario's need for task-relevant training data rather than relying on the model's general pre-trained knowledge.

Why this answer

Fine-tuning a foundation model begins with selecting an appropriate pre-trained foundation model as the starting point (D), since the whole point of fine-tuning is to adapt existing general-purpose weights rather than train from scratch. Next, you must prepare a labeled dataset specific to the target domain (B), because supervised fine-tuning requires task-relevant input-output pairs to steer the model toward the desired behavior. Then you train the model on that domain dataset with a lower learning rate (C), which is standard practice to avoid catastrophic forgetting and to gently nudge the pre-trained weights instead of overwriting them.

Option A is incorrect because deploying the model immediately without additional training is the opposite of fine-tuning—it describes using the base model as-is. Option E is incorrect because fine-tuning does not require choosing an architecture with more parameters than the base model; you typically fine-tune the same pre-trained architecture, and increasing parameter count is not a fine-tuning step.

Exam trap

AWS often tests the distinction between fine-tuning and other adaptation methods (like prompt engineering or retrieval-augmented generation), and the trap here is that candidates might think fine-tuning requires a larger model or no additional data, when in fact it requires a labeled dataset and the same architecture.

2
MCQhard

Refer to the exhibit. A developer runs the CLI command to summarize text using Claude v2 in Bedrock. The output is shorter than expected. Which change should the developer make to allow a longer response?

A.Increase 'max_tokens_to_sample' to 1000
B.Change the prompt to include 'Write a long summary'
C.Set 'stop_reason' to 'none'
D.Use a different region like us-west-2
AnswerA

The response is truncated because max_tokens_to_sample caps generated output length. Raising it to 1000 permits Claude v2 to emit more tokens before stopping, directly satisfying the requirement for a longer summary in the Bedrock CLI invocation.

Why this answer

The 'max_tokens_to_sample' parameter in the Bedrock InvokeModel API directly controls the maximum number of tokens the model can generate in its response. By default, this value is often set low (e.g., 256 tokens), which truncates the output. Increasing it to 1000 allows Claude v2 to produce a longer summary up to that token limit.

Exam trap

A common mistake is to assume that simply asking the model to produce a longer response via prompt engineering will override the API parameter 'max_tokens_to_sample'. In AWS Bedrock, the 'max_tokens_to_sample' parameter is the definitive control for response length.

How to eliminate wrong answers

Option B is wrong because simply adding 'Write a long summary' to the prompt does not override the hard token limit set by 'max_tokens_to_sample'; the model will still stop generating once the token budget is exhausted, regardless of instruction phrasing. Option C is wrong because 'stop_reason' is an output field returned by the API indicating why generation stopped (e.g., 'max_tokens', 'stop_sequence'), not an input parameter; setting it to 'none' is invalid and would not affect response length. Option D is wrong because the region (e.g., us-west-2) does not impose a different default token limit for Claude v2; the 'max_tokens_to_sample' parameter is a model-level constraint independent of regional deployment.

3
MCQmedium

A developer is using the Amazon Bedrock API to generate text. They notice that the model sometimes returns harmful content despite setting safety parameters. What is the BEST way to add an additional layer of content filtering?

A.Fine-tune the model on a curated safe dataset
B.Configure content filters in Amazon Bedrock Guardrails
C.Improve prompt engineering with more specific instructions
D.Use AWS WAF to filter API responses
AnswerB

Amazon Bedrock Guardrails applies configurable content filters that evaluate both prompts and model responses independently of the model's own safety parameters, blocking harmful categories the model may still emit and adding a deterministic policy layer.

Why this answer

Amazon Bedrock Guardrails provides a dedicated, configurable content filtering layer that can block harmful content at inference time, independent of the model's built-in safety parameters. This allows developers to enforce custom policies (e.g., hate speech, violence) without modifying the model itself, making it the best additional safeguard.

Exam trap

The AIF-C01 exam often tests the misconception that fine-tuning or prompt engineering alone can fully prevent harmful outputs, when in fact a separate, configurable guardrail layer is the recommended approach for production-grade content filtering in Amazon Bedrock.

How to eliminate wrong answers

Option A is wrong because fine-tuning the model on a curated safe dataset adjusts the model's weights to reduce harmful outputs, but it does not guarantee filtering of all harmful content at inference and requires significant retraining effort; it is not an 'additional layer' but a model modification. Option C is wrong because improving prompt engineering with more specific instructions can guide the model's behavior but cannot reliably block harmful content that the model might generate despite instructions, as it lacks enforcement at the API response level. Option D is wrong because AWS WAF is a web application firewall designed to filter HTTP requests to web applications, not to inspect or filter the content of API responses from Bedrock; it operates at the network layer, not the application content layer.

4
MCQmedium

A media company uses a foundation model on Amazon Bedrock to generate article summaries. The model occasionally omits important details. Which prompt engineering technique is most likely to improve completeness?

A.Use a lower temperature setting
B.Increase the max tokens limit
C.Include a list of required key points in the prompt
D.Add 'Be concise' to the prompt
AnswerC

Listing required key points in the prompt explicitly constrains the model's output coverage, directing it to address each specified element rather than relying on its own judgement of relevance. This directly counteracts the omission of important details by making completeness an explicit instruction.

Why this answer

Explicitly listing required key points in the prompt guides the foundation model to cover all specified elements, directly addressing the omission issue. This technique, often called 'constrained generation' or 'structured prompting,' forces the model to attend to each required detail, improving completeness without altering model parameters.

Exam trap

A common misconception is that adjusting model parameters like temperature or token limits can fix content quality issues like missing details. In practice, prompt structure and explicit instructions, such as listing required key points, are the primary tools for controlling output completeness and accuracy.

How to eliminate wrong answers

Option A is wrong because lowering temperature reduces randomness and creativity but does not guarantee inclusion of specific details; it may even make outputs more repetitive or conservative, potentially omitting important points. Option B is wrong because increasing max tokens only allows longer outputs but does not instruct the model what content to include; the model may still omit key details within the expanded token budget. Option D is wrong because adding 'Be concise' encourages brevity, which is counterproductive to completeness and may cause the model to omit even more details.

5
MCQeasy

A hospital's IT team wants a generative AI assistant that answers patient-billing questions using only the hospital's internal policy documents, without retraining the model. Which approach should they use?

A.Fine-tuning the foundation model on the full set of policy documents
B.Increasing the model's temperature setting so it explores more of its pretrained knowledge
C.Retrieval Augmented Generation (RAG) by querying a vector store of the policy documents and passing retrieved passages to the model
D.Prepending a system instruction telling the model to answer only from hospital policy
AnswerC

RAG retrieves relevant passages from an external knowledge source at inference time and includes them in the prompt, so the model grounds its answer in the hospital's own policies without any weight updates. This directly satisfies the requirement to avoid retraining while keeping answers tied to authoritative internal documents.

Why this answer

Retrieval Augmented Generation supplies the model with relevant, current policy text at query time, so responses are grounded in the hospital's own documents. Fine-tuning changes model weights and needs retraining, temperature alters sampling randomness, and a system prompt alone provides no source content. Only retrieval of the policy passages gives accurate, updatable answers without retraining.

Exam trap

The trap here is assuming that instructing a model to stay on topic is equivalent to giving it the source material it needs to answer accurately.

6
MCQhard

A company operates in a region where Amazon Bedrock is not available. They want to use generative AI but must keep data within the country. Which solution should they consider?

A.Use Amazon SageMaker to host an open-source model in the local region.
B.Wait for Bedrock to become available in their region; there is no alternative.
C.Use Amazon Bedrock in the nearest available region with cross-region inference.
D.Use an API from a third-party generative AI provider with AWS PrivateLink.
AnswerA

SageMaker lets you host an open-source model on instances inside the local region, so training and inference data never leave the country. This satisfies the data-residency constraint that rules out Amazon Bedrock, which is unavailable there.

Why this answer

Amazon SageMaker allows you to host open-source models (e.g., Llama 2, Falcon) in any AWS region, including those where Bedrock is unavailable. This satisfies the data residency requirement because the model and data never leave the local region. SageMaker provides full control over the infrastructure, enabling compliance with local data sovereignty laws.

Exam trap

The trap here is that candidates assume Bedrock is the only AWS generative AI service, overlooking SageMaker's capability to host open-source models, which is a common misconception tested in the AIF-C01 exam.

How to eliminate wrong answers

Option B is wrong because waiting for Bedrock availability is unnecessary; SageMaker offers a viable alternative today. Option C is wrong because cross-region inference would send data outside the required country boundary, violating the data residency constraint. Option D is wrong because using a third-party API, even with AWS PrivateLink, still involves data leaving the AWS network to an external provider, which may not guarantee data remains within the country.

7
MCQeasy

A company is using Amazon Bedrock to generate marketing copy. They want to ensure the model's responses are factually accurate and grounded in their proprietary knowledge base. Which feature should they use?

A.Model customization
B.Fine-tuning
C.Retrieval Augmented Generation (RAG)
D.Prompt engineering
AnswerC

Retrieval Augmented Generation queries the proprietary knowledge base and injects relevant retrieved passages into the prompt, grounding responses in the company's own content. This satisfies the factual-accuracy requirement by supplying authoritative context the base model lacks, reducing hallucination.

Why this answer

Retrieval Augmented Generation (RAG) is the correct choice because it retrieves relevant documents from the company's proprietary knowledge base and provides them as context to the foundation model at inference time. This grounds the model's responses in factual, up-to-date information without modifying the underlying model weights, ensuring accuracy and reducing hallucinations.

Exam trap

A common misconception is that you must fine-tune or customize a model to incorporate proprietary knowledge. However, with Amazon Bedrock, RAG allows you to ground responses in your knowledge base without retraining, which is more cost-effective and keeps information current.

How to eliminate wrong answers

Option A is wrong because model customization (e.g., using Amazon Bedrock's Custom Model Import or training a new base model) changes the model's weights and behavior, but it does not inherently ground responses in a specific knowledge base; it requires large amounts of labeled data and can still hallucinate. Option B is wrong because fine-tuning adjusts model parameters on a dataset, which can embed knowledge but is static, expensive, and does not allow dynamic retrieval of proprietary information; it also risks overfitting or catastrophic forgetting. Option D is wrong because prompt engineering only modifies the input prompt to guide the model's output; it does not inject external factual data and cannot guarantee grounding in a proprietary knowledge base, as the model relies solely on its pre-trained parameters.

8
MCQhard

A data scientist fine-tuned a large language model on Amazon SageMaker for financial report generation. The model produces responses that are too short and incomplete, often cutting off mid-sentence. What parameter should be adjusted first?

A.Increase the temperature parameter
B.Increase the top_p parameter
C.Increase the maximum token count
D.Switch to a different foundation model
AnswerC

Truncated, mid-sentence output indicates generation stopped at the configured output limit rather than the model finishing naturally. Raising the maximum token count lets the model complete longer financial reports, directly addressing the premature cutoff constraint in the stem.

Why this answer

The max tokens parameter limits the length of generated responses. Increasing it allows the model to produce longer completions. Temperature, top_p, and model change affect quality or diversity, but not the length cap.

9
MCQhard

A research lab is using Amazon SageMaker to fine-tune a large language model (LLM) for scientific text summarization. The training dataset contains 10 million documents, and the lab has a limited budget but needs to minimize training time. They have access to SageMaker Training with managed spot instances, which offer significant cost savings but are interruptible. The team is considering different training strategies to balance cost, time, and model quality. Which strategy should they use?

A.Use SageMaker's distributed training with data parallelism on multiple managed spot instances, and enable checkpointing.
B.Fine-tune only the last few layers of the model on a smaller subset of the data.
C.Use a single on-demand instance to avoid interruptions and maximize throughput.
D.Use a single large GPU instance to train the model from scratch.
AnswerA

Data parallelism shards the 10-million-document dataset across multiple managed spot instances, cutting wall-clock training time while spot pricing reduces cost. Checkpointing persists progress so interrupted instances resume rather than restart, preserving model quality within budget.

Why this answer

The best strategy. Using SageMaker's distributed training with data parallelism across multiple managed spot instances allows parallel processing of the 10-million-document dataset, significantly reducing training time. Spot instances offer cost savings of up to 90% compared to on-demand, and enabling checkpointing ensures that if any instance is interrupted, training can resume from the last checkpoint without losing progress.

Option B is incorrect because fine-tuning only the last few layers on a subset of data compromises model quality and does not leverage the full dataset. Option C is incorrect because a single on-demand instance is expensive and slower than distributed spot instances. Option D is incorrect because training from scratch on a single GPU is prohibitively slow and costly, especially given the budget constraints.

10
MCQeasy

Which AWS service provides a serverless experience for building and scaling generative AI applications with access to various foundation models?

A.Amazon Bedrock
B.Amazon SageMaker
C.Amazon Lex
D.AWS Lambda
AnswerA

Amazon Bedrock delivers a fully managed, serverless API that exposes multiple foundation models from providers such as Anthropic, Meta and Amazon, removing infrastructure provisioning. This directly satisfies the stem's serverless requirement for building and scaling generative AI applications, since you invoke models on demand without managing servers or capacity.

Why this answer

Amazon Bedrock is a fully managed service that provides a serverless experience for building and scaling generative AI applications. It offers access to a variety of foundation models (FMs) from providers like AI21 Labs, Anthropic, Cohere, Meta, Stability AI, and Amazon via a single API, without the need to manage underlying infrastructure.

Exam trap

The trap here is that candidates may confuse Amazon SageMaker's broad ML capabilities with the specific serverless, foundation-model-focused offering of Amazon Bedrock, or mistakenly think AWS Lambda alone provides generative AI model access when it is merely a compute trigger.

How to eliminate wrong answers

Option B is wrong because Amazon SageMaker is a comprehensive machine learning (ML) platform that requires users to manage the entire ML lifecycle, including provisioning instances, training custom models, and deploying endpoints; it is not a serverless service specifically designed for accessing pre-built foundation models. Option C is wrong because Amazon Lex is a service for building conversational interfaces (chatbots) using automatic speech recognition (ASR) and natural language understanding (NLU), and it does not provide access to foundation models for generative AI tasks. Option D is wrong because AWS Lambda is a serverless compute service that runs code in response to events, but it does not natively provide access to foundation models or a managed API for generative AI; it can be used as part of a solution but is not the primary service for building and scaling generative AI applications with FMs.

11
MCQmedium

A financial services firm is deploying a generative AI assistant that answers employee questions about internal policies. Compliance requires that every response cite the exact policy document and section used. The assistant currently relies only on the foundation model's pretrained knowledge and frequently invents policy details. Which technique should the team implement to ground responses in the firm's own documents and produce citations?

A.Implement Retrieval Augmented Generation by embedding the policy documents in a vector store and retrieving relevant passages at query time.
B.Increase the model's temperature setting so it explores more of its pretrained knowledge about corporate policy.
C.Lower the maximum token limit so the assistant gives shorter answers that are less likely to contain errors.
D.Fine-tune the foundation model on a large corpus of public financial regulations.
AnswerA

Retrieval Augmented Generation embeds source documents, retrieves the passages most semantically similar to the user's question, and passes them into the prompt so the model answers from that supplied context. Because the retrieved chunks carry document and section metadata, the assistant can cite the exact source. It also keeps answers current as policies change without retraining.

Why this answer

Grounding a model in proprietary content requires supplying that content at inference time rather than relying on pretrained weights. Retrieval Augmented Generation retrieves the most relevant document chunks and places them in the prompt, so the model's answer is conditioned on real policy text and can reference the source document and section. Sampling or length adjustments do not add knowledge.

Exam trap

The trap here is believing that fine-tuning on domain text guarantees accurate citations, when only retrieval of the actual source documents provides verifiable provenance.

12
Multi-Selectmedium

A marketing team wants to build a generative AI application that produces short promotional copy and accompanying images for product launches. They are evaluating Amazon Bedrock and want to understand which statements accurately describe its capabilities for this use case. (Choose two.)

Select 2 answers
A.Amazon Bedrock requires customers to provision and manage the underlying GPU infrastructure for each model.
B.Amazon Bedrock supports both text generation models and image generation models.
C.Amazon Bedrock only supports models hosted in a single AWS Region worldwide.
D.Amazon Bedrock provides access to multiple foundation models from different providers through a single API.
E.Amazon Bedrock automatically guarantees that generated promotional claims are legally compliant.
AnswersB, D

Bedrock includes text models for copywriting and image models such as Amazon Titan Image Generator and Stability AI models for visuals. This means the team can generate promotional copy and accompanying images from the same service, satisfying the multi-modal requirement without stitching together separate platforms.

Why this answer

Amazon Bedrock offers a unified API across many foundation models and includes both text and image generation capabilities, which fits a promotional copy and imagery workload. It is serverless with respect to model hosting, is available in many Regions, and does not guarantee legal compliance, so the two accurate capability statements are the multi-provider single API and support for text and image models.

Exam trap

The trap here is conflating Bedrock's managed model access with infrastructure management or automatic compliance guarantees, neither of which Bedrock provides.

13
MCQeasy

A media company wants to build a generative AI assistant that drafts scripts and answers questions about its own style guide. The team has no machine learning engineers and wants to avoid managing GPU infrastructure or training any models. Which approach BEST describes how they should build this solution?

A.Build a rules-based template engine that reassembles stored sentences from previous scripts.
B.Deploy an Amazon SageMaker endpoint hosting an open-source model and manually patch the underlying EC2 instances.
C.Collect a labeled dataset and train a new transformer from scratch on Amazon SageMaker training jobs.
D.Use a foundation model through Amazon Bedrock and supply the style guide as context at inference time.
AnswerD

Foundation models accessed through Amazon Bedrock are pretrained, so no training or GPU fleet management is required, and the style guide can be supplied as retrieved context in the prompt. This matches the goal of a no-ML-team, serverless generative AI solution while still grounding responses in proprietary content.

Why this answer

A pretrained foundation model delivered through a managed service removes the need to train models or operate GPU infrastructure, and the proprietary style guide can be injected as prompt context so outputs follow house rules. This combination satisfies the no-ML-team constraint while still producing grounded, task-specific generative output.

Exam trap

The trap here is assuming that any generative AI use case requires training or fine-tuning a model, when pretrained foundation models plus prompt context are usually sufficient.

14
Multi-Selectmedium

Which TWO AWS services can be used to build a chatbot that responds to customer inquiries using a company's documentation as source? (Select two.)

Select 2 answers
A.Amazon Bedrock with RAG
B.Amazon Polly
C.Amazon Q Business
D.Amazon Transcribe
E.Amazon Lex
AnswersA, C

Amazon Bedrock supplies foundation models, while Retrieval Augmented Generation grounds responses in the company's own documentation by retrieving relevant passages and passing them to the model, satisfying the requirement to answer inquiries from that source rather than relying on pretrained knowledge alone.

Why this answer

Amazon Bedrock with RAG (A) is correct because it lets you build generative AI chatbots that retrieve relevant passages from your company documentation (stored in services like Amazon S3 or OpenSearch) and pass them to a foundation model as context, so answers are grounded in the source material. Amazon Q Business (C) is correct because it is a fully managed generative AI assistant that natively connects to enterprise data sources (S3, SharePoint, Confluence, etc.), indexes the documentation, and answers customer inquiries with citations from that content. Amazon Polly (B) is only a text-to-speech service and cannot retrieve or reason over documentation, Amazon Transcribe (D) only converts speech to text, and Amazon Lex (E) builds conversational interfaces/intents but does not itself perform retrieval-augmented generation over a company's documentation.

Exam trap

AWS often tests the distinction between services that provide conversational interfaces (like Lex) versus those that enable retrieval-augmented generation from custom data sources (like Bedrock with RAG or Q Business), leading candidates to mistakenly select Lex because it is a chatbot service, even though it lacks native document retrieval capabilities.

15
MCQeasy

A developer runs this AWS CLI command to invoke a model in us-west-2 but receives an error: 'An error occurred (ModelNotFoundException) when calling the InvokeModel operation: Model not found'. What is the most likely cause?

A.The request body is not properly formatted
B.The --region parameter is missing from the command
C.The model is not available in the us-west-2 region
D.The user's IAM role lacks permissions to invoke the model
AnswerC

Model availability is region-specific in Amazon Bedrock; model IDs valid in one region return ModelNotFoundException elsewhere. Since the command targets us-west-2, the requested model simply is not offered there, so invoking it fails regardless of credentials or permissions.

Why this answer

The error 'ModelNotFoundException' specifically indicates that the model ID or ARN is not recognized in the specified region. AWS Bedrock models are region-specific; not all foundation models are available in every AWS region. The developer invoked the model in us-west-2, but the requested model may only be available in regions like us-east-1 or us-west-1, causing the service to return a ModelNotFoundException rather than a permissions or formatting error.

Exam trap

Candidates often confuse the ModelNotFoundException with permission errors (AccessDeniedException). In AWS Bedrock, this error indicates the model is not available in the specified region, not a lack of IAM permissions. Always check regional model availability before assuming a permissions issue.

How to eliminate wrong answers

Option A is wrong because a malformed request body would result in a ValidationException or MalformedRequestBody error, not a ModelNotFoundException. Option B is wrong because if the --region parameter were missing, the CLI would use the default region from the AWS config or environment variables; if no default region were set, the CLI would return a 'You must specify a region' error, not a ModelNotFoundException. Option D is wrong because insufficient IAM permissions would produce an AccessDeniedException or UnauthorizedOperation error, not a ModelNotFoundException.

16
MCQmedium

A developer is using the Amazon Bedrock InvokeModel API with the above request to summarize meeting notes. The response is a single word repeated many times. Which parameter is MOST likely causing this issue?

A.topP set to 0.9
B.stopSequences is empty
C.maxTokenCount set to 100
D.temperature set to 0
AnswerD

Temperature at 0 makes the model greedy, always picking the highest-probability token, so a repetitive loop can dominate the output. It satisfies the stem's constraint by explaining the degenerate single-word repetition, though top-p or top-k sampling would restore diversity.

Why this answer

A temperature of 0 forces the model to always select the highest-probability token at each step, which can lead to repetitive loops if the most likely token repeatedly points back to itself (e.g., the same word). This deterministic behavior eliminates randomness, causing the model to get stuck in a single-word cycle rather than generating diverse or coherent text.

Exam trap

AWS often tests the misconception that temperature only affects 'creativity' or 'randomness,' when in fact a temperature of 0 causes deterministic argmax selection, which can paradoxically produce repetitive or stuck outputs rather than simply 'less creative' text.

How to eliminate wrong answers

Option A is wrong because topP set to 0.9 (nucleus sampling) actually increases diversity by considering tokens whose cumulative probability reaches 0.9, which would reduce repetition, not cause it. Option B is wrong because an empty stopSequences list means no custom stopping conditions are applied, but this does not force repetition; the model would still generate until a natural stop (e.g., EOS token) or maxTokenCount is reached. Option C is wrong because maxTokenCount set to 100 only limits the total number of tokens generated; it does not influence token selection probability or cause a single word to repeat—it would simply stop after 100 tokens regardless of content.

17
Multi-Selecthard

A media company is evaluating foundation models for a generative AI application that produces image captions and short video summaries. The team must balance output quality, latency, and operational cost. Which TWO considerations are most important when selecting a foundation model for this multimodal task? (Choose two.)

Select 2 answers
A.Whether the model can generate outputs in multiple human languages simultaneously.
B.Whether the model was trained using a specific programming language's libraries.
C.Whether the model supports the required input modalities, such as images and video frames.
D.Whether the model's inference latency and cost align with the application's throughput requirements.
E.Whether the model's license permits the company's intended commercial distribution of outputs.
AnswersC, D

The task requires processing images and video frames, so the model must accept those modalities as input. A text-only model cannot caption or summarize visual content regardless of its quality. Verifying modality support is therefore a fundamental selection criterion, ensuring the chosen model can actually ingest the required data types before any other trade-off is evaluated.

Why this answer

For a multimodal captioning and summarization workload, the model must accept images and video frames as input, so modality support is a prerequisite. Because the team must also balance latency and cost, evaluating inference speed and pricing against throughput needs is equally critical. Together these ensure the model can handle the data and remain practical to operate at scale.

Exam trap

The trap here is gravitating toward tangential factors like licensing or language coverage while overlooking that modality support and performance economics are the decisive selection criteria.

18
MCQmedium

A company wants to personalize its generative AI model for its specific domain without sharing data with third-party model providers. Which method should they use?

A.Fine-tuning the foundation model on their proprietary data
B.Prompt engineering with domain-specific examples
C.Retrieval-augmented generation (RAG) with a domain-specific knowledge base
D.Model distillation using a larger foundation model
AnswerA

Fine-tuning adjusts a foundation model's weights using the company's proprietary domain data, and the resulting custom model remains within the organisation's Azure AI resource. No training data is shared with the third-party provider, satisfying the data-confidentiality constraint in the stem.

Why this answer

Fine-tuning the foundation model on proprietary data allows the company to adapt the model's weights to its specific domain without sharing data with third parties. This method trains the model on private datasets, enabling it to learn domain-specific patterns and terminology while keeping data in-house, which is critical for data privacy and compliance.

Exam trap

AWS often tests the distinction between methods that modify model parameters (fine-tuning) versus those that only augment input or retrieval (prompt engineering, RAG), leading candidates to mistakenly choose RAG for personalization when fine-tuning is required for deep domain adaptation.

How to eliminate wrong answers

Option B is wrong because prompt engineering with domain-specific examples does not modify the model's weights; it only influences output through input prompts, which cannot achieve the same depth of domain adaptation as fine-tuning and still relies on the base model's knowledge. Option C is wrong because retrieval-augmented generation (RAG) with a domain-specific knowledge base retrieves external information at inference time but does not train the model on proprietary data, so the model itself remains unchanged and may not fully internalize domain nuances. Option D is wrong because model distillation compresses a larger model into a smaller one for efficiency, but it does not involve training on proprietary domain data and does not address the requirement of personalization without sharing data.

19
MCQmedium

A media company is using Amazon Bedrock to generate captions for images. They have a batch processing pipeline that sends thousands of images daily to the Bedrock API using the Titan Image Generator G1 model. Recently, they started receiving ThrottlingException errors during peak hours. The team needs to process all images within 24 hours without changing the model or the application code. The current account has a default quota of 10 requests per second (RPS) for the Titan model in us-east-1. The team estimates they need 50 RPS during peak hours. They have already implemented exponential backoff in the client, but the errors persist. What is the MOST effective solution to resolve the throttling issue?

A.Request a service quota increase for the InvokeModel API for the Titan model in us-east-1
B.Use Amazon SageMaker batch transform to process images offline
C.Distribute the requests across multiple AWS Regions
D.Switch to a different foundation model that has a higher default quota
AnswerA

Raising the InvokeModel quota for Titan in us-east-1 lifts the 10 RPS ceiling to the required 50 RPS, clearing ThrottlingException at source. Exponential backoff cannot exceed a hard per-account, per-region, per-model limit, and the stem forbids changing model or code.

Why this answer

The team has already implemented exponential backoff, but the errors persist because their current quota of 10 RPS is insufficient for the required 50 RPS. Requesting a service quota increase for the InvokeModel API for the Titan Image Generator G1 model in us-east-1 directly addresses the root cause by raising the throughput limit, allowing the existing application code and model to handle the peak load without any architectural changes.

Exam trap

The trap here is that candidates may think exponential backoff or distributing across Regions solves all throttling, but the core issue is a hard service quota that must be increased, not a transient rate limit.

How to eliminate wrong answers

Option B is wrong because Amazon SageMaker batch transform is designed for offline inference on SageMaker endpoints, not for invoking Bedrock APIs; it would require changing the application code and infrastructure, which the question explicitly prohibits. Option C is wrong because distributing requests across multiple AWS Regions would require modifying the application code to route traffic to different endpoints, and it does not address the underlying quota issue in the primary region; it also introduces latency and complexity. Option D is wrong because switching to a different foundation model would require changing the application code and potentially the image generation logic, which is not allowed; moreover, other models may have different default quotas or capabilities, and the goal is to process images with the Titan model.

20
MCQeasy

A developer is building an application that generates product descriptions from images using a multimodal model. Which AWS service provides access to multimodal foundation models?

A.Amazon Rekognition
B.Amazon Textract
C.Amazon Comprehend
D.Amazon Bedrock
AnswerD

Amazon Bedrock supplies managed access to multimodal foundation models, including Amazon Titan and third-party options, through a single API. It satisfies the stem's requirement for generating text from images without infrastructure management, unlike services limited to text-only models or custom training workflows.

Why this answer

Amazon Bedrock is a managed service that provides access to a wide range of foundation models (FMs) from leading AI providers, including multimodal models that can process both images and text to generate product descriptions. This makes it the correct choice for building an application that requires multimodal capabilities.

Exam trap

The trap here is that candidates may confuse AWS AI services that handle specific modalities (Rekognition for images, Comprehend for text) with Bedrock, which is the only service that provides access to generative multimodal foundation models capable of combining both modalities in a single inference.

How to eliminate wrong answers

Option A is wrong because Amazon Rekognition is a computer vision service for image and video analysis (e.g., object detection, facial recognition), but it does not provide access to generative multimodal foundation models. Option B is wrong because Amazon Textract is an OCR service that extracts text from documents, not a platform for accessing or running multimodal generative models. Option C is wrong because Amazon Comprehend is a natural language processing (NLP) service for text analysis (e.g., sentiment, entities), and it lacks support for multimodal input or generative model access.

21
MCQmedium

An application uses this configuration to enable RAG. What is required for the knowledge base to function?

A.The agent must have internet access to retrieve documents
B.The embedding model ARN must include the account ID
C.The embedding model must be fine-tuned on the domain data
D.The knowledge base must have a vector index configured in Amazon OpenSearch Serverless
AnswerD

RAG retrieval depends on semantic similarity search, which requires vector embeddings stored in an index. Amazon OpenSearch Serverless provides that vector index, so the knowledge base can match queries to relevant document chunks before passing them to the model.

Why this answer

For a knowledge base to function in a RAG (Retrieval-Augmented Generation) setup on AWS, the knowledge base must have a vector index configured in Amazon OpenSearch Serverless. This vector index stores the embeddings generated from the source documents, enabling efficient similarity search to retrieve relevant context for the agent. Without a vector index, the knowledge base cannot perform the vector search required to fetch relevant document chunks.

Exam trap

The trap here is that candidates may think the embedding model must be fine-tuned or that internet access is needed, but the core requirement is the vector index in a vector store like Amazon OpenSearch Serverless on AWS, which is essential for the retrieval step in RAG.

How to eliminate wrong answers

Option A is wrong because the agent does not need internet access to retrieve documents; the knowledge base is hosted within AWS and accessed via internal API calls, not over the public internet. Option B is wrong because the embedding model ARN does not need to include the account ID; ARNs already contain the account ID by default, and the requirement is that the model must be accessible (e.g., via Amazon Bedrock) and the ARN must be correctly specified, but the account ID is not an additional requirement. Option C is wrong because the embedding model does not need to be fine-tuned on domain data; a pre-trained embedding model (e.g., from Amazon Bedrock) can generate embeddings for any domain, and fine-tuning is not a prerequisite for RAG functionality.

22
MCQmedium

A company deployed a chatbot using Amazon Lex integrated with a Lambda function that invokes Claude on Amazon Bedrock. The Lambda function retrieves relevant documents from an Amazon Kendra index to use as context. Users report that the chatbot's responses are often irrelevant or incorrect despite the Kendra index containing accurate information. The logs show that the Lambda function is correctly passing retrieved documents to the model. What is the most likely cause and solution?

A.Switch to a larger foundation model like Claude 3 Opus
B.The model's temperature is set too high; reduce it to 0.1
C.The maximum tokens limit is too low; increase it to 4096
D.The chunking strategy for documents is too coarse or inappropriate; refine chunking and use semantic search in Kendra
AnswerD

Coarse chunking embeds large blocks, so Kendra returns passages whose vectors dilute the specific answer, and the model receives loosely relevant context despite accurate documents. Refining chunk size and enabling semantic search sharpens retrieval precision, satisfying the requirement that passed context actually match the user's question.

Why this answer

When retrieved documents are correctly passed to the model but responses are still irrelevant, the problem is almost always upstream in retrieval quality — specifically how documents were chunked and indexed in Kendra. Coarse or poorly aligned chunks dilute the semantic signal, so the model receives context that does not actually answer the query. Refining chunking strategy and enabling semantic search in Kendra directly addresses the root cause.

Exam trap

AIF-C01 often tests the misconception that a bigger or more expensive model fixes RAG quality — candidates must recognize that retrieval (chunking, indexing, semantic search) is the usual bottleneck when context is already being passed.

How to eliminate wrong answers

Option A is wrong because switching to a larger model like Claude 3 Opus does not fix bad retrieval — a bigger model given irrelevant context will still produce irrelevant answers, and it increases cost and latency without addressing the root cause. Option B is wrong because temperature controls randomness/creativity, not factual grounding; a high temperature might cause variability but the logs show documents are being passed correctly, so the issue is retrieval relevance, not sampling. Option C is wrong because max tokens limits response length, not relevance — increasing it just allows longer (still wrong) answers and does not improve the quality of retrieved context.

23
Multi-Selectmedium

A company is building a generative AI application using Amazon Bedrock and needs to ensure that the model does not generate outputs containing personally identifiable information (PII). Which TWO actions should the company take? (Choose 2)

Select 2 answers
A.Implement a custom AWS Lambda function to scan and redact PII from inputs and outputs.
B.Use AWS Identity and Access Management (IAM) policies to restrict model access.
C.Enable Amazon CloudWatch Logs to capture and audit model outputs.
D.Configure Amazon Bedrock Guardrails to block or mask PII.
E.Place the Bedrock model endpoint within a private VPC.
AnswersA, D

A Lambda function provides deterministic, customisable inspection and redaction of PII across both request and response payloads, catching patterns Guardrails filters may not cover. It satisfies the requirement to prevent PII appearing in generated outputs.

Why this answer

Option A is correct because a custom AWS Lambda function can programmatically scan both the prompt inputs and the model responses for PII patterns (for example, using regex or Amazon Comprehend's PII detection APIs) and redact or block them before they reach the user, giving the company full control over what data enters and leaves the generative AI application. Option D is correct because Amazon Bedrock Guardrails provides a native, purpose-built sensitive information filter that can detect and either block or mask PII entities (such as names, emails, and SSNs) in both user inputs and model outputs, directly satisfying the requirement to prevent PII from appearing in generated content. Option B is not correct because IAM policies only control who can invoke which Bedrock models and actions; they do not inspect or filter the content of prompts and responses for PII.

Option C is not correct because CloudWatch Logs only captures and stores model outputs for auditing and monitoring after the fact; it does not prevent PII from being generated or returned. Option E is not correct because placing the Bedrock endpoint in a private VPC only affects network isolation and traffic routing, not the content-level detection or redaction of PII in model outputs.

Exam trap

The AIF-C01 exam often tests the distinction between network-level security controls (like VPCs) and content-level data protection mechanisms, leading candidates to mistakenly choose VPC isolation as a solution for PII redaction.

24
MCQhard

A company wants to use a large language model to generate code based on natural language descriptions. They need to minimize latency and control costs by running inference on their own infrastructure. Which approach is most suitable?

A.Use Amazon Bedrock API
B.Use Amazon SageMaker to deploy a custom LLM
C.Use Amazon Comprehend
D.Use Amazon Lex
AnswerB

SageMaker deploys the model on infrastructure the company controls, so inference runs locally rather than through a third-party API. This satisfies both constraints: reduced network latency and predictable, controllable cost, unlike fully managed model endpoints.

Why this answer

Amazon SageMaker allows you to deploy a custom large language model (LLM) on your own infrastructure, giving you full control over inference latency and cost. By using SageMaker endpoints with auto-scaling and instance selection, you can optimize for low-latency responses while avoiding per-token API charges from managed services.

Exam trap

AWS often tests the distinction between managed API services (like Bedrock) and self-managed deployment options (like SageMaker), where candidates mistakenly choose Bedrock for 'control' over costs and latency, not realizing that Bedrock is a pay-per-token managed service with no infrastructure control.

How to eliminate wrong answers

Option A is wrong because Amazon Bedrock is a managed API service that charges per-token and does not allow you to run inference on your own infrastructure, so you cannot control latency or costs at the infrastructure level. Option C is wrong because Amazon Comprehend is a natural language processing (NLP) service for tasks like sentiment analysis and entity extraction, not a generative AI service capable of code generation from natural language. Option D is wrong because Amazon Lex is designed for building conversational chatbots using intent-based models, not for deploying large language models for code generation.

25
MCQeasy

A startup wants to quickly prototype a generative AI application for summarizing news articles. They have limited ML expertise and want minimal infrastructure management. Which AWS service should they use?

A.Amazon Bedrock with a foundation model accessed via API.
B.Amazon SageMaker to build and train a custom summarization model.
C.AWS Lambda with a custom Python script using the Hugging Face Transformers library.
D.Amazon EC2 instance running a pre-trained model from AWS Marketplace.
AnswerA

Amazon Bedrock provides serverless API access to foundation models, eliminating infrastructure management and ML expertise requirements. This satisfies the startup's constraints of minimal infrastructure and limited ML skills, enabling rapid prototyping of news summarisation without training or hosting models.

Why this answer

Amazon Bedrock is the correct choice because it provides pre-trained foundation models from leading AI providers via a simple API, requiring no ML expertise or infrastructure management. The startup can quickly prototype a summarization application by sending news articles to the API and receiving summaries without training or deploying models.

Exam trap

AWS often tests the distinction between managed AI services (Bedrock) and infrastructure-heavy services (SageMaker, EC2, Lambda), where candidates mistakenly choose SageMaker for its flexibility or Lambda for its serverless nature, overlooking the specific requirement for minimal ML expertise and infrastructure management.

How to eliminate wrong answers

Option B is wrong because Amazon SageMaker is designed for building, training, and deploying custom ML models, which requires significant ML expertise and infrastructure management, contradicting the startup's need for minimal ML expertise and quick prototyping. Option C is wrong because AWS Lambda with a custom Python script using Hugging Face Transformers requires managing dependencies, cold start latency, and model loading within Lambda's memory and time limits, adding complexity and infrastructure management. Option D is wrong because Amazon EC2 requires provisioning, configuring, and managing the instance, including installing the model and handling scaling, which is not minimal infrastructure management and deviates from the goal of quick prototyping.

26
MCQmedium

A media company's generative AI writing assistant produces fluent but sometimes fabricated statistics. The team wants to reduce these hallucinations without changing the foundation model. Which action best addresses the root cause?

A.Raise the top-p sampling value to make token selection more deterministic
B.Lower the maximum token limit for each response
C.Switch the model to a smaller parameter count to reduce creativity
D.Ground the model's responses by supplying verified reference documents in the prompt and instructing it to cite them
AnswerD

Hallucinated statistics occur when the model generates plausible-sounding content without a factual anchor. Providing verified source documents in the prompt and requiring citations constrains generation to supported claims, reducing fabrication without modifying model weights, which matches the constraint of not changing the foundation model.

Why this answer

Fabricated statistics stem from the model generating text without an authoritative factual anchor. Supplying verified documents and requiring citations ties each claim to a source, which reduces invention while leaving the model untouched. Token limits, parameter count, and sampling parameters influence length, capacity, and randomness respectively, none of which ground output in verified facts.

Exam trap

The trap here is treating sampling parameters as accuracy controls when they only shape randomness and diversity of token selection.

27
Multi-Selectmedium

A company is deploying a generative AI model on Amazon Bedrock and needs to monitor for potential misuse. Which THREE measures should they implement? (Choose 3)

Select 3 answers
A.Require multi-factor authentication (MFA) for all API calls.
B.Configure Amazon Bedrock Guardrails to block harmful content.
C.Use AWS CloudTrail to log API calls and Amazon Bedrock actions.
D.Place the Bedrock endpoint in a private VPC with no internet access.
E.Enable model invocation logging in Amazon CloudWatch.
AnswersB, C, E

Amazon Bedrock Guardrails apply content filters that intercept harmful inputs and outputs at inference time, directly satisfying the requirement to monitor for potential misuse. By defining denied topics and filtering categories, the company enforces safety controls on the generative model itself, rather than relying on post-hoc logging or external review.

Why this answer

Option B is correct because Amazon Bedrock Guardrails are purpose-built to detect and block harmful or inappropriate content (e.g., hate, violence, prompt injection) in both prompts and model responses, directly addressing misuse monitoring and prevention. Option C is correct because AWS CloudTrail records all Bedrock API activity and management events, providing an auditable trail of who invoked which model, when, and from where, which is essential for detecting misuse. Option E is correct because enabling model invocation logging sends full request/response data to Amazon CloudWatch Logs, allowing continuous monitoring, alerting, and forensic analysis of model inputs and outputs.

Option A is not appropriate because MFA applies to human console sign-ins, not programmatic API calls, and Bedrock APIs use IAM credentials/signatures rather than MFA tokens. Option D is not required for misuse monitoring; a private VPC endpoint improves network isolation but does not detect or log misuse, and Bedrock endpoints are AWS-managed services accessed via VPC endpoints rather than being 'placed' in a VPC.

Exam trap

The AIF-C01 exam often tests the distinction between security controls that prevent access (like MFA or VPC isolation) versus monitoring controls that detect or block misuse at the content level, leading candidates to confuse network security with content safety.

28
MCQmedium

A company is building a chatbot using Amazon Bedrock and wants to ensure that the model generates responses consistent with its brand voice. Which technique should be used to provide the model with examples of desired responses without fine-tuning the model?

A.Fine-tune the model on a dataset of brand-compliant conversations.
B.Use prompt chaining to break down the conversation into multiple steps.
C.Implement a Retrieval Augmented Generation (RAG) system with brand documents.
D.Include few-shot examples in the system prompt to demonstrate the desired tone.
AnswerD

Few-shot examples in the system prompt steer Amazon Bedrock's model at inference time by conditioning it on demonstrations of the desired tone, satisfying the constraint of no fine-tuning. Unlike weight-updating approaches, this in-context prompting requires no training job, so brand-voice consistency is achieved immediately and cheaply.

Why this answer

Few-shot prompting allows you to provide the model with examples of desired responses directly in the system prompt, guiding the model's tone and style without modifying its underlying weights. This technique is ideal for brand voice consistency when fine-tuning is not an option, as it leverages in-context learning to influence output behavior.

Exam trap

AWS often tests the distinction between in-context learning (few-shot prompting) and fine-tuning, trapping candidates who confuse RAG (which retrieves facts) with style guidance, or who think prompt chaining is for tone control rather than task decomposition.

How to eliminate wrong answers

Option A is wrong because fine-tuning requires modifying the model's weights, which contradicts the requirement of not fine-tuning the model. Option B is wrong because prompt chaining is a technique for decomposing complex tasks into sequential steps, not for providing examples of desired tone or style. Option C is wrong because Retrieval Augmented Generation (RAG) retrieves external knowledge from documents to ground responses in facts, but it does not inherently teach the model the specific tone or brand voice; it augments context, not style.

29
MCQhard

A financial services company is subject to strict regulatory requirements. They plan to use generative AI to summarize customer interaction logs. Which combination of AWS services and configurations best ensures compliance while maintaining accuracy?

A.Deploy an open-source model on Amazon Bedrock in a local on-premises server.
B.Use Amazon Bedrock with a foundation model and public internet access without encryption.
C.Use Amazon SageMaker to host a fine-tuned model with a public API key.
D.Use Amazon Bedrock with a private VPC endpoint, AWS KMS encryption, and content filtering.
AnswerD

A private VPC endpoint keeps traffic off the public internet, KMS encryption protects data at rest and in transit, and content filtering blocks sensitive output, together meeting the strict regulatory constraints while Bedrock summarises the logs.

Why this answer

It combines a private VPC endpoint to keep all traffic within the AWS network (avoiding public internet exposure), AWS KMS encryption for data at rest and in transit, and content filtering to block sensitive or non-compliant outputs. This architecture meets strict regulatory requirements for data privacy and security while using Amazon Bedrock's managed foundation models for accurate summarization.

Exam trap

A common misconception is that encryption alone ensures compliance. However, the trap here is that public internet access (even with HTTPS) violates strict regulatory requirements that mandate private network connectivity (via VPC endpoints) and data residency controls.

How to eliminate wrong answers

Option A is wrong because deploying an open-source model on a local on-premises server does not use Amazon Bedrock (which is a fully managed AWS service) and introduces operational overhead, potential compliance gaps, and lacks AWS-native encryption and auditing. Option B is wrong because using public internet access without encryption exposes customer interaction logs to interception and violates regulatory mandates for data in transit security (e.g., TLS). Option C is wrong because using a public API key with Amazon SageMaker exposes the model endpoint to unauthorized access and lacks the private networking and encryption controls required for compliance.

30
MCQmedium

A healthcare startup is building a patient-education chatbot. They want the model to answer only using an approved set of clinical guideline documents, and they must be able to update those documents weekly without retraining any model. They plan to use Amazon Bedrock. Which approach best meets these requirements?

A.Fine-tune a foundation model in Amazon Bedrock each week on the updated guideline documents.
B.Increase the model's context window by raising the maxTokens parameter for every request.
C.Create a separate Bedrock agent for each clinical guideline document and route questions by keyword.
D.Use Retrieval Augmented Generation by storing the guidelines in a vector store and retrieving relevant passages at inference time.
AnswerD

Retrieval Augmented Generation retrieves relevant passages from an external knowledge base and supplies them to the model as context. Updating the vector store with new guideline documents refreshes the knowledge without any retraining, and grounding responses in the retrieved text improves adherence to the approved source material.

Why this answer

Retrieval Augmented Generation separates knowledge from model weights. The approved guidelines live in a vector store that can be updated weekly, and relevant passages are retrieved and passed to the model as context. This grounds answers in approved content and avoids the cost and delay of fine-tuning for every content change.

Exam trap

The trap here is treating fine-tuning as the default way to add new knowledge, when retrieval is the mechanism that supports frequent updates without retraining.

31
MCQmedium

A solutions architect is explaining why a foundation model can answer questions about topics it was never explicitly programmed for. Which characteristic of generative AI best explains this behaviour?

A.The model retrieves live answers from a search index at inference time.
B.The model re-trains itself on each user prompt before producing a response.
C.The model stores every training example verbatim and looks up the closest match when prompted.
D.The model learned statistical patterns and relationships from large-scale training data, allowing it to generalize to new prompts.
AnswerD

Foundation models are trained on very large corpora and encode statistical relationships between tokens, which lets them produce coherent responses to prompts they never saw during training. Generalization comes from those learned patterns rather than hard-coded rules, so the model can handle novel questions within the distribution of its training. This is the defining behaviour of generative AI.

Why this answer

Generative models generalize because training over massive datasets encodes statistical patterns and relationships among tokens into the model weights. At inference the prompt is processed against those learned parameters, enabling coherent responses to inputs never seen in training. Retrieval, verbatim lookup, and per-prompt retraining all describe different mechanisms that do not explain inherent generalization.

Exam trap

The trap here is conflating retrieval-augmented generation, which is an optional architecture, with the intrinsic generalization ability of a foundation model.

32
MCQmedium

A retail company wants to generate product descriptions from catalog data. The data includes structured attributes (e.g., price, brand) and unstructured reviews. The team needs to ensure factual accuracy. Which approach is most appropriate?

A.Use prompt engineering with few-shot examples
B.Fine-tune a foundation model on the entire product catalog
C.Deploy a larger foundation model with more parameters
D.Implement Retrieval-Augmented Generation (RAG) with a knowledge base
AnswerD

RAG grounds generation in retrieved catalogue records, so structured attributes such as price and brand are injected into the prompt rather than recalled from model weights. This satisfies the factual-accuracy constraint by anchoring outputs to source data, while unstructured reviews supply tone and descriptive language.

Why this answer

Retrieval-Augmented Generation (RAG) retrieves relevant documents (product attributes, reviews) and provides them as context to the model, reducing hallucinations and grounding responses in facts.

33
MCQeasy

What is a foundation model?

A.A model that only works with tabular data
B.A model that requires no additional tuning for new tasks
C.A model trained on diverse data that can be adapted to many tasks
D.A model that is specifically trained for one task, like image classification
AnswerC

Foundation models are trained on broad, diverse datasets using self-supervision, producing general-purpose representations that transfer to many downstream tasks through fine-tuning or prompting. This satisfies the stem's requirement for a base model adaptable across domains, distinguishing it from narrow task-specific models trained on single-purpose labelled data.

Why this answer

A foundation model is a large-scale AI model trained on vast, diverse datasets (e.g., text, images, code) using self-supervised learning, enabling it to be adapted to a wide range of downstream tasks through fine-tuning or few-shot learning. Option C correctly captures this core property of broad adaptability, which distinguishes foundation models from task-specific models. For example, GPT-4 and Claude are foundation models that can handle translation, summarization, and coding without being retrained from scratch.

Exam trap

AWS often tests the misconception that foundation models require no tuning at all (Option B), but the correct understanding is that they are adaptable—not that they are immediately perfect for every task without any adjustment.

How to eliminate wrong answers

Option A is wrong because foundation models are not limited to tabular data; they are typically trained on unstructured data like text, images, and audio. Option B is wrong because while foundation models can perform zero-shot tasks, they often require additional tuning (e.g., fine-tuning or prompt engineering) to achieve optimal performance on specific tasks, contradicting the claim of 'no additional tuning.' Option D is wrong because foundation models are explicitly designed to be multi-task and adaptable, not trained for a single task like image classification.

34
MCQmedium

A company deployed a question-answering system using Amazon Bedrock with a knowledge base (RAG). Users report that the model often hallucinates facts not in the knowledge base. What is the most effective way to reduce hallucinations?

A.Reduce the maximum context length to limit model input
B.Fine-tune the foundation model on a large general corpus
C.Improve the relevance of retrieved documents by refining the retrieval strategy
D.Increase the chunk size of documents in the knowledge base
AnswerC

Refining retrieval directly targets the RAG constraint: hallucinations arise when retrieved context lacks the facts needed, so the model fabricates answers. Improving retrieval relevance—through better chunking, embeddings or hybrid search—ensures grounding documents actually contain the answer, reducing fabrication without changing the foundation model itself.

Why this answer

Hallucinations in RAG systems often stem from the model receiving irrelevant or low-quality retrieved documents, which forces it to rely on its parametric knowledge rather than the provided context. By refining the retrieval strategy—such as improving embedding quality, adjusting chunk overlap, or using hybrid search—the system ensures the foundation model has the most relevant information to ground its answers, directly reducing the likelihood of fabricating facts.

Exam trap

A common misconception is that hallucinations are primarily a model training issue (fine-tuning or context length) rather than a retrieval quality issue in RAG systems, leading candidates to overlook the critical role of the retriever in grounding responses.

How to eliminate wrong answers

Option A is wrong because reducing the maximum context length limits the amount of retrieved context the model can use, which actually increases the risk of hallucinations by forcing the model to rely more on its own training data rather than the knowledge base. Option B is wrong because fine-tuning on a large general corpus would further embed general knowledge into the model, potentially exacerbating hallucinations when the model defaults to its training data instead of the knowledge base; fine-tuning is not a targeted fix for retrieval quality. Option D is wrong because increasing chunk size can lead to chunks that contain irrelevant or noisy information, reducing the precision of retrieval and potentially introducing more irrelevant context that confuses the model, rather than improving answer accuracy.

35
Multi-Selectmedium

A company is deploying a customer-facing chatbot using Amazon Bedrock. They want to reduce the risk of generating biased or harmful responses. Which TWO measures should they implement? (Choose 2.)

Select 2 answers
A.Implement a human-in-the-loop review for sensitive replies
B.Train the model exclusively on historical customer conversations
C.Use guardrails to filter content
D.Set the temperature parameter to 1.5
E.Disable logging to improve performance
AnswersA, C

Human-in-the-loop review catches biased or harmful outputs that automated guardrails miss, satisfying the requirement to reduce harmful responses in a customer-facing chatbot. Reviewers assess sensitive replies before they reach users, providing a feedback loop that can also refine prompts and filters over time.

Why this answer

Option A is correct because implementing a human-in-the-loop review for sensitive replies adds a manual verification layer that can catch biased or harmful outputs before they reach customers, which is a recommended responsible-AI practice for customer-facing generative applications. Option C is correct because Amazon Bedrock Guardrails lets you configure content filters (for hate, insults, sexual, violence, misconduct), denied topics, word filters, and contextual grounding checks that automatically block or mask harmful or biased responses at inference time. Option B is not appropriate because training exclusively on historical customer conversations can perpetuate existing biases and does not by itself mitigate harmful output.

Option D is incorrect because setting temperature to 1.5 increases randomness and creativity, making outputs less predictable and potentially more harmful. Option E is incorrect because disabling logging reduces observability and auditability, which undermines monitoring and continuous improvement of safety controls.

Exam trap

A common misconception when using Amazon Bedrock is that increasing the temperature parameter or training on raw historical data alone can improve safety, when in fact these actions increase risk or reduce oversight.

36
MCQmedium

A startup is building an AI-powered code assistant using a large language model (LLM). They want to ensure the model generates syntactically correct code and avoids security vulnerabilities. Which technique should they prioritize?

A.Augment prompts with few-shot examples of secure coding practices and unit tests
B.Deploy the model with max tokens set to 4096
C.Fine-tune the model on a large corpus of open-source code
D.Use chain-of-thought prompting to explain reasoning before code generation
AnswerA

Few-shot examples steer the LLM toward secure coding patterns and unit-test structure within the prompt itself, requiring no retraining or architectural change. This directly satisfies the startup's constraints of syntactic correctness and vulnerability avoidance, since the model conditions on concrete secure snippets rather than relying on its base pretraining alone.

Why this answer

Few-shot prompting with examples of secure coding practices and unit tests directly shapes the model's output distribution toward syntactically valid, security-conscious code by conditioning on concrete in-context demonstrations. This is a prompt-engineering technique that requires no retraining and can be updated instantly as new vulnerability patterns emerge. It simultaneously addresses both stated goals: syntactic correctness (via code examples) and security (via secure-coding exemplars).

Exam trap

AIF-C01 often tests the misconception that model configuration knobs (max tokens, temperature) or fine-tuning are the primary levers for output quality, when prompt engineering with targeted examples is usually the fastest, most controllable fix.

How to eliminate wrong answers

Option B is wrong because max tokens only caps output length and has no bearing on code correctness or security — a 4096-token limit could even truncate code mid-function. Option C is wrong because fine-tuning on generic open-source code often propagates the very vulnerabilities (e.g., hardcoded secrets, SQL injection) present in public repositories, and it is costly and slow to iterate. Option D is wrong because chain-of-thought improves reasoning transparency but does not guarantee syntactic validity or security; the model can reason correctly and still emit insecure code.

37
MCQmedium

A financial services company wants to generate personalized investment recommendations using a large language model via Amazon Bedrock. They have customer data that includes risk tolerance, portfolio holdings, and financial goals. The company is highly concerned about data privacy and must avoid exposing sensitive personally identifiable information (PII) to the model. They plan to use a foundation model to generate recommendations based on customer profiles. What is the best approach to protect customer privacy while still enabling personalization?

A.Fine-tune the model on a large dataset of investment recommendations without any customer-specific data.
B.Use prompt engineering to instruct the model to disregard any personally identifiable information.
C.Preprocess the customer data to replace sensitive fields with placeholders, then use the processed data in the prompt.
D.Include the customer data directly in the prompt and rely on the model to anonymize it.
AnswerC

Replacing sensitive fields with placeholders removes PII before any prompt reaches the foundation model, so personal identifiers never leave the company's environment. The model still receives behavioural context such as risk tolerance and goals, preserving personalisation while satisfying the privacy constraint.

Why this answer

Preprocessing customer data to replace sensitive fields with placeholders (e.g., using synthetic IDs) allows the model to generate personalized recommendations without accessing real PII. This minimizes risk. Option A is incorrect because fine-tuning on a large dataset of generic recommendations does not produce personalized outputs for individual customers.

Option B is incorrect because prompt engineering instructions are not a robust privacy control and cannot reliably prevent PII exposure. Option D is incorrect because including customer data directly in the prompt and relying on the model to anonymize it is unreliable and may still leak PII.

38
MCQhard

A bank is using Amazon Bedrock to summarize customer support transcripts. The summaries often contain factual inaccuracies (hallucinations). Which approach is most effective for reducing hallucinations?

A.Decrease the top-p to 0.1
B.Increase the model's temperature to make outputs more diverse
C.Fine-tune a smaller model on a large dataset of transcripts
D.Implement RAG by grounding summarization on retrieved transcripts
AnswerD

Grounding summaries in retrieved transcript passages constrains generation to source text, directly countering the factual drift that causes hallucinations. Retrieval narrows the model's context to verified customer dialogue, so fabricated details lack support and are suppressed. This satisfies the stem's core requirement: reducing inaccuracies in Bedrock summarisation without retraining.

Why this answer

Retrieval-Augmented Generation (RAG) grounds the model's output on actual retrieved chunks of the customer support transcripts, providing factual context that reduces the likelihood of hallucination. By retrieving relevant transcript segments and feeding them as context to the LLM, the model generates summaries based on verified source material rather than relying solely on its parametric knowledge, which is the primary cause of factual inaccuracies.

Exam trap

AWS often tests the misconception that adjusting sampling parameters (top-p, temperature) or fine-tuning alone can fix hallucinations, when in fact these methods do not provide factual grounding and RAG is the standard industry approach for reducing factual inaccuracies in generative AI.

How to eliminate wrong answers

Option A is wrong because decreasing top-p to 0.1 reduces the nucleus sampling pool to only the most likely tokens, which can actually increase repetition and factual errors by making the model overly deterministic and less able to select correct but less probable tokens. Option B is wrong because increasing temperature makes outputs more random and diverse, which typically increases hallucination risk rather than reducing it, as the model is more likely to generate plausible-sounding but incorrect content. Option C is wrong because fine-tuning a smaller model on a large dataset of transcripts may improve domain adaptation but does not directly address hallucinations; smaller models have less capacity to memorize factual details, and fine-tuning can still produce hallucinations when the model encounters queries outside its training distribution or when it must generalize beyond exact training examples.

39
MCQmedium

A machine learning engineer notices that a generative AI model occasionally produces biased outputs. Which AWS feature can automatically filter harmful content before it reaches users?

A.Amazon CloudWatch alarms
B.Amazon SageMaker Clarify
C.AWS Identity and Access Management (IAM) policies
D.Amazon Bedrock Guardrails
AnswerD

Amazon Bedrock Guardrails applies configurable content filters and denied-topic policies to model inputs and outputs, blocking harmful content before it reaches users. This automated intervention satisfies the requirement to filter biased or harmful generative output without custom moderation code.

Why this answer

Amazon Bedrock Guardrails is specifically designed to implement safeguards for generative AI applications, including the ability to filter harmful, biased, or inappropriate content before it reaches users. It allows you to define denied topics, content filters, and sensitive information filters that are applied at inference time, directly addressing the need to automatically filter biased outputs from a generative AI model.

Exam trap

The trap here is that candidates may confuse Amazon SageMaker Clarify (which detects bias in training data or model predictions) with a real-time content filtering solution, but Clarify is a static analysis tool, not a runtime guardrail for generative AI outputs.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch alarms are used for monitoring metrics and triggering notifications based on thresholds, not for filtering or modifying the content of AI model outputs. Option B is wrong because Amazon SageMaker Clarify is a tool for detecting bias in machine learning models and data during development and training, not for real-time filtering of generated content in a production generative AI application. Option C is wrong because AWS Identity and Access Management (IAM) policies control permissions and access to AWS resources, not the content or safety of outputs from a generative AI model.

40
MCQhard

A machine learning team is fine-tuning a foundation model using Amazon SageMaker. They need to optimize training time and cost. Which approach should they take?

A.Use a larger instance type with more vCPUs
B.Increase the batch size to the maximum possible
C.Use the full model weights and train on a single GPU
D.Use Parameter-Efficient Fine-Tuning (PEFT) techniques like LoRA
AnswerD

LoRA freezes the base model weights and trains small low-rank adapter matrices instead, drastically reducing the number of trainable parameters. This lowers GPU memory and compute requirements, cutting both training time and cost compared with full fine-tuning.

Why this answer

Parameter-Efficient Fine-Tuning (PEFT) techniques like LoRA (Low-Rank Adaptation) significantly reduce the number of trainable parameters by injecting low-rank matrices into the model layers, while keeping the original weights frozen. This drastically lowers memory usage and computational cost, enabling faster training and reduced GPU hours on SageMaker without sacrificing model quality.

Exam trap

A common mistake is thinking that larger instances always improve performance, but Amazon SageMaker optimization often relies on algorithmic efficiency like PEFT rather than just scaling hardware.

How to eliminate wrong answers

Option A is wrong because simply using a larger instance with more vCPUs does not optimize training time or cost if the workload is not parallelizable; it often leads to diminishing returns and higher per-hour costs without proportional speedup. Option B is wrong because increasing the batch size to the maximum possible can cause out-of-memory errors, degrade model convergence, and may require learning rate adjustments, making it an unreliable optimization strategy. Option C is wrong because using full model weights on a single GPU ignores the benefits of distributed training and parameter efficiency, leading to excessive memory consumption and longer training times, which is the opposite of optimizing cost and speed.

41
MCQeasy

A startup wants to integrate a generative AI chatbot into their mobile app with minimal latency. Which AWS service is purpose-built for deploying foundation models with low latency and high throughput?

A.AWS Lambda
B.Amazon SageMaker
C.Amazon Bedrock
D.Amazon Transcribe
AnswerC

Amazon Bedrock provides serverless access to foundation models through a single API with low-latency, high-throughput inference, meeting the chatbot's responsiveness requirement. It avoids managing infrastructure, unlike self-hosted options such as Amazon SageMaker endpoints or EC2-based deployments.

Why this answer

Amazon Bedrock is a fully managed service that provides access to foundation models (FMs) from leading AI providers via a serverless API, purpose-built for deploying generative AI applications with low latency and high throughput. It handles the underlying infrastructure, model hosting, and scaling automatically, making it ideal for integrating a generative AI chatbot into a mobile app with minimal latency.

Exam trap

Candidates often mistake Amazon SageMaker for the purpose-built service, but SageMaker requires manual infrastructure management for low-latency generative AI inference, whereas Amazon Bedrock is a fully managed, serverless service designed specifically for deploying foundation models with low latency and high throughput.

How to eliminate wrong answers

Option A is wrong because AWS Lambda is a serverless compute service for running code in response to events, not designed for hosting or serving large foundation models; it lacks the GPU acceleration and model-specific optimizations needed for low-latency generative AI inference. Option B is wrong because Amazon SageMaker is a broad machine learning platform that can deploy models, but it requires manual setup of endpoints, instance selection, and scaling configuration, adding complexity and potential latency overhead compared to Bedrock's managed FM serving. Option D is wrong because Amazon Transcribe is a speech-to-text service, not a generative AI model deployment service; it cannot host or serve foundation models for chatbot responses.

42
Multi-Selecthard

A media company is evaluating foundation models for an application that will summarize long earnings-call transcripts in English for internal analysts. The transcripts are up to 40,000 words, and the summaries must remain faithful to the source. Which TWO model characteristics are most important to evaluate for this workload? (Choose two.)

Select 2 answers
A.The model's ability to run on edge devices without network connectivity
B.The model's text-to-speech voice quality
C.The model's quality on text summarization and faithfulness to source content
D.The model's ability to generate images from text
E.The model's supported context window size
AnswersC, E

Faithfulness determines whether the summary reflects what the transcript actually said rather than inventing figures or positions. A model with weak summarization quality will produce fluent but unreliable output, which is unacceptable for financial analysis. Benchmarking summarization quality on representative transcripts is therefore essential.

Why this answer

Long transcripts demand a model whose context window can accommodate the input, and financial summarization demands strong faithfulness so the output does not distort the source. Those two characteristics together determine whether the application can produce trustworthy summaries without excessive chunking. Image generation, speech quality, and edge deployment do not address the stated requirements.

Exam trap

The trap here is selecting impressive-sounding multimodal or edge capabilities instead of the two characteristics that actually govern long-document summarization quality.

43
Multi-Selecteasy

Which TWO actions are best practices for reducing hallucinations in generative AI models? (Choose 2)

Select 2 answers
A.Increase the model size
B.Fine-tune the model on proprietary data
C.Use retrieval-augmented generation (RAG)
D.Use a smaller model to limit complexity
E.Apply prompt engineering with clear instructions and constraints
AnswersC, E

Retrieval-augmented generation grounds responses in documents fetched at inference time, so the model conditions on supplied evidence rather than relying solely on parametric memory. This directly addresses the hallucination constraint in the stem by anchoring output to verifiable source content, reducing fabricated claims when the retrieved passages are relevant and accurate.

Why this answer

Option C, retrieval-augmented generation (RAG), is correct because it grounds the model's responses in externally retrieved, authoritative documents at inference time, so the model cites or conditions on real source content rather than relying solely on parametric memory, which measurably reduces fabricated facts. Option E, prompt engineering with clear instructions and constraints, is correct because explicit directives such as "answer only from the provided context," "say 'I don't know' if unsupported," and output-format constraints steer the model away from speculative generation and make hallucinations easier to detect. Option A, increasing model size, is not a reliable fix: larger models can be more fluent and confident while still hallucinating, and scale alone does not guarantee factual grounding.

Option B, fine-tuning on proprietary data, mainly adapts style, format, and domain vocabulary and can even increase confident fabrication on facts not present in the tuning set, so it is not a primary hallucination-reduction best practice. Option D, using a smaller model, is also not a best practice for this goal, since reduced capacity generally lowers factual accuracy and reasoning rather than improving truthfulness.

Exam trap

AWS often tests the misconception that larger models or fine-tuning alone solve hallucinations, when in fact grounding techniques like RAG and explicit prompt constraints are the proven mitigations.

44
MCQhard

A company operates a customer support chatbot that uses Amazon Bedrock with a knowledge base sourced from an S3 bucket containing frequently updated product documentation. The knowledge base uses OpenSearch Serverless as the vector store and is configured to sync daily. The chatbot uses the RetrieveAndGenerate API with a custom Lambda function that applies a system prompt instructing the model to base answers solely on the retrieved context. After a major update to the product documentation, the IT team verifies that the data source sync completed successfully and the new chunks are present in the OpenSearch index. However, the chatbot continues to respond with outdated information. Further investigation reveals that the Lambda function includes a response caching mechanism using Amazon ElastiCache for Redis with a Time-To-Live (TTL) of 24 hours. The cache key is based on the user query. The team notes that no cache invalidation is performed after documentation updates. What is the most likely cause of the outdated responses?

A.The ElastiCache cache is returning stale cached responses that contain the old information.
B.The 'maximum results' parameter in the RetrieveAndGenerate API is set to a value too low to retrieve the new chunks.
C.The embedding model used by the knowledge base has not been retrained on the new documentation.
D.The IAM role for the Lambda function lacks permissions to access the new S3 objects.
AnswerA

The Lambda layer caches responses keyed by user query with a 24-hour TTL and performs no invalidation after syncs, so identical queries return pre-update answers from Redis even though OpenSearch holds fresh chunks. The retrieval layer is bypassed entirely.

Why this answer

The Lambda function caches responses in ElastiCache for Redis with a 24-hour TTL keyed on the user query, and no invalidation occurs after documentation updates. Even though the knowledge base sync succeeded and new chunks are in OpenSearch, the chatbot returns the cached stale answer for any query previously cached. This is the classic cache-staleness failure mode.

Exam trap

AIF-C01 often tests RAG troubleshooting by presenting a scenario where the data pipeline looks healthy, tempting candidates to blame retrieval parameters or IAM — when the real culprit is an application-layer cache returning stale responses.

How to eliminate wrong answers

Option B is wrong because a low 'maximum results' value would reduce the number of retrieved chunks but would not cause consistently outdated answers — and the scenario confirms new chunks are present in the index, so retrieval is not the bottleneck. Option C is wrong because Bedrock knowledge base embedding models are not 'retrained' on new documents; embeddings are generated at ingestion time, and the sync already regenerated embeddings for the new chunks. Option D is wrong because the scenario explicitly states the data source sync completed successfully and new chunks are in the OpenSearch index, which proves the Lambda's IAM role has the necessary S3 and OpenSearch permissions.

45
MCQeasy

A company wants to use a pre-trained generative AI model to analyze customer feedback. They need to adjust the model for their specific domain without retraining from scratch. Which approach is MOST suitable?

A.Fine-tuning the model on domain-specific data
B.Reinforcement Learning from Human Feedback (RLHF)
C.Training a new model from scratch on the domain data
D.Using prompt engineering to provide context
AnswerA

Fine-tuning adapts a pre-trained model's weights using domain-specific customer feedback, satisfying the requirement to specialise without training from scratch. Unlike prompt engineering or RAG, which leave parameters frozen, fine-tuning alters the model itself, embedding domain vocabulary and sentiment patterns directly. This matches the constraint of domain adaptation at lower cost than full retraining.

Why this answer

Fine-tuning is the most suitable approach because it takes a pre-trained generative AI model and updates its weights using a smaller, domain-specific dataset (e.g., customer feedback transcripts). This allows the model to adapt to the company's specific terminology, sentiment patterns, and context without the massive computational cost and data requirements of training from scratch. It preserves the general language understanding from pre-training while specializing the model for the target domain.

Exam trap

The AIF-C01 exam often tests the distinction between prompt engineering (a zero-shot or few-shot method that does not modify the model) and fine-tuning (which updates model weights), leading candidates to mistakenly choose prompt engineering as a simpler but insufficient solution for deep domain adaptation.

How to eliminate wrong answers

Option B (RLHF) is wrong because RLHF is a technique used to align model outputs with human preferences through reward modeling, not primarily for domain adaptation; it requires a separate reward model and human feedback loop, making it overkill and less direct for simply specializing on domain-specific data. Option C (training a new model from scratch) is wrong because it discards the benefits of pre-training, requiring enormous amounts of domain data and compute resources, which contradicts the requirement to avoid retraining from scratch. Option D (prompt engineering) is wrong because while it can provide context at inference time, it does not adjust the model's internal weights or permanently adapt it to the domain; it relies on the model's existing knowledge and may fail for nuanced or rare domain-specific terms.

46
Multi-Selecteasy

Which TWO of the following are key advantages of using Amazon Bedrock for building generative AI applications?

Select 2 answers
A.Automatic optimization of prompts for all models without user intervention.
B.Ability to fine-tune models using your own data without managing underlying infrastructure.
C.Eliminates the need for any data preprocessing before model invocation.
D.Guaranteed identical outputs from all models for the same prompt.
E.Access to multiple foundation models from different providers via a single API.
AnswersB, E

Bedrock lets you customise foundation models with your own labelled data through fine-tuning and continued pre-training, while AWS manages the compute and serving infrastructure, removing the operational burden of provisioning and maintaining training environments yourself.

Why this answer

Option B is correct because Amazon Bedrock is a fully managed service that lets you customize (fine-tune) foundation models with your own data while AWS handles the underlying infrastructure, so you never provision or manage servers or clusters. Option E is correct because Bedrock provides a single, unified API to access multiple foundation models from different providers such as Anthropic, AI21 Labs, Cohere, Meta, Stability AI, and Amazon, which simplifies multi-model development. Option A is wrong because Bedrock does not automatically optimize prompts for all models without user intervention; prompt engineering remains the developer's responsibility.

Option C is wrong because data preprocessing is still required for many use cases, especially when preparing fine-tuning datasets or cleaning inputs. Option D is wrong because generative models are probabilistic and do not guarantee identical outputs for the same prompt, even with low temperature settings.

Exam trap

AWS often tests the misconception that a managed AI service like Bedrock automates all data preparation and prompt engineering, when in reality these tasks remain critical user responsibilities.

47
Multi-Selectmedium

A data science team is evaluating foundation models for a code generation task. They need a model that is fine-tuned for code and can be deployed on Amazon SageMaker. Which THREE criteria are important to consider when selecting a model?

Select 3 answers
A.Licensing and usage terms
B.Cost per token for inference
C.Context window length
D.The training algorithm used
E.Model size and architecture
AnswersA, B, C

Licensing and usage terms determine whether a code generation model can legally be deployed commercially on Amazon SageMaker, covering permitted use, redistribution and derivative works. This satisfies the stem's deployment constraint, since unsuitable terms block production use regardless of model quality or code fine-tuning.

Why this answer

Option A (Licensing and usage terms) is correct because foundation models on SageMaker come with varying licenses (e.g., Apache 2.0, Llama Community License, or proprietary commercial terms), and the team must verify the license permits their intended commercial code-generation use case before deployment. Option B (Cost per token for inference) is correct because SageMaker endpoints bill based on instance hours and/or token throughput, so understanding per-token inference cost is essential for budgeting a production code-generation workload at scale. Option C (Context window length) is correct because code generation often requires ingesting large files or repositories, and a model's maximum context window (e.g., 4K, 8K, 32K, or 128K tokens) directly determines how much source code can be passed in a single prompt.

Option D is not a primary selection criterion because the training algorithm (e.g., RLHF, SFT) is an internal implementation detail that does not directly affect deployment fit or task suitability. Option E is not a primary criterion because model size and architecture are secondary considerations already reflected in measurable factors like cost, latency, and context window, rather than being a distinct decision driver for this scenario.

Exam trap

Candidates may focus on technical aspects like model architecture or training algorithm, overlooking the practical constraints of licensing, cost, and context window, which are equally critical for deployment.

48
MCQhard

A developer is using Amazon Bedrock's Converse API to build a multi-turn conversation. They notice the model forgets earlier context after a few exchanges. What is the most likely cause?

A.The API has a rate limit that truncates history
B.The model's context window is too small for the conversation
C.The model's maximum output length is set too low
D.The developer is not sending the previous messages in each request
AnswerD

The Converse API is stateless: it retains no memory between calls, so multi-turn context exists only if the developer resends prior messages each request. Omitting earlier turns means the model receives no history, causing it to forget context after a few exchanges.

Why this answer

The Converse API is stateless by design, meaning it does not retain conversation history between API calls. To maintain context across multiple turns, the developer must explicitly include the entire message history (previous user and assistant messages) in each new request. If the developer omits these previous messages, the model has no memory of earlier exchanges and will appear to forget context.

Exam trap

The trap here is that candidates may incorrectly attribute memory loss to model limitations (like context window size) rather than the developer's failure to include conversation history in each request.

How to eliminate wrong answers

Option A is wrong because rate limits control the frequency of API calls, not the content or history of messages; they do not truncate the conversation history. Option B is wrong because while a small context window can limit the total tokens the model can process, the primary issue described (forgetting after a few exchanges) is caused by not sending history, not by the window size itself. Option C is wrong because the maximum output length parameter limits the length of the model's response, not the input context or the retention of conversation history.

49
MCQeasy

A financial services company wants to deploy a generative AI assistant that summarizes internal earnings reports. The reports contain confidential, non-public financial data, and the security team requires that the data never leaves the company's AWS environment or be used to train the provider's public models. The company already uses Amazon Bedrock. Which characteristic of Amazon Bedrock directly satisfies this requirement?

A.Amazon Bedrock does not use customer inputs or outputs to train the base foundation models.
B.Amazon Bedrock restricts all inference to a single AWS Region chosen at account creation.
C.Amazon Bedrock provides server-side encryption with AWS KMS keys for data at rest.
D.Amazon Bedrock automatically redacts personally identifiable information from all prompts before inference.
AnswerA

Amazon Bedrock is architected so that customer data such as prompts and completions is not used to train the underlying base foundation models. This fulfills the security team's explicit requirement that confidential earnings content never becomes training data for a provider's public model, while still allowing the company to call the models in its own AWS account.

Why this answer

The security requirement is specifically that confidential data must not be used to train the provider's models. Amazon Bedrock's design keeps customer prompts and completions out of base foundation model training, which is the property that satisfies that constraint. Encryption, redaction, and regional pinning are useful but address different risks, so they do not meet the stated training-data isolation requirement.

Exam trap

The trap here is assuming that encryption or regional placement alone guarantees data will not be used for model training, when the relevant guarantee comes from how the managed service handles customer inputs and outputs.

50
Multi-Selectmedium

Which THREE are key capabilities of Amazon Bedrock? (Choose 3)

Select 3 answers
A.Automatic model selection based on use case
B.Model customization through fine-tuning
C.Guardrails to filter harmful content
D.Serverless inference for foundation models
E.Built-in vector database for knowledge bases
AnswersB, C, D

Fine-tuning adjusts a foundation model's weights using your labelled dataset, tailoring outputs to domain-specific tasks. This satisfies the stem's requirement for a key Bedrock capability, since Bedrock supports custom models trained on your data, alongside provisioned throughput for hosting them.

Why this answer

Amazon Bedrock provides model customization through fine-tuning (B), allowing you to adapt supported foundation models with your own labeled data to improve performance for domain-specific tasks. It also offers Guardrails for Amazon Bedrock (C), which let you define policies that filter harmful or inappropriate content and enforce topics and sensitive-information redaction across model responses. Bedrock additionally delivers serverless inference for foundation models (D), so you can invoke models via a managed API without provisioning or managing any underlying infrastructure.

Option A is not a Bedrock capability because Bedrock does not automatically choose a model for you; the developer selects the model. Option E is incorrect because Bedrock itself does not include a built-in vector database; knowledge bases for Amazon Bedrock integrate with separate vector stores such as Amazon OpenSearch Serverless or Amazon Aurora.

Exam trap

A common misconception is that Amazon Bedrock includes a built-in vector database for knowledge bases, when in fact it integrates with external vector stores such as Amazon OpenSearch Serverless or Pinecone. Another misconception is that Bedrock automatically selects the best model for the use case, whereas users must manually evaluate and choose models based on performance metrics and specific requirements.

51
MCQmedium

A media company is experimenting with an Amazon Bedrock text model to draft short news summaries. The team notices that when they ask the same question twice, the model returns noticeably different wording each time, and sometimes the summary drifts off topic. They want more deterministic, focused responses without retraining the model. Which combination of inference parameters should they adjust to reduce randomness and keep the output on topic?

A.Increase the temperature and increase the topP value.
B.Decrease the temperature and set a lower topP value.
C.Enable streaming responses and lower the topK value to zero.
D.Increase maxTokens and decrease the stop sequence length.
AnswerB

Lowering temperature sharpens the probability distribution so high-probability tokens are favored, and reducing topP narrows nucleus sampling to fewer likely tokens. Together they make output more deterministic and on topic, which matches the team's goal without any retraining or change to the model itself.

Why this answer

Randomness and topical drift in foundation model output are governed by sampling parameters. Temperature scales the probability distribution, and topP restricts sampling to the smallest set of tokens whose cumulative probability exceeds the threshold. Lowering both concentrates generation on the most likely tokens, yielding more consistent and focused summaries without any model retraining.

Exam trap

The trap here is mixing up output length and delivery controls, such as maxTokens or streaming, with sampling controls that actually govern randomness and topical focus.

52
MCQeasy

Which AWS service provides a fully managed experience for building generative AI applications with a variety of foundation models through a unified API?

A.AWS Lambda
B.Amazon SageMaker
C.Amazon Rekognition
D.Amazon Bedrock
AnswerD

Amazon Bedrock offers a fully managed, serverless experience exposing multiple foundation models from providers such as Anthropic, Meta and Amazon through one unified API, removing infrastructure management. This matches the requirement for varied models behind a single consistent interface.

Why this answer

Amazon Bedrock is the correct answer because it is a fully managed AWS service that provides access to a variety of foundation models (FMs) from providers like AI21 Labs, Anthropic, Cohere, Meta, Stability AI, and Amazon through a single, unified API. This allows developers to build and scale generative AI applications without managing underlying infrastructure or model hosting, directly addressing the requirement for a fully managed experience with diverse FMs.

Exam trap

The trap here is that candidates often confuse Amazon SageMaker's broad ML capabilities with Bedrock's specific focus on managed foundation model access, leading them to select SageMaker because they think it covers all AI workloads, but SageMaker does not provide a unified API for multiple pre-built FMs.

How to eliminate wrong answers

Option A is wrong because AWS Lambda is a serverless compute service for running code in response to events, not a service for accessing or managing foundation models for generative AI. Option B is wrong because Amazon SageMaker is a machine learning platform for building, training, and deploying custom models, but it does not provide a unified API for a variety of pre-built foundation models; it requires users to manage model hosting and inference endpoints. Option C is wrong because Amazon Rekognition is a computer vision service for image and video analysis, such as object detection and facial recognition, and does not offer generative AI capabilities or foundation model access.

53
MCQhard

A team is using Amazon Bedrock with a Claude model and wants to ensure responses adhere to a specific output format such as JSON. Which technique should be applied?

A.Use a retrieval-augmented generation (RAG) approach
B.Attach a guardrail with a JSON schema
C.Include a system prompt with explicit formatting instructions
D.Customize the model with a JSON training dataset
AnswerC

A system prompt supplies persistent instructions that condition every response, so Claude structures output as JSON rather than prose. This constrains generation at inference time, satisfying the requirement for enforcing a specific output format without retraining or fine-tuning the model.

Why this answer

Amazon Bedrock with Claude models supports system prompts that can include explicit formatting instructions, such as 'Respond in valid JSON format.' This technique directly controls the model's output structure without requiring external tools or retraining, making it the simplest and most effective method for enforcing a specific output format like JSON.

Exam trap

Candidates may confuse guardrails (which enforce content safety policies) with output formatting controls, or assume that RAG or fine-tuning are necessary for simple structural constraints, when in fact a well-crafted system prompt is the standard and most efficient approach.

How to eliminate wrong answers

Option A is wrong because retrieval-augmented generation (RAG) is designed to enhance responses with external knowledge sources, not to enforce output formatting; it does not constrain the model's output structure. Option B is wrong because guardrails in Amazon Bedrock are used to enforce content policies (e.g., blocking harmful content) and do not support JSON schema validation for output formatting; they are not designed for structural constraints. Option D is wrong because customizing the model with a JSON training dataset would require fine-tuning, which is costly, time-consuming, and unnecessary for simple formatting tasks; system prompts provide a more efficient and flexible solution.

54
MCQmedium

A media company wants a generative AI assistant that drafts scripts in a distinctive house style. The team has several thousand pages of approved scripts but no labeled input-output pairs, and they want to adapt an existing foundation model rather than train one from scratch. Which approach best matches their data and goal?

A.Reinforcement learning from human feedback using preference rankings of generated scripts
B.Supervised fine-tuning on prompt-completion pairs extracted from the scripts
C.Few-shot prompting with three example scripts inserted into each request
D.Continued pre-training on the unlabeled script corpus to adapt the model to the house style
AnswerD

Continued pre-training consumes large volumes of unlabeled domain text and adjusts the model's weights toward that distribution, which is precisely how a model absorbs a distinctive vocabulary, tone, and structure. Because the scenario has thousands of pages and no paired labels, this is the only listed method that uses the data as-is while still changing model behavior beyond what prompting alone provides.

Why this answer

Continued pre-training is designed for exactly this situation: a large body of unlabeled domain text and a desire to shift an existing foundation model toward that domain. Supervised fine-tuning and RLHF both need labels the team does not have, and few-shot prompting cannot absorb thousands of pages of stylistic signal into the model itself.

Exam trap

The trap here is reaching for supervised fine-tuning by instinct, even though the scenario states there are no labeled input-output pairs, which rules it out.

55
MCQeasy

A data science intern is reading about the transformer architecture that underpins most modern large language models. The intern asks which mechanism allows a transformer to weigh the relevance of every other word in a sentence when encoding a given word, regardless of how far apart the words appear. Which mechanism should you name?

A.Self-attention
B.Positional encoding
C.Tokenization
D.Recurrent state propagation
AnswerA

Self-attention computes a weighted relationship between every token and every other token in the sequence, so a word's representation is built from the whole context rather than a fixed local window. This directly removes the distance limitation of recurrent architectures, which is exactly the property the intern is asking about, and it is the core building block of the transformer encoder and decoder stacks.

Why this answer

Self-attention is the operation that lets each token build its representation from all other tokens in the sequence with learned weights, which is what removes the long-distance constraint of sequential models. Positional encoding, tokenization, and recurrent state propagation all play supporting or unrelated roles, but none of them performs the relevance weighting described in the scenario.

Exam trap

The trap here is assuming that positional encoding is what gives a transformer its long-range context, when positional encoding only supplies order information and self-attention does the actual cross-token weighting.

56
MCQeasy

A company wants to generate product descriptions from a few keywords without managing infrastructure. Which AWS service provides a serverless API for accessing foundation models?

A.Amazon Lex
B.Amazon SageMaker
C.Amazon Bedrock
D.Amazon Comprehend
AnswerC

Amazon Bedrock exposes foundation models from multiple providers through a fully managed, serverless API, so no infrastructure is provisioned or scaled by the customer. This matches the stem's serverless requirement for generating product descriptions from keywords.

Why this answer

Amazon Bedrock is a fully managed, serverless service that provides a single API to access and invoke foundation models (FMs) from leading AI providers like AI21 Labs, Anthropic, Cohere, Meta, and Stability AI. It eliminates the need to manage underlying infrastructure, making it the correct choice for generating product descriptions from keywords without provisioning servers.

Exam trap

The trap here is that candidates often confuse Amazon Bedrock with Amazon SageMaker, assuming SageMaker's JumpStart or hosting capabilities provide the same serverless foundation model access, but SageMaker requires explicit endpoint management and is not a serverless API for pre-built FMs.

How to eliminate wrong answers

Option A is wrong because Amazon Lex is a service for building conversational interfaces (chatbots) using automatic speech recognition and natural language understanding, not a serverless API for accessing foundation models. Option B is wrong because Amazon SageMaker is a fully managed machine learning platform that requires you to manage training jobs, endpoints, and infrastructure; it is not a serverless API for directly invoking pre-built foundation models. Option D is wrong because Amazon Comprehend is a natural language processing (NLP) service for extracting insights like sentiment, entities, and key phrases from text, not a generative AI service for producing new content from foundation models.

57
MCQhard

A company is building a chatbot that must provide accurate answers based on internal documents without retraining the model. Which approach should they use?

A.Reinforcement learning from human feedback (RLHF)
B.Fine-tuning the model on internal documents
C.Model distillation to a smaller model
D.Prompt engineering with retrieval-augmented generation (RAG)
AnswerD

Retrieval-augmented generation retrieves relevant passages from internal documents at query time and supplies them as context, so the chatbot answers from current source material without retraining. Prompt engineering shapes how that retrieved context is used, meeting the no-retraining constraint.

Why this answer

Retrieval-augmented generation (RAG) allows the chatbot to fetch relevant internal documents at inference time and incorporate them into the prompt, providing accurate, up-to-date answers without retraining the model. This approach combines prompt engineering with a retrieval step, ensuring the model's responses are grounded in the company's specific knowledge base while keeping the base model frozen.

Exam trap

The trap here is that candidates may confuse fine-tuning (which requires retraining) with RAG (which does not), or mistakenly think RLHF or distillation can inject new factual knowledge without retraining, when in fact they address alignment, efficiency, or behavior, not dynamic knowledge retrieval.

How to eliminate wrong answers

Option A is wrong because reinforcement learning from human feedback (RLHF) is a training technique used to align model behavior with human preferences, not a method for injecting new factual knowledge without retraining. Option B is wrong because fine-tuning the model on internal documents would require modifying the model's weights through additional training, which contradicts the requirement of not retraining the model. Option C is wrong because model distillation compresses a large model into a smaller one for efficiency, but it does not enable the model to answer questions based on new internal documents without retraining.

58
MCQmedium

A product team is comparing foundation models for a customer-facing FAQ bot. They need low latency for real-time chat and want to pay only for what they use without managing servers. Which combination of model characteristic and AWS consumption model best fits these requirements?

A.A large model with the highest parameter count, deployed on a provisioned throughput commitment billed hourly.
B.An open-weight model downloaded and hosted on AWS Lambda for per-request execution.
C.A multimodal model with image understanding, called through a self-managed Amazon EC2 GPU instance.
D.A smaller, lower-latency model invoked on demand so billing is based on input and output tokens processed.
AnswerD

Smaller models generally produce tokens faster, which supports responsive chat, and on-demand invocation charges per token consumed, matching the pay-only-for-what-you-use requirement with no infrastructure to manage. This aligns latency, cost model, and operational simplicity with the team's constraints for a customer FAQ workload.

Why this answer

Real-time chat favors models that emit tokens quickly, and a smaller model usually wins on latency. On-demand invocation bills by tokens processed, so a variable-traffic FAQ bot pays only for actual usage and the provider handles scaling and infrastructure. Provisioned capacity and self-hosted GPUs suit steady high volume but add cost and operational burden here.

Exam trap

The trap here is equating the largest model with the best choice, when parameter count usually increases latency and cost rather than improving fit for a latency-sensitive FAQ bot.

59
MCQhard

A data scientist is fine-tuning a large language model on Amazon SageMaker for a text summarization task. The training loss decreases steadily but the validation loss starts increasing after a few epochs. What should the scientist do to address this issue?

A.Reduce the batch size
B.Increase the learning rate
C.Increase the number of training epochs
D.Use early stopping based on validation loss
AnswerD

Early stopping halts training once validation loss stops improving, restoring the checkpoint with the lowest validation loss and preventing further overfitting. This satisfies the requirement to counter rising validation loss while training loss still falls.

Why this answer

The validation loss increasing while training loss decreases is a classic sign of overfitting. Early stopping based on validation loss halts training when the validation loss stops improving, preventing overfitting and saving computational resources. This is a standard technique in SageMaker's built-in training algorithms and custom training scripts.

Exam trap

The AIF-C01 exam often tests the distinction between overfitting and underfitting; the trap here is that candidates may mistakenly think increasing epochs (Option C) always improves performance, ignoring the validation loss divergence that signals overfitting.

How to eliminate wrong answers

Option A is wrong because reducing batch size introduces more noise into gradient estimates, which can actually worsen generalization and does not directly address overfitting. Option B is wrong because increasing the learning rate can cause the optimizer to overshoot minima, leading to divergence or unstable training, not reduced overfitting. Option C is wrong because increasing the number of training epochs would exacerbate overfitting, as the model would continue to memorize the training data beyond the point where validation loss degrades.

60
MCQmedium

A financial services firm is evaluating foundation models for a customer-facing assistant. Compliance requires that prompts and completions never leave the firm's own AWS account boundary. Which characteristic of a foundation model deployment should the team evaluate FIRST against this constraint?

A.The number of parameters in the model and its published benchmark scores.
B.The maximum number of tokens the model can accept in a single request.
C.The licensing terms that govern commercial redistribution of the model weights.
D.Whether the model is consumed as a fully managed service or deployed into infrastructure the firm controls.
AnswerD

The consumption model determines whether inference happens on shared managed endpoints or inside resources inside the firm's own account, which is the decisive factor for the boundary requirement. Fully managed APIs process prompts outside the customer account, whereas self-hosted deployments keep data within the firm's environment, so this must be settled before comparing model quality.

Why this answer

Data-residency and account-boundary requirements are determined by where inference actually executes, so the deployment and consumption model must be assessed before capability metrics. Fully managed endpoints process requests on provider-managed infrastructure, while models deployed into the firm's own account keep prompt and completion data inside that boundary.

Exam trap

The trap here is evaluating model quality metrics first, when a hard compliance constraint about where data is processed must drive the architecture decision.

61
MCQeasy

A developer is creating a generative AI application using Amazon Bedrock and needs to ensure that responses do not include toxic or harmful content. Which feature should be enabled?

A.Amazon CloudWatch Logs for prompt logging.
B.Amazon Virtual Private Cloud (VPC) for network isolation.
C.Amazon Bedrock Guardrails.
D.AWS Identity and Access Management (IAM) policies.
AnswerC

Guardrails applies configurable content filters and denied-topic policies to both prompts and responses, blocking toxic or harmful output at inference time. This satisfies the requirement that responses never include harmful content, without retraining or prompt engineering.

Why this answer

Amazon Bedrock Guardrails is the correct feature because it is specifically designed to enforce content policies, filter toxic or harmful content, and block undesirable topics in generative AI responses. It provides configurable thresholds for hate, insults, sexual content, violence, and other harmful categories, ensuring compliance with safety requirements without modifying the underlying model.

Exam trap

The trap here is that candidates often confuse monitoring/logging services (CloudWatch) or security controls (VPC, IAM) with content safety features, not realizing that Bedrock Guardrails is the only option that directly filters toxic or harmful content at the application layer.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch Logs for prompt logging captures and stores logs for monitoring and debugging, but it does not actively filter or block toxic content in responses. Option B is wrong because Amazon Virtual Private Cloud (VPC) provides network isolation and security at the infrastructure layer, but it has no mechanism to inspect or control the semantic content of AI-generated responses. Option D is wrong because AWS Identity and Access Management (IAM) policies control authentication and authorization for API calls, but they cannot enforce content safety rules or filter harmful language in model outputs.

62
MCQhard

A developer sets a low temperature value when calling a foundation model through Amazon Bedrock for a use case that extracts structured fields from invoices. A colleague argues that lowering temperature reduces hallucination and guarantees correct extraction. How should the developer respond?

A.Lowering temperature forces the model to retrieve the correct field values from the invoice image before generating text.
B.Lowering temperature disables the model's ability to generate any text that was not present in its training data.
C.Lowering temperature makes sampling more deterministic, which improves consistency but does not guarantee factual correctness or eliminate hallucination.
D.Lowering temperature guarantees the response will conform to the requested JSON schema without any prompt instructions.
AnswerC

Temperature scales the randomness of token sampling; a low value makes the model favor high-probability tokens so outputs become more repeatable. However, if the model's learned associations are wrong or the input is ambiguous, the deterministic output can still be incorrect. Consistency and accuracy are distinct properties, so the colleague's guarantee claim is unfounded.

Why this answer

Temperature governs the randomness of token sampling, so lowering it makes outputs more deterministic and repeatable. Determinism is not the same as correctness: a model can consistently produce the same wrong field value, and hallucinations can persist even at very low temperature. Accurate extraction therefore still depends on model capability, prompt design, and validation logic.

Exam trap

The trap here is equating low temperature with factual accuracy, when the parameter only makes sampling more deterministic and can consistently reproduce the same error.

63
MCQmedium

A media company wants to add a generative AI assistant to its internal knowledge portal. The assistant must answer employee questions using only the company's private policy documents, and it must cite the exact source passage for each answer. The team plans to use a foundation model hosted in Amazon Bedrock. Which approach should they implement to meet these requirements?

A.Raise the maximum token limit in the model invocation parameters so the entire policy library can be pasted into every prompt.
B.Fine-tune the foundation model on the full set of policy documents and deploy a custom model that answers from its updated weights.
C.Use Retrieval Augmented Generation by storing policy documents in a vector store and retrieving relevant passages to include in the prompt sent to the model.
D.Increase the model's temperature setting so it produces more varied and detailed answers from its pretrained knowledge.
AnswerC

Retrieval Augmented Generation retrieves the most relevant passages from an indexed knowledge base and injects them into the prompt, so the model conditions its answer on the company's own documents instead of only pretrained knowledge. Because the retrieved passages are known, the application can return them as citations. This directly satisfies both the grounding and citation requirements without retraining the model.

Why this answer

The requirements are grounding answers in private documents and citing exact source passages. Retrieval Augmented Generation satisfies both by indexing the policy corpus, retrieving the most relevant chunks at query time, and passing them into the prompt so the model can answer from supplied evidence. The application can then surface the retrieved passages as citations, keeping answers accurate and auditable without retraining the foundation model.

Exam trap

The trap here is assuming that fine-tuning or a larger context window can substitute for retrieval when the real requirement is traceable, up-to-date grounding in a private document set.

64
Multi-Selecthard

Which TWO practices help ensure responsible AI when deploying generative AI applications? (Select TWO.)

Select 2 answers
A.Deploy the model without any content filters to maximize creativity
B.Increase model size to improve accuracy at the expense of interpretability
C.Use only synthetic data for training to avoid privacy issues
D.Implement guardrails to filter harmful or inappropriate content
E.Monitor the model's outputs for bias and drift over time
AnswersD, E

Guardrails intercept prompts and responses, blocking harmful, inappropriate or policy-violating content before it reaches users. This enforces responsible AI by constraining generative output at runtime, satisfying the requirement to prevent unsafe material being surfaced in deployed applications.

Why this answer

Option D is correct because implementing guardrails—such as content filters, prompt shields, and safety classifiers—directly mitigates the risk of generative AI producing harmful, offensive, or inappropriate content, which is a core requirement of responsible AI deployment. Option E is correct because continuously monitoring model outputs for bias and drift ensures that the system remains fair and accurate as data distributions and user behavior change over time, enabling timely remediation. Option A is incorrect because removing content filters increases the risk of harmful outputs and violates responsible AI principles rather than ensuring them.

Option B is incorrect because increasing model size at the cost of interpretability reduces transparency and explainability, which are key responsible AI goals. Option C is incorrect because using only synthetic data does not by itself guarantee privacy or fairness and can introduce its own biases and quality issues.

Exam trap

AIF-C01 often tests whether candidates confuse 'model performance improvements' (size, accuracy) with 'responsible AI controls' (guardrails, monitoring, transparency), so options that sound like optimization tricks are the trap.

65
MCQmedium

Refer to the exhibit. A company sets up a knowledge base for a customer support chatbot using Amazon Bedrock. Users report that the chatbot misses relevant details from long documents. Which change to the data source configuration would most likely improve retrieval?

A.Increase the chunk size in FIXED_SIZE chunking
B.Change chunking strategy to SEMANTIC
C.Add more documents to the S3 bucket
D.Change the embedding model to a larger one
AnswerB

Semantic chunking splits documents at natural meaning boundaries rather than fixed token counts, preserving coherent context within each chunk. This improves retrieval accuracy for long documents, since relevant details are less likely to be split across chunks and missed during embedding search.

Why this answer

Semantic chunking groups text based on meaning rather than fixed token counts, preserving the natural boundaries of concepts and paragraphs. This ensures that relevant details from long documents remain intact within a single chunk, improving retrieval accuracy for the chatbot.

Exam trap

AWS often tests the misconception that simply increasing chunk size or using a larger embedding model will improve retrieval, when the real bottleneck is the chunking strategy's ability to preserve semantic coherence.

How to eliminate wrong answers

Option A is wrong because increasing the chunk size in FIXED_SIZE chunking can cause chunks to contain multiple unrelated topics, diluting the semantic focus and making retrieval less precise. Option C is wrong because adding more documents to the S3 bucket does not address the core issue of poor chunking; it may even introduce more noise if the chunking strategy remains suboptimal. Option D is wrong because changing the embedding model to a larger one may improve representation quality but does not fix the fundamental problem of how documents are split; poorly chunked content will still lose relevant details regardless of the embedding model.

66
MCQeasy

A company wants to build a generative AI application that can summarize customer support tickets. They need to ensure the model stays up-to-date with the latest product documentation without retraining. Which AWS service would best support this requirement?

A.Amazon Bedrock with Retrieval Augmented Generation (RAG)
B.Amazon Comprehend
C.Amazon Rekognition
D.Amazon SageMaker Ground Truth
AnswerA

Amazon Bedrock with RAG retrieves current product documentation from a knowledge base at inference time and injects relevant passages into the prompt, so summaries reflect the latest content without retraining. This directly satisfies the stem's constraint of staying up-to-date without retraining, since updating the underlying data store refreshes outputs immediately.

Why this answer

Amazon Bedrock with Retrieval Augmented Generation (RAG) is the correct choice because it allows the generative AI model to access and incorporate the latest product documentation from an external knowledge base without retraining. RAG works by retrieving relevant document chunks at inference time and injecting them into the model's context, ensuring responses reflect current information. This directly meets the requirement for staying up-to-date with evolving documentation while avoiding the cost and latency of full model retraining.

Exam trap

The trap here is that candidates may confuse Amazon Comprehend's text analysis capabilities (like summarization via extractive methods) with generative AI summarization, overlooking that Comprehend cannot incorporate external, dynamic knowledge sources without retraining.

How to eliminate wrong answers

Option B (Amazon Comprehend) is wrong because it is a natural language processing (NLP) service for extracting insights like sentiment, entities, and key phrases from text, but it does not provide generative AI summarization capabilities or a mechanism to dynamically incorporate updated documentation. Option C (Amazon Rekognition) is wrong because it is a computer vision service for analyzing images and videos, not for processing text-based customer support tickets or integrating with product documentation. Option D (Amazon SageMaker Ground Truth) is wrong because it is a data labeling service used to create training datasets for machine learning models, not a generative AI service that can summarize text or retrieve real-time information from external sources.

67
MCQmedium

A machine learning engineer is preparing to fine-tune a foundation model for a specialized legal summarization task. The team has only a few thousand labeled examples and a limited budget. Which fine-tuning approach is most appropriate to adapt the model efficiently under these constraints?

A.Increase the context window size so the model can read entire legal contracts at once.
B.Perform full fine-tuning that updates every parameter in the network.
C.Use parameter-efficient fine-tuning that updates only a small subset of adapter weights.
D.Train the foundation model from scratch on the legal corpus.
AnswerC

Parameter-efficient fine-tuning methods such as LoRA update only a small number of added or selected parameters while freezing the base model. This dramatically reduces compute, memory, and storage needs, making it feasible with a few thousand examples and a limited budget. It adapts the model to legal summarization effectively without the cost of full retraining.

Why this answer

Parameter-efficient fine-tuning, such as LoRA, freezes the base model and trains only a small set of additional parameters. This cuts compute and memory requirements sharply, making adaptation feasible with limited data and budget. It still specializes the model for legal summarization, delivering the task adaptation the team needs without the expense of full retraining or training from scratch.

Exam trap

The trap here is equating fine-tuning with updating all model weights, overlooking that parameter-efficient methods achieve task adaptation at a fraction of the cost.

68
MCQmedium

A financial services firm fine-tuned a generative AI model on Amazon SageMaker to summarize quarterly reports. The summaries often miss key financial metrics such as revenue and profit margins. The fine-tuning dataset contained full reports with summaries that included these metrics. The model appears to understand the reports but omits critical numbers. Which course of action would most likely improve the summaries?

A.Re-fine-tune using a carefully crafted dataset that includes explicit instructions to include key metrics and provides examples of correct summaries
B.Increase the maximum number of tokens in the summary
C.Switch to a different pre-trained model like Claude instead of the current one
D.Implement a post-processing Lambda function that extracts metrics from the original report and appends them to the summary
AnswerA

The dataset lacked supervision signalling that metrics matter, so the model learned to omit them. Re-fine-tuning with explicit instructions and exemplar summaries that retain revenue and profit figures supplies that signal, steering generation toward including critical numbers.

Why this answer

The model's failure to include key financial metrics despite having them in the training data indicates a misalignment between the training objective and the desired output. By re-fine-tuning with a dataset that explicitly instructs the model to include key metrics and provides correct examples, you directly teach the model to prioritize and extract those specific numerical values during summarization, addressing the root cause of omission.

Exam trap

AWS often tests the misconception that increasing output length or switching models will fix content omission, when the real solution lies in improving the fine-tuning data quality and instruction design.

How to eliminate wrong answers

Option B is wrong because increasing the maximum number of tokens does not force the model to include specific metrics; it only allows longer outputs, but the model may still choose to omit critical numbers if it hasn't learned to prioritize them. Option C is wrong because switching to a different pre-trained model like Claude does not guarantee that the new model will automatically include key metrics; the underlying issue is the fine-tuning data and instruction format, not the base model's architecture. Option D is wrong because implementing a post-processing Lambda function that extracts metrics and appends them is a brittle workaround that does not fix the model's behavior; it adds complexity and may produce inconsistent summaries, whereas the goal is to have the model generate complete summaries natively.

69
MCQmedium

Refer to the exhibit. A developer has attached this IAM policy to their user. When trying to invoke the Anthropic Claude v2 model using the Bedrock runtime, they receive an AccessDeniedException. Which change to the policy would resolve the issue?

A.Add the bedrock:InvokeModelWithResponseStream action
B.Change the Action to bedrock:ListFoundationModels
C.Change the Resource to arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-v2
D.Remove the Resource element and set Effect to Deny
AnswerC

The policy's Resource currently points at a provisioned throughput ARN, which only authorises invoking models via a purchased throughput commitment. Switching to the foundation-model ARN grants access to on-demand invocation of Anthropic Claude v2, satisfying the Bedrock runtime call that triggered the AccessDeniedException.

Why this answer

The IAM policy's Resource element must specify the exact ARN of the foundation model being invoked. The original policy likely used a wildcard or incorrect ARN pattern. The correct ARN for Anthropic Claude v2 in us-east-1 is arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-v2, which grants permission to invoke that specific model via the Bedrock runtime.

Exam trap

AWS often tests the nuance that Bedrock foundation model ARNs use a double colon (::) and omit the account ID, leading candidates to choose a wildcard Resource or a malformed ARN.

How to eliminate wrong answers

Option A is wrong because adding the bedrock:InvokeModelWithResponseStream action would allow streaming responses, but the error is AccessDeniedException due to an incorrect Resource ARN, not a missing action. Option B is wrong because changing the Action to bedrock:ListFoundationModels would only allow listing models, not invoking them, which does not resolve the invoke failure. Option D is wrong because removing the Resource element and setting Effect to Deny would explicitly deny all actions, making the problem worse; the Effect should be Allow, and the Resource must be correctly specified.

70
MCQmedium

A company is using Amazon Bedrock to generate code snippets. They notice the model occasionally generates code that fails to compile. What is the most effective way to improve code quality without retraining?

A.Reduce the temperature parameter to 0 for deterministic output.
B.Increase the max token limit to allow the model to complete the code fully.
C.Fine-tune the model on a dataset of correct code snippets.
D.Use few-shot prompt engineering with correct code examples and formatting instructions.
AnswerD

Few-shot prompting supplies in-context examples of compilable code plus explicit formatting rules, steering the model's token distribution toward syntactically valid output without weight updates. This directly satisfies the no-retraining constraint, unlike fine-tuning, and targets the compilation failures observed in the stem.

Why this answer

Few-shot prompt engineering provides the model with explicit examples of correct code and formatting instructions, guiding it to generate syntactically valid code without modifying the underlying model. This approach leverages in-context learning to improve output quality by conditioning the model on desired patterns, which is more effective than parameter adjustments alone for addressing compilation errors.

Exam trap

A common misconception in AWS AI services is that adjusting hyperparameters like temperature or token limits can fix output quality issues, when in fact prompt engineering techniques like few-shot learning are the primary non-retraining methods for improving model behavior.

How to eliminate wrong answers

Option A is wrong because reducing temperature to 0 makes the model deterministic but does not inherently fix code compilation errors; it only reduces randomness, which may still produce incorrect or incomplete code. Option B is wrong because increasing the max token limit allows longer outputs but does not address the root cause of compilation failures, such as syntax errors or logical mistakes. Option C is wrong because fine-tuning requires retraining the model on a dataset of correct code snippets, which contradicts the question's constraint of 'without retraining' and is not a prompt-level solution.

71
MCQeasy

A developer wants to test different prompt variations for a chatbot without making repeated API calls. Which Amazon Bedrock feature can help compare model responses?

A.Model evaluation on Amazon SageMaker
B.Amazon Bedrock Playground
C.AWS Security Token Service (STS)
D.Amazon CloudWatch Logs
AnswerB

The Playground provides an interactive console where multiple prompt variants can be run side by side and responses compared visually. This satisfies testing without repeated API calls, as experimentation happens within the interface rather than through programmatic invocation.

Why this answer

Amazon Bedrock Playground provides an interactive console interface where developers can test and compare different prompt variations, model configurations, and foundation models side-by-side without making repeated API calls. It allows real-time experimentation with parameters like temperature, top-p, and prompt engineering to observe how the model responds, making it the correct choice for this use case.

Exam trap

AWS certification exams often test the distinction between interactive experimentation tools (Playground) and backend evaluation or monitoring services (SageMaker, CloudWatch), leading candidates to mistakenly choose a service that handles model evaluation or logging rather than the one designed for real-time prompt comparison.

How to eliminate wrong answers

Option A is wrong because Model evaluation on Amazon SageMaker is a service for evaluating and monitoring the performance of machine learning models in production, not for interactively testing prompt variations in a chatbot context without API calls. Option C is wrong because AWS Security Token Service (STS) is used for generating temporary security credentials to manage access to AWS resources, not for comparing model responses or prompt testing. Option D is wrong because Amazon CloudWatch Logs is a monitoring and logging service for collecting and analyzing log data from AWS resources, not an interactive tool for testing prompt variations or comparing model outputs.

72
MCQeasy

A solutions architect is explaining the concept of a foundation model to a non-technical stakeholder. The stakeholder asks what distinguishes a foundation model from a traditional task-specific machine learning model. Which statement best describes a foundation model?

A.It is a large model pre-trained on broad data that can be adapted to many downstream tasks.
B.It is a model trained exclusively on a company's proprietary data for one specific use case.
C.It is a small model optimized to run only on edge devices with limited compute.
D.It is a rule-based system that follows hand-coded logic to produce deterministic outputs.
AnswerA

This is correct because foundation models are trained on vast, diverse datasets using self-supervised learning, giving them broad capabilities that transfer to many tasks through prompting or fine-tuning. This generality is the defining characteristic that separates them from narrow, single-purpose models, and it directly answers the stakeholder's question about what makes them distinct.

Why this answer

A foundation model is characterized by large-scale pre-training on broad data, which yields general capabilities that can be adapted to numerous downstream tasks via prompting, fine-tuning, or other techniques. This generality is what differentiates it from task-specific models built for a single narrow purpose, and it is the essence the architect needs to convey.

Exam trap

The trap here is assuming that a foundation model is defined by its size or deployment location rather than by its broad pre-training and adaptability across many tasks.

73
MCQhard

A developer wants a foundation model to reliably return a structured record with fixed fields for downstream processing. The model currently returns free-form prose that breaks parsing. Which technique most directly improves output reliability for this use case?

A.Define a JSON schema and use constrained decoding or structured output features so the model can only emit tokens valid under that schema
B.Increase the model's temperature so it generates more varied field values
C.Shorten the prompt so the model has fewer instructions to follow
D.Add the phrase 'respond in JSON' to the user prompt and retry on parse failures
AnswerA

Constrained decoding restricts the token sampling space to sequences that satisfy the schema, guaranteeing syntactically valid structured output that downstream parsers can consume. This directly targets the parsing failures caused by free-form prose, unlike prompt wording changes that merely encourage format adherence.

Why this answer

Structured output reliability comes from constraining decoding to a schema, which makes invalid tokens impossible rather than merely unlikely. Prompt-based requests for JSON and retry loops reduce but do not eliminate malformed output, raising temperature increases variability, and shortening the prompt removes guidance without adding enforcement. Schema-constrained generation is the direct fix.

Exam trap

The trap here is believing that telling a model to produce JSON in the prompt is the same as guaranteeing valid JSON.

74
Multi-Selecthard

A financial services firm is evaluating foundation models for a loan-summary assistant. They must consider both model characteristics and operational constraints. Which TWO factors most directly affect whether a candidate model can be deployed to meet their requirements? (Choose two.)

Select 2 answers
A.The popularity of the model among hobbyist developers on public forums
B.The model's maximum context window length relative to the size of the loan documents
C.The color scheme used in the model provider's marketing materials
D.Whether the model is available in the AWS Region where the firm must keep data
E.The number of parameters reported in the model's architecture diagram
AnswersB, D

Loan documents can be lengthy, and a model with a small context window cannot ingest the full text, forcing truncation or chunking that may drop critical clauses. Context length is therefore a hard feasibility constraint that determines whether the model can process the required input at all.

Why this answer

Deployability hinges on concrete constraints: whether the model can ingest the full loan document within its context window, and whether it is offered in the Region where data must remain. Parameter count, community popularity, and marketing presentation do not determine whether the model can satisfy the workload and compliance requirements.

Exam trap

The trap here is equating a model's general reputation or architecture size with its suitability for a specific regulated deployment.

75
Multi-Selecteasy

Which TWO are benefits of using Amazon SageMaker JumpStart for foundation models? (Choose 2)

Select 2 answers
A.Built-in fine-tuning scripts and notebooks
B.No coding required to fine-tune models
C.Automatic scaling without any configuration
D.Pre-trained foundation models available in the catalog
E.Free unlimited usage for all models
AnswersA, D

SageMaker JumpStart supplies pre-built fine-tuning scripts and notebooks, letting teams adapt foundation models without authoring training code from scratch. This directly satisfies the stem's benefit requirement by reducing implementation effort, since the notebooks run in SageMaker environments and expose hyperparameters for customisation.

Why this answer

Option A is correct because SageMaker JumpStart provides built-in fine-tuning scripts and example notebooks that let you adapt foundation models to your own data without writing the training pipeline from scratch. Option D is correct because JumpStart includes a catalog of pre-trained foundation models (e.g., from providers like Hugging Face, Meta, and AI21) that you can deploy or customize directly. Option B is not correct because fine-tuning still requires code or configuration, such as selecting hyperparameters and preparing datasets, so it is not fully no-code.

Option C is not correct because automatic scaling still requires configuring endpoint settings like instance counts and autoscaling policies. Option E is not correct because model usage in JumpStart is billed through SageMaker resources and is not free or unlimited.

Exam trap

AWS exams often test the misconception that 'no-code' solutions like SageMaker JumpStart eliminate all coding, when in reality they still require scripting for customization, and that built-in features like pre-trained models and scripts are distinct from automatic scaling or free usage.

Page 1 of 2 · 125 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Fundamentals of Generative AI questions.