Courseiva

CCNA Fundamentals of Generative AI Questions

50 of 125 questions · Page 2/2 · Fundamentals of Generative AI · Answers revealed

76
Multi-Selecthard

A financial services firm is deploying a generative AI assistant that answers employee questions about internal policy documents. The security team requires that answers be traceable to source text and that the model not invent policy details. Which TWO techniques should be implemented to ground responses and reduce fabricated content? (Choose two.)

Select 2 answers
A.Instruct the model to answer only from the supplied context and to state when the context is insufficient
B.Increase the temperature setting so the model explores a wider range of responses
C.Retrieve relevant passages from the policy corpus and include them in the prompt before generation
D.Remove all system instructions so the model can respond more naturally to each question
E.Expand the model's parameter count by switching to the largest available variant
AnswersA, C

Grounding instructions constrain the model to the retrieved material and give it an explicit escape hatch when evidence is missing. That behavior converts a potential hallucination into an honest abstention, which is exactly what the security team wants, and it pairs naturally with retrieval so the model has context to cite.

Why this answer

Grounding comes from combining retrieval of authoritative passages with instructions that confine the model to that evidence and permit abstention. Together they make answers traceable and suppress invention. Higher temperature, larger parameter counts, and removal of system instructions either increase variability, fail to supply evidence, or discard the constraints that keep output faithful.

Exam trap

The trap here is treating a bigger model as a fix for hallucination, when grounding requires supplying evidence and constraining the model to it rather than scaling parameters.

77
MCQeasy

A content team wants to build an internal tool that drafts blog posts from short outlines. They have no machine learning engineers on staff and want AWS to manage the underlying model infrastructure while they focus only on prompts and output. Which AWS service should they use?

A.Amazon SageMaker AI
B.Amazon Polly
C.Amazon Bedrock
D.Amazon Rekognition
AnswerC

Amazon Bedrock is a fully managed service that exposes foundation models from multiple providers through a single API, so a team with no ML engineers can send prompts and receive generated text without provisioning or tuning any infrastructure. It fits the drafting use case directly, since the team only needs prompt design and output handling rather than model hosting.

Why this answer

The team needs managed access to generative foundation models with minimal operational overhead, and Amazon Bedrock provides exactly that by serving multiple providers' models through one API without requiring the team to train or host anything. SageMaker AI offers more control but demands ML engineering, while Polly and Rekognition serve speech and vision tasks respectively.

Exam trap

The trap here is assuming that any AWS AI service can generate text, when several popular ones such as Polly and Rekognition are specialized for speech or image tasks only.

78
MCQmedium

A company wants to build a customer support chatbot that answers questions based on a large internal knowledge base. Which AWS service is most suitable for implementing RAG to retrieve relevant documents?

A.Amazon Lex
B.Amazon Polly
C.Amazon Connect
D.Amazon Kendra
AnswerD

Amazon Kendra is an ML-powered enterprise search service that indexes internal document repositories and returns semantically relevant passages, which the chatbot feeds to the model as RAG context. It satisfies the requirement for retrieving relevant documents from a large knowledge base.

Why this answer

Amazon Kendra is an intelligent enterprise search service powered by machine learning that is specifically designed for retrieving relevant documents from large knowledge bases, making it ideal for RAG implementations. It supports natural language queries and returns precise answers with citations, which is exactly what a customer support chatbot needs to retrieve relevant documents.

Exam trap

AIF-C01 often tests the confusion between services that build chatbots (Lex) and services that retrieve documents (Kendra), and candidates may pick Lex because it sounds like a chatbot service.

How to eliminate wrong answers

Option A is wrong because Amazon Lex is a service for building conversational interfaces (chatbots) using voice and text, but it does not provide document retrieval capabilities for RAG. Option B is wrong because Amazon Polly is a text-to-speech service, not a document retrieval service. Option C is wrong because Amazon Connect is a cloud contact center service, not a document retrieval or search service.

79
MCQeasy

A developer is building a customer-facing chatbot using Amazon Bedrock. To ensure the chatbot does not generate offensive or inappropriate content, which AWS feature should they implement?

A.AWS Identity and Access Management (IAM) policies
B.Amazon Bedrock Guardrails
C.Prompt engineering with system prompts
D.Increasing the model temperature parameter
AnswerB

Amazon Bedrock Guardrails applies configurable content filters and denied-topic policies that intercept harmful prompts and responses at inference time. This directly satisfies the requirement to stop offensive or inappropriate output in a customer-facing chatbot, independent of the underlying foundation model.

Why this answer

Amazon Bedrock Guardrails is the correct choice because it provides configurable safeguards that allow developers to define denied topics, content filters (e.g., hate, insults, sexual content), and sensitive information filters to prevent the model from generating offensive or inappropriate responses. Unlike prompt engineering or parameter tuning, Guardrails enforce policy-based constraints at inference time, independent of the underlying model's behavior.

Exam trap

A common pitfall in this exam is assuming that prompt engineering with system prompts or adjusting model parameters like temperature can reliably block offensive content. While these techniques can influence behavior, they do not provide enforceable, policy-based safeguards. Amazon Bedrock Guardrails must be used to define and enforce content filters, denied topics, and sensitive information filters at inference time.

How to eliminate wrong answers

Option A is wrong because IAM policies control access to AWS resources and API actions, not the content generated by a model; they cannot filter or block specific output text. Option C is wrong because prompt engineering with system prompts can guide model behavior but is not a guaranteed or enforceable mechanism—models can still bypass or ignore system instructions, especially under adversarial inputs. Option D is wrong because increasing the temperature parameter increases randomness in token selection, which may actually increase the likelihood of generating inappropriate content, not reduce it.

80
MCQmedium

A developer is building a chatbot using Amazon Bedrock and Claude. They notice that the model sometimes generates harmful or biased responses. Which AWS service can they use to implement guardrails?

A.AWS WAF
B.Amazon GuardDuty
C.AWS Shield
D.Amazon Bedrock Guardrails
AnswerD

Amazon Bedrock Guardrails applies configurable content filters, denied topics and word filters directly to model inference, blocking harmful or biased outputs before they reach users. It satisfies the stem's requirement for guardrails on a Bedrock-hosted Claude chatbot, unlike standalone moderation services that would need custom integration outside the Bedrock invocation flow.

Why this answer

Amazon Bedrock Guardrails is the correct choice because it is a native feature of Amazon Bedrock designed specifically to implement safety controls, content filters, and topic policies for foundation models like Claude. It allows developers to define denied topics, filter harmful content (e.g., hate speech, violence), and redact sensitive information, directly addressing the need to prevent harmful or biased responses in a chatbot built on Bedrock.

Exam trap

The trap here is that candidates may confuse AWS security services (WAF, GuardDuty, Shield) with AI-specific safety mechanisms, assuming any 'guard' or 'shield' service can filter model outputs, when only Amazon Bedrock Guardrails is purpose-built for content safety in generative AI.

How to eliminate wrong answers

Option A is wrong because AWS WAF is a web application firewall that protects HTTP/HTTPS APIs from common web exploits like SQL injection and cross-site scripting, not a service for implementing content guardrails on generative AI model outputs. Option B is wrong because Amazon GuardDuty is a threat detection service that monitors for malicious activity and unauthorized behavior in AWS accounts and workloads, not a tool for filtering or controlling the responses of a large language model. Option C is wrong because AWS Shield is a managed Distributed Denial of Service (DDoS) protection service that safeguards applications against DDoS attacks, and it has no capability to enforce safety policies or bias filters on AI-generated content.

81
Multi-Selecthard

A healthcare analytics team is evaluating whether a generative AI solution is appropriate for summarizing patient intake notes into structured clinical fields. They are concerned about the model producing confident but incorrect medical details. Which TWO practices best reduce the risk of fabricated content in this scenario? (Choose two.)

Select 2 answers
A.Ask the model to state its confidence as a percentage after each extracted field and trust fields scored above ninety percent.
B.Instruct the model to respond with a specific phrase such as 'not stated' when the source note does not contain a requested field.
C.Ground each summarization request by supplying the relevant intake note text in the prompt instead of relying on the model's memory.
D.Fine-tune the model on the entire historical patient record database to expand its clinical knowledge.
E.Raise the temperature to its maximum so the model produces diverse clinical phrasing across repeated runs.
AnswersB, C

Explicitly defining an abstention response gives the model a valid way to avoid filling gaps with plausible guesses. Since hallucination often arises when a field is missing and the model is expected to answer anyway, allowing an 'unknown' output removes the pressure to fabricate and makes downstream review of missing data straightforward.

Why this answer

Fabrication drops when the model is anchored to the actual source text and is given permission to report that information is absent. Supplying the intake note in the prompt keeps generation evidence-based, and defining an explicit abstention phrase prevents the model from inventing values for missing fields. Randomness and self-reported confidence do not reliably reduce hallucination.

Exam trap

The trap here is treating a model's self-reported confidence score as a trustworthy signal, when that number is generated text and is not a calibrated probability.

82
MCQhard

A developer is explaining how a large language model generates text so that a business stakeholder understands why the same prompt can yield different answers. Which description accurately captures the generation process?

A.The model queries an external search index and paraphrases the top-ranked web page for the prompt.
B.The model searches a stored database of previously written sentences and returns the closest exact match to the prompt.
C.The model predicts a probability distribution over the next token, samples from it, and appends the chosen token to repeat the process.
D.The model compiles the prompt into a set of if-then rules that deterministically map keywords to canned responses.
AnswerC

Generation is autoregressive: given the prompt and tokens produced so far, the model outputs a probability distribution over the vocabulary, a token is selected according to the sampling strategy, and that token is appended before the next prediction. Because sampling is stochastic, the same prompt can produce different completions, which explains the variability the stakeholder observed.

Why this answer

Text generation is autoregressive and probabilistic. At each step the model computes a distribution over the next token conditioned on everything before it, a token is sampled according to settings such as temperature and top-p, and the sequence grows one token at a time. Stochastic sampling is why identical prompts can yield different outputs, and why temperature changes affect variability.

Exam trap

The trap here is assuming the model stores and retrieves whole sentences, when it actually predicts one token at a time from probability distributions over its vocabulary.

83
MCQhard

Refer to the exhibit. A developer receives an error when trying to invoke the Claude Instant model from an application. The application uses the IAM role 'MyAppRole'. Which IAM policy statement should be added to the role to resolve the error?

A.{"Effect":"Allow","Action":"bedrock:GetFoundationModel","Resource":"*"}
B.{"Effect":"Allow","Action":"bedrock:InvokeModel","Resource":"arn:aws:bedrock:us-east-1::foundation-model/*"}
C.{"Effect":"Allow","Action":"bedrock:InvokeModel","Resource":"arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-instant-v1"}
D.{"Effect":"Allow","Action":"bedrock:*","Resource":"*"}
AnswerC

Bedrock authorises model invocation through the bedrock:InvokeModel action scoped to the specific foundation-model ARN. The role currently lacks this permission, so granting it on the Claude Instant ARN resolves the AccessDenied error without over-permissioning unrelated models or regions.

Why this answer

The error occurs when the application tries to invoke the Claude Instant model via the Bedrock InvokeModel API, and the IAM role 'MyAppRole' lacks the necessary permission. The required action is 'bedrock:InvokeModel', and the resource must be scoped to the specific foundation model ARN, which for Claude Instant in us-east-1 is 'arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-instant-v1'. This grants least-privilege access to invoke exactly that model.

Exam trap

The trap here is that candidates often confuse the 'InvokeModel' action with 'GetFoundationModel' or choose a wildcard resource ARN, thinking it is sufficient, but AWS specifically tests the need for a precise model ARN to enforce least-privilege access in Bedrock IAM policies.

How to eliminate wrong answers

Option A is wrong because 'bedrock:GetFoundationModel' is a read-only action used to retrieve metadata about a foundation model, not to invoke it for inference; the application needs the 'InvokeModel' action. Option B is wrong because while it allows 'bedrock:InvokeModel', the resource ARN uses a wildcard 'foundation-model/*', which grants access to all foundation models in the region, violating the principle of least privilege and potentially allowing unintended model invocations. Option D is wrong because 'bedrock:*' grants full administrative access to all Bedrock actions, which is overly permissive and not a best practice for a specific application role; it would also allow actions like model creation or deletion.

84
MCQhard

Refer to the exhibit. A developer is optimizing latency for a generative AI model deployed on SageMaker. Based on the exhibit, which change would most likely reduce per-token latency?

A.Use a CPU instance
B.Reduce model size through quantization
C.Switch to a larger instance type
D.Increase batch size to 10
AnswerB

Quantization reduces weight precision, shrinking the model so each forward pass performs fewer and cheaper memory-bound operations per token. This directly lowers per-token latency, satisfying the exhibit's constraint of reducing inference time without retraining or changing the endpoint's instance type.

Why this answer

Reducing model size through quantization directly decreases the computational and memory requirements per inference step, which lowers the time to generate each token. This is especially effective on GPU instances where smaller models fit better in GPU memory and reduce memory bandwidth bottlenecks, leading to lower per-token latency.

Exam trap

Candidates often think that larger instances always reduce latency, when in fact they may increase latency due to higher memory latency and inter-chip communication, while quantization directly addresses the memory bandwidth bottleneck in autoregressive decoding.

How to eliminate wrong answers

Option A is wrong because CPU instances lack the parallel processing capabilities needed for efficient generative AI inference, resulting in significantly higher per-token latency compared to GPU instances. Option C is wrong because switching to a larger instance type may increase throughput but does not necessarily reduce per-token latency; it can even increase latency due to higher memory access times and inter-chip communication overhead. Option D is wrong because increasing batch size to 10 increases the total tokens processed per batch, which can improve throughput but typically increases per-token latency due to longer queueing and processing times for each batch.

85
MCQhard

An enterprise wants to ensure that generative AI applications built on AWS comply with data privacy regulations. They need to prevent the model from using customer data in future training. Which feature of Amazon Bedrock should they enable?

A.Policy-based data governance
B.Opt-out of model improvement
C.Data encryption at rest
D.Model customization with customer data
AnswerB

Opting out of model improvement prevents Amazon from using the customer's inputs and outputs to train or refine foundation models. This satisfies the data privacy requirement by contractually and technically excluding customer data from future training.

Why this answer

Amazon Bedrock's opt-out of model improvement feature allows customers to prevent AWS from using their data (including prompts, completions, and associated metadata) for model training or service improvement. This is essential for compliance with data privacy regulations like GDPR or CCPA, as it ensures customer data is not retained or used beyond the immediate inference request.

Exam trap

Candidates often confuse data protection mechanisms (encryption, access control) with data usage controls (opt-out of model improvement) on Amazon Bedrock.

How to eliminate wrong answers

Option A is wrong because policy-based data governance (e.g., using AWS Lake Formation or IAM policies) controls access and permissions to data, but does not prevent the model provider from using customer data for future training. Option C is wrong because data encryption at rest (e.g., using AWS KMS or SSE) protects data confidentiality during storage but has no effect on whether the model uses that data for training. Option D is wrong because model customization with customer data (e.g., fine-tuning or continued pre-training) explicitly involves using customer data to improve the model, which is the opposite of preventing data usage for training.

86
Multi-Selecthard

Which THREE considerations are essential when deploying a generative AI application in a regulated industry such as healthcare?

Select 3 answers
A.Lowest possible inference latency for real-time responses.
B.Full audit trail of model inputs and outputs for accountability.
C.Robust content filtering to block harmful or inaccurate outputs.
D.Maximum creative freedom for the model to generate diverse responses.
E.Data privacy and compliance with regulations like HIPAA.
AnswersB, C, E

Healthcare regulators require traceability of every AI-driven decision. Logging each model input and output creates the audit trail needed for accountability, incident investigation and post-market surveillance, satisfying the stem's regulated-industry constraint that opaque inference cannot meet.

Why this answer

In a regulated industry such as healthcare, option B (full audit trail of model inputs and outputs) is essential because accountability and traceability are required to investigate incidents, demonstrate compliance during audits, and reconstruct how a clinical or patient-facing decision was produced. Option C (robust content filtering) is essential because generative models can hallucinate or emit harmful advice, and filtering/guardrails are needed to prevent unsafe or inaccurate outputs from reaching patients or clinicians. Option E (data privacy and HIPAA compliance) is essential because healthcare data is protected health information, so the deployment must enforce safeguards such as access controls, encryption, and Business Associate Agreements to avoid regulatory violations.

Option A is not essential in this context because lowest possible inference latency is a performance optimization, not a regulatory or safety requirement, and accuracy, privacy, and auditability take precedence. Option D is not appropriate because maximum creative freedom increases variability and hallucination risk, which directly conflicts with the determinism, safety, and compliance needs of regulated healthcare use.

Exam trap

The trap here is that candidates may prioritize performance metrics like latency (Option A) over compliance requirements, mistakenly assuming that speed is always critical in healthcare, whereas AWS services in regulated industries must prioritize data privacy and auditability as non-negotiable, as mandated by regulations like HIPAA.

87
MCQhard

A team is developing a real-time code completion feature using an LLM deployed on Amazon SageMaker. They observe high latency under load. Which optimization technique should they prioritize?

A.Increase batch size
B.Switch to a larger instance type
C.Increase instance count with Auto Scaling
D.Use model quantization
AnswerD

Quantization reduces weight precision, shrinking model size and memory bandwidth demand, which lowers per-token inference latency and increases throughput under concurrent load. This directly addresses the high-latency constraint for real-time code completion without retraining or architectural changes.

Why this answer

Model quantization reduces the precision of the model's weights (e.g., from FP32 to INT8), which decreases memory footprint and computational requirements, leading to lower latency per inference. This is the most effective optimization for real-time code completion because it directly reduces the time to generate each token without requiring additional infrastructure changes.

Exam trap

The AWS exam often tests the distinction between throughput optimization (batch size, scaling) and latency optimization (quantization, pruning), and the trap here is assuming that scaling out or increasing resources always solves latency issues, when in fact the bottleneck is per-inference computation time.

How to eliminate wrong answers

Option A is wrong because increasing batch size increases throughput but also increases latency for individual requests, which is counterproductive for real-time code completion that requires low per-request latency. Option B is wrong because switching to a larger instance type may reduce latency but at a higher cost and without addressing the fundamental computational bottleneck; it is a brute-force approach rather than an optimization. Option C is wrong because increasing instance count with Auto Scaling improves scalability and handles more concurrent requests, but it does not reduce the latency of individual inference calls, which is the core issue under load.

88
MCQmedium

A company is using Amazon Bedrock to summarize long documents. They notice that the summary sometimes omits key details. What is the most likely cause?

A.The model is overfitted
B.The prompt lacks examples
C.The model's context window is too small
D.The temperature parameter is too high
AnswerC

Summarisation requires the whole document plus the prompt to fit within the model's context window. When the input exceeds that token limit, content is truncated before inference, so later sections never reach the model and their details are omitted from the summary.

Why this answer

When summarizing long documents with Amazon Bedrock, the model's context window determines the maximum amount of text it can process at once. If the document exceeds this limit, the model truncates or ignores portions, leading to omitted key details. This is the most likely cause because summarization requires the model to attend to the entire input, and a small context window directly prevents full coverage.

Exam trap

AWS often tests the distinction between model capacity limits (context window) and output quality parameters (temperature, prompt engineering), leading candidates to incorrectly attribute omission errors to randomness or lack of examples rather than the fundamental constraint of input size.

How to eliminate wrong answers

Option A is wrong because overfitting refers to a model memorizing training data and failing to generalize, which is unrelated to truncation or omission of details in a summarization task. Option B is wrong because while prompt examples (few-shot learning) can improve output quality, their absence does not cause the model to physically ignore parts of the input; the core issue is capacity, not prompting technique. Option D is wrong because a high temperature parameter increases randomness and creativity in generation, which might cause irrelevant or verbose output, but it does not cause the model to skip or omit key details from the input.

89
MCQmedium

A support organization notices its generative AI assistant sometimes invents policy details that do not exist. Leadership wants a measurable way to track whether this behavior improves over time. Which action should the team take?

A.Switch to a model with a larger parameter count and assume the fabrication problem disappears.
B.Build a labeled evaluation set of questions with verified answers and score responses for factual consistency on a recurring basis.
C.Increase the model's maximum output token limit so it has more room to explain itself.
D.Rely on customer complaints as the primary signal for whether fabricated policy details are decreasing.
AnswerB

A curated evaluation set with known-correct answers turns an anecdotal complaint into a repeatable metric, allowing the team to compare prompts, retrieval strategies, or models over time. This directly addresses the request for a measurable way to track improvement in factual accuracy and supports regression testing as the system evolves.

Why this answer

Turning an observed quality problem into a tracked metric requires a fixed evaluation set with verified answers and a scoring procedure that can be rerun after each change. This makes comparisons across prompts, retrieval configurations, and models meaningful, and it establishes a regression baseline so the organization can prove whether fabrication is actually declining.

Exam trap

The trap here is assuming that newer, larger models or more output space automatically fix hallucination, when without a measurement framework no one can tell whether the problem improved.

90
MCQmedium

A financial services company wants to deploy a generative AI assistant that answers questions using its internal policy documents. The team is concerned that the model may invent policies that do not exist. Which approach best reduces this risk while keeping the model's answers tied to the company's actual documents?

A.Switch to a larger foundation model with more parameters to improve factual recall.
B.Use retrieval augmented generation to fetch relevant document passages and include them in the prompt.
C.Lower the maximum token limit so the model produces shorter answers with fewer chances to err.
D.Increase the model's temperature setting so it explores more diverse phrasings.
AnswerB

Retrieval augmented generation retrieves relevant passages from the company's own document store and supplies them as context, so the model generates answers conditioned on real policy text. This grounds responses in verifiable source material and sharply reduces fabricated content, directly addressing the concern about invented policies while keeping answers tied to actual documents.

Why this answer

Retrieval augmented generation grounds the model by supplying relevant passages from the company's own document repository at inference time. Because the model conditions its answer on retrieved, verifiable content, it is far less likely to invent policies. This directly satisfies the requirement that answers stay tied to actual internal documents rather than relying on the model's parametric memory.

Exam trap

The trap here is assuming that a bigger model or shorter answers will fix fabrication, when the real fix is supplying authoritative source content at inference time.

91
MCQeasy

A data scientist wants to fine-tune a foundation model on a specific domain dataset using Amazon SageMaker. Which built-in SageMaker feature can simplify the training process?

A.SageMaker Neo
B.SageMaker Canvas
C.SageMaker JumpStart
D.SageMaker Ground Truth
AnswerC

SageMaker JumpStart provides pre-trained foundation models with ready-made notebooks and one-click deployment, removing much of the manual configuration involved in fine-tuning. This simplifies domain-specific training by supplying the model artefacts and training scaffolding out of the box.

Why this answer

SageMaker JumpStart provides pre-trained foundation models and built-in training scripts that simplify fine-tuning on custom datasets. It handles the underlying infrastructure, hyperparameter configurations, and model deployment, allowing the data scientist to focus on the domain-specific data rather than writing custom training loops or managing SageMaker training jobs manually.

Exam trap

The trap in this question is that candidates may confuse SageMaker services that are related to aspects of ML workflows but not specifically designed for fine-tuning foundation models. SageMaker Neo handles model optimization for deployment, SageMaker Canvas is a no-code tool for building ML models without writing code, and SageMaker Ground Truth is for data labeling. Only SageMaker JumpStart provides pre-trained models and built-in training scripts suitable for fine-tuning.

How to eliminate wrong answers

Option A is wrong because SageMaker Neo is a model optimization and compilation service for deploying models on edge devices, not a tool for fine-tuning foundation models. Option B is wrong because SageMaker Canvas is a no-code visual interface for building ML models using pre-built algorithms and does not support fine-tuning of foundation models. Option D is wrong because SageMaker Ground Truth is a data labeling service used to create high-quality training datasets, not a feature for training or fine-tuning models.

92
MCQeasy

A developer is using Amazon Bedrock to create a chatbot. They want to ensure the bot does not generate toxic or offensive content. Which feature should they enable?

A.Use careful prompt engineering to avoid toxic responses.
B.Fine-tune the model on a dataset of safe responses.
C.Enable content filtering on the Bedrock model.
D.Implement external response validation using a third-party API.
AnswerC

Content filtering applies configurable thresholds that block or mask harmful categories such as hate, violence and sexual content in both prompts and responses. Enabling it on the Bedrock model directly satisfies the requirement to prevent toxic output from the chatbot.

Why this answer

Amazon Bedrock provides built-in content filtering capabilities that can be enabled at the model invocation level to automatically detect and block toxic or offensive content in both input prompts and generated responses. This feature uses predefined safety filters (e.g., hate, insults, sexual content, violence) and is the most direct and managed way to prevent harmful outputs without requiring custom development.

Exam trap

A common misconception is that prompt engineering alone is sufficient for safety, when in fact Bedrock's content filtering is the explicit, managed feature designed to enforce content policies at runtime.

How to eliminate wrong answers

Option A is wrong because careful prompt engineering can reduce but not guarantee the elimination of toxic responses, as the underlying model may still generate harmful content due to its training data or adversarial inputs. Option B is wrong because fine-tuning the model on a dataset of safe responses requires significant data preparation, cost, and expertise, and it does not provide a runtime guard against all toxic outputs, especially for edge cases. Option D is wrong because implementing external response validation using a third-party API adds latency, complexity, and potential cost, and it is not a native Bedrock feature; Bedrock already offers content filtering as a first-class, integrated service.

93
MCQeasy

A startup wants to generate product descriptions from a few keywords using a foundation model. They need a fully managed serverless solution that requires no infrastructure setup. Which AWS service should they use?

A.Amazon SageMaker
B.Amazon Comprehend
C.AWS Lambda
D.Amazon Bedrock
AnswerD

Amazon Bedrock provides serverless access to foundation models through a single API, so the startup generates product descriptions from keywords without provisioning or managing any infrastructure. It satisfies the stem's fully managed, no-setup constraint directly, unlike self-hosted alternatives requiring instance or endpoint management.

Why this answer

Amazon Bedrock is a fully managed serverless service that provides access to foundation models (FMs) from leading AI providers via a simple API, making it ideal for generating product descriptions from keywords without any infrastructure management. It directly supports generative AI tasks like text generation, unlike other AWS services that focus on different ML or NLP capabilities.

Exam trap

The trap here is that candidates may confuse Amazon SageMaker's managed ML capabilities with a serverless generative AI service, overlooking that SageMaker requires explicit infrastructure setup for model hosting, while Bedrock is purpose-built for serverless access to foundation models.

How to eliminate wrong answers

Option A is wrong because Amazon SageMaker is a fully managed machine learning platform that requires setting up training jobs, endpoints, and infrastructure for custom models, not a serverless solution for directly using pre-built foundation models. Option B is wrong because Amazon Comprehend is a natural language processing (NLP) service for tasks like sentiment analysis and entity extraction, not for generative text creation from keywords. Option C is wrong because AWS Lambda is a serverless compute service that runs custom code but does not natively provide access to foundation models; you would need to integrate it with another service like Bedrock to generate descriptions, making it not a standalone solution for this use case.

94
MCQeasy

A developer is testing different prompts for a text generation model on Amazon Bedrock. Which parameter controls the randomness of the model's output?

A.top_p
B.stop_sequences
C.temperature
D.max_tokens
AnswerC

Temperature directly scales the sampling distribution's entropy: low values sharpen probabilities toward the highest-scoring token, high values flatten them, increasing output variability. This is precisely the randomness control the developer needs when testing prompts on Amazon Bedrock, unlike top-p, which truncates the candidate token set rather than rescaling probabilities.

Why this answer

Temperature directly controls the randomness of the model's output by scaling the logits before applying the softmax function. A higher temperature (e.g., 1.5) increases randomness and creativity, while a lower temperature (e.g., 0.1) makes the output more deterministic and focused. In Amazon Bedrock, this parameter is a core setting for text generation models like Anthropic Claude and Amazon Titan.

Exam trap

Candidates often confuse top_p (nucleus sampling) with randomness control, when temperature is the primary parameter for that purpose. Both affect output diversity but through different mechanisms.

How to eliminate wrong answers

Option A is wrong because top_p controls nucleus sampling, which limits the cumulative probability of token choices, not the overall randomness of the output. Option B is wrong because stop_sequences define specific strings that halt generation, not randomness. Option D is wrong because max_tokens sets the maximum number of tokens in the generated response, not the randomness of the output.

95
MCQhard

An organization is using Amazon Bedrock to power a customer service chatbot. They notice that the chatbot occasionally generates hallucinated information about product specifications. Which strategy should be implemented to reduce hallucinations?

A.Fine-tune the model on a dataset of product specification conversations.
B.Integrate a Retrieval Augmented Generation (RAG) system with the product catalog.
C.Use more detailed prompts with explicit instructions to avoid speculation.
D.Increase the temperature parameter to make outputs more conservative.
AnswerB

Integrating RAG grounds responses in retrieved product-catalogue content, so the model conditions on factual specifications rather than relying solely on parametric memory. This directly targets the hallucination source by supplying authoritative context at inference time, satisfying the requirement to reduce fabricated product details without retraining the foundation model.

Why this answer

Retrieval Augmented Generation (RAG) grounds the model's responses in authoritative, up-to-date product catalog data, directly reducing hallucinations by ensuring the chatbot references verified facts rather than relying solely on its parametric memory. This is the most effective strategy because it provides a retrieval-based factual foundation that fine-tuning or prompt engineering alone cannot guarantee.

Exam trap

The AIF-C01 exam often tests the misconception that prompt engineering or fine-tuning alone can solve hallucination problems, when in fact they lack the dynamic, verifiable grounding that RAG provides.

How to eliminate wrong answers

Option A is wrong because fine-tuning on product specification conversations may reinforce patterns from the training data but does not prevent the model from generating plausible-sounding but incorrect details when faced with queries outside the fine-tuned distribution; it also cannot dynamically incorporate real-time catalog updates. Option C is wrong because while more detailed prompts can reduce speculation, they do not provide the model with access to external, authoritative data—hallucinations can still occur when the model's internal knowledge is incomplete or outdated. Option D is wrong because increasing the temperature parameter makes outputs more random and creative, not more conservative; decreasing temperature would make outputs more deterministic and less prone to hallucination, but even low temperature cannot eliminate hallucinations without a retrieval mechanism.

96
MCQeasy

A developer wants to test different foundation models quickly without setting up infrastructure. Which AWS service allows interactive prompting and comparison of multiple models?

A.Amazon Comprehend
B.Amazon Bedrock Playground
C.Amazon Lex
D.Amazon SageMaker Studio
AnswerB

Amazon Bedrock Playground provides a console interface for interactive prompting and side-by-side comparison of multiple foundation models without provisioning any infrastructure. This directly satisfies the stem's requirement to test different models quickly, since no servers, endpoints or deployment configuration are needed before experimentation.

Why this answer

Amazon Bedrock Playground is a feature within Amazon Bedrock that provides a web-based interface for interactive prompting and side-by-side comparison of multiple foundation models (FMs). It allows developers to test different models quickly without provisioning any infrastructure, making it ideal for rapid experimentation and evaluation.

Exam trap

The trap here is that candidates may confuse Amazon Bedrock Playground with SageMaker Studio, assuming both are for model experimentation, but SageMaker Studio requires infrastructure setup and lacks the built-in multi-model comparison interface that Bedrock Playground provides.

How to eliminate wrong answers

Option A is wrong because Amazon Comprehend is a natural language processing (NLP) service for extracting insights like sentiment, entities, and key phrases from text; it does not support interactive prompting or comparison of foundation models. Option C is wrong because Amazon Lex is a service for building conversational interfaces (chatbots) using automatic speech recognition (ASR) and natural language understanding (NLU), not for testing or comparing foundation models. Option D is wrong because Amazon SageMaker Studio is an integrated development environment (IDE) for building, training, and deploying machine learning models, but it requires setting up infrastructure (e.g., instances, kernels) and does not provide a built-in interactive playground for comparing multiple foundation models.

97
MCQeasy

A startup is building a customer support chatbot using Amazon Bedrock with the Claude foundation model. The chatbot needs to answer questions based on a knowledge base of frequently asked questions (FAQs) stored in an Amazon S3 bucket. The team wants to implement Retrieval Augmented Generation (RAG) to provide accurate and context-aware responses. They are evaluating different approaches to integrate the knowledge base. What is the most efficient way to implement RAG with Bedrock?

A.Use AWS Lambda to fetch documents from S3 and inject them into the prompt.
B.Manually extract all FAQs and include them in the prompt each time the chatbot responds.
C.Fine-tune the Claude model on the FAQs so the model memorizes the knowledge base.
D.Use Amazon Bedrock Knowledge Bases to directly connect the S3 bucket and retrieve relevant documents for the prompt.
AnswerD

Bedrock Knowledge Bases handles the full RAG pipeline natively: it ingests the S3 bucket, chunks and embeds documents, stores vectors, and retrieves relevant passages into the prompt. This removes the need to build custom chunking, embedding, and vector-store orchestration, making it the most efficient route.

Why this answer

Amazon Bedrock Knowledge Bases is a fully managed RAG capability that natively connects to Amazon S3 (and other data sources), automatically chunks and embeds documents into a vector store, and retrieves the most relevant passages at query time to augment the prompt. This eliminates the need to build custom retrieval pipelines, manage embeddings, or manually inject documents. It is the most efficient, purpose-built approach for RAG on Bedrock.

Exam trap

AIF-C01 often tests the misconception that fine-tuning is the way to inject factual knowledge into a model, when in fact RAG (via Bedrock Knowledge Bases) is the correct pattern for dynamic, grounded question answering.

How to eliminate wrong answers

Option A is wrong because using Lambda to fetch documents from S3 and inject them into the prompt is a custom, manual retrieval approach that lacks semantic search, chunking, and vector indexing — it is inefficient and unscalable. Option B is wrong because manually including all FAQs in every prompt is impractical, exceeds context window limits, increases cost, and does not scale as the knowledge base grows. Option C is wrong because fine-tuning teaches the model style and patterns, not factual retrieval; it cannot reliably memorize a dynamic knowledge base and risks hallucination and stale information.

98
Multi-Selectmedium

Which TWO factors are most important when selecting a foundation model in Amazon Bedrock for a text summarization task with strict latency requirements?

Select 2 answers
A.Average response latency per request.
B.Model size in billions of parameters.
C.Maximum input token limit.
D.Output quality and token efficiency for summarization tasks.
E.Availability of fine-tuning capability for domain adaptation.
AnswersA, D

Average response latency per request directly measures whether a candidate model can satisfy the stem's strict latency requirement, since Bedrock models vary widely in inference speed. Selecting on this metric ensures the summarisation workload meets its response-time constraint rather than optimising quality alone.

Why this answer

Option A is correct because with strict latency requirements the deciding metric is the model's average response latency per request, which directly determines whether the summarization workload can meet its service-level objective. Option D is correct because output quality and token efficiency determine whether the summary is usable and how many tokens must be generated, and fewer generated tokens translate directly into lower end-to-end latency. Option B is not decisive on its own: parameter count correlates loosely with latency but does not guarantee it, since architecture, hardware, and serving optimizations also matter.

Option C, the maximum input token limit, matters for handling long documents but does not address the strict latency constraint. Option E, fine-tuning availability, supports domain adaptation and accuracy rather than latency, so it is not one of the two most important factors here.

Exam trap

A common misconception is that model size (parameters) is the primary driver of latency, but in practice, latency depends on inference optimization, model quantization, and hardware, not just parameter count.

99
Multi-Selecteasy

Which TWO actions can help reduce the likelihood of hallucinations in a generative AI model used for question answering?

Select 2 answers
A.Increase the maximum token count to allow more complete answers.
B.Use Retrieval Augmented Generation (RAG) with a trusted knowledge base.
C.Fine-tune the model on the training data used for the application.
D.Set a lower temperature parameter (e.g., 0.1) to reduce randomness.
E.Use a larger foundation model with more parameters.
AnswersB, D

RAG grounds generation in retrieved passages from a trusted knowledge base, so answers are conditioned on verifiable source text rather than parametric memory alone. This directly reduces fabrication because the model cites retrieved evidence, satisfying the requirement to lower hallucination likelihood in question answering.

Why this answer

Option B is correct because Retrieval Augmented Generation (RAG) grounds the model's responses in documents retrieved from a trusted knowledge base, so answers are conditioned on verifiable source content rather than the model's parametric memory, which substantially reduces fabricated or unsupported statements. Option D is correct because lowering the temperature parameter (e.g., to 0.1) sharpens the next-token probability distribution, making the model select high-probability, more deterministic tokens instead of sampling low-probability alternatives that often produce invented facts. Option A does not belong because increasing the maximum token count only allows longer outputs; it does not improve factual grounding and can even give hallucinations more room to expand.

Option C does not belong because fine-tuning on the application's training data can reinforce patterns and biases in that data and does not guarantee factual accuracy, and may even increase confident hallucination. Option E does not belong because a larger foundation model with more parameters may be more fluent but is not inherently less prone to hallucination, since scale alone does not provide source grounding or reduce sampling randomness.

Exam trap

AWS often tests the misconception that simply increasing model size or output length improves answer quality, when in fact grounding through RAG and controlling randomness via temperature are the direct mechanisms to reduce hallucinations.

100
Multi-Selecthard

Which THREE are best practices for building a secure and scalable generative AI application using Amazon Bedrock? (Choose 3)

Select 3 answers
A.Implement guardrails to filter harmful content
B.Deploy models on EC2 instances for better control
C.Store API keys in source code for easy access
D.Use AWS KMS to encrypt data and model artifacts
E.Use foundation models from multiple providers via Bedrock
AnswersA, D, E

Guardrails in Amazon Bedrock apply configurable policies that filter harmful or inappropriate content in prompts and responses, denying unsafe interactions. This directly enforces the security posture the stem requires for a secure generative AI application.

Why this answer

Option A is correct because Amazon Bedrock Guardrails let you define denied topics, content filters, word filters, and PII redaction policies that are applied to both prompts and model responses, which is a core best practice for filtering harmful or non-compliant content in a generative AI application. Option D is correct because AWS KMS provides customer-managed keys and envelope encryption for data at rest, including model artifacts, knowledge bases, and logs, satisfying security and compliance requirements for protecting sensitive data. Option E is correct because Bedrock offers a unified, serverless API to access foundation models from multiple providers such as Anthropic, Meta, Mistral, Cohere, and Amazon, enabling model choice, fallback, and cost/performance optimization without managing infrastructure, which supports scalability.

Option B is not a Bedrock best practice because deploying models on self-managed EC2 instances moves you away from Bedrock's serverless, managed scaling and shifts patching, scaling, and security responsibilities to you. Option C is incorrect because hardcoding API keys in source code exposes credentials to leakage via repositories and logs; you should use AWS Secrets Manager or IAM roles instead.

Exam trap

Often, candidates mistakenly think that 'more control' (like EC2) is always better for security, when in fact managed services like Bedrock reduce attack surface and operational burden, making them the recommended approach for secure and scalable generative AI applications.

101
Multi-Selectmedium

A product team wants its generative AI assistant to answer questions about internal policy documents accurately rather than from the model's general training knowledge. Which TWO techniques should they apply? (Choose two.)

Select 2 answers
A.Increase the temperature setting so the model explores a wider range of possible answers.
B.Shorten the system prompt to a single sentence to reduce token consumption.
C.Instruct the model to answer only from the provided context and to state when the context is insufficient.
D.Ask the model to produce the longest possible answer for every question.
E.Retrieve relevant passages from the policy corpus and include them in the prompt before generation.
AnswersC, E

An explicit instruction that constrains the model to the supplied context, plus a fallback behavior for missing information, prevents the model from filling gaps with plausible but invented policy details. This complements retrieval by defining how the model should behave when retrieved evidence is absent or incomplete.

Why this answer

Grounding combines retrieval of authoritative source passages with an instruction that limits the model to that evidence and requires an explicit fallback when evidence is missing. Together these reduce reliance on parametric knowledge and make answers traceable to internal documents, which is what accurate policy question answering demands.

Exam trap

The trap here is treating generation settings such as temperature or prompt length as accuracy controls, when grounding depends on supplying evidence and constraining the model to it.

102
MCQeasy

A developer is using Amazon Bedrock's Claude model to summarize long documents. The developer notices that the summaries sometimes miss key points. Which parameter adjustment is most likely to improve summary completeness?

A.Increase the max_tokens parameter.
B.Increase the top_k parameter.
C.Increase the temperature parameter.
D.Increase the top_p parameter.
AnswerA

Truncation is the mechanism: if max_tokens is too low, the summary is cut off before covering all key points. Raising it allows the model to emit a complete summary, directly addressing the missed key points.

Why this answer

Increasing max_tokens allows the model to generate longer outputs, which is essential when summarizing long documents because the summary may need more tokens to capture all key points. If max_tokens is too low, the model truncates the response, potentially omitting important details. This directly addresses the issue of missing key points by providing sufficient output length for a complete summary.

Exam trap

The AIF-C01 exam often tests the misconception that parameters controlling randomness (temperature, top_k, top_p) affect output length or completeness, when in fact they only influence token selection diversity and creativity.

How to eliminate wrong answers

Option B is wrong because increasing top_k controls the number of highest-probability tokens considered during sampling, which affects randomness and diversity, not the length or completeness of the output. Option C is wrong because increasing temperature increases randomness in token selection, which can lead to more creative but less focused summaries, potentially worsening completeness. Option D is wrong because increasing top_p (nucleus sampling) also controls randomness by selecting tokens with cumulative probability, and does not extend the output length or guarantee inclusion of key points.

103
Multi-Selecthard

A company wants to evaluate the performance of a generative AI model before deployment. Which TWO metrics are most relevant for measuring model quality? (Select two.)

Select 2 answers
A.BLEU score
B.Response time
C.Perplexity
D.Model size
E.CPU utilization
AnswersA, C

BLEU compares n-gram overlap between generated and reference text, directly quantifying translation and summarisation fidelity. It satisfies the stem's need to measure generative output quality against expected answers, unlike latency or cost metrics that assess operational rather than quality performance.

Why this answer

BLEU score (A) is correct because it measures the quality of generated text by comparing n-gram overlap between model output and reference text, making it a standard metric for evaluating generative AI models such as translation and summarization systems. Perplexity (C) is correct because it quantifies how well a language model predicts a sample of text, with lower perplexity indicating better model confidence and language modeling quality. Response time (B) is not a quality metric but a latency/performance measure, model size (D) reflects resource footprint rather than output quality, and CPU utilization (E) is an infrastructure efficiency metric unrelated to the model's generative quality.

Exam trap

AWS exams often test the distinction between model quality metrics (like BLEU and perplexity) and operational or performance metrics (like response time or resource utilization), leading candidates to mistakenly select speed or size as relevant for quality assessment.

104
MCQhard

A data science team is comparing two approaches for customizing a foundation model in Amazon Bedrock for a domain-specific classification task. Approach one is providing a small set of labeled examples directly in the prompt for each request. Approach two is fine-tuning the model on a larger labeled dataset. The team wants the lowest operational overhead and the fastest way to start, and their label set is small and changes frequently. Which statement best describes the trade-off they should consider?

A.Prompt-based examples offer lower operational overhead and adapt quickly to changing labels, while fine-tuning can improve consistency when a larger stable dataset is available.
B.Prompt-based examples require a dedicated training job in Amazon Bedrock, so they carry the same operational overhead as fine-tuning.
C.Fine-tuning eliminates the need for any labeled data because the model learns the task from unlabeled inputs automatically.
D.Fine-tuning is preferred because it always produces higher accuracy than prompt-based examples regardless of dataset size.
AnswerA

Supplying labeled examples in the prompt, often called few-shot prompting, requires no training job and can be updated instantly as labels change. Fine-tuning is better suited when a larger, stable labeled dataset exists and consistent behavior matters. This statement correctly frames the trade-off given the team's small, changing label set.

Why this answer

The team's constraints are low operational overhead, fast start, and a small label set that changes often. Few-shot prompting embeds labeled examples in each request, needs no training job, and updates instantly. Fine-tuning is more appropriate for larger, stable datasets where consistent behavior is worth the training and maintenance effort, so it is not the best fit here.

Exam trap

The trap here is assuming fine-tuning is always superior for customization, when small or rapidly changing label sets make prompt-based examples the lower-overhead and more adaptable choice.

105
MCQeasy

A media company wants to create an internal tool that drafts short news summaries from press releases. They have no machine learning engineers, no labeled training data, and need a working prototype within two weeks. Which approach best describes how they should build this capability using generative AI concepts on AWS?

A.Use a pretrained foundation model through an API and steer its behavior with prompt engineering and few-shot examples.
B.Build a rules-based extractive summarizer that selects the first and last sentence of each press release.
C.Deploy an Amazon Rekognition custom labels project and post-process its output into summary sentences.
D.Train a small transformer model from scratch on the company's archive of press releases using Amazon SageMaker training jobs.
AnswerA

A foundation model already encodes broad language understanding from large-scale pretraining, so the team only needs to supply task instructions and a handful of example summaries in the prompt. This delivers a working prototype quickly without labeled data, GPUs, or model training, which matches the two-week constraint and the absence of ML staff.

Why this answer

Foundation models are pretrained on broad corpora and can be adapted to a new task through prompting rather than retraining. Supplying instructions plus a few example summaries in context lets a team with no ML specialists and no labeled dataset produce a functional summarizer within days. Training from scratch or using a vision service does not fit the constraint or the modality.

Exam trap

The trap here is assuming that any useful generative AI application requires custom model training, when in-context prompting of a pretrained foundation model is usually sufficient for a first prototype.

106
MCQmedium

A product team wants to compare two foundation models on their own customer-support transcripts before choosing one for a chatbot. They need a quantitative measure of how well each model's answers match reference answers. Which evaluation approach fits this need?

A.Human evaluation where reviewers rate each answer on a five-point scale
B.Automated evaluation using similarity metrics such as ROUGE or BERTScore between model outputs and reference answers
C.Reviewing each model's published model card and benchmark leaderboard rankings
D.Measuring inference latency and cost per thousand tokens for each model
AnswerB

Automated metrics compare generated text against reference answers and produce numeric scores, enabling objective side-by-side comparison across models on the same dataset. This suits the requirement for a quantitative measure over the team's own transcripts, and it scales to many examples without manual grading.

Why this answer

Automated text-similarity metrics score generated answers against reference answers and yield numbers that can be compared across models on the same transcripts. Human ratings are qualitative and costly, latency and cost measure operations rather than quality, and public benchmarks do not reflect the company's domain data. Automated evaluation on in-domain data is the fit.

Exam trap

The trap here is confusing operational metrics such as latency and cost with quality metrics that actually measure answer correctness.

107
MCQhard

A developer is prompting a foundation model to classify customer feedback into categories. The model sometimes returns extra commentary along with the category label, breaking downstream parsing. The developer wants more deterministic, tightly formatted output without retraining the model. Which technique best addresses this?

A.Set the temperature to zero and constrain the output using a structured format specification.
B.Add more few-shot examples that include verbose explanations of each category.
C.Increase the maximum output token count to give the model more room to respond.
D.Increase the top-p sampling value to broaden the token selection pool.
AnswerA

Lowering temperature to zero makes token selection greedy and far more deterministic, while constraining output with a structured format such as a JSON schema enforces the exact shape of the response. Together these yield tightly formatted labels that downstream systems can parse reliably, without any retraining, directly solving the developer's problem.

Why this answer

Setting temperature to zero makes generation greedy and deterministic, while a structured output specification constrains the response to an exact schema. This combination eliminates extraneous commentary and guarantees a parseable label, achieving the desired formatting control without any model retraining. It directly targets both the randomness and the format-enforcement gaps causing the parsing failures.

Exam trap

The trap here is reaching for sampling knobs like top-p or token limits, which affect variety and length rather than enforcing a strict, deterministic output structure.

108
MCQmedium

A solutions architect must choose a foundation model for an application that summarizes lengthy internal audit reports. The reports average 60,000 tokens, and the summaries must reflect details from the beginning, middle, and end of each document. Cost per request matters, but recall of details is the top priority. Which model characteristic should drive the selection?

A.The model's inference latency percentile
B.The number of parameters in the model
C.The model's supported output modalities
D.The model's maximum context window length
AnswerD

A document of roughly 60,000 tokens must fit inside the model's context window along with the prompt and the generated summary. If the window is smaller, content must be truncated or chunked, which risks losing the details the team cares about most. Context window length is therefore the gating characteristic for this workload.

Why this answer

When a single request must contain an entire long document, the maximum context window is the hard constraint that decides feasibility. Parameter count, output modality, and latency influence quality or experience but do not determine whether 60,000 tokens can be processed without losing the details the task requires.

Exam trap

The trap here is equating a larger parameter count with the ability to handle longer inputs, when context window size is a separate and independently configured limit.

109
MCQmedium

A data scientist is evaluating foundation models for a text summarization task and wants to use a standard metric. Which metric is commonly used to assess the quality of generated summaries?

A.F1 score
B.ROUGE
C.BLEU
D.Accuracy
AnswerB

ROUGE measures n-gram overlap between generated and reference summaries, directly satisfying the stem's requirement for a standard summarisation metric. Recall-oriented variants such as ROUGE-N and ROUGE-L capture how much reference content the model reproduced, making it the conventional benchmark for evaluating summary quality.

Why this answer

ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is the standard metric for text summarization. It measures the overlap of n-grams, word sequences, or word pairs between the generated summary and reference summaries, focusing on recall. This makes it particularly suitable for evaluating how well the generated summary captures the key content of the reference.

Exam trap

Candidates often confuse BLEU and ROUGE, mistakenly selecting BLEU for summarization because it is a well-known metric for text generation. However, BLEU is designed for machine translation precision, while ROUGE focuses on recall and n-gram overlap, making it the standard for summarization evaluation.

How to eliminate wrong answers

Option A is wrong because F1 score is a balanced measure of precision and recall, commonly used in classification tasks (e.g., binary or multi-class), not for evaluating the quality of generated text summaries. Option C is wrong because BLEU (Bilingual Evaluation Understudy) is designed for machine translation, measuring precision of n-gram overlap, and is less sensitive to recall, making it suboptimal for summarization where capturing all key points is critical. Option D is wrong because accuracy is a simple ratio of correct predictions to total predictions, used in classification or exact-match tasks, and is unsuitable for generative tasks like summarization where there is no single correct output.

110
Multi-Selectmedium

A company is using Amazon Bedrock to generate marketing copy. They want to ensure the output is safe and appropriate. Which TWO actions should they take? (Choose 2.)

Select 2 answers
A.Enable content filtering with guardrails
B.Set temperature to 0 for deterministic output
C.Use model fine-tuning with unsafe examples
D.Use a private endpoint for Bedrock
E.Implement human review of all generated content
AnswersA, E

Guardrails apply configurable content filters that detect and block harmful, hateful, or inappropriate categories in both prompts and responses, plus deny-topics and word filters. This directly enforces output safety at inference time, satisfying the requirement without changing the underlying model.

Why this answer

Option A is correct because Amazon Bedrock Guardrails provide configurable content filters (for hate, violence, sexual, insults, misconduct, and prompt-attack categories) that detect and block unsafe or inappropriate model output, directly addressing the goal of safe marketing copy. Option E is correct because implementing human review of all generated content adds a human-in-the-loop control that catches context-dependent or brand-inappropriate material that automated filters may miss, ensuring appropriateness before publication. Option B is incorrect because temperature controls randomness/creativity of sampling, not safety; setting it to 0 only makes output more deterministic and does nothing to filter harmful content.

Option C is incorrect because fine-tuning with unsafe examples would teach the model unsafe patterns, worsening rather than improving safety. Option D is incorrect because a private endpoint (VPC endpoint/PrivateLink) only secures network connectivity to Bedrock and does not evaluate or filter the safety of generated text.

Exam trap

Candidates often confuse mechanisms that control randomness (temperature=0) or network security (private endpoints) with content safety. While these settings affect output diversity and data confidentiality, they do not filter harmful or inappropriate content. Native guardrails or human review are required for content safety.

111
MCQeasy

A developer wants to generate product description images using Amazon Bedrock. They need to ensure the generated images match a specific brand style. Which feature should they primarily use?

A.Prompt engineering with detailed style descriptions.
B.Output grounding to verify brand compliance.
C.Data augmentation to increase dataset diversity.
D.Fine-tuning the image generation model on brand assets.
AnswerA

Detailed style descriptions in the prompt steer the model's output toward the required brand aesthetic, since prompt engineering is the primary control over generation. This satisfies the stem's constraint of matching a specific brand style.

Why this answer

Prompt engineering with detailed style descriptions is the primary and most direct method to guide Amazon Bedrock's image generation models (e.g., Stable Diffusion, Titan Image Generator) toward a specific brand style. By crafting precise prompts that include brand colors, design elements, and stylistic cues, the developer can influence the output without requiring additional training data or model modifications.

Exam trap

The trap here is that candidates may overestimate the necessity of fine-tuning (Option D) for style control, not realizing that prompt engineering is the primary, cost-effective feature for guiding image generation in Bedrock.

How to eliminate wrong answers

Option B is wrong because output grounding is a feature for verifying factual accuracy or source attribution in text generation (e.g., using citations), not for enforcing visual brand style compliance in image generation. Option C is wrong because data augmentation increases dataset diversity for training or fine-tuning, but it is not a feature used during inference to control the style of generated images in Bedrock. Option D is wrong because fine-tuning the image generation model on brand assets is possible but is not the primary feature; it requires additional cost, time, and expertise, whereas prompt engineering is the simplest and most immediate approach for style matching.

112
MCQmedium

A developer invoked an Amazon Bedrock model and received the following error: 'ValidationException: 1 validation error detected: Value 'claude-instant-v1' at 'modelId' failed to satisfy constraint: Member must satisfy enum value set: [ai21.j2-mid-v1, amazon.titan-text-lite-v1, anthropic.claude-v2, ...]'. What is the likely cause?

A.The Lambda function does not have the necessary IAM permissions
B.The modelId is not available in the current AWS region
C.The modelId is not part of the allowed enum of models for the account
D.The modelId is deprecated and has been renamed
AnswerC

The ValidationException explicitly reports that the supplied modelId fails the enum constraint, meaning Bedrock rejects any identifier outside its permitted model list. Since 'claude-instant-v1' is not among the enumerated values, the request never reaches inference. The cause is therefore an unsupported model identifier, not throttling, permissions, or malformed request syntax.

Why this answer

The error message explicitly states that the value 'claude-instant-v1' at 'modelId' failed to satisfy the constraint: 'Member must satisfy enum value set'. This indicates that the modelId provided is not part of the allowed list of model identifiers that the Amazon Bedrock API accepts for the invocation request. The error is a validation error from the API itself, not a permissions or availability issue, meaning the modelId string does not match any entry in the predefined enum of supported models.

Exam trap

The trap here is that candidates confuse a validation error (enum constraint) with a regional availability or permissions issue, but the specific error message 'Member must satisfy enum value set' directly points to an invalid model identifier string, not a missing resource or authorization failure.

How to eliminate wrong answers

Option A is wrong because IAM permissions errors would produce an 'AccessDeniedException' or 'AuthorizationError', not a 'ValidationException' with an enum constraint message. Option B is wrong because if the modelId were not available in the current region, the error would typically be a 'ModelNotAvailableException' or a regional availability error, not a validation error about the enum value set. Option D is wrong because a deprecated or renamed modelId would either still be accepted with a deprecation warning or produce a 'ModelNotFoundException', but the error explicitly states the value failed the enum constraint, meaning it was never a valid entry in the allowed set.

113
MCQhard

A team is deploying a generative AI model for medical report generation. They must ensure patient data privacy and comply with HIPAA. Which AWS service feature is essential for de-identifying protected health information (PHI) before sending data to a foundation model?

A.AWS CloudHSM
B.Amazon Comprehend Medical
C.Amazon Macie
D.AWS Key Management Service (AWS KMS)
AnswerB

Amazon Comprehend Medical provides dedicated PHI detection and de-identification through its DetectPHI API, identifying 18 HIPAA-defined identifier categories such as names, dates and medical record numbers. This satisfies the stem's requirement to strip protected health information before the data reaches any foundation model, unlike generic NLP services lacking healthcare-specific entity recognition.

Why this answer

Amazon Comprehend Medical is the correct service because it is specifically designed to extract and de-identify protected health information (PHI) from unstructured medical text using natural language processing (NLP). It can detect entities such as patient names, dates, and medical record numbers, and then redact or replace them before the data is sent to a foundation model, ensuring HIPAA compliance.

Exam trap

The trap here is that candidates confuse general data protection services like Macie or encryption services like KMS with the specialized PHI de-identification capability of Amazon Comprehend Medical, assuming any security service can handle HIPAA compliance for generative AI workflows.

How to eliminate wrong answers

Option A is wrong because AWS CloudHSM provides hardware security modules (HSMs) for cryptographic key storage and operations, but it does not perform data de-identification or PHI detection. Option C is wrong because Amazon Macie is a data security service that discovers and protects sensitive data using machine learning and pattern matching, but it is designed for data classification and access control, not for de-identifying PHI in unstructured text for downstream AI processing. Option D is wrong because AWS Key Management Service (AWS KMS) manages encryption keys for data at rest and in transit, but it does not have the capability to identify or remove PHI from text content.

114
MCQhard

A developer deployed this guardrail to block sensitive topics and sexual content. However, the model still generates responses about a specific sensitive topic that is not in the TopicPolicy. What should the developer do to prevent this?

A.Add a SensitiveInformationPolicy to filter PII
B.Increase the InputStrength of the content filter to MAX
C.Change the TopicPolicy Type from DENY to ALLOW
D.Add the specific topic to the TopicPolicy list
AnswerD

The guardrail only blocks topics enumerated in TopicPolicy, so an unlisted sensitive topic passes through unfiltered. Adding that specific topic to the policy list makes the filter match it, directly closing the gap without altering model weights or prompts.

Why this answer

The guardrail's TopicPolicy is designed to block specific topics by listing them. Since the model is generating responses about a sensitive topic not currently in the policy, adding that topic to the list directly extends the policy's coverage to prevent those responses. This is the precise mechanism for blocking new unwanted topics without altering other guardrail components.

Exam trap

AWS often tests the distinction between guardrail components—candidates may confuse content filters (which handle broad harmful categories) with topic policies (which handle custom, specific topics) and incorrectly choose to adjust content filter strength instead of updating the topic list.

How to eliminate wrong answers

Option A is wrong because a SensitiveInformationPolicy filters personally identifiable information (PII), not sensitive topics or sexual content; it addresses data privacy, not topic blocking. Option B is wrong because increasing the InputStrength of the content filter to MAX would only adjust the filtering of harmful content categories (e.g., hate, violence) but does not block specific topics not already covered by the TopicPolicy. Option C is wrong because changing the TopicPolicy Type from DENY to ALLOW would permit the specified topics instead of blocking them, which is the opposite of the desired outcome.

115
MCQhard

A team is using Amazon Bedrock to generate images from text prompts. The generated images often contain artifacts and do not match the prompt description. Which combination of steps should the team take to improve image quality?

A.Fine-tune the model using SageMaker Ground Truth and increase the training epochs.
B.Increase the max token count and use a larger model variant.
C.Refine the prompt with more descriptive language and adjust the CFG scale and inference steps.
D.Use a different foundation model and increase the image resolution.
AnswerC

Prompt refinement supplies the missing descriptive detail that causes prompt mismatch, while the CFG scale controls how strictly generation adheres to the prompt and inference steps govern artefact reduction. Together these directly address both stated symptoms: artefacts and poor prompt alignment.

Why this answer

Refining the prompt with more descriptive language helps the model better interpret the user's intent, while adjusting the CFG (Classifier-Free Guidance) scale controls how strictly the model adheres to the prompt, and increasing inference steps allows the diffusion process to produce higher-quality, artifact-free images. These are standard hyperparameters in diffusion-based image generation models on Amazon Bedrock, directly addressing both artifacts and prompt mismatch.

Exam trap

AWS often tests the misconception that image quality issues are best solved by model retraining or changing the model, rather than by adjusting inference-time parameters like CFG scale and inference steps, which are the immediate and correct levers for prompt adherence and artifact reduction.

How to eliminate wrong answers

Option A is wrong because fine-tuning a model using SageMaker Ground Truth and increasing training epochs is a data labeling and retraining approach that is overkill and not directly applicable to improving inference-time image quality for a pre-trained Bedrock model; it also does not address prompt adherence or artifact reduction. Option B is wrong because increasing the max token count and using a larger model variant does not fix artifacts or prompt mismatch—max token count affects text generation length, not image quality, and a larger model may not inherently improve prompt alignment without prompt engineering. Option D is wrong because using a different foundation model and increasing image resolution may change output characteristics but does not systematically address artifacts or prompt mismatch; higher resolution can even amplify artifacts if the underlying generation process is not optimized.

116
MCQeasy

A data scientist is using Amazon SageMaker to train a large language model from scratch. Which AWS service is most suitable for managing the training infrastructure, including automatic scaling and spot instance recovery?

A.AWS Lambda function.
B.Amazon SageMaker Notebook instance.
C.Amazon SageMaker Training job.
D.Amazon EC2 with a custom setup.
AnswerC

SageMaker Training jobs manage the underlying compute, handling automatic scaling and spot instance recovery natively. This satisfies the infrastructure-management requirement, unlike raw EC2 or manual cluster provisioning, which leave checkpointing and interruption handling to the data scientist.

Why this answer

Amazon SageMaker Training jobs are the most suitable service for managing training infrastructure because they provide built-in automatic scaling, managed spot instance recovery, and distributed training orchestration. This allows the data scientist to focus on model development rather than provisioning and managing EC2 instances, load balancers, or recovery scripts.

Exam trap

The AIF-C01 exam often tests the distinction between managed services (SageMaker Training) and unmanaged services (EC2 custom setup), where candidates mistakenly choose EC2 thinking they need full control, overlooking SageMaker's built-in spot recovery and scaling capabilities.

How to eliminate wrong answers

Option A is wrong because AWS Lambda functions are serverless compute services designed for short-running, event-driven tasks (max 15-minute execution time) and cannot manage long-running training jobs or infrastructure scaling. Option B is wrong because Amazon SageMaker Notebook instances are interactive development environments for prototyping and exploration, not designed to manage production training infrastructure or handle automatic scaling and spot instance recovery. Option D is wrong because Amazon EC2 with a custom setup requires manual provisioning, configuration of auto-scaling groups, and custom scripts for spot instance interruption handling, which is less efficient and more error-prone than SageMaker's managed training service.

117
MCQmedium

A company uses Amazon Bedrock to generate marketing content. They want to reduce costs while maintaining response quality. Which action is most effective?

A.Fine-tune a larger model to improve accuracy and reduce retries.
B.Increase the temperature parameter to get shorter responses.
C.Select a smaller foundation model that still meets accuracy requirements.
D.Cache previous responses to reuse for similar prompts.
AnswerC

Selecting a smaller foundation model directly reduces inference cost per token, since Bedrock charges scale with model size and capability tier. This satisfies the stem's dual constraint: cutting spend while preserving response quality, provided the smaller model still meets the stated accuracy requirements for marketing content generation.

Why this answer

The most effective cost-reduction strategy because smaller foundation models (FMs) have fewer parameters, resulting in lower compute and inference costs per request. If the smaller model still meets the required accuracy benchmarks for the marketing content task, it directly reduces operational expenditure without sacrificing quality. Amazon Bedrock offers a range of FMs (e.g., from large models like Claude 3 Opus to smaller ones like Claude 3 Haiku), allowing you to match model size to task complexity.

Exam trap

The trap here is that candidates confuse cost-reduction strategies with performance-enhancing strategies, assuming that fine-tuning or caching always saves money, when in fact the most direct lever is selecting the smallest capable model for the job.

How to eliminate wrong answers

Option A is wrong because fine-tuning a larger model increases training costs and still incurs higher per-inference costs due to the larger model size; retries are not a guaranteed cost driver, and fine-tuning does not inherently reduce inference cost. Option B is wrong because increasing the temperature parameter makes responses more random and potentially longer, not shorter; temperature controls creativity, not response length, and higher temperature often leads to more verbose or divergent outputs. Option D is wrong because caching previous responses is a latency optimization, not a cost reduction strategy; it does not reduce the per-request inference cost of generating new responses, and reusing cached responses may produce stale or irrelevant content for dynamic marketing prompts.

118
MCQhard

A healthcare organization wants to use generative AI to draft clinical notes from patient-physician conversations. They must comply with HIPAA and minimize false medical information. Which approach should they take?

A.Use Amazon SageMaker JumpStart with a publicly available clinical model and no additional modifications.
B.Use a generic open-source LLM hosted on Amazon EC2 with manual prompt engineering.
C.Use Amazon Bedrock with a HIPAA-eligible foundation model and connect it to a medical knowledge base via RAG.
D.Use Amazon Bedrock with a large foundation model and a high temperature setting for creativity.
AnswerC

Amazon Bedrock provides HIPAA-eligible models under a Business Associate Addendum, satisfying the compliance constraint. Retrieval Augmented Generation grounds responses in a curated medical knowledge base, reducing fabricated clinical content by supplying authoritative context at inference time rather than relying solely on parametric memory.

Why this answer

It combines a HIPAA-eligible foundation model via Amazon Bedrock with Retrieval-Augmented Generation (RAG) to ground responses in a curated medical knowledge base. This approach ensures compliance with HIPAA by using a service that supports Protected Health Information (PHI) processing under a Business Associate Agreement (BAA), while RAG reduces hallucination risk by retrieving factual clinical data rather than relying solely on the model's training.

Exam trap

AWS often tests the misconception that a larger or more creative model (high temperature) is better for accuracy, when in fact grounding via RAG and HIPAA-eligible infrastructure (e.g., Amazon Bedrock with a BAA) is the only safe path for regulated healthcare data.

How to eliminate wrong answers

Option A is wrong because Amazon SageMaker JumpStart provides pre-trained models but does not inherently offer HIPAA eligibility or a BAA for PHI processing, and using a publicly available clinical model without modifications fails to address false medical information through grounding. Option B is wrong because hosting a generic open-source LLM on Amazon EC2 lacks HIPAA compliance without explicit BAA and security configurations, and manual prompt engineering alone cannot reliably minimize false medical information without a retrieval mechanism. Option D is wrong because using a high temperature setting increases randomness and creativity, which directly contradicts the goal of minimizing false medical information by encouraging speculative or hallucinated outputs.

119
MCQeasy

A company uses Amazon Bedrock Agents to build an agent that interacts with users through a chat interface. The agent is configured with a knowledge base containing product documentation. Sometimes the agent fails to answer simple questions like 'What is your return policy?' and instead says it cannot find the answer. The knowledge base does contain the return policy. What is the most likely reason?

A.Increase the agent's maximum timeout for processing
B.Use a more powerful foundation model for reasoning
C.Add more documents to the knowledge base
D.Simplify and clarify the agent's instruction prompt to emphasize knowledge base usage
AnswerD

Overly complex or ambiguous instructions can cause the agent to under-utilise its knowledge base, producing false 'cannot find' responses despite the return policy existing. Clarifying the prompt to explicitly prioritise knowledge base retrieval satisfies the stem's scenario, restoring correct answers.

Why this answer

The agent's instruction prompt might be too complex or not explicitly directing the agent to use the knowledge base. Simplifying the prompt to clearly instruct the agent to first search the knowledge base can resolve the issue. Increasing timeout or adding more data is unnecessary.

A stronger model may help but is not the root cause.

120
MCQhard

A product team is designing an internal tool that helps engineers understand a large legacy codebase. They want the assistant to explain functions, suggest refactors, and answer questions about dependencies. The team is comparing a general-purpose foundation model with a model pre-trained specifically on source code. Which consideration most strongly favors the code-specialized model?

A.It is guaranteed to have a larger context window than any general-purpose model
B.Its pre-training corpus emphasized programming languages and code structure, improving performance on code comprehension tasks
C.It will always produce deterministic output for identical prompts
D.It eliminates the need for retrieval or context when answering questions about private repositories
AnswerB

Domain specialization during pre-training shapes the statistical patterns a model learns. A code-focused corpus exposes the model to syntax, idioms, library usage, and cross-file dependency conventions far more densely than general web text, which typically improves accuracy on tasks such as explaining functions or proposing refactors without any additional fine-tuning.

Why this answer

Domain specialization during pre-training is what gives a code-focused model an edge on comprehension, refactoring, and dependency reasoning, because its training data emphasized programming constructs. Context window size, determinism, and knowledge of private repositories are independent of the training domain and are not improved simply by specializing on code.

Exam trap

The trap here is assuming that a code-specialized model automatically knows a private codebase, when specialization improves general code skill but not knowledge of unreleased internal repositories.

121
MCQmedium

A financial services firm wants its generative AI assistant to answer questions using its private policy documents, which change frequently. The firm does not want to retrain the foundation model each time a document is updated. Which approach best meets these requirements?

A.Use retrieval-augmented generation to fetch relevant passages from a knowledge base and include them in the prompt.
B.Expand the model's context window by requesting a larger variant of the same model.
C.Fine-tune the foundation model on the policy documents after every update.
D.Increase the model's temperature setting so it explores more of its training knowledge.
AnswerA

Retrieval-augmented generation keeps documents in an external store, retrieves the passages most relevant to the question, and supplies them as context in the prompt. Updating a policy document only requires refreshing the index, with no model retraining, and answers can be grounded in the current text. This satisfies both the private-data and frequent-change requirements.

Why this answer

Retrieval-augmented generation decouples knowledge from model weights by retrieving relevant passages at query time and placing them in the prompt. Policy updates then require only re-indexing the documents, avoiding retraining entirely, while answers remain grounded in current private content. Temperature changes and larger context windows do not supply the missing knowledge on their own.

Exam trap

The trap here is assuming that a bigger context window or a fine-tuned model eliminates the need for retrieval when the real requirement is frequently updated private content.

122
Multi-Selectmedium

Which TWO strategies can help reduce inference costs when using Amazon Bedrock? (Select TWO.)

Select 2 answers
A.Use a higher temperature setting to generate fewer tokens
B.Increase the max tokens to allow longer responses
C.Use provisioned throughput for high-volume, predictable workloads
D.Cache frequently used responses in Amazon ElastiCache
E.Select a smaller foundation model variant
AnswersC, E

Provisioned throughput reserves dedicated model capacity for a committed term, replacing per-token on-demand pricing. For steady, high-volume workloads this committed rate is cheaper than paying on-demand for every request, directly lowering inference cost at predictable scale.

Why this answer

Option C is correct because Provisioned Throughput on Amazon Bedrock lets you purchase dedicated model units at a fixed hourly rate for steady, high-volume, predictable inference traffic, which is typically cheaper than paying on-demand per-token prices once utilization is high. Option E is correct because choosing a smaller foundation model variant (for example, a mini or lite version of a model family) reduces the number of parameters and compute required per inference, directly lowering the per-token or per-request cost while often remaining sufficient for the task. Option A is not correct because temperature controls randomness of sampling, not the number of tokens generated, so raising it does not reduce cost and may even increase output variability.

Option B is not correct because increasing max tokens allows longer responses, which increases the number of output tokens billed and therefore raises cost. Option D is not correct because Amazon ElastiCache is a caching service, not a native Bedrock cost-reduction strategy, and caching model responses is not one of the two intended Bedrock cost-optimization approaches in this scenario.

Exam trap

Candidates may mistakenly believe that caching responses in Amazon ElastiCache reduces inference costs, but this is not a built-in feature of Bedrock and would require custom implementation, still incurring costs for cache misses. Similarly, adjusting temperature or max tokens does not directly reduce per-token costs. The correct strategies are using provisioned throughput for high-volume predictable workloads and selecting a smaller foundation model variant.

123
Multi-Selecthard

A research team is using Amazon SageMaker to fine-tune a large language model. They want to optimize training cost and time without sacrificing model quality. Which THREE strategies should they implement? (Choose 3)

Select 3 answers
A.Use a larger instance type with more GPUs.
B.Apply parameter-efficient fine-tuning (PEFT) techniques like LoRA.
C.Increase the batch size to the maximum that fits in GPU memory.
D.Use managed spot training with checkpointing.
E.Enable mixed precision training (FP16).
AnswersB, D, E

Parameter-efficient fine-tuning such as LoRA freezes the base model weights and trains only small low-rank adapter matrices, cutting trainable parameters and GPU memory dramatically. This directly satisfies the stem's constraint of reducing training cost and time while preserving model quality, since the pretrained knowledge remains intact.

Why this answer

Option B is correct because parameter-efficient fine-tuning techniques such as LoRA freeze most of the pretrained model weights and train only small low-rank adapter matrices, which drastically reduces GPU memory consumption, training time, and compute cost while preserving model quality. Option D is correct because SageMaker managed spot training can cut training costs by up to 90% versus on-demand instances, and pairing it with checkpointing to Amazon S3 lets training resume from the last checkpoint after a spot interruption rather than restarting from scratch. Option E is correct because mixed precision training with FP16 reduces memory footprint and leverages GPU tensor cores for faster matrix math, speeding up training and allowing larger effective batch sizes with minimal impact on model accuracy.

Option A is not correct because simply moving to a larger multi-GPU instance raises cost and does not by itself optimize time or cost efficiency, and may be unnecessary if PEFT and FP16 already fit the workload. Option C is not correct because maximizing batch size to the GPU memory limit is not universally beneficial; overly large batches can hurt convergence and model quality and may require extensive hyperparameter retuning, so it is not a reliable cost-and-time optimization strategy.

Exam trap

The AIF-C01 exam often tests the misconception that simply scaling up hardware (larger instances) or maximizing batch size is the best optimization strategy, when in fact algorithmic efficiency (PEFT, mixed precision) and cost-saving infrastructure (spot instances) are the correct approaches for balancing cost, time, and quality.

124
MCQeasy

Refer to the exhibit. A developer wants to choose a model that can generate text (not just embeddings) and has the lowest cost. Based on the exhibit, which model should they select?

A.Titan Embed Text
B.Titan Text Express
C.Titan Text Lite
D.Need more information
AnswerC

Titan Text Lite is a text-generation model, satisfying the requirement to generate text rather than embeddings, and carries the lowest cost among the exhibit's generative options. Embeddings models such as Titan Embeddings cannot produce text, so they fail the primary constraint regardless of price.

Why this answer

Titan Text Lite is the correct choice because it is designed for text generation tasks and is explicitly positioned as the lowest-cost option among Amazon's Titan text generation models. Unlike Titan Embed Text, which only produces embeddings and cannot generate text, Titan Text Lite offers a cost-efficient solution for generating text while meeting the requirement of lowest cost.

Exam trap

The trap here is that candidates may confuse Titan Embed Text as a text generation model due to its 'Text' name, or assume that 'Express' implies lower cost, when in fact the naming convention indicates performance tier rather than cost efficiency.

How to eliminate wrong answers

Option A is wrong because Titan Embed Text is an embeddings model that converts text into numerical vectors and cannot generate text, failing the core requirement. Option B is wrong because Titan Text Express, while capable of text generation, is a higher-cost model compared to Titan Text Lite, making it not the lowest-cost option. Option D is wrong because the exhibit provides sufficient information to determine that Titan Text Lite is the correct model based on the stated requirements of text generation capability and lowest cost.

125
MCQhard

A startup is choosing between two foundation models for a summarization feature. Model X is a large general-purpose model with strong benchmark scores; Model Y is a smaller domain-specialized model with lower general benchmarks but excellent results on the startup's own document samples. Cost per token for Model Y is roughly one third of Model X. Which evaluation practice should the team follow?

A.Select Model Y solely because its cost per token is lower, since summarization quality is equivalent across models.
B.Deploy both models and let end users vote on which summaries they prefer before any offline evaluation.
C.Select Model X because higher public benchmark scores reliably predict performance on every downstream task.
D.Evaluate both models on a held-out set of representative documents, then weigh measured quality against cost and latency.
AnswerD

Task-specific evaluation on representative, held-out data is the most reliable signal of downstream performance, and combining it with cost and latency reflects the real trade-offs of production. This approach uses the startup's own evidence rather than generic benchmarks or price alone, which is the sound engineering practice for model selection.

Why this answer

Model selection should be driven by measured performance on data that resembles the production workload, combined with operational constraints such as cost and latency. Public benchmarks and price alone are weak proxies, while a held-out evaluation set gives direct evidence of whether the cheaper specialized model meets the quality bar for summarization.

Exam trap

The trap here is equating leaderboard rankings or lowest price with fitness for a specific task, when only task-specific evaluation reveals the real trade-off.

← PreviousPage 2 of 2 · 125 questions total

Ready to test yourself?

Try a timed practice session using only Fundamentals of Generative AI questions.