AIF-C01 Fundamentals of Generative AI Practice Question
Exhibit
aws bedrock-runtime invoke-model \
--model-id anthropic.claude-v2 \
--body '{"prompt":"\n\nHuman: Summarize the following text: ...\n\nAssistant:","max_tokens_to_sample":200}' \
--cli-binary-format raw-in-base64-out \
--region us-east-1 \
output.json
The output.json file contains:
{"completion": " The summary is...", "stop_reason": "stop_sequence"}Refer to the exhibit. A developer runs the CLI command to summarize text using Claude v2 in Bedrock. The output is shorter than expected. Which change should the developer make to allow a longer response?
⚠ Common exam trap
A common mistake is to assume that simply asking the model to produce a longer response via prompt engineering will override the API parameter 'max_tokens_to_sample'. In AWS Bedrock, the 'max_tokens_to_sample' parameter is the definitive control for response length.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase 'max_tokens_to_sample' to 1000
The 'max_tokens_to_sample' parameter in the Bedrock InvokeModel API directly controls the maximum number of tokens the model can generate in its response. By default, this value is often set low (e.g., 256 tokens), which truncates the output. Increasing it to 1000 allows Claude v2 to produce a longer summary up to that token limit.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Increase 'max_tokens_to_sample' to 1000
Why this is correct
The response is truncated because max_tokens_to_sample caps generated output length. Raising it to 1000 permits Claude v2 to emit more tokens before stopping, directly satisfying the requirement for a longer summary in the Bedrock CLI invocation.
- ✗
Change the prompt to include 'Write a long summary'
Why it's wrong here
Prompt wording cannot exceed the max_tokens parameter, which caps generated output length; Claude stops once that budget is reached. It is tempting because prompt engineering shapes verbosity, and this would be correct when the model's response is truncated by an instruction implying brevity rather than a token limit.
- ✗
Set 'stop_reason' to 'none'
Why it's wrong here
stop_reason is a read-only response field reporting why generation halted; it is not an input parameter and cannot be set to 'none'. It is tempting because it appears in the CLI output alongside truncation, and inspecting it would be correct when diagnosing whether max_tokens or a stop sequence ended the response.
- ✗
Use a different region like us-west-2
Why it's wrong here
Region selection affects latency, data residency and model availability, not the max_tokens ceiling governing output length. It is tempting because some models or inference profiles are only offered in certain regions, and changing region would be correct when the desired model or capacity is unavailable in the current one.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.