20+ practice questions focused on Fundamentals of Generative AI — one of the most tested topics on the AWS Certified AI Practitioner AIF-C01 exam. Each question includes a detailed explanation so you learn why the right answer is correct.
Start Fundamentals of Generative AI PracticeA company is using Amazon SageMaker JumpStart to deploy a pre-trained text generation model. After deployment, the model produces slow inference responses. Which action is most likely to improve inference latency?
Explanation: Deploying the model on a more powerful instance type with higher GPU memory directly addresses the computational bottleneck causing slow inference. A larger GPU provides more CUDA cores and memory bandwidth, enabling faster matrix operations and reducing the time per forward pass for the pre-trained text generation model.
A company is building a chatbot using Amazon Bedrock. They want to ensure the model's responses are grounded in their internal knowledge base and avoid generating information outside that scope. Which feature should they use?
Explanation: Amazon Bedrock Knowledge Bases is the correct feature because it allows you to connect a foundation model (FM) to your internal data sources, such as documents or databases, and use Retrieval Augmented Generation (RAG) to ground responses in that specific knowledge. This ensures the chatbot only generates information from the provided knowledge base, preventing hallucinations or out-of-scope content.
A company is using Amazon Bedrock to generate creative marketing copy. They want to reduce the randomness of the output while maintaining diversity. Which TWO parameters should they adjust?
Explanation: Option E (Decrease the temperature) is correct because temperature scales the randomness of token sampling in Amazon Bedrock; lowering it makes the probability distribution sharper, so the model favors high-probability tokens and produces more deterministic, less random output. Option D (Decrease the top_p value) is correct because top_p (nucleus sampling) restricts generation to the smallest set of tokens whose cumulative probability exceeds p; reducing p narrows that candidate pool, cutting off low-probability tokens and reducing randomness while still allowing varied word choices within the nucleus, which preserves diversity. Option A (Increase the temperature) is wrong because raising temperature flattens the distribution and increases randomness, the opposite of the goal. Option B (Increase the max token count) is wrong because max tokens only caps output length and does not affect sampling randomness. Option C (Increase the top_k value) is wrong because a larger top_k widens the candidate token set, increasing rather than reducing randomness.
A developer attached this IAM policy to a role used by an application that invokes Claude v2 in us-east-1. The application receives an access denied error. What is the MOST likely cause?
Explanation: The Deny statement uses a `StringNotEquals` condition on `aws:RequestedRegion` set to `us-east-1`. This means the Deny applies to any request where the requested region is NOT `us-east-1`. Since the resource ARN in the Deny statement is `arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-v2`, the condition does not match the resource's region (the resource ARN itself is in us-east-1), but the Deny is triggered when the request is made to a different region, blocking the call. The application is likely invoking the model from a region other than us-east-1, causing the Deny to take effect.
A financial services company is deploying a generative AI model on Amazon SageMaker for real-time fraud detection. The model, a fine-tuned Llama 2 7B, must respond to transaction requests within 500 milliseconds. The team has deployed the model using a SageMaker real-time endpoint with a single ml.g5.2xlarge instance. During load testing, the endpoint achieves an average latency of 450 ms at 10 requests per second (RPS), but the latency spikes to over 2 seconds at 20 RPS. The team needs to maintain sub-500 ms latency at up to 50 RPS. The model is too large to fit on a single GPU, so they are using CPU instances. They considered using a larger instance type but want to minimize cost. What should the team do to meet the latency requirement cost-effectively?
Explanation: A SageMaker multi-model endpoint (MME) allows multiple model replicas to be hosted on a fleet of instances, enabling horizontal scaling to handle increased throughput. By using multiple ml.g5.xlarge instances with auto scaling, the team can distribute the 50 RPS load across several instances, keeping per-instance latency low while minimizing cost compared to a single larger instance. This approach also leverages the fact that the model is too large for a single GPU but can be efficiently served on CPU instances with proper load distribution.
+15 more Fundamentals of Generative AI questions available
Practice all Fundamentals of Generative AI questions1. Baseline your knowledge
Start with 10 questions to gauge your current understanding of Fundamentals of Generative AI. This tells you whether you need a concept refresher or just practice.
2. Review every explanation
For each question — right or wrong — read the full explanation. Understanding why an answer is correct is more valuable than knowing the answer itself.
3. Focus on exam traps
Fundamentals of Generative AI questions on the AIF-C01 frequently use trap wording. Look for subtle differences in answers that test your precision, not just general knowledge.
4. Reach 80% consistently
Do repeated sessions until you score 80%+ three times in a row. Then move to mixed-mode practice to test cross-topic recall under realistic conditions.
The exact number varies per candidate. Fundamentals of Generative AI is tested as part of the AWS Certified AI Practitioner AIF-C01 blueprint. Practicing with targeted Fundamentals of Generative AI questions ensures you can handle any format or difficulty that appears.
Yes. Courseiva provides free AIF-C01 practice questions across all exam topics and domains. The platform includes topic-based practice, mock exams, missed-question review, bookmarked questions, and readiness tracking — no account required.
Difficulty is subjective, but Fundamentals of Generative AI is a high-priority exam concept tested in multiple ways — direct recall, scenario analysis, and command-output interpretation. Consistent practice is the best way to build confidence.
Launch a full Fundamentals of Generative AI practice session with instant scoring and detailed explanations.
Start Fundamentals of Generative AI Practice →