AIF-C01 Fundamentals of Generative AI Practice Question
A product team is comparing foundation models for a customer-facing FAQ bot. They need low latency for real-time chat and want to pay only for what they use without managing servers. Which combination of model characteristic and AWS consumption model best fits these requirements?
⚠ Common exam trap
The trap here is equating the largest model with the best choice, when parameter count usually increases latency and cost rather than improving fit for a latency-sensitive FAQ bot.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A smaller, lower-latency model invoked on demand so billing is based on input and output tokens processed.
Real-time chat favors models that emit tokens quickly, and a smaller model usually wins on latency. On-demand invocation bills by tokens processed, so a variable-traffic FAQ bot pays only for actual usage and the provider handles scaling and infrastructure. Provisioned capacity and self-hosted GPUs suit steady high volume but add cost and operational burden here.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A large model with the highest parameter count, deployed on a provisioned throughput commitment billed hourly.
Why it's wrong here
Maximum parameter count usually means higher latency per token, which conflicts with the real-time chat requirement. Provisioned throughput is billed for reserved capacity regardless of actual traffic, so a spiky FAQ bot would pay for idle capacity. This combination optimizes for peak quality and predictable volume, not the stated low-latency, pay-per-use goal.
- ✗
An open-weight model downloaded and hosted on AWS Lambda for per-request execution.
Why it's wrong here
Lambda has execution time and memory limits that make hosting a large language model impractical, and cold starts would add severe latency to chat responses. Model weights typically exceed what an ephemeral function can hold, so this architecture fails on latency and feasibility despite appearing to offer per-request billing.
- ✗
A multimodal model with image understanding, called through a self-managed Amazon EC2 GPU instance.
Why it's wrong here
Image understanding is unnecessary for a text FAQ bot, so the extra capability adds cost without benefit. Running on a self-managed EC2 GPU instance means the team patches drivers, manages scaling, and pays for the instance even when idle, which violates both the no-server-management and pay-per-use requirements.
- ✓
A smaller, lower-latency model invoked on demand so billing is based on input and output tokens processed.
Why this is correct
Smaller models generally produce tokens faster, which supports responsive chat, and on-demand invocation charges per token consumed, matching the pay-only-for-what-you-use requirement with no infrastructure to manage. This aligns latency, cost model, and operational simplicity with the team's constraints for a customer FAQ workload.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.