AIF-C01 Fundamentals of Generative AI Practice Question
A company wants to evaluate the performance of a generative AI model before deployment. Which TWO metrics are most relevant for measuring model quality? (Select two.)
⚠ Common exam trap
AWS exams often test the distinction between model quality metrics (like BLEU and perplexity) and operational or performance metrics (like response time or resource utilization), leading candidates to mistakenly select speed or size as relevant for quality assessment.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
BLEU score
BLEU score (A) is correct because it measures the quality of generated text by comparing n-gram overlap between model output and reference text, making it a standard metric for evaluating generative AI models such as translation and summarization systems. Perplexity (C) is correct because it quantifies how well a language model predicts a sample of text, with lower perplexity indicating better model confidence and language modeling quality. Response time (B) is not a quality metric but a latency/performance measure, model size (D) reflects resource footprint rather than output quality, and CPU utilization (E) is an infrastructure efficiency metric unrelated to the model's generative quality.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
BLEU score
Why this is correct
BLEU compares n-gram overlap between generated and reference text, directly quantifying translation and summarisation fidelity. It satisfies the stem's need to measure generative output quality against expected answers, unlike latency or cost metrics that assess operational rather than quality performance.
- ✗
Response time
Why it's wrong here
Response time measures latency of inference, an operational performance characteristic, not the correctness or relevance of generated output. It is tempting because slow responses degrade user experience, and would be the right metric when tuning throughput or provisioning capacity, not when assessing model quality before deployment.
- ✓
Perplexity
Why this is correct
Perplexity measures how confidently a language model predicts the next token; lower values indicate the model assigns higher probability to real text. This directly quantifies generative model quality, satisfying the stem's requirement for a relevant pre-deployment quality metric.
- ✗
Model size
Why it's wrong here
Model size measures parameter count and storage footprint, which indicate cost and latency rather than output quality. It is tempting because larger models often correlate with stronger capability, and would be relevant when planning inference infrastructure or selecting an instance type, not when evaluating generated responses.
- ✗
CPU utilization
Why it's wrong here
CPU utilization is a performance metric, not a measure of model output quality.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.