A retail company wants to compare the output quality of several foundation models available in Amazon Bedrock for a product description generation task. They need a repeatable, automated way to score responses against reference descriptions. Which AWS capability should they use?
Amazon Bedrock supports model evaluation jobs that can automatically score model outputs against reference data using metrics like BERTScore and ROUGE. This provides a repeatable, automated comparison across multiple foundation models, which matches the requirement. It removes manual scoring and produces consistent results that can guide model selection for the product description task.
Why this answer
Amazon Bedrock model evaluation jobs run automated scoring against reference datasets using metrics such as BERTScore and ROUGE, enabling consistent comparison of multiple foundation models. This directly supports selecting the best model for product description generation with repeatable, quantitative results rather than subjective manual review.
Exam trap
The trap here is confusing operational monitoring or capacity features with model quality evaluation, which requires scoring outputs against references.