Courseiva

Generative AI Leader

Full exam simulation

1:30:00
1

Fundamentals of Generative AI

medium

A team is tuning a large language model for a question-answering task. They notice the model gives high confidence scores to answers that are factually incorrect. Which evaluation metric should they primarily use to detect this overconfidence problem?

0 of 50 answered