Generative AI Leader Practice Question: Business Strategies for Generative AI Solutions
A company is using generative AI for code generation and wants to evaluate the quality of generated code for security vulnerabilities. Which metric is most appropriate?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Automatic static analysis
(Automatic static analysis) is correct because it directly scans code for security vulnerabilities, making it the most appropriate metric for this purpose. Option A (BLEU score) measures text similarity, not security. Option C (Human evaluation) is subjective and less scalable. Option D (Perplexity) measures language model confidence, not code security.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
BLEU score
Why it's wrong here
BLEU compares n-gram overlap between generated and reference text, so it scores surface similarity and cannot identify exploitable code flaws. It is tempting because it is a standard, automatic metric for evaluating generated output. Assessing security vulnerabilities instead requires static analysis or security-focused benchmarks that examine code behaviour.
- ✓
Automatic static analysis
Why this is correct
Automatic static analysis scans generated code for insecure patterns such as injection flaws, hardcoded credentials and unsafe API calls, giving an objective, repeatable security measure. Functional or similarity metrics cannot detect vulnerabilities, so static analysis directly satisfies the requirement to evaluate security quality.
- ✗
Human evaluation
Why it's wrong here
Human evaluation relies on reviewers manually reading code, which does not scale to large generated volumes and yields subjective, inconsistent security judgements. It is tempting because humans can reason about subtle logic flaws that automated tools miss, suiting small high-risk audits. Detecting vulnerabilities at scale instead requires automated static analysis or dedicated security scanners.
- ✗
Perplexity
Why it's wrong here
Perplexity measures how well a language model predicts token sequences, reflecting fluency rather than whether generated code contains exploitable flaws. It is tempting because it is cheap, automatic and widely used for comparing model outputs. Security vulnerability detection instead requires static analysis or targeted security benchmarks that inspect code semantics.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.