AI-102 Implement generative AI solutions Practice Question
You need to generate an image of a cat wearing a hat using Azure OpenAI. Which model should you use?
⚠ Common exam trap
It's easy for candidates to confuse GPT-4's multimodal capabilities (which can analyze images but not generate them) with DALL-E's generative image creation, leading them to incorrectly select GPT-4 for image generation tasks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
DALL-E
DALL-E is the Azure OpenAI model specifically designed for generating images from natural language descriptions. It uses a diffusion-based architecture to create high-quality, original images based on text prompts, making it the correct choice for generating an image of a cat wearing a hat.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Codex
Why it's wrong here
Codex translates natural language into source code; its output is program text, not rendered pixels, so no image is generated. It would be the correct model for tasks such as converting a plain-English description into a working function.
- ✓
DALL-E
Why this is correct
DALL-E is Azure OpenAI's dedicated image-generation model, so it produces a new picture from a text prompt such as a cat wearing a hat. GPT models output text, not images, so they cannot satisfy the image-generation requirement.
- ✗
GPT-4
Why it's wrong here
GPT-4 is a text and multimodal completion model; it returns language tokens, not image files, so it cannot produce the cat picture. It is the right pick for generating or reasoning over text, such as writing a caption describing that cat.
- ✗
Whisper
Why it's wrong here
Whisper performs speech-to-text transcription and cannot generate images at all, so it fails the stated requirement. It is tempting because it is a genuine Azure OpenAI model, and it would be the correct choice when the task involves transcribing or translating audio into text instead.
Go deeper
Related to this question
About these practice questions
One of 761 original AI-102 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.