Courseiva

AI-102 Implement computer vision solutions Practice Question

A company uses the Computer Vision Image Analysis API to generate captions for images. The captions are often too generic. How can they improve the descriptiveness of captions?

⚠ Common exam trap

Many candidates confuse increasing the confidence threshold with improving descriptiveness, when in reality it only reduces the number of captions returned without adding detail.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the Dense Captioning feature.

The Dense Captioning feature in the Computer Vision Image Analysis API generates one-sentence descriptions for each of up to 10 regions detected in an image, providing more specific and detailed captions than the single generic caption. This directly addresses the problem of captions being too generic by breaking the image into meaningful areas and describing each one individually.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use the Object Detection API instead.

    Why it's wrong here

    Object Detection returns bounding boxes and labels, not natural-language descriptions, so captions remain generic. It is tempting because detection identifies more objects per image, and would be correct when the requirement is locating and classifying objects rather than generating richer descriptive sentences.

  • ✗

    Train a custom model with Custom Vision.

    Why it's wrong here

    Custom Vision performs image classification and object detection, returning tags and bounding boxes, not natural-language captions, so it cannot generate descriptive sentences. It is the right choice when you need bespoke class labels for your own image categories, not caption enrichment.

  • ✗

    Increase the confidence threshold for captions.

    Why it's wrong here

    Raising the confidence threshold filters out low-scoring captions, so generic ones are suppressed rather than made descriptive; the underlying caption model is unchanged. Thresholds suit precision-tuning when false positives matter, not enriching vocabulary. Dense captioning or a richer caption feature is what adds detail.

  • ✓

    Use the Dense Captioning feature.

    Why this is correct

    Dense Captioning generates multiple descriptions for distinct regions within an image, not one caption for the whole frame. This satisfies the stem's requirement for richer, more descriptive output, since generic captions stem from whole-image summarisation. It also returns bounding boxes, adding spatial detail the standard caption endpoint cannot provide.

About these practice questions

This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.