Courseiva
mediumMultiple Choice

AI-102 Practice Question: A news organization uses Azure Video Indexer to…

A news organization uses Azure Video Indexer to generate transcripts of live broadcasts. They notice that the speaker names are not appearing in the transcript. What is the most likely cause?

⚠ Common exam trap

It's easy for candidates to confuse speaker identification with automatic diarization or assume that speaker names are automatically extracted from the video metadata, when in fact Azure Video Indexer requires explicit training of a custom Person Model with voice samples to assign names.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The speaker identification model has not been trained with voice samples.

Speaker names are missing because Azure Video Indexer's speaker identification feature requires pre-trained voice samples to match speakers to their identities. Without a custom voice model trained on known speakers' audio, the service can only label speakers as 'Speaker #1', 'Speaker #2', etc., but cannot assign actual names. This is a supervised learning process where the model must be trained with labeled voice samples before it can recognize and name speakers.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The video resolution is too low for OCR.

    Why it's wrong here

    Resolution affects OCR of on-screen text such as slates and captions, not acoustic speaker separation. Diarisation analyses audio waveforms to cluster and label voices, so pixel density is irrelevant here. Low resolution would be the culprit when burned-in text or visual insights are unreadable.

  • ✓

    The speaker identification model has not been trained with voice samples.

    Why this is correct

    Speaker identification in Azure Video Indexer requires a trained voice model built from submitted speaker audio samples; without enrolment, the service cannot map voices to named individuals. The stem's missing speaker names therefore point to an untrained model, since transcription itself still succeeds but attribution stays anonymous.

  • ✗

    The video format is not supported.

    Why it's wrong here

    Unsupported formats fail the indexing job outright, producing no transcript at all rather than one lacking speaker labels. Format support governs whether media is ingested; diarisation assigns speaker names, so the stem's symptom points elsewhere. Checking codec compatibility is valid when uploads are rejected or processing errors occur.

  • ✗

    The language is not set correctly.

    Why it's wrong here

    Language configuration drives transcription accuracy, not speaker attribution; a wrong language yields garbled or mistranslated text, yet speakers would still be separated and labelled. Diarisation requires the speaker identification setting, which is disabled by default. Language selection is the right fix when transcripts return in the wrong tongue.

About these practice questions

This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.