Courseiva

AI-102 Implement generative AI solutions Practice Question

You need to deploy a generative AI model that can be used by multiple applications within your organization. The model must support real-time inference with low latency. Which Azure service should you use?

⚠ Common exam trap

Watch out — candidates often confuse Azure Machine Learning real-time endpoints (which are for custom ML models) with Azure OpenAI Service (which is purpose-built for generative AI), overlooking the fact that Azure OpenAI provides managed, low-latency inference optimized for large language models without the overhead of containerized deployments.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Azure OpenAI Service

Azure OpenAI Service provides managed access to powerful generative AI models like GPT-4, which are optimized for real-time inference with low latency through provisioned throughput units (PTUs) and regional deployment options. This service is specifically designed for generative AI workloads, offering REST API endpoints that support streaming responses and sub-second latency for single-turn interactions, making it ideal for multiple applications requiring consistent, low-latency responses.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Azure AI Search

    Why it's wrong here

    Azure AI Search indexes and retrieves documents; it cannot host or run a generative model for inference. It is tempting because it underpins retrieval-augmented generation, but it supplies grounding context to a model hosted elsewhere, not the low-latency inference endpoint itself.

  • ✓

    Azure OpenAI Service

    Why this is correct

    Azure OpenAI Service hosts generative models behind a managed REST endpoint with provisioned throughput, giving multiple applications shared, low-latency real-time inference. It satisfies the low-latency constraint that batch-oriented or self-hosted alternatives cannot meet as directly.

  • ✗

    Azure Machine Learning real-time endpoint

    Why it's wrong here

    Real-time endpoints host custom or fine-tuned models on provisioned compute, but the scenario needs a managed generative model shared across applications, which Azure OpenAI Service provides with pay-as-you-go tokens. Real-time endpoints suit custom-trained models requiring dedicated GPU capacity and controlled scaling.

  • ✗

    Azure Functions

    Why it's wrong here

    Azure Functions runs event-driven serverless code, not model inference with GPU-backed low latency. It is tempting as a lightweight API wrapper, but it would call a model hosted elsewhere; the requirement is the hosting service itself, which Azure OpenAI Service supplies.

About these practice questions

Courseiva writes every AI-102 question from scratch — 761 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.