Courseiva

CCNA Implement generative AI solutions Questions

75 of 174 questions · Page 1/3 · Implement generative AI solutions · Answers revealed

1
Multi-Selecthard

Which TWO actions can you take to mitigate the risk of generating harmful content when using Azure OpenAI Service? (Choose two.)

Select 2 answers
A.Set a system message that instructs the model to avoid harmful outputs.
B.Fine-tune the model on a dataset of safe examples.
C.Deploy the model in multiple regions.
D.Configure Azure AI Content Safety filters.
E.Increase the maxTokens parameter to allow longer responses.
AnswersA, D

A system message sets behavioural guardrails at the prompt level, instructing the model to refuse or avoid harmful content before generation. It is a mitigation available through Azure OpenAI Service configuration, complementing rather than replacing platform-level filtering.

Why this answer

Option A is correct because a system message sets the model's behavioral guardrails at inference time, and explicitly instructing the model to refuse or avoid harmful outputs is a documented prompt-engineering mitigation for Azure OpenAI Service. Option D is correct because Azure AI Content Safety filters (the default and customizable content filters in Azure OpenAI) inspect both prompts and completions for categories such as hate, violence, sexual, and self-harm, and block or annotate harmful content before it reaches users. Option B is not correct because fine-tuning on safe examples can shape tone and style but is not a reliable content-safety control and does not replace the platform's content filtering.

Option C is not correct because multi-region deployment addresses availability and latency, not the generation of harmful content. Option E is not correct because increasing maxTokens only allows longer responses and can actually increase exposure to harmful output rather than mitigate it.

Exam trap

The trap is treating fine-tuning as a safety mechanism — candidates pick 'fine-tune on safe examples' because it sounds thorough, but the exam expects you to know that platform-level Content Safety filters plus system messages are the recognized mitigations.

2
MCQeasy

You are testing an Azure OpenAI model with the parameters shown in the exhibit. The model generates very short responses. Which parameter should you modify to allow longer responses?

A.Increase frequency_penalty
B.Increase top_p
C.Increase temperature
D.Increase max_tokens
AnswerD

Raising max_tokens directly lifts the hard ceiling on generated output length, which is the constraint producing truncated, very short responses. Temperature, top_p and frequency_penalty shape token selection, not response length, so they cannot extend output beyond the cap. Microsoft Entra ID is unrelated to inference parameters.

Why this answer

The max_tokens parameter controls the maximum number of tokens (words or subwords) the model can generate in a single response. When responses are very short, increasing max_tokens allows the model to produce longer completions up to the specified limit. The other parameters affect randomness, diversity, or probability distribution, not the length cap.

Exam trap

The trap here is that candidates confuse parameters that control output length (max_tokens) with those that control output diversity or creativity (temperature, top_p, frequency_penalty), leading them to incorrectly adjust the latter when the real issue is a token limit.

How to eliminate wrong answers

Option A is wrong because frequency_penalty reduces the likelihood of repeating the same tokens or phrases, which can actually shorten responses by discouraging repetition, not lengthen them. Option B is wrong because top_p (nucleus sampling) controls the cumulative probability threshold for token selection, affecting diversity but not the maximum output length. Option C is wrong because temperature adjusts the randomness of token selection (higher = more creative, lower = more deterministic) and does not impose or remove a length constraint.

3
MCQmedium

You are implementing a Retrieval Augmented Generation pattern with Azure AI Search and Azure OpenAI. Users complain that answers occasionally cite documents the user is not authorized to see, because the index contains all departments' content. You need to restrict retrieval so that each user only receives chunks from documents their identity permits. What should you do?

A.Create a separate index per department and have users manually select which index to search.
B.Add a system message instructing the model to ignore any content the user is not allowed to see.
C.Apply a semantic ranker to the query and rely on relevance scoring to push unauthorized content lower in the results.
D.Add a security filter on a document access field in Azure AI Search, and pass the user's group or object IDs as an OData filter in the query.
AnswerD

Azure AI Search supports filtering on fields such as a collection of allowed group IDs. By storing each document's permitted principals and passing the caller's group or object IDs from the token into an OData filter, the query returns only chunks the identity is authorized to read. This enforces security at retrieval time, which is the correct layer for RAG, before content ever reaches the model or the user.

Why this answer

Security trimming for RAG must occur at retrieval. Storing each document's allowed principals in a filterable field and passing the caller's group or object IDs as an OData filter ensures the search response never contains unauthorized chunks. Relevance ranking, prompt instructions, and manual index selection do not enforce authorization and can leak restricted content before the model or user sees it.

Exam trap

The trap here is treating relevance ranking or prompt instructions as if they provided access control.

4
MCQmedium

Your team is building a generative AI application on Azure OpenAI that must call an internal order-lookup REST API whenever a user asks about an order status. The model must decide when to call the API, and the application must execute the call and return the result to the model for the final answer. Which capability should you implement?

A.Semantic ranker enabled on an Azure AI Search index that stores order records.
B.Function calling (tools) on the chat completions request, defining the order-lookup API as a tool.
C.A system message instructing the model to answer order questions only from its training data.
D.A JSON mode response format that forces the model to emit a schema-conforming order status object.
AnswerB

Function calling lets you describe available tools in the request; the model then returns a structured tool call when it determines the API is needed. Your application executes the call and sends the result back as a tool message, after which the model composes the final answer. This exactly matches the required decide-execute-return loop.

Why this answer

The requirement is a model-driven decision to invoke a live REST API and then incorporate the returned data. Function calling provides exactly that contract: tools are declared, the model emits a structured call, the host application performs the HTTP request, and the result is fed back for final generation. The other options either constrain formatting, rely on training data, or only rank search results.

Exam trap

The trap here is equating structured output formatting, such as JSON mode, with tool invocation, when only function calling produces a call the application can execute.

5
MCQeasy

You need to provide a generative AI solution that can answer questions based on a large set of PDF documents stored in Azure Blob Storage. The solution must support natural language queries and return citations from the documents. Which Azure service combination should you use?

A.Azure Machine Learning and Azure Kubernetes Service
B.Azure Cognitive Search and Azure OpenAI Service with 'on your data'
C.Azure AI Bot Service and Azure Functions
D.Azure AI Document Intelligence and Azure AI Translator
AnswerB

Azure Cognitive Search indexes the Blob-stored PDFs and retrieves relevant passages, while Azure OpenAI Service with 'on your data' grounds responses in those results and returns citations. This pairing satisfies both the natural-language query and citation requirements.

Why this answer

Azure Cognitive Search provides the indexing and retrieval capabilities for the PDF content, while Azure OpenAI Service with the 'on your data' feature enables natural language querying and generates answers grounded in the indexed documents, including citations. This combination directly supports the requirement to query a large set of PDFs in Azure Blob Storage and return citations.

Exam trap

The trap here is that candidates often confuse Azure AI Document Intelligence (which extracts text) with the full search-and-generate pipeline, overlooking that a search index (Cognitive Search) and a generative model (Azure OpenAI) are both required to answer natural language queries with citations.

How to eliminate wrong answers

Option A is wrong because Azure Machine Learning and Azure Kubernetes Service are designed for training and deploying custom ML models, not for out-of-the-box document indexing, natural language querying, or citation generation from PDFs. Option C is wrong because Azure AI Bot Service and Azure Functions are used for building conversational bots and serverless compute, but they lack native capabilities for indexing PDF content and returning citations from documents. Option D is wrong because Azure AI Document Intelligence extracts text and structure from documents, and Azure AI Translator handles language translation, but neither service provides the search indexing or generative AI querying needed to answer natural language questions with citations.

6
MCQhard

Your team is developing an AI-powered document summarization solution using Azure OpenAI. You need to ensure that the solution complies with Microsoft's Responsible AI principles, specifically transparency. Which configuration should you implement?

A.Configure diagnostic logging to capture all model inputs and outputs.
B.Fine-tune the model on a custom dataset to improve accuracy.
C.Enable content filtering with severity levels high and medium.
D.Add a system message that informs users the summary is generated by AI.
AnswerD

A system message disclosing AI generation directly satisfies transparency by informing users they are reading machine-generated output. Unlike content filters or metadata logging, which address harm prevention and auditability, this mechanism makes the AI's role visible at the point of interaction, meeting the stem's explicit transparency requirement.

Why this answer

Transparency under Microsoft's Responsible AI principles requires that users are aware when they are interacting with an AI system. Adding a system message that explicitly states the summary is AI-generated fulfills this disclosure requirement. Diagnostic logging (A) aids in accountability and debugging but does not directly inform the user.

Fine-tuning (B) improves accuracy but does not address transparency. Content filtering (C) mitigates harmful outputs but does not disclose AI involvement.

Exam trap

The trap here is that candidates confuse 'transparency' with 'accountability' or 'safety' and select diagnostic logging or content filtering, not realizing that transparency specifically requires user-facing disclosure of AI involvement.

How to eliminate wrong answers

Option A is wrong because diagnostic logging captures inputs and outputs for auditing and debugging, but it does not communicate to the end user that the content is AI-generated, which is the core requirement of transparency. Option B is wrong because fine-tuning the model on a custom dataset enhances performance and relevance but has no role in informing users about AI authorship. Option C is wrong because enabling content filtering with severity levels high and medium is a safety measure to block harmful content, not a mechanism for disclosing AI involvement to users.

7
MCQeasy

You are using Azure OpenAI Service to generate product descriptions. The output is often too verbose. You need to reduce the length of generated text without changing the model. Which parameter should you adjust?

A.Max tokens
B.Frequency penalty
C.Temperature
D.Top-p (nucleus sampling)
AnswerA

Max tokens caps the total tokens the model may generate, directly truncating verbose completions without altering the deployed model. Setting a lower value constrains output length, satisfying the requirement to shorten product descriptions while leaving the model itself unchanged.

Why this answer

Max tokens controls the total length of the generated response by capping the number of tokens (words/subwords) the model can output. Reducing this value directly truncates the output, making descriptions shorter without altering the model or its behavior. Other parameters influence randomness or repetition but do not enforce a strict length limit.

Exam trap

The trap here is that candidates confuse parameters that affect output style (temperature, top-p, frequency penalty) with the one parameter that directly controls output length (max tokens), leading them to choose a parameter that changes how the model writes rather than how much it writes.

How to eliminate wrong answers

Option B (Frequency penalty) is wrong because it reduces the likelihood of repeating the same tokens or phrases, which can change content style but does not enforce a maximum output length. Option C (Temperature) is wrong because it controls the randomness of token selection (higher = more creative, lower = more deterministic), not the number of tokens generated. Option D (Top-p) is wrong because it limits the cumulative probability of token choices (nucleus sampling), affecting diversity but not the total token count.

8
MCQeasy

A developer wants to use Azure OpenAI to generate text from a prompt. Which parameter controls the diversity of the generated output?

A.presence_penalty
B.frequency_penalty
C.temperature
D.max_tokens
AnswerC

Temperature scales the sampling distribution's randomness: near zero the model picks the highest-probability token almost deterministically, while higher values flatten probabilities and increase diversity. It is the parameter controlling output variety, not frequency penalty or top_p alone.

Why this answer

Temperature is the parameter that directly controls the randomness or diversity of the generated output by scaling the logits before applying the softmax function. A higher temperature (e.g., 1.0) increases the probability of less likely tokens, producing more creative and varied responses, while a lower temperature (e.g., 0.1) makes the output more deterministic and focused.

Exam trap

The trap here is that candidates often confuse frequency_penalty or presence_penalty with controlling diversity, but those parameters address repetition and topic novelty, not the fundamental randomness of token selection, which is exclusively governed by temperature.

How to eliminate wrong answers

Option A is wrong because presence_penalty penalizes tokens that have already appeared in the text so far, encouraging the model to introduce new topics or avoid repetition, but it does not directly control the overall diversity or randomness of the token selection. Option B is wrong because frequency_penalty reduces the likelihood of tokens based on how frequently they have occurred in the generated text, which helps prevent repetitive phrasing but does not adjust the probability distribution's entropy like temperature does. Option D is wrong because max_tokens simply limits the maximum number of tokens in the generated response and has no effect on the diversity or randomness of the output.

9
Multi-Selectmedium

Which THREE are valid parameters when calling the Azure OpenAI Service chat completions API?

Select 3 answers
A.messages
B.index_name
C.temperature
D.max_tokens
E.embedding_model
AnswersA, C, D

Required input.

Why this answer

The `messages` parameter is required in the Azure OpenAI Service chat completions API call. It defines the conversation history and user input as an array of message objects, each with a `role` (system, user, assistant) and `content`. Without this parameter, the API cannot determine the context or prompt for generating a response.

Exam trap

Microsoft often tests the distinction between parameters for different Azure OpenAI API endpoints (chat completions vs. embeddings vs. search), so candidates may confuse `index_name` or `embedding_model` as valid chat completions parameters due to their familiarity with other Azure AI services.

10
Multi-Selecthard

You are implementing a generative AI solution using Azure OpenAI Service. The solution must generate responses that are grounded in your organization's proprietary documents and must return citations that link back to the source documents. You need to configure the deployment to meet these requirements. Which two actions should you perform? (Choose two.)

Select 2 answers
A.Create an Azure AI Search index that contains the proprietary documents with retrievable content and citation fields such as title and URL.
B.Enable diagnostic logging on the Azure OpenAI resource to capture prompt and completion text.
C.Fine-tune the model on the proprietary documents so it can reproduce them verbatim.
D.Associate the Azure AI Search index as a data source on the model deployment by using the Azure OpenAI On Your Data configuration.
E.Deploy a second model instance in a different Azure region for high availability.
AnswersA, D

Grounding and citations require a searchable index whose documents expose content plus metadata fields like title and URL. Azure AI Search provides the retrieval layer that the Azure OpenAI On Your Data feature queries. Without an index containing retrievable content and citation fields, the service cannot retrieve relevant chunks or produce source links, so this action is required.

Why this answer

Grounded responses with citations in Azure OpenAI require a retrieval source and a deployment-level data source binding. An Azure AI Search index holds the proprietary documents with content and citation metadata, and associating that index with the model deployment through On Your Data enables service-side retrieval, prompt injection, and citation generation.

Exam trap

The trap here is assuming fine-tuning can supply both grounding and citations, when citations depend on a search index and a data source binding rather than on model training.

11
MCQmedium

A company uses Microsoft Copilot for Microsoft 365 to automate email responses. They want to ensure that the Copilot responses comply with their data governance policies and do not expose sensitive information. What should they configure?

A.Microsoft Entra ID conditional access policies
B.Microsoft Purview Data Loss Prevention (DLP) policies
C.Microsoft Intune app protection policies
D.Microsoft Sentinel analytics rules
AnswerB

Microsoft Purview DLP policies inspect Copilot-generated content in Exchange Online and SharePoint, blocking responses that contain sensitive information types or sensitivity labels defined in your governance rules. This directly enforces the compliance constraint by preventing exposure at generation time, rather than relying on post-hoc auditing or user discretion.

Why this answer

Microsoft Purview Data Loss Prevention (DLP) policies are designed to detect and prevent the accidental sharing of sensitive information, such as credit card numbers or personally identifiable information (PII), across Microsoft 365 services. By configuring DLP policies, the company can scan Copilot-generated email responses for sensitive data patterns and block or warn before the email is sent, ensuring compliance with data governance policies.

Exam trap

The trap here is that candidates often confuse data governance and content inspection with access control (Entra ID) or endpoint security (Intune), leading them to select a policy that governs who can access data rather than what data is allowed to be shared.

How to eliminate wrong answers

Option A is wrong because Microsoft Entra ID conditional access policies control authentication and access to applications based on user, device, or location conditions, but they do not inspect or govern the content of Copilot-generated email responses for sensitive data. Option C is wrong because Microsoft Intune app protection policies manage how data is handled within mobile apps (e.g., copy/paste restrictions, encryption at rest), but they do not scan or enforce content-level rules on email responses generated by Copilot. Option D is wrong because Microsoft Sentinel analytics rules are used for security information and event management (SIEM) to detect threats and anomalies in logs, not to enforce data governance or prevent sensitive data exposure in email content.

12
MCQhard

You are developing a generative AI solution using Azure OpenAI Service. The solution must generate product descriptions based on a set of attributes provided by the user. You need to ensure the output adheres to a specific JSON schema. Which feature should you use to enforce the structure of the model's response?

A.Function calling with a defined function that returns JSON.
B.Temperature set to 0 to make output deterministic.
C.Prompt engineering with few-shot examples of JSON.
D.Structured outputs with a JSON schema definition.
AnswerD

Structured outputs allow you to define a JSON schema that the model's response must follow. This ensures the generated content is valid JSON and matches the specified structure, which is ideal for generating product descriptions with consistent fields. It enforces the schema at the API level, reducing the need for post-processing.

Why this answer

Structured outputs in Azure OpenAI Service enable you to specify a JSON schema that the model's response must conform to. This guarantees that the generated output is valid JSON and includes all required fields, which is essential for generating product descriptions with a consistent format. It is a robust way to enforce structure without relying on prompt engineering alone.

Exam trap

The trap here is assuming that function calling or prompt engineering can guarantee a JSON schema, when only structured outputs provide that strict enforcement.

13
MCQmedium

A team is using Azure OpenAI Service to summarize long legal contracts. They observe that summaries sometimes miss clauses located near the end of the document. The contracts are far longer than the model's maximum context window. What should they implement to improve coverage of the entire document?

A.Lower the max_tokens parameter to force the model to read further into the document.
B.Increase the temperature parameter so the model explores more of the contract content.
C.Set the top_p parameter to 1.0 and rely on nucleus sampling to expand context coverage.
D.Split the contract into overlapping chunks, summarize each chunk, then combine the partial summaries into a final summary.
AnswerD

Chunking with overlap followed by a combine step is the map-reduce pattern for documents exceeding the context window. Each chunk is summarized within limits, and overlap preserves clauses that straddle boundaries. Merging partial summaries produces coverage across the whole contract, addressing the missed end-of-document clauses.

Why this answer

Documents longer than the context window must be processed in pieces. Overlapping chunking ensures clauses spanning boundaries survive, and a combine step merges per-chunk summaries into a coherent whole. Sampling parameters and output length settings operate on generation behavior and cannot extend how much source text the model receives.

Exam trap

The trap here is assuming sampling parameters such as temperature or top_p influence how much source content the model can read, when they only affect token selection.

14
Multi-Selectmedium

You are building a generative AI chatbot using Microsoft Copilot Studio. The chatbot must answer questions from a PDF document and a SQL database. Which THREE data sources can you configure? (Choose three.)

Select 3 answers
A.Custom connector to SQL database
B.SharePoint (store the PDF document)
C.Azure AI Search index
D.Dataverse (store SQL data)
E.Azure Blob Storage
AnswersA, B, D

Custom connectors allow integration with SQL databases.

Why this answer

A is correct because Microsoft Copilot Studio supports custom connectors to access external data sources like SQL databases. This allows the chatbot to query the SQL database directly using the connector's API, enabling real-time data retrieval for generative AI responses.

Exam trap

The trap here is that candidates often assume Azure AI Search or Blob Storage are natively configurable in Copilot Studio, but they require additional middleware or custom connectors, unlike SharePoint and Dataverse which are first-party supported sources.

15
MCQhard

Refer to the exhibit. You are troubleshooting an Azure OpenAI API call that is returning incomplete responses. The response stops mid-sentence. Which parameter should you adjust?

A.Increase max_tokens to 1000.
B.Remove the stop parameter.
C.Increase temperature to 1.0.
D.Increase top_p to 1.0.
AnswerA

max_tokens caps the number of tokens generated in the completion, so a low value truncates output mid-sentence. Raising it to 1000 lets the model finish, directly resolving the incomplete response. Temperature, top_p and presence penalties alter style, not truncation length.

Why this answer

The `max_tokens` parameter controls the maximum number of tokens the model can generate in a single response. When a response stops mid-sentence, it typically means the token limit was reached before the model could complete its output. Increasing `max_tokens` to 1000 provides more room for the model to finish its generation, resolving the truncation issue.

Exam trap

The trap here is that candidates confuse parameters that control output length (`max_tokens`) with those that control output diversity (`temperature`, `top_p`) or early stopping (`stop`), leading them to pick options that change style rather than capacity.

How to eliminate wrong answers

Option B is wrong because removing the `stop` parameter would not fix mid-sentence truncation; the `stop` parameter defines sequences that halt generation early, and removing it could actually make responses longer but does not address a hard token limit. Option C is wrong because increasing `temperature` to 1.0 increases randomness and creativity in the output, but does not affect the maximum length of the response; it could even lead to more verbose or erratic completions. Option D is wrong because increasing `top_p` to 1.0 enables nucleus sampling with all tokens considered, which may alter the diversity of the output but does not extend the token budget; the model will still stop when `max_tokens` is exhausted.

16
MCQmedium

You are deploying a generative AI solution that uses Azure OpenAI Service with your own data stored in Azure AI Search. Users report that answers are sometimes pulled from documents the user is not permitted to see, because the retrieval step searches the entire index. You need to ensure each user only receives answers grounded in documents they are authorized to access, without creating a separate index per user. What should you do?

A.Add a security filter field to the Azure AI Search index and pass the user's group or tenant identifier as an OData filter on each query.
B.Enable Azure OpenAI content filtering at High severity for both prompts and completions.
C.Create a separate Azure OpenAI deployment for each department and route users to the deployment matching their department.
D.Increase the chunk size of the indexed documents so that fewer documents are returned per query.
AnswerA

Azure AI Search supports filterable fields that can be combined with the search query, so indexing an access-control field (for example, allowed groups) and passing the caller's identity as an OData filter restricts retrieval to documents that user may see. This keeps a single shared index while enforcing per-user authorization at query time, which is exactly the requirement. The model then only receives permitted grounding content, so answers cannot cite restricted documents.

Why this answer

Authorization must be enforced where the documents are selected, not where the model generates text. Adding a filterable access-control field to the Azure AI Search index and applying the caller's identity as an OData filter causes the retrieval step to return only permitted chunks, so the model can never ground an answer in a restricted document. This preserves one shared index while honoring per-user permissions.

Exam trap

The trap here is assuming that Azure OpenAI content filtering or a separate model deployment provides document-level authorization, when access control must actually be applied as a filter in the retrieval index.

17
MCQhard

You operate a RAG assistant on Azure OpenAI Service that answers questions over a product catalog. The index contains thousands of items, and users complain that answers mix details from unrelated products. You need retrieval to consider semantic meaning and handle paraphrased queries while returning only the most relevant items. What should you implement?

A.Faceted navigation that filters results by product category
B.Scoring profiles that boost documents containing a specific field value
C.Vector search using embeddings generated by an Azure OpenAI embedding model
D.Keyword search using the Azure AI Search simple query syntax
AnswerC

Vector search compares the embedding of the query against embeddings of the catalog items, so semantically similar content ranks highly even when wording differs. This handles paraphrased queries and sharpens relevance, which reduces the chance of pulling details from unrelated products. It is the appropriate retrieval mode for the described semantic matching requirement.

Why this answer

Embedding-based vector search maps both the query and the catalog items into a semantic space, so similarity reflects meaning rather than exact wording. Paraphrased questions then retrieve the right products, and ranking by vector distance keeps unrelated items out of the context. Keyword, facet, and scoring-profile approaches remain term- or attribute-driven and cannot perform semantic matching.

Exam trap

The trap here is assuming that adding filters or scoring boosts delivers semantic understanding, when only embeddings capture meaning.

18
MCQmedium

You are implementing a RAG (Retrieval-Augmented Generation) solution using Azure AI Search and Azure OpenAI Service. The solution is returning answers that are not relevant to the user query. What is the most likely cause?

A.The max_tokens parameter is set too high.
B.The chunk size is too small.
C.The index includes too many documents.
D.The relevance score threshold is set too low.
AnswerD

A low relevance score threshold lets weakly matching chunks pass into the prompt, so Azure OpenAI grounds its answer on marginal content and returns off-topic responses. Raising the threshold filters these poor matches, directly addressing the irrelevant-answer symptom described in the stem.

Why this answer

A low relevance score threshold in Azure AI Search allows documents with low semantic or vector similarity to be returned as results. When these poorly matched documents are passed to Azure OpenAI Service for answer generation, the model may produce answers that are not relevant to the user query, as the retrieved context is noisy or unrelated.

Exam trap

The trap here is that candidates often confuse the relevance score threshold with other parameters like max_tokens or chunk size, assuming that irrelevant answers stem from generation limits or indexing granularity rather than retrieval quality.

How to eliminate wrong answers

Option A is wrong because the max_tokens parameter controls the length of the generated response, not the relevance of the retrieved content; setting it too high may cause truncation or cost issues but does not directly cause irrelevant answers. Option B is wrong because a chunk size that is too small typically leads to fragmented or incomplete context, which can reduce answer quality but is less likely to cause completely irrelevant answers compared to a low relevance threshold. Option C is wrong because including too many documents in the index does not inherently cause irrelevant answers; the search query and scoring mechanism determine which documents are retrieved, and a large index can still return relevant results if the threshold and ranking are properly configured.

19
MCQmedium

You are building an internal knowledge assistant with Azure OpenAI Service. Responses must be grounded in a curated set of HR policy documents stored in Azure AI Search, and every response must include a citation to the specific source chunk. You already deployed a GPT-4o model and created the search index. You need to add the grounding layer with the least development effort. What should you do?

A.Increase the model temperature and max tokens so the model has more room to recall the HR policies from its pretrained knowledge.
B.Call the Azure AI Search REST API from the client, concatenate the top results into the system message, and post-process the model output to append source links.
C.Fine-tune the GPT-4o deployment on the HR policy documents, then rely on the model to quote the correct policy section.
D.Enable the "On Your Data" feature on the model deployment and point it at the Azure AI Search index.
AnswerD

On Your Data is a built-in Azure OpenAI capability that connects a model deployment directly to an Azure AI Search index, retrieving relevant chunks and returning citations automatically. Because the index already exists, this requires only configuration rather than custom retrieval code, satisfying the least-effort requirement while still grounding every answer in the curated HR documents.

Why this answer

Grounding a deployment in a private index with automatic citations is exactly what the On Your Data capability provides for Azure OpenAI. Since the Azure AI Search index is already built, enabling this feature on the deployment satisfies the grounding and citation requirements with configuration instead of custom code, which matches the least-effort constraint.

Exam trap

The trap here is assuming that fine-tuning or larger token limits can substitute for retrieval when the requirement is verifiable citations from a private document set.

20
MCQhard

You run the Azure CLI command shown in the exhibit. After a few minutes, the deployment fails with a quota error. What is the most likely cause?

A.The SKU name 'Standard' is invalid for Azure OpenAI deployments.
B.The model version '0613' is deprecated and no longer available.
C.The requested capacity of 10 exceeds the available quota for the gpt-4 model in that region.
D.The resource group name 'myResourceGroup' does not exist.
AnswerC

The deployment failed because the requested capacity of 10 exceeds the available quota for the gpt-4 model in that region. Azure OpenAI enforces per-model, per-region capacity quotas, so requesting more than the allotted units triggers a quota error.

Why this answer

The quota error indicates that the requested capacity (10 units) for the gpt-4 model exceeds the available quota in the target region. Azure OpenAI deployments require sufficient model-specific quota, which is region- and model-specific. The error is not related to SKU name validity, model version deprecation, or resource group existence.

Exam trap

The trap here is that candidates might confuse a quota error with a model deprecation or SKU issue, but the error message's explicit mention of 'quota' directly points to capacity limits, not configuration or availability problems.

How to eliminate wrong answers

Option A is wrong because 'Standard' is a valid SKU name for Azure OpenAI deployments; the error message specifically mentions quota, not an invalid SKU. Option B is wrong because model version '0613' is a valid and available version for gpt-4; deprecation would produce a different error (e.g., 'ModelNotFound'), not a quota error. Option D is wrong because if the resource group did not exist, the Azure CLI would fail immediately with a 'ResourceGroupNotFound' error, not after several minutes with a quota error.

21
MCQeasy

You are developing a generative AI application that must comply with responsible AI principles. Which Azure AI service should you use to detect and filter harmful content in both input prompts and output responses?

A.Microsoft Purview
B.Azure AI Content Safety
C.Azure OpenAI Service
D.Azure AI Language
AnswerB

Azure AI Content Safety provides dedicated moderation APIs that classify harmful content across categories such as hate, violence and self-harm, applying them to both prompts and generated responses. This directly satisfies the responsible AI requirement to detect and filter harmful content at input and output, which general language services do not cover.

Why this answer

Azure AI Content Safety is the dedicated Azure service for detecting and filtering harmful content such as hate speech, violence, self-harm, and sexual content in both user prompts and AI-generated responses. It provides configurable severity levels and integrates directly with generative AI workflows to enforce responsible AI policies, making it the correct choice for this requirement.

Exam trap

Microsoft often tests the distinction between a service that provides AI capabilities (Azure OpenAI Service) and a service that enforces safety policies (Azure AI Content Safety), leading candidates to mistakenly choose the model provider instead of the dedicated safety tool.

How to eliminate wrong answers

Option A is wrong because Microsoft Purview is a data governance and compliance service focused on data classification, labeling, and auditing, not on real-time content safety filtering of AI inputs and outputs. Option C is wrong because Azure OpenAI Service provides the generative AI models themselves but does not include built-in content filtering; it relies on separate services like Azure AI Content Safety or its own content filters for safety. Option D is wrong because Azure AI Language offers natural language processing capabilities such as sentiment analysis, key phrase extraction, and language understanding, but it does not specialize in detecting or filtering harmful content in generative AI contexts.

22
MCQhard

You are deploying a conversational AI solution using Microsoft Copilot Studio. The solution must comply with organizational data loss prevention (DLP) policies by preventing sensitive data from being sent to the underlying Azure OpenAI model. What should you configure?

A.Configure content filters in Azure OpenAI Studio
B.Define DLP policies in Microsoft 365 compliance center and apply to Copilot Studio
C.Enable Azure AI Content Safety in the bot's generative AI configuration
D.Set the temperature parameter to 0 to reduce variability
AnswerB

DLP policies in M365 can block sensitive data from being sent to AI models.

Why this answer

Microsoft Copilot Studio integrates with Microsoft 365 DLP policies to prevent sensitive data from being sent to the underlying Azure OpenAI model. By defining DLP policies in the Microsoft 365 compliance center and applying them to Copilot Studio, you can enforce data loss prevention rules that block or restrict the transmission of sensitive information (e.g., credit card numbers, social security numbers) to the generative AI backend. This ensures compliance with organizational security requirements without modifying the AI model itself.

Exam trap

The trap here is that candidates confuse Azure AI Content Safety (which handles harmful content moderation) with DLP policies (which handle sensitive data protection), leading them to select Option C instead of the correct DLP-based approach.

How to eliminate wrong answers

Option A is wrong because content filters in Azure OpenAI Studio are designed to filter harmful or offensive content in model outputs, not to prevent sensitive data from being sent to the model as input; they operate on the response side, not the request side. Option C is wrong because Azure AI Content Safety is a service for detecting and filtering harmful content (e.g., hate speech, violence) in both inputs and outputs, but it does not enforce DLP policies or block sensitive data based on organizational compliance rules; it focuses on safety, not data loss prevention. Option D is wrong because setting the temperature parameter to 0 reduces the randomness of the model's responses, making them more deterministic, but it has no effect on preventing sensitive data from being sent to the model; it controls output variability, not input filtering.

23
MCQeasy

Your team is building an internal knowledge assistant with Azure OpenAI Service. Legal requires that every generated answer include a traceable reference to the source document so reviewers can verify claims. The documents already live in Azure AI Search, and you want the service to return citation metadata alongside the generated text. Which feature should you configure?

A.Azure OpenAI content filtering with annotating models enabled on the deployment.
B.A system message instructing the model to always mention the file name it used to answer.
C.Logging every request and response to Azure Monitor and querying the logs after the fact.
D.The On Your Data feature configured with an Azure AI Search data source, which returns citations with the response.
AnswerD

On Your Data connects the model to an Azure AI Search index and returns the generated answer together with citation entries that identify the retrieved documents and their content. This gives reviewers the traceable references legal requires without building custom retrieval plumbing. Because the documents already reside in Azure AI Search, this configuration fits the existing environment and produces the attribution metadata in the response payload.

Why this answer

On Your Data integrates an Azure AI Search index as the grounding source and returns the model's answer together with citation information identifying the retrieved documents. Because the documents already live in Azure AI Search and the requirement is traceable references in the output, this feature directly provides the needed attribution without custom retrieval code.

Exam trap

The trap here is assuming that a prompt instruction or diagnostic logging yields verifiable citations, when citation metadata must come from the retrieval integration that actually tracks which documents grounded the answer.

24
Multi-Selecthard

You are designing a generative AI solution using Azure OpenAI Service. The solution must support multiple languages and provide consistent quality across languages. Which THREE actions should you take?

Select 3 answers
A.Fine-tune the model on a dataset of a single language
B.Use a model that supports multiple languages (e.g., GPT-4)
C.Provide examples in multiple languages in the prompt
D.Set the temperature to 0 for all requests
E.Test the solution with representative prompts in each language
AnswersB, C, E

Multilingual models handle multiple languages natively.

Why this answer

GPT-4 is a multilingual model pre-trained on diverse language corpora, enabling it to generate coherent and contextually appropriate responses across many languages without additional fine-tuning. This ensures consistent quality by leveraging the model's inherent cross-lingual capabilities, which is essential for a generative AI solution that must support multiple languages.

Exam trap

The trap here is that candidates may think fine-tuning on a single language (Option A) is sufficient for multilingual support, or that setting temperature to 0 (Option D) universally improves consistency, when in fact these actions undermine the required cross-lingual quality and flexibility.

25
MCQeasy

You need to generate a summary of a long article using Azure OpenAI. The article is 10,000 tokens long. What should you do to fit the article within the model's context window?

A.Split the article into smaller sections and summarize each section separately.
B.Increase the temperature parameter.
C.Use a model with a smaller context window.
D.Set max_tokens to a lower value.
AnswerA

Splitting the article into smaller sections keeps each request within the model's context window, since a 10,000-token article exceeds it. Summarising each chunk separately, then optionally combining the partial summaries, avoids truncation and preserves coverage of the full content.

Why this answer

The article exceeds the model's context window (typically 4096 or 8192 tokens for GPT-3.5/4). Splitting the article into smaller sections and summarizing each separately allows you to process the entire content within the token limits, then combine the summaries for a final coherent output. This is a standard chunking strategy for long documents when using Azure OpenAI.

Exam trap

The trap here is that candidates confuse parameters that control output behavior (temperature, max_tokens) with the fundamental input token limit, leading them to incorrectly believe adjusting these parameters can bypass the context window restriction.

How to eliminate wrong answers

Option B is wrong because increasing the temperature parameter affects randomness and creativity of the output, not the input token limit; it does not help fit a long article into the context window. Option C is wrong because using a model with a smaller context window would make the problem worse, as it reduces the maximum input length, not increase it. Option D is wrong because setting max_tokens to a lower value only truncates the output length, not the input; the article still exceeds the context window and will be rejected or truncated at the input stage.

26
MCQhard

You deploy a GPT-4o model in Azure OpenAI Service for an internal assistant that summarizes long contracts. Legal requires that every generated summary be traceable to the exact source passages and that the assistant never answer from general world knowledge. You need the model to ground each statement in retrieved text and expose the supporting passages to the caller. What should you configure?

A.Enable the abuse monitoring logging option on the Azure OpenAI resource
B.Increase the model's max_tokens value to fit the entire contract
C.Set temperature to 0 and top_p to 1 on the chat completions call
D.Use the Azure OpenAI On Your Own Data feature with citations enabled
AnswerD

On Your Own Data connects the deployment to an Azure AI Search index and instructs the model to answer only from retrieved chunks, returning citations that map statements to source passages. This directly provides the traceability legal requires and constrains responses to the indexed contracts rather than general knowledge. It is the supported grounding pattern for this scenario.

Why this answer

Grounding the assistant in an indexed corpus and returning citations is what makes each summary statement traceable to a specific contract passage. Azure OpenAI On Your Own Data performs retrieval against Azure AI Search, injects the retrieved chunks, and returns citation metadata with the response. Sampling parameters, output length, and monitoring logs change behavior or observability but never bind statements to sources.

Exam trap

The trap here is believing that lowering temperature removes hallucination and therefore substitutes for retrieval-based grounding.

27
MCQmedium

You are building a generative AI application with Azure OpenAI Service. The application must generate responses that are grounded in a specific set of internal documents stored in an Azure AI Search index. You want to use the built-in 'On Your Data' feature. Which configuration parameter should you set to ensure the model retrieves relevant documents before generating a response?

A.Enable the 'temperature' parameter to a low value to make responses more factual.
B.Set the model deployment to a fine-tuned model trained on the documents.
C.Use the 'max_tokens' parameter to limit the response length to match document snippets.
D.Set the data source type to Azure Cognitive Search and provide the index name.
AnswerD

Configuring the data source to Azure Cognitive Search with the correct index name enables the On Your Data feature to query the index for relevant documents before generating a response. This grounds the model's output in your proprietary data, reducing hallucinations and ensuring answers are based on the specified documents.

Why this answer

The On Your Data feature in Azure OpenAI Service allows the model to retrieve relevant documents from a specified data source, such as Azure Cognitive Search, before generating a response. By setting the data source type to Azure Cognitive Search and providing the index name, the model can ground its answers in the internal documents. This is the intended method for retrieval-augmented generation with Azure OpenAI.

Exam trap

The trap here is confusing fine-tuning with retrieval-augmented generation, assuming that training a model on documents is equivalent to dynamically retrieving them at inference time.

28
MCQmedium

You are building an internal assistant on Azure OpenAI that must stream responses to a React web app while hiding the model endpoint key from the browser. The app already authenticates users with Microsoft Entra ID. You need the least-privilege approach that keeps the key out of client code. What should you do?

A.Create an Azure API Management instance that fronts Azure OpenAI, and have the React app call API Management with the user's Entra ID token.
B.Give the React app a system-assigned managed identity and call Azure OpenAI directly from the browser.
C.Store the Azure OpenAI key in Azure Key Vault and have the React app retrieve it at runtime using the signed-in user's token.
D.Embed the Azure OpenAI key in the React build and call the completions endpoint directly from the browser.
AnswerA

API Management can validate the Entra ID token, apply rate limits and policies, and hold the Azure OpenAI key in a named value or Key Vault reference so the browser never sees it. Streaming still works because API Management supports server-sent events passthrough. This uses the existing identity provider and enforces least privilege at the gateway.

Why this answer

The key must never reach the browser, so a server-side intermediary is required. Azure API Management can authenticate the Entra ID token, enforce policies, and inject the Azure OpenAI key from a secure store while preserving streaming responses. Managed identities cannot run in a browser, and retrieving a key client-side exposes it regardless of where it was stored.

Exam trap

The trap here is treating managed identity or Key Vault as something a browser-based single-page app can use directly, when both require server-side Azure compute to function.

29
MCQmedium

You are deploying a generative AI assistant on Azure OpenAI Service that must summarize user-supplied financial reports. Security policy requires that report contents never leave your Azure tenant and that the model must not be fine-tuned. You need to provision the resource so that prompt and completion data are not retained for human review and are not used to train any shared model. What should you configure?

A.Set the Azure OpenAI resource's 'Content logging' to Disabled in Azure Monitor diagnostic settings.
B.Submit an Azure OpenAI limited access modification request to disable abuse monitoring for the subscription.
C.Create the deployment in a separate resource group and assign the Cognitive Services User role only to the application's managed identity.
D.Deploy the model with a content filter policy set to block high-severity categories and enable prompt shields.
AnswerB

Approved limited access modification for abuse monitoring removes the standard human review and retention of prompts and completions for the approved subscription. Because the requirement is that report contents are neither retained for review nor used to train shared models, this is the correct control. Fine-tuning is unaffected, and data stays within the tenant boundary.

Why this answer

The requirement is about data handling by the service itself, not about access control or content safety. The only mechanism that removes default human review and retention of prompts and completions for an Azure OpenAI resource is an approved limited access modification for abuse monitoring on the subscription. Content filters, diagnostic settings, and RBAC all address different concerns and leave the retention behavior unchanged.

Exam trap

The trap here is assuming that disabling diagnostic content logging or adding content filters changes how the service retains prompts, when data-handling behavior is governed by the abuse-monitoring modification process instead.

30
MCQeasy

You are using Azure OpenAI Service to generate marketing copy. You notice that the output sometimes contains factual inaccuracies about your company's products. Which action can you take to improve factual accuracy?

A.Lower the temperature to 0.
B.Include relevant product information in the system message.
C.Increase the maxTokens to 4000.
D.Add a stop sequence to limit output.
AnswerB

Embedding product facts in the system message grounds generation in authoritative context, directly addressing the factual-accuracy constraint. Unlike fine-tuning, which adjusts model weights, system-message grounding supplies reference material at inference time, so the model conditions its output on your company's actual product details rather than relying on potentially stale pretrained knowledge.

Why this answer

Including relevant product information in the system message provides the model with authoritative context that grounds its responses in factual data. The system message acts as a persistent instruction set that the model uses to shape its outputs, reducing reliance on its internal training data which may be outdated or incomplete. This technique, known as 'grounding,' directly improves factual accuracy by supplying the model with the specific facts it needs to generate correct marketing copy.

Exam trap

The trap here is that candidates often assume lowering temperature or increasing maxTokens will fix factual accuracy, when in reality these parameters control randomness and output length, not the correctness of the underlying information.

How to eliminate wrong answers

Option A is wrong because lowering the temperature to 0 reduces randomness and creativity but does not inject factual data; it only makes the model more deterministic in its token selection, which can still produce inaccuracies if the model lacks the correct information. Option C is wrong because increasing maxTokens to 4000 only extends the maximum length of the output, which does not address the root cause of factual errors and may even allow the model to generate more incorrect content. Option D is wrong because adding a stop sequence limits where the model stops generating text, which controls output length but does not improve the factual accuracy of the content produced.

31
Multi-Selecthard

Which THREE factors should you consider when choosing between Azure OpenAI Service and Azure Machine Learning for deploying a generative AI model?

Select 3 answers
A.Integration with Microsoft Purview for data governance.
B.Latency requirements: Azure OpenAI may offer lower latency for standard models.
C.Ability to scale to thousands of concurrent requests.
D.Need for custom model architecture: Azure ML supports custom models, Azure OpenAI uses pre-trained.
E.Operational overhead: Azure OpenAI is a fully managed service.
AnswersB, D, E

Azure OpenAI endpoints are optimized for low latency, whereas Azure ML may require additional optimization.

Why this answer

Azure OpenAI Service provides managed endpoints for pre-trained models like GPT-4, which are optimized for low-latency inference out of the box. In contrast, Azure Machine Learning requires you to deploy your own containerized model, which can introduce additional network and compute overhead, making Azure OpenAI the better choice when sub-100ms response times are critical for standard generative AI tasks.

Exam trap

The trap here is that candidates assume 'scalability' is unique to one service, but both Azure OpenAI and Azure ML can handle high concurrency; the real differentiator is latency and customizability, not raw throughput.

32
Multi-Selectmedium

Which TWO actions are required to enable a custom chatbot built with Azure OpenAI to answer questions based on a company's internal PDF documents?

Select 2 answers
A.Use Azure AI Document Intelligence to extract text from PDFs before indexing
B.Deploy Azure AI Content Safety to filter responses
C.Fine-tune the GPT model on the PDF content
D.Ingest the PDFs into an Azure Cognitive Search index
E.Configure the Azure OpenAI deployment to use 'Add your data' with the search index
AnswersD, E

Indexing enables retrieval of relevant content from PDFs.

Why this answer

Azure Cognitive Search provides the indexing and retrieval capabilities needed to make PDF content searchable. By ingesting PDFs into an Azure Cognitive Search index, the chatbot can perform vector or keyword searches over the extracted text, enabling it to retrieve relevant passages to answer user questions. This is the standard approach for grounding a custom chatbot on proprietary documents without modifying the underlying model.

Exam trap

The trap here is that candidates often confuse fine-tuning (option C) with the RAG pattern, mistakenly believing they must retrain the model on proprietary data, when in fact the 'Add your data' feature with a search index is the correct and simpler approach for question-answering over internal documents.

33
Multi-Selectmedium

Which TWO are valid ways to manage cost when using Azure OpenAI Service in a production application?

Select 2 answers
A.Fine-tune the model to reduce the number of examples needed in prompts
B.Increase the temperature parameter to 1.0
C.Use a smaller model like GPT-3.5-turbo instead of GPT-4 for simpler tasks
D.Provision more PTUs to get a lower rate per token
E.Set the max_tokens parameter to the minimum needed for the response
AnswersC, E

Smaller models have lower per-token costs.

Why this answer

Using a smaller model like GPT-3.5-turbo for simpler tasks directly reduces the per-token cost compared to GPT-4, which is significantly more expensive. Azure OpenAI Service charges based on model tier and token usage, so selecting the appropriate model for the task complexity is a primary cost management strategy.

Exam trap

The trap here is that candidates may confuse fine-tuning with prompt optimization, or assume that increasing PTUs lowers per-token cost, when in fact PTUs are a fixed-cost commitment that increases total expenditure.

34
MCQmedium

Refer to the exhibit. You are deploying an Azure AI Services resource using an ARM template. After deployment, you cannot access the resource from any client, including the Azure portal. What is the most likely cause?

A.The SKU S0 does not support network restrictions.
B.The cognitiveServiceName conflicts with an existing resource.
C.The networkAcls defaultAction is set to Deny with no allowed IP or virtual network rules.
D.The virtualNetworkRules array is empty.
AnswerC

Setting networkAcls defaultAction to Deny without any allow rules blocks all traffic, including the Azure portal, since no IP or subnet is permitted. This explains why every client fails to reach the Azure AI Services resource after deployment.

Why this answer

The networkAcls defaultAction set to Deny with no allowed IP or virtual network rules blocks all traffic to the Azure AI Services resource, including requests from the Azure portal and any client. This is because the default network access control list (ACL) denies all traffic unless explicitly permitted by IP rules or virtual network rules, rendering the resource inaccessible even for management operations.

Exam trap

The trap here is that candidates often assume an empty virtualNetworkRules array means no restrictions, but the defaultAction property controls the default behavior, and setting it to Deny with no allow rules blocks all traffic regardless of empty arrays.

How to eliminate wrong answers

Option A is wrong because the S0 SKU does support network restrictions; network ACLs are available for all paid SKUs (S0 and above), and the issue is not related to SKU limitations. Option B is wrong because a name conflict would cause a deployment error, not a scenario where the resource is deployed but inaccessible; the ARM template would fail with a conflict error during deployment. Option D is wrong because an empty virtualNetworkRules array is irrelevant when the defaultAction is Deny; the resource would still be blocked since no IP rules are specified to allow traffic, and an empty array does not implicitly permit any traffic.

35
MCQeasy

You are building a conversational AI system using Azure OpenAI Service. The system must maintain context across multiple user turns. Which parameter determines how many previous messages are considered for the next response?

A.max_tokens
B.temperature
C.top_p
D.The length of the messages array in the API call
AnswerD

The messages array carries the full conversation history sent to the model, so its length directly bounds how many prior turns inform the next completion. Truncating or extending that array controls retained context; no separate memory parameter exists in the Chat Completions API.

Why this answer

The Azure OpenAI Service API uses a `messages` array in the request body to represent the conversation history. Each entry in this array corresponds to a previous turn (with roles like 'user', 'assistant', or 'system'), and the length of this array directly determines how many prior messages are considered when generating the next response. By including more messages, you extend the context window; by truncating the array, you limit it.

Exam trap

The trap here is that candidates often confuse parameters that control output generation (like `max_tokens`, `temperature`, or `top_p`) with the mechanism for maintaining conversation history, which is explicitly managed by the structure of the API call's `messages` array.

How to eliminate wrong answers

Option A is wrong because `max_tokens` controls the maximum number of tokens (words/subwords) in the generated response, not the number of previous messages considered for context. Option B is wrong because `temperature` is a sampling parameter that influences the randomness or creativity of the output, not the conversation history length. Option C is wrong because `top_p` (nucleus sampling) sets a probability threshold for token selection, affecting output diversity, not the number of prior turns used as context.

36
Multi-Selectmedium

Which TWO actions should you take to ensure that a generative AI model deployed on Azure Machine Learning is compliant with data privacy regulations?

Select 2 answers
A.Implement data masking during preprocessing.
B.Store all training data in a separate Azure region.
C.Use differential privacy during model training.
D.Encrypt the model at rest using Azure Key Vault.
E.Log all raw input data to Azure Monitor for auditing.
AnswersA, C

Masking replaces sensitive field values such as names or card numbers with obfuscated substitutes before the data reaches training, so personally identifiable information never enters the model. This directly satisfies the regulation's requirement that personal data be de-identified at the preprocessing stage.

Why this answer

Options A and C are correct. A: Data masking during preprocessing helps anonymize sensitive data before it is used, reducing privacy risks. C: Differential privacy adds noise to the training process, preventing the model from memorizing individual data points.

B is wrong because storing data in a separate Azure region does not inherently address privacy compliance; it may help with data residency but not the core privacy protection. D is wrong because encrypting the model at rest with Azure Key Vault is a security measure, not a privacy compliance measure. E is wrong because logging raw input data to Azure Monitor could expose sensitive information, violating privacy.

37
MCQhard

Your team is using the Azure OpenAI Batch API to process 200,000 product-description rewrites overnight. The job must complete within a fixed window and cost as little as possible, and results are not needed interactively. A developer reports that the submitted batch job failed with an error about the input file. What is the most likely cause?

A.The batch job used a deployment with a lower tokens-per-minute quota than the interactive deployment.
B.The input file was uploaded as a single large JSON array containing all 200,000 requests.
C.The input file was uploaded as a JSONL file with one request object per line and a unique custom_id for each line.
D.The batch job was submitted without specifying a completion window, so it defaulted to the maximum allowed.
AnswerB

Batch API input must be a JSONL file where each line is one request object with its own custom_id. A single JSON array is not the accepted format, so the job fails validation with an input-file error. Converting the payload to line-delimited request objects and uploading it to the expected storage location resolves the failure.

Why this answer

The Batch API expects a JSONL input file in which every line is an independent request object with a unique custom_id. Submitting one large JSON array violates that contract, so the job is rejected during input validation before any processing begins.

Exam trap

The trap here is assuming any valid JSON is acceptable, when the Batch API specifically requires line-delimited JSONL with one request per line.

38
MCQmedium

A company is using Azure OpenAI to generate customer support responses. They want to ensure the model does not use any personally identifiable information (PII) in its outputs. What should they implement?

A.Fine-tune the model on anonymized data.
B.Use prompt engineering to instruct the model to redact PII.
C.Use Azure AI Content Safety to filter PII from the output.
D.Use a system message instructing the model to avoid PII.
AnswerC

Azure AI Content Safety can detect and block PII.

Why this answer

Azure AI Content Safety provides built-in PII detection and redaction capabilities that can automatically scan and filter sensitive information from model outputs. This is the most reliable approach because it operates as a post-processing filter, catching PII that the model might generate despite instructions. Fine-tuning, prompt engineering, and system messages are all fallible because they rely on the model's compliance rather than enforced filtering.

Exam trap

The trap here is that candidates confuse 'instruction-based approaches' (prompts, system messages) with 'enforcement-based approaches' (content safety filters), assuming that telling the model not to do something is as effective as actively filtering the output.

How to eliminate wrong answers

Option A is wrong because fine-tuning on anonymized data does not prevent the model from generating PII during inference; it only reduces the likelihood based on training data, and the model can still hallucinate or leak PII from its pretrained knowledge. Option B is wrong because prompt engineering is a soft instruction that the model may ignore or fail to apply consistently, especially with edge cases or adversarial inputs, and it does not provide guaranteed redaction. Option D is wrong because a system message is merely a directive to the model, not a technical enforcement mechanism; the model can still output PII if it misinterprets or overrides the instruction.

39
MCQeasy

You are using Azure OpenAI Service to summarize customer emails. The summaries must be concise and contain only key information. Which prompt engineering technique should you apply?

A.Use few-shot prompting with examples of desired summaries
B.Use chain-of-thought prompting
C.Use zero-shot prompting with a one-sentence instruction
D.Use negative prompting to avoid verbose output
AnswerA

Few-shot prompting supplies the model with paired examples of emails and their concise summaries, directly demonstrating the desired length and content selection. This satisfies the stem's constraint that summaries contain only key information, since the exemplars establish the pattern the model imitates rather than relying on vague instructions alone.

Why this answer

Few-shot prompting is the correct technique because it provides the model with explicit examples of desired input-output pairs (e.g., a verbose email and its concise summary). This guides the model to learn the exact format, tone, and level of detail required for the summaries, which is critical for consistency in a production summarization pipeline. Without examples, the model may default to its training distribution and produce overly verbose or irrelevant output.

Exam trap

The trap here is that candidates often assume a simple instruction (zero-shot) is sufficient for summarization, underestimating how much the model relies on explicit examples to enforce output structure and conciseness, especially when the task requires domain-specific key information extraction.

How to eliminate wrong answers

Option B is wrong because chain-of-thought prompting is designed for multi-step reasoning tasks (e.g., math word problems, logical deduction) where intermediate steps are needed, not for summarization where the goal is direct extraction of key information. Option C is wrong because zero-shot prompting with a one-sentence instruction lacks the concrete examples needed to constrain the model's output style and length, often resulting in summaries that are too long or miss critical details. Option D is wrong because negative prompting (e.g., 'do not be verbose') is unreliable; the model may misinterpret the negation or still produce verbose output because it lacks positive examples of the desired concise format.

40
MCQeasy

You are developing a generative AI application that uses Azure OpenAI Service to summarize large documents. The application experiences high latency when processing requests. You need to reduce the latency without changing the model. What should you do?

A.Increase the temperature parameter
B.Increase the top_p parameter
C.Reduce the max_tokens parameter in the API request
D.Increase the max_tokens parameter
AnswerC

max_tokens caps generated output length, so lowering it shortens generation time and reduces latency without altering the deployed model. The stem forbids changing the model, and this parameter is a per-request setting, making it the appropriate lever for summarisation workloads.

Why this answer

Reducing the max_tokens parameter limits the length of the generated response, which directly reduces the processing time required by the Azure OpenAI Service to produce the output. Since latency is caused by the model generating a long sequence of tokens, capping the output tokens decreases the number of autoregressive decoding steps, thereby lowering response time without altering the underlying model.

Exam trap

The trap here is that candidates often confuse parameters that affect output length (max_tokens) with those that affect output diversity (temperature, top_p), mistakenly believing that adjusting randomness can speed up generation.

How to eliminate wrong answers

Option A is wrong because increasing the temperature parameter controls the randomness of the output, not the length or speed of generation; it has no direct impact on latency. Option B is wrong because increasing the top_p parameter (nucleus sampling) affects the diversity of token selection but does not reduce the number of tokens generated or the processing time. Option D is wrong because increasing the max_tokens parameter would allow longer responses, which would increase the number of decoding steps and worsen latency, the opposite of the desired outcome.

41
MCQhard

You are building a generative AI solution with Azure OpenAI Service that uses the GPT-4 model. The solution must process user requests and call external APIs to retrieve real-time data. You need to ensure the model can invoke the correct API based on user intent and return the results in a structured format. Which feature should you implement?

A.Prompt engineering with few-shot examples of API calls.
B.Using the Assistants API with code interpreter enabled.
C.Fine-tuning the model on a dataset of API calls and responses.
D.Function calling with the 'tools' parameter in the Chat Completions API.
AnswerD

Function calling allows the model to output a JSON object specifying which function to call and with what arguments. You define functions in the 'tools' parameter, and the model decides when to invoke them. This enables integration with external APIs and structured data retrieval, directly meeting the requirement.

Why this answer

Function calling is the native feature in Azure OpenAI that allows the model to request invocation of external functions with structured arguments. By defining tools, the model can determine when to call an API and with what parameters. This provides reliable integration with external systems and returns structured responses, fulfilling the requirement.

Exam trap

The trap here is confusing function calling with code interpreter or fine-tuning, which do not provide dynamic, structured API invocation.

42
Multi-Selecteasy

Which TWO Azure AI services can be used to build a conversational chatbot that uses generative AI? (Choose two.)

Select 2 answers
A.Azure AI Search
B.Azure AI Bot Service
C.Azure AI Translator
D.Azure AI Language
E.Azure OpenAI Service
AnswersB, E

Azure AI Bot Service provides the conversational orchestration layer, handling channels, turn management and dialog state, while integrating generative AI models through its SDK. It satisfies the stem's requirement for building a chatbot, since the service supplies the bot framework and hosting that generative responses plug into.

Why this answer

Azure AI Bot Service (B) provides a dedicated platform for building, deploying, and managing conversational chatbots, integrating with channels like Microsoft Teams and Slack. Azure OpenAI Service (E) enables generative AI capabilities by providing access to large language models (e.g., GPT-4) that can generate human-like responses, making it ideal for powering the conversational intelligence of a chatbot.

Exam trap

The trap here is that candidates may confuse Azure AI Language (which includes conversational language understanding for intent recognition) with a full chatbot builder, but it lacks the generative AI and channel integration that Azure AI Bot Service and Azure OpenAI Service provide together.

43
MCQeasy

You are developing a solution that uses Azure OpenAI to generate customer support responses. You want to prevent the model from repeating the same phrases. Which parameter should you adjust?

A.top_p
B.presence_penalty
C.temperature
D.frequency_penalty
AnswerD

Frequency_penalty applies a proportional penalty to tokens each time they appear, directly reducing verbatim repetition across the generated response. It satisfies the stem's constraint of preventing repeated phrases, unlike presence_penalty, which penalises only first occurrence and encourages new topics rather than curbing recurrence.

Why this answer

The frequency_penalty parameter (option D) is correct because it directly reduces the likelihood of the model repeating the same phrases by penalizing tokens that have already appeared in the generated text. A higher frequency_penalty value (e.g., 0.5 to 1.0) decreases the probability of reusing tokens, making the output more diverse and less repetitive. This is specifically designed to address repetition in generative AI responses.

Exam trap

The trap here is that candidates often confuse presence_penalty with frequency_penalty, but presence_penalty only penalizes tokens that have appeared at least once (regardless of count), while frequency_penalty penalizes based on the actual frequency of occurrence, making it the correct choice for preventing repeated phrases.

How to eliminate wrong answers

Option A is wrong because top_p (nucleus sampling) controls the cumulative probability threshold for token selection, influencing randomness and diversity of output, but it does not specifically penalize repeated phrases. Option B is wrong because presence_penalty penalizes tokens that have appeared at least once in the text, encouraging the model to talk about new topics, but it does not target the frequency of repetition of the same phrases. Option C is wrong because temperature controls the randomness of token selection by scaling the logits before softmax, affecting creativity and variability, but it has no direct mechanism to prevent repetition of phrases.

44
MCQhard

You are reviewing an ARM template for deploying Azure OpenAI Service. The template includes a deployment for gpt-35-turbo with a capacity of 100. You need to ensure that the deployment uses provisioned throughput instead of standard. What should you modify?

A.Change the sku name to 'ProvisionedManaged'.
B.Remove the raiPolicyName property.
C.Increase the capacity to 200.
D.Change the model format to 'GPT-4'.
AnswerA

Provisioned throughput requires the deployment's SKU name to be 'ProvisionedManaged' rather than 'Standard'; capacity then represents provisioned throughput units. Changing the sku name satisfies the stem's requirement to switch from standard to provisioned throughput deployment.

Why this answer

To use provisioned throughput (PTU) with Azure OpenAI Service, you must set the SKU name to 'ProvisionedManaged' in the ARM template. The default SKU is 'Standard', which uses pay-per-token consumption. Changing the SKU name to 'ProvisionedManaged' tells the resource provider to allocate dedicated throughput capacity for the deployment, ensuring consistent latency and throughput regardless of other workloads.

Exam trap

The trap here is that candidates often think increasing capacity or changing the model version enables provisioned throughput, but the exam tests the specific SKU name 'ProvisionedManaged' as the only way to switch from standard to provisioned throughput in an ARM template.

How to eliminate wrong answers

Option B is wrong because removing the raiPolicyName property does not affect throughput provisioning; it only removes content filtering or responsible AI policies, which are unrelated to capacity allocation. Option C is wrong because increasing capacity to 200 only scales the number of tokens per minute under the current SKU (Standard), but does not change the SKU to provisioned throughput; PTU requires the SKU name change, not just a higher capacity value. Option D is wrong because changing the model format to 'GPT-4' does not enable provisioned throughput; PTU is a SKU-level setting independent of the model version, and GPT-4 can also be deployed with Standard SKU.

45
MCQmedium

You are building a generative AI application with Azure OpenAI Service. The application must use your own product catalog data stored in an Azure AI Search index to ground model responses. You need to configure the model deployment so that retrieved documents are automatically injected into the prompt without you manually assembling the context in application code. What should you configure on the model deployment?

A.Add a 'data source' configuration (your Azure AI Search index) to the model deployment by using the Azure OpenAI 'On Your Data' feature.
B.Increase the deployment's tokens-per-minute quota so that the entire product catalog can be sent in each request.
C.Enable content filtering on the deployment to restrict responses to catalog topics.
D.Set the model temperature to 0 so the model always returns deterministic answers from the catalog.
AnswerA

Configuring the model deployment with a data source such as an Azure AI Search index uses the Azure OpenAI On Your Data capability. The service automatically retrieves relevant chunks and injects them into the prompt, and can return citations, so the application does not need to manually assemble context. This matches the requirement for grounding on the product catalog index without custom retrieval code.

Why this answer

The Azure OpenAI On Your Data feature lets you attach a supported data source, such as an Azure AI Search index, directly to a model deployment. At inference time the service performs retrieval, injects the relevant chunks into the prompt, and can return citations. This avoids writing custom retrieval and prompt-assembly logic while grounding responses in the product catalog.

Exam trap

The trap here is assuming that generation settings such as temperature or quota changes can provide grounding, when grounding requires an explicit data source configuration on the deployment.

46
MCQmedium

You are developing a generative AI application using Azure OpenAI Service. The application must generate summaries of customer emails and then extract action items. You want to minimize the number of API calls and ensure the model outputs structured JSON. What should you do?

A.Fine-tune the model to output JSON with both summary and action items.
B.Use the Azure OpenAI function calling feature to define a function that returns both summary and action items.
C.Use a single prompt that asks the model to summarize the email and extract action items, and specify the output format as JSON in the prompt.
D.Use two separate prompts: one for summarization and one for action item extraction, and combine the results in code.
AnswerC

Combining both tasks into one prompt reduces API calls and latency. By instructing the model to output JSON, you get structured data that is easy to parse. This approach leverages the model's ability to handle multiple instructions in one request, which is efficient and meets the requirement for structured output.

Why this answer

The most efficient method is to use a single prompt that instructs the model to both summarize the email and extract action items, with the output specified as JSON. This reduces API calls to one, lowers latency, and provides structured data for easy parsing. It leverages the model's ability to handle multiple instructions and format outputs as requested.

Exam trap

The trap here is assuming that function calling or fine-tuning is needed for structured output, when prompt engineering with JSON instructions suffices.

47
MCQeasy

Refer to the exhibit. You are reviewing the configuration of an Azure OpenAI Service resource. The resource is configured with customer-managed keys for encryption. What is the primary benefit of this configuration?

A.Enhanced control over data encryption keys
B.Simplified deployment process
C.Improved model performance
D.Reduced operational costs
AnswerA

Customer-managed keys let you supply and rotate your own Key Vault key, so Microsoft cannot decrypt the data. This satisfies the requirement for control over the encryption key lifecycle, rather than relying on Microsoft-managed keys.

Why this answer

Customer-managed keys (CMK) allow you to control and manage the encryption keys used to protect your data at rest in Azure OpenAI Service. This provides enhanced control over who can access the keys, when they are rotated, and how they are stored, which is critical for meeting compliance and security requirements. The primary benefit is not performance, cost, or deployment simplicity, but rather the ability to enforce your own key lifecycle and access policies.

Exam trap

The trap here is that candidates often confuse customer-managed keys with platform-managed keys, assuming the primary benefit is cost savings or performance gains, when in reality the core advantage is granular control over encryption key governance and compliance.

How to eliminate wrong answers

Option B is wrong because customer-managed keys add complexity to the deployment process (you must create and manage a Key Vault, set permissions, and configure key rotation), not simplify it. Option C is wrong because encryption keys have no impact on model inference speed or accuracy; performance is determined by model size, token limits, and compute resources. Option D is wrong because CMK typically increases operational costs due to the need for additional Key Vault resources, key management overhead, and potential charges for key operations.

48
MCQeasy

You need to monitor usage and costs of your Azure OpenAI Service deployments. Which Azure tool should you use?

A.Azure Cost Management + Billing
B.Azure Monitor
C.Azure Service Health
D.Azure Advisor
AnswerA

Azure Cost Management + Billing provides native cost analysis and budgets for Azure OpenAI Service resources, satisfying the requirement to monitor both usage and spend. It aggregates consumption data per deployment, letting you track token-based charges and set alerts, which dedicated monitoring tools like Azure Monitor cannot do for billing.

Why this answer

Azure Cost Management + Billing is the correct tool for monitoring usage and costs of Azure OpenAI Service deployments because it provides detailed cost analysis, budget tracking, and usage reports across all Azure services. It allows you to set budgets, create cost alerts, and analyze spending patterns specifically for OpenAI model deployments, including per-model and per-region cost breakdowns.

Exam trap

The trap here is that candidates often confuse Azure Monitor (which tracks performance metrics like token usage and latency) with cost monitoring, but Azure Monitor does not provide billing data or cost analysis, which is the specific requirement in this question.

How to eliminate wrong answers

Option B (Azure Monitor) is wrong because it focuses on performance metrics, logs, and alerts for application health and resource utilization, not on cost tracking or billing data. Option C (Azure Service Health) is wrong because it monitors service-level issues, outages, and planned maintenance across Azure services, not usage or cost metrics. Option D (Azure Advisor) is wrong because it provides best-practice recommendations for optimizing cost, performance, and reliability, but it does not directly monitor or report on actual usage and costs in real time.

49
MCQhard

Your organization uses Azure AI Document Intelligence to extract data from invoices. The extraction accuracy for total amounts is low. You have a labeled dataset of 500 invoices. You need to improve the model's accuracy for the 'total amount' field. What should you do?

A.Add additional predefined models for invoice processing.
B.Enable OCR enhancement to improve text recognition.
C.Increase the confidence threshold for the total amount field.
D.Create a custom neural model and train it with the labeled dataset.
AnswerD

A custom neural model learns field-specific patterns from your labelled invoices, including layout and contextual cues around the total amount, which the prebuilt model handles poorly. Training it on the 500 labelled samples directly targets the low-accuracy field, satisfying the requirement to improve extraction for 'total amount'.

Why this answer

Azure AI Document Intelligence's custom neural model is specifically designed to improve extraction accuracy for fields like 'total amount' by training on labeled datasets. Unlike the prebuilt invoice model, a custom neural model learns the unique layout and variations in your invoices, directly addressing low accuracy for a specific field. Training with 500 labeled invoices provides sufficient data to fine-tune the model's extraction capabilities.

Exam trap

Microsoft often tests the misconception that adjusting confidence thresholds or adding more predefined models can improve extraction accuracy, when in fact only custom training with labeled data addresses field-specific low accuracy.

How to eliminate wrong answers

Option A is wrong because adding additional predefined models does not improve accuracy for a specific field; predefined models are fixed and cannot be retrained or customized for your data. Option B is wrong because OCR enhancement improves text recognition quality but does not address the model's ability to correctly interpret and extract the 'total amount' field from the recognized text. Option C is wrong because increasing the confidence threshold only filters out low-confidence predictions, it does not improve the underlying model's extraction accuracy; it may reduce false positives but will not correct mis-extractions.

50
MCQhard

You run the Azure CLI command shown in the exhibit to create an online endpoint for a generative AI model. The deployment fails because the selected VM instance type is not available in the East US region. Which action should you take to resolve the issue?

A.Increase the instance count to 2
B.Specify a different VM type or region that supports Standard_NC6s_v3
C.Use a batch endpoint instead of online endpoint
D.Change --compute-type to CPU
AnswerB

The failure is a regional capacity constraint, not a configuration error. Redeploying with a VM size or region where Standard_NC6s_v3 is offered resolves the unavailability, satisfying the requirement that the endpoint's compute SKU actually exists in the target location.

Why this answer

The deployment failed because the Standard_NC6s_v3 VM instance type is not available in the East US region. The correct action is to either choose a different VM type that is available in East US or deploy to a different region that supports Standard_NC6s_v3. This directly addresses the root cause of the failure, as Azure Machine Learning online endpoints require the selected VM SKU to be available in the target region.

Exam trap

The trap here is that candidates may think increasing instance count or switching to a batch endpoint will bypass the regional SKU limitation, but neither changes the underlying VM type or region, so the deployment will still fail.

How to eliminate wrong answers

Option A is wrong because increasing the instance count does not change the VM type or region; it only scales out the number of instances, which does not resolve the unavailability of the VM SKU. Option C is wrong because switching to a batch endpoint does not address the VM availability issue; batch endpoints also require compatible compute resources and are designed for asynchronous, large-scale inference, not for fixing regional SKU unavailability. Option D is wrong because changing --compute-type to CPU would not help if the model requires GPU acceleration (as implied by the NC-series VM), and it does not solve the regional availability problem for the specified VM type.

51
MCQhard

You are building a generative AI application using Azure OpenAI Service. The application must provide factual answers based on your company's internal knowledge base. You need to minimize the risk of the model generating incorrect information (hallucinations). Which approach should you take?

A.Implement Retrieval-Augmented Generation (RAG) with Azure AI Search.
B.Fine-tune the model on your company's documents.
C.Use few-shot prompting with examples of correct answers.
D.Set the max_tokens parameter to a low value.
AnswerA

RAG grounds responses in retrieved documents from Azure AI Search, so the model cites your internal knowledge base rather than relying on parametric memory. This directly minimises hallucination risk, satisfying the requirement for factual answers drawn from company content.

Why this answer

Retrieval-Augmented Generation (RAG) with Azure AI Search grounds the model's responses in your company's internal knowledge base by retrieving relevant documents in real time and injecting them into the prompt. This reduces hallucinations by ensuring the model generates answers based on retrieved facts rather than relying solely on its parametric memory. Azure AI Search provides vector and hybrid search capabilities that efficiently index and query your documents, making RAG the most effective approach for factual accuracy.

Exam trap

The AI-102 exam often tests the misconception that fine-tuning (Option B) is the best way to ground a model in proprietary data, but the trap is that fine-tuning does not provide dynamic, query-specific retrieval and can still produce hallucinations, whereas RAG explicitly forces the model to use retrieved facts.

How to eliminate wrong answers

Option B is wrong because fine-tuning adjusts the model's weights on your documents, which can lead to overfitting and does not guarantee that the model will not hallucinate; it still relies on its internal knowledge and may generate plausible-sounding but incorrect information when queried outside the fine-tuned distribution. Option C is wrong because few-shot prompting provides examples but does not ground the model in your specific knowledge base; the model can still hallucinate if the examples do not cover the exact query context or if it extrapolates incorrectly. Option D is wrong because setting max_tokens to a low value only truncates the output length and does not improve factual accuracy; it may even cause incomplete or misleading answers without addressing the root cause of hallucinations.

52
Multi-Selectmedium

You are deploying a generative AI chat application that calls an Azure OpenAI model deployment. The application must stream partial completions to the browser as tokens are produced and must reduce perceived latency for long answers. Which two request options should you configure? (Choose two.)

Select 2 answers
A.Increase the frequency_penalty value to discourage repeated tokens
B.Set stream to true on the chat completions request
C.Set the logprobs parameter to return token-level probabilities
D.Set n to a value greater than one to generate multiple candidate completions
E.Consume the response as server-sent events and forward deltas to the client
AnswersB, E

Setting stream to true makes the service return server-sent events containing incremental deltas as the model generates tokens, so the client can render partial text immediately. This directly reduces perceived latency for long answers because the user sees content before the completion finishes. Without streaming, the client waits for the entire response body, which is the behavior the scenario wants to avoid.

Why this answer

Streaming has two cooperating parts: the request must set stream to true so the service emits incremental deltas, and the client must read those server-sent events and forward them to the browser. Together they let users see text as it is generated. Parameters that alter sampling, penalties, or probabilities change content or add metadata but never change delivery timing.

Exam trap

The trap here is treating the stream parameter alone as sufficient without consuming and forwarding the event stream on the client.

53
MCQmedium

You are testing an Azure OpenAI chat completion. The response shown in the exhibit is returned. What does the finish_reason of 'content_filter' indicate?

A.The model's response was blocked by the content filter.
B.There was a system error during processing.
C.The user's prompt was flagged by the content filter.
D.The model refused to answer due to insufficient data.
AnswerA

The content_filter finish reason means Azure OpenAI's responsible AI filtering intercepted the completion, so no usable text was returned. The prompt itself was accepted; the generated output was flagged and suppressed, which is distinct from length or stop-sequence termination.

Why this answer

The 'content_filter' finish_reason indicates that the Azure OpenAI content filtering system detected that the model's generated response violated one of the configured content policies (e.g., hate, violence, self-harm, sexual content). The response was therefore blocked before being returned to the user, and the finish_reason explicitly signals this filtering action rather than a normal completion or a stop due to token limits.

Exam trap

The trap here is that candidates often confuse 'content_filter' with prompt rejection, but the finish_reason specifically indicates the model's output was blocked, not the user's input.

How to eliminate wrong answers

Option B is wrong because a system error during processing would return a different finish_reason such as 'error' or an HTTP 500 status, not 'content_filter'. Option C is wrong because the content_filter finish_reason applies to the model's response, not the user's prompt; if the prompt were flagged, the API would typically return a 400 error with a content filter violation message before any generation occurs. Option D is wrong because the model refusing to answer due to insufficient data would be indicated by a finish_reason of 'stop' (if the model generated a refusal message) or by a specific response text, not by the 'content_filter' reason.

54
Matchingmedium

Match each Azure Cognitive Search skill to its capability.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Extract text from images

Identify entities like people or organizations

Extract key phrases from text

Detect language of text

Determine sentiment of text

Why these pairings

Common Azure Cognitive Search skills include Entity Recognition, Key Phrase Extraction, OCR (for image text extraction), and Sentiment Analysis. The correct matches are A and B. Options C and D are mismatched definitions.

55
MCQeasy

A developer is building a generative AI application with Azure OpenAI Service. The application must stream partial responses to users as tokens are generated, rather than waiting for the complete response. Which parameter should the developer set in the chat completion request?

A.Set 'temperature' to 0.
B.Set 'stream' to true.
C.Set 'max_tokens' to a large value.
D.Set 'n' to a value greater than 1.
AnswerB

Setting stream to true enables server-sent events so the service returns incremental chunks as tokens are produced. The client can render partial text immediately, which improves perceived latency for long generations. This is the documented parameter for streaming chat completions in Azure OpenAI Service.

Why this answer

Streaming is controlled by the stream parameter, which when true causes the service to emit incremental delta chunks over a persistent connection. This lets the application display tokens as they are generated, reducing perceived latency while keeping the same model and prompt configuration.

Exam trap

The trap here is confusing parameters that shape output content, such as temperature or max_tokens, with the parameter that changes how output is delivered over the network.

56
MCQhard

You are using Azure OpenAI Service to generate product descriptions. You notice that the model occasionally outputs descriptions that contain factual inaccuracies about product specifications. You want to reduce these hallucinations without changing the model. What should you do?

A.Increase the frequency_penalty parameter.
B.Decrease the temperature parameter.
C.Increase the max_tokens parameter.
D.Provide the product specifications in the prompt and use the system message to instruct the model to base answers on them.
AnswerD

Grounding the model with factual data in the prompt reduces hallucinations by providing accurate context.

Why this answer

Providing the product specifications directly in the prompt and using the system message to instruct the model to base its answers on them grounds the generation in factual data, reducing hallucinations. This technique, known as 'grounding' or 'retrieval-augmented generation' (RAG), does not modify the model itself but constrains its output to the provided context, which is the only way to reduce factual inaccuracies without changing model parameters.

Exam trap

The trap here is that candidates often confuse hyperparameter tuning (like temperature or frequency_penalty) with prompt engineering techniques, mistakenly believing that adjusting randomness or repetition penalties can fix factual hallucinations, when only providing the correct context in the prompt can do so without model changes.

How to eliminate wrong answers

Option A is wrong because increasing frequency_penalty reduces repetition of tokens by penalizing tokens that have already appeared, which does not address factual accuracy or hallucinations. Option B is wrong because decreasing temperature makes the model more deterministic and less creative, but it does not prevent the model from generating plausible-sounding but factually incorrect statements about product specifications. Option C is wrong because increasing max_tokens only allows longer responses, which can actually increase the chance of hallucinations by giving the model more opportunity to generate unsupported content.

57
MCQhard

You are deploying an Azure OpenAI model for a healthcare application. You need to ensure that the model does not generate medical advice and that all responses include a disclaimer. Which configuration should you use?

A.Ground the model with your own medical documents.
B.Set max_tokens to 50 to limit response length.
C.Use Azure AI Content Safety to filter medical terms.
D.Configure a system message with instructions and enable content filtering.
AnswerD

A system message sets the model's behavioural guardrails, instructing it to refuse medical advice and append a disclaimer to every response. Content filtering adds a safety layer blocking harmful outputs. Together they satisfy both constraints: no medical advice and mandatory disclaimers in a healthcare deployment.

Why this answer

Configuring a system message with explicit instructions (e.g., 'Do not provide medical advice; always include a disclaimer') combined with Azure AI Content Safety's content filtering allows you to enforce behavioral guardrails and block harmful outputs at the application layer. The system message sets the model's behavior, while content filtering provides a secondary safety net to catch policy violations, ensuring compliance in a regulated healthcare environment.

Exam trap

The trap here is that candidates confuse content filtering with behavioral control, assuming Azure AI Content Safety can enforce custom rules like 'do not generate medical advice' when it only filters predefined harmful categories, not domain-specific instructions.

How to eliminate wrong answers

Option A is wrong because grounding the model with medical documents (e.g., via Azure OpenAI on your data) does not prevent the model from generating medical advice; it only improves factual accuracy by referencing your data, but the model can still produce advice or omit disclaimers. Option B is wrong because setting max_tokens to 50 limits response length but does not control the content or ensure a disclaimer is included; the model could still generate medical advice within that token limit. Option C is wrong because Azure AI Content Safety filters harmful content based on predefined categories (e.g., hate, violence), but it does not have a built-in 'medical terms' filter; it cannot enforce a custom rule like 'do not generate medical advice' or 'include a disclaimer'.

58
MCQeasy

You are deploying a generative AI feature that drafts marketing copy with an Azure OpenAI GPT-4o deployment. The application must stream partial responses to the browser so users see text as it is generated, rather than waiting for the full completion. What should you enable in your API call?

A.Set the stream parameter to true and consume the server-sent event chunks returned by the chat completions endpoint.
B.Enable the asynchronous batch API and retrieve results from the output file once the job finishes.
C.Set a high max_tokens value and poll the operation status endpoint until the completion is marked complete.
D.Increase the deployment's tokens-per-minute quota so responses are generated faster.
AnswerA

The chat completions API supports streaming by setting stream to true, after which the service returns incremental server-sent event chunks containing delta content. The client renders each delta as it arrives, which produces the typewriter effect the scenario requires. This is the standard, supported mechanism for partial response delivery in Azure OpenAI.

Why this answer

Streaming in Azure OpenAI is opt-in per request: setting the stream flag makes the chat completions endpoint return incremental deltas over server-sent events, which the browser can render progressively. Polling, batch processing, and quota changes all describe non-interactive or throughput-oriented behaviors that do not deliver partial tokens as they are generated.

Exam trap

The trap here is confusing throughput tuning, such as raising quota or max tokens, with response delivery mode, which only the stream parameter changes.

59
MCQmedium

You are building a generative AI application with Azure OpenAI Service that summarizes long legal contracts. Legal reviewers report that the summaries omit obligations buried in the middle of very long documents, even though the documents fit within the model's maximum context window. You need to improve recall of those mid-document clauses while minimizing added latency. What should you do first?

A.Add a system message instructing the model to read the entire document carefully before summarizing.
B.Increase the model's temperature setting so the model explores more of the document content.
C.Split the contract into overlapping chunks, summarize each chunk, then summarize the combined chunk summaries.
D.Switch the deployment to a model with a larger maximum context window and resend the full contract.
AnswerC

Chunking with overlap, followed by a map-reduce style summarization, ensures every clause is processed in a short context where attention is strongest. Overlap prevents obligations that straddle chunk boundaries from being lost. This directly improves mid-document recall and keeps each model call small, so latency stays reasonable compared with sending the entire contract repeatedly. It is the standard mitigation for weak recall in long inputs.

Why this answer

Long inputs suffer from uneven attention, so obligations in the middle can be missed even when they fit the context window. Chunking with overlap and summarizing hierarchically places every clause in a short, high-attention context and preserves clauses that cross boundaries. Raising temperature, enlarging the window, or adding instructions does not change how attention is distributed, so those approaches leave the recall gap unresolved.

Exam trap

The trap here is assuming that fitting within the context window means the model reliably uses every part of the input.

60
MCQmedium

A company uses Azure OpenAI to generate product descriptions. They want to ensure that the descriptions are consistent in style and tone. Which strategy should they use?

A.Fine-tune the model on a dataset of product descriptions.
B.Provide a few examples of desired style in the prompt (few-shot learning).
C.Set max_tokens to a small value to limit output length.
D.Increase the temperature to 1.0 for more creativity.
AnswerB

Few-shot prompting supplies concrete exemplars of the target style and tone directly in the prompt, so the model infers the desired register and structure without retraining. This satisfies the consistency constraint more reliably than zero-shot instructions alone, since patterns are demonstrated rather than described.

Why this answer

Few-shot learning (option B) is the correct strategy because it directly controls style and tone by providing examples of desired output within the prompt. This leverages the model's in-context learning ability without modifying the underlying model weights, making it ideal for enforcing consistency without the cost and complexity of fine-tuning.

Exam trap

The trap here is that candidates often confuse fine-tuning (option A) as the only way to enforce style, overlooking that few-shot learning is a lighter, more flexible method that achieves the same goal without retraining.

How to eliminate wrong answers

Option A is wrong because fine-tuning requires a large, curated dataset and retraining the model, which is overkill for simple style consistency and introduces risks of catastrophic forgetting or overfitting to narrow patterns. Option C is wrong because setting max_tokens to a small value only truncates the output length; it does not influence the style, tone, or content of the generated text. Option D is wrong because increasing temperature to 1.0 increases randomness and creativity, which would actually reduce consistency in style and tone, not enforce it.

61
MCQmedium

You are building a solution to generate product descriptions using Azure OpenAI Service. You need to ensure that the output adheres to a specific tone (professional, friendly) and length (50-100 words). Which parameter should you adjust?

A.Configure max_tokens to limit response length.
B.Modify the top_p parameter.
C.Set the system message with instructions about tone and length.
D.Adjust the temperature parameter.
AnswerC

The system message sets persistent behavioural instructions applied before user turns, so tone and word-count guidance there governs every generated description. Prompt-level parameters such as temperature or max_tokens cannot reliably enforce a professional-friendly tone or a 50-100 word range.

Why this answer

The system message in Azure OpenAI Service is specifically designed to set the overall behavior and context for the model, including tone and length constraints. By providing instructions like 'Respond in a professional and friendly tone, and keep the output between 50 and 100 words,' the model will adhere to these guidelines throughout the conversation. This is the primary mechanism for controlling qualitative aspects of the output, as opposed to parameters that control randomness or token limits.

Exam trap

The trap here is that candidates often confuse parameters that control output randomness (temperature, top_p) or length (max_tokens) with the system message's role in defining qualitative constraints like tone and style, leading them to select A or D instead of C.

How to eliminate wrong answers

Option A is wrong because max_tokens only caps the total number of tokens (words/punctuation) in the response, but it does not enforce a specific tone or guarantee the output will be between 50-100 words—it simply cuts off at the limit, which can result in incomplete sentences. Option B is wrong because top_p (nucleus sampling) controls the diversity of word choices by limiting the cumulative probability of token selection; it does not influence tone or enforce a word count range. Option D is wrong because temperature adjusts the randomness of the model's output (higher values produce more creative/random responses, lower values produce more deterministic ones), but it cannot enforce a specific tone or a precise length range.

62
MCQmedium

You are developing a customer support chatbot using Azure OpenAI Service. The chatbot must only answer questions related to the company's product catalog and policies. You want to minimize the risk of the chatbot generating harmful or off-topic responses. Which approach should you use?

A.Set the max_tokens parameter to 100.
B.Use a system message that instructs the model to only answer product-related questions.
C.Set the temperature parameter to 0.
D.Set the top_p parameter to 0.1.
AnswerB

A system message constrains the model's behaviour across all turns, instructing it to answer only product-catalogue and policy questions, which reduces harmful or off-topic output. Content filtering alone cannot enforce topical scope, since it targets categories rather than subject relevance.

Why this answer

A system message sets the foundational behavior of the model by providing high-level instructions that guide all subsequent responses. By explicitly instructing the model to only answer product-related questions, you establish a clear boundary that minimizes off-topic or harmful outputs. This approach leverages the model's instruction-following capability, which is more effective than parameter tuning alone for content restriction.

Exam trap

The trap here is that candidates often confuse content filtering parameters (temperature, top_p, max_tokens) with instruction-based control, assuming that reducing randomness or output length can prevent off-topic responses, when in fact only explicit system-level instructions can enforce domain constraints.

How to eliminate wrong answers

Option A is wrong because setting max_tokens to 100 only limits the length of the response, not the content or topic; the model could still generate harmful or off-topic text within that token limit. Option C is wrong because setting temperature to 0 makes the model deterministic and reduces randomness, but it does not prevent the model from generating off-topic or harmful content if the prompt or context leads it there. Option D is wrong because setting top_p to 0.1 narrows the probability distribution for token selection, which reduces diversity but does not constrain the model to a specific domain or topic.

63
Multi-Selecteasy

Which TWO features of Azure AI Content Safety can help you moderate user-generated content in a social media application?

Select 2 answers
A.Self-harm content detection.
B.Hate speech severity detection.
C.PII redaction.
D.Groundedness detection.
E.Prompt injection detection.
AnswersA, B

Self-harm content detection identifies text, images and videos depicting suicide or self-injury, letting moderators filter or escalate such posts. This directly addresses the social media scenario's need to remove harmful user-generated content before it reaches vulnerable users.

Why this answer

Self-harm content detection (A) is a feature of Azure AI Content Safety that specifically identifies text or images related to self-harm, which is a critical category for moderating user-generated content in social media to prevent harm and comply with safety policies. Hate speech severity detection (B) is another core feature that classifies hate speech into severity levels (e.g., low, medium, high), enabling nuanced moderation of offensive content.

Exam trap

The trap here is that candidates may confuse Azure AI Content Safety's features with those of other Azure AI services (like Azure AI Language for PII or Azure OpenAI for prompt injection), leading them to select options that are technically valid in Azure but not part of Content Safety's core moderation capabilities.

64
MCQhard

You are building a generative AI feature in an application using the Azure OpenAI SDK. The feature must stream tokens to the user as they are generated and must also capture the full response for logging. Which implementation approach satisfies both requirements?

A.Disable streaming and set n=2 so the model returns two completions, displaying one and logging the other.
B.Make two separate non-streaming requests: one to display the response and one to log it.
C.Set stream=true and log only the first chunk received from the stream.
D.Set stream=true on the chat completions request, iterate over the streamed chunks to display deltas, and accumulate the delta content to form the complete response for logging.
AnswerD

Streaming returns incremental chat completion chunk objects containing delta content. Displaying each delta gives the user progressive output, while concatenating deltas reconstructs the full message for logging. This single request satisfies both real-time display and complete capture without a second call.

Why this answer

Streaming chat completions deliver delta content across multiple chunk objects. Rendering those deltas satisfies the progressive display requirement, and concatenating them rebuilds the exact full message for logging. A single streaming request therefore meets both needs, whereas duplicate calls, partial logging, or multiple completions fail one requirement or the other.

Exam trap

The trap here is believing that a streaming response cannot also be captured in full, when accumulating the delta chunks reconstructs the complete message.

65
Multi-Selecthard

You are deploying a generative AI model using Azure AI Foundry. The model must be accessible only from within a specific virtual network. Additionally, you need to monitor all API calls for auditing. Which two configurations are required? (Choose two.)

Select 2 answers
A.Assign a managed identity to the model deployment.
B.Configure CORS to allow only the VNet's domain.
C.Enable public network access from selected IP addresses.
D.Enable diagnostic settings to send logs to a Log Analytics workspace.
E.Disable public network access and configure a private endpoint.
AnswersD, E

Diagnostic settings stream control-plane and data-plane request logs from the Azure AI Foundry resource to a Log Analytics workspace, satisfying the auditing requirement for all API calls. This directly addresses the stem's monitoring constraint, complementing the private endpoint needed for virtual network isolation.

Why this answer

Enabling diagnostic settings to send logs to a Log Analytics workspace allows you to capture and audit all API calls made to the model deployment. This is essential for monitoring, security auditing, and compliance, as it records detailed telemetry such as request timestamps, caller IPs, and operation names. Option E is correct because disabling public network access and configuring a private endpoint ensures that the model is only accessible from within the specified virtual network, meeting the isolation requirement.

Exam trap

The trap here is that candidates often confuse network-level access controls (like IP whitelisting or CORS) with true VNet isolation via private endpoints, and they overlook that diagnostic settings are the standard Azure mechanism for auditing API calls, not managed identities or CORS.

66
MCQmedium

You are developing a solution that uses Azure Document Intelligence to extract data from invoices and then uses Azure OpenAI to summarize the extracted data. The solution occasionally produces summaries that omit key fields like the invoice total. What should you do to improve accuracy?

A.Set temperature to 0 to make the output more deterministic
B.Use a larger model like GPT-4 instead of GPT-3.5
C.Increase the max_tokens parameter
D.Define a structured prompt that explicitly requests each field and provide examples
AnswerD

Document Intelligence output can be summarised loosely by a free-form prompt, causing omissions. A structured prompt naming each required field, with few-shot examples showing the expected format, constrains the model to include the invoice total and other key values.

Why this answer

The issue is that the summarization prompt lacks explicit instructions for which fields to include. By defining a structured prompt that explicitly requests each key field (e.g., invoice total, date, vendor) and providing examples, you guide the Azure OpenAI model to consistently extract and include those fields in the summary, reducing omission errors. This approach leverages prompt engineering to improve output reliability without changing model parameters or size.

Exam trap

The trap here is that candidates often assume that model size or parameter tuning (temperature, max_tokens) is the primary fix for content omission, when in fact prompt engineering—specifically structured prompts with explicit field requests—is the correct solution for ensuring specific data is included in the output.

How to eliminate wrong answers

Option A is wrong because setting temperature to 0 makes the output more deterministic but does not force the model to include specific fields; it only reduces randomness in token selection, not the likelihood of omitting requested content. Option B is wrong because using a larger model like GPT-4 instead of GPT-3.5 improves general reasoning but does not guarantee that key fields are included unless the prompt explicitly requests them; the omission is a prompt design issue, not a model capability issue. Option C is wrong because increasing max_tokens only allows longer responses but does not influence which content the model chooses to include; the model may still omit fields even with a larger token budget.

67
Multi-Selectmedium

You are designing a generative AI solution using Azure OpenAI Service with your own data indexed in Azure AI Search. Which THREE components are essential for the retrieval-augmented generation (RAG) pattern?

Select 3 answers
A.Data ingestion pipeline to Azure AI Search
B.Azure AI Search index
C.Azure Functions for orchestration
D.Azure API Management for rate limiting
E.Azure OpenAI model
AnswersA, B, E

Data must be ingested into the search index.

Why this answer

A data ingestion pipeline is essential to load and index your data into Azure AI Search, enabling the retrieval step in RAG. Without this pipeline, the search index would have no data to query, breaking the retrieval-augmented generation pattern.

Exam trap

The trap here is that candidates often confuse optional production components (like Azure Functions for orchestration or API Management for rate limiting) with the core, mandatory components of the RAG pattern, which are the data source, search index, and the LLM model.

68
MCQmedium

You are developing a generative AI solution that uses Azure OpenAI Service. The solution must generate product descriptions in multiple languages. You need to ensure that the model consistently follows specific formatting rules, such as including a bullet list of features. Which strategy should you use?

A.Fine-tune the model with a dataset containing formatted examples.
B.Set a system message with explicit formatting instructions.
C.Increase the max_tokens parameter to allow longer outputs.
D.Adjust the temperature parameter to a lower value.
AnswerB

A system message sets persistent instructions that the model applies across every completion, so formatting rules such as the bullet list of features are followed consistently in each generated language. This satisfies the constraint of consistent formatting without repeating instructions per request.

Why this answer

System messages in Azure OpenAI Service allow you to set persistent instructions that guide the model's behavior across the entire conversation. By including explicit formatting rules—such as requiring a bullet list of features—in the system message, you enforce consistent output structure without retraining the model. This approach is efficient, cost-effective, and directly leverages the API's design for controlling response format.

Exam trap

Microsoft often tests the misconception that fine-tuning is the only way to enforce output structure, when in fact system messages provide a lightweight, zero-shot alternative for formatting control.

How to eliminate wrong answers

Option A is wrong because fine-tuning requires a large, curated dataset and significant compute resources; it is overkill for simple formatting rules and introduces risk of overfitting or losing generality, whereas a system message achieves the same goal with zero training overhead. Option C is wrong because increasing max_tokens only extends the maximum length of the response, not the structure or format; it does not enforce bullet lists or any specific formatting rules. Option D is wrong because lowering the temperature parameter reduces randomness and makes outputs more deterministic, but it does not impose explicit formatting constraints like bullet lists; it controls creativity, not structure.

69
MCQeasy

You are developing a generative AI feature with Azure OpenAI Service. The feature must generate a structured JSON object that conforms to a specific schema so it can be consumed by a downstream application without post-processing. What should you use to constrain the model output to the schema?

A.A higher frequency_penalty to discourage extra tokens.
B.Setting n to a value greater than 1 to generate multiple candidate responses.
C.Structured outputs with a JSON schema supplied in the request.
D.A system message that politely asks the model to return JSON.
AnswerC

Structured outputs let you supply a JSON schema so the model is constrained to produce valid JSON that matches it. This removes the need for downstream parsing or repair and guarantees schema conformance for the application. It is the feature designed specifically to enforce structured response formats in Azure OpenAI, matching the requirement.

Why this answer

Structured outputs in Azure OpenAI allow a JSON schema to be provided with the request so the model's response is constrained to conform to that schema. This guarantees valid, parseable JSON with the expected fields, eliminating brittle post-processing and making the output directly consumable by the downstream application.

Exam trap

The trap here is believing that a prompt instruction or a sampling parameter can guarantee JSON schema conformance, when only a formal schema constraint enforces it.

70
MCQeasy

You need to generate a poem using Azure OpenAI. The poem should be about nature and have a cheerful tone. Which parameter should you adjust to influence the tone?

A.top_p
B.temperature
C.max_tokens
D.system message
AnswerD

The system message sets the model's behavioural instructions and persona before the user prompt, so specifying a cheerful tone there reliably steers the generated poem's style. Temperature, max tokens and top_p affect randomness or length, not the requested emotional tone.

Why this answer

The system message (D) is the correct parameter to influence the tone of a generated poem because it acts as a high-level instruction that sets the behavior, persona, and style of the model. By including a directive like 'You are a cheerful poet writing about nature,' you directly control the tone without altering randomness or output length.

Exam trap

The trap here is that candidates confuse parameters that control randomness (temperature, top_p) with those that control instruction-following and style (system message), leading them to incorrectly select temperature as the primary tone influencer.

How to eliminate wrong answers

Option A is wrong because top_p controls nucleus sampling—the cumulative probability threshold for token selection—and does not directly set tone; it affects diversity of output, not style. Option B is wrong because temperature adjusts the randomness of token probabilities (higher values increase creativity, lower values make output more deterministic), but it does not specify a cheerful tone; it only influences how likely the model is to choose less probable tokens. Option C is wrong because max_tokens limits the length of the generated response and has no impact on the emotional tone or style of the poem.

71
MCQeasy

You are developing a generative AI application that uses Azure OpenAI Service. You want to ensure that the application does not generate offensive content. Which Azure service should you use?

A.Azure AI Bot Service
B.Azure AI Content Safety
C.Azure AI Language
D.Azure AI Search
AnswerB

Azure AI Content Safety provides dedicated moderation APIs that detect and filter offensive, violent, hateful, and self-harm content in prompts and completions. Integrating it with the Azure OpenAI application enforces the stem's requirement to prevent offensive output, unlike prompt engineering alone or built-in model filters.

Why this answer

Azure AI Content Safety is the correct service because it is specifically designed to detect and filter offensive, inappropriate, or harmful content in text and images. For a generative AI application using Azure OpenAI, this service can be integrated to review prompts and completions in real time, ensuring that generated outputs comply with content policies and do not contain hate speech, violence, or other offensive material.

Exam trap

The trap here is that candidates often confuse Azure AI Content Safety with Azure AI Language's moderation features, but Azure AI Language does not include a dedicated content safety API for offensive content detection, whereas Content Safety is purpose-built for this task.

How to eliminate wrong answers

Option A is wrong because Azure AI Bot Service is a platform for building conversational agents, not a content moderation or safety service; it lacks native capabilities to detect offensive content. Option C is wrong because Azure AI Language provides natural language processing features like sentiment analysis and entity recognition, but it does not include dedicated content safety filters for offensive or harmful content. Option D is wrong because Azure AI Search is a cognitive search service for indexing and querying data, not a content moderation tool; it cannot filter generated content for offensiveness.

72
MCQmedium

A company uses Azure OpenAI Service to generate product descriptions. They notice that the descriptions sometimes contain factually incorrect information. Which strategy should they use to reduce hallucinations?

A.Increase the temperature parameter to 1.0.
B.Implement Retrieval-Augmented Generation (RAG) by grounding prompts with a knowledge base.
C.Reduce the max_tokens parameter to limit output length.
D.Add a system message instructing the model to be more careful.
AnswerB

RAG grounds the model's responses in retrieved, verifiable content from a knowledge base, so generated descriptions reference actual product data rather than relying solely on parametric memory. This directly reduces fabricated facts, satisfying the requirement to cut hallucinations.

Why this answer

Retrieval-Augmented Generation (RAG) grounds the model's output in a trusted, external knowledge base, providing factual context that directly reduces hallucinations. By retrieving relevant documents and injecting them into the prompt, the model generates responses based on verified information rather than relying solely on its parametric memory, which is the primary cause of factual inaccuracies in Azure OpenAI Service.

Exam trap

The trap here is that candidates often confuse hyperparameter tuning (temperature, max_tokens) or prompt engineering (system messages) as solutions for factual accuracy, when in fact only grounding with external data (RAG) directly addresses the hallucination problem by providing a verifiable source of truth.

How to eliminate wrong answers

Option A is wrong because increasing the temperature parameter to 1.0 increases randomness and creativity in the output, which actually exacerbates hallucinations by encouraging the model to generate less predictable and potentially more fabricated content. Option C is wrong because reducing max_tokens only truncates the output length; it does not address the root cause of factual inaccuracies and may even cut off critical context or reasoning. Option D is wrong because adding a system message to 'be more careful' is a vague instruction that the model cannot reliably interpret to correct factual errors; it lacks the concrete, grounded data source that RAG provides.

73
Multi-Selectmedium

You are building a generative AI assistant on Azure OpenAI that must call internal REST APIs to look up order status. You decide to use function calling. Which TWO actions are required to make the assistant reliably invoke the correct API and return a coherent answer? (Choose two.)

Select 2 answers
A.Set the model temperature to zero so the function call arguments are always deterministic.
B.Execute the function in your application code and send the result back to the model as a tool message so it can generate the final response.
C.Define each API as a function with a name, a description, and a JSON schema for its parameters, and pass these definitions in the tools parameter of the chat completion request.
D.Enable the model's built-in web browsing tool so it can reach the internal APIs directly.
E.Increase the max_tokens value to ensure the function call arguments fit in the response.
AnswersB, C

The model does not execute functions; it only emits a request to call one with arguments. Your application must run the actual API, then append a message with role tool that includes the function result, keyed by the tool call ID. The model uses that result to produce the final natural-language answer. Skipping this round trip leaves the conversation incomplete and the user without an answer.

Why this answer

Function calling is a two-part contract. First, the application declares available functions with names, descriptions, and JSON schemas in the tools parameter so the model can choose one and emit structured arguments. Second, the application executes the chosen function and returns the result as a tool message, allowing the model to compose the final answer.

Temperature, browsing, and token limits do not enable this loop.

Exam trap

The trap here is assuming the model executes the function itself rather than only emitting a structured call request.

74
MCQhard

Your organization uses Azure OpenAI Service with a data source configured as 'Azure OpenAI on your data'. You notice that the responses include outdated information even though the underlying data source has been updated. What is the most likely cause?

A.The model is using a cached version of the prompt
B.The index in Azure Cognitive Search has not been refreshed
C.The data source is configured to sync only daily
D.The Azure CDN is caching the responses
AnswerB

Azure OpenAI on your data retrieves grounding content from the Azure Cognitive Search index, not the source directly. If that index is not refreshed after the source changes, stale documents are returned despite the underlying data being current.

Why this answer

When using Azure OpenAI Service with 'Azure OpenAI on your data', the responses are generated by querying an Azure Cognitive Search index that contains your data. If the underlying data source has been updated but the responses still include outdated information, the most likely cause is that the index in Azure Cognitive Search has not been refreshed to reflect those updates. The model itself does not store or cache the data; it relies on the index at query time, so an outdated index directly leads to outdated responses.

Exam trap

The trap here is that candidates may confuse the model's lack of awareness of data updates with caching mechanisms (like CDN or prompt caching), when the real issue is the decoupled indexing pipeline in Azure Cognitive Search that requires explicit refresh.

How to eliminate wrong answers

Option A is wrong because the model does not cache the prompt; each request is processed independently, and caching would not cause outdated information from the data source—it would only affect repeated identical prompts. Option C is wrong because while a sync schedule could cause delays, the question states the data source 'has been updated' and the responses are outdated, implying the index is not refreshed regardless of schedule; the default sync behavior is not the core issue. Option D is wrong because Azure CDN is used for static content delivery, not for caching Azure OpenAI responses, which are dynamic and not routed through CDN.

75
MCQeasy

You are developing a generative AI solution that uses Azure OpenAI Service. You need to control the creativity of the generated responses. Which parameter should you adjust?

A.top_p
B.max_tokens
C.temperature
D.frequency_penalty
AnswerC

Temperature directly scales the randomness of token sampling in Azure OpenAI Service, so lowering it narrows probability distribution toward likely tokens and raising it broadens creativity. This satisfies the stem's requirement to control response creativity, unlike max_tokens or top_p, which govern length or nucleus sampling respectively.

Why this answer

The temperature parameter directly controls the randomness of token selection in the model's output. Lower values (e.g., 0.2) make the model more deterministic and focused, while higher values (e.g., 0.8) increase creativity and variability. This is the primary parameter for adjusting creativity in Azure OpenAI Service.

Exam trap

Azure often tests the distinction between temperature (creativity/randomness) and top_p (nucleus sampling diversity), leading candidates to confuse top_p as the creativity control when it actually controls the cumulative probability cutoff for token selection.

How to eliminate wrong answers

Option A is wrong because top_p (nucleus sampling) controls the cumulative probability threshold for token selection, not the overall creativity; it can be used alongside temperature but does not directly control creativity. Option B is wrong because max_tokens limits the length of the generated response, not its creativity or randomness. Option D is wrong because frequency_penalty reduces repetition by penalizing tokens that have already appeared, which affects diversity but not the core creativity or randomness of the output.

Page 1 of 3 · 174 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Implement generative AI solutions questions.