Courseiva

CCNA Implement generative AI solutions Questions

75 of 174 questions · Page 2/3 · Implement generative AI solutions · Answers revealed

76
Multi-Selectmedium

You are preparing an Azure OpenAI deployment for a production generative AI application that must stream responses to a web front end and must limit the cost of overly long conversations. You need to configure parameters that control response length and streaming behavior. Which two parameters should you set? (Choose two.)

Select 2 answers
A.frequency_penalty
B.top_p
C.temperature
D.stream
E.max_tokens
AnswersD, E

Setting stream to true causes the service to return the completion incrementally as server-sent events rather than as one blocking payload. This is exactly what a web front end needs to display tokens as they are produced, improving perceived latency. It does not change token accounting or cost, so it pairs naturally with a length limit.

Why this answer

Two distinct needs are stated: bounding the cost of long conversations and streaming responses to a web front end. The max_tokens parameter limits how many tokens each completion may contain, providing the cost ceiling. The stream parameter switches the API to incremental server-sent delivery, satisfying the front-end streaming requirement.

Sampling parameters such as temperature, top_p, and frequency_penalty influence output content rather than length or delivery.

Exam trap

The trap here is treating sampling parameters like temperature or top_p as cost controls, when only the completion-length limit actually bounds how many tokens are billed.

77
MCQeasy

You are designing a solution that uses Azure AI Document Intelligence to extract data from invoices. The solution must classify invoices by vendor and extract line items. Which prebuilt model should you use?

A.Prebuilt invoice model
B.Custom extraction model
C.Prebuilt receipt model
D.Prebuilt layout model
AnswerA

The prebuilt invoice model is trained specifically to extract invoice fields such as vendor name, invoice number, line items, amounts and due dates from invoice documents, satisfying both the vendor classification and line-item extraction requirements without custom training.

Why this answer

The Prebuilt invoice model (Option A) is specifically designed to extract common fields from invoices, including vendor details and line items, without requiring custom training. This model is optimized for invoice documents and provides out-of-the-box extraction of structured data such as vendor name, invoice date, and line-item descriptions, quantities, and amounts.

Exam trap

The trap here is that candidates often confuse the Prebuilt layout model with the Prebuilt invoice model, assuming layout extraction is sufficient for invoice data, but layout only provides raw text and table positions without the semantic understanding needed for vendor classification and line-item extraction.

How to eliminate wrong answers

Option B is wrong because a custom extraction model requires labeled training data and is used when prebuilt models do not meet specific document needs, but here the prebuilt invoice model already covers vendor classification and line-item extraction. Option C is wrong because the Prebuilt receipt model is designed for receipts (e.g., from stores or restaurants), not invoices, and does not extract vendor classification or line-item details in the same structured format. Option D is wrong because the Prebuilt layout model extracts text, tables, and selection marks but does not perform semantic classification or vendor-specific field extraction; it lacks the prebuilt understanding of invoice-specific fields.

78
MCQhard

Refer to the exhibit. You are configuring an Azure OpenAI Service deployment for document summarization. The current parameters produce summaries that are often too verbose. You need to make the summaries more concise while maintaining factual accuracy. Which parameter change should you make?

A.Increase top_p to 1.0
B.Increase frequency_penalty to 0.5
C.Decrease max_tokens to 100
D.Increase temperature to 0.7
AnswerC

Lowering max_tokens caps the response length, directly curbing verbosity by truncating generation before padding accumulates. However, it constrains output size rather than steering style, so factual accuracy is preserved only if essential content fits within 100 tokens; summarisation quality may suffer if key facts are cut off.

Why this answer

Decreasing max_tokens to 100 directly limits the maximum length of the generated summary, forcing the model to produce shorter output. This addresses the verbosity issue without altering the model's factual accuracy, as max_tokens controls output length, not content selection or creativity.

Exam trap

Microsoft often tests the distinction between parameters that control output length (max_tokens) versus those that control creativity or diversity (temperature, top_p, frequency_penalty), leading candidates to mistakenly adjust the latter when the issue is simply excessive length.

How to eliminate wrong answers

Option A is wrong because increasing top_p to 1.0 makes the model consider a wider set of possible tokens, which can increase diversity and potentially lead to even more verbose or less focused summaries. Option B is wrong because increasing frequency_penalty to 0.5 penalizes tokens that have already appeared, reducing repetition but not directly controlling summary length; it may even cause the model to use more unique words, increasing verbosity. Option D is wrong because increasing temperature to 0.7 increases randomness in token selection, which can produce more creative but less consistent summaries, potentially harming factual accuracy and not reliably reducing length.

79
MCQmedium

You are deploying a chatbot using Azure OpenAI Service with a custom dataset indexed in Azure AI Search. Users report that the chatbot frequently responds with 'I don't know' for questions that the dataset should cover. What is the most likely cause?

A.The search scope is limited to a small number of documents.
B.The confidence threshold in the retrieval configuration is set too low, filtering out relevant chunks.
C.The temperature setting in the model deployment is set too high.
D.The chunk size in the index is too large, causing irrelevant chunks to be retrieved.
AnswerB

Low confidence threshold causes relevant chunks to be excluded, leading to 'I don't know' responses.

Why this answer

When the confidence threshold is set too low, Azure AI Search filters out retrieved chunks that don't meet the minimum confidence score, even if those chunks are relevant. This causes the chatbot to respond with 'I don't know' because no sufficiently confident context is passed to the Azure OpenAI model for answer generation. The issue is specifically in the retrieval configuration, not in the model's generation parameters.

Exam trap

The trap here is that candidates often confuse the confidence threshold in retrieval with the temperature parameter in generation, assuming a high temperature causes the model to refuse answers, when in fact temperature controls creativity, not retrieval filtering.

How to eliminate wrong answers

Option A is wrong because limiting the search scope to a small number of documents would reduce the pool of potential matches, but the chatbot would still return answers from those documents if they contain relevant information; the symptom of 'I don't know' for covered questions points to a filtering issue, not a scope limitation. Option C is wrong because a high temperature setting affects the randomness and creativity of the generated response, not the model's ability to retrieve or use context; it might cause verbose or off-topic answers, but not a refusal to answer. Option D is wrong because large chunk sizes can cause retrieval of irrelevant chunks due to lower precision, but this would lead to incorrect or hallucinated answers, not the model saying 'I don't know'; the model would still attempt to answer using the retrieved context.

80
MCQmedium

You are building a generative AI solution with Azure OpenAI Service. Prompts must be assembled from a system instruction, retrieved document chunks, and the user's question. You need a mechanism that automatically inserts the retrieved chunks into a designated placeholder in a prompt template before the request is sent to the model. What should you use?

A.A content filter configured on the Azure OpenAI deployment
B.A system-assigned managed identity on the Azure OpenAI resource
C.The max_tokens parameter on the chat completions request
D.Prompt flow with a Jinja prompt template node
AnswerD

Jinja templating in prompt flow renders placeholders such as {{context}} and {{question}} by substituting runtime inputs, so retrieved document chunks are injected into the template before the chat call executes. This provides deterministic prompt assembly, supports conditional logic, and keeps the orchestration inside a deployable flow, which matches the requirement to fill a designated placeholder automatically.

Why this answer

Prompt flow's Jinja prompt template node is the supported way to author a reusable template whose placeholders are filled from flow inputs at runtime, which is exactly what injecting retrieved chunks requires. The other choices address safety filtering, authentication, or output length, none of which assemble prompt text. Using a template also keeps prompt logic versioned and testable within the flow.

Exam trap

The trap here is assuming that any Azure OpenAI feature which touches prompts, such as content filtering, performs prompt assembly.

81
MCQeasy

A developer wants to integrate a pre-built AI model that can extract key information from invoices, such as vendor name, invoice date, and total amount. Which Azure AI service should they use?

A.Azure AI Language
B.Azure OpenAI Service
C.Azure AI Document Intelligence
D.Azure Cognitive Search
AnswerC

Azure AI Document Intelligence provides pre-built invoice models that extract structured fields such as vendor name, invoice date and total amount from documents, requiring no custom training. This directly matches the requirement for a pre-built extraction model.

Why this answer

Azure AI Document Intelligence (formerly Form Recognizer) is the correct service because it is specifically designed to extract structured data from documents like invoices, including fields such as vendor name, invoice date, and total amount. It uses pre-built models trained on common document types, making it ideal for this use case without requiring custom training.

Exam trap

The trap here is that candidates may confuse Azure AI Language's entity extraction capabilities with document-specific extraction, not realizing that Azure AI Document Intelligence is purpose-built for semi-structured documents like invoices and forms.

How to eliminate wrong answers

Option A is wrong because Azure AI Language focuses on text analytics, sentiment analysis, and entity recognition from unstructured text, not on extracting structured fields from semi-structured documents like invoices. Option B is wrong because Azure OpenAI Service provides generative AI models (e.g., GPT-4) for text generation and conversation, not a pre-built model optimized for invoice data extraction. Option D is wrong because Azure Cognitive Search is a search and indexing service for building search experiences over data, not a document processing service for extracting key information from invoices.

82
MCQhard

Refer to the exhibit. You are deploying a generative AI model as an online endpoint in Azure Machine Learning. You receive complaints that the endpoint returns 503 errors during peak hours. What is the most likely cause?

A.The manual scale setting with 2 instances may be insufficient for peak traffic.
B.The request timeout of 30 seconds is too short.
C.The environment variable MODEL_CACHE_SIZE is set too low.
D.The model version is not specified correctly.
AnswerA

A manual scale setting pinned at two instances cannot absorb peak-hour request volume, so the endpoint's instances saturate and Azure Machine Learning returns 503 errors. Configuring autoscale, or raising the instance count, matches capacity to the traffic constraint described in the stem.

Why this answer

503 errors during peak hours indicate that the endpoint is overwhelmed by the request volume. With manual scaling set to only 2 instances, the compute capacity is insufficient to handle the increased traffic, causing the service to reject requests. Azure Machine Learning online endpoints require sufficient instance count or autoscaling to absorb traffic spikes.

Exam trap

Microsoft often tests the distinction between HTTP status codes (503 vs. 408/504) to mislead candidates into confusing timeout-related errors with capacity-related errors.

How to eliminate wrong answers

Option B is wrong because a 30-second request timeout would cause timeout errors (e.g., 408 or 504), not 503 Service Unavailable errors; 503 specifically signals resource exhaustion, not slow processing. Option C is wrong because the MODEL_CACHE_SIZE environment variable affects model loading performance and caching, not the endpoint's ability to handle concurrent requests; a low cache size might cause slower inference but not 503 errors. Option D is wrong because an incorrect model version would result in deployment failures (e.g., 404 or 500 errors) during endpoint creation or update, not intermittent 503 errors during peak traffic.

83
MCQhard

You are deploying a generative AI assistant with Azure OpenAI Service. Company policy requires that the assistant must never output profanity, and that any attempt to elicit such content must be blocked at the service level rather than filtered in application code. You need the least administrative effort to enforce this. What should you do?

A.Add a system message instructing the model to never use profanity.
B.Enable Azure API Management policies to inspect and rewrite responses containing profanity.
C.Set the model temperature to 0 and top_p to 1 to reduce the chance of profanity.
D.Create a custom content filter with a blocklist for profanity terms and associate it with the deployment.
AnswerD

Azure OpenAI content filters support custom blocklists of terms that are blocked at the service level when associated with a deployment. Creating a blocklist for profanity and attaching it to the deployment enforces the policy without application-side filtering. This is the least-effort service-level control that directly targets the prohibited terms and blocks elicitation attempts.

Why this answer

Custom content filters in Azure OpenAI let you define blocklists of specific terms and apply them to a deployment. Terms in the blocklist are blocked at the service level, which aligns with the policy that prohibited content must be stopped by the service rather than by application code. Associating the filter with the deployment enforces the rule for all calls.

Exam trap

The trap here is treating a system message or sampling parameter as a content-safety control, when only service-level filtering such as a custom blocklist reliably blocks prohibited terms.

84
MCQeasy

You are developing a chat application that uses Azure OpenAI Service to answer customer queries. The solution must ensure that the model does not generate harmful or offensive content. Which Azure AI service should you configure?

A.Azure AI Bot Service
B.Azure AI Search
C.Azure AI Content Safety
D.Azure AI Language
AnswerC

Azure AI Content Safety provides dedicated harm classification across categories such as hate, violence and self-harm, filtering both prompts and completions. Configuring it on the Azure OpenAI endpoint enforces the requirement that the model never generates harmful or offensive content.

Why this answer

Azure AI Content Safety is specifically designed to detect and filter harmful or offensive content in text and images, making it the appropriate service to integrate with an Azure OpenAI chat application to enforce content safety policies. It provides APIs for content moderation, including severity-based filtering for hate, self-harm, sexual, and violence categories, which directly addresses the requirement to prevent the model from generating harmful output.

Exam trap

The trap here is that candidates often confuse Azure AI Language's text analytics capabilities (like sentiment analysis) with content safety, assuming that language understanding inherently includes harm detection, but Azure AI Content Safety is a separate, specialized service for content moderation.

How to eliminate wrong answers

Option A is wrong because Azure AI Bot Service is a platform for building, testing, and deploying conversational agents (bots), not a content moderation or safety service; it does not natively filter harmful content from model outputs. Option B is wrong because Azure AI Search is a search-as-a-service solution for indexing and querying data, with no built-in content safety or moderation capabilities for generated text. Option D is wrong because Azure AI Language provides natural language processing features like sentiment analysis, key phrase extraction, and question answering, but it does not include content safety filters for detecting harmful or offensive content.

85
MCQhard

You are building a generative AI solution that uses Azure OpenAI function calling to let the model invoke backend APIs. During testing, the model sometimes invents parameter values that cause API errors. You need to make function invocation more reliable. What should you do?

A.Define clear function schemas with typed parameters, required fields, and descriptions, and validate arguments before executing the API call.
B.Set tool_choice to auto and increase the temperature to encourage more creative argument generation.
C.Disable parallel tool calls and force the model to call only one function per turn.
D.Include the full OpenAPI specification of every backend API in the system message.
AnswerA

Function calling relies on the model producing arguments that match the declared JSON schema. Precise types, required fields, and descriptions guide the model, while server-side validation rejects invalid arguments before they reach the API. Together they reduce invented values and prevent downstream errors.

Why this answer

Reliable function calling depends on well-defined schemas and validation. Typed parameters, required fields, and descriptive names help the model map user intent to correct arguments, and validating before execution prevents invalid values from reaching the API. Higher temperature, oversized specifications, and parallel-call limits do not enforce argument correctness.

Exam trap

The trap here is believing that adding more API documentation or creative sampling improves argument accuracy.

86
MCQeasy

A development team is building an Azure OpenAI chat application and wants the model to return structured JSON that matches a defined schema containing an 'intent' field and a 'confidence' field. The application will parse the response directly. Which deployment parameter should the team use to constrain the output format?

A.Set the 'temperature' parameter to 0.
B.Set the 'top_p' parameter to 1.
C.Configure a structured output format, such as a JSON schema response format, on the request.
D.Increase the 'max_tokens' value to allow the full JSON document to be generated.
AnswerC

Structured outputs, exposed through the response format parameter with a JSON schema, constrain generation so the completion conforms to the supplied schema, including the required intent and confidence fields. This is the purpose-built mechanism for machine-parseable responses and removes the need for fragile prompt-only formatting instructions or post-processing repairs in the application.

Why this answer

Constraining a model to emit machine-parseable JSON with specific fields is accomplished with structured outputs, where a JSON schema is supplied in the request's response format. This guarantees schema conformance far more reliably than sampling parameters or length limits, which only influence variability or truncation.

Exam trap

The trap here is confusing determinism with structure: lowering temperature or raising max_tokens does not make a model emit schema-valid JSON.

87
MCQmedium

Your company uses Microsoft Copilot for Microsoft 365. You need to ensure that Copilot only accesses data from approved SharePoint sites and does not use any other organizational data. What should you configure?

A.Configure Conditional Access policies in Microsoft Entra ID.
B.Use Microsoft Purview Data Map to catalog the approved sites.
C.Apply sensitivity labels to the approved SharePoint sites.
D.Set up data retention policies in Microsoft Purview.
AnswerC

Sensitivity labels can be used to restrict Copilot's data sources.

Why this answer

Sensitivity labels can be configured to restrict Microsoft Copilot for Microsoft 365 to only access content from approved SharePoint sites. By applying a sensitivity label that includes the 'mark content' or 'encrypt content' setting with a specific scope, you can use the 'Microsoft Copilot for Microsoft 365' condition under 'Access control' to block Copilot from processing data from unlabeled or non-approved sites. This ensures Copilot respects the label's policy and excludes all other organizational data.

Exam trap

The trap here is that candidates often confuse data access control with authentication or data lifecycle management, leading them to select Conditional Access or retention policies instead of recognizing that sensitivity labels provide the specific Copilot data source restriction capability.

How to eliminate wrong answers

Option A is wrong because Conditional Access policies in Microsoft Entra ID control user authentication and access to applications, not the data sources that Copilot can index or retrieve content from. Option B is wrong because Microsoft Purview Data Map is a metadata catalog for data governance and discovery, not a mechanism to enforce access restrictions on Copilot's data sources. Option D is wrong because data retention policies in Microsoft Purview manage how long data is kept and when it is deleted, not which data Copilot is allowed to access.

88
MCQhard

You have configured a system message for an Azure OpenAI chat completion deployment as shown in the exhibit. Users are reporting that the assistant sometimes refuses to answer questions that are clearly within the scope of the provided data. What is the most likely issue?

A.The system message encourages the model to err on the side of caution, leading to false-negative refusals.
B.The system message explicitly prohibits making up information, which is correct behavior.
C.The system message does not include instructions to use the provided data.
D.The temperature parameter is set too high, causing the model to hallucinate.
AnswerA

Overly defensive system message wording instructs the model to decline when uncertain, so legitimate in-scope questions fall below its confidence threshold and trigger refusals. Softening that instruction, while keeping grounding constraints, restores answers without permitting out-of-scope responses.

Why this answer

The system message likely contains overly cautious language (e.g., 'only answer if you are certain' or 'do not speculate'), which causes the model to refuse answering even when the data clearly supports the response. This is a known behavior in Azure OpenAI chat completions where the system message's tone and constraints directly influence refusal rates, leading to false-negative refusals.

Exam trap

Microsoft often tests the misconception that refusal issues are caused by missing data instructions or high temperature, when in fact the root cause is the system message's overly cautious phrasing that induces false-negative refusals.

How to eliminate wrong answers

Option B is wrong because explicitly prohibiting making up information is a standard best practice to reduce hallucination, not a cause of false-negative refusals; it does not inherently make the model overly cautious. Option C is wrong because the system message in the exhibit (as described) does include instructions to use the provided data, so the issue is not a missing directive but the cautious phrasing. Option D is wrong because a high temperature parameter increases randomness and creativity, leading to hallucinations or off-topic responses, not systematic refusal to answer within scope.

89
Multi-Selectmedium

Which TWO actions can you take to reduce the cost of using Azure OpenAI Service for a chat application?

Select 2 answers
A.Increase the frequency penalty.
B.Enable content filtering.
C.Set the max_tokens parameter to a lower value.
D.Increase the max_tokens parameter to allow longer responses.
E.Use a smaller model like GPT-3.5 instead of GPT-4.
AnswersC, E

Reduces token count per response.

Why this answer

Reducing the max_tokens parameter directly limits the number of tokens generated per API call, which lowers the cost since Azure OpenAI Service bills per token (both input and output). By capping the response length, you avoid paying for unnecessarily long completions.

Exam trap

The trap here is that candidates often confuse cost-saving techniques with performance-tuning parameters, mistakenly thinking that adjusting penalty settings or content filtering reduces token consumption, when in fact only limiting token output or using a cheaper model directly lowers the bill.

90
MCQhard

You are building a generative AI feature that drafts marketing copy in English and must also produce accurate French and German versions. The team wants a single deployment that handles all three languages without maintaining separate prompt templates per language, and the translations must preserve the marketing tone. Which Azure OpenAI capability should the solution rely on?

A.Use a single model deployment and instruct it in the system message to generate in the requested language while preserving tone.
B.Translate the English output using Azure AI Translator and then post-process with a grammar model.
C.Deploy a separate model instance per language and route requests based on the target locale.
D.Fine-tune the model on parallel marketing copy in all three languages before deploying it.
AnswerA

Azure OpenAI chat models are multilingual and can follow instructions that specify the target language and desired style. A single deployment with a system message describing the marketing tone and the requested output language avoids separate templates per language. This directly satisfies the goal of one deployment handling English, French, and German with consistent tone.

Why this answer

Azure OpenAI chat models are trained on multilingual data and respond well to explicit language and style instructions. Specifying the target language and the marketing tone in the system message lets one deployment serve English, French, and German without maintaining separate templates. Separate deployments, an external translation pipeline, or fine-tuning all add complexity that the scenario does not require and do not better guarantee tone preservation.

Exam trap

The trap here is assuming multilingual output requires separate deployments or a dedicated translation service, when a single deployment with clear language and tone instructions is sufficient.

91
Multi-Selectmedium

You are building a generative AI feature where an Azure OpenAI model must call internal functions, such as looking up an order status and issuing a refund, based on the user's natural language request. You need the model to decide which function to invoke and to incorporate the function result into its answer. Which two actions should you take? (Choose two.)

Select 2 answers
A.Execute the returned function call in your application and send the result back in a subsequent message with the tool role.
B.Attach the internal API documentation as an 'On Your Data' source so the model can retrieve and call the endpoints.
C.Fine-tune the model on historical order and refund conversations so it memorizes the internal APIs.
D.Raise the temperature so the model explores multiple possible functions before answering.
E.Define each function with a name, description, and JSON parameter schema and pass them in the tools parameter of the chat request.
AnswersA, E

The model only proposes a function call; it does not execute internal code. The application must run the requested lookup or refund and then return the output to the model in a tool-role message so the model can compose a grounded final answer. This round trip is what connects the model's decision to real business data and completes the interaction.

Why this answer

Function calling requires two coordinated steps: declaring the callable functions with their parameter schemas in the request, and then executing the model's requested call and returning the result in a tool-role message. Together these let the model decide on the correct internal operation and produce a final answer grounded in real system data.

Exam trap

The trap here is believing fine-tuning or document retrieval can invoke internal APIs, when only declared tools plus application-side execution return usable results.

92
Multi-Selecteasy

You are using Azure OpenAI Service to generate code snippets. The output must be safe and free of security vulnerabilities. Which TWO practices should you follow? (Select TWO.)

Select 2 answers
A.Increase the temperature parameter to encourage diversity
B.Rely on the model's built-in safety features
C.Use Azure AI Content Safety to filter outputs
D.Fine-tune the model with a dataset of secure code examples
E.Include a system message instructing the model to follow secure coding practices
AnswersC, E

Content safety can detect and block harmful code.

Why this answer

Azure AI Content Safety is a dedicated service that provides an additional layer of filtering for harmful or inappropriate content, including security vulnerabilities, beyond what the model itself offers. It allows you to define custom severity thresholds and blocklists, ensuring that generated code snippets are safe before they reach the user. This is a recommended practice for production deployments to mitigate risks like injection attacks or exposure of sensitive patterns.

Exam trap

The trap here is that candidates often assume the model's built-in safety features are sufficient (Option B) or that increasing temperature (Option A) improves safety by adding randomness, when in fact Azure AI Content Safety is the explicit, exam-tested tool for output filtering in generative AI solutions.

93
Multi-Selecthard

Which THREE components are required to build a custom chat application using Azure OpenAI Service that can answer questions based on your own private data?

Select 3 answers
A.The 'Add your data' feature configured in Azure OpenAI Studio.
B.Azure AI Search index.
C.An Azure OpenAI Service deployment.
D.A fine-tuned custom model.
E.An Azure OpenAI embeddings model deployment.
AnswersA, B, C

Enables grounding on private data.

Why this answer

The 'Add your data' feature in Azure OpenAI Studio provides a no-code interface to connect your private data sources (e.g., Azure Blob Storage, local files) to an Azure OpenAI chat model. It automatically chunks the data, creates an Azure AI Search index, and configures the retrieval-augmented generation (RAG) pipeline, enabling the model to answer questions grounded in your proprietary content without fine-tuning.

Exam trap

The trap here is that candidates often confuse fine-tuning (Option D) with retrieval-augmented generation, assuming that custom data requires model retraining, when in fact the 'Add your data' feature uses a RAG approach that does not modify the base model.

94
MCQmedium

Your company wants to build a custom generative AI model that generates architectural designs. The model should be trained on the company's proprietary dataset of floor plans and designs. Which Azure service should you use?

A.Azure OpenAI Service
B.Azure Machine Learning
C.Azure AI Vision
D.Azure AI Document Intelligence
AnswerB

Azure Machine Learning provides the training pipelines, compute and model registration needed to fine-tune a generative model on proprietary floor-plan data. It satisfies the custom-training requirement that prebuilt Azure OpenAI models cannot meet with private datasets.

Why this answer

Azure Machine Learning is the correct service because it provides a comprehensive platform for training custom generative AI models using your own proprietary datasets. It supports deep learning frameworks like PyTorch and TensorFlow, enabling you to build, train, and deploy a custom generative model for architectural designs, whereas Azure OpenAI Service is limited to pre-trained models and fine-tuning, not custom training from scratch.

Exam trap

The trap here is that candidates often confuse 'fine-tuning' (offered by Azure OpenAI Service) with 'custom training from scratch' (offered by Azure Machine Learning), leading them to incorrectly select Azure OpenAI Service when the question explicitly requires training a model on proprietary data with custom architecture.

How to eliminate wrong answers

Option A is wrong because Azure OpenAI Service only allows fine-tuning of existing pre-trained models (like GPT-4) on your data, not training a custom generative model from scratch on proprietary floor plans; it lacks the flexibility to define custom architectures. Option C is wrong because Azure AI Vision is designed for image analysis tasks (e.g., object detection, OCR) and does not support generative model training or generation of new designs. Option D is wrong because Azure AI Document Intelligence is specialized for extracting structured data from documents (e.g., forms, invoices) and cannot be used to train or generate architectural designs.

95
MCQhard

You have deployed a generative AI model using Azure Machine Learning. The model is used for generating financial reports. You need to monitor the model's performance and detect data drift in the input data. What should you use?

A.Azure Machine Learning data drift monitoring
B.Azure Monitor
C.Application Insights
D.Azure AI Language
AnswerA

Azure Machine Learning data drift monitoring directly satisfies the requirement to detect drift in the model's input data. It computes statistical divergence between a baseline dataset and recent inference data, covering numerical and categorical features, and raises alerts when distributions shift. Model performance metrics alone would not surface input-distribution changes for the financial reports.

Why this answer

Azure Machine Learning data drift monitoring is the correct choice because it is specifically designed to detect statistical changes in input data over time, which is critical for generative AI models used in financial reporting where data distributions can shift due to market changes or new regulations. It compares the current input data distribution against a baseline dataset using metrics like Wasserstein distance or Population Stability Index, and triggers alerts when drift exceeds a threshold. This ensures the model's outputs remain reliable and compliant, which is a core requirement for generative AI solutions in regulated industries.

Exam trap

The trap here is that candidates confuse general monitoring tools (Azure Monitor, Application Insights) with the specialized data drift detection capability in Azure Machine Learning, assuming any monitoring tool can detect statistical drift in input data.

How to eliminate wrong answers

Option B is wrong because Azure Monitor is a general-purpose observability service for infrastructure and application metrics (e.g., CPU, memory, request rates), not for detecting statistical data drift in model inputs. Option C is wrong because Application Insights focuses on application performance monitoring (APM) and telemetry (e.g., request failures, dependencies), not on comparing data distributions or detecting drift in input features. Option D is wrong because Azure AI Language is a service for natural language processing tasks (e.g., sentiment analysis, key phrase extraction), not for monitoring model performance or detecting data drift in input data.

96
MCQeasy

Your company uses Azure AI Content Safety to moderate user-generated content in a chat application. You need to detect and block sexual content in multiple languages. Which pre-built category should you configure?

A.Sexual
B.Self-harm
C.Hate
D.Violence
AnswerA

The Sexual category directly targets sexual content, satisfying the multilingual detection requirement because Azure AI Content Safety's pre-built categories are language-agnostic, applying trained classifiers across supported languages without per-language configuration. Configuring it blocks such material at the severity threshold you set, meeting the stem's need to detect and block sexual content across multiple languages.

Why this answer

Azure AI Content Safety provides pre-built severity-based categories for content moderation. The 'Sexual' category is specifically designed to detect and block explicit sexual content, including text and images, across multiple languages. This makes it the correct choice for your requirement to moderate sexual content in a chat application.

Exam trap

The trap here is that candidates may confuse 'Sexual' with broader categories like 'Hate' or 'Violence', not realizing that Azure AI Content Safety has a dedicated pre-built category for sexual content with specific detection capabilities.

How to eliminate wrong answers

Option B (Self-harm) is wrong because it focuses on content related to self-injury or suicide, not sexual material. Option C (Hate) is wrong because it targets hate speech based on protected attributes like race or religion, not sexual content. Option D (Violence) is wrong because it detects violent acts or threats, which are distinct from sexual content.

97
MCQeasy

You need to generate realistic synthetic data using Azure OpenAI Service to train a machine learning model. The data must be diverse and cover edge cases. Which approach should you use?

A.Use prompt engineering with detailed instructions to generate varied examples.
B.Fine-tune the model on a small dataset of real examples.
C.Use Azure OpenAI embeddings to generate similar data points.
D.Set a high temperature parameter only.
AnswerA

Prompt engineering lets you specify tone, format, demographic spread and explicit edge-case scenarios within the instruction, so the model generates varied synthetic examples covering the required diversity. Fine-tuning needs existing data, and templates alone cannot produce the breadth the training set demands.

Why this answer

Prompt engineering with detailed instructions allows you to explicitly control the diversity, structure, and edge-case coverage of generated synthetic data without requiring a pre-existing dataset. By crafting system messages and user prompts that specify variations in attributes, formats, and boundary conditions, you can produce a wide range of realistic examples that mimic real-world distributions, which is essential for training robust machine learning models.

Exam trap

The trap here is that candidates often assume fine-tuning or embeddings are the only ways to generate realistic data, overlooking that prompt engineering with detailed instructions is the most direct and flexible method for producing diverse synthetic data without requiring a pre-existing labeled dataset.

How to eliminate wrong answers

Option B is wrong because fine-tuning on a small dataset of real examples would bias the model toward the limited patterns in that dataset, reducing diversity and failing to generate edge cases, which contradicts the requirement for varied synthetic data. Option C is wrong because Azure OpenAI embeddings are used to measure semantic similarity between text inputs, not to generate new data points; they can retrieve or cluster existing data but cannot produce novel synthetic examples. Option D is wrong because setting a high temperature parameter alone increases randomness in token selection but does not provide the structured control needed to ensure coverage of specific edge cases or diverse scenarios; it may produce incoherent or irrelevant outputs without detailed prompt guidance.

98
Multi-Selecthard

You are deploying a solution that uses Azure OpenAI Service to generate financial reports. You need to ensure the outputs are accurate and consistent. Which TWO parameters should you adjust? (Choose two.)

Select 2 answers
A.Set presence_penalty to 0.5.
B.Set temperature to 0.
C.Set max_tokens to 2000.
D.Set frequency_penalty to 0.7.
E.Set top_p to 0.1.
AnswersB, E

Low temperature makes output more deterministic.

Why this answer

Setting temperature to 0 makes the model deterministic, always choosing the most likely next token. This is critical for financial reports where consistency and reproducibility are required, as it eliminates randomness in the output.

Exam trap

The trap here is that candidates often confuse parameters that control creativity (temperature, top_p) with those that control repetition (frequency_penalty, presence_penalty) or output length (max_tokens), leading them to select options that do not directly address accuracy and consistency.

99
MCQhard

You are a generative AI engineer at a financial services company. The company uses Azure OpenAI Service to generate investment summaries. You have deployed a GPT-4 model with a content filter set to 'Low' for hate speech. The model frequently generates summaries that include biased language against certain demographics. You need to reduce biased outputs while maintaining the ability to generate detailed financial analysis. You cannot afford to retrain the model. You have the following options: A) Change the content filter severity to 'High' for all categories, B) Add a system message instructing the model to avoid bias and provide examples of unbiased summaries in the prompt, C) Use the Azure AI Language service to detect bias in the output and regenerate if bias is found, D) Deploy a different model like GPT-3.5 which has less bias. Which course of action should you take?

A.Use the Azure AI Language service to detect bias in the output and regenerate if bias is found.
B.Add a system message instructing the model to avoid bias and provide examples of unbiased summaries in the prompt.
C.Change the content filter severity to 'High' for all categories.
D.Deploy a different model like GPT-3.5 which has less bias.
AnswerB

Adding a system message with explicit anti-bias instructions and unbiased examples steers the model's outputs through prompt engineering, requiring no retraining. Raising the filter to High may block legitimate financial content, while detection and regeneration adds latency and cost.

Why this answer

Adding a system message that explicitly instructs the model to avoid bias and providing few-shot examples of unbiased summaries is the most direct and cost-effective way to steer the model's behavior without retraining. This approach leverages prompt engineering to influence the model's output distribution while preserving the ability to generate detailed financial analysis.

Exam trap

The trap is assuming that content filters can address bias; candidates often pick 'High' severity thinking it will filter biased language, but content filters target harmful content categories, not bias, and may over-block legitimate content.

How to eliminate wrong answers

Option A is wrong because using Azure AI Language to detect bias and regenerate adds latency, cost, and complexity, and it does not prevent biased outputs from being generated in the first place. Option C is wrong because setting content filters to 'High' for all categories may block legitimate financial content and does not specifically target bias; content filters are designed for harmful content categories, not nuanced bias. Option D is wrong because deploying a different model like GPT-3.5 does not guarantee less bias and may degrade the quality of detailed financial analysis.

100
MCQeasy

You are developing a custom chatbot using Azure AI Bot Service and Language Understanding (CLU). The chatbot needs to escalate to a human agent when the user's sentiment is negative. Which component should you use to detect sentiment?

A.Azure AI Language sentiment analysis
B.Azure Cognitive Search
C.Orchestration workflow
D.QnA Maker
AnswerA

Azure AI Language sentiment analysis returns per-utterance sentiment scores and confidence values, which the bot can evaluate to trigger escalation. CLU handles intent and entity extraction only, not sentiment, so it cannot satisfy the negative-sentiment escalation condition. This component directly meets the requirement to detect sentiment within the conversation flow.

Why this answer

Azure AI Language sentiment analysis is the correct component because it provides pre-built sentiment detection capabilities that analyze text and return sentiment labels (positive, negative, neutral) and confidence scores. This directly meets the requirement to detect negative user sentiment in chatbot conversations, enabling escalation to a human agent when needed.

Exam trap

The trap here is that candidates may confuse Azure Cognitive Search (a search service) with AI Language services, or assume that QnA Maker includes sentiment analysis, when in fact only Azure AI Language provides dedicated sentiment detection.

How to eliminate wrong answers

Option B (Azure Cognitive Search) is wrong because it is designed for indexing and searching documents, not for analyzing sentiment in real-time chat messages. Option C (Orchestration workflow) is wrong because it manages routing between multiple language models or skills, but does not perform sentiment analysis itself. Option D (QnA Maker) is wrong because it is a service for creating question-and-answer knowledge bases from FAQ-like content, and lacks native sentiment detection capabilities.

101
MCQeasy

You are implementing a generative AI solution using Azure OpenAI. You need to ensure that the model's outputs do not contain certain inappropriate words or phrases. Which feature should you configure?

A.System message instructions
B.Grounding with your data
C.Content filters
D.Max tokens limit
AnswerC

Content filters in Azure OpenAI evaluate both prompts and completions against configurable severity thresholds across categories such as hate, violence, and sexual content, blocking inappropriate words or phrases. This directly enforces the required output restriction without retraining or prompt engineering.

Why this answer

Content filters in Azure OpenAI are specifically designed to detect and block inappropriate words or phrases in both prompts and completions. They operate at the service level, applying configurable severity thresholds for categories like hate, violence, sexual content, and self-harm, ensuring model outputs adhere to policy without requiring prompt engineering or data modifications.

Exam trap

Microsoft often tests the misconception that system messages (Option A) are sufficient for content safety, when in fact they are only behavioral guidelines and lack the enforcement mechanism of dedicated content filters.

How to eliminate wrong answers

Option A is wrong because system message instructions guide model behavior and tone but cannot reliably enforce content restrictions; they are advisory and can be overridden by the model, especially in edge cases. Option B is wrong because grounding with your data (using Azure Cognitive Search) augments prompts with your own data for relevance and accuracy, but it does not filter or block inappropriate content from the model's generated responses. Option D is wrong because the max tokens limit controls the length of the output, not its content; it cannot prevent the model from generating inappropriate words or phrases within the allowed token count.

102
MCQmedium

A developer uses the Azure OpenAI API to generate code. They want to ensure that the generated code is in Python. Which parameter should they set?

A.temperature
B.top_p
C.system message
D.max_tokens
AnswerC

The system message sets the model's behavioural context before any user turn, so instructing it there to respond only in Python reliably constrains the generated code's language. A user message or temperature setting cannot enforce this persistent constraint.

Why this answer

The system message is used to set the behavior and context of the AI model, including specifying the desired output format or language. By setting the system message to 'You are a helpful assistant that always writes code in Python', the developer can instruct the model to generate Python code consistently. This parameter is part of the chat completions API and directly influences the model's persona and constraints.

Exam trap

Microsoft often tests the distinction between parameters that control output randomness (temperature, top_p) and those that control output structure or behavior (system message), leading candidates to mistakenly choose temperature or top_p for language specification.

How to eliminate wrong answers

Option A is wrong because temperature controls the randomness of the output, not the language or format of the generated code. Option B is wrong because top_p (nucleus sampling) controls the cumulative probability threshold for token selection, which affects diversity but does not specify the output language. Option D is wrong because max_tokens limits the length of the generated response, not the programming language or content type.

103
MCQmedium

You are building a generative AI solution with Azure OpenAI Service that must produce concise, factual answers using only a supplied knowledge base. During testing, the model sometimes invents details not present in the knowledge base. You need to reduce hallucinations without retraining the model. What should you do?

A.Increase the max_tokens setting so the model has more room to explain its reasoning.
B.Switch to a larger model deployment with a bigger context window.
C.Lower the temperature and include a system message that instructs the model to answer only from the provided context and to say it does not know when the answer is absent.
D.Fine-tune the base model on the knowledge base so it memorizes the facts.
AnswerC

Lowering temperature reduces sampling randomness, and a system message that constrains the model to the supplied context encourages grounded answers and an explicit 'I don't know' fallback. Together these prompt-engineering techniques reduce hallucination without retraining, and they work within the existing deployment. This directly addresses invented details while keeping the model unchanged.

Why this answer

Hallucinations in grounded generation are commonly reduced with prompt engineering: lower temperature to make output more deterministic, and a system message that restricts answers to the provided context and instructs the model to admit when the answer is not present. These changes need no retraining and can be applied immediately to the deployment.

Exam trap

The trap here is reaching for fine-tuning or a larger model to fix factual errors, when the scenario explicitly rules out retraining and the issue is prompt grounding.

104
MCQmedium

A company wants to generate personalized product descriptions for its e-commerce site using Azure OpenAI. They need to ensure the model's output adheres to brand guidelines and does not generate prohibited content. Which approach should they use?

A.Use a system message with brand guidelines and apply content filtering.
B.Use prompt engineering with negative prompts and ignore content filtering.
C.Provide few-shot examples in the user message and rely on the model's training.
D.Fine-tune the model with brand guidelines and disable content filtering for performance.
AnswerA

A system message sets persistent behavioural instructions, so brand tone and prohibited-topic rules apply to every completion, while Azure OpenAI's content filter blocks harmful categories independently. Together they satisfy both the brand-guideline and prohibited-content constraints.

Why this answer

Using a system message allows you to embed brand guidelines directly into the conversation context, instructing the model on tone, style, and prohibited content. Azure OpenAI's content filtering provides an additional safety layer by automatically detecting and blocking harmful or policy-violating outputs, ensuring compliance with both brand and regulatory requirements.

Exam trap

Microsoft often tests the misconception that fine-tuning or prompt engineering alone is sufficient for safety and compliance, when in reality Azure OpenAI requires explicit content filtering and system messages to enforce brand guidelines reliably.

How to eliminate wrong answers

Option B is wrong because ignoring content filtering removes the safety guardrails that prevent prohibited content, and negative prompts alone are unreliable for enforcing brand guidelines. Option C is wrong because few-shot examples in the user message do not guarantee consistent adherence to brand guidelines across all outputs, and relying solely on the model's training ignores the need for explicit content filtering. Option D is wrong because disabling content filtering for performance sacrifices safety and compliance, and fine-tuning alone cannot dynamically enforce brand guidelines as effectively as a system message combined with content filtering.

105
MCQmedium

Refer to the exhibit. { "content_filters": [ { "type": "hate", "action": "block", "severity": "high" }, { "type": "sexual", "action": "block", "severity": "medium" }, { "type": "self_harm", "action": "block", "severity": "low" } ] } You deploy an Azure OpenAI model with the above content filter configuration. A user submits a prompt that the system rates as "hate" at severity level "medium". What happens?

A.The prompt is allowed because the severity is below the threshold.
B.The prompt is blocked because hate content is detected.
C.The prompt is blocked because the severity is medium.
D.The prompt is allowed because the hate filter is not configured for medium.
AnswerA

The hate filter blocks at severity high, so medium-severity hate content falls below that threshold and passes. The action applies only when the rated severity meets or exceeds the configured level, so the prompt is not blocked.

Why this answer

The content filter configuration blocks 'hate' content only at severity level 'high'. Since the user's prompt was rated as 'hate' at severity 'medium', it falls below the configured threshold and is allowed. Azure OpenAI content filters evaluate severity levels (low, medium, high) and apply the configured action only when the detected severity meets or exceeds the specified threshold.

Exam trap

A common mistake in Azure exams is assuming that any detection of a content type (e.g., hate) automatically triggers the configured action, ignoring that only severity levels meeting or exceeding the threshold are blocked.

How to eliminate wrong answers

Option B is wrong because the filter does not block all hate content indiscriminately; it only blocks hate content at severity 'high' or above. Option C is wrong because the severity 'medium' is below the configured threshold of 'high' for the hate filter, so it is not blocked. Option D is wrong because the hate filter is indeed configured (with action 'block' and severity 'high'), but it simply does not apply to severity 'medium'.

106
MCQhard

You are deploying a generative AI solution that must produce structured JSON order confirmations from free-text customer emails. Downstream systems reject any response that is not valid JSON matching a fixed schema. Which approach best guarantees schema-conformant output from an Azure OpenAI chat completion?

A.Set the temperature to zero and include an example of the exact JSON output in the prompt.
B.Use a structured outputs configuration with a strict JSON schema that defines required fields and types.
C.Set response_format to json_object and describe the desired schema in the system message.
D.Request the output in JSON and run a regex-based validator that repairs any deviations before forwarding.
AnswerB

Structured outputs constrain decoding so the generated tokens conform to a supplied JSON schema, including required properties and allowed types. This provides a strong guarantee that the response validates against the schema, which is exactly what downstream systems demand. Combined with a clear prompt, it removes the need for retry loops or brittle post-processing, making it the most reliable option for strict schema conformance.

Why this answer

Structured outputs constrain token generation to a supplied JSON schema, enforcing required fields and correct types so responses validate against the contract. JSON mode only guarantees syntactic validity, and few-shot examples or post-processing repairs do not provide a hard structural guarantee. For downstream systems that reject any schema deviation, constrained decoding is the most reliable approach.

Exam trap

The trap here is equating JSON mode with schema enforcement, when JSON mode only guarantees valid JSON syntax.

107
MCQhard

You are building a generative AI solution using Azure OpenAI Service that must generate code snippets. The solution must ensure that the generated code does not include any known vulnerable patterns, such as SQL injection. You need to implement a safeguard. What should you do?

A.Implement a post-processing step that analyzes the generated code with a static analysis tool and rejects or sanitizes unsafe snippets.
B.Add a system message instructing the model to never generate vulnerable code.
C.Fine-tune the model on a dataset of secure code to reduce the likelihood of generating vulnerable patterns.
D.Use Azure AI Content Safety to scan the generated code for vulnerabilities and block unsafe outputs.
AnswerA

A static analysis tool can detect known vulnerable patterns like SQL injection in generated code. By post-processing the model's output, you can enforce security policies before the code is used. This is a robust safeguard that complements the generative model and ensures only safe code is accepted. It is a standard practice in secure code generation.

Why this answer

The most reliable safeguard is to post-process the generated code with a static analysis tool that detects vulnerable patterns like SQL injection. This provides deterministic enforcement, rejecting or sanitizing unsafe outputs before use. While prompt instructions and fine-tuning can help, they do not guarantee security.

Static analysis is a standard, effective method for ensuring code safety.

Exam trap

The trap here is assuming that content safety filters or prompt instructions can reliably prevent code vulnerabilities, when dedicated static analysis is needed.

108
MCQeasy

You are deploying a generative AI application using Azure OpenAI Service. The application must generate responses that are creative and varied for a marketing campaign. Which parameter should you adjust to increase the randomness of the model's output?

A.Temperature
B.Presence penalty
C.Top-p (nucleus sampling)
D.Frequency penalty
AnswerA

Temperature controls the randomness of the model's output. Higher values (e.g., 1.0 or above) make the output more varied and creative, while lower values make it more deterministic. For a marketing campaign requiring creative and varied responses, increasing the temperature is the correct adjustment.

Why this answer

Temperature is the parameter that directly controls the randomness of the model's output. Higher temperatures lead to more diverse and creative responses, which is suitable for a marketing campaign that requires varied content. Other parameters like top-p and penalties affect diversity in different ways but are not the primary control for randomness.

Exam trap

The trap here is confusing parameters that increase diversity (like top-p or penalties) with the primary parameter for randomness, which is temperature.

109
Multi-Selecthard

Which THREE components are required to build a custom copilot using Microsoft Copilot Studio that can answer questions from a SharePoint document library?

Select 3 answers
A.A Microsoft Copilot Studio copilot
B.A knowledge source configured in Copilot Studio (e.g., Azure Cognitive Search)
C.A Power Automate flow to trigger the copilot
D.A SharePoint site with the documents
E.An Azure OpenAI Service deployment
AnswersA, B, D

The copilot is the conversational interface.

Why this answer

A Microsoft Copilot Studio copilot is the core conversational interface that orchestrates the interaction. Without the copilot itself, there is no runtime environment to process user queries, manage conversation state, or invoke configured knowledge sources. It is the mandatory host for all custom copilot functionality.

Exam trap

A common misconception is that you must provision an Azure OpenAI Service deployment to use generative AI in Copilot Studio, when in fact the platform provides its own managed GPT models and only requires a knowledge source like SharePoint.

110
MCQhard

A company uses Azure OpenAI to generate product descriptions. They notice that the model occasionally produces descriptions that include false claims about product features. The company needs to reduce the frequency of these inaccuracies without changing the training data. Which parameter adjustment would be most effective?

A.Increase the top_p parameter
B.Increase the max_tokens parameter
C.Decrease the temperature parameter
D.Increase the frequency_penalty parameter
AnswerC

Lowering temperature reduces sampling randomness, so the model favours higher-probability tokens and produces more grounded, less speculative text. This directly addresses the false product claims without altering training data, satisfying the stem's constraint of no training-data changes.

Why this answer

Decreasing the temperature parameter reduces the randomness of the model's output, making it more deterministic and less likely to generate creative but factually incorrect statements. This directly addresses the need to reduce false claims without modifying training data, as lower temperature forces the model to rely on its most probable (and typically more accurate) token predictions.

Exam trap

The trap here is that candidates often confuse temperature with creativity or length control, assuming that increasing randomness (higher temperature) or extending output length (max_tokens) will somehow improve accuracy, when in fact lower temperature is the standard parameter for reducing hallucinations.

How to eliminate wrong answers

Option A is wrong because increasing top_p (nucleus sampling) expands the set of candidate tokens considered, which increases output diversity and can actually worsen factual inaccuracies by allowing less probable tokens. Option B is wrong because increasing max_tokens only extends the maximum length of the generated text; it does not influence the factual accuracy or creativity of the content. Option D is wrong because increasing frequency_penalty penalizes tokens that have already appeared, reducing repetition but not addressing the root cause of hallucinated or false claims.

111
MCQhard

You are building a generative AI application that must process large volumes of PDF documents and generate summaries using Azure OpenAI. The solution must be cost-effective and handle variable workloads. Which architecture should you recommend?

A.Use Azure Kubernetes Service (AKS) with a persistent node pool of GPU nodes.
B.Use Azure Functions with a consumption plan to trigger processing jobs and call Azure OpenAI.
C.Deploy a GPU-enabled virtual machine and run the summarization jobs sequentially.
D.Use Azure Logic Apps to iterate through documents and call Azure OpenAI.
AnswerB

Azure Functions on a consumption plan scales automatically and bills per execution, matching the stem's cost-effectiveness and variable-workload constraints. It triggers PDF processing jobs that call Azure OpenAI, avoiding idle capacity charges that always-on compute would incur.

Why this answer

Azure Functions with a consumption plan provides a serverless, event-driven architecture that scales automatically to handle variable workloads, ensuring cost-effectiveness by charging only for compute time used. This architecture is ideal for processing large volumes of PDFs, as each document can trigger a function execution that calls Azure OpenAI for summarization, without the need for always-on infrastructure.

Exam trap

Microsoft often tests the misconception that GPU or specialized compute is required for AI workloads, but in this scenario, the heavy lifting is done by Azure OpenAI's API, so the focus should be on cost-effective, scalable compute for orchestration, not local GPU processing.

How to eliminate wrong answers

Option A is wrong because Azure Kubernetes Service (AKS) with a persistent node pool of GPU nodes incurs continuous costs even during idle periods, making it less cost-effective for variable workloads, and the GPU nodes are unnecessary since Azure OpenAI is called via API, not run locally. Option C is wrong because deploying a GPU-enabled virtual machine and running summarization jobs sequentially introduces a single point of failure, lacks auto-scaling, and wastes resources on GPU hardware that is not required for API calls. Option D is wrong because Azure Logic Apps is designed for workflow orchestration and integration, not for high-throughput, cost-effective batch processing of large document volumes, and it would incur higher costs per execution compared to Azure Functions.

112
MCQmedium

You are a cloud solution architect at a legal firm. The firm needs to automate the summarization of legal documents. They have a large corpus of past case summaries and legal documents stored in Azure Blob Storage. They want to use Azure OpenAI to generate summaries for new documents. The solution must ensure that the generated summaries are accurate and do not contain hallucinated legal facts. The firm also requires that the solution be serverless and minimize operational overhead. You need to design the solution. Option A: Use Azure OpenAI with a system message that instructs the model to be accurate. Deploy the model as a web app on Azure App Service and call it from Azure Functions triggered by new blob uploads. Option B: Use Azure OpenAI with Retrieval-Augmented Generation (RAG) by indexing the past case summaries in Azure AI Search. Use Azure Functions to process new documents, retrieve relevant cases, and pass them as context to the model. Store summaries in Azure Cosmos DB. Option C: Fine-tune an Azure OpenAI model on the past case summaries and deploy it as a managed endpoint. Use Azure Logic Apps to trigger summarization when new blobs are added. Option D: Use Azure OpenAI with the chat API and provide the entire document in the prompt. Use Azure Container Instances to run a service that calls the API and writes summaries back to Blob Storage. Which option should you choose?

A.Option B
B.Option A
C.Option D
D.Option C
AnswerA

RAG grounds responses in retrieved documents, reducing hallucination.

Why this answer

It uses Retrieval-Augmented Generation (RAG) with Azure AI Search to ground the model's output in verified past case summaries, directly addressing the requirement to avoid hallucinated legal facts. The serverless architecture is achieved via Azure Functions triggered by blob uploads, minimizing operational overhead, while storing summaries in Azure Cosmos DB provides a scalable, low-latency output store.

Exam trap

The trap here is that candidates may assume fine-tuning (Option C) or a simple system message (Option A) is sufficient to ensure factual accuracy, but Azure OpenAI models require grounded context via RAG to reliably avoid hallucination in domain-specific tasks like legal summarization.

How to eliminate wrong answers

Option A is wrong because a system message alone cannot prevent hallucination; the model may still fabricate legal facts without grounded context. Option C is wrong because fine-tuning on past case summaries does not guarantee factual accuracy for new, unseen documents and introduces operational overhead with a managed endpoint, contradicting the serverless requirement. Option D is wrong because providing the entire document in the prompt without retrieval augmentation does not anchor the model to verified facts, and Azure Container Instances adds operational overhead compared to a serverless trigger.

113
MCQmedium

A company is building a chatbot using Azure OpenAI Service to answer customer queries. The chatbot must not generate harmful or offensive content. Which Azure AI service should be integrated to filter inappropriate content?

A.Azure Bot Service
B.Azure Cognitive Search
C.Azure AI Content Safety
D.Azure Form Recognizer
AnswerC

Azure AI Content Safety provides dedicated moderation models that detect harmful, violent, hateful and sexual content in both prompts and completions. Integrating it filters inappropriate generated output, satisfying the requirement that the chatbot must not produce offensive responses.

Why this answer

Azure AI Content Safety is the correct service because it provides built-in content moderation APIs that detect and filter harmful or offensive text and images, including hate speech, violence, self-harm, and sexual content. Integrating this service with the Azure OpenAI chatbot ensures that user inputs and model outputs are screened in real time, preventing the generation of inappropriate responses.

Exam trap

The trap here is that candidates often confuse Azure Bot Service's ability to 'manage conversations' with built-in content filtering, but it actually lacks native moderation and requires explicit integration with a dedicated content safety service.

How to eliminate wrong answers

Option A is wrong because Azure Bot Service is a framework for building, deploying, and managing bots, but it does not include native content filtering capabilities; it would require integration with a separate content moderation service. Option B is wrong because Azure Cognitive Search is used for indexing and searching over structured and unstructured data, not for filtering harmful content in real-time chat interactions. Option D is wrong because Azure Form Recognizer (now Azure AI Document Intelligence) is designed to extract information from forms and documents, not to moderate or filter offensive language or imagery.

114
MCQhard

A healthcare startup is developing a chatbot that uses Azure OpenAI to answer patient questions. They need to ensure that the chatbot only uses information from their verified medical database and does not generate unsupported medical advice. What is the best approach?

A.Fine-tune a model on the medical database and deploy it.
B.Embed the entire medical database in the system message.
C.Rely on Azure OpenAI's content filtering to block unsupported advice.
D.Use Azure AI Search with vector search to retrieve relevant documents and pass them as context.
AnswerD

Retrieval-augmented generation grounds responses in your own data: Azure AI Search with vector search returns the most semantically relevant verified medical documents, which are injected into the prompt as context. This constrains the model to the medical database, preventing unsupported advice.

Why this answer

It uses Azure AI Search with vector search to retrieve only relevant, verified documents from the medical database and passes them as context to the Azure OpenAI model. This grounds the model's responses in authoritative data, preventing it from generating unsupported medical advice. The retrieval-augmented generation (RAG) pattern ensures the chatbot answers are based on the provided context rather than the model's internal knowledge.

Exam trap

Microsoft often tests the misconception that fine-tuning or content filtering alone can control factual accuracy, when in reality retrieval-augmented generation (RAG) with Azure AI Search is the correct pattern for grounding responses in specific, verified data.

How to eliminate wrong answers

Option A is wrong because fine-tuning a model on a medical database does not guarantee it will avoid generating unsupported advice; the model can still hallucinate or produce information not present in the training data, and fine-tuning does not enforce retrieval of specific verified documents at inference time. Option B is wrong because embedding the entire medical database in the system message would exceed the token limit (typically 4,096 or 8,192 tokens for most models), making it impractical and inefficient, and it would not allow dynamic retrieval of the most relevant information. Option C is wrong because Azure OpenAI's content filtering is designed to block harmful or offensive content, not to verify the factual accuracy or medical validity of the model's responses; it cannot prevent the generation of unsupported medical advice that appears plausible.

115
MCQmedium

You are building a generative AI assistant with Azure OpenAI Service that must answer questions using a large corpus of internal product manuals. The manuals are updated weekly, and the assistant must reflect changes without retraining the model. You need to implement a retrieval-augmented generation (RAG) pattern. What should you use to index and retrieve relevant manual content?

A.Use Azure Cosmos DB with the built-in vector search and connect it directly to the model via a function calling tool.
B.Create an Azure AI Search index with vector embeddings of the manual chunks and use it as a data source in the Azure OpenAI 'on your data' configuration.
C.Store the manuals as plain text in Azure Blob Storage and pass the entire corpus in each prompt.
D.Fine-tune the base model on the manuals each week using Azure OpenAI fine-tuning.
AnswerB

This correctly implements RAG by using vector search to retrieve semantically relevant manual chunks, which are then passed to the model. Azure AI Search supports vector indexes and integrates with Azure OpenAI 'on your data', enabling weekly updates without retraining. It is the standard, supported approach for grounding generative responses in dynamic internal content.

Why this answer

The correct approach is to use Azure AI Search with vector embeddings as a data source for Azure OpenAI 'on your data'. This enables retrieval-augmented generation, where relevant manual chunks are retrieved and supplied to the model, ensuring responses reflect weekly updates without retraining. It is the supported, scalable method for grounding generative AI in dynamic internal documents.

Exam trap

The trap here is assuming that fine-tuning or passing all documents in the prompt can substitute for a proper retrieval index.

116
Multi-Selecthard

You are deploying a retrieval-augmented generation chat solution on Azure OpenAI. During testing, users report that the model confidently answers questions using facts that are not in the indexed documents, and that retrieved chunks sometimes come from the wrong department's SharePoint site. You need to reduce these ungrounded answers without retraining the model. (Choose two.)

Select 2 answers
A.Set the temperature parameter to 2 to make the model more factual.
B.Increase the max_tokens parameter so the model has more room to explain its reasoning before answering.
C.Set the message role "system" content to instruct the model to answer only from the provided context and to state when the answer is not found.
D.Enable content filtering at the highest severity threshold to block ungrounded statements.
E.Apply security trimming and metadata filters so retrieval only returns chunks the user is allowed to see and that match the requested department.
AnswersC, E

A system message sets behavioral boundaries for the whole chat session, so telling the model to restrict itself to supplied context and to admit when it cannot find an answer directly reduces fabricated responses. This is the standard, non-retraining grounding control used with Azure OpenAI chat completions, and it works alongside retrieval quality fixes rather than replacing them.

Why this answer

Ungrounded answers in a RAG pattern usually come from two sources: the model not being constrained to the retrieved context, and the retrieval step returning irrelevant or unauthorized chunks. A system message that enforces answer-only-from-context behavior, combined with metadata filtering and security trimming at the search layer, addresses both causes without retraining. Length and randomness parameters, plus harmful-content filters, do not influence factual grounding.

Exam trap

The trap here is assuming that lowering or raising sampling parameters such as temperature or max_tokens will fix hallucinations, when grounding depends on prompt constraints and retrieval quality instead.

117
MCQeasy

Your company wants to build a custom Copilot for customer support using Microsoft Copilot Studio. The support team needs to query a backend CRM system securely using the Copilot. Which authentication method should you configure for the custom connector to the CRM?

A.Microsoft Entra ID OAuth 2.0
B.API key authentication
C.Basic authentication
D.Client certificate authentication
AnswerA

Microsoft Entra ID OAuth 2.0 lets Copilot Studio authenticate to the CRM through delegated, token-based authorisation, so the support team queries CRM data securely without embedding credentials. This satisfies the stem's secure backend CRM authentication requirement.

Why this answer

Microsoft Entra ID OAuth 2.0 is the correct authentication method because it provides secure, token-based delegated access to the backend CRM system without exposing long-lived credentials. Copilot Studio connectors that require user context and secure resource access should use OAuth 2.0 with Microsoft Entra ID, which supports modern authentication flows like authorization code grant and can enforce conditional access policies.

Exam trap

The trap here is that candidates often choose API key authentication because it seems simpler, but Microsoft Copilot Studio connectors for secure backend systems require OAuth 2.0 to support user delegation and token lifecycle management, not static keys.

How to eliminate wrong answers

Option B (API key authentication) is wrong because API keys are static, shared secrets that lack user context, token expiration, and granular scoping, making them unsuitable for secure, delegated access to a CRM in a Copilot scenario. Option C (Basic authentication) is wrong because it transmits credentials in plaintext (Base64-encoded) over HTTP, which is insecure and violates modern security best practices; it also does not support token refresh or scoped permissions. Option D (Client certificate authentication) is wrong because while it provides strong mutual TLS authentication, it is typically used for server-to-server or machine-to-machine scenarios, not for user-delegated access from a Copilot that needs to act on behalf of a support agent.

118
MCQmedium

You are building an Azure OpenAI solution that must generate product descriptions from a small set of 40 curated examples that demonstrate your brand voice. You want the model to imitate the style without changing the model weights or incurring the cost of a fine-tuning job. What should you do?

A.Set the temperature parameter to 2 and the top_p parameter to 0.1 on every request.
B.Upload the examples to an Azure AI Search index and enable semantic ranker on the index.
C.Include several of the curated examples directly in the system and user messages of each chat completion request.
D.Create a fine-tuning job in Azure OpenAI Studio using the 40 examples and deploy the resulting custom model.
AnswerC

Placing curated examples in the system and user messages is few-shot prompting, which conditions the model on your brand voice at inference time without modifying weights. It is the appropriate approach when you have a small set of examples and want to avoid the cost and latency of a fine-tuning job. The examples guide tone and structure for each response.

Why this answer

Few-shot prompting embeds a handful of curated examples in the prompt so the model mimics their tone, structure, and vocabulary on each call. Because the requirement is a small example set and no weight modification or fine-tuning cost, in-context examples are the correct mechanism. Sampling parameters and retrieval indexes do not substitute for style conditioning.

Exam trap

The trap here is assuming that any style customization requires fine-tuning, when a small example set is better handled with in-context few-shot prompting.

119
MCQmedium

You are implementing a generative AI chat solution on Azure OpenAI Service. The solution must ground its answers in a large internal knowledge base and return inline citations that link back to the exact source chunks. You want to minimize custom code and use a managed retrieval pipeline. Which Azure OpenAI feature should you configure?

A.Use the On Your Data feature with an Azure AI Search index and set the citation output to include document references.
B.Deploy a fine-tuned model trained on the internal knowledge base and use it for chat completions.
C.Set the temperature parameter to a low value and enable content filtering on the deployment.
D.Create a system message that instructs the model to answer only from the knowledge base and to include source names.
AnswerA

On Your Data with Azure AI Search is the managed retrieval-augmented generation feature of Azure OpenAI. It chunks, indexes, and retrieves content from your data source and returns citations that reference the source documents, with minimal custom code. Setting citations on ensures answers link back to the retrieved chunks.

Why this answer

Grounding answers in private data with verifiable citations requires a retrieval pipeline. Azure OpenAI On Your Data with an Azure AI Search index provides managed ingestion, chunking, retrieval, and citation output, so answers link to the actual source chunks. Prompt instructions or model tuning alone cannot retrieve documents or guarantee accurate references.

Exam trap

The trap here is assuming that a prompt or fine-tuned model can reliably cite internal documents without a retrieval index.

120
MCQhard

You are developing a generative AI application using Azure OpenAI. The application must log all prompts and completions for auditing, but sensitive data such as personal health information (PHI) must be redacted before logging. Which approach should you use?

A.Use Azure AI Language's PII detection to identify and redact PHI in prompts and completions before writing to your logging store.
B.Enable Azure OpenAI diagnostic settings to send logs to Azure Monitor, and rely on built-in PHI redaction.
C.Configure Azure OpenAI to not log prompts and completions, and instead log only token counts and model names.
D.Store logs in an Azure Storage account with encryption at rest and enable soft delete.
AnswerA

Azure AI Language provides PII detection that can identify and redact personal health information and other sensitive data. By calling this service on prompts and completions before logging, you ensure that only redacted text is stored. This meets auditing needs while protecting sensitive data.

Why this answer

To audit prompts and completions while protecting PHI, integrate Azure AI Language PII detection into your logging pipeline. Redact sensitive entities before persisting logs. This satisfies auditing requirements and compliance mandates, whereas relying on diagnostic logs or storage encryption alone would leave PHI exposed.

Exam trap

The trap here is assuming that Azure OpenAI diagnostic logs or storage encryption automatically redact PHI, when redaction must be performed explicitly before logging.

121
MCQhard

You have deployed a chatbot using Azure OpenAI with a system message as shown. The chatbot sometimes provides incorrect answers that are not supported by the sources. What is the most likely cause?

A.Content filters are incorrectly configured, allowing harmful content.
B.The data ingestion pipeline has errors, so the sources are not available.
C.The temperature parameter is too low, causing repetitive answers.
D.The system message does not guarantee grounding; the model may still hallucinate.
AnswerD

System messages steer tone and boundaries but cannot enforce factual grounding; the model still generates tokens probabilistically, so unsupported claims remain possible. Since the stem's constraint is answers not supported by sources, hallucination persists regardless of prompt wording. Retrieval augmentation or grounding data is required to constrain outputs to source content.

Why this answer

A system message in Azure OpenAI provides instructions and context but does not enforce factual grounding. The model can still generate responses that are not supported by the provided sources (hallucination), especially if the system message is not explicitly designed to restrict the model to only use the given data. Grounding requires additional techniques like retrieval-augmented generation (RAG) or explicit constraints in the prompt.

Exam trap

The trap here is that candidates often assume a system message is sufficient to enforce factual accuracy, confusing instruction-following with grounded generation, and overlook the need for retrieval-augmented generation or explicit source constraints.

How to eliminate wrong answers

Option A is wrong because content filters control the safety and appropriateness of outputs, not the factual accuracy or grounding of responses; they would not cause unsupported answers. Option B is wrong because if the data ingestion pipeline had errors making sources unavailable, the chatbot would likely fail to retrieve data or return errors, not produce plausible but incorrect answers. Option C is wrong because a low temperature parameter makes outputs more deterministic and repetitive, but it does not cause hallucination; in fact, lower temperature often reduces randomness and can improve consistency with training data, but it does not guarantee grounding to specific sources.

122
Multi-Selectmedium

Which TWO actions should you take to reduce the cost of using Azure OpenAI for a chatbot that handles high traffic?

Select 2 answers
A.Increase max_tokens to reduce the number of requests.
B.Use the batch API for non-real-time requests.
C.Implement caching for frequently asked questions.
D.Lower the temperature to 0 to reduce token usage.
E.Fine-tune the model to reduce prompt length.
AnswersB, C

The batch API processes asynchronous, non-real-time workloads at a 50% discount compared with standard global deployments, directly satisfying the stem's cost-reduction constraint for high-traffic chatbots. Offloading tolerant requests frees real-time capacity, lowering overall spend while preserving interactive latency for genuine user conversations.

Why this answer

Option B is correct because the Azure OpenAI Batch API processes asynchronous, non-real-time workloads at a 50% discount compared to standard global pricing, so routing any chatbot requests that don't need immediate responses through batch jobs directly lowers per-token cost. Option C is correct because caching responses to frequently asked questions (for example with Azure Cache for Redis or a semantic cache) means repeated identical or similar prompts are served without calling the model at all, eliminating token charges for that traffic and reducing overall request volume. Option A is not correct because increasing max_tokens raises the maximum output length and therefore can increase, not decrease, token consumption and cost.

Option D is not correct because temperature controls randomness in sampling, not the number of tokens generated, so setting it to 0 does not reduce token usage or cost. Option E is not correct because fine-tuning does not shorten the prompt by itself and adds training and hosting costs; prompt length is reduced through techniques like prompt compression or shorter system messages, not fine-tuning alone.

Exam trap

The trap here is that candidates confuse token-related parameters (max_tokens, temperature) with cost-saving mechanisms, when in reality cost reduction for high-traffic chatbots relies on architectural patterns like batching and caching, not on tweaking inference parameters.

123
MCQeasy

You are deploying a conversational AI solution using Microsoft Copilot Studio. The solution must provide responses based on data from an internal knowledge base stored in SharePoint. Which feature should you configure?

A.Configure Azure AI Language with custom question answering.
B.Create an Azure Bot Service with a QnA Maker knowledge base.
C.Enable Generative Answers and add SharePoint as a data source.
D.Use Azure OpenAI Service with data integration (Bring Your Own Data).
AnswerC

Generative Answers retrieves and synthesises responses from configured enterprise data sources, and SharePoint is a natively supported source in Microsoft Copilot Studio. Adding it grounds the conversational agent in the internal knowledge base, satisfying the requirement to answer from SharePoint content rather than relying on the model's pretrained knowledge.

Why this answer

Microsoft Copilot Studio natively supports Generative Answers, which allows the AI to dynamically generate responses from specified data sources without requiring a separate QnA or custom question-answering service. By adding SharePoint as a data source, the solution can directly query the internal knowledge base stored in SharePoint and produce contextually relevant answers. This is the most straightforward and integrated approach within Copilot Studio for this scenario.

Exam trap

The trap here is that candidates often confuse the need for a separate QnA service (like Azure AI Language or QnA Maker) with Copilot Studio's built-in Generative Answers capability, not realizing that Copilot Studio can directly connect to SharePoint as a live data source without additional services.

How to eliminate wrong answers

Option A is wrong because Azure AI Language with custom question answering is a separate service that requires manual ingestion and management of Q&A pairs, whereas Copilot Studio's Generative Answers feature handles this automatically from SharePoint. Option B is wrong because Azure Bot Service with QnA Maker is a legacy approach that is now deprecated in favor of Copilot Studio and Generative Answers; QnA Maker also requires explicit Q&A pair extraction and does not natively integrate with SharePoint as a live data source. Option D is wrong because Azure OpenAI Service with data integration (Bring Your Own Data) is designed for custom GPT models and requires more complex setup and indexing, while Copilot Studio provides a simpler, out-of-the-box solution for connecting to SharePoint.

124
MCQeasy

You are building a generative AI solution using Azure OpenAI. The solution must generate responses that include factual information from a set of internal documents that are updated frequently. You need to ensure the model uses the latest document versions without retraining. What should you do?

A.Fine-tune the Azure OpenAI model each time the documents are updated.
B.Use Azure OpenAI with a retrieval-augmented generation pattern that queries an up-to-date search index of the documents.
C.Increase the model's temperature setting to encourage more creative and up-to-date responses.
D.Embed the documents directly into the model's prompt by concatenating all of them in each request.
AnswerB

RAG retrieves relevant passages from an index at query time, so when documents are updated and reindexed, the model automatically uses the latest content. This avoids retraining and ensures responses are grounded in current information. It is the standard approach for dynamic knowledge bases.

Why this answer

Retrieval-augmented generation with a search index allows the model to access current document content at inference time. When documents change, you update the index, and the model retrieves the latest passages. This avoids fine-tuning and ensures responses reflect the most recent information, which is critical for frequently updated knowledge bases.

Exam trap

The trap here is thinking that fine-tuning or higher temperature can keep a model current, when in fact retrieval from a refreshed index is what provides up-to-date grounding.

125
MCQeasy

You are designing a generative AI assistant that must produce deterministic, repeatable answers for a compliance workflow. The assistant uses Azure OpenAI chat completions and must return the same output for identical prompts. Which parameter setting should you apply?

A.Set frequency_penalty to 2 and presence_penalty to 2.
B.Set temperature to 0 and top_p to 1.
C.Increase the max_tokens value and enable streaming responses.
D.Set temperature to 1 and top_p to 0.95.
AnswerB

Temperature 0 makes the model select the highest-probability token at each step, and top_p 1 disables nucleus sampling restrictions. Together they minimize randomness, producing nearly deterministic and repeatable outputs for identical prompts and parameters, which matches the compliance requirement for consistent answers.

Why this answer

Deterministic output requires minimizing sampling randomness. Temperature 0 selects the most probable token each step, and top_p 1 disables nucleus truncation, so identical inputs produce nearly identical outputs. Creative temperatures, penalties, or length and streaming controls do not provide repeatability and can introduce variation across runs.

Exam trap

The trap here is confusing output length or streaming controls with parameters that actually govern randomness.

126
MCQeasy

Refer to the exhibit. You are deploying a GPT-4 model using Azure OpenAI Service. The deployment uses the Standard scale type. Which statement is true about this deployment?

A.The model version is not specified and will default to the latest.
B.Content filtering is disabled for this deployment.
C.The deployment uses provisioned throughput with reserved capacity.
D.The deployment uses pay-as-you-go pricing and global rate limits.
AnswerD

Standard deployments bill per token on a pay-as-you-go basis and draw on the shared, globally pooled quota for that model, so throughput varies with regional demand. Provisioned scale types instead reserve dedicated capacity with predictable latency and a fixed hourly charge.

Why this answer

The Standard scale type in Azure OpenAI Service uses pay-as-you-go pricing, where you are billed based on the number of tokens processed. It also enforces global rate limits (e.g., tokens per minute) that apply across all deployments in the region, rather than providing dedicated capacity. Option D correctly identifies these characteristics.

Exam trap

Microsoft often tests the distinction between Standard (pay-as-you-go with rate limits) and Provisioned (reserved capacity) scale types, and candidates mistakenly associate 'Standard' with default model versions or disabled content filtering.

How to eliminate wrong answers

Option A is wrong because when deploying a GPT-4 model, the model version must be explicitly specified; it does not default to the latest version. Option B is wrong because content filtering is enabled by default for all Azure OpenAI Service deployments and cannot be disabled through the scale type setting. Option C is wrong because provisioned throughput with reserved capacity is a feature of the Provisioned scale type, not the Standard scale type.

127
MCQeasy

You are deploying an Azure OpenAI model for a public-facing FAQ assistant. The assistant must answer only questions covered by a fixed set of approved topics, and any off-topic question should receive a polite refusal. Which approach most directly enforces this behavior?

A.Raise the frequency_penalty parameter to reduce repeated off-topic phrases.
B.Deploy the model with the lowest available quota tier to limit how many questions users can ask.
C.Set the model deployment's content filter to block the 'violence' category at high severity.
D.Configure a system message that defines the allowed topics and instructs the model to refuse anything outside them.
AnswerD

The system message sets the model's operating instructions and persona for every turn in the conversation. Defining the permitted topics and the refusal behavior there gives the model a consistent boundary to apply across all user turns. This is the most direct, low-overhead way to constrain scope for a simple FAQ assistant.

Why this answer

The system message is the model's persistent instruction set, making it the natural place to declare allowed topics and refusal behavior. For a fixed FAQ scope, this directly shapes every response. The other settings influence sampling, safety categories, or capacity, none of which restrict the assistant to approved topics.

Exam trap

The trap here is confusing safety content filters with topical scope control, when filters only block harmful categories rather than off-topic questions.

128
MCQmedium

You are deploying a generative AI application using Azure OpenAI Service. The application must generate responses in multiple languages while maintaining high accuracy. You need to minimize token usage. Which approach should you recommend?

A.Use a base model with a system message to output in the desired language
B.Translate all input to English and then translate output back
C.Fine-tune a model for each target language
D.Use a separate deployment for each language
AnswerA

A system message steers the base model to respond in the target language, avoiding separate per-language models or lengthy translation prompts. This keeps prompt tokens minimal, satisfying the requirement to minimise token usage while retaining accuracy.

Why this answer

Azure OpenAI Service base models (e.g., GPT-4) natively support multilingual generation via a system message that sets the desired output language. This approach avoids the overhead of translation pipelines or fine-tuning, directly minimizing token usage while maintaining high accuracy through the model's inherent multilingual capabilities.

Exam trap

Azure AI-102 often tests the misconception that translation pipelines or fine-tuning are necessary for multilingual support, when in fact a system message on a base model achieves the goal with lower token usage and complexity.

How to eliminate wrong answers

Option B is wrong because translating all input to English and then back adds significant token overhead (doubling input/output tokens) and introduces translation errors that degrade accuracy, contradicting the requirement to minimize token usage. Option C is wrong because fine-tuning a separate model for each target language is resource-intensive, requires large labeled datasets per language, and does not reduce token usage compared to a single base model with a system message. Option D is wrong because using a separate deployment for each language multiplies infrastructure costs and management complexity without any token savings, as each deployment still processes tokens for the same base model.

129
MCQhard

Refer to the exhibit. You are configuring a system message for an Azure OpenAI deployment. The assistant is still generating harmful code despite the instruction. Which additional measure should you implement?

A.Fine-tune the model on safe code examples.
B.Lower the temperature parameter to 0.
C.Add more examples to the prompt.
D.Enable Azure AI Content Safety with a custom blocklist for harmful code.
AnswerD

Prompt instructions alone cannot reliably block harmful code generation. Azure AI Content Safety filters model inputs and outputs against configurable harm categories, and a custom blocklist adds scenario-specific terms, enforcing the constraint at the platform layer regardless of what the system message says.

Why this answer

Azure AI Content Safety provides a dedicated content filtering layer that can block harmful code generation at the inference level, regardless of the system message. A custom blocklist allows you to define specific patterns (e.g., code snippets for malware) that the model is prohibited from outputting, enforcing safety beyond prompt instructions.

Exam trap

The trap here is that candidates often assume prompt engineering (system messages or few-shot examples) is sufficient for safety, but Azure OpenAI requires explicit content filtering via Azure AI Content Safety to reliably block harmful outputs at scale.

How to eliminate wrong answers

Option A is wrong because fine-tuning on safe code examples would require retraining the model, which is costly, time-consuming, and not a quick mitigation for an existing deployment; it also does not guarantee blocking of harmful code at inference time. Option B is wrong because lowering the temperature to 0 makes the model more deterministic but does not prevent it from generating harmful code if that code is the most likely completion. Option C is wrong because adding more examples to the prompt (few-shot prompting) can guide behavior but is unreliable for safety enforcement, as the model may still generate harmful code if the examples are not exhaustive or if the model overfits to the instruction.

130
MCQeasy

A company wants to generate product descriptions for thousands of items using an Azure OpenAI GPT-4 model. They need to ensure the descriptions match a consistent brand voice. Which approach is most efficient and cost-effective?

A.Write a separate prompt for each product category
B.Use Azure OpenAI on your data with a vector database of brand guidelines
C.Set a system message with brand voice guidelines and use few-shot examples
D.Fine-tune a base model on existing product descriptions
AnswerC

A system message plus few-shot examples conditions one GPT-4 deployment to produce consistent brand-voice output across thousands of items, avoiding the cost and latency of fine-tuning or per-item prompt engineering. This satisfies the consistency and cost-efficiency constraints simultaneously.

Why this answer

Setting a system message with brand voice guidelines and providing few-shot examples allows the GPT-4 model to consistently apply the desired tone and style across all product descriptions without retraining. This approach is efficient and cost-effective as it avoids the high compute and data preparation costs of fine-tuning, while still enabling precise control over output through in-context learning.

Exam trap

Microsoft often tests the misconception that fine-tuning is always the best approach for consistency, but in Azure OpenAI, in-context learning via system messages and few-shot examples is more efficient and cost-effective for tasks like brand voice adherence, as fine-tuning is reserved for deep customization of model behavior.

How to eliminate wrong answers

Option A is wrong because writing a separate prompt for each product category would be highly inefficient and inconsistent, as it requires manual effort for thousands of items and does not leverage the model's ability to generalize from a single system message. Option B is wrong because using Azure OpenAI on your data with a vector database of brand guidelines is overkill for this task; vector databases are designed for retrieval-augmented generation (RAG) to ground responses in external data, but brand voice guidelines are better conveyed via system messages and examples, not as searchable documents. Option D is wrong because fine-tuning a base model on existing product descriptions is costly, requires significant labeled data, and risks overfitting to the training set, whereas in-context learning with a system message and few-shot examples achieves the same goal with far less expense and complexity.

131
MCQhard

Refer to the exhibit. A developer runs this PowerShell script to call Azure OpenAI. The script fails with an authentication error. What is the most likely cause?

A.The script uses the wrong HTTP header; it should use 'api-key' instead of 'Authorization: Bearer'.
B.The script uses the wrong HTTP method; it should use GET.
C.The API version is incorrect.
D.The endpoint URI is missing the resource name.
AnswerA

Azure OpenAI's data-plane REST API authenticates with the 'api-key' header carrying the resource key, not an OAuth bearer token. Sending 'Authorization: Bearer' therefore triggers a 401 authentication error, so switching the header to 'api-key' resolves the failure.

Why this answer

The script uses 'Authorization: Bearer' header, but Azure OpenAI requires the API key to be passed in the 'api-key' header. The 'Authorization: Bearer' header is used for Azure AD token-based authentication, not for direct API key authentication. Since the script is using an API key (as indicated by the PowerShell script), the correct header is 'api-key'.

Exam trap

The trap here is that candidates confuse Azure OpenAI's API key authentication with Azure AD token authentication, assuming 'Authorization: Bearer' is always correct, when in fact the header name differs based on the authentication method.

How to eliminate wrong answers

Option A is correct because Azure OpenAI API key authentication requires the 'api-key' header, not 'Authorization: Bearer'. Option B is wrong because the Azure OpenAI chat completions endpoint requires a POST method, not GET, to send the prompt and parameters in the request body. Option C is wrong because the API version is specified in the URI (e.g., '2023-12-01-preview') and an incorrect version would return a '400 Bad Request' or '404 Not Found', not an authentication error.

Option D is wrong because the endpoint URI includes the resource name (e.g., 'https://<resource>.openai.azure.com'), and a missing resource name would cause a DNS resolution failure or '404 Not Found', not an authentication error.

132
MCQhard

Your team ships a generative assistant that calls a custom function to look up live order status. Users report that the assistant sometimes invents an order status instead of calling the function, and that when it does call the tool the arguments are occasionally malformed. You need to make tool invocation more reliable while keeping the existing model. What should you do?

A.Define the function with a clear name, description, and parameter schema, set tool_choice to require a tool call, and validate returned arguments before executing.
B.Switch the deployment to a larger model and keep the existing function definitions unchanged.
C.Add the phrase "always call the function" to the system message and leave tool_choice at its default value.
D.Raise the frequency_penalty so the model is discouraged from repeating previously invented order statuses.
AnswerA

Function calling reliability depends on precise tool metadata and an explicit tool_choice setting. Requiring a tool call stops the model from answering from memory, and a well-described parameter schema reduces malformed arguments. Validating arguments before execution prevents bad data from reaching the backend, which addresses both reported symptoms without changing the model.

Why this answer

Reliable function calling comes from three things working together: descriptive tool metadata so the model understands when and how to call it, an explicit tool_choice that forces a call when the intent is known, and client-side validation of the returned arguments before execution. Penalties, larger models, and prompt-only pleading do not enforce tool invocation or argument correctness.

Exam trap

The trap here is assuming that a stronger model or a sterner system message will force tool use, when the decision is governed by tool_choice and the argument shape by the parameter schema.

133
MCQeasy

Refer to the exhibit. You are using Microsoft Graph to retrieve user information for use in a Microsoft 365 Copilot extension. The response shows that the mail and mobilePhone fields are null. What is the most likely reason?

A.The user has not configured those properties in their Microsoft Entra ID profile.
B.The API call required additional permissions.
C.The user is a guest user.
D.The user does not exist in the tenant.
AnswerA

Microsoft Graph returns only values populated in the directory; null mail and mobilePhone indicate those attributes were never set on the user object in Microsoft Entra ID. Missing consent or permissions would typically produce an error, not null fields.

Why this answer

The mail and mobilePhone fields are null because the user has not populated these attributes in their Microsoft Entra ID (formerly Azure AD) profile. Microsoft Graph returns the actual stored values for these properties; if they are empty or unset, the API response will show null. This is the most common and straightforward reason for null values in user profile fields.

Exam trap

Microsoft often tests the misconception that missing data in a successful API response is due to permission issues or user type, when in reality the most likely cause is that the data simply hasn't been configured.

How to eliminate wrong answers

Option B is wrong because if the API call required additional permissions, the response would return a 403 Forbidden error or an insufficient privileges message, not a successful response with null fields. Option C is wrong because guest users can have mail and mobilePhone properties configured; being a guest does not inherently cause these fields to be null—they would only be null if the guest user's profile lacks those values. Option D is wrong because if the user did not exist in the tenant, the API would return a 404 Not Found error, not a successful response with null fields.

134
MCQeasy

You need to deploy a generative AI model that can generate images from text descriptions. Which Azure service should you use?

A.Azure OpenAI Service
B.Azure Machine Learning
C.Azure AI Vision
D.Azure AI Language
AnswerA

Azure OpenAI Service hosts DALL-E models that generate images directly from text prompts, satisfying the text-to-image requirement. Unlike Azure AI Vision, which only analyses or classifies existing images, it performs true generative synthesis. This makes it the appropriate service for producing original visuals from descriptions.

Why this answer

Azure OpenAI Service provides access to advanced generative AI models like DALL-E, which are specifically designed to generate images from natural language text descriptions. This service offers pre-trained models that can create high-quality images based on textual prompts, making it the correct choice for this task.

Exam trap

The trap here is that candidates may confuse Azure AI Vision (which analyzes images) with image generation, or assume that Azure Machine Learning is the only way to implement generative AI, overlooking the purpose-built Azure OpenAI Service for this task.

How to eliminate wrong answers

Option B is wrong because Azure Machine Learning is a platform for building, training, and deploying custom machine learning models, not a pre-built service for generating images from text; it would require you to develop and train your own image generation model from scratch. Option C is wrong because Azure AI Vision is designed for analyzing and extracting information from images (e.g., object detection, OCR), not for generating images from text descriptions. Option D is wrong because Azure AI Language focuses on natural language processing tasks such as text analysis, translation, and sentiment analysis, and does not include image generation capabilities.

135
MCQmedium

You are deploying a conversational AI chatbot using Azure AI Language service. The chatbot must be able to switch between multiple intents in a single conversation without restarting the session. Which feature should you enable?

A.Active learning
B.Orchestration workflow
C.Dynamic entity extraction
D.Prebuilt domain components
AnswerB

Orchestration workflow connects multiple Azure AI Language projects, such as conversational language understanding and question answering, into one bot. It routes each utterance to the appropriate intent mid-conversation, letting the chatbot switch intents without restarting the session.

Why this answer

Orchestration workflow in Azure AI Language service allows a chatbot to switch between multiple intents within a single conversation by routing requests to different language models or custom question answering knowledge bases. This enables the chatbot to handle diverse intents seamlessly without restarting the session, as each intent can be processed by the most appropriate component.

Exam trap

The trap is confusing active learning (a feedback loop for model improvement) with orchestration workflow (a routing mechanism for multi-intent conversations). Orchestration is the correct feature for switching intents without restarting.

How to eliminate wrong answers

Option B is wrong because orchestration workflow is used to connect multiple language models or services (e.g., combining LUIS with QnA Maker) but does not inherently enable switching between intents within a single conversation without restarting; it manages routing between different models. Option C is wrong because dynamic entity extraction handles the identification of entities that vary in value (e.g., dates, numbers) but does not affect the ability to switch between intents mid-conversation. Option D is wrong because prebuilt domain components provide ready-made models for common scenarios (e.g., booking flights) but do not enable dynamic intent switching; they are static and require retraining to adapt to new intents.

136
Multi-Selecthard

Your organization is deploying a generative AI chatbot using Azure OpenAI Service. The chatbot must answer questions based on internal documents stored in Azure Blob Storage. You need to implement a retrieval-augmented generation (RAG) solution. Which THREE components are required? (Select THREE.)

Select 3 answers
A.Azure Functions for preprocessing
B.Azure AI Search index
C.Azure OpenAI On Your Data configuration
D.Azure SQL Database for metadata
E.Embedding model deployment in Azure OpenAI
AnswersB, C, E

Stores embeddings and enables vector search.

Why this answer

Azure AI Search is the core indexing and retrieval engine in a RAG solution. It ingests documents from Azure Blob Storage, creates a searchable index, and enables vector or hybrid search to retrieve relevant chunks. The chatbot then uses these retrieved chunks as context for the Azure OpenAI model to generate grounded answers.

Exam trap

The trap here is that candidates often confuse optional preprocessing components (like Azure Functions) or auxiliary storage (like Azure SQL Database) as mandatory, when the three essential pillars are the search index, the embedding model, and the Azure OpenAI On Your Data integration that ties retrieval to generation.

137
MCQhard

You are a machine learning engineer at a large retail company. The company has thousands of product descriptions that need to be updated regularly. They currently use a manual process. You propose using Azure OpenAI to generate new descriptions based on product attributes. You have a dataset of existing product descriptions and attributes stored in an Azure SQL Database. The solution must be cost-effective, scalable, and must not require retraining the model. You need to design the solution. You have the following options: Option A: Use Azure OpenAI with few-shot learning by including examples in the prompt for each product. Deploy the model on an Azure Kubernetes Service (AKS) cluster for high throughput. Option B: Use Azure OpenAI with prompt templates that include product attributes and call the API for each product. Use Azure Logic Apps to orchestrate the workflow and store results back to Azure SQL Database. Option C: Fine-tune a custom model on the existing product descriptions and deploy it as a managed endpoint. Use Azure Data Factory to batch process all products. Option D: Use Azure OpenAI with the batch API to generate descriptions for all products at once, using a single prompt that lists all products and attributes. Store the batch output in Azure Blob Storage and then import into Azure SQL Database. Which option should you choose?

A.Option C
B.Option D
C.Option A
D.Option B
AnswerD

Prompt templates with attributes are cost-effective and scalable.

Why this answer

(Azure Logic Apps) is the correct choice. It uses Azure OpenAI with prompt templates that insert product attributes, making individual API calls per product. This approach is scalable because Azure Logic Apps can handle high volumes with built-in retry and concurrency, and it is cost-effective as you only pay per API call and execution.

It does not require model retraining. In contrast, Option A (AKS) introduces unnecessary infrastructure complexity; Option C (fine-tuning) requires retraining; and Option D (batch API) risks exceeding prompt size limits and is less suitable for incremental updates.

Exam trap

The trap is that candidates may mistakenly choose Option A (AKS) thinking it provides better scalability, Option C (fine-tuning) for customization, or Option D (batch API) for efficiency, but they overlook that Option B (Azure Logic Apps) offers the right balance of cost-effectiveness, scalability, and no retraining for incremental updates.

How to eliminate wrong answers

Option A is wrong because few-shot learning with examples in the prompt for each product is not cost-effective for thousands of products—it increases token usage and latency, and deploying on AKS adds unnecessary infrastructure complexity without addressing the need for batch processing. Option B is wrong because using Azure Logic Apps to call the API for each product individually is not scalable for thousands of products—it would result in high latency, cost, and potential throttling, and it does not leverage batch processing for efficiency. Option C is wrong because fine-tuning a custom model requires retraining, which violates the requirement that the solution must not require retraining the model, and deploying as a managed endpoint adds ongoing cost and complexity.

138
MCQhard

You are a data scientist at a healthcare company. You have deployed a GPT-4 model using Azure OpenAI to answer patient inquiries about medical conditions. The model is configured with temperature=0.3 and max_tokens=200. Recently, the compliance team flagged that some responses contain contradictory information compared to the official medical guidelines. You need to ensure the model's answers align strictly with the provided medical documents (stored as PDFs in Azure Blob Storage). You have access to Azure Cognitive Search and Azure AI Document Intelligence. The solution must minimize hallucinations and not require retraining the model. What should you do?

A.Use prompt engineering to add a system message that tells the model to only answer based on the uploaded PDFs. Keep the current deployment.
B.Index the medical PDFs into Azure Cognitive Search. Configure the Azure OpenAI deployment to use 'Add your data' pointing to this index. Set the system message to instruct the model to base answers only on the retrieved context.
C.Fine-tune GPT-4 on the medical documents using Azure OpenAI fine-tuning capabilities. Use the fine-tuned model for the chatbot.
D.Deploy Azure AI Content Safety to filter responses that contradict guidelines. Set up a custom content filter using a list of approved phrases.
AnswerB

Indexing the PDFs in Azure Cognitive Search and wiring the deployment's "Add your data" to that index implements retrieval-augmented generation: relevant guideline passages are injected into the prompt, grounding each response. The system message restricting answers to retrieved context directly satisfies the no-retraining constraint while minimising contradictory output.

Why this answer

It uses Azure Cognitive Search to index the medical PDFs and then configures the Azure OpenAI deployment with 'Add your data' to retrieve relevant context from that index at inference time. This retrieval-augmented generation (RAG) approach grounds the model's answers in the official documents without retraining, directly addressing the compliance team's requirement to align responses with the provided guidelines and minimize hallucinations.

Exam trap

Microsoft often tests the distinction between prompt engineering (which is lightweight but unreliable for grounding) and RAG with a search index (which provides verifiable, document-grounded responses), leading candidates to choose the simpler prompt-only solution without considering its inability to enforce factual accuracy.

How to eliminate wrong answers

Option A is wrong because prompt engineering alone cannot guarantee that the model will only use the uploaded PDFs; the model's internal knowledge may still produce contradictory information, and there is no mechanism to enforce retrieval of the actual document content. Option C is wrong because fine-tuning GPT-4 on the medical documents would require retraining the model, which contradicts the requirement to not retrain, and fine-tuning does not inherently prevent hallucinations when the model encounters out-of-distribution queries. Option D is wrong because Azure AI Content Safety with a custom filter of approved phrases is a post-hoc filtering approach that cannot ensure the model's responses are grounded in the specific PDFs; it would only block or flag responses that match a predefined list, not align answers with dynamic document content.

139
MCQmedium

You deploy a chat application using Azure OpenAI Service. Users report that the model sometimes generates inappropriate content. You need to implement a safety system that can be customized for your organization's policies. What should you use?

A.Use the content filter system in Azure OpenAI Studio
B.Use Azure AI Content Safety with custom categories and severity thresholds
C.Apply responsible AI templates from Azure AI Studio
D.Configure Microsoft Entra ID Conditional Access policies
AnswerB

Azure AI Content Safety provides configurable harm categories with adjustable severity thresholds, letting you align filtering with your organisation's own policies rather than relying on the model's fixed default filters. This directly satisfies the requirement for a customisable safety system.

Why this answer

Azure AI Content Safety provides a customizable content moderation service that allows you to define custom categories and severity thresholds aligned with your organization's specific policies. This enables you to filter inappropriate content beyond the default filters, giving you granular control over what the model generates.

Exam trap

The trap here is that candidates often confuse the built-in content filters in Azure OpenAI Studio (which are not customizable) with Azure AI Content Safety (which is a separate, customizable service), leading them to choose option A.

How to eliminate wrong answers

Option A is wrong because the content filter system in Azure OpenAI Studio provides only predefined content categories (e.g., hate, violence) with fixed severity levels, and cannot be customized to match an organization's unique policies. Option C is wrong because responsible AI templates in Azure AI Studio are design patterns and guidelines for building ethical AI, not a runtime content safety system that can filter generated outputs. Option D is wrong because Microsoft Entra ID Conditional Access policies control authentication and access to resources, not the content generated by the AI model.

140
MCQhard

You are designing a generative AI solution that uses Azure OpenAI GPT-4 to answer customer support questions. The solution must comply with Microsoft's Responsible AI principles, particularly transparency and accountability. Which implementation approach best meets these requirements?

A.Use the model without any modifications, and have a human review all responses.
B.Fine-tune the model on a curated dataset of support tickets and disable content filtering.
C.Enable content filtering, log all interactions, and include a disclaimer that responses are AI-generated.
D.Use the default model deployment and rely on the model's inherent safety.
AnswerC

Logging every interaction creates the audit trail that accountability demands, while the AI-generated disclaimer delivers the transparency requirement by telling users they are not reading human output. Content filtering addresses harm prevention rather than the two named principles, but the logging and disclosure pairing directly satisfies the stem's stated constraints.

Why this answer

It directly addresses Microsoft's Responsible AI principles of transparency and accountability. Enabling content filtering (via Azure AI Content Safety) ensures harmful outputs are blocked, logging all interactions provides an audit trail for accountability, and including a disclaimer that responses are AI-generated satisfies transparency by clearly informing users they are interacting with an AI system.

Exam trap

The trap here is that candidates assume human review (Option A) or model fine-tuning (Option B) alone satisfy Responsible AI principles, but Microsoft explicitly requires automated content filtering, logging, and transparency disclaimers as part of a comprehensive compliance strategy.

How to eliminate wrong answers

Option A is wrong because using the model without modifications fails to implement content filtering or logging, leaving the solution non-compliant with accountability and safety requirements; human review alone is insufficient for real-time compliance and does not provide automated transparency. Option B is wrong because disabling content filtering violates safety principles, and fine-tuning on a curated dataset does not guarantee compliance with transparency or accountability; it also risks overfitting or introducing bias without proper oversight. Option D is wrong because relying solely on the model's inherent safety is insufficient; Azure OpenAI's default deployment does not enforce logging or disclaimers, and the model can still produce harmful or non-transparent outputs without explicit content filtering and audit mechanisms.

141
MCQhard

An organization is deploying a conversational AI solution using Azure OpenAI. They want to ensure the model's responses are grounded in their own knowledge base documents to reduce hallucinations. Which approach should they implement?

A.Integrate Azure Cognitive Search for retrieval-augmented generation (RAG)
B.Fine-tune the model on the knowledge base documents
C.Implement Azure AI Content Safety filters
D.Use prompt engineering to instruct the model to only use the knowledge base
AnswerA

Retrieval-augmented generation queries Azure Cognitive Search to fetch relevant passages from the organisation's own documents, then passes them into the Azure OpenAI prompt as grounding context. This directly satisfies the requirement to reduce hallucinations by anchoring responses in the knowledge base rather than relying solely on parametric model memory.

Why this answer

Retrieval-Augmented Generation (RAG) with Azure Cognitive Search allows the model to dynamically retrieve relevant chunks from the organization's knowledge base documents at inference time. This grounds responses in authoritative, up-to-date content, directly reducing hallucinations by providing factual context rather than relying solely on the model's parametric memory.

Exam trap

The trap here is that candidates often confuse fine-tuning (B) as a way to 'teach' the model the knowledge base, not realizing that RAG is the recommended pattern for grounding responses in external, query-specific data without retraining.

How to eliminate wrong answers

Option B is wrong because fine-tuning adjusts the model's weights on a static dataset, which can lead to overfitting and does not guarantee that the model will reference the knowledge base for every query; it also fails to incorporate new or updated documents without retraining. Option C is wrong because Azure AI Content Safety filters only block harmful or inappropriate content after generation; they do not provide factual grounding or reduce hallucinations. Option D is wrong because prompt engineering alone cannot enforce factual adherence; the model may still generate plausible-sounding but incorrect information from its training data, as it lacks a retrieval mechanism to verify claims against the knowledge base.

142
Multi-Selecthard

You are deploying a generative AI model using Azure AI Foundry. The model must be accessible only from a specific virtual network. Additionally, you need to monitor all API calls for auditing. Which two configurations are required? (Choose two.)

Select 2 answers
A.Enable diagnostic settings to send logs to a Log Analytics workspace.
B.Assign a managed identity to the model deployment.
C.Enable public network access from selected IP addresses.
D.Disable public network access and configure a private endpoint.
E.Configure CORS to allow only the VNet's domain.
AnswersA, D

Logs enable auditing of all API calls.

Why this answer

Enabling diagnostic settings to send logs to a Log Analytics workspace captures all API call details (e.g., request URI, response status, caller IP) for auditing and monitoring. This is the standard Azure method for collecting resource-level logs, and it works with Azure AI Foundry deployments to meet compliance and security requirements.

Exam trap

The trap here is that candidates confuse network access controls (private endpoints) with authentication mechanisms (managed identities) or browser-level restrictions (CORS), leading them to select B or E instead of the correct pairing of D and A.

143
MCQeasy

A marketing team uses Azure OpenAI Service to draft product descriptions. They report that outputs vary widely in tone and sometimes ignore the required brand voice. You need to make responses follow a fixed set of style rules consistently across all requests without changing the model deployment. What should you do?

A.Create a second deployment of the same model with a different name
B.Raise the temperature value to increase output creativity
C.Add a system message that defines the brand voice and formatting rules
D.Set the presence_penalty parameter to a large positive value
AnswerC

The system message is a high-priority instruction that shapes the model's behavior across the conversation, making it the correct place to encode tone, persona, and formatting constraints. Applied to every request, it enforces the brand voice consistently without retraining or redeploying the model. This is the standard, low-cost way to steer generative output in Azure OpenAI Service.

Why this answer

System messages carry instructions the model treats with high priority and apply to the whole request, so they are the right mechanism for enforcing brand voice, tone, and formatting rules. Sampling parameters such as temperature and presence_penalty alter randomness and repetition, and additional deployments of the same model behave identically, so none of them can encode style policy.

Exam trap

The trap here is reaching for a sampling parameter to control style, when style guidance belongs in the system message.

144
MCQmedium

Your company uses Microsoft 365 Copilot to generate meeting summaries. Some users report that summaries include information from meetings they did not attend. What is the most likely cause?

A.The meeting organizer granted everyone in the organization view access.
B.Users have access to meeting artifacts via shared calendars or transcripts.
C.Copilot is using Bing search results to augment summaries.
D.Copilot is incorrectly configured to ignore meeting permissions.
AnswerB

Copilot surfaces content from meeting artefacts the user can already reach through Microsoft 365 permissions, such as shared calendars, transcripts and recordings. Summaries therefore draw on meetings the user did not attend but still has access to, which is the most likely cause of the reported behaviour.

Why this answer

Microsoft 365 Copilot generates meeting summaries by aggregating content from meeting artifacts such as transcripts, recordings, and shared calendars. If a user has access to a meeting's transcript or recording (e.g., via a shared calendar or because the meeting was recorded and stored in a location the user can access), Copilot can include that meeting's information in summaries even if the user did not attend. This behavior is by design, as Copilot respects existing permissions on the underlying data.

Exam trap

The trap here is that candidates often assume Copilot uses meeting attendance or organizer permissions to filter summaries, when in fact it relies on the underlying permissions of the meeting's artifacts (transcripts, recordings, calendar items), which can be broader than the attendee list.

How to eliminate wrong answers

Option A is wrong because granting everyone view access to a meeting would allow users to see the meeting details, but Copilot does not automatically include meetings in summaries based solely on view access; it requires access to the meeting's artifacts like transcripts or recordings. Option C is wrong because Copilot does not use Bing search results to augment meeting summaries; it relies on the user's Microsoft Graph data and permissions, not external web searches. Option D is wrong because there is no configuration setting in Copilot to 'ignore meeting permissions'; Copilot strictly adheres to the permissions set on the meeting artifacts and does not have a mode that bypasses them.

145
MCQmedium

Your organization is building a chatbot using Azure OpenAI Service. The chatbot must provide citations from a set of internal documents stored in Azure Blob Storage. You need to configure the solution to minimize token usage while ensuring citations are accurate. Which approach should you use?

A.Embed all document content into the system prompt
B.Fine-tune a model on the documents so it can recall them from memory
C.Use a large context window model (e.g., 32K) and include all documents in the prompt
D.Use Azure OpenAI on your data with Azure Cognitive Search for hybrid retrieval
AnswerD

Azure OpenAI on your data with Cognitive Search hybrid retrieval combines keyword and vector search, returning only the most relevant document chunks as grounding context. That narrows the prompt payload, minimising tokens while preserving accurate citations from the Blob Storage documents.

Why this answer

Azure OpenAI on your data with Azure Cognitive Search for hybrid retrieval combines vector search and keyword search to efficiently find relevant document chunks from Azure Blob Storage, minimizing token usage by only sending the most pertinent content to the model for citation generation. This approach ensures accurate citations without embedding all documents into the prompt or relying on model memory.

Exam trap

The trap here is that candidates often confuse fine-tuning with retrieval-augmented generation (RAG), assuming fine-tuning can store factual knowledge for citation, when in reality RAG with a search index is required for accurate, token-efficient document grounding.

How to eliminate wrong answers

Option A is wrong because embedding all document content into the system prompt would consume an enormous number of tokens, exceeding context limits and incurring high costs, while also being impractical for large document sets. Option B is wrong because fine-tuning a model on documents does not enable it to recall specific citations accurately; fine-tuning adjusts model behavior but does not store document content for retrieval, leading to hallucinations or incorrect references. Option C is wrong because using a large context window model (e.g., 32K) and including all documents in the prompt still wastes tokens on irrelevant content, increases latency and cost, and does not guarantee accurate citations as the model may lose focus on the specific source material.

146
MCQmedium

You are using Azure OpenAI to generate product descriptions. You notice that the descriptions are often too similar to each other. Which parameter should you adjust to increase diversity?

A.Increase the temperature value.
B.Decrease the top_p value.
C.Increase the max_tokens value.
D.Increase the frequency_penalty value.
AnswerA

Raising temperature flattens the model's probability distribution over next tokens, so lower-probability words are sampled more often, producing varied outputs. This directly satisfies the stem's requirement to increase diversity in the generated product descriptions, countering the repetitive similarity observed.

Why this answer

Increasing the temperature parameter makes the model more creative by raising the probability of sampling lower-probability tokens, which increases diversity in the generated text. A higher temperature (e.g., 0.9) flattens the probability distribution, so the model is less likely to always pick the most probable next word, resulting in more varied outputs.

Exam trap

Microsoft often tests the distinction between temperature (which controls randomness/creativity) and frequency_penalty (which controls repetition), leading candidates to mistakenly choose frequency_penalty when the question asks for diversity in content rather than just avoiding repetition.

How to eliminate wrong answers

Option B is wrong because decreasing top_p (nucleus sampling) reduces the cumulative probability mass considered for token selection, which actually makes outputs less diverse by focusing only on the most likely tokens. Option C is wrong because increasing max_tokens only extends the maximum length of the generated response; it does not affect the randomness or diversity of token choices. Option D is wrong because increasing frequency_penalty reduces the likelihood of repeating the same tokens or phrases, which can increase lexical diversity but does not directly control the overall creativity or randomness of the output like temperature does.

147
MCQmedium

Refer to the exhibit. An administrator runs this Azure CLI command to deploy a GPT-4 model in Azure AI Foundry. The command fails with an error that the deployment name already exists. What should the administrator do to resolve the issue?

A.Use a different deployment name or delete the existing deployment.
B.Specify a different resource group.
C.Remove the --sku-name parameter.
D.Use a different model version.
AnswerA

Deployment names must be unique within an Azure AI Foundry resource. Since the CLI command failed because that name is already taken, the administrator must either supply a new unique name or remove the existing deployment before retrying.

Why this answer

The error message indicates that a deployment with the same name already exists in the Azure AI Foundry workspace. In Azure AI Foundry, deployment names must be unique within a workspace. The correct resolution is to either choose a different deployment name or delete the existing deployment before re-running the command.

This aligns with the Azure CLI behavior where resource names (including AI model deployments) must be unique per scope.

Exam trap

The trap here is that candidates may think the error is about model availability or SKU constraints, but the error explicitly states 'deployment name already exists,' which is a naming conflict, not a capacity or version issue.

How to eliminate wrong answers

Option B is wrong because specifying a different resource group does not resolve a deployment name conflict within the same workspace; the deployment name uniqueness is scoped to the workspace, not the resource group. Option C is wrong because removing the --sku-name parameter would change the pricing tier or capacity, but does not address the duplicate name error. Option D is wrong because using a different model version does not change the deployment name; the conflict is on the name, not the model version.

148
Multi-Selectmedium

Which THREE components are required to implement a Retrieval-Augmented Generation (RAG) solution with Azure OpenAI Service? (Choose three.)

Select 3 answers
A.An embedding model (e.g., text-embedding-ada-002)
B.A fine-tuned model
C.An Azure OpenAI Service model (LLM)
D.Azure AI Content Safety
E.A vector database (e.g., Azure AI Search)
AnswersA, C, E

Embedding models convert documents into vector representations.

Why this answer

An embedding model like text-embedding-ada-002 is essential for converting user queries and document chunks into dense vector representations. These vectors enable semantic similarity search in a vector database, which is the core retrieval step in RAG. Without embeddings, the system cannot match user intent to relevant content.

Exam trap

The trap here is that candidates often confuse optional safety or tuning components (like Content Safety or fine-tuning) as mandatory, when the core RAG triad is strictly retrieval (embeddings + vector DB) plus generation (LLM).

149
MCQmedium

You are deploying a generative AI solution by using Azure OpenAI Service. The solution must ground model responses in a private Azure AI Search index that contains product manuals. You need to configure the model deployment so that the service automatically retrieves relevant chunks from the index and includes them in the prompt at inference time. Which configuration should you use?

A.Upload the manuals to the model's training dataset and enable continuous training on the deployment.
B.Enable the 'Azure OpenAI On Your Data' feature by specifying the Azure AI Search data source in the model's chat completion request.
C.Fine-tune the base model on the contents of the product manuals and deploy the fine-tuned model for inference.
D.Create a custom retrieval pipeline by using Azure Functions that queries the index and appends results to each user message before calling the model.
AnswerB

Azure OpenAI On Your Data lets you attach an Azure AI Search index directly to a chat completion call, and the service handles chunk retrieval, embedding, and prompt assembly server-side, so relevant manual passages are injected before the model generates a response. This is the supported mechanism for grounding without building custom retrieval orchestration.

Why this answer

Azure OpenAI On Your Data is designed to connect a deployment to an Azure AI Search index so that retrieval and prompt augmentation happen automatically at request time. It handles embedding, chunk ranking, and citation generation, which matches the requirement to ground responses in private manuals without custom orchestration.

Exam trap

The trap here is assuming that fine-tuning or uploading documents trains the model to know private content, when runtime grounding actually requires a data source connection such as Azure OpenAI On Your Data.

150
MCQeasy

You need to create a chatbot that uses Azure OpenAI to answer questions about your company's internal policies. The responses must be based only on the provided policy documents. Which approach should you use?

A.Use the model's pre-existing knowledge about common policies.
B.Fine-tune a GPT model on the policy documents.
C.Use prompt engineering to instruct the model to only use policy knowledge.
D.Use Retrieval-Augmented Generation (RAG) with an Azure AI Search index of the documents.
AnswerD

RAG retrieves relevant passages from an Azure AI Search index and grounds Azure OpenAI responses in those documents, satisfying the constraint that answers derive only from the supplied policy content rather than the model's pretrained knowledge.

Why this answer

RAG (Retrieval-Augmented Generation) is the correct approach because it grounds the model's responses in your actual policy documents by retrieving relevant chunks from an Azure AI Search index at query time and injecting them into the prompt. This ensures the chatbot answers only from the provided documents, avoids hallucination, and keeps the source of truth external and updatable without retraining. Azure OpenAI's 'On Your Data' feature implements exactly this pattern using Azure AI Search as the vector/keyword store.

Exam trap

AI-102 often tests the misconception that fine-tuning 'teaches' a model new factual knowledge, when in reality fine-tuning shapes behavior and style while RAG is the correct pattern for grounding responses in specific, updatable documents.

How to eliminate wrong answers

Option A is wrong because the model's pre-existing knowledge is generic, may be outdated, and cannot be verified against your internal policies — it will hallucinate or answer from public data. Option B is wrong because fine-tuning teaches style, format, and tone rather than reliably injecting factual content; the model can still hallucinate, and updating policies would require re-running expensive training jobs. Option C is wrong because prompt engineering alone cannot guarantee the model has access to the actual policy text — without retrieval, the model has no way to know your specific internal documents and will fabricate answers.

← PreviousPage 2 of 3 · 174 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Implement generative AI solutions questions.