Courseiva

CCNA Implement generative AI solutions Questions

24 of 174 questions · Page 3/3 · Implement generative AI solutions · Answers revealed

151
MCQeasy

A team is prototyping an Azure OpenAI chat application and wants to iterate quickly on prompt wording without changing application code or redeploying. They also want to compare outputs from different prompt variants side by side. Which Azure AI Foundry feature should they use?

A.The Azure OpenAI resource's keys and endpoint blade
B.The model deployment quota page in the Azure portal
C.The prompt flow authoring experience in Azure AI Foundry
D.The content filter configuration page for the deployment
AnswerC

Prompt flow provides a visual authoring surface where prompts, inputs, and model configurations are defined as a flow and can be edited and re-run without changing application code. It supports comparing variants and evaluating outputs, which matches the goal of fast prompt iteration and side-by-side comparison. This is the intended tool for that workflow.

Why this answer

The requirement is to iterate on prompts and compare variants without code changes. Prompt flow in Azure AI Foundry is the authoring and orchestration surface designed for exactly that, letting prompts be edited, run, and evaluated visually. Quota pages, key and endpoint blades, and content filter settings are administrative or safety configuration surfaces and provide no prompt editing or comparison capability.

Exam trap

The trap here is confusing administrative configuration surfaces, such as quota or content filter pages, with the authoring environment where prompts are actually developed and evaluated.

152
MCQeasy

You are building a generative AI solution using Azure Machine Learning prompt flow. The solution must allow business analysts without coding experience to modify prompts and evaluate different model versions. What should you do?

A.Provide the analysts with a Jupyter notebook using the OpenAI Python SDK
B.Deploy a chatbot in Microsoft Copilot Studio and let analysts configure it
C.Implement a custom web UI using Azure Static Web Apps and Azure Functions
D.Use Azure Machine Learning prompt flow with the visual designer and variant management
AnswerD

The visual designer gives non-coders a drag-and-drop canvas to edit prompts, while variant management lets them compare model versions side by side. Together these satisfy the requirement that business analysts modify prompts and evaluate model versions without writing code.

Why this answer

Azure Machine Learning prompt flow provides a visual designer that enables non-technical users to modify prompts without coding, and its variant management feature allows them to evaluate different model versions side-by-side. This directly addresses the requirement for business analysts to iteratively refine prompts and compare model outputs in a controlled, no-code environment.

Exam trap

The trap here is that candidates may confuse the no-code visual designer in Azure Machine Learning prompt flow with other low-code tools like Copilot Studio or assume that a custom web UI is simpler, but the question specifically requires a solution that allows prompt modification and model evaluation without coding, which only prompt flow's variant management and visual designer provide.

How to eliminate wrong answers

Option A is wrong because Jupyter notebooks require Python coding skills and familiarity with the OpenAI SDK, which business analysts without coding experience cannot use. Option B is wrong because Microsoft Copilot Studio is designed for building conversational agents with pre-built templates, not for modifying prompts or evaluating different model versions in a generative AI pipeline. Option C is wrong because implementing a custom web UI with Azure Static Web Apps and Azure Functions requires significant development effort and coding, which defeats the purpose of enabling non-technical analysts to modify prompts directly.

153
MCQmedium

A company uses Azure OpenAI to generate marketing copy. They want to ensure that the generated content does not contain offensive language. Which feature should they enable?

A.Use DALL-E to generate images instead of text.
B.Use a system message instructing the model to avoid offensive language.
C.Enable diagnostic logging to review all outputs.
D.Enable content filtering at the deployment level.
AnswerD

Content filtering at the deployment level applies Azure OpenAI's classification models to both prompts and completions, blocking hate, violence, sexual and self-harm categories. This directly satisfies the requirement that generated marketing copy must not contain offensive language, enforcing the restriction on every request routed through that deployment.

Why this answer

Azure OpenAI provides built-in content filtering at the deployment level that automatically detects and blocks offensive or harmful language in both input prompts and generated outputs. This feature uses Microsoft's Responsible AI models to enforce safety policies without requiring custom code or manual review, making it the most reliable and scalable solution for preventing offensive content in marketing copy.

Exam trap

The trap here is that candidates often assume prompt engineering (system messages) is sufficient for safety, but Azure OpenAI requires explicit content filtering at the deployment level to enforce policies reliably and prevent bypassing via prompt injection.

How to eliminate wrong answers

Option A is wrong because DALL-E is an image generation model, not a text filtering mechanism; switching to images does not address the requirement to prevent offensive language in text outputs. Option B is wrong because a system message is a prompt engineering technique that provides guidance to the model but does not guarantee enforcement; the model may still generate offensive content if the instruction is not followed or if the model is manipulated. Option C is wrong because diagnostic logging only records outputs for review after generation, not preventing offensive content in real-time; it is a monitoring tool, not a content filter.

154
MCQeasy

You are developing a generative AI application that uses Azure OpenAI Service. The application must generate responses that are grounded in a specific set of documents stored in Azure Blob Storage. You want to use the simplest approach that allows the model to reference these documents without building a custom retrieval pipeline. What should you use?

A.Fine-tune the model with the documents and deploy the fine-tuned model.
B.Implement a custom RAG pipeline using Azure Cognitive Search and the Chat Completions API.
C.Azure OpenAI On Your Data with Azure Blob Storage as the data source.
D.Use the Completions API with a prompt that includes the full text of all documents.
AnswerC

Azure OpenAI On Your Data natively supports Azure Blob Storage as a data source. It handles indexing, retrieval, and citation generation automatically, requiring minimal custom code. This is the simplest way to ground responses in documents stored in Blob Storage without building a retrieval pipeline.

Why this answer

Azure OpenAI On Your Data is a managed feature that integrates with Azure Blob Storage, enabling the model to retrieve and cite documents without custom code. It handles indexing and retrieval automatically, making it the simplest solution for grounding responses in a specific document set. Other options require more effort or are not designed for this purpose.

Exam trap

The trap here is assuming that fine-tuning or manual prompt stuffing can achieve grounding, when a managed retrieval service is specifically designed for this.

155
MCQmedium

You are using Azure OpenAI Service to generate code snippets for a development team. You notice that the generated code sometimes contains security vulnerabilities. You need to minimize the risk of generating insecure code while maintaining productivity. What should you do?

A.Use system messages to instruct the model to prioritize security
B.Fine-tune the model on a dataset of secure code
C.Set the temperature parameter to 0
D.Disable content filtering to allow more flexibility
AnswerA

System messages set persistent behavioural instructions applied to every request, so embedding a security-first directive steers the model away from vulnerable patterns such as unsanitised input handling, satisfying the requirement to reduce insecure output without adding review overhead.

Why this answer

System messages in Azure OpenAI Service allow you to set the context and behavior of the model, including instructing it to prioritize security when generating code. This approach directly influences the model's output without requiring retraining or sacrificing flexibility, making it the most effective way to reduce security vulnerabilities while maintaining productivity.

Exam trap

The trap here is that candidates may overestimate the effectiveness of fine-tuning (Option B) for security, not realizing that system messages are a simpler, more practical first-line defense in Azure OpenAI Service, while fine-tuning is better suited for domain-specific style or knowledge rather than real-time safety constraints.

How to eliminate wrong answers

Option B is wrong because fine-tuning requires a curated dataset of secure code and significant computational resources, which is time-consuming and may not generalize well to all scenarios; it also reduces the model's flexibility for other tasks. Option C is wrong because setting the temperature parameter to 0 makes the model deterministic and less creative, which can hinder code generation quality and does not inherently address security vulnerabilities. Option D is wrong because disabling content filtering removes safety guardrails that help block harmful or insecure outputs, increasing the risk of generating vulnerable code rather than reducing it.

156
MCQhard

You are using Azure AI Foundry to fine-tune a GPT-3.5 model on a dataset of customer service conversations. The fine-tuning job fails with an error indicating that the training data format is invalid. What is the most likely issue?

A.The training data is not in JSONL format with the correct structure.
B.The training data is in CSV format instead of JSON.
C.The training data contains only one conversation example.
D.The training data does not include the assistant's responses.
AnswerA

Fine-tuning requires training data as JSONL, with each line holding a valid chat completion object containing messages with role and content pairs. Any deviation, such as plain JSON arrays or CSV, triggers the invalid format error.

Why this answer

Azure AI Foundry requires fine-tuning data to be in JSONL format with a specific structure: each line must be a JSON object containing a 'messages' array with 'role' and 'content' fields for system, user, and assistant turns. The error indicates the training data format is invalid, and the most likely cause is that the data is not in this required JSONL structure, as JSONL is the only accepted format for GPT-3.5 fine-tuning in Azure OpenAI Service.

Exam trap

The trap here is that candidates confuse the general requirement for 'JSON format' with the specific requirement for 'JSONL format with a messages array,' leading them to incorrectly select CSV or plain JSON as the issue, when the real problem is the lack of the correct conversational structure.

How to eliminate wrong answers

Option B is wrong because CSV format is not supported for fine-tuning GPT-3.5 models in Azure AI Foundry; the service requires JSONL, not JSON or CSV, and CSV lacks the nested 'messages' structure needed for conversational data. Option C is wrong because having only one conversation example does not cause a format error; it may lead to poor model performance but the format itself would still be valid if structured correctly. Option D is wrong because while missing assistant responses would make the data unusable for training, the error specifically indicates a format issue, not a content issue; the JSONL structure could still be technically valid without assistant responses.

157
MCQhard

You are building a generative AI solution using Azure OpenAI Service. The application must retrieve information from a large private knowledge base. You need to ensure the model uses only relevant documents from the knowledge base to generate answers. Which feature should you configure?

A.Implement a custom prompt flow
B.Use Azure OpenAI On Your Data with vector search
C.Configure a content filter
D.Fine-tune the model with the knowledge base
AnswerB

Azure OpenAI On Your Data with vector search indexes the private knowledge base and retrieves semantically relevant chunks, grounding responses in those documents rather than the model's parametric knowledge. This constrains generation to the supplied content.

Why this answer

B is correct because Azure OpenAI On Your Data with vector search enables the model to retrieve only the most semantically relevant documents from a private knowledge base by converting both the user query and the documents into high-dimensional vectors and performing similarity search. This ensures the model's responses are grounded in the specific, relevant information without exposing the entire knowledge base to the model.

Exam trap

The trap here is that candidates often confuse fine-tuning (D) with retrieval-augmented generation (RAG), assuming that training the model on the knowledge base is the best way to ground answers, when in fact RAG with vector search is the correct pattern for dynamic, relevant document retrieval without modifying the base model.

How to eliminate wrong answers

Option A is wrong because implementing a custom prompt flow does not inherently include a retrieval mechanism; it only orchestrates the sequence of calls and prompts, so it cannot ensure that only relevant documents are used from the knowledge base. Option C is wrong because configuring a content filter is a safety mechanism to block harmful or inappropriate content, not a retrieval or grounding feature to select relevant documents. Option D is wrong because fine-tuning the model with the knowledge base would bake the entire knowledge into the model's weights, which is inefficient, costly, and does not allow dynamic retrieval of only relevant documents per query; it also risks overfitting and cannot handle updates to the knowledge base without retraining.

158
MCQmedium

You are using Azure OpenAI Service to generate marketing copy. The marketing team reports that the generated content sometimes contains factual inaccuracies. You need to improve the factual accuracy of the generated content. What should you do?

A.Increase the max_tokens parameter
B.Include relevant context and facts in the prompt
C.Decrease the temperature parameter
D.Disable content filtering
AnswerB

Grounding the model with relevant facts and context in the prompt constrains generation to supplied information, reducing hallucinated claims. This directly addresses the factual inaccuracy problem by giving the model authoritative source material rather than relying on parametric knowledge alone.

Why this answer

The most effective way to improve factual accuracy in Azure OpenAI generations is to ground the model with relevant context and facts in the prompt — this is the core of Retrieval-Augmented Generation (RAG). By supplying authoritative source content, the model conditions its output on verified information rather than relying solely on parametric memory, which reduces hallucinations. This directly addresses the marketing team's complaint about factual inaccuracies.

Exam trap

AI-102 often tests the misconception that lowering temperature or increasing tokens improves factual accuracy — candidates must recognize that grounding with context (RAG) is the correct approach to reduce hallucinations.

How to eliminate wrong answers

Option A is wrong because max_tokens only controls the length of the generated output, not its factual correctness — increasing it can even allow more room for fabricated content. Option C is wrong because lowering temperature reduces randomness and makes outputs more deterministic, but a deterministic model can still confidently state incorrect facts; temperature does not inject or remove knowledge. Option D is wrong because disabling content filtering removes safety guardrails against harmful content and has no bearing on factual accuracy — it may even increase risk of inappropriate outputs.

159
MCQeasy

You are prototyping a chat experience on Azure OpenAI and want the model to produce structured JSON that matches a schema your application can deserialize reliably. Which feature should you configure?

A.Enable content filtering and set the severity threshold to high.
B.Increase top_p to 1 and rely on the model's instruction-following ability.
C.Set response_format to json_schema with a strict schema definition on the chat completions call.
D.Set temperature to 0 and add the word JSON to the user prompt.
AnswerC

Structured outputs with a strict JSON schema constrain generation so the response conforms to the supplied schema, which makes deserialization dependable. This is the purpose-built mechanism for schema-constrained JSON in chat completions and is more reliable than prompt-only instructions, because the model is constrained during decoding rather than merely asked to comply.

Why this answer

Schema-constrained generation is the only option that enforces structure at decode time. Configuring response_format with a strict JSON schema makes the model emit payloads that conform to the defined fields and types, so the application can deserialize without defensive parsing. Sampling parameters, prompt hints, and content filters do not enforce schema conformance.

Exam trap

The trap here is believing that setting temperature to zero or naming JSON in the prompt guarantees valid JSON, when only schema-constrained decoding enforces the contract.

160
Multi-Selecteasy

You are deploying a chat application using Azure OpenAI. The application should only answer questions based on a specific set of internal documents. Which THREE features should you use?

Select 3 answers
A.Azure AI Search index with the internal documents
B.Grounding with your data in Azure OpenAI Studio
C.Content filters to block out-of-domain questions
D.System message to limit the assistant's scope
E.Fine-tuning the model on the internal documents
AnswersA, B, D

The index provides the data source for grounding.

Why this answer

Azure AI Search indexes allow you to ingest internal documents and perform vector or hybrid search over them. When integrated with Azure OpenAI, the search results are used as grounding context for the model, ensuring responses are based solely on your data.

Exam trap

Microsoft often tests the distinction between content filtering (which handles safety) and domain restriction (which requires retrieval or prompt engineering), leading candidates to incorrectly select content filters for limiting question scope.

161
MCQeasy

You need to generate an image of a cat wearing a hat using Azure OpenAI. Which model should you use?

A.Codex
B.DALL-E
C.GPT-4
D.Whisper
AnswerB

DALL-E is Azure OpenAI's dedicated image-generation model, so it produces a new picture from a text prompt such as a cat wearing a hat. GPT models output text, not images, so they cannot satisfy the image-generation requirement.

Why this answer

DALL-E is the Azure OpenAI model specifically designed for generating images from natural language descriptions. It uses a diffusion-based architecture to create high-quality, original images based on text prompts, making it the correct choice for generating an image of a cat wearing a hat.

Exam trap

The trap here is that candidates may confuse GPT-4's multimodal capabilities (which can analyze images but not generate them) with DALL-E's generative image creation, leading them to incorrectly select GPT-4 for image generation tasks.

How to eliminate wrong answers

Option A is wrong because Codex is a model specialized for generating code from natural language, not for image generation. Option C is wrong because GPT-4 is a large language model focused on text generation and reasoning, lacking native image generation capabilities. Option D is wrong because Whisper is a speech-to-text model designed for audio transcription, not image generation.

162
MCQeasy

You need to deploy a generative AI model that can be used by multiple applications within your organization. The model must support real-time inference with low latency. Which Azure service should you use?

A.Azure AI Search
B.Azure OpenAI Service
C.Azure Machine Learning real-time endpoint
D.Azure Functions
AnswerB

Azure OpenAI Service hosts generative models behind a managed REST endpoint with provisioned throughput, giving multiple applications shared, low-latency real-time inference. It satisfies the low-latency constraint that batch-oriented or self-hosted alternatives cannot meet as directly.

Why this answer

Azure OpenAI Service provides managed access to powerful generative AI models like GPT-4, which are optimized for real-time inference with low latency through provisioned throughput units (PTUs) and regional deployment options. This service is specifically designed for generative AI workloads, offering REST API endpoints that support streaming responses and sub-second latency for single-turn interactions, making it ideal for multiple applications requiring consistent, low-latency responses.

Exam trap

The trap here is that candidates often confuse Azure Machine Learning real-time endpoints (which are for custom ML models) with Azure OpenAI Service (which is purpose-built for generative AI), overlooking the fact that Azure OpenAI provides managed, low-latency inference optimized for large language models without the overhead of containerized deployments.

How to eliminate wrong answers

Option A is wrong because Azure AI Search is a retrieval service for indexing and querying vector and keyword data, not a generative AI model deployment service; it lacks native model inference capabilities. Option C is wrong because Azure Machine Learning real-time endpoints are designed for custom ML model deployment, but they introduce higher latency due to container startup times and lack the optimized inference infrastructure (e.g., PTUs) that Azure OpenAI provides for generative models. Option D is wrong because Azure Functions is a serverless compute service for event-driven code execution, not a model hosting platform; it would require manual integration with a model endpoint and cannot guarantee the low latency needed for real-time generative AI inference.

163
MCQmedium

Refer to the exhibit. A developer received this response from an Azure OpenAI chat completion call. The prompt was "What is the capital of France?". The finish_reason is "stop". What does this indicate?

A.The response was truncated due to content filtering.
B.The model completed the response naturally.
C.The model stopped generating before the response was complete.
D.The response reached the max_tokens limit.
AnswerB

The finish_reason field reports why generation halted. A value of "stop" means the model emitted its natural end-of-sequence token, producing a complete answer rather than being cut off by the max_tokens limit or filtered by a content policy.

Why this answer

The finish_reason 'stop' indicates that the model completed the response naturally, meaning it generated a complete answer to the prompt and reached a logical stopping point (e.g., the end of a sentence or the end of the generated text). This is the standard behavior for a successful completion where the model did not encounter any content filter, token limit, or other interruption.

Exam trap

Microsoft often tests the distinction between finish_reason values, and the trap here is that candidates confuse 'stop' with 'length' or assume any non-error finish_reason means truncation, when in fact 'stop' explicitly signals a natural and complete generation.

How to eliminate wrong answers

Option A is wrong because 'stop' specifically means the model finished generating on its own, not that content filtering truncated the response; content filtering would return a finish_reason of 'content_filter'. Option C is wrong because 'stop' indicates the model completed the response, not that it stopped prematurely; a premature stop would be indicated by a finish_reason of 'length' (if max_tokens hit) or 'null' (if interrupted). Option D is wrong because reaching the max_tokens limit would result in a finish_reason of 'length', not 'stop'.

164
Multi-Selecteasy

Which TWO statements about Azure OpenAI Service content filters are true?

Select 2 answers
A.They can be configured with severity levels (low, medium, high)
B.They only filter the output of the model
C.They cannot be customized for specific use cases
D.They are bypassed when using PTU deployments
E.They include categories such as hate, sexual, violence, and self-harm
AnswersA, E

Severity levels allow granular control over filtering.

Why this answer

Azure OpenAI Service content filters can be configured with severity levels (low, medium, high) to control the strictness of filtering for each content category. This allows administrators to fine-tune the filter sensitivity based on their application's risk tolerance and compliance requirements.

Exam trap

The trap here is that candidates often assume content filters only apply to model outputs (Option B) or that PTU deployments offer a way to bypass safety controls (Option D), but Azure enforces filters uniformly across all deployment types.

165
Multi-Selectmedium

You are building an application that uses the Azure OpenAI Assistants API. The assistant must maintain conversational context across multiple user turns and use a code interpreter to analyze uploaded CSV files. Which two actions should you perform? (Choose two.)

Select 2 answers
A.Create a new assistant for each user turn to avoid context conflicts.
B.Create a thread and add user messages to it, then create a run that references the assistant and the thread.
C.Set the 'stream' parameter to true on every run to preserve conversation state between turns.
D.Use the completions endpoint with a manually maintained message array instead of the Assistants API.
E.Enable the code_interpreter tool on the assistant and upload the CSV files as files that the run can access.
AnswersB, E

The Assistants API persists conversation state in a thread. Adding messages to the thread and running the assistant against both the assistant ID and thread ID gives the model access to prior turns, satisfying the multi-turn context requirement without manually resending history in each call.

Why this answer

The Assistants API stores multi-turn context in threads, so messages are appended to a thread and runs execute against the assistant and thread. Code interpreter must be enabled as a tool, and CSV files must be uploaded and attached so the tool can read them during the run. Together these actions satisfy both the context and analysis requirements.

Exam trap

The trap here is treating streaming or per-turn assistant creation as a way to keep context, when conversation state is actually held in threads and tools must be explicitly enabled and given file access.

166
MCQhard

You are building a generative AI solution using Azure AI Foundry. The solution must meet compliance requirements that require all model inputs and outputs to be auditable for a minimum of one year. What should you enable?

A.Azure Monitor alerts for unusual activity.
B.Azure Monitor metrics for the Azure AI Foundry resource.
C.Azure Monitor workbooks to visualize usage.
D.Diagnostic settings to capture request and response logs and store them in a storage account.
AnswerD

Diagnostic settings stream request and response logs to a storage account, giving durable, queryable records that satisfy the one-year audit retention requirement. This captures actual model inputs and outputs rather than metrics alone, which cannot reconstruct individual interactions.

Why this answer

Enabling diagnostic settings for the Azure AI Foundry resource allows you to capture detailed request and response logs for all model interactions. By routing these logs to a storage account, you retain the data for the required one-year audit period, meeting compliance needs for full traceability of inputs and outputs.

Exam trap

The trap here is that candidates confuse monitoring features (alerts, metrics, workbooks) with data retention capabilities, assuming any Azure Monitor feature can satisfy audit requirements without understanding that only diagnostic settings provide the raw log capture needed for compliance.

How to eliminate wrong answers

Option A is wrong because Azure Monitor alerts are designed to notify on unusual activity or anomalies, not to provide long-term audit storage of model inputs and outputs. Option B is wrong because Azure Monitor metrics capture aggregated performance data like latency or request counts, not the detailed request/response payloads needed for auditing. Option C is wrong because Azure Monitor workbooks are visualization tools for metrics and logs, not a storage mechanism for raw audit data.

167
Multi-Selecthard

Which THREE factors should you consider when selecting a model for a generative AI solution on Azure?

Select 3 answers
A.Cost per token and deployment options.
B.Model capability and modality (text, code, image).
C.Latency and throughput requirements.
D.Number of transformer layers in the model.
E.Training data source and licensing.
AnswersA, B, C

Token pricing directly determines running cost at scale, while deployment options (serverless versus provisioned throughput) govern quota and capacity planning. Both are explicit selection factors for generative AI models on Azure, letting architects balance budget against the throughput the workload demands.

Why this answer

Option A is correct because cost per token and deployment options (such as pay-as-you-go versus provisioned throughput) directly affect the total cost and scalability of a generative AI solution on Azure. Option B is correct because the model's capability and modality determine whether it can handle the required task, such as text generation, code completion, or image creation. Option C is correct because latency and throughput requirements dictate whether the chosen model and deployment type can meet the application's performance and concurrency needs.

Option D is not a primary selection factor because the number of transformer layers is an internal architectural detail that influences capability but is not a decision criterion by itself. Option E is not a primary selection factor because training data source and licensing are legal and compliance considerations, not core factors for selecting a model for a generative AI solution on Azure.

Exam trap

The trap here is that candidates confuse internal model architecture (like transformer layers) with selection criteria, when in fact Azure abstracts those details and you only need to consider cost, capability, latency, and deployment options.

168
MCQeasy

You need to generate realistic synthetic data for training a machine learning model while ensuring the data does not contain personally identifiable information (PII). Which Azure service should you use?

A.Azure AI Search
B.Azure AI Document Intelligence
C.Azure OpenAI Service
D.Azure AI Language
AnswerC

Azure OpenAI Service can generate synthetic text or data through its language models, and its content filtering and data-handling controls help avoid reproducing PII. This satisfies the requirement to produce realistic training data without embedding personally identifiable information.

Why this answer

Azure OpenAI Service provides access to powerful generative AI models (e.g., GPT-4) that can create realistic synthetic data by learning patterns from training data. Crucially, these models can be configured to avoid memorizing or reproducing PII, and you can apply content filters and data masking to ensure the generated output is free of personally identifiable information.

Exam trap

The trap here is that candidates confuse Azure AI Language's text generation capabilities (e.g., summarization, question answering) with the full generative AI power of Azure OpenAI Service, but Azure AI Language does not offer the same level of flexible, high-fidelity synthetic data generation.

How to eliminate wrong answers

Option A is wrong because Azure AI Search is a search-as-a-service solution for indexing and querying data, not a generative AI service capable of creating synthetic data. Option B is wrong because Azure AI Document Intelligence (formerly Form Recognizer) is designed to extract structured information from documents (e.g., OCR, key-value pairs), not to generate new synthetic datasets. Option D is wrong because Azure AI Language provides pre-built and custom NLP capabilities (e.g., sentiment analysis, entity recognition) but does not include generative models for creating realistic synthetic data from scratch.

169
MCQhard

A financial services firm wants to use Azure OpenAI to generate investment advice summaries. They must ensure that the model does not produce any advice that could be interpreted as personalized financial advice. What is the most effective strategy?

A.Set temperature to 0 and top_p to 0 to make outputs deterministic.
B.Use a system message that instructs the model to avoid personalized advice and apply strict content filtering.
C.Provide few-shot examples of disclaimers in the prompt.
D.Fine-tune the model on a dataset of generic financial summaries.
AnswerB

A system message sets persistent behavioural boundaries, instructing the model to decline personalised financial advice, while content filtering blocks prohibited outputs. Together they satisfy the stem's constraint that no output be interpretable as personalised investment advice.

Why this answer

Azure OpenAI's system messages allow you to set the model's behavior and constraints at the conversation level, which is the most direct and effective way to enforce a policy like avoiding personalized financial advice. Combined with Azure's content filtering (which can block harmful or restricted content), this approach provides both instruction-based and filter-based guardrails without requiring model retraining or relying solely on example-based prompting.

Exam trap

The trap here is that candidates often assume deterministic parameters (temperature=0, top_p=0) guarantee safe outputs, but they only control randomness, not content compliance—Azure's system message and content filtering are the correct tools for enforcing content policies.

How to eliminate wrong answers

Option A is wrong because setting temperature to 0 and top_p to 0 makes outputs deterministic but does not prevent the model from generating personalized financial advice; it only reduces randomness, not content compliance. Option C is wrong because few-shot examples of disclaimers in the prompt can be ignored or overridden by the model if the underlying training data biases it toward personalized responses; system messages have higher priority in the instruction hierarchy. Option D is wrong because fine-tuning on generic financial summaries requires significant labeled data and compute, and it does not guarantee the model will avoid personalized advice—it may still generate such content if the fine-tuning dataset is not carefully curated to exclude it.

170
MCQmedium

You are configuring an Azure OpenAI deployment for a generative AI solution that summarizes long legal contracts. Users report that summaries sometimes omit clauses near the end of documents. The documents are up to 120 pages. You need to improve completeness without changing the model. What should you do?

A.Increase the max_tokens parameter of the completion request to its maximum value.
B.Enable streaming responses so the client receives partial summaries as they are generated.
C.Split each document into overlapping chunks, summarize each chunk, and then combine the summaries in a final aggregation step.
D.Raise the temperature so the model explores more of the document content.
AnswerC

Chunking with overlap ensures content near boundaries is not lost, and a map-reduce style aggregation summarizes each part before combining them. This keeps each request within the context window while covering the entire document, which directly addresses omitted clauses near the end without changing the model.

Why this answer

Long documents can exceed the model context window, causing later sections to be dropped. Chunking with overlap preserves boundary content, and summarizing chunks before aggregating covers the whole document. Output length, temperature, and streaming do not affect how much input text the model can attend to, so they cannot fix missing clauses.

Exam trap

The trap here is assuming that a larger output token limit lets the model consider more of the input document.

171
MCQmedium

A legal team wants an assistant that drafts contract summaries. Their policy requires that every generated summary include traceable references to the exact clauses used and that reviewers be able to see which source passages informed each statement. You are using Azure OpenAI with your own document index. Which approach best meets the traceability requirement?

A.Fine-tune the base model on previously approved contract summaries so its style matches legal expectations.
B.Enable the model's logprobs parameter and expose the token probabilities to reviewers as evidence of reliability.
C.Increase the model deployment's tokens-per-minute quota so longer contracts fit in a single prompt.
D.Use the On Your Data pattern with Azure AI Search, return document chunks with their IDs and titles, and instruct the model to cite the retrieved chunk identifiers in its output.
AnswerD

Grounding the completion in Azure AI Search chunks and returning citations lets the model reference the exact retrieved passages, and the service can return the citation metadata alongside the response. Reviewers can then map each statement to a chunk and its source clause, which directly satisfies the traceability policy without retraining.

Why this answer

Traceability requires linking generated text back to specific source passages. Using Azure AI Search as the grounding source with the On Your Data pattern returns citation metadata for the retrieved chunks, and prompting the model to cite those chunks produces summaries where each statement can be traced to a clause. Quota, fine-tuning, and logprobs do not create that link.

Exam trap

The trap here is confusing model confidence signals such as logprobs with source attribution, when only retrieved citation metadata can identify the clause behind a statement.

172
MCQhard

You are designing a generative AI solution that uses Azure OpenAI Service. The solution must generate code snippets in Python and JavaScript. You need to ensure the model reliably outputs code in the correct language based on user input. Which approach should you use?

A.Set the top_p parameter to a low value.
B.Use a system message to specify the desired language.
C.Fine-tune the model on a dataset of code in both languages.
D.Set the temperature to 0 to make the model deterministic.
AnswerB

A system message sets persistent behavioural instructions applied to every turn, so it constrains the model to emit Python or JavaScript according to the user's request. This satisfies the reliability requirement better than per-prompt phrasing, which the model may inconsistently honour across requests.

Why this answer

System messages in Azure OpenAI Service allow you to set the context or behavior of the model, such as specifying the desired programming language for code generation. This approach is lightweight, requires no retraining, and reliably guides the model to output code in the correct language based on the user's request, leveraging the model's existing training on both Python and JavaScript.

Exam trap

The trap here is that candidates often confuse hyperparameters like temperature and top_p with content control mechanisms, mistakenly believing they can enforce output language, when in fact they only affect randomness and token selection probability.

How to eliminate wrong answers

Option A is wrong because setting top_p to a low value reduces the pool of tokens considered for sampling, which can make outputs more focused but does not control the language of the generated code; it is a nucleus sampling parameter, not a language selector. Option C is wrong because fine-tuning the model on a dataset of code in both languages is overkill for this requirement, as the base model already understands both languages; fine-tuning is typically used for specialized tasks or to adapt to a specific domain, not for simple language switching. Option D is wrong because setting temperature to 0 makes the model deterministic by always choosing the most likely token, but it does not enforce the output language; it can still produce code in the wrong language if the prompt is ambiguous, and it reduces creativity but does not guarantee language adherence.

173
Multi-Selecthard

You are building a generative AI assistant on Azure OpenAI Service that must invoke backend business functions, such as checking order status and issuing refunds, in response to natural-language requests. You want the model to decide when a function is needed and to supply structured arguments, while your application retains control over execution. Which two actions should you take? (Choose two.)

Select 2 answers
A.Define the available functions and their parameters in the tools parameter of the chat completions request.
B.Execute the returned function call in your application, then send the function result back to the model in a subsequent request so it can compose the final answer.
C.Fine-tune the base model on historical order and refund conversations so it learns to call the correct endpoints.
D.Grant the Azure OpenAI resource a managed identity with permission to call the backend APIs directly on behalf of the model.
E.Set the temperature parameter to 0 and increase the max_tokens value to guarantee deterministic function selection.
AnswersA, B

Supplying function definitions through the tools parameter is how you tell the model which operations exist and what arguments each expects. The model then returns a structured tool call naming the function and a JSON argument payload when it judges a function is needed, rather than fabricating an answer. This is the mechanism that lets the model choose actions based on the conversation while your code decides whether and how to run them.

Why this answer

Function calling works in two phases: you declare callable operations and their parameter schemas in the request, and the model returns a structured call when a function is appropriate. Your application then executes that function and returns the output so the model can finish the response. Both declaring the tools and executing plus returning results are required for a working implementation.

Exam trap

The trap here is believing the model itself executes backend functions or that fine-tuning teaches it to call endpoints, when in reality the model only proposes calls that your application must run and report back.

174
MCQmedium

You are building a generative AI assistant that must summarize long technical manuals stored in Azure Blob Storage. The manuals are often 300 pages, and the model must produce a concise summary with references to page numbers. Which approach should you use?

A.Use Azure AI Document Intelligence to extract text, then send each page as a separate request to Azure OpenAI and concatenate the summaries.
B.Use Azure OpenAI fine-tuning to train a custom model on the manuals, then ask the model to summarize any manual.
C.Use Azure OpenAI GPT-4o with a single prompt containing the entire manual text.
D.Use Azure AI Search to index the manuals with page metadata, then use a retrieval-augmented generation (RAG) pattern with Azure OpenAI to generate summaries and citations.
AnswerD

Indexing the manuals in Azure AI Search with page-level metadata enables retrieval of relevant chunks and supports citation of page numbers. The RAG pattern lets Azure OpenAI generate a summary grounded in retrieved content, avoiding context-window limits. This is the recommended approach for long documents requiring references.

Why this answer

Long documents exceed model context windows, so retrieval-augmented generation is needed. Azure AI Search indexes content with metadata, allowing Azure OpenAI to generate grounded summaries and cite page numbers. This combination handles length, improves accuracy, and provides traceability, which is essential for technical manuals.

Exam trap

The trap here is assuming that a large-context model can ingest an entire long manual in one prompt, ignoring token limits and the need for verifiable page citations.

← PreviousPage 3 of 3 · 174 questions total

Ready to test yourself?

Try a timed practice session using only Implement generative AI solutions questions.