Courseiva

AI-102 Implement generative AI solutions Practice Question

You are deploying a retrieval-augmented generation chat solution on Azure OpenAI. During testing, users report that the model confidently answers questions using facts that are not in the indexed documents, and that retrieved chunks sometimes come from the wrong department's SharePoint site. You need to reduce these ungrounded answers without retraining the model. (Choose two.)

⚠ Common exam trap

The trap here is assuming that lowering or raising sampling parameters such as temperature or max_tokens will fix hallucinations, when grounding depends on prompt constraints and retrieval quality instead.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set the message role "system" content to instruct the model to answer only from the provided context and to state when the answer is not found.

Ungrounded answers in a RAG pattern usually come from two sources: the model not being constrained to the retrieved context, and the retrieval step returning irrelevant or unauthorized chunks. A system message that enforces answer-only-from-context behavior, combined with metadata filtering and security trimming at the search layer, addresses both causes without retraining. Length and randomness parameters, plus harmful-content filters, do not influence factual grounding.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set the temperature parameter to 2 to make the model more factual.

    Why it's wrong here

    Temperature controls randomness of token sampling; raising it to 2 increases diversity and unpredictability, which typically makes ungrounded or invented text more likely, not less. Lowering temperature can make output more deterministic but does not by itself force the model to stay within retrieved documents, so this value works against the stated goal of reducing fabricated answers.

  • ✗

    Increase the max_tokens parameter so the model has more room to explain its reasoning before answering.

    Why it's wrong here

    max_tokens only caps the length of the generated completion. Allowing a longer response does not supply missing facts and can even give the model more space to elaborate on an incorrect premise. It does nothing to improve retrieval relevance or to instruct the model to stay grounded, so it fails to resolve the reported behavior.

  • ✓

    Set the message role "system" content to instruct the model to answer only from the provided context and to state when the answer is not found.

    Why this is correct

    A system message sets behavioral boundaries for the whole chat session, so telling the model to restrict itself to supplied context and to admit when it cannot find an answer directly reduces fabricated responses. This is the standard, non-retraining grounding control used with Azure OpenAI chat completions, and it works alongside retrieval quality fixes rather than replacing them.

  • ✗

    Enable content filtering at the highest severity threshold to block ungrounded statements.

    Why it's wrong here

    Azure OpenAI content filters classify harmful categories such as hate, violence, and self-harm; they do not detect whether a claim is supported by retrieved documents. An ungrounded but benign factual statement passes content filtering easily, so raising severity thresholds would not stop fabricated answers and could block legitimate responses instead.

  • ✓

    Apply security trimming and metadata filters so retrieval only returns chunks the user is allowed to see and that match the requested department.

    Why this is correct

    When retrieved chunks come from the wrong department, the grounding data itself is wrong, so the model answers from irrelevant content. Filtering by document metadata and enforcing per-user permissions in the search index restricts context to authorized, relevant chunks, which directly addresses the cross-department retrieval symptom and improves answer fidelity without any fine-tuning.

About these practice questions

This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.