Oracle · Free Practice Questions · Last reviewed May 2026
48real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
12% of exam · 6 sample questions below
A developer is building a RAG application using Oracle Cloud Infrastructure (OCI) Document Understanding and OCI Generative AI. After chunking documents and generating embeddings, the developer observes that the retrieval step often returns chunks that are semantically unrelated to the query. Which action is MOST likely to improve retrieval relevance?
Switch from a dense embedding model to a sparse embedding model.
Adjust the chunk size and chunk overlap to better capture coherent passages.
Chunk size and overlap determine whether each embedding captures a coherent passage; poorly sized chunks split context, producing vectors that match unrelated queries. Tuning these parameters improves semantic alignment between chunks and queries, addressing the retrieval relevance problem.
Increase the chunk size to capture more context.
Reduce the number of retrieved chunks (k) in the vector search.
An organization stores its knowledge base in Oracle Autonomous Database and wants to build a RAG chatbot using OCI Generative AI. The chatbot must retrieve the most relevant documents based on user queries. Which indexing approach is BEST suited for efficient similarity search on text embeddings?
Create an ANN index on the embedding vector column.
An approximate nearest neighbour index on the embedding vector column enables fast similarity search across large document sets, avoiding exhaustive comparisons. This satisfies the requirement for efficient retrieval of the most relevant documents by user query.
Create a bitmap index on the embedding vector column.
Create an inverted index on the document text column.
Create a B-tree index on the document text column.
A data scientist is building a RAG application that processes PDF invoices. The extraction step uses OCI Document Understanding to convert PDFs to text. The scientist then splits the text into chunks and generates embeddings using OCI Generative AI. However, the retrieval often misses critical fields like invoice numbers and dates. Which preprocessing step would MOST likely improve retrieval of these specific fields?
Increase the chunk size to include entire invoices.
Apply stemming and lemmatization to the text before chunking.
Tag each chunk with metadata such as invoice number, date, and vendor, and use metadata filtering during retrieval.
Invoice numbers and dates are structured fields that semantic similarity alone retrieves poorly, since chunk embeddings blur exact tokens. Tagging each chunk with extracted metadata and applying metadata filters during retrieval lets the query constrain candidates to matching invoice numbers, dates or vendors, directly satisfying the need to surface these critical fields.
Switch from dense embeddings to sparse embeddings for better exact match.
A team is deploying a RAG system that uses OCI Generative AI to answer questions about internal HR policies. The system must comply with data residency requirements: all data processing must stay within a specific OCI region. The team uses OCI Data Science for orchestration. Which architecture BEST meets the data residency requirement?
Deploy the generative AI model endpoints within the same OCI region as the data and compute.
Data residency requires that processing, not merely storage, stays inside the specified region. Hosting the OCI Generative AI model endpoints in the same OCI region as the data and OCI Data Science compute ensures prompts, embeddings and outputs never cross regional boundaries, satisfying the stem's residency constraint.
Use OCI Generative AI endpoints in a different region but store data in the required region.
Use an external third-party LLM endpoint that guarantees data residency.
Store embeddings in a different region but run inference in the required region.
A developer notices that the RAG system returns irrelevant chunks when the user query contains typos or abbreviations. Which technique would BEST improve retrieval robustness for such queries?
Decrease the chunk size to focus on smaller units.
Increase the number of retrieved chunks to cover more variations.
Use a spell-checker on the retrieved chunks.
Implement query rewriting or expansion using a language model before embedding.
Typos and abbreviations produce embeddings that drift from the indexed chunks, so similarity search misses relevant passages. Rewriting or expanding the query with a language model normalises spelling and resolves abbreviations into canonical terms before embedding, restoring lexical and semantic overlap with the corpus.
Which TWO are best practices for building a RAG application on OCI? (Choose two.)
Use a vector database such as OCI OpenSearch with ANN indexes for storing embeddings.
Storing embeddings in OCI OpenSearch with approximate nearest neighbour indexes makes similarity search scale to large corpora, returning nearest chunks in sub-linear time rather than exhaustive comparison. This is the standard retrieval backbone for OCI RAG pipelines, keeping latency acceptable as document volume grows.
Generate embeddings for documents at query time to ensure freshness.
Pre-index the documents and update the index periodically to reflect new content.
Pre-indexing documents and refreshing the index periodically keeps retrieval aligned with current content without rebuilding embeddings per query. Stale indexes return outdated chunks, so scheduled updates ensure newly added or revised documents remain retrievable, which is a core operational practise for production RAG on OCI.
Store the source documents only in OCI Object Storage and retrieve them at query time using full-text search.
Use a different embedding model for documents and queries to capture distinct semantics.
Want more Building LLM Applications with RAG and Vector Search practice?
Practice this domain12% of exam · 6 sample questions below
A company is deploying a large language model for a customer service chatbot. The model needs to understand industry-specific jargon and maintain low latency. Which approach best balances these requirements?
Employ retrieval-augmented generation (RAG) with a general model
Rely solely on prompt engineering with a general model
Use a large general-purpose LLM with zero-shot prompting
Fine-tune a small open-source LLM on domain-specific data
Fine-tuning a small open-source model on domain data teaches industry jargon through weight updates while keeping inference cheap and fast. A smaller model serves responses with lower latency than a large general-purpose LLM, satisfying both the terminology and latency constraints.
A data scientist observes that their fine-tuned LLM performs well on training data but generates repetitive and dull responses in production. What is the most likely cause and best solution?
The model is overfitted; apply stronger regularization
The temperature is set too low; increase temperature during inference
Low temperature sharpens the softmax distribution, so the model repeatedly selects high-probability tokens, producing repetitive, dull output. Raising temperature during inference flattens that distribution, increasing token diversity. This directly addresses the production-only symptom: the model generalises fine, so the fault lies in decoding, not training.
The training data lacks diversity; add more varied examples
The model has too many layers; reduce model size
An organization wants to use an LLM to summarize legal documents. Which consideration is most important for ensuring accurate summaries?
Fine-tune the model on a curated legal corpus
Fine-tuning on a curated legal corpus adapts the model's weights to domain terminology, citation conventions and reasoning patterns, improving factual fidelity for legal text. This satisfies the accuracy requirement more directly than prompt engineering alone, which cannot instil specialised legal knowledge.
Use the largest available general-purpose model
Rely on zero-shot summarization with careful prompting
Pre-train a new model from scratch on legal texts
A developer is building a code generation assistant. The model occasionally produces syntactically correct but semantically wrong code. Which technique directly addresses semantic correctness?
Expand the token vocabulary
Lower the temperature to 0
Apply RLHF using human-validated code examples
RLHF with human-validated code examples optimises the model against human judgements of correctness, not just syntax, directly targeting the semantic errors described. This addresses the stem's requirement that generated code be semantically, not merely syntactically, correct.
Increase beam search width
A company fine-tunes an LLM on internal support tickets. After deployment, the model hallucinates company-specific product names. What is the most effective mitigation?
Switch to a smaller model to reduce hallucination risk
Use prompt engineering to remind the model to be accurate
Implement RAG with a verified product database
RAG retrieves verified product names from the database at inference time and injects them into the prompt, grounding generation in authoritative facts rather than relying on weights that may have absorbed incorrect ticket text. This directly prevents hallucinated product names without retraining the fine-tuned model.
Fine-tune further with more ticket data
A team wants to evaluate an LLM's performance on a text classification task. Which metric is most appropriate for a balanced dataset?
BLEU score
Perplexity
Accuracy
Accuracy divides correct predictions by total predictions, giving a directly interpretable score. On a balanced dataset each class contributes equally, so accuracy is not skewed by majority-class dominance, satisfying the stem's balanced-data constraint without needing precision, recall or F1.
ROUGE score
Want more Fundamentals of Large Language Models practice?
Practice this domainA data scientist is designing a prompt to generate a structured report with sections for Summary, Findings, and Recommendations. Which output format specification in the prompt would be MOST effective?
"Write a report with three sections: Summary, Findings, Recommendations."
"Provide the output in JSON format with keys: 'summary', 'findings', and 'recommendations'."
JSON enforces explicit key-value pairs, so the model must populate 'summary', 'findings', and 'recommendations' as discrete fields rather than free prose. This directly satisfies the stem's requirement for a structured report with three defined sections, giving deterministic, machine-parseable output that plain text or Markdown headings cannot guarantee.
"Return the report in bullet points."
"Output the report as a markdown document with headings."
A developer notices that a Cohere Command model occasionally generates contradictory statements in the same response when asked to reason step-by-step. Which technique is designed to address inconsistency by generating multiple reasoning paths and selecting the most consistent answer?
Few-shot prompting
Temperature tuning
Tree-of-thought prompting
Self-consistency
Self-consistency samples multiple diverse reasoning chains from the model and aggregates their final answers, typically by majority vote, selecting the most frequently reached conclusion. This reduces contradictions arising from any single flawed chain, directly addressing the inconsistent step-by-step reasoning described.
Which parameter controls the creativity and randomness of a model's output by adjusting the probability distribution before sampling the next token?
Frequency penalty
Max tokens
Temperature
Temperature scales the logits before the softmax, flattening or sharpening the probability distribution prior to sampling the next token. Lower values make output more deterministic; higher values increase randomness and creativity, directly controlling the diversity of generated text.
Top-k
An engineer is using the ReAct pattern to build a reasoning agent. The agent should first reason about the user query, then call an external API, and finally incorporate the API result into a final answer. Which prompt structure best implements this pattern?
"You have access to an API. For each step, output 'Thought:', then 'Action:', then 'Observation:' before the final answer."
The ReAct pattern interleaves reasoning and acting, so the prompt must expose explicit Thought, Action and Observation steps. This structure lets the agent reason, invoke the external API, then fold the returned observation into its final answer, matching the stem's required sequence.
"You are a helpful assistant. Answer the user's question based on your knowledge."
"First, call the API. Then, output the result."
"Think step by step, then provide the final answer."
A prompt engineer is tasked with reducing hallucinations in a document-grounded generation task using Cohere Command R. Which system prompt component is MOST effective for enforcing that the model only uses provided documents?
"You are an expert in the field. Use your extensive knowledge to answer."
"Provide a detailed answer with references."
"Use a temperature of 0.0 for factual answers."
"Answer the question based solely on the provided document. If the document does not contain the answer, say 'I don't know'."
An explicit grounding constraint plus a refusal fallback directly suppresses pre-trained knowledge leakage: instructing the model to answer solely from the supplied document and to return "I don't know" when the answer is absent removes the incentive to hallucinate, satisfying the stem's document-grounded requirement.
In few-shot prompting, what is the primary purpose of including examples in the prompt?
To reduce the need for a system prompt
To provide a template for the desired output format and reasoning pattern
Including examples demonstrates the expected output structure and reasoning style, letting the model infer the task pattern from context rather than explicit instructions. This satisfies few-shot prompting's core constraint: guiding behaviour through in-context demonstration, which is precisely how the technique conditions responses without weight updates.
To increase the model's vocabulary
To decrease the computational cost of inference
Want more Prompt Engineering practice?
Practice this domain12% of exam · 6 sample questions below
A company is deploying a generative AI service on OCI using the OCI Data Science service with a large language model (LLM) in a VCN. The model inference endpoint must be accessible only from a private subnet within the same VCN. Which networking component should be configured to enable this?
NAT Gateway
Dynamic Routing Gateway (DRG)
Internet Gateway
Service Gateway
Service gateway enables private subnet access to OCI services like Data Science.
A company is using OCI Generative AI service with a dedicated AI cluster for text generation. They notice that the latency is higher than expected. The cluster is in the Ashburn region, and users are distributed globally. What is the most effective way to reduce latency?
Enable the OCI Generative AI inference optimizer
Deploy dedicated AI clusters in regions closer to the users
Latency is dominated by physical distance between users and the inference endpoint. Deploying dedicated AI clusters in regions nearer those users shortens network round-trip time, which is more effective than tuning the model or increasing cluster capacity in Ashburn.
Increase the number of nodes in the dedicated AI cluster
Use a content delivery network (CDN) to cache responses
A developer wants to call the OCI Generative AI service from a Python application running on an OCI Compute instance. Which method is the most secure for authenticating the API calls?
Use a resource principal
Use the OCI CLI with a config file containing credentials
Use instance principals with a dynamic group and policy
Instance principals let the compute instance authenticate through its dynamic group membership, so no API keys or credentials are stored on the instance. A policy then authorises that dynamic group to call the Generative AI service, eliminating long-lived secrets that could be exfiltrated.
Use an API signing key stored on the instance
Which THREE components are required to deploy a custom generative AI model on OCI Data Science model deployment?
A load balancer to distribute traffic
An inference script (e.g., score.py) to handle prediction requests
OCI Data Science model deployment requires an inference script defining the predict function that loads the model and transforms incoming request payloads into responses. Without it, the deployment runtime cannot serve prediction requests, so the endpoint fails to initialise.
A model artifact containing the model files
The model artifact packages the serialised model files, weights and dependencies that the deployment loads at runtime. OCI Data Science cannot instantiate the model without this artifact, making it a mandatory component alongside the inference script and deployment configuration.
An API signing key for authentication
A deployment configuration specifying resources and environment
The deployment configuration specifies the compute shape, instance count, bandwidth and environment variables the endpoint uses. OCI Data Science needs these resource parameters to provision and scale the inference container, so deployment cannot proceed without it.
Which TWO are valid methods to monitor the performance of a generative AI model deployed on OCI Data Science?
Use OCI Notifications to receive alerts on model drift
Use OCI Monitoring service to track custom metrics like latency and throughput
OCI Monitoring ingests custom metrics emitted by the model endpoint, letting you chart latency and throughput against thresholds and trigger alarms. This directly satisfies the stem's requirement to monitor deployed generative AI performance continuously rather than only reviewing logs after the fact.
Use OCI Logging service to collect inference logs
OCI Logging captures the inference request and response records produced by the deployed model, giving an auditable trail for diagnosing errors and usage patterns. It complements metric-based monitoring by satisfying the stem's need to observe individual predictions rather than only aggregate performance figures.
Use OCI Events service to trigger retraining on low accuracy
Use OCI Audit service to review API call logs
An administrator runs the above CLI command to check the status of a dedicated AI cluster. The cluster is ACTIVE with capacity 10. However, a user reports that inference requests to this cluster are failing with a '429 Too Many Requests' error. What is the most likely cause?
The cluster is hitting the maximum inference requests per minute limit
A 429 response is a throttling signal, not a capacity or health failure. Since the cluster is ACTIVE with capacity 10, the requests are exceeding the configured inference requests per minute ceiling. Raising that limit, or reducing call frequency, resolves the failures.
The cluster does not have enough nodes to handle the load
The user is not in the same compartment as the cluster
The cluster is not in ACTIVE state
Want more Deploying and Managing Generative AI on OCI practice?
Practice this domain13% of exam · 6 sample questions below
You need to convert a set of customer support tickets into vector embeddings for a similarity search application. Which OCI Generative AI model should you use?
Cohere Rerank
Cohere Embed (e.g., embed-english-v3.0)
Cohere Embed models are purpose-built text embedding models on OCI Generative AI, converting support tickets into dense vectors for similarity search. This satisfies the stem's constraint of generating vector embeddings, unlike generation models such as Cohere Command or Llama, which produce text rather than embeddings.
Meta Llama 3
Cohere Command R
A developer is using the OCI Generative AI Chat API with Cohere Command R+ to build a multi-turn conversational agent. They want the agent to always respond in a formal tone, regardless of the user's phrasing. Which parameter should they set in the API request to achieve this consistently?
Set the 'stop_sequences' parameter to include periods
Set the 'temperature' parameter to 0.1
Use the 'system' parameter (or preamble_override) to provide a system message like 'You are a formal assistant'
The system parameter (or preamble_override) injects a persistent instruction before the conversation, shaping tone across all turns. Setting it to 'You are a formal assistant' enforces formal responses regardless of user phrasing, satisfying the consistent-tone requirement.
Set the 'max_tokens' parameter to 200
A company uses OCI Generative AI Agents to build a RAG application that answers questions from documents stored in OCI Object Storage. The knowledge base is updated daily. Which step is necessary to ensure the agent incorporates the latest documents?
Re-sync the knowledge base after updating the documents in Object Storage
The knowledge base indexes a snapshot of the Object Storage data, so newly added or changed documents are invisible until ingestion runs again. Re-syncing after each daily update refreshes the index, ensuring the agent retrieves the latest documents.
Use the Embedding API to manually index each new document
Set the session API to refresh automatically
Recreate the agent each time documents change
You want to test different prompts and parameters (temperature, max tokens) for a summarization task using a foundation model without writing any code. Which OCI tool should you use?
OCI Console -> Generative AI -> Models
OCI Generative AI Playground
The OCI Generative AI Playground provides a no-code console interface for experimenting with prompts and parameters such as temperature and max tokens against foundation models. This satisfies the requirement to test summarisation prompts without writing code.
OCI CLI with the 'oci generative-ai' commands
Python SDK with InferenceClient
A team is fine-tuning a Llama 3 model using OCI Generative AI. The training dataset contains 10,000 prompt-completion pairs in JSONL format. After submitting the fine-tuning job, it fails with a 'Data validation error'. What is the most likely cause?
The base model selected does not support fine-tuning
The dataset is in the wrong compartment
The fine-tuning job does not have enough model units allocated
The JSONL file uses 'input' and 'output' as key names instead of 'prompt' and 'completion'
OCI Generative AI fine-tuning expects JSONL records keyed as 'prompt' and 'completion'. Supplying 'input' and 'output' instead means the required fields are absent, so schema validation rejects the dataset before training begins, producing the reported data validation error.
An application needs to generate embeddings for text in multiple languages (English, Spanish, French). Which OCI Generative AI embedding model should be used?
embed-english-v3.0
Cohere Command R
embed-multilingual-v3.0
embed-multilingual-v3.0 maps English, Spanish and French text into one shared vector space, so semantically equivalent phrases across languages land close together. That cross-lingual alignment satisfies the stem's multilingual requirement, which a monolingual English-only embedding model cannot deliver regardless of dimension or context length.
Meta Llama 3
Want more OCI Generative AI Service practice?
Practice this domainWhat is the primary purpose of the self-attention mechanism in a Transformer model?
To generate token embeddings in parallel
To reduce the dimensionality of token embeddings
To encode positional information of tokens
To compute a weighted sum of all token representations based on pairwise relevance
Self-attention projects each token into query, key and value vectors, then computes pairwise relevance scores whose softmax weights produce a weighted sum of all token representations, giving every position direct contextual access regardless of distance.
Which of the following best describes the difference between an encoder-only model (e.g., BERT) and a decoder-only model (e.g., GPT)?
Encoder-only uses bidirectional attention and is suited for classification or NER; decoder-only uses causal attention and is suited for text generation
Encoder-only architectures apply bidirectional self-attention, letting each token attend to all others, which suits classification and NER. Decoder-only architectures apply causal masking so tokens attend only to prior tokens, enabling autoregressive text generation. This attention-direction difference is the defining architectural distinction.
Encoder-only is trained for text generation; decoder-only is trained for classification
Both use the same attention pattern but differ in number of layers
Encoder-only uses causal attention; decoder-only uses bidirectional attention
A practitioner wants to evaluate an LLM-generated summary against a human-written reference using a metric that focuses on recall of key information. Which metric is most appropriate?
BLEU
Perplexity
Cosine similarity
ROUGE
ROUGE measures n-gram overlap between generated and reference summaries, with recall-oriented variants such as ROUGE-L and ROUGE-N counting how much reference content the candidate captures. This directly satisfies the stem's requirement to focus on recall of key information against a human-written reference.
A company needs to generate embeddings for a large corpus of legal documents to enable semantic search. Which type of model should they use?
An encoder-only embedding model like Cohere Embed
Encoder-only models map text into dense vector embeddings, which is exactly what semantic search requires for similarity comparison. Cohere Embed produces these vectors without generating text, matching the legal corpus embedding task rather than a generation workload.
A decoder-only generation model like GPT
A text-to-speech model
A machine translation model
Which of the following sampling strategies selects tokens based on a cumulative probability threshold from the highest probability tokens?
Top-p (nucleus) sampling
Top-p sampling sorts tokens by probability, sums them cumulatively, and retains the smallest set whose cumulative probability reaches the threshold p; sampling then occurs within that nucleus. This directly matches the stem's cumulative probability threshold criterion, unlike top-k's fixed token count.
Top-k sampling
Greedy decoding
Temperature sampling
An OCI Generative AI practitioner observes that a Cohere Command model generates responses with outdated information about a recent event. The model was fine-tuned six months ago. Which technique should be applied to incorporate new knowledge without retraining the model?
Use a longer context window and include all new articles in the prompt
Fine-tune the model again with the new data
Implement a RAG pipeline that indexes the latest documents into a vector store and retrieves relevant passages at query time
Retrieval-augmented generation retrieves relevant passages from a vector store of freshly indexed documents and supplies them as context at query time. This injects current knowledge without retraining or fine-tuning, satisfying the requirement to avoid model retraining.
Increase the temperature parameter to encourage more creative outputs
Want more LLM Fundamentals practice?
Practice this domain13% of exam · 6 sample questions below
In LangChain, which component is responsible for connecting a language model to a retriever and a prompt template to answer questions based on retrieved documents?
RetrievalQA chain
The RetrievalQA chain wires a retriever and a prompt template to a language model, fetching relevant documents and feeding them into the prompt for grounded answers. This satisfies the stem's requirement to answer questions from retrieved documents.
LLMChain
SequentialChain
AgentExecutor
A developer is building a RAG pipeline using LangChain and Oracle AI Vector Search. After loading and splitting PDF documents, they generate embeddings and store them in Oracle Database using OracleVS. Which method should they call on the vector store object to create a retriever that uses similarity search with a configurable number of results?
as_retriever()
Calling `as_retriever()` on the OracleVS vector store returns a VectorStoreRetriever, which performs similarity search against Oracle AI Vector Search. Passing `search_kwargs={"k": n}` satisfies the stem's requirement for a configurable number of results, unlike direct `similarity_search()` calls that return a fixed count.
from_texts()
max_marginal_relevance_search()
similarity_search()
A company uses LangChain with OCI Generative AI. They notice that their agent-based application occasionally exceeds the rate limits of the OCI Generative AI service, causing errors. Which strategy is MOST effective for handling rate limits in a production LangChain application?
Implement a retry mechanism with exponential backoff when calling the model
Exponential backoff retries the model call after progressively longer delays, spreading requests so the application recovers from OCI Generative AI throttling without hammering the service. This directly addresses the intermittent rate-limit errors described, making it the most effective production strategy.
Increase the k value in the retriever to reduce the number of API calls
Switch to a smaller model to reduce token consumption
Reduce the chunk_size parameter in text splitters
Which LangChain memory type stores the entire conversation history as a list of messages and is best for simple, short conversations?
ConversationTokenBufferMemory
ConversationSummaryMemory
ConversationBufferMemory
ConversationBufferMemory retains every user and AI message verbatim in an in-memory list, passing the complete history to the model each turn. This unbounded retention suits short, simple conversations where full context matters more than token efficiency.
ConversationBufferWindowMemory
A developer wants to index a large corpus of HTML web pages for a RAG pipeline using LangChain. They need to load the content from URLs, split the text into chunks, and generate embeddings. Which combination of LangChain components should they use?
PDFLoader, RecursiveCharacterTextSplitter, OCIGenAIEmbeddings
WebBaseLoader, RecursiveCharacterTextSplitter, OCIGenAIEmbeddings
WebBaseLoader fetches and parses HTML from URLs, satisfying the requirement to load content directly from web pages. RecursiveCharacterTextSplitter then chunks the extracted text, and OCIGenAIEmbeddings generates the vectors for the RAG pipeline, completing the load-split-embed sequence the scenario demands.
CSVLoader, RecursiveCharacterTextSplitter, OCIGenAIEmbeddings
WebBaseLoader, TokenTextSplitter, OCIGenAI
In LangChain, what is the purpose of the LCEL (LangChain Expression Language) | operator?
To pipe the output of one component into the next component
LCEL's pipe operator composes runnables into a sequence, feeding each component's output as the next component's input. This creates a declarative chain supporting streaming, batching and async execution, so prompt templates, models and parsers connect without custom glue code.
To run two chains in parallel
To concatenate two prompt templates
To compare outputs of two different models
Want more LangChain and AI Application Development practice?
Practice this domain12% of exam · 6 sample questions below
A data scientist is using OCI Generative AI Service to generate product descriptions. They notice that the output often repeats phrases. Which parameter adjustment would MOST directly address this issue?
Increase the temperature
Increase the max tokens
Increase the frequency penalty
Frequency penalty scales down the logits of tokens already generated in the output, directly discouraging the model from repeating the same phrases. Raising it satisfies the stem's repetition constraint more directly than temperature or top-p, which affect randomness rather than penalising prior token reuse.
Decrease the top-p value
A company needs to integrate OCI Generative AI Service with an existing application that uses OCI IAM for authentication. They want to use resource principal to allow the application to call the service without storing API keys. Which step is REQUIRED?
Create an OCI API key for the application
Enable the Generative AI Service for resource principal in the tenancy
Assign the application to a group with admin privileges
Create a dynamic group and a policy granting access to the Generative AI Service
Resource principal authentication requires the calling application to be a member of a dynamic group, matched by instance or resource rules. A policy then grants that dynamic group permission to call the Generative AI Service, eliminating stored API keys.
Refer to the exhibit. A user runs the command shown and receives the error: 'ServiceError: NotAuthorizedOrNotFound'. What is the MOST likely cause?
The CLI is not configured with OCI credentials
The user does not have the 'inspect' permission on the model
NotAuthorizedOrNotFound is returned when the caller lacks the inspect permission on the model resource; OCI masks authorisation failures as not-found, so the missing inspect verb on the model is the most likely cause of the error.
The model ID is incorrectly formatted
The model is in a different region than iad
An enterprise deployed a custom fine-tuned model for generating financial reports. After the first month, the model's outputs began to include outdated information and occasional factual errors. The team suspects data drift. What is the best course of action?
Switch to a newer base model like Llama 3.1 without retraining.
Decrease the temperature parameter to 0.1 to reduce model creativity.
Retrain the model on the latest financial data and monitor for drift.
Retraining on current financial data directly addresses drift by refreshing the model's learned patterns, correcting outdated outputs and factual errors. Ongoing drift monitoring then detects future distribution shifts, satisfying the stem's requirement to restore accuracy and prevent recurrence.
Increase the max tokens value to allow longer responses.
A developer is using OCI Generative AI Service to generate product descriptions. The outputs are often too generic and lack brand-specific tone. The developer has a small set of 20 high-quality example descriptions. What is the most efficient approach to improve output quality?
Fine-tune a base model on the 20 examples.
Use few-shot prompting by including the 20 examples in the prompt.
Few-shot prompting places the 20 curated examples directly in the context window, letting the model infer the desired brand tone through in-context learning. This satisfies the constraint of a small example set without retraining or fine-tuning, making it the most efficient route to less generic output.
Use a more detailed system prompt describing the brand tone.
Use chain-of-thought prompting to guide the model step by step.
A team is using OCI Generative AI Agents to build a customer support bot. The bot sometimes generates answers that contradict the knowledge base. What is the most likely cause?
The chunking strategy for the knowledge base does not capture enough context overlap.
Insufficient chunk overlap splits related sentences across separate chunks, so retrieval returns fragments lacking the surrounding context needed for a grounded answer. The model then fills gaps from its parametric knowledge, producing responses that contradict the knowledge base.
The max tokens value is too low, truncating the response.
The temperature parameter is set too high, causing the model to hallucinate.
The model's repetition penalty is too high.
Want more Using OCI Generative AI Service practice?
Practice this domainThe 1Z0-1127-25 exam has 40 questions and must be completed in 90 minutes. The passing score is 65/1000.
Scenario-based questions covering exam objectives with detailed answer explanations.
The exam covers 8 domains: Building LLM Applications with RAG and Vector Search, Fundamentals of Large Language Models, Prompt Engineering, Deploying and Managing Generative AI on OCI, OCI Generative AI Service, LLM Fundamentals, LangChain and AI Application Development, Using OCI Generative AI Service. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official Oracle 1Z0-1127-25 exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.