Courseiva
Design Applications →mediumMultiple Choice

Databricks-GenAI-Assoc Design Applications Practice Question

An organization needs to build a RAG application on Databricks that minimizes data egress and maximizes security by keeping all data within the workspace perimeter. Which architectural pattern best satisfies this requirement?

⚠ Common exam trap

Candidates often suggest external API-based embedding models, ignoring the requirement to keep data within the workspace perimeter to minimize egress and satisfy strict residency requirements.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy an embedding model on Mosaic AI Model Serving and utilize Databricks Vector Search.

Utilizing Mosaic AI Model Serving with private endpoints and leveraging Vector Search indexes ensures that both the embedding model and the retrieval process occur within the Databricks control plane. By avoiding external API calls to third-party providers, the organization maintains strict governance, data residency compliance, and lower latency for inference, which is critical for enterprise-grade generative AI applications handling sensitive corporate documents.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Call external LLM APIs from a standard Python notebook without VPC constraints.

    Why it's wrong here

    Relying on external APIs introduces significant data egress risks and potential compliance violations for sensitive data. Without VPC constraints, data traffic is not contained within the secure workspace perimeter, making it unsuitable for organizations requiring strict internal governance and data residency control for their RAG pipeline components.

  • ✗

    Export data to an external vector database and use a cloud-hosted LLM.

    Why it's wrong here

    Exporting data outside the Databricks environment adds unnecessary complexity and increases the attack surface for sensitive information. Maintaining separate infrastructure for vector storage also complicates the governance model and lifecycle management of the RAG application, which should ideally be centralized within the Databricks ecosystem for optimal performance.

  • ✓

    Deploy an embedding model on Mosaic AI Model Serving and utilize Databricks Vector Search.

    Why this is correct

    This approach keeps all data processing, storage, and inference within the Databricks workspace perimeter. By using internal serving for embeddings and native vector search capabilities, the architecture minimizes egress traffic and simplifies security policy enforcement, ensuring that data never leaves the protected environment during the RAG retrieval process.

  • ✗

    Use a public LLM endpoint with a public bucket to store the vector index.

    Why it's wrong here

    Storing vector indexes in public buckets exposes them to unauthorized access, violating data security standards. Using public LLM endpoints also necessitates sending proprietary data over public networks, creating significant egress concerns and potential data leakage, which contradicts the fundamental requirements for a secure and private RAG application architecture.

About these practice questions

Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.