Courseiva

CCNA Data Sharing Federation Questions

20 questions · Data Sharing Federation topic · All types, answers revealed

1
MCQmedium

What is the primary difference between sharing data via Delta Sharing compared to sharing data via Databricks-to-Databricks sharing?

A.Delta Sharing is only for real-time streaming data.
B.Databricks-to-Databricks sharing requires the recipient to download a credential file.
C.Delta Sharing allows recipients outside of the Databricks ecosystem to access data.
D.Databricks-to-Databricks sharing is limited to within the same cloud region.
AnswerC

Delta Sharing is an open-source protocol that decouples the data provider from the recipient's environment. This enables organizations to share data securely with partners who may be using different cloud providers or query engines, as long as they can consume the open Delta Sharing protocol via a credential file.

Why this answer

Delta Sharing is an open protocol that allows sharing data with any client, including those outside the Databricks ecosystem, by using credential files. Databricks-to-Databricks sharing is a proprietary feature that allows seamless data access across different Unity Catalog metastores within the Databricks ecosystem, leveraging built-in authentication and governance features without requiring external credential management or token distribution.

Exam trap

Candidates often assume Delta Sharing only works within the Databricks platform, confusing it with internal workspace sharing, and overlook its core capability as an open protocol for external recipients.

2
MCQeasy

A data engineer is using Lakehouse Federation to query data from an external PostgreSQL database. The engineer has created a foreign catalog and foreign schema. Which statement accurately describes how the data is accessed during a query?

A.The query is translated into PostgreSQL SQL, executed on the external database, and the results are streamed back to Databricks for further processing.
B.The external data is cached in the Databricks workspace's local storage for the duration of the query.
C.The query is executed entirely within the external PostgreSQL database, and only the results are returned to Databricks.
D.The data is copied into Delta Lake tables in Unity Catalog before the query runs.
AnswerA

Lakehouse Federation uses query pushdown to translate parts of the query into the external database's SQL dialect, executes them remotely, and streams the results back to Databricks. Databricks then performs any remaining operations, such as joins with local data or final aggregations. This minimizes data transfer and leverages the external database's capabilities.

Why this answer

Lakehouse Federation enables querying external databases without copying data. It uses query pushdown to translate and execute parts of the query on the external system, then streams results back to Databricks for any remaining processing. This approach minimizes data movement and leverages the external database's compute.

The other options incorrectly describe data copying or full remote execution.

Exam trap

The trap here is assuming that Lakehouse Federation copies data into Databricks or executes the entire query externally, when it actually uses a hybrid pushdown approach.

3
MCQmedium

A Data Engineer needs to share a Delta table with a partner organization using Delta Sharing. The partner does not use Databricks. Which component must the Data Engineer generate to facilitate this secure connection?

A.A shared Unity Catalog metastore link
B.A personal access token with REST API permissions
C.A sharing credential file
D.A cross-account IAM role in the recipient's cloud
AnswerC

The sharing credential file is the fundamental mechanism for Delta Sharing. It contains the short-lived access token and the endpoint URL required by the external client to authenticate and authorize requests. This file acts as the bridge between the provider's data and the recipient's consumption tool, ensuring secure access.

Why this answer

Delta Sharing uses a credential file to authenticate external recipients. The provider generates a sharing token encapsulated in a JSON file, which the recipient imports into their Delta Sharing client. This ensures that the data provider maintains full control over access permissions while the recipient can query the data using standard tools without needing a Databricks account, maintaining a secure and decoupled architecture.

Exam trap

Candidates often mistakenly believe external non-Databricks partners need a Databricks account or shared JDBC connection strings to access Delta Sharing data.

4
MCQmedium

When using Delta Sharing to share data with a recipient, what is the best way to handle updates to the shared data?

A.The provider must recreate the share object.
B.The recipient must request a new credential file.
C.Updates are automatically visible to the recipient.
D.The provider must run an 'REFRESH SHARE' command.
AnswerC

Because Delta Sharing queries the live Delta table, the recipient always sees the most recent committed state of the data. This provides a seamless, real-time data sharing experience, eliminating the manual overhead of exporting files or synchronizing data snapshots between the provider and the recipient organizations.

Why this answer

Delta Sharing automatically reflects the state of the table at the time of the query. Because the recipient accesses the Delta table directly (or via a managed share), any new data written to the table becomes immediately available to the recipient. There is no need for the provider to manually refresh the share, which ensures data consistency and reduces the maintenance burden for the Data Engineer.

Exam trap

Candidates often assume that Delta Sharing requires a manual 'push' or refresh mechanism to sync data, failing to realize that Delta Sharing provides a direct, live view of the underlying table.

5
MCQmedium

Which statement correctly describes the relationship between Unity Catalog and Databricks SQL Warehouses when using Lakehouse Federation?

A.The SQL Warehouse performs the data ingestion into the Unity Catalog metastore.
B.The SQL Warehouse pushes down query predicates to the external database.
C.Unity Catalog must store the external data in a managed Delta table.
D.All federated queries are executed solely within the Databricks control plane.
AnswerB

A key benefit of Lakehouse Federation is query pushdown. The SQL Warehouse optimizes the execution plan by sending filters, aggregations, and joins to the external database engine. This reduces data movement across the network and leverages the external database's compute power, significantly improving query performance for federated sources.

Why this answer

Lakehouse Federation allows users to query external databases directly from Databricks SQL Warehouses. The SQL Warehouse acts as the compute engine that translates Spark SQL queries into the native dialect of the external source. Unity Catalog acts as the central governance layer that stores the connection information, credentials, and schema mapping, allowing users to query external data as if it were local tables.

Exam trap

Candidates often think the data is copied into Databricks first, missing that Lakehouse Federation allows pushdown execution directly against the source database.

6
MCQhard

A data engineer is managing a Delta Share that includes a table with frequent updates. The engineer wants to ensure that recipients always see the latest version of the data without manual intervention. Which statement is correct regarding how Delta Sharing handles updates to shared tables?

A.Recipients must manually refresh their local copy of the shared table to see updates.
B.Updates are automatically propagated to recipients, and they see the latest data on their next query.
C.The shared table is versioned, and recipients must specify a version to query; otherwise, they see the initial snapshot.
D.Recipients receive change data capture (CDC) streams and must apply them to a local table.
AnswerB

Delta Sharing serves the current version of the shared table at query time. When the provider updates the table, the changes are immediately available to recipients on their next query. No manual action is needed. This is a key benefit of Delta Sharing: it provides live access to the latest data without copying or synchronization.

Why this answer

Delta Sharing provides live access to the shared table's current state. When the provider updates the table, recipients automatically see the changes on their next query without any manual refresh or CDC application. This is because the Delta Sharing server reads the latest Delta transaction log and serves the current snapshot.

The other options incorrectly suggest manual steps or default to stale data.

Exam trap

The trap here is assuming that recipients must refresh or apply CDC to see updates, when Delta Sharing actually serves the latest data on each query.

7
MCQeasy

A data engineer wants to share a Delta table with an external partner who does not have a Databricks account. The partner needs to access the data using Python. Which method should the engineer recommend to the partner for accessing the shared data?

A.Mount the shared table as an external table in their local Spark cluster using the Delta Sharing connector.
B.Use the Databricks SQL Connector for Python with the partner's Databricks personal access token.
C.Install the delta-sharing Python library and use the provided credential file to read the shared table.
D.Download the shared table as a Parquet file from the Databricks workspace and load it into their Python environment.
AnswerC

The delta-sharing Python library is specifically designed for recipients to access Delta Shares without needing a Databricks account. The partner can install the library via pip, use the credential file provided by the data engineer, and read the shared table as a pandas DataFrame or Apache Spark DataFrame. This method is secure, supports the Delta Sharing protocol, and is the recommended approach for non-Databricks recipients.

Why this answer

For recipients without a Databricks account, the delta-sharing Python library is the standard way to access shared data. It uses the credential file to authenticate and allows reading the shared table into Python. This method is secure, supports live data, and does not require Databricks infrastructure on the recipient side.

Other methods either require Databricks access or are not part of the Delta Sharing protocol.

Exam trap

The trap here is assuming that the partner needs Databricks-specific tools like the SQL Connector, when the open Delta Sharing protocol provides a dedicated Python library for non-Databricks users.

8
Multi-Selectmedium

A Data Engineer is setting up Lakehouse Federation for a PostgreSQL database. Which TWO steps are required to ensure that users can securely query the data using Unity Catalog?

Select 2 answers
A.Create a connection object with appropriate credentials in Unity Catalog.
B.Ingest all PostgreSQL data into a S3 bucket first.
C.Create a foreign catalog that references the PostgreSQL connection.
D.Install a custom JDBC driver on every user's local machine.
E.Configure a VPC peering connection to the external database.
AnswersA, C

Creating a connection object is the first step. It encapsulates the connection details (URL, driver, etc.) and credentials (e.g., username/password or secret) in a secure manner. This object is stored in Unity Catalog, allowing administrators to manage access centrally and providing compute clusters with necessary information to access.

Why this answer

To set up Lakehouse Federation, you must first establish a connection object that stores credentials securely in Unity Catalog. Then, you must create a foreign catalog that maps the remote database to a local Unity Catalog structure. These two steps enable Unity Catalog to manage permissions and translate queries for the federated source without exposing raw credentials to end users.

Exam trap

Candidates often think querying external databases requires manual table replication or creating local views, skipping the required Unity Catalog connection and foreign catalog steps.

9
MCQeasy

A data engineer is setting up a Delta Share to provide a partner with access to a specific table. The partner uses Databricks and wants to query the shared data using their own Databricks workspace. What is the correct sequence of actions for the data engineer to enable this Databricks-to-Databricks sharing?

A.Create a share, add the table, create a recipient of type Databricks, and grant the recipient access to the share.
B.Export the table to a cloud storage location and provide the partner with the storage credentials.
C.Create a share, add the table, generate a credential file, and send it to the partner.
D.Grant the partner's Databricks workspace SELECT privileges on the table, and they can query it directly.
AnswerA

For Databricks-to-Databricks sharing, the provider creates a share, adds the table, creates a recipient with the Databricks sharing identifier, and grants the recipient access to the share. The recipient then mounts the share in their workspace. This sequence ensures proper authentication and access control.

Why this answer

Databricks-to-Databricks sharing uses Unity Catalog to create a share, add tables, and define a recipient using the partner's Databricks sharing identifier. The recipient is then granted access to the share. The partner mounts the share in their workspace, enabling them to query the data with their own compute.

This method avoids credential files and ensures secure, real-time access.

Exam trap

The trap here is thinking that a credential file is needed for Databricks-to-Databricks sharing, but that is only for non-Databricks recipients.

10
MCQmedium

A data engineer at a healthcare company needs to share a Delta table containing patient records with an external research partner. The partner must only see aggregated statistics, not individual patient rows. The engineer wants to enforce this at the data sharing layer without creating a separate physical copy of the table. Which approach should the engineer take?

A.Create a Delta Share and add the table with a partition filter that excludes sensitive columns.
B.Export the aggregated data to a CSV file and provide it to the partner via a secure file transfer.
C.Use Lakehouse Federation to create a foreign catalog pointing to the partner's database and grant SELECT on the aggregated columns.
D.Create a view that aggregates the data and share the view through Delta Sharing.
AnswerD

Delta Sharing supports sharing views, including views that aggregate data. By creating a view that computes the required statistics, the engineer can share only the aggregated results. The partner queries the view as if it were a table, but the underlying raw data remains protected. This enforces the aggregation logic at the data sharing layer without duplicating data, and the view definition is managed centrally in Unity Catalog.

Why this answer

Delta Sharing allows sharing views that can perform aggregation and filtering, enabling fine-grained control over what data is exposed. Sharing an aggregated view ensures the partner sees only summary statistics, not raw patient records. This approach avoids data duplication and maintains centralized governance.

The other options either do not support the required granularity or are not designed for external sharing.

Exam trap

The trap here is assuming that Delta Sharing supports column-level security or partition filters to hide sensitive columns, when in fact it requires a view to achieve that level of control.

11
MCQmedium

A data engineering team maintains a Unity Catalog metastore in a Databricks workspace. They need to provide an external partner with read-only access to a specific Delta table, but the partner's analytics platform is not Databricks and does not support the Delta Lake protocol. The partner can consume Parquet files over a REST API. Which Unity Catalog feature should the team use to share the table?

A.Unity Catalog Volumes
B.Databricks SQL Connector for Python
C.Delta Sharing
D.Lakehouse Federation
AnswerC

Delta Sharing is an open protocol that allows sharing Delta tables with external recipients, including non-Databricks clients, via a REST API. It automatically serves the shared data in Parquet format to recipients that do not support Delta, enabling seamless access. This matches the partner's requirement exactly and is the native Unity Catalog sharing feature for this scenario.

Why this answer

Delta Sharing is the correct choice because it is the only Unity Catalog feature designed to share Delta tables externally using an open REST protocol. It supports recipients on non-Databricks platforms by serving data as Parquet, which aligns with the partner's capabilities. The other options serve different purposes: Volumes store files, Lakehouse Federation queries external systems, and the SQL Connector is a client library.

Exam trap

The trap here is confusing Delta Sharing with Lakehouse Federation, which is used to query external data sources from Databricks, not to share data outward.

12
MCQmedium

A data engineer at a healthcare company needs to share a Delta table containing patient records with an external research partner. The partner uses a non-Databricks platform and requires read-only access to the latest data, with updates reflected in near real-time. The engineer creates a Delta Share and adds the table. Which additional step is required to allow the partner to access the shared data?

A.Create a recipient object and provide the recipient with the activation link or credential file.
B.Configure a Databricks SQL warehouse and share its connection string with the recipient.
C.Grant the recipient's user account SELECT privileges on the shared table.
D.Generate a bearer token for the recipient and provide it along with the share name.
AnswerA

To share data with an external partner, you must create a recipient in Unity Catalog, which generates an activation link or a credential file. The recipient uses this to authenticate and access the share. This is the standard procedure for Delta Sharing to non-Databricks users, enabling secure, read-only access to the shared tables.

Why this answer

For a non-Databricks recipient to access a Delta Share, the provider must create a recipient object in Unity Catalog. This generates an activation link or credential file that the recipient uses to authenticate. The share must contain the table, and the recipient is granted access to the share.

This process ensures secure, read-only access without requiring the recipient to have a Databricks account.

Exam trap

The trap here is assuming that granting SELECT privileges on the table directly to an external user is sufficient, but Delta Sharing requires a recipient object and credential file for non-Databricks access.

13
MCQmedium

A data engineer is using Lakehouse Federation to query an external PostgreSQL database. The engineer creates a connection with the PostgreSQL JDBC URL and credentials, and then creates a foreign catalog. Users report that queries against foreign tables are slow and sometimes fail with connection timeouts. The engineer checks the connection and confirms the credentials are correct. What is the most likely cause of the performance and timeout issues?

A.The foreign catalog is not configured with a read-only mode, causing write attempts that time out.
B.The foreign catalog must be refreshed to update statistics, causing slow queries.
C.The Databricks cluster lacks the necessary network connectivity to the PostgreSQL database, causing timeouts.
D.The PostgreSQL database is not configured with the correct Unity Catalog connection parameters.
AnswerC

Lakehouse Federation queries are executed by the Databricks cluster, which must have network access to the external database. If the cluster is in a VPC without proper peering or firewall rules, connections may be slow or time out. Ensuring network connectivity, such as VPC peering or allowing Databricks IPs, is crucial for performance and reliability.

Why this answer

Lakehouse Federation queries run on Databricks compute, which must connect to the external database over the network. If the cluster cannot reach the database due to network restrictions, queries will be slow or time out. Ensuring proper network connectivity, such as VPC peering or firewall rules, is essential for reliable federation.

Exam trap

The trap here is focusing on database configuration or catalog refresh, when the underlying issue is network connectivity between the Databricks cluster and the external database.

14
MCQmedium

What is the primary benefit of using Unity Catalog for data federation?

A.It eliminates the need for any network configuration.
B.It provides a single security model for all data sources.
C.It automatically converts all external data to Delta format.
D.It allows the SQL Warehouse to run entirely on the external database.
AnswerB

Unity Catalog abstracts the security differences between various data sources. Whether you are querying a PostgreSQL database, a Snowflake instance, or internal Delta tables, you use the same Unity Catalog permission syntax. This dramatically reduces the administrative overhead and potential for configuration errors across a complex multi-source data landscape.

Why this answer

Unity Catalog acts as a centralized governance layer that provides a unified namespace for both local and external data. By using Unity Catalog, organizations can apply consistent security policies, such as column-level masking or row-level filtering, to federated data sources. This allows users to access disparate systems through a single, secure interface without learning unique security models for each data source.

Exam trap

Candidates often assume data federation improves raw query performance or optimizes storage costs, missing that its primary value lies in centralized security governance.

15
MCQhard

A data engineer is using Lakehouse Federation to query a PostgreSQL database. The engineer notices that a query filtering on a column with a high cardinality is performing poorly, even though the remote database has an index on that column. What is the most likely reason for the poor performance?

A.The PostgreSQL database is not configured with the correct statistics, causing the query planner to choose a sequential scan.
B.The foreign catalog is using a JDBC connection with a small fetch size, causing many round trips to the database.
C.The foreign catalog is not configured to allow predicate pushdown for the PostgreSQL database.
D.The filter condition uses a function or expression that cannot be pushed down to PostgreSQL, causing a full table scan.
AnswerD

Lakehouse Federation pushes down simple predicates like equality and range filters. However, if the filter uses a function or expression that PostgreSQL cannot evaluate, such as a complex UDF or a non-deterministic function, the pushdown fails. Databricks then retrieves all rows and applies the filter locally, ignoring the remote index. This results in a full table scan and poor performance. The engineer should rewrite the query to use pushdown-compatible expressions.

Why this answer

Poor performance on a filtered query against a federated PostgreSQL database often indicates that the filter was not pushed down. If the filter uses an expression that PostgreSQL cannot evaluate, Databricks retrieves all data and filters locally, bypassing the remote index. The engineer should simplify the filter or use pushdown-compatible expressions to leverage the index and improve performance.

Exam trap

The trap here is blaming the remote database's statistics or configuration, when the issue is that the filter expression prevents predicate pushdown from Databricks to PostgreSQL.

16
Multi-Selecthard

A data engineer is setting up Lakehouse Federation to query an external MySQL database from Databricks. The engineer creates a connection using the MySQL connector and a foreign catalog. Users in the 'analysts' group report that they can see the foreign catalog but cannot query any tables. The engineer has granted USAGE on the connection to the 'analysts' group. Which TWO additional permissions must be granted to the 'analysts' group to allow them to query tables in the foreign catalog? (Choose two.)

Select 2 answers
A.BROWSE on the foreign catalog
B.CREATE on the foreign catalog
C.USE CATALOG on the foreign catalog
D.USE SCHEMA on the schemas containing the tables
E.SELECT on the foreign catalog
AnswersC, D

In Unity Catalog, to query objects within a catalog, a user must have USE CATALOG on that catalog. Even if the user has USAGE on the connection, without USE CATALOG on the foreign catalog, they cannot access any schemas or tables within it. This is a fundamental privilege for catalog-level access.

Why this answer

To query foreign tables in a foreign catalog, users need USE CATALOG on the foreign catalog and USE SCHEMA on the specific schemas. USAGE on the connection only allows the catalog to use the connection; it does not grant access to the catalog's objects. These two privileges are the minimum required for read access.

Exam trap

The trap here is assuming that USAGE on the connection is sufficient for querying foreign tables, overlooking the need for catalog and schema-level privileges.

17
MCQhard

A data engineer is managing a Delta Share that includes a table with customer transactions. The share is used by multiple recipients. The engineer needs to update the shared data daily with new transactions and also remove data for customers who have requested deletion (right to be forgotten). The engineer wants to ensure recipients see the updated data without having to recreate the share. What is the best approach?

A.Use a materialized view to capture changes and share the view instead of the table.
B.Update the underlying Delta table with new data and use DELETE to remove records; recipients will see changes on their next query.
C.Use Delta Sharing's built-in history sharing to automatically propagate updates and deletions.
D.Recreate the share with the updated table and notify recipients to update their credentials.
AnswerB

Delta Sharing shares the live table. When the provider updates the Delta table (e.g., with MERGE or INSERT) and deletes records, recipients querying the shared table will see the latest version. There is no need to recreate the share. The recipients' queries will reflect the current state of the table, including deletions, as long as they query the latest version.

Why this answer

Delta Sharing shares the current state of the Delta table. When the provider performs updates or deletes on the table, recipients see the changes on their next query. There is no need to recreate the share or use materialized views.

This approach ensures recipients always access the latest data, including deletions for compliance.

Exam trap

The trap here is thinking that updates require recreating the share or that deletions are not visible, when in fact Delta Sharing reflects the live table state.

18
Multi-Selectmedium

A data engineer is setting up a Delta Share to provide an external partner with access to a subset of data. The partner will use a non-Databricks client that supports the Delta Sharing protocol. Which two actions are required to enable the partner to access the shared data? (Choose two.)

Select 2 answers
A.Create a recipient in Unity Catalog and provide the partner with the activation link or credential file.
B.Configure a Databricks SQL warehouse for the recipient to query the shared data.
C.Add the table to the share and grant the recipient access to the share.
D.Enable the Delta Sharing server on the Databricks workspace.
E.Grant the recipient SELECT privileges on the shared table.
AnswersA, C

Creating a recipient in Unity Catalog is essential to establish the sharing relationship. The recipient is issued a credential file or activation link that the partner uses to authenticate to the Delta Sharing server. Without this, the partner cannot access the share. This step is a core part of the Delta Sharing setup process, ensuring secure and controlled access.

Why this answer

To enable an external partner to access a Delta Share, two key steps are required: creating a recipient and providing them with the credential file or activation link, and adding the table to a share while granting the recipient access to that share. These actions establish the secure sharing relationship and define the data scope. Other options are either not applicable to non-Databricks clients or not part of the Delta Sharing setup process.

Exam trap

The trap here is thinking that standard Unity Catalog SQL grants like SELECT are used for Delta Sharing recipients, when access is actually controlled at the share level.

19
MCQeasy

A Data Engineer is using Lakehouse Federation to query a Snowflake database from Databricks. The engineer has created a connection and a foreign catalog. Which statement correctly describes how data is accessed when a user queries a table in the foreign catalog?

A.The data is cached in the Databricks workspace after the first query.
B.The data is copied into Delta Lake before the query runs.
C.The query is executed on Databricks, and the data is streamed from Snowflake.
D.The query is executed on Snowflake, and only the results are returned to Databricks.
AnswerD

Lakehouse Federation pushes down queries to the external database when possible. For Snowflake, the query is executed on Snowflake, and only the result set is returned to Databricks. This minimizes data movement and leverages Snowflake's compute. The foreign catalog provides a unified interface, but the execution happens remotely.

Why this answer

Lakehouse Federation enables querying external databases without moving data. When a user queries a foreign catalog table, Databricks pushes the query down to the external database, such as Snowflake. The external database executes the query and returns only the results.

This approach minimizes data transfer and leverages the external system's compute. The foreign catalog provides a seamless experience, but the data remains in the source system.

Exam trap

The trap here is assuming that federation copies or caches data in Databricks, when it actually pushes down queries to the external system.

20
MCQhard

A Data Engineer is using Lakehouse Federation to query an external PostgreSQL database from Databricks. The engineer creates a foreign catalog named 'pg_catalog' using a connection that specifies the host, port, and credentials. Users report that queries against the foreign catalog fail with a permission error, even though the connection works when tested. The engineer confirms that the connection has the correct credentials and that the PostgreSQL user has SELECT privileges on the required tables. What is the most likely cause?

A.The PostgreSQL database is not in the same region as the Databricks workspace.
B.The users have not been granted USE on the foreign catalog and SELECT on the tables.
C.The connection uses a service principal instead of a user account.
D.The foreign catalog was created without specifying the 'postgresql' database type.
AnswerB

In Lakehouse Federation, after creating a foreign catalog, users must be granted USE on the catalog and SELECT on the tables within it. Without these Unity Catalog privileges, queries will fail with permission errors, even if the underlying PostgreSQL user has access. The connection credentials are used by Databricks to access the external database, but users still need Unity Catalog permissions.

Why this answer

In Lakehouse Federation, the connection stores credentials to access the external database, but end users must still have Unity Catalog privileges to query the foreign catalog. Specifically, they need USE on the foreign catalog and SELECT on the tables. Without these, queries fail with permission errors.

The connection test succeeds because it uses the stored credentials directly, bypassing user-level Unity Catalog checks. The correct answer addresses the missing Unity Catalog grants.

Exam trap

The trap here is assuming that if the connection works and the external user has SELECT, end users can query without additional Unity Catalog privileges.

Ready to test yourself?

Try a timed practice session using only Data Sharing Federation questions.