Courseiva

Databricks-DE-Assoc Databricks Intelligence Platform Practice Question

A data engineer is designing a Delta Lake table that will be used for both batch and streaming reads. The table must support upserts from a streaming source and maintain ACID transactions. The engineer wants to ensure that the table can be efficiently queried by downstream consumers using SQL while minimizing storage costs. Which TWO actions should the engineer take to meet these requirements? (Choose two.)

⚠ Common exam trap

The trap here is assuming that partitioning by a high-cardinality column always improves performance, when it often leads to small file problems and higher costs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the MERGE INTO command to perform upserts from the streaming source into the Delta table.

To support streaming upserts with ACID transactions and efficient SQL queries while minimizing storage costs, the engineer should use MERGE INTO for upserts and regularly run OPTIMIZE with Z-ORDER. MERGE INTO ensures atomic upserts, and OPTIMIZE with Z-ORDER compacts files and improves data skipping, reducing storage and query costs. Other options either add overhead, degrade performance, or remove essential Delta Lake features.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the MERGE INTO command to perform upserts from the streaming source into the Delta table.

    Why this is correct

    MERGE INTO is the correct SQL command to perform upserts (inserts, updates, deletes) in a single atomic operation on a Delta table. It supports ACID transactions and is efficient for streaming upserts when used with foreachBatch. This meets the requirement for upserts from a streaming source while maintaining ACID guarantees.

  • ✗

    Enable change data feed on the Delta table to allow downstream consumers to read change events.

    Why it's wrong here

    Change data feed is useful for capturing row-level changes, but it is not required for upserts or ACID transactions. It adds storage overhead and is typically used for CDC scenarios. Enabling it does not directly minimize storage costs or improve SQL query performance for the base table. It is not a primary requirement for the described use case.

  • ✗

    Partition the Delta table by a high-cardinality column to improve query performance.

    Why it's wrong here

    Partitioning by a high-cardinality column can lead to the small file problem, increasing storage costs and degrading query performance due to excessive metadata. It is generally recommended to partition by low-cardinality columns (e.g., date) or use liquid clustering. This action would not minimize storage costs and could harm performance.

  • ✓

    Run OPTIMIZE with Z-ORDER on the Delta table regularly to compact small files and improve data skipping.

    Why this is correct

    OPTIMIZE compacts small files into larger ones, reducing storage overhead and improving read performance. Z-ORDER clusters data by specified columns, enhancing data skipping for SQL queries. This directly addresses minimizing storage costs and improving query efficiency, which are key requirements for the described scenario.

  • ✗

    Convert the Delta table to a Parquet table to reduce storage costs.

    Why it's wrong here

    Converting to Parquet would remove ACID transactions, time travel, and schema enforcement, which are required for upserts and streaming reads. Parquet does not support MERGE INTO or transactional guarantees. This action would violate the core requirements and is not appropriate for a Delta Lake table.

About these practice questions

This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.