Databricks-DE-Assoc Databricks Intelligence Platform Practice Question
A data engineer is designing a Delta Lake table that will be used for both batch and streaming reads. The table must support upserts from a streaming source and maintain ACID transactions. The engineer wants to ensure that the table can be efficiently queried by downstream consumers using SQL while minimizing storage costs. Which TWO actions should the engineer take to meet these requirements? (Choose two.)
⚠ Common exam trap
The trap here is assuming that partitioning by a high-cardinality column always improves performance, when it often leads to small file problems and higher costs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the MERGE INTO command to perform upserts from the streaming source into the Delta table.
To support streaming upserts with ACID transactions and efficient SQL queries while minimizing storage costs, the engineer should use MERGE INTO for upserts and regularly run OPTIMIZE with Z-ORDER. MERGE INTO ensures atomic upserts, and OPTIMIZE with Z-ORDER compacts files and improves data skipping, reducing storage and query costs. Other options either add overhead, degrade performance, or remove essential Delta Lake features.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use the MERGE INTO command to perform upserts from the streaming source into the Delta table.
Why this is correct
MERGE INTO is the correct SQL command to perform upserts (inserts, updates, deletes) in a single atomic operation on a Delta table. It supports ACID transactions and is efficient for streaming upserts when used with foreachBatch. This meets the requirement for upserts from a streaming source while maintaining ACID guarantees.
- ✗
Enable change data feed on the Delta table to allow downstream consumers to read change events.
Why it's wrong here
Change data feed is useful for capturing row-level changes, but it is not required for upserts or ACID transactions. It adds storage overhead and is typically used for CDC scenarios. Enabling it does not directly minimize storage costs or improve SQL query performance for the base table. It is not a primary requirement for the described use case.
- ✗
Partition the Delta table by a high-cardinality column to improve query performance.
Why it's wrong here
Partitioning by a high-cardinality column can lead to the small file problem, increasing storage costs and degrading query performance due to excessive metadata. It is generally recommended to partition by low-cardinality columns (e.g., date) or use liquid clustering. This action would not minimize storage costs and could harm performance.
- ✓
Run OPTIMIZE with Z-ORDER on the Delta table regularly to compact small files and improve data skipping.
Why this is correct
OPTIMIZE compacts small files into larger ones, reducing storage overhead and improving read performance. Z-ORDER clusters data by specified columns, enhancing data skipping for SQL queries. This directly addresses minimizing storage costs and improving query efficiency, which are key requirements for the described scenario.
- ✗
Convert the Delta table to a Parquet table to reduce storage costs.
Why it's wrong here
Converting to Parquet would remove ACID transactions, time travel, and schema enforcement, which are required for upserts and streaming reads. Parquet does not support MERGE INTO or transactional guarantees. This action would violate the core requirements and is not appropriate for a Delta Lake table.
About these practice questions
This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.