Courseiva

Databricks-DE-Pro Data Ingestion and Acquisition Practice Question

Which approach is most appropriate for ingesting data from a JDBC source into Delta Lake where the source table has no 'updated_at' or 'version' column for incremental loading?

⚠ Common exam trap

Candidates often suggest using MERGE or incremental ingestion logic even when no watermark column exists, forgetting that these techniques require a reliable way to identify new or modified records.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Perform a full overwrite of the Delta table for every load.

When a source table lacks a watermark column, you cannot use standard incremental loading techniques. The most robust approach is a full overwrite of the destination table, ensuring that the target remains a faithful copy of the source. While this can be resource-intensive for large tables, it is the only way to ensure data integrity without primary keys or timestamps to track changes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use the 'partitionColumn' parameter with a random UUID.

    Why it's wrong here

    Partition columns are used for parallelizing reads from large JDBC tables. Using a random UUID as a partition column provides no benefit for incremental extraction. Furthermore, without a logical timestamp or sequence, you cannot identify which records are new or updated, rendering this approach ineffective for incremental data ingestion.

  • ✓

    Perform a full overwrite of the Delta table for every load.

    Why this is correct

    Since there is no mechanism to identify changed data, performing a full overwrite ensures the target table always matches the source. This is the standard pattern for handling tables without watermark columns. You should balance the frequency of the load with the size of the table to manage compute costs.

  • ✗

    Enable streaming ingestion using the JDBC source readStream API.

    Why it's wrong here

    Streaming reads from JDBC require a watermark column (like a timestamp or incrementing ID) to track the progress of the stream. Without such a column, the readStream API cannot determine which records have been processed and which are new, leading to failure or incorrect data loading behaviors.

  • ✗

    Use the 'fetchSize' parameter to optimize the load.

    Why it's wrong here

    The 'fetchSize' parameter improves performance by controlling how many rows are retrieved in each network round-trip. While it improves the efficiency of the data transfer, it does not solve the fundamental challenge of tracking changes in the source table without a watermark or version column for incremental processing.

About these practice questions

One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.