Databricks-DE-Assoc Data Transformation and Modeling Practice Question
A data engineer is building a Delta Live Tables pipeline that ingests streaming data from a Kafka topic into a bronze table, then applies a series of transformations to produce a silver table. The engineer notices that the pipeline is reprocessing all data from the beginning of the Kafka topic on each run, causing high latency. The Kafka topic has a retention period of 7 days, and the pipeline is configured to use the default settings. What is the most likely cause of this behavior?
⚠ Common exam trap
The trap here is assuming that streaming sources require manual checkpoint configuration, but DLT handles checkpoints automatically; the real issue is the persistence of the storage location.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The pipeline is using a fresh checkpoint location on each run because the storage location for the pipeline is not persistent or is being cleared.
The root cause is that the pipeline's checkpoint location is not persistent, causing DLT to lose track of processed offsets. Delta Live Tables automatically manages checkpoints, but they must be stored in a durable location. If the storage location is ephemeral or cleared, the pipeline will treat each run as a new stream, reprocessing all available data from the Kafka topic. Configuring a persistent storage location resolves this issue.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The pipeline is configured to run in continuous mode, but the cluster is being restarted, causing the checkpoint to be lost.
Why it's wrong here
In continuous mode, DLT maintains checkpoints across cluster restarts because they are stored in the pipeline's storage location. Cluster restarts do not cause checkpoint loss. If the pipeline were in triggered mode, it would process new data since the last run. The symptom of reprocessing all data from the beginning indicates that the checkpoint is either not being used or is being invalidated, which is not typical of continuous mode with proper storage configuration.
- ✓
The pipeline is using a fresh checkpoint location on each run because the storage location for the pipeline is not persistent or is being cleared.
Why this is correct
Delta Live Tables stores checkpoints in the pipeline's storage location. If that location is not persistent (e.g., a temporary directory) or is being cleared between runs, the pipeline cannot resume from the last processed offset. As a result, it will reprocess all data from the beginning of the Kafka topic. Ensuring a persistent storage location is critical for maintaining streaming state and avoiding redundant processing.
- ✗
The pipeline is not configured with a checkpoint location, so it cannot track the offset of the last processed record.
Why it's wrong here
Delta Live Tables automatically manages checkpoints for streaming sources when you use the STREAMING LIVE TABLE syntax. While a checkpoint location is crucial for streaming, DLT handles this internally. The issue described is not due to missing checkpoint configuration; rather, it is likely related to how the source is defined or how the pipeline is triggered. Without a checkpoint, the pipeline would fail entirely, not reprocess from the beginning.
- ✗
The Kafka source is defined using the readStream method without specifying the startingOffsets option, so it defaults to 'earliest' on each run.
Why it's wrong here
In Delta Live Tables, streaming sources are defined using the STREAMING LIVE TABLE syntax, not directly with readStream. DLT abstracts the streaming configuration and manages offsets via checkpoints. Specifying startingOffsets is not required; DLT will continue from where it left off. The behavior of reprocessing from the beginning suggests that the checkpoint is being reset or not persisted, which is not caused by missing startingOffsets.
Visual reference
About these practice questions
This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.