Courseiva

Databricks-DE-Assoc Data Transformation and Modeling Practice Question

You are tasked with handling late-arriving data in a streaming pipeline that performs windowed aggregations. Which approach ensures that the output remains accurate while balancing memory usage?

⚠ Common exam trap

Candidates frequently confuse watermarking with general window durations or drop requirements. They assume watermarking alone will trigger output actions without pairing it with proper window aggregations.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implement a watermark on the timestamp column and aggregate over a window.

Watermarking is the primary mechanism to handle late-arriving data. It allows the system to define a threshold for how long state should be maintained for late updates. By setting an appropriate watermark duration, you balance the requirement for accuracy (including late data) with the need for memory efficiency (cleaning up expired state). This is a foundational concept for robust streaming ETL on Databricks.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the state TTL to infinity to ensure no data is ever dropped.

    Why it's wrong here

    Setting TTL to infinity will lead to state store exhaustion, eventually crashing the streaming application. State storage is finite, and keeping every record forever prevents the system from cleaning up old event windows. This approach is not scalable and will cause the pipeline to fail in production scenarios.

  • ✓

    Implement a watermark on the timestamp column and aggregate over a window.

    Why this is correct

    Watermarks allow Spark to understand how much data is allowed to be late. By associating a watermark with a windowed aggregation, Spark can safely discard old state for time windows that have passed the threshold, maintaining both correctness and memory efficiency as the stream processes data over time.

  • ✗

    Use the 'complete' output mode to ensure all historical data is reprocessed.

    Why it's wrong here

    Complete mode requires keeping the entire historical dataset in memory to recalculate the state. This is extremely inefficient and not intended for windowed aggregations on streaming data. It will cause massive performance degradation and eventually crash the executor nodes due to out-of-memory errors on large datasets.

  • ✗

    Filter out records with timestamps older than the current processing time.

    Why it's wrong here

    Filtering based on processing time ignores the event time of the data. In distributed systems, data often arrives out of order; using processing time as a filter will discard valid late data, leading to incorrect business metrics. You must use event time to ensure accuracy in streaming analytics.

About these practice questions

This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.