Courseiva

Databricks-DE-Assoc Data Transformation and Modeling Practice Question

Exhibit

{
  "source": "Kafka",
  "processing": "Structured Streaming",
  "issue": "State store is growing indefinitely.",
  "current_code": "df.groupBy('user_id').count().writeStream..."
}

Refer to the exhibit. The streaming pipeline is experiencing memory issues because the state store size increases continuously. What is the most effective way to address this while maintaining aggregation accuracy?

⚠ Common exam trap

Candidates often try to increase cluster memory or timeout settings to fix state issues, failing to realize that watermarking is the architectural fix for cleaning up stale state.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Add a watermark to the stream and use a windowed aggregation instead of a global count.

Continuous state growth is a common issue when grouping by high-cardinality keys like 'user_id' in streaming without time-bound constraints. By introducing a watermark and a window, you inform Spark which state is eligible for eviction. This allows the engine to periodically clean up expired state, ensuring the pipeline remains stable and memory usage stays within the bounds of the allocated cluster resources.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the cluster memory and executor count indefinitely.

    Why it's wrong here

    Increasing cluster resources is just masking the underlying problem. Because the state is growing continuously, it will eventually exhaust even the largest cluster. This approach is not scalable and leads to massive infrastructure costs without actually fixing the memory leak caused by unbound state in the streaming application.

  • ✓

    Add a watermark to the stream and use a windowed aggregation instead of a global count.

    Why this is correct

    Adding a watermark allows the state store to expire old data. By moving from a global count to a windowed count, you limit the state to the size of the window duration. This is the correct pattern to ensure long-term stability in streaming pipelines dealing with continuous, unbounded event data.

  • ✗

    Change the output mode to 'Append' to flush the state to the sink immediately.

    Why it's wrong here

    Append mode is not even supported for aggregations in Structured Streaming because it does not allow for updates to previous records. This will result in an AnalysisException. The output mode must be 'Update' or 'Complete', and neither will solve the issue of the state store growing in the absence of watermarks.

  • ✗

    Use the 'dropDuplicates' method instead of 'groupBy' to simplify the state.

    Why it's wrong here

    dropDuplicates also maintains state and will grow indefinitely if it is not time-bound. It does not resolve the memory issue. In fact, if the goal is to count occurrences per user, dropDuplicates doesn't provide the same functionality as a group-by aggregation, so it's both incorrect and ineffective.

About these practice questions

One of 276 original Databricks-DE-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.