Databricks-Spark-Assoc Using Spark SQL Practice Question
A developer is using Spark SQL to process a streaming DataFrame from a Kafka source. They need to perform a stateful operation that maintains state across micro-batches to count occurrences of each key over a sliding window of 10 minutes, sliding every 5 minutes. Which Spark SQL operation should they use?
⚠ Common exam trap
Many candidates confuse tumbling windows with sliding windows; tumbling windows do not overlap, while sliding windows require both window and slide durations.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use groupBy with window('timestamp', '10 minutes', '5 minutes') and then count.
Structured streaming in Spark SQL supports windowed aggregations through the window function. Specifying a window duration and slide duration enables sliding windows. Grouping by the window and key, then applying an aggregation like count, maintains state across micro-batches and produces results for each window. This is the built-in, fault-tolerant approach for stateful stream processing.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use the foreachBatch sink to manually maintain state in an external database.
Why it's wrong here
foreachBatch allows custom logic per micro-batch, but it does not provide built-in state management for sliding windows. Manually maintaining state in an external database is complex, error-prone, and not the intended use of Spark SQL for stateful streaming. It bypasses Spark's state store and checkpointing mechanisms, making fault tolerance difficult.
- ✓
Use groupBy with window('timestamp', '10 minutes', '5 minutes') and then count.
Why this is correct
Spark SQL structured streaming supports windowed aggregations using the window function. Specifying a window duration of 10 minutes and slide duration of 5 minutes creates overlapping windows. Grouping by the window and key, then counting, maintains state across micro-batches to produce counts per window. This is the correct approach for sliding window aggregations in structured streaming.
- ✗
Use a tumbling window with window('timestamp', '10 minutes') and then count.
Why it's wrong here
A tumbling window of 10 minutes does not slide; it creates non-overlapping windows. This would not produce counts every 5 minutes. The requirement is for a sliding window with a 5-minute slide, so a tumbling window is incorrect. It would miss the overlapping nature and produce different results.
- ✗
Use the dropDuplicates operator to remove duplicates within 10 minutes, then count.
Why it's wrong here
dropDuplicates is used for deduplication, not for windowed aggregation. It does not maintain state for counting over sliding windows. While it can be used in streaming for deduplication with a watermark, it does not perform the required counting operation. It would not produce counts per key per window.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.