Databricks-Spark-Assoc Structured Streaming Practice Question
A developer writes a Structured Streaming query that reads from a Kafka topic with `spark.readStream.format("kafka")` and then calls `.writeStream.format("console").start()`. The query runs, but after a few minutes the driver logs show that the query is only processing newly arriving offsets and older messages in the topic are never read. What is the most likely cause?
⚠ Common exam trap
The trap here is assuming the Kafka source always reads the full topic backlog by default, when in fact its default start position is the latest offset.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The Kafka source defaults to `startingOffsets` = "latest", so the first micro-batch begins at the tail of each partition and earlier records are skipped.
When a Kafka-backed streaming query starts without an existing checkpoint, the source resolves its first offsets from `startingOffsets`. Because the default is "latest", the initial micro-batch aligns to the end of each partition, so any records produced before the query started are ignored. Setting `startingOffsets` to "earliest" or a JSON offset map is required to consume the backlog.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Kafka consumers always begin reading from the committed consumer-group offset, and Structured Streaming reuses that group offset as its start position.
Why it's wrong here
Databricks Structured Streaming does not join or use a Kafka consumer group to track its position. Progress is persisted in the checkpoint location, and the initial position is governed by `startingOffsets`. Reusing a consumer-group offset is a classic Spark Streaming (DStreams) behavior, not how the Structured Streaming Kafka source works here.
- ✗
The console sink silently discards any record whose Kafka `offset` value is lower than the current micro-batch ID.
Why it's wrong here
The console sink writes every row of each micro-batch to stdout using the query's output mode; it has no logic that compares Kafka offsets to micro-batch identifiers. It cannot filter older records, so this cannot explain why earlier topic data is skipped. The skipping is decided earlier, at the source when offsets are resolved.
- ✗
The topic's retention policy removed all historical segments before the query started, so only new records remained available to read.
Why it's wrong here
Kafka retention deletes whole log segments once they exceed the configured time or size limit, and it would remove records permanently. If retention had already purged the older data, the developer could not claim the records were skipped but still present. Retention alone does not set a query's initial read position, which is what determines whether old offsets are consumed.
- ✓
The Kafka source defaults to `startingOffsets` = "latest", so the first micro-batch begins at the tail of each partition and earlier records are skipped.
Why this is correct
The Kafka Structured Streaming source uses `startingOffsets` to decide where the very first query starts when no checkpoint exists. Its default value is "latest", which means the query begins at the newest offset per partition. Older records already present in the topic are therefore never delivered to the first micro-batch, exactly matching the observed behavior.
Visual reference
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.