Courseiva
Structured Streaming →hardMultiple Choice

Databricks-Spark-Assoc Structured Streaming Practice Question

You are running a Structured Streaming query on Databricks that reads from a Kafka topic and writes to a Delta table. The query uses `option("maxOffsetsPerTrigger", 10000)` to limit the number of records per micro-batch. During a peak, the Kafka topic accumulates a large backlog. You notice that the query is processing data but the backlog is not decreasing. What is the most likely cause?

⚠ Common exam trap

The trap here is assuming that `maxOffsetsPerTrigger` is a soft limit or that it only applies to the initial load, when in fact it strictly caps the number of records processed per micro-batch.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The `maxOffsetsPerTrigger` value is lower than the rate at which new data is arriving.

The backlog is not decreasing because the query is limited to processing 10000 records per micro-batch, which is likely less than the incoming rate. This cap is set by `maxOffsetsPerTrigger`. To reduce the backlog, you must increase this limit or the processing capacity. The other options do not address the rate mismatch: the option is not ignored, triggers do not overlap, and starting offsets only affect initial reads.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The Kafka source is not configured with `startingOffsets` set to `latest`.

    Why it's wrong here

    The `startingOffsets` option only determines where to start reading when the query is first started or when there is no checkpoint. It does not affect the ongoing processing rate. If the query is already running and processing data, changing `startingOffsets` would not help with the backlog. The backlog issue is due to the rate limit, not the starting point.

  • ✗

    The trigger interval is too short, causing the query to start new micro-batches before the previous one finishes.

    Why it's wrong here

    Structured Streaming processes micro-batches sequentially; a new micro-batch does not start until the previous one completes. A short trigger interval would not cause overlapping batches. If the trigger interval is too short, the query may simply wait for the next trigger, but it does not cause the backlog to remain constant. The issue is more likely related to the processing rate versus the incoming rate.

  • ✓

    The `maxOffsetsPerTrigger` value is lower than the rate at which new data is arriving.

    Why this is correct

    If `maxOffsetsPerTrigger` is set to 10000 and the Kafka topic receives more than 10000 new records per trigger interval, the query will only process 10000 records per micro-batch. The backlog will grow because the processing rate is capped below the arrival rate. To reduce the backlog, you would need to increase `maxOffsetsPerTrigger` or increase the trigger frequency, provided the cluster can handle the load.

  • ✗

    The `maxOffsetsPerTrigger` option is ignored when writing to Delta Lake.

    Why it's wrong here

    The `maxOffsetsPerTrigger` option is a Kafka source option and is respected regardless of the sink. It limits the number of offsets processed per trigger. Writing to Delta Lake does not disable this option. The backlog not decreasing suggests that the processing rate is lower than the incoming rate, but the option itself is not ignored. This option is not the cause of the issue.

About these practice questions

Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.