Courseiva
Structured Streaming →mediumMultiple Select

Databricks-Spark-Assoc Structured Streaming Practice Question

You are building a Structured Streaming job that reads from a Kafka source and writes to a Delta table. You need to ensure that the job can recover from failures and process data exactly once. Which TWO of the following are required to achieve exactly-once semantics? (Choose two.)

⚠ Common exam trap

The trap here is thinking that Kafka offset settings or rate limits provide exactly-once semantics, when actually checkpointing and an idempotent sink are the key requirements.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a checkpoint location to store progress information.

Exactly-once semantics in Structured Streaming require a reliable checkpoint location to track progress and an idempotent sink like Delta Lake that can handle replays without duplicating data. Checkpointing ensures the query can recover from failures, while Delta's transactional writes guarantee that each batch is applied exactly once.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable `spark.sql.streaming.forceDeleteTempCheckpointLocation` to clean up temporary files.

    Why it's wrong here

    This configuration is about managing temporary checkpoint files, but it does not ensure exactly-once semantics. In fact, incorrectly deleting checkpoint data can break fault tolerance. The property is unrelated to the core requirements for exactly-once, which are checkpointing and an idempotent sink.

  • ✗

    Configure the Kafka source with `startingOffsets` set to `earliest`.

    Why it's wrong here

    `startingOffsets` only determines where to start reading when the query first starts or when there is no checkpoint. It does not by itself guarantee exactly-once semantics. For exactly-once, you need checkpointing and idempotent writes; starting from earliest is a separate concern.

  • ✗

    Set the `maxOffsetsPerTrigger` option to limit the number of records per batch.

    Why it's wrong here

    `maxOffsetsPerTrigger` controls the rate of data ingestion, helping with backpressure and resource management. It does not contribute to exactly-once semantics. While it can prevent overwhelming the system, it is not a requirement for fault tolerance or exactly-once processing.

  • ✓

    Use a checkpoint location to store progress information.

    Why this is correct

    A checkpoint location is essential for fault tolerance and exactly-once semantics. It stores the progress of the streaming query, including offsets processed and state information. On restart, the query resumes from where it left off, ensuring no data is lost or duplicated. Without a checkpoint, the job cannot recover correctly.

  • ✓

    Use a Delta table as the sink with idempotent writes.

    Why this is correct

    Delta Lake supports idempotent writes through its transaction log, which allows the streaming query to write data exactly once. When combined with checkpointing, Delta ensures that each micro-batch is written atomically and that reprocessing a batch does not duplicate data. This is crucial for exactly-once semantics.

About these practice questions

Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.