Courseiva
Structured Streaming →mediumMultiple Select

Databricks-Spark-Assoc Structured Streaming Practice Question

A developer is building a Structured Streaming job that reads from a Delta table as a stream and writes to another Delta table. The job must support exactly-once processing and allow the output to be updated incrementally. Which two options are required to achieve exactly-once semantics? (Choose two.)

⚠ Common exam trap

The trap here is thinking that a specific output mode or trigger setting alone can guarantee exactly-once, when checkpointing and sink idempotency are the real requirements.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Ensure the sink is idempotent or transactional, such as Delta Lake.

Exactly-once in Structured Streaming relies on two pillars: a reliable checkpoint to track progress and an idempotent or transactional sink to handle replays. A checkpoint location stores offsets and state, enabling recovery without data loss. A sink like Delta Lake ensures that re-executed batches do not duplicate output. Together, they provide end-to-end exactly-once. Other options affect output mode, custom sink logic, or trigger frequency, none of which guarantee exactly-once.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set the output mode to `complete`.

    Why it's wrong here

    Output mode controls how results are written, not the delivery guarantee. `Complete` mode rewrites the entire result table each time, which is inefficient and not related to exactly-once semantics. Exactly-once is achieved through checkpointing and idempotent sinks. Therefore this option does not contribute to the requirement and is incorrect.

  • ✓

    Ensure the sink is idempotent or transactional, such as Delta Lake.

    Why this is correct

    Exactly-once semantics require that the sink can handle replays without duplicating data. Delta Lake provides transactional writes and idempotent merges, which allow the streaming query to retry a micro-batch without creating duplicates. In this scenario, writing to a Delta table as the sink ensures that even if a batch is reprocessed after failure, the result remains consistent, thus satisfying exactly-once.

  • ✓

    Configure a checkpoint location for the query.

    Why this is correct

    A checkpoint location is essential for exactly-once semantics because it stores the progress information, including which offsets have been processed and the state for stateful operations. Without a checkpoint, the query cannot recover from failures and may reprocess data. In this scenario, setting a checkpoint location ensures that the streaming query can resume from where it left off, providing exactly-once guarantees.

  • ✗

    Use the `foreachBatch` sink to write to the Delta table.

    Why it's wrong here

    While `foreachBatch` can be used for custom logic, it does not by itself guarantee exactly-once semantics. The sink must be idempotent or transactional. Delta Lake's native sink already provides exactly-once when used with a checkpoint. Using `foreachBatch` adds complexity and requires manual handling of idempotency, so it is not a required option for exactly-once in this scenario.

  • ✗

    Use `trigger(processingTime='0 seconds')` for continuous processing.

    Why it's wrong here

    The trigger interval controls how often micro-batches are executed, not the delivery guarantee. Continuous processing mode is experimental and does not support all operations. It does not provide exactly-once by itself. Therefore this option is not required and does not help achieve the stated goal. It is a distractor related to performance tuning rather than fault tolerance.

About these practice questions

This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.