Courseiva
Structured Streaming →hardMultiple Select

Databricks-Spark-Assoc Structured Streaming Practice Question

An analytics engineer is designing a Structured Streaming pipeline that reads from an Apache Kafka topic and writes JSON-formatted output to cloud object storage. The pipeline must maintain exactly-once processing semantics and support automatic recovery from cluster restarts. Which TWO actions are mandatory to achieve these requirements? (Choose 2)

⚠ Common exam trap

Candidates often forget that exactly-once semantics requires BOTH a transactional/idempotent sink AND a persistent checkpoint location, selecting only one of these mandatory components.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure a valid checkpoint location using the .option("checkpointLocation", path) method.

Achieving end-to-end exactly-once guarantees in Spark Structured Streaming requires idempotent or transactional sinks combined with persistent checkpoint directories. The checkpoint mechanism saves the exact state and offsets, allowing the streaming query to resume seamlessly after an interruption without duplicating data processing.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Configure a valid checkpoint location using the .option("checkpointLocation", path) method.

    Why this is correct

    A checkpoint location records Kafka offsets and commit metadata durably, letting Structured Streaming resume exactly where it stopped after a restart. Without it, the query cannot recover progress, so exactly-once delivery and automatic restart recovery are impossible.

  • ✗

    Enable Delta Lake format as the sink to leverage ACID transactions and transactional metadata.

    Why it's wrong here

    Delta Lake is not mandatory: exactly-once and restart recovery come from checkpointing plus a replayable source and idempotent sink, and JSON files in object storage can satisfy this with the right write mode. Delta is tempting because its transaction log gives ACID guarantees, which is the correct choice when the sink itself must support concurrent updates or merges.

  • ✗

    Set the output mode of the streaming query to complete mode.

    Why it's wrong here

    Complete mode rewrites the entire result table on every trigger, which is unsuitable for append-only append sinks or large-scale aggregations. The output mode depends entirely on business logic rather than being a prerequisite for processing guarantees.

  • ✓

    Use an idempotent or transactional sink combined with proper source offset management.

    Why this is correct

    Exactly-once requires the sink to deduplicate replays and the source offsets to be committed atomically with the write. An idempotent or transactional sink paired with managed offsets prevents duplicate output when the query restarts and reprocesses.

  • ✗

    Increase the driver memory allocation to at least 64GB to store all Kafka offset metadata.

    Why it's wrong here

    Driver memory does not govern offset durability or recovery; offsets and progress are committed to a checkpoint location, which survives restarts regardless of heap size. Raising driver memory is tempting because large Kafka offset volumes can pressure the driver, but that is a tuning concern, not a requirement for exactly-once semantics.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.