Courseiva

Databricks-DE-Pro Developing Code (Python/SQL) Practice Question

When designing a streaming pipeline using Structured Streaming, which THREE of the following are necessary to ensure 'exactly-once' processing semantics in Databricks?

⚠ Common exam trap

Candidates often focus only on the sink and ignore the source. They forget that 'exactly-once' is impossible if the source system cannot replay data during a failure recovery scenario.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Using a source system that supports replaying data (e.g., Kafka or Delta).

Exactly-once processing requires that the source system, the processing engine, and the sink system all support fault-tolerant checkpoints and idempotency. Databricks handles this through checkpointing and write-ahead logs. Understanding these components is critical for data engineers to ensure data consistency in critical financial or operational systems, preventing the common pitfalls of duplicate entries or missed data during cluster restarts or transient failures.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Using a source system that supports replaying data (e.g., Kafka or Delta).

    Why this is correct

    Exactly-once requires the ability to re-read the stream from a specific offset if a failure occurs. Kafka and Delta provide the necessary offset management to ensure that data can be re-processed accurately after a system interruption, which is the foundational requirement for guaranteeing consistent results in streaming pipelines.

  • ✓

    Writing to the sink using an idempotent operation.

    Why this is correct

    Idempotency ensures that multiple attempts to write the same data result in the same final state in the destination. This is crucial for handling retries that occur during a failure; without idempotency, a retry could lead to duplicate rows, violating the exactly-once processing guarantee required by many downstream systems.

  • ✗

    Setting 'spark.sql.shuffle.partitions' to 1.

    Why it's wrong here

    Setting shuffle partitions to 1 destroys parallelism and is not related to exactly-once semantics. It would severely limit throughput and cause performance bottlenecks in any production streaming environment, regardless of the consistency guarantees provided by the underlying streaming engine or the data source storage system being used.

  • ✓

    Maintaining a checkpoint directory in a reliable storage location.

    Why this is correct

    The checkpoint directory stores the state and metadata of the streaming query. If the query fails, it uses this information to recover from the exact point of failure. This mechanism is mandatory for maintaining the state across failures and ensuring that progress tracking remains consistent throughout the pipeline lifecycle.

  • ✗

    Disabling the Delta Lake write-ahead log.

    Why it's wrong here

    Disabling the write-ahead log would make it impossible for Delta Lake to guarantee transactional integrity during commits. This would lead to data corruption or inconsistency if a failure occurred during a write operation, making it impossible to achieve exactly-once processing or reliable recovery in any production pipeline.

About these practice questions

Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.