Databricks-Spark-Assoc Structured Streaming Practice Question
You are designing a Structured Streaming job that must read from a file source and write to a Delta table. You want the job to be resilient to failures and to continue processing only new files after a restart. Which two actions should you take? (Choose two.)
⚠ Common exam trap
The trap here is focusing on performance-oriented file source options such as maxFilesPerTrigger or latestFirst, which tune behavior but do not persist progress or guarantee idempotent writes.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure a unique checkpoint location for the query.
Resilient incremental file ingestion requires two things: durable progress tracking and a sink that can commit batches atomically. A unique checkpoint location records which files and offsets have been processed, and a Delta sink provides transactional, idempotent writes that prevent duplicate effects when batches are retried. Throughput and ordering options do not affect recovery correctness.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Configure a unique checkpoint location for the query.
Why this is correct
The checkpoint location stores progress information, including which files or offsets have been processed and the state of stateful operations. On restart, the engine reads this metadata to resume exactly where it left off, preventing reprocessing of already-handled files. Without a checkpoint, the query cannot recover its position and will either fail to restart or reprocess from the beginning, so this is essential for resilient incremental processing.
- ✗
Enable the option latestFirst on the file source.
Why it's wrong here
latestFirst changes the order in which files are processed, prioritizing the most recently modified ones. It is useful for scenarios where recent data matters most, but it does not persist progress or guarantee exactly-once behavior. Progress tracking still depends on the checkpoint. Choosing this option affects ordering only and does not help the job resume correctly after a restart.
- ✗
Set the query to use trigger(once=True) so it stops after each run.
Why it's wrong here
trigger(once=True) processes all currently available data and then terminates the query. While this can be useful for scheduled batch-style ingestion, it does not by itself provide resilience or incremental continuation; that still depends on the checkpoint. It also prevents continuous processing. As a resilience measure it is irrelevant, and it may conflict with the goal of a continuously running streaming job.
- ✗
Set the file source option maxFilesPerTrigger to a high value to process all files at once.
Why it's wrong here
maxFilesPerTrigger controls how many files are included in a single micro-batch. Increasing it may speed up ingestion but does not provide fault tolerance or track progress across restarts. It also risks large batches and memory pressure. This option addresses throughput tuning, not resilience, so it does not contribute to resuming from the correct position after a failure.
- ✓
Use a Delta table sink so the write is idempotent and transactional.
Why this is correct
Delta Lake provides ACID transactions and idempotent writes when combined with the checkpoint metadata. If a batch is retried after a failure, the engine can use the checkpoint to determine which batches committed, and Delta's transaction log prevents duplicate application of the same batch. This pairing is what enables exactly-once semantics for the sink, making it a key part of a resilient pipeline.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.