Which component in Structured Streaming is responsible for providing fault tolerance and ensuring data is processed exactly once?
Checkpointing records the metadata and offsets of the micro-batches in a persistent store. This allows the streaming query to recover from failures and resume processing exactly where it left off, which is the cornerstone of providing the exactly-once fault-tolerance guarantees expected in enterprise data pipelines.
Why this answer
Checkpointing is the mechanism that stores the query's metadata and progress in durable storage (like DBFS or S3). By recording the offset of the processed data, Spark can recover from failures and restart from the exact point where it left off. This is fundamental to ensuring that streaming applications are reliable and maintain exactly-once processing guarantees across restarts or cluster crashes.
Exam trap
Candidates often confuse checkpointing with logging. They think checkpointing is just for debugging output, whereas it is actually the critical mechanism for state recovery and exactly-once processing guarantees.