Courseiva

Databricks-DE-Pro Data Ingestion and Acquisition Practice Question

A data engineer is building a streaming ingestion pipeline using Databricks Auto Loader to ingest JSON files from cloud storage into a Delta Bronze table. The pipeline must handle schema evolution without failing and must minimize the number of files that require reprocessing when the schema changes. The engineer wants to configure Auto Loader appropriately. Which two configuration settings should be used to achieve these requirements? (Choose two.)

⚠ Common exam trap

The trap here is assuming that enabling file notification or type inference alone will handle schema evolution and reduce reprocessing, when in fact schema evolution mode and schema location are the critical settings.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set cloudFiles.schemaLocation to a dedicated directory in cloud storage.

To handle schema evolution without failing and minimize reprocessing, Auto Loader must be configured to add new columns dynamically and to store schema metadata persistently. Setting the schema evolution mode to addNewColumns ensures new fields are incorporated, while specifying a schema location allows Auto Loader to track schema changes and avoid reprocessing already-ingested files. Together, these settings provide robust schema evolution with efficient incremental processing.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Set cloudFiles.schemaLocation to a dedicated directory in cloud storage.

    Why this is correct

    The cloudFiles.schemaLocation specifies where Auto Loader stores the inferred schema and metadata about schema evolution. By providing a persistent location, Auto Loader can track schema changes over time and avoid reprocessing files that were already ingested with an older schema. This is essential for minimizing reprocessing and ensuring consistent schema management across pipeline restarts.

  • ✗

    Set cloudFiles.inferColumnTypes to true.

    Why it's wrong here

    Setting cloudFiles.inferColumnTypes to true enables Auto Loader to infer column data types more accurately, but it does not control schema evolution mode or minimize reprocessing. While it can be useful for initial schema inference, it does not handle new columns automatically nor does it persist schema information across runs. Therefore, it does not satisfy the need for schema evolution without reprocessing.

  • ✗

    Set cloudFiles.useNotifications to true.

    Why it's wrong here

    Enabling cloudFiles.useNotifications configures Auto Loader to use cloud provider notification services (e.g., AWS SNS/SQS, Azure Event Grid) to detect new files more efficiently. While this improves file discovery performance and reduces listing costs, it does not affect schema evolution behavior or the reprocessing of files when the schema changes. Therefore, it does not meet the stated requirements for schema handling.

  • ✓

    Set cloudFiles.schemaEvolutionMode to 'addNewColumns'.

    Why this is correct

    Setting cloudFiles.schemaEvolutionMode to 'addNewColumns' allows Auto Loader to automatically add new columns to the schema as they appear in the data, without failing the stream. This directly supports schema evolution. It also ensures that only new columns are added, and existing data remains intact, which is the desired behavior for a Bronze ingestion layer that must handle evolving JSON structures without manual intervention.

  • ✗

    Set cloudFiles.format to 'parquet'.

    Why it's wrong here

    Setting cloudFiles.format to 'parquet' would tell Auto Loader to expect Parquet files, but the scenario specifies JSON files. This would cause ingestion to fail or misinterpret the data. The format setting is about the source file type, not schema evolution or reprocessing. Thus, it is incorrect for this scenario and does not address the requirements.

About these practice questions

This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.