Courseiva
Data Ingestion and Loading →mediumMultiple Choice

Databricks-DE-Assoc Data Ingestion and Loading Practice Question

A data engineer is using Databricks Auto Loader to stream CSV files from an ADLS Gen2 container into a Delta table. The source directory contains a mix of files, but only files with the prefix 'sales_' should be ingested. The engineer wants Auto Loader to ignore all other files without moving or deleting them. Which Auto Loader option should the engineer configure to achieve this?

⚠ Common exam trap

Many candidates confuse options that control file discovery or schema handling with those that filter based on file paths, leading to ingestion of unwanted files.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

cloudFiles.pathGlobFilter

Auto Loader provides the cloudFiles.pathGlobFilter option to filter ingested files using glob patterns. By setting it to 'sales_*', only files with the desired prefix are processed, and others are ignored without being moved or deleted. This directly satisfies the requirement to ingest a subset of files based on filename.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    cloudFiles.format

    Why it's wrong here

    This option tells Auto Loader the format of the source files, such as csv, json, or parquet. It is required for parsing but does not filter files. All files in the directory, regardless of prefix, would be read as long as they match the format. It does not help exclude files that lack the 'sales_' prefix.

  • ✗

    cloudFiles.schemaLocation

    Why it's wrong here

    This option specifies the directory where Auto Loader stores inferred schema information and metadata. It is used to persist schema evolution and does not filter files by name or path. Configuring this option would not prevent non-matching files from being ingested, so it fails to meet the filtering requirement.

  • ✗

    cloudFiles.includeExistingFiles

    Why it's wrong here

    This option controls whether Auto Loader processes files that already exist in the source directory when the stream starts. Setting it to true would ingest all existing files regardless of their names, which does not filter by prefix. It does not provide any path-based filtering, so files without the 'sales_' prefix would still be picked up, violating the requirement.

  • ✓

    cloudFiles.pathGlobFilter

    Why this is correct

    This option accepts a glob pattern to filter files based on their path. Setting it to 'sales_*' ensures Auto Loader only ingests files whose names start with 'sales_'. Other files are ignored without being moved or deleted. It is the correct way to apply a filename prefix filter during Auto Loader ingestion.

About these practice questions

Courseiva writes every Databricks-DE-Assoc question from scratch — 276 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.