Databricks-DE-Assoc · domain
Data Ingestion and Loading
This domain covers getting data into Delta Lake on Databricks: Auto Loader, COPY INTO, and streaming ingestion patterns. Questions present realistic ingestion scenarios and ask you to choose the right mode, trigger, schema-evolution behavior, or rescue-data handling for CSV, JSON, and partitioned cloud storage paths.
Focused practice
Practice Data Ingestion and Loading questions
Scored sessions drawing only from this domain — pick a length below.
What this domain covers
What to know about Data Ingestion and Loading
Be able to pick Auto Loader versus COPY INTO, choose the correct trigger for batch-style incremental runs, and configure schema inference, evolution, and rescued data. The most important thing: match the ingestion mode and trigger to the stated cost and freshness requirement.
Auto Loader directory listing versus file notification mode with cloud notification services
Rescued data column for capturing malformed records during CSV and JSON ingestion
Schema inference and schema evolution from date-partitioned paths in cloud object storage
Trigger.AvailableNow for incremental jobs that process all available data then stop
Watch out for
Common Data Ingestion and Loading exam traps
- ▸Assuming file notification mode needs no cloud setup; it requires configuring notification services and IAM permissions on the storage account.
- ▸Forgetting that malformed records land in the rescued data column only when rescue mode is enabled, not by default in every read path.
- ▸Confusing Trigger.Once with Trigger.AvailableNow, or using continuous triggers when the requirement is to process all data and shut down.
Question index
All Data Ingestion and Loading questions (23)
Click any question to see the full explanation, or start a practice session above.
A data engineer wants to run an incremental ingestion job every six hours. They want to ensure that each run processes all available data and then shuts down the cluster to save costs. Which Trigger should be used in the Structured Streaming code?
Medium2Refer to the exhibit. A data engineer is configuring an Auto Loader stream with the provided options. What will happen if a new JSON file arrives containing a field that is not currently in the target table schema?
Medium3When ingesting data from cloud storage, what is the most important reason to use a Service Principal or IAM Role instead of individual user credentials?
Medium4A data engineer is using Auto Loader to ingest data from a Kafka topic into a Delta table. The engineer wants to ensure that the ingestion handles late-arriving data and provides exactly-once semantics. Which combination of features should the engineer use?
Hard5A data engineer is using Auto Loader to stream data from Kafka into a Delta table. The Kafka topic receives messages in Avro format, and the schema is stored in a Confluent Schema Registry. The engineer wants Auto Loader to automatically fetch the schema from the registry and evolve it as new versions are registered. Which configuration should the engineer use?
Hard6A data engineer is configuring a Databricks Auto Loader stream to ingest CSV files from a cloud storage location into a Delta table. The CSV files have a header row, and the engineer wants to automatically infer the schema and store the inferred schema in a specified location for consistency across restarts. Which Auto Loader option should be used to persist the inferred schema?
Medium7A data engineer is designing a pipeline using Structured Streaming to ingest data into Delta Lake. Which THREE benefits are provided by using checkpoints in this scenario?
Medium8A data engineer is using Auto Loader and wants to handle a situation where a column 'user_id' is sometimes an integer and sometimes a string in the source JSON files. Which TWO strategies can be used to manage this schema conflict?
Hard9A data engineer needs to ingest a large CSV file from cloud storage into a Delta table using Databricks SQL. The engineer wants to perform a one-time load and ensure that the operation is atomic. Which command should be used?
Easy10A data engineer is using Databricks Auto Loader to stream CSV files from an ADLS Gen2 container into a Delta table. The source directory contains a mix of files, but only files with the prefix 'sales_' should be ingested. The engineer wants Auto Loader to ignore all other files without moving or deleting them. Which Auto Loader option should the engineer configure to achieve this?
Medium11Refer to the exhibit. A data engineer is using this command to reload data into a Delta table after a schema correction. What is the primary effect of setting the 'force' option to 'true' in this context?
Medium12When using Auto Loader to ingest data, why is it recommended to provide a 'schemaLocation' rather than manually defining the entire schema in the code?
Easy13A data engineer is using Databricks Auto Loader to ingest JSON files from a directory. The stream is configured with `cloudFiles.schemaLocation` set to a specific path. The engineer notices that the stream fails when a new file contains an additional column. What is the most likely reason for the failure?
Hard14A data engineer is designing an ingestion pipeline using Databricks Auto Loader to process JSON files from an S3 bucket. The pipeline must handle schema evolution and ensure that data is ingested exactly once. Which two features of Auto Loader support these requirements? (Choose two.)
Hard15A data engineer is using Auto Loader to ingest data from a directory that receives thousands of files every hour. They are considering switching from the default directory listing mode to file notification mode. What is the primary reason for making this change?
Hard16While ingesting CSV files using Auto Loader, a data engineer notices that some records have malformed data that does not match the inferred schema. How can the engineer capture these records without failing the entire ingestion stream?
Medium17A data engineer is using Databricks Auto Loader to ingest JSON files from an Azure Data Lake Storage Gen2 container into a Delta table. The engineer notices that the ingestion is slow and wants to optimize file discovery. The directory contains millions of files, and new files are added frequently. Which Auto Loader option should be used to improve file discovery performance?
Hard18A data engineer is using Auto Loader to ingest files from an S3 bucket into a Delta table. The files are partitioned by date in the path, e.g., s3://bucket/data/2023-01-01/file1.json. The engineer wants to automatically add a column 'date' to the ingested data based on the file path. Which Auto Loader feature should be used?
Medium19A data engineer is evaluating whether to use Auto Loader or the COPY INTO command for a new ingestion pipeline. Which TWO features are unique to Auto Loader compared to COPY INTO?
Hard20A data engineer is using Auto Loader to ingest JSON files from a cloud storage directory into a Delta table. The directory receives files continuously, and the engineer wants to ensure that the ingestion process can handle schema drift where new columns are added to the JSON files over time. The engineer also wants to minimize the need to reprocess all data when the schema changes. Which configuration should the engineer use?
Hard21A data engineer needs to ingest data from a legacy SQL Server database into a Delta Lake bronze table. The ingestion must be performant and support parallel reads from the source table. What is the best practice for configuring the JDBC connection in this scenario?
Medium22When designing an ingestion strategy for a Delta Lake architecture, which TWO advantages does Delta Lake provide over traditional Parquet tables for incoming data?
Medium23A data engineering team needs to ingest millions of small JSON files from an S3 bucket into a Delta Lake table. The solution must provide incremental loading, support schema evolution, and automatically scale to handle increasing file volumes without manual tracking of processed files. Which tool is best suited for this requirement?
MediumOther domains
All Databricks-DE-Assoc exam domains
Frequently asked questions
- What does the Data Ingestion and Loading domain cover on the Databricks-DE-Assoc exam?
- Be able to pick Auto Loader versus COPY INTO, choose the correct trigger for batch-style incremental runs, and configure schema inference, evolution, and rescued data. The most important thing: match the ingestion mode and trigger to the stated cost and freshness requirement.
- How many questions are in this domain?
- This page lists all 23 Data Ingestion and Loading questions in the Databricks-DE-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Data Ingestion and Loading questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.