Be able to pick Auto Loader versus COPY INTO, choose the correct trigger for batch-style incremental runs, and configure schema inference, evolution, and rescued data. The most important thing: match the ingestion mode and trigger to the stated cost and freshness requirement.
Start practicing
Data Ingestion and Loading — choose a session length
Free · No account required
Domain overview
This domain covers getting data into Delta Lake on Databricks: Auto Loader, COPY INTO, and streaming ingestion patterns. Questions present realistic ingestion scenarios and ask you to choose the right mode, trigger, schema-evolution behavior, or rescue-data handling for CSV, JSON, and partitioned cloud storage paths.
Exam objectives
Auto Loader directory listing versus file notification mode with cloud notification services
Rescued data column for capturing malformed records during CSV and JSON ingestion
Schema inference and schema evolution from date-partitioned paths in cloud object storage
Trigger.AvailableNow for incremental jobs that process all available data then stop
Assuming file notification mode needs no cloud setup; it requires configuring notification services and IAM permissions on the storage account.
Forgetting that malformed records land in the rescued data column only when rescue mode is enabled, not by default in every read path.
Confusing Trigger.Once with Trigger.AvailableNow, or using continuous triggers when the requirement is to process all data and shut down.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A data engineering team needs to ingest millions of small JSON files from an S3 bucket into a Delta Lake table. The solution must provide incremental loading, support schema evolution, and automatically scale to handle increasing file volumes without manual tracking of processed files. Which tool is best suited for this requirement?
2A data engineer is evaluating whether to use Auto Loader or the COPY INTO command for a new ingestion pipeline. Which TWO features are unique to Auto Loader compared to COPY INTO?
3Refer to the exhibit. A data engineer is configuring an Auto Loader stream with the provided options. What will happen if a new JSON file arrives containing a field that is not currently in the target table schema?
4When using Auto Loader to ingest data, why is it recommended to provide a 'schemaLocation' rather than manually defining the entire schema in the code?
5A data engineer is designing a pipeline using Structured Streaming to ingest data into Delta Lake. Which THREE benefits are provided by using checkpoints in this scenario?
6While ingesting CSV files using Auto Loader, a data engineer notices that some records have malformed data that does not match the inferred schema. How can the engineer capture these records without failing the entire ingestion stream?
7A data engineer needs to ingest data from a legacy SQL Server database into a Delta Lake bronze table. The ingestion must be performant and support parallel reads from the source table. What is the best practice for configuring the JDBC connection in this scenario?
8Refer to the exhibit. A data engineer is using this command to reload data into a Delta table after a schema correction. What is the primary effect of setting the 'force' option to 'true' in this context?
9When designing an ingestion strategy for a Delta Lake architecture, which TWO advantages does Delta Lake provide over traditional Parquet tables for incoming data?
10A data engineer is using Auto Loader to ingest data from a directory that receives thousands of files every hour. They are considering switching from the default directory listing mode to file notification mode. What is the primary reason for making this change?
11A data engineer wants to run an incremental ingestion job every six hours. They want to ensure that each run processes all available data and then shuts down the cluster to save costs. Which Trigger should be used in the Structured Streaming code?
12A data engineer is using Auto Loader and wants to handle a situation where a column 'user_id' is sometimes an integer and sometimes a string in the source JSON files. Which TWO strategies can be used to manage this schema conflict?
13When ingesting data from cloud storage, what is the most important reason to use a Service Principal or IAM Role instead of individual user credentials?
14A data engineer is configuring a Databricks Auto Loader stream to ingest CSV files from a cloud storage location into a Delta table. The CSV files have a header row, and the engineer wants to automatically infer the schema and store the inferred schema in a specified location for consistency across restarts. Which Auto Loader option should be used to persist the inferred schema?
15A data engineer is using Databricks Auto Loader to stream CSV files from an ADLS Gen2 container into a Delta table. The source directory contains a mix of files, but only files with the prefix 'sales_' should be ingested. The engineer wants Auto Loader to ignore all other files without moving or deleting them. Which Auto Loader option should the engineer configure to achieve this?
16A data engineer is using Databricks Auto Loader to ingest JSON files from an Azure Data Lake Storage Gen2 container into a Delta table. The engineer notices that the ingestion is slow and wants to optimize file discovery. The directory contains millions of files, and new files are added frequently. Which Auto Loader option should be used to improve file discovery performance?
17A data engineer needs to ingest a large CSV file from cloud storage into a Delta table using Databricks SQL. The engineer wants to perform a one-time load and ensure that the operation is atomic. Which command should be used?
18A data engineer is using Auto Loader to ingest JSON files from a cloud storage directory into a Delta table. The directory receives files continuously, and the engineer wants to ensure that the ingestion process can handle schema drift where new columns are added to the JSON files over time. The engineer also wants to minimize the need to reprocess all data when the schema changes. Which configuration should the engineer use?
19A data engineer is using Databricks Auto Loader to ingest JSON files from a directory. The stream is configured with `cloudFiles.schemaLocation` set to a specific path. The engineer notices that the stream fails when a new file contains an additional column. What is the most likely reason for the failure?
20A data engineer is designing an ingestion pipeline using Databricks Auto Loader to process JSON files from an S3 bucket. The pipeline must handle schema evolution and ensure that data is ingested exactly once. Which two features of Auto Loader support these requirements? (Choose two.)
21A data engineer is using Auto Loader to ingest files from an S3 bucket into a Delta table. The files are partitioned by date in the path, e.g., s3://bucket/data/2023-01-01/file1.json. The engineer wants to automatically add a column 'date' to the ingested data based on the file path. Which Auto Loader feature should be used?
22A data engineer is using Auto Loader to ingest data from a Kafka topic into a Delta table. The engineer wants to ensure that the ingestion handles late-arriving data and provides exactly-once semantics. Which combination of features should the engineer use?
23A data engineer is using Auto Loader to stream data from Kafka into a Delta table. The Kafka topic receives messages in Avro format, and the schema is stored in a Confluent Schema Registry. The engineer wants Auto Loader to automatically fetch the schema from the registry and evolve it as new versions are registered. Which configuration should the engineer use?
Be able to pick Auto Loader versus COPY INTO, choose the correct trigger for batch-style incremental runs, and configure schema inference, evolution, and rescued data. The most important thing: match the ingestion mode and trigger to the stated cost and freshness requirement.
The Courseiva Databricks-DE-Assoc question bank contains 23 questions in the Data Ingestion and Loading domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Data Ingestion and Loading domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included