Courseiva

Databricks-DE-Assoc · domain

Data Ingestion and Loading

This domain covers getting data into Delta Lake on Databricks: Auto Loader, COPY INTO, and streaming ingestion patterns. Questions present realistic ingestion scenarios and ask you to choose the right mode, trigger, schema-evolution behavior, or rescue-data handling for CSV, JSON, and partitioned cloud storage paths.

23 questions2 easy12 medium9 hard

Focused practice

Practice Data Ingestion and Loading questions

Scored sessions drawing only from this domain — pick a length below.

What this domain covers

What to know about Data Ingestion and Loading

Be able to pick Auto Loader versus COPY INTO, choose the correct trigger for batch-style incremental runs, and configure schema inference, evolution, and rescued data. The most important thing: match the ingestion mode and trigger to the stated cost and freshness requirement.

Auto Loader directory listing versus file notification mode with cloud notification services

Rescued data column for capturing malformed records during CSV and JSON ingestion

Schema inference and schema evolution from date-partitioned paths in cloud object storage

Trigger.AvailableNow for incremental jobs that process all available data then stop

Watch out for

Common Data Ingestion and Loading exam traps

  • ▸Assuming file notification mode needs no cloud setup; it requires configuring notification services and IAM permissions on the storage account.
  • ▸Forgetting that malformed records land in the rescued data column only when rescue mode is enabled, not by default in every read path.
  • ▸Confusing Trigger.Once with Trigger.AvailableNow, or using continuous triggers when the requirement is to process all data and shut down.

Question index

All Data Ingestion and Loading questions (23)

Click any question to see the full explanation, or start a practice session above.

1

A data engineer wants to run an incremental ingestion job every six hours. They want to ensure that each run processes all available data and then shuts down the cluster to save costs. Which Trigger should be used in the Structured Streaming code?

Medium
2

Refer to the exhibit. A data engineer is configuring an Auto Loader stream with the provided options. What will happen if a new JSON file arrives containing a field that is not currently in the target table schema?

Medium
3

When ingesting data from cloud storage, what is the most important reason to use a Service Principal or IAM Role instead of individual user credentials?

Medium
4

A data engineer is using Auto Loader to ingest data from a Kafka topic into a Delta table. The engineer wants to ensure that the ingestion handles late-arriving data and provides exactly-once semantics. Which combination of features should the engineer use?

Hard
5

A data engineer is using Auto Loader to stream data from Kafka into a Delta table. The Kafka topic receives messages in Avro format, and the schema is stored in a Confluent Schema Registry. The engineer wants Auto Loader to automatically fetch the schema from the registry and evolve it as new versions are registered. Which configuration should the engineer use?

Hard
6

A data engineer is configuring a Databricks Auto Loader stream to ingest CSV files from a cloud storage location into a Delta table. The CSV files have a header row, and the engineer wants to automatically infer the schema and store the inferred schema in a specified location for consistency across restarts. Which Auto Loader option should be used to persist the inferred schema?

Medium
7

A data engineer is designing a pipeline using Structured Streaming to ingest data into Delta Lake. Which THREE benefits are provided by using checkpoints in this scenario?

Medium
8

A data engineer is using Auto Loader and wants to handle a situation where a column 'user_id' is sometimes an integer and sometimes a string in the source JSON files. Which TWO strategies can be used to manage this schema conflict?

Hard
9

A data engineer needs to ingest a large CSV file from cloud storage into a Delta table using Databricks SQL. The engineer wants to perform a one-time load and ensure that the operation is atomic. Which command should be used?

Easy
10

A data engineer is using Databricks Auto Loader to stream CSV files from an ADLS Gen2 container into a Delta table. The source directory contains a mix of files, but only files with the prefix 'sales_' should be ingested. The engineer wants Auto Loader to ignore all other files without moving or deleting them. Which Auto Loader option should the engineer configure to achieve this?

Medium
11

Refer to the exhibit. A data engineer is using this command to reload data into a Delta table after a schema correction. What is the primary effect of setting the 'force' option to 'true' in this context?

Medium
12

When using Auto Loader to ingest data, why is it recommended to provide a 'schemaLocation' rather than manually defining the entire schema in the code?

Easy
13

A data engineer is using Databricks Auto Loader to ingest JSON files from a directory. The stream is configured with `cloudFiles.schemaLocation` set to a specific path. The engineer notices that the stream fails when a new file contains an additional column. What is the most likely reason for the failure?

Hard
14

A data engineer is designing an ingestion pipeline using Databricks Auto Loader to process JSON files from an S3 bucket. The pipeline must handle schema evolution and ensure that data is ingested exactly once. Which two features of Auto Loader support these requirements? (Choose two.)

Hard
15

A data engineer is using Auto Loader to ingest data from a directory that receives thousands of files every hour. They are considering switching from the default directory listing mode to file notification mode. What is the primary reason for making this change?

Hard
16

While ingesting CSV files using Auto Loader, a data engineer notices that some records have malformed data that does not match the inferred schema. How can the engineer capture these records without failing the entire ingestion stream?

Medium
17

A data engineer is using Databricks Auto Loader to ingest JSON files from an Azure Data Lake Storage Gen2 container into a Delta table. The engineer notices that the ingestion is slow and wants to optimize file discovery. The directory contains millions of files, and new files are added frequently. Which Auto Loader option should be used to improve file discovery performance?

Hard
18

A data engineer is using Auto Loader to ingest files from an S3 bucket into a Delta table. The files are partitioned by date in the path, e.g., s3://bucket/data/2023-01-01/file1.json. The engineer wants to automatically add a column 'date' to the ingested data based on the file path. Which Auto Loader feature should be used?

Medium
19

A data engineer is evaluating whether to use Auto Loader or the COPY INTO command for a new ingestion pipeline. Which TWO features are unique to Auto Loader compared to COPY INTO?

Hard
20

A data engineer is using Auto Loader to ingest JSON files from a cloud storage directory into a Delta table. The directory receives files continuously, and the engineer wants to ensure that the ingestion process can handle schema drift where new columns are added to the JSON files over time. The engineer also wants to minimize the need to reprocess all data when the schema changes. Which configuration should the engineer use?

Hard
21

A data engineer needs to ingest data from a legacy SQL Server database into a Delta Lake bronze table. The ingestion must be performant and support parallel reads from the source table. What is the best practice for configuring the JDBC connection in this scenario?

Medium
22

When designing an ingestion strategy for a Delta Lake architecture, which TWO advantages does Delta Lake provide over traditional Parquet tables for incoming data?

Medium
23

A data engineering team needs to ingest millions of small JSON files from an S3 bucket into a Delta Lake table. The solution must provide incremental loading, support schema evolution, and automatically scale to handle increasing file volumes without manual tracking of processed files. Which tool is best suited for this requirement?

Medium

Frequently asked questions

What does the Data Ingestion and Loading domain cover on the Databricks-DE-Assoc exam?
Be able to pick Auto Loader versus COPY INTO, choose the correct trigger for batch-style incremental runs, and configure schema inference, evolution, and rescued data. The most important thing: match the ingestion mode and trigger to the stated cost and freshness requirement.
How many questions are in this domain?
This page lists all 23 Data Ingestion and Loading questions in the Databricks-DE-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Data Ingestion and Loading questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
databricks-data-engineer-associate DATABRICKS-DATA-ENGINEER-ASSOCIATE data ingestion loading Practice Questions