Be able to choose the right ingestion mechanism for a source, configure Auto Loader or DLT correctly, and land data in Delta. The single most important thing: mask or tokenize sensitive identifiers before they reach Bronze storage, not after.
Start practicing
Data Ingestion and Acquisition — choose a session length
Free · No account required
Domain overview
This domain covers acquiring data into Databricks from files, message buses, and external systems, then landing it in Bronze. Expect questions on Delta Live Tables versus hand-rolled Structured Streaming, Auto Loader ingestion modes, Delta Lake as a sink, and enforcing PII masking during ingestion.
Exam objectives
Auto Loader directory listing versus file notification modes for incremental cloud file ingestion
Delta Live Tables expectations, streaming tables, and medallion Bronze/Silver/Gold pipeline declarations
Structured Streaming sources and sinks including Kafka, Event Hubs, and Delta Lake
Unity Catalog and column masking or row filter policies applied to sensitive ingestion data
Assuming DLT is just syntactic sugar over Structured Streaming, missing managed orchestration, expectations, and lineage benefits
Choosing file notification mode by default when directory listing is simpler and sufficient for low-volume or irregular bursts
Masking PII only in Silver or Gold, leaving raw identifiers persisted in Bronze tables and files
Click any question to see the full explanation and answer options, or start a focused practice session above.
Refer to the exhibit. You are using Auto Loader to ingest data with evolving schemas. After running the job for a week, you realize that new columns added to the source JSON are not being captured in the destination table. What must you add to the configuration?
2Which approach is most appropriate for ingesting data from a JDBC source into Delta Lake where the source table has no 'updated_at' or 'version' column for incremental loading?
3Your organization is ingesting sensitive PII data. You need to ensure that personal identifiers are masked during the ingestion process before they are stored in the Bronze layer of your Medallion architecture. What is the best practice for this?
4Which TWO of the following are primary benefits of using Delta Live Tables (DLT) for data ingestion over standard Structured Streaming pipelines?
5When ingesting data using Auto Loader, what is the purpose of the 'cloudFiles.schemaLocation' parameter?
6Which THREE of the following are essential components of an effective ingestion monitoring strategy in Databricks?
7When ingesting data from a Kafka topic into Delta Lake, what is the best way to handle out-of-order data arriving in the stream?
8You are ingesting data from multiple source systems with varying file formats (JSON, CSV, Parquet) into a centralized Bronze landing zone. Which architecture pattern is the most scalable for maintaining this ingestion layer?
9You are designing an ingestion pipeline that must handle massive bursts of data at irregular intervals. Which feature should you prioritize to ensure the ingestion process remains cost-effective?
10What is the primary advantage of using Delta Lake as the sink for your data ingestion pipelines compared to raw Parquet files?
11A data engineer is building a streaming ingestion pipeline using Databricks Auto Loader to ingest JSON files from cloud storage into a Delta Bronze table. The pipeline must handle schema evolution without failing and must minimize the number of files that require reprocessing when the schema changes. The engineer wants to configure Auto Loader appropriately. Which two configuration settings should be used to achieve these requirements? (Choose two.)
12A data engineer is building a streaming ingestion pipeline from Apache Kafka to a Delta table. The pipeline must perform deduplication on a unique event_id field and handle late-arriving data. The engineer wants to use Structured Streaming with a watermark of 10 minutes. Which of the following approaches correctly implements deduplication and watermarking?
13A data engineer is ingesting data from an Apache Kafka topic into a Delta Lake table using Structured Streaming. The Kafka topic receives messages with a timestamp field in the value payload, but the ingestion must handle late-arriving data and produce correct aggregations. The engineer wants to ensure that watermarks are applied correctly. Which approach should be used?
14A data engineer is configuring an Auto Loader stream to ingest JSON files from an S3 bucket into a Bronze Delta table. The source bucket contains both .json and .json.gz files, and the engineer wants to ensure that only .json files are processed. Which parameter should be set to achieve this?
15A data engineer is ingesting data from an Apache Kafka topic into a Delta Lake table using Structured Streaming. The Kafka topic receives messages with a timestamp field in the value payload, but the messages can arrive out of order by up to 10 minutes. The engineer wants to perform time-windowed aggregations on the ingested data while minimizing state store overhead. Which approach should be used to handle the out-of-order data correctly?
16A data engineer is using Databricks Auto Loader to ingest CSV files into a Delta table. The engineer notices that some files have a different delimiter (semicolon instead of comma). Which option should be used to handle this variation?
17A data engineer is using Auto Loader to ingest JSON files from cloud storage into a Delta table. The files contain a nested field 'address' with subfields 'city' and 'zip'. The engineer wants to flatten the nested structure during ingestion so that 'city' and 'zip' become top-level columns in the Bronze table. Which Auto Loader feature should be used to achieve this?
18A data engineer is ingesting streaming data from Apache Kafka into a Delta table using Databricks Structured Streaming. The engineer wants to ensure exactly-once processing and handle late-arriving data. Which combination of features should the engineer use?
19A data engineer is ingesting data from an Azure SQL Database into a Delta Lake table using the JDBC connector in a Databricks notebook. The source table contains millions of rows, and the engineer wants to optimize the ingestion by reading the data in parallel. The source table has a numeric primary key column named 'id' that is evenly distributed. Which approach should the engineer use to achieve parallel reads?
Be able to choose the right ingestion mechanism for a source, configure Auto Loader or DLT correctly, and land data in Delta. The single most important thing: mask or tokenize sensitive identifiers before they reach Bronze storage, not after.
The Courseiva Databricks-DE-Pro question bank contains 19 questions in the Data Ingestion and Acquisition domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Data Ingestion and Acquisition domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included