Databricks-DE-Pro · domain
Data Ingestion and Acquisition
This domain covers acquiring data into Databricks from files, message buses, and external systems, then landing it in Bronze. Expect questions on Delta Live Tables versus hand-rolled Structured Streaming, Auto Loader ingestion modes, Delta Lake as a sink, and enforcing PII masking during ingestion.
Focused practice
Practice Data Ingestion and Acquisition questions
Scored sessions drawing only from this domain — pick a length below.
What this domain covers
What to know about Data Ingestion and Acquisition
Be able to choose the right ingestion mechanism for a source, configure Auto Loader or DLT correctly, and land data in Delta. The single most important thing: mask or tokenize sensitive identifiers before they reach Bronze storage, not after.
Auto Loader directory listing versus file notification modes for incremental cloud file ingestion
Delta Live Tables expectations, streaming tables, and medallion Bronze/Silver/Gold pipeline declarations
Structured Streaming sources and sinks including Kafka, Event Hubs, and Delta Lake
Unity Catalog and column masking or row filter policies applied to sensitive ingestion data
Watch out for
Common Data Ingestion and Acquisition exam traps
- ▸Assuming DLT is just syntactic sugar over Structured Streaming, missing managed orchestration, expectations, and lineage benefits
- ▸Choosing file notification mode by default when directory listing is simpler and sufficient for low-volume or irregular bursts
- ▸Masking PII only in Silver or Gold, leaving raw identifiers persisted in Bronze tables and files
Question index
All Data Ingestion and Acquisition questions (19)
Click any question to see the full explanation, or start a practice session above.
What is the primary advantage of using Delta Lake as the sink for your data ingestion pipelines compared to raw Parquet files?
Medium2You are ingesting data from multiple source systems with varying file formats (JSON, CSV, Parquet) into a centralized Bronze landing zone. Which architecture pattern is the most scalable for maintaining this ingestion layer?
Hard3A data engineer is building a streaming ingestion pipeline from Apache Kafka to a Delta table. The pipeline must perform deduplication on a unique event_id field and handle late-arriving data. The engineer wants to use Structured Streaming with a watermark of 10 minutes. Which of the following approaches correctly implements deduplication and watermarking?
Medium4A data engineer is ingesting data from an Apache Kafka topic into a Delta Lake table using Structured Streaming. The Kafka topic receives messages with a timestamp field in the value payload, but the ingestion must handle late-arriving data and produce correct aggregations. The engineer wants to ensure that watermarks are applied correctly. Which approach should be used?
Medium5A data engineer is building a streaming ingestion pipeline using Databricks Auto Loader to ingest JSON files from cloud storage into a Delta Bronze table. The pipeline must handle schema evolution without failing and must minimize the number of files that require reprocessing when the schema changes. The engineer wants to configure Auto Loader appropriately. Which two configuration settings should be used to achieve these requirements? (Choose two.)
Hard6When ingesting data using Auto Loader, what is the purpose of the 'cloudFiles.schemaLocation' parameter?
Medium7A data engineer is ingesting data from an Apache Kafka topic into a Delta Lake table using Structured Streaming. The Kafka topic receives messages with a timestamp field in the value payload, but the messages can arrive out of order by up to 10 minutes. The engineer wants to perform time-windowed aggregations on the ingested data while minimizing state store overhead. Which approach should be used to handle the out-of-order data correctly?
Hard8A data engineer is ingesting streaming data from Apache Kafka into a Delta table using Databricks Structured Streaming. The engineer wants to ensure exactly-once processing and handle late-arriving data. Which combination of features should the engineer use?
Medium9A data engineer is using Databricks Auto Loader to ingest CSV files into a Delta table. The engineer notices that some files have a different delimiter (semicolon instead of comma). Which option should be used to handle this variation?
Easy10A data engineer is using Auto Loader to ingest JSON files from cloud storage into a Delta table. The files contain a nested field 'address' with subfields 'city' and 'zip'. The engineer wants to flatten the nested structure during ingestion so that 'city' and 'zip' become top-level columns in the Bronze table. Which Auto Loader feature should be used to achieve this?
Easy11Which approach is most appropriate for ingesting data from a JDBC source into Delta Lake where the source table has no 'updated_at' or 'version' column for incremental loading?
Medium12Your organization is ingesting sensitive PII data. You need to ensure that personal identifiers are masked during the ingestion process before they are stored in the Bronze layer of your Medallion architecture. What is the best practice for this?
Medium13Refer to the exhibit. You are using Auto Loader to ingest data with evolving schemas. After running the job for a week, you realize that new columns added to the source JSON are not being captured in the destination table. What must you add to the configuration?
Hard14Which THREE of the following are essential components of an effective ingestion monitoring strategy in Databricks?
Medium15A data engineer is ingesting data from an Azure SQL Database into a Delta Lake table using the JDBC connector in a Databricks notebook. The source table contains millions of rows, and the engineer wants to optimize the ingestion by reading the data in parallel. The source table has a numeric primary key column named 'id' that is evenly distributed. Which approach should the engineer use to achieve parallel reads?
Medium16You are designing an ingestion pipeline that must handle massive bursts of data at irregular intervals. Which feature should you prioritize to ensure the ingestion process remains cost-effective?
Medium17When ingesting data from a Kafka topic into Delta Lake, what is the best way to handle out-of-order data arriving in the stream?
Medium18A data engineer is configuring an Auto Loader stream to ingest JSON files from an S3 bucket into a Bronze Delta table. The source bucket contains both .json and .json.gz files, and the engineer wants to ensure that only .json files are processed. Which parameter should be set to achieve this?
Medium19Which TWO of the following are primary benefits of using Delta Live Tables (DLT) for data ingestion over standard Structured Streaming pipelines?
HardOther domains
All Databricks-DE-Pro exam domains
Frequently asked questions
- What does the Data Ingestion and Acquisition domain cover on the Databricks-DE-Pro exam?
- Be able to choose the right ingestion mechanism for a source, configure Auto Loader or DLT correctly, and land data in Delta. The single most important thing: mask or tokenize sensitive identifiers before they reach Bronze storage, not after.
- How many questions are in this domain?
- This page lists all 19 Data Ingestion and Acquisition questions in the Databricks-DE-Pro question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Data Ingestion and Acquisition questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.