Courseiva

Databricks-DE-Pro · domain

Data Ingestion and Acquisition

This domain covers acquiring data into Databricks from files, message buses, and external systems, then landing it in Bronze. Expect questions on Delta Live Tables versus hand-rolled Structured Streaming, Auto Loader ingestion modes, Delta Lake as a sink, and enforcing PII masking during ingestion.

19 questions2 easy12 medium5 hard

Focused practice

Practice Data Ingestion and Acquisition questions

Scored sessions drawing only from this domain — pick a length below.

What this domain covers

What to know about Data Ingestion and Acquisition

Be able to choose the right ingestion mechanism for a source, configure Auto Loader or DLT correctly, and land data in Delta. The single most important thing: mask or tokenize sensitive identifiers before they reach Bronze storage, not after.

Auto Loader directory listing versus file notification modes for incremental cloud file ingestion

Delta Live Tables expectations, streaming tables, and medallion Bronze/Silver/Gold pipeline declarations

Structured Streaming sources and sinks including Kafka, Event Hubs, and Delta Lake

Unity Catalog and column masking or row filter policies applied to sensitive ingestion data

Watch out for

Common Data Ingestion and Acquisition exam traps

  • ▸Assuming DLT is just syntactic sugar over Structured Streaming, missing managed orchestration, expectations, and lineage benefits
  • ▸Choosing file notification mode by default when directory listing is simpler and sufficient for low-volume or irregular bursts
  • ▸Masking PII only in Silver or Gold, leaving raw identifiers persisted in Bronze tables and files

Question index

All Data Ingestion and Acquisition questions (19)

Click any question to see the full explanation, or start a practice session above.

1

What is the primary advantage of using Delta Lake as the sink for your data ingestion pipelines compared to raw Parquet files?

Medium
2

You are ingesting data from multiple source systems with varying file formats (JSON, CSV, Parquet) into a centralized Bronze landing zone. Which architecture pattern is the most scalable for maintaining this ingestion layer?

Hard
3

A data engineer is building a streaming ingestion pipeline from Apache Kafka to a Delta table. The pipeline must perform deduplication on a unique event_id field and handle late-arriving data. The engineer wants to use Structured Streaming with a watermark of 10 minutes. Which of the following approaches correctly implements deduplication and watermarking?

Medium
4

A data engineer is ingesting data from an Apache Kafka topic into a Delta Lake table using Structured Streaming. The Kafka topic receives messages with a timestamp field in the value payload, but the ingestion must handle late-arriving data and produce correct aggregations. The engineer wants to ensure that watermarks are applied correctly. Which approach should be used?

Medium
5

A data engineer is building a streaming ingestion pipeline using Databricks Auto Loader to ingest JSON files from cloud storage into a Delta Bronze table. The pipeline must handle schema evolution without failing and must minimize the number of files that require reprocessing when the schema changes. The engineer wants to configure Auto Loader appropriately. Which two configuration settings should be used to achieve these requirements? (Choose two.)

Hard
6

When ingesting data using Auto Loader, what is the purpose of the 'cloudFiles.schemaLocation' parameter?

Medium
7

A data engineer is ingesting data from an Apache Kafka topic into a Delta Lake table using Structured Streaming. The Kafka topic receives messages with a timestamp field in the value payload, but the messages can arrive out of order by up to 10 minutes. The engineer wants to perform time-windowed aggregations on the ingested data while minimizing state store overhead. Which approach should be used to handle the out-of-order data correctly?

Hard
8

A data engineer is ingesting streaming data from Apache Kafka into a Delta table using Databricks Structured Streaming. The engineer wants to ensure exactly-once processing and handle late-arriving data. Which combination of features should the engineer use?

Medium
9

A data engineer is using Databricks Auto Loader to ingest CSV files into a Delta table. The engineer notices that some files have a different delimiter (semicolon instead of comma). Which option should be used to handle this variation?

Easy
10

A data engineer is using Auto Loader to ingest JSON files from cloud storage into a Delta table. The files contain a nested field 'address' with subfields 'city' and 'zip'. The engineer wants to flatten the nested structure during ingestion so that 'city' and 'zip' become top-level columns in the Bronze table. Which Auto Loader feature should be used to achieve this?

Easy
11

Which approach is most appropriate for ingesting data from a JDBC source into Delta Lake where the source table has no 'updated_at' or 'version' column for incremental loading?

Medium
12

Your organization is ingesting sensitive PII data. You need to ensure that personal identifiers are masked during the ingestion process before they are stored in the Bronze layer of your Medallion architecture. What is the best practice for this?

Medium
13

Refer to the exhibit. You are using Auto Loader to ingest data with evolving schemas. After running the job for a week, you realize that new columns added to the source JSON are not being captured in the destination table. What must you add to the configuration?

Hard
14

Which THREE of the following are essential components of an effective ingestion monitoring strategy in Databricks?

Medium
15

A data engineer is ingesting data from an Azure SQL Database into a Delta Lake table using the JDBC connector in a Databricks notebook. The source table contains millions of rows, and the engineer wants to optimize the ingestion by reading the data in parallel. The source table has a numeric primary key column named 'id' that is evenly distributed. Which approach should the engineer use to achieve parallel reads?

Medium
16

You are designing an ingestion pipeline that must handle massive bursts of data at irregular intervals. Which feature should you prioritize to ensure the ingestion process remains cost-effective?

Medium
17

When ingesting data from a Kafka topic into Delta Lake, what is the best way to handle out-of-order data arriving in the stream?

Medium
18

A data engineer is configuring an Auto Loader stream to ingest JSON files from an S3 bucket into a Bronze Delta table. The source bucket contains both .json and .json.gz files, and the engineer wants to ensure that only .json files are processed. Which parameter should be set to achieve this?

Medium
19

Which TWO of the following are primary benefits of using Delta Live Tables (DLT) for data ingestion over standard Structured Streaming pipelines?

Hard

Frequently asked questions

What does the Data Ingestion and Acquisition domain cover on the Databricks-DE-Pro exam?
Be able to choose the right ingestion mechanism for a source, configure Auto Loader or DLT correctly, and land data in Delta. The single most important thing: mask or tokenize sensitive identifiers before they reach Bronze storage, not after.
How many questions are in this domain?
This page lists all 19 Data Ingestion and Acquisition questions in the Databricks-DE-Pro question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Data Ingestion and Acquisition questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
databricks-data-engineer-professional DATABRICKS-DATA-ENGINEER-PROFESSIONAL data ingestion acquisition Practice Questions