Courseiva

Databricks-DE-Assoc · topic practice

Data Ingestion and Loading practice questions

This domain covers getting data into Delta Lake on Databricks: Auto Loader, COPY INTO, and streaming ingestion patterns. Questions present realistic ingestion scenarios and ask you to choose the right mode, trigger, schema-evolution behavior, or rescue-data handling for CSV, JSON, and partitioned cloud storage paths.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Data Ingestion and Loading

What the exam tests

What to know about Data Ingestion and Loading

Be able to pick Auto Loader versus COPY INTO, choose the correct trigger for batch-style incremental runs, and configure schema inference, evolution, and rescued data. The most important thing: match the ingestion mode and trigger to the stated cost and freshness requirement.

Auto Loader directory listing versus file notification mode with cloud notification services

Rescued data column for capturing malformed records during CSV and JSON ingestion

Schema inference and schema evolution from date-partitioned paths in cloud object storage

Trigger.AvailableNow for incremental jobs that process all available data then stop

Watch out for

Common Data Ingestion and Loading exam traps

  • ▸Assuming file notification mode needs no cloud setup; it requires configuring notification services and IAM permissions on the storage account.
  • ▸Forgetting that malformed records land in the rescued data column only when rescue mode is enabled, not by default in every read path.
  • ▸Confusing Trigger.Once with Trigger.AvailableNow, or using continuous triggers when the requirement is to process all data and shut down.

Practice set

Data Ingestion and Loading questions

20 questions · select your answer, then reveal the explanation

A company is ingesting large Parquet files into a Delta table. They notice that while the ingestion is fast, subsequent queries on the table are slow because each ingestion creates a few very large files. Which feature should be enabled to optimize the file size during ingestion?

A data engineer is configuring an Auto Loader stream to ingest JSON files from an Azure Blob Storage container into a Delta table. The source files frequently have schema evolutions, such as new nested fields being added. Which setting should be enabled in the Auto Loader readStream configuration to automatically capture and incorporate these schema changes into the target Delta table without failing the stream?

A data engineer is designing an ingestion pipeline to load data from a Kafka topic into a Delta Lake table using Databricks Structured Streaming. The pipeline must ensure exactly-once processing semantics and handle schema changes in the incoming JSON messages. Which TWO configurations or features should the engineer use? (Choose two.)

A data engineer needs to ingest a large CSV file from cloud storage into a Delta table using Databricks SQL. The file has a header row, and the engineer wants to create the table in one command while inferring the schema automatically. Which SQL command should be used?

A data engineer is using Databricks Auto Loader to stream CSV files from an Azure Data Lake Storage Gen2 container into a Delta table. The files are constantly appended, and the engineer needs to ensure that the ingested data includes the file path and ingestion timestamp for auditing. Which Auto Loader feature should be used to add these metadata columns?

A data engineer is building a nightly ingestion pipeline that reads CSV files from a cloud storage location into a Delta table. The files are known to have a header row, and the engineer wants to ensure that the header is not treated as data. The engineer also wants to automatically infer the schema from the file contents. Which combination of options should be used with the COPY INTO command to achieve this?

A data engineer needs to load a large CSV file from cloud storage into a Delta table using Databricks SQL. The file has a header row and uses a pipe delimiter. The engineer wants to use the COPY INTO command. Which SQL statement will correctly load the data?

A data engineer is designing an ingestion pipeline using Databricks Auto Loader to ingest Parquet files from cloud storage into a Delta table. The engineer wants to ensure that the pipeline can handle schema evolution and rescue malformed records. Which TWO options should be configured? (Choose two.)

A data engineer is designing a pipeline to ingest data from a REST API that returns JSON responses. The API supports pagination and rate limiting. The engineer needs to load the data into a Delta table and wants to ensure that the ingestion is incremental and can handle API rate limits gracefully. Which two approaches should the engineer use? (Choose two.)

A data engineer is loading a batch of CSV files into a Delta table using the COPY INTO command. The source directory contains both new files and files that were loaded successfully in a previous run. The engineer notices that the second run reprocesses all files, causing duplicate records. Which COPY INTO option should the engineer use to avoid reprocessing already-loaded files?

A data engineer is using Databricks Auto Loader to ingest CSV files from an AWS S3 bucket into a Delta table. The engineer wants to ensure that the ingestion process automatically handles files that arrive with a different delimiter, such as semicolon instead of comma. Which Auto Loader option should be used to specify the delimiter?

A data engineer is using the COPY INTO command to load Parquet files from a cloud storage location into a Delta table. The engineer notices that some files have been modified since the last load and wants to ensure that only new or updated files are ingested in the next run. Which COPY INTO option should be used to achieve this?

A data engineer is using Auto Loader to ingest CSV files from a cloud storage location into a Delta table. The engineer wants to ensure that any malformed records that do not match the schema are captured for later analysis instead of causing the stream to fail. Which Auto Loader option should the engineer configure to achieve this?

A data engineer is using Databricks Auto Loader to ingest data from a cloud storage directory that contains a mix of CSV and JSON files. The engineer wants to ingest only the CSV files and ignore the JSON files. Which approach should be used to filter the files?

A data engineer is using Auto Loader to ingest CSV files from a cloud storage location into a Delta table. The engineer wants to ensure that the ingestion process handles schema evolution gracefully when new columns are added to the CSV files. Which TWO actions should the engineer take to enable schema evolution and capture any data that does not match the evolving schema? (Choose two.)

A data engineer is building a Databricks Structured Streaming pipeline that ingests new files from an S3 bucket into a Delta table. The pipeline must process each file exactly once, even if the stream is stopped and restarted, and must not reprocess files that were already ingested. The engineer is not allowed to modify the source files or use an external tracking database. Which approach should the engineer use?

A data engineering team needs to ingest millions of small JSON files from an S3 bucket into a Delta Lake table. The solution must provide incremental loading, support schema evolution, and automatically scale to handle increasing file volumes without manual tracking of processed files. Which tool is best suited for this requirement?

A data engineer is evaluating whether to use Auto Loader or the COPY INTO command for a new ingestion pipeline. Which TWO features are unique to Auto Loader compared to COPY INTO?

Refer to the exhibit. A data engineer is configuring an Auto Loader stream with the provided options. What will happen if a new JSON file arrives containing a field that is not currently in the target table schema?

Exhibit

{
  "cloudFiles.format": "json",
  "cloudFiles.schemaLocation": "/mnt/bronze/schemas/orders",
  "cloudFiles.inferColumnTypes": "true",
  "cloudFiles.schemaEvolutionMode": "addNewColumns"
}

When using Auto Loader to ingest data, why is it recommended to provide a 'schemaLocation' rather than manually defining the entire schema in the code?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Data Ingestion and Loading sessions

Start a Data Ingestion and Loading only practice session

Every question in these sessions is drawn from the Data Ingestion and Loading domain — nothing else.

Related practice questions

Related Databricks-DE-Assoc topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-DE-Assoc exam test about Data Ingestion and Loading?
Be able to pick Auto Loader versus COPY INTO, choose the correct trigger for batch-style incremental runs, and configure schema inference, evolution, and rescued data. The most important thing: match the ingestion mode and trigger to the stated cost and freshness requirement.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Data Ingestion and Loading questions in a focused session?
Yes — the session launcher on this page draws every question from the Data Ingestion and Loading domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-DE-Assoc topics?
Use the topic links above to move to related areas, or go back to the Databricks-DE-Assoc question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-DE-Assoc exam covers. They are not copied from any real exam or dump site.