Courseiva

Databricks-DA-Assoc · topic practice

Importing Data practice questions

This domain covers getting external files into Databricks tables: the Add Data UI, COPY INTO, Auto Loader (cloudFiles), and Delta Lake as the target format. Questions are scenario-based, asking you to pick the right ingestion path for CSV or JSON in cloud object storage, configure delimiters, headers, and schemas, and reason about schema evolution and incremental loads.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Importing Data

What the exam tests

What to know about Importing Data

Be able to pick and configure the right ingestion method: Add Data UI for quick imports, COPY INTO for idempotent batch loads, and Auto Loader for incremental streaming with schema evolution. The key is matching file format options and target Delta table behavior to the scenario.

Choosing Delta Lake over Parquet for ACID transactions, schema enforcement, and time travel on ingested tables

Using COPY INTO with options like format, delimiter, header, and mergeSchema for idempotent batch file ingestion

Configuring Auto Loader (cloudFiles) with schema evolution and checkpointing for streaming ingestion from cloud storage

Applying the Add Data UI to preview, set column types, and create a Delta table from CSV or JSON files

Watch out for

Common Importing Data exam traps

  • ▸Assuming COPY INTO re-loads every file each run; it tracks ingested files and skips already-processed ones unless forced.
  • ▸Setting the wrong delimiter or header option, so the schema inference produces a single garbled column instead of parsed fields.
  • ▸Confusing schema evolution with schema enforcement: evolution adds new columns, while enforcement rejects mismatched writes.

Practice set

Importing Data questions

20 questions · select your answer, then reveal the explanation

Question 1mediummultiple choice
Read the full Importing Data explanation →

Refer to the exhibit. The analyst finds that the resulting DataFrame has column names like '_c0', '_c1'. What is missing in the configuration?

Exhibit

spark.read.format('csv').option('header', 'true').load('path')
Question 2mediummultiple choice
Study the full Python automation breakdown →

A data analyst is creating a Databricks SQL dashboard that must ingest a Parquet file stored in cloud object storage. The analyst wants to define the table using pure SQL in the Databricks SQL editor, without using Python or notebooks. Which SQL command should be used to create a table that reads directly from the Parquet file?

A data analyst wants to upload a small CSV file from their local machine to the Databricks workspace to analyze it in a notebook. They need a quick, one-time upload without writing code. Which Databricks feature should they use?

Question 4mediummultiple choice
Read the full Importing Data explanation →

A data analyst is ingesting a stream of JSON event files from cloud object storage into a Bronze Delta table using Auto Loader in a Databricks notebook. The source directory receives new, uniquely named files continuously, and the analyst sets the schema location to a Unity Catalog volume path. After several days, the analyst notices that some newly arriving files are being ignored entirely, while others are processed. Which Auto Loader behavior most likely explains why a subset of files is never ingested?

Question 5mediummultiple choice
Read the full Importing Data explanation →

A data analyst is creating a Delta table from a Parquet file stored in DBFS. They use the following SQL command: CREATE TABLE my_table USING DELTA LOCATION '/mnt/data/parquet_files/'. After running the command, they notice the table contains no data. What is the most likely cause?

A data analyst must load a one-time snapshot of a JDBC-accessible PostgreSQL database into a Delta table in Unity Catalog using a Databricks notebook. The analyst wants to minimize cluster memory pressure and enable downstream predicate pushdown on the resulting Delta table. Which two configuration choices should the analyst apply? (Choose two.)

A data analyst is loading a CSV dataset into a Delta table using spark.read.csv with header=true and inferSchema=true. The CSV contains an ID column whose values are numeric but must be preserved as strings to avoid losing leading zeros. The analyst notices that the resulting Delta table has the ID column typed as bigint, and downstream joins against a string-typed source are failing. Which change should the analyst make to the load to produce the correct schema?

Question 8mediummultiple choice
Read the full Importing Data explanation →

A data analyst is importing a 30 GB CSV dataset from an ADLS Gen2 location into a Bronze Delta table in Unity Catalog. The CSV has a header row, uses commas as delimiters, and contains quoted fields with embedded commas. The analyst runs spark.read.option("header", "true").option("inferSchema", "true").csv("abfss://raw@storage.dfs.core.windows.net/events/") and then writes the DataFrame to a Delta table. The resulting table has several columns typed as strings that should be numeric, and a few rows are misaligned with values shifted into the wrong columns. Which change should the analyst make to correct BOTH the type inference and the misaligned rows?

Question 9mediummultiple choice
Read the full Importing Data explanation →

A data analyst needs to ingest a large volume of CSV files from an external S3 bucket into a Delta table. Which method provides the most efficient, fault-tolerant, and incremental loading approach in Databricks?

Question 10easymultiple choice
Read the full Importing Data explanation →

An analyst is preparing to upload a small local CSV file to a Databricks workspace. Which tool is most appropriate for a quick, one-time upload without requiring infrastructure setup?

Question 11mediummultiple choice
Read the full Importing Data explanation →

Refer to the exhibit. An analyst is attempting to read a CSV file using Spark. Why is the code failing?

Exhibit

Error: AnalysisException: Path does not exist: dbfs:/mnt/data/sales_2023.csv

Which TWO of the following scenarios are valid use cases for using the COPY INTO command? (Choose two)

Question 13easymultiple choice
Read the full Importing Data explanation →

When importing data into Databricks using the 'Add Data' UI, what is the default file format for the created table if the source is a CSV file?

Question 14mediummultiple choice
Read the full Importing Data explanation →

An analyst is using the Spark DataFrame API to read a large JSON dataset. The dataset contains nested fields that are causing schema inference to fail. Which approach best resolves this?

Which THREE actions occur when using Auto Loader with schema evolution enabled? (Choose three)

Question 16mediummultiple choice
Read the full Importing Data explanation →

When importing data into Databricks, what is the primary benefit of using Delta Lake over standard Parquet files?

Which THREE of the following are valid sources for importing data into a Databricks Delta table? (Choose three)

Question 18mediummultiple choice
Read the full Importing Data explanation →

A data analyst needs to ingest a daily batch of comma-separated value (CSV) files stored in cloud object storage into a Databricks Delta table. The files contain a header row, use a semicolon (;) as the delimiter, and occasionally include embedded newlines within quoted fields. Which approach ensures the parser correctly handles the multi-line fields and missing values?

Question 19easymultiple choice
Read the full Importing Data explanation →

A junior data analyst attempts to load a large JSON file into a Databricks DataFrame using spark.read.json(path), but the resulting DataFrame contains numerous null values across critical columns. Upon inspection, the raw JSON records have inconsistent nesting structures and missing attributes. What is the most effective way for the analyst to inspect the inferred schema before transforming the data?

Question 20mediummultiple choice
Read the full Importing Data explanation →

A data analyst needs to load a daily batch of 500 CSV files from an S3 bucket into a Delta table. The files are consistently formatted, and the analyst wants to ensure that files already processed are not re-imported in subsequent runs. Which approach is most efficient for this idempotency requirement?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Importing Data sessions

Start a Importing Data only practice session

Every question in these sessions is drawn from the Importing Data domain — nothing else.

Related practice questions

Related Databricks-DA-Assoc topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-DA-Assoc exam test about Importing Data?
Be able to pick and configure the right ingestion method: Add Data UI for quick imports, COPY INTO for idempotent batch loads, and Auto Loader for incremental streaming with schema evolution. The key is matching file format options and target Delta table behavior to the scenario.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Importing Data questions in a focused session?
Yes — the session launcher on this page draws every question from the Importing Data domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-DA-Assoc topics?
Use the topic links above to move to related areas, or go back to the Databricks-DA-Assoc question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-DA-Assoc exam covers. They are not copied from any real exam or dump site.