Databricks-DA-Assoc · domain
Importing Data
This domain covers getting external files into Databricks tables: the Add Data UI, COPY INTO, Auto Loader (cloudFiles), and Delta Lake as the target format. Questions are scenario-based, asking you to pick the right ingestion path for CSV or JSON in cloud object storage, configure delimiters, headers, and schemas, and reason about schema evolution and incremental loads.
Focused practice
Practice Importing Data questions
Scored sessions drawing only from this domain — pick a length below.
What this domain covers
What to know about Importing Data
Be able to pick and configure the right ingestion method: Add Data UI for quick imports, COPY INTO for idempotent batch loads, and Auto Loader for incremental streaming with schema evolution. The key is matching file format options and target Delta table behavior to the scenario.
Choosing Delta Lake over Parquet for ACID transactions, schema enforcement, and time travel on ingested tables
Using COPY INTO with options like format, delimiter, header, and mergeSchema for idempotent batch file ingestion
Configuring Auto Loader (cloudFiles) with schema evolution and checkpointing for streaming ingestion from cloud storage
Applying the Add Data UI to preview, set column types, and create a Delta table from CSV or JSON files
Watch out for
Common Importing Data exam traps
- ▸Assuming COPY INTO re-loads every file each run; it tracks ingested files and skips already-processed ones unless forced.
- ▸Setting the wrong delimiter or header option, so the schema inference produces a single garbled column instead of parsed fields.
- ▸Confusing schema evolution with schema enforcement: evolution adds new columns, while enforcement rejects mismatched writes.
Question index
All Importing Data questions (23)
Click any question to see the full explanation, or start a practice session above.
A data analyst needs to quickly upload a small CSV file from their local machine into Databricks to explore its contents. They want to avoid writing code and prefer a graphical interface. Which Databricks feature should they use?
Easy2A data analyst needs to ingest a large volume of CSV files from an external S3 bucket into a Delta table. Which method provides the most efficient, fault-tolerant, and incremental loading approach in Databricks?
Medium3A data analyst is using the Databricks SQL Connector for Python to query a large Delta table. The query returns millions of rows, and the analyst wants to process the results in batches to avoid memory issues. Which method should the analyst use to fetch rows in chunks?
Hard4A data analyst is setting up a recurring ingestion job that must load only new files from a cloud storage directory into a Delta table, must track which files have already been processed, and must fail clearly if the source schema changes unexpectedly. Which TWO capabilities should the analyst rely on? (Choose two.)
Hard5Which THREE of the following are valid sources for importing data into a Databricks Delta table? (Choose three)
Medium6A data analyst has a 5 MB pipe-delimited text file with a header row on their local laptop and needs to load it into a Databricks Unity Catalog table for a one-time analysis. They want the fastest path with the least configuration. Which approach should they use?
Easy7A data analyst needs to perform a one-time upload of a 15 MB compressed Parquet file from a local laptop into a Unity Catalog volume in a Databricks workspace. The analyst has the Databricks SQL editor open and no cluster running. Which approach should the analyst take to place the file into the volume path?
Easy8An analyst is preparing to upload a small local CSV file to a Databricks workspace. Which tool is most appropriate for a quick, one-time upload without requiring infrastructure setup?
Easy9Which TWO of the following scenarios are valid use cases for using the COPY INTO command? (Choose two)
Medium10A data analyst runs COPY INTO to load new CSV files from an S3 path into a Delta table every night. One night the job completes successfully but the row count does not change, even though new files were placed in the source path. The files have the same names as files loaded previously. What is the most likely cause?
Medium11An analyst is loading Parquet files from cloud storage into a Delta table using COPY INTO. The source directory contains files with several different schemas because upstream teams added columns over time. The analyst wants the table to accept the union of all columns without manual intervention. Which COPY INTO behavior should the analyst configure?
Hard12Refer to the exhibit. An analyst is attempting to read a CSV file using Spark. Why is the code failing?
Medium13A data analyst must load a Parquet file from an S3 bucket into a Databricks DataFrame, but the bucket is in a different AWS account and requires temporary credentials. The analyst has an AWS access key ID, secret access key, and session token. Which code snippet correctly configures Spark to read the file securely?
Medium14A junior data analyst attempts to load a large JSON file into a Databricks DataFrame using spark.read.json(path), but the resulting DataFrame contains numerous null values across critical columns. Upon inspection, the raw JSON records have inconsistent nesting structures and missing attributes. What is the most effective way for the analyst to inspect the inferred schema before transforming the data?
Easy15A data analyst needs to load a daily batch of 500 CSV files from an S3 bucket into a Delta table. The files are consistently formatted, and the analyst wants to ensure that files already processed are not re-imported in subsequent runs. Which approach is most efficient for this idempotency requirement?
Medium16When importing data into Databricks using the 'Add Data' UI, what is the default file format for the created table if the source is a CSV file?
Easy17An analyst is loading a CSV file where one column contains values like 007, 012, and 003. After creating a table through the Add Data UI, those values appear as 7, 12, and 3. The analyst needs the leading zeros preserved for downstream reporting. What should the analyst do?
Medium18When importing data into Databricks, what is the primary benefit of using Delta Lake over standard Parquet files?
Medium19A data analyst is using the Databricks Add Data UI to import a CSV file from cloud storage into a Delta table. The analyst wants to ensure the import process handles the data correctly. Which two actions can the analyst perform directly in the Add Data UI? (Choose two.)
Hard20Which THREE actions occur when using Auto Loader with schema evolution enabled? (Choose three)
Hard21An analyst is using the Spark DataFrame API to read a large JSON dataset. The dataset contains nested fields that are causing schema inference to fail. Which approach best resolves this?
Medium22A data analyst needs to ingest a daily batch of comma-separated value (CSV) files stored in cloud object storage into a Databricks Delta table. The files contain a header row, use a semicolon (;) as the delimiter, and occasionally include embedded newlines within quoted fields. Which approach ensures the parser correctly handles the multi-line fields and missing values?
Medium23A user wants to import a local 50MB CSV file into a Databricks workspace for a quick ad-hoc analysis. Which TWO methods are available within the Databricks UI to accomplish this task directly?
HardOther domains
All Databricks-DA-Assoc exam domains
Frequently asked questions
- What does the Importing Data domain cover on the Databricks-DA-Assoc exam?
- Be able to pick and configure the right ingestion method: Add Data UI for quick imports, COPY INTO for idempotent batch loads, and Auto Loader for incremental streaming with schema evolution. The key is matching file format options and target Delta table behavior to the scenario.
- How many questions are in this domain?
- This page lists all 23 Importing Data questions in the Databricks-DA-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Importing Data questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.