Databricks-DA-Assoc Importing Data Practice Question
A data analyst needs to ingest a daily batch of comma-separated value (CSV) files stored in cloud object storage into a Databricks Delta table. The files contain a header row, use a semicolon (;) as the delimiter, and occasionally include embedded newlines within quoted fields. Which approach ensures the parser correctly handles the multi-line fields and missing values?
⚠ Common exam trap
Candidates frequently overlook the 'multiLine' option. They assume standard CSV parsing handles embedded newlines automatically, which leads to corrupted data rows and schema errors during the ingestion process.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply the spark.read.format("csv").option("delimiter", ";").option("header", "true").option("multiLine", "true").load(path) command.
Configuring the Spark CSV reader with explicit options like multiLine set to true ensures that records containing embedded newlines are not split prematurely across rows. Data analysts frequently encounter messy raw external datasets where standard single-line parsing fails, leading to corrupted row counts and schema mismatches. Explicitly defining formatting options guarantees reliable ingestion before downstream business intelligence reporting takes place.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use spark.read.table() with default parameters because Spark automatically infers delimiters and multiline configurations for all incoming CSV files.
Why it's wrong here
Spark's default CSV reader assumes comma delimiters and multiLine disabled; it does not infer semicolons or quoted embedded newlines, so rows are split incorrectly. It is tempting because schema inference handles headers and types, and would be correct for standard comma-separated files without embedded newlines.
- ✓
Apply the spark.read.format("csv").option("delimiter", ";").option("header", "true").option("multiLine", "true").load(path) command.
Why this is correct
Setting the delimiter option to a semicolon, enabling header parsing, and setting multiLine to true allows Spark to correctly interpret complex rows. This combination handles formatting variations common in enterprise data dumps without dropping records.
- ✗
Convert the CSV files into Parquet format using a local Python script outside Databricks before uploading them to the workspace file system.
Why it's wrong here
Pre-converting outside Databricks adds an external pipeline step and still requires a parser to read the semicolon-delimited, multi-line CSV correctly, so the parsing problem simply moves. It is tempting because Parquet ingestion is fast and schema-stable, and would be correct if the source were already columnar and no CSV parsing were needed.
- ✗
Execute a standard SQL COPY INTO command without specifying any options, as Delta tables automatically resolve non-standard delimiters.
Why it's wrong here
COPY INTO without options assumes comma delimiters and does not enable multiLine parsing, so semicolon-separated rows and quoted embedded newlines are mis-split. It is tempting because COPY INTO is the standard idempotent ingestion path, and would be correct once delimiter, header and multiLine options are supplied.
About these practice questions
One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.