Databricks-GenAI-Assoc Data Preparation Practice Question
A team stores raw documents in a Unity Catalog volume and needs to build a training set for fine-tuning a large language model. The raw files are in mixed formats including PDF, DOCX, and plain text. The team wants a single Delta table where every row is one document with its extracted text and source path, and wants the extraction to run in parallel across the cluster. Which approach should the team use?
⚠ Common exam trap
The trap here is assuming a generic text reader or external table can parse PDF and DOCX, when binary formats require the binaryFile data source plus explicit extraction logic.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use binaryFile data source to read the volume, then apply a UDF that extracts text per file and writes results to a Delta table.
The binaryFile data source is designed to read arbitrary files as rows with path and content, enabling distributed extraction with a UDF. This yields one row per document with text and source path in a Delta table. Driver loops and plain text readers cannot handle mixed binary formats at scale or in parallel.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Mount the volume as a local filesystem path and use a Python for-loop to read and extract each file on the driver.
Why it's wrong here
Looping over files on the driver serializes all extraction work on a single node and will not scale to a large document set. It also bypasses Spark's parallelism entirely, which contradicts the requirement to run extraction in parallel. The resulting data would still need conversion into a Delta table afterward.
- ✓
Use binaryFile data source to read the volume, then apply a UDF that extracts text per file and writes results to a Delta table.
Why this is correct
The binaryFile data source reads each file as a row with path and binary content, which lets Spark distribute file processing across executors. Applying an extraction UDF per row produces one output row per document with text and source path. Writing to Delta gives the unified table the team wants for fine-tuning.
- ✗
Use spark.read.text on the volume, which automatically parses PDF and DOCX content into text columns.
Why it's wrong here
spark.read.text reads plain text files line by line and does not parse binary formats like PDF or DOCX. It would produce garbled or empty content for those formats. It also does not capture the source path per document in the way the team needs for a unified training table.
- ✗
Create an external table over the volume with a schema that includes columns for each document format's fields.
Why it's wrong here
An external table over mixed-format files does not extract text; it only exposes files according to a reader, and there is no single schema that natively parses PDF, DOCX, and text together. It would not produce a row per document with extracted text. Extraction logic still has to be applied separately.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.