Databricks-GenAI-Assoc Data Preparation Practice Question
A Databricks team is building a retrieval corpus from mixed-format documents stored in a Unity Catalog volume. They need a preparation pipeline that preserves document structure for later chunking and that records which source file each chunk came from so retrieval results can cite evidence. Which two design choices best meet these requirements? (Choose two.)
⚠ Common exam trap
The trap here is treating normalization to plain text as harmless, when discarding layout markers actually removes the structure needed for coherent chunking and provenance.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Carry the source file path as a metadata column on each chunk and store it in the Delta table alongside the chunk text and embedding.
Preserving document structure requires a parser that extracts layout elements rather than flattening everything to plain text, and ai_parse_document provides that for PDFs and images. Provenance requires carrying the source file path on each chunk so retrieval results can cite the originating document. Together these choices keep structure intact and make every chunk traceable.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Concatenate all documents into a single large text field before chunking to reduce the number of rows.
Why it's wrong here
Concatenating documents destroys the boundaries between source files, so chunks can span multiple documents and lose their provenance. It also prevents attaching a meaningful source path per chunk, which breaks the citation requirement. Reducing row count this way trades away the structural and traceability guarantees the pipeline needs.
- ✓
Carry the source file path as a metadata column on each chunk and store it in the Delta table alongside the chunk text and embedding.
Why this is correct
Attaching the source file path as a metadata column lets retrieval results reference the originating document, enabling citations and traceability. Delta tables support arbitrary metadata columns, so the path can travel with each chunk through embedding generation and into the vector index without additional joins.
- ✗
Convert every document to plain text and discard page and heading markers to normalize the corpus.
Why it's wrong here
Discarding page and heading markers removes the structural signals that make chunking coherent, so passages may merge unrelated sections or split mid-topic. While plain text is simpler to process, the loss of structure undermines retrieval quality and makes it harder to reconstruct document context for citations.
- ✗
Generate embeddings for entire documents rather than chunks to avoid storing multiple rows per file.
Why it's wrong here
Embedding whole documents produces a single vector that averages many topics, which reduces retrieval precision because the vector cannot localize the relevant passage. It also conflicts with the requirement to preserve structure for chunking, since chunk-level granularity is what enables precise evidence citation.
- ✓
Use the ai_parse_document function to extract text and layout elements from PDFs and images, retaining the document structure for downstream chunking.
Why this is correct
The ai_parse_document function extracts text and structural elements such as headings, tables, and reading order from PDFs and images, which preserves the layout information needed to chunk coherently. Retaining that structure prevents chunks from splitting across section boundaries and gives later stages the context required to build meaningful retrieval passages.
About these practice questions
Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.