Be able to build a governed pipeline from raw volume files to a Delta training or retrieval table, using Unity Catalog for lineage and DLT plus ai_query for labeling. The key is materializing expensive model outputs so they are not recomputed on every run.
Start practicing
Data Preparation — choose a session length
Free · No account required
Domain overview
Data Preparation covers how Databricks practitioners turn raw, mixed-format source material into clean, governed training and retrieval datasets. Expect questions on Unity Catalog volumes and lineage, Delta Live Tables bronze/silver/gold pipelines, ai_query for labeling, chunking strategies before embedding, and assembling instruction or fine-tuning sets at scale.
Exam objectives
Using Unity Catalog lineage and Delta table history to trace training-set provenance
Building DLT pipelines that call ai_query to label or classify raw text
Choosing chunk size and overlap for text before vectorization
Reading PDF, DOCX, and text from Unity Catalog volumes into a unified dataset
Treating ai_query as cheap: rerunning it on every pipeline refresh instead of caching or materializing labeled results in a Delta table
Picking chunking parameters by intuition rather than matching them to embedding model context limits and retrieval query shape
Assuming raw files in a volume are queryable directly, skipping parsing and normalization into Delta tables
Click any question to see the full explanation and answer options, or start a focused practice session above.
A Data Engineer needs to ensure that PII data in a Delta table is masked before serving it to non-privileged users. Which Databricks feature provides the most efficient, centralized control for this requirement?
2You need to ingest data from an external JSON source into a Delta table. The source schema is inconsistent. Which strategy is most effective for preparing this data?
3Which TWO of the following are benefits of using Delta Lake for data preparation over standard Parquet files on cloud storage?
4Which Databricks feature is best suited for maintaining the lineage of data used during the preparation of training sets for Generative AI?
5When preparing a dataset for fine-tuning an LLM, you need to ensure the data is representative of the target domain. What is the most effective approach to detect and mitigate sampling bias in your training set using Databricks?
6Which of the following describes the purpose of using a 'Feature Store' when preparing data for Generative AI applications?
7You are preparing a large dataset for fine-tuning a model using Databricks Delta Live Tables (DLT). Which configuration is best for ensuring data quality and lineage in this pipeline?
8Which TWO factors are most important when selecting a chunking strategy for text data prior to vectorization?
9A team is preparing data for a RAG system and needs to remove duplicates from a large collection of PDF text extracts. What is the most efficient way to perform de-duplication in Databricks?
10Which of the following describes the 'Gold' layer in a Medallion architecture, and why is it important for GenAI data preparation?
11You need to store embeddings generated by an LLM in a Delta table. Which data type is most efficient for storing these high-dimensional vector arrays in Databricks?
12When preparing data for a fine-tuning task, you realize the dataset is severely imbalanced. Which Databricks technique should you use to create a more balanced dataset?
13A GenAI engineer is preparing a large text corpus stored in a Unity Catalog volume for fine-tuning a chat model. The raw files are JSON Lines, each containing a 'conversation' array with alternating 'user' and 'assistant' turns, but many records have malformed turns or missing roles. The engineer needs to convert this into a Delta table with a schema of (conversation_id STRING, messages ARRAY<STRUCT<role:STRING, content:STRING>>) while filtering out records where any turn has a null role or empty content. Which approach uses the appropriate Databricks-native capability for this transformation?
14You are preparing a large corpus of customer support emails stored as Parquet files in a Unity Catalog volume. The emails must be cleaned by removing boilerplate signatures and disclaimers before tokenization for a fine-tuning dataset. Your team wants to enforce this transformation as a declarative, testable pipeline stage that fails fast if any cleaned record still contains a known boilerplate marker. Which Databricks capability should you use to implement this cleaning step with built-in data quality expectations?
15A data engineer is preparing a dataset of product descriptions for embedding generation. The source table in Unity Catalog contains a column 'description' with mixed languages, and the team wants to filter to English-only text before vectorization. They need a scalable, built-in Databricks function that can detect the language of each description without external API calls. Which function should they use?
16A GenAI engineer is preparing a corpus of HTML product pages stored in a Unity Catalog volume for a RAG application. The pages contain navigation bars, script tags, and boilerplate footers that add noise to embeddings. Which Databricks-native approach best removes this noise while preserving the main article text before chunking?
17You are preparing a large text corpus for fine-tuning a generative AI model. The corpus is stored in a Delta table with columns: doc_id, raw_text, and metadata. You need to create a cleaned dataset that removes personally identifiable information (PII) and normalizes whitespace, while preserving document boundaries for training. Which two actions should you perform to achieve this in a scalable and maintainable way? (Choose two.)
18A GenAI engineer must build a fine-tuning dataset from 40 TB of raw JSONL conversation logs stored in cloud object storage. The logs are immutable and only ever read once, in full, during preprocessing. Which storage and access configuration should be used to minimize cost while keeping the data readable by Spark on Databricks?
19A Generative AI engineer is building a Delta Live Tables pipeline that ingests raw JSON event logs into a bronze table, then uses ai_query to classify each event's free-text field. The classification call is expensive, so the engineer wants to avoid re-running it on events that have already been processed in previous pipeline updates. The source table is append-only and new events arrive continuously. Which Delta Live Tables feature should the engineer configure on the bronze table to prevent reprocessing of previously ingested rows?
20A team is preparing a pretraining corpus and must filter out near-duplicate documents before tokenization. The corpus contains 500 million short text records in a Delta table. They want a scalable, deterministic deduplication signal that can be computed per record and compared across the dataset. Which approach best fits this requirement?
21A data engineer is preparing a large corpus of customer support transcripts stored as Parquet in a Unity Catalog volume. Before generating embeddings, each transcript must be tokenized and truncated to a maximum token length. The engineer wants to use a Databricks-native approach that runs distributed across the cluster and avoids pulling the full corpus into a single node. Which approach best satisfies these requirements?
22A Generative AI engineer is preparing a Delta table of product reviews for a retrieval-augmented generation application. The reviews contain HTML tags, inconsistent casing, and occasional very long paragraphs that exceed the embedding model's context window. The engineer wants to clean and normalize the text before chunking and embedding. Which two actions should the engineer take to directly address the stated quality issues? (Choose two.)
23A data engineer needs to persist a prepared instruction-tuning dataset so that downstream fine-tuning jobs can read it with ACID guarantees, time travel, and schema enforcement, and so that Unity Catalog can track column-level lineage. Which storage format and registration should be used?
24A team stores raw documents in a Unity Catalog volume and needs to build a training set for fine-tuning a large language model. The raw files are in mixed formats including PDF, DOCX, and plain text. The team wants a single Delta table where every row is one document with its extracted text and source path, and wants the extraction to run in parallel across the cluster. Which approach should the team use?
25You are building a data preparation pipeline for a generative AI application. The raw data is stored in a Unity Catalog volume as JSON files with nested fields. You need to flatten the nested structure and extract specific fields into a Delta table for downstream embedding. The JSON schema may evolve, with new fields added occasionally. Which approach provides the most robust and maintainable solution?
26An engineer is cleaning a Delta table of customer support transcripts stored in Unity Catalog. The table has a nested array column named messages, where each element has role and content fields. They need to keep only rows where at least one message has role equal to 'user' and content longer than 20 characters. Which transformation correctly expresses this filter?
27A Databricks team is building a retrieval corpus from mixed-format documents stored in a Unity Catalog volume. They need a preparation pipeline that preserves document structure for later chunking and that records which source file each chunk came from so retrieval results can cite evidence. Which two design choices best meet these requirements? (Choose two.)
28A Generative AI engineer is preparing a Delta table of support tickets for embedding generation. The pipeline computes embeddings with ai_query and stores them in a separate embeddings table keyed by ticket_id. Tickets are frequently updated by agents, and the engineer wants the embeddings table to reflect the latest ticket text without recomputing embeddings for tickets whose text has not changed. Which design best achieves this?
29A team is building a RAG application using Databricks Vector Search with a Delta table as the source. They need the index to automatically reflect new and updated chunks as the source table changes, without rebuilding the entire index each time. Which configuration should they use?
30A team is preparing a large text corpus for embedding generation with a foundation model endpoint on Databricks. They must reduce token cost and improve retrieval quality before vectorization. Which two preprocessing steps should be applied to the raw text? (Choose two.)
31A data engineer needs to prepare a Delta table of customer reviews for embedding generation. The reviews contain HTML tags, inconsistent whitespace, and mixed casing that hurt embedding quality. Which preparation step should be applied before generating embeddings?
32A data engineer is preparing a large corpus of support tickets stored in a Unity Catalog volume for fine-tuning a Llama model on Databricks. They must remove personally identifiable information (PII) before the data reaches the training cluster. Which TWO approaches are appropriate for detecting and redacting PII at scale in this pipeline? (Choose two.)
33A GenAI engineer is preparing a large text corpus for fine-tuning an LLM. The corpus contains many near-duplicate documents and documents in multiple languages. They need to reduce redundancy and ensure language consistency. Which two steps should be performed during data preparation? (Choose two.)
34A team is preparing a Delta table of product descriptions for a RAG application. The table receives continuous upserts from a streaming pipeline, and the embedding job reads the table every hour. Engineers notice the embedding job reprocesses every row on each run even though only a few rows change. Which change should be made to the source table to let the embedding job process only new or updated rows?
35A data engineer is preparing a large Delta table of conversation logs for embedding generation. The table has frequent small appends, and the engineer needs to reduce file fragmentation and improve read throughput before the embedding job runs. Which two actions should the engineer take? (Choose two.)
36A data engineer is assembling a fine-tuning dataset from a Delta table of conversation transcripts. Each transcript contains a list of message objects with a role and a content field, and the training job requires one row per conversation with the messages serialized into the expected format. Which transformation should the engineer apply?
37A data engineer is preparing a dataset for fine-tuning a chat model. The dataset contains conversations with alternating user and assistant messages. They need to format each conversation into a single string with special tokens indicating roles. Which approach is most appropriate in Databricks?
38A GenAI engineer prepares a Delta table of product descriptions that will feed a chunking and embedding pipeline. The descriptions are written in mixed languages, and the embedding model supports only English. The engineer needs to ensure that non-English rows are detected and routed for translation before embedding. Which approach is most appropriate?
39A team is preparing a fine-tuning dataset from customer reviews stored in a Delta table. They need to filter out reviews shorter than 20 tokens and reviews flagged as spam by a classifier, then write the result to a Unity Catalog table for training. Which approach best fits Databricks best practices?
40A GenAI engineer is preparing a fine-tuning dataset from a Delta table in Unity Catalog that contains raw user feedback. The feedback text includes irregular capitalization, HTML tags, and excessive punctuation. The engineer needs to normalize the text using Spark NLP within a Databricks notebook, ensuring the pipeline is reproducible and scalable. Which approach should the engineer use to apply this transformation?
Be able to build a governed pipeline from raw volume files to a Delta training or retrieval table, using Unity Catalog for lineage and DLT plus ai_query for labeling. The key is materializing expensive model outputs so they are not recomputed on every run.
The Courseiva Databricks-GenAI-Assoc question bank contains 40 questions in the Data Preparation domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Data Preparation domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included