Courseiva

Databricks-GenAI-Assoc · topic practice

Data Preparation practice questions

Data Preparation covers how Databricks practitioners turn raw, mixed-format source material into clean, governed training and retrieval datasets. Expect questions on Unity Catalog volumes and lineage, Delta Live Tables bronze/silver/gold pipelines, ai_query for labeling, chunking strategies before embedding, and assembling instruction or fine-tuning sets at scale.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Data Preparation

What the exam tests

What to know about Data Preparation

Be able to build a governed pipeline from raw volume files to a Delta training or retrieval table, using Unity Catalog for lineage and DLT plus ai_query for labeling. The key is materializing expensive model outputs so they are not recomputed on every run.

Using Unity Catalog lineage and Delta table history to trace training-set provenance

Building DLT pipelines that call ai_query to label or classify raw text

Choosing chunk size and overlap for text before vectorization

Reading PDF, DOCX, and text from Unity Catalog volumes into a unified dataset

Watch out for

Common Data Preparation exam traps

  • ▸Treating ai_query as cheap: rerunning it on every pipeline refresh instead of caching or materializing labeled results in a Delta table
  • ▸Picking chunking parameters by intuition rather than matching them to embedding model context limits and retrieval query shape
  • ▸Assuming raw files in a volume are queryable directly, skipping parsing and normalization into Delta tables

Practice set

Data Preparation questions

20 questions · select your answer, then reveal the explanation

A team is preparing data for a machine learning model and needs to handle missing values in a feature vector. Which TWO techniques are standard practices in Databricks for handling nulls in large-scale feature engineering pipelines?

A Data Engineer needs to handle schema evolution in a Delta Lake pipeline where source systems frequently add new columns. Which configuration ensures that the write operation automatically updates the target table schema without manual intervention?

Question 3mediummultiple choice
Read the full Data Preparation explanation →

When preparing data for a high-frequency streaming dashboard, a Data Engineer finds that the downstream Delta table is being flooded with small files, causing slow reads. Which tool should be used to rectify this without stopping the streaming job?

You are migrating a legacy ETL pipeline to Databricks. Which THREE of the following steps are considered best practices for optimizing data preparation in a Delta Lake environment?

Refer to the exhibit. You are implementing a data quality framework using Delta Live Tables. What happens when a record arrives with a null value in the 'id' column?

Exhibit

{
  "rule": "enforce_schema",
  "action": "drop_nulls",
  "target_col": "id",
  "fail_on_error": true
}
Question 6mediummultiple choice
Read the full Data Preparation explanation →

During data preparation, you observe severe data skew in a join operation. Which strategy is best for mitigating this skew?

Refer to the exhibit. The join fails due to ambiguous column names. How can you resolve this while keeping the 'id' column in the final result?

Exhibit

Error: AnalysisException: 'Multiple sources found for column id.'

val df1 = spark.table("table_a")
val df2 = spark.table("table_b")
val joined = df1.join(df2, "id")
Question 8mediummultiple choice
Read the full Data Preparation explanation →

A Data Engineer is using Databricks to prepare data for a machine learning model. The source data has significant imbalance. Which technique is most effective for addressing this imbalance during the preparation stage?

Question 9mediummultiple choice
Read the full Data Preparation explanation →

You are preparing a dataset by calculating a rolling average over a window of 30 days. Which Window function specification is required to ensure the average correctly considers only the preceding 30 days?

Question 10mediummultiple choice
Read the full Data Preparation explanation →

A data engineer is preparing a dataset for a LLM fine-tuning task. The raw data consists of millions of JSON files stored in Unity Catalog. Which approach provides the highest performance for loading this data into a PyTorch-based training pipeline?

Which TWO of the following steps are essential when cleaning unstructured text data for RAG applications within Databricks?

Question 12mediummultiple choice
Read the full Data Preparation explanation →

Refer to the exhibit. A data engineer is using a JSON-based configuration to prepare log data. Given the policy is 'strict_schema', what happens if an incoming log entry contains a field not defined in the Delta table schema?

Exhibit

{
  "policy": "strict_schema",
  "data_source": "s3://raw-logs/",
  "transformation": "regex_filter",
  "sink": "delta_table"
}

Which THREE techniques are commonly used to optimize data preparation for embedding generation in Databricks?

An engineer notices that a feature vector table in Delta is becoming excessively fragmented after frequent 'upserts'. Which action should be taken to optimize the table for future reads during inference?

Refer to the exhibit. You are appending new, streaming JSON logs to an existing Delta table. What is the most appropriate way to handle this schema evolution?

Exhibit

Error: PySparkException: [CANNOT_MERGE_INCOMPATIBLE_DATATYPES] Cannot merge incompatible data types 'String' and 'Integer'.
Question 16mediummultiple choice
Read the full Data Preparation explanation →

An enterprise engineering team is preparing unstructured PDF invoices stored in a Unity Catalog volume for a Retrieval-Augmented Generation pipeline. Which ingestion approach best aligns with Databricks-recommended data preparation practices for generative AI?

Question 17mediummultiple choice
Read the full Data Preparation explanation →

A data engineer is preparing a large set of policy PDFs stored in a Unity Catalog volume for a RAG application. Before generating embeddings, they must remove page headers, footers, and page numbers that repeat across nearly every page. Which approach is most appropriate for cleaning this text at scale on Databricks?

Question 18mediummultiple choice
Read the full Data Preparation explanation →

A GenAI team is preparing a fine-tuning dataset stored as a Delta table in Unity Catalog. They need to ensure that each training example is a valid JSON string conforming to a specific schema before feeding it to the model. They want to validate and optionally quarantine malformed records in a single pass. Which Databricks SQL function is most appropriate for this task?

A data engineer is preparing a large corpus of text for a RAG application. The text is stored in a Delta table with columns doc_id and content. The engineer needs to split the content into chunks of approximately 512 tokens with a 50-token overlap, and then compute embeddings for each chunk. They want to use Databricks-native capabilities to minimize data movement and cost. Which two approaches are valid and efficient for this task? (Choose two.)

A data engineer is preparing a dataset for training a generative AI model. The dataset contains a column 'text' with raw user reviews. Before tokenization, the engineer needs to remove HTML tags and special characters, convert text to lowercase, and expand common contractions. Which Databricks feature or approach is best suited for this text normalization step?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Data Preparation sessions

Start a Data Preparation only practice session

Every question in these sessions is drawn from the Data Preparation domain — nothing else.

Related practice questions

Related Databricks-GenAI-Assoc topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-GenAI-Assoc exam test about Data Preparation?
Be able to build a governed pipeline from raw volume files to a Delta training or retrieval table, using Unity Catalog for lineage and DLT plus ai_query for labeling. The key is materializing expensive model outputs so they are not recomputed on every run.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Data Preparation questions in a focused session?
Yes — the session launcher on this page draws every question from the Data Preparation domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-GenAI-Assoc topics?
Use the topic links above to move to related areas, or go back to the Databricks-GenAI-Assoc question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-GenAI-Assoc exam covers. They are not copied from any real exam or dump site.