Courseiva

CCNA Data Preparation Questions

40 questions · Data Preparation · All types, answers revealed

1
Multi-Selectmedium

A data engineer is preparing a large corpus of support tickets stored in a Unity Catalog volume for fine-tuning a Llama model on Databricks. They must remove personally identifiable information (PII) before the data reaches the training cluster. Which TWO approaches are appropriate for detecting and redacting PII at scale in this pipeline? (Choose two.)

Select 2 answers
A.Apply a Spark NLP or presidio-based pandas UDF that detects entities such as names, emails, and phone numbers and replaces them with placeholders.
B.Run a Databricks job that uses the ai_query() function with a foundation model endpoint to classify and rewrite each ticket, removing PII.
C.Set the table's retention policy to 0 days so that raw tickets are deleted immediately after ingestion.
D.Use Unity Catalog column masks on the raw text column so that anyone querying the table sees redacted values.
E.Enable server-side encryption on the Unity Catalog volume and rely on the storage layer to anonymize the text.
AnswersA, B

A pandas UDF running a PII detection library like Presidio or Spark NLP processes partitions in parallel and can redact entities before data is written to the training table. This keeps the redaction inside the lakehouse, scales with the cluster, and integrates with Unity Catalog governance. It is a common pattern for pre-training data sanitization.

Why this answer

PII must be removed from the content itself before training. Distributed detectors such as Presidio or Spark NLP in pandas UDFs, and model-based rewriting with ai_query(), both transform the text at scale inside Databricks. Column masks, encryption, and retention policies govern access or lifecycle but leave the underlying tokens intact, so they do not satisfy the preprocessing requirement.

Exam trap

The trap here is confusing access controls like column masks or encryption with actual content redaction, which must alter the text before training.

2
MCQhard

You are building a data preparation pipeline for a generative AI application. The raw data is stored in a Unity Catalog volume as JSON files with nested fields. You need to flatten the nested structure and extract specific fields into a Delta table for downstream embedding. The JSON schema may evolve, with new fields added occasionally. Which approach provides the most robust and maintainable solution?

A.Use the explode function on the nested arrays and then select the required fields.
B.Use the from_json function with a predefined schema and select the required fields.
C.Use the schema_of_json function to infer the schema at runtime and then apply from_json.
D.Use the variant data type to store the JSON and then use variant_get to extract fields.
AnswerD

The variant data type in Databricks can store semi-structured JSON without a predefined schema, and variant_get allows extracting fields by path. This handles schema evolution gracefully because new fields are automatically included in the variant. It is a robust, maintainable solution for evolving JSON, and it integrates with Delta tables for downstream processing.

Why this answer

The variant data type stores JSON without a rigid schema, and variant_get extracts fields dynamically, accommodating schema evolution. This avoids brittle predefined schemas and manual updates. Other options either require static schemas, rely on sampling, or only handle arrays, making them less robust for evolving nested JSON.

Exam trap

The trap here is assuming that inferring a schema at runtime solves schema evolution, when only a flexible type like variant truly accommodates new fields without pipeline changes.

3
MCQhard

A Generative AI engineer is preparing a Delta table of support tickets for embedding generation. The pipeline computes embeddings with ai_query and stores them in a separate embeddings table keyed by ticket_id. Tickets are frequently updated by agents, and the engineer wants the embeddings table to reflect the latest ticket text without recomputing embeddings for tickets whose text has not changed. Which design best achieves this?

A.Add a hash column of the ticket text and use a QUALIFY clause to filter unchanged rows within the same query.
B.Use change data capture with APPLY CHANGES to merge ticket updates into the embeddings table, recomputing embeddings only for changed rows.
C.Enable Auto Optimize on the embeddings table so only modified files are rewritten during updates.
D.Run a full refresh of the embeddings table on every pipeline update so all ticket embeddings stay current.
AnswerB

APPLY CHANGES processes the change feed and applies inserts, updates, and deletes to the embeddings table. By deriving embeddings from the changed rows only, unchanged tickets keep their existing vectors and are not re-embedded. This matches the requirement to reflect updates while avoiding redundant embedding computation.

Why this answer

To update embeddings only for changed tickets, the pipeline must consume a change feed and apply changes to the embeddings table. APPLY CHANGES merges inserts, updates, and deletes while allowing embedding derivation to run only on changed rows. Full refreshes recompute everything, and file-level or query-level filters cannot persist decisions across updates.

Exam trap

The trap here is assuming a hash or file-optimization feature can skip unchanged rows across runs, when only a change data capture flow persists that decision.

4
MCQmedium

A GenAI engineer is preparing a large text corpus stored in a Unity Catalog volume for fine-tuning a chat model. The raw files are JSON Lines, each containing a 'conversation' array with alternating 'user' and 'assistant' turns, but many records have malformed turns or missing roles. The engineer needs to convert this into a Delta table with a schema of (conversation_id STRING, messages ARRAY<STRUCT<role:STRING, content:STRING>>) while filtering out records where any turn has a null role or empty content. Which approach uses the appropriate Databricks-native capability for this transformation?

A.Load the files as a text RDD, apply a regular expression to extract role and content pairs, and then convert the RDD to a DataFrame with the desired schema using createDataFrame.
B.Use the from_json function with a defined StructType schema to parse each JSON line into the conversation structure, then apply a filter using a higher-order function such as filter on the messages array to retain only records where all turns have non-null role and non-empty content.
C.Use Databricks Auto Loader with schema evolution to ingest the files into a Bronze table, then use Delta Live Tables expectations to drop rows where the role column is null or the content column is empty.
D.Use the spark.read.json method with inferSchema enabled to automatically detect the schema, then use a UDF written in Python to iterate over each turn and remove records with null roles or empty content.
AnswerB

from_json with an explicit StructType accurately parses nested JSON into the required ARRAY<STRUCT<role:STRING, content:STRING>> schema. Applying a higher-order filter (e.g., array filtering) or an exists/forall expression on the parsed array lets you discard records with malformed turns. This leverages Spark SQL's native JSON and array functions, which are designed for scalable, schema-enforced transformations on large corpora, exactly matching the need to enforce structure and filter invalid records.

Why this answer

The correct approach uses from_json with an explicit schema to parse the nested JSON into the required array of structs, then applies a higher-order function to filter out records with malformed turns. This leverages Spark SQL's native, optimized JSON parsing and array operations, which scale well and enforce the target schema. It avoids the performance pitfalls of UDFs, the fragility of regex, and the indirectness of ingestion frameworks not designed for immediate nested transformation.

Exam trap

The trap here is assuming that schema inference or UDF-based cleanup is necessary when native from_json and higher-order functions can both parse and validate the nested structure in a single, optimized pass.

5
Multi-Selecthard

A Databricks team is building a retrieval corpus from mixed-format documents stored in a Unity Catalog volume. They need a preparation pipeline that preserves document structure for later chunking and that records which source file each chunk came from so retrieval results can cite evidence. Which two design choices best meet these requirements? (Choose two.)

Select 2 answers
A.Concatenate all documents into a single large text field before chunking to reduce the number of rows.
B.Carry the source file path as a metadata column on each chunk and store it in the Delta table alongside the chunk text and embedding.
C.Convert every document to plain text and discard page and heading markers to normalize the corpus.
D.Generate embeddings for entire documents rather than chunks to avoid storing multiple rows per file.
E.Use the ai_parse_document function to extract text and layout elements from PDFs and images, retaining the document structure for downstream chunking.
AnswersB, E

Attaching the source file path as a metadata column lets retrieval results reference the originating document, enabling citations and traceability. Delta tables support arbitrary metadata columns, so the path can travel with each chunk through embedding generation and into the vector index without additional joins.

Why this answer

Preserving document structure requires a parser that extracts layout elements rather than flattening everything to plain text, and ai_parse_document provides that for PDFs and images. Provenance requires carrying the source file path on each chunk so retrieval results can cite the originating document. Together these choices keep structure intact and make every chunk traceable.

Exam trap

The trap here is treating normalization to plain text as harmless, when discarding layout markers actually removes the structure needed for coherent chunking and provenance.

6
Multi-Selecthard

You are preparing a large text corpus for fine-tuning a generative AI model. The corpus is stored in a Delta table with columns: doc_id, raw_text, and metadata. You need to create a cleaned dataset that removes personally identifiable information (PII) and normalizes whitespace, while preserving document boundaries for training. Which two actions should you perform to achieve this in a scalable and maintainable way? (Choose two.)

Select 2 answers
A.Convert the Delta table to Parquet files and manually inspect each file for PII.
B.Use the ai_classify function to label each document as 'clean' or 'dirty' and drop dirty ones.
C.Use the ai_analyze_sentiment function to filter out negative documents before training.
D.Apply a regular expression to replace all whitespace sequences with a single space in the raw_text column.
E.Use the ai_mask function to redact PII entities from the raw_text column.
AnswersD, E

Normalizing whitespace with a regular expression such as regexp_replace(raw_text, '\\s+', ' ') collapses multiple spaces, tabs, and newlines into single spaces. This standardizes the text for tokenization and reduces noise without altering document boundaries. It is a scalable and deterministic step that can be applied in a SELECT or withColumn transformation.

Why this answer

ai_mask redacts PII in place, and regexp_replace normalizes whitespace, together achieving the cleaning goals at scale. Both are declarative, repeatable, and preserve document boundaries. Sentiment filtering, manual inspection, and classification are either irrelevant or destructive to the dataset, and they do not address PII removal and whitespace normalization.

Exam trap

The trap here is thinking that any AI function can clean text, when only ai_mask targets PII and regex handles whitespace; other AI functions like sentiment or classification do not perform the required transformations.

7
MCQmedium

You are preparing a large corpus of customer support emails stored as Parquet files in a Unity Catalog volume. The emails must be cleaned by removing boilerplate signatures and disclaimers before tokenization for a fine-tuning dataset. Your team wants to enforce this transformation as a declarative, testable pipeline stage that fails fast if any cleaned record still contains a known boilerplate marker. Which Databricks capability should you use to implement this cleaning step with built-in data quality expectations?

A.Use Spark Structured Streaming with a foreachBatch function that calls a Python UDF to strip signatures and raises an exception if a marker remains.
B.Create a Databricks SQL query that applies regexp_replace to the email body and schedule it as a SQL warehouse alert when a marker is detected.
C.Write a notebook that reads the Parquet files, applies a pandas UDF to clean the text, and writes the result back to the volume, then manually review the output.
D.Define the cleaning logic in a Delta Live Tables pipeline and attach an EXPECT constraint that asserts no boilerplate marker remains in the cleaned column.
AnswerD

Delta Live Tables pipelines support declarative expectations that can drop, fail, or quarantine records based on a condition. By attaching an EXPECT constraint to the cleaned table, the pipeline fails fast if any record still contains a boilerplate marker, and the expectation metrics are automatically captured in the event log for auditing and testing.

Why this answer

Delta Live Tables expectations are the declarative, testable mechanism for enforcing data quality during preparation. Attaching an EXPECT constraint to the cleaned table ensures the pipeline fails fast if any boilerplate marker remains, while capturing metrics for auditing. This approach provides lineage, observability, and reproducibility, unlike imperative notebooks or asynchronous alerts.

Exam trap

The trap here is assuming that any code that raises an exception on bad data is equivalent to a declarative data quality expectation, when only Delta Live Tables expectations provide built-in fail-fast semantics and metrics.

8
MCQhard

A team is preparing a pretraining corpus and must filter out near-duplicate documents before tokenization. The corpus contains 500 million short text records in a Delta table. They want a scalable, deterministic deduplication signal that can be computed per record and compared across the dataset. Which approach best fits this requirement?

A.Sort the corpus by document length and drop every record whose length falls within one standard deviation of another record.
B.Compute an MD5 hash of each full document and drop rows sharing the same hash value.
C.Compute a MinHash signature per document with a fixed number of hash permutations and compare signatures with a locality-sensitive hashing banding scheme.
D.Train a sentence embedding model and cluster documents with k-means, then drop all but one document per cluster.
AnswerC

MinHash with LSH banding produces a fixed-length signature per document and a scalable candidate-generation step, so near-duplicates are found without all-pairs comparison. It is deterministic given fixed seeds and permutations, distributes cleanly across Spark partitions, and directly targets near-duplicate detection rather than exact matches, fitting the 500 million record corpus.

Why this answer

Near-duplicate detection at scale needs a compact per-record signature plus a candidate-generation strategy, which MinHash with locality-sensitive hashing banding provides. Exact hashing misses near-duplicates, embedding clustering is costly and nondeterministic, and length-based filtering has no content signal, so none of those satisfy the deterministic, scalable requirement.

Exam trap

The trap here is equating deduplication with exact-match hashing, when near-duplicate removal requires a similarity-preserving signature rather than a cryptographic digest.

9
MCQeasy

A data engineer is preparing a dataset of product descriptions for embedding generation. The source table in Unity Catalog contains a column 'description' with mixed languages, and the team wants to filter to English-only text before vectorization. They need a scalable, built-in Databricks function that can detect the language of each description without external API calls. Which function should they use?

A.regexp_extract
B.ai_detect_language
C.ai_classify
D.ai_analyze_sentiment
AnswerB

ai_detect_language is a built-in Databricks SQL function that returns the detected language code for a given text column. It is designed for scalable language identification directly in SQL or PySpark, requiring no external API calls. Filtering on the returned language code allows the team to keep only English descriptions before embedding, exactly matching the requirement.

Why this answer

The built-in ai_detect_language function is purpose-built for language identification at scale within Databricks. It returns a language code that can be used to filter the dataset to English-only records before embedding generation. Other AI functions like ai_analyze_sentiment or ai_classify serve different purposes, and regex cannot reliably determine language.

Exam trap

The trap here is confusing general-purpose AI functions such as ai_classify with a dedicated language detection function, when only ai_detect_language directly returns a language code for filtering.

10
MCQeasy

Which Databricks feature is best suited for maintaining the lineage of data used during the preparation of training sets for Generative AI?

A.Delta Lake Time Travel
B.Unity Catalog
C.Cluster Policies
D.MLflow Experiments
AnswerB

Unity Catalog captures fine-grained lineage information as data moves from raw ingestion to final training sets. It offers a unified view of dependencies, which is essential for auditing the training process. This visibility is crucial for ensuring that training data meets safety and compliance standards in enterprise environments.

Why this answer

Unity Catalog provides centralized governance, including automated data lineage tracking. This allows engineers to trace the source of data used for model training, which is vital for reproducibility and regulatory compliance. Knowing the exact provenance of training data helps in troubleshooting model bias and ensures that sensitive data sources are correctly identified and managed throughout the lifecycle of the model training process.

Exam trap

Candidates often suggest external logging tools or manual documentation, failing to recognize that Unity Catalog provides native, automated lineage tracking that is integrated directly into the data platform.

11
MCQhard

You need to store embeddings generated by an LLM in a Delta table. Which data type is most efficient for storing these high-dimensional vector arrays in Databricks?

A.Store as a Base64 encoded string.
B.Store as an array of doubles.
C.Store as separate columns for each dimension (e.g., dim1, dim2, ...).
D.Store as a compressed binary blob (e.g., serialized byte array).
AnswerB

An array of doubles is the standard format for vector embeddings in Delta Lake. It is natively supported, memory-efficient, and easily accessible by Spark's vectorized query engine and popular libraries like FAISS or MosaicML Vector Search. This format provides the best balance between storage performance and computational efficiency for retrieval.

Why this answer

Storing embeddings as an 'Array' of 'Floats' (or 'Doubles') in a Delta table is the most efficient and native way to handle them. This allows Spark to utilize vectorized operations for similarity search calculations and enables compatibility with various indexing libraries. Using this approach ensures high performance during retrieval, minimizing the latency of finding the most relevant context for the LLM during generation.

Exam trap

Candidates often assume complex custom objects or string-serialized JSON are necessary. Delta natively supports arrays, which are significantly more efficient for vector math and similarity searches.

12
MCQmedium

A GenAI engineer is preparing a corpus of HTML product pages stored in a Unity Catalog volume for a RAG application. The pages contain navigation bars, script tags, and boilerplate footers that add noise to embeddings. Which Databricks-native approach best removes this noise while preserving the main article text before chunking?

A.Register the volume as a Delta table and run a SQL query using regexp_replace to strip all tags, then tokenize the result with the built-in Databricks tokenizer.
B.Increase the chunk size to 4096 tokens so that boilerplate is a smaller proportion of each chunk, and rely on the embedding model to ignore the noise.
C.Load the HTML files with spark.read.text and apply a VectorAssembler to convert the raw strings into dense vectors, then cluster and drop outliers.
D.Use an ai_parse_document or BeautifulSoup-style parser in a PySpark UDF to extract the main content, then write the cleaned text as a Delta table with a content column.
AnswerD

Parsing HTML with a document-aware parser lets you target the main article region and discard navigation, script, and footer nodes before text is written. Storing the cleaned content in Delta makes downstream chunking and embedding reproducible. This is the Databricks-native pattern for turning raw HTML in a volume into analysis-ready text.

Why this answer

Cleaning HTML before chunking requires a parser that understands document structure, not a regex or a numeric feature transformer. Extracting the main content with a document-aware parser and persisting the cleaned text to Delta produces a reliable, lineage-tracked input for chunking and embedding, which is the goal of data preparation for RAG on Databricks.

Exam trap

The trap here is assuming that a simple regex tag-stripping step or a larger chunk size is equivalent to semantic HTML cleaning.

13
MCQhard

A team is building a RAG application using Databricks Vector Search with a Delta table as the source. They need the index to automatically reflect new and updated chunks as the source table changes, without rebuilding the entire index each time. Which configuration should they use?

A.Enable Delta Lake time travel on the source table and point the Vector Search index at a specific version using VERSION AS OF.
B.Use an external HNSW index built on a Databricks cluster and register it in Unity Catalog as a model, then query it with the model serving endpoint.
C.Create a Vector Search index with sync mode set to TRIGGERED or CONTINUOUS and specify a Delta table with Change Data Feed enabled as the source.
D.Create a standard Vector Search index and schedule a nightly job that drops the index and recreates it from the full Delta table.
AnswerC

Databricks Vector Search supports Delta Sync indexes that read Change Data Feed from the source Delta table to incrementally update embeddings. Enabling CDF and choosing a triggered or continuous sync keeps the index fresh without full rebuilds. This is the supported pattern for production RAG pipelines where source documents are frequently updated.

Why this answer

Delta Sync indexes in Databricks Vector Search rely on Change Data Feed to pick up inserts, updates, and deletes incrementally. Choosing a triggered or continuous sync mode keeps embeddings current without full recomputation. Enabling CDF on the source Delta table is a prerequisite, and this pattern is the documented way to keep a RAG index fresh as documents change.

Exam trap

The trap here is thinking that Vector Search automatically tracks changes to any Delta table, when incremental sync requires Delta Sync and Change Data Feed to be configured.

14
MCQhard

A GenAI engineer prepares a Delta table of product descriptions that will feed a chunking and embedding pipeline. The descriptions are written in mixed languages, and the embedding model supports only English. The engineer needs to ensure that non-English rows are detected and routed for translation before embedding. Which approach is most appropriate?

A.Translate every description into English before chunking, regardless of its original language, to guarantee uniform embeddings.
B.Add a language detection UDF that runs ai_classify or a language-identification library on each description, write the detected language to a column, and split the table into English and non-English subsets.
C.Drop all rows where the description contains non-ASCII characters, since those rows are likely non-English.
D.Configure the embedding model endpoint to accept a language parameter and pass the detected locale from Unity Catalog tags on the table.
AnswerB

Detecting language per row and persisting the result creates an auditable routing column. Splitting the table lets English rows proceed directly to chunking and embedding while non-English rows go to a translation step. This is deterministic, testable data preparation and it keeps the pipeline extensible if more languages are added later.

Why this answer

When the embedding model is English-only, the pipeline must identify which rows need translation. Detecting language per row and storing the result as a column gives a deterministic routing signal, so English rows embed directly while non-English rows are sent to translation. This keeps the preparation step auditable and avoids both data loss and unnecessary translation cost.

Exam trap

The trap here is assuming that a table-level tag or a character-range filter can substitute for per-row language detection.

15
MCQeasy

Which of the following describes the 'Gold' layer in a Medallion architecture, and why is it important for GenAI data preparation?

A.It contains raw, unprocessed data for initial exploration.
B.It stores transient data used only for debugging intermediate steps.
C.It is the final, curated state of data, optimized for consumption by ML models.
D.It is a temporary cache for speeding up cluster startup times.
AnswerC

The Gold layer provides clean, validated data. For GenAI, this means the text is pre-processed, chunks are optimized, and PII is scrubbed. By consuming data from the Gold layer, engineers ensure that their models are learning from the highest quality sources, which significantly improves overall application performance and reliability.

Why this answer

The Gold layer represents highly refined, business-level aggregates or prepared datasets ready for consumption. In GenAI, this is where the final, cleaned, and curated training sets (or vector-ready documents) reside. Having a Gold layer ensures that models are trained on validated, high-quality data, which is fundamental to building reliable, production-grade Generative AI applications that meet organizational standards for accuracy and data governance.

Exam trap

Candidates frequently mistake the Gold layer for the 'Silver' layer, which is cleaned but not necessarily aggregated or business-ready for final consumption by GenAI applications.

16
MCQmedium

A GenAI engineer must build a fine-tuning dataset from 40 TB of raw JSONL conversation logs stored in cloud object storage. The logs are immutable and only ever read once, in full, during preprocessing. Which storage and access configuration should be used to minimize cost while keeping the data readable by Spark on Databricks?

A.Convert the JSONL to Parquet with an external tool, upload the Parquet back, then read it from the bucket.
B.Run COPY INTO to ingest the JSONL into a managed Delta table, then read the Delta table for preprocessing.
C.Register the cloud path as an external location in Unity Catalog and read the JSONL files directly with Spark.
D.Mount the object storage bucket to DBFS with a legacy mount and read files through the /mnt path.
AnswerC

Reading the immutable JSONL directly from its object-storage path through a Unity Catalog external location avoids copying 40 TB and avoids paying for a second persisted copy. Unity Catalog governs access while Spark reads the files in place, which matches the single-pass, read-only access pattern and minimizes both storage and egress cost.

Why this answer

Because the logs are immutable and read once in full, the lowest-cost approach is to leave them where they are and read them in place under Unity Catalog governance. Copying or converting the data creates a redundant full-size copy and extra compute for no reuse benefit, while legacy mounts sacrifice governance without reducing storage or access cost.

Exam trap

The trap here is assuming that ingesting raw logs into Delta is always the right first step, when single-pass immutable reads are cheaper and better governed when left in place.

17
Multi-Selecthard

A GenAI engineer is preparing a large text corpus for fine-tuning an LLM. The corpus contains many near-duplicate documents and documents in multiple languages. They need to reduce redundancy and ensure language consistency. Which two steps should be performed during data preparation? (Choose two.)

Select 2 answers
A.Use the language detection function from the spark-nlp library to filter documents to a single target language.
B.Train a custom tokenizer on the entire corpus to better handle multilingual text.
C.Use Delta Lake time travel to revert to a previous version of the dataset if duplicates are found.
D.Compute MinHash signatures and apply locality-sensitive hashing to identify and remove near-duplicate documents.
E.Apply a fixed-size chunking strategy to split all documents into 512-token segments before deduplication.
AnswersA, D

Language detection identifies the language of each document, allowing you to filter out documents that do not match the target language. This ensures consistency and prevents the model from learning from irrelevant languages, which is critical when fine-tuning for a specific language task.

Why this answer

To reduce redundancy and ensure language consistency, the engineer should use MinHash with LSH to remove near-duplicate documents and apply language detection to filter to the target language. These steps directly address the stated goals and are scalable for large corpora.

Exam trap

The trap here is confusing tokenization or versioning features with actual data cleaning steps, leading to choices that do not remove duplicates or filter languages.

18
Multi-Selectmedium

Which TWO of the following are benefits of using Delta Lake for data preparation over standard Parquet files on cloud storage?

Select 2 answers
A.ACID compliance for reliable write operations.
B.Automatic conversion of JSON to XML format.
C.Time travel capability for auditing and debugging.
D.Native support for cross-cloud streaming ingestion.
E.Automatic hardware upgrades for compute clusters.
AnswersA, C

ACID compliance ensures that concurrent reads and writes are managed correctly, preventing data corruption during simultaneous operations. In standard Parquet, a failing write operation could leave orphan files or partial data, whereas Delta Lake ensures an 'all-or-nothing' consistency model that is critical for production pipelines.

Why this answer

Delta Lake provides ACID transactions and time travel, which are essential for robust data engineering. ACID transactions ensure that data preparation pipelines never leave tables in a partially written state, while time travel allows for auditing and reverting changes. These features significantly simplify data lifecycle management, reduce the need for custom retry logic, and enhance the overall reliability of data preparation workflows in complex, multi-user production environments.

Exam trap

Candidates often conflate Delta Lake with general data formats, failing to identify that ACID transactions and time travel are specific, functional advantages that distinguish Delta from standard Parquet files.

19
MCQmedium

A team is preparing data for a RAG system and needs to remove duplicates from a large collection of PDF text extracts. What is the most efficient way to perform de-duplication in Databricks?

A.Use a Python 'set' in a local loop on the driver node.
B.Execute a distributed 'dropDuplicates' operation in Spark.
C.Manually compare every document against every other document using nested loops.
D.Ignore duplicates, as vector databases automatically handle them during insertion.
AnswerB

Spark's 'dropDuplicates' is optimized for large, distributed datasets. It efficiently identifies and removes duplicate rows across the entire cluster, making it the standard approach for large-scale de-duplication. This ensures high-quality training sets and efficient vector indices without requiring complex custom code for distributed data handling.

Why this answer

Using Spark's 'dropDuplicates()' on the text content column or calculating a hash (e.g., MD5) of the text to identify duplicates is highly scalable. This is crucial because redundant information in the vector database can cause the retriever to prioritize duplicate entries, leading to biased results and inefficiency. Removing duplicates ensures that the search index remains focused and that the retrieved context is diverse and informative.

Exam trap

Candidates often attempt single-node Python loops or pandas-based string matching on massive text corpora, resulting in out-of-memory errors on large datasets.

20
MCQmedium

A GenAI engineer is preparing a fine-tuning dataset from a Delta table in Unity Catalog that contains raw user feedback. The feedback text includes irregular capitalization, HTML tags, and excessive punctuation. The engineer needs to normalize the text using Spark NLP within a Databricks notebook, ensuring the pipeline is reproducible and scalable. Which approach should the engineer use to apply this transformation?

A.Use Delta Live Tables to define a streaming pipeline that applies the lower() and trim() functions to the text column.
B.Use the spark-nlp library's DocumentAssembler and Normalizer annotators in a Spark ML Pipeline, and save the pipeline to MLflow.
C.Use Databricks SQL's regexp_replace function in a SELECT statement to clean the text, and create a new table with the cleaned data.
D.Use pandas UDFs with Python's re module to apply custom cleaning functions to each row, and cache the resulting DataFrame.
AnswerB

Spark NLP provides DocumentAssembler and Normalizer annotators that can be combined in a Spark ML Pipeline, enabling scalable and reproducible text normalization. Saving the pipeline to MLflow ensures versioning and reproducibility. This approach leverages Spark's distributed processing and integrates with Databricks workflows, making it ideal for preparing large-scale fine-tuning datasets.

Why this answer

The correct approach uses Spark NLP's DocumentAssembler and Normalizer within a Spark ML Pipeline, which provides distributed, reproducible text normalization. Saving the pipeline to MLflow ensures version control and reusability. This method is scalable and integrates well with Databricks, making it suitable for preparing large fine-tuning datasets with complex cleaning needs.

Exam trap

The trap here is assuming that simple string functions or regex alone suffice for comprehensive text normalization, overlooking the need for specialized NLP annotators and pipeline reproducibility.

21
MCQmedium

A data engineer is preparing a dataset for fine-tuning a chat model. The dataset contains conversations with alternating user and assistant messages. They need to format each conversation into a single string with special tokens indicating roles. Which approach is most appropriate in Databricks?

A.Use the format_string function to insert role tokens into a template string for each message.
B.Use the explode function to flatten the messages and then collect_list to reassemble them with role tokens.
C.Use the ai_query function to call an LLM that formats the conversation into a single string.
D.Use the concat_ws function to join messages with a delimiter, and prefix each message with a role token.
AnswerD

concat_ws can join an array of strings with a delimiter. By first transforming each message into a string with a role prefix, you can create a formatted conversation string. This is a straightforward and scalable way to prepare chat data for fine-tuning.

Why this answer

The most appropriate approach is to use concat_ws to join messages after prefixing each with a role token. This creates a single formatted string per conversation efficiently and deterministically, which is ideal for fine-tuning chat models.

Exam trap

The trap here is overcomplicating the task by using an LLM or misusing aggregation functions, when a simple string concatenation function suffices.

22
MCQmedium

You need to ingest data from an external JSON source into a Delta table. The source schema is inconsistent. Which strategy is most effective for preparing this data?

A.Use inferSchema = true in the read configuration.
B.Define a rigid schema at the ingestion point.
C.Implement a bronze-silver medallion pattern.
D.Convert the JSON to CSV before loading into Delta.
AnswerC

The bronze-silver pattern enables ingestion of raw, unstructured data into a staging area, followed by rigorous cleaning and validation in a downstream silver table. This approach prevents pipeline failures during ingestion and ensures that cleaning logic is decoupled from the data acquisition layer, allowing for better maintainability.

Why this answer

The 'bronze-to-silver' pattern is the industry standard in Databricks for handling inconsistent data. By loading raw JSON into a 'Bronze' table with a schema of type 'string' (or a single JSON column), you preserve all information for audit. You then use 'Silver' tables to perform schema enforcement, data cleaning, and type casting.

This separation of concerns allows for robust error handling and iterative refinement of the data cleaning logic.

Exam trap

Candidates often suggest cleaning data directly into a final table, overlooking the importance of the Bronze layer for preserving raw historical data before applying schema enforcement in Silver.

23
MCQeasy

A data engineer needs to persist a prepared instruction-tuning dataset so that downstream fine-tuning jobs can read it with ACID guarantees, time travel, and schema enforcement, and so that Unity Catalog can track column-level lineage. Which storage format and registration should be used?

A.Write the dataset as an ORC table in the Hive metastore.
B.Write the dataset as a Delta table registered in Unity Catalog.
C.Write the dataset as CSV files to an external location registered in Unity Catalog.
D.Write the dataset as Parquet files in a Unity Catalog volume.
AnswerB

Delta Lake provides ACID transactions, time travel, and schema enforcement on the prepared dataset, and registering the table in Unity Catalog enables column-level lineage and fine-grained governance. This combination directly satisfies every stated requirement, including the ability for downstream fine-tuning jobs to read a consistent, versioned table.

Why this answer

Delta Lake on Unity Catalog is the only option that simultaneously delivers ACID transactions, time travel, schema enforcement, and column-level lineage for the prepared dataset. Parquet, CSV, and ORC files lack transactional table semantics, and the Hive metastore does not provide the Unity Catalog lineage the team needs.

Exam trap

The trap here is treating a governed file path as equivalent to a governed table, when lineage and ACID properties come from the table format and catalog registration.

24
MCQmedium

A team is preparing a fine-tuning dataset from customer reviews stored in a Delta table. They need to filter out reviews shorter than 20 tokens and reviews flagged as spam by a classifier, then write the result to a Unity Catalog table for training. Which approach best fits Databricks best practices?

A.Load the table into a pandas DataFrame on the driver, apply filtering, and write it back using the Databricks SQL connector row by row.
B.Create a view that filters short and spam reviews, then point the fine-tuning job at the view without materializing a new table.
C.Use Delta Live Tables with a streaming table that ingests all reviews and applies expectations to drop short and spam rows at write time.
D.Use a Spark DataFrame with a token-count UDF and a filter on the spam flag, then write the result to a Unity Catalog table with .write.mode("overwrite").saveAsTable().
AnswerD

Filtering with Spark transformations keeps the work distributed and lets Catalyst optimize the plan. A token-count UDF computes length per row, and a simple filter removes spam-flagged rows. Writing with saveAsTable into Unity Catalog creates a governed training table that downstream fine-tuning jobs can read with lineage intact.

Why this answer

Batch preprocessing of an existing Delta table is best done with Spark DataFrame transformations, which distribute token counting and filtering across executors. Writing the cleaned result to a Unity Catalog table creates a governed, reproducible training artifact. Driver-side pandas and row-by-row writes do not scale, views recompute on every read, and streaming pipelines are overkill for a one-time preparation step.

Exam trap

The trap here is choosing a view or a streaming pipeline for a one-time batch cleaning task, when a materialized Spark write is the appropriate pattern.

25
Multi-Selectmedium

A Generative AI engineer is preparing a Delta table of product reviews for a retrieval-augmented generation application. The reviews contain HTML tags, inconsistent casing, and occasional very long paragraphs that exceed the embedding model's context window. The engineer wants to clean and normalize the text before chunking and embedding. Which two actions should the engineer take to directly address the stated quality issues? (Choose two.)

Select 2 answers
A.Drop every review shorter than 100 characters to reduce noise in the corpus.
B.Store the raw HTML in a separate column so it remains available for future parsing needs.
C.Split long paragraphs into overlapping chunks sized to the embedding model's token limit.
D.Strip HTML tags and normalize whitespace using a Spark transformation before chunking.
E.Convert all review text to uppercase to standardize casing across the corpus.
AnswersC, D

The reviews contain paragraphs that exceed the model context window, so chunking with overlap ensures no content is silently truncated and context is preserved across boundaries. Sizing chunks to the token limit directly addresses the stated overflow problem. Overlap helps retrieval quality by preventing answers that straddle a boundary from being split apart.

Why this answer

The reviews need markup removed and length controlled before embedding. Stripping HTML and normalizing whitespace cleans the text, while chunking with overlap sized to the model token limit prevents truncation of long paragraphs. Uppercasing, dropping short reviews, and archiving raw HTML do not resolve the stated noise and overflow problems.

Exam trap

The trap here is treating generic data hygiene steps like casing changes or row filtering as fixes, when the stated defects are markup noise and context-window overflow.

26
MCQmedium

When preparing data for a fine-tuning task, you realize the dataset is severely imbalanced. Which Databricks technique should you use to create a more balanced dataset?

A.Run a standard 'SELECT * FROM table' query.
B.Use the sampleBy() method to perform stratified sampling.
C.Increase the number of epochs during the training process.
D.Use a UDF to delete all rows that belong to majority classes.
AnswerB

Stratified sampling allows you to select specific proportions from each category, enabling effective oversampling of minority classes or undersampling of majority classes. This technique is standard in Spark for preparing balanced training datasets, ensuring that the model learns effectively from all classes, regardless of their original prevalence in the data.

Why this answer

Using Spark's 'sample()' or 'sampleBy()' transformation allows you to perform stratified oversampling or undersampling to balance classes. This ensures that the fine-tuned model doesn't become biased toward the most frequent categories in the dataset. Proper balancing is essential for ensuring robust model performance across all target classes, preventing the model from underperforming on rare but critical edge cases in real-world applications.

Exam trap

Candidates often suggest manual filtering or simple random sampling. Random sampling does not solve class imbalance, whereas stratified sampling specifically ensures minority classes are adequately represented in the training set.

27
MCQmedium

You are preparing a large dataset for fine-tuning a model using Databricks Delta Live Tables (DLT). Which configuration is best for ensuring data quality and lineage in this pipeline?

A.Use standard Spark SQL jobs without DLT.
B.Implement DLT Expectations to flag or drop invalid records.
C.Write all data to a temporary folder and clean it after training.
D.Rely on the LLM to automatically filter out low-quality data during training.
AnswerB

Expectations provide a declarative way to enforce quality in DLT pipelines. They ensure that the training data meets specific criteria, such as length or content validity. This reduces noise in the model training process, leading to better results and more reliable behavior, which is essential for enterprise GenAI deployments.

Why this answer

Delta Live Tables (DLT) with Expectations allows you to define constraints on data quality. By tagging data that fails these constraints, engineers ensure only clean data reaches the final training table. This provides built-in lineage and automatic quality monitoring.

In the context of GenAI, clean data is paramount, as noise in the training set leads to poor model performance and unpredictable behavior in downstream applications.

Exam trap

Candidates often overlook 'Expectations' as a quality tool, viewing them as optional. In a GenAI pipeline, they are essential for ensuring that only high-quality, validated data enters the model training.

28
MCQmedium

A Generative AI engineer is building a Delta Live Tables pipeline that ingests raw JSON event logs into a bronze table, then uses ai_query to classify each event's free-text field. The classification call is expensive, so the engineer wants to avoid re-running it on events that have already been processed in previous pipeline updates. The source table is append-only and new events arrive continuously. Which Delta Live Tables feature should the engineer configure on the bronze table to prevent reprocessing of previously ingested rows?

A.Configure the bronze table as a streaming table so it processes only new data on each pipeline update.
B.Add a QUALIFY clause that filters out rows where the classification column is already populated.
C.Enable change data capture by setting pipelines.cdcEnabled to true on the bronze table.
D.Set the table property pipelines.autoOptimize.managed to true on the bronze table.
AnswerA

Streaming tables in Delta Live Tables read incrementally, so on each pipeline update only newly arrived source rows flow through the ai_query classification. Previously ingested events are not re-read, which is exactly what avoids paying for repeated classification calls. This is the standard pattern for append-only ingestion with expensive downstream transformations.

Why this answer

An append-only bronze ingestion table should be defined as a streaming table so each pipeline update consumes only new source records. This makes expensive operations like ai_query run once per new event rather than on the full history. Materialized views and batch-style definitions re-read all source data, and storage properties like Auto Optimize only affect file layout, not incremental read behavior.

Exam trap

The trap here is assuming that a table property or filter clause can prevent reprocessing, when incremental read semantics in Delta Live Tables come from declaring the dataset as a streaming table.

29
Multi-Selectmedium

A data engineer is preparing a large Delta table of conversation logs for embedding generation. The table has frequent small appends, and the engineer needs to reduce file fragmentation and improve read throughput before the embedding job runs. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Run VACUUM with a retention of zero hours to delete old files and free space before the embedding job.
B.Increase the number of shuffle partitions to the maximum supported value so that each task writes a single small file.
C.Enable auto compaction and optimized writes on the Delta table so small writes are coalesced automatically.
D.Run OPTIMIZE on the table to compact small files, optionally with ZORDER on columns frequently used in filters.
E.Convert the table to a Parquet directory and rely on the file system to merge small files during reads.
AnswersC, D

Auto compaction merges small files after a write, and optimized writes shuffle data so fewer, larger files are produced in the first place. Together they keep fragmentation low between manual maintenance runs, which suits a table receiving frequent small appends. This directly addresses the file-count problem before the embedding job reads the data.

Why this answer

File fragmentation from frequent appends is addressed by consolidating small files. OPTIMIZE with optional ZORDER compacts and clusters existing data, while auto compaction and optimized writes prevent new fragmentation from accumulating. Together they reduce the number of files the embedding job must open and can improve predicate pushdown, directly improving read throughput.

Exam trap

The trap here is confusing VACUUM, which deletes unreferenced files, with OPTIMIZE, which rewrites live data into fewer files.

30
MCQmedium

A Data Engineer needs to ensure that PII data in a Delta table is masked before serving it to non-privileged users. Which Databricks feature provides the most efficient, centralized control for this requirement?

A.Apply Row-Level Security filters using Delta table constraints.
B.Create materialized views with hard-coded redacted values.
C.Utilize Unity Catalog column-level masking functions.
D.Implement Spark UDFs to perform in-memory data masking.
AnswerC

Unity Catalog enables defining dynamic masks on columns using SQL functions. This centralizes security policy enforcement, applying masking rules consistently across all users and compute clusters. It is the most robust method for PII protection because it prevents unauthorized visibility without altering the underlying raw data storage files.

Why this answer

Unity Catalog's dynamic views and column-level masking policies are the standard for securing PII. By applying functions like MASK or current_user() within a view definition, administrators decouple security logic from physical table storage. This approach is essential for compliance, ensuring that sensitive data is hidden at query time without duplicating datasets, thus maintaining a single source of truth while enforcing granular access control across all Databricks workspaces and compute resources.

Exam trap

Candidates often suggest creating separate tables or filtering via application code. Unity Catalog masking is the centralized, efficient way to handle this without duplicating data or creating security gaps.

31
Multi-Selecthard

Which TWO factors are most important when selecting a chunking strategy for text data prior to vectorization?

Select 2 answers
A.The maximum input token length of the target embedding model.
B.The file format of the source documents (e.g., PDF vs. Word).
C.The desired level of semantic granularity for retrieval.
D.The total number of documents in the corpus.
E.The storage cost of the resulting vector database.
AnswersA, C

Embedding models have strict limits on the number of tokens they can process in a single sequence. Exceeding this limit leads to truncation, which destroys semantic context and makes the resulting vectors inaccurate. Aligning chunk size with model token limits is a critical technical requirement for successful data preparation.

Why this answer

Selecting a chunking strategy requires balancing semantic context and technical limitations like context window size. If chunks are too small, they lack meaning; if too large, they exceed the limits of the embedding model and introduce irrelevant noise. Proper chunking is vital for maximizing the accuracy of RAG systems, as it determines the granularity of the information available for the retriever to present to the LLM.

Exam trap

Candidates often focus solely on the embedding model's limits while forgetting that the chunking strategy must also align with the business goal of retrieving semantically meaningful information.

32
MCQeasy

A team stores raw documents in a Unity Catalog volume and needs to build a training set for fine-tuning a large language model. The raw files are in mixed formats including PDF, DOCX, and plain text. The team wants a single Delta table where every row is one document with its extracted text and source path, and wants the extraction to run in parallel across the cluster. Which approach should the team use?

A.Mount the volume as a local filesystem path and use a Python for-loop to read and extract each file on the driver.
B.Use binaryFile data source to read the volume, then apply a UDF that extracts text per file and writes results to a Delta table.
C.Use spark.read.text on the volume, which automatically parses PDF and DOCX content into text columns.
D.Create an external table over the volume with a schema that includes columns for each document format's fields.
AnswerB

The binaryFile data source reads each file as a row with path and binary content, which lets Spark distribute file processing across executors. Applying an extraction UDF per row produces one output row per document with text and source path. Writing to Delta gives the unified table the team wants for fine-tuning.

Why this answer

The binaryFile data source is designed to read arbitrary files as rows with path and content, enabling distributed extraction with a UDF. This yields one row per document with text and source path in a Delta table. Driver loops and plain text readers cannot handle mixed binary formats at scale or in parallel.

Exam trap

The trap here is assuming a generic text reader or external table can parse PDF and DOCX, when binary formats require the binaryFile data source plus explicit extraction logic.

33
MCQhard

A team is preparing a Delta table of product descriptions for a RAG application. The table receives continuous upserts from a streaming pipeline, and the embedding job reads the table every hour. Engineers notice the embedding job reprocesses every row on each run even though only a few rows change. Which change should be made to the source table to let the embedding job process only new or updated rows?

A.Partition the Delta table by the product category column so the embedding job reads fewer files.
B.Enable Change Data Feed on the Delta table and have the embedding job read the change feed since the last processed version.
C.Convert the table to a view that filters rows by the current timestamp so only recent rows appear.
D.Run OPTIMIZE with Z-ORDER on the product identifier column to compact small files before each embedding run.
AnswerB

Change Data Feed records row-level inserts, updates, and deletes with commit versions, so the embedding job can read only changes since its last checkpoint. This avoids reprocessing unchanged rows on every run. It is the intended mechanism for incremental downstream consumption of a Delta table that receives continuous upserts.

Why this answer

Change Data Feed captures row-level changes with commit versions, allowing a downstream job to read only the inserts, updates, and deletes that occurred since its last checkpoint. This directly solves the problem of reprocessing unchanged rows. Partitioning, compaction, and timestamp filtering improve read characteristics but do not expose which rows changed, so they cannot deliver incremental processing.

Exam trap

The trap here is confusing read-performance optimizations such as partitioning or OPTIMIZE with change tracking, when only Change Data Feed exposes which rows were inserted, updated, or deleted.

34
MCQmedium

An engineer is cleaning a Delta table of customer support transcripts stored in Unity Catalog. The table has a nested array column named messages, where each element has role and content fields. They need to keep only rows where at least one message has role equal to 'user' and content longer than 20 characters. Which transformation correctly expresses this filter?

A.Cast the messages array to a string and apply a LIKE pattern matching the literal text 'user' followed by any content.
B.Use the higher-order function exists on the messages array with a lambda testing role = 'user' AND length(content) > 20.
C.Use the transform higher-order function on messages and then check whether the resulting array is non-empty.
D.Explode the messages array, filter on role and content length, then group by the row identifier to rebuild the rows.
AnswerB

The exists higher-order function returns true when any array element satisfies the lambda predicate, which matches the requirement of at least one qualifying message. Applying it inside a filter keeps rows where the condition holds, and lambda access to role and content works directly on the nested struct elements without exploding the array.

Why this answer

Filtering rows based on whether any nested element meets a predicate is exactly what the exists higher-order function provides, and its lambda can test both role and content length in one pass. Exploding reshapes and shuffles data, string matching is ambiguous, and transform preserves array length so it cannot signal absence of a match.

Exam trap

The trap here is reaching for explode-and-regroup out of habit, when a higher-order array function expresses the same predicate without changing the table shape.

35
MCQmedium

When preparing a dataset for fine-tuning an LLM, you need to ensure the data is representative of the target domain. What is the most effective approach to detect and mitigate sampling bias in your training set using Databricks?

A.Increase the batch size during the fine-tuning process.
B.Perform statistical profiling on feature distributions and re-sample if necessary.
C.Change the model architecture to a larger parameter count.
D.Use a random seed to shuffle the data before splitting.
AnswerB

Statistical profiling reveals imbalances, allowing engineers to apply techniques like oversampling or undersampling to correct them. This quantitative approach is objective and standard for ensuring high-quality machine learning training sets. It prevents the model from favoring specific patterns simply because they were more frequent in the dataset.

Why this answer

Using exploratory data analysis (EDA) with Databricks SQL or Spark to analyze distribution statistics is the most effective way to identify bias. By comparing the distribution of the training set to a representative sample of real-world production data, engineers can identify under-represented categories. This ensures that the fine-tuned model performs reliably across all expected input scenarios, preventing the model from developing blind spots.

Exam trap

Candidates tend to focus exclusively on model hyperparameter tuning while ignoring raw dataset distributions, missing structural imbalances present in the training inputs.

36
Multi-Selecthard

A team is preparing a large text corpus for embedding generation with a foundation model endpoint on Databricks. They must reduce token cost and improve retrieval quality before vectorization. Which two preprocessing steps should be applied to the raw text? (Choose two.)

Select 2 answers
A.Remove stop words and punctuation from every passage before generating embeddings.
B.Deduplicate near-identical passages and drop documents that fail a minimum content-length threshold.
C.Truncate every document to its first 512 characters to guarantee uniform chunk size.
D.Normalize whitespace and strip boilerplate such as navigation menus, headers, and repeated legal footers.
E.Convert all text to uppercase before tokenization to standardize casing.
AnswersB, D

Near-duplicate passages waste embedding calls and crowd the index with redundant vectors, so removing them lowers cost and improves result diversity. Dropping documents below a minimum length eliminates fragments and stubs that produce low-information embeddings, which raises overall retrieval precision without discarding meaningful content.

Why this answer

Removing boilerplate and normalizing whitespace cuts wasted tokens while sharpening the text signal, and deduplicating plus dropping undersized fragments avoids paying to embed redundant or low-information content. Uppercasing, stop-word stripping, and blind truncation either damage semantics or discard content, so they do not improve cost or retrieval quality.

Exam trap

The trap here is assuming aggressive token reduction like stop-word removal or truncation always lowers cost, when it can degrade embedding quality without meaningful savings.

37
MCQeasy

A data engineer needs to prepare a Delta table of customer reviews for embedding generation. The reviews contain HTML tags, inconsistent whitespace, and mixed casing that hurt embedding quality. Which preparation step should be applied before generating embeddings?

A.Apply a hash of the raw review text as the embedding input to guarantee deterministic vectors.
B.Increase the embedding vector dimension to capture the additional characters in the raw text.
C.Store the reviews in a Parquet file instead of Delta to improve text compression before embedding.
D.Normalize the text by stripping HTML tags, collapsing whitespace, and applying consistent casing.
AnswerD

Removing HTML tags, collapsing whitespace, and normalizing casing reduces noise that would otherwise dilute the embedding signal. Embedding models encode the literal tokens they receive, so markup and erratic spacing consume vector capacity without adding semantic value. Cleaning the text first yields more consistent and comparable vectors across the corpus.

Why this answer

Embedding models encode the exact tokens they receive, so HTML markup, irregular whitespace, and inconsistent casing add noise that dilutes semantic signal. Normalizing the text before embedding removes that noise and produces vectors that better reflect meaning. Storage format and vector dimension do not change the input text, so they cannot solve the quality problem.

Exam trap

The trap here is assuming that a storage or model configuration change can compensate for dirty input text, when the embedding model only sees the characters it is given.

38
MCQmedium

A data engineer is assembling a fine-tuning dataset from a Delta table of conversation transcripts. Each transcript contains a list of message objects with a role and a content field, and the training job requires one row per conversation with the messages serialized into the expected format. Which transformation should the engineer apply?

A.Use the explode function to expand each message object into a separate row before writing the dataset.
B.Use the collect_list aggregate to group the message objects back into an array per conversation and serialize that array into the required format.
C.Use the pivot function to turn message roles into columns with content values in the cells.
D.Use the flatten function to collapse the nested message array into a single string per row.
AnswerB

collect_list aggregates the individual message rows back into an ordered array grouped by conversation identifier, which is exactly the nested structure the training format expects. Serializing that array into the required representation produces one training example per conversation, matching the job's input contract.

Why this answer

The training job requires one row per conversation with the message sequence preserved, so the individual message rows must be grouped back together. collect_list aggregates messages into an ordered array per conversation, and serializing that array yields the nested format the job expects. Exploding, pivoting, or flattening would alter the structure and lose the required per-conversation layout.

Exam trap

The trap here is reaching for explode out of habit when flattening data, even though this scenario requires the opposite operation of aggregating messages back into one row per conversation.

39
MCQhard

A data engineer is preparing a large corpus of customer support transcripts stored as Parquet in a Unity Catalog volume. Before generating embeddings, each transcript must be tokenized and truncated to a maximum token length. The engineer wants to use a Databricks-native approach that runs distributed across the cluster and avoids pulling the full corpus into a single node. Which approach best satisfies these requirements?

A.Use spark.read.parquet to load the corpus and apply a pandas_udf that tokenizes and truncates each transcript.
B.Use dbutils.fs.cp to copy the Parquet files to a local path, then run a Python tokenizer loop over each file.
C.Create a Delta table from the Parquet files and query it with a SQL UDF that calls a tokenizer from a Python library.
D.Read the Parquet files with pandas.read_parquet on the driver, tokenize, then write the truncated text back to the volume.
AnswerA

A pandas UDF runs vectorized tokenization inside Spark executors, so the corpus stays distributed and each partition processes its own rows. This avoids collecting the data to the driver and scales with cluster size. Tokenizing and truncating in the same distributed step is exactly the preparation needed before embedding generation.

Why this answer

Distributed tokenization of a large corpus requires the work to run inside Spark executors, which a pandas UDF provides by operating vectorized on each partition. Reading with pandas on the driver or looping over files locally centralizes the workload and breaks at scale. A pandas UDF keeps the corpus distributed while applying the tokenizer and truncation per row.

Exam trap

The trap here is equating native Python tokenization with distributed execution, when only executor-side constructs like pandas UDFs keep the corpus distributed.

40
MCQmedium

Which of the following describes the purpose of using a 'Feature Store' when preparing data for Generative AI applications?

A.To store raw text files for long-term archival.
B.To provide a unified interface for consistent feature computation during training and inference.
C.To serve as a high-performance database for LLM model weights.
D.To replace the need for data cleaning and preprocessing entirely.
AnswerB

The primary goal of a feature store is to eliminate inconsistencies in feature logic. By centralizing this, organizations ensure that the features the model was trained on are computed in exactly the same way when deployed. This is foundational for building stable, repeatable machine learning pipelines in Databricks.

Why this answer

The Feature Store acts as a centralized repository for standardized, reusable features. It ensures that the same logic is used for data preparation during both training and inference. This consistency is critical for preventing training-serving skew, where the model's performance in production differs from its training performance because the input data was transformed differently, ensuring the model remains accurate and reliable in real-world scenarios.

Exam trap

Candidates often mistakenly believe the Feature Store is primarily for model storage, ignoring its primary role in ensuring identical feature computation logic between training and real-time inference pipelines.

Ready to test yourself?

Try a timed practice session using only Data Preparation questions.