Courseiva
Data Preparation →hardMultiple Choice

Databricks-GenAI-Assoc Data Preparation Practice Question

A GenAI engineer prepares a Delta table of product descriptions that will feed a chunking and embedding pipeline. The descriptions are written in mixed languages, and the embedding model supports only English. The engineer needs to ensure that non-English rows are detected and routed for translation before embedding. Which approach is most appropriate?

⚠ Common exam trap

The trap here is assuming that a table-level tag or a character-range filter can substitute for per-row language detection.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Add a language detection UDF that runs ai_classify or a language-identification library on each description, write the detected language to a column, and split the table into English and non-English subsets.

When the embedding model is English-only, the pipeline must identify which rows need translation. Detecting language per row and storing the result as a column gives a deterministic routing signal, so English rows embed directly while non-English rows are sent to translation. This keeps the preparation step auditable and avoids both data loss and unnecessary translation cost.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Translate every description into English before chunking, regardless of its original language, to guarantee uniform embeddings.

    Why it's wrong here

    Translating all rows adds cost and latency and can degrade English descriptions through unnecessary round trips that introduce translation artifacts. It also discards the original text, which may be needed for display or compliance. The requirement is to detect and route non-English rows, not to overwrite the entire corpus, so this over-processes the data.

  • ✓

    Add a language detection UDF that runs ai_classify or a language-identification library on each description, write the detected language to a column, and split the table into English and non-English subsets.

    Why this is correct

    Detecting language per row and persisting the result creates an auditable routing column. Splitting the table lets English rows proceed directly to chunking and embedding while non-English rows go to a translation step. This is deterministic, testable data preparation and it keeps the pipeline extensible if more languages are added later.

  • ✗

    Drop all rows where the description contains non-ASCII characters, since those rows are likely non-English.

    Why it's wrong here

    Non-ASCII characters appear in English text as accented brand names, currency symbols, and emoji, so this filter removes valid English rows. It also fails to catch non-English text written entirely in ASCII, such as romanized content. Dropping data by character range is neither accurate language detection nor an acceptable way to preserve corpus coverage.

  • ✗

    Configure the embedding model endpoint to accept a language parameter and pass the detected locale from Unity Catalog tags on the table.

    Why it's wrong here

    Embedding endpoints do not read Unity Catalog table tags, and a table-level tag cannot describe mixed-language rows within the same table. There is no language parameter that makes an English-only model embed non-English text meaningfully. This approach misuses governance metadata as a runtime signal and would not route any rows for translation.

About these practice questions

Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.