Databricks-GenAI-Assoc Data Preparation Practice Question
A data engineer is preparing a dataset of product descriptions for embedding generation. The source table in Unity Catalog contains a column 'description' with mixed languages, and the team wants to filter to English-only text before vectorization. They need a scalable, built-in Databricks function that can detect the language of each description without external API calls. Which function should they use?
⚠ Common exam trap
Watch out — candidates often confuse general-purpose AI functions such as ai_classify with a dedicated language detection function, when only ai_detect_language directly returns a language code for filtering.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ai_detect_language
The built-in ai_detect_language function is purpose-built for language identification at scale within Databricks. It returns a language code that can be used to filter the dataset to English-only records before embedding generation. Other AI functions like ai_analyze_sentiment or ai_classify serve different purposes, and regex cannot reliably determine language.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
regexp_extract
Why it's wrong here
regexp_extract uses regular expressions to extract substrings from text. It cannot reliably detect language because language is not determined by simple character patterns. Attempting to filter English-only text with regex would be brittle and inaccurate, especially for short product descriptions. It does not provide a language label and is not a language detection tool.
- ✓
ai_detect_language
Why this is correct
ai_detect_language is a built-in Databricks SQL function that returns the detected language code for a given text column. It is designed for scalable language identification directly in SQL or PySpark, requiring no external API calls. Filtering on the returned language code allows the team to keep only English descriptions before embedding, exactly matching the requirement.
- ✗
ai_classify
Why it's wrong here
ai_classify assigns custom labels to text based on a provided list of categories. While it could be prompted to classify language, it is not a dedicated language detection function and would require manual label definitions and potentially produce inconsistent results. It is not the built-in, efficient solution for filtering by language in a scalable data preparation pipeline.
- ✗
ai_analyze_sentiment
Why it's wrong here
ai_analyze_sentiment returns sentiment scores (positive, negative, neutral) for text. It does not detect the language of the text. Using it to filter English-only records would not work because sentiment is independent of language, and the function does not return a language label. This would incorrectly include non-English descriptions in the embedding dataset.
About these practice questions
Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.