Courseiva
Data Preparation →mediumMultiple Choice

Databricks-GenAI-Assoc Data Preparation Practice Question

You need to ingest data from an external JSON source into a Delta table. The source schema is inconsistent. Which strategy is most effective for preparing this data?

⚠ Common exam trap

Candidates often suggest cleaning data directly into a final table, overlooking the importance of the Bronze layer for preserving raw historical data before applying schema enforcement in Silver.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implement a bronze-silver medallion pattern.

The 'bronze-to-silver' pattern is the industry standard in Databricks for handling inconsistent data. By loading raw JSON into a 'Bronze' table with a schema of type 'string' (or a single JSON column), you preserve all information for audit. You then use 'Silver' tables to perform schema enforcement, data cleaning, and type casting. This separation of concerns allows for robust error handling and iterative refinement of the data cleaning logic.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use inferSchema = true in the read configuration.

    Why it's wrong here

    Inferring schema on inconsistent data is dangerous because Spark may guess types incorrectly, leading to data loss or conversion errors. Furthermore, it requires a full scan of the dataset, which is inefficient. Manual schema definition or the bronze-silver pattern is significantly more reliable and performant.

  • ✗

    Define a rigid schema at the ingestion point.

    Why it's wrong here

    Defining a rigid schema at ingestion will cause the pipeline to fail immediately when encountering inconsistent records. Inconsistent data sources require a more flexible ingestion layer that allows for late-binding schema decisions, which is best achieved through the bronze-silver medallion architecture pattern.

  • ✓

    Implement a bronze-silver medallion pattern.

    Why this is correct

    The bronze-silver pattern enables ingestion of raw, unstructured data into a staging area, followed by rigorous cleaning and validation in a downstream silver table. This approach prevents pipeline failures during ingestion and ensures that cleaning logic is decoupled from the data acquisition layer, allowing for better maintainability.

  • ✗

    Convert the JSON to CSV before loading into Delta.

    Why it's wrong here

    Converting to CSV introduces extra overhead and potential formatting issues, such as delimiter conflicts. JSON is natively supported in Spark and contains structural information that is lost during CSV conversion. Keeping data in its native format until after the initial ingestion is the superior architectural choice.

About these practice questions

One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.