Courseiva

Databricks-DE-Assoc Data Ingestion and Loading Practice Question

A data engineer is using Auto Loader and wants to handle a situation where a column 'user_id' is sometimes an integer and sometimes a string in the source JSON files. Which TWO strategies can be used to manage this schema conflict?

⚠ Common exam trap

Candidates often suggest changing the source file structure as the primary fix, overlooking that Auto Loader provides built-in mechanisms like schema hints and rescued data columns to handle type mismatches gracefully.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Providing a schema hint to force 'user_id' to be treated as a string.

Handling type mismatches is a common challenge in data ingestion. Auto Loader provides schema hints to force a specific type and the rescued data column to capture records that fail to meet that type. These tools allow the engineer to maintain data integrity and prevent the ingestion pipeline from failing due to inconsistent source data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Providing a schema hint to force 'user_id' to be treated as a string.

    Why this is correct

    By using a schema hint, the engineer can tell Auto Loader to treat the 'user_id' column as a string regardless of its format in individual files. Since integers can be safely cast to strings, this approach ensures that all records are ingested successfully into a single, consistently typed column.

  • ✗

    Using the 'mergeSchema' option to create two separate columns for the types.

    Why it's wrong here

    The mergeSchema option does not create separate columns based on data types for the same field name. Delta Lake and Spark do not support multiple types for a single column name. Attempting to merge conflicting types would typically result in an error or require manual intervention to resolve the conflict.

  • ✓

    Enabling the rescued data column to capture the conflicting records.

    Why this is correct

    If a schema is defined or hinted, and a record arrives with a type that cannot be safely cast, Auto Loader will put the original record into the _rescued_data column. This allows the pipeline to continue running while the engineer can later analyze the rescued data to fix the source issue.

  • ✗

    Setting 'cloudFiles.inferColumnTypes' to false to ignore all types.

    Why it's wrong here

    Setting inference to false would result in all columns being treated as strings by default if no schema is provided. While this might avoid the conflict, it doesn't provide a way to manage the 'user_id' specifically and would lose the benefits of having proper data types for other columns.

  • ✗

    Deleting the checkpoint and restarting the stream to re-infer the types.

    Why it's wrong here

    Deleting the checkpoint would cause the stream to lose its progress and potentially re-process all data, but it would not solve the underlying conflict of inconsistent types within the source files. The inference process would still encounter both integers and strings, leading to the same schema conflict again.

About these practice questions

One of 276 original Databricks-DE-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.