Courseiva
Data Ingestion and Loading →mediumMultiple Choice

Databricks-DE-Assoc Data Ingestion and Loading Practice Question

While ingesting CSV files using Auto Loader, a data engineer notices that some records have malformed data that does not match the inferred schema. How can the engineer capture these records without failing the entire ingestion stream?

⚠ Common exam trap

Candidates often suggest dropping rows or using standard dropMalformed modes which discard data, rather than utilizing the dedicated rescued data column.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

By enabling the '_rescued_data' column in the Auto Loader configuration.

The rescued data column is a powerful feature in Auto Loader that prevents data loss during ingestion. By capturing malformed or unexpected data in a dedicated column, engineers can ensure that the main pipeline continues to run while providing a way to audit and correct data quality issues after the fact.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    By setting the 'mode' option to 'FAILFAST' in the reader.

    Why it's wrong here

    The FAILFAST mode is designed to stop the entire ingestion process immediately upon encountering a single malformed record. This is the opposite of what is required in this scenario, as the goal is to keep the pipeline running and capture the problematic records for later analysis and remediation.

  • ✓

    By enabling the '_rescued_data' column in the Auto Loader configuration.

    Why this is correct

    When the rescued data column is enabled, Auto Loader automatically places any data that cannot be parsed into the expected schema into a special JSON column. This allows the rest of the record's valid fields to be processed normally while preserving the malformed content for future debugging.

  • ✗

    By using a TRY_CAST function in a transformation after the read.

    Why it's wrong here

    While TRY_CAST can prevent errors during data type conversion, it cannot recover data that was already corrupted or dropped during the initial ingestion phase. The rescued data column is unique because it captures the raw source information before it is discarded by the Spark reader's schema enforcement logic.

  • ✗

    By increasing the 'maxFilesPerTrigger' to handle more errors.

    Why it's wrong here

    The maxFilesPerTrigger option is used to control the rate of data ingestion and the size of each micro-batch. It has no effect on how individual malformed records are handled or whether they are captured for auditing purposes, making it an irrelevant setting for solving data quality or parsing issues.

About these practice questions

Courseiva writes every Databricks-DE-Assoc question from scratch — 276 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.