Databricks-DE-Pro Data Ingestion and Acquisition Practice Question
You are ingesting data from multiple source systems with varying file formats (JSON, CSV, Parquet) into a centralized Bronze landing zone. Which architecture pattern is the most scalable for maintaining this ingestion layer?
⚠ Common exam trap
Candidates often suggest a 'monolithic' pipeline to handle all formats, failing to recognize that isolating source-specific logic is necessary for scalability and maintenance.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A separate ingestion pipeline or notebook for each source system.
A modular ingestion architecture that decouples source-specific logic from the common landing layer is the industry standard. By using parameterized notebooks or DLT pipelines for each source, you isolate potential failures, enable independent scaling, and simplify the management of schema mappings. This pattern ensures that changes in one source system do not impact the ingestion pipelines of others, creating a highly resilient and maintainable data acquisition ecosystem.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A single monolithic notebook with a massive if/else block for each source.
Why it's wrong here
Monolithic notebooks are difficult to test, debug, and maintain. As the number of sources increases, the code becomes unmanageable and prone to errors. If one part of the pipeline fails, the entire job might crash, making it an anti-pattern for scalable data engineering in production environments.
- ✓
A separate ingestion pipeline or notebook for each source system.
Why this is correct
Isolating each source in its own pipeline allows for independent configuration, scheduling, and error handling. This modularity is essential for scalability, as it minimizes the blast radius of any individual pipeline failure and allows teams to manage and optimize ingestion logic based on the specific requirements of each source system.
- ✗
Ingest everything into a single raw table before applying transformations.
Why it's wrong here
While the Bronze layer often contains raw data, throwing everything into one table without regard for source-specific schemas or ingestion logic creates a data swamp. It becomes nearly impossible to track lineage, manage schema drift, or enforce quality, leading to significant challenges in downstream Silver and Gold layer processing.
- ✗
Use a third-party tool exclusively and bypass Databricks for ingestion.
Why it's wrong here
While third-party tools (like Fivetran) are useful, they may not always meet the complex transformation requirements or cost targets of an enterprise. A Databricks-centric approach using Auto Loader is generally preferred for its deep integration, performance, and ability to handle both simple and complex ingestion scenarios within the same environment.
About these practice questions
One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.