Courseiva
Data Ingestion and Loading →mediumMultiple Choice

Databricks-DE-Assoc Data Ingestion and Loading Practice Question

A data engineer needs to ingest data from a legacy SQL Server database into a Delta Lake bronze table. The ingestion must be performant and support parallel reads from the source table. What is the best practice for configuring the JDBC connection in this scenario?

⚠ Common exam trap

Candidates often forget the required set of four parameters ('partitionColumn', 'lowerBound', 'upperBound', 'numPartitions') and mistakenly think a single parameter is enough to enable parallel reads.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure 'partitionColumn', 'lowerBound', 'upperBound', and 'numPartitions'.

Parallelizing JDBC reads is essential for moving large datasets from relational databases to Databricks. By providing partitioning columns and boundaries, Spark can spawn multiple executors to read different segments of the data simultaneously. This significantly reduces the time required for the initial load and improves the overall throughput of the ingestion pipeline.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a single executor to ensure data consistency during the read.

    Why it's wrong here

    Using a single executor for a large JDBC read creates a significant bottleneck, as all data must pass through a single network connection. This approach does not scale and fails to take advantage of the distributed nature of the Databricks platform, leading to very long ingestion times for large tables.

  • ✓

    Configure 'partitionColumn', 'lowerBound', 'upperBound', and 'numPartitions'.

    Why this is correct

    These four parameters allow Spark to split the JDBC query into multiple smaller queries that can be executed in parallel. By defining the range and number of partitions, the data engineer enables multiple workers to fetch data at the same time, which is critical for high-performance data ingestion from external databases.

  • ✗

    Rely on the default JDBC settings for automatic parallelization.

    Why it's wrong here

    By default, a JDBC connection in Spark uses only one partition and one task to read the entire dataset. There is no automatic parallelization based on the source table's structure. Therefore, the engineer must explicitly provide the partitioning parameters to enable parallel reads and avoid a single-threaded ingestion process.

  • ✗

    Use the COPY INTO command to read directly from the JDBC source.

    Why it's wrong here

    The COPY INTO command is designed specifically for ingesting data from cloud object storage locations like S3 or ADLS. It does not support direct ingestion from JDBC sources. To read from a relational database, the engineer must use the Spark JDBC reader or a specialized connector provided by Databricks.

About these practice questions

This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.