Databricks-DE-Pro Data Ingestion and Acquisition Practice Question
A data engineer is ingesting data from an Azure SQL Database into a Delta Lake table using the JDBC connector in a Databricks notebook. The source table contains millions of rows, and the engineer wants to optimize the ingestion by reading the data in parallel. The source table has a numeric primary key column named 'id' that is evenly distributed. Which approach should the engineer use to achieve parallel reads?
⚠ Common exam trap
The trap here is thinking that increasing fetchsize or enabling adaptive query execution will parallelize the JDBC read, when in fact only explicit partitioning options create multiple concurrent tasks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Specify the 'partitionColumn', 'lowerBound', 'upperBound', and 'numPartitions' options in the JDBC read configuration.
To read a large JDBC table in parallel, you must configure the JDBC connector with partitioning options: partitionColumn, lowerBound, upperBound, and numPartitions. The partitionColumn must be numeric, and the bounds define the range for partitioning. Spark then creates multiple tasks to read the data concurrently. This is the standard method for parallelizing JDBC reads in Databricks.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set the 'fetchsize' option to a large value to increase the number of rows retrieved per round trip.
Why it's wrong here
The fetchsize option controls how many rows are fetched per database round trip, which can improve performance for single-threaded reads, but it does not enable parallel reads. It reduces network overhead but does not create multiple concurrent connections. For parallelism, you need to partition the query. This option alone would not achieve the desired parallel ingestion.
- ✗
Use the 'query' option with a custom SQL statement that includes a 'WHERE' clause with modulo arithmetic to split the data.
Why it's wrong here
While you could manually split the query using modulo arithmetic and run multiple reads, this is cumbersome and not the standard approach. The JDBC connector's built-in partitioning options are designed for this purpose and handle the splitting automatically. Using a custom query with modulo would require manual orchestration and may not be as efficient. It is not the recommended method for parallel reads.
- ✓
Specify the 'partitionColumn', 'lowerBound', 'upperBound', and 'numPartitions' options in the JDBC read configuration.
Why this is correct
The JDBC connector supports parallel reads by partitioning the data based on a numeric column. By specifying partitionColumn (e.g., 'id'), along with lowerBound, upperBound, and numPartitions, Spark divides the query into multiple partitions that can be read concurrently. This significantly speeds up ingestion for large tables. The column must be numeric and evenly distributed, which 'id' satisfies. This is the correct approach for parallel JDBC reads.
- ✗
Enable 'spark.sql.adaptive.enabled' to automatically parallelize the JDBC read.
Why it's wrong here
Adaptive query execution (AQE) optimizes query plans at runtime, but it does not automatically parallelize JDBC reads. JDBC reads are executed as a single task unless partitioning is specified. AQE can improve join and shuffle performance, but it does not change the parallelism of the initial read from an external database. This option does not address the need for parallel ingestion.
About these practice questions
Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.