Databricks-DA-Assoc Importing Data Practice Question
A data analyst is using the Databricks SQL Connector for Python to query a large Delta table. The query returns millions of rows, and the analyst wants to process the results in batches to avoid memory issues. Which method should the analyst use to fetch rows in chunks?
⚠ Common exam trap
The trap here is assuming that fetchall() is acceptable for large datasets or that fetchone() is efficient, when actually fetchmany() is designed for batch processing.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
cursor.fetchmany(size)
The fetchmany(size) method allows fetching a specified number of rows at a time, enabling batch processing of large result sets. This is the standard way to avoid loading all rows into memory. Other methods either fetch all rows, one row, or require manual query modifications.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
cursor.fetchone()
Why it's wrong here
fetchone() returns a single row at a time. While this avoids memory issues, it is inefficient for processing millions of rows due to the overhead of many round trips. The analyst would need to loop millions of times, which is slow. Batch fetching is more efficient for large datasets.
- ✗
cursor.fetchall()
Why it's wrong here
fetchall() retrieves all rows at once, which can exhaust memory when dealing with millions of rows. It does not support batching and is unsuitable for large result sets. The analyst needs a method that allows incremental retrieval to manage memory efficiently.
- ✓
cursor.fetchmany(size)
Why this is correct
fetchmany(size) returns the next set of rows up to the specified size, enabling batch processing. This method is ideal for large result sets because it allows the analyst to iterate over chunks, reducing memory overhead. It is part of the Python DB API 2.0 specification and supported by the Databricks SQL Connector.
- ✗
cursor.execute() with a LIMIT clause
Why it's wrong here
Using execute() with LIMIT would require multiple queries with different offsets to simulate batching, which is inefficient and can lead to inconsistent results if data changes. It is not a built-in batching mechanism and adds complexity. The connector provides fetchmany for this purpose.
About these practice questions
One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.