Courseiva

Databricks-DE-Pro · topic practice

Developing Code (Python/SQL) practice questions

This domain covers writing PySpark and Spark SQL on Databricks: reading cloud files with the DataFrameReader, schema inference and explicit schemas, Delta Lake reads/writes, Delta Live Tables expectations, and tuning shuffle, broadcast, and memory behavior on clusters. Questions present realistic code or job symptoms and ask you to pick the correct API, config, or DLT expectation behavior.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Developing Code (Python/SQL)

What the exam tests

What to know about Developing Code (Python/SQL)

Be able to write correct PySpark and Spark SQL for reading cloud files, transforming data, and writing Delta tables, plus configure cluster and shuffle settings. The single most important thing is choosing the right API or config for the stated symptom, especially Delta Live Tables expectation behavior.

Reading CSV/Parquet/JSON from cloud storage with spark.read and inferSchema or explicit StructType schemas

Tuning shuffle and memory with spark.sql.shuffle.partitions, spark.sql.autoBroadcastJoinThreshold, and executor memory settings

Writing and merging Delta tables with DeltaTable, MERGE, and schema evolution options

Defining Delta Live Tables pipelines in Python with @dlt.table and @dlt.expect_or_drop quality expectations

Watch out for

Common Developing Code (Python/SQL) exam traps

  • ▸Assuming inferSchema is free: it triggers an extra pass over the file, so production code should supply an explicit schema instead
  • ▸Raising executor memory first for OOM during shuffle when the real fix is increasing shuffle partitions or reducing broadcast size
  • ▸Confusing DLT expectation decorators: expect_or_drop drops failing rows while expect_or_fail stops the pipeline, so using the wrong one breaks requirements

Practice set

Developing Code (Python/SQL) questions

20 questions · select your answer, then reveal the explanation

Question 1mediummultiple choice
Study the full Python automation breakdown →

A data engineer is implementing a Delta Lake table with streaming ingestion. To ensure high-concurrency writes while maintaining data integrity, what configuration must be enabled for the table to avoid 'Optimistic Concurrency Control' conflicts?

Which TWO of the following statements accurately describe the behavior of the 'MERGE' operation in Delta Lake regarding schema evolution and concurrency?

Question 3mediummultiple choice
Study the full Python automation breakdown →

Refer to the exhibit. A data engineer encounters this error while running a SQL query on a Delta table. Based on the error message, what is the most likely cause?

Exhibit

Error: AnalysisException: [UNRESOLVED_COLUMN.WITH_SUGGESTION] A column or function parameter with name 'user_id' cannot be resolved. Did you mean one of the following? ['userId', 'user_id_raw']
Question 4mediummultiple choice
Study the full Python automation breakdown →

You are performing a 'Vacuum' operation on a large Delta table to remove expired files. You notice that the operation is running for an extended period and consuming significant cluster resources. What is the most likely reason for this performance overhead?

You have a large Delta table and you need to perform an update on a single row based on a complex condition. You have noticed that 'UPDATE' statements are slow. What is the most effective way to optimize this operation in a production pipeline?

Question 6mediummultiple choice
Study the full Python automation breakdown →

A Data Engineer is developing a Delta Live Tables (DLT) pipeline. They need to ensure that rows containing null values in the 'customer_id' column are dropped during the ingestion process. Which constraint syntax is correct?

A Data Engineer is optimizing a PySpark application. They notice significant data skew on the 'region_id' column during a join operation. Which TWO techniques can the engineer use to mitigate this skew?

An engineer wants to capture data lineage and metrics for a custom Python transformation function that processes a Spark DataFrame outside of DLT. What is the most effective way to track this?

Question 9mediummultiple choice
Study the full Python automation breakdown →

When using Auto Loader to ingest files from cloud storage, what is the primary purpose of the 'cloudFiles.schemaLocation' option?

A Data Engineer needs to perform a full refresh of a Delta table while maintaining the table's existing permissions and history. Which approach is best?

Question 11mediummultiple choice
Study the full Python automation breakdown →

An engineer needs to ensure that a Delta table is only updated if a record with the same ID does not exist in the target table. Which MERGE clause should be used?

Question 12mediummultiple choice
Study the full Python automation breakdown →

Refer to the exhibit. An engineer is trying to run a VACUUM command on a Delta table but receives an error stating the retention threshold is too low. What is the implication of setting the configuration shown in the exhibit?

Exhibit

{
  "cluster_id": "0521-123456-abcde789",
  "spark_version": "13.3.x-scala2.12",
  "spark_conf": {
    "spark.databricks.delta.retentionDurationCheck.enabled": "false"
  }
}
Question 13mediummultiple choice
Study the full Python automation breakdown →

You have a Delta table that is frequently queried by ID. To optimize performance, you decide to implement Z-Ordering. Which column is the best candidate for Z-Ordering?

Question 14mediummultiple choice
Study the full Python automation breakdown →

A data engineer is writing a PySpark job that reads from a Delta table, applies several transformations, and writes the result to another Delta table. The engineer notices that the job occasionally fails with a FileNotFoundException during the write phase. The source table is frequently updated by another job. Which action should the engineer take to resolve this issue?

A data engineer is writing a PySpark script in a Databricks notebook to process a DataFrame. They need to cache the DataFrame to avoid recomputation across multiple actions. Which method should they use to cache the DataFrame in memory and on disk?

A data engineer is working with a Delta table that has a column 'id' as the primary key. The table is updated frequently with new records, and the engineer needs to ensure that only the latest record for each 'id' is kept. The engineer decides to use MERGE to upsert data from a staging table. Which MERGE condition will correctly update existing records and insert new ones?

A Data Engineer needs to implement a PySpark transformation that processes a large DataFrame partitioned by date. The transformation must apply a complex Python function to each row, but the function requires loading a large lookup table (about 2 GB) that should be reused across all executors. The engineer wants to minimize serialization overhead and avoid reloading the lookup table for every row. Which approach is most appropriate?

A data engineer needs to create a Delta table that enforces a constraint to ensure that the 'email' column contains a valid email format. The engineer wants to reject any rows that do not match the pattern. Which SQL statement should be used to add this constraint?

Question 19mediummultiple choice
Study the full Python automation breakdown →

A Data Engineer is developing a PySpark job that reads from a Delta table, performs several transformations, and writes the result to another Delta table. The engineer notices that the job is producing many small files in the target table, which degrades read performance. The engineer wants to reduce the number of files written without changing the partitioning of the target table. Which optimization technique should be applied?

Question 20mediummultiple choice
Study the full Python automation breakdown →

A data engineer is using Databricks Repos to manage a Python project. The project contains multiple notebooks and a Python module in a separate file. The engineer wants to import the module into a notebook. Which approach should be used to ensure the module is importable within the notebook?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Developing Code (Python/SQL) sessions

Start a Developing Code (Python/SQL) only practice session

Every question in these sessions is drawn from the Developing Code (Python/SQL) domain — nothing else.

Related practice questions

Related Databricks-DE-Pro topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-DE-Pro exam test about Developing Code (Python/SQL)?
Be able to write correct PySpark and Spark SQL for reading cloud files, transforming data, and writing Delta tables, plus configure cluster and shuffle settings. The single most important thing is choosing the right API or config for the stated symptom, especially Delta Live Tables expectation behavior.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Developing Code (Python/SQL) questions in a focused session?
Yes — the session launcher on this page draws every question from the Developing Code (Python/SQL) domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-DE-Pro topics?
Use the topic links above to move to related areas, or go back to the Databricks-DE-Pro question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-DE-Pro exam covers. They are not copied from any real exam or dump site.