Be able to write correct PySpark and Spark SQL for reading cloud files, transforming data, and writing Delta tables, plus configure cluster and shuffle settings. The single most important thing is choosing the right API or config for the stated symptom, especially Delta Live Tables expectation behavior.
Start practicing
Developing Code (Python/SQL) — choose a session length
Free · No account required
Domain overview
This domain covers writing PySpark and Spark SQL on Databricks: reading cloud files with the DataFrameReader, schema inference and explicit schemas, Delta Lake reads/writes, Delta Live Tables expectations, and tuning shuffle, broadcast, and memory behavior on clusters. Questions present realistic code or job symptoms and ask you to pick the correct API, config, or DLT expectation behavior.
Exam objectives
Reading CSV/Parquet/JSON from cloud storage with spark.read and inferSchema or explicit StructType schemas
Tuning shuffle and memory with spark.sql.shuffle.partitions, spark.sql.autoBroadcastJoinThreshold, and executor memory settings
Writing and merging Delta tables with DeltaTable, MERGE, and schema evolution options
Defining Delta Live Tables pipelines in Python with @dlt.table and @dlt.expect_or_drop quality expectations
Assuming inferSchema is free: it triggers an extra pass over the file, so production code should supply an explicit schema instead
Raising executor memory first for OOM during shuffle when the real fix is increasing shuffle partitions or reducing broadcast size
Confusing DLT expectation decorators: expect_or_drop drops failing rows while expect_or_fail stops the pipeline, so using the wrong one breaks requirements
Click any question to see the full explanation and answer options, or start a focused practice session above.
You are optimizing a PySpark job that reads from a Delta table. You notice skewed data distribution on the 'customer_id' column, causing Task-level stragglers. Which transformation should you apply to the DataFrame to mitigate this skew during a join operation?
2You are processing sensitive PII data in a Delta table. You need to ensure that specific columns containing PII are not readable by general data analysts while maintaining the ability to perform aggregate analysis on those rows. Which feature should you implement?
3A data engineer is tuning a Spark job and decides to use 'Z-Ordering' on a Delta table. Which THREE of the following are valid considerations when selecting columns for Z-Ordering?
4You are migrating a legacy ETL process to Delta Live Tables (DLT). You have an existing table defined with a complex transformation that involves a custom Python function using a third-party library. How should you structure this in DLT to ensure the function is available and correctly applied?
5When designing a streaming pipeline using Structured Streaming, which THREE of the following are necessary to ensure 'exactly-once' processing semantics in Databricks?
6You have a large Spark DataFrame that you need to filter and save as multiple smaller Parquet files based on the values in a 'region' column. Which method should you use to optimize the file layout for subsequent queries?
7Your team is using a shared cluster for development. A user reports that their job is slow because the cluster memory is frequently filled by large data broadcasts. What configuration adjustment should you make to prevent this issue across all jobs on the cluster?
8Refer to the exhibit. You are appending data to an existing Delta table. What is the most likely cause of this error, and how should you resolve it?
9Which THREE of the following are benefits of using Delta Lake over standard Parquet files for your data lake storage?
10Which command should be used to display the history of transactions performed on a Delta table, including operations like overwrites and updates?
11Which THREE of the following are benefits of using Delta Lake over standard Parquet files in Databricks?
12You are writing a Spark application that uses a Broadcast Hash Join. You want to force the join to use broadcast optimization for a specific table. How do you implement this in PySpark?
13Which of the following describes the behavior of the 'OPTIMIZE' command in Databricks?
14When partitioning a large Delta table, what is the best practice regarding the number of unique values in the partition column?
15When running a PySpark job, you receive an 'Out of Memory (OOM)' error during a shuffle operation. Which configuration is the most appropriate to address this first?
16Which TWO of the following are valid ways to trigger a job in Databricks?
17A Data Engineer is developing a Delta Live Tables (DLT) pipeline using Python. They need to ensure that records failing a specific data quality check are dropped, but the pipeline continues to process the remaining valid records. Which expectation syntax should the engineer implement?
18Which TWO of the following are valid ways to pass configuration parameters to a Databricks Job task? (Select TWO)
19When using the Unity Catalog, how should an engineer properly reference a table named 'sales' located in the 'finance' schema within the 'prod_catalog' catalog using Spark SQL?
20Which THREE features are provided by Delta Lake when compared to standard Parquet files? (Select THREE)
21A Data Engineer needs to perform an 'upsert' operation on a Delta table. Which command provides the most efficient way to merge new data with existing records in a single transactional step?
22A Data Engineer is using Databricks Asset Bundles (DABs) to manage a project. Where should the engineer define the job settings, such as clusters, schedules, and task dependencies?
23Which property must be set to ensure a Spark Structured Streaming query can handle changes to the source data schema, such as adding a new column?
24An engineer is writing a Python function to process data in a Databricks Notebook. Which command should they use to ensure that secrets, such as API keys, are not hardcoded or exposed in the plain text of the notebook?
25Which TWO of the following are benefits of using the Databricks Delta Lake 'Optimize' command? (Select TWO)
26You are writing a PySpark script to join two large tables. You want to ensure the join operation is optimized for performance by broadcasting the smaller table. Which configuration property should you adjust, or code construct should you use, to force this behavior?
27You are developing a Delta Live Tables (DLT) pipeline and need to ensure high data quality. Which TWO of the following statements correctly describe how Expectations work within DLT?
28You are using the Databricks CLI to automate workspace tasks. Which THREE of the following statements correctly identify capabilities of the Databricks CLI?
29You are debugging a PySpark job that is experiencing severe memory pressure during a join on a massive column. You suspect data skew. Which approach is best to mitigate this issue?
30Which of the following describes the correct use of a UDF (User Defined Function) in PySpark for production pipelines?
31When designing a production-grade data pipeline in Databricks, what is the recommended approach for managing secrets such as database credentials?
32A data engineer is building a Structured Streaming job that reads from a Kafka topic and writes to a Delta table. The pipeline must tolerate late-arriving data up to 10 minutes and update aggregations accordingly. The engineer wants to use a watermark on the event-time column. Which code snippet correctly applies the watermark and performs a 5-minute tumbling window aggregation?
33A data engineer is developing a PySpark job that reads a large Delta table, performs a groupBy on a high-cardinality column, and writes the result to another Delta table. The job is experiencing performance issues due to data skew. The engineer wants to optimize the shuffle by using salting. Which approach correctly implements salting to distribute the skewed keys evenly?
34A data engineer is using Delta Live Tables (DLT) to build a pipeline that ingests data from a streaming source. The pipeline must ensure that the target table is updated incrementally and that data quality constraints are enforced. The engineer wants to use expectations to drop invalid records while maintaining pipeline performance. Which TWO of the following are true regarding DLT expectations and their behavior? (Choose two.)
35A data engineer is optimizing a PySpark job that processes a large DataFrame and writes the result to a Delta table. The job currently uses repartition(100) before writing, but the output consists of many small files. The engineer wants to reduce the number of output files without shuffling the entire dataset again. Which approach should be used?
36A Data Engineer is building a Structured Streaming pipeline that reads from a Kafka topic and writes to a Delta table. The pipeline must handle late-arriving data up to 2 hours and produce correct aggregations per 10-minute window. The engineer wants the streaming query to automatically clean up old state so the job does not accumulate unbounded state. Which combination of Structured Streaming features should be used?
37A data engineer is implementing a Structured Streaming job that reads from a Kafka topic and writes to a Delta table. The engineer needs to ensure that the job can recover from failures without data loss or duplication. The job uses foreachBatch to perform upserts into the Delta table. Which checkpointing configuration is required to achieve exactly-once semantics?
38A data engineer is using Delta Live Tables (DLT) to create a pipeline that processes streaming data. They need to ensure that the pipeline only processes new data since the last run and that the pipeline can recover from failures without reprocessing all data. Which combination of features should they use to achieve this?
39A data engineer needs to read a CSV file from cloud storage into a Spark DataFrame in Databricks. The file has a header row and uses commas as delimiters. The engineer wants to infer the schema automatically. Which code snippet correctly reads the file?
40A data engineer is building a Structured Streaming pipeline that reads from a Kafka topic and writes to a Delta table. The pipeline must handle late-arriving data and update already processed aggregates. Which watermark strategy should be used to allow updates to aggregates while bounding state store growth?
41A data engineer is using PySpark to process a large DataFrame and needs to reduce the number of partitions before writing to a Delta table to avoid creating too many small files. The DataFrame currently has 2000 partitions, each about 10 MB. The engineer wants to reduce the number of partitions to approximately 200 while minimizing data shuffling. Which approach is most appropriate?
42A data engineer is developing a PySpark job that reads from a Delta table and performs a series of transformations. The engineer notices that the job is slow and suspects that the query plan is not optimized because statistics are outdated. Which command should the engineer run to update the statistics for the Delta table to improve query performance?
Be able to write correct PySpark and Spark SQL for reading cloud files, transforming data, and writing Delta tables, plus configure cluster and shuffle settings. The single most important thing is choosing the right API or config for the stated symptom, especially Delta Live Tables expectation behavior.
The Courseiva Databricks-DE-Pro question bank contains 42 questions in the Developing Code (Python/SQL) domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Developing Code (Python/SQL) domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included