Databricks-DE-Pro · domain
Developing Code (Python/SQL)
This domain covers writing PySpark and Spark SQL on Databricks: reading cloud files with the DataFrameReader, schema inference and explicit schemas, Delta Lake reads/writes, Delta Live Tables expectations, and tuning shuffle, broadcast, and memory behavior on clusters. Questions present realistic code or job symptoms and ask you to pick the correct API, config, or DLT expectation behavior.
Focused practice
Practice Developing Code (Python/SQL) questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Developing Code (Python/SQL)
Be able to write correct PySpark and Spark SQL for reading cloud files, transforming data, and writing Delta tables, plus configure cluster and shuffle settings. The single most important thing is choosing the right API or config for the stated symptom, especially Delta Live Tables expectation behavior.
Reading CSV/Parquet/JSON from cloud storage with spark.read and inferSchema or explicit StructType schemas
Tuning shuffle and memory with spark.sql.shuffle.partitions, spark.sql.autoBroadcastJoinThreshold, and executor memory settings
Writing and merging Delta tables with DeltaTable, MERGE, and schema evolution options
Defining Delta Live Tables pipelines in Python with @dlt.table and @dlt.expect_or_drop quality expectations
Watch out for
Common Developing Code (Python/SQL) exam traps
- ▸Assuming inferSchema is free: it triggers an extra pass over the file, so production code should supply an explicit schema instead
- ▸Raising executor memory first for OOM during shuffle when the real fix is increasing shuffle partitions or reducing broadcast size
- ▸Confusing DLT expectation decorators: expect_or_drop drops failing rows while expect_or_fail stops the pipeline, so using the wrong one breaks requirements
Question index
All Developing Code (Python/SQL) questions (42)
Click any question to see the full explanation, or start a practice session above.
You are using the Databricks CLI to automate workspace tasks. Which THREE of the following statements correctly identify capabilities of the Databricks CLI?
Medium2A Data Engineer is building a Structured Streaming pipeline that reads from a Kafka topic and writes to a Delta table. The pipeline must handle late-arriving data up to 2 hours and produce correct aggregations per 10-minute window. The engineer wants the streaming query to automatically clean up old state so the job does not accumulate unbounded state. Which combination of Structured Streaming features should be used?
Medium3You have a large Spark DataFrame that you need to filter and save as multiple smaller Parquet files based on the values in a 'region' column. Which method should you use to optimize the file layout for subsequent queries?
Medium4A Data Engineer needs to perform an 'upsert' operation on a Delta table. Which command provides the most efficient way to merge new data with existing records in a single transactional step?
Medium5Which TWO of the following are valid ways to pass configuration parameters to a Databricks Job task? (Select TWO)
Medium6A data engineer is optimizing a PySpark job that processes a large DataFrame and writes the result to a Delta table. The job currently uses repartition(100) before writing, but the output consists of many small files. The engineer wants to reduce the number of output files without shuffling the entire dataset again. Which approach should be used?
Medium7Which TWO of the following are benefits of using the Databricks Delta Lake 'Optimize' command? (Select TWO)
Medium8A data engineer is developing a PySpark job that reads from a Delta table and performs a series of transformations. The engineer notices that the job is slow and suspects that the query plan is not optimized because statistics are outdated. Which command should the engineer run to update the statistics for the Delta table to improve query performance?
Medium9A Data Engineer is using Databricks Asset Bundles (DABs) to manage a project. Where should the engineer define the job settings, such as clusters, schedules, and task dependencies?
Medium10You are writing a Spark application that uses a Broadcast Hash Join. You want to force the join to use broadcast optimization for a specific table. How do you implement this in PySpark?
Medium11You are optimizing a PySpark job that reads from a Delta table. You notice skewed data distribution on the 'customer_id' column, causing Task-level stragglers. Which transformation should you apply to the DataFrame to mitigate this skew during a join operation?
Medium12Which of the following describes the behavior of the 'OPTIMIZE' command in Databricks?
Medium13You are developing a Delta Live Tables (DLT) pipeline and need to ensure high data quality. Which TWO of the following statements correctly describe how Expectations work within DLT?
Hard14Which THREE of the following are benefits of using Delta Lake over standard Parquet files in Databricks?
Medium15A data engineer is tuning a Spark job and decides to use 'Z-Ordering' on a Delta table. Which THREE of the following are valid considerations when selecting columns for Z-Ordering?
Hard16A data engineer is implementing a Structured Streaming job that reads from a Kafka topic and writes to a Delta table. The engineer needs to ensure that the job can recover from failures without data loss or duplication. The job uses foreachBatch to perform upserts into the Delta table. Which checkpointing configuration is required to achieve exactly-once semantics?
Hard17A data engineer is building a Structured Streaming job that reads from a Kafka topic and writes to a Delta table. The pipeline must tolerate late-arriving data up to 10 minutes and update aggregations accordingly. The engineer wants to use a watermark on the event-time column. Which code snippet correctly applies the watermark and performs a 5-minute tumbling window aggregation?
Medium18When partitioning a large Delta table, what is the best practice regarding the number of unique values in the partition column?
Easy19Which of the following describes the correct use of a UDF (User Defined Function) in PySpark for production pipelines?
Medium20You are debugging a PySpark job that is experiencing severe memory pressure during a join on a massive column. You suspect data skew. Which approach is best to mitigate this issue?
Hard21Which THREE features are provided by Delta Lake when compared to standard Parquet files? (Select THREE)
Medium22A data engineer is using PySpark to process a large DataFrame and needs to reduce the number of partitions before writing to a Delta table to avoid creating too many small files. The DataFrame currently has 2000 partitions, each about 10 MB. The engineer wants to reduce the number of partitions to approximately 200 while minimizing data shuffling. Which approach is most appropriate?
Hard23A data engineer is using Delta Live Tables (DLT) to create a pipeline that processes streaming data. They need to ensure that the pipeline only processes new data since the last run and that the pipeline can recover from failures without reprocessing all data. Which combination of features should they use to achieve this?
Hard24A data engineer is building a Structured Streaming pipeline that reads from a Kafka topic and writes to a Delta table. The pipeline must handle late-arriving data and update already processed aggregates. Which watermark strategy should be used to allow updates to aggregates while bounding state store growth?
Medium25A data engineer is using Delta Live Tables (DLT) to build a pipeline that ingests data from a streaming source. The pipeline must ensure that the target table is updated incrementally and that data quality constraints are enforced. The engineer wants to use expectations to drop invalid records while maintaining pipeline performance. Which TWO of the following are true regarding DLT expectations and their behavior? (Choose two.)
Hard26You are migrating a legacy ETL process to Delta Live Tables (DLT). You have an existing table defined with a complex transformation that involves a custom Python function using a third-party library. How should you structure this in DLT to ensure the function is available and correctly applied?
Hard27When using the Unity Catalog, how should an engineer properly reference a table named 'sales' located in the 'finance' schema within the 'prod_catalog' catalog using Spark SQL?
Hard28You are processing sensitive PII data in a Delta table. You need to ensure that specific columns containing PII are not readable by general data analysts while maintaining the ability to perform aggregate analysis on those rows. Which feature should you implement?
Medium29Which command should be used to display the history of transactions performed on a Delta table, including operations like overwrites and updates?
Easy30Which property must be set to ensure a Spark Structured Streaming query can handle changes to the source data schema, such as adding a new column?
Hard31A data engineer needs to read a CSV file from cloud storage into a Spark DataFrame in Databricks. The file has a header row and uses commas as delimiters. The engineer wants to infer the schema automatically. Which code snippet correctly reads the file?
Easy32Your team is using a shared cluster for development. A user reports that their job is slow because the cluster memory is frequently filled by large data broadcasts. What configuration adjustment should you make to prevent this issue across all jobs on the cluster?
Hard33When running a PySpark job, you receive an 'Out of Memory (OOM)' error during a shuffle operation. Which configuration is the most appropriate to address this first?
Hard34When designing a production-grade data pipeline in Databricks, what is the recommended approach for managing secrets such as database credentials?
Medium35A data engineer is developing a PySpark job that reads a large Delta table, performs a groupBy on a high-cardinality column, and writes the result to another Delta table. The job is experiencing performance issues due to data skew. The engineer wants to optimize the shuffle by using salting. Which approach correctly implements salting to distribute the skewed keys evenly?
Medium36A Data Engineer is developing a Delta Live Tables (DLT) pipeline using Python. They need to ensure that records failing a specific data quality check are dropped, but the pipeline continues to process the remaining valid records. Which expectation syntax should the engineer implement?
Medium37When designing a streaming pipeline using Structured Streaming, which THREE of the following are necessary to ensure 'exactly-once' processing semantics in Databricks?
Medium38Which THREE of the following are benefits of using Delta Lake over standard Parquet files for your data lake storage?
Hard39Which TWO of the following are valid ways to trigger a job in Databricks?
Medium40You are writing a PySpark script to join two large tables. You want to ensure the join operation is optimized for performance by broadcasting the smaller table. Which configuration property should you adjust, or code construct should you use, to force this behavior?
Medium41An engineer is writing a Python function to process data in a Databricks Notebook. Which command should they use to ensure that secrets, such as API keys, are not hardcoded or exposed in the plain text of the notebook?
Medium42Refer to the exhibit. You are appending data to an existing Delta table. What is the most likely cause of this error, and how should you resolve it?
MediumOther domains
All Databricks-DE-Pro exam domains
Frequently asked questions
- What does the Developing Code (Python/SQL) domain cover on the Databricks-DE-Pro exam?
- Be able to write correct PySpark and Spark SQL for reading cloud files, transforming data, and writing Delta tables, plus configure cluster and shuffle settings. The single most important thing is choosing the right API or config for the stated symptom, especially Delta Live Tables expectation behavior.
- How many questions are in this domain?
- This page lists all 42 Developing Code (Python/SQL) questions in the Databricks-DE-Pro question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Developing Code (Python/SQL) questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.