Courseiva

CCNA Troubleshooting and Tuning DataFrame Apps Questions

41 questions · Troubleshooting and Tuning DataFrame Apps · All types, answers revealed

1
MCQhard

A PySpark job reads a Delta table, applies a filter on a timestamp column, and then performs a window function partitioned by customer_id ordered by event_time. The job is slow, and the physical plan shows the filter is applied after the window. The developer wants the filter to reduce data before the window shuffle. Which action should the developer take?

A.Increase spark.sql.windowExec.buffer.spill.threshold so the window operator can spill more data to disk.
B.Repartition the DataFrame by customer_id before applying the filter so the window avoids a shuffle.
C.Apply the filter on the timestamp column immediately after reading the Delta table, before calling the window function, so predicate pushdown and early filtering reduce rows entering the shuffle.
D.Cache the DataFrame before the window and call count() to materialize it, then apply the filter.
AnswerC

Filtering right after the read lets Spark push the predicate into the Delta scan where possible and reduces the number of rows that must be shuffled for the window. This lowers shuffle size, memory pressure, and runtime, and it is the direct way to get the filter to act before the window partitioning.

Why this answer

The physical plan shows the filter after the window, meaning the window shuffle processes all rows. Moving the filter to immediately after the read lets Spark push it into the Delta scan where possible and shrink the dataset before the window shuffle, reducing shuffle volume, memory use, and overall runtime.

Exam trap

The trap here is tuning window spill or repartitioning by the window key, when the real win is applying the filter before the window so fewer rows ever enter the shuffle.

2
MCQmedium

When analyzing a Spark job's execution plan, what does a 'BroadcastHashJoin' indicate compared to a 'SortMergeJoin'?

A.BroadcastHashJoin requires both tables to be pre-sorted.
B.BroadcastHashJoin avoids a shuffle operation.
C.SortMergeJoin is always faster than BroadcastHashJoin.
D.BroadcastHashJoin only works for inner joins.
AnswerB

Because the smaller table is sent to every executor node, the larger table remains in its original partition structure, and no data is shuffled across the network. This makes it a much faster join strategy compared to SortMergeJoin, which requires full shuffles of both tables to align the keys.

Why this answer

A 'BroadcastHashJoin' indicates that Spark has decided to send the smaller table to all executors to perform the join in memory, avoiding a shuffle. A 'SortMergeJoin' involves shuffling both tables, which is significantly more expensive. Understanding when Spark chooses one over the other helps developers identify if they need to explicitly broadcast small tables or adjust configuration thresholds to optimize join performance and minimize costly cluster-wide shuffle operations.

Exam trap

Candidates often think BroadcastHashJoin speeds up execution by sorting data faster, missing that its primary advantage is completely eliminating network shuffling.

3
MCQmedium

A Databricks notebook job calls a PySpark UDF built on a Python function that performs string parsing. The job completes successfully on a small sample, but on the full production dataset it fails with a PythonException and the executor logs show high garbage collection time. Which change is the most appropriate first step to make the job reliable without changing business logic?

A.Set spark.sql.execution.arrow.pyspark.enabled to true so the UDF uses Arrow for every row it processes.
B.Mark the UDF with the deterministic flag and register it with spark.udf.register so Spark can cache its results automatically.
C.Replace the Python UDF with a Spark SQL built-in function such as regexp_extract or split, which runs inside the JVM and avoids Python serialization and per-row interpreter overhead.
D.Increase spark.executor.memory and spark.executor.cores so each executor has more headroom to run the Python workers.
AnswerC

Built-in Spark SQL functions execute in the JVM with code generation, so they avoid Python worker startup, serialization, and interpreter overhead per row. For string parsing, regexp_extract, split, or transform are usually equivalent to the Python logic. This directly reduces executor memory pressure and GC time while preserving the same results. It is the least invasive change and the standard tuning step for Python UDF hotspots.

Why this answer

A scalar Python UDF forces each row through a Python worker, incurring serialization, interpreter startup, and per-row overhead that shows up as GC pressure at scale. Rewriting the logic with Spark SQL built-in functions keeps execution in the JVM with whole-stage code generation, eliminating that overhead and typically resolving both the reliability and performance symptoms. It preserves the business logic while removing the root cause rather than masking it with more resources.

Exam trap

The trap here is assuming that more executor memory or cores will fix a Python UDF hotspot, when the real cost is the per-row Python worker and serialization overhead.

4
MCQmedium

A Spark job is experiencing data skew during a join operation on a key column. Which strategy is most effective for mitigating this issue without changing the business logic?

A.Increase the spark.sql.shuffle.partitions configuration value significantly.
B.Broadcast the larger table to all executor nodes.
C.Add a random prefix to the join key of the skewed table and replicate the join key of the other table.
D.Cache the skewed table in memory before performing the join.
AnswerC

Salting distributes the skewed keys across different partitions by creating a composite key. By replicating the join key in the smaller table, you ensure that the original join condition is still satisfied while allowing Spark to process the previously bottlenecked key across several parallel tasks simultaneously.

Why this answer

Salting the join key by appending a random integer helps distribute the skewed keys across multiple partitions. This prevents a single executor from handling the bulk of the data, which is the primary cause of long-running tasks in skewed joins. Understanding how to repartition data based on a salted key is critical for developers tasked with optimizing performance in distributed systems where key distribution is inherently uneven across the cluster nodes.

Exam trap

Candidates often forget that when salting a join key on one table, the matching table must also be expanded or replicated to join all the salted variants.

5
MCQmedium

Which THREE factors should a developer consider when choosing a partition count for a shuffle operation?

A.The total size of the data being shuffled.
B.The total number of available cores in the cluster.
C.The memory limit of the driver node.
D.The specific task being performed (e.g., aggregation vs. join).
E.The number of rows in the source file header.
AnswerA, B, D

Data size directly influences how much data each task must handle. Larger datasets require more partitions to keep task sizes manageable and prevent OOM errors, whereas small datasets can be processed with fewer partitions to reduce the overhead associated with launching and managing tasks in the Spark cluster.

Why this answer

Selecting the right partition count is a balancing act between parallelism and overhead. Too few partitions lead to underutilization of the cluster, while too many partitions create excessive overhead for the scheduler. Developers must balance the data volume, the available cluster resources (cores), and the specific requirements of the operation to ensure optimal performance and avoid common performance pitfalls like task scheduling bottlenecks or executor memory issues.

Exam trap

Candidates often pick static rules of thumb like 'always use 200 partitions' instead of evaluating data volume, cluster cores, and operation types.

6
MCQhard

Refer to the exhibit. What is the most likely cause of the repeated ExecutorLostFailure messages in the logs?

A.The driver node is experiencing a network latency issue.
B.The Spark job is attempting to use too many partitions.
C.The executor process reached its memory limit and was killed.
D.The shuffle partition count is set too low for the data.
AnswerC

When a Spark executor consumes more memory than allocated, the underlying container (like YARN or Kubernetes) will kill it, leading to the ExecutorLostFailure log. This often occurs during heavy data processing or shuffles where memory consumption exceeds the configured JVM heap limits or container memory limits.

Why this answer

ExecutorLostFailure indicates that an executor has died, often due to exceeding memory limits or heartbeats failing. When processing large data, the executor's heap memory can be exhausted, leading to a crash. This signifies that the cluster resources are insufficient for the current workload, requiring either more memory per executor or more efficient data processing strategies to prevent the executor from being terminated by the operating system or the cluster manager.

Exam trap

Candidates often mistake 'ExecutorLostFailure' for a network issue, ignoring the fact that it is frequently the result of the OOM killer terminating the executor due to excessive memory usage.

7
MCQmedium

A developer runs a PySpark job on Databricks that filters a Delta table and then calls count() and show(). The Spark UI shows two separate scans of the same table. The developer wants to avoid the second scan without changing the filter logic. Which action should be taken?

A.Call checkpoint() on the filtered DataFrame to write it to a reliable path before count().
B.Enable spark.sql.adaptive.enabled and set spark.sql.adaptive.coalescePartitions.enabled to true.
C.Call persist() with StorageLevel.MEMORY_AND_DISK on the filtered DataFrame before count() and show().
D.Set spark.sql.files.maxPartitionBytes to 256 MB so each scan reads fewer, larger partitions.
AnswerC

persist() with MEMORY_AND_DISK materializes the filtered DataFrame on the first action and reuses the stored partitions for the second action, eliminating the duplicate scan. It is equivalent in effect to cache() but with an explicit storage level, and it directly targets the repeated scan shown in the Spark UI without altering the filter logic.

Why this answer

Persisting the filtered DataFrame with MEMORY_AND_DISK materializes it on the first action and lets the second action read the stored partitions instead of rescanning the Delta table. AQE, maxPartitionBytes, and checkpointing do not provide in-memory reuse across actions; they change plan shape or write to disk rather than caching the intermediate result for repeated access.

Exam trap

The trap here is assuming that adaptive execution or read-partition tuning will prevent a second scan, when only an explicit persist or cache keeps the intermediate DataFrame available across actions.

8
Multi-Selecthard

A Spark job is running slower than expected due to excessive shuffling. Which TWO of the following techniques would directly reduce the volume of data transferred over the network?

Select 2 answers
A.Broadcast join for small tables.
B.Increase spark.sql.shuffle.partitions.
C.Filter and select columns early.
D.Enable dynamic resource allocation.
E.Increase spark.driver.memory.
AnswersA, C

Broadcast joins send the entire small table to every executor, allowing the join to occur locally. This eliminates the need to shuffle the large table, which is the most expensive part of a join, thereby reducing network overhead and significantly improving the performance of the overall job.

Why this answer

Reducing network traffic is critical in Spark performance tuning. By utilizing features like broadcast joins, you avoid the shuffle phase entirely when one table is small. Additionally, using filter and select operations early in the transformation pipeline ensures that only the necessary rows and columns are transmitted, significantly decreasing the total data volume shuffled during wide transformations.

Exam trap

Candidates often suggest increasing cluster size to solve shuffling issues. While this helps performance, it does not reduce the *volume* of data shuffled; only filtering or broadcasting does that.

9
Multi-Selecthard

A developer is tuning a Spark job that performs a join between a large fact table and a medium-sized dimension table. The job suffers from data skew, with a few keys having a disproportionately large number of rows. The developer wants to mitigate the skew. Which two actions are most effective? (Choose two.)

Select 2 answers
A.Enable Adaptive Query Execution (AQE) and set spark.sql.adaptive.skewJoin.enabled to true.
B.Broadcast the medium-sized dimension table to all executors.
C.Use a repartition on the join key before the join to evenly distribute the data.
D.Manually salt the skewed keys in the fact table by adding a random suffix and replicate the dimension table accordingly.
E.Increase spark.sql.shuffle.partitions to a very high number.
AnswersA, D

AQE with skew join optimization automatically detects skewed partitions during the shuffle and splits them into smaller sub-partitions. This balances the workload across tasks without manual intervention. It is a built-in feature that directly addresses skew by dynamically handling oversized partitions, making it an effective solution.

Why this answer

Data skew during a join occurs when a few keys have many more rows than others, causing some tasks to run much longer. Adaptive Query Execution with skew join optimization dynamically splits skewed partitions, while manual salting distributes hot keys across multiple tasks. Both approaches balance the workload and reduce the impact of skew.

Exam trap

The trap here is assuming that increasing shuffle partitions or repartitioning on the join key will fix skew, when they can actually worsen it by concentrating hot keys.

10
MCQmedium

Which configuration parameter should be adjusted to change the default number of partitions when reading from a shuffle-heavy operation?

A.spark.executor.memory
B.spark.default.parallelism
C.spark.sql.shuffle.partitions
D.spark.driver.memory
AnswerC

This parameter explicitly defines the number of partitions to use when shuffling data for joins or aggregations. By tuning this value, developers can control the level of parallelism, which is crucial for balancing the trade-off between task overhead and executor utilization during resource-intensive stages of a Spark job.

Why this answer

The `spark.sql.shuffle.partitions` configuration is the primary setting used to control the parallelism of shuffle-based operations in Spark SQL. Adjusting this value allows developers to match the level of parallelism to the specific data volume and cluster resources. Understanding how this setting impacts performance is essential for fine-tuning Spark applications to prevent task contention and ensure that the cluster is fully utilized during data-intensive stages of a job.

Exam trap

Candidates often confuse spark.sql.shuffle.partitions with spark.default.parallelism or source-specific reading options when trying to control shuffle stages.

11
MCQeasy

A developer runs a PySpark job that joins a 10 GB DataFrame with a 50 MB lookup DataFrame. The job takes far longer than expected, and the Spark UI shows a SortMergeJoin with a large shuffle read and write for both sides. The developer wants the smallest change that most improves performance. Which action should the developer take?

A.Increase spark.sql.shuffle.partitions so the sort-merge join shuffle uses more tasks and completes faster.
B.Broadcast the 50 MB lookup DataFrame using broadcast() or by raising spark.sql.autoBroadcastJoinThreshold above 50 MB.
C.Repartition the 50 MB lookup DataFrame by the join key before the join to co-locate matching rows.
D.Cache the 10 GB DataFrame with persist(StorageLevel.MEMORY_AND_DISK) before the join.
AnswerB

Broadcasting the small side sends a copy of the 50 MB lookup to every executor and converts the join to a BroadcastHashJoin, eliminating the shuffle of both DataFrames. This is the minimal change that removes the expensive sort-merge shuffle and typically yields the largest speedup for a small-to-large join.

Why this answer

The Spark UI shows a SortMergeJoin with large shuffle on both sides, which is unnecessary when one input is only 50 MB. Broadcasting the small lookup table turns the join into a BroadcastHashJoin, removes the shuffle and sort of both DataFrames, and requires only a hint or a threshold adjustment.

Exam trap

The trap here is tuning shuffle partitions or caching when the join strategy itself is the problem and a broadcast would eliminate the shuffle entirely.

12
MCQhard

A Spark job reads a large Parquet dataset, performs a groupBy on a high-cardinality column, and writes the result to a Delta table. The job fails with a FetchFailedException on a particular executor. The Spark UI shows that the executor had sufficient memory but the shuffle fetch failed due to a connection reset. Which configuration change is most likely to resolve this issue?

A.Increase spark.executor.memory to prevent the executor from being killed during shuffle.
B.Increase spark.reducer.maxSizeInFlight to allow larger shuffle blocks to be fetched.
C.Set spark.shuffle.io.maxRetries and spark.shuffle.io.retryWait to higher values to handle transient network issues.
D.Set spark.sql.adaptive.enabled=false to disable adaptive query execution and avoid shuffle re-computation.
AnswerC

FetchFailedException due to connection reset often indicates transient network problems or shuffle service timeouts. Increasing shuffle I/O retries and retry wait allows the reducer to retry fetching blocks after a failure, improving resilience. This directly addresses the connection reset by giving the fetch more attempts and time to succeed.

Why this answer

A FetchFailedException with connection reset during shuffle fetch is typically caused by transient network issues or shuffle service timeouts. Increasing spark.shuffle.io.maxRetries and spark.shuffle.io.retryWait makes the shuffle fetch more resilient by retrying failed attempts. This is the most direct configuration change to handle temporary network glitches without altering the overall job logic or resource allocation.

Exam trap

The trap here is assuming that memory or query planning changes will fix a shuffle fetch failure, when the error is network-related and requires retry tuning.

13
MCQmedium

A streaming DataFrame job on Databricks writes to a Delta table every 10 seconds. Over several hours, the number of files in the target directory grows into the hundreds of thousands, and downstream reads slow dramatically. The job uses foreachBatch with a write that produces many small files per micro-batch. Which action should be taken to reduce the small-files problem for this streaming write?

A.Call repartition(1) on the streaming DataFrame before the write in foreachBatch.
B.Configure the write with optimizeWrite enabled (or set the target table property delta.autoOptimize.optimizeWrite = true) so Spark coalesces output files per partition before committing.
C.Increase the trigger interval from 10 seconds to 10 minutes so each micro-batch writes more data.
D.Set spark.sql.shuffle.partitions to 1 so all output is written by a single task.
AnswerB

Optimize Write coalesces the many small files produced by each micro-batch into fewer, larger files before they are committed to the Delta table. This directly attacks file proliferation at write time, reducing the file count and improving downstream read performance without changing the streaming logic.

Why this answer

The streaming job creates many small files because each micro-batch writes multiple files per partition. Optimize Write coalesces those files into fewer, larger ones at commit time, directly reducing file proliferation and improving downstream read performance without sacrificing parallelism or increasing latency.

Exam trap

The trap here is reaching for repartition(1) or a single shuffle partition to reduce file count, which fixes file count by destroying write parallelism.

14
MCQmedium

Which file format is best suited for performance-critical Spark applications that require efficient schema enforcement and column pruning?

A.CSV
B.JSON
C.Parquet
D.XML
AnswerC

Parquet is a highly optimized columnar format that allows Spark to skip reading unnecessary columns and push down filter operations to the storage layer. This minimizes I/O and CPU overhead, making it the industry standard for performance-critical, large-scale data processing workflows within Databricks and the Apache Spark ecosystem.

Why this answer

Parquet is a columnar storage format that natively supports predicate pushdown and column pruning, allowing Spark to read only the necessary data from disk. This drastically reduces I/O throughput and improves query performance significantly. Understanding why columnar formats are superior to row-based formats for analytics is a foundational concept for Databricks developers designing high-performance data lakes and efficient ETL pipelines at scale.

Exam trap

Candidates often select CSV or JSON formats, confusing human readability with performance-critical analytics features like native columnar pruning.

15
MCQhard

Refer to the exhibit. What is the best way to resolve this error?

A.Increase the spark.driver.maxResultSize setting.
B.Write the results to a data lake instead of collecting them.
C.Disable the Spark broadcast join threshold.
D.Increase the number of executor instances.
AnswerB

Writing results to persistent, distributed storage (like Parquet files on S3/ADLS) is the architecture-correct way to handle large outputs. This avoids sending all data to the driver, allowing the Spark executors to perform the work in parallel and keeping the driver's memory footprint small and stable throughout the execution.

Why this answer

This error occurs when the result set returned to the driver from the executors exceeds the limit defined by `spark.driver.maxResultSize`. The best practice is to stop trying to bring huge datasets back to the driver, and instead write the results to a distributed storage system like S3 or ADLS. This prevents the driver from becoming a memory bottleneck and ensures the application remains scalable for large-scale data processing tasks.

Exam trap

Candidates often attempt to increase 'spark.driver.maxResultSize' to fix the error, rather than changing their code pattern to avoid bringing massive data back to the driver node.

16
MCQeasy

A developer notices that a PySpark DataFrame transformation chain runs a full scan of a Delta table each time a new action is invoked, even though the source data has not changed. The developer wants to persist the intermediate DataFrame in memory across actions. Which method should be used?

A.Call DataFrame.collect() after each transformation to materialize the results on the driver.
B.Call DataFrame.cache() before the first action so the DataFrame is stored in memory on first computation.
C.Set spark.sql.adaptive.enabled to true so the optimizer reuses prior scan results automatically.
D.Call DataFrame.checkpoint() to truncate the lineage and write the data to a reliable file system.
AnswerB

cache() marks the DataFrame for in-memory persistence with the default MEMORY_AND_DISK storage level, so the first action materializes it and subsequent actions reuse the cached partitions instead of rescanning the Delta table. This directly addresses the repeated full scans described in the scenario and is the idiomatic way to persist across multiple actions.

Why this answer

cache() persists the DataFrame across actions using the default MEMORY_AND_DISK level, so the first action computes and stores partitions and later actions read from the cache rather than rescanning the Delta table. Checkpointing writes to disk and cuts lineage but does not keep data in memory, while AQE and collect() do not provide reusable in-memory persistence.

Exam trap

The trap here is conflating checkpointing, which truncates lineage to disk, with caching, which keeps the materialized DataFrame available in memory for repeated actions.

17
MCQeasy

Which action should be taken to optimize a Spark application that performs multiple operations on the same DataFrame and shows evidence of redundant re-computations in the Spark UI DAG visualization?

A.Increase the number of shuffle partitions.
B.Use the cache() or persist() method on the DataFrame.
C.Implement a custom partitioner for all joins.
D.Convert the DataFrame to a RDD.
AnswerB

Caching stores the materialized result of a DataFrame. When the application accesses the data again, Spark retrieves it from the cache rather than re-computing the entire lineage. This is the correct technique for eliminating redundant work when multiple downstream operations depend on the same intermediate DataFrame.

Why this answer

Persisting (or caching) a DataFrame instructs Spark to store the computed results of the DataFrame in memory or on disk. This is vital when the same DataFrame is accessed multiple times across different stages, as it prevents Spark from re-executing the entire lineage graph from the source, significantly reducing processing time and resource consumption in complex, multi-step analytical pipelines.

Exam trap

Candidates often mistake cache() for an action. They forget that cache() is a lazy transformation and must be followed by an action like count() or write() to actually materialize the data in memory.

18
MCQhard

Which TWO of the following techniques effectively reduce the shuffle volume in a Databricks Spark job?

A.Apply filter transformations as early as possible in the DataFrame lineage.
B.Increase the memory allocated to the Spark driver.
C.Select only necessary columns before performing wide transformations.
D.Enable dynamic allocation of executors.
E.Set the spark.sql.shuffle.partitions to 1.
AnswerA, C

Early filtering, or predicate pushdown, reduces the total number of rows processed. By discarding irrelevant records before any shuffle operation occurs, you significantly decrease the amount of data written to disk and transferred over the network, leading to faster execution times and lower resource consumption during the shuffle phase.

Why this answer

Reducing shuffle volume is essential for performance, as shuffling moves data across the network, which is the most expensive operation in Spark. Techniques like predicate pushdown and column pruning minimize the amount of data read and processed before the shuffle phase occurs. Mastering these optimizations ensures that jobs remain scalable as datasets grow, preventing network congestion and I/O saturation during complex transformations or aggregate operations on large distributed DataFrames.

Exam trap

Candidates mistakenly believe that adding a repartition command or caching early reduces shuffle volume, confusing memory persistence with network transfer reduction.

19
MCQeasy

A Databricks job fails with an OutOfMemoryError on the driver. The job collects a large DataFrame to the driver for local processing. Which action should you take to resolve this?

A.Increase the executor memory.
B.Increase spark.driver.maxResultSize.
C.Set spark.sql.shuffle.partitions to a higher value.
D.Replace collect() with take(n) or write the DataFrame to storage.
AnswerD

collect() transfers the entire DataFrame to the driver, which can cause an OutOfMemoryError if the data is large. Using take(n) retrieves only a limited number of rows, or writing the DataFrame to storage avoids bringing all data to the driver. This directly addresses the driver memory issue by reducing the amount of data transferred. It is the correct approach.

Why this answer

The OutOfMemoryError on the driver is caused by collecting a large DataFrame. The correct fix is to avoid transferring all data to the driver by using take() for a sample or writing the results to storage. Increasing driver or executor memory or adjusting shuffle partitions does not address the fundamental problem of excessive data transfer to the driver.

Exam trap

The trap here is increasing driver memory or maxResultSize instead of reducing the amount of data collected to the driver.

20
MCQhard

When a Spark job is stuck in a shuffle phase, what is the most effective first step to identify the root cause of the performance bottleneck?

A.Restart the cluster to clear any cached data.
B.Check the Spark UI 'Stages' tab for uneven task durations.
C.Increase the spark.executor.memory value.
D.Enable speculative execution immediately.
AnswerB

The Spark UI allows you to see the duration and data size of every task in a stage. If one task takes significantly longer than others, it is a clear indicator of data skew. This allows the developer to isolate the specific partition causing the delay and implement a remediation strategy.

Why this answer

The Spark UI provides a detailed breakdown of stages, tasks, and memory usage. Examining the 'Stages' and 'SQL' tabs reveals which tasks are taking the longest, which is indicative of skew or resource contention. Learning to interpret the Spark UI is the single most important skill for a Databricks developer, as it turns opaque 'stuck' jobs into actionable data regarding partition distribution and executor behavior.

Exam trap

Candidates immediately restart the cluster or rewrite business logic instead of inspecting task metrics in the Spark UI to isolate data skew.

21
MCQmedium

A Spark application reads a large CSV file and performs a series of transformations, including a filter and a join with a small lookup table. The job is running slowly, and the Spark UI shows that the CSV parsing stage is taking a long time. Which action would most improve performance?

A.Increase the number of partitions when reading the CSV by setting a lower spark.sql.files.maxPartitionBytes.
B.Set spark.sql.autoBroadcastJoinThreshold to -1 to disable broadcast join.
C.Convert the CSV to Parquet and read the Parquet file instead.
D.Cache the CSV DataFrame immediately after reading.
AnswerC

Parquet is a columnar format that supports predicate pushdown and column pruning, which drastically reduces I/O and parsing overhead compared to CSV. CSV parsing is row-based and requires reading and parsing every field, even if only a few columns are needed. Converting to Parquet allows Spark to read only the necessary columns and skip irrelevant data, significantly speeding up the initial read and subsequent transformations.

Why this answer

Converting the CSV to Parquet addresses the slow CSV parsing by leveraging Parquet's columnar storage, which enables column pruning and predicate pushdown. This reduces the amount of data read and parsed, directly improving the performance of the initial stage. The other options either do not target the parsing bottleneck or could worsen performance.

Exam trap

The trap here is assuming that increasing partitions or caching will fix slow CSV parsing, when the format itself is the primary bottleneck.

22
MCQhard

A Spark Structured Streaming job on Databricks reads from a Delta table and writes micro-batches to another Delta table with a 30-second trigger. After several hours, the batch duration grows from 4 seconds to over 60 seconds and the job falls behind. The source table is compacted regularly, and the cluster has enough CPU. Which tuning action is most likely to restore the original batch duration?

A.Increase spark.sql.shuffle.partitions to a very high value so each micro-batch uses more tasks and finishes faster.
B.Reduce the number of shuffle partitions and disable adaptive query execution so the streaming plan stays deterministic across micro-batches.
C.Enable Delta Lake optimized writes and tune the trigger interval so the job processes larger, less frequent micro-batches.
D.Check for stateful aggregation or deduplication without a watermark and add a watermark with a bounded state cleanup so state does not grow unbounded.
AnswerD

Unbounded state from a stateful operation without a watermark causes each micro-batch to scan an ever-growing state store, which steadily increases batch duration and can make the job fall behind. Adding a watermark lets Spark drop state older than the watermark, keeping state bounded and batch times stable. This directly addresses a gradual, hours-long degradation pattern.

Why this answer

A steady increase in batch duration over hours with a compacted source and adequate CPU strongly indicates unbounded state growth in a stateful streaming operation. Without a watermark, Spark retains all keys indefinitely, so each micro-batch reads and writes a larger state store. Adding a watermark with a bounded delay lets Spark evict old state, keeping per-batch work roughly constant and restoring stable batch times.

Exam trap

The trap here is treating a slowly growing streaming batch time as a parallelism problem and adding shuffle partitions instead of investigating state accumulation.

23
MCQhard

A structured streaming DataFrame writes to a Delta table with a foreachBatch function that performs an upsert. After a cluster restart, the stream reprocesses some micro-batches and duplicate rows appear in the target table. The foreachBatch code already uses MERGE keyed on a unique id. Which change best prevents duplicates after restart?

A.Set the trigger to processingTime='1 minute' so micro-batches are larger and less likely to overlap.
B.Add a checkpointLocation to the writeStream options so the stream records its progress and resumes from the last committed offset.
C.Switch the output mode from append to complete so each micro-batch is fully replaced in the target table.
D.Add a deduplication step using dropDuplicates on the unique id inside the foreachBatch before the MERGE.
AnswerB

Structured streaming relies on the checkpoint location to persist progress, offsets, and state. Without it, a restarted query has no record of which micro-batches were committed and reprocesses data, producing duplicates even when the MERGE key is unique. Setting checkpointLocation gives exactly-once semantics for the sink when combined with an idempotent operation such as MERGE. This is the direct fix for reprocessing after restart.

Why this answer

Structured streaming tracks processed offsets and state in the checkpoint location. Without a checkpoint, a restarted query cannot know what it already committed and replays earlier micro-batches, creating duplicates even with an idempotent MERGE. Supplying checkpointLocation lets the query resume from the last committed offset, so each batch is processed once.

Combined with the existing MERGE on a unique id, this yields the intended exactly-once behavior on restart.

Exam trap

The trap here is believing a MERGE on a unique key alone guarantees exactly-once, when restart reprocessing is actually governed by the streaming checkpoint.

24
MCQeasy

A developer notices that a DataFrame transformation chain is executed twice: once for a count action used for logging and again for a write action. The source is a large Delta table and the repeated scan adds several minutes. Which action avoids the duplicate computation with the least risk?

A.Convert the DataFrame to a Pandas DataFrame for the count and then write from the original Spark DataFrame.
B.Call cache or persist on the transformed DataFrame before the count so the second action reuses the materialized data.
C.Set spark.sql.shuffle.partitions equal to the number of executor cores to speed up each execution.
D.Increase spark.sql.autoBroadcastJoinThreshold so more joins are broadcast and the plan becomes cheaper.
AnswerB

Caching the transformed DataFrame materializes it once, so the subsequent write action reads the cached partitions instead of recomputing the entire lineage from the Delta table. This directly eliminates the duplicate scan. It is a targeted change with predictable memory cost and no change to results. For a DataFrame reused across multiple actions, caching is the standard remedy.

Why this answer

When the same DataFrame lineage feeds two actions, Spark recomputes it for each action unless the intermediate result is materialized. Calling cache or persist after the expensive transformations stores the result so the later write reads it directly, removing the second full scan of the Delta table. The change is local, reversible, and does not alter results, making it the lowest-risk option among those presented.

Exam trap

The trap here is trying to tune shuffle or join settings for a problem that is actually caused by recomputing the same lineage across two separate actions.

25
MCQhard

A Databricks job joins a 500 GB sales table with a 300 GB returns table on a customer_id key. A few customer_id values account for a large fraction of rows on both sides, and the job fails with executor OOM during the join. The developer wants to distribute the hot keys across more partitions without changing the query logic. Which technique should be used?

A.Broadcast the 300 GB returns table to every executor to avoid shuffling the skewed keys.
B.Enable Adaptive Query Execution with skew join handling so Spark splits skewed partitions at runtime.
C.Increase spark.sql.shuffle.partitions to 8000 to spread the skewed keys across more tasks.
D.Repartition both DataFrames by customer_id with 8000 partitions before the join.
AnswerB

AQE skew join handling detects partitions that are much larger than the median after the shuffle and splits them into smaller sub-partitions, each processed as a separate task. This distributes the hot customer_id values across multiple tasks and relieves the executor OOM without changing the join logic. It is the built-in mechanism for exactly this scenario.

Why this answer

AQE skew join handling identifies partitions that are disproportionately large after the shuffle and splits them into smaller sub-partitions, each handled by a separate task. This spreads the hot customer_id values across executors and resolves the OOM. Raising shuffle partitions or repartitioning by the join key keeps each hot key in one partition, and broadcasting a 300 GB table is infeasible.

Exam trap

The trap here is believing that increasing shuffle partitions or repartitioning by the join key will split a hot key, when hash partitioning always routes all rows for one key to a single partition.

26
Multi-Selectmedium

A PySpark DataFrame job on Databricks runs slowly. Inspection of the Spark UI shows that a shuffle stage writes 200 partitions but downstream stages process only a few, and the physical plan shows an Exchange before a filter. Which two changes are most likely to improve performance? (Choose two.)

Select 2 answers
A.Increase spark.sql.shuffle.partitions to a much larger number so each task handles fewer rows.
B.Repartition the DataFrame on the join key before the shuffle to colocate matching rows.
C.Reduce the number of shuffle partitions or coalesce the output so fewer, larger partitions are produced.
D.Cache the shuffled DataFrame with persist so downstream stages reuse the same partitions.
E.Apply the filter before the join or aggregation that triggers the Exchange so less data is shuffled.
AnswersC, E

When a shuffle produces many near-empty partitions, consolidating them reduces task launch overhead and improves per-task efficiency. Coalescing after the shuffle or lowering the partition count creates fewer, better-sized partitions for downstream stages. This matches the observed pattern of many partitions with little data. It addresses the scheduling waste visible in the UI.

Why this answer

The plan shows an Exchange feeding stages that process only a few partitions, meaning shuffle volume and partition sizing are both inefficient. Filtering earlier cuts the rows that ever reach the shuffle, and consolidating partitions removes the overhead of many tiny tasks. Together they reduce both the data moved across the network and the number of tasks scheduled, which is what the UI evidence points to.

Resource or caching changes would not target either cause.

Exam trap

The trap here is reaching for more shuffle partitions by reflex, when the UI actually shows too many near-empty partitions and excess shuffle input.

27
MCQmedium

A Databricks job reads a Parquet dataset, applies a chain of `withColumn` transformations, and writes it back with `.write.mode("overwrite").parquet(path)`. The Spark UI shows 6000 small output files totaling 50 GB, and a downstream reader is slow because of per-file overhead. You want to reduce the file count while keeping the write as a Spark-native operation. What should you do?

A.Convert the DataFrame to an RDD and call `saveAsTextFile` to let Spark choose the file layout.
B.Set `spark.sql.files.maxPartitionBytes` to a very large value before the write.
C.Add `.option("maxRecordsPerFile", 100000)` to the write call to merge small files.
D.Call `df.repartition(200)` immediately before the write to consolidate data into fewer output partitions.
AnswerD

Each output partition produces one file, so reducing the partition count directly reduces the file count. Repartitioning performs a full shuffle that evenly distributes rows across the target partitions, which yields larger, more balanced files that downstream readers can scan with far less per-file overhead.

Why this answer

Output file count equals the number of partitions at write time, so consolidating partitions with a shuffle-based repartition before writing produces fewer, larger files. Read-side settings and per-file record caps do not change the write fan-out, and dropping to RDD text output sacrifices the columnar format that makes the downstream reads fast.

Exam trap

The trap here is confusing a read-side tuning knob such as the maximum partition bytes with a write-side control over output file count.

28
MCQmedium

Refer to the exhibit. You are reviewing the logs for a Spark application and notice the warning regarding broadcasting a large task binary. What is the most likely cause and mitigation?

A.The partition size is too large; increase spark.sql.files.maxPartitionBytes.
B.A large variable is captured in a closure; use a Broadcast Variable.
C.The executors have insufficient memory; increase spark.executor.memory.
D.The cluster is out of network bandwidth; enable compression.
AnswerB

When a large object is referenced inside a transformation, Spark tries to serialize it with the task. A broadcast variable provides a mechanism to distribute the object efficiently once per node rather than once per task, preventing the warning and reducing the serialization burden on the driver.

Why this answer

The warning indicates that a large object, likely a high-dimensional collection or a large variable, is being captured in a closure and broadcast to all executors. This increases network pressure and memory usage. The mitigation is to use a broadcast variable or remove the reference to the large object from the closure to prevent serialization overhead and potential performance degradation.

Exam trap

Candidates often assume the solution is to increase the executor memory. However, the error is caused by serializing a large object in a closure, which remains an issue regardless of total heap size.

29
MCQmedium

Refer to the exhibit. Which action is the most likely cause of the error shown in the Spark job logs?

A.The executor nodes have insufficient memory for the shuffle operation.
B.The driver is attempting to aggregate a massive dataset into its memory.
C.The cluster configuration uses too many small partitions.
D.The broadcast join threshold is set too low for the current job.
AnswerB

The collect() method triggers the transfer of all partitions from worker nodes to the driver node. If the combined data exceeds the driver's allocated memory, the JVM throws an OOM error. This is a common architectural mistake when debugging or extracting large-scale distributed data to a single location.

Why this answer

The collect() action attempts to pull the entire DataFrame into the driver node's memory. When the dataset size exceeds the heap space allocated to the driver, a Java heap space error occurs. Developers must avoid collecting large datasets and instead use take() or head() for previews, or write results to cloud storage to maintain application stability when working with big data at scale.

Exam trap

Candidates often assume that calling 'collect()' is a safe way to inspect data, failing to realize it pulls the entire result set into the driver's memory, causing OOM errors.

30
MCQmedium

A job is reading a huge amount of data from a table, but only uses three columns. Which optimization technique will provide the most significant I/O performance benefit?

A.Caching the DataFrame.
B.Column pruning.
C.Increasing the number of partitions.
D.Broadcasting the table.
AnswerB

Column pruning forces Spark to read only the columns required by the transformation. By ignoring unused columns at the storage layer, you significantly reduce the amount of data transferred from disk to memory, which is the primary performance gain for wide tables in distributed analytics and ETL applications.

Why this answer

Column pruning is the practice of reading only the required columns from a data source. In formats like Parquet, this reduces the total amount of data read from disk and transferred across the network. This is a crucial optimization for Databricks developers to reduce I/O bottlenecks and improve overall pipeline speed, especially when dealing with wide tables containing hundreds of unused columns in analytical query workloads.

Exam trap

Candidates often select repartitioning or caching instead of column pruning, misunderstanding that I/O bottlenecks depend on the amount of data read from disk.

31
MCQhard

A developer is troubleshooting a Spark job that fails with an OutOfMemoryError on the driver. The job collects a large DataFrame to the driver using .collect() and then processes it locally. The developer wants to avoid the driver OOM while still obtaining the results. Which approach is most appropriate?

A.Replace .collect() with .take(1000) to limit the number of rows returned.
B.Increase spark.driver.memory to a very large value to accommodate the collected data.
C.Use .foreach() to process each row on the driver instead of collecting the entire DataFrame.
D.Write the DataFrame to a distributed storage system and then read it back in smaller chunks for local processing.
AnswerD

Writing the DataFrame to distributed storage (e.g., Delta Lake, Parquet) and then reading it in batches avoids bringing the entire dataset to the driver at once. This leverages Spark's distributed nature for the heavy lifting and allows the driver to process manageable chunks. It is a scalable pattern that prevents driver OOM while still enabling access to all data.

Why this answer

Collecting a large DataFrame to the driver is an anti-pattern because it moves all data to a single node, risking OOM. The scalable solution is to persist the DataFrame to distributed storage and then read it in smaller partitions or batches for local processing. This maintains distribution and avoids overwhelming the driver.

Exam trap

The trap here is believing that increasing driver memory is a sustainable fix, when the real solution is to avoid collecting large data to the driver altogether.

32
MCQhard

A Spark job performing a join between a 10 GB table and a 5 MB lookup table is running slowly, and the physical plan shows a SortMergeJoin. You want to avoid the shuffle. What should you do?

A.Repartition both DataFrames on the join key before the join.
B.Increase spark.sql.shuffle.partitions to 2000.
C.Use a cross join and filter afterward.
D.Set spark.sql.autoBroadcastJoinThreshold to a value larger than 5 MB, such as 10 MB, and ensure the small table is broadcast.
AnswerD

The auto broadcast join threshold controls the maximum size of a table that can be broadcast to all executors. The default is 10 MB, but if the small table is slightly above the threshold or statistics are missing, Spark may choose SortMergeJoin. Explicitly setting the threshold higher or using broadcast() hint forces a BroadcastHashJoin, eliminating the shuffle of the large table. This is the correct approach to avoid the shuffle in this scenario.

Why this answer

Broadcasting the small table eliminates the need to shuffle the large table, converting the join to a BroadcastHashJoin. This is achieved by ensuring the small table's size is below the auto broadcast join threshold or by using an explicit broadcast hint. The other options either do not change the join strategy or introduce unnecessary shuffles.

Exam trap

The trap here is thinking that increasing shuffle partitions or repartitioning will avoid the shuffle, when in fact they only change how the shuffle is performed.

33
MCQmedium

A developer runs a PySpark job on Databricks that reads a large Delta table, filters on a timestamp column, and writes results to another Delta table. The job takes 45 minutes, but the Spark UI shows that 90% of task time is spent reading from the source table. The developer wants to reduce the read time. Which action should the developer take?

A.Increase the number of shuffle partitions by setting spark.sql.shuffle.partitions to a higher value.
B.Enable Delta Lake data skipping by ensuring the timestamp column is in the table's partitioning or Z-ORDER BY columns.
C.Cache the source DataFrame in memory using .cache() before applying the filter.
D.Repartition the source DataFrame by the timestamp column before filtering.
AnswerB

Delta Lake data skipping uses file-level statistics to skip reading files that do not contain relevant data. If the timestamp column is a partition column or has been optimized with Z-ORDER BY, the query engine can prune files based on the filter predicate, drastically reducing I/O and read time. This directly addresses the observed bottleneck.

Why this answer

The Spark UI indicates that the bottleneck is reading from the source Delta table. Delta Lake data skipping leverages file statistics to avoid reading irrelevant files when a filter predicate is applied. Ensuring the timestamp column is a partition column or has Z-ORDER BY applied allows the engine to prune files effectively, reducing I/O and overall job time.

Exam trap

The trap here is assuming that caching or repartitioning will speed up a slow read, when the real solution is to reduce the amount of data read via data skipping.

34
MCQmedium

A developer notices a Spark job is failing with an OutOfMemoryError during a join operation on two large tables. The join key is highly skewed, causing one task to process significantly more data than others. Which technique should be applied to resolve this skew?

A.Increase the spark.driver.memory configuration.
B.Implement salting on the join key to distribute the skewed keys.
C.Enable AQE and increase the spark.sql.shuffle.partitions value.
D.Convert the larger table into a broadcast variable.
AnswerB

Salting distributes rows with the same join key across multiple partitions by appending a random integer. This ensures that the massive volume of data associated with a single key is processed in parallel by different executors, effectively eliminating the bottleneck that triggers OutOfMemoryError in skewed join scenarios.

Why this answer

Salting involves adding a random prefix or suffix to the join key to redistribute the data across multiple partitions. This prevents a single executor from handling a disproportionate amount of data. This approach is a standard industry pattern for resolving skew-related OOM errors, as it breaks the hot key into smaller, manageable chunks that can be processed in parallel across the cluster without bottlenecking at a single node.

Exam trap

Candidates frequently suggest repartitioning by the skewed key itself. This is ineffective because it simply moves the same skewed data into a single partition, failing to alleviate the bottleneck on that specific task.

35
MCQmedium

Your Spark application is experiencing severe data skew while performing a join between a large fact table and a small dimension table. Which technique should you apply to optimize performance?

A.Increase the number of partitions using repartition() on the large table.
B.Enable AQE and increase the spark.sql.shuffle.partitions configuration.
C.Force a broadcast join for the small dimension table.
D.Apply a salting technique to the join keys of the small table.
AnswerC

Broadcasting the smaller DataFrame eliminates the need for a sort-merge join, which is where skewed data causes bottlenecks. By copying the small table to every executor, you perform a map-side join, effectively avoiding the shuffle of the large table and preventing skewed keys from overloading specific worker nodes.

Why this answer

Broadcast hash joins are the most effective solution for skew caused by a large table join when one side is small enough to fit in memory. By broadcasting the small table to all executor nodes, Spark avoids the expensive shuffle operation that causes data skew. This prevents specific partitions from becoming hotspots, which is critical for maintaining stable and performant ETL pipelines in production Databricks environments.

Exam trap

Candidates often suggest salting or repartitioning as the first step, ignoring that broadcasting is the most efficient and direct way to handle skew when a small table is involved.

36
MCQhard

A job joining a large fact table with a small dimension table runs out of memory on executors during the join. The dimension table is about 40 MB after filtering and the configured spark.sql.autoBroadcastJoinThreshold is 10 MB. The join key is highly skewed in the fact table. Which action is most appropriate?

A.Raise spark.sql.autoBroadcastJoinThreshold above the dimension table size so the small side is broadcast and no shuffle of the fact table occurs.
B.Repartition the fact table by the join key with a high partition count before the join to distribute the skew.
C.Increase spark.executor.memory so each executor can hold the skewed partitions of the fact table during the sort-merge join.
D.Salt the join key on both sides of the join so the skewed values are spread across many partitions.
AnswerA

Broadcasting the filtered 40 MB dimension table eliminates the shuffle of the large fact table and turns the join into a map-side operation. This avoids the memory pressure and shuffle associated with a sort-merge join on a skewed key. The threshold is a size guardrail, so raising it to cover the known small side is the intended control. This directly removes the source of the executor memory failure.

Why this answer

The dimension table is small enough to broadcast once the threshold is raised, which converts the join into a map-side operation and removes the shuffle of the large, skewed fact table. That eliminates the executor memory pressure caused by concentrating hot-key rows in a few reduce tasks. Salting and repartitioning address skew only in large-to-large joins and add cost here, while adding executor memory merely postpones the failure without changing the underlying plan.

Exam trap

The trap here is treating a large-to-small join as a skew problem to be salted, when broadcasting the small side removes the shuffle entirely.

37
MCQhard

A developer needs to optimize a Spark application that performs repetitive filtering and grouping on the same large DataFrame. Which feature should they implement to improve performance?

A.Increase the number of cores per executor.
B.Use the cache() or persist() method on the DataFrame.
C.Switch the file format from Parquet to CSV.
D.Enable speculative execution.
AnswerB

Caching pins the DataFrame in memory or disk, allowing subsequent actions to read the computed data directly rather than re-running the entire lineage. This significantly improves performance for iterative processing tasks, reducing latency and avoiding repeated data reads and transformations that occur in non-cached, re-evaluated Spark DataFrames.

Why this answer

Caching (or persisting) the DataFrame keeps it in memory or on disk for subsequent actions. This is essential for iterative algorithms or workloads where the same data is reused, as it avoids recomputing the lineage from scratch. Understanding the trade-offs between memory and disk persistence is fundamental for building performant Databricks pipelines that maximize resource reuse and minimize redundant computation time.

Exam trap

Candidates frequently confuse dataframe lineage optimization or broadcast hints with caching, missing that reused DataFrames need explicit persistence.

38
MCQhard

A Spark job writes a large DataFrame to a Delta table partitioned by date. The job is taking much longer than expected, and the Spark UI shows that many tasks are writing very small files. You have already set spark.sql.shuffle.partitions to 200. What is the most effective way to reduce the number of small files written?

A.Set spark.sql.files.maxRecordsPerFile to a high value to combine records into fewer files.
B.Use coalesce(1) before writing to reduce the number of output files.
C.Increase spark.sql.shuffle.partitions to 2000.
D.Enable optimized writes by setting spark.databricks.delta.optimizeWrite.enabled to true.
AnswerD

Optimized writes automatically coalesce small files during the write operation by adding a shuffle step that reduces the number of output files based on the data size. This is specifically designed to mitigate the small file problem in Delta Lake on Databricks. It balances file sizes without manual repartitioning, making it the most effective solution for reducing small files while maintaining parallelism.

Why this answer

Enabling optimized writes on Databricks triggers an automatic shuffle before writing to Delta, which coalesces data into fewer, larger files. This directly addresses the small file problem without sacrificing parallelism or requiring manual tuning. The other options either worsen the issue, introduce bottlenecks, or do not target the root cause of many small files from partitioned writes.

Exam trap

The trap here is thinking that increasing shuffle partitions or coalescing to one partition will solve small files, when optimized writes are the Databricks-specific feature designed for this.

39
MCQmedium

A PySpark job on Databricks repeatedly calls `df.count()` and `df.show()` inside a loop across 40 iterations, and the Spark UI shows the identical lineage being recomputed on every iteration even though the source Delta table is unchanged. You want to avoid re-executing the upstream transformations without materializing the data to disk. What should you do?

A.Wrap the DataFrame in `spark.createDataFrame(df.rdd)` so the derived object keeps the computed rows in the driver.
B.Call `df.persist(StorageLevel.MEMORY_AND_DISK)` before the loop and `df.unpersist()` after it.
C.Set `spark.sql.adaptive.enabled` to true so the optimizer reuses prior query results automatically.
D.Add `.repartition(200)` to the DataFrame before the loop so each iteration reads a different partition set.
AnswerB

Caching the DataFrame before the loop stores the computed partitions in executor memory (spilling to disk when needed), so each subsequent action reads the cached blocks instead of walking the lineage again. Since the source is unchanged, the cached result stays valid for all 40 iterations, and unpersisting releases the memory once the loop finishes.

Why this answer

Because the same DataFrame is acted on repeatedly and the underlying Delta table does not change, persisting the DataFrame keeps the computed partitions available across iterations, so the lineage is evaluated once rather than 40 times. Repartitioning, enabling adaptive execution, or round-tripping through RDDs all leave the recomputation behavior intact.

Exam trap

The trap here is assuming that enabling an optimizer feature automatically reuses results across actions, when only an explicit cache or persist call stores computed partitions.

40
MCQeasy

A developer notices that a Spark DataFrame job on Databricks is running slowly and the Spark UI shows that many tasks are reading from a Delta table with a large number of small files. The job performs a filter on a date column and then aggregates results. Which optimization technique will most directly improve read performance in this scenario?

A.Increase spark.sql.files.maxPartitionBytes to read more data per task.
B.Run OPTIMIZE on the Delta table to compact small files into larger ones.
C.Set spark.sql.autoBroadcastJoinThreshold to -1 to disable broadcast joins.
D.Enable Databricks Delta Cache to cache the table files on local SSDs.
AnswerB

OPTIMIZE compacts small files into larger, more efficient files, reducing the number of file opens and improving read throughput. This directly addresses the small files problem, which is a common cause of slow reads in Delta Lake. After compaction, the job will read fewer files, leading to faster scan and aggregation.

Why this answer

The small files problem in Delta Lake causes slow reads because each file requires a separate open and read operation. Compacting small files into larger ones with OPTIMIZE reduces the number of files, improving I/O efficiency. This is the most direct fix for the described symptom of many tasks reading small files, leading to faster filter and aggregation performance.

Exam trap

The trap here is thinking that caching or increasing partition size will fix slow reads caused by many small files, when the core issue is file count and compaction is needed.

41
MCQmedium

You are debugging a PySpark DataFrame job on Databricks that performs multiple transformations and actions on a large delta table. You notice that the execution plan shows redundant computations where the same upstream DataFrame is evaluated repeatedly. Which transformation should you apply to optimize this workflow and avoid recomputing the upstream lineage?

A.Call df.broadcast() on the DataFrame before each downstream join operation.
B.Call df.repartition() to distribute the data evenly across partitions before execution.
C.Call df.cache() or df.persist() before referencing the DataFrame in multiple downstream actions.
D.Call df.coalesce() to reduce the number of shuffle partitions prior to the final write operation.
AnswerC

Caching serializes or stores the evaluated DataFrame partitions in memory or disk storage. When subsequent actions trigger execution, Spark retrieves the cached blocks directly instead of re-evaluating the entire upstream transformation DAG from source files.

Why this answer

Persisting or caching a DataFrame instructs Spark to store intermediate results in memory or disk across actions, preventing expensive recomputation of the entire lineage graph. This optimization is critical for iterative algorithms or workflows where a single DataFrame is referenced multiple times downstream.

Exam trap

Candidates often confuse caching with broadcasting or assume Spark automatically caches all reused DataFrames. Spark does not evaluate reference frequency and will recompute lineage unless explicitly told to persist.

Ready to test yourself?

Try a timed practice session using only Troubleshooting and Tuning DataFrame Apps questions.