Courseiva

CCNA Troubleshooting Questions

42 questions · Troubleshooting topic · All types, answers revealed

1
MCQmedium

A data engineer runs a nightly Databricks job that reads from a large Delta table and writes results to another Delta table. The job has been taking progressively longer each night. The engineer examines the Spark UI and sees that the 'SQL' tab shows a single stage with many small tasks, each processing very few records, and the 'Storage' tab shows that the source table has thousands of small files. Which action should the engineer take to improve performance?

A.Enable Adaptive Query Execution (AQE) by setting spark.sql.adaptive.enabled to true.
B.Increase the number of shuffle partitions by setting spark.sql.shuffle.partitions to a higher value.
C.Run OPTIMIZE on the source Delta table to compact small files.
D.Repartition the source DataFrame to a smaller number of partitions before writing.
AnswerC

OPTIMIZE compacts many small files into larger ones, reducing the number of tasks and improving read throughput. In this scenario, the Spark UI shows thousands of small files and many small tasks, which is a classic small-file problem. Compaction directly addresses this by merging files, leading to fewer, larger files that are more efficient to read.

Why this answer

The Spark UI reveals many small tasks and thousands of small files in the source Delta table, which is a small-file problem. OPTIMIZE compacts these small files into larger ones, reducing task overhead and improving read performance. Other options either do not address the root cause or could exacerbate the issue by creating more partitions.

Exam trap

The trap here is assuming that increasing shuffle partitions or enabling AQE will solve small-file problems, when in fact compaction is required.

2
MCQmedium

A data engineer notices that a production Delta Lake table is experiencing slow read performance during concurrent write operations. The table contains millions of small files. Which action should the engineer take to resolve this performance degradation?

A.Run the VACUUM command on the table.
B.Increase the number of worker nodes in the cluster.
C.Execute the OPTIMIZE command on the table.
D.Enable Z-Ordering on all columns in the table.
AnswerC

OPTIMIZE compacts small files into larger, optimally sized files, which significantly improves read performance. This process is essential for maintaining high-performance Delta tables that experience frequent streaming writes or high-frequency batch inserts, as it reduces the number of files the query engine must track and scan.

Why this answer

Compacting small files into larger files using the OPTIMIZE command reduces metadata overhead and improves I/O efficiency. This is a critical maintenance task in Databricks because small files cause excessive object storage requests and metadata listing latency. By consolidating these files, the engine can scan data much more effectively during read operations, even while concurrent writes are occurring, as Delta Lake ensures ACID transactions remain consistent and performant.

Exam trap

Candidates often suggest 'dropping and recreating the table' or 'increasing the cluster size', which are destructive or costly workarounds for a simple file management issue that OPTIMIZE is designed to solve.

3
Multi-Selecthard

A data engineer is optimizing a Databricks job that performs a join between a large fact table and a small dimension table. The job is slow, and the engineer suspects that the join strategy is not optimal. Which TWO actions should the engineer take to improve performance? (Choose two.)

Select 2 answers
A.Enable Adaptive Query Execution (AQE) by setting spark.sql.adaptive.enabled to true
B.Broadcast the small dimension table using a broadcast hint
C.Cache the large fact table in memory
D.Increase spark.sql.shuffle.partitions to a very high value
E.Repartition the large fact table by the join key before the join
AnswersA, B

AQE can automatically convert a sort-merge join to a broadcast join at runtime if it detects that one side is small enough. It also optimizes shuffle partitions and handles skew. Enabling AQE allows Spark to adapt the join strategy based on actual data sizes, which can improve performance without manual hints. This is a recommended best practice in Databricks.

Why this answer

Broadcasting the small dimension table avoids shuffling the large fact table, which is a major performance win. Enabling Adaptive Query Execution allows Spark to automatically choose a broadcast join at runtime if the dimension table is small enough, and to optimize other aspects of the query. Together, these actions address the inefficient join strategy and improve performance.

Exam trap

The trap here is assuming that increasing shuffle partitions or caching always helps, when the real issue is the join strategy and the solution is to avoid the shuffle altogether.

4
Multi-Selectmedium

A data engineer is investigating a slow-running SQL query. Which TWO metrics from the Spark UI should the engineer examine to identify data skew?

Select 2 answers
A.Task duration
B.Executor CPU usage
C.Task input size
D.Shuffle read size
E.Driver memory usage
AnswersA, C

Task duration highlights significant variations in processing time across different executors. If most tasks complete in seconds but a few take minutes, it is a clear indicator of data skew or resource contention, prompting further investigation into the specific data distribution of the underlying partition.

Why this answer

Data skew occurs when one partition is significantly larger than others, causing one task to take much longer than the rest. By examining task duration and task input size, an engineer can spot outliers where a single executor handles a disproportionate amount of data. Identifying these metrics allows the engineer to apply remediation techniques like salting or repartitioning to balance the workload across the cluster effectively.

Exam trap

Candidates often choose metrics like 'Shuffle read size' or 'JVM memory usage', which are related to performance but do not directly expose the uneven distribution of data that characterizes skew.

5
MCQhard

A data engineer notices that a Databricks job processing a Delta table with 10,000 partitions runs slowly. The job filters on a column that is not the partition column, and the query plan shows that all partitions are being scanned. The engineer wants to improve performance without repartitioning the table. Which feature should be used?

A.Enabling Change Data Feed on the table
B.Setting delta.dataSkippingNumIndexedCols to 0
C.Increasing the number of shuffle partitions
D.Z-ORDER BY on the filter column during OPTIMIZE
AnswerD

Z-ORDER BY co-locates related data in the same set of files, allowing Delta Lake to skip files based on min/max statistics for the Z-ordered column. When the filter column is Z-ordered, the query can skip many files even if it is not the partition column, dramatically reducing the amount of data scanned and improving performance without changing the partitioning scheme.

Why this answer

When a filter is on a non-partition column, Delta Lake can still skip files if it has min/max statistics for that column and the data is clustered. Z-ORDER BY reorganizes data so that related values are stored together, improving the effectiveness of data skipping. Running OPTIMIZE with Z-ORDER BY on the filter column allows the query to read only the relevant files, avoiding a full scan of all partitions.

Exam trap

The trap here is assuming that partitioning is the only way to achieve data skipping, overlooking Z-ORDER BY as a complementary technique for non-partition columns.

6
MCQmedium

Which Spark configuration property can be used to enable Adaptive Query Execution (AQE) in Databricks?

A.spark.sql.adaptive.enabled
B.spark.sql.autoBroadcastJoinThreshold
C.spark.databricks.adaptive.execution
D.spark.sql.shuffle.partitions
AnswerA

Setting spark.sql.adaptive.enabled to 'true' is the required configuration to turn on the Adaptive Query Execution framework. Once enabled, Spark will dynamically optimize query execution plans based on runtime statistics, leading to more efficient processing and faster performance for complex SQL queries in Spark 3.x environments.

Why this answer

Adaptive Query Execution (AQE) is a critical optimization framework in Spark that re-optimizes query plans at runtime based on statistics collected during shuffle stages. It enables features like coalescing shuffle partitions and dynamically switching join strategies. Enabling AQE is a best practice for modern Databricks workloads as it provides significant performance gains without manual intervention, automatically adapting to varying data distributions in production pipelines.

Exam trap

Candidates often confuse AQE with general Spark configurations like executor memory or shuffle partitions. They might pick settings that control static partitioning instead of the dynamic runtime optimization framework.

7
MCQmedium

A data engineer notices that a scheduled Delta Lake maintenance pipeline is running significantly slower than expected. Upon checking the table history, they see that hundreds of tiny, fragmented data files have accumulated due to frequent streaming micro-batches. Which specific optimization command should the data engineer run first to resolve this file-size bottleneck?

A.Run VACUUM table_name RETAIN 0 HOURS to instantly purge all historical file references and force file compaction across the entire dataset.
B.Execute REFRESH TABLE table_name to clear the local metastore cache and force Spark to rebuild the underlying file index from scratch.
C.Invoke the OPTIMIZE table_name command to compact small files into larger, optimized data files and improve subsequent query execution efficiency.
D.Execute ALTER TABLE table_name SET TBLPROPERTIES (delta.autoOptimize.optimizeWrite = true) to rewrite past historical micro-batches retroactively.
AnswerC

OPTIMIZE compacts the hundreds of small fragmented files produced by frequent streaming micro-batches into larger, right-sized data files. This directly resolves the file-size bottleneck, reducing per-file overhead and improving subsequent query execution efficiency without altering the underlying data.

Why this answer

Running OPTIMIZE reorganizes the layout of Delta Lake data by compacting small files into larger, uniform files of roughly 1 GB in size. This significantly reduces metadata overhead and improves scan performance for subsequent read queries. While VACUUM removes old physical files, it does not compact active files.

OPTIMIZE directly addresses the core symptom of small file proliferation caused by frequent streaming micro-batches.

Exam trap

Candidates often confuse OPTIMIZE with VACUUM, believing that cleaning up old versions will automatically combine active small files into larger ones, which is not what VACUUM does.

8
MCQhard

Refer to the exhibit. A job fails with the provided error message. What is the most likely cause of this failure in a Databricks Delta Lake environment?

A.The Delta table has reached the maximum file limit.
B.The VACUUM retention threshold was set too low.
C.The cluster running the job lacks proper IAM permissions.
D.The transaction log is corrupted and requires manual deletion.
AnswerB

Setting the VACUUM retention period shorter than the time a long-running process needs to access old data leads to failures. If a process is mid-read while a concurrent VACUUM deletes files, the reader will be unable to locate the necessary data files, resulting in the reported error.

Why this answer

This error occurs when the Delta log entries point to files that no longer exist, typically caused by a VACUUM operation running with a retention period shorter than the time required for long-running streaming jobs to finish. When a file is physically deleted by VACUUM, but a reader still references it via the transaction log, the read operation fails because the file object is inaccessible in the underlying storage.

Exam trap

Candidates often assume the error is related to network connectivity or cluster permissions. They overlook the relationship between long-running jobs and the aggressive deletion of files by the VACUUM command.

9
MCQmedium

A job is failing with a 'Disk Space' error on the worker nodes. The code performs several large joins. Which configuration should the engineer adjust to mitigate the disk space usage?

A.Increase spark.sql.shuffle.partitions.
B.Decrease spark.driver.memory.
C.Set spark.databricks.io.cache.enabled to false.
D.Increase the number of task retries.
AnswerA

Increasing the number of shuffle partitions creates smaller, more manageable chunks of data. By reducing the size of individual partitions being joined, it becomes more likely that the operation can complete in memory, thereby avoiding the need to spill data to the local disk of worker nodes.

Why this answer

Large joins often require spilling to disk when the shuffle data exceeds the available executor memory. Adjusting the spark.sql.shuffle.partitions configuration can reduce the size of individual partitions, potentially fitting them within memory. Alternatively, increasing the instance type memory or optimizing the join condition reduces the reliance on local disk storage.

Managing shuffle partitions is a standard practice to balance memory pressure against the risk of disk overflow during shuffles.

Exam trap

Candidates often suggest 'increasing executor memory' as the first step, which is a costly infrastructure change compared to the performance tuning of shuffle partitions for memory-intensive joins.

10
MCQmedium

A data engineer runs a nightly Databricks job that ingests data into a Delta table using multiple concurrent write streams. The engineer notices that some write transactions are failing with a ConcurrentAppendException. The job writes to the same partition of the table from different tasks. Which action should the engineer take to resolve this issue?

A.Enable auto compaction on the Delta table to reduce the number of small files.
B.Partition the Delta table by a column that distributes writes across different partitions.
C.Increase the Spark shuffle partitions to allow more parallel writes.
D.Use the SQL `SET spark.databricks.delta.serializable` to enforce serializable isolation.
AnswerB

ConcurrentAppendException occurs when multiple writers attempt to add data to the same partition. By partitioning the table on a column that separates the write streams, each writer targets distinct partitions, avoiding conflicts. This is a recommended strategy for concurrent writes. It allows Delta Lake's optimistic concurrency control to commit transactions without conflict, as they modify disjoint sets of files.

Why this answer

The ConcurrentAppendException arises when concurrent transactions try to modify the same partition. Partitioning the table by a column that distributes the writes ensures that each transaction operates on separate partitions, eliminating the conflict. This leverages Delta Lake's ability to handle concurrent writes to different partitions.

Other options do not address the root cause: auto compaction and shuffle partitions do not prevent conflicts, and the serializable isolation setting is nonexistent.

Exam trap

The trap here is assuming that increasing parallelism or enabling auto compaction will resolve write conflicts, but the real solution is to isolate writes to different partitions.

11
Multi-Selectmedium

A data engineer is troubleshooting a slow-running Spark job on Databricks. Which TWO metrics in the Spark UI are most useful for identifying data skew?

Select 2 answers
A.Task Duration (max vs median)
B.Shuffle Read Size
C.Executor CPU Usage
D.Total Cluster Memory Usage
E.Driver Memory Consumption
AnswersA, B

Comparing the maximum task duration to the median or minimum duration reveals tasks that are processing significantly more data than others. A large gap between the median and max duration is a classic indicator that specific tasks are encountering skewed data partitions and slowing down the entire stage.

Why this answer

Identifying data skew requires looking for significant disparities in task execution times and data distribution across partitions. When one or two tasks take drastically longer than the others while processing similar data volumes, or when partition sizes are highly uneven, skew is likely present. Monitoring these metrics allows engineers to implement techniques like salting or broadcast joins to redistribute the data and balance the workload across executors.

Exam trap

Candidates often look at 'Executor Memory' or 'CPU usage' as primary indicators of skew. These metrics are symptoms, but the Spark UI specifically identifies skew via task distribution metrics.

12
MCQeasy

A data engineer is monitoring a Databricks job and notices that the job's duration has gradually increased over the past week. The job reads a large Delta table, performs aggregations, and writes results to another Delta table. The engineer wants to identify the stage that is taking the most time. Which Spark UI tab should the engineer use to quickly identify the slowest stage?

A.Environment tab
B.Stages tab
C.Storage tab
D.Jobs tab
AnswerB

The Stages tab lists all stages for a job and displays their duration, number of tasks, and other metrics. By sorting stages by duration, the engineer can quickly identify the slowest stage. This tab provides the necessary detail to pinpoint performance bottlenecks at the stage level, making it the correct choice.

Why this answer

The Spark UI Stages tab provides a detailed list of all stages, including their duration, number of tasks, and shuffle read/write metrics. By examining this tab, the engineer can sort stages by duration and identify the one consuming the most time. This is the most efficient way to locate the bottleneck stage and then investigate further.

Exam trap

The trap here is confusing the Jobs tab with the Stages tab; the Jobs tab gives high-level job durations, but stage-level timing is found in the Stages tab.

13
Multi-Selecthard

A data engineer is investigating why a Databricks job that writes to a Delta table is experiencing performance degradation over time. The job performs frequent small appends. Which TWO actions should the engineer take to improve write performance? (Choose two.)

Select 2 answers
A.Use partitioning on a high-cardinality column.
B.Increase the number of shuffle partitions to 2000.
C.Run OPTIMIZE on the Delta table to compact small files.
D.Set spark.sql.adaptive.enabled to true.
E.Enable auto compaction on the Delta table.
AnswersC, E

Frequent small appends create many small files, which degrade read and write performance. Running OPTIMIZE compacts these small files into larger ones, reducing the number of files and improving I/O efficiency. This is a recommended maintenance operation for Delta tables with many small files.

Why this answer

Frequent small appends create many small files, which degrade performance. Running OPTIMIZE compacts these files into larger ones, and enabling auto compaction automates this process. Together, they reduce the number of files and improve write and read efficiency.

These are standard Delta Lake maintenance practices.

Exam trap

The trap here is assuming that general Spark tuning parameters like shuffle partitions or AQE will solve Delta-specific small file issues, when the real fix is Delta Lake maintenance operations.

14
Multi-Selectmedium

An engineer needs to identify the root cause of a job failure. Which THREE of the following are valid locations or methods to investigate the logs?

Select 3 answers
A.Cluster event logs
B.Notebook workspace browser history
C.Spark Driver logs
D.Spark Executor logs
E.Databricks account profile settings
AnswersA, C, D

Cluster event logs track infrastructure activities, such as node additions, removals, and configuration changes. These logs are essential for determining if a job failed due to node provisioning issues or cluster lifecycle events, which are distinct from code-level errors occurring during the Spark job execution.

Why this answer

Troubleshooting in Databricks requires access to various log levels. The cluster event log captures infrastructure-level changes, driver logs provide application execution details, and executor logs offer task-specific debugging info. Using these three sources provides a holistic view, covering everything from cluster startup issues to specific code-level exceptions.

This multi-layered approach is essential for isolating whether a problem is infrastructure-related, configuration-based, or rooted in the logic of the transformation code itself.

Exam trap

Candidates often overlook 'Cluster event logs' as a source of information, focusing only on Spark logs, which misses infrastructure-level failures like spot instance terminations or node provisioning errors.

15
MCQmedium

A data engineer needs to troubleshoot a job that is failing during the 'shuffle' phase. Which Spark UI tab should the engineer examine to analyze the shuffle partitions and identify potential imbalances?

A.The 'Executors' tab.
B.The 'SQL' tab.
C.The 'Stages' tab.
D.The 'Environment' tab.
AnswerC

The 'Stages' tab provides granular metrics for shuffle read and write operations at the task level. By reviewing the distribution of data across shuffle partitions, the engineer can identify if specific tasks are processing significantly more data, which is the root cause of most shuffle-related performance issues.

Why this answer

The 'Stages' tab in the Spark UI is the primary place to analyze shuffle performance. Within this tab, an engineer can view the 'Shuffle Read' and 'Shuffle Write' metrics for each task. By examining the partition-level distribution, the engineer can detect if the data is being shuffled unevenly, which is a common cause of stage-level bottlenecks and job failures during complex transformations like joins or window functions.

Exam trap

Candidates inspect the 'Jobs' or 'Executors' tabs instead of the 'Stages' tab, missing the task-level shuffle read and write partition distribution metrics.

16
MCQmedium

An engineer has a large table that is frequently joined with small dimension tables. To optimize this, which optimization technique should be applied to the join operation?

A.Use a Broadcast Hash Join.
B.Implement a Shuffle Hash Join.
C.Increase the number of shuffle partitions.
D.Enable Z-Ordering on the join key.
AnswerA

A broadcast join sends the smaller table to all worker nodes. This eliminates the need to shuffle the large table, which is the most expensive part of a join operation in Spark. This strategy is highly effective when one side of the join is significantly smaller than the other.

Why this answer

Broadcasting small tables prevents the 'shuffle' of the larger table. By sending a copy of the dimension table to every executor, Spark can perform the join locally, significantly reducing network overhead. This is a fundamental optimization for star schemas in Databricks.

Knowing when and how to force broadcast joins allows engineers to dramatically speed up analytical queries by minimizing the movement of massive datasets across the cluster network.

Exam trap

Candidates confuse broadcast hash joins with sort-merge joins, failing to leverage small dimension tables to eliminate expensive network shuffle overhead entirely.

17
MCQmedium

A data engineer is using Databricks Jobs to run a nightly ETL pipeline. The job occasionally fails due to a transient network error when writing to an external database. The engineer wants to automatically retry the job a few times before marking it as failed. What is the most efficient way to configure this in Databricks?

A.Use a Databricks notebook to catch exceptions and loop until the write succeeds.
B.Configure the job to retry on failure with a specified number of retries and interval.
C.Set the job's timeout to a higher value to give the write operation more time to complete.
D.Set the job's maximum concurrent runs to 3 to allow multiple attempts.
AnswerB

Databricks Jobs support automatic retries on failure. By configuring the retry policy with a maximum number of retries and an interval between attempts, the job will automatically re-run if it fails due to transient errors. This is the most efficient and native way to handle transient failures without manual intervention.

Why this answer

Databricks Jobs offer a built-in retry policy that automatically re-runs a failed job a specified number of times with a defined interval. This is ideal for handling transient errors like network issues. Configuring retries at the job level is efficient, requires no code changes, and integrates with job monitoring and alerting.

Exam trap

The trap here is confusing concurrent runs with retries; concurrent runs allow parallel executions, while retries automatically re-run a failed job sequentially.

18
MCQhard

A data engineer is monitoring a Databricks job that runs a Structured Streaming query. The engineer notices that the query's input rate is high, but the processing rate is low, and the batch duration is increasing over time. The query uses a Delta table as a source and writes to another Delta table. Which action should the engineer take to improve the streaming query's performance?

A.Increase the trigger interval to process larger batches less frequently.
B.Enable Delta Lake optimized writes by setting spark.databricks.delta.optimizeWrite.enabled to true.
C.Increase the number of shuffle partitions to improve parallelism during processing.
D.Tune the source Delta table by running OPTIMIZE to compact small files.
AnswerD

A common cause of low processing rate in streaming queries reading from Delta is the accumulation of many small files, which increases the overhead of reading each micro-batch. Running OPTIMIZE compacts these files, reducing the number of files to read per batch and improving throughput. This directly addresses the input side, allowing the query to process data faster and reduce batch duration.

Why this answer

For a Structured Streaming query reading from a Delta table, many small files can cause high read overhead per micro-batch, leading to low processing rates and increasing batch durations. Compacting the source table with OPTIMIZE reduces the number of files, improving read efficiency. Other options either do not address the root cause or could exacerbate the issue.

Exam trap

The trap here is assuming that increasing trigger interval or shuffle partitions will solve streaming lag, when the real issue is often small files in the source Delta table.

19
MCQmedium

An engineer notices that a specific notebook job is consistently taking longer to start. They observe high 'initialization' times in the job logs. Which action should the engineer take to improve startup time?

A.Enable Auto-scaling on the cluster.
B.Use a Databricks Pool for the job cluster.
C.Increase the number of worker nodes.
D.Update the notebook code to use RDDs.
AnswerB

Databricks Pools maintain a set of idle, ready-to-use instances. By configuring the job to use a pool, the cluster can allocate nodes nearly instantaneously, bypassing the time spent requesting new instances from the cloud provider and reducing the overall initialization and startup phase for the job.

Why this answer

Initialization time in Databricks jobs often stems from the time required to pull container images, install cluster-scoped libraries, or initialize the Spark context on a new cluster. By using an existing cluster (pool) or pre-warming the cluster, the overhead of provisioning hardware and installing dependencies is removed. This optimization is crucial for meeting strict SLAs in production pipelines where every minute of latency impacts downstream processes.

Exam trap

Candidates often suggest increasing cluster size or changing instance types. While this might mask the problem, it is not the most efficient way to reduce startup latency.

20
Multi-Selecthard

A junior data engineer notices that a scheduled Databricks job running a heavy ETL notebook is failing intermittently due to cluster driver out-of-memory errors. Which TWO configuration changes or architectural adjustments should be implemented to resolve this issue? (Select exactly TWO)

Select 2 answers
A.Increase the maximum number of worker nodes in the autoscaling cluster configuration.
B.Select a larger driver node instance type with more memory capacity.
C.Refactor the PySpark notebook code to avoid using collect() on large DataFrames.
D.Enable Delta caching on the cluster workers to offload data from the driver.
E.Decrease the Spark SQL shuffle partitions default setting to a lower number.
AnswersB, C

Driver out-of-memory errors stem from insufficient memory on the driver node. Selecting a larger driver instance type with more memory capacity directly addresses that constraint, giving the notebook enough headroom to complete its ETL workload.

Why this answer

Driver out-of-memory errors typically occur when the driver node collects too much data into local memory using actions like collect() or handles excessive broadcast joins. Upgrading to a driver node with more RAM provides immediate headroom, while refactoring code to avoid pulling massive datasets to the driver prevents memory exhaustion fundamentally.

Exam trap

Many candidates assume scaling out worker nodes fixes driver memory issues. However, worker nodes process distributed data partitions independently, whereas the driver coordinates execution and collects results, meaning worker scaling has no direct impact on driver memory pressure.

21
MCQmedium

A data engineer is monitoring a Databricks job and notices that the job's tasks are spending a significant amount of time in garbage collection (GC). The job processes large amounts of data with many small objects. Which action should the engineer take to reduce GC overhead?

A.Set spark.executor.extraJavaOptions to -XX:+UseG1GC.
B.Use the Kryo serializer and optimize data structures to reduce object creation.
C.Decrease the number of partitions to reduce task overhead.
D.Increase the executor memory and enable off-heap memory.
AnswerB

The Kryo serializer is more efficient than Java serialization, reducing the size and number of objects created during shuffling and caching. Additionally, optimizing data structures to avoid unnecessary object creation can significantly lower GC pressure. This directly addresses the root cause of high GC overhead.

Why this answer

High garbage collection overhead often results from creating many short-lived objects. Using the Kryo serializer reduces the memory footprint of serialized data, and optimizing data structures to minimize object creation directly reduces the number of objects that the GC must manage. This leads to less frequent and shorter GC pauses.

Exam trap

The trap here is focusing on JVM tuning flags like G1GC, which can help but do not address the root cause of excessive object allocation.

22
MCQmedium

A data engineer runs a Structured Streaming job that writes to a Delta table. The job processes data from a Kafka topic and uses a 10-minute watermark. After a few hours, the engineer notices that the streaming query's input rate is steady, but the processing rate has dropped significantly, and the batch duration has increased from 5 seconds to over 2 minutes. The job is running on a cluster with autoscaling enabled. Which action should the engineer take FIRST to diagnose the performance degradation?

A.Increase the number of shuffle partitions by setting spark.sql.shuffle.partitions to a higher value.
B.Restart the streaming job with a larger driver node to handle the increased metadata load.
C.Repartition the Kafka source topic to increase parallelism in the streaming query.
D.Examine the Streaming Query progress metrics in the Spark UI, focusing on the 'addBatch' and 'walCommit' durations.
AnswerD

The Spark UI's Structured Streaming tab provides per-batch timing breakdowns, including addBatch (processing) and walCommit (write-ahead log commit). These metrics pinpoint whether time is spent in computation or in Delta Lake transaction commits. Examining them first is the correct diagnostic step because it directly reveals where the latency originates without guessing.

Why this answer

The Streaming Query progress metrics in the Spark UI are the primary tool for diagnosing Structured Streaming performance. They break down batch time into components like addBatch and walCommit, showing whether the delay is in data processing or Delta Lake commit operations. This targeted diagnosis avoids premature tuning and guides the engineer to the actual bottleneck.

Exam trap

The trap here is assuming that increased batch duration always means a need for more shuffle partitions or cluster resources, rather than first checking the built-in streaming metrics that isolate the slow stage.

23
MCQeasy

Which Databricks feature should a data engineer use to view the execution plan, including information about the physical operators and data lineage, to troubleshoot a slow-running SQL query?

A.The Delta Lake History logs.
B.The Spark UI SQL tab.
C.The Cluster Metrics dashboard.
D.The Data Explorer.
AnswerB

The Spark UI's SQL tab provides detailed information on how Spark parses, optimizes, and executes a query. It allows engineers to inspect the physical plan, identify time-consuming operators, and check for issues such as skewed joins or excessive data shuffling, making it the primary tool for query troubleshooting.

Why this answer

The Spark UI (specifically the SQL tab) provides a comprehensive graphical and textual representation of the query execution plan. By examining the 'Analyzed Plan' and 'Physical Plan', engineers can identify bottlenecks like expensive joins, full table scans, or lack of pruning. Mastering these diagnostic tools is essential for optimizing query performance and ensuring that execution logic aligns with the intended data processing requirements in the Databricks environment.

Exam trap

Candidates look at the general cluster metrics or job logs instead of navigating directly to the Spark UI SQL tab for physical execution plans.

24
MCQmedium

A data engineer is optimizing a Delta table that suffers from slow read performance due to small file sizes. Which command should the engineer execute to consolidate these small files into larger, more efficient files without altering the underlying table data?

A.ALTER TABLE table_name REORGANIZE
B.OPTIMIZE table_name
C.VACUUM table_name
D.ANALYZE TABLE table_name COMPUTE STATISTICS
AnswerB

The OPTIMIZE command packs small files into larger ones to improve scan speed. It leverages Delta Lake's file-level metadata to identify small files and rewrite them into larger, optimized files. This is the standard practice for maintaining performance in tables that receive frequent small write operations.

Why this answer

The OPTIMIZE command is specifically designed for compacting small files into larger ones, which improves read performance by reducing metadata overhead and increasing scan efficiency. This operation is essential in Databricks environments where frequent streaming or batch writes create many small files, leading to 'small file syndrome.' By restructuring the data files while maintaining the table's logical state, OPTIMIZE ensures that downstream queries scan fewer, more optimized data blocks.

Exam trap

Candidates often select 'VACUUM' because it is a common maintenance command. However, VACUUM deletes files rather than consolidating them, which would not solve a small file performance issue.

25
MCQeasy

A data engineer needs to monitor the costs associated with specific projects running on a shared Databricks workspace. Which feature should the engineer use to attribute these costs accurately?

A.Cluster Tags
B.Notebook Versioning
C.Delta Lake Audit Logs
D.Databricks SQL Alerts
AnswerA

Cluster tags provide a mechanism to label compute resources. These tags propagate to cloud provider billing reports, allowing admins to group costs by project. This is the standard Databricks mechanism for cost attribution in shared environments where multiple teams share the same underlying infrastructure and workspace resources.

Why this answer

Tags allow for granular cost tracking by applying metadata to Databricks resources like clusters. By using custom tags, organizations can categorize spending by project, department, or cost center. This is vital for financial accountability in multi-tenant environments.

When billing records are exported, these tags are included, enabling precise allocation of expenditures back to the specific project owners or internal clients, ensuring transparency and efficient budget management across the organization.

Exam trap

Candidates often suggest 'Workspace folders' or 'Job names' as cost-tracking methods, which lack the programmatic link to billing records that Cluster Tags provide for financial reporting.

26
MCQhard

A data engineer is optimizing a Databricks job that reads from a large Delta table and performs a join with a smaller table. The job is experiencing performance issues due to shuffling. The engineer wants to reduce the amount of data shuffled during the join. Which technique should the engineer use?

A.Broadcast the smaller table using a broadcast hint.
B.Increase the number of shuffle partitions to 2000.
C.Enable Adaptive Query Execution (AQE) and set spark.sql.adaptive.enabled to true.
D.Repartition the larger table on the join key before the join.
AnswerA

Broadcasting the smaller table avoids shuffling the larger table. The smaller table is sent to all worker nodes, and the join is performed locally on each node. This eliminates the shuffle of the large table, significantly reducing network overhead and improving performance. This is a standard optimization for joins where one table is small enough to fit in memory.

Why this answer

Broadcasting the smaller table eliminates the need to shuffle the larger table. The smaller table is replicated to all worker nodes, and the join is performed locally, which drastically reduces network I/O and improves performance. This is the most effective technique when one side of the join is small enough to fit in memory.

Exam trap

The trap here is assuming that increasing shuffle partitions or repartitioning will reduce shuffle volume, when in fact they can increase it; the key is to avoid shuffling the large table entirely.

27
MCQmedium

A data engineer is analyzing a Spark job that is failing with 'Out of Memory' (OOM) errors. Which configuration parameter should be tuned to increase the amount of memory allocated to the execution of joins and aggregations?

A.spark.memory.fraction
B.spark.executor.cores
C.spark.sql.shuffle.partitions
D.spark.driver.memory
AnswerA

This parameter controls the fraction of the heap space used for execution and storage. Increasing this value gives more memory to the execution pool relative to the storage pool, which directly helps in processing large joins and aggregations without running into OOM errors during the shuffle phase.

Why this answer

Spark's memory management divides heap memory into storage and execution. Joins and aggregations occur in the execution memory. When these operations process large datasets that exceed the allocated execution memory, Spark may fail with OOM.

Tuning the memory fraction configuration allows the engineer to shift the balance between storage (for caching) and execution, providing more headroom for complex shuffle-heavy operations like joins and group-by aggregations.

Exam trap

Candidates often suggest increasing 'spark.driver.memory' or 'spark.executor.memory'. While these help with total capacity, they do not manage the internal division between storage and execution memory.

28
MCQeasy

Which tool in Databricks provides real-time monitoring of cluster resource usage, including CPU, memory, and network throughput, for an active job?

A.The Spark UI.
B.The Ganglia UI.
C.The Query Profile.
D.The Databricks Job Run History.
AnswerB

Ganglia is the built-in monitoring tool in Databricks that provides deep, time-series insights into cluster-level health metrics like CPU load, memory utilization, and network traffic. It is the primary resource for troubleshooting hardware-level bottlenecks and determining if a cluster is appropriately sized for the workload it is executing.

Why this answer

The Metrics tab in the Databricks cluster UI offers a visual dashboard showing resource utilization. Monitoring these metrics helps engineers identify if a job is CPU-bound, memory-bound, or network-bound. This visibility is crucial for rightsizing clusters.

By observing these trends, engineers can optimize costs by reducing cluster size for underutilized jobs or improve performance by scaling up resources for jobs that encounter significant hardware contention.

Exam trap

Candidates often confuse the Spark UI with the Ganglia UI, incorrectly believing Spark metrics pages provide direct operating system-level CPU, memory, and network hardware monitoring.

29
MCQmedium

A data engineer wants to monitor the health and performance of Databricks Jobs over time. Which feature should they use to visualize trends, such as job success rates and average execution times, across multiple runs?

A.The Jobs Run History dashboard.
B.The Spark SQL UI.
C.The Delta Lake Time Travel feature.
D.Cluster Ganglia Metrics.
AnswerA

The Job Run History dashboard allows users to view the execution history of a job, including status, duration, and failure logs. It is the primary tool for analyzing performance trends over time, helping engineers identify when a job began failing or slowing down, which is essential for long-term pipeline stability.

Why this answer

The Databricks Jobs UI provides a comprehensive view of job history, which is essential for identifying long-term performance trends and reliability issues. By analyzing these metrics, engineers can detect regressions, troubleshoot intermittent failures, and optimize costs by adjusting schedules or resources. This historic view is a cornerstone of operational maintenance in Databricks, ensuring that pipelines remain stable and performant as data volumes evolve.

Exam trap

Candidates often suggest using the Spark UI or Ganglia metrics. While these provide technical cluster details, they do not provide the high-level job success and trend visualization requested.

30
MCQeasy

A data engineer is investigating a Databricks job that failed overnight. The job's status in the Jobs UI shows 'Failed', and the engineer needs to view the error message and stack trace to determine the cause. Where should the engineer look to find the detailed error information for the failed run?

A.In the driver logs, accessible from the run's detail page in the Jobs UI.
B.In the cluster's event log, which records all cluster scaling and termination events.
C.In the Databricks audit logs, which record user actions and API calls.
D.In the Spark UI's SQL tab, which shows the query plan and execution metrics.
AnswerA

The driver logs contain the standard output and error streams from the Spark driver, including exception stack traces and error messages. In the Jobs UI, each run has a detail page with links to logs. This is the primary location to find why the job failed, as it captures the exact error thrown during execution.

Why this answer

Driver logs are the definitive source for application-level errors in Databricks jobs. They capture stdout and stderr from the driver, including full stack traces. The Jobs UI provides direct access to these logs from the run's detail page, making it the first place to check when a job fails due to an exception.

Exam trap

The trap here is confusing infrastructure logs like cluster event logs with application logs; cluster events won't show the Python or Scala exception that caused the job to fail.

31
MCQmedium

A data engineer runs a nightly Databricks job that reads a large Delta table and writes aggregated results to another Delta table. The cluster logs show many small files in the source table, and the job runtime has increased steadily over weeks. The engineer wants to reduce the number of files without rewriting the entire table. Which command should be used?

A.ANALYZE TABLE sales COMPUTE STATISTICS
B.OPTIMIZE sales
C.ALTER TABLE sales SET TBLPROPERTIES ('delta.autoOptimize.optimizeWrite' = 'true')
D.VACUUM sales RETAIN 0 HOURS
AnswerB

OPTIMIZE compacts small files into larger ones (bin-packing) and can be run on the entire table or with a WHERE clause on a subset. It does not require rewriting the whole table and is the standard Delta Lake maintenance command to address the small-file problem, directly improving read performance for the nightly aggregation job.

Why this answer

The small-file problem is a common cause of slow reads in Delta Lake. OPTIMIZE performs bin-packing to combine small files into larger ones, reducing the number of files that must be opened during a read. It can be run on the entire table or a subset using a WHERE clause, and it does not require rewriting the entire table, making it the appropriate maintenance command for this scenario.

Exam trap

The trap here is confusing file cleanup (VACUUM) or statistics collection (ANALYZE) with file compaction (OPTIMIZE), when only compaction directly reduces the number of small files.

32
Multi-Selectmedium

Which THREE strategies are recommended to improve the performance of reading from a Delta table in Databricks?

Select 3 answers
A.Use Z-Ordering on columns frequently used in WHERE clauses.
B.Partition the table by every column used in the query.
C.Use OPTIMIZE to consolidate small files into larger files.
D.Always set the broadcast join threshold to -1.
E.Ensure statistics are kept up to date using ANALYZE TABLE.
AnswersA, C, E

Z-Ordering co-locates related data in the same files, which dramatically improves the efficiency of data skipping. When combined with filters, Delta Lake can skip entire files that do not contain the required data, significantly reducing the amount of I/O required for query execution.

Why this answer

Optimizing reads is crucial for analytics. Techniques like Z-Ordering, Data Skipping, and Partitioning allow Spark to ignore irrelevant data files, reducing I/O. Proper file sizing ensures that Spark can efficiently read data in parallel.

These strategies together minimize the amount of data scanned and transferred across the network, leading to significantly faster query results in large-scale data lake environments.

Exam trap

Candidates frequently include 'VACUUM' as a performance improvement strategy. While VACUUM is a maintenance task, it does not improve read performance; it actually removes historical data files.

33
MCQmedium

Which Databricks command should you use to recover storage space by removing files that are no longer referenced by a Delta table and are older than the retention period?

A.DELETE FROM table_name WHERE date < current_date() - 30;
B.TRUNCATE TABLE table_name;
C.VACUUM table_name RETAIN 30 DAYS;
D.DROP TABLE table_name;
AnswerC

VACUUM is the specific command used to delete data files that are no longer part of the Delta table's current state and are older than the specified retention threshold. This operation is necessary to reclaim storage space in cloud object storage and maintain compliance with data management policies.

Why this answer

VACUUM is the standard tool for cleaning up orphan files in Delta Lake. Because Delta Lake uses versioning, old files are kept for 'Time Travel' functionality until specifically vacuumed. Managing this storage is essential for cost control and compliance with data retention policies.

Engineers must understand that once VACUUM is run, the history of the table is permanently pruned, making point-in-time recovery to those older versions impossible.

Exam trap

Candidates confuse OPTIMIZE with VACUUM, incorrectly believing compaction removes old historical data files when only VACUUM handles data retention removal.

34
MCQhard

Refer to the exhibit. A data engineer receives this error when collecting data from a large transformation back to the driver node. Which approach should be used to fix this issue?

A.Increase the spark.driver.maxResultSize setting.
B.Rewrite the job to write the results to a Delta table.
C.Reduce the number of tasks in the Spark job.
D.Enable Spark dynamic allocation.
AnswerB

Writing the results to a Delta table ensures the data is persisted in a distributed format on storage. This avoids overwhelming the driver node and allows subsequent tasks to consume the data in parallel, which is the standard, scalable pattern for handling large datasets in Databricks.

Why this answer

The error occurs because the result set being pulled to the driver exceeds the configured limit. Collecting massive data to the driver is an anti-pattern in distributed computing as it bypasses the cluster's parallel processing capabilities. Instead of forcing data into the driver's memory, the engineer should write the output to cloud storage or a Delta table, allowing downstream processes to handle the data in a distributed, scalable manner without memory pressure.

Exam trap

Candidates frequently suggest increasing the driver memory or the Spark memory configuration, which is a temporary fix that ignores the fundamental architectural flaw of pulling large datasets into the driver node.

35
MCQhard

Refer to the exhibit. An engineer applies these configurations to a cluster. What is the primary benefit of enabling the Databricks IO Cache for a workload that involves repeatedly reading the same Delta tables?

A.It enables ACID transactions for non-Delta tables.
B.It speeds up repeated reads by caching data on local SSDs.
C.It automatically scales the cluster based on disk usage.
D.It forces the cluster to store all data in memory.
AnswerB

The Databricks IO cache stores frequently accessed data on the local SSDs of the worker nodes. When the same data is needed for subsequent queries, Spark retrieves it from the local cache rather than the cloud storage, drastically reducing latency and increasing overall query throughput for repeated read workloads.

Why this answer

The Databricks IO Cache (also known as the disk cache) accelerates data reads by caching remote data on the local SSDs of the worker nodes. For workloads that frequently query the same tables, this eliminates the latency and network overhead of repeatedly fetching data from cloud object storage. This is particularly effective for read-heavy analytical workloads, enabling significantly faster query execution times by leveraging high-speed local disk I/O.

Exam trap

Candidates often confuse the Databricks IO Cache with Spark RDD caching (cache() or persist()). They assume it applies to memory-based caching of DataFrames rather than disk-based caching of storage files.

36
Multi-Selectmedium

A data engineer is optimizing a Databricks job that processes a large dataset. The job performs a join between a large Delta table and a small dimension table, then writes the result to a Delta table. The engineer notices that the join is causing a large shuffle and wants to reduce shuffle overhead. Which two actions should the engineer take to improve performance? (Choose two.)

Select 2 answers
A.Enable Adaptive Query Execution (AQE) to automatically convert the join to a broadcast join if applicable.
B.Broadcast the small dimension table to avoid shuffling the large table.
C.Repartition the large Delta table on the join key before the join.
D.Increase the number of shuffle partitions to distribute the join workload more evenly.
E.Cache the large Delta table in memory before the join to speed up reading.
AnswersA, B

AQE can dynamically switch a sort-merge join to a broadcast join if it detects that one side of the join is small enough after initial stages. This reduces shuffle overhead by avoiding the shuffle of the large table. Enabling AQE allows Spark to optimize the join at runtime, complementing manual broadcast hints.

Why this answer

Broadcasting the small dimension table eliminates the need to shuffle the large table, as the small table is replicated to each executor for local joins. Enabling Adaptive Query Execution allows Spark to automatically convert sort-merge joins to broadcast joins when one side is small, further reducing shuffle. Together, these actions minimize shuffle overhead and improve join performance.

Exam trap

The trap here is thinking that repartitioning or increasing shuffle partitions will reduce shuffle overhead, when the real solution is to avoid shuffling the large table by broadcasting the small one.

37
MCQhard

A data engineer is troubleshooting a Databricks job that intermittently fails with a `SparkException: Job aborted due to stage failure: Task not serializable`. The job reads from a Parquet file, performs a transformation using a custom function defined in a Python class, and writes to a Delta table. The engineer suspects that the custom function is causing the issue. Which action should the engineer take to resolve the serialization error?

A.Mark the custom function with the `@staticmethod` decorator to avoid serializing the enclosing class.
B.Increase the driver memory to ensure that the serialized function fits in memory.
C.Refactor the custom function to avoid referencing any non-serializable objects from the enclosing class or module.
D.Set the Spark configuration `spark.serializer` to `org.apache.spark.serializer.KryoSerializer` to enable Kryo serialization.
AnswerC

The `Task not serializable` error occurs when a closure captures objects that cannot be serialized and sent to executors. By refactoring the function to avoid referencing non-serializable objects (e.g., database connections, file handles, or large objects), the function becomes serializable and the error is resolved. This directly addresses the root cause.

Why this answer

The `Task not serializable` error arises when Spark tries to serialize a closure for execution on executors but encounters objects that cannot be serialized. This often happens when a function references a non-serializable object from its enclosing scope, such as a database connection or a large data structure. Refactoring the function to avoid such references ensures that only serializable data is captured, resolving the error.

Exam trap

The trap here is assuming that changing the serializer or increasing memory will fix the issue, when the real problem is the capture of non-serializable objects in the closure.

38
MCQeasy

A data engineer is investigating why a Databricks job that reads from a Delta table is slow. The job performs a simple SELECT with a filter on a partition column. The engineer suspects that the table has many small files. Which Spark UI tab should be examined to confirm the number of files read?

A.SQL tab
B.Stages tab
C.Environment tab
D.Storage tab
AnswerB

The Stages tab in the Spark UI shows the number of tasks, input size, and records read for each stage. For a scan operation, the number of tasks often corresponds to the number of files or partitions read. By examining the input size and task count, the engineer can confirm whether many small files are being read, which indicates the small-file problem.

Why this answer

The Stages tab provides per-stage metrics including the number of tasks, input size, and records read. For a scan stage, the number of tasks is often proportional to the number of files or partitions. A high task count with small input size per task suggests many small files.

This makes the Stages tab the most direct place to confirm the small-file issue.

Exam trap

The trap here is assuming the SQL tab is always the best place for query diagnostics, when for file-level metrics the Stages tab gives clearer input and task counts.

39
Multi-Selectmedium

When configuring a Databricks Job, which TWO factors most directly influence the choice between a 'Job Cluster' and an 'All-Purpose Cluster'?

Select 2 answers
A.The need for interactive debugging and notebook testing.
B.The need for minimal compute cost for scheduled production workflows.
C.The requirement for high availability during query execution.
D.The number of users concurrently using the cluster.
E.The type of data source (e.g., S3 vs. ADLS).
AnswersA, B

All-Purpose clusters are designed for interactive use, allowing engineers to run code snippets, inspect data, and debug notebooks in real-time. They offer faster startup times for subsequent commands, making them the appropriate choice for development environments where quick iteration and constant developer feedback are required.

Why this answer

Choosing the right cluster type is vital for balancing cost and performance. Job Clusters are ephemeral, optimized for production tasks, and significantly cheaper, whereas All-Purpose Clusters are persistent and intended for interactive development. Understanding the cost implications and lifecycle management of these clusters is fundamental for any Databricks engineer responsible for managing operational budgets and ensuring efficient resource utilization across various environments.

Exam trap

Candidates often select All-Purpose clusters for production tasks to save time, ignoring the significantly higher costs and lack of proper lifecycle isolation compared to ephemeral Job clusters.

40
MCQmedium

Which action is recommended to resolve a scenario where a Databricks Job is failing due to excessive metadata operations on a Delta table with millions of files?

A.Increase the driver node instance size.
B.Run OPTIMIZE to consolidate files.
C.Disable the Delta transaction log.
D.Add more worker nodes to the cluster.
AnswerB

Running OPTIMIZE reduces the number of files by merging small files into larger ones. This directly reduces the number of entries in the Delta log and the number of metadata calls required to resolve the table state, significantly improving the performance of subsequent queries and avoiding the metadata bottleneck.

Why this answer

When a Delta table contains millions of small files, the transaction log and metadata operations become a bottleneck. The 'list' operations required to build the state of the table consume significant time and driver memory. Implementing partition pruning or using Delta Lake's table property 'delta.enableChangeDataFeed' are not the primary solutions here.

Instead, running OPTIMIZE to consolidate files is the standard way to reduce metadata overhead and improve table performance.

Exam trap

Candidates often suggest partitioning the table as a fix. While partitioning helps with data skipping, it does not fix the metadata overhead caused by having millions of small files.

41
Multi-Selectmedium

A Databricks job is failing due to 'Out of Memory' (OOM) errors during a join operation on two large datasets. Which TWO actions could help mitigate this issue?

Select 2 answers
A.Increase the cluster's worker node instance type with more memory.
B.Decrease the number of partitions in the Spark cluster.
C.Disable the spark.sql.autoBroadcastJoinThreshold configuration.
D.Enable dynamic allocation for the cluster.
E.Use the cache() method on every DataFrame in the pipeline.
AnswersA, C

Scaling up to worker nodes with higher memory capacity provides the Spark executors with more heap space. This allows them to process larger partitions and handle complex shuffle operations without spilling to disk or triggering OOM exceptions. This is the most direct hardware-level fix for memory-intensive join operations.

Why this answer

OOM errors during joins often occur because the driver or worker nodes lack enough memory to handle the shuffle or broadcast operations. Increasing the cluster memory allows for larger data partitions to reside in memory, while adjusting the broadcast join threshold prevents the optimizer from attempting to broadcast datasets that exceed available node capacity. These configurations are essential for stabilizing large-scale ETL pipelines that process high-volume data.

Exam trap

Candidates often try to optimize joins by enabling broadcast thresholds on massive datasets, accidentally triggering Out of Memory errors instead of disabling the threshold.

42
Multi-Selecthard

A streaming job using Structured Streaming is lagging significantly behind the 'current time'. Which THREE of the following could be the root cause of this processing latency?

Select 3 answers
A.The trigger interval is set to a time shorter than the processing time of the batch.
B.Lack of proper watermarking on stateful aggregations.
C.Using a fixed-size cluster with no auto-scaling enabled.
D.Inefficient shuffling due to data skew.
E.Using too few partitions in the input source.
AnswersA, B, D

If the time taken to process a batch exceeds the trigger interval, the streaming job cannot keep up. This leads to a queue of pending batches, causing the watermark and processing time to drift further behind, effectively indicating that the cluster is undersized for the current workload volume.

Why this answer

Streaming latency often stems from resource contention, inefficient state management, or micro-batch trigger configurations. When the input rate exceeds the processing rate, lag accumulates. Identifying these bottlenecks requires analyzing the Spark UI for processing time vs. batch time.

Understanding these factors is vital for maintaining low-latency pipelines and ensuring that SLAs are met in production-grade streaming environments deployed on Databricks.

Exam trap

Candidates often overlook trigger interval misconfigurations or forget watermarking requirements, mistakenly assuming streaming lag is always caused solely by insufficient cluster sizing.

Ready to test yourself?

Try a timed practice session using only Troubleshooting questions.