Courseiva
← Back to Databricks Certified Data Engineer Professional questions

Scenario-based practice

Hard Difficulty Questions

Practise Databricks Certified Data Engineer Professional practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

20
scenario questions
Databricks-DE-Pro
exam code
Databricks
vendor

Scenario guide

How to approach hard difficulty questions

These are the questions most candidates get wrong. They require connecting multiple concepts, reading tricky output, or knowing edge-case behaviour that isn't on most study cards. Practising them trains you to operate under uncertainty — a necessary skill on the real exam.

Quick answer

Hard Difficulty Questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Related practice questions

Related Databricks-DE-Pro topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmulti select
Full question →

A data engineer is investigating a job failure that occurred only in the production environment. Which TWO features in Databricks help in comparing the production environment to the development environment?

Question 2hardmultiple choice
Full question →

A data engineer is configuring a Delta Live Tables (DLT) pipeline that processes streaming data from Apache Kafka. The pipeline performs a series of transformations and writes to a Delta table. The engineer notices that the pipeline is experiencing high latency and wants to optimize it for cost and performance. The pipeline is set to continuous mode. Which configuration change is most effective to reduce cost while maintaining acceptable latency?

Question 3hardmultiple choice
Full question →

Refer to the exhibit. If a job is failing because the cluster is reaching the 'max_workers' limit too quickly, which monitoring metric should the engineer verify to confirm if the cluster is actually utilizing these nodes efficiently?

Exhibit

{
  "cluster_policy": {
    "autoscale": {
      "min_workers": 2,
      "max_workers": 8
    }
  }
}
Question 4hardmultiple choice
Full question →

Your team is using a shared cluster for development. A user reports that their job is slow because the cluster memory is frequently filled by large data broadcasts. What configuration adjustment should you make to prevent this issue across all jobs on the cluster?

Question 5hardmultiple choice
Full question →

Which TWO of the following are primary benefits of using Delta Live Tables (DLT) for data ingestion over standard Structured Streaming pipelines?

Question 6hardmultiple choice
Full question →

When running a PySpark job, you receive an 'Out of Memory (OOM)' error during a shuffle operation. Which configuration is the most appropriate to address this first?

Question 7hardmultiple choice
Full question →

An engineer notices that a SQL warehouse is frequently hitting 'Max Concurrency' limits. Which log should they consult to identify which specific queries are consuming most of the warehouse resources?

Question 8hardmulti select
Full question →

A data engineering team is modeling a large Delta Lake fact table that stores clickstream events. Analysts frequently run queries that filter by event_date and then aggregate by user_id, and the table receives continuous appends plus occasional late-arriving corrections. The team wants to reduce bytes scanned and improve join performance. Which two design choices are most appropriate? (Choose two.)

Question 9hardmulti select
Full question →

A data engineer is implementing fine-grained access control on a Delta table in Unity Catalog that contains sensitive customer data. The requirement is to mask the `credit_card` column for all users except members of the `finance` group, and to filter out rows where the `region` column is not in the user's allowed regions. Which two Unity Catalog features should the engineer use? (Choose two.)

Question 10hardmultiple choice
Full question →

A Data Engineer needs to encrypt data at rest within a Databricks workspace that uses a customer-managed key (CMK). What is the primary purpose of this configuration?

Question 11hardmultiple choice
Full question →

You are performing a complex data transformation involving a self-join on a large, skewed table. Which technique is most effective for preventing data skew and improving join performance?

Question 12hardmultiple choice
Full question →

Refer to the exhibit. A Databricks job fails with a 403 Forbidden error when trying to write to the S3 bucket. Why does this happen?

Exhibit

JSON policy: {
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": ["s3:GetObject"],
      "Resource": ["arn:aws:s3:::my-bucket/*"]
    }
  ]
}
Question 13hardmultiple choice
Full question →

A data engineer is designing a solution to share a Delta table with an external partner organization. The partner uses a different Databricks account and must be able to read the table, but the data must not be copied outside the provider's cloud storage. The provider uses Unity Catalog and wants to minimize operational overhead while ensuring the partner sees only the shared table. Which Unity Catalog feature should the engineer use?

Question 14hardmulti select
Full question →

A data engineer is tasked with reducing compute costs for an interactive SQL analytics workspace that runs sporadic, highly unpredictable queries. The jobs experience cold start delays and occasional out-of-memory errors due to sudden concurrency spikes. Which TWO strategies should the engineer implement to balance cost efficiency and performance?

Question 15hardmultiple choice
Full question →

A data engineer is using PySpark to process a large DataFrame and needs to reduce the number of partitions before writing to a Delta table to avoid creating too many small files. The DataFrame currently has 2000 partitions, each about 10 MB. The engineer wants to reduce the number of partitions to approximately 200 while minimizing data shuffling. Which approach is most appropriate?

Question 16hardmultiple choice
Full question →

A data engineer is implementing a data quality check on a streaming Delta table that receives late-arriving events. The engineer needs to detect records where the event timestamp is more than 2 hours behind the current watermark. Which approach correctly identifies such records for quarantine?

Question 17hardmultiple choice
Full question →

A Data Engineer is implementing a data quality framework for a Delta table that receives streaming updates. The engineer needs to ensure that any record with a negative `amount` is not written to the table, but also wants to capture these invalid records in a separate quarantine table for later analysis. The pipeline must not fail on invalid records. Which approach using Delta Live Tables achieves this?

Question 18hardmultiple choice
Full question →

Which THREE of the following are benefits of using Delta Lake over standard Parquet files for your data lake storage?

Question 19hardmulti select
Full question →

Which TWO of the following are true regarding Unity Catalog's ability to govern external locations?

Question 20hardmulti select
Full question →

A data engineer is optimizing a Spark job that reads from a large Delta table and performs a join with a smaller dimension table. The job is running slowly, and the engineer suspects data skew and shuffle overhead are the main issues. Which two techniques should the engineer apply to improve performance and reduce cost? (Choose two.)

These Databricks-DE-Pro practice questions are part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style Databricks-DE-Pro questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.