Courseiva
← Back to Databricks Certified Data Engineer Professional questions

Scenario-based practice

Refer to the Exhibit Practice Questions

Practise Databricks Certified Data Engineer Professional practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

15
scenario questions
Databricks-DE-Pro
exam code
Databricks
vendor

Scenario guide

How to approach refer to the exhibit practice questions

Practise exhibit-style questions that ask you to read a topology, table, command output or diagram before choosing the best answer.

Quick answer

Exhibit-style questions test whether you can read a topology, command output, diagram or table before choosing the best answer.

How to extract the relevant detail from an exhibit.

How topology, command output or routing information affects the answer.

How to avoid answering from memory before reading the evidence.

How to map the exhibit back to the exam objective.

Related practice questions

Related Databricks-DE-Pro topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmultiple choice
Full question →

Refer to the exhibit. A Databricks job fails with a 403 Forbidden error when trying to write to the S3 bucket. Why does this happen?

Exhibit

JSON policy: {
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": ["s3:GetObject"],
      "Resource": ["arn:aws:s3:::my-bucket/*"]
    }
  ]
}
Question 2hardmultiple choice
Full question →

Refer to the exhibit. A data engineer creates an instance pool to reduce cluster startup times for development teams. However, finance reports indicate unexpected cloud infrastructure charges. Based on the configuration shown in the exhibit, what is the primary driver of these unexpected costs?

Exhibit

{
  "instance_pool_id": "pool-0412-182230-cried5",
  "min_idle_instances": 2,
  "max_capacity": 10,
  "node_type_id": "i3.xlarge"
}
Question 3hardmultiple choice
Full question →

Refer to the exhibit. Which action is the most appropriate to resolve this memory-related failure during the job execution?

Exhibit

Error Log: org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 1.0 failed 4 times, most recent failure: Lost task 0.3 in stage 1.0: ExecutorLostFailure (executor 2 exited caused by one of the running tasks) Reason: Container killed by YARN for exceeding memory limits. 10.2 GB of 10 GB physical memory used.
Question 4hardmultiple choice
Full question →

Refer to the exhibit. The alert is intended to trigger if the data in 'my_table' is older than one hour. Which query modification correctly implements this check?

Exhibit

{
  "alert": {
    "name": "Data Freshness",
    "query": "SELECT max(updated_at) FROM my_table",
    "threshold": "now() - interval 1 hour"
  }
}
Question 5mediummultiple choice
Full question →

Refer to the exhibit. You are appending data to an existing Delta table. What is the most likely cause of this error, and how should you resolve it?

Exhibit

Error: Py4JJavaError: An error occurred while calling o68.save. : org.apache.spark.sql.AnalysisException: Cannot write incompatible data to table 'sales_data': Column 'price' (decimal(10,2)) cannot be cast to 'price' (decimal(8,2)).
Question 6hardmultiple choice
Full question →

Refer to the exhibit. An engineer created this alert for a query. Under what condition will the alert status change to 'Triggered'?

Exhibit

{
  "alert_config": {
    "name": "Latency Alert",
    "query_id": "12345",
    "options": {
      "column": "duration",
      "op": ">",
      "value": "3600"
    }
  }
}
Question 7hardmultiple choice
Full question →

Refer to the exhibit. A Databricks administrator wants to restrict access to a specific SQL Alert. Based on the JSON policy, which statement accurately describes the current permission model for this alert?

Exhibit

{
  "action": "allow",
  "principal": "user@example.com",
  "resource": "/alerts/12345",
  "permission": "CAN_VIEW"
}
Question 8hardmultiple choice
Full question →

Refer to the exhibit. A Data Engineer is attempting to merge data into a table with these constraints defined. If the incoming batch contains rows that violate these rules, what is the default behavior of the Delta Lake engine during the merge operation?

Exhibit

{
  "type": "Table",
  "constraints": {
    "check": "price > 0",
    "not_null": ["id", "transaction_date"]
  }
}
Question 9mediummultiple choice
Full question →

Refer to the exhibit. The engineer wants to replace only one specific partition in the 'orders' table. What is the best method in Databricks?

Exhibit

{
  "table": "orders",
  "strategy": "overwrite",
  "partition": "date"
}
Question 10hardmultiple choice
Full question →

Refer to the exhibit. You are using Auto Loader to ingest data with evolving schemas. After running the job for a week, you realize that new columns added to the source JSON are not being captured in the destination table. What must you add to the configuration?

Exhibit

{"cloudFiles.format": "json", "cloudFiles.schemaLocation": "/tmp/schema", "cloudFiles.inferColumnTypes": "true"}
Question 11mediummultiple choice
Full question →

Refer to the exhibit. An engineer has configured the cluster settings as shown. What is the expected impact on the Delta table's performance and write operations?

Exhibit

{"cluster_config": {"spark.databricks.delta.optimizeWrite.enabled": "true", "spark.databricks.delta.autoCompact.enabled": "true"}}
Question 12mediummultiple choice
Full question →

Refer to the exhibit. An engineer observes that queries filtering on 'customer_id' are running slowly despite Z-Ordering. What is the most likely cause?

Exhibit

{"table_name": "sales_data", "partition_columns": ["region", "date"], "z_order_columns": ["customer_id"], "file_format": "delta"}
Question 13mediummultiple choice
Full question →

Refer to the exhibit. A data engineer is deploying a production pipeline that references a table in the default schema. The job fails with the provided error. What is the root cause?

Exhibit

{
  "error": "AnalysisException",
  "message": "Table or view not found: default.sales_data",
  "trace": "at org.apache.spark.sql.errors.QueryCompilationErrors$.tableOrViewNotFoundError"
}
Question 14mediummultiple choice
Full question →

Refer to the exhibit. Why is Task D marked as 'Skipped'?

Exhibit

Task A: [Success]
Task B: [Failed]
Task C: [Pending]
Task D: [Skipped]
Question 15hardmultiple choice
Full question →

Refer to the exhibit. A user encounters this error when running a query. What is the correct action to resolve this issue while maintaining the security model?

Exhibit

{"error": "PERMISSION_DENIED", "message": "User does not have USE CATALOG privilege on catalog 'finance'", "operation": "SELECT * FROM finance.revenue.q1_data"}

These Databricks-DE-Pro practice questions are part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style Databricks-DE-Pro questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.