Courseiva

Databricks-DE-Assoc · topic practice

Working with Lakeflow Jobs practice questions

This domain covers orchestrating data pipelines with Lakeflow Jobs in Databricks: notebook and task configuration, dependencies, retries, scheduling, and failure recovery. Questions present failure scenarios or multi-task DAGs and ask you to pick the correct setting, dependency behavior, or repair outcome, so you must know how task-level options interact.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Working with Lakeflow Jobs

What the exam tests

What to know about Working with Lakeflow Jobs

Be able to configure a multi-task Lakeflow Job: set task dependencies, add task-level Retries for transient failures, and use Repair and Rerun to recover only failed portions. The key is knowing default dependency behavior: downstream tasks wait for all upstream parents to succeed.

Configuring task-level Retries and timeouts for notebook tasks hitting transient external API errors

Using 'Repair and Rerun' to rerun only failed and downstream tasks while preserving successful task results

Setting task dependencies so downstream tasks run only when all upstream parents succeed

Understanding how Lakeflow Jobs orchestrate multi-task workflows with run-if conditions and scheduling

Watch out for

Common Working with Lakeflow Jobs exam traps

  • ▸Assuming Repair and Rerun reruns the entire job; it only reruns failed tasks and their downstream dependents, reusing prior successful task outputs.
  • ▸Confusing task-level Retries with job-level retries or cluster restart policies, leading to wrong answers about where retry behavior is configured.
  • ▸Believing downstream tasks run by default when a parent fails; default dependency requires all parents to succeed unless a run-if condition is set.

Practice set

Working with Lakeflow Jobs questions

20 questions · select your answer, then reveal the explanation

A data engineer wants to schedule a Databricks Job to run every Tuesday at 8:30 AM UTC. Which cron expression correctly represents this schedule?

When configuring a Databricks Job, which notification option ensures that a team is alerted specifically when a job fails?

A data engineer is configuring a Lakeflow Job with multiple tasks. The engineer wants to ensure that the job can be repaired and that only failed tasks are re-run without re-executing successful tasks. Which TWO features or configurations enable this behavior? (Choose two.)

A data engineer is configuring a Lakeflow Job with a task that runs a Python wheel. The engineer needs to ensure that the task can access a specific set of Python libraries that are not included in the Databricks Runtime. The engineer wants to avoid installing these libraries at runtime using %pip. Which TWO of the following approaches are valid for making these libraries available to the task? (Choose two.)

A data engineer is building a Lakeflow Job with three tasks: bronze_ingest, silver_transform, and gold_aggregate. The engineer wants gold_aggregate to run only when silver_transform succeeds, and wants the job to skip gold_aggregate without marking the whole job as failed when silver_transform is skipped. Which two configuration choices support this behavior? (Choose two.)

A data engineer configures a Lakeflow Job with a single notebook task. The task sometimes runs longer than expected because of upstream source latency. The engineer wants the task to be automatically retried only when it fails, but they do not want to pay for idle compute between retries. Which configuration should they use?

A data engineer needs to run a Lakeflow Job every weekday at 06:00 in the company's local time zone, which is America/New_York. The engineer configures a scheduled trigger with a cron expression. Which cron expression correctly represents this schedule in Quartz format?

A data engineer has a Lakeflow Job with a notebook task that reads from a streaming source and is expected to run indefinitely. The engineer wants the task to be restarted automatically if the cluster is terminated or the notebook fails due to an out-of-memory error, but does not want the entire job to be marked as failed. The engineer also wants to limit the number of restart attempts to 3. Which configuration should be used for this task?

A data engineer is configuring a Lakeflow Job with a parameter named 'env' that has a default value of 'dev'. The engineer wants to override this value to 'prod' when triggering the job via the Databricks REST API. The engineer also wants to ensure that the parameter is available to all tasks in the job. Which approach should be used?

A data engineer needs to configure a Databricks Job containing multiple tasks where downstream tasks should only execute if all upstream parent tasks complete successfully. Which task dependency setting should be configured?

Which TWO of the following statements are true regarding the behavior and capabilities of Databricks Jobs parameters and values?

A data engineer has configured a Databricks Job with multiple dependent tasks forming a linear pipeline. Task A extracts data, Task B transforms it, and Task C loads it into a gold table. The pipeline runs daily. The team notices that if Task B fails due to an intermittent schema validation issue, the entire job run fails, but they want Task C to execute conditionally only if Task B succeeds, while alerting the on-call engineer immediately upon any failure. How should the task dependencies and conditional execution be configured?

A data engineer needs to configure a Databricks Job containing three distinct tasks: ingest, transform, and report. The transform task must only execute if the ingest task completes successfully, but the report task should execute regardless of whether the transform task succeeds or fails. How should the task dependencies be configured?

A data engineer wants to ensure that a Databricks Job task only runs if the preceding task completes successfully, but needs to add a specific timeout threshold for this individual task. Where should this configuration be applied?

A data engineer is designing a Databricks Job workflow. Which TWO of the following are valid ways to trigger a Databricks Job?

Refer to the exhibit. If 'task1' fails due to a timeout, what happens to 'task2'?

Exhibit

{
  "tasks": [
    {
      "task_key": "task1",
      "notebook_task": { "notebook_path": "/nb1" },
      "timeout_seconds": 3600
    },
    {
      "task_key": "task2",
      "depends_on": [{ "task_key": "task1" }],
      "notebook_task": { "notebook_path": "/nb2" }
    }
  ]
}

Which of the following is the primary benefit of using a 'Job Cluster' rather than an 'All-Purpose Cluster' for running scheduled data pipelines?

A data engineer needs to pass the execution date to a job task dynamically. Which feature should they use?

Which THREE of the following are benefits of using Delta Live Tables (DLT) for managing your data pipelines?

What is the primary function of the 'Retries' setting in a Databricks Job task?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Working with Lakeflow Jobs sessions

Start a Working with Lakeflow Jobs only practice session

Every question in these sessions is drawn from the Working with Lakeflow Jobs domain — nothing else.

Related practice questions

Related Databricks-DE-Assoc topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-DE-Assoc exam test about Working with Lakeflow Jobs?
Be able to configure a multi-task Lakeflow Job: set task dependencies, add task-level Retries for transient failures, and use Repair and Rerun to recover only failed portions. The key is knowing default dependency behavior: downstream tasks wait for all upstream parents to succeed.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Working with Lakeflow Jobs questions in a focused session?
Yes — the session launcher on this page draws every question from the Working with Lakeflow Jobs domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-DE-Assoc topics?
Use the topic links above to move to related areas, or go back to the Databricks-DE-Assoc question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-DE-Assoc exam covers. They are not copied from any real exam or dump site.